Research paper · COM 885 · extended with a computational pipeline

“That’s kino anon” — reading charts as gamified devices on 4chan’s /lit/.

A communication-research study of the “charts” circulating on 4chan’s Literature board. How do these playful visual reading lists frame and diffuse literary radicalism inside an anonymous online counter-public? The original qualitative analysis has since been extended into a reproducible pipeline that measures the whole corpus — and the numbers complicate the story.

264 charts
Corpus, fully processed
34 581
Works extracted by OCR
75 %
Are catalogues, not paths
56 %
Density of the canon network
Research question

When a reading list becomes political communication.

On /lit/, users share charts — small PNG or JPEG images that rank authors and books along logics of progression and hierarchy. The paper argues that these visuals do more than recommend references: they structure paths of initiation and distinction within an anonymous public, turning literature into an indirect political language.

Central claim — “How is radicalism circulated through the form of the chart?”

Reading is reframed as a quest: you climb from tier to tierProgression, unlock authors and become “proficient”Reward, follow arrows, levels and colour codesSimplification, and enter the closed circle of “real readers”Community.

Theoretical framework

Reading /lit/ as a digital counter-public.

The analysis mobilises theories of the public sphere and its fragmentation to make sense of an anonymous, affective and highly symbolic discursive arena.

01

Fragmented public spheres

Bennett & Pfetsch (2018): political communication now unfolds across dispersed micro-publics rather than one common arena.

02

Counter-publics

Fraser (2001): subordinate groups build parallel discursive arenas to circulate their own counter-discourses — /lit/ as a contemporary illustration.

03

Pseudo-public sphere

De Zeeuw (2024): spaces that imitate the mechanisms of public debate without honouring its deliberative rules — where nothing is quite serious, nothing quite play.

04

Gamification & ludification

Deterding et al. (2011) and Raessens (2014): game-design logics applied to non-game contexts, reconfiguring culture as a playful, identity-driven pursuit.

05

Framing

Entman (2004): to frame is to select and make salient. The chart delimits what should be read, thought and believed before any argument is made.

06

Trolling & transgression

Nagle (2017): a “culture of transgression” that normalised reactionary discourse behind cynicism and memetic play.

Results · the whole corpus, measured

What 264 charts actually do.

The original study read a sample closely. The pipeline now reads all of it: every chart was OCR’d locally, its blocks and sections rebuilt geometrically, its text screened for a lexicon of gamification, and its form coded on the image. The findings below refine the paper’s argument rather than confirm it — which is what measurement is for.

Figure 1 · form
Bar chart of chart forms: plain grid 110, sectioned grid 76, flowchart 29, tier list 24, text list 11, collage 11, map 3.
Three quarters of the corpus are catalogues — grids, with or without named sections — that juxtapose works without prescribing any route. Structurally prescriptive forms, flowcharts and tier lists, account for only 20 %, and just 13 % of charts draw arrows at all. 75 % / 20 % Against the imaginary of “start with the Greeks”, the dominant form on /lit/ is not the guided path but the ranked list.
Figure 2 · gamified markers
Share of charts carrying each marker: prescribed order 26%, imperative 24%, entry point 17%, prerequisites 14%, graded difficulty 12%, optional 6%, reward 6%, named tiers 4%, value hierarchy 2%.
Nine registers were searched by regular expression across every chart’s full text — a deterministic, checkable method, with the matching lines kept as evidence. No mechanic is in the majority. The most widespread are discursive: a quarter of charts prescribe an order or address the reader in the imperative. The most explicitly video-game-like — named tiers (4 %) and the God-Tier / pleb value hierarchy (2 %) — turn out to be rare. 26 % vs 4 % Gamification on /lit/ is first a matter of prescriptive language, not of video-game decorum.
Figure 3 · the central result
Heatmap of form against marker. Flowcharts: 66% entry point, 62% imperative. Tier lists: 25% named tiers, 29% graded difficulty. Text lists: 73% imperative. Grids: near zero across the board.
Crossing form with language reveals a division of prescriptive labour. Flowcharts prescribe twice over: two thirds name an entry point and two thirds command the reader. Tier lists barely command but they grade — they hold almost all the named tiers and value hierarchies. Most telling, text lists, the visually poorest form, compensate entirely through language: 73 % use imperatives, the highest lexical score in the corpus. Grids do neither, with a median score of zero. A distributed function Gamification is not a property of the chart. It is shared between layout and utterance, and each form picks its channel.
Figure 4 · typology, put to the test
Stacked bars comparing two emergent clusters against the paper's three theoretical types, adjusted Rand index 0.12.
Unsupervised clustering on the 19 measured variables — never on the coded labels — returns two groups, not three: 222 mute catalogues and 42 talkative, prescriptive guides. Confronted with the paper’s three types, the adjusted Rand index is 0.12; against the wiki’s thematic categories it falls to 0.02. The partition therefore recovers neither visual appearance nor subject matter. The theoretical distinction that survives measurement is the one between guiding and cataloguing; “playful cartographies” do not separate out as a third empirical family. ARI 0.12 The data isolate a dimension of their own — the regime of utterance — irreducible to both style and topic.
Figure 5 · the canon as a network
Co-occurrence network of the 124 core authors, forming a single dense central mass.
Two authors are linked when they appear on the same chart. Of the 3 006 authors recovered after cleaning, 78 % appear on a single chart — but the 124 who reach five or more form a block whose density is 56 %: more than half of all possible pairs actually co-occur. There is one central mass, not a federation of specialised canons. One block, not many The chart does not recommend books. It reproduces a solidary list — which is precisely how a device of conformity works.
Figure 6 · the canon in numbers
Bar chart of the twenty authors present on the most charts: Kafka and McCarthy 18, Dostoevsky and Plato 17, Herbert, Hesse, Nietzsche, Orwell and Shakespeare 16.
The twenty authors present on the largest number of charts. Kafka and McCarthy lead with 18 each, ahead of Dostoevsky and Plato at 17 — and the list is remarkably consistent with what /lit/ says about itself: modernist difficulty, Russian pessimism, one ancient philosopher, and Frank Herbert as the board’s permitted genre author. Note the ceiling: no author reaches even 7 % of the corpus. Recognition on /lit/ is concentrated in a very small circle, but no single name is obligatory.
Analysis · gamification

Four recurring gamified mechanics.

Charts were first coded against the game-design elements described by Deterding, Dixon, Khaled & Nacke (2011). Each mechanic is given below with what the corpus-wide measurement now says about it.

Mechanic · 01

Progression

From novice to expert: the reader advances level by level along a literary path, choosing to stop, continue or branch off.

Measured · 26 % prescribe an order, 17 % name an entry point

Mechanic · 02

Symbolic reward

Mastery of a corpus is staged as an achievement — becoming “proficient”, earned through the “grind” of reading.

Measured · rarer than expected — 6 % promise a status, 4 % name tiers

Mechanic · 03

Cognitive simplification

Ideological complexity is turned into clear visual schemas: arrows, tiers, categories and colour codes.

Measured · only 13 % actually draw arrows; simplification is mostly typographic

Mechanic · 04

Community identification

A ready-made “panoply” of authors that rewards the status of the initiate — the “real reader” inside a closed circle.

Measured indirectly · a 56 %-dense canon block is what “the panoply” looks like

Analysis · framing structures

Three types of chart — and what survived the test.

Beyond their shared mechanics, the charts were read as falling into three framing structures. Clustering the full corpus supports the first two and questions the third.

Type 01 · learning & hierarchy

Initiatory guides

Teaching codes — arrows, chronological order, practical advice. A graded, gamified learning experience where each reading unlocks the next and builds a community of self-taught experts.

Confirmed · 53 charts, and the pole that structures cluster 1

Type 02 · prescription

Cultural canons

A closed ideological library (traditionalist, libertarian, communist…) whose works reinforce one another. The chart becomes a framing device (Entman): visual hierarchy replaces rational deliberation.

Confirmed & dominant · 197 charts, 75 % of the corpus

Type 03 · exploration & simulation

Playful cartographies

Decision trees and RPG-style skill trees where the user “chooses” a route. The apparent freedom of the player masks an orientation framed in advance by the chart’s author.

Not separable · 14 charts; striking individually, not a statistical family

The object · a few charts

Three charts, one gamified form.

The three examples below span the ideological spectrum — libertarian, philosophical and purely literary — yet share the same gamified grammar. That is precisely the argument: the message lies as much in the form of the chart as in its content. Click a chart to open it full size.

Charts reproduced from the study’s corpus for the sole purpose of academic analysis. They are user-made artefacts collected on 4chan’s /lit/ board and do not reflect the author’s views.

Corpus & method · built and run

A pipeline that runs on a laptop.

Everything below was written, run and measured — not planned. The constraint was deliberate: no remote API, no key, no per-use cost, so the whole analysis can be replayed by anyone on an ordinary machine. Where a step turned out not to work, it was measured and dropped rather than quietly kept.

Step · 01

Text extraction — 264 charts in 11 minutes

Charts are huge images (up to 7 600 px) whose text is small. Running OCR on a downscaled image loses the captions entirely, so each chart is cut into horizontal bands and read at full resolution. Two engines are supported: the OS-native one by default, and Surya, fully open source, as a control that reproduces the same result on any system — 75× slower, which is why it is a control and not the default.

Step · 02

Structure by geometry, not by model

Which lines belong to the same block, and which heading governs them, is a question of geometry: the OCR coordinates already contain the answer. Reconstructing sections deterministically takes under a second per chart, can hallucinate nothing, and returns exactly the same result on every run. The same task handed to a local vision model took eleven minutes per image and mistook section headings for books.

Step · 03

Gamification markers — a versioned lexicon

Nine registers are searched by regular expression over each chart’s full text, and every hit is stored with the line that produced it. The lexicon is a plain editable file, not code: it is the study’s instrument, so it must be open to challenge. The evidence trail matters — on a Dante chart, the 27 “imperatives” turn out to be quoted verses of the Commedia, not the chart addressing its reader.

Step · 04

Emergent typology, not an imposed one

Asking a model “what type of chart is this?” imposes a taxonomy decided in advance. Instead each chart is described by 19 measurable variables — proportions, text density, row and column regularity, typographic dispersion, line length, plus the nine markers — and the groups are allowed to emerge. The coded forms and the wiki’s categories are held back to judge the result, never to build it.

Step · 05

Author networks, after a cleaning pass

Author names come from an “Author: Title” pattern that also produces false authors — History, ECONOMICS — and split forms, where Kafka and Franz Kafka count as two nodes and halve each other’s centrality. Both are corrected before any measurement, and every discarded form and merged variant is logged for inspection.

Step · 06

Reliability, reported rather than assumed

Automated form classification was tested against a full manual coding of all 264 charts and rejected on the evidence: Krippendorff’s α = 0.078 for layout type, barely above an annotator who answers the same label every time. The same model detects drawn arrows at α = 0.842 and is kept for that alone. A negative result, measured and reported, is worth more than an unexamined variable.

Known limits

Limit 01

Books that are only covers

Only 24 % of extracted works carry an identifiable author. On charts where books appear solely as cover images with no caption, no OCR can recover a title — those charts are structurally under-represented in the network.

Limit 02

Provenance of the form coding

The 264 forms were coded by a large vision-language model, not by a human and not by the local pipeline. It is external validation and is declared as such; the reproducible pipeline itself stays fully local.

Limit 03

No repost counts

The corpus comes from the wiki, not from threads, so actual recirculation cannot be measured. Content overlap between the annual “Top 100” editions is a floor, not a ranking.

Limit 04

Text volume bias

Lexical marker scores grow mechanically with the amount of text on a chart. Comparisons between forms of very different verbosity require normalisation first.

Text extraction
Vision (macOS) Surya · open-source control Pillow band tiling
Structure & markers
pure geometry versioned lexicon Qwen2.5-VL · Ollama Llama 3.1 · local
Analysis
pandas scikit-learn networkx matplotlib
Validation & reproducibility
Krippendorff’s α adjusted Rand index fixed seeds conda · Jupyter · git
Conclusion

Findings and an opening.

Findings

Radicalism as an economy of symbols

The visual, gamified form of the charts makes literary radicalism seductive, shareable and culturally valued. Diffusion owes less to explicit militancy than to an economy of symbols and provocation that rewards subcultural belonging — radicalism that “exposes without imposing”.

Findings · measured

Conformity before initiation

The corpus-wide measurement shifts the emphasis. Charts classify far more than they guide: three quarters are catalogues, and the canon behaves as a single 56 %-dense block where citing one author entails the others. Where prescription does appear, it travels through language as much as through layout.

Method

What a negative result is worth

Two claims were tested and revised on the evidence: a local vision model cannot classify chart forms (α = 0.078), and the third theoretical type does not separate out empirically (ARI = 0.12). Reporting both is what makes the remaining findings credible.

Opening

Next research

Almost no scholarship targets /lit/ directly. Two threads follow: analysing 4chan’s own structure — code, UI/UX, affordances — and tracking the annual “Top 100” series as a time span, to see whether the canon actually moves.