Infrastructure as Ideology
Every dataset is a border.

The tokenizer isn't a technical detail. It is the moment where language becomes readable to a statistical machine — and the moment where centuries of unequal access to writing, publication, and infrastructure get encoded as efficiency. The corpus is not neutral. It is the residue.
A byte-pair encoding table has a class politics. A subword vocabulary is a curriculum. What the model finds legible, the model rewards. What it doesn't, it renames.

Statistical efficiency is redistribution disguised as neutrality.
Radix, Donori, 2021. At the village's Contemporaryfestival, a Raspberry Pi running the RiTa library generated its first autonomous sentences on a table in the main square. Training set: the recorded history of Donori, spliced with a large corpus of spam email. The output — hybrids of parish record and enlargement offer, of local saint and Nigerian prince — read as poetry, and made everyone laugh.
Later, a series of notebooks containing false stories about the village were left in a local bar, waiting for the older residents to find them. Some did. Some corrected the dates.
In 2021 this was a joke. Four years later, the same procedure runs behind every consumer-grade chatbot at industrial scale — trained on the same intact history, seeded with the same spam — and nobody is laughing anymore.

From Radix (2021), Contemporary festival, Donori (IT). Raspberry Pi + RiTa library, trained on the parish records of Donori and a spam corpus.