Are Words Enough to Understand a Language?
A 3Blue1Brown video on cross-entropy and gzip language trees, and why measuring languages by their spelling might be measuring the wrong thing.
I watched a 3Blue1Brown video today: But what is cross-entropy?, the second in Grant Sanderson’s “Compression is Intelligence” series. It left me with an idea I can’t put down. Not about the main topic, which is how large language models are trained. About a throwaway example at the very beginning.
Zipping your way to a family tree of languages
The video opens with a genuinely strange 2002 paper, Language Trees and Zipping. The setup: take a pile of text documents in different languages, and, using nothing but gzip (the same compression you’d use to shrink a file on your laptop), automatically recover which languages are related to which, well enough to reconstruct something like a family tree. No linguistics. No dictionaries. Just a general-purpose compressor.
The trick is simple to state. Take two documents, A and B. Compress A on its own, and note the size. Then append a small snippet of B to A and compress that. If B’s patterns are similar to A’s, the combined file barely grows: the compressor, already tuned to A, finds B’s structure only mildly surprising. If B is very different, the file grows a lot. That extra growth becomes a distance metric between A and B, and applying it across many language pairs is enough to recover a plausible tree of language lineage, and even to identify who wrote an anonymous document.
Grant’s real point in bringing this up is that this trick is a crude stand-in for a much more precise idea: cross-entropy. If you train a code optimized for one distribution, Q, and then use it on data that actually follows a different distribution, P, the average number of bits you’ll need per symbol is the cross-entropy of Q relative to P, and that’s minimized exactly when Q equals P. Gzip’s LZ77 compressor is nowhere near an optimal code, and a document isn’t a probability distribution, but the co-compression trick is roughly asking the same question cross-entropy asks: how well does something tuned to A explain B?
The video then spends its second half showing that this exact formula, cross-entropy, is what LLM pretraining loss actually is: the average surprise, per token, of the model’s predicted distribution against the true next token drawn from the training data. Same math, two contexts, which Grant treats as math “winking” at you about a hidden connection.
What the zipping trick can’t see
Here’s where I got stuck, in a good way. The gzip trick works entirely on the written form of a language: the sequence of characters, whatever alphabet or orthography they happen to use. That’s fine for the paper’s purposes, and Grant is upfront that it’s a rough approximation. But it means the “distance” it measures is really a distance between writing systems and spelling conventions, not between the languages as spoken.
Two languages can sound almost identical and still look completely unrelated on paper: different scripts, different transliteration rules, different orthographic reforms layered on top of the same underlying sounds. And the reverse happens too: languages using the same alphabet can drift phonetically while their spelling stays superficially close, or converge in pronunciation while spelling conventions keep them looking distinct. A compression distance computed on raw text is, in a real sense, measuring the wrong layer. It can pick up genuine structure (cognates, shared morphology, loanword patterns), but it’s entangled with a historical accident: which script a community happened to standardize on.
What I want to know is what happens if you strip that accident away. If you first transcribe everything into a single universal phonetic representation (the International Phonetic Alphabet, or some other consistent phone-level encoding) and then run the same kind of compression-distance or cross-entropy analysis, do you get the same tree? A different one? Does it sharpen the boundaries the text-based version blurred, or blur the ones it sharpened?
A rough shape for the project
I don’t have this worked out, but here’s the shape it’s taking in my head:
Get phonetic data, not just text. The Language Trees and Zipping paper used Universal Declaration of Human Rights translations because they’re parallel and freely available. I’d need the phonetic equivalent: either existing IPA-transcribed corpora (something like WikiPron or PHOIBLE) or a phonemizer tool (espeak-ng and similar) run over the same parallel text, accepting that automated phonemization is itself an approximation and a source of noise to be honest about.
Redefine the distance on phone sequences. The same co-compression idea (compress A, compress A-plus-a-snippet-of-B, look at the difference), but now A and B are streams of phonemes or phonetic feature vectors instead of raw characters. This is where it starts to genuinely rhyme with the video’s second half: a compressor tuned to one language’s phonotactics (which sound sequences it allows and how often) meeting another language’s phoneme stream is a fairly direct phonetic analogue of cross-entropy between two distributions.
Compare the resulting trees. Run both versions (text-based and phonetic-based) against a known linguistic family tree (or against each other) and see where they agree, where they diverge, and whether the divergences are interpretable: script-driven false negatives, sound-shift-driven false positives, that kind of thing.
And if that holds up: a phonetic language model. This is the part I’m least sure about, but it’s the part I find most interesting. If next-token prediction over text is what makes cross-entropy loss the training signal for an LLM, what does next-phoneme prediction look like? A model that predicts the next sound rather than the next character or subword token would, in principle, be learning something closer to the actual acoustic-phonological structure of language rather than a particular community’s spelling conventions. Whether that’s useful for anything beyond satisfying curiosity (better cross-lingual transfer, maybe, or a genuinely phonetic notion of “how close are these two languages”), I don’t know yet. That’s exactly why I want to build the first, smaller piece and see what it says before speculating further.
None of this is close to a plan yet, more like the first hour of an idea, written down before I lose the thread. If the phonetic version of the zipping trick produces something interesting, I’ll write the follow-up.