The Book Nobody Can Read: I Gave the Voynich Manuscript a Lie Detector Test
There is a book in a library at Yale University that has been driving people slightly mad for over a century. It is about 600 years old, written on real parchment, in an alphabet that exists nowhere else on Earth. It is full of drawings of plants that no botanist can identify, star charts that match no sky, and small naked women relaxing in tubs of green liquid connected by strange plumbing. Nobody can read a single word of it.
It is called the Voynich manuscript, and I have been obsessed with it for decades. I love old writing and I love a good mystery, and this thing is both cranked up to eleven. The world's best codebreakers tried to crack it. The team that broke Japanese naval codes in World War II spent years on it and got nowhere. Every few months someone announces they have finally solved it, and every few months that solution quietly falls apart.
This summer I decided to stop dreaming about translating it and try something else. If you cannot read a suspect's diary, you can still put the suspect through a lie detector test. So that is what I did: together with an AI assistant I ran a full statistical interrogation of the manuscript, using the same math that powers file compression and language models. It took three long evenings, a couple hundred pages of results, and it ended somewhere I genuinely did not expect.
You cannot translate it, but you can measure it
Here is the trick that makes this possible. Since the 1990s, researchers have painstakingly typed out the entire manuscript letter by letter, using a system that assigns a normal Latin character to each Voynich symbol. The words come out looking like daiin, chedy, qokeedy and shol. You still cannot read them, but a computer can count them. All 37,000 of them.
And once you can count, you can ask sneaky questions. Not "what does this word mean?" but "how does this text behave?" Because real language, any real language, has a statistical fingerprint. It does not matter if it is Dutch, Latin, Arabic or Klingon fan fiction: languages obey certain laws of information, the same way falling objects obey gravity.
To have something to compare against, we also processed a small library of texts from roughly the same era: the Latin Bible, Chaucer's Canterbury Tales in Middle English, a German knight epic, an Old French chanson, Hebrew, Arabic, and Dante in medieval Italian. Same measurements, same code, fair fight.
The weirdest statistic in any book on Earth
Imagine playing Wheel of Fortune with a text. I show you a letter and you guess the next one. In English, if I show you a "q", you will guess "u" and you will usually be right. Information theory puts a precise number on this guessability. It is called entropy: high entropy means surprising, low entropy means predictable.
Every language we tested lands in the same comfortable zone. Knowing one letter, the next letter carries a bit more than 3 bits of surprise. That is true for Latin, English, German, Hebrew, all of them. It is one of those quietly universal things about human language.
The Voynich manuscript scores 2.1. That is not just low, that is absurdly low. The letters are so predictable that the script almost writes itself, like a song where you can hum the next note before you hear it. And here is the paradox that has haunted researchers for fifty years: the words are completely normal. The vocabulary size, the way word frequencies are distributed, the rate at which new words appear as the text goes on: all of it lands exactly where a real language should land. Normal words, built from letters that are far too obedient. No natural language does that. Not one.
We also tested the most popular escape hatch for this paradox: maybe it is Latin written in heavily abbreviated medieval shorthand. So we built a simulated shorthand Latin and measured it. Its letters stayed stubbornly surprising, at around 3.2 bits. The escape hatch slammed shut.
Words that refuse to talk to each other
The next test is my favorite, because it is so simple and so brutal. In real language, words gossip about each other over long distances. If the word "doctor" appears in a sentence, the odds that "hospital" or "patient" shows up a few words later go way up. You can measure this chatter mathematically at distance 1, distance 2, distance 10, distance 50, and you get a smooth curve: strong for neighbors, slowly fading as words get further apart. Every language produces that curve. It is the sound of sentences meaning something.
The Voynich curve is like nothing else. Words do influence their direct neighbor, a little, about five times weaker than in Latin. And then, at distance 2, the signal falls off a cliff. Latin still has strong chatter at distance 2. The manuscript has almost none, fifteen times less. There is no middle range at all, no zone where grammar hands over to storyline. Whatever this text is doing, its words are not building sentences. Each word barely knows its neighbor exists, and has never heard of the word ten places back.
That single graph, in my view, is the end of the dream of translation. You cannot decode a message from word order if the word order carries no message.
The scribe's wandering eye
So if the words are not sentences, where do they come from? The data has a strong opinion about that too, and this is where it gets almost creepy, because you can feel a human being on the other side of the parchment.
Voynich words love to repeat. Not just exact repeats like the famous daiin daiin daiin, but near-repeats: chedy, then shedy, then chedy again, small variations circling each other like moths. We measured where these echoes happen, and they cluster tightly around each other on the page. Then we found the smoking gun: at the exact same word distance, the echo rate drops by a third the moment you cross a page boundary. The words do not know about pages. Pages are physical objects. But the echoes stop at the edge of the page anyway.
There is a very human explanation. Picture a scribe writing, dipping the quill, and glancing up at what he just wrote. His eye lands on a word a few lines back. He copies it, changes a stroke here, swaps a symbol there, writes it down, and moves on. When he turns to a fresh page, yesterday's words are out of sight, and out of the statistics. We tested this idea hard, with randomized controls, and it held: the reuse is anchored to what was visible on the page, not to how long ago it was written.
Testing the fun theories
We also took two beloved fan theories and gave them the same treatment.
Theory one: the book is a practical health guide, maybe a bathhouse manual or a women's medicine handbook. If that were true, the bathing section should be the most recipe-like part of the book, full of repeated formulas like "take X, boil, apply". We measured exactly that, and found precisely the opposite. The bathing pages are the least formulaic section of the entire manuscript. The plant pages score highest, and they do so for a boring reason: one page, one plant, one writing session.
Theory two: the text was generated with a mechanical gadget, like a volvelle, a rotating paper disk with symbols on it that some suspect Renaissance hoaxers used. Here is the cool part: a rotating disk would leave a rhythm in the text, a repeating pattern tied to how often the disk turns per line. That is measurable. We ran the frequency analysis, the same kind you would use to find a hidden beat in a noisy recording. There is no beat. No rhythm, no period, nothing. Whatever the scribe used, it was not a machine turning on a schedule. It was a mind, drifting.
The twist nobody saw coming
By this point I was fairly convinced we were looking at an elaborate fake, and I expected the remaining tests to be a victory lap. Instead, the manuscript pulled one more rabbit out of its hood.
The text is built from about 426 recurring building blocks, little chunks like qok, che, dy, that snap together in fixed positions: openers, middles, enders. We compared the statistics of these blocks to the syllables of real languages, chopped up with exactly the same tools. The match with medieval Italian is uncanny. The information content per block: 7.86 bits for Voynich, 7.89 for Dante. The network of which block follows which, viewed as a mathematical graph, sits closer to Italian than Italian sits to Latin.
Read that again, because it stunned me: inside each word, the manuscript behaves like a Romance language. Between words, it behaves like nothing at all. It is a body with perfect bones and no heartbeat. Whoever wrote this had absorbed the sound structure of a real language, the rhythm of syllables, the feel of word-building, and then wrote 200 pages of it without sentences.
So we built a ghost writer
A theory is only worth something if you can make it produce evidence. So we built a small program that writes Voynichese the way our data says the scribe did: it knows the 426 blocks and their snapping rules, it glances back at its own page and copies recent words with small mutations, it avoids repeating a word twice in a row, and its taste drifts slowly from page to page. Then we measured its output with the same 13 statistical tests as the real thing.
It matched nine of them, several to the second decimal. The word entropy, the percentage of words that appear exactly once, the rate of near-twin words, the strange fingerprint in that lie detector chart: all reproduced. Here is a line it wrote:
dqokain chedy ycthy ol sarol shcthy
Show that to anyone who has stared at the manuscript and they will nod: yes, that is Voynichese. No secret message went into it. None can come out of it.
So what is this thing?
Here is where the evidence points. The Voynich manuscript is, in all likelihood, not a language and not a code. It is a performance. Somebody in the early 1400s trained themselves in an invented script until they could write it fluently, with the syllable-feel of a real tongue in their fingers, and then filled 200 expensive parchment pages with it. Not lazily: the discipline is astonishing, and it holds up under statistical tests its maker could not have imagined. Fake writing, done with total mastery, is still a masterpiece of something.
Why would anyone do that? Parchment cost a fortune, so "just for fun" feels thin. A mysterious book full of secret knowledge was worth real money to the right buyer, and there is documentation that an emperor later paid a fortune for this one. Or maybe it was something stranger and more sincere, the written equivalent of speaking in tongues. The statistics cannot tell us what was in the writer's head. They can only tell us what is on the page, and what is not.
Will a future super-AI ever crack it? I asked myself the same question, and I think the honest answer is no, not because AI will not get smarter, but because you cannot extract a message that was never put in. Our tests show the word order carries almost no information. Decoding it is like trying to unbake a loaf of bread back into a recipe that never existed. If richer scans someday reveal hidden text, or the few remaining loose ends surprise us, I will happily reopen the case. The manuscript has embarrassed confident people for a century, and I do not intend to be the last one.
But I will admit something. Part of me spent decades hoping to one day read its secrets. And when the final numbers came in, I was not disappointed for long. A 600-year-old book that contains no message, yet is built so well that it takes information theory, seven medieval reference texts and a synthetic ghost writer to prove it: that might be a better story than any translation could have been. The mystery was never in the words. It was in the hand that wrote them.