Skip to main content

The Book Nobody Can Read: I Gave the Voynich Manuscript a Lie Detector Test

There is a book in a library at Yale University that has been driving people slightly mad for over a century. It is about 600 years old, written on real parchment, in an alphabet that exists nowhere else on Earth. It is full of drawings of plants that no botanist can identify, star charts that match no sky, and small naked women relaxing in tubs of green liquid connected by strange plumbing. Nobody can read a single word of it.

It is called the Voynich manuscript, and I have been obsessed with it for decades. I love old writing and I love a good mystery, and this thing is both cranked up to eleven. The world's best codebreakers tried to crack it. The team that broke Japanese naval codes in World War II spent years on it and got nowhere. Every few months someone announces they have finally solved it, and every few months that solution quietly falls apart.

This summer I decided to stop dreaming about translating it and try something else. If you cannot read a suspect's diary, you can still put the suspect through a lie detector test. So that is what I did: together with an AI assistant I ran a full statistical interrogation of the manuscript, using the same math that powers file compression and language models. It took three long evenings, a couple hundred pages of results, and then a few more evenings after two reviewers took my write-up apart and made me redo things. It ended somewhere I genuinely did not expect, and not everything I believed halfway through survived.

A botanical page from the Voynich manuscript showing a plant with blue and red star-shaped flowers, surrounded by undeciphered handwriting
A typical botanical page. Nobody has ever identified this plant, and nobody has ever read the text next to it. Image: Beinecke Library, Yale University (public domain).

You cannot translate it, but you can measure it

Here is the trick that makes this possible. Since the 1990s, researchers have painstakingly typed out the entire manuscript letter by letter, using a system that assigns a normal Latin character to each Voynich symbol. The words come out looking like daiin, chedy, qokeedy and shol. You still cannot read them, but a computer can count them. Just over 38,000 of them, of which 34,000 sit in running paragraphs rather than in labels stuck next to a drawing.

And once you can count, you can ask sneaky questions. Not "what does this word mean?" but "how does this text behave?" Because real language, any real language, has a statistical fingerprint. It does not matter if it is Dutch, Latin, Arabic or Klingon fan fiction: languages obey certain laws of information, the same way falling objects obey gravity.

To have something to compare against, we also processed a small library of texts from roughly the same era: the Latin Bible, Chaucer's Canterbury Tales in Middle English, a German knight epic, an Old French chanson, Hebrew, Arabic, and Dante in medieval Italian. Same measurements, same code, fair fight.

The weirdest statistic in any book on Earth

Imagine playing Wheel of Fortune with a text. I show you a letter and you guess the next one. In English, if I show you a "q", you will guess "u" and you will usually be right. Information theory puts a precise number on this guessability. It is called entropy: high entropy means surprising, low entropy means predictable.

Every language we tested lands in the same comfortable zone. Knowing one letter, the next letter carries a bit more than 3 bits of surprise. That is true for Latin, English, German, Hebrew, all of them. It is one of those quietly universal things about human language.

The Voynich manuscript scores 2.1. That is not just low, that is absurdly low. The letters are so predictable that the script almost writes itself, like a song where you can hum the next note before you hear it. And here is the paradox that has haunted researchers for fifty years: the words are completely normal. The vocabulary size, the way word frequencies are distributed, the rate at which new words appear as the text goes on: all of it lands exactly where a real language should land. Normal words, built from letters that are far too obedient. No natural language does that. Not one.

We also tested the most popular escape hatch for this paradox: maybe it is Latin written in heavily abbreviated medieval shorthand. So we built a simulated shorthand Latin, using seventeen genuine abbreviations out of Cappelli, the dictionary that scholars use to decode medieval scribbles, and measured it. The result went the wrong way for the theory. Shortening the words pushed the letter surprise up, from 3.31 to 3.40 bits, not down toward 2.1. That makes sense once you think about it: abbreviation packs more meaning into fewer symbols, so each symbol has to work harder and becomes harder to guess. The manuscript does the exact opposite. The escape hatch slammed shut.

Words that refuse to talk to each other

The next test is my favorite, because it is so simple and so brutal. In real language, words gossip about each other over long distances. If the word "doctor" appears in a sentence, the odds that "hospital" or "patient" shows up a few words later go way up. You can measure this chatter mathematically at distance 1, distance 2, distance 10, distance 50, and you get a smooth curve: strong for neighbors, slowly fading as words get further apart. Every language produces that curve. It is the sound of sentences meaning something.

Chart comparing how word correlation fades with distance in the Voynich manuscript, Latin, Middle English and two computer models
The lie detector chart. Latin (blue) and Middle English (green) fade out slowly and smoothly, like real conversations. The Voynich manuscript (black) collapses almost instantly after distance 1. The red line is our generator, which we will get to. Measured on the ZL3b transliteration.

The Voynich curve is like nothing else. Words do influence their direct neighbor, a little, about five times weaker than in Latin. And then, at distance 2, the signal falls off a cliff. Latin still has strong chatter at distance 2. The manuscript has almost none, fifteen times less. There is no middle range at all, no zone where grammar hands over to storyline. Whatever this text is doing, its words are not building sentences. Each word barely knows its neighbor exists, and has never heard of the word ten places back.

Notice that the two oddities point in opposite directions, and that is not a contradiction, it is the signature. At the letter level the manuscript is far too predictable compared to any language. At the word-to-word level it is far too disconnected. Those are two different measurements on two different scales, and a real language sits in the middle on both.

One more thing had to be checked before I trusted any of this. Every number above depends on someone having typed the manuscript out, and there is no guarantee that two readers would interpret the same loops and strokes the same way. So I repeated the main measurements on three more versions of the same book: one typed out by a different researcher, one written in a rival alphabet that treats certain symbol pairs as single characters, and one where I merged those pairs myself to see what that alone would do. The letter predictability shifted a bit, from 2.13 to 2.50 bits depending on the alphabet, which is what you would expect and still nowhere near the 3 bits of a real language. The collapse of the word connections after distance 1 showed up in all four versions. The independent typist's version actually gave the weakest neighbor signal of the four, which reassured me, because if my headline result had come out of one person's guesses about where words end, that is exactly where it would have surfaced.

The chart above, in my view, kills the dream of translating this book by reading its sentences. You cannot decode a message from word order if the word order carries no message. There is one loophole, and it is interesting enough that I have given it its own section further down.

The scribe's wandering eye

So if the words are not sentences, where do they come from? The data has a strong opinion about that too, and this is where it gets almost creepy, because you can feel a human being on the other side of the parchment.

Voynich words love to repeat. Not just exact repeats like the famous daiin daiin daiin, but near-repeats: chedy, then shedy, then chedy again, small variations circling each other like moths. We measured where these echoes happen, and they cluster tightly around each other on the page. Then we found the smoking gun: at the exact same word distance, the echo rate is about 40 percent higher inside a single page than it is across a page boundary. The words do not know about pages. Pages are physical objects. But the echoes stop at the edge of the page anyway.

There is a very human explanation. Picture a scribe writing, dipping the quill, and glancing up at what he just wrote. His eye lands on a word a few lines back. He copies it, changes a stroke here, swaps a symbol there, writes it down, and moves on. When he turns to a fresh page, yesterday's words are out of sight, and out of the statistics. We tested this idea hard, with randomized controls, and it held: the reuse is anchored to what was visible on the page, not to how long ago it was written. And it does not depend on any single individual, which matters, because the manuscript was written by more than one person, as we will get to.

The physical page shapes the text in other ways too. On the botanical pages you can see that the drawings came first: the text squeezes itself around stems and leaves, into whatever room is left over. In a normal book the words are the point and the layout serves them. Here it looks like the reverse, and our measurements agree: the line and the page behave like hard units of production, more real to the writers than any sentence.

Credit where it belongs: the core idea that each Voynich word is a modified copy of an earlier word was proposed by Torsten Timm and Andreas Schinner in a paper published in 2019, and other independent tinkerers have reproduced it since. What our investigation adds is the evidence that the copying is anchored to the visible page rather than to time, the syllable grammar of the building blocks, and the boundary tests that tie it all to the physical book.

A page from the Voynich manuscript showing small human figures bathing in green pools connected by tube-like structures
The famous bathing section. Charming, weird, and according to our tests not a spa manual: these pages show the least recipe-like structure in the whole book. Image: Beinecke Library, Yale University (public domain).

Testing the fun theories

We also took two beloved fan theories and gave them the same treatment.

Theory one: the book is a practical health guide, maybe a bathhouse manual or a women's medicine handbook. If that were true, the bathing section should be the most recipe-like part of the book, full of repeated formulas like "take X, boil, apply". We measured exactly that, and found precisely the opposite. The bathing pages are the least formulaic section of the entire manuscript. The plant pages score highest, and they do so for a boring reason: one page, one plant, one writing session.

Theory two: the text was generated with a mechanical gadget, like a volvelle, a rotating paper disk with symbols on it. Rotating paper instruments did exist by then, Ramon Llull was building them around 1300, although the famous cipher disk of Alberti only appeared around 1467, decades after the parchment was made. So the theory already has a dating problem in its strongest form. But we did not need the history books to settle it. A rotating disk would leave a rhythm in the text, a repeating pattern tied to how often the disk turns per line. That is measurable. We ran the frequency analysis, the same kind you would use to find a hidden beat in a noisy recording. There is no beat. No rhythm, no period, nothing. Whatever the writers used, it was not a machine turning on a schedule. It was a mind, drifting.

The twist nobody saw coming

By this point I was fairly convinced we were looking at an elaborate fake, and I expected the remaining tests to be a victory lap. Instead, the manuscript pulled one more rabbit out of its hood.

The text is built from about 426 recurring building blocks, little chunks like qok, che, dy, that snap together in fixed positions: openers, middles, enders. That count lands somewhere between 420 and 460 depending on which alphabet you count in, which does not change anything that follows. We compared the statistics of these blocks to the syllables of real languages, chopped up with exactly the same tools. They behave like syllables and not like letters. The information content per block is 7.86 bits for Voynich, against 7.88 for Dante's Italian and 7.94 for Latin. On the network of which block follows which, viewed as a mathematical graph, the manuscript leans toward Italian rather than Latin, but it sits outside both languages by a fair margin.

I have to flag a correction here, because the first version of this article claimed more. It said the manuscript sat closer to Italian than Italian sits to Latin, which would have been remarkable. When a reviewer asked whether that depended on how I was splitting the reference languages into syllables, I redid it with proper phonological rules instead of the quick pattern I had used. Latin changed a lot, the two real languages moved close together, and that particular comparison flipped. The finding that survives is the weaker one above: syllable-sized blocks, leaning Romance, not drawn from any syllabary I can match.

Even in that weaker form it stopped me. Inside each word, the manuscript has the texture of a real language. Between words, it has nothing at all. It is a body with good bones and no heartbeat. Whoever wrote this had absorbed the sound structure of a spoken tongue, the rhythm of syllables, the feel of word-building, and then wrote 200 pages of it without sentences.

So we built a ghost writer

A theory is only worth something if you can make it produce evidence. So we built a small program that writes Voynichese the way our data says the scribes did: it knows the 426 blocks and their snapping rules, it glances back at its own page and copies recent words with small mutations, it avoids repeating a word twice in a row, and its taste drifts slowly from page to page. It even writes on the real page layout, the same number of lines per page and words per line as the manuscript. Then we measured its output with the same 14 statistical tests as the real thing.

It matched nine of them, several to the second decimal: the word entropy, the percentage of words that appear exactly once, the rate of near-twin words, the average word length. It also produced that sharp drop after distance 1 in the lie detector chart, although that particular match turned out to be the shakiest one on the list, for reasons I will come to. Here is a line it wrote:

dqokain chedy ycthy ol sarol shcthy

Show that to anyone who has stared at the manuscript and they will nod: yes, that looks like Voynichese. And this line, at least, provably contains no message, because we know exactly how it was made.

I should be precise about what this does and does not prove. "Looks like Voynichese" is not a matter of taste here, it is a scorecard: fourteen measurable properties, from entropy to word repetition patterns, checked by the same code that measured the real manuscript.

Then a reviewer asked an awkward question. A program like this uses random numbers, so how much of that score was luck? Fair question, so I ran it a hundred more times with different random starting points. Nine matches turned out to be the most common outcome, which was a relief, but the hundred runs also drew a much sharper picture of where the program is good and where it is not. Six of the fourteen properties it gets right every single time, and they are all about the shape of words and the shape of the vocabulary. The ones it gets wrong are the ones about sequence. And the strength of the link between neighboring words, the match I was proudest of, turned out to sit at the very edge of the program's own range: a typical run overshoots the real manuscript by about a third. That is worth saying out loud, because it means my little scribe imitates handwriting habits well and imitates whatever governs word order badly.

And a matching generator does not prove the manuscript is meaningless. Statistics can never prove that, and I will not pretend otherwise. What it proves is that no hidden message is needed to explain what is on the page. If there is meaning in this book, our measurements say it is not living in the order of the words. Where it could live instead is the subject of the next section.

A dense text page from the recipes section of the Voynich manuscript, with star symbols marking paragraphs
The recipes section: pure text, paragraph after paragraph, each marked with a little star. Our generator produces pages with the same statistical fingerprint. Image: Beinecke Library, Yale University (public domain).

The loophole I cannot close

Here I have to argue against myself, because there is one kind of code that survives every test above, and leaving it out would be exactly the sort of overreach I have been complaining about the whole way down this page.

Picture a cipher where every letter of the hidden text has several possible stand-ins, and you pick one at random each time you write it. An "a" becomes qok on this line and chedy on the next. These are called verbose homophonic ciphers, they existed in fifteenth century Italy, they can be worked by hand from a set of tables, and in 2025 a researcher showed that one built for Latin and Italian produces text with several of the manuscript's strangest statistics.

Now think about what such a cipher does to my measurements. The same plaintext word comes out different every time it is written, so the vocabulary looks huge and full of near-twins. The links between words get shredded, because the connection lives in the hidden text and my tools only ever see the disguise. In other words, a verbose cipher predicts the empty word order that I measured. My strongest finding is not evidence against that theory. It is a prediction of it, and it took me a while to admit that.

So the sentence I would have loved to write, that no cipher can hide a message in a text whose word order carries nothing, is simply false. It is not a free win for the cipher camp either. To fill 38,000 Voynich words the hidden text would have to be roughly a third that length, and someone would have needed to run an elaborate multi-table procedure by hand, consistently, across 200 pages and several different scribes, without slipping. An independent study in 2026 also found that this cipher misses other fingerprints of the manuscript that have to be matched at the same time. But it is not excluded, and I am not going to pretend it is.

What survives is narrower than the headline I wanted, and I have come to prefer it: this book does not build sentences out of word order. If a message is in there, it is hidden inside the words, not between them.

So what is this thing?

Here is where the evidence points, with the loophole above kept firmly in view. The Voynich manuscript is, in all likelihood, not a language, and not a code of the ordinary sort where one symbol stands for one letter. It reads as a performance, and probably a group performance. Handwriting experts who studied the manuscript, most recently the paleographer Lisa Fagin Davis, found that it was written by several different hands, possibly as many as five. That makes the statistics stranger, not simpler: multiple people, fluent in the same invented script, producing text with the same statistical fingerprint. This was not one obsessive with a quill. It was a system that could be learned, shared and taught, with the syllable-feel of a real tongue built into it, and a small team filled 200 parchment pages with it. The discipline is astonishing, and it holds up under tests its makers could not have imagined. Fake writing, done with total mastery by a group, is still a masterpiece of something.

Why would anyone do that? The vellum itself is workmanlike rather than luxurious, with flaws and rough repairs, so this was no royal commission on spotless calfskin. But 200 pages of parchment plus months or years of skilled labor by several writers was still a serious investment for whoever made it. A mysterious book full of secret knowledge was worth real money to the right buyer, and one emperor reportedly paid a fortune for this one, though that story reaches us secondhand through a letter written decades later. Or maybe it was something stranger and more sincere, the written equivalent of speaking in tongues, practiced by a small circle. The statistics cannot tell us what was in the writers' heads. They can only tell us what is on the page, and what is not.

Will a future super-AI ever crack it? I think the honest answer is that it will not, and not because AI will stop getting smarter. The information a translator needs is not in the order of the words, and no amount of intelligence recovers information that is not there. The only route left is the verbose cipher, and that one is nastier than it looks: the machine would be guessing at a hidden text with almost nothing to check its guesses against, which is how you get a thousand confident and mutually contradictory solutions. If richer scans someday reveal writing we cannot currently see, or the loose ends surprise me, I will happily reopen the case. The manuscript has embarrassed confident people for a century and I have no wish to be the last one.

But I will admit something. Part of me spent decades hoping to one day read its secrets. And when the final numbers came in, I was not disappointed for long. A 600-year-old book whose sentences turn out not to be sentences, yet which is built so precisely that it takes information theory, seven medieval reference texts and a synthetic ghost writer to show it: that might be a better story than any translation could have been. The mystery was never in the words. It was in the hand that wrote them.

The full technical paper: everything above is written up formally, with the exact measurements, the noise floors, the permutation controls and the generative model, in Word-Order Information, Page-Anchored Repetition, and a Data-Estimated Generative Model for the Voynich Manuscript (PDF, 15 pages).