The cross-section test: a topic-independent layer exists — but it's a skeleton, not prose grammar
GPT-5.5's closing question: are the rigid word forms an encoding of grammar, or the grammar of the generator itself? To attack it we needed a test we hadn't run — every prior test was within one section. The real grammar test is across sections.
Plain-English first:
A real language's glue words (the, and, of, et, in) show up everywhere — evenly on a cooking page, a war page, an astronomy page. They don't care about the topic. So: does Voynich have any words that appear evenly across all of its sections at once?
We normalized the words five ways — raw, merge-near-twins, strip-prefix, strip-suffix, and keep-only-the-core — and for each, counted how many frequent tokens spread evenly across all 8 sections.
The answer is "some, but modest, and not the top words." ~12–16% of frequent tokens are genuinely topic-independent. A dumb page-by-page generator would produce almost none of these — so there is shared, cross-topic machinery. But the manuscript's most common words stay glued to their own sections, which is backwards from a real language, where the commonest words are exactly the topic-independent glue.
The most striking bit: when you strip the affixes and keep only the cores, short forms like ok, sh, r, yk recur evenly across every section. That looks like a structural skeleton — topic-independent core elements, with the topic-specific variation carried by the prefixes/suffixes wrapped around them.
The numbers
normalization vocab freq>=30 cross-section-uniform examples (count)
raw 8129 195 24 r(167) okal(153) saiin(119) kar(59)
lev-stem 3475 140 16 okchey(155) ches(123) oiin(104)
prefix-strip 5761 151 22 r(167) kar(92) chey(84) kal(62)
suffix-strip 5808 165 24 r(167) chcth(91) chea(82) dal(61)
core-only 4162 127 20 yk(209) r(167) sh(119) ok(99) ii(86)
(derived suffixes: keey hedy eedy aiin edy dy ol y ...; prefixes: qok cho she qo ch o .... "uniform" = spread across the 8 sections no more than chance, z<2.)
Three findings:
1. A topic-independent layer exists. ~20–24 frequent tokens spread evenly across all sections; a strict per-section generator would yield ≈0. So this is not a dumb section-conditioned generator — there is shared machinery.
2. It is not shaped like natural language. In Latin the top words are the uniform function words. In Voynich the top words (daiin, ol, chedy) are section-bound, and the topic-independent layer is a mid-frequency set. That's backwards from prose.
3. Core-only is the tell. Stripping affixes exposes short cores (ok, sh, r, yk) that recur across every section — a structural skeleton, with topic variation in the affixes around it.
Caveats: affixes are empirically derived (crude); z<2 is a generous "uniform" threshold; the shortest cores (yk, ii) may be partly stripping artifacts. So, stated carefully: a modest topic-independent layer exists across all normalizations — strongest as short cores — but it does not sit where natural-language function words sit.
Answer to the question: there's a real topic-independent layer, but it's structural-skeleton-shaped, not prose-grammar-shaped — which is what a cipher / constructed / rule-mediated system looks like, not ordinary writing. The matrix needle holds at rule-mediated: not gibberish, not a dumb generator, not transparent prose; a system with global structural elements plus heavy section-conditioning, whose word-machinery stays too rigid and self-similar to read as ordinary language.
Sources - ZL EVA transliteration — voynich.nu/data/ZL3b-n.txt - Stolfi, prefix–core–suffix ("crust–mantle–core") word paradigm - Currier (1976), A/B languages; section labels per the ZL metadata








