The Rohonc Codex

Using AI to extend and test the Király and Tokai dictionary

Are the readings true?

The tests check the project's own readings. They were designed to break the readings. The bar for each test was declared first, before the test ran.

Nine tests were designed by the project. Then two frontier models from other labs, Gemini 2.5 Pro and Grok 4.7, were shown the write-up and asked what tests they would require. Those were run as specified.

testwhat it asksresultverdict
The transcription every other test reads, checked against an independent one:
Test 16a second, independent transcription91.1% of glyphs agree where both can be compared, against a 12.9% control, 458.5 sigma; 83.8% of the words identicalPASS
The readings:
Test 1source presence, held-out foliosA+B 16.4 sigma; 99% of K&T's own ratePASS
Test 5K&T's own sentence translationsA+B 8.6 sigma; 153% of K&TPASS
Test 6word order, held-out foliosA+B 7.1 sigmaPASS
Test 2part of speech from contextthe instrument fails on K&T's own words (68.3% against a needed 70%; 4.7 sigma against a needed 5); the readings were never scoredNO VERDICT
Test 3the blindfold, run clean4 of 24 strict, 16.7%; the declared band for that result was 15-30%, and its declared consequence, tier C passage readings become tier D, was appliedIN BAND
Test 7K&T's words removedreads 11.6% from ours alone, by design: the readings extend the dictionary. Recovery of a hidden K&T word: 10.3 sigma against a declared 5, on a hundred-shuffle control; the ten-shuffle runs gave 4.4 before the variant fix and 8.4 after, and all three are keptPASS
Test 8bootstrap from a random 30%passage 37.6 sigma (6.1% absolute); recovery 15.3 sigma (1.0% absolute)PASS
Test 9hidden 70% validated by the 30%A+B 6.9 sigma; 115% of K&T's own ratePASS
After the outside review by Gemini 2.5 Pro and Grok 4.7, same day:
Test 11passage map from K&T's words alonemedian rank 25 of 1,334; top-1 12.4%; p 2e-28PASS
Test 10the search replayed on null books7 / 215 / 2 signs kept at three scales; the count rule has no power at any, so it cannot tell a search from a deciphermentNO VERDICT
Test 10, held outthe same runs, scored where the readings were not derivedreal book 8 sigma over shuffle; null books nonereported, not a verdict
Test 12Test 5 rescored under the strict rule73.7%, 8.4 sigma; but +22.9 points over K&T's own headwords trips the declared leakage clauseFAIL on that clause
Test 12btwo independent re-glossers, Gemini 2.5 Pro and Grok 4.7given K&T's words only and the sign blanked, they recover our gloss 68.4% and 71.1% of the timePASS
Test 13the underdetermination census61 of 94 A/B readings have a common verb present in every chapter their folios cite, against a bar of 40%FAIL
Test 13, the same ruleapplied to the chosen glosses0 of the 94 pass it. The rivals exist; the method did not choose them. The FAIL measures the rival space, not what the method didreported, not a verdict
Test 14the blind rotated run, outside readerrotated book 0 fills, 0 matches; real book 3.8 fills a page, 13 of 18PASS
Test 15passage identification, outside readerchapter level 9 of 20 against 0 of 20 shuffled; p 0.002PASS
PASS and FAIL are verdicts on the readings against a bar declared before the run. NO VERDICT means the instrument failed its own check on Kiraly and Tokai's words, or has no power to tell a search from a decipherment; it says nothing about the readings either way, and it is kept on the page because it was specified. IN BAND is a test with declared bands rather than a bar: the result fell in the band it was predicted to, and the consequence declared for that band was applied.

The passes and failures add up to a checked attempt, not a settled result: checked by the project's own tests and by two models from other labs, not yet by anyone who reads the manuscript independently.

The passage map can be recovered from Király and Tokai's words alone. It does not depend on anything this project read. The two instruments Grok 4.7 specified for the search itself gave no verdict either way. They do not model the search as it was run.

The saved run behind every line is in the data, one file per test, each stating its bar at the top. The reviewers' own specifications are there too, quoted whole: Gemini 2.5 Pro and Grok 4.7. The programs are at The programs.

This rests on the Király and Tokai dictionary; their translation is unpublished; the readings here are this project's own.