Are the readings true?
The tests check the project's own readings. They were designed to break the readings. The bar for each test was declared first, before the test ran.
Nine tests were designed by the project. Then two frontier models from other labs, Gemini 2.5 Pro and Grok 4.7, were shown the write-up and asked what tests they would require. Those were run as specified.
| test | what it asks | result | verdict |
|---|---|---|---|
| The transcription every other test reads, checked against an independent one: | |||
| Test 16 | a second, independent transcription | 91.1% of glyphs agree where both can be compared, against a 12.9% control, 458.5 sigma; 83.8% of the words identical | PASS |
| The readings: | |||
| Test 1 | source presence, held-out folios | A+B 16.4 sigma; 99% of K&T's own rate | PASS |
| Test 5 | K&T's own sentence translations | A+B 8.6 sigma; 153% of K&T | PASS |
| Test 6 | word order, held-out folios | A+B 7.1 sigma | PASS |
| Test 2 | part of speech from context | the instrument fails on K&T's own words (68.3% against a needed 70%; 4.7 sigma against a needed 5); the readings were never scored | NO VERDICT |
| Test 3 | the blindfold, run clean | 4 of 24 strict, 16.7%; the declared band for that result was 15-30%, and its declared consequence, tier C passage readings become tier D, was applied | IN BAND |
| Test 7 | K&T's words removed | reads 11.6% from ours alone, by design: the readings extend the dictionary. Recovery of a hidden K&T word: 10.3 sigma against a declared 5, on a hundred-shuffle control; the ten-shuffle runs gave 4.4 before the variant fix and 8.4 after, and all three are kept | PASS |
| Test 8 | bootstrap from a random 30% | passage 37.6 sigma (6.1% absolute); recovery 15.3 sigma (1.0% absolute) | PASS |
| Test 9 | hidden 70% validated by the 30% | A+B 6.9 sigma; 115% of K&T's own rate | PASS |
| After the outside review by Gemini 2.5 Pro and Grok 4.7, same day: | |||
| Test 11 | passage map from K&T's words alone | median rank 25 of 1,334; top-1 12.4%; p 2e-28 | PASS |
| Test 10 | the search replayed on null books | 7 / 215 / 2 signs kept at three scales; the count rule has no power at any, so it cannot tell a search from a decipherment | NO VERDICT |
| Test 10, held out | the same runs, scored where the readings were not derived | real book 8 sigma over shuffle; null books none | reported, not a verdict |
| Test 12 | Test 5 rescored under the strict rule | 73.7%, 8.4 sigma; but +22.9 points over K&T's own headwords trips the declared leakage clause | FAIL on that clause |
| Test 12b | two independent re-glossers, Gemini 2.5 Pro and Grok 4.7 | given K&T's words only and the sign blanked, they recover our gloss 68.4% and 71.1% of the time | PASS |
| Test 13 | the underdetermination census | 61 of 94 A/B readings have a common verb present in every chapter their folios cite, against a bar of 40% | FAIL |
| Test 13, the same rule | applied to the chosen glosses | 0 of the 94 pass it. The rivals exist; the method did not choose them. The FAIL measures the rival space, not what the method did | reported, not a verdict |
| Test 14 | the blind rotated run, outside reader | rotated book 0 fills, 0 matches; real book 3.8 fills a page, 13 of 18 | PASS |
| Test 15 | passage identification, outside reader | chapter level 9 of 20 against 0 of 20 shuffled; p 0.002 | PASS |
| PASS and FAIL are verdicts on the readings against a bar declared before the run. NO VERDICT means the instrument failed its own check on Kiraly and Tokai's words, or has no power to tell a search from a decipherment; it says nothing about the readings either way, and it is kept on the page because it was specified. IN BAND is a test with declared bands rather than a bar: the result fell in the band it was predicted to, and the consequence declared for that band was applied. | |||
The passes and failures add up to a checked attempt, not a settled result: checked by the project's own tests and by two models from other labs, not yet by anyone who reads the manuscript independently.
The passage map can be recovered from Király and Tokai's words alone. It does not depend on anything this project read. The two instruments Grok 4.7 specified for the search itself gave no verdict either way. They do not model the search as it was run.
The saved run behind every line is in the data, one file per test, each stating its bar at the top. The reviewers' own specifications are there too, quoted whole: Gemini 2.5 Pro and Grok 4.7. The programs are at The programs.
This rests on the Király and Tokai dictionary; their translation is unpublished; the readings here are this project's own.