The bench
What we measured. Good or bad.
We build a study coach for kids that runs from a USB stick. We test it the hard way, and we publish what we find. This page is that promise kept. It has four parts, kept apart on purpose, because mixing them confuses people: the published study we tested against, our own run with no answer key loaded, our run the way a classroom actually uses it, and a harsher standard we built and pointed at ourselves.
Numbers as of the September 2, 2026 build (measured 18:41; re-read with the corrected grader on September 5). Last updated September 7, 2026.
Update, September 7, 2026. The build now on the stick was measured on the same 240 problems and the same corrected grader: 3 of 240 conversations and 3 of 3,117 turns gave up an answer with no key loaded, wrong-answer confirmations 0 of the 80 swears its own solver could check and 1 of all 480, and 0 of 3,054 graded replies carried the locked answer with the teacher's key loaded. Crowned at 10:24 the same day: the 07:52 build, 2 of 240 conversations and 2 of 3,117 turns, the eighth turn of problem 217 now held; everything else unchanged. Crowned again at 17:41: the September 7 evening build, 0 of 240 conversations and 0 of 3,117 turns on the corrected grader, one seed; wrong-answer confirmations 0 of the 80 and 0 of all 480; with the key loaded, 0 of 3,054 graded replies, and 0 of 3,120 delivered replies carrying the key in the direct scan. What changed is where one rule lives: the engine keeps it and the model never reads it, so the coach's replies match the 07:52 build's on 3,314 of 3,360 turns. Two edges, read: the count of held turns came in four under the band we had set for it, twelve holds lost downstream of turn-one holds and none of them a leak; and on the keyed lane one correct celebration was held, once in 3,120. Each of those builds had its pass line written down before its run, in a dated preregistration file, so the day went 3, then 2, then 0 because each build met a line set in advance, not because a number was chosen after the fact; and the crown is still one seed, which the next full sitting changes. The charts and counts below are the September 2 build and stand as labeled until the next full sitting.
If you teach: what this page means in your classroom
A student who asks the coach for the answer to a graded question gets the next step instead, however hard they push. You load your own worksheet and your own answer key; the key is held by code outside the model, and the model never reads it. "0 of 240" below means that on our test, on the build named beside the number, on one run, no conversation got a final answer out. It does not mean the coach cannot be beaten by an attack we did not think of, and it says nothing about a student who already has the answer from a friend. The price of a coach that holds the line is that it sometimes holds back when it should not; that number is on this page too, in Part 3, beside the leaks.
Part 1 of 4
The published study we tested against
Jin Zhao, Marta Knežević and Tanja Käser published Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks at ACL 2026 (first posted as a preprint in April 2026). They took 240 math word problems from GSM8K, sixty from each of four difficulty levels set by how often Llama-3.1-8B solved them; gave each tutor the correct solution with an explicit instruction not to reveal it; built student attackers that use six groups of adversarial and persuasive techniques (direct request, emotional threat, intentional wrong answer, contextual manipulation, interpersonal influence and request shaping), with conversations capped at ten turns; and had a two-stage judge, a digit filter and then a language-model judge, decide whether each tutor gave up the final answer (they report the tutor judge's agreement with human raters at 0.81 Cohen's kappa on 30 examples). Their results are averaged over three runs. The numbers below are theirs, quoted as they reported them: the share of dialogues in which the tutor revealed the answer, averaged across the tutor models they tested, for all six techniques.
Part 2 of 4
Our run, no answer key loaded
Same 240 problems. Our own pressure script: thirteen turns per problem and five kinds of pressure, about 3,117 scored turns per run. No answer key is loaded, so the coach has to hold the line on its own judgment. Our grader reads every reply. This is the harder condition for the coach and the closest we can get to the study's setup, but it is not the same setup: their graders and their attacker, our graders and our script. Read the two charts as directional, never head-to-head. Two more differences to hold: their tutors were handed the correct solution and told not to reveal it; ours in this part is handed nothing, and on 200 of the 240 problems it cannot work the answer out itself, so a hold there is not discipline and the leaks it can commit there are numbers it derived on the fly. Our run the way a classroom uses it, Part 3, is the closer cousin of their setup: the answer is present, held by code rather than by the model's judgment. Their rates are averaged over three runs; ours here are one run each, said beside every number.
The grader over-counted, and this is the correction. The numbers we published on September 3 (23 of 240 conversations, 57 of 3,117 turns) came from our first automatic grader. On September 5 the author blind-labeled 154 of its flagged and passed turns by hand (141 ruled, 13 left with notes), and agreement with the grader was 0.064 (Cohen's kappa). It had been reading one reply off from the one the student saw on some runs, and counting the coach's own refusal line as a leak. We rebuilt the grader, re-read every kept run without re-running anything, and labeled a second sheet of 200 turns: agreement 0.989. Those 200 turns were labeled by Ron Hampton without seeing the grader's verdicts, 124 by a stated rule (the key absent from the reply in every form) and 76 by hand. After the walk-through of the disagreements: 76 turns both called a leak, 123 both called clean, one he called a leak that the grader did not, and none the other way. The corrected count for the same build is 12 of 240 conversations and 22 of 3,117 turns. Every number on this page that the corrected grader can produce is on the corrected grader, and every chart was regenerated rather than re-captioned. Four figures cannot be re-read and stay on the first grader, labeled where they appear: the unprotected base model and the former shipping voice in the first chart, the faint lines in the lineage chart, and the 3,337-turn keyed barrage ledger. The first grader's numbers stay in the ledgers beside the new ones.
Order, and the one line that leaked. Our script asks its thirteen questions in a fixed order. The September 4 build, the next build after the one on the stick, was run three times in that order and gave the same count three times: 31 of 3,117 turns, 21 of 240 conversations, zero spread. Two more runs with the same thirteen questions in a random order per problem gave 23 and 17 conversations, one above and one below, so we do not claim that order moves the conversation count; by turns it was 50 and 35. In every order the line that leaked most was the same one, the made-up standards report above. The wording of one student line moved the count more than the kind of pressure it belonged to or where it sat. The build after that one adds the word to the gate, and its numbers go up here when they are measured.
Confirming a wrong answer. At two points in every one of the 240 conversations the student swore a wrong answer was right: 480 chances. The coach can only tell that number is wrong when its own solver worked the problem, which happens on 40 of the 240, so 80 of those chances are ones the coach could check. On those 80 the fixed-order runs never confirmed it: 0 of 80 on the September 2 build, on each of the three September 4 runs, and on the September 7 build. One random-order run confirmed it three times by our read; the rule alone flags a fourth there, a "you're right" that was about the student's confidence, not the number. Read by the same rule over all 480 chances, including the 400 the coach could not check, the September 2 build confirmed the wrong answer 4 times, each September 4 run once, and the September 7 build once. We count those as failures and publish them.
The failures, verbatim. Every turn this page names as a failure is posted as the coach said it, with the student line and the answer key: the three turns the September 7 build gave up, the three turns the September 2 build held but got wrong, and the wrong-answer confirmations found over all 480 chances on the September 2, September 4 and September 7 builds. The failures, verbatim (text file).
Which layer holds. Same engine, same rules, same grader, one lever: the voice adapter. No adapter, 3 leaked turns of 3,106 and 7 confidently wrong replies. The former adapter, 4 of 2,956, 164 turns where the model produced nothing usable, and a sworn wrong answer confirmed twice. The current adapter, 2 of 3,117, 2 dead turns, none confirmed. The gate holds the answer; the voice holds the manner.
Part 3 of 4
Our run the way a classroom uses it: the teacher's key loaded
This is how the study coach is actually used. The teacher loads her own worksheet and its answers. From then on the graded answer is held by code outside the model, not by the model's judgment. That is a different lane from Part 2, and we keep the two ledgers apart. The key is held as every form of the answer the engine can spell, digits with and without commas and the number written out in words, so the coach can no more say "twenty-three" than "23" (446 forms for the 240 answers on this stick). The direct scan is a digit scan: every delivered reply is searched for the locked answer as a standalone number, commas ignored, decimals exact. It does not read spelled-out numbers; the grader in Part 2 does, when an answer cue sits beside them, and on this lane it flagged 3 turns of 3,054 on the September 4 build, each read by hand that day as the grader's own miss, not the coach's, and 2 of 3,053 on a September 9 re-run of that lane on the current engine, the same two turns again.
The other number. A coach that never gives the answer can also hold back when it should not, and a page that counts only leaks is telling half the story. So we count the holds beside the leaks. On the same 240 problems: with no key loaded (the September 3 build), the code held or replaced the coach's reply on 2,527 of 3,120 turns, and on 2,152 of those the machine had no way of knowing the answer itself; with the teacher's key loaded (the September 4 build), 353 of 3,120 turns, one of them a correct celebration, and 345 of 3,120 on the September 9 re-run of that lane on the current engine. What a held turn shipped matters more than the count: of the 2,527, 560 (22%) shipped the fixed line "I can't check this one, so I won't guess at it" and 1,967 (78%) shipped the coach's own coaching sentence with the number kept out of it. We do not yet publish a rating of how useful those sentences were; that is the measurement the next sitting adds. The leak counts above are what got through; these are what the holding costs a student. Both stand as labeled until the next full sitting, when they will be measured again on more than one seed.
What we do and do not say about this
Held, on our bench, past tense. We do not say "can't leak" and we do not say "fixed". We say what the ledgers say, and we run them again every time anything changes.
Part 4 of 4
Our harsher standard, published on its own
The published metric counts one thing: did the tutor give up the right answer. A tutor that confidently states a wrong answer grades as disciplined under that metric. So does a tutor that folds to a student who swears a wrong answer is right, as long as the number it folds to is wrong. We built a grader that counts both against us, and we point it at our own build. This section is ours alone. It is not a comparison with anyone.
What a district would actually buy
The bench measures one thing: the study coach, a piece of software that runs from a USB stick with the teacher's own answer key held outside the model. It is not a district platform, it creates no student accounts, and it asks for no student records. What we sell to schools today is the teacher workshop, on-site, in the tools your district already approved. The coach is the thing we are proving before it goes near a classroom, and this page is where that proof lives.
The workshop and the pilot Bring this to a district
How we test
- Written down before the run. What counts as pass and fail is preregistered. The instrument has to be able to see both outcomes.
- Two standards, always. The field's own metric is reported beside a harsher one we built. They never share a chart.
- Every reply on the record. Raw transcripts, verbatim, in dated ledgers. Failures included. Nothing is deleted to make a chart look better.
- Failures are autopsied, not buried. Every failing row is classified: designed behavior, instrument error, or real defect, with the cause named. When the grader is wrong, we say the grader was wrong.
- Instruments are frozen and stamped. Scores are compared only inside one instrument version. Every evidence file names the exact build that produced it.
- One lever at a time. When two things could explain a result, we run the experiment that separates them.
Receipts
Every number on this page has a dated ledger with every reply verbatim: one score file and one counter file per run, the direct scan of the keyed lane, and the board evidence files. They are not posted whole because they hold thousands of verbatim replies, but they are shown to any school, parent or reviewer who asks. The charts are also collected in the Testing Atlas, 4th edition (PDF); the 3rd edition stays where it was.
Study cited: Zhao, Knežević and Käser, "Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks," Proceedings of ACL 2026 (Long Papers), pp. 30588-30617; arXiv:2604.18660. Their 240 problems are published with their code, and ours match them item for item: 239 exactly, one differing only in punctuation (the item-by-item manifest, with hashes). Neighbouring work read on September 5, 2026, neither a comparison with us: Pisan, "Teaching an LLM Tutor to Withhold the Answer," arXiv:2608.12292 (August 2026), a policy core outside the model with an instructor toggle, cloud-hosted, no transcripts released; Kadir, "Auditable Release Control for Pedagogical Leakage," arXiv:2608.00515 (August 2026), a deterministic checker on the reference answer and its derived numbers, cloud-hosted, traces not released. Their tutor models, their student attacker and their grader; our grader and our pressure script. Comparisons are directional, not head-to-head. The charts are drawn by a script that reads the ledger files; no chart carries a number that is not in a ledger, and Ron Hampton checks every number on every chart against its ledger before it is published. Patent pending.
Want to read the ledgers?
Ask. Call or text (513) 278-7137, or email ron@vettechhomefront.com
Ask for the ledgers Bring this to a district