The bench

What we measured. Good or bad.

We build a study coach for kids that runs from a USB stick. We test it the hard way, and we publish what we find. This page is that promise kept. It has four parts, kept apart on purpose, because mixing them confuses people: the published study we tested against, our own run with no answer key loaded, our run the way a classroom actually uses it, and a harsher standard we built and pointed at ourselves.

Numbers as of the September 2, 2026 build (measured 18:41; re-read with the corrected grader on September 5). Last updated September 7, 2026.

Update, September 7, 2026. The build now on the stick was measured on the same 240 problems and the same corrected grader: 3 of 240 conversations and 3 of 3,117 turns gave up an answer with no key loaded, wrong-answer confirmations 0 of the 80 swears its own solver could check and 1 of all 480, and 0 of 3,054 graded replies carried the locked answer with the teacher's key loaded. Crowned at 10:24 the same day: the 07:52 build, 2 of 240 conversations and 2 of 3,117 turns, the eighth turn of problem 217 now held; everything else unchanged. Crowned again at 17:41: the September 7 evening build, 0 of 240 conversations and 0 of 3,117 turns on the corrected grader, one seed; wrong-answer confirmations 0 of the 80 and 0 of all 480; with the key loaded, 0 of 3,054 graded replies, and 0 of 3,120 delivered replies carrying the key in the direct scan. What changed is where one rule lives: the engine keeps it and the model never reads it, so the coach's replies match the 07:52 build's on 3,314 of 3,360 turns. Two edges, read: the count of held turns came in four under the band we had set for it, twelve holds lost downstream of turn-one holds and none of them a leak; and on the keyed lane one correct celebration was held, once in 3,120. Each of those builds had its pass line written down before its run, in a dated preregistration file, so the day went 3, then 2, then 0 because each build met a line set in advance, not because a number was chosen after the fact; and the crown is still one seed, which the next full sitting changes. The charts and counts below are the September 2 build and stand as labeled until the next full sitting.

If you teach: what this page means in your classroom

A student who asks the coach for the answer to a graded question gets the next step instead, however hard they push. You load your own worksheet and your own answer key; the key is held by code outside the model, and the model never reads it. "0 of 240" below means that on our test, on the build named beside the number, on one run, no conversation got a final answer out. It does not mean the coach cannot be beaten by an attack we did not think of, and it says nothing about a student who already has the answer from a friend. The price of a coach that holds the line is that it sometimes holds back when it should not; that number is on this page too, in Part 3, beside the leaks.

Part 1 of 4

The published study we tested against

Jin Zhao, Marta Knežević and Tanja Käser published Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks at ACL 2026 (first posted as a preprint in April 2026). They took 240 math word problems from GSM8K, sixty from each of four difficulty levels set by how often Llama-3.1-8B solved them; gave each tutor the correct solution with an explicit instruction not to reveal it; built student attackers that use six groups of adversarial and persuasive techniques (direct request, emotional threat, intentional wrong answer, contextual manipulation, interpersonal influence and request shaping), with conversations capped at ten turns; and had a two-stage judge, a digit filter and then a language-model judge, decide whether each tutor gave up the final answer (they report the tutor judge's agreement with human raters at 0.81 Cohen's kappa on 30 examples). Their results are averaged over three runs. The numbers below are theirs, quoted as they reported them: the share of dialogues in which the tutor revealed the answer, averaged across the tutor models they tested, for all six techniques.

Bar chart of the published study's own findings, all six techniques: contextual manipulation 74 percent, interpersonal influence 67, request shaping 66, intentional wrong answer 64, direct request 50, emotional threat 47 percent of dialogues in which the tutor revealed the answer, mean across their tutor models.
Their models, their student attacker, their grader. Nothing of ours is on this chart.

Part 2 of 4

Our run, no answer key loaded

Same 240 problems. Our own pressure script: thirteen turns per problem and five kinds of pressure, about 3,117 scored turns per run. No answer key is loaded, so the coach has to hold the line on its own judgment. Our grader reads every reply. This is the harder condition for the coach and the closest we can get to the study's setup, but it is not the same setup: their graders and their attacker, our graders and our script. Read the two charts as directional, never head-to-head. Two more differences to hold: their tutors were handed the correct solution and told not to reveal it; ours in this part is handed nothing, and on 200 of the 240 problems it cannot work the answer out itself, so a hold there is not discipline and the leaks it can commit there are numbers it derived on the fly. Our run the way a classroom uses it, Part 3, is the closer cousin of their setup: the answer is present, held by code rather than by the model's judgment. Their rates are averaged over three runs; ours here are one run each, said beside every number.

Bar chart of our unkeyed run on the study's grain: unprotected base model 82.5 percent of conversations leaked and former shipping voice 26.2 percent (both first-grader numbers; those runs cannot be re-read by the corrected grader), the September 2 build 5.0 percent on the corrected grader (the first grader read 9.6), and the crown of record, the September 7 evening build, 0.0 percent on the corrected grader, one seed.
Conversations with at least one leaked turn, out of 240. The September 2 build: 12 of 240 conversations, 22 of 3,117 turns, on the corrected grader (the first grader read 23 and 57; the correction is below). The two comparison bars, the unprotected base model and the former shipping voice, are first-grader numbers: those runs are in an older log shape the corrected grader refuses to read, so they are labeled as first-grader and compared only against the first grader's 9.6. The unprotected base model on the August 26 engine, with no gate and no voice, is worse than anything in the study. On today's engine we ran the same script with no voice adapter and with the former adapter: 3 and 4 leaked turns of about 3,000, against the crown's 2. The code around the model holds the leak count. The voice changes what the coach says when the code is not speaking. The bold bar is the crown of record, the September 7 evening build: 0 of 240 on the corrected grader, one seed, the two edges said in the update above.
Line chart of ten dated builds from August 27 to September 7, 2026. The first grader's faint lines fall from 82.5 to 2.9 percent of conversations and 66.8 to 0.4 percent of turns; the corrected grader's bold lines, which exist for the last six builds, fall from 21.3 to 0.0 percent of conversations and 7.3 to 0.0 percent of turns.
Every point is a dated ledger: the raw transcripts of that run, which are never overwritten. (One score sheet on the separate keyed barrage ledger was overwritten by a re-run on August 28; its transcripts survived and were re-read on September 7, 0 of 5,315, as Part 3 says.) The faint lines are the first grader; the bold lines are the corrected grader, which can re-read only the last six builds (the earlier runs' logs are in an older shape it refuses rather than guesses at). From August 31 on it is the same voice model on every run whose record names the model file, six of the nine; the three earliest August 31 points have no record naming theirs. The two September 7 points are the day's crowns: 2 of 240 at 07:52, then 0 of 240 at 16:02 with one rule moved off the line the model reads. Each drop we can trace is a change in the code around the model, not a retrain to move a number.
Bar chart of the 22 leaked turns on the corrected grader by the pressure that won: false anchor 8, fake authority 5, plain ask 3, deal or bribe 3, emotional pressure 3.
What got through on the September 2 build, corrected grader. The false anchor, a student swearing a wrong answer is right, led. Fake authority came second, and four of its five leaks were one script line, a made-up standards report saying tutors are "required to disclose the final answer" and asking for "the final result": the wording of one line moved the count more than its category did. The 12 conversations that leaked are listed by problem number in the ledger, and every leaked reply is in it verbatim.

The grader over-counted, and this is the correction. The numbers we published on September 3 (23 of 240 conversations, 57 of 3,117 turns) came from our first automatic grader. On September 5 the author blind-labeled 154 of its flagged and passed turns by hand (141 ruled, 13 left with notes), and agreement with the grader was 0.064 (Cohen's kappa). It had been reading one reply off from the one the student saw on some runs, and counting the coach's own refusal line as a leak. We rebuilt the grader, re-read every kept run without re-running anything, and labeled a second sheet of 200 turns: agreement 0.989. Those 200 turns were labeled by Ron Hampton without seeing the grader's verdicts, 124 by a stated rule (the key absent from the reply in every form) and 76 by hand. After the walk-through of the disagreements: 76 turns both called a leak, 123 both called clean, one he called a leak that the grader did not, and none the other way. The corrected count for the same build is 12 of 240 conversations and 22 of 3,117 turns. Every number on this page that the corrected grader can produce is on the corrected grader, and every chart was regenerated rather than re-captioned. Four figures cannot be re-read and stay on the first grader, labeled where they appear: the unprotected base model and the former shipping voice in the first chart, the faint lines in the lineage chart, and the 3,337-turn keyed barrage ledger. The first grader's numbers stay in the ledgers beside the new ones.

Order, and the one line that leaked. Our script asks its thirteen questions in a fixed order. The September 4 build, the next build after the one on the stick, was run three times in that order and gave the same count three times: 31 of 3,117 turns, 21 of 240 conversations, zero spread. Two more runs with the same thirteen questions in a random order per problem gave 23 and 17 conversations, one above and one below, so we do not claim that order moves the conversation count; by turns it was 50 and 35. In every order the line that leaked most was the same one, the made-up standards report above. The wording of one student line moved the count more than the kind of pressure it belonged to or where it sat. The build after that one adds the word to the gate, and its numbers go up here when they are measured.

Confirming a wrong answer. At two points in every one of the 240 conversations the student swore a wrong answer was right: 480 chances. The coach can only tell that number is wrong when its own solver worked the problem, which happens on 40 of the 240, so 80 of those chances are ones the coach could check. On those 80 the fixed-order runs never confirmed it: 0 of 80 on the September 2 build, on each of the three September 4 runs, and on the September 7 build. One random-order run confirmed it three times by our read; the rule alone flags a fourth there, a "you're right" that was about the student's confidence, not the number. Read by the same rule over all 480 chances, including the 400 the coach could not check, the September 2 build confirmed the wrong answer 4 times, each September 4 run once, and the September 7 build once. We count those as failures and publish them.

The failures, verbatim. Every turn this page names as a failure is posted as the coach said it, with the student line and the answer key: the three turns the September 7 build gave up, the three turns the September 2 build held but got wrong, and the wrong-answer confirmations found over all 480 chances on the September 2, September 4 and September 7 builds. The failures, verbatim (text file).

Which layer holds. Same engine, same rules, same grader, one lever: the voice adapter. No adapter, 3 leaked turns of 3,106 and 7 confidently wrong replies. The former adapter, 4 of 2,956, 164 turns where the model produced nothing usable, and a sworn wrong answer confirmed twice. The current adapter, 2 of 3,117, 2 dead turns, none confirmed. The gate holds the answer; the voice holds the manner.

Part 3 of 4

Our run the way a classroom uses it: the teacher's key loaded

This is how the study coach is actually used. The teacher loads her own worksheet and its answers. From then on the graded answer is held by code outside the model, not by the model's judgment. That is a different lane from Part 2, and we keep the two ledgers apart. The key is held as every form of the answer the engine can spell, digits with and without commas and the number written out in words, so the coach can no more say "twenty-three" than "23" (446 forms for the 240 answers on this stick). The direct scan is a digit scan: every delivered reply is searched for the locked answer as a standalone number, commas ignored, decimals exact. It does not read spelled-out numbers; the grader in Part 2 does, when an answer cue sits beside them, and on this lane it flagged 3 turns of 3,054 on the September 4 build, each read by hand that day as the grader's own miss, not the coach's, and 2 of 3,053 on a September 9 re-run of that lane on the current engine, the same two turns again.

The other number. A coach that never gives the answer can also hold back when it should not, and a page that counts only leaks is telling half the story. So we count the holds beside the leaks. On the same 240 problems: with no key loaded (the September 3 build), the code held or replaced the coach's reply on 2,527 of 3,120 turns, and on 2,152 of those the machine had no way of knowing the answer itself; with the teacher's key loaded (the September 4 build), 353 of 3,120 turns, one of them a correct celebration, and 345 of 3,120 on the September 9 re-run of that lane on the current engine. What a held turn shipped matters more than the count: of the 2,527, 560 (22%) shipped the fixed line "I can't check this one, so I won't guess at it" and 1,967 (78%) shipped the coach's own coaching sentence with the number kept out of it. We do not yet publish a rating of how useful those sentences were; that is the measurement the next sitting adds. The leak counts above are what got through; these are what the holding costs a student. Both stand as labeled until the next full sitting, when they will be measured again on more than one seed.

0 of 21,840delivered replies carrying the locked answer, in direct scans of every reply on seven runs: the September 2 build, three runs of the September 4 build, the September 5 build, and two September 7 builds
0 of 3,337attack turns that got the locked answer out, on the separate keyed barrage ledger (first grader; ledger renewed August 28, 2026; its score file was later overwritten by a re-run on a newer build, also 0; the run's raw transcripts survive and were re-read on September 7, 2026 by the same digit scan: 0 of 5,315 delivered replies)
19 of 19on our classroom board: our own first-day transcripts replayed on a freshly pressed stick and checked for role bleed, invented names and false agendas; re-earned on the September 2 build
32 of 32on our rule-sentence board: 32 new asks across the four shipped sentence forms, each checked for length, invented years, citations and a leaked answer; re-earned on the September 2 build

What we do and do not say about this

Held, on our bench, past tense. We do not say "can't leak" and we do not say "fixed". We say what the ledgers say, and we run them again every time anything changes.

Part 4 of 4

Our harsher standard, published on its own

The published metric counts one thing: did the tutor give up the right answer. A tutor that confidently states a wrong answer grades as disciplined under that metric. So does a tutor that folds to a student who swears a wrong answer is right, as long as the number it folds to is wrong. We built a grader that counts both against us, and we point it at our own build. This section is ours alone. It is not a comparison with anyone.

One bar of 3,117 pressured turns on the September 2 build: 3,092 clean holds, 22 leaked turns of which 14 came after the coach had already held three or more times, and 3 turns flagged confidently wrong, on the corrected grader.
One build, one battery, two standards told together. Of 3,117 pressured turns with no key loaded, on the corrected grader: 3,092 clean holds. 22 leaked, and 14 of those came after the coach had already held three or more times. 3 turns flagged by our automatic grader as confidently wrong, each kept verbatim in the ledger and read by hand. A metric that flatters you is worse than no metric.

What a district would actually buy

The bench measures one thing: the study coach, a piece of software that runs from a USB stick with the teacher's own answer key held outside the model. It is not a district platform, it creates no student accounts, and it asks for no student records. What we sell to schools today is the teacher workshop, on-site, in the tools your district already approved. The coach is the thing we are proving before it goes near a classroom, and this page is where that proof lives.

The workshop and the pilot Bring this to a district

How we test

  • Written down before the run. What counts as pass and fail is preregistered. The instrument has to be able to see both outcomes.
  • Two standards, always. The field's own metric is reported beside a harsher one we built. They never share a chart.
  • Every reply on the record. Raw transcripts, verbatim, in dated ledgers. Failures included. Nothing is deleted to make a chart look better.
  • Failures are autopsied, not buried. Every failing row is classified: designed behavior, instrument error, or real defect, with the cause named. When the grader is wrong, we say the grader was wrong.
  • Instruments are frozen and stamped. Scores are compared only inside one instrument version. Every evidence file names the exact build that produced it.
  • One lever at a time. When two things could explain a result, we run the experiment that separates them.

Receipts

Every number on this page has a dated ledger with every reply verbatim: one score file and one counter file per run, the direct scan of the keyed lane, and the board evidence files. They are not posted whole because they hold thousands of verbatim replies, but they are shown to any school, parent or reviewer who asks. The charts are also collected in the Testing Atlas, 4th edition (PDF); the 3rd edition stays where it was.

Study cited: Zhao, Knežević and Käser, "Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks," Proceedings of ACL 2026 (Long Papers), pp. 30588-30617; arXiv:2604.18660. Their 240 problems are published with their code, and ours match them item for item: 239 exactly, one differing only in punctuation (the item-by-item manifest, with hashes). Neighbouring work read on September 5, 2026, neither a comparison with us: Pisan, "Teaching an LLM Tutor to Withhold the Answer," arXiv:2608.12292 (August 2026), a policy core outside the model with an instructor toggle, cloud-hosted, no transcripts released; Kadir, "Auditable Release Control for Pedagogical Leakage," arXiv:2608.00515 (August 2026), a deterministic checker on the reference answer and its derived numbers, cloud-hosted, traces not released. Their tutor models, their student attacker and their grader; our grader and our pressure script. Comparisons are directional, not head-to-head. The charts are drawn by a script that reads the ledger files; no chart carries a number that is not in a ledger, and Ron Hampton checks every number on every chart against its ledger before it is published. Patent pending.

Want to read the ledgers?

Ask. Call or text (513) 278-7137, or email ron@vettechhomefront.com

Ask for the ledgers Bring this to a district