Case Study: Auditing the Grader — When ‘Prompt Sensitivity’ Turned Out to Be a Regex Bug
🌟 Overview: The Finding That Wasn’t
“My benchmark found prompt sensitivity. Then I audited the grader — and the story fell apart.”
I ran a small model-evaluation experiment: 480 generations across two small open-weight models (Qwen 2.5 1.5B Instruct vs. Llama 3.2 1B Instruct), 40 grade-school math problems I wrote myself, three instruction wordings, two stochastic runs each. Machine-graded. The first results were exciting — Qwen beat Llama by 11.3 points, and the gap moved with the prompt wording (15.0 / 8.8 / 10.0 points), with statistical significance flipping between wordings. A clean little demonstration of prompt sensitivity.
Then I checked the grader. The “sensitivity” lived in my regex, not in the models’ arithmetic. This case study is the full post-mortem: the audit procedure, the failure mechanism, the corrected numbers, and the measurement lessons I took from it.
❓ The Question
When two models are compared under different prompt wordings and the gap between them moves, is that prompt sensitivity in the models — or grading sensitivity in the pipeline? My grading pipeline depended on output format, my prompts changed output format, and I was attributing the resulting score movement to reasoning differences. That is a confound, not a finding — until proven otherwise.
🔬 Method
Design. Fully crossed: 2 models × 3 instruction wordings × 40 hand-written arithmetic tasks × 2 stochastic trials = 480 generations. V1 and V2 are instruction paraphrases; V3 adds a “careful math tutor” persona (a different treatment, disclosed as such). Temperature 0.7, max 256 tokens, local quantized GGUF checkpoints on CPU. (Up front: 1.5B vs 1B is a 50% size advantage for Qwen — nothing here is a controlled architecture comparison.)
The grader under audit. A two-stage parser: first try to match the requested “Answer:” format; on failure, fall back to extracting the last integer in the output. Correctness = extracted integer equals the gold answer.
Statistics. 10,000-replicate task-clustered bootstrap (task IDs resampled with replacement, multiplicities preserved; one shared resample per replicate for per-variant gaps and difference-of-differences). Random 20-task subset diagnostic. Trial-to-trial correctness disagreement.
🔍 The Audit
The audit ran in two stages. First, an AI assistant performed a preliminary triage: it screened all 36 parser-failures where the gold number appeared in the output text, proposing 20 candidate flips (and classifying the other 16 as genuine model errors), while finding zero false positives across a random sample of 50 parser passes. I then personally adjudicated all 20 flip candidates in full — confirming 19, overturning 1.
Two things about the adjudication belong on the record. First, the criterion wasn’t uniform: on the 20 hand-adjudicated candidates I applied a stricter rule — the gold integer stated with no false arithmetic in support — while the 403 parser-passed records were scored on final-answer extraction alone. Only ~4% of correct labels were ever examined for reasoning quality. The labels mix two standards, and the analysis says which records got which. The single overturn: the model answered 6 full boxes correctly but computed the remainder as 150 − 24 = 126 (correct: 6 × 24 = 144, so 6 short — a quarter box). A conservative call, kept wrong.
Second, I got one call wrong and fixed it before publishing. I initially overturned a record where the model answered 154 but explained it as “adding the number to itself 10 times” — then reversed myself, because the parenthetical “(since 11 = 10 + 1)” makes it the distributive law: 14 + (14 × 10). Clumsily worded, mathematically valid. The transcripts are public; the call is checkable.
The mechanism was almost comically mundane. Eighteen of the nineteen flips are the same failure — a premise-echoing recap: the model derives the right answer, then closes with a sentence restating it alongside a number from the problem premise, and the last-integer fallback grabs the premise. “The factory makes 875 toys in 7 days” → scored as 7. “24 cups from 3 full jugs” → scored as 3. “560 pages in 2 weeks” → scored as 2. Llama narrates after answering more than Qwen does, so the parser punished Llama’s verbosity as if it were Llama’s innumeracy — hardest under the prompts that elicited the most narration. Fifteen of the 19 flips were Llama’s; 10 were under the persona variant.
📊 What Survived
| Parser-graded gap (Qwen − Llama) | Audited gap | |
|---|---|---|
| Wording 1 | 15.0 pp, 95% CI [3.8, 26.3] | 6.25 pp, 95% CI [−3.7, 16.2] |
| Wording 2 | 8.8 pp, 95% CI [0.0, 18.8] | 6.25 pp, 95% CI [−1.3, 16.2] |
| Persona (V3) | 10.0 pp, 95% CI [0.0, 20.0] | 7.5 pp, 95% CI [−1.3, 17.5] |
| Overall | 11.3 pp | 6.67 pp, 95% CI [0.4, 14.2] |
Under a pure final-answer criterion the single overturn reverts to correct: the V3 gap rejoins the others at 6.25pp and the overall gap is 15/240 = 6.25pp. The headline doesn’t move — which is why disclosing the criterion asymmetry matters: the substantive conclusion is robust to whether you enforce clean intermediate reasoning or score raw final answers.
The difference-of-differences — the direct test of whether the gap varies by prompt — is centered near zero (0.1, −1.0, −1.2), but the intervals span ±10 to ±15 points: far too small a study to rule out real prompt interactions. What remains is a 6.67-point Qwen lead whose interval barely excludes zero, 10.0% trial-to-trial disagreement, and an honest summary: maybe a modest Qwen lead, and we can’t say much more than that.
💡 Key Insights
💡 Key Insight #1: Audit the grader before you believe the comparison.
- What I Found: The parser agreed with audited labels on 96.0% of records — and still manufactured a headline effect, because every confirmed error was a false negative falling overwhelmingly on one model.
- The “So What?”: Agreement rates are the wrong metric — asymmetric error is. A grader that fails one side’s outputs more than the other’s adds bias, not noise, and no downstream bootstrapping removes it.
- Why It Matters: Any automated grader, from regex to LLM-judge, has an error profile. When that profile correlates with the thing being compared, it manufactures effects. Characterize the instrument before trusting its outputs.
💡 Key Insight #2: Apparent prompt sensitivity can be grading sensitivity.
- What I Found: The 6.2-point swing in the model gap across wordings — plus flipping significance labels — vanished once the grader was fixed. This mirrors Hua et al. (2025): rigid answer-matching mistakes formatting variation for capability variation.
- The “So What?”: Before theorizing about model behavior under prompt perturbations, check whether the measurement apparatus is what’s moving.
- Why It Matters: Standard harness implementations of grade-school-math benchmarks used essentially this kind of regex extraction for years. The failure mode isn’t confined to toy studies.
💡 Key Insight #3: The significant-vs-not-significant comparison is a trap.
- What I Found: My first framing treated wording one as “decisive” and the others as “can’t rule out zero” — as if crossing p < 0.05 between conditions meant the conditions differed. The interaction intervals all covered zero.
- The “So What?”: Gelman and Stern’s warning holds: the difference between “significant” and “not significant” is not itself statistically significant.
- Why It Matters: If an evaluation story depends on which side of an arbitrary threshold each condition lands on, there is no evaluation story.
💡 Key Insight #4: Say what the estimand is.
- What I Found: Writing the estimand down — the difference in expected final-answer accuracy between these two specific quantized checkpoints on these 40 tasks, averaged over three wordings and decoding noise — exposed every overgeneralization the first draft wanted to make.
- The “So What?”: The task-clustered bootstrap treats 40 hand-written items as exchangeable, licensing inference about a task population they were never sampled from. The intervals are a lower bound on uncertainty, not a generalization license.
- Why It Matters: The estimand sentence should be written before anything runs. I wrote it after, and it still caught things.
🧪 Why This Is a Measurement Problem
Catching the bug took two instincts at once: software-testing instincts (your parser is code; test it like code) and measurement instincts (your grader is an instrument; characterize its errors, check whether they’re symmetric, define the estimand). That intersection — treating an evaluation pipeline as an empirical instrument that requires calibration and failure-mode profiling before its outputs can be trusted — is what turns model evaluation from ad-hoc scripting into a real measurement science.
🚧 Challenges & Learnings
- Challenge 1: Unblinded, single-rater adjudication. I adjudicated knowing the model, variant, and narrative stakes.
- What I Learned: Publish the transcripts. Every one of the 20 candidate flips — confirmations, the overturn, and the reversed call — is public so anyone can re-adjudicate and check me for confirmation bias. Transparency is the only defense I had, so I used all of it.
- Challenge 2: Triage instead of a census. 86 of 480 records were read in full; the rest are retained/imputed labels. Among 158 fallback-extracted passes, 20 were sampled with zero false positives — a bound, not a proof.
- What I Learned: State the sampling bound plainly (≈20 possible false positives one-sided) and let it constrain the claim: the 16-success lead could theoretically be erased under a sufficiently adversarial arrangement. The price of triage is a weaker headline.
- Challenge 3: Reproducibility gaps. Top-p/top-k weren’t logged; Python’s
hash(task_id)seeding isn’t pinned across processes; one output was truncated at the 256-token cap.- What I Learned: Log everything the sampler touches. A study about measurement rigor shouldn’t have measurement gaps in its own methods section.
📚 Code & Data
- Materials: Code, the 40 tasks, raw outputs, adjudications, and all 20 flip transcripts.
- Prior art: Sclar et al. (2024) on format sensitivity · Hua et al. (2025), “Flaw or Artifact?” · Romanou et al. (2026), BrittleBench · Jacobs & Wallach (2021) · Wallach et al. (2025) · Gelman & Stern on the significance-comparison fallacy
Methods note: Qwen 2.5 1.5B Instruct and Llama 3.2 1B Instruct (quantized GGUF, llama-cpp-python, CPU), 40 author-written arithmetic tasks, 3 instruction wordings (V3 adds a persona), 2 trials at temperature 0.7. 10,000-replicate task-clustered bootstrap. Nothing here generalizes to frontier scale — the value is concreteness, including the fragility of the author’s own first conclusion.
