Evaluation

How I treat AI evaluation as measurement: constructs, rubrics, rater reliability, and honest uncertainty. Mental-health evaluation of frontier models is a research direction, not work I have already done.

A benchmark score is a claim about a construct. It needs the same evidence any measurement does: a clear construct, an anchored rubric, rater reliability, validity checks, and honest uncertainty. I use that standard for AI models and for the people deciding with them. Human factors, psychophysics, and Bayesian modeling are the reason the measurement can be defended, not a separate headline.

What I work on

Evaluation is measurement. If a score moves, I want to know whether the behavior moved or the instrument did. That means writing down the construct, anchoring the scale to observable behavior, stating when an item cannot be assessed, and checking the grader before interpreting the effect. The public example is a regex audit, not a claim about a product model: Auditing the Grader.

Methods I have used

These are methods that already appear somewhere on this site. I am not listing plans as finished studies.

  • Psychophysics and signal detection, on the Research page.
  • Hierarchical Bayesian models in brms and Stan, used to separate evidence quality from response caution in that same program.
  • Inter-rater reliability (κ, ICC) for annotation quality on surgical video, under Research Assistant I on About.
  • A grader audit of a rule-based regex. The outputs came from Qwen 2.5 1.5B Instruct and Llama 3.2 1B Instruct. The grader was not an LLM and not a human. Case study.
  • Calibration with Platt scaling, on the surgeon dashboard proof of concept. That dashboard was validated on synthetic sessions. Clinical effectiveness remains untested.
  • Equivalence tests and validated nulls. On the Research page, effort shifted how much people responded and did not measurably change perceptual sensitivity. That null was bounded with two one-sided tests, not left as “not significant.” Measure: perception under arousal.

LLM-as-judge validation is work I am planning: reliability on ordinal ratings, human versus model agreement, and whether an automated judge can stand in for a human annotator. I do not have a public agreement result to cite. I am not listing it as a completed method.

Open questions

These are research interests, aimed at evaluation and benchmarking for high-stakes domains, especially mental health, for frontier-lab models. I have no clinical mental-health training, and I have not built or run a mental-health benchmark.

  • When a mental-health benchmark scores a model response as “safe,” what construct is that score measuring, and does it agree with what clinicians would judge?
  • How well do LLM judges agree with clinician raters on ordinal rubrics, and where do they diverge systematically (length, hedging, boilerplate crisis-line insertion)?
  • Can latent-variable models detect benchmark gaming: models that learn the rubric’s surface features rather than the behavior it is meant to capture?
  • How sensitive are benchmark rankings to reasonable alternative grading choices?

This work needs clinical expertise I don’t have. I am interested in collaborating with clinicians and clinical researchers.

Selected evidence

  • Auditing the Grader. An apparent prompt-sensitivity gap was the regex, not the models.
  • Annotation-quality audits using κ and ICC, summarized on About.
  • A bounded null from the dissertation program: effort did not tune perceptual sensitivity, within a pre-stated equivalence bound. Research.
Back to top