Peeking at a Biased Judge

Sequential Testing · Anytime-Valid Inference for LLM Evaluation with Model Judges

Overview

Two practices are near-universal in LLM evaluation: stopping early, and scoring examples with an LLM judge rather than a human. Anytime-valid methods like e-processes exist to make the first one safe. This paper shows the combination is not merely inefficient. An e-process run on judge scores is a valid test of the mean of the judge's scores, not the human estimand. When the judge's bias opposes the true gap, the procedure certifies a winner the human data cannot support with probability tending to one, and no sample size repairs it.

Paper (PDF) GitHub

Key Findings