Peeking at a Biased Judge
Overview
Two practices are near-universal in LLM evaluation: stopping early, and scoring examples with an LLM judge rather than a human. Anytime-valid methods like e-processes exist to make the first one safe. This paper shows the combination is not merely inefficient. An e-process run on judge scores is a valid test of the mean of the judge's scores, not the human estimand. When the judge's bias opposes the true gap, the procedure certifies a winner the human data cannot support with probability tending to one, and no sample size repairs it.
Key Findings
- Judge bias is sign-varying, not a constant offset. Across the 30 (pair, turn) cells of MT-Bench the GPT-4 judge's bias ranges from −0.392 to +0.429 and changes sign in 16 of 30 cells, while pooled bias is +0.001. Any evaluation reporting a single aggregate agreement or bias number is reporting the quantity in which this failure is invisible.
- Bias converts into a decision-relevant threshold: the Elo separation a judge with opposing bias can overturn, Δmin(b) = 400·log10((1+b)/(1−b)), regardless of how much data is collected.
- Across ten judges spanning four vendors, a 4B–32B scale ladder, and one generation step (36,280 judgments), every judge requires 40–71 Elo of separation, with 95% intervals typically 30–80.
- Scale does not fix it. Bias is flat across the Qwen3 ladder (spread 0.022) while the spread across vendors at 20–31B is 3.2× larger (F = 13.98, p = 0.029). Three of nine modern open-weights judges are worse than the 2023 GPT-4 baseline (median 0.161 → 0.137).
- Practically: detecting that a judge is biased at all takes about 90 labels; resolving Δmin to ±20 Elo takes roughly 540. Below that separation, a judge-only e-process cannot resolve two models at any sample size — the fix is a prediction-powered rectifier run on the labels you do have.