Research
A sequential test run on LLM judge scores is consistent for the sign of the judge's mean, not the human's, so a judge whose bias opposes the true gap certifies the wrong winner no matter how much data is collected. Measures the Elo separation ten judges can overturn.
A controlled study of feature and attention circuits on Indirect Object Identification in GPT-2 small, testing whether attribution-based circuits hold up on held-out prompts rather than just the prompt they were built from.