Why Your LLM Judge Grades Its Own Homework: Measuring 23.6% Preference Leakage in Automated Evals
When automated evaluations use the same model family for generation and scoring, measured scores inflate dramatically. Here is empirical data across 140 test runs and how multi-model grading solves it.