AI ENGINEERING / RESEARCH
Evaluating an
AI judging panel
I built an AI judging panel for an industry awards program and compared it with human experts. The evaluation favored the AI so consistently that I wanted to check whether the evaluator was working.
Read the case studyLLM evaluation · Python · Study design
WINNER AGREEMENT WITH HUMAN EXPERTS
7/7
category winners matched
on held-out applications
✓
✓
✓
✓
✓
✓
✓
Tested retrospectively on 29 held-out applications.
The winners matched; individual scores still differed.
The winners matched; individual scores still differed.
330 → 10human grading hours
↗