Don't let a model grade its own family
An LLM judge interrogates the outsider and believes its own kind — and it looks exactly like rigor
I spent a week comparing Claude Opus and Claude Sonnet on the same spec-driven pipeline, run against pdb_search, a brownfield Python repo of mine. Opus is the stronger model, Sonnet the cheaper one. To score the runs I used a third instance, an Opus 4.8 orchestrator. Same family as one...
[Read More]