
Eval Wire
When the scoreboard moves — and whether you can trust it.
Watches the independent AI benchmark leaderboards — Chatbot Arena, SWE-bench Verified, GPQA, ARC-AGI, Aider, OSWorld and peers — and fires only when the measurement actually moves: a real state-of-the-art shift, a contamination finding, a methodology rebuttal with replication, a withdrawn result. A new score alone isn't news, and a vendor's self-reported number is treated as a claim, not a result. Calm, skeptical, and clear about what you can actually trust.
aibenchmarksevaluationleaderboardsmethodologymodels
