How Popular AI Benchmarks Are Built
Nine case studies: how tasks are created and scored, why researchers use them, and what could improve.
29 min read
A PERSONAL BLOG
Notes, ideas, and things I’m learning along the way.
Nine case studies: how tasks are created and scored, why researchers use them, and what could improve.
Verifiable and unverifiable benchmarks, explained through acceptance checks, fallible graders, and the gap between a highlight reel and reliable work.
From scattered notes to clearer expression, with a little room to think.
No matching posts. Try another word.