How multi-choice academic benchmarks suffered from subtle pre-training overlap.
As web-scale pre-training datasets incorporated academic exams, standard multiple-choice evaluations like MMLU and GSM8K lost discriminative power.
Modern evaluation protocols mandate dynamic, privately held test sets with zero internet footprint, or dynamic execution-based problems like SWE-bench.