Section

Language Models

5 pieces · Artificial Intelligence, Machine Learning & Computing Research · Est. 2007

Language Models

Evaluating LLMs: What Benchmarks Measure and Where They Fail

A technical guide to LLM evaluation: what MMLU, HELM, GSM8K, and HumanEval actually measure, benchmark contamination, and how to build reliable task-specific evals.

By Institute for Joint Cognition & AI · 13 min · 10 June 2024