SectionLanguage Models
5 pieces · Artificial Intelligence, Machine Learning & Computing Research · Est. 2007
Language ModelsBenchmark contamination makes MMLU and HumanEval unreliable for frontier models. What a credible, contamination-resistant evaluation pipeline requires.
By Institute for Joint Cognition & AI · 9 min · 26 June 2026
Language ModelsMulti-agent deployments fail mostly on orchestration, not weak models — race conditions, memory gaps, timeouts. Evaluation must look beyond task success.
By Institute for Joint Cognition & AI · 10 min · 26 June 2026
Language ModelsA technical guide to LLM evaluation: what MMLU, HELM, GSM8K, and HumanEval actually measure, benchmark contamination, and how to build reliable task-specific evals.
By Institute for Joint Cognition & AI · 13 min · 10 June 2024
Language ModelsWhy bigger context windows don't fix retrieval quality, what 'lost in the middle' means in practice, and when to use RAG, long context, or both.
By Institute for Joint Cognition & AI · 12 min · 25 April 2024
Language ModelsA technical guide to retrieval augmented generation: chunking, embeddings, vector search, reranking, failure modes, and how to evaluate your RAG pipeline.
By Institute for Joint Cognition & AI · 13 min · 22 January 2024