Can AI Agents Do AI Research? ResearchGym and the Push to Benchmark Agentic Science
ResearchGym withholds five papers' methods and asks agents to rediscover them. Agents beat the baseline in 1 of 15 evaluations and finished 26.5% of sub-tasks.
By Institute for Joint Cognition & AI · 5 min · 6 October 2026

ResearchGym (arXiv:2602.15112) takes five oral and spotlight papers from ICML, ICLR, and ACL, keeps their datasets, evaluation harnesses, and baseline implementations, and withholds the method each paper proposed. An agent is dropped into the resulting containerised environment and asked to do what the authors did: form hypotheses, run experiments, and beat a strong human baseline on the paper’s own metrics.
The design is more honest than most agentic-science evaluations, and the results are correspondingly less flattering. Agents improved over baselines in 1 of 15 evaluations — 6.7%, by a margin of 11.5% — and completed 26.5% of sub-tasks on average.
Why the construction matters
Most claims about AI systems doing research rest on evaluations with a structural weakness: either the task is small enough to be a coding problem, or the grading is done by a language model asked whether the output looks like good research.
ResearchGym avoids both. The five environments are drawn from Materials Tokenization (ACL Spotlight), Cross-Modal Retrieval with Query Shift (ICLR Spotlight), TIMING temporality-aware integrated gradients (ICML Spotlight), SD-LoRA continual learning (ICLR Oral), and Prioritized Generative Replay (ICLR Oral), subdivided into 39 sub-tasks. Grading is execution-based, derived from the original papers’ own evaluation scripts, with task-native metrics — F1, accuracy, recall — plus a normalised score against the paper’s result. No LLM judge is in the loop.
That means a reported improvement is an improvement on the metric the original authors were optimising, measured by their code. It is a much harder thing to game than a rubric.
The withholding is the clever part. Because the proposed method is removed but everything around it is preserved, the agent faces the actual research problem the authors faced, with the same tooling and the same yardstick. Success means independently arriving at something competitive with a paper that cleared oral or spotlight review.
The capability-reliability gap
The headline finding is not that agents cannot do research. It is that they can occasionally, and not dependably.
In a single run, an agent surpassed the withheld solution on an ICML 2025 Spotlight task. Best@3 normalised performance ranged from 0.34 to 1.07 across tasks — the upper end meaning an agent exceeded the reference solution. That is a real result, and it is not compatible with a flat claim that agents cannot contribute to research.
It is also not compatible with the claim that they can be relied upon to. A 6.7% rate of improving on baselines, with 26.5% of sub-tasks completed, describes a system that produces a usable result sometimes and burns the budget otherwise. The distribution matters more than the maximum: a method that occasionally reaches state of the art but cannot be trusted to finish is an expensive lottery ticket, not a research programme.
Proprietary scaffolds did not escape the pattern. Claude Code (Opus-4.5) and Codex (GPT-5.2) display a similar gap, alongside the paper’s own ReAct-based rg-agent with GPT-5, AI-Scientist-v2, and ML-Master. The shortfall is not an artifact of one team’s harness.
The failure modes are about process, not knowledge
The recurring long-horizon failures the authors identify are worth listing precisely, because they are not failures of domain understanding:
- Impatience — abandoning runs before they produce signal
- Poor time and resource management — spending the budget on unproductive directions
- Overconfidence in weak hypotheses — committing to an idea the early evidence does not support
- Difficulty coordinating parallel experiments — losing track of what is running and what it showed
- Hard limits from context length — the experimental record outgrowing what the agent can hold
These are the skills a research advisor spends years transferring to a graduate student, and none of them are knowledge deficits. An agent that knows the literature and can write correct experimental code still fails on the meta-level task of allocating finite attention across a long uncertain search. That points at scaffolding and memory architecture as the binding constraint rather than model capability, which is a more tractable target than it might first appear.
Read the budget before reading the score
The results are conditioned on a specific and modest allocation: 12 GPU hours plus $10 of API budget per task, with the best runs receiving an additional 12 hours and $10, all on a single NVIDIA A100 80GB. Token consumption averaged roughly 0.18M input and 0.14M output per run.
This is a meaningful caveat in both directions. It means the benchmark measures research effectiveness under tight constraints, which is arguably the realistic setting and which rewards exactly the resource management the agents lack. It also means the numbers do not establish what the same agents would do with a hundred times the compute. Anyone citing the 6.7% figure as a ceiling on agentic research capability is over-reading it; anyone citing the single ICML success as evidence of arrival is over-reading it in the other direction.
What the benchmark deliberately excludes
The authors are explicit about scope. Research problems requiring multi-modal reasoning are excluded on hardware and data-transfer grounds. Training smaller models is out of reach, since only frontier LLMs achieve non-trivial results. Purely theoretical, analysis-driven, or proof-based papers are excluded because evaluating them requires subjective judgement.
That last exclusion defines the benchmark more than the others. ResearchGym measures empirical ML research with a runnable metric — the subset of science where success is a number that goes up. Whether an agent can do research that is not of that shape remains unmeasured, and the construction that makes ResearchGym rigorous is the same one that makes it inapplicable to theory.
The useful conclusion
ResearchGym is best read as an instrument for measuring reliability rather than capability. The capability question has an answer: frontier agents can sometimes match or exceed spotlight-level empirical work under a small budget. The reliability question has a worse answer, and it is the one that determines whether any of this becomes useful infrastructure.
For anyone building in this space, the failure list is the roadmap. Nothing on it requires a better model to address — they are all problems of experiment management, budget discipline, and keeping a long experimental record accessible. Those are engineering problems, which is the most encouraging thing in the paper.