AI Reliability Research

What agent benchmarks can tell us

A benchmark score is evidence about a task, not a guarantee about an agent.

A reader’s guide to tool-using AI evaluations: what gets tested, what counts as success, and what a result leaves out.

Read the guide   See the sources

A person working with code on a laptop

Photo: Alicia Christin Gerald

Start with the evidence

Read evaluation papers with the same care you would bring to a measurement instrument: inspect the task and the yardstick first.

This short edition draws on the sources listed in Source notes. Each page says what its evidence supports and where interpretation begins.