AI Reliability Research

AI Reliability Research

Source notes

These linked sources support the statements in the guide. Accessed 5 October 2026. See each publisher for its own date, methods, licensing and later corrections.

  1. WebArena: A Realistic Web Environment for Building Autonomous AgentsZhou et al., arXiv · Research paper

    Introduces a reproducible web environment and task-based evaluation.

  2. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer EnvironmentsXie et al., arXiv · Research paper

    Presents tasks in real computer environments spanning applications and file operations.

  3. OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World TasksarXiv preprint · Research paper, June 2026

    Describes 108 longer workflows and reports both completion and partial-credit measures; preprint findings are not a universal ranking.

  4. AI RMF Core: MeasureU.S. National Institute of Standards and Technology · Framework guidance

    Calls for documented methods, metrics, uncertainty, benchmark comparison and performance assessment.

How to read this edition

The guide summarizes only the claims described beside each source. Research findings, technical standards and organizational policies have different purposes; they are labeled accordingly.

Image

Homepage and social-sharing image: A person working with code on a laptop. Photo by Alicia Christin Gerald. View the original photograph, used under the Unsplash License. The image illustrates the subject and is not evidence for the guide’s claims.