AI Reliability Research
Source notes
These linked sources support the statements in the guide. Accessed 5 October 2026. See each publisher for its own date, methods, licensing and later corrections.
- WebArena: A Realistic Web Environment for Building Autonomous AgentsZhou et al., arXiv · Research paper
Introduces a reproducible web environment and task-based evaluation.
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer EnvironmentsXie et al., arXiv · Research paper
Presents tasks in real computer environments spanning applications and file operations.
- OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World TasksarXiv preprint · Research paper, June 2026
Describes 108 longer workflows and reports both completion and partial-credit measures; preprint findings are not a universal ranking.
- AI RMF Core: MeasureU.S. National Institute of Standards and Technology · Framework guidance
Calls for documented methods, metrics, uncertainty, benchmark comparison and performance assessment.
How to read this edition
The guide summarizes only the claims described beside each source. Research findings, technical standards and organizational policies have different purposes; they are labeled accordingly.
Image
Homepage and social-sharing image: A person working with code on a laptop. Photo by Alicia Christin Gerald. View the original photograph, used under the Unsplash License. The image illustrates the subject and is not evidence for the guide’s claims.