Install
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool
1+ week, 1+ hour ago (173+ words) Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar, Liqun Cheng, Ming Liu, Parthasarathy Ranganathan, et al. Samuel Kushnir, Kimia Noorbakhsh and colleagues at MIT, Google, Google DeepMind and Stanford keep a performance-modeling library whose main branch contains almost no code: the repository…...
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
1+ week, 18+ hour ago (238+ words) Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera, Adeline Kassler, Dmitrii Troitskii, Alexandra Souly, et al. Axel Ahlqvist and colleagues at the UK AI Security Institute, Meridian and Anthropic attack evaluation awareness, the problem that capable models can tell when they are…...
Terminal-Bench 3.0
2+ week, 2+ day ago (613+ words) measures agent abilities at the frontier Built by the team behind earlier Terminal-Bench releases and Harbor, Terminal-Bench 3.0 is an on-going large-scale community effort to source the most diverse, difficult and high-quality tasks. Our first release contains 74 tasks across 7 domains. The…...
Build your own agent benchmark
3+ week, 6+ day ago (896+ words) Veris is a simulation sandbox for building custom AI agent benchmarks around your own policies, workflows, and systems. Evaluate the agents you buy and the agents you build, end to end, against the job they actually have to do. Useful…...
Optima tackles AI benchmarking's biggest flaw by letting users test models against their own data
4+ week, 1+ day ago (405+ words) Optima lets users build custom benchmarks and compare AI models on quality, cost, and speed for their specific use cases. Users can build their own benchmarks using their own data, workflows, or descriptions of a use case, then run them…...