Install
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Your Eval Has Failure Modes Too
3+ week, 5+ day ago (77+ words) Part 3 of a review series on production agentic systems....
Your evals pass. That doesn't mean they work.
2+ week, 6+ day ago (628+ words) better than I expected, and one question kept coming up in different forms: "Okay — but wouldn't my eval...The agent said "payment successful," and every eval passed....Here's the thing about a green eval run: it tells you your system…...
Interactive AI Eval Dashboards with Data Studio
6+ day, 1+ hour ago (420+ words) Welcome to the final entry of our series about designing, analyzing and visualizing AI evals!...
Claude Code Just Shipped Evals for Plugins.
1+ day, 11+ hour ago (945+ words) On September 11, 2026, Claude Code 2.1.269 added one line to the changelog: “Added claude plugin eval...: run a plugin's eval suite against Claude Code and get scored, reproducible results (JSON + HTML report...
AI Evals at a Glance: Heatmaps for Stakeholders
2+ week, 6+ day ago (20+ words) Visualizing AI evals with Inspect Viz Welcome back to our blog series on running,......
Why Enterprise AI Needs A New Approach To Evals
1+ week, 6+ day ago (642+ words) performing tasks, teams must continually rebuild their evaluation layer (commonly known as their “evals...I focus on the following key criteria: • Lead With Specs: Detailed specs become even more important for...agentic evals....multiple calls to configured tools to complete…...
How to build your first AI Eval
3+ week, 2+ day ago (373+ words) Learn what an eval is: a test with answers you already know....Use evals as the foundation for AI systems that can improve their own work....
LLM evals are a parameter sweep — use a parameter sweep tool
2+ week, 4+ day ago (713+ words) The LLM eval space rebuilt that layer as SaaS, with a meter on it....Two methodological points before code, because they're where evals usually go wrong....would add cost, latency and sampling noise to a comparison that was already deterministic....
The Authentication Mistake Even Senior Developers Still Make
3+ week, 6+ day ago (31+ words) Not a vulnerability....
Same-Prompt Green Is Not a Gate: Five Eval Myths
1+ week, 3+ day ago (917+ words) I keep seeing the same broken eval loop. A model drafts the feature in one burst....This is an eval problem, not a prompt trick. "The tests passed, so the feature works."...