Dress rehearsal
Test runs grade the performance, not the costume.
Automate evals after agents run in Evalab. Define judging criteria and grade performance for every commit or PR.
Saved tests, test runs, and PR test runs live in the same workspace as twins and scenarios. Judging stays on the trace — every provider call is in the prompt book.