Dress rehearsal

Test runs grade the performance, not the costume.

Automate evals after agents run in Evalab. Define judging criteria and grade performance for every commit or PR.

PR TEST RUNS RUN-2048 · Passed 4 calls · 0 forbidden · 266ms POST /github/check-runs BLOCKED

Saved tests, test runs, and PR test runs live in the same workspace as twins and scenarios. Judging stays on the trace — every provider call is in the prompt book.

Read the docs