workspace.evaluations.run. Final CTA: Launch evaluation.
Wizard by dataset
Dataset cards
Simulation · Autonomous · Production Then choose mode: Agentic eval · Single-turn eval · Multi-turn eval, and env: Development · UAT · Production (Production restricted to Superadmin/QA).Simulation case groups
Exact field labels come from
GET /eval/config.
Autonomous cases
User message · User variable · OutcomeProduction path
- Sessions — pick observability sessions (or filters that resolve to sessions).
- Metadata — fields from
/eval/config(includes ids such asexpected_outcome). - Metrics → Review → Launch evaluation.
/eval/run-production-sessions and mixed-metric run streams.
Metrics step
- Catalog metrics with threshold sliders
- Custom metric:
metric_id, modeagentic/single_turn/multi_turn, score typesnumeric/boolean/category - Context: Global context / Metric context

