02 · Evaluate & testComing soon
Experiment Agent
Compares prompts, models, and configs against your eval criteria
Runs N candidates, different prompts, models, and tool configs, against your dataset and evaluator, then outputs a ranked comparison with cost, latency, and quality per candidate. Pick a baseline, prove a candidate beats it.
Use it to pick a baseline before optimisation, validate a prompt mutation against a held-out set, or pressure-test a model swap before committing.
What it does
- Runs candidates across prompts, models, and configs
- Scores each against your eval criteria
- Compares cost, latency, and quality side by side
Inputs
- Prompt candidates
- Eval datasets
- Model provider credentials
Outputs
- Experiment harness wired to dataset and evaluator
- Ranked comparison report per run
- A decision artifact: which candidate ships, and why
Works with
Get early access to Experiment Agent
Join the early-access list and we will reach out the moment this agent ships.