Derive datasets that cover every case.
The Dataset Agent derives a holistic test set straight from your production traces and real failures, so it covers the cases your agent actually meets and gives a score you can trust.
Available today as a guided design-partner engagement.
A green test suite is only as honest as the set behind it.
Most eval sets are written once and left to rot, so they drift away from production.
From production traces to a living test set
It builds the set your evals run on, and keeps it current.
Read your history
*discover-datasetIt pulls the high-signal runs, wins and failures, from your traces.
Group by scenario
Cases are held out and grouped so the set matches what your agent really meets, hard cases weighted.
Seed when thin
*build-datasetWith few traces, it seeds a starting set from examples you and your team agree on.
Anchor the hard cases
A short review pins the ambiguous cases to human judgment, the ones a rubric misses.
Grow from failures
Every production failure becomes a permanent case, so the set compounds over time.
Built from the traces you already have.
It reads your production history and distills the set your evals run on.
- A held-out regression set
- Hard cases weighted
- A seed set on cold start
- A set that only grows
What you can do with it
Real-trace set
Mined from what your agent actually did.
Matches production
Grouped by scenario, hard cases weighted.
Starts day one
Seeded from a few examples you agree on.
Never goes stale
Every failure folds in, so it only grows.
Dataset Agent vs A static or synthetic eval set
Most teams reach for a hand-written or LLM-generated eval set. It ships once, then quietly stops matching production. The Dataset Agent is the living alternative, sourced from your own traces.
Questions
How is this different from the Evaluator?
They are siblings inside the Evaluate stage. The Dataset Agent owns the set: the living collection of cases mined from traces and grown from every failure. The Evaluator is the judge that scores an agent against that set. The dataset is the input; the verdict is the output. You need both, and they are separate agents so neither grades its own work.
Where do the cases actually come from?
From your real production traces and the failures diagnosed in them, read from your observability platform. discover-dataset distills a held-out regression set out of that history, stratified by scenario with edge cases over-weighted. Nothing is guessed up front: the set is sourced from what your agent really did, not a synthetic sample.
What if we do not have many traces yet?
That is the cold-start case. build-dataset bootstraps a seed set from confirmed human anchors, examples you and your reviewers agree on, so you have something to evaluate against on day one. Be clear on what this is: a seed frame that compounds as real traces arrive, not a finished set. It gets stronger the more your agent runs.
Why not just write the eval set by hand or generate it synthetically?
Because a hand-written or synthetic set is built on curated data that already differs from production, and it drifts further with every change. That is the exact mechanism behind tests pass but quality drops: the set stops covering the edge cases and failure modes your agent meets in the wild. Distilling from real traces keeps the set honest.
How does it keep the set from going stale?
The set grows monotonically. Every production failure gets codified into it as a new case, criteria and cases compound over time, and the suite never shrinks. So a green run stays anchored to a bar that still matches how your agent actually behaves, instead of a frozen set that drifted the day it shipped.
What does it not do?
It analyzes traces, it does not produce them: trace collection and structured logging are upstream observability work, and this agent assumes those traces already exist. It curates the set, it does not judge, that is the Evaluator. And on cold start the bootstrap gives you a seed frame that compounds over time, not a complete regression set on the first run.
How does human judgment enter the set?
Through a human-in-the-loop seed interview that calibrates the set to human labels, and through the confirmed human anchors used to bootstrap a cold-start set. Human judgment is where it matters most: pinning down what good looks like on the cases a blank rubric would never capture.
How is this delivered?
It is a research preview, built and running against real traces, not a self-serve download. Book a walkthrough and we will wire up your trace sources, distill your first held-out regression set, and show the set growing from your own failures. You see it on your own traces before you commit.
Give your evals a set that still looks like production.
Book a custom demo and we will distill your first set from your traces.