Derive datasets that cover every case.

The Dataset Agent derives a holistic test set straight from your production traces and real failures, so it covers the cases your agent actually meets and gives a score you can trust.

Available today as a guided design-partner engagement.

Integrates with every observability platform and agent framework

Vercel AI SDK
OpenAI
LangChain
LangGraph
Mastra
Langfuse
LangSmith
Braintrust
Phoenix
Datadog
Claude Code
Cursor
Codex
OpenCode
Vercel AI SDK
OpenAI
LangChain
LangGraph
Mastra
Langfuse
LangSmith
Braintrust
Phoenix
Datadog
Claude Code
Cursor
Codex
OpenCode

A green test suite is only as honest as the set behind it.

Most eval sets are written once and left to rot, so they drift away from production.

Tests stay green while quality slips.
The set is mined from your real traces, not guessed.
Your eval set no longer looks like production.
It is grouped by scenario, with the hard cases weighted.
No eval ever gets off the ground.
A starter set is seeded from a few examples you agree on.
A frozen set drifts the day you ship it.
Every failure gets folded in, so the set only grows.

From production traces to a living test set

It builds the set your evals run on, and keeps it current.

Your history becomes a labelled set with criteria and a stable score.
01

Read your history

*discover-dataset

It pulls the high-signal runs, wins and failures, from your traces.

02

Group by scenario

Cases are held out and grouped so the set matches what your agent really meets, hard cases weighted.

03

Seed when thin

*build-dataset

With few traces, it seeds a starting set from examples you and your team agree on.

04

Anchor the hard cases

A short review pins the ambiguous cases to human judgment, the ones a rubric misses.

05

Grow from failures

Every production failure becomes a permanent case, so the set compounds over time.

Built from the traces you already have.

It reads your production history and distills the set your evals run on.

Reads from
LangfuseLangSmithOpenTelemetryPhoenix
What you get
  • A held-out regression set
  • Hard cases weighted
  • A seed set on cold start
  • A set that only grows

What you can do with it

Real-trace set

Mined from what your agent actually did.

Matches production

Grouped by scenario, hard cases weighted.

Starts day one

Seeded from a few examples you agree on.

Never goes stale

Every failure folds in, so it only grows.

Dataset Agent vs A static or synthetic eval set

Most teams reach for a hand-written or LLM-generated eval set. It ships once, then quietly stops matching production. The Dataset Agent is the living alternative, sourced from your own traces.

A static or synthetic eval setDataset Agent
Where cases come fromHand-written or LLM-synthesized, guessed up frontMined from your real production traces and diagnosed failures
Match to productionDiffers from the real distribution the day it shipsStratified from the distribution your agent actually meets, edge cases over-weighted
Over timeFrozen; goes stale as prompt and traffic changeGrows monotonically: every failure codified in, the suite never shrinks
Cold startBlank page; no eval ever gets off the groundBootstraps a seed set from confirmed human anchors on day one
Anchor for labelsAuto-labeled or skipped, so the answer is a guessCalibrated to human labels through a human-in-the-loop seed
Who scores itWhatever harness happens to read the fileHanded to the Evaluator, a separate agent so neither grades its own work

Questions

How is this different from the Evaluator?

They are siblings inside the Evaluate stage. The Dataset Agent owns the set: the living collection of cases mined from traces and grown from every failure. The Evaluator is the judge that scores an agent against that set. The dataset is the input; the verdict is the output. You need both, and they are separate agents so neither grades its own work.

Where do the cases actually come from?

From your real production traces and the failures diagnosed in them, read from your observability platform. discover-dataset distills a held-out regression set out of that history, stratified by scenario with edge cases over-weighted. Nothing is guessed up front: the set is sourced from what your agent really did, not a synthetic sample.

What if we do not have many traces yet?

That is the cold-start case. build-dataset bootstraps a seed set from confirmed human anchors, examples you and your reviewers agree on, so you have something to evaluate against on day one. Be clear on what this is: a seed frame that compounds as real traces arrive, not a finished set. It gets stronger the more your agent runs.

Why not just write the eval set by hand or generate it synthetically?

Because a hand-written or synthetic set is built on curated data that already differs from production, and it drifts further with every change. That is the exact mechanism behind tests pass but quality drops: the set stops covering the edge cases and failure modes your agent meets in the wild. Distilling from real traces keeps the set honest.

How does it keep the set from going stale?

The set grows monotonically. Every production failure gets codified into it as a new case, criteria and cases compound over time, and the suite never shrinks. So a green run stays anchored to a bar that still matches how your agent actually behaves, instead of a frozen set that drifted the day it shipped.

What does it not do?

It analyzes traces, it does not produce them: trace collection and structured logging are upstream observability work, and this agent assumes those traces already exist. It curates the set, it does not judge, that is the Evaluator. And on cold start the bootstrap gives you a seed frame that compounds over time, not a complete regression set on the first run.

How does human judgment enter the set?

Through a human-in-the-loop seed interview that calibrates the set to human labels, and through the confirmed human anchors used to bootstrap a cold-start set. Human judgment is where it matters most: pinning down what good looks like on the cases a blank rubric would never capture.

How is this delivered?

It is a research preview, built and running against real traces, not a self-serve download. Book a walkthrough and we will wire up your trace sources, distill your first held-out regression set, and show the set growing from your own failures. You see it on your own traces before you commit.

Give your evals a set that still looks like production.

Book a custom demo and we will distill your first set from your traces.