Turn a dropped eval score into the exact fix.

Mutagent clusters the failing traces, names each failure what, why, and where, and drafts the PR for you to approve.

Free, runs locally in your own coding agent. Nothing leaves your machine.

Integrates with your stack

Sound familiar?

You see the symptom. Mutagent names the cause.

At thousands of traces, the cause is buried. Mutagent clusters them and pins it.

Your eval score dropped and the dashboard stops there
The exact failure behind the drop, ranked as a fix
It passed every eval, then quietly broke in prod
The failures your evals missed, clustered and named
A confident wrong answer, and nothing errored
The hallucination traced to stale context, with a freshness check
A green pass hid a trajectory that drifted off policy
The drift named what, why, and where, as a regression test
See it run

A real 66-second diagnosis

Point it at your traces, then watch it cluster, find the root cause, and draft the fix.

From a wall of traces to a named root cause

Run it in your own coding agent, and it turns your whole trace history into a short list of named, located causes. Scroll through the steps.

  1. 01

    Run it

  2. 02

    Scan every trace

  3. 03

    Sort signal from noise

  4. 04

    Group and fan out

  5. 05

    Walk to the root cause

  6. 06

    Name it and propose the fix

No platform switch

Works with your stack. Fixes your stack.

Mutagent reads the traces wherever they live and lands the fix on whatever your agent runs on.

Sources in
Langfuse
OpenTelemetry
Datadog
{ }Raw JSONL
Claude Code logs
Codex logs
MUTAGENT/diagnostics
Targets out
Claude skills & subagents
OpenCode
Codex agents
Vercel AI SDK
Mastra & LangGraph
Deep Agents
</>Prompts
Proof

What it found

96%
of the failures we read were prompt or behavior, not the model
70%
of messages rejected from one 3-line contradiction
3 axes
name every failure: what, why, where
+11.6%
real uplift where two optimizers delivered 0

From the same recruiting-AI team teardown: 7,524 agent executions across a 24-hour production window on Langfuse, March 2026, anonymized at their request. The uplift is from our own variance-floor benchmark.

You can only fix what you can name. A failure with no name is a vibe.

Questions

Is it really free?

Yes. You install Mutagent into your own coding agent and run the diagnosis on your own traces. No card, no gate.

Isn't this just observability with extra steps?

No. Observability shows you what happened. Mutagent names why, on three axes, and opens the fix. It is the layer that acts, not another dashboard.

Does my data leave my machine?

No. It runs in your own environment, through your own coding agent. Your traces stay where they are.

What do I need to run it?

Traces in Langfuse, OpenTelemetry, raw JSONL, or your Claude Code or Codex session logs, plus a coding agent like Claude Code, Codex, Cursor, or OpenCode.

Will it change my code on its own?

No. It proposes; you own the judgment. It lands a change only after you approve it, as a PR on an isolated branch.

What happens on the call?

We walk through your own report together, show you the regression tests it would add, and scope what the full engineer would catch on your stack. No deck.

Turn a moved score into the change you make.

Run the diagnosis on your own agent, free. See the first finding, then book a call to see the full engineer.