Claude Haiku 4.5
+77With Tokonoma
127 of 140 solved
90.7%
Without Tokonoma
50 of 140 solved
35.7%
Tokonoma makes that expertise accessible so that teams solve more problems with no variance.
| Model | Without Tokonoma | With Tokonoma |
|---|---|---|
| Claude Haiku 4.5 | 50 of 140 solved | 127 of 140 solved |
| Claude Sonnet 5 | 74 of 140 solved | 138 of 140 solved |
| Claude Opus 4.8 | 69 of 140 solved | 133 of 140 solved |
With Tokonoma
127 of 140 solved
90.7%
Without Tokonoma
50 of 140 solved
35.7%
With Tokonoma
138 of 140 solved
98.6%
Without Tokonoma
74 of 140 solved
52.9%
With Tokonoma
133 of 140 solved
95.0%
Without Tokonoma
69 of 140 solved
49.3%
Two cases started with the same alarming error. One was a harmless maintenance blip. The other was a real outage. Tokonoma gave the agent useful past experience without stopping it from checking what was happening now.
0%
recognized without Tokonoma
100%
recognized with Tokonoma
Without the team's past experience, the agent treated the familiar false alarm like a new problem every time. With Tokonoma, all three models recognized the harmless pattern in every valid run.
59 / 60
real outages caught
Up to
15 fewerUp to 15 fewer
tool calls per investigation
Tokonoma did not make the agents overconfident. They still checked the current evidence, caught 59 of 60 real outages, and needed up to 15 fewer tool calls per investigation.
We included cases where past experience should help, where it should be ignored, and where it was intentionally wrong. That tests both whether the agent solves more problems and whether it can still think for itself.
Past experience points past a false clue
Past team knowledge is required
Tests whether the agent still raises the alarm
Significant help from Tokonoma
Tests whether the agent still raises the alarm
Tokonoma did not surface unrelated investigation history
Unrelated history
Unrelated history stayed out of the investigation.
The control had no recalled history. Tokonoma did not surface the unrelated record. Both groups solved 20 of 20.
Deliberately wrong records
A record is a lead to check, not an order to follow.
We planted a past diagnosis that looked related but was wrong. In 20 of 20 runs, the agent set it aside and solved the case from the current evidence.
Your incidents next
Bring us incidents your team has already solved. We will run the same coding agent with and without Tokonoma, then show the difference in plain numbers.