Which hypothesis best explains all of the evidence observed so far? Weigh every observation, including any that contradict a hypothesis.
The model judges. Plain code decides.
A root-cause analysis agent built on decision models. It names the component at fault, shows what failed inside it, and says when it isn't sure.
- Right component
- 58 of 60
- Wrong, of 12 verified
- 0
- Per investigation
- ~8 s, ~$0.004
On 60 staged fault injections it had never seen, from the RCAEval Online Boutique benchmark. Time is the run below; cost is the average.
How it works
- 1
Give it an alert
Point it at an alert time and your metrics. The services that moved become the suspects.
- 2
It checks one thing at a time
Each round it runs one check, within a budget you set, and asks a decision model which suspect fits. The answers are probabilities, not prose.
- 3
Code decides when it's proven
Verified only when the evidence supports it and every other plausible suspect is ruled out. Then it looks inside the culprit for what failed, and code writes the report.
A real recorded incident: checkout got slow. Watch it narrow 12 suspects to one.
Most likely cause
Proven? Code decides
Checkoutservice led from the first check. Frontend had been plausible, so code held the verdict until the model said a check had ruled frontend out with at least 0.50. It took until round 7.
Root cause: checkoutservice (verified) What failed: - checkoutservice_dashboard: latency 0.231 -> 0.478 2.07x +35s CPU 0.374 -> 16.35 43.68x +9s memory 10 MiB -> 246 MiB 24.98x +9s Where it showed: - latency_by_service: checkoutservice 2.07x, left normal range at +35s Ruled out: emailservice, frontend, paymentservice, productcatalogservice, redis, shippingservice Not used as evidence (failed or empty): checkoutservice_disk Checks run: 8
How it's different from LLM agents.
Many AI root-cause agents let a large language model run the whole investigation: what to look at, when to stop, and what to conclude. TypeRCA splits the job. The model only answers narrow questions; code does the rest.
See the exact questions the model answers
Does the evidence observed so far rule out frontend as the root cause? Yes only if a specific observed result is inconsistent with it. Not having tested it is not ruling it out.
Where does this fall on an ordered scale? One distribution over the levels.
Tested on incidents it had never seen.
60 Online Boutique fault injections from RCAEval, built twice: without the version 2 checks (A) and with them (B). Same model, same loop, at most 8 checks, preregistered.
| Out of 60 | A | B | Paired |
|---|---|---|---|
| Right service | 57 | 58 | p = 1 |
| Supported diagnosis | 44 | 58 | p = 0.001 |
| Verified by the strict rule | 0 | 12 | p < 0.001 |
| Verified but wrong | 0 | 0 | |
| Mean checks run | 7.98 | 6.80 |
Model: Jev 1.13. Estimated model cost per investigation with B: $0.0041.
Disk faults with a supported diagnosis, up from 0 of 10.
Both were network faults blamed on a neighbouring service. Still open.
The baseline run replays exactly in CI on every change.
What a reviewer would find.
Fixed 1
- Disk faults without evidence. Per-suspect dashboards made all 10 fresh cases supported.
Built, not yet measured 2
- Open world. Agent version 2 can answer "none of these suspects", widen into a reserve, and escalate.
- Live incidents. Investigates straight from Prometheus. Tested against a fake server only.
Open 3
- Network faults still get blamed on a neighbour. Call-graph data from traces is next.
- Metrics only, from one benchmark family, on staged single-fault incidents.
- One provider. A second adapter waits on a verifiable API spec.
Run it in a minute.
# tests run offline, no API key python3 -m unittest discover -s tests -t . # investigate recorded incidents export TYPESAFE_API_KEY=... python3 -m typerca run --scenarios bench/re2-test2 \ --model jev --out runs/jev.jsonl # or a live alert from Prometheus python3 -m typerca investigate \ --prometheus http://prometheus:9090 \ --at 2026-10-01T03:12:00Z
from typerca.models import BaseModel, Capabilities
class MyModel(BaseModel):
name = "my-model"
capabilities = Capabilities(max_questions=20)
def _ask(self, request):
# send request.state and request.questions,
# return one probability per option
...
Bring any decision model with one method. Validation, request splitting and cost accounting are shared.