TypeRCA
Open source MIT Python standard library only

The model judges. Plain code decides.

A root-cause analysis agent built on decision models. It names the component at fault, shows what failed inside it, and says when it isn't sure.

Right component
58 of 60
Wrong, of 12 verified
0
Per investigation
~8 s, ~$0.004

On 60 staged fault injections it had never seen, from the RCAEval Online Boutique benchmark. Time is the run below; cost is the average.

How it works

  1. 1

    Give it an alert

    Point it at an alert time and your metrics. The services that moved become the suspects.

  2. 2

    It checks one thing at a time

    Each round it runs one check, within a budget you set, and asks a decision model which suspect fits. The answers are probabilities, not prose.

  3. 3

    Code decides when it's proven

    Verified only when the evidence supports it and every other plausible suspect is ruled out. Then it looks inside the culprit for what failed, and code writes the report.

A real recorded incident: checkout got slow. Watch it narrow 12 suspects to one.

typerca run --cases rec_ob_025a43e8ef --model jev

Most likely cause

Proven? Code decides

Waiting for evidence
0 of 8 checks, 0.0 s

How it's different from LLM agents.

Many AI root-cause agents let a large language model run the whole investigation: what to look at, when to stop, and what to conclude. TypeRCA splits the job. The model only answers narrow questions; code does the rest.

LLM runs the investigationTypeRCA
Next stepThe LLM writes what to look at nextThe model picks from a fixed list of checks, and code runs them
When it stopsWhen the LLM says it's doneWhen a written rule in code is met, inside a fixed budget
"Confident" meansThe LLM's own wordingEvery plausible alternative was ruled out by a check that actually ran
OutputProse you have to trustA report built by code from measured numbers, labelled verified or unverified
No good answerDepends on the promptCan answer "none of these suspects" and escalate
AuditHard to reproduce a runEvery answer is recorded, and runs replay exactly
CostLong prompts and long outputs at every step22 short decisions and about 8 s in the run above; about $0.004 per investigation on average
See the exact questions the model answers
Choicebest_explanation
Which hypothesis best explains all of the evidence observed so far? Weigh every observation, including any that contradict a hypothesis.
After latency_by_service: checkoutservice 0.97, redis 0.02
Yes or noruled_out:frontend
Does the evidence observed so far rule out frontend as the root cause? Yes only if a specific observed result is inconsistent with it. Not having tested it is not ruling it out.
Round 7: yes 0.54, clears the 0.50 rule
Scoreordered rubric
Where does this fall on an ordered scale? One distribution over the levels.
Part of the interface every provider implements

Tested on incidents it had never seen.

60 Online Boutique fault injections from RCAEval, built twice: without the version 2 checks (A) and with them (B). Same model, same loop, at most 8 checks, preregistered.

Out of 60ABPaired
Right service5758p = 1
Supported diagnosis4458p = 0.001
Verified by the strict rule012p < 0.001
Verified but wrong00
Mean checks run7.986.80

Model: Jev 1.13. Estimated model cost per investigation with B: $0.0041.

10 of 10

Disk faults with a supported diagnosis, up from 0 of 10.

2 misses

Both were network faults blamed on a neighbouring service. Still open.

120 replayed

The baseline run replays exactly in CI on every change.

Read the full report

What a reviewer would find.

Fixed 1

  • Disk faults without evidence. Per-suspect dashboards made all 10 fresh cases supported.

Built, not yet measured 2

  • Open world. Agent version 2 can answer "none of these suspects", widen into a reserve, and escalate.
  • Live incidents. Investigates straight from Prometheus. Tested against a fake server only.

Open 3

  • Network faults still get blamed on a neighbour. Call-graph data from traces is next.
  • Metrics only, from one benchmark family, on staged single-fault incidents.
  • One provider. A second adapter waits on a verifiable API spec.

Run it in a minute.

shell
# tests run offline, no API key
python3 -m unittest discover -s tests -t .

# investigate recorded incidents
export TYPESAFE_API_KEY=...
python3 -m typerca run --scenarios bench/re2-test2 \
    --model jev --out runs/jev.jsonl

# or a live alert from Prometheus
python3 -m typerca investigate \
    --prometheus http://prometheus:9090 \
    --at 2026-10-01T03:12:00Z
my_model.py
from typerca.models import BaseModel, Capabilities

class MyModel(BaseModel):
    name = "my-model"
    capabilities = Capabilities(max_questions=20)

    def _ask(self, request):
        # send request.state and request.questions,
        # return one probability per option
        ...

Bring any decision model with one method. Validation, request splitting and cost accounting are shared.