Which hypothesis best explains all of the evidence observed so far? Weigh every observation, including any that contradict a hypothesis.
The model judges. Plain code decides.
An on-call teammate that investigates every alert in seconds. It finds what broke (the service, the deploy, the config change), shows the evidence, says "verified" only when it's sure, and fixes the safe cases itself.
Today it finds the faulty service from your Prometheus metrics. Coming: deploys and config changes, traces, logs, Slack, incident memory and safe fixes. See the roadmap
- Right service
- 58/60
- Passed verification; none of those 12 were wrong
- 12/60
- Per investigation
- ~7 s, ~$0.004
On 60 staged fault injections it had never seen, from the RCAEval Online Boutique benchmark. Model: TypeSafe Jev 1.13. Time and cost are averages.
A recorded investigation of a staged benchmark incident: checkout slowed down. Watch it narrow 12 suspects to one.
Most likely cause
Verification rule, enforced by code
Checkoutservice led from the first check. Frontend had been plausible, so code held the verdict until the model said a check had ruled frontend out with at least 0.50. It took until round 7.
Root cause: checkoutservice (verified) What failed: - checkoutservice_dashboard: latency 0.231 -> 0.478 2.07x +35s CPU 0.374 -> 16.35 43.68x +9s memory 10 MiB -> 246 MiB 24.98x +9s Where it showed: - latency_by_service: checkoutservice 2.07x, left normal range at +35s Ruled out: emailservice, frontend, paymentservice, productcatalogservice, redis, shippingservice Not used as evidence (failed or empty): checkoutservice_disk Checks run: 8
Where it's going.
From "which service broke" to "what changed, why, and the fix", with the same rule throughout: it only says verified when the evidence holds.
Live today
- Finds the faulty service from Prometheus metrics.
typerca initsets it up in one command. - Finds when it broke from a rough time like "20 min ago".
- Sees network slowdowns on the right service, by measuring each service from its callers' side.
- Says "verified" only when sure. Plain code checks the evidence; everything else is labelled a lead.
- About 7 seconds and under a cent per investigation.
- Every investigation can be replayed step by step.
Known limits today: on network problems it sometimes still blames the wrong service, and it only reads metrics for now.
Next
- Deploys and config changes as causes. "The checkout deploy at 03:10", not just "checkout".
- Runs on every alert, answers in Slack. Posted before anyone opens a laptop, with a one-click "right / wrong".
- Incident memory. "This looks like last month's outage, fixed by rolling back the cart deploy."
Later
- The whole call chain. It already measures each service from its callers; next it follows the full chain, so a slow network is never blamed on the caller.
- Logs. Which error messages are new since things broke.
- Suggested fixes, then safe automatic ones. Roll back or restart only when verified and your policy allows it.
- Prove it on your own system. Inject test faults in staging and get a score before you trust it.
- More sources beyond Prometheus, such as Datadog and CloudWatch.
How it works
- 1
Tell it when something broke
Give it the time of the alert. It looks at your metrics and lists the services that changed around then.
- 2
It follows one clue at a time
It looks at one thing, such as a service's memory or error rate, then asks an AI decision model which service best explains what it has seen so far. It stops after a set number of looks. The AI decision model today is TypeSafe's Jev.
- 3
It only says "verified" when the evidence holds
The AI weighs each clue, but plain code makes the final call. An answer is marked verified only if a clue points to it and every other likely service has been ruled out. Otherwise it's marked unverified: a strong lead, not proof.
What the model does, and what code does.
The AI decision model only answers small, fixed questions. Plain code makes every move and every final call, so the rules are readable, testable and the same every time.
See the exact questions the model answers
Does the evidence observed so far rule out frontend as the root cause? Yes only if a specific observed result is inconsistent with it. Not having tested it is not ruling it out.
Where does this fall on an ordered scale? One distribution over the levels.
Tested on incidents it had never seen.
60 Online Boutique fault injections from RCAEval, built twice: without the timeline and per-service dashboard checks (A) and with them (B). Same model, same loop, at most 8 checks, preregistered.
| Out of 60 | A | B | Paired |
|---|---|---|---|
| Right service | 57 | 58 | p = 1 |
| Supported diagnosis | 44 | 58 | p = 0.001 |
| Verified by the strict rule | 0 | 12 | p < 0.001 |
| Verified but wrong | 0 | 0 | |
| Mean checks run | 7.98 | 6.80 |
Model: Jev 1.13, from TypeSafe. Estimated model cost per investigation with B: $0.0041.
- Right service
- Named the service where the fault was injected.
- Supported diagnosis
- Right service, and its cited evidence holds up: it cited only checks it ran, no failed or empty query, and at least one check whose result shows the injected fault. A per-service dashboard counts, since it holds every metric.
- Verified
- Passed code's strict rule: a check supports the answer and every once-plausible alternative is judged ruled out.
Disk faults with a supported diagnosis, up from 0 of 10.
Both were network faults blamed on a neighbouring service. Still open.
The baseline run replays exactly in CI on every change.
Also tried on a live system: 6 test faults, 4 found, 1 verified, none verified wrong. Both misses were network faults again. Six cases, not a benchmark.
Two more live rounds, measuring each service from its callers: 5 of 8 found, none verified wrong. Network problems now show on the right service, but it doesn't always pick it yet. Round 2, round 3. Step through all 14 live incidents.
Run it yourself.
Python 3.10 or newer, standard library only. It runs on TypeSafe's Jev decision model, with decision thresholds tuned for it.
| Provider | --model | Status |
|---|---|---|
| TypeSafe Jev | jev | Supported, used for every result on this page |
git clone https://github.com/mingleiw/typerca && cd typerca # tests run offline, no API key python3 -m unittest discover -s tests -t . # investigate recorded incidents (TypeSafe key for Jev) export TYPESAFE_API_KEY=... python3 -m typerca run --scenarios bench/re2-test2 \ --model jev --out runs/jev.jsonl # or a live alert: detect your metrics once, then investigate python3 -m typerca init --prometheus http://prometheus:9090 python3 -m typerca investigate \ --prometheus http://prometheus:9090 \ --around "20 min ago" --config typerca-config.json
init recognises Istio, Linkerd, OpenTelemetry, Prometheus HTTP histograms and cAdvisor, writes the queries, and shows which services have which data. Only if it finds nothing it knows, adapt examples/prometheus-config.json by hand. --around finds when things started; if you know the exact time, use --at.