Incident Intelligence

One object.
All the evidence.

Six heuristics watch your error rates, uptime checks, latency, and deploys on a 60-second cycle. When one fires, the signals become a first-class Incident — with a likely cause already attached.

SEV·CRITICALINC-241 · openox-coffee · production
Error-rate spike · POST /checkout/payment
opened 14:32 UTC·4 signals·heuristic: error-rate ≥3× baseline
confidence
78%
Likely cause: deploy v2.14.1 — 4 min before first spike
14:28deployv2.14.1 deployed · GitHub Actions
14:31errorStripe::CardError ×412 · charge.rb:81
14:32uptimemonitor api.ox-coffee.io failed ×3
14:33apmp95 4,823 ms · 2.4× baseline
How it works

Signal to incident in four moves.

No configuration. The engine evaluates every project, every minute — and only promotes what crosses a threshold.

01 · Watch

Six heuristics, 60s cycle

New high-volume groups, uptime failures, error-rate spikes, latency anomalies, resource saturation, deploy correlation — evaluated every minute.

02 · Merge

Same key, same incident

Signals sharing an environment, top frame, host, monitor, route, or release SHA merge into one incident instead of paging you six times.

03 · Score

Deterministic likely cause

Backtrace overlap, time correlation, and similar prior incidents feed a signal-weighted confidence score. No LLM involved.

04 · Summarize

AI only when you ask

The evidence panel is complete without AI. Want prose? Hit Summarize — it runs behind an explicit action, never automatically.

The trigger conditions

Six heuristics. Real thresholds.

Each condition is documented, deterministic, and tuned to fire on incidents — not noise. Uptime failures enter as critical; everything else is scored by severity.

  • Evaluated on a 60-second cycle, per project
  • Manual incidents for what the heuristics don’t catch
  • Open → Acknowledged → Resolved lifecycle
Read the heuristics docs →
heuristicfires when
new high-volume group≥ threshold in first hour
uptime failureconsecutive check failures
error-rate spike≥ 3× baseline
p95 latency≥ 2× baseline
CPU / memory≥ 90% sustained
deploy correlationspike follows a release
Deterministic first, AI on demand

A confidence score, not a guess.

The likely-cause panel runs the moment evidence loads. Each contributing signal is shown with its weight — you can audit the score, not just trust it.

  • Signal-weighted confidence, contributing signals shown
  • AI summary only behind an explicit Summarize → action
likely cause · deterministicno AI credits used
deploy v2.14.1 · confidence 78%
backtrace overlap with deploy diff
+42
time correlation · 4 min gap
+24
similar prior incident · INC-187
+12
AI summary runs only when you ask
Merge keys

Evidence merges. Pages don’t multiply.

Six merge keys decide whether a new signal joins an open incident or opens a new one. One bad deploy is one incident — not an error page, an uptime page, and a latency page.

  • Keys: environment, top frame, host, monitor, route, release SHA
  • Evidence timeline keeps every merged signal inspectable
signal merge keyssame key → same incident
environmenttop_framehostmonitorrouterelease_sha
14:31 Stripe::CardError route=/checkout/payment → INC-241
14:33 p95 spike route=/checkout/payment → merged
14:35 CPU 91% host=web-2 → INC-242 (new)

Close the errorgap.

Start free, no credit card. Turn scattered production signals into one investigable incident in under five minutes. Cancel anytime, keep your data.