Gathering incident evidence so on-call engineers start from facts
The platform team at a software company, whose on-call engineers respond to alerts across many services, dashboards and runbooks.
- Challenge
- The first stretch of every incident goes on opening dashboards, searching logs, checking recent deployments and finding the right runbook. Much of the know-how sits with a few senior engineers, and incident notes are pieced together from memory afterward.
- Approach
- 1Turn past incidents with known causes into an evaluation set, graded by senior engineers
- 2Give an agent read-only tools for metrics, logs, traces, deployment history and runbooks
- 3Post a timeline and likely causes to the incident channel, each linked to the query or log line behind it
- 4Let the agent propose runbook steps such as a rollback, which run only after the on-call engineer approves them
- 5Mask customer data before model calls, cap steps and spend per incident, and trace every tool call
- Outcome
- On-call engineers start from a timeline with the evidence attached instead of a blank screen. Every suggestion links to the data behind it, fixes stay in human hands, and each change to the agent is replayed against past incidents before release.
Services involved





