Free operations resource
AI Agent Incident Response Checklist: the first 15 minutes
When a tool-using agent does something unexpected, the first job is not to explain the model. It is to stop harmful effects, preserve enough evidence to reconstruct what happened, and identify the smallest boundary you can safely restore.
No signup required. This page is designed to be printed or saved as a PDF.
0–3 minutes
Contain the effect without destroying the evidence.
☐ Freeze the specific action path that is causing harm.Stop the affected tool, worker, browser session, credential, queue, or deployment path. Avoid shutting down unrelated systems unless the incident is propagating.
☐ Separate reasoning from real-world custody.If the agent can still issue actions, revoke or isolate only the action channel first. Preserve read-only access to logs, messages, receipts, and state when possible.
☐ Mark the incident boundary.Record the first known bad effect, the last known good effect, affected accounts/systems, and whether the failure is still producing new external effects.
☐ Prevent blind retries.Pause automatic retry of any side effect whose success is ambiguous. A duplicate charge, message, order, deletion, or deployment is often worse than a delayed one.
3–7 minutes
Capture evidence from independent planes.
☐ Preserve the agent-visible conversation.Capture the relevant user request, assistant reasoning-visible outputs available to operators, tool requests, and final response. Do not assume the API trace alone contains the decision context.
☐ Capture tool and provider state.Record request IDs, provider receipts, HTTP responses, transaction IDs, filenames, commit IDs, browser URLs, and account-side postconditions. Distinguish “request accepted” from “effect actually happened.”
☐ Capture host/runtime state.Record the process/session identity, deployment or commit version, relevant environment/config version, resource pressure, and any restart/crash around the incident window.
☐ Capture a visual state when UI semantics matter.A screenshot can resolve ambiguity that DOM text, logs, or transport receipts cannot—especially around authentication, confirmation screens, stale pages, and browser custody.
7–11 minutes
Classify the failure before changing the system.
| Question | What it distinguishes |
|---|---|
| Did the agent choose the wrong action, or did the right action execute incorrectly? | Decision failure vs. execution failure |
| Was the effect duplicated, omitted, stale, or applied to the wrong target? | Idempotency / identity / freshness failure |
| Did the operator view match provider reality? | Observability failure vs. actual provider failure |
| Was the selected runtime the code you thought was running? | Source/runtime convergence failure |
| Did another agent, daemon, browser, or retry loop share the same custody? | Coordination/cross-custody failure |
11–15 minutes
Recover the smallest safe boundary.
☐ Choose one reversible recovery.Prefer a narrow rollback, credential rotation, queue hold, session replacement, or known-good deployment over a full rebuild.
☐ Prove the repair at the level that failed.If the failure was a provider effect, verify the provider. If it was browser custody, verify the exact browser/session. If it was runtime skew, verify the selected runtime—not just the repository.
☐ Restore one path first.Bring back the minimum useful capability and observe one real or harmless canary before restoring concurrency, retries, or broad automation.
☐ Record the causal chain.Write: trigger → decision → action mechanism → observed effect → detection → recovery → durable prevention. Avoid replacing the chain with a generic “agent hallucinated” label.
Need the reusable production kit?
Agentic AI Incident Response Kit 2026
The paid kit expands this checklist into reusable incident playbooks, observability schemas, postmortem templates, failure classification, and commercial internal-use materials for a team operating tool-using agents.
Commercial deployment rights
Enterprise commercial license
For a legal entity that needs perpetual commercial-use rights to Cognilode-authored autonomous-operations infrastructure and designated operational documentation.
Buy enterprise license — $10,000