AI Agent Incident Response Plan Template
Use this when a tool-using AI agent, LLM workflow, browser agent, or autonomous service causes—or may have caused—an incorrect external effect. It is designed to keep containment, evidence, recovery, and recurrence prevention in one compact operational record.
Enough structure to act without turning the incident into paperwork.
What actually happened outside the model: provider objects, money movement, messages, files, deployments, or account state.
Decision, tool execution, stale context, identity/custody, retry/idempotency, runtime skew, or observability mismatch.
The narrow repair and the independent evidence that the failed layer is actually healthy again.
Copyable incident-response plan
# AI Agent Incident Response Plan ## 1. Incident identity - Incident ID: - Start time / detection time: - Owner: - Affected agent/workflow: - Affected provider/account/system: - Severity: SEV-1 / SEV-2 / SEV-3 / SEV-4 ## 2. Known external effect - Last known good effect: - First known bad or ambiguous effect: - Is the harmful effect still propagating? yes/no/unknown - Money, messages, files, deployments, credentials, or accounts affected: - Provider-native IDs / URLs / transaction IDs: ## 3. Immediate containment - Exact action path frozen: - What remains intentionally online/read-only: - Automatic retries paused for ambiguous side effects: - Credentials or sessions revoked/isolated, if necessary: ## 4. Independent evidence captured - Agent-visible request/context: - Tool/API request and response: - Provider-side state: - Runtime/commit/config identity: - Screenshot/UI state, when relevant: - Other independent source: ## 5. Failure classification - Decision failure? - Execution/tool failure? - Stale/freshness failure? - Wrong target / identity / custody failure? - Duplicate/idempotency failure? - Runtime/source divergence? - Observability/reporting mismatch? - Evidence supporting the leading hypothesis: - Evidence that could disprove it: ## 6. Recovery - Smallest reversible repair: - Canary or first restored action: - Provider/runtime/UI evidence that recovery worked: - Remaining uncertainty: ## 7. Customer / counterparty impact - Who was affected: - What needs correction/refund/reversal/notification: - Exact completion state: ## 8. Recurrence prevention - Trigger → decision → action → effect → detection → recovery chain: - Durable code/config/process change: - What should be deleted or simplified: - What signal would catch recurrence earlier: - Owner and completion evidence: ## 9. Closure - External effects reconciled: - Recovery independently verified: - Follow-up work separated from incident closure: - Closure time:
Agentic AI Incident Response Kit 2026
The paid kit adds reusable playbooks, observability schemas, postmortem templates, failure classification, and commercial internal-use materials for teams running tool-using agents.