Trace the request path
- Propagate a request or correlation ID through gateway, feature retrieval, preprocessing, inference, postprocessing, and downstream actions.
- Record the model version, artifact digest, serving release, configuration revision, and feature-schema version for each rollout.
- Keep enough structured failure context to reproduce representative errors without leaking sensitive inputs.
Measure system and model behavior
- Track request volume, error rate, latency percentiles, queue depth, restarts, CPU, memory, accelerator use, and dependency failures.
- Track prediction distributions, confidence or score distributions, input drift indicators, and quality metrics when labels become available.
- Compare canary or new-version behavior directly with the current production baseline rather than judging metrics in isolation.
- Surface fallback and retry behavior explicitly so successful HTTP responses cannot hide degraded service.
Make observability actionable
- Tie alerts to an owner action such as traffic reduction, rollback, dependency failover, or investigation of a specific failure boundary.
- Preserve the last known-good reference values and release identity needed for rollback decisions.
- After recovery, verify both operational metrics and representative prediction outputs before closing the incident.
- Delete noisy evidence that never changes decisions and automate the signals that repeatedly do.
Continue with the complete triage workflow
This focused checklist is one failure boundary. The complete production ML deployment triage checklist covers classification, observability, reversible hypotheses, rollback, recovery verification, and runbook conversion in one sequence.
Production ML Deployment Triage expands the field workflow into a longer engineering guide. Price, checkout, delivery, refund, and tax terms remain provider-native on Leanpub.