Agent Production-Readiness Checklist
Test whether an agent is ready for a stated operating boundary across evaluations, permissions, observation, human oversight, and recovery.
Start with the actual decision
An agent is not only a model. Its production behavior depends on tools, data access, orchestration, permissions, evaluations, human checkpoints, logging, and recovery paths.
Readiness is always readiness for a particular scope. A system may be suitable for supervised drafting and unsuitable for autonomous execution in the same workflow.
Intended user
Product owners, platform leaders, applied-AI teams, enterprise architects, and control functions reviewing a demonstration or pilot before its operating boundary expands.
Collect before comparing
- 01Workflow map, business owner, user groups, and proposed autonomy level
- 02Tool inventory, data access, permission boundaries, and side effects
- 03Evaluation set, acceptance thresholds, failure examples, and test dates
- 04Human approval, escalation, override, and fallback paths
- 05Logs, traces, alerts, incident procedures, and change controls
- 06Expected volume, consequence of error, and rollback capability
Test every material criterion against the workload
Task and boundary
Is the agent’s allowed work, prohibited work, and escalation threshold explicit?
An undefined operating boundary cannot be evaluated or governed.
Evaluation evidence
Do tests represent real tasks, edge cases, tool calls, refusals, and unacceptable outcomes?
A persuasive demo is not evidence across the intended workload distribution.
Least privilege
Can each user and tool perform only the actions required for the approved workflow?
Capability without authorization boundaries expands the impact of errors and misuse.
Human control
Are approvals, review queues, overrides, and escalations placed at the consequences that need them?
Human involvement should be designed around risk, not added as a decorative step.
Observation and recovery
Can an operator reconstruct what happened, detect material failure, stop execution, and recover safely?
Non-deterministic behavior requires traceable inputs, actions, outputs, and ownership.
Change governance
What happens when prompts, models, tools, permissions, or upstream data change?
Readiness expires when a material component changes without reevaluation.
Do not expand autonomy when a high-consequence failure lacks either a tested prevention control or a credible detection, containment, and recovery path. Constrain the agent to the largest boundary supported by current evidence, then validate the next boundary separately.
This checklist is not a security assessment, red-team report, compliance attestation, or assurance of safe behavior. It does not replace domain-specific risk review. Use stricter review when an agent can move money, change records, affect rights, access sensitive data, or take irreversible action.
Editorial version 1.0
Initial public checklist reviewed August 2026. Re-review after material changes to model, prompt, tools, permissions, data, workflow, or operating scope.
StandardsReview the readiness engagement
See the bounded inputs, deliverables, and exclusions for an Agent Production-Readiness Review.
Continue