From AI Idea to Operating Evidence: An Eight-Stage Decision Loop
Move one AI opportunity from observed workflow friction to an owned, bounded operating decision. The eight-stage IMPAKT loop connects discovery, workload definition, complete alternatives, hard gates, matched tests, release evidence, operation, and renewal so prototype momentum cannot substitute for decision-changing proof.
A field note by Edgar Domínguez Llanos for IMPAKT.
An artificial intelligence (AI) initiative should advance only when the previous stage produced evidence that changes a named decision. IMPAKT's eight-stage loop begins with the lived workflow before any model demo: discover, frame, bound, compare, test, gate, operate, and learn and reopen. Each stage has an owner, output, boundary, and stop condition. The loop prevents a promising prototype from acquiring budget or authority through momentum alone, while giving useful experiments a clear path to bounded operation and renewal.
Add an evidence loop to the innovation funnel
A conventional innovation funnel narrows many ideas into a few investments. Enterprise AI needs an additional property: evidence must remain connected to the exact workload, configuration, authority, and consequence it supports.
A demo can prove that a model produced a persuasive result once. Representative task acceptance, an eligible data path, bounded tools, service behavior, recovery, and economics remain separate tests. The next stage should therefore ask what decision the demo changed and what remains unknown.
The NIST AI Risk Management Framework organizes voluntary risk-management work around Govern, Map, Measure, and Manage. IMPAKT's loop is a separate operating synthesis and carries no certification. It uses the same discipline of connecting context, measurement, governance, and ongoing management instead of treating evaluation as a final checkbox.
Stage one: discover the lived workflow
Observe the current work before proposing automation. Identify users, affected people, handoffs, exceptions, delays, failure consequences, incentives, existing evidence, and the manual or deterministic baseline.
The output is a current-state record showing where the friction or obligation occurs and who bears the outcome. If the process is poorly understood, an AI layer can hide the uncertainty rather than remove it.
Stop when no accountable owner or meaningful baseline can be named. Continue when the problem exists independently of a preferred model.
Stage two: frame the decision
State the decision in one sentence. Name the owner, alternatives, deadline, consequence of error, non-goals, and the evidence that could change the choice. Include a non-AI or manual alternative; otherwise the comparison is already biased.
The output is a decision brief. It separates exploratory curiosity from a capital, architecture, release, or operating decision. A vague goal such as “adopt agents” fails this stage because it has no accepted outcome or counterfactual.
Stage three: bound the workload envelope
Assign a stable workload ID. Record the start and accepted end state, actors, data path, case distribution, quality threshold, severe failures, latency, demand, authority, tools, service continuity, recovery, operating capacity, and economic unit.
Mark each material input as observed, sourced, assumed, or unknown. Give every unknown an owner, test, and date. Discovery can proceed with uncertainty, but an unknown that can reverse eligibility or severe consequence blocks release.
Split the envelope when two cohorts have materially different data, authority, consequence, or service requirements. Give each material decision its own workload identity.
Stage four: compare complete eligible options
Generate complete configurations rather than model names. Consider managed APIs, managed open-weight endpoints, private or edge execution, hybrid stage decomposition, deterministic software, process redesign, and continued manual work where each is eligible.
For every candidate, name model or artifact, execution venue, provider or operator, prompt, retrieval, tools, policy, evaluator, region, fallback, and owner. Exclude incomplete candidates until their missing fields have owners and tests.
Divergence remains inside the workload envelope and its hard data and authority boundaries.
Apply hard gates before weighted scoring
Apply capability, data, jurisdiction, authority, service, recovery, supply, and ownership gates before weighted scoring. A failed gate eliminates the candidate for the stated scope. Record and surface that reason separately from any average.
The result is a set of eligible options and a list of owned unknowns. If none survives, narrow the workload, add capacity, find an eligible partner, retain manual service, or stop while preserving the evidence bar.
The live IMPAKT workload-placement matrix expands this stage for API, self-hosted, and hybrid decisions.
Stage five: run the smallest decision-changing test
Test eligible complete configurations on matched cases, acceptance thresholds, service conditions, authority, and cost units. Preserve raw prompts, outputs, traces, timings, evaluator decisions, failures, and recovery evidence.
The smallest test is defined by its ability to change the named decision. A latency comparison without output acceptance leaves architecture unresolved. A task evaluation without the tool path leaves action authority unresolved.
Use the same baseline and disclose asymmetries. Expire the result when a material system component or workload distribution changes.
Stage six: gate the operating boundary
The gate chooses among advance, remain limited, redesign, or stop. It records the permitted users, data, actions, tools, capacity, service level, human checkpoints, monitoring, incident response, recovery, owner, evidence date, expiry, and promotion conditions.
The NIST Generative AI Profile offers cross-sectoral suggested actions for generative-AI risks. Use it for failure discovery and control design; specific release approval remains an enterprise decision.
Passing a gate records a bounded decision with risk, population, tool, and workflow limits intact.
Stage seven: operate inside the boundary
Observe accepted completion, severe failures, interventions, latency, cost, capacity, overrides, user and business outcomes, displaced work, incidents, recourse, and recovery. Compare with the archived baseline and counterfactual.
Stage eight: learn and reopen
End each review with renew, reroute, resize, or retire. Re-enter the earliest affected stage after a change to owner, outcome, data, action, model, prompt, tool, supplier, law, evaluator, cost, or recovery. Trace the changed dependency and renew only what it invalidated.
Use an eight-output evidence packet
Keep one artifact from each stage:
- current-state workflow record;
- named decision brief;
- workload envelope;
- hard-gate shortlist of complete options;
- matched test and raw evidence;
- bounded release card;
- operating evidence record;
- learning and renewal decision.
Use shared IDs across the packet. A reviewer should be able to trace one accepted or failed workflow from the initiating request through model interactions, evidence, tool actions, evaluations, human decisions, and final disposition.
Trace one workload through all eight outputs
Synthetic worked workload — supplier-clause extraction. Workload PROC-CLAUSE-01 extracts termination dates, renewal terms, and supporting spans from approved supplier contracts into a fixed schema. Discover archives the current manual review, affected procurement users, exception path, and review time in minutes. Frame asks whether a bounded assistant should prepare a clause record for human approval, compared with the current process. Bound fixes document classes, source rights, schema, prohibited unsupported citations, latency unit, demand range, and manual recovery.
Compare produces a hard-gate shortlist of complete API, private, deterministic, and manual configurations with unknowns owned. Test runs the same representative contracts and acceptance rubric, preserving raw outputs, cited spans, invalid schemas, review minutes, and cost per accepted extraction. Gate can permit preparation for named users while retaining authenticated approval and a manual path. The live production-readiness gate deepens that release record; the evaluation-oracle article deepens regression evidence.
Operate records accepted extractions, severe citation failures, overrides, incidents, and recovery against the baseline. Learn and reopen renews, reroutes, resizes, or retires after a material change. This loop defines the evidence method; the enterprise roadmap sets when capital reviews recur. All values remain to measure, and the qualified authority ends at preparation for authenticated review.
Use the workload identifier across the decision brief, configuration manifest, case set, release card, traces, incident record, and renewal decision. Give each artifact an owner, version, evidence date, and expiry. The measurable system effect is accepted extraction with its citation and recovery burden visible; the enterprise variables are review capacity, consequence, service, and cost. The rule is to fund only the next evidence-producing stage. The boundary follows the exact document cohort, user group, schema, and preparation authority that passed.
Decision rule
Fund the next stage only when the previous stage produced inspectable evidence and an explicit boundary. After change, return to the earliest affected stage and end the review in renew, reroute, resize, or retire.
What this does not prove
The eight-stage loop is a new IMPAKT synthesis from distributed methods and articles. It does not prove that IMPAKT operated every artifact in production, that following the loop guarantees value or safety, or that voluntary NIST guidance grants approval. Legal, security, privacy, workforce, domain, and affected-party decisions remain with accountable specialists.
Editorial process
This article was extracted from the IMPAKT LLM Operating Playbook with AI-assisted structure, drafting, editing, and metadata preparation. It underwent an independent critique and substantive revision loop against IMPAKT's publication rubric; primary sources are linked beside supported claims, and synthesis, recommendations, and evidence boundaries remain explicit.
Sources
- NIST, Artificial Intelligence Risk Management Framework, version 1.0, January 26, 2023; accessed August 29, 2026.
- NIST, Artificial Intelligence Risk Management Framework: Generative AI Profile, NIST AI 600-1, July 2024; page updated April 8, 2026; accessed August 29, 2026.