A 20-Question Screen for LLM Progress Claims
Use a 20-question screen to evaluate claims about scaling, benchmarks, context, retrieval, tools, quantization, serving, work effects, adoption, security, governance, and frontier opacity. Every question carries an evidence class, countercondition, local test, owner, and expiry so a dated result stays inside its decision boundary.
A field note by Edgar Domínguez Llanos for IMPAKT.
Put large language model (LLM) progress claims through 20 bounded questions before they influence architecture, capital, or authority. Training loss, task utility, context use, tool authority, serving economics, adoption, and governance can move at different speeds. This screen is secondary IMPAKT synthesis: every bold question is an inference or recommendation, while each linked source supports only the cited premise. A countercondition limits the inference, and a local test determines whether it transfers to the named workload.
Questions one through five: separate capability signals
1. Does smooth training loss imply smooth workload utility? The scaling-laws paper reports smooth empirical loss relationships in its setting; IMPAKT infers that discrete, multi-condition acceptance may move differently. Countercondition: a continuous task metric may track the source curve. Test: plot the same versioned cases against both loss and accepted completion.
2. Does an observed plateau define a ceiling? Changing data and compute allocation altered results in the Chinchilla study; IMPAKT infers that some plateaus belong to a method rather than the task. Countercondition: data, compute, architecture, or evaluation can still impose a durable limit. Test: change one binding layer and preserve the cases.
3. Does a benchmark gain transfer to the operating task? This is an IMPAKT evidence recommendation: exposure, task-specific optimization, prompts, or tools can change the comparison. Countercondition: a matched method and maintained hidden workload may preserve transfer. Test: replay versioned executable, time-split, and adversarial cases before using the score.
4. Does an average cover calibration and severe tails? IMPAKT infers that an average can conceal prohibited outcomes or unreliable confidence. Countercondition: a metric explicitly designed and validated for calibration or tail risk may cover the named claim. Test: measure confidence, abstention, severe cases, and repeated failure separately.
5. Do demonstrations establish task acquisition? The GPT-3 paper reports few-shot behavior in specified evaluations; IMPAKT treats demonstrations as configuration. Countercondition: a stable task and prompt may generalize within a bounded distribution. Test: vary demonstration order, labels, counterexamples, and held-out cases.
Use this first group to separate a changed research metric from a changed workload decision. Preserve the metric, cases, configuration, and acceptance threshold before drawing the bridge.
Questions six through ten: test behavior, context, and authority
6. Does post-training add the missing knowledge or permission? The InstructGPT paper reports behavioral change under its methods; IMPAKT separates that result from current evidence and authority. Countercondition: training can encode stable knowledge for a bounded task. Test: evaluate knowledge, obedience, factual support, and permission independently.
7. Does accepted context length establish effective evidence use? Lost in the Middle reports position-dependent behavior in tested settings. Countercondition: some trained tasks use long context well. Test: vary evidence position, distractors, synthesis, and output budget using the context-path method.
8. Does retrieval resolve factuality? The retrieval-augmented generation paper establishes one retrieval-generation method; IMPAKT infers that the evidence burden moves into corpus rights, freshness, ranking, poisoning, citation, and entailment. Countercondition: a controlled corpus can reduce specific errors. Test: separate retrieval recall from claim support.
9. Does useful tool access also expand consequence? OWASP's excessive-agency guidance identifies functionality, permissions, and autonomy as risk drivers. Countercondition: read-only or reversible tools can keep consequence narrow. Test: bind identity, resource, operation, authorization, verification, and rollback using the least-dynamic pattern.
10. Does component accuracy compose into accepted completion? tau-bench is one bounded record of multi-turn, tool-mediated evaluation; IMPAKT infers at system level. Countercondition: checkpoints and repair can improve completion. Test: run repeated trajectories with state, permissions, severe events, interventions, and recovery visible through the composition test.
Run questions 6–10 on one end-to-end trace. The same evidence packet should explicitly expose supplied knowledge, context construction, retrieved passages, tool authority, checkpoints, and final recovery so improvements cannot hide a shifted failure.
Questions eleven through fifteen: qualify deployment claims
11. Do artifact rights determine execution venue? This is an IMPAKT operating taxonomy: source availability, weight access, license, provider access, and execution location are separate fields. Countercondition: a license or provider may restrict eligible venues. Test: record each right and route independently in the placement matrix.
12. Does a smaller numerical representation preserve the workload? The GPTQ paper reports a post-training quantization method in specified evaluations. Countercondition: some converted artifacts preserve the required behavior. Test: compare the exact artifact on target hardware, matched cases, severe slices, and accepted cost.
13. Can the model name explain deployed performance? The PagedAttention paper is one bounded record showing that serving design matters. Countercondition: one dominant model or hardware constraint may overwhelm software differences. Test: pin cache, scheduler, kernels, batch, device topology, quality, and demand through the inference-stage record.
14. Does lower unit cost determine total demand or spend? This is an IMPAKT economic inference; lower unit cost can change both use and spend. Countercondition: capped demand may make the direction simple. Test: observe demand bands and calculate the completed-workflow cost without assuming elasticity.
15. Do observed work effects transfer across tasks and populations? The final Generative AI at Work study reports heterogeneous effects in one customer-support setting. Countercondition: a closely matched task and population may reproduce part of the effect. Test: preserve the local baseline, population, outcome, horizon, and attribution method.
Archive one configuration manifest across questions 11–15. It should join artifact rights, execution route, precision, serving stack, demand, accepted-workflow economics, task, and population before procurement compares alternatives.
Questions sixteen through twenty: separate adoption and governance clocks
16. Which adoption curve does a percentage measure? This is an IMPAKT measurement recommendation: access, experimentation, integration, repeat use, and accountable operation use different denominators. Countercondition: a statistic with one explicit stage and denominator can be decision-useful. Test: reconstruct numerator, denominator, cohort, date, and operating boundary.
17. What binds after model access becomes adequate? IMPAKT infers that data rights, evaluations, integration permissions, exception handling, operating skill, or process change may become limiting. Countercondition: capability can remain the binding constraint. Test: hold the qualified model stable and measure each unresolved complement against accepted completion.
18. How do authority, untrusted context, and irreversibility change security exposure? OWASP's prompt-injection guidance treats injection as a system risk. Countercondition: isolated, read-only work can limit consequence. Test: exercise untrusted inputs, downstream authorization, denial, containment, and recovery at the exact authority boundary.
19. Do governance layers move on one schedule? The NIST AI Risk Management Framework is a voluntary operating reference; IMPAKT infers that law, contracts, procurement, insurers, boards, and internal policy can move separately. Countercondition: one binding rule may dominate. Test: map owners, dates, precedence, expiry, and reopen triggers in the evidence loop.
20. What causal claim survives limited disclosure? This is an IMPAKT epistemic rule: an external evaluation can establish an observed result under a configuration while leaving causes unresolved. Countercondition: controlled ablations or fuller disclosure can isolate mechanisms. Test: separate the observed comparison from claims about architecture, data, post-training, contamination, or inference.
These five questions end in accountable dates and owners. Keep adoption denominators, complementary investment, authority, governance precedence, and disclosure limits in the same renewal record so a current signal cannot become permanent permission.
Use the 20-question claim screen
When a new claim reaches an architecture or board discussion, record:
- the exact metric, workload, population, system boundary, date, and source;
- whether the claim is observed, sourced, inferred, recommended, synthetic, or unknown;
- the relevant question and its countercondition;
- the enterprise variable the mechanism could change;
- the local test that would confirm or reject transfer;
- the decision owner, evidence expiry, and reversal trigger.
For example, “a larger context window solves our document workflow” invokes question 7. Run evidence-position and synthesis cases on the exact system. “A new quantization cuts cost” invokes questions 12 through 14. Measure accepted-workflow cost, capacity, quality, and demand before translating artifact size into savings.
Synthetic claim packet — supplier-clause extraction. An architecture review asserts that a larger context window can replace retrieval for PROC-CLAUSE-01. Record the vendor result as a sourced observation and the transfer claim as inference. Question 7 supplies the countercondition: effective use can vary with evidence position and distractors. The local test uses identical approved contracts, output budgets, citation rules, and accepted schemas across direct-context and retrieval configurations. The procurement workflow owner owns the decision; the evaluation owner owns raw evidence; the record expires after a model, prompt, corpus, or context-policy change. Fund the larger-window route only if accepted extractions, severe citation failures, latency in milliseconds, review minutes, and cost per accepted extraction clear the same gate. This packet qualifies one document cohort and leaves other languages, contract types, and authority boundaries unresolved.
Keep contradictory evidence instead of averaging it away
Two findings can both be valid when they use different tasks, populations, dates, or methods. Preserve the boundary rather than selecting the more convenient headline. When evidence conflicts within the same scope, mark the decision unresolved and name the discriminating test.
History is useful for identifying mechanisms and renewal triggers. It is weak authority for scheduling an inevitable outcome. A paper date is not a roadmap date, and a benchmark improvement is not a production-readiness event.
Decision rule
Use each question as a bounded diagnostic. Match the claim's metric, system layer, task, population, and date; then require a local decision-changing test before funding or expanding authority.
What this does not prove
The 20 questions are secondary synthesis rather than empirical laws or a complete evidence review. The cited sources use different models, methods, tasks, and populations. This article does not establish a historical rate of enterprise AI progress, forecast 2026–2031, or claim that every plateau, deployment gap, adoption pattern, or policy change follows the same mechanism.
Editorial process
This article was extracted from the IMPAKT LLM Operating Playbook with AI-assisted structure, drafting, editing, and metadata preparation. It underwent an independent critique and substantive revision loop against IMPAKT's publication rubric; primary sources are linked beside supported claims, and synthesis, recommendations, and evidence boundaries remain explicit.
Sources
- Kaplan et al., Scaling Laws for Neural Language Models, submitted January 23, 2020; accessed August 29, 2026.
- Hoffmann et al., Training Compute-Optimal Large Language Models, submitted March 29, 2022; accessed August 29, 2026.
- Brown et al., Language Models are Few-Shot Learners, submitted May 28, 2020; accessed August 29, 2026.
- Ouyang et al., Training Language Models to Follow Instructions with Human Feedback, submitted March 4, 2022; accessed August 29, 2026.
- Liu et al., Lost in the Middle, submitted July 6, 2023; accessed August 29, 2026.
- Lewis et al., Retrieval-Augmented Generation, submitted May 22, 2020; accessed August 29, 2026.
- Yao et al., tau-bench, submitted June 17, 2024; accessed August 29, 2026.
- Frantar et al., GPTQ, submitted October 31, 2022; accessed August 29, 2026.
- Kwon et al., PagedAttention, submitted September 12, 2023; accessed August 29, 2026.
- Brynjolfsson, Li, and Raymond, Generative AI at Work, Quarterly Journal of Economics, 2025; accessed August 29, 2026.
- OWASP, LLM01: Prompt Injection, 2025 edition; accessed August 29, 2026.
- OWASP, LLM06: Excessive Agency, project page; accessed August 29, 2026.
- NIST, Artificial Intelligence Risk Management Framework, version 1.0, January 26, 2023; accessed August 29, 2026.