Do not use AI as the final decision-maker when a plausible error could cause lasting harm, when the person affected cannot challenge the result, when the input is not permitted to leave its current boundary, or when nobody can test, monitor, and reverse the system. In those cases, keep the task human or restrict AI to a clearly defined support role.
This is not a claim that people always outperform models. A rules engine can route routine records more consistently than a tired operator, and a model can help a qualified reviewer find patterns in a large file. The real question is whether this particular use has an accountable owner, evidence that matches its consequences, and a workable exit when the system fails.
The three lanes are a screening device, not a substitute for professional judgment or the rules that govern a particular setting. A hospital, employer, lender, public agency, and family photo archive do not share one risk threshold. The rest of this guide shows how to identify the consequence, data boundary, reviewer, evidence, and exit for the actual job in front of you.
The short answer: stop, assist, or automate
Put the proposed use into one of three lanes before choosing a product. Stop means the AI is not used for that task. Choose it when the data cannot be shared, the output cannot be checked by someone competent, the action would be unsafe to reverse, or the organization lacks permission to process the input in that way.
Assist means the system may organize material, draft options, flag records, or summarize evidence, but a named person examines the original sources and owns the decision. This lane is appropriate only when the reviewer has enough knowledge, time, authority, and information to disagree. A human who clicks approve on every recommendation is not a control.
Automate is the narrowest lane. It fits a stable, low-consequence task with defined inputs, a measurable correct result, logging, monitoring, and a tested rollback. NIST's discussion of human and AI configurations makes the same contextual distinction: some uses may not need human oversight, while others specifically require it. The label AI does not settle which lane applies. NIST is revising AI RMF 1.0, so check the live framework before approving a deployment record.
Start with consequence and reversibility
Capability demos begin with what a model can produce. A deployment review should begin with what happens when the output is wrong. NIST's AI RMF Core calls for teams to identify the likelihood and magnitude of impacts and to define human oversight, roles, and responsibilities. That turns an abstract accuracy discussion into a concrete consequence map.
Write down who can be affected, the worst credible error, how quickly it could spread, and whether the original state can be restored. A bad internal label that is corrected before anyone acts is different from an incorrect rejection, public allegation, payment, treatment suggestion, or safety instruction. Frequency matters, but a rare failure can still be unacceptable when recovery is impossible.
If the task touches a person's health, legal rights, access to work or finance, or physical safety, route the design to qualified people who understand the applicable professional and regulatory duties. AI may support an approved process, but a general chatbot output is not a substitute for that process. The more serious the consequence, the stronger the evidence, review, record, and appeal path must be.
Employment is one example of why the use case matters more than the product label. The US Equal Employment Opportunity Commission says in its AI and algorithmic fairness initiative that anti-discrimination law applies to technology used in employment decisions. That does not turn this article into a US compliance guide. It shows why a team cannot approve a general “AI for HR” project without defining the exact decision, affected people, evidence, review, and governing rules.
Ten stop signs that take automation off the table
- No person owns the outcome. If the answer to “who is accountable?” is the vendor, the model, or the team in general, stop. Name one operational owner and one escalation route before a pilot. Responsibility cannot be delegated to software.
- The reviewer cannot verify the work. A polished answer is not evidence. NIST's Generative AI Profile describes confabulation, including false content and citations presented with confidence. If nobody can inspect the underlying facts, calculation, code, or domain judgment, human review is only decoration.
- A mistake cannot be repaired in time. Do not automate an action whose damage would arrive before detection, or whose target state cannot be restored. Add approval, delay, a transaction limit, a reversible draft state, or remove AI from the action entirely.
- The affected person cannot understand or contest the result. A decision process needs a route for correction when it affects an individual. If the organization cannot explain what happened, receive additional context, and provide meaningful reconsideration, the process is not ready for autonomous use.
- The required facts are missing. A model can fill gaps with plausible language, but plausibility is not the same as a complete record. Stop when the task depends on current documents, private events, physical inspection, or context that the system does not possess. Retrieve the evidence first.
- The input is confidential or unapproved. Do not paste customer records, employee data, source code, contracts, credentials, health information, or unreleased plans into a service until its permitted use, retention, access, training, deletion, and location rules have been checked for that account and workflow.
- Untrusted content can trigger a real action. A document, webpage, email, or support ticket may contain instructions intended to redirect an AI agent. The UK's National Cyber Security Centre identifies prompt injection and data poisoning as routes to unintended behavior in its secure AI guidance. Keep untrusted text away from permissions to send, publish, pay, delete, or execute until the threat model and controls have been tested.
- There is no representative baseline. A few friendly examples do not establish reliability. If the team cannot assemble ordinary cases, difficult cases, known failures, and inputs from the environment where the system will operate, it cannot tell whether the pilot is better than the current process.
- The metric rewards the wrong behavior. Faster replies can hide more corrections. More flagged cases can hide a higher false-positive rate. A single average can hide damage concentrated in a smaller group. If the metric is disconnected from the real outcome, do not let it authorize automation.
- A simpler method is clearer and sufficient. A form with required fields, a deterministic calculation, a search filter, a template, or an ordinary database rule may be easier to test and explain. NIST's Manage guidance explicitly asks whether an AI system is an appropriate solution. Complexity is a cost, not evidence of quality.
Human review must be real, not ceremonial
“A human is in the loop” says little by itself. The reviewer needs the original material, relevant expertise, a visible statement of model limits, enough time to inspect the recommendation, and permission to override it without being punished for slowing the queue. The process should record what the reviewer changed and why.
The UK's Information Commissioner's Office gives a jurisdiction-specific example in its guidance on AI and individual rights. It says meaningful human review is active rather than a rubber stamp, and that reviewers need the authority and competence to go against a recommendation. Organizations elsewhere must check their own law, but the operational test is broadly useful.
Ask the reviewer to explain the decision without repeating the model's answer. Can they point to the underlying evidence? Can they identify information the model did not consider? Can they describe a condition that would change the result? If not, the organization has inserted a person into the workflow without creating independent judgment.
Review capacity belongs in the design, not in a footnote. If one person is expected to clear hundreds of recommendations an hour, the interface and target will push them toward agreement. Sample decisions for a second review, track how often people override the system, examine disagreements, and test whether reviewers notice planted errors. When a reviewer repeatedly cannot reconstruct the basis for an output, reduce the system's role instead of treating extra training as the only remedy.
Private data and untrusted input change the answer
A safe use with synthetic text can become an unsafe use with real customer data. Before a tool receives sensitive material, identify the account type, provider, subprocessors, storage location, retention period, deletion route, training setting, access controls, and logging. Check the contract and live product controls, not a marketing summary copied months ago.
The US Federal Trade Commission's article on privacy and confidentiality commitments by AI companies warns against changing data practices without clear notice and appropriate consent. That page is US enforcement context, not a universal compliance checklist, but it makes a practical point: a new AI purpose does not erase the promises already made about the data.
Security boundaries matter even when the data is permitted. The NCSC's secure deployment guidance calls for access controls, audit logs, incident procedures, evaluation, and disclosure of known limitations. An agent with broad tool permissions needs stricter controls than a chat window that produces a draft. Limit credentials, separate environments, require approval for consequential actions, and retain a way to investigate what happened.
Replace intuition with a written go or no-go record
A short decision record keeps the debate tied to one use case. It is not a broad corporate policy. HUMAI's separate AI policy guide covers organization-wide rules; this record captures the facts needed to approve or reject a particular workflow.
Write these items before the pilot:
- the exact task, its owner, and what remains outside scope;
- whether AI will stop at a draft, recommend an action, or perform one;
- the people, systems, money, rights, and relationships that could be affected;
- approved input classes and the data that must never enter the system;
- known failure modes, including missing facts, confabulation, bias, and malicious input;
- the current non-AI baseline and the acceptance threshold;
- the human reviewer, override authority, appeal route, logs, and incident owner;
- rollback conditions, a review date, and the person allowed to pause the system.
A blank field is a decision signal. If the team cannot name an affected group, a baseline, or a rollback owner, it should not fill the space with “to be determined” and deploy anyway. Move the use back to assist or stop until the missing control exists.
Test the smallest version that can fail
Run the pilot in a sandbox with no live side effect. Freeze the model or product version, system instructions, tools, data permissions, and evaluation set. Compare it with the current process and, where useful, with a simpler non-AI method. Set acceptance thresholds and stop conditions before looking at results.
The baseline should use the same inputs and the same definition of success. Count the labor required to prepare data, review outputs, correct errors, and handle exceptions, not only the seconds taken to generate an answer. A workflow that appears faster before review can be slower after rework. Keep every test artifact so another reviewer can reproduce the comparison rather than trusting a slide with aggregate scores.
Include normal inputs, edge cases, incomplete records, deliberately misleading instructions, conflicting evidence, and examples that should be refused or escalated. For generated text, inspect every factual claim and citation. For classification or recommendations, report false positives and false negatives separately, then examine whether error rates or consequences differ across relevant groups. A single accuracy number does not describe those tradeoffs.
Record corrections, overrides, incidents, and complaints after launch. Model behavior, product settings, upstream data, and the surrounding workflow can change. The NIST AI Resource Center describes testing, evaluation, verification, and validation as part of operationalizing the framework, not as a one-time demo. If the team cannot observe failure or suspend the workflow, automation is premature.
Define the pause rule in plain language. It might be a serious incident, a rise in a particular error, a missing log, an unreviewed product update, or a volume that exceeds reviewer capacity. Pausing is not proof the experiment failed. It is evidence that the control worked before the workflow created a larger problem.
Keep the decision attached to an owner
The useful question is not “Can an AI produce an answer?” It is “Can this organization justify this use, detect a bad result, protect the people and data involved, and recover when the system fails?” If any part of that sentence has no owner or evidence, the answer is not yet. Keep the work human, narrow it to support, or choose the simpler tool.