Ask an organization how it manages AI risks and the first answer is almost always the same: a human reviews the outcomes. Human-in-the-loop is the most cited control in the AI risk assessments we review. It is also the control that most often fails in practice. An employee who has to approve hundreds of AI suggestions per day approves everything after a week. A reviewer without a mandate, or without insight into how the model reaches an outcome, is not oversight but a formality. Both ISO 42001 and the EU AI Act set requirements for human oversight, and both expect more than a human sitting somewhere in the middle. In this article we explain what the requirements entail, why oversight fails in practice, and how to design it so that it works and can be demonstrated to an auditor.
## What ISO 42001 and the EU AI Act require
ISO 42001 approaches human oversight from the management system. The standard asks you to determine, per AI system, what degree of human involvement is appropriate, to substantiate that choice from the risk assessment and impact assessment, and to assign the corresponding roles, responsibilities and competences. Oversight is thus not a loose control but a design decision you must be able to justify: why does this system require review before outcomes take effect, and why is monitoring after the fact sufficient for that other system?
The EU AI Act is more concrete for high-risk AI systems. Article 14 of the regulation requires that these systems are designed so that natural persons can effectively oversee them. The overseeing person must be able to understand the operation and limitations of the system, remain aware of the tendency to automatically rely on its output (automation bias), correctly interpret the output, be able to decide not to use the system or to disregard its outcome, and be able to intervene or stop the system. For deployers this means oversight must be assigned to people with the right competence, training and authority. That is a substantially higher bar than "someone looks at it".
## Three forms of human involvement
In practice it helps to distinguish three forms. With human-in-the-loop, a human reviews each individual outcome before it takes effect: the AI proposes, the human decides. Think of a credit officer who confirms or rejects every automated recommendation. With human-on-the-loop, the system operates autonomously, but a human monitors its behavior and can intervene: think of fraud detection that blocks transactions, with an analyst watching patterns and outliers. With human-in-command, the human steers at the system level: when is the system used, within which boundaries, and when is it switched off?
The choice between these forms is a risk decision. Case-by-case review is the heaviest regime and fits decisions with major impact on individuals. Monitoring fits high volumes with lower impact per case. What matters is that you make the choice explicit and document it. What we regularly see in audits: the risk assessment claims human-in-the-loop, while practice has long shifted to on-the-loop because the volume makes case-by-case review impossible. Then your control on paper does not match reality, and that is exactly what an auditor will pick apart.
## Why oversight fails in practice
The biggest risk to human oversight is not unwillingness but habituation. Automation bias is well documented: the more often a system is right, the less critically people examine its outcomes. If a model makes a good proposal in 98 out of 100 cases, it is human to approve all 100. The reviewer becomes a rubber stamp, and the oversight exists only administratively.
In addition, we see four recurring design flaws. Volume: the number of outcomes to review cannot be handled within the available time, so reviewing becomes ticking boxes. Information: the reviewer only sees the outcome, without context, confidence indication or reasoning, and therefore cannot make a substantive judgment. Mandate: the reviewer is formally allowed to deviate, but every deviation must be explained to a manager, while going along with the system raises no questions. The incentive then works the wrong way. Competence: oversight is assigned to someone who knows the domain or the system too little to recognize a wrong outcome.
## Designing oversight that works
Oversight that works is designed around the question: what does this person need to be able to substantively contradict an AI outcome? At least five ingredients belong to that.
First, authority: the overseer must be able to deviate, escalate or stop the system without friction, and that authority must be formally assigned. Second, information: show not only the outcome but also the relevant input, an indication of uncertainty and, where possible, the main factors behind the outcome. Third, time and volume: size the number of reviews per person on actual review time, and consider risk-based selection where only outcomes above a risk threshold are reviewed case by case. Fourth, training: teach reviewers what the system can and cannot do, what its known weaknesses are and how automation bias works. This connects directly to the AI literacy obligation in Article 4 of the EU AI Act. Fifth, measurability: measure how often reviewers deviate from the system. A deviation rate of zero is rarely a sign that the model is perfect; it is usually a sign that the oversight is not functioning.
## What the auditor tests
In an ISO 42001 audit, the question is not whether you promised human oversight, but whether it exists and works. Concretely, we look at the substantiation of the chosen form of oversight in the risk assessment, the formal assignment of roles and authority, evidence that reviewers are trained, and records showing that oversight actually takes place: logging of reviews, documented deviations and their follow-up. Interventions must be traceable: who changed or blocked which outcome, when, and what happened with it? An organization that can show that deviations are recorded, analyzed and fed back to the model owner demonstrates not only oversight but also the improvement cycle that ISO 42001 expects as a management system standard.
Human oversight is not a band-aid you can put over every AI risk. Well designed, it is one of the strongest controls there is; poorly designed, it is false assurance that falls apart in an audit. The difference lies in design, mandate and measurability.
Secure Audit helps organizations set up AI governance under ISO 42001, from risk assessment and control design to certification preparation. The certification audit itself is performed by an accredited certification body. Get in touch for a no-obligation conversation.
Frequently asked questions
Is human-in-the-loop mandatory under the EU AI Act?+
Not in all cases. Article 14 requires effective human oversight for high-risk AI systems, but does not prescribe that every individual outcome is reviewed by a human beforehand. The form of oversight must fit the risks of the system and the context of use.
What is the difference between human-in-the-loop and human-on-the-loop?+
With human-in-the-loop, a human reviews each outcome before it takes effect. With human-on-the-loop, the system operates autonomously and a human monitors its behavior, with the ability to intervene. The first fits decisions with major impact per case, the second fits high volumes with lower impact.
How do you demonstrate effective human oversight in an audit?+
With the substantiation of the chosen form of oversight in the risk assessment, formally assigned roles and authority, training evidence for reviewers, and records of reviews and deviations. An auditor also looks at deviation rates: zero deviations usually indicates oversight that is not functioning.
What is automation bias?+
The human tendency to trust the outcomes of an automated system and to review them less and less critically as the system is right more often. It is the main reason human oversight erodes in practice, and the EU AI Act explicitly requires overseers to remain aware of it.
Does human oversight also apply to purchased AI?+
Yes. As a deployer you are responsible for appropriate oversight of the use of the system, even if a vendor built it. You do need information from the vendor for that, such as documentation on the operation and limitations of the system.
Need help with compliance?
Need to comply with ISO 27001, ISO 42001, NEN 7510, NIS2 or DORA, or do you need a SOC 2 report? We guide you through the entire process: from gap analysis to implementation.
Explore Compliance ServicesAbout the author
Partner | IT Auditor