The certificate is in and the celebration is deserved. But for an AI management system, the certification date is not a finish line. AI models are not static software: their performance shifts as the world around them changes, and vendors modify models without asking. ISO 42001 therefore expects you to keep monitoring, demonstrably, after certification. In this article: why models degrade, what the standard concretely requires and how to set up monitoring that survives a surveillance audit.
Why AI models degrade after certification
With traditional software, what worked yesterday works tomorrow, unless someone changes something. That assumption does not hold for AI models. Four mechanisms undermine performance after go-live, and none of them requires anyone to touch the system.
The first is data drift: the input the model receives shifts away from the data it was trained or validated on. A credit model built on customer behavior from before an economic downturn will face structurally different applications afterwards. The model keeps running, but the assumptions underneath it no longer hold.
The second is concept drift: the relationship between input and outcome itself changes. Fraud patterns are the classic example. Fraudsters adapt their behavior as soon as detection becomes effective, so the very signals the model was trained on lose their predictive value.
The third is context change within your own organization. A new product line, a different customer segment, a changed process around the model: all of these alter the population the model is applied to, without anyone recognizing it as a model change.
The fourth mechanism is the vendor. Anyone using a hosted model or an AI feature inside SaaS software deals with updates that happen out of sight. A new model version can respond materially differently to the same prompts. In practice we see organizations notice this only when users start complaining, which is exactly the moment you want to be able to show you could have known earlier.
What ISO 42001 requires
ISO 42001 has no chapter titled monitoring after certification, but the requirement follows directly from the structure of the standard. Clause 9 obliges the organization to determine what needs to be monitored and measured, with which methods, when and by whom, and to analyze and evaluate the results. Clause 10 requires handling nonconformities and continually improving the management system. The controls in the annex add requirements for operating AI systems throughout their lifecycle, including tracking performance, recording events and reassessing impact when significant changes occur.
Then there is the certification cycle itself. An ISO 42001 certificate is valid for three years, with annual surveillance audits. In those audits the certification body does not retest the entire system; it tests whether the management system is alive: is performance being tracked, are deviations being addressed, does monitoring information feed into the management review. An organization that stopped measuring after the certificate will fail that test immediately.
What to monitor concretely
The standard does not prescribe a fixed set of indicators. The task is to determine, per AI system, what is proportionate, in line with the risk assessment. In practice we work with five categories.
First, performance indicators. For some systems these are model metrics such as precision and recall, but process indicators are often more meaningful: the percentage of outcomes overruled by staff, the number of complaints about automated decisions, or the throughput time of cases the model handles. Establish a baseline at go-live, otherwise a later shift cannot be interpreted.
Second, drift indicators: statistical monitoring of the distribution of inputs and outputs. This does not require an advanced MLOps platform. For many applications, a periodic comparison of key characteristics of incoming data against the reference period is sufficient.
Third, bias and fairness, for systems that make or prepare decisions about people. A one-off bias test at go-live ages just as fast as the model itself. Repeat the analysis periodically and after every material change in model or population.
Fourth, usage and scope creep. Is the system still being used for what it was intended and assessed for? New user groups, new applications of the same output and sharply increased volumes are signals that the original risk assessment no longer fits.
Fifth, vendor and model changes. Actively follow release notes from AI vendors, record which model version is in use and arrange contractually that material changes are reported. Without that agreement, this part of monitoring is practically impossible.
How to set it up
Start at go-live of each system with a monitoring plan of one or two pages: which indicators, which thresholds, who reviews them, how often, and what happens when a threshold is breached. That last element is the crux. A dashboard without a follow-up procedure is decoration. A threshold breach must trigger a defined action: further investigation, risk reassessment, retraining, or in the extreme case taking the system offline.
Then connect monitoring to three existing parts of the management system. To the incident process, so a serious deviation is handled as an AI incident. To the risk assessment, so significant changes in model or context lead to reassessment. And to the management review, so leadership periodically sees how the AI systems perform and can make decisions about them. That closes the loop of clauses 9 and 10.
What the auditor wants to see in the surveillance audit
Concretely, an auditor will ask for the monitoring plan per system, the actual measurement results over the past period, evidence of follow-up on threshold breaches, updated risk assessments after model changes, and the treatment of all of this in the management review. The strongest evidence is a case where monitoring flagged a problem that was then handled properly. That demonstrates a working system better than a stack of green dashboards.
Common mistakes
We repeatedly encounter five patterns. Monitoring limited to technical availability, as if a model that runs is also a model that performs. Thresholds without agreed follow-up. No visibility whatsoever into vendor model changes. A bias analysis done once at go-live and never again. And monitoring results that stay in a team channel and never reach management. Each of these is a finding waiting to happen in a surveillance audit.
Conclusion
Certification tests a snapshot; monitoring determines whether the AI management system stays sound afterwards. Organizations that maintain a proportionate monitoring plan per system, tie thresholds to follow-up and feed the results back into risk assessment and management review keep not only their certificate, but actual control over their AI.
Secure Audit supports organizations in building and assessing AI governance under ISO 42001, from initial inventory to certification and the years after. Contact us for a no-obligation conversation.
Frequently asked questions
Does ISO 42001 require continuous, automated monitoring of AI models?+
No. The standard requires you to determine what to monitor, how and how often, and to demonstrate that it happens and is followed up. For a high-risk system that may amount to near-continuous monitoring; for a low-risk application a periodic manual review can suffice. The justification must align with the risk assessment.
What is the difference between data drift and concept drift?+
With data drift, the input the model receives changes compared to the data it was built on. With concept drift, the relationship between input and outcome itself changes, for example because fraudsters adapt their behavior. Both cause performance loss but require different detection and different measures.
What should I do when my AI vendor changes the model?+
Arrange contractually that material changes are reported, record which model version you use, and treat a reported change as a trigger for reassessment: check performance against your own indicators and update the risk assessment where needed. Without a notification clause in the contract this is practically unworkable.
How much weight does monitoring carry in the annual surveillance audit?+
A lot. Surveillance audits specifically test whether the management system is alive, and monitoring, nonconformity handling and management review are the main gauges for that. An organization that stopped measuring after certification should expect a finding.
Do I need to retest for bias after go-live?+
Yes, if the system makes or prepares decisions about people. A bias analysis ages through drift and population changes. Repeat the analysis periodically and after every material change in model, data or usage context.
Need help with compliance?
Need to comply with ISO 27001, ISO 42001, NEN 7510, NIS2 or DORA, or do you need a SOC 2 report? We guide you through the entire process: from gap analysis to implementation.
Explore Compliance ServicesAbout the author
Partner | IT Auditor