Bias and data governance under ISO 42001: what auditors expect

Compliance9 min read·
K

Kees van der Vlies

Partner | IT Auditor

Also available in:Nederlands

Ask an organisation about its biggest AI risk and the answer is often a data breach or a hallucinating chatbot. But the risks that cause the most damage in practice sit one layer deeper: in the data that AI systems are trained on and fed with, and in the bias hidden inside it. A model that systematically disadvantages certain groups does not stand out in a demo. It only stands out once the harm is done, affecting customers, applicants or citizens. That is why both ISO 42001 and the EU AI Act set explicit requirements for data governance and the handling of bias. In this article: what those requirements entail, where bias comes from and how to demonstrably control it.

## What ISO 42001 requires of data for AI

Annex A of ISO 42001 contains a dedicated cluster of controls for data used by AI systems. The common thread: an organisation must understand and control the full data journey of its AI systems. That means recording where data for development and training comes from (provenance), on what grounds data was acquired, which quality requirements apply and how data is prepared: cleaning, labelling, enriching and selecting. For each of those steps the standard expects documented criteria and demonstrable execution.

Important to realise: within the scope of your AIMS these requirements also apply if you do not train models yourself. Anyone who procures an AI application and feeds it with their own business data, for example a RAG application on internal documents or a scoring model on customer data, has data governance obligations for exactly that data. The auditor's question then is not how the base model was trained, but whether you know which data enters your application, whether that data is accurate and current, and who is responsible for it.

## The link with Article 10 of the EU AI Act

For high-risk AI systems, the EU AI Act turns data governance into a legal obligation. Article 10 requires training, validation and testing datasets to be subject to data governance practices, including an assessment of availability, suitability and representativeness, and an examination of possible biases that could lead to discrimination or other harm. Datasets must be relevant, sufficiently representative and, to the best extent possible, free of errors and complete for the intended purpose.

Implementing ISO 42001 builds the foundation on which Article 10 compliance can rest: the data governance process already exists, and for high-risk systems you increase the depth. The other way around does not work: a standalone Article 10 exercise without an underlying management system goes stale within a year.

## Where bias comes from

Bias is not a programming error but a data phenomenon, and it creeps in at multiple points. Historical bias: the data reflects past decisions, including their distortions. A recruitment model trained on ten years of hiring decisions learns the preferences of that era, including the undesirable ones. Representation bias: certain groups are underrepresented in the data, causing the model to perform worse for those groups. Measurement and label bias: the measured attribute is a flawed proxy for what you actually want to know, for example arrests as a proxy for crime, or labels assigned by human reviewers with their own assumptions. And finally drift: a model that performed well at go-live develops bias because the world changes and today's data no longer resembles the training data.

That last category is missed most often. Bias testing as a one-off hurdle before go-live is false comfort; without monitoring you only see drift when the incident happens.

## Data governance in practice

A workable approach starts with a data inventory per AI system: which data sources feed this system, who owns those sources, and what is their quality and provenance? Record a short profile per dataset, covering purpose, origin, processing steps, known limitations and the legal basis if it contains personal data. This does not have to become a bureaucratic monster: one concise, current document per dataset suffices, as long as it is honest about the limitations.

Then assign the roles. Every dataset that feeds an AI system has an owner who is accountable for quality and currency. For procured AI, data governance belongs in the vendor assessment: ask the vendor about the provenance of training data, bias testing performed and how the model is updated. A vendor that cannot say anything about this is telling you something too.

## Testing for bias demonstrably

Bias testing does not start with a tool but with a definition. Determine per application what fair means: equal error rates across groups, equal probability of a positive outcome given equal suitability, or something else. That choice is context-dependent and partly a governance decision, not a purely technical one. Record the chosen definition and the acceptance thresholds in advance, otherwise every test result gets rationalised after the fact.

Then test at three moments. Before go-live: performance broken down by relevant subgroups, not just the average. A model with 95 percent accuracy can sit at 70 percent for a specific group; the average hides that. Upon changes: every retraining or significant data change is a trigger to repeat the subgroup analysis. And continuously: monitor production outcomes for shifts between groups, and define thresholds that trigger human review or escalation.

## What the auditor wants to see

In a certification audit or internal audit, it comes down to demonstrable operation. Concretely: the data inventory and dataset profiles, quality and fairness criteria defined in advance, test results with subgroup analyses, and decision-making when deviations occur. That last item is where things most often go wrong. A test result outside the thresholds that nobody acted on is a finding for an auditor. A documented decision to temporarily accept a deviation with compensating measures, such as additional human review, is not. The difference is not the deviation but the control.

## Common mistakes

Steering only on average accuracy and never breaking results down by subgroup. Testing for bias once at go-live and never again. Dismissing data governance as not applicable because the organisation does not train models, while its own business data does feed AI applications. Not defining fairness criteria in advance, leaving tests without consequences. And not asking vendors about the provenance of procured models and datasets, creating a blind spot in the risk picture.

## Conclusion

Data governance and bias control are not an academic footnote to AI governance but its core. The organisations that do this well need no impressive tooling: they know which data feeds their AI systems, have determined in advance what is acceptable, test against it and intervene when results deviate. That is exactly what ISO 42001 means by control, and what Article 10 of the EU AI Act legally anchors for high-risk systems.

Frequently asked questions

Does data governance apply if we only procure AI and do not develop it?+

Yes. The data you feed into a procured AI application, such as internal documents or customer data, is your responsibility. In addition, the vendor's handling of training data and bias belongs in the vendor assessment.

Which fairness definition should I choose?+

There is no universally correct definition; different definitions can even be mathematically incompatible. Choose per application based on context and potential harm, record the choice and acceptance thresholds in advance, and have management or the risk owner ratify the choice.

Does ISO 42001 require specific bias detection tools?+

No. The standard requires control and demonstrability, not specific technology. For many applications a structured subgroup analysis in existing analytics environments suffices; tooling becomes relevant with larger model portfolios.

How often should we test for bias?+

At minimum before go-live, at every retraining or significant data change, and continuously through production monitoring. Testing once at go-live is insufficient, because bias can also emerge later through changing data.

What is the difference between data quality and data governance?+

Data quality concerns the properties of the data itself: accuracy, completeness, currency and representativeness. Data governance is the management process around it: ownership, criteria, provenance, decision-making and control. Quality is an outcome of good governance.

Need help with compliance?

Need to comply with ISO 27001, ISO 42001, NEN 7510, NIS2 or DORA, or do you need a SOC 2 report? We guide you through the entire process: from gap analysis to implementation.

Explore Compliance Services

About the author

K
Kees van der Vlies

Partner | IT Auditor

Back to knowledge base

Have a question?

Get in touch for advice on IT audit, compliance and information security.

Contact us