GAICC AI Conference & Awards 2026 "Governing the Future – Building Responsible, Safe and Human-centric AI"

ISO 42001 Annex A.7: Data Quality, Provenance and Preparation Controls in Practice

Last Updated : October 7, 2026
iso-42001-annex-a7-data-quality-provenance-controls

On this page

Annex A.7 of ISO/IEC 42001 is the control area for data in AI systems. Its five controls govern how data is managed, sourced, checked, traced and prepared. The controls run from A.7.2 to A.7.6, because A.7.1 is the objective. Each control asks you to define and document a practice, then keep records proving you followed it. A team that never trains a model still holds A.7 data: retrieval corpora, evaluation sets, prompt libraries and production inputs.

Annex A.7 in ISO/IEC 42001 and the controls numbered 7 in ISO/IEC 27001 cover different ground. In the 2022 edition of ISO/IEC 27001, the 7.x controls deal with physical security, such as perimeters and entry. Data documentation also appears once more in the AI standard, in the A.4.3 control on data resources, which this page covers where it meets A.7.

GAICC restates the requirements here in its own words and quotes nothing beyond control IDs and short titles. Before building records from this page, confirm each control’s wording in a licensed copy of the published standard.

The five A.7 controls at a glance

Annex A.7 of ISO/IEC 42001 holds five controls under one objective. In plain terms, the objective is to understand what role data plays in your AI systems and what effects it has. That understanding must hold across the whole life cycle. The register gives each control’s ID, short title, what it asks, the evidence that usually proves it and the person who usually owns it.

ControlShort titleRequirement in plain wordsTypical evidenceUsual owner
A.7.2Data for development and enhancement of AI systemDefine, document and run data management processes for building and improving AI systemsData management procedure; design records naming the training approach and data quality expectedData owner for each AI system
A.7.3Acquisition of dataDecide and write down where each dataset is sourced and why it was chosenAcquisition record per dataset: categories, quantity, source, rights, prior handlingData owner, with legal counsel for rights
A.7.4Quality of data for AI systemsWrite data quality requirements and show that the data behind development and operation meets themQuality specification per data category; validation results with datesData steward
A.7.5Data provenanceDefine a process that records each dataset’s origin and later history, for as long as the data and the system existProvenance log per dataset version; supplier statementsData engineering lead
A.7.6Data preparationSet criteria for picking preparation methods, then log which methods were appliedPreparation specification per AI task; versioned pipeline code and run logsLead data scientist or ML engineer

Both the evidence column and the owner column reflect GAICC’s editorial view; neither comes from the standard. Annex B, which is also normative, repeats the same five controls as B.7.2 to B.7.6 and adds implementation guidance for each. The full set of areas sits in the control objectives page, which lists all nine Annex A areas.

Why the numbering starts at A.7.2

A.7.1 is the objective statement, so the first data control is A.7.2. Every Annex A area follows that pattern: the objective takes the .1 slot (Annex B prints it as B.7.1) and the controls follow. AI answers still get the A.7 numbers wrong. GAICC captured US Google results in early October 2026. Google’s AI summary for the A.7 query listed A.7.3 to A.7.6 and left out A.7.2. A Google AI Mode answer about the Statement of Applicability labelled A.7.2 as data quality, which is A.7.4. One ranking Annex A guide renumbers the areas and files data under A.6.

Use the order in the register, from management through sourcing, quality and provenance to preparation. The same order mirrors how a dataset moves through a project.

A.7.2 Data management for development and enhancement

Control A.7.2 asks for documented data management processes covering how AI systems are developed and improved. The Annex B guidance lists five topics such processes can address.

  • Privacy and security, because some of the data is sensitive.
  • Threats that data-dependent development opens, for security and for safety.
  • Transparency, since teams may need to show provenance and explain how data shapes an output.
  • Representativeness of training data for the domain where the system will operate.
  • Accuracy and integrity, which errors and tampering both degrade.

Evidence for A.7.2 is a written procedure plus proof that teams follow it, because a procedure on its own shows intent and rarely satisfies a certification auditor. Pair it with the design record that A.6.2.3 already asks for. Among the design choices to document, that guidance lists how the model will be trained and what data quality it needs.

Document data at the design stage

Implementers often ask whether data belongs in the record before any data is collected. In practice it does, because data decisions first appear in the design record, which states the training approach and the quality the data must meet. The acquisition and quality controls then test real data against that plan. The A.7.2 guidance points to ISO/IEC 22989 for life cycle and data management concepts, which helps teams name the stages consistently.

“Enhancement” covers any later change to a system’s data: retraining on fresh records, fine-tuning, adding a new source or refreshing a retrieval index. Each of those events should reopen the A.7 records for the affected dataset.

A.7.3 Acquisition records for every dataset

Acquisition is the subject of control A.7.3, which requires you to determine and document how the data used in AI systems is acquired and selected. The Annex B guidance lists the details an acquisition record can hold. The table turns that list into fields, with notes for US teams.

FieldWhat to recordNote for US teams
Data categoriesWhich kinds of data the system needs; ISO/IEC 19944-1 offers a category structureFlag personal information early
QuantityVolume needed and volume obtainedRecord sampling choices
SourceInternal, purchased, shared, open or syntheticKeep the contract or license reference
Source characteristicsStatic, streamed, gathered or machine generatedStreams need a refresh rule
Data subjectsDemographics and characteristics, including known or likely biasesFeeds the bias checks under A.7.4
Prior handlingEarlier uses and whether privacy and security requirements were metAsk suppliers in writing
Data rightsPersonal information, copyright and license termsTraining data copyright is still being litigated
MetadataLabelling and enrichment detailsName the labelling vendor, if any
ProvenanceLink to the A.7.5 recordOne ID shared across both records

Rights deserve their own line because the legal position is unsettled in the United States. GAICC’s case note on scraped video used as training data shows how a missing acquisition record turns into a governance question. Where personal information is in the data, state privacy laws such as California’s can also apply.

A.7.4 Data quality requirements per data category

Control A.7.4 requires documented data quality requirements and proof that data used to develop and operate the AI system meets them. Clause 3.25 of ISO/IEC 42001 treats data quality as the degree to which data satisfies the requirements an organization sets for a given context. The Annex B guidance also cites the ISO/IEC 25024 definition, which ties quality to stated and implied needs under specified conditions. Both definitions lead to the same task: write the requirements down before you measure anything.

For organizations using supervised or semi-supervised learning, the guidance asks that the quality of four data categories be defined, measured and improved. Those categories are training, validation, test and production data. The specification below shows one way to set requirements per category.

Data categoryPurposeExample dimensionsExample acceptance checkWhen checked
TrainingTeach the intended behaviorCompleteness, label accuracy, representativeness of the operating domain, duplicatesLabel audit on a random sample meets the agreement rate the data steward setBefore each training run
ValidationTune and select modelsIndependence from training data, label accuracy, coverageNo record overlaps with training dataBefore tuning starts
TestEstimate real performanceIndependence, subgroup and edge case coverage, currentnessEach subgroup the impact assessment names has enough records to measureBefore release
ProductionInputs the system receives in useSchema conformance, missing values, drift from training data, timelinessDrift metric stays inside the threshold set at releaseContinuously, with alerts

The table is illustrative. Dimensions and thresholds are decisions your organization makes and records, and ISO/IEC 42001 sets no values for anyone. For measures, the guidance points to the ISO/IEC 5259 series on data quality for analytics and machine learning. Part 1 (2024) sets terms and examples. Part 2 (2024) defines data quality measures, Part 3 (2024) sets data quality management requirements and guidelines, and Part 4 (2024) gives a process framework. Part 5 (2025) covers data quality governance, and Part 6 (2026) is a technical report on visualization.

A.7.5 A provenance record built from ISO 8000-2

Provenance under control A.7.5 needs a defined, documented process for recording where data came from across the life cycles of both the data and the AI system. The Annex B guidance points to ISO 8000-2, the data quality vocabulary, for what a provenance record can hold. The current edition dates from 2022, and ISO lists a revised edition at proof stage. The guidance names six kinds of event: creation, update, transcription, abstraction, validation and transfer of control. It adds two more to consider: sharing data without handing over control, and transformations.

Provenance eventWhat to recordIllustrative entry for a support assistant’s fine-tuning set
CreationWho or what produced the data, when, and under what termsTickets exported from the help desk system, March 2026
UpdateEach change, with version and dateVersion 1.2 relabels 400 tickets
TranscriptionCopies or format conversions between systemsAttachments converted from PDF to text, extractor version noted
AbstractionSummaries, aggregation, sampling or embeddingsTickets chunked and embedded with a named model version
ValidationChecks run and their resultsPersonal data scan passed; label audit report attached
Transfer of controlChanges of owner or custodian, with the agreementDataset delivered to a labelling vendor under contract
Sharing without transferWho else can read the dataCopy shared with an evaluation partner, access ends in June
TransformationPreparation steps applied, linked to A.7.6Deduplication and redaction scripts at a tagged commit

The guidance also asks organizations to consider whether provenance needs verifying, depending on the source of the data, its content and the context of use. A dataset bought from a broker for a hiring model deserves more checking than internal server telemetry feeding a capacity planning model, because the stakes differ so much. NIST’s Generative AI Profile, NIST AI 600-1, suggests that inventory entries carry provenance information such as source, signatures, versioning and watermarks. Those four items fit the creation and update rows above.

Datasheets and model cards as working artifacts

Datasheets for datasets give A.7 records a ready format. Gebru and colleagues proposed that every dataset carry a datasheet, with questions grouped by stage. Their stages are motivation, composition, collection process, preprocessing and labeling, uses, distribution and maintenance. Each stage lines up with an A.7 control.

Datasheet sectionA.7 control it supports
MotivationA.7.3, the reason a dataset was selected
CompositionA.7.4, quality and representativeness
Collection processA.7.3 and A.7.5, source and creation events
Preprocessing, cleaning and labelingA.7.6, preparation methods
UsesA.7.2, intended and excluded uses of the data
DistributionA.7.5, sharing and transfer of control
MaintenanceA.7.5 update events and A.7.2 enhancement

Model cards, proposed by Mitchell and colleagues, describe a trained model and its evaluation. Their template includes sections on evaluation data and training data. The authors note that training data details may not be possible to provide in practice. That candor suits vendor models whose training data is undisclosed. Record what is known and what is not, then add evaluation evidence that compensates. ISO/IEC 42001 requires neither format, but both give auditors a familiar structure.

A.7.6 Data preparation criteria and methods

Preparation methods, and the criteria for choosing them, must be defined and documented under control A.7.6. Raw data rarely suits a model as it arrives, so the Annex B guidance lists common preparation methods.

  • Statistical exploration of distributions, ranges and samples
  • Cleaning
  • Imputation of missing entries by a stated method
  • Normalization and scaling
  • Labelling of target variables for supervised learning
  • Encoding of categories into numbers a model can use

For each AI task, record the criteria you used to pick methods and the exact methods and transforms applied. Versioned pipeline code with run logs tied to a dataset version is the strongest evidence. For language model work, GAICC treats chunking, deduplication, redaction of personal data and prompt templating as preparation methods too. The guidance points to the ISO/IEC 5259 series and ISO/IEC 23053 for more on machine learning preparation.

One drafting detail helps in audit discussions. The A.7.6 entry in Annex B uses “shall”, where the other Annex B controls use “should”. The Annex A wording is the requirement either way.

A.7 for teams that do not train models

Annex A.7 applies to an organization that never trains a model whenever data flows into an AI system it builds, configures or runs. Many organizations deploying generative AI fall into this group, and their data lives in retrieval stores, test sets and prompt libraries. The table maps each data type to the A.7 controls that usually apply.

Data you holdWhere it sitsControls that usually applyMinimum evidence
Retrieval corpus for RAGDocument store or vector index behind an assistantA.7.3, A.7.4, A.7.5, A.7.6Source list with owners and rights; freshness rule; chunking and embedding specification; index version log
Fine-tuning setAdapter or fine-tuned modelAll fiveFull acquisition and provenance records; label checks; preparation specification
Evaluation setTest prompts, reference answers, red team promptsA.7.4, A.7.5Proof of independence from tuning data; coverage notes; version history
Prompt and few-shot librarySystem prompts and worked examplesA.7.5, A.7.6, and A.6 as a built componentVersion control, approver, source of each example
Vendor and third-party dataFoundation model training data; bought datasets; enrichment servicesA.7.3 and A.7.5, with A.10.3 suppliersSupplier documentation; contract terms on data use; known gaps recorded
Production inputs and logsUser prompts, uploaded files, stored outputsA.7.4, A.7.5, with A.6.2.8 event logsInput checks; retention rule; record of who can read logs

NIST’s Generative AI Profile makes a matching point. One of its suggested actions is to verify the provenance of training data and of test, evaluation, verification and validation data. The same action asks teams to confirm that fine-tuning or retrieval augmented generation data is grounded. A practitioner replying to a question in an online AI governance community in August 2026 drew the boundary plainly. “It is data going into the system, which is A.7, and that is what gets you provenance, quality and lineage.” The same reply put a few-shot example library under A.6 too, since engineers build and version that library like code.

Production inputs need one more note. The Annex B guidance on intended use says input data may need to match the system’s documentation for performance to hold. A retrieval assistant documented for policy documents will drift if staff paste in customer emails. For the wider threat picture, including prompt injection and leakage, see the risks specific to generative systems.

[TRAINER INSIGHT NEEDED: When GAICC instructors review retrieval augmented generation deployments, which A.7 record is most often missing for the retrieval corpus, and what did the organization offer as evidence instead?]

Bias and representativeness, stated with care

Bias checks under A.7.4 begin with a modest instruction in the guidance. Consider how bias affects system performance and fairness, then adjust the model and data until both are acceptable for the use case. Nothing in the control asks for bias-free data, and no record should claim it.

NIST SP 1270, published in March 2022, sorts bias in AI into three categories: systemic, statistical and human. The report observes that current efforts against harmful bias remain focused on computational factors, such as how representative a dataset is. It argues that human and systemic factors get overlooked. It also warns that “even when datasets are representative, they may still exhibit entrenched historical and systemic biases.” A representativeness check is therefore necessary evidence. It cannot prove fairness on its own.

Write bias evidence as measurements with limits. A defensible record says which subgroups were measured on which test set version, what gaps appeared and who accepted them against which threshold. An indefensible record says the data is unbiased or fair. The guidance points to ISO/IEC TR 24027 for the forms bias can take. Harms that arise from how a system is used, as opposed to its data, belong in the impact assessment.

Privacy and ISO/IEC 27001 next to A.7

Privacy runs through Annex A.7, though the area does not replace a privacy program. The A.7.2 guidance lists privacy implications as a data management topic. The A.7.3 guidance lists personal data among the data rights to record. Clause 4.1 notes that an organization’s role can follow from whether it acts as a controller or processor of personal data. The guidance on allocating responsibilities with third parties points to ISO/IEC 29100 and ISO/IEC 27701 for privacy controls. The 2025 edition of ISO/IEC 27701 is titled as a requirements standard for privacy information management systems, no longer as an extension to ISO/IEC 27001.

Memorization, membership inference and other model-specific privacy threats sit outside A.7’s records, and GAICC covers them in AI data and privacy risk management.

Teams certified to ISO/IEC 27001 already run several data controls. The table shows what each covers and what A.7 adds for AI.

ISO/IEC 27001:2022 controlWhat it coversWhat A.7 adds
5.9 Inventory of information and other associated assetsA list of information assets and ownersData category per AI use (training, validation, test, production), labelling process and known bias, through A.4.3
5.12 Classification of informationConfidentiality, integrity and availability classesQuality requirements and representativeness, which classification does not address
5.33 Protection of recordsRecords kept safe from loss or tamperingProvenance events across the data’s life, including transformations
5.34 Privacy and protection of PIILegal privacy dutiesRights beyond personal data, such as copyright and license terms
8.10 Information deletion and 8.11 Data maskingRemoving or hiding dataDocumented preparation methods that change data for each AI task

The comparison is GAICC’s own analysis. For the full set of overlaps, see how the two standards fit together.

How A.7 evidence feeds risk, impact and the SoA

Annex A.7 records feed the planning clauses directly and belong in the same evidence trail. Annex C of ISO/IEC 42001 names data quality and the data collection process as risk sources for machine learning, with data poisoning as an example. It also lists the availability and quality of training and test data as a potential organizational objective. The flow below shows how one dataset’s records move through the management system.

  1. The acquisition record (A.7.3) and the data resource documentation (A.4.3) place the dataset in your register of AI systems and resources.
  2. Quality results (A.7.4) and provenance gaps (A.7.5) become entries in the clause 6.1.2 risk assessment, which also weighs consequences for individuals and societies.
  3. The clause 6.1.4 impact assessment draws on the same records. ISO/IEC 42005 includes a section on data information and quality in the assessment record, and the results feed back into 6.1.2.
  4. Risk treatment under clause 6.1.3 decides which A.7 controls are needed. The Statement of Applicability then lists each included or excluded control with its reason.
  5. Internal audit traces a sample dataset through all of the above before the certification body does.

Register work starts with building an AI system inventory, a word ISO/IEC 42001 does not use but clause 4.1 and the A.4 resource controls imply. The impact side is covered in the AI impact assessment guide.

Exclusion wording for each data control belongs in the Statement of Applicability. A buyer of AI features might exclude A.7.6 because it prepares no data, then reopen that decision the day it adds a retrieval store.

What an auditor samples for A.7

At Stage 2, certification auditors look for proof that the A.7 controls operate. A common test design picks one in-scope AI system and traces one dataset end to end. The trace runs from the acquisition record to rights, quality results, provenance events, preparation steps, the model or index version and production monitoring. Gaps usually show up where one team hands data to another, for example when data engineers pass a cleaned set to a separate modeling team.

Run the same trace yourself first. An internal audit program for AI should sample at least one dataset per audit cycle and record what it followed.

Maturity levelWhat existsWhat it shows an auditor
DocumentedProcedures for all five controls; a dataset list with sourcesIntent, with no proof of operation yet
OperatingAcquisition, quality and provenance records for every dataset in current model or index versionsThat the controls run for the systems in scope
Incident readyRecords linked by dataset version, so one query answers which systems used a datasetThat the organization can respond when a source is withdrawn, poisoned or challenged

[TRAINER INSIGHT NEEDED: Among the five data controls, which one do organizations find hardest to prove at Stage 2, and which record closed the auditor’s question?]

Where A.7 work begins

Annex A.7 priorities differ by what your organization does with data.

Teams that train or fine-tune models should write A.7.3 and A.7.5 records for every dataset in the current model version before adding new ones. Retrieval on a vendor model is the next case. Treat the corpus as a dataset, with an owner, a freshness rule and an index version log.

Buyers of AI features inside packaged software focus on A.7.3 and A.7.5 through supplier documentation and record the gaps they cannot close. Organizations that already hold ISO/IEC 27001 can extend the asset inventory with data categories and provenance and keep a single register.

Candidates sitting the Lead Implementer exam should know the five titles in order and that A.7.1 is the objective, then practice tracing one dataset into clauses 6.1.2, 6.1.3 and 6.1.4. The ISO 42001 Lead Implementer online course covers that ground in its module on clauses 6.1.1 to 6.1.4. Video lessons and exam simulator access come with it.

Frequently asked questions

What is the ISO standard for data quality?

ISO 8000 is the data quality series, and ISO/IEC 42001 cites its Part 2 vocabulary for provenance records. ISO/IEC 25024 supplies the definition the A.7.4 guidance quotes. For machine learning, the ISO/IEC 5259 series covers data quality for analytics and ML, with parts published from 2024 to 2026.

Are the A.7 data controls mandatory?

Because Annex A is normative, each A.7 control belongs in your Statement of Applicability with a justification for its inclusion or exclusion under clause 6.1.3. A buyer that prepares no data might exclude A.7.6, for example, until it adds a retrieval store. The wider question, whether an organization has to adopt ISO/IEC 42001 at all, is covered in who needs ISO 42001 and why.

Which data quality dimensions does ISO 42001 require?

ISO/IEC 42001 sets no list of dimensions. Clause 3.25 ties data quality to the organization’s own data requirements in a given context, which leaves the dimensions and thresholds to you. For supervised or semi-supervised learning, the A.7.4 guidance asks that training, validation, test and production data quality be defined, measured and improved.

How do we handle in-scope and out-of-scope data in the same data warehouse?

Start from the AIMS scope set under clause 4.3. The A.7 controls attach to data used in AI systems, so tag each table with the in-scope systems that read it and keep A.7 records for those. Tables no in-scope system uses need no A.7 record until one starts reading them, which your change log should catch.

What A.7 evidence should we ask an AI vendor for?

Ask for the details an acquisition record needs: data categories, sources, rights and prior handling, plus any provenance the vendor can share. Control A.10.3 expects a process that checks supplied services, products and materials, datasets among them, against the organization’s approach to responsible AI. Where a vendor will not disclose training data, record the gap and add your own evaluation evidence.

Does A.7 cover synthetic data?

Yes. The A.7.3 guidance lists synthetic data among the data sources an acquisition record can name, next to internal, purchased, shared and open data. Treat a synthetic set like any other dataset, logging the generator and its version as the creation event and applying the quality checks you use for collected data.

When do A.7 records need updating?

A.7.5 asks for provenance across the life cycles of both the data and the AI system, so records stay live until both are retired. Update them at every enhancement event, such as retraining, a new source or an index refresh. Clause 8.2 also reruns the risk assessment when significant changes are proposed or occur, and that rerun should read the latest A.7 records.

Share it :
About the Author

Dr Faiz Rasool

Director at the Global AI Certification Council (GAICC) and PM Training School

A globally certified instructor in ISO/IEC, PMI®, TOGAF®, SAFe®, and Scrum.org disciplines. With over three years’ hands-on experience in ISO/IEC 42001 AI governance, he delivers training and consulting across New Zealand, Australia, Malaysia, the Philippines, and the UAE, combining high-end credentials with practical, real-world expertise and global reach.

About the Author

Latha Karthigaa

Head of AI Governance at the Global AI Certification Council (GAICC)

A PhD-qualified AI governance leader in Software Engineering from the University of Auckland, she brings hands-on experience founding and exiting AI companies, and leading real-world AI solutions for finance and legal firms across the USA, UK, Australia, and New Zealand, combining governance, risk, compliance, and commercial expertise.

Start Your ISO/IEC 42001 Lead Implementer Training Today

4.8 / 5.0 Rating

Related Post