Annex A.7 of ISO/IEC 42001 is the control area for data in AI systems. Its five controls govern how data is managed, sourced, checked, traced and prepared. The controls run from A.7.2 to A.7.6, because A.7.1 is the objective. Each control asks you to define and document a practice, then keep records proving you followed it. A team that never trains a model still holds A.7 data: retrieval corpora, evaluation sets, prompt libraries and production inputs.
Annex A.7 in ISO/IEC 42001 and the controls numbered 7 in ISO/IEC 27001 cover different ground. In the 2022 edition of ISO/IEC 27001, the 7.x controls deal with physical security, such as perimeters and entry. Data documentation also appears once more in the AI standard, in the A.4.3 control on data resources, which this page covers where it meets A.7.
GAICC restates the requirements here in its own words and quotes nothing beyond control IDs and short titles. Before building records from this page, confirm each control’s wording in a licensed copy of the published standard.
The five A.7 controls at a glance
Annex A.7 of ISO/IEC 42001 holds five controls under one objective. In plain terms, the objective is to understand what role data plays in your AI systems and what effects it has. That understanding must hold across the whole life cycle. The register gives each control’s ID, short title, what it asks, the evidence that usually proves it and the person who usually owns it.
| Control | Short title | Requirement in plain words | Typical evidence | Usual owner |
| A.7.2 | Data for development and enhancement of AI system | Define, document and run data management processes for building and improving AI systems | Data management procedure; design records naming the training approach and data quality expected | Data owner for each AI system |
| A.7.3 | Acquisition of data | Decide and write down where each dataset is sourced and why it was chosen | Acquisition record per dataset: categories, quantity, source, rights, prior handling | Data owner, with legal counsel for rights |
| A.7.4 | Quality of data for AI systems | Write data quality requirements and show that the data behind development and operation meets them | Quality specification per data category; validation results with dates | Data steward |
| A.7.5 | Data provenance | Define a process that records each dataset’s origin and later history, for as long as the data and the system exist | Provenance log per dataset version; supplier statements | Data engineering lead |
| A.7.6 | Data preparation | Set criteria for picking preparation methods, then log which methods were applied | Preparation specification per AI task; versioned pipeline code and run logs | Lead data scientist or ML engineer |
Both the evidence column and the owner column reflect GAICC’s editorial view; neither comes from the standard. Annex B, which is also normative, repeats the same five controls as B.7.2 to B.7.6 and adds implementation guidance for each. The full set of areas sits in the control objectives page, which lists all nine Annex A areas.
Why the numbering starts at A.7.2
A.7.1 is the objective statement, so the first data control is A.7.2. Every Annex A area follows that pattern: the objective takes the .1 slot (Annex B prints it as B.7.1) and the controls follow. AI answers still get the A.7 numbers wrong. GAICC captured US Google results in early October 2026. Google’s AI summary for the A.7 query listed A.7.3 to A.7.6 and left out A.7.2. A Google AI Mode answer about the Statement of Applicability labelled A.7.2 as data quality, which is A.7.4. One ranking Annex A guide renumbers the areas and files data under A.6.
Use the order in the register, from management through sourcing, quality and provenance to preparation. The same order mirrors how a dataset moves through a project.
A.7.2 Data management for development and enhancement
Control A.7.2 asks for documented data management processes covering how AI systems are developed and improved. The Annex B guidance lists five topics such processes can address.
- Privacy and security, because some of the data is sensitive.
- Threats that data-dependent development opens, for security and for safety.
- Transparency, since teams may need to show provenance and explain how data shapes an output.
- Representativeness of training data for the domain where the system will operate.
- Accuracy and integrity, which errors and tampering both degrade.
Evidence for A.7.2 is a written procedure plus proof that teams follow it, because a procedure on its own shows intent and rarely satisfies a certification auditor. Pair it with the design record that A.6.2.3 already asks for. Among the design choices to document, that guidance lists how the model will be trained and what data quality it needs.
Document data at the design stage
Implementers often ask whether data belongs in the record before any data is collected. In practice it does, because data decisions first appear in the design record, which states the training approach and the quality the data must meet. The acquisition and quality controls then test real data against that plan. The A.7.2 guidance points to ISO/IEC 22989 for life cycle and data management concepts, which helps teams name the stages consistently.
“Enhancement” covers any later change to a system’s data: retraining on fresh records, fine-tuning, adding a new source or refreshing a retrieval index. Each of those events should reopen the A.7 records for the affected dataset.
A.7.3 Acquisition records for every dataset
Acquisition is the subject of control A.7.3, which requires you to determine and document how the data used in AI systems is acquired and selected. The Annex B guidance lists the details an acquisition record can hold. The table turns that list into fields, with notes for US teams.
| Field | What to record | Note for US teams |
| Data categories | Which kinds of data the system needs; ISO/IEC 19944-1 offers a category structure | Flag personal information early |
| Quantity | Volume needed and volume obtained | Record sampling choices |
| Source | Internal, purchased, shared, open or synthetic | Keep the contract or license reference |
| Source characteristics | Static, streamed, gathered or machine generated | Streams need a refresh rule |
| Data subjects | Demographics and characteristics, including known or likely biases | Feeds the bias checks under A.7.4 |
| Prior handling | Earlier uses and whether privacy and security requirements were met | Ask suppliers in writing |
| Data rights | Personal information, copyright and license terms | Training data copyright is still being litigated |
| Metadata | Labelling and enrichment details | Name the labelling vendor, if any |
| Provenance | Link to the A.7.5 record | One ID shared across both records |
Rights deserve their own line because the legal position is unsettled in the United States. GAICC’s case note on scraped video used as training data shows how a missing acquisition record turns into a governance question. Where personal information is in the data, state privacy laws such as California’s can also apply.
A.7.4 Data quality requirements per data category
Control A.7.4 requires documented data quality requirements and proof that data used to develop and operate the AI system meets them. Clause 3.25 of ISO/IEC 42001 treats data quality as the degree to which data satisfies the requirements an organization sets for a given context. The Annex B guidance also cites the ISO/IEC 25024 definition, which ties quality to stated and implied needs under specified conditions. Both definitions lead to the same task: write the requirements down before you measure anything.
For organizations using supervised or semi-supervised learning, the guidance asks that the quality of four data categories be defined, measured and improved. Those categories are training, validation, test and production data. The specification below shows one way to set requirements per category.
| Data category | Purpose | Example dimensions | Example acceptance check | When checked |
| Training | Teach the intended behavior | Completeness, label accuracy, representativeness of the operating domain, duplicates | Label audit on a random sample meets the agreement rate the data steward set | Before each training run |
| Validation | Tune and select models | Independence from training data, label accuracy, coverage | No record overlaps with training data | Before tuning starts |
| Test | Estimate real performance | Independence, subgroup and edge case coverage, currentness | Each subgroup the impact assessment names has enough records to measure | Before release |
| Production | Inputs the system receives in use | Schema conformance, missing values, drift from training data, timeliness | Drift metric stays inside the threshold set at release | Continuously, with alerts |
The table is illustrative. Dimensions and thresholds are decisions your organization makes and records, and ISO/IEC 42001 sets no values for anyone. For measures, the guidance points to the ISO/IEC 5259 series on data quality for analytics and machine learning. Part 1 (2024) sets terms and examples. Part 2 (2024) defines data quality measures, Part 3 (2024) sets data quality management requirements and guidelines, and Part 4 (2024) gives a process framework. Part 5 (2025) covers data quality governance, and Part 6 (2026) is a technical report on visualization.
A.7.5 A provenance record built from ISO 8000-2
Provenance under control A.7.5 needs a defined, documented process for recording where data came from across the life cycles of both the data and the AI system. The Annex B guidance points to ISO 8000-2, the data quality vocabulary, for what a provenance record can hold. The current edition dates from 2022, and ISO lists a revised edition at proof stage. The guidance names six kinds of event: creation, update, transcription, abstraction, validation and transfer of control. It adds two more to consider: sharing data without handing over control, and transformations.
| Provenance event | What to record | Illustrative entry for a support assistant’s fine-tuning set |
| Creation | Who or what produced the data, when, and under what terms | Tickets exported from the help desk system, March 2026 |
| Update | Each change, with version and date | Version 1.2 relabels 400 tickets |
| Transcription | Copies or format conversions between systems | Attachments converted from PDF to text, extractor version noted |
| Abstraction | Summaries, aggregation, sampling or embeddings | Tickets chunked and embedded with a named model version |
| Validation | Checks run and their results | Personal data scan passed; label audit report attached |
| Transfer of control | Changes of owner or custodian, with the agreement | Dataset delivered to a labelling vendor under contract |
| Sharing without transfer | Who else can read the data | Copy shared with an evaluation partner, access ends in June |
| Transformation | Preparation steps applied, linked to A.7.6 | Deduplication and redaction scripts at a tagged commit |
The guidance also asks organizations to consider whether provenance needs verifying, depending on the source of the data, its content and the context of use. A dataset bought from a broker for a hiring model deserves more checking than internal server telemetry feeding a capacity planning model, because the stakes differ so much. NIST’s Generative AI Profile, NIST AI 600-1, suggests that inventory entries carry provenance information such as source, signatures, versioning and watermarks. Those four items fit the creation and update rows above.
Datasheets and model cards as working artifacts
Datasheets for datasets give A.7 records a ready format. Gebru and colleagues proposed that every dataset carry a datasheet, with questions grouped by stage. Their stages are motivation, composition, collection process, preprocessing and labeling, uses, distribution and maintenance. Each stage lines up with an A.7 control.
| Datasheet section | A.7 control it supports |
| Motivation | A.7.3, the reason a dataset was selected |
| Composition | A.7.4, quality and representativeness |
| Collection process | A.7.3 and A.7.5, source and creation events |
| Preprocessing, cleaning and labeling | A.7.6, preparation methods |
| Uses | A.7.2, intended and excluded uses of the data |
| Distribution | A.7.5, sharing and transfer of control |
| Maintenance | A.7.5 update events and A.7.2 enhancement |
Model cards, proposed by Mitchell and colleagues, describe a trained model and its evaluation. Their template includes sections on evaluation data and training data. The authors note that training data details may not be possible to provide in practice. That candor suits vendor models whose training data is undisclosed. Record what is known and what is not, then add evaluation evidence that compensates. ISO/IEC 42001 requires neither format, but both give auditors a familiar structure.
A.7.6 Data preparation criteria and methods
Preparation methods, and the criteria for choosing them, must be defined and documented under control A.7.6. Raw data rarely suits a model as it arrives, so the Annex B guidance lists common preparation methods.
- Statistical exploration of distributions, ranges and samples
- Cleaning
- Imputation of missing entries by a stated method
- Normalization and scaling
- Labelling of target variables for supervised learning
- Encoding of categories into numbers a model can use
For each AI task, record the criteria you used to pick methods and the exact methods and transforms applied. Versioned pipeline code with run logs tied to a dataset version is the strongest evidence. For language model work, GAICC treats chunking, deduplication, redaction of personal data and prompt templating as preparation methods too. The guidance points to the ISO/IEC 5259 series and ISO/IEC 23053 for more on machine learning preparation.
One drafting detail helps in audit discussions. The A.7.6 entry in Annex B uses “shall”, where the other Annex B controls use “should”. The Annex A wording is the requirement either way.
A.7 for teams that do not train models
Annex A.7 applies to an organization that never trains a model whenever data flows into an AI system it builds, configures or runs. Many organizations deploying generative AI fall into this group, and their data lives in retrieval stores, test sets and prompt libraries. The table maps each data type to the A.7 controls that usually apply.
| Data you hold | Where it sits | Controls that usually apply | Minimum evidence |
| Retrieval corpus for RAG | Document store or vector index behind an assistant | A.7.3, A.7.4, A.7.5, A.7.6 | Source list with owners and rights; freshness rule; chunking and embedding specification; index version log |
| Fine-tuning set | Adapter or fine-tuned model | All five | Full acquisition and provenance records; label checks; preparation specification |
| Evaluation set | Test prompts, reference answers, red team prompts | A.7.4, A.7.5 | Proof of independence from tuning data; coverage notes; version history |
| Prompt and few-shot library | System prompts and worked examples | A.7.5, A.7.6, and A.6 as a built component | Version control, approver, source of each example |
| Vendor and third-party data | Foundation model training data; bought datasets; enrichment services | A.7.3 and A.7.5, with A.10.3 suppliers | Supplier documentation; contract terms on data use; known gaps recorded |
| Production inputs and logs | User prompts, uploaded files, stored outputs | A.7.4, A.7.5, with A.6.2.8 event logs | Input checks; retention rule; record of who can read logs |
NIST’s Generative AI Profile makes a matching point. One of its suggested actions is to verify the provenance of training data and of test, evaluation, verification and validation data. The same action asks teams to confirm that fine-tuning or retrieval augmented generation data is grounded. A practitioner replying to a question in an online AI governance community in August 2026 drew the boundary plainly. “It is data going into the system, which is A.7, and that is what gets you provenance, quality and lineage.” The same reply put a few-shot example library under A.6 too, since engineers build and version that library like code.
Production inputs need one more note. The Annex B guidance on intended use says input data may need to match the system’s documentation for performance to hold. A retrieval assistant documented for policy documents will drift if staff paste in customer emails. For the wider threat picture, including prompt injection and leakage, see the risks specific to generative systems.
[TRAINER INSIGHT NEEDED: When GAICC instructors review retrieval augmented generation deployments, which A.7 record is most often missing for the retrieval corpus, and what did the organization offer as evidence instead?]
Bias and representativeness, stated with care
Bias checks under A.7.4 begin with a modest instruction in the guidance. Consider how bias affects system performance and fairness, then adjust the model and data until both are acceptable for the use case. Nothing in the control asks for bias-free data, and no record should claim it.
NIST SP 1270, published in March 2022, sorts bias in AI into three categories: systemic, statistical and human. The report observes that current efforts against harmful bias remain focused on computational factors, such as how representative a dataset is. It argues that human and systemic factors get overlooked. It also warns that “even when datasets are representative, they may still exhibit entrenched historical and systemic biases.” A representativeness check is therefore necessary evidence. It cannot prove fairness on its own.
Write bias evidence as measurements with limits. A defensible record says which subgroups were measured on which test set version, what gaps appeared and who accepted them against which threshold. An indefensible record says the data is unbiased or fair. The guidance points to ISO/IEC TR 24027 for the forms bias can take. Harms that arise from how a system is used, as opposed to its data, belong in the impact assessment.
Privacy and ISO/IEC 27001 next to A.7
Privacy runs through Annex A.7, though the area does not replace a privacy program. The A.7.2 guidance lists privacy implications as a data management topic. The A.7.3 guidance lists personal data among the data rights to record. Clause 4.1 notes that an organization’s role can follow from whether it acts as a controller or processor of personal data. The guidance on allocating responsibilities with third parties points to ISO/IEC 29100 and ISO/IEC 27701 for privacy controls. The 2025 edition of ISO/IEC 27701 is titled as a requirements standard for privacy information management systems, no longer as an extension to ISO/IEC 27001.
Memorization, membership inference and other model-specific privacy threats sit outside A.7’s records, and GAICC covers them in AI data and privacy risk management.
Teams certified to ISO/IEC 27001 already run several data controls. The table shows what each covers and what A.7 adds for AI.
| ISO/IEC 27001:2022 control | What it covers | What A.7 adds |
| 5.9 Inventory of information and other associated assets | A list of information assets and owners | Data category per AI use (training, validation, test, production), labelling process and known bias, through A.4.3 |
| 5.12 Classification of information | Confidentiality, integrity and availability classes | Quality requirements and representativeness, which classification does not address |
| 5.33 Protection of records | Records kept safe from loss or tampering | Provenance events across the data’s life, including transformations |
| 5.34 Privacy and protection of PII | Legal privacy duties | Rights beyond personal data, such as copyright and license terms |
| 8.10 Information deletion and 8.11 Data masking | Removing or hiding data | Documented preparation methods that change data for each AI task |
The comparison is GAICC’s own analysis. For the full set of overlaps, see how the two standards fit together.
How A.7 evidence feeds risk, impact and the SoA
Annex A.7 records feed the planning clauses directly and belong in the same evidence trail. Annex C of ISO/IEC 42001 names data quality and the data collection process as risk sources for machine learning, with data poisoning as an example. It also lists the availability and quality of training and test data as a potential organizational objective. The flow below shows how one dataset’s records move through the management system.
- The acquisition record (A.7.3) and the data resource documentation (A.4.3) place the dataset in your register of AI systems and resources.
- Quality results (A.7.4) and provenance gaps (A.7.5) become entries in the clause 6.1.2 risk assessment, which also weighs consequences for individuals and societies.
- The clause 6.1.4 impact assessment draws on the same records. ISO/IEC 42005 includes a section on data information and quality in the assessment record, and the results feed back into 6.1.2.
- Risk treatment under clause 6.1.3 decides which A.7 controls are needed. The Statement of Applicability then lists each included or excluded control with its reason.
- Internal audit traces a sample dataset through all of the above before the certification body does.
Register work starts with building an AI system inventory, a word ISO/IEC 42001 does not use but clause 4.1 and the A.4 resource controls imply. The impact side is covered in the AI impact assessment guide.
Exclusion wording for each data control belongs in the Statement of Applicability. A buyer of AI features might exclude A.7.6 because it prepares no data, then reopen that decision the day it adds a retrieval store.
What an auditor samples for A.7
At Stage 2, certification auditors look for proof that the A.7 controls operate. A common test design picks one in-scope AI system and traces one dataset end to end. The trace runs from the acquisition record to rights, quality results, provenance events, preparation steps, the model or index version and production monitoring. Gaps usually show up where one team hands data to another, for example when data engineers pass a cleaned set to a separate modeling team.
Run the same trace yourself first. An internal audit program for AI should sample at least one dataset per audit cycle and record what it followed.
| Maturity level | What exists | What it shows an auditor |
| Documented | Procedures for all five controls; a dataset list with sources | Intent, with no proof of operation yet |
| Operating | Acquisition, quality and provenance records for every dataset in current model or index versions | That the controls run for the systems in scope |
| Incident ready | Records linked by dataset version, so one query answers which systems used a dataset | That the organization can respond when a source is withdrawn, poisoned or challenged |
[TRAINER INSIGHT NEEDED: Among the five data controls, which one do organizations find hardest to prove at Stage 2, and which record closed the auditor’s question?]
Where A.7 work begins
Annex A.7 priorities differ by what your organization does with data.
Teams that train or fine-tune models should write A.7.3 and A.7.5 records for every dataset in the current model version before adding new ones. Retrieval on a vendor model is the next case. Treat the corpus as a dataset, with an owner, a freshness rule and an index version log.
Buyers of AI features inside packaged software focus on A.7.3 and A.7.5 through supplier documentation and record the gaps they cannot close. Organizations that already hold ISO/IEC 27001 can extend the asset inventory with data categories and provenance and keep a single register.
Candidates sitting the Lead Implementer exam should know the five titles in order and that A.7.1 is the objective, then practice tracing one dataset into clauses 6.1.2, 6.1.3 and 6.1.4. The ISO 42001 Lead Implementer online course covers that ground in its module on clauses 6.1.1 to 6.1.4. Video lessons and exam simulator access come with it.
Frequently asked questions
What is the ISO standard for data quality?
ISO 8000 is the data quality series, and ISO/IEC 42001 cites its Part 2 vocabulary for provenance records. ISO/IEC 25024 supplies the definition the A.7.4 guidance quotes. For machine learning, the ISO/IEC 5259 series covers data quality for analytics and ML, with parts published from 2024 to 2026.
Are the A.7 data controls mandatory?
Because Annex A is normative, each A.7 control belongs in your Statement of Applicability with a justification for its inclusion or exclusion under clause 6.1.3. A buyer that prepares no data might exclude A.7.6, for example, until it adds a retrieval store. The wider question, whether an organization has to adopt ISO/IEC 42001 at all, is covered in who needs ISO 42001 and why.
Which data quality dimensions does ISO 42001 require?
ISO/IEC 42001 sets no list of dimensions. Clause 3.25 ties data quality to the organization’s own data requirements in a given context, which leaves the dimensions and thresholds to you. For supervised or semi-supervised learning, the A.7.4 guidance asks that training, validation, test and production data quality be defined, measured and improved.
How do we handle in-scope and out-of-scope data in the same data warehouse?
Start from the AIMS scope set under clause 4.3. The A.7 controls attach to data used in AI systems, so tag each table with the in-scope systems that read it and keep A.7 records for those. Tables no in-scope system uses need no A.7 record until one starts reading them, which your change log should catch.
What A.7 evidence should we ask an AI vendor for?
Ask for the details an acquisition record needs: data categories, sources, rights and prior handling, plus any provenance the vendor can share. Control A.10.3 expects a process that checks supplied services, products and materials, datasets among them, against the organization’s approach to responsible AI. Where a vendor will not disclose training data, record the gap and add your own evaluation evidence.
Does A.7 cover synthetic data?
Yes. The A.7.3 guidance lists synthetic data among the data sources an acquisition record can name, next to internal, purchased, shared and open data. Treat a synthetic set like any other dataset, logging the generator and its version as the creation event and applying the quality checks you use for collected data.
When do A.7 records need updating?
A.7.5 asks for provenance across the life cycles of both the data and the AI system, so records stay live until both are retired. Update them at every enhancement event, such as retraining, a new source or an index refresh. Clause 8.2 also reruns the risk assessment when significant changes are proposed or occur, and that rerun should read the latest A.7 records.

