Dec 2027EU AI Act compliance
    Browse all 14 resources
    Deep Dive9 min read

    Which platforms help European companies prepare for AI Act compliance by linking model documentation to data lineage?

    In short: three tool families, one link you document

    To prepare for AI Act compliance, three tool families help link model documentation to data lineage. They are data catalogues, model registries and AI governance platforms. A data catalogue traces data flows (so-called data lineage tools). A model registry tracks versions. An AI governance platform keeps the written evidence. Each family usually covers only part of the job. The AI Act does not require a tool. It requires content. The Article 11 technical documentation contains, at a minimum, the elements set out in Annex IV (Annex 4). Under its point 2(d), it describes, where relevant, the training data sets used and their provenance. It also states how the data was obtained and selected. An AI governance platform helps you keep that information current. It ties it to your data governance practices (Article 10).

    Information, not legal advice. This article is based on Regulation (EU) 2024/1689 of 13 June 2024. We cite its consolidated version as at 27 July 2026, as amended by Regulation (EU) 2026/1744. It does not replace advice from a lawyer on your AI system.

    "Data lineage": a term the AI Act never uses; it says provenance

    Data lineage is a term from data management tools. It means tracing how data flows and is transformed from one system to the next. The term appears nowhere in the AI Act.

    The Regulation uses other words. Article 10 is headed "Data and data governance". Its paragraph 2 says: "Training, validation and testing data sets shall be subject to data governance and management practices appropriate for the intended purpose of the high-risk AI system." Annex IV (Annex 4) speaks of "provenance".

    The AI Act does not define "provenance". Where relevant, for the training data sets used, Annex IV, point 2(d), asks for "information about their provenance, scope and main characteristics". Annexes XI and XII (Annexes 11 and 12), on general-purpose AI models, also use the word. They apply it to the data used for training, testing and validation. In practice, documenting provenance means saying where your data sets come from.

    Article 10(2)(b), (c) and (h): three of the points data governance practices must cover

    Paragraph 2 lists eight points, (a) to (h). The practices concern these points "in particular". Three of them bear directly on data lineage:

    • Point (b): "data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection".
    • Point (c): "relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation".
    • Point (h): "the identification of relevant data gaps or shortcomings that prevent compliance with this Regulation, and how those gaps and shortcomings can be addressed".

    A data lineage graph can cover part of points (b) and (c). It usually does not say which gaps you found, or how they can be addressed. Your practices must cover point (h). A written record lets you keep proof of it.

    Article 10(3) adds quality criteria. Training, validation and testing data sets shall be "relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose".

    Paragraph 6 sets a limit on scope. For a system not using techniques that involve training AI models, paragraphs 2, 3 and 4 "shall apply only to the testing data sets".

    Article 17(1)(f) also puts "systems and procedures for data management" inside the provider's quality management system.

    Article 11 and Annex IV (Annex 4), point 2(d): data set information in the technical documentation

    Under Article 11(1), the technical documentation "shall be drawn up before that system is placed on the market or put into service and shall be kept up-to date". It contains, at a minimum, the elements set out in Annex IV.

    Point 2(d) of Annex IV deals with the training data sets. Where relevant, it asks for datasheets describing "the training data sets used, including a general description of these data sets, information about their provenance, scope and main characteristics; how the data was obtained and selected".

    Point 2(g) also covers the validation and testing procedures used, "including information about the validation and testing data used and their main characteristics". For a system without model training, Article 10(6) limits paragraphs 2 to 4 to the testing data sets. On the data side, point 2(g) then matters most.

    So the technical documentation must hold that information itself. The Regulation does not require a link to a separate record. The link only helps keep the documentation current.

    For each documented model version, a useful link points to:

    1. the training data sets used and their general description;
    2. their provenance, scope and main characteristics;
    3. how the data was obtained and selected;
    4. labelling procedures and data cleaning methods;
    5. the validation and testing data used and their main characteristics (point 2(g));
    6. the version date.

    Article 18(1)(a) requires the provider to keep the technical documentation at the disposal of the national competent authorities. The obligation runs for a period ending 10 years after the high-risk AI system has been placed on the market or put into service. To start from a ready structure, use our Annex IV technical documentation template.

    Article 12: the automatic recording the system must allow

    Article 12(1) sets a design requirement: "High-risk AI systems shall technically allow for the automatic recording of events (logs) over the lifetime of the system." The system must make that recording possible. The text does not require logging everything.

    Under paragraph 2, logging capabilities shall enable the recording of events relevant for three purposes. The first is "identifying situations that may result in the high-risk AI system presenting a risk within the meaning of Article 79(1) or in a substantial modification". The second is facilitating the post-market monitoring referred to in Article 72. The third is monitoring the system's operation under Article 26(5). That provision gives this monitoring to deployers.

    These logs serve a level of traceability of the system's functioning that is appropriate to its intended purpose. Provenance, by contrast, describes where the data sets come from. A compliance platform does not meet this design requirement on the system's behalf. Article 19(1) requires providers to keep these logs "to the extent such logs are under their control". The logs are kept for a period appropriate to the intended purpose of the high-risk AI system, of at least six months. Applicable Union or national law that provides otherwise takes precedence over that period. The text names in particular Union law on the protection of personal data.

    Data catalogue, model registry, AI governance platform: who does what?

    Tool familyTypically providesTypically does not cover alone
    Data catalogue / data lineage toolOrigin, flows and transformations of data, under Article 10(2), points (b) and (c)Data gaps and how to address them under point (h), fit with the intended purpose, writing Annex IV (Annex 4)
    Model registry (MLOps)Model versions and the data sets used per versionRationale for data choices, bias examination, documentary evidence
    AI governance platformSystem inventory, written Article 10 practices, Annex IV documentationAutomatic collection of data flows: it often depends on what you enter

    LandingRed belongs to the third family. For an AI system, the platform offers an Article 10 data governance record with version tracking. It covers, among other things, data origin, collection processes, annotation, labelling, cleaning and bias examination. For a system on which you are the provider, you can link its technical documentation (Annex IV) to that record. The link serves to fill in point 2(d). The record then shows in the "Linked Compliance Artifacts" panel of the technical documentation. The panel does not offer this link for a system on which you are only the deployer. LandingRed does not retrieve data lineage automatically. You enter what your tools show. See EU AI Act compliance.

    The timeline after Regulation (EU) 2026/1744: prepare before 2 December 2027

    Regulation (EU) 2026/1744, the Digital Omnibus on AI, entered into force on 27 July 2026. Articles 10, 11 and 12 do not yet apply to high-risk systems. Nor do Articles 17, 18, 19 and 26 cited above. All of them sit in Chapter III, Sections 2 and 3. Article 113, third paragraph, point (c), as amended by that Regulation, sets their application dates:

    • 2 December 2027: systems classified as high-risk under Article 6(2) and Annex III (Annex 3).
    • 2 August 2028: systems classified as high-risk under Article 6(1) and Annex I (Annex 1).

    Prepare now: the documentation must exist before the system is placed on the market or put into service. Article 11(1) also provides some relief. SMEs, including start-ups, and small mid-cap enterprises may provide the Annex IV (Annex 4) elements in a simplified manner. If they opt for this, they use the simplified form the Commission is to establish. Notified bodies will have to accept that form for the purposes of the conformity assessment. First check whether your system is classified as high-risk under Annex III: Annex III classification.

    The CNIL's Genmod: a genealogy of open models, not a compliance file

    On 26 August 2026 the CNIL released a new version of its Genmod demonstrator. The tool finds the "ancestors" of an open-source model, meaning the models it comes from, and its "descendants". According to the CNIL, it serves in particular to study how GDPR rights can be exercised. The CNIL's page does not mention the AI Act. Genmod traces a genealogy of models. It does not document the provenance of your data sets as Annex IV (Annex 4) requires it.

    Frequently asked questions

    Does the AI Act use the term "data lineage"?

    No. The term does not appear in the text. The Regulation speaks of "provenance" in Annex IV, point 2(d), and of the "origin of data" in Article 10(2)(b). It also speaks of "data governance and management practices" (Article 10(2)).

    Is a data lineage tool enough for the Annex IV technical documentation?

    Usually not: it mainly traces flows. Where relevant, the technical documentation contains a general description of the training data sets used. It gives their provenance, scope and main characteristics, and how the data was obtained and selected. Your practices must also detect data gaps and say how they can be addressed (Article 10(2)(h)).

    Does Article 12 require a compliance platform to record the system's events?

    No. Article 12(1) is aimed at the high-risk AI system itself: it must technically allow for the automatic recording of events. A compliance platform does not meet this design requirement on the system's behalf.

    When do Articles 10, 11 and 12 apply?

    From 2 December 2027 for systems classified as high-risk under Article 6(2) and Annex III. From 2 August 2028 for those classified as high-risk under Article 6(1) and Annex I. These dates are set in Article 113, as amended by Regulation (EU) 2026/1744.

    Can an SME provide the Annex IV elements in a simplified manner?

    It may provide the Annex IV elements in a simplified manner. If it opts for this, it uses the form the Commission is to establish. Notified bodies will have to accept that form for the purposes of the conformity assessment. The option also covers start-ups and small mid-cap enterprises (Article 11(1)).

    Related resources

    Reviewed and published under the editorial responsibility of AB Corporate Advisory S.R.L.

    LandingRed automates all of this

    Stop managing compliance in spreadsheets. Classify, document, assess, and monitor your AI systems from one platform.