The AI Asset Register: The Problem and the Boundary - Part one of four
Most AI registers look complete and are not. This article sets out the three decisions, what counts as an AI system, what a record represents, and what classification is for, that determine whether yours reflects your actual estate.
Every AI register I have been handed looks complete. That is the problem with them.
It usually arrives as a spreadsheet, or lately as a module in a GRC platform. Forty entries, sometimes sixty. Each one has a named owner, a risk rating and a review date. Somebody senior has signed it off.
Then you start asking questions, and within an hour you have found a summarisation tool the customer service team adopted in March, a credit model that appears once in the register but sits behind four different lending decisions, and a translation feature that arrived inside a case management system nobody thought of as an AI purchase.
The register is not wrong. It is just not the estate.
This has little to do with diligence. The people building these registers are usually careful and working hard. What defeats them is that three decisions get made by accident, before anyone realises a decision is being made at all.
What counts as an AI system. What a single record represents. What classification is actually for.
All three are cheap to settle on day one and expensive to change afterwards, because changing them means rebuilding the register rather than editing it. This series is about making those choices deliberately, finding the systems you didn't know existed, and keeping the whole thing true once the project team has moved on.
Why this is harder than an application inventory
Most organisations already know how to inventory software. Procurement record, project, change request, CMDB entry. The discipline is decades old.
AI defeats that machinery in three specific ways.
It arrives without a procurement event. Nobody buys AI. They buy a CRM, and eighteen months later the CRM ships a feature. No requisition to intercept and no project to gate.
It changes underneath you. A vendor updates a model, and your entry stays exactly as it was, describing a system that no longer exists in the form you assessed.
And the same component carries different risk in different uses. A shared foundation model serving four teams is not one thing with one rating. It is four risk positions that happen to share an API key.
In the 2026 Cost of a Data Breach Report, IBM found that 68 per cent of the breached organisations studied lacked governance to manage AI or detect shadow AI. Thirty-five per cent had no policy at all and a further 33 per cent had one still in development. Over the same period, the share of security incidents involving shadow AI more than doubled, reaching 43 per cent against 20 per cent the year before.
Two further findings from that report are worth carrying into this series. Lack of visibility into the number and location of applications carried a measurable price, adding just over USD 200,000 to the average breach cost. Among organisations that suffered an AI-related breach, the root causes were mostly structural rather than model-related: compromised connected APIs, applications, and cloud misconfigurations. Not knowing what you have is not a documentation problem. It is where the incidents are.
What the series covers
Four parts, weekly. This one settles the boundary and the record schema. Part two covers classification, including how much can be automated. Part three is discovery, where the shadow AI question gets answered. Part four is the operating model that keeps the register true after go-live.
What counts as an AI system
Start here, because everything downstream inherits this decision.
The instinct is to find a definition and apply it. That instinct is right and also not sufficient, because every available definition leaves a band of genuinely arguable cases. What works better is to apply three lenses together and accept that some answers will be judgements you have to write down.
The Technology lens
Article 3(1) of the AI Act contains the definition most organisations will cite, and the Commission’s February 2025 guidelines break it into seven elements. A machine-based system. Designed to operate with varying levels of autonomy. That may exhibit adaptiveness after deployment. For explicit or implicit objectives, it infers from the input it receives how to generate outputs. Outputs include predictions, content, recommendations, or decisions. Capable of influencing physical or virtual environments.
Three things about that definition catch people out.
Adaptiveness after deployment is not a necessary condition. The Act says may, not shall. A model frozen at deployment is still in scope, which captures much of the traditional machine learning that teams assume falls outside.
The meaning of inference is where most in-house definitions go wrong. The natural instinct is to draw the line between behaviour learned from data and behaviour specified by a developer, which would put rule engines safely outside. Recital 12 does not support that. It describes the techniques enabling inference as including machine learning approaches and also logic- and knowledge-based approaches that infer from encoded knowledge or a symbolic representation of the task. The line the Act actually draws is that the capacity to infer transcends basic data processing by enabling learning, reasoning or modelling. That is broader than machine learning. If your inclusion rule says “uses ML”, you have written a rule that does not match the instrument you are citing.
And nobody is going to hand you a clean boundary. The guidelines state plainly that no automatic determination or exhaustive list of systems inside or outside the definition is possible, and that assessment has to rest on the specific architecture and functionality of the system in front of you. The Commission declined to draw a bright line. You will not draw one either, and the sooner you stop trying, the sooner you can start documenting judgements instead.
This lens should produce not a boundary but a deployment taxonomy, because the deployment pattern determines who holds which obligation. Built in-house. Fine-tuned on a base model. Consumed through an API. Embedded in a purchased product. NIST asks the same question at MAP 2.1, which wants the specific tasks and the methods used to implement them.
IBM’s 2026 data gives that taxonomy some weight. AI-related breaches occurred at broadly similar rates whether the model was open source, delivered by a vendor as SaaS, or deployed by a vendor on-premises. What differed was the cost. Open source models averaged USD 5.63 million per breach, and models trained in-house averaged USD 4.98 million. Same category of incident, different exposure depending on where the thing came from. That is a provenance effect, and a register that does not record provenance cannot see it.
The data lens
The second lens asks what the system consumes and produces. Training data, if any. Inference time inputs. Whether it generates new data that persists. Whether inputs cross a tenancy boundary on their way to a model.
This lens earns its place because it links to the data governance you already have. Most organisations have classified their data, and almost nobody has connected that work to their AI estate. Part two builds directly on this link.
The risk and decision role lens
The third lens asks what the output does. Informs, recommends, or decides.
Then the harder question. Can a human practically intervene? Not: Is there a review step in the process diagram, but does the reviewer have the time, the information and the standing to overturn the output? MAP 2.2 asks for the system’s knowledge limits and how its output may be used and overseen by humans, which is the right pairing: what the system cannot do, then what the human is expected to catch. A review step nobody exercises is not oversight.
The line that matters most in scoping work is this. A simple model in a consequential decision path outranks a sophisticated one in a sandbox. Sophistication is not risk. Consequence is risk. I have seen far more governance attention spent on an impressive model doing nothing important than on a logistic regression quietly deciding who gets a callback.
What you write down
Two artefacts, and the second is the one people skip.
An inclusion rule, in writing. ISO 42001 Clause 4.3 already requires the scope of the management system to be available as documented information, so for anyone heading towards certification this is not optional housekeeping.
And an exclusion list, with reasoning attached to each entry. Every organisation excludes things. Rules engines, statistical scoring, forecasting models, autocomplete. Those exclusions are what a regulator or an auditor will ask about, and reconstructing the reasoning eighteen months later from memory is a bad afternoon.
Article 3(13) defines reasonably foreseeable misuse, and MAP 3.3 asks for a targeted application scope documented against the system’s capability. Both point the same way. Record what the system is for and what it is not for, because the second is what gets tested when somebody extends it next year.
What a record represents
Now the schema. The structural question comes first, because the field list depends on the answer.
One record, or two layers
Most registers start flat. One row per thing. It works until the first shared component appears.
Consider a foundation model serving four use cases. A customer service assistant, an internal knowledge search, a marketing copy generator, and a triage tool that routes complaints. Same model, same API, same vendor. One flat row.
But the triage tool touches complaint outcomes, and the copy generator does not. The knowledge search reads internal documents and the assistant reads customer records. Four different risk positions, four affected populations, four control sets. A single row cannot hold four answers, so it holds an average, which is true of nothing.
The opposite failure is just as common. Register only use cases and you get a list the business can govern, and the technical team cannot act on, because the dependency telling you a vendor update affects eleven entries is nowhere in the record.
The structure that survives both is two layers. The use case is the parent. The model or component is the child. The relationship is many-to-many.
This matters more than it sounds. Moving from a flat register to a two-layer one later is a migration, not an edit, because every classification decision attached to a flat row has to be made again rather than mapped across.
The field set
Six groups. Anything that does not fall into one of them should be challenged before it goes in.
- Identity and ownership. Business owner, technical owner, accountable executive, lifecycle status. A team name in an owner field means nobody.
- Technical. Model or service, version, provider, deployment pattern, integration points, dependencies both ways. Versions and providers change underneath you, so you need a last-verified date rather than just a value. ISO 42001 A.4.2 requires the resources used to develop, deploy, operate and maintain AI systems to be identified and documented.
- Data. Inputs, training data lineage, sensitivity inherited from the existing classification scheme, retention, tenancy boundary.
- Purpose and population. What decision the system serves, who is affected, volume and frequency. Article 3(12) defines intended purpose by the provider’s specified context and conditions of use, and A.9.4 requires intended use to be documented, communicated and enforced, with use beyond it identified and controlled.
- Classification. Dimensional scores, resulting tier, rationale. Deliberately a placeholder here. Part two fills it.
- Governance. Applied control set, assessments completed, evidence pointers, reassessment trigger and date. This is what turns a description into something testable.
The field that does the work
Of all of those, the one most often omitted is the rationale attached to the classification, and it is the one I would keep if I could keep only one.
An outcome without its reasoning cannot be reviewed, inherited by a similar system, or challenged. It is an assertion wearing the clothes of a record.
The test is simple. Could somebody who was not in the room perform the judgement again from the record alone?
What to leave out
The instinct when designing a schema is to capture everything that might one day be useful. Resist it. Fields nobody maintains do more damage than fields that do not exist, because a reader cannot tell a stale value from a current one. A register with twelve reliable fields is worth more than one with forty of uncertain age.
One test for every proposed field. Does anything downstream consume it. If nothing does, leave it out and add it when something does.
Where this leaves you
Two things here cannot be retrofitted cheaply. The boundary, because changing it changes the population. And the schema, because changing it invalidates the work already done against the old one. Everything else in the series is recoverable. These two are not, which is why they come first.
The classification block in that field set is deliberately empty. Next week it gets filled, along with how much of it a machine can do for you, and exactly where that stops working.