The AI Asset Register: Classification - Part two of four
Most AI registers use one classification field: high, medium, low. This piece proposes five dimensions instead, three that combine into a score and two that can override it outright, so the reasoning behind every tier stays visible.
Last part ended with a gap. The record schema had six field groups and one of them, the classification block, was left deliberately empty. This article fills it.
If you are arriving here first, the short version is that a register needs three decisions settled before it is built: what counts as an AI system, what a single record represents, and what classification is for. This is the third of those, and it is the one that everything downstream depends on, because classification is what decides which controls apply to which system.
One thing to say plainly before we start. What follows is a proposal. The frameworks tell you what has to be assessed. None of them tells you how to score it, and the gap between those two things is where every organisation ends up inventing something. This is my version of that invention, offered because it has held up, not because a standard requires it.
Why a single tier fails
Most registers carry one classification field. High, medium, low. Or tier one through four. It is the obvious design and it breaks in two predictable ways.
The first is that one label collapses distinctions that drive genuinely different controls. Two systems can both land in your top tier for entirely unrelated reasons. One handles special category data about a large population and makes recommendations a human reviews carefully. The other touches ordinary data about forty people and decides outcomes with no practical override. Both are high. They need almost nothing in common by way of treatment.
The second failure is slower and more expensive. A single label carries no information about why it was assigned. So when a new obligation arrives, and one always does, you cannot map the existing labels onto it. You have to go back to every system and assess it again from scratch, because the reasoning that produced the label was never recorded in a form anything else can read.
I have watched an organisation that classified its estate cleanly against the EU AI Act tiers discover that this told them almost nothing about the questions ISO 42001 asks. The Act wants to know whether a system falls in Annex III. A.5.4 and A.5.5 want to know how outputs may affect individual rights, autonomy and wellbeing, and what the broader societal effects might be. Same estate, different questions, and the existing labels contained none of the answers.
The single tier register is the one that gets rebuilt in year two.
Five dimensions
The alternative is to score a small number of dimensions that move independently, and let the tier fall out of them rather than being assigned directly.
Before the list, an honest accounting of what is borrowed and what is not.
The AI Act defines risk in Article 3(2) as the combination of the probability of harm occurring and the severity of that harm. NIST asks at MAP 5.1 for the likelihood and magnitude of each identified impact, drawing on expected use, comparable systems and public incident reports rather than internal estimate alone. MAP 1.1 wants intended purposes and the applicable legal landscape documented together. MAP 4.1 wants technology and legal risks mapped across components including third party data and software. ISO 42001 A.5.2 requires a documented impact assessment methodology that defines scope, frequency and how results inform control selection.
So the raw material is largely framework derived. The composition into five scored dimensions is mine.
- Data sensitivity and provenance. What the system sees, where that data came from, and whether the organisation had the right to use it that way. This dimension is usually already answered somewhere in the organisation, which matters later.
- Purpose and decision autonomy. What decision the output serves, and whether it informs, recommends or decides. Then the practical question from part one: whether a human can actually intervene, which MAP 2.2 frames as documenting the system’s knowledge limits alongside how its output is overseen.
- Impact. Three sub factors that should not be collapsed into one number: how many people are affected, how badly, and whether the harm can be undone. That third one needs flagging. Reversibility is not in the Act’s definition of risk, which stops at probability and severity. I score it anyway, because probability and severity together cannot distinguish between a wrong credit decision that a complaint process corrects in a fortnight and a wrong safeguarding referral that follows someone for years. If you disagree, drop it. But drop it deliberately rather than by inheriting a definition that never contemplated the question.
- Regulatory exposure. Which obligations attach, in which jurisdictions, and under which role. This is the dimension most likely to change while the system itself sits perfectly still.
- Technical provenance. Built in house, fine tuned, or bought. This determines what evidence you can obtain directly and what you will have to request from a supplier under A.10.3. Part one noted that IBM’s 2026 data shows provenance moving breach cost even where incident rates are similar across deployment types. That is the empirical case for scoring it rather than treating it as metadata.
What the dimensions do that a tier cannot
The tier still exists. You need something short that drives the control set, and nobody is going to look up five scores before deciding whether a system needs adversarial testing. The dimensions sit underneath it.
Resist the urge to publish a formula. How the five combine is a risk appetite decision, and a weighted average invented by whoever built the spreadsheet will quietly encode an appetite nobody agreed to.
What you should decide explicitly is which dimensions can force the top tier on their own. In most organisations at least two can. A system that decides rather than recommends, with no practical override, belongs in the top tier whatever the other four say. So does one whose harms cannot be undone. Name those overrides in the methodology, because otherwise an averaging rule will bury exactly the systems you built the model to catch.
The objection to all this is obvious and fair. It is more work than assigning one label. The answer is that the work happens once. A single tier has to be derived again every time a new framework arrives. Five visible dimensions get mapped to it, which is an afternoon rather than a project.
That is not an abstract claim. Regulatory exposure answers Article 6 and Annex III. Impact answers A.5.4 and A.5.5 and MAP 5.1. Purpose and autonomy answer MAP 1.1 and the Article 14 human oversight requirements when they bite. Technical provenance answers A.10.3 and MAP 4.1. Data sensitivity answers most of what your existing information security regime already asks. The dimensions were not designed backwards from those obligations, but they land on them, which is what you would expect given where the raw material came from.
A worked example
Take a CV screening tool used to shortlist candidates. Annex III point 4 covers employment and worker management, including recruitment and candidate selection, so this is a system with a clear regulatory anchor. Under the Digital Omnibus the obligations for this category now apply from 2 December 2027 rather than August 2026, which changes the deadline and nothing else.
Scores are illustrative. The point is the reasoning, not the numbers.
- Data sensitivity and provenance: high. CVs carry personal data, frequently reveal protected characteristics whether or not you asked for them, and the training data may be historical hiring decisions with whatever pattern those decisions contained.
- Purpose and decision autonomy: high. The recruiter nominally reviews the shortlist. In practice the tool decides, because a candidate filtered out at stage one is never seen. The override exists and is not exercised, which is precisely the condition part one warned about.
- Impact: population is large and turns over constantly. Severity is significant, since this is access to employment. Reversibility is poor, because a rejected candidate does not know they were filtered and has nothing to appeal.
- Regulatory exposure: high. This one looks arguable and then stops being arguable, which makes it the most instructive score in the example.
Article 6(3) lets a provider conclude that an Annex III system is not high risk where it does not pose a significant risk of harm to health, safety or fundamental rights, including by not materially influencing the outcome of decision making, and where one of four conditions is met. The first is that the system performs a narrow procedural task. Article 6(4) requires that assessment to be documented before the system is placed on the market.
So a tool that only extracts education history into database fields, or deduplicates applications, has a genuine case. Reach for that reasoning with a ranking tool, though, and you run into the subparagraph that follows the four conditions: an Annex III system is always high risk where it performs profiling of natural persons. Scoring candidates on inferred suitability is profiling. The derogation is not narrowly available here. It is unavailable.
That is the shape of a lot of classification work. The dimension looks like a judgement call, and the judgement dissolves once you read one subparagraph further. Which is why the rationale field matters more than the score. A record that says high tells you nothing. A record that says high, derogation considered and excluded by the profiling bar, is a decision somebody can check. - Technical provenance: bought. Which means the fairness testing evidence, the training data description and the performance breakdown by demographic group all have to come from the vendor, and that is a contractual question you would rather have discovered before signing.
Composed, this is a top tier system. But look at what the dimensions preserve that the tier throws away. The autonomy score tells you the first control to fix is the review step, not the model. The provenance score tells you the second is the contract. A tier of high tells you neither.
Inheritance, or how to get moving
An estate of two hundred systems is not going to get five dimensions scored by hand in a quarter. It does not have to.
Two of the five are largely answerable from things the organisation already holds.
Data sensitivity comes from the existing classification scheme, which most organisations maintain under ISO 27001 A.5.12, and the asset inventory under A.5.9. Where a system consumes a dataset already marked as confidential or special category, the AI system inherits that floor automatically.
Regulatory exposure is partly derivable too. Records of processing under GDPR Article 30 already say what personal data is processed for what purpose in what jurisdiction. Existing impact assessments say more.
Technical provenance is usually in procurement records and the asset register.
Run those joins and you have a provisional score on three dimensions across the whole estate on day one, generated from governance you have already paid for. That is a genuinely defensible starting position, and it is also the argument that gets the work funded, because it costs almost nothing.
Where the automation stops
Now the part that matters more than the automation.
Inherited data sensitivity sets a floor. It does not set the classification. It tells you the answer cannot be lower than a certain point. It says nothing about what the answer is.
The two dimensions that most strongly drive the tier, purpose and decision autonomy, and impact, cannot be derived from any existing register. They are facts about how a system is used, not about what it processes. No join produces them. No amount of tooling reaches them. A model reading ordinary customer records to decide who gets a debt referral is more consequential than one reading medical records to suggest article tags, and no data classification scheme in the world will tell you that.
So the rule is this. Auto classification produces a provisional tier. A provisional tier governs nothing until a person has confirmed it, and the act of confirming is where the rationale gets written.
A.5.2 and A.5.3 point the same way. What is required is a documented assessment process and documented results. An unconfirmed automated output is neither.
The failure mode here is worth naming precisely, because it is not the obvious one. The risk is not that automated classification is inaccurate. It is that a register full of unconfirmed provisional tiers looks finished. Every field is populated. Every system has a rating. Nobody goes looking for gaps in something that appears complete, and the unconfirmed entries sit there being treated as decisions for years. An incomplete register at least advertises its own incompleteness.
If you build inheritance, build the confirmation state alongside it in the same sprint. Provisional and confirmed have to be visibly different in the record, and your reporting has to count them separately.
Where this leaves you
The classification block from part one now has a shape. Five dimensions, individually visible, composing into a tier that drives controls, with a rationale field that makes the whole thing reviewable.
I would be interested to know where you would cut or add. The dimensions are a proposal and the weighting between them is a risk appetite question that no article can answer for you.
Next week, discovery. You now know how to classify a system. Part three is about finding the ones nobody told you about, which turns out to be the harder half.