← All posts
EU AI Act· · 8 min read

The AI Asset Register: Discovery - Part three of four

A boundary and a classification model are only as good as the population you run them against. Where to find the AI your organisation does not know it has, including the vendor defaults that connect your data to a model without anyone deciding that.

audio-thumbnail
AI Inventory Part 3
0:00
/116.15866666666666

If you have come to this article first, here is what it assumes. A register needs a boundary rule that decides what counts as an AI system, a schema that decides what a record represents, and a classification model that decides how a system is scored. Parts one and two covered those. This one covers a harder problem: finding the systems the boundary rule and the schema were never applied to, because nobody knew they existed.

A boundary and a classification model are only as good as the population you run them against. Get discovery wrong and the first two parts of this series become an excellent process for governing the twenty percent of your AI estate that happened to go through a project.

Here is why that twenty percent figure is not rhetorical. IBM's 2026 Cost of a Data Breach Report found that security incidents involving an organisation's own shadow AI, meaning AI adopted without going through any approval process, more than doubled in a year, from 20 per cent of incidents to 43 per cent. Those incidents cost more too, averaging USD 5.39 million against USD 4.63 million the year before. And 68 per cent of the breached organisations in that study had no governance in place to manage AI or detect shadow AI at all, a little over a third with nothing and the rest with something still being built.

That is the gap this article is about closing.

Mine before you ask

The instinct when starting discovery is to send a survey. Please list any AI tools your team uses. This fails for a reason that has nothing to do with honesty. It asks people to self identify against a definition they have not read, in a category many of them do not believe applies to them, with an implied compliance consequence attached to answering truthfully. You get a partial list from your most conscientious teams and silence from everyone else.

Do the opposite. Mine what you already have before you ask anyone anything, and arrive at the conversation with a draft rather than a blank form. A draft invites correction. A blank form invites silence.

Five sources are available before a single interview:

Procurement and contract records. Anything bought as software, including anything with an AI feature nobody flagged at purchase.

Cloud and API spend. Finds paid usage cleanly. Misses free tiers entirely, which is where a great deal of adoption starts.

Model registries, where they exist, such as MLflow or a cloud provider's model catalogue. Finds what was built. Says nothing about what was bought.

Code repositories. SDK imports and API calls to model providers. Finds built systems, misses purchased ones and no code integrations completely.

Existing impact assessments and security reviews. Anything that already went through a formal process, which by definition is not the population you are worried about, but it tells you what good documentation looks like when you find the gaps.

No single source sees the whole estate. That is not a flaw in the method. It is the reason the next two sections exist.

The path that has no gate

Here is the structural problem underneath all of this. Software procurement has a gate. Someone requisitions, someone approves, a record gets created as a side effect of the process working normally. AI adoption mostly does not go through that gate, and it is worth seeing exactly why.

Four paths, and only one of them reliably produces a register entry.

Built internally. Goes through a project, usually has a budget line, usually gets noticed.

Procured as AI. Somebody bought a model API or an AI platform knowingly. Also usually gets noticed, because the word AI appeared somewhere in the buying conversation.

Shipped into a product you already own. No procurement event at all. The purchase already happened, for a CRM or a project tool or a document platform, and the AI feature arrived later as an update. Nobody goes back and reapproves a product you already pay for just because it got a new capability.

Adopted directly by staff. A free tier, a browser extension, a personal account used for work. No procurement event, no project, frequently no expense trail.

Three of four paths produce nothing for a register to catch. That is the whole discovery problem in one sentence, and it is why the next two sections matter more than any tool you buy.

When the vendor decides for you

The third path deserves more attention than it usually gets, because it is not hypothetical and it is not slow moving. Three examples, each currently live, each illustrating a different version of the same problem.

Google's Workspace Intelligence, launched on 22 April 2026, gives Gemini a standing connection across Gmail, Drive, Docs, Calendar, Chat and Meet so that generative features can draw on an organisation's own content without anyone pasting it in. Google's own admin documentation is explicit that the default setting for each data source is on, and that a change to that setting can take up to 48 hours to propagate. An administrator who does not go looking for this setting has, by default, connected their organisation's email and documents to an AI feature they never separately approved.

Atlassian's policy is a cleaner illustration of a different mechanism: the same vendor shipping opposite defaults to different customers of the same product. From 17 August 2026, Atlassian began using Jira and Confluence content to train its AI models. On Free and Standard tiers, metadata collection cannot be turned off at all, and in app content collection is on by default with an opt out available. On Premium, metadata collection is likewise mandatory, though in app content defaults to off. Only on Enterprise are both categories off by default with a genuine opt out. So the answer to "does our data train Atlassian's models" depends entirely on a contract tier, which is exactly the kind of detail a governance team rarely holds and a procurement system rarely surfaces on its own.

OpenAI's ChatGPT products complete the pattern. Per OpenAI's own help documentation, connectors and apps are enabled by default on ChatGPT Business workspaces and disabled by default on Enterprise and Edu workspaces, where an administrator has to switch each one on deliberately. Two workspaces at the same organisation, on different plans, can have entirely different data exposure through the identical product, and nothing about the product name tells you which one you are looking at.

The pattern across all three is the same. The decision to connect organisational data to a model has already been made somewhere other than inside your organisation, and finding out which way it was made requires reading a settings page or a contract, not asking a colleague what tools they use.

Two consequences follow directly.

The sub processor point is easy to miss. Whichever model sits behind a vendor's AI feature is now part of your data's path, whether or not your register or your data processing agreement was updated to say so. The feature shipping is not the same event as your contract being updated to reflect it, and the two can be months apart.

And this cannot be a one time check. Defaults change with releases. What was off in your last review can be on today. Vendor release monitoring, a recheck at every contract renewal, and a standing AI question in vendor risk reassessment are the only durable answer, because a survey completed in March tells you nothing about a default flipped in June.

The tools people bring themselves

The fourth path is the one most people mean when they say shadow AI. Free tiers, browser extensions, personal accounts used for work convenience. No expense trail, no administrative record, no vendor decision to trace, because there was no vendor relationship at all in any formal sense.

Worth saying plainly what this usually is. Competent people routing around a process that is slower than their deadline. Treating that as a discipline problem produces a lot of stern memos and very little behaviour change, because the underlying incentive, get the work done faster, has not moved.

It is also not one problem. A summarisation tool reading public web pages and a coding assistant with write access to production repositories carry entirely different risk, and a discovery programme that treats every unsanctioned tool identically will spend its credibility flagging the harmless ones while the consequential ones sit unexamined in the same pile.

What you can actually detect

Six sources, each partial in a different direction.

SourceFindsMisses
Egress and DNS monitoringTraffic to known AI service endpointsAnything embedded inside an already approved SaaS domain
CASB or SSPMSanctioned and unsanctioned app usage, configuration driftTools used on personal devices off the corporate network
OAuth grant and SSO app reviewThird party apps granted access to corporate dataTools that never request an OAuth scope
Expense and card dataPaid individual and team subscriptionsFree tiers, which is where most adoption starts
Browser extension inventoryExtensions with page content accessAnything used outside the managed browser
Repository scanningSDK imports, model calls, exposed API keysPurchased systems and no code integrations

The table is not a shopping list. It is an argument. Every source is partial, and the gaps do not line up neatly with each other, which is exactly why reconciliation across sources matters more than any single tool. Buying the best egress monitoring available still leaves you blind to a free tier used on a personal laptop, and no amount of CASB coverage finds an SDK import buried in a repository nobody reviews.

One more thing worth saying, because it is easy to skip past. Several of these sources monitor individual behaviour, not just application traffic. A programme that deploys them without telling people creates a second governance problem in the process of investigating the first one, and the proportionality question deserves the same explicit answer as the boundary question from part one: what you are watching, and why.

Making reconciliation continuous

A one time discovery sweep produces a snapshot that is already stale by the time it is presented. What you actually want is a standing reconciliation loop: pull from the sources above, compare against the register, and surface only the differences.

Some of that loop is genuinely mechanical. Pulling data from each source, normalising identifiers so the same tool is not counted twice under two names, matching against existing register entries, and producing a list of deltas. That is reconciliation logic, and there is no reason a person needs to do it by hand every quarter.

What is not mechanical is everything downstream of the delta. Deciding whether a newly detected tool falls inside the boundary rule from part one is a judgement. Deciding what its classification should be is the judgement from part two. Both were argued in this series to require a person, and nothing about automating the detection step changes that argument. A tool that produces a clean list of deltas has not produced an assessment. It has produced a queue.

The failure mode worth naming directly: a reconciliation report that looks comprehensive creates false confidence precisely because it draws on several sources at once. Comprehensive coverage of six partial sources is still partial. The report tells you what changed among the things you can see. It cannot tell you what you still cannot see, and treating its completeness as evidence of the estate's completeness is the mistake this whole article has been arguing against.

Where detection stops being enough

Detection alone is a race you lose slowly. New tools appear faster than any endpoint list gets maintained, and a security team chasing signatures is always one release behind.

The register only stays accurate for as long as the sanctioned path is genuinely easier than the workaround. That means intake measured in days rather than months, an approved tool catalogue people can actually find and actually use, and a triage tier that does not put a browser extension through the same process as a production model deployment.

Read shadow AI adoption as a signal about the approval process rather than only about the people using it. An organisation with a lot of unsanctioned tools usually has a sanctioned path that is too slow to compete with, and no amount of detection tooling fixes that on its own.

Where this leaves you

You now have a boundary, a schema, a classification model, and a way to find what the first two were never applied to. Four parts of the register are built. What is missing is what keeps all of it true after the project team that built it moves on to something else, which is where part four picks up.

Najwan Hudaihed
Najwan Hudaihed

25+ years in IT audit and cybersecurity. Writes on auditing AI systems for technology risk, internal audit, and governance practitioners.