Two years ago, nobody asked our clients where their AI outputs came from. In August 2026, three of our last five enterprise deals included an AI questionnaire in the procurement pack - model inventory, data flows, human oversight, incident process. The teams that can answer in a week win the deal. The teams that cannot spend a quarter reconstructing history from Slack threads.
Why AI governance became urgent this year
Three forces converged. The EU AI Act's obligations for high-risk and general-purpose systems moved from published text to enforceable practice, so any product touching EU users now needs documentation rather than intentions. Enterprise procurement caught up, and AI questionnaires became as routine as security ones. And - quietly the biggest driver - AI stopped being one chatbot and became forty agents wired into CRM, billing, HR and support, each one an unlogged path to production data.
None of that requires a compliance department. It requires the same discipline you already apply to change management, applied to models.
Step one: an honest AI inventory
Every engagement starts the same way, and it is never glamorous: we list every place a model is called. Vendor features count. The marketing team's Copilot licence counts. The support macro that quietly calls an API counts. On a typical 300-person company we find between 18 and 40 AI touchpoints; leadership usually estimates six.
For each entry we capture: owner, purpose, model and provider, data sent, data retained, who sees the output, and what happens when it is wrong. That last column is the one that changes the conversation.
You cannot govern what you have not counted. The inventory is not paperwork - it is the first time most leadership teams see the actual size of their AI surface.
Classify by risk, not by hype
We sort every entry into three practical tiers rather than arguing about regulatory categories:
- Assistive. A human reads the output before anything happens. Drafting, summarising, research. Light controls, log the prompts, move on.
- Operational. The model writes to a system of record - creates a ticket, updates a CRM field, routes a case. Needs evaluation, logging and a rollback path.
- Consequential. The output materially affects a person: credit, hiring, pricing, medical or safety context. Human decision-maker in the loop, documented evaluation, retention of every decision record. No exceptions.
Most organisations discover one or two consequential systems they had filed as "just an automation". That reclassification is usually the highest-value hour of the whole project.
The five-layer control stack

1. Policy that fits on one page
What data may go to which providers, what always requires human sign-off, and who to call when something goes wrong. If your AI policy is twelve pages, nobody has read it and it protects nothing.
2. Access and data boundaries
Agents get scoped, short-lived credentials to the specific systems their job needs - never a shared admin token. Retrieval is filtered by the requesting user's permissions, so an AI assistant cannot become a permission-bypass machine. This is the single most common defect we find in home-built copilots.
3. Evaluation before and after launch
A held-out set of real cases with known-good answers, run on every prompt or model change. Ten minutes of CI beats a week of anecdotes about whether the new model is "better".
4. Observability and logging
Prompt, retrieved context, model version, output, tool calls, cost, latency, and the human decision that followed - stored with a retention policy that matches your data rules. This is what turns "we think it behaved" into evidence.
5. Human review where it counts
Not a rubber-stamp checkbox: a named role, a queue, a service level, and the authority to reject. Review that cannot say no is theatre.
Evaluation: proving the thing works
The teams that struggle are the ones treating evaluation as a launch gate rather than a habit. We build a small golden set - 50 to 200 real, messy cases from the client's own history - and score every release against it for accuracy, refusal behaviour, tone and safety. Then we sample live traffic weekly and add new failure cases to the set. After six months the golden set is the most valuable asset in the project; it encodes everything the organisation has learned about where its models break.
Cost and latency belong on the same dashboard. A model that is 3% more accurate and 4x more expensive is a business decision, not an engineering one, and it should be visible to the person who owns the budget. Our AI and automation practice ships this dashboard as part of every build.
The audit trail auditors actually want
Having sat through several of these now, the questions are remarkably consistent. Be able to produce, within a day:
- 01System inventoryEvery AI system, its owner, its purpose and its risk tier - dated and version-controlled.
- 02Data flow recordWhat personal data reaches which provider, under what contract, retained for how long, in which region.
- 03Evaluation evidenceThe test set, the scores at launch, and the scores at each material change since.
- 04Human oversight designWho reviews what, with what authority, and evidence that reviews actually happen.
- 05Incident logWhat went wrong, who noticed, what changed. An empty log is less credible than an honest one.
- 06Change historyModel versions, prompt versions, and the approvals attached to each.
If your data protection posture is still catching up, pair this with the groundwork in our cybersecurity practice - the two programmes share most of their evidence.
A realistic 90-day rollout
Weeks 1-3: inventory, risk tiering, one-page policy signed by an executive who will actually enforce it.
Weeks 4-7: credentials and data boundaries tightened on the operational and consequential systems. Logging turned on everywhere, even where it is only assistive.
Weeks 8-11: golden sets built for the top three systems, evaluation wired into CI, dashboards live for accuracy, cost and latency.
Week 12: a dry-run audit with somebody internal playing the assessor. Whatever you cannot answer in that room is your Q4 backlog.
Ninety days is enough for a mid-market organisation. It is not enough if you start it the week a customer's questionnaire lands.
Where RanWebs fits
We run this as a fixed-scope engagement for mid-market and enterprise teams across the US, UK, EU and APAC: inventory and risk tiering in the first three weeks, then the control stack built into your existing pipelines rather than bolted alongside them. You keep the dashboards, the golden sets and the evidence pack; we hand over and step out. It pairs naturally with our AI agents and copilot development and custom software work.
For the integration side of the same problem, read our companion piece on MCP in the enterprise, and on the delivery side, vibe coding without wrecking production. First call is free and goes straight to a senior consultant - email info@ranwebs.com or use the contact page.

