For a government agency or a regulated enterprise, the build vs buy AI decision inverts the usual answer, and it inverts before cost enters the conversation. Classify the data first. Then check what you must be able to reconstruct for an auditor. Then check what your contract vehicle can actually purchase. Only after those three do you compare price — and by then, price often isn’t deciding anything. If a workload touches data that cannot leave your perimeter, or produces decisions a regulator can demand you explain in two years, you’re building the control surface regardless of what the vendor comparison spreadsheet says.

Everything on page one of that search is written for a private-sector SaaS buyer with no residency constraint, no procurement code, and no regulator. Change those three inputs and the framework doesn’t just shift — it reverses.

The standard framework and the three things it quietly assumes

The usual advice is sound where it applies. Buy commodity, build differentiation, don’t reinvent infrastructure. Fine. But it rests on three assumptions that are invisible until they’re false.

One: your data can leave your perimeter. The entire SaaS calculus assumes you can send the workload somewhere else. Remove that and half the vendor market disappears from the comparison, not because the products are bad but because they can’t be deployed in a shape you’re permitted to run.

Two: procurement is a purchase order. In the private sector, “buy” means someone with a card and a budget line. In the public sector, “buy” means a competitive process, a contract vehicle with defined scope, and possibly a framework agreement that predates the product category you’re trying to acquire. Buying can genuinely be slower and more expensive than building, which is a sentence that makes no sense to a private-sector reader.

Three: nobody audits the decision later. Regulated buyers get asked to justify the choice, sometimes years afterwards, sometimes to someone hostile.

As context, the market really has moved toward buying. Menlo Ventures’ 2025 State of Generative AI in the Enterprise found 76% of enterprise AI use cases were purchased rather than built in-house, up from 53% the year before. That’s a big swing in twelve months and it reflects something real — ready-made tools reach production faster.

But that number is measured across enterprises generally. It tells you what the median buyer with no residency constraint decided. It’s a weather report from a different climate, and I’d be careful about reading your own decision out of it.

Before anything else: yes, I run a custom shop

 

Let me get the obvious objection out of the way, because if I don’t, you’ll be reading the rest of this waiting for the sales pitch.

Buying wins constantly in regulated contexts, and I’d say so to a client’s face. Identity and access management: buy it. Document storage: buy it. Transcription, OCR, translation: buy it, in a compliant deployment shape. Foundation models themselves: you are not training one, and any vendor suggesting you should for a typical agency workload is selling you a research project.

The commodity layer is enormous and getting bigger, and building any of it is waste dressed up as sovereignty. What follows is about the narrow band where the standard advice actively misleads — not a general argument that regulated buyers should build more.

The four constraints that override cost

 

Data residency and sovereignty

 

This is the one that gets discussed most and understood least. Everyone knows to ask where the data is stored. Fewer buyers trace all four places it actually goes.

At rest — the obvious one, and usually the only one in the vendor’s compliance PDF.

In transit — where does it route? A vendor with an EU data centre may still route through infrastructure elsewhere for a subset of operations.

In the model provider’s logs — this is the one buyers miss most often. If your vendor is a wrapper around a third-party model, your data reaches that third party. Whether it’s retained, for how long, in what jurisdiction, and whether it can be used for training are four separate questions with four separate answers, and the vendor may not know all of them without asking.

In the vendor’s support tooling — the quietest leak of all. When a support engineer reproduces your bug, what do they see, from where, and is that session logged? I’ve seen an otherwise airtight residency posture defeated by a support portal that mirrored customer records into a helpdesk in a different jurisdiction. Nobody was being careless. Nobody had asked.

Which brings up the distinction that matters more than any other in this article: the vendor’s compliance is not your deployment’s compliance. A vendor holding certifications tells you about their organisation and their reference architecture. It tells you very little about the configuration you’re running, the data you’re feeding it, and the obligations that attach to your use of it. Those obligations are yours. They do not transfer with a purchase order.

Get data residency for custom AI written into the contract as a testable requirement, not a representation. “Data will be processed and stored within [jurisdiction], including in provider logs, backups, and support access, and Customer may audit this annually” is a requirement. “Vendor maintains compliance with applicable data protection law” is a sentence.

Auditability — can you reconstruct why the system said that, two years later

 

Not “is the model explainable” in the academic sense. The practical question: eighteen months from now, someone challenges a decision. Can you produce what the system received, what version of the model and prompt produced the output, what context was retrieved, and what the human did with it?

That’s a logging and versioning discipline, and it’s mostly independent of which model you use. But it has to be designed in from the start, and most bought platforms don’t expose enough of their own internals to let you reconstruct it. You end up able to say “the system recommended X” and unable to say why — which in a regulated setting is roughly equivalent to having no answer.

Procurement structure — what your contract vehicle can actually buy

 

Sometimes the deciding constraint has nothing to do with technology. Consumption-based token pricing may not fit a fixed-price framework. A capability that changes materially every quarter may not fit a contract written around defined deliverables. An existing vehicle may let you buy development services next month while buying a subscription takes three quarters.

I’ve watched a build decision get made purely on this, and it was the correct call. The technology comparison favoured buying. The contract vehicle made buying an eighteen-month process and building a six-week one.

Regulatory exposure

 

For anyone deploying AI agents in the public sector inside the EU or selling into it, the ground moved a month ago and a lot of published guidance is now wrong.

Here’s the current state, and please verify against the regulation text and your own counsel rather than taking a blog post’s word for it. The AI Omnibus, Regulation (EU) 2026/1744, entered into force on 27 July 2026. It deferred the obligations for stand-alone high-risk systems under Annex III — which covers employment, essential services, education, law enforcement, and border management among others — from 2 August 2026 to 2 December 2027. High-risk AI embedded in regulated products under Annex I moved to 2 August 2028.

What did not move: the Article 50 transparency obligations still applied from 2 August 2026. Disclosure that a person is interacting with an AI system, labelling of AI-generated content, notification for emotion recognition and biometric categorisation. A limited grace period applies to machine-readable marking for generative systems already on the market.

On penalties, be precise, because the headline number gets misattached constantly. The €35 million or 7% of global annual turnover ceiling applies to prohibited practices under Article 5 — not to high-risk compliance failures. Breaching high-risk obligations tops out at €15 million or 3%. Supplying incorrect information to authorities: €7.5 million or 1%.

If you’ve been planning around an August 2026 high-risk deadline, you have sixteen more months than you thought. I’d use them rather than bank them — the deferral changed the date, not the requirements, and conformity assessment plus technical documentation plus registration is not a quarter of work.

The rule across all four: any single one of these at full strength moves the workload to build, and cost stops being the deciding variable. Not “cost matters less.” Cost stops deciding.

What “build” actually means now

 

Here’s where regulated buyers most often mis-price the decision, and it’s usually the thing that flips a stalled evaluation.

Building does not mean from scratch. Nobody is training a foundation model. Building means orchestration, evaluation, and control surfaces on top of bought foundations — the retrieval layer, the routing logic, the audit trail, the human review step, the guardrails, the integration into the systems where the work happens. The model is a bought component, often a commodity one, sometimes swappable.

That reframing usually cuts the estimate substantially, because buyers picture a research programme and the actual work is systems integration with an unusual dependency. Custom AI for regulated industries is mostly plumbing, and plumbing is a well-understood cost.

AI-assisted development has moved that floor further down over the past two years, which changes which workloads are worth building at all — but that’s its own article and I’ll leave it there.

The decision framework

 

Run it as a sequence, not a matrix. The order is doing the work.

  1. Classify the data. What’s the most sensitive category this workload touches, and can it leave your perimeter? If no, you’ve eliminated most of the vendor market. Stop comparing them.
  2. Check the audit requirement. Must you reconstruct a specific decision months later? If yes, can the candidate expose enough internals to let you? Most can’t. Ask for a worked example, not an assurance.
  3. Check the contract vehicle. What can you actually buy, in what pricing shape, in what timeframe? This eliminates options that pass steps 1 and 2.
  4. Now compare cost — across whatever survived. Include exit cost. Include the integration work, which exists in both directions.

Most regulated evaluations I see run this backwards: cost comparison first, then a compliance review that invalidates the winner. Running it in order is faster even though it feels slower, because it never spends effort evaluating options that were disqualified from the beginning.

The specification work behind steps 1 through 3 is where the decision is genuinely made, and it’s cheap relative to what it prevents. If you’re mid-evaluation and the residency or audit answer isn’t settled yet, that’s the point at which an outside read on the specification is worth more than another vendor demo — not because you’re ready to buy anything, but because a requirement you discover after selection is the expensive kind.

The procurement trap nobody warns you about

Buying a platform whose exit cost is unknowable.

Three questions to ask before signing, all of which should be answerable in writing:

Model lock-in. If the underlying model is deprecated or repriced, what happens? Who absorbs it? Can you substitute?

Prompt and tuning portability. Your prompts, retrieval configuration, and any tuning represent real accumulated work. Can you export them in a form that means something outside the platform, or do they only exist as rows in the vendor’s database?

Data export in a usable format. Not “we provide a data export.” In what format, including what — conversation history, retrieved context, decision logs, the audit trail you’ll need after you’ve left? A CSV of final outputs with no lineage is technically an export and practically nothing.

If you can’t get clear answers, price the ambiguity as a risk rather than assuming it’s small. The five-year TCO your procurement team wants is unknowable in its token-cost component — but it’s boundable. Model a ceiling from your worst-case volume at current list pricing, then negotiate a cap. A bounded worst case is a defensible number for a procurement file. A confident point estimate for 2031 is not.

Where I disagree with the standard advice

 

“Buy the commodity, build the differentiator.” Correct in the private sector. Misleading in government, and I think it misleads in a specific and expensive way.

In a commercial business, the differentiator is the product — the thing customers choose you for. Compliance is overhead. So the advice cleanly separates: build the special part, buy the boring part.

In government and heavily regulated enterprise, the differentiator frequently is the compliance surface. The audit trail, the residency architecture, the review workflow, the retention rules, the way an appeal gets handled. That’s the boring part by private-sector standards, and it’s the part your obligations attach to.

It’s also the part vendors will not customise for one buyer. A platform serving four hundred customers will not rebuild its logging to satisfy your regulator, and it shouldn’t — that’s the business model working as designed. So the standard advice tells regulated buyers to buy exactly the component they most need to control.

Invert it for your context. Buy the capability. Build the accountability. The model is the commodity now; what you do around it is the part with your name on it.

Frequently asked questions

Can government agencies use commercial AI models at all? Usually yes, in the right deployment shape — private endpoints, VPC deployment, or a regional instance with contractual guarantees on logging and training use. The constraint is rarely the model itself; it’s the default multi-tenant configuration. Ask vendors what deployment shapes they support before eliminating them, and get the logging answer in writing.

Does on-premise AI mean worse performance? Usually somewhat, and it’s better to say so plainly. On-premise or air-gapped inference typically means a smaller open-weight model and more engineering around it. For well-scoped tasks — classification, extraction, retrieval-grounded answering — the gap is often narrower than expected. For open-ended reasoning it’s wider. Scope the task before assuming the tradeoff is unacceptable.

What does the EU AI Act require from a custom AI system? It depends on classification. Article 50 transparency duties applied from 2 August 2026. Stand-alone high-risk obligations under Annex III — risk management, data governance, technical documentation, human oversight, conformity assessment, registration — were deferred to 2 December 2027 by Regulation (EU) 2026/1744. Building custom doesn’t reduce obligations; it usually makes meeting them easier because you control the internals. Confirm classification with counsel.

Is it cheaper to build or buy AI agents? Over three years, buying is usually cheaper for commodity workloads and building is often cheaper for anything with heavy compliance, integration, or customisation requirements — mainly because “buy” in those contexts still carries substantial integration cost that comparisons omit. In regulated settings, cost frequently isn’t the deciding variable anyway; residency and auditability are.

How do you write data residency requirements into an AI contract? Make them testable and specific. Name the jurisdiction, and cover all four surfaces: storage at rest, transit routing, model-provider logging and training use, and vendor support access. Add an audit right and a breach remedy. Reject compliance representations that describe the vendor’s certifications rather than your deployment’s behaviour.

Receive A Complimentary Consultation

Book Now