Every prompt sent to a third-party model is a data-egress decision. The moment a clause from a contract, a patient record, or an unreleased product spec leaves your boundary and lands in someone else's inference stack, you have made a sovereignty call, whether or not anyone signed off on it. Most organizations are making that call hundreds of times a day without noticing.
The custody problem hiding inside the prompt box
Data custody is simply the question of who physically holds your data, who can read it, and under whose law and policy it sits at any given moment. It is not the same as ownership. You can own data outright and still lose custody of it the instant it is pasted into a tool you do not control. In the pre-LLM world, custody changes were deliberate: an integration, a contract, a data-sharing agreement. Each one passed through someone whose job was to ask where the data would end up. Generative AI collapsed that friction to a text box, removed the gatekeeper, and the gap between policy and behavior opened wide. The tooling is now faster than the governance meant to contain it, and that mismatch is the real exposure.
The numbers are blunt about how wide. In Cisco's 2024 Data Privacy Benchmark Study, 27% of organizations had banned generative-AI use at least temporarily, 69% feared harm to their legal and intellectual-property rights, and 68% worried that information entered could be disclosed publicly or to competitors. Yet in the same study, 48% admitted to entering non-public company information into GenAI tools and 45% had entered employee information. That is the self-inflicted control gap in one line: the fear is correct, and people are doing the feared thing anyway. Cisco also found that 91% believe data stored in their own country or region is inherently safer, which tells you residency is already an instinct, not just a regulation.
Why this is a board-level issue, not an IT footnote
The cost of getting custody wrong is no longer abstract. The global average data-breach cost reached $4.88M in 2024, up 10%, according to IBM's Cost of a Data Breach Report. The same report found that 40% of breaches involved data spread across multiple environments, that those breaches cost more than $5M on average and took 283 days to contain, and that more than a third involved shadow data. Unmanaged AI usage is a near-perfect engine for producing exactly that condition: copies of sensitive data scattered across environments nobody is tracking.
The threat picture compounds it. Verizon's 2024 Data Breach Investigations Report found that 68% of breaches involved a non-malicious human element, that vulnerability exploitation as an initial access vector surged about 180%, and that 15% of breaches involved a third party. AI does not invent these failure modes; it widens the surface for all three. And the incidents are mounting: Stanford HAI's 2025 AI Index Report recorded 233 AI-related incidents in 2024, a 56.4% increase over the prior year.
Then there is the law, which has put a price on the question. When personal data is transferred outside the EEA, the European Commission's rules on international transfers require that the protection travel with the data: a lawful transfer needs an adequacy decision or appropriate safeguards such as Standard Contractual Clauses or Binding Corporate Rules. Routing a prompt through a model hosted in another jurisdiction can be a transfer. And the penalties under Article 99 of the EU AI Act are sized to get a board's attention: up to EUR 35,000,000 or 7% of worldwide annual turnover for prohibited practices, up to EUR 15,000,000 or 3% for operator-obligation breaches, and up to EUR 7,500,000 or 1% for supplying incorrect or misleading information.
The solution space: three honest options
Keeping custody does not mean abandoning AI. It means choosing a deployment posture deliberately. In practice there are three:
- On-premise, self-hosted. Open-weight models run on infrastructure you own. Data never leaves your boundary. Maximum control, maximum responsibility.
- Private cloud and VPC isolation. Dedicated, contractually isolated environments in a provider's cloud, with no shared inference and no training on your inputs. A middle path that trades some sovereignty for managed scale.
- Contractual zero-retention tiers. Frontier models accessed through enterprise agreements that contractually prohibit retention and training on your data. The lightest lift, and the one that rests entirely on the contract and the vendor's enforcement of it.
The self-hosted option used to carry an asterisk that no longer applies: that open models were good enough only for toy work. That objection has largely expired. McKinsey reports that more than 50% of organizations now use open-source AI technologies, with the leading barriers being security and compliance at 56% and uncertainty about long-term support at 45%, according to its analysis in "Open source in the age of AI." Open is now the mainstream, not the fringe.
The tradeoffs, stated plainly
We will not pretend self-hosting is free or automatically safer. Two things are true at once. First, the capability gap has closed dramatically. On the Chatbot Arena Leaderboard, the gap between the top closed-weight model and the top open-weight model narrowed from 8.04% in early January 2024 to 1.70% by February 2025, per Stanford HAI's 2025 AI Index Report. For most enterprise work, the open option is now genuinely competitive.
Second, running your own AI carries real and recurring costs: security and compliance ownership, uncertain long-term support, ongoing maintenance, scarce talent, and hardware. The same McKinsey barriers that hold organizations back are the bills you take on when you self-host. And owning the stack does not, by itself, confer security. You also own the patching. That 180% surge in vulnerability exploitation in the Verizon data is a description of what happens to infrastructure that is stood up and then left alone. Self-hosting moves the risk; it does not delete it. The organizations that do it well treat the model as one more production system with an owner, a patch cadence, and an audit trail, not as a science project that ships once and is forgotten.
So the honest framing is conditional, not absolutist. On-premise or private deployment is worth it when data is regulated, genuinely sensitive, or residency-bound, when the workload is steady enough to justify owned capacity, and when leakage would be existential. It is the wrong call when the data is non-sensitive, when demand is bursty and unpredictable, or when the work genuinely depends on frontier capability you cannot yet reproduce in-house. A serious practice will tell you which case you are in, even when the answer is "use the hosted model and write a tight contract."
How to decide
The decision is not a vibe; it is a sequence. We work it in this order:
- Classify the data. What sensitivity tiers actually flow through the use case? Most AI risk lives in a small fraction of the data.
- Map regulatory and residency obligations. GDPR transfer rules, AI Act exposure, and sector rules decide what is even permissible before cost enters the conversation.
- Define the required capability. Be specific about what the model must do well, then test whether an open or private option already clears that bar.
- Weigh latency and sovereignty. Where the inference physically happens shapes both performance and your legal position.
- Total cost and operational maturity. Be candid about whether your team can run, patch, and govern what you stand up.
We anchor the governance around all of it to the NIST AI Risk Management Framework, the voluntary framework released on January 26, 2023, whose four functions, Govern, Map, Measure, and Manage, give a board a common language for AI risk that does not depend on any one vendor's roadmap.
Keep the intelligence. Keep custody.
We architect deployments so your data never has to leave your environment, with governance built to match the risk, not the hype. If data custody is on your agenda, let's have an executive conversation about where your AI should actually run.
Get in Touch