Insights

Private AI

When private deployment is justified — and when it is not

Victory Harbour Operating Team · · 8 min read

We build private AI infrastructure. We also tell most organisations that ask for it that they do not need it.

That is not modesty. It is that “private AI” is usually the answer people arrive at before the question has been made precise, and the gap between the requirement and the solution is where a lot of money goes.

The conversation typically starts with a real concern — our data cannot go to a public model — and jumps immediately to the most expensive available response. Between those two points sit four or five options that are cheaper, faster and sufficient for most of the situations that prompt the question.

The four reasons that hold up

Regulatory obligation. A specific rule requires data to remain within a jurisdiction or within your control. Not a general sense that regulators would disapprove — an actual provision, in an actual instrument, that someone can cite.

Contractual commitment. You have told clients their data stays within defined boundaries. This is the most common legitimate reason and the one most often discovered late, when someone finally reads the master services agreement.

The data is the business. Not confidential — constitutive. A pricing model, a formulation, a proprietary dataset, a case archive that took a decade to accumulate. If the asset itself is the thing that walks out, control of the environment is a strategic question rather than a risk-management one.

Volume economics. At sustained high throughput, dedicated infrastructure can cost less per task than API pricing. This one is arithmetic, and it is the reason most likely to be asserted without being calculated. Do the calculation, and include the operating cost below.

If none of these apply, what is usually being expressed is discomfort. Discomfort is legitimate — it is often an accurate instinct about a real risk — but it should be converted into a specific requirement before it is converted into infrastructure.

“Discomfort should be converted into a specific requirement before it is converted into infrastructure.”

The reasons that do not survive examination

“We don't trust the cloud.” Your email, your accounting, your CRM and your customer records are already in it. If the concern is specific to AI, name what is different: usually it is a fear that inputs become training data, which is a contractual question with a contractual answer.

“It's more secure.” Only if you operate it well. A self-hosted model on an unpatched server with over-broad access is considerably less secure than a major provider's managed service. Private deployment moves the security burden onto you; it does not reduce it.

“We want control.” Control over what, exactly? Model selection, data residency, retention, audit logging and access scoping can each be obtained separately, and most of them without private infrastructure.

“We might need it later.” Build for the requirement you have. Migration between deployment models is real work, but it is far less work than operating infrastructure you did not need for two years.

The costs that get underestimated

The comparison people make is hardware or hosting against API spend. That is not the comparison.

Model operations. Someone has to keep it running, patch it, monitor it, handle capacity. This is a role, not a task.

Evaluation. Managed providers improve models continuously and invisibly. Self-hosted, you own the upgrade decision — which means you need a test suite that tells you whether a new version is better for your work, and you need to maintain it. Most organisations discover this after the first upgrade goes badly.

Capability lag. Frontier capability appears in managed services first. Depending on your use case, running a generation behind may be irrelevant or may be the whole ballgame.

Talent. The people who can run this well are scarce and expensive, and you are competing for them with organisations for whom it is the core product.

None of these are arguments against private deployment. They are arguments for including them in the comparison, because the ones who regret the decision are almost always the ones who compared hosting cost to API cost and stopped there.

The middle options most people skip

Between “public API” and “we run our own models” there is a range, and it is where most requirements are actually met:

Contractual terms — no training on your data, defined retention, deletion guarantees

Solves: The most common underlying concern, at essentially zero cost

Regional deployment — provider infrastructure in a specified jurisdiction

Solves: Most data residency requirements

Private networking — dedicated endpoints, no traffic over the public internet

Solves: Network exposure concerns

VPC / dedicated capacity — provider models inside your cloud boundary

Solves: Isolation without operating the models

Selective routing — sensitive work to a private model, everything else to managed services

Solves: The realistic answer for most organisations

That last row deserves emphasis. The premise that all AI work must go to one place is rarely examined. In practice a minority of tasks touch genuinely sensitive material. Routing by sensitivity gives you the protection where it is needed and the capability everywhere else — and it is what we most often end up building.

A decision sequence

  1. Name the specific requirement. Which rule, which contract clause, which asset. If it cannot be named, the requirement is not yet defined.
  2. Ask which work is actually affected. Usually a minority. Quantify it.
  3. Work up the ladder, not down. Start with contractual terms; move up only when the option genuinely fails to meet the named requirement.
  4. Cost the full picture — including operations, evaluation and the upgrade path, over three years.
  5. Then decide, and write down what would have to change for the decision to be revisited.

What we would tell you

If you have a named regulatory or contractual requirement, or your data is the asset itself, private deployment is justified and we will build it. In our own Operating Labs we run multi-model routing across three model families under continuous internal workloads with per-task cost tracking, and we have evaluated permission-scoped private retrieval on a 10,000-document internal corpus. Internal lab conditions, not client results.

If you cannot yet name the requirement, we will say so, and we will suggest the narrower option that meets the concern you actually have — even though it is the smaller engagement. A private platform built on an undefined requirement becomes an expensive system nobody can justify at the next budget review, and that outcome is worse for us than the smaller project.

Start with the workflow that matters

If the question is whether your data can safely be used with AI, that is worth answering precisely before it is answered expensively. This is the assessment the Private AI & Knowledge Platform engagement starts with.