In your cloud tenancy
Open-weight models running inside your own Azure, AWS or Google subscription, in a UK or EU region, under your existing security controls, network policy and logging.
For some organisations "we do not send that to a third party" is not a preference, it is the whole constraint. Private AI puts capable open models inside your own boundary - your cloud tenancy, your servers, or fully air-gapped - and keeps the data there.
These are the constraints that bring organisations to a private deployment.
The right pattern depends on your obligation, not on fashion. We will tell you when a public API with the right contract is genuinely sufficient.
Open-weight models running inside your own Azure, AWS or Google subscription, in a UK or EU region, under your existing security controls, network policy and logging.
Self-hosted inference on GPUs you own or rent, sized to your actual concurrency. The most economical option at sustained volume, and the only one for some estates.
Fully offline deployment with an update path that does not require a live connection - built for environments where connectivity is not permitted.
Grounding on your documents with citation, access-aware retrieval so a user only ever sees passages they are entitled to, and honest refusal when the answer is not there.
Adaptation to your domain, format and tone when prompting and retrieval have been exhausted - with an evaluation set that proves it helped.
Sensitive work stays private, general work uses a commercial API where that is cheaper or stronger, with the classification rule enforced in code rather than in a policy document.
The engagement runs with the people who do the work, not around them. Sessions are short, scheduled around delivery, and every stage ends with something you can act on.
You get a named consultant for the whole engagement - the person in the room is the person doing the work.
What exactly cannot leave, under which obligation, and who has to be satisfied. Frequently the constraint is narrower than the organisation assumes - and occasionally much wider.
Candidate open models evaluated on your own tasks, with concurrency, latency and cost modelled honestly against a commercial API baseline.
Infrastructure as code inside your environment, wired to your identity provider, logging and monitoring.
Runbooks, upgrade path and capacity plan - operated by your team, by us, or jointly.
Open models have closed much of the gap but not all of it. On the hardest reasoning tasks a frontier commercial model is still ahead, and we will show you the difference on your own work rather than argue about it.
Private is not automatically cheaper. Below a certain sustained volume, a commercial API with the right data-processing terms costs less and carries less operational burden. We model both.
Private is not automatically compliant, either. Running a model yourself removes the third-party transfer question and leaves every other obligation - lawful basis, transparency, retention, human oversight - exactly where it was.
The audit establishes what genuinely cannot leave your boundary - and what only feels like it cannot.