AI Services

Capable AI, without your data leaving

For some organisations "we do not send that to a third party" is not a preference, it is the whole constraint. Private AI puts capable open models inside your own boundary - your cloud tenancy, your servers, or fully air-gapped - and keeps the data there.

  • UK / EU data residency
  • No training on your data
  • Runs without internet where required
Private AI in practice
When private is the requirement

The cases where public APIs are simply not available to you

These are the constraints that bring organisations to a private deployment.

  • Special category or clinical data that cannot be processed by a third-party model
  • Contractual or client obligations that prohibit external processing
  • Classified, sensitive or air-gapped environments with no route to the internet
  • Legal privilege or commercial confidentiality that survives no exceptions
  • Volumes at which per-token pricing has become the dominant line in the budget
What we deploy

Deployment patterns, chosen to fit the constraint

The right pattern depends on your obligation, not on fashion. We will tell you when a public API with the right contract is genuinely sufficient.

In your cloud tenancy

Open-weight models running inside your own Azure, AWS or Google subscription, in a UK or EU region, under your existing security controls, network policy and logging.

On your own hardware

Self-hosted inference on GPUs you own or rent, sized to your actual concurrency. The most economical option at sustained volume, and the only one for some estates.

Air-gapped

Fully offline deployment with an update path that does not require a live connection - built for environments where connectivity is not permitted.

Retrieval over your own corpus

Grounding on your documents with citation, access-aware retrieval so a user only ever sees passages they are entitled to, and honest refusal when the answer is not there.

Fine-tuning where it earns it

Adaptation to your domain, format and tone when prompting and retrieval have been exhausted - with an evaluation set that proves it helped.

Hybrid routing

Sensitive work stays private, general work uses a commercial API where that is cheaper or stronger, with the classification rule enforced in code rather than in a policy document.

In practice

What working with us looks like

The engagement runs with the people who do the work, not around them. Sessions are short, scheduled around delivery, and every stage ends with something you can act on.

You get a named consultant for the whole engagement - the person in the room is the person doing the work.

A laptop on a quiet desk
Photo: Unsplash
How it runs

Constraint first, hardware last

1

Establish the real constraint

What exactly cannot leave, under which obligation, and who has to be satisfied. Frequently the constraint is narrower than the organisation assumes - and occasionally much wider.

  • Obligation traced to the actual clause or regulation
  • Sign-off from DPO, security and the business
2

Model and sizing

Candidate open models evaluated on your own tasks, with concurrency, latency and cost modelled honestly against a commercial API baseline.

  • Benchmarked on your work, not public leaderboards
  • Three-year total cost compared to the API option
3

Deploy and integrate

Infrastructure as code inside your environment, wired to your identity provider, logging and monitoring.

  • Reproducible deployment your platform team can rebuild
  • Access, retention and logging aligned to your policy
4

Operate or hand over

Runbooks, upgrade path and capacity plan - operated by your team, by us, or jointly.

  • Model upgrade path with regression testing
  • Capacity plan tied to measured demand
Deliverables

What you leave with

What you get

  • Private inference running in your chosen environment
  • Model selection evidenced on your own tasks
  • Infrastructure as code, reproducible by your team
  • Access-aware retrieval over your own documents
  • Cost model compared honestly against commercial APIs
  • Security, logging and retention aligned to your policy
  • Runbooks, upgrade path and capacity plan

The honest trade-offs

Open models have closed much of the gap but not all of it. On the hardest reasoning tasks a frontier commercial model is still ahead, and we will show you the difference on your own work rather than argue about it.

Private is not automatically cheaper. Below a certain sustained volume, a commercial API with the right data-processing terms costs less and carries less operational burden. We model both.

Private is not automatically compliant, either. Running a model yourself removes the third-party transfer question and leaves every other obligation - lawful basis, transparency, retention, human oversight - exactly where it was.

Questions

Questions about private ai

Whichever open-weight family evaluates best on your tasks and fits your hardware - the leading options change every few months, which is precisely why we benchmark at the time rather than committing in advance.
It depends on model size and concurrency, and it is often less than expected for document and drafting workloads. We size it from your measured demand and give you the rented-versus-owned comparison.
The deployment can be built to meet your accreditation requirements and we produce the evidence, but the certification is awarded to you, not to us. We have delivered into environments with formal assurance regimes and design for that scrutiny from the start.
It resolves the international transfer and third-party processing questions. Lawful basis, transparency, data minimisation, retention and human oversight still apply, and our governance work covers those.

Test the constraint before you buy hardware

The audit establishes what genuinely cannot leave your boundary - and what only feels like it cannot.