AI Managed Services

Run your own AI. Prove it behaved.

For companies that will not send their work to somebody else's model. We specify the hardware, deploy open models on it, connect them to your systems, and keep the whole thing running - with the evidence trail that makes the decision defensible to a customer, an auditor, or a regulator.

Read this before you buy anything

Running AI locally is almost never the cheaper option. A hosted API is subsidised, elastic, and someone else's problem to operate. If cost is your reason, the numbers will probably not support it, and we would rather tell you that now than after the invoice.

Local wins when you have a constraint: data that contractually or legally cannot leave your boundary, a customer who audits where their information is processed, a regulator who will ask, a network with no route out, or a volume so sustained that per-token pricing stops making sense. If one of those is true, the rest of this page is for you.

Sizing

Four sizes, described by what you need rather than what we would sell.

Sized by how many people are using it at the same moment and what they are doing, not by headcount. Hardware is quoted at build time against current pricing, because that market moves.

S

Pilot

5 to 10 at once

roughly 10-25 people on light use

One team, one use case

Hardware
Single-GPU workstation, 24 to 32GB VRAM
Models
One capable model, quantised for reliable tool calls
Stand-up
2 to 3 weeks, parts in hand
  • One document set indexed for retrieval
  • Chat interface your team can actually reach
  • Governance: the deployment is registered and every request is recorded

Proving the thing works on your data before anyone signs a bigger cheque.

M

Department

25 to 40 at once

roughly 25-75 people on light use

A department that depends on it daily

Hardware
Rack server, one or two datacentre GPUs
Models
A fast model for routine work and a capable one for the hard asks
Stand-up
4 weeks on used gear, 6 to 12 new
  • Retrieval across several systems, refreshed on a schedule
  • Tool calling into the systems you already run
  • Single sign-on, so access follows your directory
  • Governance: model inventory, usage evidence, and data-handling proof

The first deployment people would complain about losing.

L

Company

100+ at once

roughly 75-250 people on light use

Company-wide, with uptime expectations

Hardware
Redundant pair, multi-GPU, load balanced
Models
Multiple models served concurrently, routed by task
Stand-up
8 to 16 weeks, GPU-dependent
  • Failover, so a driver update does not take AI down with it
  • Per-team access control and usage accounting
  • Retrieval pipelines with owners and review cycles
  • Governance: full control coverage mapped to ISO 42001 and the EU AI Act

AI has become infrastructure and is now somebody's job to keep up.

XL

Regulated

Sized to the constraint

the environment sets this one, not the headcount

Air-gapped, multi-site, or under a regulator

Hardware
Quoted against your constraints, not a catalogue
Models
Pinned versions, provenance recorded, changes gated by review
Stand-up
Scoped per site
  • Fully disconnected operation where the environment demands it
  • Documented disaster recovery, tested rather than described
  • Evidence packaged for the audit you are actually facing
  • Governance: the whole programme, not just the AI

The network is the constraint, and 'send it to a vendor' was never an option.

Every size is a fixed-price build plus a monthly managed line. No per-seat licence on the models, because the models are yours.

How we size it

Headcount is the wrong question.

Two things decide what you need, and they move independently. A thirty-person team running agents all day needs more than a hundred-person firm asking the handbook a question twice a week - and if the first team stops when it stops, they need redundancy the second one does not.

Dial one

How much compute

Set by how many people are using it at the same moment, multiplied by how heavy each of those sessions is. The difference between the lightest and heaviest work is roughly fivefold.

  • Asking questionsabout 1 session per 15 people

    Looking something up in documents. Short prompts, short answers, and everyone does it at different moments.

  • Drafting and summarisingabout 1 session per 8 people

    Producing real output. Longer responses, held for longer, and steady through the working day.

  • Agents doing workabout 1 session per 3 people

    Multi-step tasks that read, decide and write back. One person can occupy a session for minutes at a time.

Dial two

How much redundancy

Set by what happens in the hour it is unavailable. This buys resilience rather than speed, and it is most of the price difference between a department build and a company one.

  • Inconvenient

    People go back to doing it by hand for an afternoon. One machine, with a spare part on the shelf.

  • Expensive

    Work queues up and somebody notices by lunchtime. A warm spare and a same-day swap.

  • Stops the business

    Customers feel it. Two machines, either one able to fail without anyone outside noticing.

The two questions we will ask on the first call

  1. How many people will be using it at the same moment, and what are they doing with it? That sets the compute.
  2. What happens if it is unavailable for an hour on a Tuesday? That sets the redundancy.

Most answers land on one of the four sizes. The ones that do not are exactly the deployments a headcount quote would have got wrong, in one direction or the other.

About those timelines

The accelerator is the long pole in every build. Everything else - chassis, CPUs, memory, storage, networking - is a stocked commodity that arrives in days. A current datacentre GPU ordered new can run eight to sixteen weeks or longer, and that window moves with demand we do not control.

So we quote two dates: when you can be running, and when the final configuration lands. Where it makes sense we start you on previous-generation hardware, which is available now and often at a fraction of list, then move the workload onto new gear when it arrives. You get to a working system in weeks instead of waiting a quarter to find out whether the thing was worth building.

In every size

What you are actually buying.

The hardware is the easy part. What makes a local deployment survive its first year is everything around it.

We build it and we keep it running

Hardware specified and procured, the stack installed, and a named person on the other end of the phone when a driver update breaks something at an awkward hour. Monitoring, patching and model refreshes are part of the monthly line, not a separate conversation.

Governance is in the box

Every deployment is registered, every request is recorded as signed evidence, and the model itself is inventoried with its version and licence. You went local to be able to prove something. This is the part that proves it.

Your data stays where it is

Retrieval runs against your systems in place. Nothing is copied to a vendor, and nothing leaves the boundary you drew. We can show you the record that says so.

Open weights, licence checked

We deploy models whose licence permits what you are actually doing with them, and we tell you which is which. Some popular open models carry usage conditions worth reading before you build on them.

How it learns your business

Training a model on your data is usually the wrong answer.

It is the request we get most often, and it is worth understanding why we will normally talk you out of it.

What training actually does

Fine-tuning teaches a model style, format and vocabulary. It is poor at teaching facts, and worse at teaching facts that change. Train it on this quarter's numbers and you get a model that states last quarter's with total confidence.

What we do instead

Retrieval over your documents, a memory of past conversations, and live lookups into the systems that hold the truth. The model reads the current answer at the moment it is asked, rather than remembering an old one.

Where training earns its place

One narrow job: teaching a smaller model to call your specific systems reliably. A light tune on your own naming and call patterns turns a model that is nearly right into one you can depend on. That is a real use, and it is the only one we will quote without argument.

Why this comes from Joopler

Local AI has more governance to do, not less.

Moving the model in-house does not remove the obligations. It removes the vendor who was quietly meeting some of them for you.

Nobody else is keeping the record

Record-keeping duties apply wherever the model runs. With a hosted provider there is at least a vendor log to point at. On your own hardware there is nothing at all unless something is recording it, and the reason you went local was to be able to prove something in the first place.

The claim needs evidence

"Our data never leaves the building" is the sentence that justifies the whole project, and it is the one a customer will ask you to demonstrate. Joopler turns it into a signed, independently checkable record rather than an assurance.

A local model that acts is an agent

The moment it reads a ticket and writes to a system, it is an agent operating unattended on your estate. It gets a register entry, an owner, declared limits, a trace of what it did, and a stop button that has been tested - the same as any agent you bought.

How agent governance works

Model provenance is a supply-chain question

Which weights, which version, which licence, and what any tuning was trained on. Nobody tracks this today, and it is the first thing an assessor will ask once local deployments are common enough to notice. We record it from day one.

And one for the house

The same stack, built for a home: a private assistant that answers in the room you are standing in, controls what is already on your network, and never sends a word of it anywhere. No subscription, no account, no microphone reporting to a platform. We build these too, and they are a genuinely good way to see what local AI feels like before committing a company to it.

Ask about a home build

Ready to see verifiable compliance?

Book a demo. Connect your stack. Share auditor-defensible evidence in days, not months.