Ask any vendor what AI costs and you will get a number that sounds small: a fraction of a cent per request. The trouble with that number is the word per. A cloud AI bill is a meter, and meters do not stop. A local AI appliance is the opposite shape entirely — you pay once for a box you own, then a thin subscription to keep it governed and updated. The same workload, two completely different cost curves.
This is not a story about one being universally cheaper. It is about which shape fits your business — and, just as importantly, what each shape does to your data along the way. Because with cloud AI, the price is not the only thing that leaves the building.
The two cost shapes
Strip away the spreadsheets and there are really only two ways to pay for AI.
Cloud AI is a meter that never stops. You are billed per token, per seat, or both — every question, every month, for as long as you use it. There is no point at which the meter switches off. The price you pay this year is also not the price you are guaranteed next year: the provider sets it, and they can change the model behind it whenever they like. And every billable request is also a request that carried your data to a third party to be answered.
A local appliance is a one-time box you own. You buy the hardware once. After that, the marginal cost of answering one more question is essentially the electricity to run the box. On top of the hardware you carry a thin governance-and-updates subscription — what we call Cloud Shield — but that is a flat monthly base, not a usage meter. The data is answered where it sits; nothing crosses a border to get a reply.
| Local appliance | Cloud AI API | |
|---|---|---|
| Cost shape | One-time hardware + thin subscription | Per-token / per-seat, forever |
| Marginal cost per question | ≈ electricity | Billed every time |
| Where data lives | Your building | A third-party cloud |
| Who controls the model | You | The provider (can change) |
| Cost at scale | Flat — usage is free once owned | Grows with every user |
The middle two rows are the ones people forget when they reach for a calculator. Cost is not the only axis on this table — where your data lives and who controls the model are line items too, even though no invoice prints them.
The hidden costs of the cloud meter
The per-token price is the headline. The bill is bigger than the headline, and most of the extra lives off the rate card.
- Egress and integration. Moving your knowledge into and out of a provider's infrastructure has its own cost, and it tends to grow with you rather than shrink.
- Compliance overhead. Every cross-border transfer of personal data is a question you now have to answer — transfer mechanisms, sub-processor chains, residency clauses that get renegotiated when the legal ground shifts. That work is real, recurring, and rarely on anyone's AI budget. It is also exactly the category of work a local deployment can reduce — by removing the cross-border transfer that triggers much of it — which we unpack in sovereign AI, explained.
- Lock-in. Once your prompts, your integrations, and your team's habits are built around one provider's API, moving is expensive — which is leverage the provider holds, not you.
- Price changes. A meter you do not control is a budget line you cannot forecast. The rate can move, the model behind it can be swapped, and you absorb both.
None of these show up when you multiply tokens by a unit price. All of them show up on the year-end total.
Why local AI gets cheaper the more you use it
Here is the structural twist that flips the comparison. With cloud AI, every new user, every new use case, every busy month adds to the bill — usage and cost rise together, forever. The thing you most want (heavy adoption) is the thing that costs you most.
A local appliance inverts that. Once the box is paid for, an extra thousand questions cost you almost nothing. There is no per-token meter to feed, so internal usage is effectively free. The more your team and your customers lean on it, the better the economics look — because you are spreading a fixed cost over a growing pile of value instead of paying a toll on every interaction.
That is why the line crosses. For light, occasional use, the cloud's pay-as-you-go shape can be perfectly sensible — you are renting, and renting suits low usage. But for a business that uses AI daily, across a team, on knowledge it owns, the flat curve of an owned box eventually sits below the rising curve of a meter. We will not pretend to print you a guaranteed breakeven month — prices are too volatile in 2026 for that to be honest — but the direction is not in doubt: the heavier and more permanent your use, the more an owned box favours you.
What box, and what does it actually cost?
The hardware ranges from a paperback-sized mini PC to a multi-tenant workstation, and you pick the one that matches how big a model you need to run.
| Machine | Runs | Memory | Compute | Price |
|---|---|---|---|---|
| Minisforum AI X1 Pro | 7–14B + long context | up to 96 GB (upgradeable) | Ryzen AI 9 HX 370 · Radeon 890M · 50 TOPS | from €739 |
| GEEKOM A8 | 7–14B | up to 64 GB (upgradeable) | Ryzen 9 8945HS · Radeon 780M | from €759 |
| Minisforum MS-S1 MAX | 30–70B (clusters higher) | 128 GB unified | Ryzen AI Max+ 395 · Radeon 8060S | from €2,679 |
| Apple Mac Studio M3 Ultra | 30–100B+ | 96–256 GB unified | M3 Ultra · 819 GB/s bandwidth | from €4,799 |
For a single knowledge base and a 7–14B model — enough to answer customer questions from your documents — an entry mini PC starts around the low hundreds of euros (indicatively ~€490 at the floor). Step up to a 128 GB box that runs a 70B model and you are into the performance tier (indicatively ~€2,400+). Both numbers are indicative only; for live figures, see current prices on the hardware page, which is re-verified monthly. You can pick the box on the hardware page once you know which tier fits, and our mini-PC buying guide walks through the field in detail.
The point is that this is capital, paid once — not a subscription that compounds. The TCO question becomes: how many months of a cloud meter equal one box you keep for years?
The honest caveat: hardware is capital, and someone maintains it
It would be dishonest to sell the owned box as free of effort. It is not.
- Hardware is paid up front. A capital outlay is a different conversation with your finance team than a small monthly line item, even when the total over time is lower. For some businesses, the cash-flow shape of a subscription is genuinely the better fit — and that is a legitimate reason to choose it.
- You maintain it — or someone does for you. A box in your building needs updates, the occasional patch, and someone to keep the governance layer current. That is exactly what the thin Cloud Shield subscription is for: a flat monthly base that keeps the model updated and the trust boundary governed, without re-introducing a per-token meter. If you would rather not touch it at all, a consultant can run it for you — see our local AI hosting guide for SMEs.
So the real comparison is not "free box vs. expensive cloud." It is a one-time cost plus a thin, flat subscription versus a meter that grows with every user and carries your data abroad on every request. For the actual subscription tiers, see the pricing page — the base is a thin monthly figure, not a usage charge.
The cost you cannot put on an invoice
There is one more line that belongs in any honest TCO and never appears on either bill: control.
With the cloud meter, your knowledge leaves the building on every request, the model can change underneath you, and the price is set by someone else. With an owned, governed box, none of that is true — and the thing that makes that box genuinely safe to put in front of customers is not the hardware at all. It is the governance layer: the boundary between your internal knowledge and the public agent, which decides in advance what a customer can ever see. We argue that boundary, not the box, is the real product in the product is governance.
Local hosting does not guarantee compliance — nothing does — but it strengthens your DSGVO and EU AI Act posture by design, because the hardest questions (where is the data, who can change the model, can I prove it) are answered the same way: nothing left the building. That is not a number on a quote. For a German SME, it is often the deciding one.
Ready to do your own maths? Pick a box on the hardware page, check the thin monthly base on the pricing page, and try the live demo to see a governed agent answer on real knowledge — with nothing leaving the building.