A capable AI assistant no longer needs a data centre. In 2026 a mini PC the size of a paperback runs a useful large language model entirely on your own desk — no cloud account, no per-token meter, nothing leaving the building. The hard part is no longer whether you can self-host; it's which box to buy.
This guide compares the mini PCs and small workstations worth running a local AI core on, grouped by what they can actually do. If you just want the shortlist with current prices and where to buy, jump to the hardware page — it stays up to date. If you want to understand the trade-offs first, read on.
What actually makes a mini PC good for local AI
One number decides almost everything: memory. A language model has to fit in fast memory to run well, and the size of the model you can host is capped by how much memory the chip can address.
- A 7–14 billion parameter model, quantized to 4-bit, needs roughly 8–16 GB. Any modern mini PC with 32 GB handles it comfortably.
- A 70 billion parameter model needs around 40 GB — which is why the interesting boxes ship 128 GB of unified memory, of which up to 96 GB can be handed to the GPU.
Three things matter after capacity, in order: memory bandwidth (how fast tokens come out), the software stack (CUDA, Metal or ROCm/Vulkan), and only then raw compute. One thing that matters less than the marketing suggests: the NPU. The "50 TOPS" Copilot+ figure sells laptops, but local LLM inference runs on the integrated GPU, not the NPU. We unpack the silicon trade-offs in AMD vs Intel vs Apple vs NVIDIA for local LLMs, and how to size memory to a model in what hardware you need to run an LLM locally.
The entry tier: one knowledge base, 7–14B
This is where most small businesses should start. A single NPU/iGPU mini PC, 32–96 GB of RAM, one knowledge base, one public-facing agent. These are the boxes you can buy today and have running in an afternoon.
| Machine | Runs | Memory | Compute | Price |
|---|---|---|---|---|
| Minisforum AI X1 Pro | 7–14B + long context | up to 96 GB (upgradeable) | Ryzen AI 9 HX 370 · Radeon 890M · 50 TOPS | from €739 |
| GEEKOM A8 | 7–14B | up to 64 GB (upgradeable) | Ryzen 9 8945HS · Radeon 780M | from €759 |
| Minisforum UM890 Pro | 7–14B | up to 96 GB (upgradeable) | Ryzen 9 8945HS · Radeon 780M | from €489 |
| GMKtec EVO-X1 | 7–14B | 32 GB LPDDR5X-7500 | Ryzen AI 9 HX 370 · Radeon 890M · 50 TOPS | from €880 |
| Beelink SER9 Pro | 7–14B | 32 GB LPDDR5X-7500 | Ryzen AI 9 HX 370 · Radeon 890M · 50 TOPS | ~ €1,000 |
| ASUS NUC 14 Pro AI | 7–14B | up to 96 GB | Intel Core Ultra · Arc · 48 TOPS | ~ €1,250 |
A few notes on the field:
- The Minisforum AI X1 Pro is our default recommendation — it takes up to 96 GB of upgradeable memory, which gives the most headroom in the class for the price.
- The GEEKOM A8 is the dependable value pick, with a three-year warranty and EU fulfilment.
- The Minisforum UM890 Pro is the cheapest box that still qualifies — a fine choice for a single SME with one knowledge base.
- The GMKtec EVO-X1 has the highest memory bandwidth in the tier, which translates into more tokens per second.
- The Beelink SER9 Pro is the most turnkey — it ships ready to run a local model out of the box.
- The ASUS NUC 14 Pro AI is the pick for Intel/vPro shops that want fleet manageability.
Any of these runs a 7–14B model well — which is enough to answer customer questions from your documents, draft replies, and search your own knowledge. See the live specs and disclosed buy links on the hardware page.
The performance tier: run a 70B in one box
When you want several knowledge bases, or a single larger 70B-class model for sharper answers, you step up to a 128 GB unified-memory machine or a fast discrete GPU.
| Machine | Runs | Memory | Compute | Price |
|---|---|---|---|---|
| Minisforum MS-S1 MAX | 30–70B (clusters higher) | 128 GB unified | Ryzen AI Max+ 395 · Radeon 8060S | from €2,679 |
| Apple Mac Studio M3 Ultra | 30–100B+ | 96–256 GB unified | M3 Ultra · 819 GB/s bandwidth | from €4,799 |
| NVIDIA DGX Spark | 30–70B (infer to ~200B; ~405B linked) | 128 GB unified | GB10 Grace Blackwell · full CUDA | ~ €4,800 |
| RTX 5090 SFF Workstation | ~30B, very fast | 32 GB GDDR7 + system RAM | GeForce RTX 5090 · CUDA · 1,792 GB/s | from €3,999 |
| Beelink GTR9 Pro | 30–70B | 128 GB unified | Ryzen AI Max+ 395 · Radeon 8060S | ~ €4,000 |
| Framework Desktop | 30–70B | 128 GB unified | Ryzen AI Max+ 395 'Strix Halo' | ~ €2,400 |
Here the choice splits by software ecosystem. The Minisforum MS-S1 MAX and Beelink GTR9 Pro are AMD "Strix Halo" boxes that load a 70B in unified memory and can be clustered. The Apple Mac Studio M3 Ultra has by far the highest memory bandwidth, making it the fastest pure-inference box (though it runs via Metal, not CUDA). The NVIDIA DGX Spark trades raw speed for full CUDA and the strongest software stack. The RTX 5090 SFF is the fastest option for models that fit in its 32 GB of VRAM. We put these head to head in can a mini PC run a 70B model?.
For 70B-plus models served to many tenants at once, you move past mini PCs into workstation territory — the hardware page lists those too, on a quote basis.
How to choose, in one minute
- One knowledge base, tight budget → an entry mini PC. Start with the Minisforum AI X1 Pro or the GEEKOM A8.
- Several knowledge bases, or you want a 70B → a 128 GB performance box like the MS-S1 MAX, or a Mac Studio if silent, fast inference matters most.
- Many tenants or maximum throughput → a discrete-GPU workstation (talk to us for a quote).
Not sure which side of the line you're on? The honest answer is usually "start smaller than you think." A 14B model on an entry-tier box (from ~€490, indicative) surprises most people. You can always add a bigger box later — the knowledge and the governance move with you.
One caveat: prices move
AI hardware prices have been unusually volatile through 2026 thanks to the memory shortage — some boxes have moved by double digits in a single month. Treat any figure you read (including ours) as indicative, and confirm the live price before you order. We re-verify the hardware page monthly and stamp it with a date, and some of its links are affiliate links (clearly disclosed) — buying through them costs you nothing extra and helps keep the open core free.
The box is only half the product
A mini PC running a raw model is a confidential database that will read its own secrets aloud to anyone who asks. What makes a local AI safe to put in front of customers is the governance around it — the boundary between your internal knowledge and the public agent, and the loop that improves it. That, not the hardware, is the actual product; we make the case in the product is governance and explain the sovereignty model in sovereign AI, explained.
So pick the box that fits your knowledge — then put a governed core on it. If you want the whole picture, including the real cost versus cloud and compliance, start with our guide to local AI hosting for SMEs, or see the full hardware shortlist.