CryptoGPUs
Local AI

Best computer for local AI: DGX Spark vs Mac Studio vs Strix Halo

NVIDIA DGX Spark, Mac Studio M5 Ultra and Ryzen AI Max+ 395 compared on memory, bandwidth, real benchmarks and price, plus when a GPU tower is the better buy.

Three unbranded compact desktop computers on a grey concrete desk in soft light with a warm orange glow

Key takeaways

  • Memory capacity decides which models fit; memory bandwidth sets the ceiling on tokens per second.
  • A 128GB Ryzen AI Max+ 395 box gives the most memory per dollar and roughly matches the DGX Spark on generation, but reads prompts about half as fast.
  • The Mac Studio M5 Ultra's 1.2 TB/s makes it the fastest small box for generation, with no CUDA and 96GB in the $5,499 base model.
  • Buy a DGX Spark for CUDA and clustering, not value: the 128GB model now sells for over $6,000.
  • A tower with an RTX PRO 6000 is far faster for models that fit in its 96GB, at several times the cost and power.

If you want the most model for the money, the best computer for local AI in late 2026 is a 128GB Ryzen AI Max+ 395 box such as the Framework Desktop: it holds 120B-class models for about half the price of a 128GB DGX Spark. Pick the Mac Studio M5 Ultra if you want the fastest single-user generation in a quiet box and can live without CUDA, and the DGX Spark only if you need CUDA, faster prompt processing than AMD offers, or plan to cluster. A tower with a big NVIDIA card still beats all three on speed, but only for models that fit in its VRAM, and at several times the price.

Capacity decides what fits, bandwidth decides how fast

Memory capacity decides which models load at all. At 4-bit quantization, weights take roughly 0.5 to 0.6GB per billion parameters, so a dense 70B model needs about 40GB before you add the KV cache for your context and the operating system's share. OpenAI's gpt-oss-120b in its native MXFP4 format is a 59GiB file in the llama.cpp benchmark logs. That is why 128GB is the number to aim for if you want 100B-plus models with room for long conversations.

Memory bandwidth decides generation speed. To produce each token, the GPU reads every active weight from memory once, so the ceiling is roughly bandwidth divided by the size of the active weights. One tester measured about 215 GB/s of real GPU bandwidth on the Ryzen AI Max+ 395; divide by a 40GB 70B model and you get a ceiling a little over 5 tokens per second. The measured result in the same Framework community test series was 5.0. Fine for batch jobs, painful for chat.

Mixture-of-experts (MoE) models are the reason these boxes make sense. gpt-oss-120b has 117B parameters but only 5.1B active per token, according to OpenAI's model card, so it needs the memory of a big model and generates at the speed of a small one. On the same Ryzen AI Max+ 395, a 30B MoE model with 3B active parameters generated 72 tokens per second, while a dense 24B model managed 14.

Prompt processing is a separate bottleneck

Before the first token appears, the machine has to read your whole prompt. This step, called prefill, is limited by compute rather than bandwidth, which is where NVIDIA's tensor cores earn their keep. In an October 2025 owner comparison running the same llama.cpp command on gpt-oss-120b, a DGX Spark processed 1,816 prompt tokens per second against 1,000 on a GMKtec EVO-X2, while generation was nearly tied at 44.7 and 47.5 tokens per second. With 32K tokens already in context, the AMD box dropped to 349 tokens per second of prefill.

That gap is easy to feel. At 1,000 tokens per second, a 30,000-token agent context takes half a minute before anything comes back; at 2,400 it takes about 13 seconds. If you use coding agents or paste long documents, weigh prefill as heavily as generation speed.

The options side by side

MachineMemoryBandwidthPower (spec)SoftwareOur price
NVIDIA DGX Spark 128GB128GB LPDDR5x unified273 GB/s140W chip, 240W supplyDGX OS (Ubuntu on Arm), CUDA$6,950
NVIDIA DGX Spark 64GB64GB LPDDR5x unified273 GB/s140W chip, 240W supplyDGX OS (Ubuntu on Arm), CUDA$4,999
Mac Studio M5 Ultra96GB base, 256GB or 512GB options1.2 TB/s480W max continuousmacOS, MLX, Metal$5,499
Ryzen AI Max+ 395 (Framework Desktop, GMKtec EVO-X2)128GB LPDDR5X-8000; 96GB VRAM in BIOS, about 120GB on Linux256 GB/s (about 215 measured)120W sustained, 140W boostWindows or Linux, ROCm, Vulkan$3,449 / $3,499
Tower with RTX PRO 6000 (CG AI Pro 96)96GB GDDR7 ECC VRAM plus system RAM1,792 GB/s600W for the card aloneCUDA, Windows or Linux$21,499

Prices are ours at the time of writing. Bandwidth and power figures are manufacturer specs except the measured AMD number.

NVIDIA DGX Spark: CUDA and fast prefill at a rising price

The DGX Spark pairs NVIDIA's GB10 chip, a Blackwell GPU joined to a 20-core Arm CPU, with 128GB of LPDDR5x on a 256-bit bus. NVIDIA quotes 273 GB/s of bandwidth and up to 1 petaFLOP at FP4, rates the 128GB version for inference on models up to 200B parameters, and builds in a ConnectX-7 NIC at 200 Gbps so up to four units can be linked.

Bandwidth is its weak spot. 273 GB/s is close to AMD's chip, so generation on big models lands in the same range. The llama.cpp maintainers' February 2026 numbers show gpt-oss-120b at 58.7 tokens per second of generation and 2,444 of prefill, falling to 42.8 and 1,567 with 32K tokens of context. That generation figure is about 30 percent better than the owner above measured a week after launch, so software updates matter here. On dense models it falls well behind the Mac, as the next section shows.

Price has moved against it. The Spark launched at $3,999 in October 2025, NVIDIA raised it to $4,699 in February 2026 citing memory supply, and Wccftech reports that the 128GB model now sells for over $6,000, while a new 64GB version arrives through NVIDIA's partners (Acer, ASUS, Dell, Gigabyte, HP and MSI) at $4,999. We list the DGX Spark 128GB at $6,950 and the 64GB model at $4,999 at the time of writing.

Our view: buy the Spark for CUDA, not for value. It puts NVIDIA's full software stack and 128GB of memory in a 1.2kg box, its prefill keeps long agent sessions responsive, and clustering two units over the built-in NIC takes a single cable. The 64GB version is harder to recommend: NVIDIA rates it for models up to about 100B parameters, and the same $4,999 buys a 128GB AMD system. It also runs DGX OS, an Ubuntu build on Arm, so Windows and x86-only software are off the table.

Mac Studio M5 Ultra: the fastest generation in a small box

Apple announced the Mac Studio with M5 Ultra on 25 August 2026, with deliveries from 22 September. The number that matters for local AI is 1.2 TB/s of memory bandwidth, 50 percent more than the M3 Ultra and about 4.4 times the Spark. Apple also added Neural Accelerators to every GPU core, aimed at the Mac's old weakness, and claims up to 4x faster LLM prompt processing in LM Studio than the M3 Ultra.

Independent testing backs the bandwidth arithmetic. In the Tom's Hardware review, the M5 Ultra generated tokens on a dense 27B model at about twice the rate of an M4 Max and almost four times the DGX Spark, and its prompt processing also beat the Spark. A 4.4x bandwidth advantage turning into a near-4x speed advantage is exactly what the formula above predicts.

The catch is memory per dollar. The $5,499 base model has a 30-core CPU, a 64-core GPU and 96GB, less than the 128GB boxes. Going to 256GB adds $4,000, and the 512GB option arrives in late October with the 80-core GPU only. Memory is soldered and macOS needs its own share, so size up rather than down. There is no CUDA: you work in MLX, llama.cpp's Metal backend, LM Studio or Ollama, which covers most LLM inference but not every image, video or training tool. Our Mac Studio M5 Ultra is $5,499 at the time of writing.

Ryzen AI Max+ 395: the most memory per dollar

AMD's Ryzen AI Max+ 395, codenamed Strix Halo, combines 16 Zen 5 cores with a 40-CU Radeon 8060S GPU and up to 128GB of LPDDR5X-8000 on a 256-bit bus. That works out to 256 GB/s on paper and about 215 GB/s measured. It is an ordinary x86 PC, so Windows and any Linux distribution work.

Two details matter before you buy. First, the GPU does not get all 128GB by default. AMD's own cluster guide notes that the BIOS caps dedicated VRAM at 96GB on a 128GB Framework Desktop, then shows a Linux kernel parameter that raises the GPU's share to 120GB. For models over 90GB, plan on Linux. Second, prefill is the soft spot: generation on gpt-oss-120b is roughly level with the Spark, but prompt processing runs at about half the speed, and the AMD chip loses more of it as context grows. AMD's May 2026 tests, as summarised by AIMultiple, claim a 4 to 14 percent token throughput lead over the Spark, but those are vendor figures measured with a 100-token context.

Which box? The Framework Desktop 128GB ($3,449 at the time of writing) is the one we'd put on a desk: a 4.5-litre case, a standard 400W FlexATX supply, a 120mm fan, 5Gbit Ethernet and a PCIe x4 slot on the board for a faster NIC later. The GMKtec EVO-X2 128GB ($3,499) runs the same chip at the same 120W sustained, 140W boost in a 193 x 186 x 77mm box with an external adapter of about 230W and 2.5GbE; owners on GMKtec's own page praise its speed, but several complain about fan noise under CPU load. If you want it assembled, tested and covered for three years, our CG AI Halo 128 workstation ($3,799) is our own build in this 128GB class and can ship with Ubuntu, Ollama, llama.cpp and Open WebUI installed.

AMD's successor, the Ryzen AI Max+ PRO 495, now comes with up to 192GB, but Framework's 192GB configuration starts at $6,799. At that price it competes with the Spark and the Mac, not the $3,500 boxes, so the 128GB 395 remains the value point.

The tower alternative: discrete GPUs

Everything above is a compromise forced by memory capacity. If your model fits in VRAM, a discrete GPU is in another league because its memory is so much faster. NVIDIA's RTX PRO 6000 Blackwell carries 96GB of GDDR7 with ECC at 1,792 GB/s, about 6.5 times the Spark's bandwidth, and the RTX 5090 puts 32GB on a 512-bit bus. In the LMSYS dataset summarised by AIMultiple, an RTX 5090 generated gpt-oss-20b at 205.5 tokens per second against 60.9 on a DGX Spark, and read prompts about four times faster (8,519 against 2,054 tokens per second).

The costs are money, power and noise. NVIDIA rates the RTX PRO 6000 at 600W on its own and the 5090 at 575W, with a 1,000W system recommended for the latter. A model that spills out of VRAM into system RAM slows to a crawl, so a 32GB card is a 30B-class machine, not a 120B one. Our CG AI Pro 96 workstation with an RTX PRO 6000 is $21,499 at the time of writing. We spec it for 70B models at FP8 and 120B MoE models entirely on the GPU, and it serves several users at once and handles image generation and fine-tuning far faster than the small boxes. For one person chatting with a model, that is hard to justify; for a small team sharing one machine, it often is the right call.

Software: CUDA, MLX, ROCm and Vulkan

For plain LLM inference with llama.cpp, Ollama or LM Studio, all three platforms work in 2026. The differences show up at the edges.

  • NVIDIA (CUDA): the Spark ships with DGX OS, CUDA and NVIDIA's container toolkit preinstalled. Most research code, fine-tuning scripts and image or video pipelines target CUDA first. The trade-off is the Arm CPU: x86-only binaries won't run, and the October 2025 owner test found gpt-oss-120b took 1 minute 40 seconds to load against 22 seconds on Strix Halo.
  • Apple (MLX and Metal): MLX is Apple's open-source framework for running and fine-tuning models on Apple silicon. Thunderbolt 5 with RDMA lets you pool several Macs; Apple claims four Mac Studios run inference up to 3x faster than one.
  • AMD (ROCm and Vulkan): ROCm 7 supports the 395's gfx1151 GPU, and AMD points to the Lemonade project's prebuilt llama.cpp binaries. In the Framework community tests, Vulkan was usually fastest for generation while ROCm often won on prefill, so expect to try both and to edit kernel parameters for large models.

Power draw and noise

All three small boxes sip power next to a GPU tower. NVIDIA rates the GB10 chip at 140W with a 240W supply and declares 29 dB(A) at the operator position under full GPU load, quiet enough for a desk. The Strix Halo machines run at 120W sustained and 140W boost; Framework's Noctua fan is rated at 28.8 dBA. Apple lists 480W maximum continuous power for the Mac Studio, a ceiling rather than typical draw, and Tom's Hardware found the M5 Ultra largely quiet, though the fan can make itself heard.

If the machine serves models around the clock, do the sums. 140W for 24 hours a day is about 1,230 kWh a year. A single 600W card at full load is about 5,260 kWh before the rest of the tower, and it needs a big power supply, real airflow and a room where fan noise doesn't matter.

What we'd buy

  • Best value for one user running big MoE models: a 128GB Ryzen AI Max+ 395. The Framework Desktop at $3,449, or the CG AI Halo 128 at $3,799 if you want it set up with a 3-year parts warranty. Run Linux so the GPU can use up to 120GB.
  • Fastest single-user generation, quiet, no CUDA needed: the Mac Studio M5 Ultra. The $5,499 96GB model handles dense 70B models and gpt-oss-120b; for 200B-class MoE models with long context, budget for 256GB.
  • CUDA development, long agent prompts, clustering later: the DGX Spark 128GB at $6,950. Skip the 64GB model unless CUDA on models well under 100B parameters is all you need.
  • A team, image and video work, or fine-tuning: a tower with an RTX PRO 6000, such as the CG AI Pro 96 at $21,499.
  • Models under 30B only: don't buy any of these. A desktop with a 16GB or 32GB NVIDIA card generates several times faster on models that fit. And if you only need a model occasionally, paying per token in the cloud is usually cheaper; local hardware pays off when privacy, control or steady heavy use matter.

Everything above is in our AI workstations collection, ships insured in an unbranded box, and can be paid in Bitcoin, Monero, USDT and more than 20 other coins (see how to pay).

Questions and answers

Is the DGX Spark faster than a Mac Studio for local LLMs?

For generating text, no. The M5 Ultra has about 4.4 times the memory bandwidth, and Tom's Hardware measured almost four times the Spark's speed on a dense 27B model. The Spark's advantages are CUDA software support and easy clustering over its built-in 200 Gbps NIC.

How much of the 128GB can the Ryzen AI Max+ 395 GPU use?

Up to 96GB can be set as dedicated VRAM in the BIOS. On Linux, AMD documents a kernel parameter that raises the GPU's usable share to about 120GB, which is what you want for 100B-plus models.

Can a 128GB mini PC run a 70B model?

Yes. A 70B model at 4-bit needs about 40GB, so it fits easily, but bandwidth limits a dense 70B to around 5 tokens per second on a Ryzen AI Max+ 395. Mixture-of-experts models such as gpt-oss-120b run far faster, around 45 to 60 tokens per second on these boxes.

Is the 64GB DGX Spark worth buying?

Only if you specifically need CUDA in a small box. NVIDIA rates it for models up to about 100B parameters, and at $4,999 it costs more than a 128GB Ryzen AI Max+ 395 system with twice the memory.

Why does memory bandwidth matter more than TFLOPS for local AI?

Generating each token means reading all active weights from memory, so tokens per second is capped by bandwidth divided by the active model size. Compute matters more for prompt processing, which is why NVIDIA hardware reads long prompts faster.

Sources

  1. NVIDIA DGX Spark: specificationsNVIDIA
  2. NVIDIA's 64 GB DGX Spark launches this month for $4999, while the 128 GB Spark jumps past $6000Wccftech
  3. NVIDIA DGX Spark sees $700 price hike due to ongoing memory shortages, $4699 new pricingWccftech
  4. llama.cpp benchmarks on NVIDIA DGX Sparkggml-org / GitHub
  5. Apple introduces new Mac Studio with M5 Max and M5 UltraApple Newsroom
  6. Mac Studio technical specificationsApple
  7. Apple Mac Studio (M5 Ultra) review: Local model citizen outpaces DGX Spark and ThreadripperTom's Hardware
  8. AMD Ryzen AI Max+ 395 specificationsAMD
  9. How to run a one trillion-parameter LLM locally: an AMD Ryzen AI Max+ cluster guideAMD
  10. AMD Strix Halo (Ryzen AI Max+ 395) GPU LLM performance testsFramework Community
  11. DGX Spark vs. Strix Halo: initial impressionsFramework Community
  12. DGX Spark alternatives: RTX, Ryzen AI Halo and Mac StudioAIMultiple
  13. Framework Desktop specificationsFramework
  14. GMKtec EVO-X2 AMD Ryzen AI Max+ 395 AI Mini PCGMKtec
  15. NVIDIA RTX PRO 6000 Blackwell Workstation EditionNVIDIA
  16. GeForce RTX 5090 specificationsNVIDIA
  17. gpt-oss-120b model cardOpenAI / Hugging Face

Prices, fees and specifications were checked on October 5, 2026 and change over time. Product prices are ours at the time of writing.

Products in this article

All products

Keep reading

All articles
Local AI

How much VRAM for local LLM work in 2026: a practical sizing guide

A practical 2026 sizing guide: add up weights, KV cache and overhead, check real GGUF sizes for current open models, and see what fits on 16GB to 512GB.

Self-custody

Best hardware wallet 2026: Trezor, Ledger, Coldcard and more compared

Eleven wallets compared on firmware, secure chips, air-gapping and price, plus what the 2026 Coldcard seed bug changed and how to receive a wallet safely.

Payments

Best crypto to pay with online: Bitcoin, Monero or USDT

Real October 2026 fees, confirmation times, privacy and freeze risk for Bitcoin, Monero, USDT, USDC and Litecoin, and which coin to pick for each priority.