For under $2,000, the best hardware for running AI models at home is a used 24GB graphics card if you already own a PC, or a 48GB Mac mini if you want a quiet all-in-one box. Everything else is either too small to be useful or costs more than your budget.

This is a snapshot of prices and speeds as of mid-October 2026, and it's an unusual year to be shopping, so some of the usual advice doesn't apply. Prices are volatile, so treat the numbers as a guide and check live listings before you buy.

The short answer


- You already have a desktop PC: buy a used RTX 3090 (24GB). It's still the best memory-per-dollar GPU, and it runs 27 to 32 billion parameter models comfortably.
- You're starting from nothing and want quiet and simple: a Mac mini with the M4 Pro chip and 48GB of memory.
- You have under $1,000: a 16GB graphics card such as the RTX 5060 Ti in an existing PC, or a 32GB Mac mini.
- You want to experiment with the biggest models: an AMD Ryzen AI Max+ 395 mini PC with 64GB, if you can find one in stock. The 128GB versions are over budget.

Why this is a strange year to buy


Memory has gotten expensive. A 32GB DDR5 kit that cost around $90 in summer 2025 sits at roughly $430 to $500 now, because chip makers are putting their capacity into memory for AI data centers. Graphics cards followed: the RTX 5090 is selling for roughly 2.4 to 2.5 times its $1,999 list price, and the 16GB RTX 5070 Ti and RTX 5080 are running about 1.5 times theirs.

Analysts disagree on when this eases, with some pointing to late 2026 and others to sometime in 2027. So the sensible moves are to reuse parts you already own, buy used where it's safe, or buy a machine where the memory comes included, like a Mac. Building a new PC from scratch with new DDR5 is the worst place to spend a tight budget this year.

The one number that matters


For running a model, the spec that matters most is how much fast memory the model can use: graphics card memory (VRAM) on a PC, or unified memory on a Mac or AMD's new mini PCs. If a model doesn't fit, it either won't run or crawls. Memory bandwidth then decides how fast it generates text. If you want the full explanation, our guide to running an open-weight LLM on your own computer covers memory sizes by model.

The contenders


OptionTypical price (Oct 2026)MemoryBest forThe catch
Used RTX 3090about $1,290 to $1,450 (card only)24GB VRAM27B to 32B models at roughly 30 to 50 tokens per secondNeeds a PC and a strong power supply, used card with little or no warranty
Mac mini M4 Pro, 48GBabout $1,800 to $2,00048GB unified30B-class models, near-silent, easy setup70B models run at single-digit speeds, memory can't be upgraded
Mac mini M4, 32GBabout $1,15032GB unified7B to 14B models at roughly 28 to 35 tokens per secondLimited for larger models
RTX 5060 Ti, 16GBabout $430 to $58516GB VRAM13B to 14B models, small mixture-of-experts models16GB ceiling
RX 9070 XT, 16GBabout $700 to $80016GB VRAMAMD option with official Linux ROCm supportSmaller software ecosystem than NVIDIA
Ryzen AI Max+ 395 mini PC, 64GBabout $1,960 to $2,20064GB unifiedLarge mixture-of-experts modelsStock is patchy, dense models are slow, 128GB versions cost $3,300 and up

Prices come from tracker pages and retailer listings dated early October 2026, and some sources disagree by a few hundred dollars.

What the speeds feel like


Speed is measured in tokens per second, and a rough rule of thumb is that 5 feels slow but readable, 20 or more feels comfortable, and 40 or more feels snappy.

Community benchmarks for a 27-billion-parameter dense model (Qwen3.8 27B, 4-bit):

- RTX 3090: about 40 tokens per second with standard llama.cpp, and over 100 with a tuned setup.
- RTX 4090 and 5090, for reference: about 46 and 75, which shows how much you'd pay for each step up.
- Ryzen AI Max+ 395: about 11, rising to roughly 24 to 30 with a speed-up technique called multi-token prediction.

The picture changes with "mixture-of-experts" models, which only use part of their weights per word. These run much faster on memory-rich machines: around 72 to 75 tokens per second for a 30B-class model on Ryzen AI Max+ 395 mini PCs, and a similar 40 to 80 range reported on a 48GB Mac mini M4 Pro. Dense 70B models are the weak spot for every machine in this price range, with reports of roughly 3 to 12 tokens per second depending on the chip and quantization.

All of these numbers come from different people, different settings, and different dates, so use them for ranking, not for promises.

What I would skip


- The RTX 5090. At $3,800 to $5,000 and up, it's out of this budget by a wide margin.
- New 16GB cards at $1,100 to $1,650. The RTX 5070 Ti and RTX 5080 cost much more than the 5060 Ti for the same 16GB of memory, which is the number that limits you.
- 128GB mini PCs. They're appealing on paper, but the cheapest listings are $3,300 and up.
- Paying extra for the fastest card when your models don't fit anyway. Memory first, speed second.

Buying tips


1. Count VRAM, not total RAM. A PC with 64GB of system memory and an 8GB graphics card is still an 8GB machine for model purposes.
2. Check the power requirement. A high-end card needs a capable power supply, and the 3090 is a power-hungry card, so confirm what the card needs before you buy.
3. Be careful with used cards. Buy from a seller with a return window, test it under load as soon as it arrives, and check the temperatures.
4. Decide your memory now if you buy a Mac. It can't be upgraded afterward, and the 48GB version is the one that opens up 30B-class models comfortably.
5. Try before you buy. Run the model you want on a rented cloud GPU for an hour first, so you know what size you really need.

Is it even worth buying hardware?


Often not. If you only send a few prompts a day, a subscription or an API is cheaper than $2,000 of hardware. Local hardware makes sense when you want privacy, when you run a lot of requests, or when you simply enjoy tinkering. For the models these machines run and how they compare with hosted ones, see our top 10 LLMs right now, and for what to pay for a hosted assistant, our AI assistant comparison.

Common questions


Is 16GB enough? For 13B to 20B-class models, yes. For the 27B to 32B models that currently give the best results at this price, you want 24GB or more.

Do I need an NVIDIA card? It's the safest choice for software support, but AMD cards work well on Linux, and Macs and AMD's mini PCs run these models without any graphics card at all.

Mac or PC? A Mac is quieter, smaller and simpler. A PC with a used 3090 is faster per dollar on models that fit in 24GB, but louder and less convenient.

Should I wait for prices to fall? Some analysts expect memory prices to ease in late 2026, and others say 2027. If you need the machine now, buy for what you need and avoid paying a premium for more than that.

Sources


- Used RTX 3090 price tracker, October 2026: https://gpudojo.com/rtx-3090
- RTX 5090 price history, October 2026: https://bestvaluegpu.com/history/new-and-used-rtx-5090-price-history-and-specs/
- Why GPU prices are climbing again in 2026: https://pcgamecheck.com/blog/why-gpus-expensive-2026-memory-shortage-explained
- DDR5 RAM prices 2026: https://tech-insider.org/ddr5-ram-prices-2026-pc-builders/
- Ryzen AI Max+ 395 mini PCs compared: https://computingforgeeks.com/ryzen-ai-max-395-mini-pc-comparison/
- Strix Halo tokens per second: https://datahardware.ai/blog/strix-halo-tokens-per-second-2026
- Mac mini LLM performance in 2026: https://www.popularai.org/p/mac-mini-llm-performance-in-2026
- Mac mini M4 vs mini PC vs GPU for local LLMs: https://computingforgeeks.com/mac-mini-vs-mini-pc-vs-gpu-local-llm/
- RTX 3090 local LLM benchmarks: https://www.hardware-corner.net/gpu-llm-benchmarks/rtx-3090/