Running a large language model on your own computer is no longer a weekend project. In 2026 it takes three steps: check how much memory you have, install one free app, and download a model that fits. This guide walks through each step, with the memory numbers you need to pick the right model size and the mistakes that cause most first-time failures.

Why Run a Model Locally?


- Privacy: prompts and documents never leave your machine.
- Cost: after the hardware, there is no per-token bill, which matters if you send thousands of requests.
- Offline use: it works without a connection, and it won't change under you when a provider updates a model.
- Control: you choose the model, the settings, and what gets logged.

There is a trade-off. The best models in the world are still closed and hosted, and the open-weight models you can run on a single desktop sit below them. If you only send a few prompts a day, an API or a subscription is usually cheaper than buying hardware. For a comparison of where open models stand, see our [top 10 LLMs right now](https://directory.drveri.com/blog/the-top-10-llms-right-now-october-2026-including-open and [open-weight LLMs in 2026](https://directory.drveri.com/blog/open-weight-llms-in-2026-how-close-are-they-to-gpt-5-and

Step 1: Check Your Memory Budget


The single number that decides what you can run is how much fast memory the model can use: video memory (VRAM) on an NVIDIA or AMD graphics card, or unified memory on an Apple Silicon Mac. A model is a big file of weights, and it needs to fit in memory with room left over for the conversation itself.

Most people run models in a compressed ("quantized") form. The usual starting point is a 4-bit format called Q4_K_M, which cuts memory use by roughly 70 percent compared with the full-precision file while keeping most of the quality. These are approximate figures from several 2026 guides, including room for context:
Model sizeMemory for the weightsWith context and overheadTypical hardware
8 billion parametersabout 4.5 GBabout 5 to 7 GB8 GB graphics card
14 billionabout 8 GBabout 9 to 11 GB12 to 16 GB card
32 billionabout 18 to 19 GBabout 20 to 23 GB24 GB card
70 billionabout 39 to 40 GBabout 40 to 46 GB48 GB or more, or 64 GB unified memory

Two things move these numbers. A longer context window uses much more memory, so a 32K context can add several gigabytes over an 8K one. And exact figures vary by model, so a safe habit is to look at the actual file size on the model's download page and add about 20 percent for headroom.

If you have no dedicated GPU, a model can still run from ordinary system RAM, but it will be much slower. A recent Mac with 16 GB or more of unified memory is one of the easiest ways to start.

Step 2: Pick a Runner


Three free tools do the same job at different levels of control. They all read the same GGUF model files.
ToolBest forHow you use it
LM StudioBeginners and anyone who prefers a windowDesktop app with a built-in model browser and chat screen
OllamaDevelopers and anyone who wants a local APIOne-line commands, plus an API server other apps can call
llama.cppMaximum control and tuningCommand-line program you build or download

A reasonable default: start with LM Studio if you want to click through a chat, or with Ollama if you are comfortable in a terminal. You can use both. Speed comparisons between them disagree from machine to machine, so treat any single benchmark as a rough hint, not a rule.

Step 3: Install and Run Your First Model


With Ollama. Download the installer from ollama.com for Windows or macOS, or on Linux and macOS run the install script shown on their site. Then open a terminal and start a model, for example:

ollama run qwen3:8b

Ollama downloads the model the first time and then opens a chat. To see what is loaded and whether it is using your GPU, run:

ollama ps

With LM Studio. Install it from lmstudio.ai, open the Discover tab, search for a model, download one that fits your memory, then open the Chat tab and load it.

Model names change quickly. Always check the library page of the tool you use for the current names and sizes before downloading, and prefer the official model page over a blog post's spelling.

Which Model Should You Choose?


Recent 2026 roundups point to a few sensible starting points. Sources disagree on the exact best pick, so take these as places to begin, then try two models on your own prompts.

- About 8 GB of memory: an 8-billion-parameter class model, or one of the small Gemma 4 sizes.
- About 16 GB: gpt-oss-20b is the most common recommendation, and Gemma 4 26B-A4B is a stronger option if you accept a tighter fit.
- About 24 GB: Qwen3.8-27B is the most common default, with Gemma 4 31B as the main rival. Both fit at 4-bit.
- 64 GB or more: a 70-billion-class model becomes practical.

Check each model's license too. Several of the strongest open models use custom licenses with commercial limits, while others are MIT or Apache 2.0.

Step 4: Check That It Is Running Well


- Good: answers stream at a comfortable reading speed.
- Too slow, word by word: the weights have spilled out of GPU memory into system RAM. Use a smaller model or a smaller quantization.
- Out-of-memory errors: lower the context length, close other apps that use the GPU, or choose a smaller model.
- Quality seems poor: try a larger quantization if you have the memory, since very aggressive compression costs accuracy.

Step 5: Use It From Other Apps


Ollama runs a local server, by default on port 11434, that speaks an OpenAI-compatible API. That means many editors, chat front ends and scripts can point at your own machine instead of a paid service just by changing the address. LM Studio also has a server mode. For more on using local models for programming, see our guide to [home computers for running coding LLMs locally](https://directory.drveri.com/blog/best-home-computers-for-running-coding-llms-locally-in-2026

Common Mistakes


- Choosing a model by parameter count alone. A smaller model that fits comfortably beats a bigger one that spills into slow memory.
- Forgetting the context. A model that loads fine can still run out of memory on a long document.
- Trusting any single benchmark number for speed. Try it on your own hardware.
- Ignoring the license when building a product.

Quick Summary


Check your memory, install LM Studio or Ollama, download a 4-bit model that fits with about 20 percent to spare, and test it on your own prompts. If it is slow, go one size down. That is all most people need to get a capable assistant running privately on their own machine.

Sources


- Running local LLMs in 2026 (SitePoint): https://www.sitepoint.com/run-local-llms-2026-complete-developer-guide/
- Ollama vs LM Studio vs llama.cpp (MachineLearningMastery): https://machinelearningmastery.com/ollama-vs-lm-studio-vs-llama-cpp-which-local-ai-runtime-should-you-use-in-2026/
- Local LLM hardware requirements 2026: https://www.promptquorum.com/local-llms/local-llm-hardware-guide-2026
- Best local LLMs by VRAM tier: https://llmconfigurator.com/en/guides/best-local-llm-by-vram
- Ollama: https://ollama.com and LM Studio: https://lmstudio.ai