How to Run an Open-Weight LLM on Your Own Computer (Step by Step, 2026)
Running a large language model on your own computer is no longer a weekend project. In 2026 it takes three steps: check how much memory you have, install one free app, and download a model that fits. This guide walks through each step, with the memory numbers you need to pick the right model size and the mistakes that cause most first-time failures.
- Privacy: prompts and documents never leave your machine.
- Cost: after the hardware, there is no per-token bill, which matters if you send thousands of requests.
- Offline use: it works without a connection, and it won't change under you when a provider updates a model.
- Control: you choose the model, the settings, and what gets logged.
There is a trade-off. The best models in the world are still closed and hosted, and the open-weight models you can run on a single desktop sit below them. If you only send a few prompts a day, an API or a subscription is usually cheaper than buying hardware. For a comparison of where open models stand, see our [top 10 LLMs right now](https://directory.drveri.com/blog/the-top-10-llms-right-now-october-2026-including-open and [open-weight LLMs in 2026](https://directory.drveri.com/blog/open-weight-llms-in-2026-how-close-are-they-to-gpt-5-and
The single number that decides what you can run is how much fast memory the model can use: video memory (VRAM) on an NVIDIA or AMD graphics card, or unified memory on an Apple Silicon Mac. A model is a big file of weights, and it needs to fit in memory with room left over for the conversation itself.
Most people run models in a compressed ("quantized") form. The usual starting point is a 4-bit format called Q4_K_M, which cuts memory use by roughly 70 percent compared with the full-precision file while keeping most of the quality. These are approximate figures from several 2026 guides, including room for context:
Two things move these numbers. A longer context window uses much more memory, so a 32K context can add several gigabytes over an 8K one. And exact figures vary by model, so a safe habit is to look at the actual file size on the model's download page and add about 20 percent for headroom.
If you have no dedicated GPU, a model can still run from ordinary system RAM, but it will be much slower. A recent Mac with 16 GB or more of unified memory is one of the easiest ways to start.
Three free tools do the same job at different levels of control. They all read the same GGUF model files.
A reasonable default: start with LM Studio if you want to click through a chat, or with Ollama if you are comfortable in a terminal. You can use both. Speed comparisons between them disagree from machine to machine, so treat any single benchmark as a rough hint, not a rule.
With Ollama. Download the installer from ollama.com for Windows or macOS, or on Linux and macOS run the install script shown on their site. Then open a terminal and start a model, for example:
ollama run qwen3:8b
Ollama downloads the model the first time and then opens a chat. To see what is loaded and whether it is using your GPU, run:
ollama ps
With LM Studio. Install it from lmstudio.ai, open the Discover tab, search for a model, download one that fits your memory, then open the Chat tab and load it.
Model names change quickly. Always check the library page of the tool you use for the current names and sizes before downloading, and prefer the official model page over a blog post's spelling.
Recent 2026 roundups point to a few sensible starting points. Sources disagree on the exact best pick, so take these as places to begin, then try two models on your own prompts.
- About 8 GB of memory: an 8-billion-parameter class model, or one of the small Gemma 4 sizes.
- About 16 GB: gpt-oss-20b is the most common recommendation, and Gemma 4 26B-A4B is a stronger option if you accept a tighter fit.
- About 24 GB: Qwen3.8-27B is the most common default, with Gemma 4 31B as the main rival. Both fit at 4-bit.
- 64 GB or more: a 70-billion-class model becomes practical.
Check each model's license too. Several of the strongest open models use custom licenses with commercial limits, while others are MIT or Apache 2.0.
- Good: answers stream at a comfortable reading speed.
- Too slow, word by word: the weights have spilled out of GPU memory into system RAM. Use a smaller model or a smaller quantization.
- Out-of-memory errors: lower the context length, close other apps that use the GPU, or choose a smaller model.
- Quality seems poor: try a larger quantization if you have the memory, since very aggressive compression costs accuracy.
Ollama runs a local server, by default on port 11434, that speaks an OpenAI-compatible API. That means many editors, chat front ends and scripts can point at your own machine instead of a paid service just by changing the address. LM Studio also has a server mode. For more on using local models for programming, see our guide to [home computers for running coding LLMs locally](https://directory.drveri.com/blog/best-home-computers-for-running-coding-llms-locally-in-2026
- Choosing a model by parameter count alone. A smaller model that fits comfortably beats a bigger one that spills into slow memory.
- Forgetting the context. A model that loads fine can still run out of memory on a long document.
- Trusting any single benchmark number for speed. Try it on your own hardware.
- Ignoring the license when building a product.
Check your memory, install LM Studio or Ollama, download a 4-bit model that fits with about 20 percent to spare, and test it on your own prompts. If it is slow, go one size down. That is all most people need to get a capable assistant running privately on their own machine.
- Running local LLMs in 2026 (SitePoint): https://www.sitepoint.com/run-local-llms-2026-complete-developer-guide/
- Ollama vs LM Studio vs llama.cpp (MachineLearningMastery): https://machinelearningmastery.com/ollama-vs-lm-studio-vs-llama-cpp-which-local-ai-runtime-should-you-use-in-2026/
- Local LLM hardware requirements 2026: https://www.promptquorum.com/local-llms/local-llm-hardware-guide-2026
- Best local LLMs by VRAM tier: https://llmconfigurator.com/en/guides/best-local-llm-by-vram
- Ollama: https://ollama.com and LM Studio: https://lmstudio.ai
Why Run a Model Locally?
- Privacy: prompts and documents never leave your machine.
- Cost: after the hardware, there is no per-token bill, which matters if you send thousands of requests.
- Offline use: it works without a connection, and it won't change under you when a provider updates a model.
- Control: you choose the model, the settings, and what gets logged.
There is a trade-off. The best models in the world are still closed and hosted, and the open-weight models you can run on a single desktop sit below them. If you only send a few prompts a day, an API or a subscription is usually cheaper than buying hardware. For a comparison of where open models stand, see our [top 10 LLMs right now](https://directory.drveri.com/blog/the-top-10-llms-right-now-october-2026-including-open and [open-weight LLMs in 2026](https://directory.drveri.com/blog/open-weight-llms-in-2026-how-close-are-they-to-gpt-5-and
Step 1: Check Your Memory Budget
The single number that decides what you can run is how much fast memory the model can use: video memory (VRAM) on an NVIDIA or AMD graphics card, or unified memory on an Apple Silicon Mac. A model is a big file of weights, and it needs to fit in memory with room left over for the conversation itself.
Most people run models in a compressed ("quantized") form. The usual starting point is a 4-bit format called Q4_K_M, which cuts memory use by roughly 70 percent compared with the full-precision file while keeping most of the quality. These are approximate figures from several 2026 guides, including room for context:
| Model size | Memory for the weights | With context and overhead | Typical hardware |
|---|---|---|---|
| 8 billion parameters | about 4.5 GB | about 5 to 7 GB | 8 GB graphics card |
| 14 billion | about 8 GB | about 9 to 11 GB | 12 to 16 GB card |
| 32 billion | about 18 to 19 GB | about 20 to 23 GB | 24 GB card |
| 70 billion | about 39 to 40 GB | about 40 to 46 GB | 48 GB or more, or 64 GB unified memory |
Two things move these numbers. A longer context window uses much more memory, so a 32K context can add several gigabytes over an 8K one. And exact figures vary by model, so a safe habit is to look at the actual file size on the model's download page and add about 20 percent for headroom.
If you have no dedicated GPU, a model can still run from ordinary system RAM, but it will be much slower. A recent Mac with 16 GB or more of unified memory is one of the easiest ways to start.
Step 2: Pick a Runner
Three free tools do the same job at different levels of control. They all read the same GGUF model files.
| Tool | Best for | How you use it |
|---|---|---|
| LM Studio | Beginners and anyone who prefers a window | Desktop app with a built-in model browser and chat screen |
| Ollama | Developers and anyone who wants a local API | One-line commands, plus an API server other apps can call |
| llama.cpp | Maximum control and tuning | Command-line program you build or download |
A reasonable default: start with LM Studio if you want to click through a chat, or with Ollama if you are comfortable in a terminal. You can use both. Speed comparisons between them disagree from machine to machine, so treat any single benchmark as a rough hint, not a rule.
Step 3: Install and Run Your First Model
With Ollama. Download the installer from ollama.com for Windows or macOS, or on Linux and macOS run the install script shown on their site. Then open a terminal and start a model, for example:
ollama run qwen3:8b
Ollama downloads the model the first time and then opens a chat. To see what is loaded and whether it is using your GPU, run:
ollama ps
With LM Studio. Install it from lmstudio.ai, open the Discover tab, search for a model, download one that fits your memory, then open the Chat tab and load it.
Model names change quickly. Always check the library page of the tool you use for the current names and sizes before downloading, and prefer the official model page over a blog post's spelling.
Which Model Should You Choose?
Recent 2026 roundups point to a few sensible starting points. Sources disagree on the exact best pick, so take these as places to begin, then try two models on your own prompts.
- About 8 GB of memory: an 8-billion-parameter class model, or one of the small Gemma 4 sizes.
- About 16 GB: gpt-oss-20b is the most common recommendation, and Gemma 4 26B-A4B is a stronger option if you accept a tighter fit.
- About 24 GB: Qwen3.8-27B is the most common default, with Gemma 4 31B as the main rival. Both fit at 4-bit.
- 64 GB or more: a 70-billion-class model becomes practical.
Check each model's license too. Several of the strongest open models use custom licenses with commercial limits, while others are MIT or Apache 2.0.
Step 4: Check That It Is Running Well
- Good: answers stream at a comfortable reading speed.
- Too slow, word by word: the weights have spilled out of GPU memory into system RAM. Use a smaller model or a smaller quantization.
- Out-of-memory errors: lower the context length, close other apps that use the GPU, or choose a smaller model.
- Quality seems poor: try a larger quantization if you have the memory, since very aggressive compression costs accuracy.
Step 5: Use It From Other Apps
Ollama runs a local server, by default on port 11434, that speaks an OpenAI-compatible API. That means many editors, chat front ends and scripts can point at your own machine instead of a paid service just by changing the address. LM Studio also has a server mode. For more on using local models for programming, see our guide to [home computers for running coding LLMs locally](https://directory.drveri.com/blog/best-home-computers-for-running-coding-llms-locally-in-2026
Common Mistakes
- Choosing a model by parameter count alone. A smaller model that fits comfortably beats a bigger one that spills into slow memory.
- Forgetting the context. A model that loads fine can still run out of memory on a long document.
- Trusting any single benchmark number for speed. Try it on your own hardware.
- Ignoring the license when building a product.
Quick Summary
Check your memory, install LM Studio or Ollama, download a 4-bit model that fits with about 20 percent to spare, and test it on your own prompts. If it is slow, go one size down. That is all most people need to get a capable assistant running privately on their own machine.
Sources
- Running local LLMs in 2026 (SitePoint): https://www.sitepoint.com/run-local-llms-2026-complete-developer-guide/
- Ollama vs LM Studio vs llama.cpp (MachineLearningMastery): https://machinelearningmastery.com/ollama-vs-lm-studio-vs-llama-cpp-which-local-ai-runtime-should-you-use-in-2026/
- Local LLM hardware requirements 2026: https://www.promptquorum.com/local-llms/local-llm-hardware-guide-2026
- Best local LLMs by VRAM tier: https://llmconfigurator.com/en/guides/best-local-llm-by-vram
- Ollama: https://ollama.com and LM Studio: https://lmstudio.ai