What Happened When I Switched to a Local LLM for Technical Writing
The appeal of running a language model locally is easy to understand. No API costs, no data leaving your machine, no rate limits, and full control over which model you're using. For teams dealing with sensitive client information or just trying to reduce ongoing costs, local LLMs have become a serious option.
The reality is more qualified. Local models are genuinely capable for some content work tasks, significantly worse than frontier models for others, and come with setup overhead that isn't always worth it depending on your situation.
What "local LLM" Means in Practice
Running a model locally means downloading model weights to your machine and using software that loads those weights and runs inference on your own hardware, without sending anything to an external API.
The main tools for this are Ollama, LM Studio, and llama.cpp. Ollama is probably the easiest entry point: you install it, run a single command to pull a model, and interact with it through a simple API or a basic chat interface. LM Studio has a graphical interface that makes it more accessible for non-technical users.
The model you run depends on your hardware. A 7-billion-parameter model (7B) will run on most modern machines with 8GB of RAM, though slowly if you're using CPU only. A 13B model benefits from a GPU with at least 8GB VRAM. Models larger than 13B require significant GPU memory and run slowly on consumer hardware.
What Hardware You Actually Need
The single biggest factor in local LLM performance is whether you're running inference on a GPU or CPU.
CPU-only inference is possible but slow. A 7B model might generate 3-8 tokens per second on a modern CPU. That's readable but slow enough to feel painful for longer content generation tasks — 1,000 words at 5 tokens per second takes more than three minutes.
GPU inference is dramatically faster. The same 7B model running on a GPU with the full model loaded in VRAM will generate 50-100 tokens per second. For Apple Silicon Macs (M1 and later), the unified memory architecture allows models to use the GPU efficiently even without dedicated VRAM, which makes M-series Macs surprisingly capable for local inference.
For most content work use cases, a modern M-series Mac or a Windows machine with a mid-range NVIDIA GPU (16GB VRAM or more) gives a good experience.
Which Models Are Worth Using
The open-source model landscape changes quickly. As of writing, the models that consistently perform well for content tasks in their size class are:
Llama 3.1 8B — Meta's current small model is capable enough for many content tasks: summarization, editing, reformatting, extracting key points, drafting short-form content.
Mistral 7B and variants — Fast, capable, and well-supported by local inference tools.
Qwen 2.5 — Strong multilingual capability and competitive performance on reasoning tasks for its size.
Phi-3.5 — Microsoft's efficient small model, performs well on instruction-following tasks relative to its size.
None of these match GPT-4o or Claude 3.5 Sonnet for complex writing, nuanced editing, or content that requires sophisticated reasoning. They're good enough for many practical tasks, but the gap with frontier models is real and noticeable on demanding work.
Content Tasks Where Local Models Work Well
Summarization. Feeding an article and asking for a concise summary or key points is a task where 7B-13B models perform reliably well. The context length is manageable and the task is straightforward.
Reformatting. Asking a model to convert a bulleted list into prose, or restructure content into a different format, works consistently. These tasks don't require sophisticated reasoning — mostly pattern-following.
First drafts of simple content. Email templates, brief product descriptions, FAQ answers, social media posts for known topics. These don't require the depth of knowledge or writing sophistication where frontier models pull ahead significantly.
Extracting structured information. Giving the model a chunk of unstructured text and asking it to extract specific fields in JSON format works well with instruction-tuned models.
Checking for specific issues. Reading a draft and asking "are there any factual inconsistencies?" or "does this use passive voice too frequently?" — local models handle these review tasks reasonably.
Content Tasks Where Local Models Underperform
Complex, long-form writing. Coherence across a 3,000-word article is harder for smaller models. Arguments lose track of themselves, examples get repeated, and the prose quality drops noticeably compared to frontier models.
Sophisticated editing. Asking a model to identify structural problems in a piece of writing, suggest a significantly different approach to an argument, or identify where tone and evidence are misaligned — these tasks require a level of judgment that smaller models struggle with.
Domain-specific accuracy. Local models are more likely to hallucinate or make confident factual errors on specialized topics because their training data is smaller and less curated than frontier models. Always verify specific facts.
Consistency across large projects. Maintaining consistent tone, terminology, and framing across a series of articles is something local models find difficult because they have limited context windows and no persistent memory.
RAG: the Approach That Extends Local Model Usefulness
Retrieval-Augmented Generation (RAG) is a technique where you give the model access to a specific knowledge base — your existing articles, your style guide, your product documentation — at query time, by retrieving relevant chunks and including them in the context window.
This extends local model usefulness significantly because it addresses one of the main limitations: lack of specific knowledge. With RAG, you can ask a local model to "write a description consistent with our existing product descriptions" and provide the relevant examples in context, producing output that's more aligned with your existing content than a pure generation approach.
Tools like LlamaIndex and LangChain make it relatively straightforward to build a simple RAG pipeline against local files.
The Setup Cost Is Real
Local LLMs require time to set up, and models occasionally need updating when better versions are released. The software ecosystem is changing quickly enough that tutorials from six months ago may need adaptation.
For a solo content creator or a small team that generates a lot of content, the setup investment can pay off. For someone who uses AI assistance occasionally, cloud API costs are probably lower than the time investment of maintaining a local setup.
Evaluate honestly which situation you're in before committing.