Running AI at Home

I run a large language model and an AI agent on hardware I own, at home. The model runs entirely on my own machine, so there is no per-token bill and my prompts never go to a model provider. I built it to learn how these systems behave when they have to work every day, and because a private, free-to-run assistant is useful. This page describes the setup, what it costs in speed, and what went wrong along the way.

At a glance

How it is laid out

How fast it is

Speed depends heavily on what you ask for. The model uses speculative decoding: it guesses several words ahead and checks them in one pass, which pays off when the text is predictable. These are measured single-user speeds on my machine:

Kind of outputTokens per second
Prose36
Python code47–49
Config files (YAML)52–56
JSON58–60

Those match what the engine's authors publish: 32–34 tokens per second on mixed text and about 59 on repetitive text. The fastest community figure I found was 82 on code, from an experimental build that manages only about 22 on prose, so I kept the stable engine.

How the speed has changed

I have swapped the main model several times since August. These are the speeds I recorded at the time, on the same machine. They are not a clean comparison: the engine, the test prompts and the context length changed along the way, and the older runs used short prompts.

WhenModel and engineTokens/s
Aug 16Qwen3.8 27B, dense (community fine-tune), 4-bitllama.cpp41 (31 at 6-bit; about 17 past 60K tokens of context)
Aug 17Qwen3-Coder 30B-A3B, 4-bitllama.cppabout 94
Aug 19Ornith 1.5 35B-A3Bllama.cpp73–74
Sep 23Qwen3.6 35B-A3B (community fine-tune), 4-bitllama.cpp with Vulkanabout 98 (87–106), short prompts
Sep 24 to nowQwen3.8 Flash-Next, about 125BGufo36 prose, 47–60 code and data

Two things stand out. First, I traded speed for size. The August favorites were 30–35 billion-parameter models that use only about 3 billion parameters per word and ran at 70–100 tokens per second. The current model is three to four times larger and uses about twice as many parameters per word, so it runs at roughly 35–60. Second, long conversations hurt dense models most: the 27B dense model fell from 41 to about 17 tokens per second once the chat passed 60,000 tokens, because it reads all of its weights for every word.

What the agent does

It has written rules: never delete the NAS or the backups without asking, log every self-repair before making it, and get my sign-off on anything I would judge by eye or ear. Most of the setup itself was built with AI assistants (Claude Code and Sable) under my direction.

What I learned

Written October 2026. Speeds are my own measurements on one machine and will change as the software does.