I run a large language model and an AI agent on hardware I own, at home. The model runs entirely on my own machine, so there is no per-token bill and my prompts never go to a model provider. I built it to learn how these systems behave when they have to work every day, and because a private, free-to-run assistant is useful. This page describes the setup, what it costs in speed, and what went wrong along the way.
At a glance
- Hardware: one small PC with an AMD Ryzen AI Max+ 395 (“Strix Halo”) and 128 GB of memory shared by the processor and graphics
- Model: Qwen3.8 Flash-Next, an open-weights mixture-of-experts model (about 125 billion parameters, roughly 6 billion used per word), 4-bit, with image input and a 64K-token context
- Engine: Gufo, an open-source inference engine built for this chip
- Agent: Hermes, an open-source agent framework, in its own container (I call the agent Sable)
- Cost: free and open-source software only; no paid APIs
How it is laid out
- One machine does one job. The Strix Halo box only serves models. Nothing else runs on it, so a crashed app can never take the model down.
- The agent lives elsewhere. Sable runs in a container on a separate Proxmox host, together with my projects and tools, and calls the model over the home network.
- A small helper model does the chores. A 4-billion-parameter model writes titles and summaries so the big model's time goes to real work. Speech-to-text (Whisper) and text-to-speech (Kokoro) run separately.
- A NAS holds the files and the backups.
How fast it is
Speed depends heavily on what you ask for. The model uses speculative decoding: it guesses several words ahead and checks them in one pass, which pays off when the text is predictable. These are measured single-user speeds on my machine:
| Kind of output | Tokens per second |
|---|---|
| Prose | 36 |
| Python code | 47–49 |
| Config files (YAML) | 52–56 |
| JSON | 58–60 |
Those match what the engine's authors publish: 32–34 tokens per second on mixed text and about 59 on repetitive text. The fastest community figure I found was 82 on code, from an experimental build that manages only about 22 on prose, so I kept the stable engine.
How the speed has changed
I have swapped the main model several times since August. These are the speeds I recorded at the time, on the same machine. They are not a clean comparison: the engine, the test prompts and the context length changed along the way, and the older runs used short prompts.
| When | Model and engine | Tokens/s |
|---|---|---|
| Aug 16 | Qwen3.8 27B, dense (community fine-tune), 4-bitllama.cpp | 41 (31 at 6-bit; about 17 past 60K tokens of context) |
| Aug 17 | Qwen3-Coder 30B-A3B, 4-bitllama.cpp | about 94 |
| Aug 19 | Ornith 1.5 35B-A3Bllama.cpp | 73–74 |
| Sep 23 | Qwen3.6 35B-A3B (community fine-tune), 4-bitllama.cpp with Vulkan | about 98 (87–106), short prompts |
| Sep 24 to now | Qwen3.8 Flash-Next, about 125BGufo | 36 prose, 47–60 code and data |
Two things stand out. First, I traded speed for size. The August favorites were 30–35 billion-parameter models that use only about 3 billion parameters per word and ran at 70–100 tokens per second. The current model is three to four times larger and uses about twice as many parameters per word, so it runs at roughly 35–60. Second, long conversations hurt dense models most: the 27B dense model fell from 41 to about 17 tokens per second once the chat passed 60,000 tokens, because it reads all of its weights for every word.
What the agent does
- Updates the machines every night at 4 AM and checks their health every hour.
- Writes The Lagging every week, and helps keep this website running (see how this site is run).
- Backs up my study notes every 30 minutes with restic, which stores each file once and keeps a long history in a few hundred megabytes.
It has written rules: never delete the NAS or the backups without asking, log every self-repair before making it, and get my sign-off on anything I would judge by eye or ear. Most of the setup itself was built with AI assistants (Claude Code and Sable) under my direction.
What I learned
- Silence is not success. A backup job quietly disappeared and nobody noticed for two weeks. Now a check alerts when a job goes quiet, not only when it errors. A multi-gigabyte model download also stalled without a single error message.
- Test an agent under its real conditions. I tried a community-modified variant of the model. It passed simple tool-call tests, then started printing its tool calls into the chat once it ran with the agent's full prompt. I rolled it back the same night. The old weights stay on disk until a new model has earned the swap.
- An agent can misdiagnose confidently. When a photo request timed out, the agent concluded the model server was down. The server was fine; the request had been waiting in line behind longer jobs. The logs told the story.
- Speed numbers need a workload. “30 tokens a second” and “60 tokens a second” were both true on the same machine.
Written October 2026. Speeds are my own measurements on one machine and will change as the software does.