Frontier Model on Your Mac

Dirk Harms-Merbitz · September 2026 · linuxtoaster.com

toast is a Unix tool that happens to be a language model. You pipe text through it the way you pipe through grep. It does not care where the intelligence comes from: a cloud provider, our own models, or a daemon on the machine. Until this year the daemon on the machine was toasted, our engine for Qwen3-Coder-Next, a 30-billion parameter model that does honest work on a 64 GB Mac mini.

This post is about a bigger daemon. DwarfStar puts DeepSeek V4 Flash, 284 billion parameters, on a 128 GB MacBook. toast finds it, hands it tools, files and images, and the whole thing costs nothing to run. This is what it is, what it adds, and where our credits, slices and the cloud providers still fit.

What DwarfStar is

In May Salvatore Sanfilippo, antirez, the person who wrote Redis, released ds4, later renamed DwarfStar. It is an inference engine in C, self-contained, MIT licensed, with Metal, CUDA and ROCm backends. It does not link GGML; some of llama.cpp's kernels and quant layouts are adapted inside it, and the ggml authors' copyright stays in the license file because of that, which he says plainly.

It is deliberately narrow. It is not a general GGUF runner. It runs a short list of models, using quantized files the project itself produces, and it tests everything in integration: model loading, prompt rendering, tool calls, KV state, the HTTP server and the coding agent are built and tested together. Narrow is why it is fast and why it works.

The models:

These are mixture-of-experts models. Every token routes through a handful of experts, so the arithmetic per token is what a 13 to 18 billion parameter dense model costs, while the knowledge behind it is ten to twenty times that. That asymmetry is what makes a laptop possible.

The quantization is the other half. antirez's GGUFs are asymmetric: routed experts, which are most of the parameter count but each see a fraction of tokens, are squeezed to 2-bit; the router, attention projections, shared experts and output stay at 8-bit or half precision, because errors there touch every token. The q2 file is about 80 GB and is the one for a 128 GB Mac. A 256 GB machine takes q4. Smaller Macs can stream experts from SSD, slower but workable. Multi-token prediction gives speculative decoding on top: a small predictor proposes, the model verifies, and the words come faster than the memory bandwidth would otherwise allow.

What the server offers, which is what toast talks to:

Start it, leave it running. It is a daemon with an HTTP port, which is the shape toast was built around.

What it adds to toast

toast already spoke to DwarfStar the day it shipped, because DwarfStar speaks OpenAI and toastd, our daemon, forwards to a table of providers. DwarfStar is one more row. The work was in making the combination good, and in what the combination makes possible.

It is found, not configured. toast probes for local inference before it goes to the cloud: Apple's on-device model, then toasted, then DwarfStar on port 8000. If a ds4-server is listening, toast "why did it crash" goes there. Nothing to set. -p dwarfstar forces it, -p toast forces ours, DWARFSTAR_PORT moves the port.

Tools, the trained way. A .tools file in the directory lists the commands the model may run, one per line, with a comment after the # as its description. toast turns that file into native functions, one per line, and DwarfStar renders them into the model's own tool format. A command that is not in the file has no function to call. Every stage of every pipeline is still checked, every command is still printed before it runs, and .tools itself is never writable by the model. This is the same allowlist toast has always had; with DwarfStar it is enforced by the model's training as well as by our parser, and DeepSeek V4 Flash is, by antirez's own testing, very reliable at it.

$ printf 'ps       # processes\nlsof     # open files\n' > .tools
$ toast "what is holding port 8000"
[toast] lsof -i :8000
ds4-server, pid 41022, since 09:14.

For models without trained tool calling, toast still has its text protocol, and it now sends a stop sequence so the model cannot invent a tool's output. With DwarfStar you will not see that path.

Files. STRINGAPPEND and STRINGREPLACE in .tools become append and replace functions. toast server.c "bounds-check parse_header" reads the file, the model calls replace with an exact search and replacement, toast writes it atomically and prints what changed. .crumbs, the notebook the model keeps about your machine and project, works the same way, offline.

Images. toast screenshot.png "what's the error". With ds4-server --vision, toastd inlines the image into the message and DeepSeek reads it. Screenshots, diagrams, photos of whiteboards, all local.

Thinking, or not. DeepSeek thinks by default and the reasoning is separate from the answer, so toast prints only the answer. DWARFSTAR_NOTHINK=1 asks for direct replies when you want a one-liner back in a second.

All six steps. The tutorial climbs from asking a question to leaving a loop running for a week. Steps one to three, ask, chat, review, were always free. Steps four to six, loop, schedule, delegate, need jam and ito, which come with Local, our package for this: DwarfStar built and pinned to a tested commit, an installer that checks your memory, picks the quant and resumes an 80 GB download, start when toast asks and stop when idle, the vision encoder wired in, jam and ito included. $49 once, for a 128 GB Mac. Or clone the repository and build it yourself; toast will find that too.

Steering, which is new. DwarfStar can edit the model's activations at runtime. Research on refusal found that many coarse behaviors are not smeared through the weights but sit along a single direction in the residual stream, the vector that flows through the layers. DwarfStar's tools extract such a direction from two sets of prompts, target and contrast, and the server applies it per layer:

y = y - scale * d * dot(d, y)

Scale 1 removes the behavior, -1 amplifies it. antirez's bundled example is succinct versus verbose: at -1 the same answer is one paragraph, at 2 it has sections. toast's persona has always asked, in prose, for as few words as the question allows. A "toast" direction, built from the replies we want against chatty assistant replies and shipped with Local, gives DeepSeek that preference at the activation level. A knob, not a fine-tune; for what toast wants, a knob is enough.

Why the two fit

DwarfStar wants to be one long-running process with a port and a big model in memory. toast wants to be a reflex, grep with a brain, invoked a hundred times a day for a second each. The daemon holds the model; the tool comes and goes. That is the Unix split, and it is why toast never had to become an application.

Both are small C programs with strong opinions about what they do not do. DwarfStar does not run every model; toast does not manage sessions, prompts or plugins. DwarfStar tests tool calling as part of the build; toast treats the allowlist as the product. Neither needs Python, a package manager, or a hidden directory of blobs.

And the privacy story finally has a local answer that is not a compromise. The log, the source tree, the mail, the screenshot: none of it leaves the machine, and the model reading it is a frontier model, not a small one chosen because it fits.

When credits

DwarfStar is the free tier's engine. It is not the whole product.

Some work wants a bigger or a different model. A 2-bit quant of a 284-billion parameter model is remarkable, but Claude, GPT, Gemini and the full-precision hosted versions of DeepSeek and GLM are stronger, and sometimes the difference is the difference. Some work wants a specialist. Some work happens on a Mac with 16 GB.

Credits cover all of it, on one balance, metered by the token, from $20. Every major provider through linuxtoaster: Anthropic, OpenAI, Google, xAI, Moonshot, DeepSeek, Alibaba, Zhipu, Mistral. Our own models. And the slices.

A slice is a word after toast that names a job: toast reviewer, toast editor, toast paranoid, toast kimi. A persona plus a model behind it, chosen and tested by us and replaced when a better one appears. Slices always run on our side; that is where the persona lives, and toast sends them there whatever is running locally.

Which means one pipeline uses both:

grep -h ERROR /var/log/*.log | sort | uniq -c | sort -rn | head -50 \
  | toast "group these by fault, one line each" \
  | toast paranoid "which of these is a break-in"

grep, sort and uniq do what they have always done and turn the logs into fifty lines. The first toast runs on DwarfStar, free, and turns fifty lines into six. The second is a slice and sees six lines. Local for the reading, slice for the judgement. The logs never leave the machine; the credits go to the part that needs a specialist. The other way around works too: toast editor rewrites a chapter on our side, and the local model checks that it kept the plot.

Your own API keys still work and still cost nothing through us. Hosted and Colocation include credits every month; a colocated Mac Studio runs Local at q4 with speculative decoding and meters slices on top.

Getting started

$ git clone https://github.com/antirez/ds4 && cd ds4
$ ./download_model.sh q2        # 128 GB Mac; q4 for 256 GB
$ make                          # Metal
$ ./ds4-server --ctx 32768
$ toast "hello there"           # found on port 8000

Or install Local and skip the first four lines. Either way, toast --balance tells you what you have, and the first time you type toast reviewer with no credits it will say so and show you where to add them.