Lightweight toast talks to a local toastd, which keeps
an HTTP/2 connection pool to linuxtoaster.com. Written in C to minimize latency.
With BYOK, toastd connects directly to your provider — your traffic never touches
our servers. Shell is the default client, not the requirement: a Python script, a C
program or a cron job can talk to toastd the same way.
A specialist called by name — toast editor, toast security,
toast paranoid. The word after toast picks a server-side config;
toast --add editor drops a symlink if you want one word. The name
selects a hosted, tuned configuration that spans providers. --balance
lists the ones your account can reach, --list the ones you have. Plain
toast uses your local .persona file instead. Only one of
them is ever read, so they cannot fight.
toast exits 1 when the model prints DONE, so
while toast ... ends when the work is finished, not when a counter
runs out. 20 while caps it anyway. Tool calls run through
jam, five rounds max, against a
.tools allowlist you write yourself.
Locally. Context in .crumbs, conversations in .chat,
tool permissions in .tools, your own slice in .persona.
Version them, grep them, delete them. Memory you cannot read is memory you cannot
correct.
On MacBooks, the installer downloads appled — a local inference
provider using Apple Intelligence. No API key and no cost per token — prompts never
leave your machine, and running it spends none of your credits. It still needs an
account: toast brings up a Stripe page the first time you run it.
Yes — appled, toasted, Ollama, MLX, LM Studio, KoboldCpp, llama.cpp, vLLM, LocalAI, or Jan. No internet, no API keys, full privacy.
Got a PROVIDER_API_KEY set for Anthropic, Cerebras, Google Gemini,
Groq, OpenAI, OpenRouter, Together, Mistral, Perplexity, or xAI? Use
toast -p provider. Zero config, zero cost from us.
A from-scratch local inference daemon for Apple Silicon (Pro tier). Loads Qwen3-Next-Coder — a 30B coding model — via C++ against Apple's MLX API. ~100 tok/s generation, ~400 tok/s prefill, 0.6s to first token. 128 GB supports 8/6/4-bit quantization, 64 GB supports 4-bit.
squawkd joins a multicast group on the LAN and squawk bot
Paranoid puts a model in the room as another participant, so a Mini, a Studio
and your laptop answer in one conversation. Off the LAN it tunnels over
ssh; nothing listens on an inbound port. Business ships the other kind
of cluster: Thunderbolt, one model larger than any one box.
Start the daemon as toastd -l and it logs locally. The log lives on
your machine and never comes to us. Paired with a local provider or your own API
key — where the traffic never touches our servers either — that makes toastd usable
as an inference gateway for a practice that has to answer for its records: a dental
office, a law firm, anyone whose prompts are privileged.
Yes. A seminar is $2995 for up to eight people, run by a forward-deployed engineer — in San Mateo, or at your office with travel added. Your repositories, your workflows, and a plain answer about which of the six levels each of your tasks belongs on. Book one.
The binaries are macOS and Linux. On Windows the simplest path is SSH: a
Hosted shell puts toast, jam and ito on our metal, reachable
from PowerShell, Windows Terminal or anything else that speaks ssh. WSL
is Linux, so the installer works there too. Or press ` on this page and
talk to toast in the browser.
Everything runs under an account. Tools is $129 a year and includes the four binaries with local and BYOK inference. Solo is $20 a month and adds credits for inference we host and the slices; inference is charged by use, and unused credits roll over and expire a year after they were bought. Local providers and your own API keys never reach our servers, so they spend no credits at all. We collect anonymized usage (model, token count) — never your prompts. Seminars, consulting and FDE work are priced separately.
LinuxToaster sells AI labor the way cloud computing sells compute: specialized capacity, invoked when needed, composed by the customer, and paid for according to use.
LinuxToaster is built by Dirk Harms-Merbitz, who has run stochastic software next to deterministic software for thirty years: a style transfer for comparative literature at twenty; a black box that tried ten thousand parameter combinations in under a second to fit itself to whatever sensor it was bolted to, and flew in Dutch army helicopters and sat on Mercedes and BMW test benches; a datacenter in Los Angeles; routers, because internet traffic is stochastic and packets behave like particles. Then the pipe, the loop, a shell that survives what a model writes, and a local inference engine. The book in the week above was the test case. LinuxToaster is what fell out.
AI is the marketing word. Stochastic is the engineering one.
Abstractions are how you climb. Each layer up is one more layer between you and the machine. Tools are how you touch the thing itself, and that is where the fun is.
Reality through tools over status through abstractions.