FAQ

How does it work?

Lightweight toast talks to a local toastd, which keeps an HTTP/2 connection pool to linuxtoaster.com. Written in C to minimize latency. With BYOK, toastd connects directly to your provider — your traffic never touches our servers. Shell is the default client, not the requirement: a Python script, a C program or a cron job can talk to toastd the same way.

What's a slice?

A specialist called by name — toast editor, toast security, toast paranoid. The word after toast picks a server-side config; toast --add editor drops a symlink if you want one word. The name selects a hosted, tuned configuration that spans providers. --balance lists the ones your account can reach, --list the ones you have. Plain toast uses your local .persona file instead. Only one of them is ever read, so they cannot fight.

How does a loop know when to stop?

toast exits 1 when the model prints DONE, so while toast ... ends when the work is finished, not when a counter runs out. 20 while caps it anyway. Tool calls run through jam, five rounds max, against a .tools allowlist you write yourself.

Where's my data stored?

Locally. Context in .crumbs, conversations in .chat, tool permissions in .tools, your own slice in .persona. Version them, grep them, delete them. Memory you cannot read is memory you cannot correct.

What's appled?

On MacBooks, the installer downloads appled — a local inference provider using Apple Intelligence. No API key and no cost per token — prompts never leave your machine, and running it spends none of your credits. It still needs an account: toast brings up a Stripe page the first time you run it.

Can I run it fully offline?

Yes — appled, toasted, Ollama, MLX, LM Studio, KoboldCpp, llama.cpp, vLLM, LocalAI, or Jan. No internet, no API keys, full privacy.

What's BYOK?

Got a PROVIDER_API_KEY set for Anthropic, Cerebras, Google Gemini, Groq, OpenAI, OpenRouter, Together, Mistral, Perplexity, or xAI? Use toast -p provider. Zero config, zero cost from us.

What's toasted?

A from-scratch local inference daemon for Apple Silicon (Pro tier). Loads Qwen3-Next-Coder — a 30B coding model — via C++ against Apple's MLX API. ~100 tok/s generation, ~400 tok/s prefill, 0.6s to first token. 128 GB supports 8/6/4-bit quantization, 64 GB supports 4-bit.

Several machines?

squawkd joins a multicast group on the LAN and squawk bot Paranoid puts a model in the room as another participant, so a Mini, a Studio and your laptop answer in one conversation. Off the LAN it tunnels over ssh; nothing listens on an inbound port. Business ships the other kind of cluster: Thunderbolt, one model larger than any one box.

Can I keep an audit trail?

Start the daemon as toastd -l and it logs locally. The log lives on your machine and never comes to us. Paired with a local provider or your own API key — where the traffic never touches our servers either — that makes toastd usable as an inference gateway for a practice that has to answer for its records: a dental office, a law firm, anyone whose prompts are privileged.

Do you train teams?

Yes. A seminar is $2995 for up to eight people, run by a forward-deployed engineer — in San Mateo, or at your office with travel added. Your repositories, your workflows, and a plain answer about which of the six levels each of your tasks belongs on. Book one.

macOS? Windows?

The binaries are macOS and Linux. On Windows the simplest path is SSH: a Hosted shell puts toast, jam and ito on our metal, reachable from PowerShell, Windows Terminal or anything else that speaks ssh. WSL is Linux, so the installer works there too. Or press ` on this page and talk to toast in the browser.

How does billing work?

Everything runs under an account. Tools is $129 a year and includes the four binaries with local and BYOK inference. Solo is $20 a month and adds credits for inference we host and the slices; inference is charged by use, and unused credits roll over and expire a year after they were bought. Local providers and your own API keys never reach our servers, so they spend no credits at all. We collect anonymized usage (model, token count) — never your prompts. Seminars, consulting and FDE work are priced separately.

Who built it

LinuxToaster sells AI labor the way cloud computing sells compute: specialized capacity, invoked when needed, composed by the customer, and paid for according to use.

LinuxToaster is built by Dirk Harms-Merbitz, who has run stochastic software next to deterministic software for thirty years: a style transfer for comparative literature at twenty; a black box that tried ten thousand parameter combinations in under a second to fit itself to whatever sensor it was bolted to, and flew in Dutch army helicopters and sat on Mercedes and BMW test benches; a datacenter in Los Angeles; routers, because internet traffic is stochastic and packets behave like particles. Then the pipe, the loop, a shell that survives what a model writes, and a local inference engine. The book in the week above was the test case. LinuxToaster is what fell out.

AI is the marketing word. Stochastic is the engineering one.

Abstractions are how you climb. Each layer up is one more layer between you and the machine. Tools are how you touch the thing itself, and that is where the fun is.

Reality through tools over status through abstractions.

Keep me in the loop

Product updates, new features, the occasional blog post. No spam.

Unsubscribe anytime.

Launchpadly Startup Directory Featured on tools.cafe