Local AI for Apple silicon, written from the arithmetic up.
Kvad runs language, image and video models on your own Mac, from one install. It has its own
kernels for the M5's GPU, an OpenAI-compatible API, and a codebase you can read from the first
matrix multiply to the last.
v0.11.0macOS on Apple silicon, Linux on the CPUMIT
One engine behind a browser, a terminal, a shell and an API.
A web UI that is part of the binary
Twelve pages over the same server the API runs on: chat, a playground, images, videos, training, evals, benchmarks and monitoring. Nothing to install beside it, and it is there when the service starts.
Search Hugging Face, download, load and chat without leaving the terminal. One key cycles a model through six backends, which is the quickest way to feel what quantisation costs and what the GPU buys.
Models Chat Qwen/Qwen2.5-1.5B-Instruct · 1544M · metal q8 1641MB · chat
┌ conversation ────────────────────────────────────────────────────────────────────┐
│you ▸ │
│ Why is decoding faster with 8-bit weights? Answer in three sentences. │
│ │
│ai ▸ │
│ Decoding is faster with 8-bit weights because it reduces the amount of floating │
│point operations required, leading to a more efficient computation process. │
│Additionally, 8-bit weights use fewer bits per channel, which allows for a more │
│compact storage format, reducing memory usage. This results in faster data │
│processing and decoding times. │
│ │
│ │
│ │
│ │
│ │
└──────────────────────────────────────────────────────────────────────────────────┘
┌ message ─────────────────────────────────────────────────────────────────────────┐│▌│└──────────────────────────────────────────────────────────────────────────────────┘ loaded Qwen/Qwen2.5-1.5B-Instruct [128.4 tok/s] enter send esc stop ^l clear pgup/pgdn scroll tab models ^c quit
A command line that knows about the server
Every route the API has is a command. With the service running, the command line sends its work there instead of loading a second copy of the weights, and the first line of output says which happened.
$ kvad search qwen2.5MODEL DOWNLOADS SIZE ARCH STATUSQwen/Qwen2.5-7B-Instruct 8891777 14.2 GB llama chat · fits at f32Qwen/Qwen2.5-0.5B-Instruct 8661907 942.3 MB llama chat · fits at f32Qwen/Qwen2.5-1.5B-Instruct 7509634 2.9 GB llama downloadedQwen/Qwen2.5-32B-Instruct 1920796 61.0 GB llama chat · fits at q8$ kvad load Qwen/Qwen3-14Bloaded Qwen/Qwen3-14B on metal q8 llama · 40 layers · 40 heads (8 KV heads, 5x grouped) · 5120 embd · 40960 ctx 14768.3M parameters · weights 14.6 GB · instruction-tuned$ kvad psMODEL CHARGED CONTEXT CACHEDQwen/Qwen3-14B@gpu-q8 19.6 GB 40960 0 toolsmemory: 16.4 GB of 36.0 GB left · each charged for 32768 tokens of KV cache
OpenAI's API, on your own address
Chat completions with streaming, reasoning and tool calls, plus images and video, in the shapes existing SDKs already send. Point a client at port 5823 and change nothing else.
/v1/images/generationswith steps, seed, LoRAs and a preview per step
/v1/videosa job you start, follow and fetch
/v1/modelssays which models can take tools
What it runs
Words, pictures and video, from the weights people already publish.
Language
Six architectures, loaded straight from Hugging Face safetensors: the Llama family (Llama,
Mistral, Qwen 2 to 3, SmolLM2), DeepSeek V2 and V3, Qwen3.5 and Qwen3-Next, and GPT-2.
On the CPU or on Metal, at f32, bf16, q8 or q4, with a KV cache that survives between turns.
Made by Kvad for this page: SDXL at 1024², 30 steps, about a minute and a half each on an M5
Pro. The prompts and seeds are in the image guide.
Playground
See what the model was choosing between.
Generation computes a probability for every token in the vocabulary at every step, and then
throws all but one away. The playground keeps them. Each token is coloured by how sure the
model was, and a click shows what else was in the running.
Beside it: the tokeniser, and two models answering the same prompt.
127.0.0.1:5823
Benchmarks
Numbers you can reproduce, with a button.
Three of Kvad's own published numbers were once wrong, each from timing a single run. So
the server has a benchmark built in, and it is strict: rounds interleave the variants, every
sample is kept, the median is shown with its range, and it refuses to measure a machine that
is busy with something else.
Each loaded model is charged for its weights and a full context of KV cache, so a model that
was admitted can always finish its conversation. A load that does not fit is refused with the
name of what is in the way. Nothing is ever unloaded behind your back.
A character-level GPT on any text file, by a training loop written out in the repository
with no framework under it. The loss curve is drawn as it falls, with what the model writes
at each checkpoint beside it.
The M5's GPU has matrix units. Most software never reaches them.
Ordinary Metal shaders do not use them. Kvad has kernels that do, for dense and for
quantised weights, and falls back to the standard path on earlier chips without being told.
Kvad has two jobs. One is to be worth running. The other is to explain itself: every matrix
multiply, every derivative and every attention head is code in the repository, with the
reasoning for each constant next to it.
/// RMSNorm: like LayerNorm, but without the centring step and without a bias.////// LayerNorm subtracts the mean, then divides by the standard deviation./// Someone eventually checked whether the subtraction was doing any work, and/// it was not: only the rescaling matters. Dropping it saves a pass over the/// data and a whole bias vector per normalisation, of which a transformer has/// two per layer.////// This is worth noticing as a pattern. A lot of "architecture progress" is/// finding out that a piece everyone inherited from the previous paper was/// never load-bearing.pubfnrms_norm(x: &[f32], weight: &[f32], eps: f32) ->Vec<f32> {
letmean_square = x.iter().map(|v| v * v).sum::<f32>() / x.len() asf32;
letinv = 1.0 / (mean_square + eps).sqrt();
x.iter().zip(weight.iter()).map(|(&v, &w)| v * inv * w).collect()
}