An LLM engine in Rust

Local AI for Apple silicon, written from the arithmetic up.

Kvad runs language, image and video models on your own Mac, from one install. It has its own kernels for the M5's GPU, an OpenAI-compatible API, and a codebase you can read from the first matrix multiply to the last.

v0.11.0 macOS on Apple silicon, Linux on the CPU MIT

The chat page of Kvad's web UI: Qwen3-14B answering a question, with the decode rate and cached tokens under each replyThe chat page of Kvad's web UI: Qwen3-14B answering a question, with the decode rate and cached tokens under each reply
215 tok/s
Qwen2.5-1.5B decoding on Metal at q4. Median of five interleaved rounds.
2.5×
Faster prefill on the M5's matrix units: 2,079 tokens in 0.39 s, down from 0.99 s.
14.6 GB
Qwen3-14B in memory at q8. Quantised once, then mapped from disk on every load.
92 s
An SDXL picture at 1024², 30 steps, including the decode.

Measured on an M5 Pro with 48 GB. How, and what did not get faster

Four ways in

One engine behind a browser, a terminal, a shell and an API.

A web UI that is part of the binary

Twelve pages over the same server the API runs on: chat, a playground, images, videos, training, evals, benchmarks and monitoring. Nothing to install beside it, and it is there when the service starts.

The web UI
The dashboard: the loaded model, decode rate, time to first token, queue depth, memory and diskThe dashboard: the loaded model, decode rate, time to first token, queue depth, memory and disk

A terminal app for when you are already there

Search Hugging Face, download, load and chat without leaving the terminal. One key cycles a model through six backends, which is the quickest way to feel what quantisation costs and what the GPU buys.

The terminal app
 Models   Chat          Qwen/Qwen2.5-1.5B-Instruct · 1544M · metal q8 1641MB · chat 
┌ conversation ────────────────────────────────────────────────────────────────────┐
│you ▸                                                                             │
│  Why is decoding faster with 8-bit weights? Answer in three sentences.           │
│                                                                                  │
│ai  ▸                                                                             │
│  Decoding is faster with 8-bit weights because it reduces the amount of floating │
│point operations required, leading to a more efficient computation process.       │
│Additionally, 8-bit weights use fewer bits per channel, which allows for a more   │
│compact storage format, reducing memory usage. This results in faster data        │
│processing and decoding times.                                                    │
│                                                                                  │
│                                                                                  │
│                                                                                  │
│                                                                                  │
│                                                                                  │
└──────────────────────────────────────────────────────────────────────────────────┘
┌ message ─────────────────────────────────────────────────────────────────────────┐
│▌                                                                                 │
└──────────────────────────────────────────────────────────────────────────────────┘
 loaded Qwen/Qwen2.5-1.5B-Instruct  [128.4 tok/s]
 enter send esc stop ^l clear pgup/pgdn scroll tab models ^c quit

A command line that knows about the server

Every route the API has is a command. With the service running, the command line sends its work there instead of loading a second copy of the weights, and the first line of output says which happened.

Command reference
$ kvad search qwen2.5
MODEL                             DOWNLOADS      SIZE  ARCH   STATUS
Qwen/Qwen2.5-7B-Instruct            8891777   14.2 GB  llama  chat · fits at f32
Qwen/Qwen2.5-0.5B-Instruct          8661907  942.3 MB  llama  chat · fits at f32
Qwen/Qwen2.5-1.5B-Instruct          7509634    2.9 GB  llama  downloaded
Qwen/Qwen2.5-32B-Instruct           1920796   61.0 GB  llama  chat · fits at q8

$ kvad load Qwen/Qwen3-14B
loaded Qwen/Qwen3-14B on metal q8
  llama · 40 layers · 40 heads (8 KV heads, 5x grouped) · 5120 embd · 40960 ctx
  14768.3M parameters · weights 14.6 GB · instruction-tuned

$ kvad ps
MODEL                  CHARGED  CONTEXT  CACHED
Qwen/Qwen3-14B@gpu-q8  19.6 GB  40960    0       tools

memory: 16.4 GB of 36.0 GB left · each charged for 32768 tokens of KV cache

OpenAI's API, on your own address

Chat completions with streaming, reasoning and tool calls, plus images and video, in the shapes existing SDKs already send. Point a client at port 5823 and change nothing else.

API guide
$ curl http://127.0.0.1:5823/v1/chat/completions \
    -H 'content-type: application/json' \
    -d '{
      "model": "Qwen/Qwen3-14B",
      "messages": [{"role": "user", "content": "What is a KV cache?"}],
      "stream": true
    }'
  • /v1/chat/completionsstreaming, reasoning_content, tool calls
  • /v1/images/generationswith steps, seed, LoRAs and a preview per step
  • /v1/videosa job you start, follow and fetch
  • /v1/modelssays which models can take tools

What it runs

Words, pictures and video, from the weights people already publish.

Language

Six architectures, loaded straight from Hugging Face safetensors: the Llama family (Llama, Mistral, Qwen 2 to 3, SmolLM2), DeepSeek V2 and V3, Qwen3.5 and Qwen3-Next, and GPT-2.

On the CPU or on Metal, at f32, bf16, q8 or q4, with a KV cache that survives between turns.

Models and backends

Images

Stable Diffusion 1.5 and SDXL with their fine-tunes, FLUX.1-schnell and Qwen-Image, including the community's GGUF files.

LoRAs are applied per request, to a model that stays loaded. An SDXL LoRA can be trained on your own pictures.

Images

Video

LTX-2.5, with sound, from a prompt or from a picture, up to 120 frames a second.

A video is a job on the server: start it, close the page, and fetch it when it is done.

Video

Playground

See what the model was choosing between.

Generation computes a probability for every token in the vocabulary at every step, and then throws all but one away. The playground keeps them. Each token is coloured by how sure the model was, and a click shows what else was in the running.

Beside it: the tokeniser, and two models answering the same prompt.

The playground: a completion with each token coloured by the model's confidence, beside the sampler's settingsThe playground: a completion with each token coloured by the model's confidence, beside the sampler's settings

Benchmarks

Numbers you can reproduce, with a button.

Three of Kvad's own published numbers were once wrong, each from timing a single run. So the server has a benchmark built in, and it is strict: rounds interleave the variants, every sample is kept, the median is shown with its range, and it refuses to measure a machine that is busy with something else.

Benchmarks and evals
The benchmarks page: Qwen2.5-1.5B on three backends, with the median decode rate and the range of five rounds for eachThe benchmarks page: Qwen2.5-1.5B on three backends, with the median decode rate and the range of five rounds for each

Memory

Several models, one honest budget.

Each loaded model is charged for its weights and a full context of KV cache, so a model that was admitted can always finish its conversation. A load that does not fit is refused with the name of what is in the way. Nothing is ever unloaded behind your back.

What is in memory
The models page: Qwen3-14B in memory, charged 19.6 of 36 GB, above the list of models on the machineThe models page: Qwen3-14B in memory, charged 19.6 of 36 GB, above the list of models on the machine

Training

Train something small and watch it learn.

A character-level GPT on any text file, by a training loop written out in the repository with no framework under it. The loss curve is drawn as it falls, with what the model writes at each checkpoint beside it.

Train your own
The training page: a finished run with its training and validation loss curves, and the text the model wrote at each checkpointThe training page: a finished run with its training and validation loss curves, and the text the model wrote at each checkpoint

Apple silicon

The M5's GPU has matrix units. Most software never reaches them.

Ordinary Metal shaders do not use them. Kvad has kernels that do, for dense and for quantised weights, and falls back to the standard path on earlier chips without being told.

On an M5 ProStandard MetalKvad's kernels
Prefill, Qwen2.5-1.5B, 2,079 tokens0.99 s0.39 s2.5×
FLUX.1-schnell, a step at 1024², q812.9–13.7 s8.7–9.2 s1.5×
SDXL, a step at 1024², f163.90–3.98 s3.09–3.15 s1.26×
Decode, Qwen2.5-1.5B, q8, a short context114–117 tok/s126–129 tok/s1.1×
Kvad on Apple silicon

How it works

An engine you can read.

Kvad has two jobs. One is to be worth running. The other is to explain itself: every matrix multiply, every derivative and every attention head is code in the repository, with the reasoning for each constant next to it.

  1. 1nervusA neural network and backpropagation, from scratch. No dependencies.
  2. 2kvadTransformer inference by hand: six architectures, quantisation, the kernels.
  3. 3kvad-gpuThe same forward passes on Metal, and the image and video models.
  4. 4kvad-tuiThe terminal app.
  5. 5kvad-serveThe server and the web UI.
Reading the engine
/// RMSNorm: like LayerNorm, but without the centring step and without a bias.
///
/// LayerNorm subtracts the mean, then divides by the standard deviation.
/// Someone eventually checked whether the subtraction was doing any work, and
/// it was not: only the rescaling matters. Dropping it saves a pass over the
/// data and a whole bias vector per normalisation, of which a transformer has
/// two per layer.
///
/// This is worth noticing as a pattern. A lot of "architecture progress" is
/// finding out that a piece everyone inherited from the previous paper was
/// never load-bearing.
pub fn rms_norm(x: &[f32], weight: &[f32], eps: f32) -> Vec<f32> {
    let mean_square = x.iter().map(|v| v * v).sum::<f32>() / x.len() as f32;
    let inv = 1.0 / (mean_square + eps).sqrt();
    x.iter().zip(weight.iter()).map(|(&v, &w)| v * inv * w).collect()
}
crates/llm/src/tensor.rs, as it is in the repository.

Not yet

What Kvad does not do.

You should know before you install it, not after.

  • One generation at a time. Requests queue. There is no continuous batching and no paged KV cache, so it serves one person well and a team badly.
  • The GPU is Metal only. Linux runs language models on the CPU. There is no CUDA build, and no images or video off a Mac.
  • Tool calls cannot be forced. tool_choice is auto or none, and only the Hermes-style format that Qwen uses is parsed.
  • No embeddings endpoint. Chat, images and video are the three that exist.
The roadmap

Install it, and ask it something.

Then kvad run --prompt "Why is the sky blue?", or open http://127.0.0.1:5823. Other ways to install

esc

Try , or .