Colibrì is a small inference engine that runs very large mixture-of-experts models on ordinary hardware. The engine is one C program per model family. At runtime it links nothing else: no PyTorch, no BLAS, no Python. Python is used once, to convert a checkpoint, and again only if you want the optional API gateway. The program that writes the tokens is C.
The project is github.com/JustVugg/colibri, Apache-2.0, by Vincenzo Fornaro. The site is justvugg.github.io/colibri. This article follows the method published there. Storage, RAM, and video memory are treated as one hierarchy. A token stages only the weights its router actually picks. A smaller machine is slower. The router’s choices, and the precision of the weights, stay the same.
What “running the model” actually is
A reply is not produced in one shot. The model writes one token at a time, then feeds that token back in and writes the next. A token is a piece of text: a word, a piece of a word, or a mark of punctuation. GLM-family vocabularies here hold 154,880 of them.
Each step looks like this:
- The newest token id is turned into a vector, a list of 4,096 numbers, by looking up one row of the embedding table.
- That vector walks every layer. Each layer mixes it with what the conversation has already said, then runs it through a small feed-forward network. In the later layers that network is a handful of specialists, not the whole file.
- The vector that comes out of the last layer is scored against every word in the vocabulary. The winning id is the next token. It is appended, and the walk repeats.
The first walk, over your whole prompt, is called prefill. Every walk after that, one new token at a time, is decode. Prefill is heavy because many positions are live at once. Decode is where a small machine spends its life, and decode is where the disk shows up: each new token can hire specialists that were not hired by the token before it.
Why the file is huge and the work is small
A mixture-of-experts layer does not run every feed-forward network it contains. A small router scores the specialists and keeps the top few. The others are real parameters. They sit in the checkpoint for the tokens that will need them later.
Think of a hospital. A few doctors are on every case: that is the shared expert, plus the attention block, which every token passes through. Most specialists stay in their offices until a case names them. Eight hires out of a few hundred is why a file of hundreds of billions of parameters behaves, while it is writing, like a much smaller network.
The reference model in Colibrì is GLM-5.2: about 744 billion parameters, with about 40 billion active on a given token. That is roughly 5.4% of the file. Of those active weights, only about 11 GB changes from one token to the next. That changing slice is the routed experts. Attention, the shared expert, and the embedding table stay where they were loaded.
Only a small slice of a 744B model is active per token. Diagram from the Colibrì README.
Count the specialists on GLM-5.2 and the size stops being abstract. There are 75 mixture-of-experts layers, 256 routed experts in each, plus the experts on the extra multi-token head: 19,456 routed experts. At 4-bit packing each one is about 19 MB. Together they are about 370 GB on disk. The dense trunk — attention, shared experts, embeddings, about 17 billion parameters — stays resident in RAM at 4-bit, about 9.9 GB.
One expert is three matrices, not a mystery blob. They are usually called gate, up, and down. The hidden vector is multiplied by the first two, passed through a gated activation (a SwiGLU-style nonlinearity: one branch is a gate that opens and closes the other), then multiplied by the third matrix back to the hidden width. Those three matrices are stored next to each other in the file, so a cache miss is one read, not three round trips.
The model therefore does not have to fit in fast memory. It has to be placed. Fast memory holds what every token touches. The drive holds the specialists. A cache in between holds the specialists this conversation has touched lately.
A JIT compiler, except the program is the weights
A just-in-time compiler does not compile a whole program before it starts. It watches what actually runs and compiles those hot paths when they are needed. Colibrì makes the same bet about a huge parameter space. Weights are data to be staged across video memory, RAM, and the SSD at the moment the router proves they are needed.
The bet is safe only because routing has structure. The project maps that structure. On GLM-5.2, 13,260 experts have been characterised into regions such as Python, SQL, mathematics, poetry, law, and Chinese. Position in that map is measured routing affinity, not a learned embedding. A Python-heavy chat keeps hiring the Python region. A cache can learn that. A cache cannot learn a coin flip, and expert choice is not a coin flip.
The engine writes the record to .coli_usage, updated every turn, and pins the hottest experts into RAM so they stop taking a disk trip. That pin set is a policy, not a promise: it helps a repeated workload, and it can overfit one prompt. The project treats it as something you measure, which is why it can be turned down and why a held-out chat can look different from the chat the pins were trained on.
The range of machines is the same code path. On a 25 GB laptop almost every specialist streams from disk. On a large host the entire expert set can be pinned (PIN_GB=all, and CUDA_EXPERT_GB=auto when a GPU is there) and the disk leaves the decode path. Both produce the router’s tokens. The laptop produces them more slowly.
The five steps, more slowly
Every layer of every token walks the same five steps. Placement decides speed. It does not decide which expert fires, and it does not decide how many bits that expert was stored with.
Route, union, place, overlap, learn. Diagram from the Colibrì README.
1. Route
The layer already holds a small scoring matrix in RAM, because the router runs on every token and is tiny next to an expert. It turns the hidden vector into one score per specialist. GLM-5.3-Flash uses a sigmoid score, a correction bias (e_score_correction_bias, the no-aux load-balancing trick), and a routed scaling factor of 2.5, then keeps the top 8 out of 288. GLM-5.2’s layers hire from 256. The shared expert is not in that contest. It runs every time, so common knowledge is not left to whichever specialist won the lottery.
The output of the layer is a weighted mix of those hired experts, added back into a residual stream. On these GLM and DeepSeek-style models that residual is richer than a single skip wire: several streams travel side by side (an expansion of 4), get mixed by a small learned matrix, and are pulled back onto a legal mixing by 20 Sinkhorn iterations. That is manifold-constrained hyper-connections. Every token pays that cost. It is part of the dense trunk, so it stays in RAM.
2. Union
During prefill, and during speculative verification, many positions are live in the same layer. Each position names up to eight experts. Lots of those names repeat. The engine builds the set of unique experts for the whole batch and reads each one once. That is the batch-union. Eight tokens that all hire expert 17 do not cause eight disk reads of expert 17.
The same idea shows up if you spread the work across machines. A coordinator can keep token generation, routing, and the attention state on one computer, and send a layer’s whole union to another machine as one request. A token does not pay a network round trip per expert. The single-machine path is unchanged when no workers are configured.
3. Place
For each expert in the union the engine asks a simple question: where is the fastest copy I already have?
- If it is in video memory, use that.
- If it is in the RAM cache, use that.
- If a pinned hot set holds it, use that.
- Otherwise read it from the SSD into a cache slot, then use it.
The RAM cache is per layer, and it is an LRU: the specialist that has gone longest without being hired is the one that gets evicted when a new one arrives. Pins sit beside that LRU. They are the experts the usage file says this workload always wants, so a cold LRU does not throw them out. On a multi-socket machine, COLI_NUMA=1 spreads the resident weights across memory controllers so one socket is not the bottleneck for the trunk that every token touches.
4. Overlap
A miss costs milliseconds. A matrix multiply on weights that are already hot costs less than that wait if the disk is the slow part. The engine tries not to do the wait and the math one after the other.
- The three matrices of an expert are adjacent and come in with one
pread. - A bounded async pool (
PIPE=1, the default) loads the missing experts while the experts already in RAM are multiplying. - A lookahead thread (
PILOT=1) runs the next layer’s router before the current layer has finished, and starts fetching what that router names. On the published measurement, routing is about 71.6% predictable one layer ahead. Lookahead can lose on some hosts, so it stays a switch. - On a GPU, a resident pipeline (
COLI_CUDA_PIPE=2) keeps the residual stream on the device across layers, so the CPU expert loop does not keep hauling the hidden state back and forth. Apple Silicon can run the batched expert math on the unified-memory GPU through Metal. Vulkan covers the expert tier, the dense projections, and the attention core on any GPU with a Vulkan 1.2 driver, including AMD cards through Mesa, and older cards whose vendor stack is gone.
Direct I/O is the other disk switch. DIRECT=1 uses O_DIRECT, which skips the operating-system page cache and talks to the drive more directly. On a drive with its own DRAM cache and spare bandwidth, the project measured about +34% decode with PIPE=1 on one Blackwell Windows box, and a jump from 4.25 to 9.69 GB/s in a raw I/O test on an NVIDIA GB10. On a cheap QLC drive, a DRAM-less drive, or a virtual disk, the same switch can do nothing or slow you down. You try it on the machine you have.
5. Learn
After the turn, the engine writes which experts fired. The next turn’s pin set and, if you ask for it, the next partial copy of the model on a second drive, are ranked from that history. The first few conversations of a new topic are the cold ones. Later conversations of the same kind find their specialists already in RAM.
Three shelves, and a second copy of the library
VRAM, RAM, and NVMe hold the same experts at different speeds. Diagram from the Colibrì README.
NVMe holds the full set. It is the source of truth and the slow shelf. RAM holds the dense trunk plus the LRU and the pins. VRAM is optional and holds whichever experts you budget for it. CUDA, Metal, and Vulkan are backends on that same runtime. A GPU never changes the token. It changes how fast a resident expert’s multiply finishes, and how many experts can avoid the drive.
When decode is waiting on one drive, a second SSD can hold another copy. You point at it with COLI_MODEL_MIRROR. At startup the engine checks each file’s size and safetensors header against the primary. A file that differs, or that the second disk does not have, stays on the first disk, so a smaller second SSD can serve only the shards it holds. The mirror is never written: .coli_usage, .coli_kv, and the other sidecars stay on the primary. A read error on the mirror falls back to the primary with a warning. Both copies are byte-identical, so which drive served an expert cannot change the token.
The engine measures the drives and weights the split. Buffered reads assign experts to drives deterministically. Eligible direct reads can stripe one expert across the copies. On the project’s measurements, two independent NVMe drives gave about +37.5% decode on GLM-5.2. A slower third drive, after the weighting, was neutral. Shared controllers, a warm page cache, and a machine that is already compute-bound all eat the gain. Bandwidth headroom is the thing you bought. A guaranteed multiplier is not.
If the second drive cannot hold the whole model, the usage file already knows which shards are worth copying. coli mirror plan reads the safetensors headers, ranks shards by the hottest experts, and respects a free-space reserve. coli mirror stage copies through temporary files, checks every shard with SHA-256, leaves any existing mirror shard in place, and publishes a receipt only when that selection is complete. The primary model is not modified.
Attention, and why the conversation state is small
The feed-forward specialists are most of the disk. Attention is most of the “remember the chat” cost, and these models spend a lot of design on making that cost small.
Full attention lets a token look at earlier tokens. Naively that means storing a key vector and a value vector for every previous token, at the full hidden width, in every layer. Multi-head latent attention (MLA) does not store that. It stores a compressed latent, 576 numbers per token instead of 32,768, about 57 times smaller. The layer reconstructs what it needs when it scores. Colibrì writes that compressed state to .coli_kv. A conversation can reopen with no re-prefill, and the reopened session matches an uninterrupted one byte for byte.
Sparse attention (the lightning indexer, DSA) goes further. Even with a compressed cache, scoring every past token is expensive once the chat is long. The indexer scores pools of keys, keeps the top pools, and only then runs attention on the tokens inside them. Forcing the indexer to keep every key reproduces ordinary dense attention exactly, which is how the project checks that the sparse path is the same function with a shortlist, not a different model.
Linear attention is the other pattern, used heavily in GLM-5.3-Flash. Instead of a growing cache, the layer keeps a fixed-size state and updates it as each token arrives: a short convolution, a decay, and a small recurrent update. The cost per token stays flat as the prompt gets longer. Flash alternates them. The pattern is three linear layers, then one full-attention layer, repeating. Across the stack that is 34 linear layers and 11 full layers, and the engine follows config.layer_types rather than guessing. Position lives in the linear layers’ decay and short convolution. The full layers are NoPE: they do not add rotary position on top.
The forward pass is checked against a Transformers reference. Teacher-forcing on the tiny oracle is typically 30 to 32 tokens out of 32. The two misses are floating-point near-ties that depend on the compiler. A faster kernel has to earn its place by matching tokens on a real run, not by winning a microbenchmark.
Guessing the next tokens, then checking them
Decode’s painful part is the serial loop: one token, one full walk, then the next token. Speculative decoding tries to walk several tokens in one visit. GLM-5.2 has a native multi-token prediction head that drafts a few future tokens. The main model verifies that draft in one batched forward. When the drafts are good, the project measures 2.2 to 2.8 tokens per forward.
Two rules come from broken experiments, and they ship as defaults. The draft head has to be 8-bit. A 4-bit head collapsed to roughly 0–4% acceptance, which means the draft was rejected almost every time and the extra work was pure cost. Draft and verify also have to compute the same function: SPEC_PIN=1 pins both to one kernel family, because a “faster” verify kernel that rounds differently will reject drafts that were actually right. Grammar-forced drafts, for a JSON schema or another constrained language, add acceptance that is nearly free because many tokens are already determined by the grammar.
Whether this is a win depends on how warm the expert cache is. A rejected draft still has to be rolled back, and the experts it touched may have been misses. Around an 85% expert hit rate, MTP has also measured a loss of about 32%. DRAFT=0 turns it off. GLM-5.3’s published container ships without that head, so speculation stays off there. DeepSeek V4 has drafters too, and they default off for the same reason: on the measured multi-turn chat, replaying the recurrent state after a rejection cost more than the accepted drafts saved.
Inside GLM-5.3-Flash, layer by layer
The same placement runs nine families. Flash is the one people reach for when the 744B checkpoint is more disk than they have: about 321 billion parameters, about 40 billion active, vision included. The converted container is about 195 GB. The official RAM note is 25 GB, roughly 12 GB of dense weights at 4-bit plus room for an expert cache.
Read off the released checkpoint, the text stack is:
- 45 text layers, plus one multi-token layer at index 45. Hidden size 4096. Vocabulary 154,880.
- Hybrid attention. The repeating pattern is linear, linear, linear, full. The full layers sit at indexes 3, 7, 11, and so on through 43. Each full layer carries its own lightning indexer. Keys are grouped into pools of 4, the pools are scored, the winning pools are expanded back into token indexes, and the incomplete tail pool is always kept.
- Mixture of experts from layer 3. 288 routed experts, top 8, plus one shared expert. Each expert’s inner width is 2048. The first three layers are dense, with inner width 12288, so they have no router and no disk traffic. They are part of the resident trunk.
- Hyper-connections at both the attention site and the feed-forward site of every layer, expansion 4, 20 Sinkhorn steps.
- Vision, when you send a picture. A 24-block tower, hidden size 1024, 14-pixel patches on a 448-pixel image, then a spatial merge of 2. Patches are projected to 4096 and written into the text stream where the image tokens sit. A text-only prompt never enters the tower. The tower is about 0.1–0.2 GB, small next to the experts, and it can be left unloaded.
A single Flash expert is smaller than a GLM-5.2 expert, because the inner width is 2048 rather than the larger GLM-5.2 expert. In the 4-bit group-scaled container a Flash cache slot is about 14.2 MB. That is the unit the LRU deals in. A 1 GB cache is on the order of seventy of those slots, spread across the 42 sparse layers, which is only a few specialists remembered per layer. That is why RAM above the dense trunk turns so directly into tokens per second: every extra slot is an expert that does not wait on the SSD for the next token that wants it.
How the weights are packed
Full precision would multiply the disk copy several times over. The converter splits the checkpoint on purpose:
- Routed experts are stored as 4-bit integers with a scale per group of 64 weights (int4-gs64). These are almost all of the bytes. Against the original 8-bit floating-point experts, cosine similarity on real weights measures about 0.994. Group scales matter: older per-row 4-bit packs of the GLM-5.2 family measured about 9 points worse on quality and were tied to generations that would not stop. The group-scaled container is the one the project tells you to run.
- Everything else — attention, shared experts, embeddings, the output head — stays in bfloat16 on disk and is quantized when the engine loads. Those tensors are about 3% of the bytes. Keeping them exact on disk costs on the order of 14 GB and means a later decision about dense precision does not require downloading the checkpoint again.
Quantization means each weight is stored with fewer bits, plus a scale that rebuilds an approximation at multiply time. The layer still multiplies matrices. It multiplies small integers and rescales, instead of hauling a 16-bit or 32-bit float for every specialist on the drive.
The download itself is shard by shard. The public Flash checkpoint is on the order of 171 GB of expert shards. The converter fetches one shard, writes the Colibrì container, and can drop the source shard. A crash in the middle is resumable. You do not need the raw checkpoint and the converted model fully resident at the same time, though you do want on the order of 180 GB free while both exist.
Fitting Flash into 16 GB
The stock Flash engine already streams experts. What it keeps resident is still large: about 9.4 GB of dense weights at 4-bit (attention, shared experts, the three dense layers), about 0.36 GB for the output head, and — this is the surprise — the embedding table in 32-bit floats, 2.54 GB, because a word lookup was never run through the quantized matrix path. Add the compressed attention state (about 0.30 GB at an 8,192-token context), the linear-attention recurrent state (about 0.15 GB), and the vision tower, and the resident set is about 11.6 GB.
Loading makes it worse for a moment. The old loader expanded a whole tensor into 32-bit floats, quantized it, and only then freed the temporary. That spike is about 2.5 GB sitting beside the tensor it is replacing, so peak memory lands near 14 GB. On a 16 GB machine the operating system still needs a share. The expert cache is whatever is left, and “whatever is left” can be one slot per layer — eight hires, one slot — which means the disk does almost all of the work.
A separate low-memory build of the same engine, c/lowram, keeps the router and the math and spends less RAM on the trunk. The first three changes reproduce the stock numbers. The last two are optional and say so, because they change accuracy.
- The word table becomes a quantized matrix. At 8 bits per row it is 0.63 GB instead of 2.54 GB. A single row, 16 KB, is expanded when that word is needed. 8-bit is the default on purpose: this table is never multiplied, so its error is not averaged across 4,096 columns. It is the input to layer one. Setting the bit width back to 32 reproduces the stock table exactly. 4-bit would shrink it to 0.36 GB and is the aggressive setting.
- Load in row blocks. The loader quantizes a block of rows and releases it before the next block. The 2.5 GB spike goes away. The bytes that land in the matrix match the old loader bit for bit, because quantization was already per row.
- Skip vision for a text-only service. That is a few hundred megabytes and one less tower to fault in.
- Optional 3-bit experts. As a cache slot fills, the engine can repack that expert from 4-bit to 3-bit. A slot drops from 14.16 MB to 11.01 MB, about 29% more slots in the same RAM. On one measured expert, relative error versus full precision rises from about 0.10 to about 0.27. Off unless you set it.
- Optional lattice packing. A different container at about 3.06 bits per weight shrinks a slot to 9.63 MB. On the same measurement it errs less than plain 3-bit (relative error about 0.19). Repacking on a miss is slower. Also off unless you set it.
After the first two changes, a 14 GB budget on a 16 GB host is a realistic working set, and the leftover RAM is cache slots. The specialists still live on the SSD. The dense trunk is just small enough that the cache is a cache, not a single chair.
What the speed looks like when the shelves change
Same engine, same 4-bit container. The hardware changes where the experts live. Diagram from the Colibrì README.
Published highlights for that GLM-5.2 container:
- Six RTX 5090s, every expert resident: about 5.8–6.8 tokens per second, time to first token around 13 seconds.
- A 128 GB CPU desktop, cache warm: about 1.8 tokens per second.
- One RTX 5070 Ti, GPU-resident pipeline: about 1.07 tokens per second.
- The 25 GB laptop where the project started, cold: about 0.05–0.1 tokens per second. Slow, and the tokens are still the model’s tokens. That floor is the point of the design.
Disk speed dominates whenever the working set of experts does not fit in RAM. A hard drive or a network mount is a different machine from an NVMe SSD, even when the CPU is identical. Expect a fraction of a token per second while the cache is cold, and a few tokens per second once the pins and the LRU are warm and the drive is fast.
What you actually run
You need the program, a few hundred kilobytes, and the model, hundreds of gigabytes. coli chat is a terminal conversation. coli web opens the dashboard: the chat, Brio, the Brain page, and a profile of the last turns. coli plan prints where this machine will put the experts before you spend an afternoon loading them. coli doctor checks that the files you have are the files the engine expects.
Brio is worth a paragraph because it uses the same forward pass for a different job. You give the model a document and a closed list of answers. It reads the probability of each allowed answer and generates nothing. The reply includes an entropy, so “the model is not sure” is a number. On the published Qwen3.6 example, “request changes” landed at 99.9% with entropy 0.005, four tokens read, zero tokens written. Chat is unchanged for any request that does not ask for this.
The Brain page is the expert map made visible: one cell per specialist, colour for the shelf it currently sits on, a white flash for every expert hired during a turn. Profiling splits a turn into phases. One published Qwen3.6 trace on a CPU box shows 19 seconds of wall time for 36 prompt tokens and 55 generated tokens, 2.9 tokens per second, with 11.4 seconds of disk service overlapped with compute. That overlap is step 4 of the path, measured on a real reply.
Nine families share the front end and not the math: GLM-5.2 and 5.3, GLM-5.3-Flash, Inkling, Kimi K3, DeepSeek V4 Flash, DeepSeek V4.1 Flash, Qwen3.8-Flash-Next, Qwen3.6, and OLMoE. Each architecture is its own C file. The safetensors reader, the quantization decoders, the tokenizer, and the expert cache live in headers they all include, so a fix in the cache reaches every family. coli reads the model’s config.json and picks the binary. The command line does not change when the model does.
The short version is the whole method. A frontier mixture-of-experts model is large because it contains thousands of specialists. Each token hires a few of them, mixes them with a compressed memory of the chat, and writes one piece of text. Colibrì keeps the shared brain in RAM, leaves the specialists on disk, and stages the hired ones through RAM or a GPU just in time — the same way a JIT compiler only compiles the code that actually runs.
Source, measurements, and diagrams: github.com/JustVugg/colibri.






(0) Comments on "How a Frontier AI 321-billion-parameter model runs on a normal computer without GPU?"
* Most comments will be posted if that are on-topic and not abusive