theophilusjohn.com
#enrg-0001/042026 · Sole engineer

Inference
in the browser

Browser-native LLM inference engine in WGSL compute shaders. Zero server cost.

43.1
tok/s at 2048 ctx
334.9 MiB
weights resident
5.63x
compression
0
server cost

Enargeia runs Qwen2.5-0.5B-Instruct entirely in the browser. The visitor’s GPU does the inference, so serving it costs nothing beyond static files — no server, no API key, no per-token cost. Every kernel is WGSL running on WebGPU. Generating a token is one pass up through the model’s layers, dispatched in order, and every one of those dispatches runs on the machine in front of you.

The model is quantized to int4, except the tied embedding and LM head which ship at int8 — 334.9 MiB resident against 1884.6 at fp32, for 13.5% worse perplexity. Decode runs at 45.5 tokens/second at 512 context and 43.1 at 2048, flat across context length. Time to first token is 219 ms on a short prompt. Cold load over the CDN is 14.6 seconds; warm is 1.4.

Measured on an Apple M2 with a 10-core GPU, headless and with nothing else on the GPU; a foreground tab shares it with the compositor and measures lower.