coconutlabs

13 commits this week10 repos trackedlatest result 2026-08-08kvwarden v0.1.6 on pypi

systems engineering · applied ai

A quiet tenant keeps its latency under load. I build the schedulers that make that true and measure them on rented hardware. The parts that failed are published too.

quiet-tenant p99 TTFT

flooder 32 rps · 300 s

solo
53.9 ms
kvwarden
61.5 ms
fifo
1,585

1× A100 · Llama-3.1-8B · vLLM 0.19.1 · n=311 post-warmup · FIFO bar off scale. full provenance

ratio to solo

1.14×

vs fifo tail

26×

harness

public

engineers

two

Coconut Labs works on the shared layer of inference: scheduling, fairness, cache pressure, and the measurements that keep claims honest.

The lab is small by design. Fewer abstractions between the benchmark, the note, and the code.

1,585 ms of waiting for a prompt that costs 53.9 ms with nothing else on the box.

the lab

One engineer, close to the work.

Coconut Labs is the name I publish under. The work happens in the open at github.com/coconut-labs and shows up here when there is a result worth standing behind. Jay Patel has built alongside me on several of these and is credited on them.

How I work

Building something at this layer? Write us.