whipbench whipbench GitHub
whipbench WHIP · WHEP · native clients, no browser · v0.0.1

whipbench

a vendor-neutral WebRTC benchmark over WHIP and WHEP

whipbench publishes a synthetic clip over WHIP (RFC 9725), plays it back with any number of WHEP viewers, and writes a report of what an operator cares about — join time, packet transit, packet loss, jitter, keyframe interval and run stability — with the same definitions whichever server sits in the middle. The clients are native Go on pion/webrtc v4, with no browser, so one machine can drive many viewers, and the report is meant to be laid next to another server's.

Status: 0.0.1, not released. The clients, the report and the in-process test relay work, and the first live run against MediaMTX is in evals/. MediaMTX does not negotiate the header extension 0.0.1's stamp travels in, so against it packet transit is reported unavailable — by design, never estimated. The delay method of the next milestone is decided (WB-1): one-way delay per frame, by frame fingerprint, which needs no header extension (WB-38).

Install

From source, with Go 1.27 or later (pure Go, CGO_ENABLED=0):

go install github.com/Allan-Nava/whipbench/cmd/whipbench@latest
whipbench version

The two synthetic clips are embedded in the binary; nothing else has to sit beside it. Release binaries for Linux and macOS (amd64 and arm64) with checksums will be attached to each tagged release.

Quick start

A MediaMTX on this machine, in Docker. The extra host makes it advertise 127.0.0.1 as an ICE candidate, which a client outside the container can reach through the published UDP port:

docker run --rm -d --name mediamtx -e MTX_WEBRTCADDITIONALHOSTS=127.0.0.1 \
  -p 127.0.0.1:8889:8889 -p 127.0.0.1:8189:8189/udp bluenviron/mediamtx:latest

# One publisher, 50 viewers arriving over 10 s, held for 30 s, then a report.
whipbench run examples/mediamtx-local-50.json --out reports/

Or the two halves separately:

whipbench publish --whip http://127.0.0.1:8889/bench/whip --include-loopback
whipbench view    --whep http://127.0.0.1:8889/bench/whep -n 20 --ramp 5s --duration 30s --include-loopback

run and view print the verdict and write whipbench-<UTC time>-<name>.json and .md. The exit status is 0 for a valid run, 3 for a run with no verdict, 1 for an error and 2 for a usage error, so a CI job can gate on it.

What it measures

Every viewer records its own numbers; the report adds aggregates over the viewers that joined. Definitions, because a benchmark number without one cannot be compared:

measure definition
join: first RTP packet WHEP POST sent → first RTP packet received
join: first keyframe WHEP POST sent → last packet of the first keyframe received complete (every sequence number from its first packet to its marker). No decoder runs, so this is the earliest moment a frame could be decoded, not a decoded frame. A viewer has joined when it gets here within the join timeout (10 s by default).
signalling WHEP POST sent → SDP answer received
loss expected = highest − first extended sequence number + 1 (RFC 3550 A.1); lost = expected − received; duplicates are not counted as received
jitter RFC 3550 §6.4.1 interarrival jitter, J += (|D| − J)/16 on every packet, in milliseconds
keyframe interval spacing of keyframe starts in RTP time — the GOP the server actually delivers
packet transit arrival time − the publisher's send-time stamp, per packet — not the one-way delay of a frame; see the next section
stability stalls (a gap of 500 ms or more between packets), viewers dropped after joining, publisher schedule slips, and a per-second timeline of active viewers and packets

Percentiles are nearest-rank: p95 is a value some viewer actually had. Join, loss and jitter take one value per joined viewer; packet transit pools every packet of every viewer in a histogram with 1% buckets (min, max and mean exact).

The publisher streams a pre-encoded clip, 640×360 at 30 fps, 4 s, a keyframe every 30 frames, in VP8 or constrained-baseline H.264. scripts/make-clips.sh records the exact ffmpeg command and remakes them. Frames go on the wire byte for byte, paced at the frame rate, and the loop is seamless: frame k has RTP timestamp base + 3000·k whichever pass it belongs to. NACK is negotiated, RTX is not, so a retransmission arrives on the original sequence space and the loss a viewer reports is what was never recovered.

Latency, and its limits

No figure whipbench reports is called "latency": the word covers too many different numbers to rank two servers by. WB-1 decided which ones whipbench measures; this section says what each is and what it is not.

Packet transit — what 0.0.1 reports. The publisher writes the wall-clock time each RTP packet is handed to the stack into the abs-capture-time header extension (http://www.webrtc.org/experiments/rtp-hdrext/abs-capture-time, a 64-bit NTP timestamp; pion implements its payload as rtp.AbsCaptureTimeExtension). A viewer subtracts that stamp from the arrival time of the same packet. With a pre-encoded clip the moment of capture is the moment of sending, which is why the field carries the send time.

What that number is, and what it is not:

One-way delay — the headline, decided by WB-1 and not built yet. Per frame: from the publisher writing the frame's first packet to the viewer receiving its last, one sample per complete frame. It is still network plus server forwarding, still not glass-to-glass, and it is the figure servers are ranked by. Its send instant has two sources, side by side and never averaged: the fingerprint — each viewer hashes the frame it reassembled and looks up when the publisher sent those bytes, which needs no header extension and so works through a server like MediaMTX (WB-38, the figure v0.1.0's report is built on) — and the abs-capture-time stamp, once per frame instead of per packet, which replaces packet transit (WB-39). Every figure will carry its topology, clock method, uncertainty and whether it can be compared with another report (WB-40), and its sample window and retransmissions (WB-41); between two hosts the clock offset comes from WB-3.

Capture-to-decode — later, in v0.6.0. A frame index drawn into the clips and read back from decoded VP8 keyframes on a sample of viewers (WB-2). It differs from one-way delay only where a server transcodes. A timestamp drawn live into the frames would need a live encoder, which whipbench does not have.

Scenarios

A scenario is a declarative JSON file, in the spirit of a k6 options block. Unknown fields are an error, so a typo cannot fall back to a default:

{
  "name": "mediamtx-local-50",
  "whip": "http://127.0.0.1:8889/bench/whip",
  "whep": "http://127.0.0.1:8889/bench/whep",
  "codec": "vp8",
  "viewers": 50,
  "rampSeconds": 10,
  "holdSeconds": 30,
  "warmupSeconds": 2,
  "joinTimeoutSeconds": 10,
  "includeLoopback": true,
  "metrics": "127.0.0.1:9464"
}

The publisher connects, the warmup passes, viewer i of n starts at i·ramp/n, and all of them stop together holdSeconds after the ramp ends. bearerEnv names an environment variable holding a bearer token; the token is sent and never written anywhere. Leave whip out to watch a stream something else publishes.

The ramp is deterministic. When its step is a multiple of the clip's 1 s GOP, every viewer arrives at the same point of the GOP and the join times cluster (ten viewers over ten seconds all wait about one second); a step that is not, such as 50 viewers over 10 s, samples the GOP evenly.

Reports

Each run writes JSON (schema whipbench.report/v0) and a Markdown rendering of it. A report records the whipbench version, the scenario, the publisher's figures, every viewer's figures, the aggregates, error counts by kind, a per-second timeline and the definitions above.

First live numbers

evals/2026-10-01-mediamtx-local.md: MediaMTX v1.21.1 in Docker on the same laptop, 10 and 50 viewers, VP8 and H.264. All viewers joined, no packet was lost, and packet transit (the 0.0.1 reports' latency key) was unavailable because MediaMTX's answers did not negotiate abs-capture-time. One machine, one server, loopback: it shows the tool works end to end and nothing about how MediaMTX compares with anything.

How it compares

Two open tools sit closest. Both are good at what they set out to do; whipbench sets out to do something narrower.

whipbench is the combination neither aims at: both standard endpoints, native clients cheap enough to run hundreds from one machine, and one report format with fixed definitions — including a no-verdict rule and an explicit "unavailable" — so that the same scenario against MediaMTX, OvenMediaEngine, LiveKit, Janus or a managed service produces results that can be put side by side.

Limits

Stated plainly, because a benchmark that hides them is worse than none:

Load-testing etiquette

Run whipbench against servers you operate, or against a managed service only on your own account and within its terms, or with the written permission of whoever runs it. A load test against someone else's service is indistinguishable from an attack on it. The live numbers in this repository come from local servers only. The longer note is WB-6.

Development

go vet ./... && go test -race ./...     # unit tests and real round trips through an in-process relay
golangci-lint run ./...
./scripts/check-repo.sh                 # version, changelog and README invariants
node scripts/leakcheck.mjs              # nothing private in the tracked files or the history
npm run backlog && npm run build:site   # backlog lint, roadmap freshness, the Pages site

The end-to-end tests need no network and no external server: internal/testserver is a small WHIP/WHEP relay on pion, run on loopback inside the test, that forwards the publisher's packets to every viewer and can strip the header extension, refuse viewers past a limit or drop packets. CONTRIBUTING.md has the conventions and the release runbook.

Roadmap

BACKLOG.md is the plan and ROADMAP.md is generated from it. In short: v0.1.0 measures one-way delay by frame fingerprint (WB-1 decided it, WB-38 builds it) and records a live run against MediaMTX with the first report; v0.2.0 adds simulcast, layer switches and metrics, and the stamp, clock and comparability around one-way delay (WB-39 to WB-41); v0.3.0 is the comparative report across four servers and a write-up; v0.4.0 spreads the load over several machines and reports the client's own ceiling; v0.5.0 runs whipbench in CI with assertions and a GitHub Action; v0.6.0 measures what the viewer sees — capture-to-decode (WB-2), freezes, picture quality, impaired networks; and v1.0.0 freezes the report schema, the CLI and the scenario keys, with every published number reproducible.