# The latency of our phone streams, and the engineering behind it

> Measured on 3,724 production streams: where the time goes between a phone's screen and your browser tab, and what we built to keep it short.

1 October 2026 · Distilled

## What we measured

Every phone on Distilled is a full Android system running on one of our nodes, and you watch its screen as live H.264 video in your browser. Two things decide how fast that feels: how soon the first frame appears when you open a phone, and how little delay each new frame picks up on the way.

- **92 ms**: Median time to first key frame when the phone is already streaming
- **1.4 ms**: Median time a frame waits for the GPU encoder
- **0.35%**: Of one CPU core spent encoding, down from about 70%
- **9 kbit/s**: Median bandwidth of a phone in a grid

The stream numbers come from the API's logs, which record each stream's timings and byte counts when it ends. We read all 3,724 streams from 18:04 UTC on 30 September to 20:16 UTC on 1 October. The encoder numbers come from runs on a node with the same GPU, scrolling the Settings app.

Figure: Share of streams with their first key frame sent, by milliseconds after the API admitted the viewer. With the encoder running, the median is 92 ms and nearly all are under 150 ms. When it has to start, the median is 304 ms.

*Share of streams whose first key frame had been sent by each moment after the API admitted the viewer. The dots mark the medians. With the encoder already running (783 streams), almost every stream is under 150 ms. When it has to start (1,757 streams), the median is 304 ms.*

When a phone is already streaming, 99% of new viewers had their first key frame within 172 ms. Starting an encoder adds about 200 ms at the median. The rest of this post follows a frame from the phone to your tab, then covers opening phones quickly, showing many at once, and staying live on a slow link.

## Where the time goes

A frame makes three hops. The node encodes it, the API checks your ticket and relays it, and the browser decodes it with WebCodecs and paints it onto a canvas. All streams to a node share one tunnel, so a grid of 60 phones costs one connection, not 60.

Figure: A frame goes from the phone to the node, through one tunnel per node to the API, and over a WebSocket to the browser

*Every frame takes this path. The labels on the arrows say what carries it between hops.*

No step waits to collect more work. Android composes a frame only when the screen changes, so a phone that sits still sends nothing.

Figure: One frame's path: the phone composes it, hands it to the GPU in 1.4 ms at the median, the video engine encodes it in about 1.1 ms, it is sent alone, and the browser paints it as soon as it is decoded

*What happens to one frame on the GPU path. The two timed steps are measured. The others are design choices that add no wait.*

- **No batching.** Each frame leaves the node as its own message as soon as it is encoded.
- **No reordering.** The GPU encoder makes only forward-predicted frames, so the decoder never holds one back to wait for a later one.
- **No playout buffer.** On a steady link the browser paints each frame as soon as it is decoded.

Opening a phone adds one cost: joining. The API admits a viewer in 27 ms at the median, and the node's first message arrives 92 ms later (216 ms at the 99th percentile). What follows depends on whether the phone's encoder is already running.

## Encoding on the node's GPU

A phone's frames come from one of two encoders. The first runs inside the phone, where Android's software encoder shares the CPU with the app on screen. The second hands each finished frame to the node's GPU, which encodes it without the CPU touching a pixel.

Figure: Encoding inside the phone uses about 70% of a core; encoding on the node's GPU uses 0.35%

*The two encoders side by side. Moving the encode off the phone cuts its CPU cost by a factor of about 200.*

Inside the phone, the encoder competes with the app it is filming. Playing a full-screen animation at 960 px and 6 Mbit/s, the phone dropped to 7 fps with half-second stalls. At 640 px and 0.8 Mbit/s the same animation ran at 20 fps, so those are the defaults on that path.

On the GPU path, the phone's compositor shares each frame with the GPU as a buffer. The video engine converts it from RGBA and encodes it, and the encoded bytes are written out straight from GPU memory. A frame waits 1.4 ms at the median and 1.9 ms at the 99th percentile to be taken, and the result matched the phone's own screenshot at 49 to 50 dB PSNR.

A node's phones share one video engine. When another phone would not fit, every GPU-encoded phone on the node drops to the next frame cap together.

## Opening a phone quickly

A decoder can only start at a key frame, the one kind of frame that holds a complete picture. A viewer almost always arrives between two, so the node keeps every frame since the last key frame and sends them first.

Figure: A row of frames. Tall ones are key frames. A viewer that joins is sent every frame since the last key frame, then the live ones.

*What a new viewer is sent. Tall bars are key frames and short ones are deltas. The replay starts at the last key frame, so the viewer's decoder can rebuild the current screen before live frames arrive.*

A still phone sends nothing, so its last key frame can be minutes old and a replay may grow to 1 MiB. For those phones the node asks the GPU encoder for a fresh key frame every 45 frames, which keeps the replay small enough to decode in one burst. The player then skips past the burst, so the view opens on the present.

Starting an encoder is the slow part of joining, so both ends avoid doing it twice. The node keeps a phone's encoder running for 30 seconds after its last viewer leaves, and the console keeps a phone's stream open for 30 seconds after it scrolls out of view. In the window above, 31% of the streams that showed a frame found the encoder already running.

> Browsers disagree on the bytes they accept. Chromium decodes the stream as the phone sends it. Firefox and Safari only decode MP4 framing, so the console rewraps each frame for them in the tab.

## Many phones on one page

On the Phones page every tile is its own stream. Three things keep that cheap: most tiles see only key frames, tiles open their streams in parallel, and a tile that comes back shows its last frame right away.

### Previews

A grid tile streams key frames only, at most two a second. When the screen stops changing, the tile is sent the rest of the replay so it shows where the screen came to rest. Hovering a tile switches it to every frame. At most 64 tiles stream at once, the best placed first, and none start while the page is scrolling.

The median tile used 9 kbit/s and the 90th percentile 13 kbit/s, because most phones in a grid are still. Across the window, 49 hours of preview tiles moved 130 MB in total. A phone someone was actively using moved up to 2.5 Mbit/s at 25 fps.

### Lanes

Chromium opens one WebSocket at a time to each host and port, and queues the rest, so with every tile on port 443 a grid filled one tile at a time. The API also answers on five other HTTPS ports, and the console treats each port as a lane with one stream opening on it at a time.

Figure: Grid tiles open sockets on six API ports at once instead of queueing on one

*Grid streams spread over six ports, so about six tiles open at once instead of queueing behind one.*

Before turning lanes on, we opened three real preview streams on each port. Every lane was as quick as 443.

Figure: Time to first frame per port, fastest to slowest of three tries: 443 took 348 to 861 ms, and the other five ports fall in the same range

*Time to first frame on each port, fastest to slowest of three tries. The extra ports (champagne) fall in the same range as 443 (grey).*

## Staying live on a slow link

On a live screen, a late frame is worse than a missing one, so no stage lets a queue grow. Each one drops frames instead, always to the next key frame, because a delta decoded after a gap paints garbage.

Figure: Four stages drop frames to the next key frame when a viewer falls behind: the node, the API relay, the browser's decoder and its player

*The four stages a frame passes, and what each does when the viewer falls behind. The champagne line is when it acts; the last line is its limit.*

The node eases the encoder first, because that costs the least. Every second it checks whether any viewer is behind. If one is, a GPU-encoded phone steps down, giving up detail before motion, and climbs back a step after five seconds with everyone caught up. The frame cap stops at 15 fps: at 10 fps the encoder would hold a frame for 101 ms, past the 100 ms the phone allows, and the phone would drop it.

Figure: Encoder steps: 6 Mbit/s at 30 fps, 3 at 30, 1.5 at 24, 0.75 at 20, 0.5 at 15

*The steps the node can ease a GPU encoder down to. Each step halves the bit rate, to a floor of 0.5 Mbit/s; from step 2 the frame cap comes down too.*

In the window we measured, the relay never had to skip a frame or close a viewer for being too slow. When a connection drops, the tab notices within eight seconds even if it never closed, and reconnects after a jittered wait of one to eight seconds. A restarted encoder starts a new session, so the browser starts a fresh decoder rather than predicting from a stream that ended.

## Summary

- **Short path.** Frames are sent one by one, with no batching, reordering or buffering, so a frame reaches the tab about as fast as it is encoded.
- **GPU encoding.** Moving the encode off the phone cut its CPU cost from about 70% of a core to 0.35% and left the phone to the app on screen.
- **Fast joins.** A replay from the last key frame and an encoder kept warm give a median first key frame of 92 ms.
- **Cheap grids.** Key-frame previews and parallel lanes keep a tile at a median of 9 kbit/s.
- **No queues.** Every stage drops to the next key frame instead of falling behind.

The API that starts and drives these phones is documented at [/docs/](https://distilled.cx/docs/).

Source: https://distilled.cx/blog/streaming-phone-screens/
