API Reference
Overview
Conventions shared by every endpoint — transport, how a room is bootstrapped, and what the server keeps.
Conventions
| Base path | /v1 |
| Response body | JSON, except /v1/stream, which is text/event-stream |
| CORS | Enabled on all routes, so browsers may call the API directly |
| Auth | None. The server holds no credentials of its own |
Query strings and request bodies are decoded with a schema at the boundary.
Anything that does not match returns 400 INVALID_REQUEST with the
validation message.
The endpoints
| Endpoint | Purpose |
|---|---|
GET /v1/stream |
Subscribe to a video's chat over Server-Sent Events |
POST /v1/probe |
Diagnose which innertube clients work for a video |
That is the whole API. There is no bootstrap call and no polling call — opening the stream is the entire client-side protocol.
The server owns the polling loop
This is the single most important property to understand, because it shapes everything else.
┌──────────────────────────────┐
YouTube ◄──────────┤ one poller per videoId │
│ owns the continuation token │
└───────────────┬──────────────┘
│ persist, then publish
┌───────────────▼──────────────┐
│ message archive │
└───────────────┬──────────────┘
│
┌─────────────────┼─────────────────┐
SSE viewer SSE viewer SSE viewerWhen the first viewer of a video connects, the server bootstraps a room and starts polling YouTube. Every later viewer of the same video attaches to that same room. Consequences worth planning around:
- Viewers are free. Ten viewers of one stream cost exactly the same number of YouTube requests as one. Your load on YouTube scales with distinct videos, not with users.
- Credentials never reach you. The innertube API key, the client version and the continuation token stay on the server. Earlier versions of this API returned all three; nothing on the wire carries them now.
- The room outlives your connection. Disconnecting does not stop the poller immediately — it lingers for about 90 seconds so a page refresh rejoins rather than re-bootstraps.
- Run one instance. Two servers would each poll the same video and assign conflicting cursors. There is no leader election.
What happens during a room bootstrap
Understanding this explains most of the error codes.
1. Resolve the video id
The input is accepted as a full URL or a bare id and reduced to an
11-character id. Anything else fails with INVALID_VIDEO_ID.
2. Fetch innertube credentials
The server requests https://www.youtube.com/sw.js and extracts
INNERTUBE_API_KEY and INNERTUBE_CLIENT_VERSION from it. These are fetched
live rather than hardcoded, because YouTube rotates them. A failure here is
SW_JS_FAILED.
3. Find the chat renderer
The server loads the popout chat page first
(/live_chat?is_popout=1&v=<id>) and falls back to the watch page
(/watch?v=<id>). From the HTML it brace-matches the ytInitialData JSON
object, then walks it looking for a liveChatRenderer (live) or
liveChatReplayRenderer (replay) that carries a continuation token. Which
one was found is reported in the stream's init frame as chatType.
If neither page yields a renderer with a token, the result is NO_LIVE_CHAT —
the video is not live, has chat disabled, has no replay chat, or is
unavailable.
4. Pick a working client
A continuation token alone is not enough; it has to be redeemed against an
innertube client that YouTube will accept for this video. The server tries
WEB, then TVHTML5, and for each tries the get_live_chat endpoint before
get_live_chat_replay, stopping at the first combination that returns a
parseable response. That pair is then fixed for the room's lifetime.
If every combination fails, the result is NO_WORKING_CLIENT.
5. Poll
From here the room polls on its own, at the interval YouTube asks for (clamped to 500 ms – 15 s), until the chat ends, the errors pile up, or the last viewer leaves.
Where errors appear
The bootstrap above happens before your response body starts, so all of
its failures are ordinary JSON errors with real status codes. Once the stream
is open the status line has already been sent, so anything that goes wrong
after that arrives as an event: error frame carrying the same code.
That split is the one structural thing to handle in a consumer. See Errors.
Upstream pacing
The server spaces the requests it makes to YouTube at least 250 ms apart, runs at most 4 concurrently, collapses identical in-flight requests into a single call, and pauses every room when YouTube throttles it. See Upstream Pacing.
Reading YouTube's failures
YouTube returns HTTP 200 with an error object in the body when an
innertube call fails. The server treats such a response as a failure, not a
success, and inspects it further: a quota signal is a rate limit, anything
else is an expired token, which the poller repairs by re-bootstrapping the
room without your connection noticing.