Search documentation

Search pages and headings

API Reference

Overview

Conventions shared by every endpoint — transport, how a room is bootstrapped, and what the server keeps.

Conventions

Base path /v1
Response body JSON, except /v1/stream, which is text/event-stream
CORS Enabled on all routes, so browsers may call the API directly
Auth None. The server holds no credentials of its own

Query strings and request bodies are decoded with a schema at the boundary. Anything that does not match returns 400 INVALID_REQUEST with the validation message.

The endpoints

Endpoint Purpose
GET /v1/stream Subscribe to a video's chat over Server-Sent Events
POST /v1/probe Diagnose which innertube clients work for a video

That is the whole API. There is no bootstrap call and no polling call — opening the stream is the entire client-side protocol.

The server owns the polling loop

This is the single most important property to understand, because it shapes everything else.

                    ┌──────────────────────────────┐
YouTube  ◄──────────┤ one poller per videoId        │
                    │ owns the continuation token   │
                    └───────────────┬──────────────┘
                                    │ persist, then publish
                    ┌───────────────▼──────────────┐
                    │ message archive               │
                    └───────────────┬──────────────┘
                                    │
                  ┌─────────────────┼─────────────────┐
              SSE viewer        SSE viewer        SSE viewer

When the first viewer of a video connects, the server bootstraps a room and starts polling YouTube. Every later viewer of the same video attaches to that same room. Consequences worth planning around:

  • Viewers are free. Ten viewers of one stream cost exactly the same number of YouTube requests as one. Your load on YouTube scales with distinct videos, not with users.
  • Credentials never reach you. The innertube API key, the client version and the continuation token stay on the server. Earlier versions of this API returned all three; nothing on the wire carries them now.
  • The room outlives your connection. Disconnecting does not stop the poller immediately — it lingers for about 90 seconds so a page refresh rejoins rather than re-bootstraps.
  • Run one instance. Two servers would each poll the same video and assign conflicting cursors. There is no leader election.

What happens during a room bootstrap

Understanding this explains most of the error codes.

1. Resolve the video id

The input is accepted as a full URL or a bare id and reduced to an 11-character id. Anything else fails with INVALID_VIDEO_ID.

2. Fetch innertube credentials

The server requests https://www.youtube.com/sw.js and extracts INNERTUBE_API_KEY and INNERTUBE_CLIENT_VERSION from it. These are fetched live rather than hardcoded, because YouTube rotates them. A failure here is SW_JS_FAILED.

3. Find the chat renderer

The server loads the popout chat page first (/live_chat?is_popout=1&v=<id>) and falls back to the watch page (/watch?v=<id>). From the HTML it brace-matches the ytInitialData JSON object, then walks it looking for a liveChatRenderer (live) or liveChatReplayRenderer (replay) that carries a continuation token. Which one was found is reported in the stream's init frame as chatType.

If neither page yields a renderer with a token, the result is NO_LIVE_CHAT — the video is not live, has chat disabled, has no replay chat, or is unavailable.

4. Pick a working client

A continuation token alone is not enough; it has to be redeemed against an innertube client that YouTube will accept for this video. The server tries WEB, then TVHTML5, and for each tries the get_live_chat endpoint before get_live_chat_replay, stopping at the first combination that returns a parseable response. That pair is then fixed for the room's lifetime.

If every combination fails, the result is NO_WORKING_CLIENT.

5. Poll

From here the room polls on its own, at the interval YouTube asks for (clamped to 500 ms – 15 s), until the chat ends, the errors pile up, or the last viewer leaves.

Where errors appear

The bootstrap above happens before your response body starts, so all of its failures are ordinary JSON errors with real status codes. Once the stream is open the status line has already been sent, so anything that goes wrong after that arrives as an event: error frame carrying the same code.

That split is the one structural thing to handle in a consumer. See Errors.

Upstream pacing

The server spaces the requests it makes to YouTube at least 250 ms apart, runs at most 4 concurrently, collapses identical in-flight requests into a single call, and pauses every room when YouTube throttles it. See Upstream Pacing.

Reading YouTube's failures

YouTube returns HTTP 200 with an error object in the body when an innertube call fails. The server treats such a response as a failure, not a success, and inspects it further: a quota signal is a rate limit, anything else is an expired token, which the poller repairs by re-bootstrapping the room without your connection noticing.