API Reference
Upstream Pacing
How the server paces, coalesces and throttle-gates its outgoing YouTube requests so one shared IP cannot get itself blocked.
The server sits between your clients and YouTube using a single shared egress IP. Every request that leaves it for YouTube — chat polls and the page fetches a bootstrap needs — passes through one gate with three mechanisms.
Pacing
Outgoing calls are spaced at least 250 ms apart and run at most 4 concurrently. A burst of incoming requests does not translate into a burst of outgoing ones; it becomes a queue that drains at a steady rate.
Fan-out does most of the work here before the gate ever sees anything: one poller serves every viewer of a video, so a thousand viewers of one stream produce the same upstream traffic as one. The gate handles what is left — many distinct videos being bootstrapped at once.
Deduplication
Concurrent polls that are byte-for-byte equivalent at the upstream level — same token, client, endpoint, API key and host — collapse into a single YouTube call. Every caller receives the same response.
This is safe because polling a continuation token does not consume it: two callers racing on the same token would have received the same batch anyway.
Two boundaries worth knowing:
- In-flight only. The moment a shared call finishes, the next identical request makes a fresh YouTube call. Nothing is cached between completed requests — this is admission pacing, not a response cache.
- Shared outcomes. All callers waiting on one upstream call share its result including failures.
Page fetches are paced but not deduplicated — they have nothing to deduplicate against.
The throttle cooldown
YouTube throttles by IP address, not by request. So when one room is rate-limited, every other room is too, whether it knows it or not — and a single room backing off alone achieves nothing while the others keep the block open.
The gate therefore holds a process-wide cooldown. A rate-limited response pushes it forward, and no outgoing call — poll or page fetch — claims a slot until it expires.
| Signal | Treated as |
|---|---|
HTTP 429 or 403 |
Rate limit |
HTTP 200 with RESOURCE_EXHAUSTED in the body |
Rate limit |
HTTP 200 with a quota / rate-limit message |
Rate limit |
HTTP 200 with any other error object |
Expired token |
The last two rows matter: without inspecting the body, a throttle is indistinguishable from an expired token, and the server's response to an expired token is to re-bootstrap — that is, to make more requests of an endpoint that has just asked for fewer.
Retry-After is honoured when present, in both of its legal forms
(delta-seconds and an HTTP date), clamped to five minutes. YouTube is
inconsistent about sending it, so an absent or unusable value falls back to
30 seconds rather than to no wait at all.
What a throttled room does
Resuming at the old rate right after a throttle walks straight back into it. So a throttled room also widens its own poll interval, doubling up to 60 seconds — deliberately above the normal 15-second ceiling, because under sustained throttling the right interval is one YouTube's own pacing hint would never suggest. It halves back on each success rather than resetting.
A rate limit is not counted as a stream error. It has its own budget, so a room that is merely waiting its turn is not killed after three of them, while a permanent block still ends the room eventually.
What a client sees
A rate limit that happens before your stream opens is a 429 with a
Retry-After header. One that happens after it opens is an event: error
frame with code RATE_LIMITED. Either way, wait — reconnecting immediately
makes it worse, since your reconnect drives another bootstrap.
The client library implements this: it reads
Retry-After, waits it out, and counts rate limits against a separate budget
from ordinary errors.
Tuning
The knobs live in apps/api/src/lib/constant.ts alongside everything else:
| Constant | Default | Meaning |
|---|---|---|
UPSTREAM_MIN_INTERVAL_MS |
250 |
Minimum spacing between outgoing requests |
UPSTREAM_MAX_CONCURRENCY |
4 |
Cap on concurrent outgoing requests |
THROTTLE_COOLDOWN_MS |
30000 |
Cooldown when Retry-After is absent or unusable |
THROTTLE_MAX_COOLDOWN_MS |
300000 |
Ceiling clamping a hostile Retry-After |
THROTTLE_BACKOFF_FACTOR |
2 |
How fast a throttled room widens its interval |
THROTTLE_MAX_INTERVAL_MS |
60000 |
Ceiling for a widened poll interval |
MAX_THROTTLE_ERRORS |
8 |
Consecutive rate limits before a room gives up |
Raising the first two increases throughput toward YouTube at the cost of rate-limit headroom; lowering them protects the IP further at the cost of queueing latency.