The bouncer at the door of every API.
"Design a rate limiter. I need to cap users at N requests per time window across a fleet of API servers. Token bucket or sliding window — which, and how do you make it work distributed?"
A rate limiter is a bouncer with a clicker counter. Every request asks "am I allowed?", the limiter checks a per-key counter against the rule, and answers yes or 429 Too Many Requests — ideally in under a millisecond, on the request's hot path. The two classic algorithms differ in one question: token bucket asks "do I have a token saved up?" (bursts allowed, smooths average rate), sliding window asks "how many did you send in the last N seconds?" (stricter, no bursts).
The interview sentence: "Token bucket when bursts are legitimate, sliding window when the limit is a hard ceiling; either way, keep the counters in Redis with a Lua script so the check-and-decrement is atomic."
"Rate limiting is just if count > limit: reject with a counter that resets every minute." That's fixed window limiting, and it has a famous hole: a user can send limit requests at 12:00:59 and limit more at 12:01:01 — double the intended rate across the boundary, completely "legally." Sliding window closes the hole by always looking back a full window from now, and token bucket closes it differently by making bursts spend saved-up tokens. If your design has a counter that resets on a clock tick, an interviewer will drive a truck through the boundary.
429 with Retry-After when over the limit; include X-RateLimit-Remaining headers on success.Analogy first: the limiter is a toll booth. The math tells you how many lanes (Redis ops) you need so the booth never becomes the traffic jam.
| Quantity | Assumption | Result |
|---|---|---|
| Decisions/sec | 50k API requests/sec fleet-wide | 50k limiter checks/sec — one Redis op per check, pipelined |
| Redis load | 1 Lua script eval per check, ~0.2ms | A single Redis handles ~100k+ evals/sec — one primary covers it with room |
| Memory for counters | 1M active keys × ~200 bytes (token bucket state) | ~200 MB — trivial; sliding-window logs cost more per key |
| Sliding-window log cost | 1M keys × 100 req/min window × 16 bytes/timestamp | ~1.6 GB — 8× the bucket; this is why buckets win on memory |
| Fail-open risk | Limiter down 60s at 50k rps | 3M uncounted requests — acceptable vs. 3M rejected requests if you fail closed |
Takeaway after the math: the limiter's cost is one Redis round-trip per request and a few hundred MB of state. The expensive mistake isn't the algorithm — it's failing closed and turning a Redis blip into a full outage.
Picture the toll booth again: every car (request) stops, the attendant (Lua script) checks the ledger (Redis) and either waves it through or holds up the red paddle — all in one atomic motion so two booths can't both wave through the last token.
flowchart LR
A["Client"] --> B["API gateway
extract key: user/IP/API key"]
B --> C["Rate limiter
middleware"]
C --> D[("Redis
counters")]
C -- "allow" --> E["Upstream service"]
C -- "429 + Retry-After" --> A
F["Config service"] -- "rule updates" --> C
Two API servers checking the same key at the same microsecond could both read tokens=1, both decrement, and both allow — the classic check-then-act race. A Lua script runs atomically inside Redis: refill calculation, comparison, and decrement happen as one indivisible step. No locks, no races, one round trip. Every production rate limiter does the decision in a single atomic script; if your design has the app server doing read-then-write, that's the bug the interviewer is hunting for.
Each key has a jar holding up to capacity tokens, refilling at rate per second. A request takes one token; empty jar means 429. A full jar lets a burst of capacity requests through instantly — that's the feature: it smooths the average while tolerating bursts.
sequenceDiagram
participant G as Gateway
participant R as Redis (Lua)
G->>R: EVAL check(key, now)
R->>R: refill = (now - last) * rate
R->>R: tokens = min(capacity, tokens + refill)
alt tokens >= 1
R->>R: tokens -= 1
R->>G: ALLOW, remaining=tokens
else tokens < 1
R->>G: DENY 429, retry_after
end
Keep the timestamps of recent requests (a sorted set in Redis). On each request, drop timestamps older than window, count what's left: under the limit → allow and record now; at the limit → 429. No bursts, no boundary exploits — the window always looks back a full N seconds from this instant.
sequenceDiagram
participant G as Gateway
participant R as Redis (Lua)
G->>R: EVAL check(key, now)
R->>R: remove timestamps older than now - window
R->>R: count = remaining timestamps
alt count < limit
R->>R: add now to the log
R->>G: ALLOW, remaining = limit - count - 1
else count >= limit
R->>G: DENY 429, retry_after
end
# allowed
HTTP 200
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 73
X-RateLimit-Reset: 1696264800
# denied
HTTP 429 Too Many Requests
Retry-After: 42
{ "error": "rate_limited",
"message": "100 requests per minute exceeded" }
{tokens, last_refill} — 2 numbersrl:{rule}:{key} + TTL{limit, window, algorithm}TTL every key at 2× the window so stale keys evaporate. Tiered rules ("100/min + 1000/hr") are just two keys checked in one Lua script — deny if either trips.
| Decision | Token bucket | Sliding window | Verdict |
|---|---|---|---|
| Burst handling | Bursts of capacity allowed — good for humans clicking fast | No bursts — every second is policed equally | Bucket for user APIs, window for abuse-prone endpoints |
| Memory per key | Two numbers — tiny | One timestamp per recent request — 8× or more | Bucket wins at 1M+ keys |
| Precision | Average-rate guarantee; short spikes pass | Exact: never more than limit per window | Window when the limit is contractual (billing tiers) |
| Distributed state | Central Redis + Lua — one source of truth | Same — both need shared state | Tie; never do per-process counters across a fleet |
Retry-After headers + exponential backoff on the client; the 429 must teach, not just punish.rl:{rule}:{key} with TTL = 2× window.Retry-After, exponential backoff on 429s — enforced in our own SDKs, documented for everyone else.Both algorithms, same traffic, side by side — running the real logic in JS. Tune the knobs, then fire bursts and watch one bouncer wave people through while the other holds the line. Try "Auto-burst ×50" to see the difference a burst makes.
allowed (green) vs 429s (red) per second, last 20s
allowed (green) vs 429s (red) per second, last 20s
At the API gateway / edge for coarse limits (per IP, per API key) — it sheds load before it touches your fleet. Finer limits (per user per endpoint, billing tiers) live as middleware in the service, where the key and the rule context exist. Big systems run both: the edge stops the flood, the service enforces the contract. And the counters? Always shared (Redis) — per-process counters across a fleet are just N independent limiters pretending to be one.