How to Design a Rate Limiter
A rate limiter caps the number of requests a client, IP, or API key can make in a time window, protecting backend services from abuse and traffic spikes. It is one of the most frequently asked system design questions because every large system needs one and the answer exposes your understanding of caching, distribution, and edge architecture.
1. Requirements Clarification
Clarify what is being limited and with what granularity before proposing algorithms.
Functional Requirements
- Limit requests per client, API key, user ID, or IP address
- Reject excess requests with a clear error (429)
- Configurable limits per rule (e.g. 100 req/min for free tier)
- Support variable kinds: requests/sec, bytes/sec, concurrent connections
Non-Functional Requirements
- Latency: add <1ms overhead per check
- Accuracy: do not allow burst bypass of the limit
- Availability: limiter failure must not block the whole API (fail-open vs fail-closed)
- Distribution: consistent limits across all API nodes
2. The Four Classic Algorithms
Fixed Window Counter
Count requests in a fixed bucket (e.g. per minute). Simple but allows double-burst at boundaries : 60 requests in the last second of one minute and 60 in the first second of the next pass 120 in 2 seconds.
Sliding Window Log
Keep a timestamp per request; drop the ones older than the window on each check. Precise, but O(n) memory per key because every timestamp is stored.
Sliding Window Counter
Blend the previous window's count with the current one: weighted = prev * overlapRatio + current. Smooth rate without storing every request.
Token Bucket
The interview favorite. A bucket holds up to capacity tokens and refills at refillRate tokens/sec. Each request consumes one token; if empty, reject. It allows bounded bursts yet enforces the long-run average.
class TokenBucket {
private tokens: number;
private lastRefill: number;
constructor(private capacity: number, private refillPerSec: number) {
this.tokens = capacity;
this.lastRefill = Date.now();
}
allow(): boolean {
const now = Date.now();
this.tokens = Math.min(this.capacity, this.tokens + (now - this.lastRefill) / 1000 * this.refillPerSec);
this.lastRefill = now;
if (this.tokens < 1) return false;
this.tokens -= 1;
return true;
}
}3. High-Level Architecture
The limiter sits in front of the API servers, typically at the API gateway or an edge layer, so unneeded work never reaches the backend:
graph LR
C["Clients"] --> G["API Gateway / Edge"]
G --> RL["Rate Limiter"]
RL --> R["Redis (counters / buckets)"]
RL -->|"allow"| API["API Servers"]
RL -->|"429 reject"| C
style RL fill:#D97A2B,stroke:#B86418,color:#fff
style R fill:#FAF6EE,stroke:#E8DFC84. Distributed Implementation (Redis)
Every API node must agree on the shared count, so the counters live in Redis. Use atomic operations to avoid an allow-double-spend race:
-- Fixed window via Redis INCR + EXPIRE
-- (Lua makes read-check-write atomic)
local key = KEYS[1]
local limit = tonumber(ARGV[1])
local current = redis.call('INCR', key)
if current == 1 then redis.call('EXPIRE', key, ARGV[2]) end
if current > limit then redis.call('DECR', key) return 0 end
return 1This is one round trip per request, which stays well under the 1ms budget when Redis is colocated. For extreme scale, shard the keys by client hash across a small Redis cluster, mirroring the consistent hashing you would use for cache keys.
5. Where to Put the Limiter
- Client-side: cosmetic; clients can be tampered with.
- API gateway: the standard answer : central, language-agnostic, protects every service.
- Per-service (in-process): less consistent across nodes, ideal only if requests are hash-sticky.
- Reverse proxy (nginx): the coarsest layer, often combined with an app-level limiter.
6. Response Handling and Error Design
Rejected requests get 429 Too Many Requests plus headers the client can act on: X-RateLimit-Limit, X-RateLimit-Remaining, and Retry-After. Logging which rule fired and returning a stable error code in the body make client retry loops sane.
7. Common Interview Mistakes
- Picking only fixed window and getting grilled on the boundary burst : preempt with sliding window or token bucket.
- Ignoring distribution : a single-node limiter is not a distributed-system answer.
- Not mentioning fail-open vs fail-closed for limiter failures.
- Skipping Redis atomicity : a separate check-then-increment races under concurrency.
- Forgetting per-tier rules (free vs paid) and per-IP vs per-key distinction.
8. Summary: Key Decisions
Answer in this order: clarify granularity and scale, pick token bucket (backed by sliding window where smoothness matters), place it at the gateway, store counters in Redis with atomic Lua, return 429 with Retry-After. Then name the trade-offs for each.
Frequently Asked Questions
Is rate limiting in-memory or do you always need Redis?
For a single server, in-process buckets are fine and zero-latency. For multiple API nodes, shared Redis gives one global count. The classic answer is Redis for distribution, but saying "in-process when sticky, Redis when shared" shows nuance interviewers reward.
How does the token bucket allow bursts?
The bucket starts full at capacity, so a client can immediately spend up to capacity tokens at once. Afterward it refills at the steady rate. This is why token bucket is preferred for APIs where legitimate clients occasionally burst.
How is rate limiting related to caching?
Both are pre-computation layers and both use Redis-style stores, so the memory, eviction, and consistency lessons carry over. A rate-limiter read is essentially a cache lookup with atomic accounting, which is why it pairs with the cache patterns in our caching tutorial.
Related Tutorials
- Caching Strategies : the Redis foundation a limiter reuses.
- Consistent Hashing : sharding limiter keys across a Redis cluster.
- Load Balancing : where the gateway/limiter front-end fits.
- System design case studies : rate limiting shows up in most of them.
Put it into practice
Ready to practice?
Practice this design in a live system design mock interview with InterviewSkool's AI interviewer.
Start a System Design Interview →