Distributed rate limiter architecture.
Every design here is a trade between how accurate the limit is and how much latency you are willing to add to every single request to achieve it.
The components
Every row below is read from the graph that produced the diagram above, so the two cannot disagree.
| Component | Tier | Why it is there |
|---|---|---|
| Client | Client | Supporting component |
| Edge proxy | Edge | First rejection point, cheapest place to say no |
| API gateway | Edge | Supporting component |
| Limiter middleware | Application | Token bucket, evaluated per request |
| Local counter | Application | In-process, absorbs most checks |
| Sync worker | Application | Reconciles local drift with the shared store |
| Policy service | Application | Per plan and per route quotas |
| Upstream services | Application | Supporting component |
| Redis | Data | Shared counters, Lua for atomicity |
| Postgres | Data | Supporting component |
| Prometheus | Infrastructure | Rejection rate is the signal to watch |
Design decisions worth arguing about
A diagram shows what was chosen. It does not show what it cost, and that is usually the part that matters in a review or an interview.
Reject as early as possible
The cheapest rejection is the one that never reaches your application. Pushing coarse limits to the edge proxy means abusive traffic costs you a connection rather than a request path, which is the difference between absorbing an attack and being taken down by one. The edge cannot see per-plan quotas or user identity, so this only ever handles the blunt cases and a second layer is still required.
Local counters trade accuracy for latency
Checking a shared store on every request adds a network round trip to your p50. Counting locally and reconciling periodically removes that, at the cost of allowing more through than the limit strictly permits, bounded by the number of nodes times the sync interval. For protecting a system from overload, approximate is entirely fine. For a quota someone is billed against, it is not, and that distinction should decide the design rather than a preference for one algorithm.
Token bucket, because bursts are normal
A fixed window rejects the eleventh request in a second even when the previous nine seconds were idle, which is a poor description of how clients actually behave, and it produces a stampede at every window boundary. A token bucket permits a burst up to the bucket size and then enforces a sustained rate, which matches both real traffic and what users expect. It costs two values per key rather than one and needs atomic updates, which is what the Lua script is for.
Policy separate from enforcement
Hardcoding limits in the middleware means a plan change is a deploy. Loading them from a policy service makes limits data, at the cost of a lookup that must be cached aggressively and a decision about what happens when the policy service is unavailable. Failing open there is usually right, since a limiter that blocks everything because it cannot read its config is a worse outage than briefly unlimited traffic.
How it changes with scale
Cost is per request rather than per user, so it grows with total traffic including the traffic you are rejecting. Key cardinality is the thing that surprises people: limiting per user per route multiplies out quickly, and the shared store ends up holding far more keys than expected. Short TTLs are what keep that bounded.
Where it breaks first
The shared store becoming unavailable. Every request now needs a decision with no shared state, and the choice made in advance determines whether you fail open, accepting unlimited traffic, or fail closed, rejecting everything. Both are bad; the failure mode here is not having decided, so the behaviour is whatever the client library's default timeout does.
Draw this yourself
Open the Diagram tab and describe the system. The agent emits a semantic graph rather than coordinates, so you can edit the components and the layout re-solves instead of drifting.
When the shape is right, Implement in code turns the canvas into a markdown specification, every component, every relationship and the notes, and starts a real turn in the Code tab with it.
Questions about this design
Token bucket or leaky bucket?
Token bucket for APIs, because it permits the bursts real clients produce. Leaky bucket enforces a perfectly smooth output rate, which is what you want when protecting something that genuinely cannot absorb a burst, such as a downstream with fixed concurrency.
What should a rejected request return?
429, with a Retry-After header. Without that header clients retry immediately and often, which is the single largest source of load during a rejection event.
Should rate limit state be replicated across regions?
Usually not. Cross-region consistency costs more latency than the accuracy is worth. Per-region limits sized to the regional share are simpler and fail more gracefully.
How do you rate limit unauthenticated traffic?
By IP, knowing it is a poor identifier: shared NATs mean many users behind one address and attackers rotate addresses cheaply. It is a mitigation rather than a control, which is why unauthenticated limits should be tight and authenticated ones generous.
More templates
- Observability Stack ArchitectureThe architecture is mostly about cost control. Collecting everything is technically easy and financially ruinous, so the interesting decisions are all about what to throw away.
- Authentication Service DesignStateless tokens make verification free and revocation hard, and that single tradeoff explains most of the components in this diagram.
- Event-Driven Microservices DesignDistributed transactions do not exist here, so every consistency guarantee you want has to be rebuilt out of events, retries and compensation.
- LLM Inference Service DesignGPUs are the budget, so almost every decision here is about keeping them busy without letting the queue destroy tail latency.
Last updated