Feature flag service architecture.
Evaluation has to be local and instant, because a flag check sits in the hot path of code that would otherwise not make a network call at all.
The components
Every row below is read from the graph that produced the diagram above, so the two cannot disagree.
| Component | Tier | Why it is there |
|---|---|---|
| Client SDKs | Client | Evaluate locally, never block on a network call |
| CDN | Edge | Serves the ruleset, cached at the edge |
| Streaming endpoint | Edge | SSE push, so a kill switch is instant |
| Flag API | Application | Supporting component |
| Rule compiler | Application | Targeting rules to a payload SDKs can evaluate |
| Audit log | Application | Who flipped what, and when |
| Exposure events | Application | Which variant each user actually saw |
| Postgres | Data | Supporting component |
| Redis | Data | Supporting component |
| Kafka | Data | Supporting component |
| Warehouse | Data | Experiment analysis, not the serving path |
Design decisions worth arguing about
A diagram shows what was chosen. It does not show what it cost, and that is usually the part that matters in a review or an interview.
Evaluate in the SDK, never over the network
A flag check happens in code paths that run thousands of times a second, so a network call per evaluation is out of the question. The SDK holds the compiled ruleset in memory and evaluates locally in microseconds. The cost is that every client is briefly out of date after a change, and that the ruleset must be small enough to hold and cheap enough to evaluate.
Streaming, because a kill switch is the point
Polling every thirty seconds is fine for a gradual rollout and useless for the case the product actually exists to serve, which is turning something off immediately during an incident. A streaming connection pushes changes in under a second. The cost is holding an open connection per client process and a fallback path for networks that will not permit it.
The CDN is the availability story
If the flag service is down and SDKs cannot fetch a ruleset, applications either fail or run on stale defaults. Serving the ruleset as a static artifact from a CDN means the read path survives an outage of everything else you operate. It also means a change is not live until the cache expires or is purged, so purge behaviour becomes load-bearing.
Exposure events are a separate system
Knowing which variant a user actually saw is what makes experiments analysable, and it is high-volume telemetry with nothing in common with flag serving. Keeping it on its own path means an analytics outage cannot affect evaluation. The consequence is that experiment results lag, sometimes by hours, which surprises people expecting a live dashboard.
How it changes with scale
Evaluation cost is zero to the service, since it happens in the client. What scales with usage is exposure event volume, which grows with traffic times flag count and quickly exceeds every other write in the system. Ruleset size grows with flag count and matters because it is held in memory in every process.
Where it breaks first
A flag that never gets cleaned up. The technical system stays healthy while the ruleset accumulates hundreds of permanently-on flags, each one a branch in the code nobody removes. The failure is organisational and shows up as evaluation latency and unreadable code rather than as an outage, which is why flag expiry dates and an owner per flag matter more than any part of this diagram.
Draw this yourself
Open the Diagram tab and describe the system. The agent emits a semantic graph rather than coordinates, so you can edit the components and the layout re-solves instead of drifting.
When the shape is right, Implement in code turns the canvas into a markdown specification, every component, every relationship and the notes, and starts a real turn in the Code tab with it.
Questions about this design
Should flag evaluation happen server-side or in the client?
In the SDK, wherever it runs, using a ruleset fetched in advance. A network call per evaluation is too slow for the code paths flags actually sit in.
How do you avoid leaking upcoming features to browser clients?
You cannot fully: a ruleset in the browser is readable. Anything genuinely sensitive should be evaluated server-side and the client told only the outcome.
Build or buy?
Buy, unless you have unusual requirements. The serving path is simple; the parts that take time are the SDKs for every language you use and the experiment analysis on the other end.
How long should a flag live?
Give it an expiry date when it is created and treat a passed date as a bug. Flags that outlive their rollout are the main long-term cost of the whole system.
More templates
- Webhook Delivery System DesignYou are making requests to servers you do not control, which are frequently slow, sometimes wrong, and occasionally gone.
- Audit Logging System DesignAn audit log that can be edited is not an audit log, which makes this one of the few systems where the absence of features is the design.
- Vector Search Service DesignThe failure mode here is not an error. It is recall dropping slowly as the index grows, with nothing in your metrics saying so.
- AI Agent Orchestration DesignThe orchestrator owns the loop, and almost every hard problem here is about knowing when to stop.
Last updated