At a glance
Now or scheduled.
Opt-out · quiet hours.
Ack ≠ delivered.
Queues · workers · budgets.
Stable IDs · mark after send.
Breaker · bulkhead · failover.
Build a working send path first, then isolate priority traffic and harden delivery semantics. Don't open with Kafka topology before FR/NFR and entities exist.
Whiteboard order
- Requirements — FR above the line; NFR with surge math + two traffic classes.
- Entities — Notification · Campaign · Segment · User (+ prefs / devices).
- API — POST /notifications, POST /campaigns, PUT preferences/devices; priority + idempotency.
- HLD v1 — sync send → scheduler → campaigns → prefs (working system).
- Deep dives — durable queue → priority isolation → dedup → brownout.
Resist solving 5k/s and OTP SLA before a single email can leave the building.
What is a notification system?
A notification system is an internal platform other teams use to message users across channels (push, email, SMS). Product services hand it a message and a recipient; the platform owns routing, scheduling, preferences, retries, and bookkeeping.
Customers are engineering teams, not end users: auth sending OTPs, orders confirming purchases, marketing blasting a promo to a million users. That shapes auth (service API keys), SLAs (OTP vs promo), and APIs (act-on-behalf-of userId).
Requirements
Functional (steer here)
- Send a notification to a user via push, email, or SMS — immediately or scheduled.
- Send campaigns: same message to a segment, immediate or scheduled.
- Users set preferences: channel opt-outs and quiet hours.
Below the line (skip if low on time)
- Template management / rich authoring.
- Delivery analytics dashboards (open/CTR).
- In-app notification feeds and badge counts.
- Frequency capping across types (nice follow-up later).
Capacity gut-check. 10M notifications/day ≈ 100/s if even — almost nothing. But a 1M campaign in 5 minutes is >3,000/s. Design for surges of ~5,000/s (~50× sustained).
Non-functional
- At-least-once delivery with best-effort deduplication (prefer duplicate over drop).
- Handle surges of ~5,000 notifications/sec.
- High priority (OTP, security) delivered within 5 seconds of acceptance even during a surge.
NFR below the line
- Strict ordering across channels.
- Full compliance/spam legal workflows.
- Monitoring/CI as first-class design (mention, don't boil the ocean).
Functional:
1. Send to a user (now | scheduled) via push | email | sms
2. Campaigns to a segment (now | scheduled)
3. User preferences (opt-outs, quiet hours)
Non-functional:
1. At-least-once + best-effort dedup
2. ~5,000 notif/s surge
3. Critical ≤5s even during surge
Core entities
Start broad — columns come later. Talk these through with the interviewer.
API / system interface
Start simple; grow idempotency keys as you narrate evolution.
POST /notifications → 202 Accepted { id } # after async; early MVP may be 200
Body: {
userId,
channel, // push | email | sms
priority, // high | standard ← required for SLA isolation
content, // { title, body }
scheduledAt, // optional, default now
idempotencyKey
}
GET /notifications/{id} → { id, status, updatedAt }
POST /campaigns → 202 Accepted { id }
Body: {
segmentId, channel, priority, templateId,
scheduledAt, idempotencyKey
}
PUT /users/{userId}/preferences → Preferences
Body: { optOuts: ["sms"], quietHours: { start, end, tz } }
PUT /users/{userId}/devices → Device
Body: { platform: ios|android, pushToken }
HLD 1 — Single send (now + scheduled)
Build the direct path first: one email, right now. SMS is the same shape; push adds device tokens.
- API Gateway — authenticate service API keys; rate-limit runaways so auth OTPs survive.
- Notification Service — lookup address, persist row, call provider.
- Postgres — Notification rows + synced slice of user contact data.
- Email/SMS provider — last mile (SMTP/API). You don't own inbox delivery.
Send-now flow
- Upstream POSTs /notifications.
- Gateway authenticates and forwards.
- Service loads email/phone; writes Notification PENDING; calls provider.
- Updates SENT/FAILED; returns id + status (MVP 200).
Push: app calls PUT /users/{id}/devices on login/token rotate. Send path looks up token instead of email.
Scheduled: write status SCHEDULED; stop. A scheduler cron (~1 min) finds due rows and pushes them through the same send path. OTP never uses cron — it arrives as an immediate high-priority request.
HLD 2 — Campaigns
POST /campaigns writes a Campaign row and returns 202. Expand the scheduler to poll due campaigns too.
- Cron finds Campaign with scheduledAt ≤ now.
- Snapshot segment membership → recipient list.
- For each user, enqueue/send via the same path as 1:1.
- Mark campaign complete (or FAILED with checkpoint — deep dive).
HLD 3 — Preferences
No new components. Preferences live on User; PUT via gateway → notification service.
Working system: sync OTP path, campaign fan-out via cron, preferences enforced at dispatch. NFRs still unsolved — that's the deep-dive half.
Deep dive — Never drop an accepted notification
Pattern: multi-step process (accept → enqueue → deliver → record). Each step can fail alone. Temporal-level orchestration is usually waved off — show the mechanics.
Approach A — retries in-process. Keep sync call; backoff on provider fail.
- Fixes flaky single responses only.
- Crash/deploy still loses in-memory retries.
- Holds caller open during outage → auth timeouts.
- Caller can't tell give-up vs never-got-it → effectively at-most-once.
Approach B — Postgres as queue. Write PENDING, 202, pollers SELECT … FOR UPDATE SKIP LOCKED, retry with attempt/next_run.
- Survives restarts.
- You reinvent a mediocre queue (retries, visibility, hot table of history).
- Latency ≥ poll interval — sluggish for OTP.
Approach C — durable queue (preferred). Enqueue to SQS (or similar); 202 when queue acks. Workers: receive → prefs → provider → write outcome → delete.
Deep dive — OTP ≤5s during a campaign blast
This is the crux. 1M promos enqueue instantly; providers allow ~5k/s → ~3 min drain — fine for marketing. An OTP landing a million deep waits ~200s. Fail.
Autoscaling workers alone?
- Scale-out takes seconds–minutes; OTP SLA is 5s.
- Clearing 1M in <5s needs ~200k/s — fantasy account limits and cost.
- Don't bet the SLA on elastic magic.
Split queues by priority. High vs standard; drain high first.
Still shared downstream: Twilio budget and worker threads. Standard can starve HP at the provider; a worker stuck on a slow SMS holds the pool.
Surviving provider brownout (p99 1s → 30s):
- Circuit breaker — fail fast on bad error/timeout rate; shed standard load.
- Bulkhead — semaphore cap in-flight calls per provider.
- Provider router — primary/backup (SMS down → email OTP path).
Deep dive — Dedup on top of at-least-once
At-least-once creates duplicates: worker sends then dies before delete; caller retries POST; cron republishes a half-finished segment.
Stable IDs first. Caller idempotency key → notification id. Campaign messages: campaignId:userId. Campaign create needs its own idempotency key or retries mint a new campaignId and double every recipient.
Anti-pattern: SET NX lock before send. Worker claims key, dies before send → redelivery sees lock → drops forever. You converted at-least-once into at-most-once. Yikes.
# Before dispatch
GET sent:{notificationId} # present → duplicate, ack+drop
# Only after provider accepts
SET sent:{notificationId} 1 EX 172800 # 48h covers DLQ replay
# also write SENT row in Postgres (source of truth)
Final design
Additional deep dives & cases
Interviewers often pull one of these — prepare a crisp take.
Backend depth — what staff loops probe
Related Lattice: HLD delivery · multi-step / sagas · scaling writes · rate limiter · long-running tasks.