Notification platform with priority queues and multi-channel delivery

Design a notification system

Full HLD: notification platform for push/email/SMS — campaigns, preferences, durable queues, priority isolation for OTPs, and best-effort dedup under 5k/s surges.

At a glance

Central post office analogy
Teams drop letters — the platform sorts, schedules, and hands off to couriers.
011:1 + campaigns

Now or scheduled.

02Preferences

Opt-out · quiet hours.

03Durable queue

Ack ≠ delivered.

04Priority tiers

Queues · workers · budgets.

05Dedup

Stable IDs · mark after send.

06Brownout

Breaker · bulkhead · failover.

Takeaway

Build a working send path first, then isolate priority traffic and harden delivery semantics. Don't open with Kafka topology before FR/NFR and entities exist.

Whiteboard order

  1. Requirements — FR above the line; NFR with surge math + two traffic classes.
  2. Entities — Notification · Campaign · Segment · User (+ prefs / devices).
  3. API — POST /notifications, POST /campaigns, PUT preferences/devices; priority + idempotency.
  4. HLD v1 — sync send → scheduler → campaigns → prefs (working system).
  5. Deep dives — durable queue → priority isolation → dedup → brownout.
Takeaway

Resist solving 5k/s and OTP SLA before a single email can leave the building.

What is a notification system?

A notification system is an internal platform other teams use to message users across channels (push, email, SMS). Product services hand it a message and a recipient; the platform owns routing, scheduling, preferences, retries, and bookkeeping.

Customers are engineering teams, not end users: auth sending OTPs, orders confirming purchases, marketing blasting a promo to a million users. That shapes auth (service API keys), SLAs (OTP vs promo), and APIs (act-on-behalf-of userId).

Requirements

Functional (steer here)

  • Send a notification to a user via push, email, or SMS — immediately or scheduled.
  • Send campaigns: same message to a segment, immediate or scheduled.
  • Users set preferences: channel opt-outs and quiet hours.

Below the line (skip if low on time)

  • Template management / rich authoring.
  • Delivery analytics dashboards (open/CTR).
  • In-app notification feeds and badge counts.
  • Frequency capping across types (nice follow-up later).
Surge math
Daily average hides the campaign cliff — design for surge.

Capacity gut-check. 10M notifications/day ≈ 100/s if even — almost nothing. But a 1M campaign in 5 minutes is >3,000/s. Design for surges of ~5,000/s (~50× sustained).

Traffic classes
High priority vs standard — opposite SLAs on one platform.

Non-functional

  • At-least-once delivery with best-effort deduplication (prefer duplicate over drop).
  • Handle surges of ~5,000 notifications/sec.
  • High priority (OTP, security) delivered within 5 seconds of acceptance even during a surge.

NFR below the line

  • Strict ordering across channels.
  • Full compliance/spam legal workflows.
  • Monitoring/CI as first-class design (mention, don't boil the ocean).
Functional:
1. Send to a user (now | scheduled) via push | email | sms
2. Campaigns to a segment (now | scheduled)
3. User preferences (opt-outs, quiet hours)

Non-functional:
1. At-least-once + best-effort dedup
2. ~5,000 notif/s surge
3. Critical ≤5s even during surge

Core entities

Start broad — columns come later. Talk these through with the interviewer.

API / system interface

Start simple; grow idempotency keys as you narrate evolution.

POST /notifications → 202 Accepted { id }   # after async; early MVP may be 200
Body: {
  userId,
  channel,      // push | email | sms
  priority,     // high | standard   ← required for SLA isolation
  content,      // { title, body }
  scheduledAt,  // optional, default now
  idempotencyKey
}

GET /notifications/{id} → { id, status, updatedAt }

POST /campaigns → 202 Accepted { id }
Body: {
  segmentId, channel, priority, templateId,
  scheduledAt, idempotencyKey
}

PUT /users/{userId}/preferences → Preferences
Body: { optOuts: ["sms"], quietHours: { start, end, tz } }

PUT /users/{userId}/devices → Device
Body: { platform: ios|android, pushToken }

HLD 1 — Single send (now + scheduled)

Build the direct path first: one email, right now. SMS is the same shape; push adds device tokens.

MVP direct send
Synchronous accept + deliver — fine for learning, fragile for NFRs.
  • API Gateway — authenticate service API keys; rate-limit runaways so auth OTPs survive.
  • Notification Service — lookup address, persist row, call provider.
  • Postgres — Notification rows + synced slice of user contact data.
  • Email/SMS provider — last mile (SMTP/API). You don't own inbox delivery.

Send-now flow

  1. Upstream POSTs /notifications.
  2. Gateway authenticates and forwards.
  3. Service loads email/phone; writes Notification PENDING; calls provider.
  4. Updates SENT/FAILED; returns id + status (MVP 200).
Multi-channel
Router + bookkeeper; providers own carrier/device connections.

Push: app calls PUT /users/{id}/devices on login/token rotate. Send path looks up token instead of email.

Scheduled: write status SCHEDULED; stop. A scheduler cron (~1 min) finds due rows and pushes them through the same send path. OTP never uses cron — it arrives as an immediate high-priority request.

HLD 2 — Campaigns

POST /campaigns writes a Campaign row and returns 202. Expand the scheduler to poll due campaigns too.

Campaign fan-out
Definition → per-user notification instances.
  1. Cron finds Campaign with scheduledAt ≤ now.
  2. Snapshot segment membership → recipient list.
  3. For each user, enqueue/send via the same path as 1:1.
  4. Mark campaign complete (or FAILED with checkpoint — deep dive).

HLD 3 — Preferences

No new components. Preferences live on User; PUT via gateway → notification service.

Takeaway

Working system: sync OTP path, campaign fan-out via cron, preferences enforced at dispatch. NFRs still unsolved — that's the deep-dive half.

Deep dive — Never drop an accepted notification

Pattern: multi-step process (accept → enqueue → deliver → record). Each step can fail alone. Temporal-level orchestration is usually waved off — show the mechanics.

Approach A — retries in-process. Keep sync call; backoff on provider fail.

  • Fixes flaky single responses only.
  • Crash/deploy still loses in-memory retries.
  • Holds caller open during outage → auth timeouts.
  • Caller can't tell give-up vs never-got-it → effectively at-most-once.

Approach B — Postgres as queue. Write PENDING, 202, pollers SELECT … FOR UPDATE SKIP LOCKED, retry with attempt/next_run.

  • Survives restarts.
  • You reinvent a mediocre queue (retries, visibility, hot table of history).
  • Latency ≥ poll interval — sluggish for OTP.

Approach C — durable queue (preferred). Enqueue to SQS (or similar); 202 when queue acks. Workers: receive → prefs → provider → write outcome → delete.

Durable queue
Ack on enqueue; visibility timeout + DLQ give at-least-once.

Deep dive — OTP ≤5s during a campaign blast

This is the crux. 1M promos enqueue instantly; providers allow ~5k/s → ~3 min drain — fine for marketing. An OTP landing a million deep waits ~200s. Fail.

Autoscaling workers alone?

  • Scale-out takes seconds–minutes; OTP SLA is 5s.
  • Clearing 1M in <5s needs ~200k/s — fantasy account limits and cost.
  • Don't bet the SLA on elastic magic.

Split queues by priority. High vs standard; drain high first.

Still shared downstream: Twilio budget and worker threads. Standard can starve HP at the provider; a worker stuck on a slow SMS holds the pool.

Priority isolation
Queues + worker pools + Redis token budgets.

Surviving provider brownout (p99 1s → 30s):

  • Circuit breaker — fail fast on bad error/timeout rate; shed standard load.
  • Bulkhead — semaphore cap in-flight calls per provider.
  • Provider router — primary/backup (SMS down → email OTP path).

Deep dive — Dedup on top of at-least-once

At-least-once creates duplicates: worker sends then dies before delete; caller retries POST; cron republishes a half-finished segment.

Stable IDs first. Caller idempotency key → notification id. Campaign messages: campaignId:userId. Campaign create needs its own idempotency key or retries mint a new campaignId and double every recipient.

Anti-pattern: SET NX lock before send. Worker claims key, dies before send → redelivery sees lock → drops forever. You converted at-least-once into at-most-once. Yikes.

Dedup flow
Check cache/DB before send; write marker only after provider accepts.
# Before dispatch
GET sent:{notificationId}     # present → duplicate, ack+drop

# Only after provider accepts
SET sent:{notificationId} 1 EX 172800   # 48h covers DLQ replay
# also write SENT row in Postgres (source of truth)

Final design

Final design: upstream to API gateway to notification service, high and standard queues, workers with Redis token bucket and dedup, SMTP APNs/FCM Twilio, and Postgres schemas
Final design — inspired by the classic whiteboard: priority queues, Redis (token bucket + dedup), schema cards, and provider diamonds.

Additional deep dives & cases

Interviewers often pull one of these — prepare a crisp take.

Backend depth — what staff loops probe

Related Lattice: HLD delivery · multi-step / sagas · scaling writes · rate limiter · long-running tasks.

← Lattice