Learning path

Foundations

Numbers, APIs, data, distributed theory

  1. Numbers to know 2026 hardware ballparks for caches, databases, app servers, and queues — and the interview mistakes that come from using 2015 numbers (premature sharding, fake write bottlenecks, unnecessary caches).
  2. System design concepts roadmap The foundations learning path in one page — what to read, in what order, and the default choice for each topic. Skim here; open the linked articles for depth.
  3. Networking essentials The networking calls that actually move your whiteboard: TCP vs UDP, REST vs gRPC, SSE vs WebSockets vs WebRTC, L4 vs L7 load balancing, and the failure vocabulary — timeouts, backoff, idempotency, circuit breakers. One analogy family throughout: Indian Railways and India Post.
  4. API design The 5-minute API slice, run through one Ticketmaster/IRCTC booking flow: REST resources, GraphQL and RPC when they're actually earned, pagination, versioning, and JWT vs. API keys — with the definitional fluff cut.
  5. Data modeling What “good enough” schema design looks like in 45 minutes: Postgres by default, entities and keys tied to your APIs, indexes for real queries, when to denormalize, and when not to reach for Mongo, Cassandra, or a graph DB.
  6. Database indexing How indexes turn table scans into lookups: B-trees as the default, when LSM trees win for writes, hash vs range, geospatial options, inverted indexes for text, plus composite and covering patterns you’ll actually defend on a whiteboard.
  7. Proximity search Why B-trees fail on nearby queries, spatial trees (quadtree, k-d/BKD, R-tree) vs encoded keys (geohash, S2, H3), Haversine post-filters, Redis/PostGIS/ES production patterns, and how Uber-scale dispatch actually shards.
  8. Scalability fundamentals Vertical vs horizontal scale, bottlenecks, and how to talk about growth without hand-waving — from IPL finals to coalition arithmetic.
  9. Sharding When one database hits the ceiling: partitioning vs sharding, how to pick a shard key, range vs hash vs directory distribution, then hot spots, cross-shard queries, and consistency — without sharding before the napkin says so.
  10. Consistent hashing Why hash(key) % N explodes when you add a node, how a hash ring fixes remapping, virtual nodes for even failure load, and when to deep-dive vs just name DynamoDB/Cassandra in interviews.
  11. CAP theorem What CAP really means in interviews: partition tolerance is required, so you choose consistency or availability when the network splits — with examples, hybrid designs, and how to open the NFR phase.
  12. Consistency models Strong, eventual, and read-your-writes — when each is enough, grounded in EVM counts, WhatsApp read receipts, and the read-your-writes bug that breaks shopping carts.
  13. Specialized data structures at scale When hash tables won't fit: Bloom filters for membership, Count-Min Sketch for frequency, HyperLogLog for cardinality, and histogram buckets for quantiles — what they buy you, when to reach for them, and when simple scaling is enough.

Lattice