Scarcity

Scarcity: semaphores & pools

Manage limited resources in LLD: semaphores for concurrent ops and budgets, blocking-queue pools for connections/GPUs, timeouts, and utilization techniques.

Not enough to go around

Connection pool
Five connections in use — sixth waits; leaks make everyone wait forever.
Scarcity is demand exceeding supply: finite DB connections, memory budgets, expensive objects created once. The failure mode isn’t silent corruption — it’s hanging, timeouts, or overload while the downstream system looks “fine.”
Five connections checked out and never returned → new requests block forever. Monitoring may show no crashes; users just spin. You must track in-use capacity, block or fail when empty, and wake waiters on release — always release in finally.

Analogy: limited parking permits

Scarcity permits
Only N permits exist — everyone else waits.

Scarcity is a building with N parking permits: a semaphore hands out permits and takes them back. A connection pool is a set of loaner bikes — acquire one, ride, return it (or replace it if the chain breaks).

Semaphores — limit concurrent operations

A counting lock: N permits. Acquire before work, release in finally. Nightclub bouncer analogy: capacity tokens. Perfect when you don’t need to hand out a specific object — only “at most N in flight.”
from threading import Semaphore


class APIClient:
    def __init__(self, max_in_flight: int = 5):
        self._permits = Semaphore(max_in_flight)

    def request(self, endpoint: str):
        self._permits.acquire()
        try:
            return self._http.get(endpoint)
        finally:
            self._permits.release()
import java.util.concurrent.Semaphore;

class APIClient {
    private final Semaphore permits;

    APIClient(int maxInFlight) {
        this.permits = new Semaphore(maxInFlight);
    }

    Object request(String endpoint) throws InterruptedException {
        permits.acquire();
        try {
            return http.get(endpoint);
        } finally {
            permits.release();
        }
    }
}
  • Download manager — Semaphore(3)
  • Image/video pipeline — cap CPU-heavy jobs
  • External API wrapper — respect concurrent request limits

Resource pooling — hand out real objects

Connections/GPUs have state. A semaphore limits count but doesn’t give you the object. Use a bounded blocking queue preloaded with resources: take/get to acquire, put to release.
from queue import Empty, Queue
import time


class ConnectionPool:
    def __init__(self, size: int, timeout_s: float = 0.5):
        self._q: Queue = Queue(maxsize=size)
        self._timeout = timeout_s
        for _ in range(size):
            self._q.put(self._create())

    def acquire(self):
        try:
            return self._q.get(timeout=self._timeout)
        except Empty:
            raise TimeoutError("no connection available")

    def execute(self, query: str):
        conn = self.acquire()
        try:
            return conn.execute(query)
        finally:
            self._q.put(conn)
import java.util.concurrent.*;

class ConnectionPool {
    private final BlockingQueue<Connection> q;
    private final long timeoutMs;

    ConnectionPool(int size, long timeoutMs) {
        this.q = new ArrayBlockingQueue<>(size);
        this.timeoutMs = timeoutMs;
        for (int i = 0; i < size; i++) {
            q.add(create());
        }
    }

    Connection acquire() throws InterruptedException {
        Connection c = q.poll(timeoutMs, TimeUnit.MILLISECONDS);
        if (c == null) throw new TimeoutException("no connection available");
        return c;
    }

    Object execute(String query) throws Exception {
        Connection conn = acquire();
        try {
            return conn.execute(query);
        } finally {
            q.put(conn);
        }
    }
}
  • Always set queue capacity = pool size (never unbounded).
  • Request paths: poll/get(timeout) — fail with 503, don’t wait forever.
  • Optional: validate before handoff; discard & replace stale connections.

Limit aggregate consumption

Constraint is total units (MB bandwidth, buffer memory), not op count. Semaphore where each permit is 1 MB; acquire size, release size.
MB = 1024 * 1024


class DiskWriter:
    def __init__(self, budget_mb: int = 100):
        self._budget = Semaphore(budget_mb)

    def write(self, data: bytes, path: str) -> None:
        permits = max(1, (len(data) + MB - 1) // MB)
        # acquire N units of budget (Python Semaphore supports n)
        for _ in range(permits):
            self._budget.acquire()
        try:
            open(path, "wb").write(data)
        finally:
            self._budget.release(permits)
import java.util.concurrent.Semaphore;

class DiskWriter {
    private static final int MB = 1024 * 1024;
    private final Semaphore budget;

    DiskWriter(int budgetMb) {
        this.budget = new Semaphore(budgetMb);
    }

    void write(byte[] data, String path) throws Exception {
        int permits = Math.max(1, (data.length + MB - 1) / MB);
        budget.acquire(permits);
        try {
            java.nio.file.Files.write(java.nio.file.Path.of(path), data);
        } finally {
            budget.release(permits);
        }
    }
}

Reuse expensive objects

DB pools, GPU contexts, scarce file handles — BlockingQueue of the real objects. Say: “Creation is expensive, so I’ll pool and return in finally.”

Maximizing utilization (advanced follow-up)

Governance stops overload. Infra / trading / AI loops may ask how to keep scarce resources busy:
  • Work stealing — per-worker queues; idle workers steal (uneven task lengths).
  • Batching — amortize acquire/release; trade latency for throughput.
  • Adaptive sizing — grow/shrink pool with load (tune carefully).

Decision tree

Scarcity decision tree
Ops vs budget vs objects — then utilization techniques.

Worked mini-example: connection pool

pool = Semaphore(10)  # max 10 DB connections

def query(sql):
    pool.acquire()
    try:
        conn = checkout()
        return conn.execute(sql)
    finally:
        release(conn)
        pool.release()
Semaphore pool = new Semaphore(10);  // max 10 DB connections

Object query(String sql) throws InterruptedException {
    pool.acquire();
    Connection conn = null;
    try {
        conn = checkout();
        return conn.execute(sql);
    } finally {
        release(conn);
        pool.release();
    }
}

Scarcity anti-patterns

  • Creating a new DB connection per request with no pool.
  • Semaphore limit without timeout — thread pileup.
  • Pool size = thread count blindly.
  • Ignoring utilization metrics when asked “is 10 enough?”

Scarcity checklist

  1. What resource is finite?
  2. What’s the max concurrent users of it?
  3. Acquire timeout / fail-fast policy?
  4. Release on all error paths?
  5. How do you tune under load?

Common interview pitfalls

These mistakes show up constantly on this prompt. Name the trap, then show the fix in your design — don’t wait for the interviewer to catch you.

  • No limit — resource exhaustion.
  • Limit without timeout.
  • Pool sized randomly.
  • Forgetting release in finally.
  • Mixing scarcity with correctness-only answers.

Interview script (say this)

Read this once out loud before a mock. It’s the spine of a strong answer — not a script to recite robotically.

  1. Scarcity: finite sockets, GPUs, memory, rate budget.
  2. Semaphore or pool caps concurrency.
  3. Acquire with timeout; fail fast under overload.
  4. Always release in finally.
  5. Utilization guides sizing — not vibes.
  6. Example: 10 DB connections shared by 200 handlers.
  7. Reuse objects in pool to avoid allocate churn if relevant.
  8. Tie to rate limiter as scarcity of permits.

Extra verification traces

Walk these three traces on the board. If you can narrate them cleanly, your implementation section usually follows.

Staff-level follow-ups

At staff+, they twist the prompt. Answer in one sentence that names the seam — don’t redesign the whole board.

  • Adaptive pool? — Controller scales pool with latency/CPU signals.
  • Per-tenant budgets? — Hierarchy of semaphores: global + tenant.
  • GPU jobs? — Same pool idea; queue + lease timeout.

Complete solution: connection pool

from queue import Queue, Empty
from threading import Lock

class Connection:
    def __init__(self, cid: int):
        self.id = cid
        self.closed = False

class ConnectionPool:
    def __init__(self, size: int, factory):
        self._factory = factory
        self._q: Queue = Queue(maxsize=size)
        self._created = 0
        self._lock = Lock()
        self._size = size
        for i in range(size):
            self._q.put(factory(i))
            self._created += 1

    def acquire(self, timeout: float = 2.0) -> Connection:
        try:
            return self._q.get(timeout=timeout)
        except Empty:
            raise TimeoutError("pool exhausted")

    def release(self, conn: Connection) -> None:
        if conn.closed:
            # replace broken connection
            with self._lock:
                conn = self._factory(self._created)
                self._created += 1
        self._q.put(conn)

# Semaphore(size) also works if connections are identical and recreate-on-error
# is simple — Queue is clearer when you hand out real objects.
import java.util.concurrent.*;
import java.util.concurrent.locks.ReentrantLock;
import java.util.function.IntFunction;

class Connection {
    final int id;
    boolean closed;

    Connection(int cid) { this.id = cid; }
}

class ConnectionPool {
    private final IntFunction<Connection> factory;
    private final BlockingQueue<Connection> q;
    private int created = 0;
    private final ReentrantLock lock = new ReentrantLock();
    private final int size;

    ConnectionPool(int size, IntFunction<Connection> factory) {
        this.size = size;
        this.factory = factory;
        this.q = new ArrayBlockingQueue<>(size);
        for (int i = 0; i < size; i++) {
            q.add(factory.apply(i));
            created++;
        }
    }

    Connection acquire(long timeoutMs) throws Exception {
        Connection c = q.poll(timeoutMs, TimeUnit.MILLISECONDS);
        if (c == null) throw new TimeoutException("pool exhausted");
        return c;
    }

    void release(Connection conn) throws InterruptedException {
        if (conn.closed) {
            // replace broken connection
            lock.lock();
            try {
                conn = factory.apply(created);
                created++;
            } finally {
                lock.unlock();
            }
        }
        q.put(conn);
    }
}

// Semaphore(size) also works if connections are identical and recreate-on-error
// is simple — BlockingQueue is clearer when you hand out real objects.

← Lattice