On HelpDigiSchool, every meaningful event — a grade published, an assignment posted, a report card ready, a payment past due — has to reach the right people: school founders and administrators, teachers, secondary school students checking their homework, university students, parents. Over 1,000 people can be connected at the same time, and each of them needs to hear about it across several channels at once: an in-app bell, a browser push, an email, sometimes WhatsApp.
The classic trap is sending all of that synchronously, inside the request thread that triggered the event. It works great — until the day an admin clicks "remind all overdue accounts" for 1,000 students at once, or a poorly targeted query reaches far more recipients than intended. That's exactly the kind of incident that pushed me to rethink the sending architecture instead of stacking patches on top of it.
Staying fast before a single notification is even sent
With that many concurrent users, the first risk isn't sending notifications — it's the perceived slowness of simply connecting to the app.
The clearest symptom: a page taking over 30 seconds to load grades for a 170-student class, with requests timing out. The cause is almost always the same: N+1 queries hitting the database student by student instead of fetching the data in one round trip, combined with missing or poorly targeted caching. Eliminating those N+1 queries and adding caching on the most-requested data brought that load time under 3 seconds.
The second symptom only shows up under real load: a connection pool (database, cache) sized for normal usage runs dry the moment an unusual number of users connect at once, triggering a cascade of 503 errors. Diagnosing that exhaustion and resizing the pool accordingly cut the failure rate under heavy load from 75% to under 10%.
Neither fix has anything to do with notifications themselves — but without them, no queue architecture would matter: there's no point keeping sends from slowing down the app if simply connecting to it is already the bottleneck.
Never send inline
The first rule I set for myself: a user action should never wait on an external send to finish. Creating a notification is a database write — fast, inside the same transaction as the business action. The actual delivery (push, email, WhatsApp) is handed off elsewhere, and the request returns immediately.
Concretely, that means separating two things we tend to conflate: accepting a send request and executing it. Accepting is instant. Executing can take time, fail, need a retry — and must never block anyone.
A queue instead of a thread pool
For the least forgiving channels — WhatsApp above all, where sending too fast risks getting the number banned — I built a real queue: every message to send becomes a database row, with a status (pending, processing, sent, failed) and a priority. A scheduled worker picks one message at a time, in priority order, and actually sends it.
Using a table instead of a conventional message broker (RabbitMQ, Kafka...) isn't a fallback compromise: on infrastructure running several replicas, the database is already the shared source of truth, and a row-level lock at claim time (SELECT ... FOR UPDATE SKIP LOCKED) is enough to guarantee a message is never processed twice, with no extra coordination between instances.
A deliberately slow throughput
The counter-intuitive part: the fix isn't sending faster, it's sending slower — and irregularly. A channel like WhatsApp penalizes anything that looks too mechanical. So the scheduled worker adds a random delay between sends, a longer pause every few messages, a daily quota per sender, and a "human" sending window (nothing at 3am). On repeated provider errors, a temporary circuit breaker halts the whole channel instead of hammering it.
For more forgiving channels — email, browser push — the calculation is different: they run as async tasks on a dedicated thread pool with a bounded queue capacity, absorbing a spike without it ever reaching the threads serving HTTP requests.
What I take from it
The natural instinct when facing a load spike is to add threads, bigger batches, more parallelism. Here, the right answer was the opposite: decouple acceptance from execution, then deliberately slow down execution to a pace the downstream system — a database shared across several replicas, a WhatsApp API that bans suspicious behavior — can sustain indefinitely. A slow, reliable queue always beats a fast send that eventually falls over.