Skip to content
Kempton Watch Tip

Posts

Built for the Worst Day

The rule I set in advance said don't switch. I did anyway: under a sudden rush, the slowest page went from 1.3 seconds to 14 milliseconds.

By Piet Speurder · · updated · 7 min read

About how it's builtinfrastructureperformancescalingseries

When a story breaks and a crowd arrives in a minute, does the site hold? What it runs on, why it moved to warm workers against its own rule (for the spike, not the stopwatch), the benchmark numbers, and a forecast of what breaks first.

Bar chart "Under a sudden rush", cold start versus warm workers, waiting time in milliseconds: typical 404 vs 6, slow 1,312 vs 14, worst 1,810 vs 23. Lower is better.

Kempton Watch

A record like this one gets read on ordinary days by a handful of people. Then a story breaks, someone with a large following links to it, and a crowd arrives inside a minute. That is exactly the moment the site cannot afford to fall over. So the question for this post is narrow and practical: when everyone comes at once, does it hold?

This is the second of four posts on how the site was built. The first was about what it refuses to publish. This one is about keeping it standing.

What it runs on

Nothing exotic. The site runs on managed hosting, which means a company operates the machines and I operate the software. There is one small application server, a MySQL database for the records, a Redis cache for the things that are asked for over and over, an object store that holds the images, a mail service for the few emails the site sends, and an HTTPS edge in front of all of it that carries the images and the scripts close to the reader. A separate background worker handles the slow jobs so a reader never waits on them.

The pages themselves are built fresh on the server for each visit; they are not held ready at the edge. What makes that quick is the Redis cache: almost all of the traffic is people reading public pages that rarely change, so the cache answers most of what a page needs without the database waking at all. The servers are sized small and most of them sleep when no one is around, then wake in well under a second. The hosting region is Frankfurt, the shortest round trip to Gauteng among the ones on offer.

On top of the public pages there is a machine-readable interface that AI assistants can query directly, and a pipeline on a separate machine that refreshes the data roughly every two hours. Reads are cheap. That is the whole reason a small setup can carry a big day.

The one decision that mattered

There is a choice in how the application server handles a request. The old way starts the program fresh for every request and throws it away afterwards. The newer way keeps a few workers warm and feeds requests through them, so the startup cost is paid once instead of every time. The warm approach is called a persistent-worker runtime; I tested it against the cold one before changing anything.

I set a rule in advance, in writing, so the result could not be argued into whatever I already wanted: switch only if the warm runtime made the slowest page in twenty at least 30 percent faster under normal load. At normal load it made it about 16 percent faster. By my own rule, that was a no. A 16 percent gain on a page that already renders in a few hundredths of a second is real, but no reader would ever feel it.

What changed the decision was not the normal day. It was the bad day. Under a sudden rush, the cold setup queues: requests pile up waiting for a free worker, and the slowest ones go from milliseconds to over a second. The warm setup barely moved. That is a resilience argument, not a speed one, and resilience is the point of this exercise. So I switched, against the letter of my own rule, for the spike, not the stopwatch.

Bar chart "My own rule said no": p95 improvement against a 30% rule set for normal load; normal load 16% fails it, busy hour 44% and sudden rush 99% would clear it.
Kempton Watch

The benchmark, in plain numbers

Everything below is a waiting time in milliseconds, so lower is better. I measured three situations: a normal hour, a busy hour with plenty of room to spare, and a sudden rush. I ran each one repeatedly and took the middle result. No request failed in any run.

At normal load (about twenty readers plus a couple of AI clients), the slowest page in twenty went from 41 ms to 35 ms. The machine-readable interface improved more, from 35 ms to 24 ms. Both setups were comfortable; the server sat at roughly a quarter to a third of one processor.

At a busy hour (a hundred readers), the gap widened. That slow page fell from 25 ms to 14 ms, and the worst one in a hundred fell from 216 ms to 22 ms. More telling was the processor: the cold setup was pinned near 90 percent and running out of room, while the warm one sat at about half. Oddly, the busy hour was quicker than the quiet one for both setups, probably because at a trickle more requests land on a cold path.

Dumbbell chart "Faster when it is busy": p95 response time by page at a hundred readers, cold start to warm workers, sorted by improvement. Lower is better.
Kempton Watch

Under a sudden rush, zero to two hundred readers in thirty seconds, which is roughly what a news mention does, the difference stopped being subtle. The chart at the top of this post shows it. The cold setup's slow page climbed to about 1.3 seconds and its worst to nearly 1.8 seconds, because requests were queuing. The warm setup held at 14 milliseconds. Same code, same data, same small machine. That single comparison is why the switch was worth doing.

The test held the request rate fixed, so throughput cannot separate the two runtimes. This is not about serving more requests a second; it is about not falling apart when they all arrive together. The warm setup uses a little more memory, roughly 180 to 200 MB against 105 MB, which leaves ample headroom on a one-gigabyte machine.

What happens as more people come

Now the forecast. Everything here is an estimate, extrapolated from the measurements above: a planning aid, not a promise. At the busy-hour test the warm server handled about a hundred requests a second at roughly half of one processor, so one server should hold somewhere around a hundred and fifty to a hundred and ninety requests a second before the processor becomes the limit. Readers are not requests. With a few seconds of reading between clicks, one reader is a fraction of a request a second, so that ceiling is several hundred people reading at once.

Scenario Readers at once One small server Rough cost a month First thing to upgrade
Quiet weeks a handful mostly asleep, comfortable US$15–35 nothing
A story breaks a few hundred, briefly holds (tested: 14 ms, nothing failed) US$35–60 hold pages at the edge
Sustained attention several hundred, for weeks meets its limit around 150–190 requests a second US$75–200+ a bigger server, then more of them, and a larger database

The costs are order-of-magnitude estimates, not quotes. The published prices are a United States floor and the Frankfurt region runs higher, so read them as shape, not sums. The spending alarm is set at US$40 a month, which quiet weeks sit under.

Quiet weeks. A few readers, the two-hourly refresh, some crawlers and AI clients. The server mostly sleeps. Nothing needs changing.

A story breaks. A few hundred people arrive over a minute or two, read two or three pages, and leave. This is the rush I tested, and one small server held it. What feels the day first is not the processor; it is image traffic leaving the object store and the database waking on the first request. The edge carries the images, so only the first visitor pays their full cost, and the pages are rebuilt quickly from the warm cache. The first upgrade is to start holding the pages themselves at the edge, which the site does not do yet, not a bigger machine.

Sustained national attention. Weeks of raised traffic, not a single burst. Here one small server does eventually meet its limit. The first hardware upgrade is a bigger server, which is a setting, not a rebuild. Beyond that, running several servers in parallel needs a higher hosting tier, and the database would want a size up as well, with storage and image traffic growing alongside.

In short: the processor on the application server runs out first, and not until well past any traffic this site has seen. Each fix is ordered and cheap: hold pages at the edge, then a bigger server, then more servers. None of it is a redesign.


Showing the Working — how this site was built: What This Site Refuses to Publish · Built for the Worst Day · The Bill, Itemised · Don't Trust Me. Check the Record.

Written by the people who keep this record, under its rules: names only as published by officials or on-the-ground or wire outlets, suspects described and never named, locations at suburb precision. Corrections to hello@kempton.watch.