Let's talk
Reliable infrastructure Media & Publishing · 20 January 2025

Making a 20 million visit platform boring again

A high-traffic content platform kept falling over during its busiest moments. Here's how we made deployments calm and traffic spikes a non-event.

Industry Media & Publishing
Focus Reliable infrastructure
Outcome Deploys stopped being an event and traffic spikes stopped being a crisis

The problem

This was a content platform doing more than 20 million visits a month, big enough that a bad hour costs real money and real reputation. The trouble was, “big traffic day” and “site’s about to fall over” had become the same sentence.

Certain stories or campaigns would send traffic through the roof, and instead of that being good news, it was a fire drill. Pages loaded slowly, some visitors got errors, and the team ended up firefighting on the exact days they most needed things to work. On top of that, releasing new code had become something people quietly dreaded. Deployments happened late at night or in narrow windows because nobody trusted them to go smoothly during the day. That’s a rough way to build software: your best people spending their energy managing risk instead of building things.

Underneath it, the infrastructure had grown the way a lot of successful products do. Bits added under pressure, workarounds nobody had time to revisit, and a general sense that nobody had a full picture of how it all held together. It worked, mostly, until it didn’t.

What we did

I started by figuring out where the platform actually broke under load, not where people assumed it broke. That meant properly stress-testing it against realistic traffic patterns rather than guessing, and it turned up a few surprises: some of the slowest, most fragile parts weren’t the obvious ones.

From there, the work fell into a few strands. We rebuilt how the platform handled sudden surges in visitors, so it could absorb a spike without buckling, rather than relying on someone spotting trouble and reacting fast enough. We changed how content and images got served to the busiest pages, taking load off the core system so it wasn’t doing unnecessary work for every single visitor. And we put proper monitoring in place, the kind that tells you a problem is forming before your visitors notice, not after your phone starts ringing.

Just as importantly, we rebuilt how new code got released. Deployments went from a nervy, manual process to something automated, tested, and reversible. If something did go wrong after a release, the team could roll it back in minutes rather than scrambling to work out what broke and why. That single change did more for morale than almost anything else, because it meant releasing software stopped being something people had to brace for.

None of this was about ripping everything out and starting again. Big platforms with real traffic can’t afford a risky rewrite, so it was a case of finding the weak points and reinforcing them one at a time, while the site kept running and kept serving readers.

Where things landed

Big traffic days stopped being an emergency. When a story took off, the platform handled it, and the team found out about it from the analytics rather than from an alert going off. That’s a genuinely different way to run a media business: growth in attention becomes something to welcome, not something to fear.

Deployments changed character too. What used to be a late-night, hold-your-breath process became routine enough that it could happen during the working day, with people confident enough in the rollback process that a bad release was an inconvenience rather than a crisis.

The wider effect was on the engineering team’s time. Less of it went on firefighting and reacting, more of it went on the product itself. For a business built on getting content in front of readers fast, having infrastructure you can trust rather than merely tolerate makes a real difference to what the team can actually get done.

Facing something similar?

If this sounds like where your business is right now, let's talk about it.

Get in touch