This post is about a practical checklist for surviving Black Friday traffic. It started, as most of our posts do, with an incident that did not go the way we expected.
Lessons for your own runbook
Know who can change DNS at two in the morning. Know your origin IPs and who can rotate them. Know which routes are expensive, and have a rate limit ready for each of them.
Most outages during attacks are not caused by the attack itself but by rushed changes made while under pressure.
What we changed
We moved the decision from a single threshold to a continuous score, added an explanation to every decision and made every rule testable against historical traffic before it goes live.
The result is fewer late-night pages for our analysts and — more importantly — fewer real users challenged by mistake.
Observability first
Every mitigation decision is logged with its reasons, and every log line can be traced to the rule and score components that produced it. When a customer asks why a request was challenged, the answer is a link, not a guess.
Logs stream to the customer’s SIEM within seconds, which also means their security team sees attacks in the same tools they use for everything else.
What actually happens during a flood
The first thing to fail during a layer 7 flood is almost never bandwidth. It is connection slots, worker processes or database connections on the origin — resources measured in hundreds or thousands, not gigabits. An attacker who can make each request expensive only needs a few thousand requests per second.
That is why we score requests before they are proxied. By the time a request reaches your origin, it has been attributed to a client, compared against that client’s history and weighed against the current load on the route it targets.
The economics behind it
A booter service rents out a 100 Gbps attack for less than the price of a pizza. Defending against it with bandwidth alone is a race you lose. Defending against it by making each malicious request cost more than it earns is a race you win.
This is the entire idea behind never metering attack traffic: our costs scale with filtering, not with your invoice.
{ "match": { "path": "/login", "risk": ">= 57" }, "action": "challenge" }
The full incident data behind this post is available to customers in the dashboard under Reports.
Do you publish the edge IP ranges in a machine-readable format?
Any plans to support per-tenant limits keyed on a JWT claim?
Solid runbook advice. The DNS-at-2am point hit home.
Thanks! Yes — the risk score and its components are included in every log record.
Thanks — sharing this with our on-call team.
The point about origin IPs leaking through certificate transparency logs is underrated.
Good question. We will cover that in a follow-up post.
Do you publish the edge IP ranges in a machine-readable format?
Thanks! Yes — the risk score and its components are included in every log record.
Could you share the dataset behind the percentages?
Nice to read a vendor blog that admits what went wrong.
Nice to read a vendor blog that admits what went wrong.