We get asked about tail latency is the only latency that matters more than almost anything else, so here is the long answer.
Latency budget
Our budget for the whole filtering pipeline is one millisecond at the 99th percentile. Anything that cannot be decided within that budget runs asynchronously and influences the next request from the same client, not the current one.
That constraint shapes everything: data structures, where state lives and which signals we are willing to compute inline.
What actually happens during a flood
The first thing to fail during a layer 7 flood is almost never bandwidth. It is connection slots, worker processes or database connections on the origin — resources measured in hundreds or thousands, not gigabits. An attacker who can make each request expensive only needs a few thousand requests per second.
That is why we score requests before they are proxied. By the time a request reaches your origin, it has been attributed to a client, compared against that client’s history and weighed against the current load on the route it targets.
The numbers
Across the last quarter, 71% of challenged clients never attempted a solution, 24% solved one challenge and then behaved normally, and 5% solved challenges repeatedly while continuing to attack — the last group is where analysts spend their time.
Median added latency for legitimate visitors that were challenged was 280 ms on desktop and 410 ms on mid-range Android devices.
Observability first
Every mitigation decision is logged with its reasons, and every log line can be traced to the rule and score components that produced it. When a customer asks why a request was challenged, the answer is a link, not a guess.
Logs stream to the customer’s SIEM within seconds, which also means their security team sees attacks in the same tools they use for everything else.
Why per-route baselines matter
A thousand requests per second to your homepage is Tuesday. A thousand requests per second to your password-reset endpoint is an attack. Global rate limits cannot tell the difference; per-route baselines can.
We learn the normal shape of traffic per route and per hour of the week, so a surge on a sensitive endpoint raises the risk score long before it approaches a global threshold.
As always, questions and corrections are welcome at [email protected].
How does the proof-of-work challenge behave on older Android devices? Any numbers below Android 10?
Could you share the dataset behind the percentages?
Verified good bots and allow-listed partners bypass challenges entirely.
Great write-up. We saw almost the same pattern on our login endpoint last month.
Our auditors asked for exactly this kind of incident evidence under DORA.
Would love a follow-up on how you handle HTTP/3 fingerprinting.
Any plans to support per-tenant limits keyed on a JWT claim?
How does the proof-of-work challenge behave on older Android devices? Any numbers below Android 10?
Our auditors asked for exactly this kind of incident evidence under DORA.
Machine-readable ranges are at /ips.json and via the API.
The point about origin IPs leaking through certificate transparency logs is underrated.