On Monday 4 March 2024 at 03:20 UTC, a cache-busting query flood targeted a B2B SaaS customer in Latin America. The attack peaked at 314.7 million requests per second and lasted 149 minutes. Traffic originated from 316 autonomous systems in 23 countries, predominantly a headless-browser farm.
| Vector | cache-busting query flood |
| Peak | 314.7 million requests per second |
| Duration | 149 min |
| Time to mitigation | 0.554 s |
| Attack traffic reaching origin | 0.046% |
| Legitimate traffic challenged | 0.02% |
Timeline
The attack was preceded by a hacktivist channel announcing the target. Request rates on the targeted routes exceeded their hourly baseline by a factor of 666 within 26 seconds. The risk score of participating clients crossed the challenge threshold automatically and proof-of-work difficulty rose with origin load.
What the customer saw
A brief increase in p99 latency of 138 ms during the first minute, then normal service.
Recommendations
- Enable authenticated origin pulls.
- Add a dedicated rate limit for the targeted route.
- Keep origin IPs out of public DNS history.
We moved from a scrubbing provider to always-on last year; time to mitigation went from minutes to basically nothing.
We had the exact false-positive issue with CGNAT carriers. The weighting change makes sense.
Thanks! Yes — the risk score and its components are included in every log record.
Is the risk score exposed in the logs so we can build our own dashboards on it?
How do you avoid challenging uptime monitors and partners?