On Wednesday 18 August 2021 at 21:30 UTC, a cache-busting query flood targeted a electronics retail customer in Southeast Asia. The attack peaked at 358.5 million requests per second and lasted 78 minutes. Traffic originated from 1621 autonomous systems in 53 countries, predominantly compromised cloud VMs.
| Vector | cache-busting query flood |
| Peak | 358.5 million requests per second |
| Duration | 78 min |
| Time to mitigation | 0.566 s |
| Attack traffic reaching origin | 0.055% |
| Legitimate traffic challenged | 0.46% |
Timeline
The attack was preceded by no stated motive. Request rates on the targeted routes exceeded their hourly baseline by a factor of 294 within 14 seconds. The risk score of participating clients crossed the challenge threshold automatically and proof-of-work difficulty rose with origin load.
What the customer saw
No customer-visible impact. The on-call engineer was notified and acknowledged the incident from the dashboard.
Recommendations
- Lower challenge thresholds on authentication endpoints during high-risk events.
- Enable authenticated origin pulls.
- Review allow-listed partner ranges quarterly.
Any plans to support per-tenant limits keyed on a JWT claim?
Great to hear, thanks for sharing your experience.
The billing model is what got our finance team on board, honestly.
How does the proof-of-work challenge behave on older Android devices? Any numbers below Android 10?
Great write-up. We saw almost the same pattern on our login endpoint last month.
Is the risk score exposed in the logs so we can build our own dashboards on it?
We had the exact false-positive issue with CGNAT carriers. The weighting change makes sense.
Could you share the dataset behind the percentages?
Nice to read a vendor blog that admits what went wrong.