How to Rate-Limit Stolen Large-Model API Calls: Stop the Bleeding First, Then Set Tiered Thresholds by Token

When large-model APIs are abused, IP-based RPS limits often fail because cost lies in tokens and inference time; this article gives the first-hour containment sequence and explains the four dimensions rate limiting must cover, plus gateway composite-key configurations.