Your API is live. Traffic is growing. Suddenly someone hits your endpoint 10,000 times in one minute. What happens?
Without rate limiting: Your server goes down, your real users can't access the service, and you're in for a rough morning.
With rate limiting: The abuser gets a polite 429 Too Many Requests, and everyone else carries on as normal.
What is rate limiting?
Rate limiting restricts how many requests a client can send within a time period. It's your first line of defense against:
- DoS attacks - attempts to overload your server
- Brute force - login attempts guessing passwords
- Scraping - bots harvesting your data
- Accidental loops - buggy client code spamming your API
The four classic algorithms
1. Fixed Window
The simplest approach. Count requests in fixed time intervals (e.g., per minute).
Pros:
- Super simple to implement
- Low memory overhead
Cons:
- "Burst problem" - a user can send 100 requests at 12:59 and 100 more at 13:00, meaning 200 requests in 2 minutes while the limit is 100/minute
2. Sliding Window
Solves the burst problem by tracking requests over a sliding time period.
Pros:
- Smoother distribution of requests
- No edge-case at window boundaries
Cons:
- Requires more memory (timestamps per request)
- Slightly more complex to implement
This is our preferred choice for most use cases.
3. Token Bucket
Think of it as a bucket of tokens. Each request consumes one token. The bucket refills at a constant rate.
Pros:
- Allows bursts (if the bucket is full)
- Models realistic user behavior
Cons:
- More complex state management
Perfect for APIs where users naturally have periods of inactivity followed by bursts.
4. Leaky Bucket
Requests enter a queue and are processed at a constant rate, like water dripping from a bucket.
Pros:
- Guaranteed even load on the server
- No sudden spikes
Cons:
- Can feel slow to the user
- Requests can queue up and create latency
Often used for payment APIs and other critical systems.
Which algorithm should you choose?
| Use Case | Recommendation |
|---|---|
| General API protection | Sliding Window |
| Allow bursts from legitimate users | Token Bucket |
| Critical endpoints (payments, login) | Leaky Bucket |
| Simple prototype | Fixed Window |
Implementation in practice
HTTP Response Headers
Always tell the client what's happening:
HTTP/1.1 200 OK
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 42
X-RateLimit-Reset: 1704830400When the limit is reached
HTTP/1.1 429 Too Many Requests
Retry-After: 60
Content-Type: application/json
{
"error": "rate_limit_exceeded",
"message": "You have sent too many requests. Try again in 60 seconds.",
"retry_after": 60
}Tiered Rate Limits
Not all users are equal. Differentiate your limits:
| Tier | Requests/minute | Use case |
|---|---|---|
| Anonymous | 10 | Unauthorized calls |
| Free | 60 | Free users |
| Pro | 300 | Paying customers |
| Enterprise | 3000+ | Large customers with SLA |
Per-Endpoint Limits
Some endpoints are more expensive than others:
const limits = {
'GET /users': { max: 100, window: '1m' },
'POST /upload': { max: 10, window: '1m' },
'POST /ai/generate': { max: 5, window: '1m' },
};Where do you implement it?
1. API Gateway (recommended)
Cloudflare, AWS API Gateway, Kong, or Nginx handle it before your code even runs.
Pros: Scalable, distributed, no code changes required.
2. Middleware
Express, Fastify, or your framework's middleware layer.
import rateLimit from 'express-rate-limit';
const limiter = rateLimit({
windowMs: 60 * 1000, // 1 minute
max: 100,
standardHeaders: true,
message: { error: 'Too many requests' }
});
app.use('/api/', limiter);3. Database/Redis
Use Redis for distributed rate limiting across multiple servers:
// Simple sliding window with Redis
const key = `ratelimit:${userId}:${endpoint}`;
const count = await redis.incr(key);
if (count === 1) {
await redis.expire(key, 60);
}
if (count > 100) {
throw new RateLimitError();
}2026 trend: AI-driven rate limiting
With more AI agents calling APIs, we're seeing a new approach:
- Behavior analysis - don't just count requests, analyze patterns
- Threat scoring - dynamic adjustment based on risk
- Adaptive limits - automatic scaling under load
This is still early, but keep an eye on tools from Cloudflare and Kong.
Our recommendation
For most projects:
- Start with an API gateway (Cloudflare is free for basics)
- Use Sliding Window as the default algorithm
- Implement tiered limits from day one
- Log everything - you'll thank yourself later
Rate limiting isn't just security. It's also documentation of your API's capacity and a way to protect your infrastructure budget.