Skip to content
TBH
Back to blog
January 9, 20264 min read

API Rate Limiting: How to Protect Your Backend

Your API is live. Traffic is growing. Suddenly someone hits your endpoint 10,000 times in one minute. What happens?

Without rate limiting: Your server goes down, your real users can't access the service, and you're in for a rough morning.

With rate limiting: The abuser gets a polite 429 Too Many Requests, and everyone else carries on as normal.

What is rate limiting?

Rate limiting restricts how many requests a client can send within a time period. It's your first line of defense against:

  • DoS attacks - attempts to overload your server
  • Brute force - login attempts guessing passwords
  • Scraping - bots harvesting your data
  • Accidental loops - buggy client code spamming your API

The four classic algorithms

1. Fixed Window

The simplest approach. Count requests in fixed time intervals (e.g., per minute).

Pros:

  • Super simple to implement
  • Low memory overhead

Cons:

  • "Burst problem" - a user can send 100 requests at 12:59 and 100 more at 13:00, meaning 200 requests in 2 minutes while the limit is 100/minute

2. Sliding Window

Solves the burst problem by tracking requests over a sliding time period.

Pros:

  • Smoother distribution of requests
  • No edge-case at window boundaries

Cons:

  • Requires more memory (timestamps per request)
  • Slightly more complex to implement

This is our preferred choice for most use cases.


3. Token Bucket

Think of it as a bucket of tokens. Each request consumes one token. The bucket refills at a constant rate.

Pros:

  • Allows bursts (if the bucket is full)
  • Models realistic user behavior

Cons:

  • More complex state management

Perfect for APIs where users naturally have periods of inactivity followed by bursts.


4. Leaky Bucket

Requests enter a queue and are processed at a constant rate, like water dripping from a bucket.

Pros:

  • Guaranteed even load on the server
  • No sudden spikes

Cons:

  • Can feel slow to the user
  • Requests can queue up and create latency

Often used for payment APIs and other critical systems.


Which algorithm should you choose?

Use CaseRecommendation
General API protectionSliding Window
Allow bursts from legitimate usersToken Bucket
Critical endpoints (payments, login)Leaky Bucket
Simple prototypeFixed Window

Implementation in practice

HTTP Response Headers

Always tell the client what's happening:

HTTP/1.1 200 OK
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 42
X-RateLimit-Reset: 1704830400

When the limit is reached

HTTP/1.1 429 Too Many Requests
Retry-After: 60
Content-Type: application/json
 
{
  "error": "rate_limit_exceeded",
  "message": "You have sent too many requests. Try again in 60 seconds.",
  "retry_after": 60
}

Tiered Rate Limits

Not all users are equal. Differentiate your limits:

TierRequests/minuteUse case
Anonymous10Unauthorized calls
Free60Free users
Pro300Paying customers
Enterprise3000+Large customers with SLA

Per-Endpoint Limits

Some endpoints are more expensive than others:

const limits = {
  'GET /users': { max: 100, window: '1m' },
  'POST /upload': { max: 10, window: '1m' },
  'POST /ai/generate': { max: 5, window: '1m' },
};

Where do you implement it?

1. API Gateway (recommended)

Cloudflare, AWS API Gateway, Kong, or Nginx handle it before your code even runs.

Pros: Scalable, distributed, no code changes required.

2. Middleware

Express, Fastify, or your framework's middleware layer.

import rateLimit from 'express-rate-limit';
 
const limiter = rateLimit({
  windowMs: 60 * 1000, // 1 minute
  max: 100,
  standardHeaders: true,
  message: { error: 'Too many requests' }
});
 
app.use('/api/', limiter);

3. Database/Redis

Use Redis for distributed rate limiting across multiple servers:

// Simple sliding window with Redis
const key = `ratelimit:${userId}:${endpoint}`;
const count = await redis.incr(key);
if (count === 1) {
  await redis.expire(key, 60);
}
if (count > 100) {
  throw new RateLimitError();
}

2026 trend: AI-driven rate limiting

With more AI agents calling APIs, we're seeing a new approach:

  • Behavior analysis - don't just count requests, analyze patterns
  • Threat scoring - dynamic adjustment based on risk
  • Adaptive limits - automatic scaling under load

This is still early, but keep an eye on tools from Cloudflare and Kong.


Our recommendation

For most projects:

  1. Start with an API gateway (Cloudflare is free for basics)
  2. Use Sliding Window as the default algorithm
  3. Implement tiered limits from day one
  4. Log everything - you'll thank yourself later

Rate limiting isn't just security. It's also documentation of your API's capacity and a way to protect your infrastructure budget.

See what we have built

Nordvec, nævn.dk, Matematik i Måneby in the browser, Semantika and open source.

See the projects