Rate Limiting Implementation
This document outlines the rate limiting strategy implemented for the SparkyFitness application, focusing on protecting sensitive authentication endpoints.
Purpose
Rate limiting is crucial for:
- Preventing Brute-Force Attacks: Limiting the number of login attempts from a single IP address within a given time frame.
- Mitigating Denial-of-Service (DoS) Attacks: Restricting the rate of requests to prevent server overload.
- Preventing Account Creation Spam: Limiting the rate of new user registrations.
- Protecting Password Reset Flows: Preventing abuse of email-based recovery mechanisms.
- Securing MFA Endpoints: Limiting attempts to bypass multi-factor authentication.
Implementation Layers
SparkyFitness rate-limits in two places:
- Nginx (frontend container) — a coarse cap across all
/api/auth/traffic. - The application (Better Auth) — per-IP caps on sign-in and per-key caps on API keys. These are the meaningful controls; see the sections below.
Nginx Configuration
Defined in docker/nginx.conf, which is copied into the frontend image as /etc/nginx/templates/default.conf.template and rendered by docker-entrypoint.sh.
limit_req_zone $binary_remote_addr zone=login_signup_zone:10m rate=${NGINX_RATE_LIMIT};The rate is set by NGINX_RATE_LIMIT (default 5r/s), substituted at container start by docker/docker-entrypoint.sh. Set it to a very high value such as 10000r/s to effectively disable this layer.
A single blanket location covers every auth endpoint:
location ^~ /api/auth/ {
limit_req zone=login_signup_zone burst=10 nodelay;
proxy_pass http://$backend;
# ... proxy headers ...
}burst=10: allows a burst of 10 requests beyond the configured rate.nodelay: excess requests are rejected immediately with429rather than queued.
A caveat worth knowing
WARNING
The zone keys on $binary_remote_addr, which is the address of whatever proxy sits in front of the container. If you run behind a reverse proxy, CDN or tunnel (Nginx Proxy Manager, Cloudflare Tunnel), that is a single address for every visitor — so all users share one bucket rather than getting one each. The per-IP protection that actually distinguishes clients is the application-layer sign-in limit described below, which reads X-Forwarded-For.
API Key Rate Limiting
In addition to Nginx rate limiting on auth endpoints, API key authentication has its own per-key rate limit enforced at the application layer by Better Auth.
Defaults
| Setting | Value |
|---|---|
| Time window | 60,000 ms (1 minute) |
| Max requests | 100 per window |
These defaults apply to all newly created API keys.
Configuration
The limits can be overridden via environment variables:
| Variable | Description | Default |
|---|---|---|
SPARKY_FITNESS_API_KEY_RATELIMIT_WINDOW_MS | Time window in milliseconds | 60000 |
SPARKY_FITNESS_API_KEY_RATELIMIT_MAX_REQUESTS | Max requests per window | 100 |
Behavior
When the rate limit is exceeded, the server returns:
- HTTP 429 with
{"error": "Rate limit exceeded."} - A
Retry-Afterheader indicating seconds until the window resets (at most 60s with default settings)
This rate limit is per API key, not per IP. Cookie-based browser sessions are not affected.
Sign-In Rate Limiting
Password sign-in, two-factor verification and email-OTP verification are rate-limited per client IP at the application layer, using Better Auth's customRules. This replaces Better Auth's built-in default for those paths.
Defaults
| Setting | Value |
|---|---|
| Time window | 60 seconds |
| Max attempts | 4 per window |
Applied to /api/auth/sign-in/email, /api/auth/two-factor/* and /api/auth/email-otp/verify-email. Each path keeps its own budget, so one MFA login (sign-in, then send-OTP, then verify-OTP) does not exhaust a single counter.
Why the default is 4 per minute, not a tighter burst
Better Auth's own default for /sign-in is 3 requests per 10 seconds. That is a tight burst, but it resets every 10 seconds — so it permits 18 failed logins per minute, indefinitely.
A failed sign-in answers 401, and intrusion-detection tooling reads sustained 401s as a brute force. CrowdSec's generic 401 scenario, for example, is a leaky bucket of capacity 5 that drains one event every 10 seconds. At 18/minute that bucket overflows in about twenty seconds, and the client IP is banned at the edge — commonly for hours. A legitimate user fumbling their password manager can reach that on their own.
Capping the sustained rate fixes it. The 5th attempt in a minute receives 429, which those rules do not count, so failures can never accumulate to the ban threshold. Compared to Better Auth's default this is one request looser on the initial burst (4 immediate attempts versus 3) but far stricter on sustained guessing (4/minute versus 18/minute) — and the burst difference is immaterial next to the sustained cap that actually prevents the ban.
Configuration
| Variable | Description | Default |
|---|---|---|
SPARKY_FITNESS_SIGN_IN_RATELIMIT_WINDOW | Time window in seconds | 60 |
SPARKY_FITNESS_SIGN_IN_RATELIMIT_MAX | Max attempts per window | 4 |
Behavior worth knowing
- Successful sign-ins count too. Better Auth counts every request, not just failures, so this is a ceiling on all sign-in traffic from one address. Instances behind a shared office IP or carrier-grade NAT should raise
..._MAX. The attempt is counted as the request arrives, in one atomic check-and-increment, so simultaneous attempts cannot slip past the limit together. - The window resets only after silence. Each counted request refreshes the timer, so a client that keeps retrying holds itself blocked until it pauses for a full window. This is why the window is deliberately short — it self-heals after a minute of quiet.
- A missing client IP now shares one bucket, it no longer fails open. If the IP cannot be determined (no
X-Forwarded-For), Better Auth logs a warning and falls back to a single shared per-path bucket rather than skipping the limit. So a reverse-proxy misconfiguration throttles everyone together against one budget instead of disabling rate limiting — the symptom is unexplained 429s under light load. Fix the proxy so the client address is forwarded;SPARKY_FITNESS_REAL_IP_HEADERandSPARKY_FITNESS_TRUSTED_PROXY_HOPScover the usual cases. (Before Better Auth 1.7 this case skipped rate limiting entirely.) - IPv6 is grouped by
/64by default. Better Auth'snormalizeIPmasks IPv6 addresses to a/64(its defaultipv6Subnet) before keying, so an entire/64shares one budget and a client rotating addresses within its prefix does not get a fresh budget. This is stricter, not looser; the trade-off is that a lockout applies to the whole/64. Override withadvanced.ipAddress.ipv6Subnetin the Better Auth config if you need a different mask. - The demo one-click login (
/api/auth/demo-login) is a separate Express route with its own limiter (10 per minute per IP) and is not affected by these variables.
Testing the Rate Limiting
Send repeated sign-in attempts with a deliberately wrong password and watch the status codes. Run them sequentially, not in parallel, so you are measuring the application limit rather than Nginx's burst:
for i in $(seq 1 6); do
curl -k -s -o /dev/null -w "%{http_code}\n" \
-X POST -H "Content-Type: application/json" \
-d '{"email":"test@example.com","password":"wrong-password"}' \
https://your-domain.com/api/auth/sign-in/email
doneWith the defaults you should see four 401s followed by 429s. If you see 429 on the fourth request instead of the fifth, the custom rules are not being applied and Better Auth's 3-per-10s default is in force.
WARNING
Do not run this against a production instance that sits behind fail2ban or CrowdSec unless your own IP is allow-listed. The 401s this generates are exactly what those tools ban for, and you can lock yourself out at the CDN edge for hours.
Note on 503 Errors During Testing
When load-testing through Nginx, 503 Service Unavailable errors might be observed instead of 429s. This indicates that Nginx is indeed applying the rate limit (as confirmed by Nginx error logs showing limiting requests), but it's also encountering issues connecting to or receiving timely responses from the backend server. While the rate limiting itself is functional, consistent 503s suggest an underlying issue with the backend's stability or readiness under load, which is outside the scope of the rate limiting implementation itself.
