Cookie-based auth is fine — until it isn't.
TL;DR — Cookie sessions are a perfectly reasonable default for a single web app and a small permission model. They fail the moment your permission set grows past a few dozen entries or your integrations want to share the same session. The failure looks like silent 400/431s, mysteriously dropped requests, and a help desk that doesn't know what to tell users. The fix is a server-side RBAC store with a small bearer token referencing it — not a bigger cookie.
The day the requests stopped going through
A power user at a Jobscope client started getting random 400 responses in the middle of the workday. Not all requests, just some. No stack trace, no application log line — the request was being rejected before it reached application code. The team's first instinct was a load balancer issue. Then a CORS issue. Then they restarted everything. Nothing fixed it.
Two days later we found it. The user belonged to an admin role that had grown to 240+ permissions. Each permission was being serialised into the session cookie. The cookie had crossed 8 KB. IIS was rejecting the request at the front door because the request header total exceeded its limit, before the application stack got a look. Other users had smaller permission sets and didn't hit the limit. The bug was effectively a permission-load bomb, exploding only on whoever had the worst luck of having the most rights.
That's the cookie-bloat failure mode in one sentence: your authorisation model grew, your authentication model didn't, and your web server is now silently dropping your most senior users' requests.
Why this fails specifically
Cookies travel on every request. There's a hard ceiling on how big they can be — somewhere between 4 KB and 16 KB depending on the browser, the proxy, and the web server. Most web servers default to 8 KB for the entire request header set; once the cookie alone is 6 KB, you've left ~2 KB for everything else. That gets blown by a long URL or a few standard headers.
The pathology has three parts:
- Permissions in the cookie at all. Once auth-z lives in the cookie, the cookie scales with your permission model, not with the user.
- Permissions denormalised at issue time. If you serialise the full permission list rather than a role reference, the cookie grows linearly with the role's grants.
- No validation on cookie size. Nothing on the way in or out tells you when the cookie is approaching the request-size ceiling. The first signal is a user complaint.
The reason this often goes undetected for years is that admin users — the ones with the largest permission set — are usually internal staff. Internal staff don't tend to file external bug reports. They restart the browser, clear cookies, escalate to IT, and gradually adapt around the broken parts of the app.
How to spot it in your platform in 30 minutes
Three checks:
Check 1 — Cookie size for your most-permissioned user
Log in as your highest-permissioned account. Open DevTools → Application → Cookies. Sum the size of cookies on your domain. If the total is over 4 KB, you're in the danger zone. Over 6 KB, you're already failing for some users somewhere.
Check 2 — IIS / Nginx error log for 400/431
# Nginx — look for "400 Bad Request" with empty request line
grep -E "400 .* -$" /var/log/nginx/access.log | head
# IIS — look in HTTPERR for "BadRequest" or
# "Request_Header_Or_Cookie_Too_Long"
findstr /i "BadRequest Header_Or_Cookie" %windir%\System32\LogFiles\HTTPERR\*.log
If those logs have entries from the last 90 days that you can correlate with specific users, the failure mode is already happening in production.
Check 3 — Your role table
Run a count of permissions per role. Any role with more than 50 permissions, you should assume is at risk. Any role over 100 is a fire.
The fix that worked
The instinct most teams have is "raise the IIS / Nginx header limit". Don't. You're treating a symptom — and you'll re-hit the ceiling as the next admin role grows. The structural fix is to take authorisation out of the cookie entirely.
The pattern I shipped on Jobscope:
- Cookie carries identity, not permissions. A signed token referencing the user id and a session id. Fixed size, ~80 bytes encoded.
- Server-side session store holds a reference to the user's role assignments — not the permissions themselves, just the role ids.
- Permissions resolve server-side, on request. A small middleware looks up the role's permission set from a cached table and exposes it to the controller via the request context.
- Permission cache is per-process, with a short TTL. A role's permission list almost never changes, so a 60-second in-process cache covers nearly all reads with no DB hit.
The cookie shrinks from 6+ KB to under 200 bytes. The permission lookup happens once per request, not once per session — which sounds slower but is actually faster, because the cookie no longer has to travel 6 KB on every request. Egress dropped, header parsing got cheaper, request latency improved on the order of single-digit milliseconds.
Code shape for the middleware
The actual implementation depends on your stack. For a .NET app, the relevant change is roughly:
public class PermissionContextMiddleware
{
private readonly RequestDelegate _next;
private readonly IPermissionResolver _resolver;
public async Task InvokeAsync(HttpContext ctx)
{
var userId = ctx.User?.GetUserId();
if (userId is not null)
{
// Permission set resolved server-side, cached
// for 60s per role. Role membership comes from
// session store, NOT from the cookie.
var perms = await _resolver.ForUserAsync(userId);
ctx.Items["Permissions"] = perms;
}
await _next(ctx);
}
}
For Spring Boot, the same pattern fits a HandlerInterceptor or a custom filter chained ahead of SecurityContextPersistenceFilter. For FastAPI, it's a dependency that runs ahead of every protected route.
Why JWT alone isn't the whole answer
The popular alternative pitch is "use a JWT". JWT is fine — better than cookies for cross-domain integration, easier to validate without a session lookup. But a JWT that carries the full permission list has the exact same failure mode as a fat cookie, just shifted to the Authorization header. I've seen JWTs that exceeded 12 KB because someone serialised every permission into the claims set.
The rule, regardless of token shape:
Token → carries identity (user id, session id, expiry, signing)
DB → carries authorisation (role assignments + permission grants)
Cache → carries the resolved permission set, in-process, short TTL
That separation is what scales. The token never grows with the permission model, because the permission model lives where it should — in your authorisation store.
What this fix doesn't solve
This fix solves the cookie-bloat failure mode and gives you a clean place to add row-level access controls on top. It doesn't solve other auth problems you might have:
- If you're sharing sessions with integration partners, that's a separate problem. The right pattern is OAuth ROPC or client_credentials for service-to-service. Covered in detail here.
- If you don't have audit logging on permission changes, this fix gives you a logical place to add it (the permission resolver) but doesn't add it for you.
- If your roles are growing by 50 permissions a quarter, you have a permission-model problem, not an auth problem. Roles aren't supposed to grow forever — they're supposed to be replaced when the business model shifts.
The 2-week plan if you're hitting this today
- Today: turn on the diagnostics above. Confirm the failure mode. Identify the worst-affected roles.
- Day 2–3: ship the server-side permission resolver behind a feature flag. Run it in shadow mode — resolve permissions, compare to cookie-derived permissions, log mismatches.
- Day 4–5: when shadow mode shows zero mismatches for two consecutive weeks, flip the flag for the worst-affected user cohort first. Then expand.
- Day 6 onward: remove the permission payload from the cookie entirely; reduce cookie size; raise the cookie TTL because it's now small enough that frequent re-issuance isn't useful.
That sequence has shipped on a real engagement without an outage. The Jobscope rollout took roughly 7-10 Days end-to-end, with the shadow-mode period giving the team enough confidence to flip the switch without holding a war room open.
Suspect this is happening on your platform but can't find it in the logs?