a map of backend systems
◀ Back to the map

Redis line

Caching patterns done right

Caching is the reason most people reach for Redis — and the easiest thing to get subtly, expensively wrong.


Redis deep-dive · Part 5 of 10. Previous: Atomicity in depth. Next: Messaging: Pub/Sub vs Streams.

Caching is the front door to Redis for almost everyone: the query is slow, so you stash the result and serve it from memory next time. It works on the first try, which is exactly why the failure modes stay hidden until traffic finds them. The hard part isn’t putting a value in the cache — it’s deciding who writes it, when it becomes wrong, and what happens the instant it disappears. This chapter is about doing all three deliberately, on a familiar stack: a backend API with Postgres behind it and Redis alongside.

We’ll build up from the default pattern, sharpen the one write-path rule that trips everyone up, then face the failure mode that turns a cache from a shield into a loaded gun pointed at your database.

The strategies: who writes the cache, and when

There are only a few ways to arrange a cache and a database, and they differ on one axis: who is responsible for keeping the cache populated, and at what moment.

Cache-aside (lazy loading) is the default, and it’s what “using Redis as a cache” almost always means. Your application manages the cache directly; Redis knows nothing about Postgres. The two paths look like this:

READ:   val = GET key
        if hit  -> return val
        if miss -> val = SELECT ... from Postgres
                   SET key val EX ttl
                   return val

WRITE:  UPDATE ... in Postgres
        DEL key            # invalidate, don't rewrite

Three properties fall out of this shape, and they’re the reason it dominates. Only requested data is ever cached — a key exists in Redis only because someone actually asked for it, so you never waste memory on rows nobody reads. A miss costs exactly one extra round-trip — check Redis, miss, hit Postgres, backfill. And crucially, if Redis is down you degrade, you don’t break — every miss falls through to the database. Slower, but still correct. That graceful fallback is what makes cache-aside the safe default.

Write-through inverts the responsibility: the cache sits in front of the database, and every write goes through it. The app writes to the cache, and the cache (or the app, synchronously) writes to the database in the same step. The data in the cache is always fresh, because a write can’t complete without updating it. The costs are the mirror image of the benefits: every write pays the cache-write latency, and you populate entries for data that may never be read again — you’re caching on write, not on demand.

Write-behind (write-back) is the asynchronous cousin: write to the cache, acknowledge immediately, and flush to the database later in the background. It’s the fastest of the three on the write path, and the most dangerous — if Redis dies before the flush, those writes are simply gone. You’re trading durability for latency, which is only acceptable when the data can afford to be lost.

PatternWho writes the cacheFreshnessMain cost
Cache-asideApp, lazily on read missBounded by TTL + invalidationExtra round-trip on a miss
Write-throughCache, synchronously on every writeAlways freshEvery write pays cache cost; caches unread data
Write-behindCache, async flush to DBFresh in cache, lagging in DBCan lose writes if Redis dies pre-flush

Write-through and write-behind are rarer, and usually appear only when a caching library or a specialized layer implements them for you. For hand-rolled application caching, reach for cache-aside first and don’t look further unless you have a specific reason.

Invalidate, don’t update (why DEL beats SET on writes)

Look again at the cache-aside write path: on a database write, it does DEL key, not SET key newValue. That choice looks pedantic — why throw away a value you already have in hand? — but it’s the single most important rule in this chapter, and it’s a direct descendant of the lost-update race from Part 4.

Picture two requests updating the same row, each trying to keep the cache fresh with SET:

Writer A: UPDATE row = 10           Writer B: UPDATE row = 20
Writer A: reads row -> 10           Writer B: reads row -> 20
                                    Writer B: SET key 20
Writer A: SET key 10        <-- lands last, but is older

Both writers read from the database and then set the cache, but nothing forces those two steps to happen in the same order for both. If A’s SET lands after B’s — even though A’s value is older — the cache now holds 10 while the database holds 20. And it stays wrong until the TTL expires. A stale write has won the race and poisoned the cache indefinitely.

DEL sidesteps the whole problem. It doesn’t matter which writer’s DEL lands last, because they all do the same thing: remove the key. The next read finds nothing, re-fetches from the authoritative database, and repopulates with the current value. The worst case is a redundant cache miss — never a permanently wrong value.

TTL strategy

Even with disciplined invalidation, always put a TTL on cached data. The two mechanisms are not redundant — they cover different failures. Invalidation handles the writes you know about; the TTL handles the ones you miss. Sooner or later a code path will update a row and forget to DEL the key, or an invalidation event will get dropped, or a cache entry will drift out of sync for a reason no one predicted. The TTL is the backstop that guarantees every stale entry eventually self-heals, whether or not your invalidation logic was perfect.

Choosing the number is a straight trade-off. A short TTL means fresher data and more frequent database hits as entries expire and get rebuilt. A long TTL means less database load but more tolerance for staleness. There’s no universal right answer — it depends on how wrong the data is allowed to be and how expensive it is to recompute.

There’s one more rule that isn’t optional, and it becomes obvious the moment we look at how caches fail under load: jitter your TTLs. Never expire many keys at the same instant. Why that matters is the next section.

The failure mode: cache stampede (thundering herd)

Here’s the scenario that takes down databases. A hot key expires — or, worse, ten thousand keys that were all written with the same TTL expire in the same second. Every request that would have been a cache hit is now a miss, and all of them turn to the database at once to recompute the identical value. Redis was absorbing that read traffic; the instant the key vanishes, that protection evaporates for every concurrent reader simultaneously, and the full weight of production traffic lands on Postgres in a single spike. The database falls over, which makes recomputation slower, which makes the pileup worse. This is the cache stampede, or thundering herd.

The insidious part is that the cache was doing its job right up until the microsecond it expired. Nothing was slow, no alarm fired — and then a shared expiry converted thousands of cheap cache hits into thousands of expensive database queries at the same timestamp.

There are three standard mitigations, and they stack:

MitigationWhat it doesWhen to use
Jittered TTLAdd randomness to each TTL so expirations spread out instead of firing togetherAlways — the cheap default
Recompute lockOn a miss, one requester takes a lock and rebuilds; the rest wait or serve staleHot keys where one DB hit is fine but a thousand isn’t
Early recomputationRefresh the value before it expires, so it’s never actually absentVery expensive values read constantly

1. Jittered / randomized TTL. Instead of EX 3600, use 3600 + rand(0..300). Now the expirations that would have bunched up in one second are smeared across a five-minute window, and no single instant sees the whole herd. This is the cheapest fix and the most important one — do it by default on every cached key, and most stampedes never form in the first place.

2. Recompute lock (“one recomputes”). On a miss, the first requester atomically grabs a short lock — SET lock:key 1 NX EX 5 — and only the holder recomputes the value and repopulates the cache. Everyone else either waits a beat and re-reads, or serves the last-known value. Ten thousand concurrent misses collapse into a single database hit. The NX flag (“set only if it doesn’t already exist”) is the atomic primitive from Part 4’s family that makes “exactly one winner” possible without an application-side lock.

3. Early recomputation. Refresh the value before its TTL runs out — probabilistically as expiry approaches, or from a background job on a schedule — so the key is never actually missing under load. The stampede can’t happen because there’s never a moment with nothing to serve.

Worked example: the homepage feed

Make it concrete. Your API serves a “homepage feed” assembled from an expensive Postgres query — joins, ranking, the works. The naive cache is one line:

SET homepage:feed <json> EX 60

It works beautifully until traffic arrives. Every 60 seconds, on the dot, the key expires and every in-flight request stampedes Postgres to rebuild the same feed at the same instant. The cache that was supposed to protect the database has scheduled a synchronized attack on it, once a minute, forever.

Here’s the same feed done right, combining all three tools:

# 1. Jittered TTL so feeds across the fleet don't expire in lockstep
SET homepage:feed <json> EX 75        # conceptually 60 + rand(0..15)

# 2. On a miss, one request wins the recompute lock and rebuilds
SET homepage:feed:lock 1 NX EX 10     # only the NX winner recomputes
#    winner: run the query, SET homepage:feed with a fresh jittered TTL
#    losers: serve last-known feed, or wait ~50ms and re-read the key

# 3. On a content change (a new post is published), invalidate
DEL homepage:feed                     # next read rebuilds from scratch

Note point three: when the content genuinely changes, you DEL the whole key and let the next read rebuild it. You do not try to surgically splice the new post into the cached JSON — that’s the SET-on-write race from earlier, wearing a different hat. Drop it, rebuild it, move on.

Everything in this chapter reduces to a few disciplined defaults: cache-aside as the pattern, DEL not SET on writes, a TTL on every key, jitter on every TTL, and a recompute lock on the hot ones. None of it is exotic. It’s just the difference between a cache that quietly protects your database and one that periodically fires the whole herd straight through it.

Caching correctly isn’t about keeping the cache perfectly in sync with the database — it’s about bounding how stale it can get while never letting it vanish all at once. Invalidate, don’t update; TTL everything; jitter always.