Most backend services hit the same wall eventually: the database keeps answering the same expensive questions, and latency climbs with traffic even though the data barely changes. Putting Redis in front of the database is the standard fix. It also creates a second copy of your data, and nearly every caching bug I have debugged came down to that copy drifting from the source of truth or vanishing at the wrong moment.
In this post I cover the patterns I use with Redis and Node.js: cache-aside in TypeScript with ioredis and Postgres, invalidation, write-through, write-behind, read-through, and the production concerns of stampedes, key design, eviction, and hit ratio. By the end you should be able to pick a strategy for a piece of data and know which failure modes come with it.
Why cache, and what belongs in one
A cache trades freshness for speed: you accept slightly old data in exchange for skipping a slow query or a call to a rate-limited API. That trade pays off for data with three properties:
- Read-heavy. Reads far outnumber writes. A profile read on every page view and edited once a month is ideal; a shopping cart that changes on every click is not.
- Expensive to produce. Aggregations, multi-table joins, and third-party API responses qualify. A fast indexed lookup usually does not: you add a network hop and a consistency problem to save very little.
- Tolerant of staleness. Someone has to be able to say "a minute out of date is fine." If the honest answer is "never", as with an account balance, read from the database.
When unsure, I measure first: in Postgres, pg_stat_statements shows which queries dominate total execution time.
Cache-aside in TypeScript
Cache-aside, also called lazy loading, is the pattern I reach for first. The application owns the logic: check the cache, and on a miss, load from the database and store the result with a TTL. Redis never talks to Postgres directly.
The database side is a plain repository on top of pg:
import type { Pool } from 'pg';
export interface User {
id: string;
email: string;
displayName: string;
updatedAt: string; // ISO 8601 string, so it survives a JSON round trip
}
interface UserRow {
id: string;
email: string;
display_name: string;
updated_at: Date;
}
const toUser = (row: UserRow): User => ({
id: row.id,
email: row.email,
displayName: row.display_name,
updatedAt: row.updated_at.toISOString(),
});
export class UserRepository {
constructor(private readonly pool: Pool) {}
async findById(id: string): Promise<User | null> {
const { rows } = await this.pool.query<UserRow>(
'SELECT id, email, display_name, updated_at FROM users WHERE id = $1',
[id],
);
return rows[0] ? toUser(rows[0]) : null;
}
async updateDisplayName(id: string, displayName: string): Promise<User | null> {
const { rows } = await this.pool.query<UserRow>(
`UPDATE users SET display_name = $2, updated_at = now()
WHERE id = $1
RETURNING id, email, display_name, updated_at`,
[id, displayName],
);
return rows[0] ? toUser(rows[0]) : null;
}
}The service implements get, miss, load, and set:
import type Redis from 'ioredis';
import { withJitter } from '../cache/ttl';
import type { User, UserRepository } from './user-repository';
const USER_TTL_SECONDS = 600;
export const userKey = (id: string) => `app:user:v1:${id}`;
export class UserService {
constructor(
private readonly redis: Redis,
private readonly users: UserRepository,
) {}
async getUser(id: string): Promise<User | null> {
const key = userKey(id);
// 1. Get. A Redis error or a corrupt entry is treated as a miss.
try {
const cached = await this.redis.get(key);
if (cached !== null) return JSON.parse(cached) as User | null;
} catch (err) {
console.warn('cache read failed', { key, err });
}
// 2. Miss: load from the source of truth.
const user = await this.users.findById(id);
// 3. Set with a TTL, so even an entry that is never invalidated ages out.
if (user !== null) {
try {
await this.redis.set(key, JSON.stringify(user), 'EX', withJitter(USER_TTL_SECONDS));
} catch (err) {
console.warn('cache write failed', { key, err });
}
}
return user;
}
}Both Redis calls have their own try block, so an outage turns into cache misses instead of failed requests. That only works if the client fails fast. By default, ioredis queues commands while it reconnects, which turns an outage into hanging requests, so I configure it explicitly:
import Redis from 'ioredis';
import { Pool } from 'pg';
import { UserRepository } from './users/user-repository';
import { UserService } from './users/user-service';
export const redis = new Redis(process.env.REDIS_URL ?? 'redis://localhost:6379', {
enableOfflineQueue: false, // reject commands immediately while disconnected
maxRetriesPerRequest: 1, // fail pending commands after one reconnect attempt
commandTimeout: 200, // ms; a slow cache is treated like a missing one
});
redis.on('error', (err) => console.error('redis error:', err.message));
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
export const userService = new UserService(redis, new UserRepository(pool));TTLs and jitter
Every key gets a TTL. It bounds staleness when everything else goes wrong: a missed invalidation, a buggy write path, a race. I pick it by asking how long a stale value could be shown before it causes real harm, then go shorter. For entity caches that usually means minutes; reference data can live for hours.
The less obvious problem is synchronized expiry. If a warm-up job writes 50,000 keys within a few seconds, they all expire together one TTL later, and the database absorbs the whole reload at once. Random jitter spreads the expirations out:
// withJitter(600) returns a value between 540 and 660 (plus or minus 10%).
export function withJitter(baseSeconds: number, ratio = 0.1): number {
const spread = baseSeconds * ratio;
const ttl = baseSeconds - spread + Math.random() * 2 * spread;
return Math.max(1, Math.round(ttl)); // EX must be a positive integer
}Ten percent is my default, and it costs nothing.
Invalidating on writes
TTLs bound staleness, but a user who changes their display name expects to see it on the next page load. A write can update the cached value (write-through, covered below) or delete the key so the next read repopulates it. With cache-aside, I delete:
// Inside UserService
async renameUser(id: string, displayName: string): Promise<User | null> {
// Commit to Postgres first. Touch the cache only after the write is durable.
const user = await this.users.updateDisplayName(id, displayName);
try {
await this.redis.del(userKey(id));
} catch (err) {
// The TTL is now the only bound on staleness for this user. Log it loudly.
console.error('cache invalidation failed', { id, err });
}
return user;
}Deleting is usually safer. With updates, two concurrent writers each commit and then write their own version to Redis, and nothing guarantees those writes land in commit order, so the cache can keep the older value until the TTL expires. Two deletes in any order leave the same result: no key. Deleting also means the write path never rebuilds the cached shape, which is often assembled from several tables.
Order matters too: delete only after the transaction commits. Delete earlier and a concurrent reader can repopulate the key with the old row in the gap.
The race that delete-on-write leaves open
Even with the right order, there is a window:
Reader A Writer B
t1 GET app:user:v1:42 (miss)
t2 SELECT ... FROM users (old row)
t3 UPDATE users ... COMMIT
t4 DEL app:user:v1:42
t5 SET app:user:v1:42 (old row)
The old row now sits in the cache until its TTL expires.Reader A has to query before B commits and write after B deletes. That is rare, but under real load rare happens daily. Three mitigations, cheapest first:
- Short TTLs. The stale value dies on its own. Often this is enough.
- Delayed double delete. Delete after the commit, then again after a delay longer than a typical read-and-populate cycle. It narrows the window rather than closing it.
- Versioned keys. Keep a version counter per entity and build the data key from it. Bumping the version replaces the delete.
import type Redis from 'ioredis';
// Delayed double delete: the second DEL removes a stale value written by a reader
// that loaded the old row before the commit and cached it after the first DEL.
export async function deleteTwice(redis: Redis, key: string, delayMs = 1_000): Promise<void> {
await redis.del(key);
setTimeout(() => {
redis.del(key).catch((err) => console.warn('delayed delete failed', { key, err }));
}, delayMs);
}
// Versioned keys: readers build the data key from a per-user version counter.
const versionKey = (id: string) => `app:user:v1:${id}:version`;
export async function versionedUserKey(redis: Redis, id: string): Promise<string> {
const version = (await redis.get(versionKey(id))) ?? '0';
return `app:user:v1:${id}@${version}`;
}
// Writers bump the version after the commit. The counter outlives the data keys.
export async function bumpUserVersion(redis: Redis, id: string): Promise<void> {
await redis.multi().incr(versionKey(id)).expire(versionKey(id), 86_400).exec();
}With versioned keys, a slow reader writes its stale value under a version nobody reads anymore. The cost is a second round trip per read, so I save this for data where brief staleness after a write really hurts.
Write-through, write-behind, and read-through
Cache-aside fills the cache on reads. The other strategies move that work elsewhere.
Write-through
Write-through updates the cache on every write: commit to Postgres, then set the fresh value in Redis. In the textbook version the cache writes to the database itself; with Redis, your application does both.
// Write-through variant of renameUser: set the fresh value instead of deleting.
async renameUser(id: string, displayName: string): Promise<User | null> {
const user = await this.users.updateDisplayName(id, displayName);
if (user !== null) {
await this.redis
.set(userKey(id), JSON.stringify(user), 'EX', withJitter(USER_TTL_SECONDS))
.catch((err) => console.error('cache write-through failed', { id, err }));
}
return user;
}The first read after a write is a hit, which suits data read right after it changes, like a settings page. The costs are slower writes, memory spent on values nobody may read, and the out-of-order problem above, so I keep these TTLs short.
Write-behind
Write-behind, or write-back, reverses the order: the request writes only to Redis, and a background worker flushes changes to Postgres later, usually in batches. Writes run at Redis speed, and a burst of increments to one row collapses into a single UPDATE.
The price is durability: until the flush, Redis holds the only copy. With RDB snapshots alone, a crash loses everything since the last snapshot; with AOF and appendfsync everysec, you can still lose about a second of writes. Replication is asynchronous, so a failover can drop acknowledged writes, and eviction can remove pending writes outright.
Read-through
Read-through moves loading into the cache layer: callers ask the cache for a key, and the cache loads from the database on a miss. Redis knows nothing about Postgres, so in practice this is a wrapper that owns the loader, called as something like cache.get(key, loader). The data flow matches cache-aside; the benefit is that TTLs, serialization, error handling, and stampede protection live in one place. Once a codebase has more than a few cached entities, I wrap cache-aside this way.
Choosing between them
| Strategy | Who fills the cache | First read after a write | Main risk | Good fit |
|---|---|---|---|---|
| Cache-aside | App, on a read miss | Miss, then fresh | Stale value from the read/write race | Default for read-heavy data |
| Read-through | Cache layer, on a miss | Miss, then fresh | Same as cache-aside | Many entities sharing one policy |
| Write-through | App, on every write | Hit | Slower writes, cold data in memory | Data read right after it changes |
| Write-behind | App, on every write (DB later) | Hit | Losing writes before the flush | Counters and loss-tolerant writes |
Preventing cache stampedes
A stampede, or thundering herd, happens when a hot key expires and hundreds of concurrent requests miss at once. They all run the same expensive query, the database slows down, and the cache that protected it becomes the reason it falls over. I use three fixes, often together.
A lock with SET NX PX
Only one caller rebuilds the value; the rest wait briefly and re-check the cache. The lock is a key written with NX, so only one caller can create it, and PX, so it expires if the holder crashes. Its value is a random token, and release goes through a Lua script so a process only deletes a lock it still owns:
import { randomUUID } from 'node:crypto';
import { setTimeout as sleep } from 'node:timers/promises';
import type Redis from 'ioredis';
// Delete the lock only if it still holds our token. A separate GET and DEL could
// delete a lock that expired and was acquired by another process in between.
const RELEASE_LOCK = `
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("DEL", KEYS[1])
end
return 0`;
export async function getWithLock<T>(
redis: Redis,
key: string,
ttlSeconds: number,
load: () => Promise<T>,
): Promise<T> {
const lockKey = `${key}:lock`;
const token = randomUUID();
let ownsLock = false;
try {
for (let attempt = 0; attempt < 20 && !ownsLock; attempt++) {
const cached = await redis.get(key);
if (cached !== null) return JSON.parse(cached) as T;
ownsLock = (await redis.set(lockKey, token, 'PX', 5_000, 'NX')) === 'OK';
if (!ownsLock) await sleep(25 + Math.random() * 25); // someone else is loading
}
} catch (err) {
console.warn('cache unavailable, loading directly', { key, err });
return load();
}
// We hold the lock, or we waited long enough and load anyway.
try {
const value = await load();
await redis.set(key, JSON.stringify(value), 'EX', ttlSeconds).catch(() => undefined);
return value;
} finally {
if (ownsLock) await redis.eval(RELEASE_LOCK, 1, lockKey, token).catch(() => undefined);
}
}Pick a lock TTL longer than the slowest expected load. If a holder stalls past it anyway, two processes rebuild at once: a duplicate query, nothing worse. That makes it an efficiency optimization, not a correctness guarantee.
Coalescing requests in-process
The lock costs round trips and makes callers wait. Before a request even gets that far, I deduplicate inside the process: if a load for a key is already in flight, later callers await the same promise.
const inFlight = new Map<string, Promise<unknown>>();
export function singleflight<T>(key: string, fn: () => Promise<T>): Promise<T> {
const existing = inFlight.get(key);
if (existing) return existing as Promise<T>;
const promise = fn().finally(() => inFlight.delete(key));
inFlight.set(key, promise);
return promise;
}
// 500 concurrent requests for user 42 on this instance share a single load:
// const user = await singleflight(`user:${id}`, () => userService.getUser(id));This is the idea behind Go's singleflight package. It does not coordinate across instances, so 20 pods can still produce 20 loads, but that beats one per request and needs no network calls. Callers share one object, so treat it as read-only.
Probabilistic early expiration
Locks and coalescing react to a miss; probabilistic early expiration avoids it. As a key nears expiry, each reader has a small but growing chance of refreshing it early while everyone else keeps getting hits. This version follows XFetch, from the paper "Optimal Probabilistic Cache Stampede Prevention" by Vattani, Chierichetti, and Lowenstein, which scales the head start by how long the value took to compute:
import type Redis from 'ioredis';
interface Entry<T> {
value: T;
computeMs: number; // how long the last recomputation took
expiresAt: number; // epoch milliseconds
}
const BETA = 1; // above 1 refreshes earlier, below 1 later
export async function xfetch<T>(
redis: Redis,
key: string,
ttlSeconds: number,
compute: () => Promise<T>,
): Promise<T> {
const raw = await redis.get(key).catch(() => null); // Redis errors count as misses
if (raw !== null) {
const entry = JSON.parse(raw) as Entry<T>;
// -log(random) is usually small but occasionally large, so a few readers act
// as if it were later than it is and refresh before the key actually expires.
const headStart = entry.computeMs * BETA * -Math.log(Math.random());
if (Date.now() + headStart < entry.expiresAt) return entry.value;
}
const started = Date.now();
const value = await compute();
const entry: Entry<T> = {
value,
computeMs: Date.now() - started,
expiresAt: Date.now() + ttlSeconds * 1000,
};
await redis.set(key, JSON.stringify(entry), 'EX', ttlSeconds).catch(() => undefined);
return value;
}It needs no coordination, but it only helps while the key exists. A cold or just-deleted key still needs the lock or coalescing, which is why I layer them.
Designing keys and values
Negative caching
As written, getUser never caches a miss, so a scraper hammering nonexistent IDs sends every request to Postgres. Negative caching stores the absence too, with a shorter TTL:
const NOT_FOUND_TTL_SECONDS = 30;
// Steps 2 and 3 of getUser, now caching misses as well:
const user = await this.users.findById(id);
const ttl = user === null ? NOT_FOUND_TTL_SECONDS : withJitter(USER_TTL_SECONDS);
try {
// JSON.stringify(null) is the string "null", which parses back to null on a hit.
await this.redis.set(key, JSON.stringify(user), 'EX', ttl);
} catch (err) {
console.warn('cache write failed', { key, err });
}
return user;The read path needs no changes. Keep the negative TTL short or delete the key on create; otherwise a user who signs up right after you cached their absence stays invisible. Validate IDs first, too, so malformed input never becomes a key.
Key design
- Namespace everything, from general to specific:
app:user:v1:42. The prefix separates services on a shared instance and lets you inspect a group withSCANand a pattern likeapp:user:*. Never runKEYSin production; it blocks the server while it walks the whole keyspace. - Version the format. When the cached shape changes, bump
v1tov2. During a rolling deploy, old and new code then use different keys instead of parsing each other's payloads. - Keep the key space bounded. Keys built from raw input, like search strings, grow without limit and evict useful entries. Normalize inputs, and remember that hashing a long input shortens the key but does not reduce how many keys exist.
Serialization
I default to JSON: it is readable in redis-cli and every language decodes it. Its sharp edge is types. A Date becomes a string and stays one after JSON.parse, a BigInt throws, and a Map becomes an empty object. That is why the repository converts updated_at to an ISO string: a hit must return exactly what a miss returns, or you get bugs that only appear when the cache is warm.
When payloads grow or serialization shows up in CPU profiles, I switch to MessagePack and read with getBuffer so ioredis returns raw bytes. That trades debuggability for size and speed, so measure first.
Running Redis as a cache in production
Memory limits and eviction policies
Two settings decide how a cache forgets:
maxmemory 4gb
maxmemory-policy allkeys-lrumaxmemory caps the dataset; on 64-bit systems the default is no limit. Set it below the machine's RAM to leave room for fragmentation, buffers, and persistence forks. maxmemory-policy decides what happens at the limit. The default, noeviction, is wrong for a cache: once memory is full, writes fail with OOM errors and the cache silently stops accepting new data. The alternatives that matter:
allkeys-lruevicts the least recently used keys. It is my default for a dedicated cache.volatile-lruevicts only keys with a TTL, for instances that also hold keys that must survive. If no key has a TTL, writes fail as withnoeviction.allkeys-lfuevicts the least frequently used keys. It protects a stable hot set from one-off sweeps, like a crawler visiting every product page once, which can flush popular keys under LRU.
Redis approximates LRU and LFU by sampling, which works well in practice; the eviction docs cover the other policies. I run one instance per job: an allkeys cache, and a separate noeviction instance with persistence for anything that must not disappear.
Measuring hit ratio
Redis counts successful and failed key lookups in INFO:
redis-cli INFO stats | grep -E '^(expired_keys|evicted_keys|keyspace_hits|keyspace_misses):'expired_keys:1733120
evicted_keys:0
keyspace_hits:48213377
keyspace_misses:5120944Hit ratio is keyspace_hits / (keyspace_hits + keyspace_misses), about 90% in this sample. The counters are cumulative since startup or the last CONFIG RESETSTAT, so compute the ratio from the difference between two samples. They also cover every key on the instance, so I emit per-prefix hit and miss counters from the application too. There is no universally good number; watch each cache's trend against its own baseline.
What to alert on
- A sustained drop in hit ratio. Often a deploy changed key names or TTLs, or the working set outgrew memory.
- Steadily rising
evicted_keys. The working set no longer fits. - Memory near
maxmemoryon anoevictioninstance. Writes are about to fail. - Cache errors and timeouts in the application. The code fails open, so Redis can be down while every request succeeds. The first symptom is usually database load, so watch cache errors and database query rate together.
Key takeaways
- Cache data that is read-heavy, expensive to produce, and tolerant of staleness.
- Default to cache-aside with a jittered TTL on every key, and treat Redis errors as misses.
- Invalidate by deleting after the commit; bound the remaining race with short TTLs, a delayed second delete, or versioned keys.
- Use write-through when reads closely follow writes, and write-behind only for data you can afford to lose.
- Protect hot keys with in-process coalescing, a
SET NX PXlock, or probabilistic early expiration. - Configure
maxmemoryand an eviction policy, and track hit ratio per cache as a trend.
None of these patterns is complicated on its own. What makes caching hard is that it fails quietly: a stale value looks like valid data, and an unreachable Redis looks like a slow database. Start with cache-aside, make every failure path fall back to Postgres, and measure from day one.