Splitting a system into services turns function calls into network calls, and the network changes how things fail. A slow dependency makes its callers slow, retries pile onto a struggling service, and an order gets saved while its event quietly disappears.

Most of these failures trace back to one decision, made over and over: when a service needs something from another, does it wait for an answer or send a message and move on? Each choice fails differently and needs its own defenses.

This article covers both: timeouts, retries, circuit breakers, and bulkheads for synchronous calls; idempotent consumers, the transactional outbox, sagas, and dead-letter queues for messaging; plus tracing across both. Examples use TypeScript on Node 20+, PostgreSQL with pg, and a broker-agnostic publish() function, and I finish with a decision guide and my defaults.

Who needs the answer, and when

Before choosing REST, gRPC, or a broker, I ask one question about each interaction: who needs the answer, and when?

If the caller cannot continue without the result, the interaction is synchronous by nature. Checkout needs a price and a payment authorization before it can respond, and a queue would only rebuild request/response with more moving parts.

If the caller only needs to know the work will happen, it is asynchronous. The order service shouldn't wait for the confirmation email; it should record the order and announce it.

Some interactions are both: a report export needs an immediate 202 Accepted, not an immediate file. Let the question decide, not the tools you already run.

Synchronous calls without cascading failures

REST and gRPC calls are the simplest thing that works: the code reads top to bottom and errors come straight back to the caller.

The cost is temporal coupling. A request succeeds only if every service in its path is healthy at that moment, so availability multiplies: four dependencies at 99.9% each give roughly 99.6%, assuming independent failures. Latency compounds too, and a parallel fan-out is only as fast as its slowest call.

The worst case is a cascade: a dependency slows down, callers wait, their in-flight requests pile up holding sockets and memory, and they become slow too. Four patterns stop that chain reaction.

Put a timeout on every call

Node's built-in fetch will wait minutes for a response that should take milliseconds. AbortSignal.timeout() fixes that in one line:

src/clients/inventory.tsTypeScript
const INVENTORY_URL = process.env.INVENTORY_URL ?? "http://inventory:3000";
 
export class HttpError extends Error {
  readonly status: number;
 
  constructor(status: number) {
    super(`Upstream responded with HTTP ${status}`);
    this.name = "HttpError";
    this.status = status;
  }
}
 
export async function getStock(sku: string): Promise<number> {
  const res = await fetch(`${INVENTORY_URL}/stock/${encodeURIComponent(sku)}`, {
    signal: AbortSignal.timeout(800), // give up after 800 ms instead of minutes
  });
  if (!res.ok) throw new HttpError(res.status);
 
  const body = (await res.json()) as { available: number };
  return body.available;
}

On timeout, fetch rejects with an error named TimeoutError. I set timeouts slightly above the dependency's observed p99 latency, and budgets must shrink with depth: a handler with two seconds to answer cannot give its dependency five. If the incoming request has its own deadline, combine both with AbortSignal.any() (Node 20.3+).

Retry only what is safe to retry

Retries turn transient failures into successes, but only for operations that are safe to repeat. GET, PUT, and DELETE are idempotent by HTTP semantics, if your API honors them; retry a POST only when the server deduplicates on an idempotency key, like Stripe's Idempotency-Key header.

Retries also add load exactly when a dependency is struggling. Exponential backoff spreads attempts out, and full jitter randomizes each delay so clients don't retry in synchronized waves:

src/lib/retry.tsTypeScript
import { setTimeout as sleep } from "node:timers/promises";
 
export interface RetryOptions {
  maxAttempts: number; // total attempts, including the first
  baseDelayMs: number;
  maxDelayMs: number;
  isRetryable: (err: unknown) => boolean;
}
 
export async function withRetry<T>(
  operation: () => Promise<T>,
  { maxAttempts, baseDelayMs, maxDelayMs, isRetryable }: RetryOptions,
): Promise<T> {
  for (let attempt = 1; ; attempt++) {
    try {
      return await operation();
    } catch (err) {
      if (attempt >= maxAttempts || !isRetryable(err)) throw err;
      // Full jitter: a random delay between 0 and the exponential ceiling
      const ceiling = Math.min(maxDelayMs, baseDelayMs * 2 ** (attempt - 1));
      await sleep(Math.random() * ceiling);
    }
  }
}
src/checkout/stock.tsTypeScript
import { getStock, HttpError } from "../clients/inventory";
import { withRetry } from "../lib/retry";
 
// GET is idempotent, so retrying it is safe
export function getStockWithRetry(sku: string): Promise<number> {
  return withRetry(() => getStock(sku), {
    maxAttempts: 3,
    baseDelayMs: 100,
    maxDelayMs: 1_000,
    isRetryable: (err) =>
      err instanceof TypeError || // network failure: fetch rejects with a TypeError
      (err instanceof Error && err.name === "TimeoutError") ||
      (err instanceof HttpError && [429, 502, 503, 504].includes(err.status)),
  });
}

Retry at one layer only: if three layers each make three attempts, the bottom service can receive 27 requests for one user action.

Circuit breakers

A circuit breaker stops calls that are almost certain to fail, protecting your service and the struggling dependency. It has three states:

  • Closed: calls pass through and failures are counted.
  • Open: after too many failures, calls fail immediately without touching the network.
  • Half-open: after a cooldown, one trial call goes through. Success closes the circuit; failure opens it again.
src/lib/circuit-breaker.tsTypeScript
type State = "closed" | "open" | "half-open";
 
export class CircuitOpenError extends Error {
  constructor() {
    super("Circuit open: failing fast");
    this.name = "CircuitOpenError";
  }
}
 
export class CircuitBreaker {
  private state: State = "closed";
  private failures = 0;
  private openedAt = 0;
  private readonly failureThreshold: number;
  private readonly resetTimeoutMs: number;
 
  constructor(options: { failureThreshold: number; resetTimeoutMs: number }) {
    this.failureThreshold = options.failureThreshold;
    this.resetTimeoutMs = options.resetTimeoutMs;
  }
 
  async call<T>(fn: () => Promise<T>): Promise<T> {
    if (this.state === "open") {
      if (Date.now() - this.openedAt < this.resetTimeoutMs) throw new CircuitOpenError();
      this.state = "half-open"; // cooldown is over: this call is the trial
    } else if (this.state === "half-open") {
      throw new CircuitOpenError(); // a trial call is already in flight
    }
 
    try {
      const result = await fn();
      this.state = "closed";
      this.failures = 0;
      return result;
    } catch (err) {
      this.failures++;
      if (this.state === "half-open" || this.failures >= this.failureThreshold) {
        this.state = "open";
        this.openedAt = Date.now();
      }
      throw err;
    }
  }
}

I put the breaker inside the retry (withRetry wraps breaker.call), so every attempt counts and a CircuitOpenError, which isn't retryable, ends the loop at once. This version counts consecutive failures for clarity; in production I use opossum, which adds rolling windows, fallbacks, and metrics events. Count only failures that signal an unhealthy dependency, like timeouts and 5xx responses, never a 404.

Bulkheads

A bulkhead gives each dependency its own slice of resources, such as a separate connection pool and concurrency cap, so a slow one can only exhaust its own share. If the recommendations service hangs, only its slots fill up and payment calls keep their capacity. When a bulkhead is full, fail fast with a fallback instead of queueing without limit.

Asynchronous messaging

With messaging, the sender hands a message to a broker and moves on. The receiver doesn't need to be up at that moment, spikes get absorbed, and new consumers can subscribe without changing the producer. The price is eventual consistency, harder debugging, and a broker to operate. My examples hide broker-specific code behind one small interface:

src/messaging/types.tsTypeScript
export interface Message {
  id: string; // assigned by the producer, stable across redeliveries
  type: string; // "OrderPlaced", "ReserveStock", ...
  key: string; // partition or routing key, e.g. the order ID
  payload: unknown;
  headers: Record<string, string>;
}
 
// Resolves only once the broker has durably accepted the message
// (publisher confirms in RabbitMQ, acks=all in Kafka).
export type Publish = (topic: string, message: Message) => Promise<void>;

Queues, pub/sub, and event streams

A queue distributes work: each message goes to one of several competing consumers, which acknowledges it after processing so the broker can delete it. RabbitMQ is the classic example; for pub/sub, each subscriber binds its own queue to an exchange and gets its own copy.

An event stream is an append-only, partitioned log, with Kafka the best-known example. Messages stay for a retention period whether or not anyone has read them, and each consumer group tracks its own offset, so many consumers can read the same events and a new one can replay history. Order holds only within a partition, and the partition count caps a group's parallelism.

Commands vs. events

A command asks one service to do something, like ReserveStock. It is imperative, has exactly one handler, and can be rejected. An event states a fact, like OrderPlaced. It is past tense, has any number of subscribers, and can only be reacted to. Commands make the sender depend on the receiver; with events, consumers depend on the publisher's schema instead.

Watch for commands disguised as events: if things break when one particular consumer ignores SendWelcomeEmailRequested, it's a command and belongs on the email service's queue.

Delivery guarantees and idempotent consumers

Delivery guarantees come in three flavors:

  • At-most-once: acknowledge before processing. A crash loses the message, but it never runs twice.
  • At-least-once: acknowledge after processing. A crash in between causes redelivery, so duplicates are normal. This is the practical default.
  • Exactly-once: read the fine print.

The realistic target is at-least-once delivery with idempotent consumers. Some handlers are naturally idempotent, like setting a status to SHIPPED. For the rest, I record each message ID in the same transaction as the business change:

migrations/003_processed_messages.sqlSQL
CREATE TABLE processed_messages (
  consumer     text        NOT NULL,
  message_id   text        NOT NULL,
  processed_at timestamptz NOT NULL DEFAULT now(),
  PRIMARY KEY (consumer, message_id)
);
src/db/transaction.tsTypeScript
import type { Pool, PoolClient } from "pg";
 
export async function withTransaction<T>(
  pool: Pool,
  work: (client: PoolClient) => Promise<T>,
): Promise<T> {
  const client = await pool.connect();
  try {
    await client.query("BEGIN");
    const result = await work(client);
    await client.query("COMMIT");
    return result;
  } catch (err) {
    await client.query("ROLLBACK");
    throw err;
  } finally {
    client.release();
  }
}
src/messaging/handle-once.tsTypeScript
import type { Pool, PoolClient } from "pg";
import { withTransaction } from "../db/transaction";
 
export async function handleOnce(
  pool: Pool,
  consumer: string,
  messageId: string,
  handle: (client: PoolClient) => Promise<void>,
): Promise<void> {
  await withTransaction(pool, async (client) => {
    const { rowCount } = await client.query(
      `INSERT INTO processed_messages (consumer, message_id)
       VALUES ($1, $2)
       ON CONFLICT (consumer, message_id) DO NOTHING`,
      [consumer, messageId],
    );
    if (rowCount === 0) return; // duplicate: already processed, just ack it
 
    await handle(client); // business writes share the same transaction
  });
}

A duplicate hits the primary key and is skipped. If the handler throws, everything rolls back and the redelivery starts clean; acknowledge only after the commit. The message ID must come from the producer and stay stable across redeliveries, and this only protects writes to the same database. For external APIs, pass the message ID as their idempotency key.

Dead-letter queues and poison messages

A poison message fails every time: a malformed payload, an unknown schema version, a bug triggered by one record. Under at-least-once delivery it's redelivered forever, burning CPU and, on an ordered stream, blocking everything behind it. Bound the retries, then move it to a dead-letter queue (DLQ):

  • Retry transient errors, like timeouts, with backoff; dead-letter permanent ones, like validation failures, right away.
  • Store the source, error, attempt count, and timestamp with the message.
  • Give the DLQ an owner, alert on its depth, and redrive messages with a tested tool once the cause is fixed.

RabbitMQ supports dead-lettering natively; with Kafka, the consumer produces the failed record to a separate topic and commits the offset so the partition keeps moving. A DLQ nobody watches is just a slower way to lose data.

The dual-write problem and the transactional outbox

This bug is easy to ship in a first event-driven service:

TypeScript
// Two writes to two systems, with no transaction spanning them
await pool.query(
  "INSERT INTO orders (id, customer_id, status) VALUES ($1, $2, 'PENDING')",
  [orderId, customerId],
);
await publish("order-events", orderPlaced); // crash before this line and the event is lost

If the process crashes or the broker is down after the insert, the order exists but nobody hears about it. Publish first and you get the reverse: an event for an order that was never saved.

The transactional outbox turns two writes into one: the service inserts the event into an outbox table in the same transaction as the business change, and a separate relay publishes it later. The order and its event commit together or not at all.

The outbox table

migrations/004_outbox.sqlSQL
CREATE TABLE outbox (
  id             uuid        PRIMARY KEY DEFAULT gen_random_uuid(), -- built in since PostgreSQL 13
  aggregate_type text        NOT NULL,              -- 'order'
  aggregate_id   text        NOT NULL,              -- becomes the message key
  event_type     text        NOT NULL,              -- 'OrderPlaced'
  payload        jsonb       NOT NULL,
  headers        jsonb       NOT NULL DEFAULT '{}', -- trace context and other metadata
  created_at     timestamptz NOT NULL DEFAULT now(),
  published_at   timestamptz
);
 
-- The relay only ever reads unpublished rows
CREATE INDEX outbox_unpublished_idx ON outbox (created_at) WHERE published_at IS NULL;

The partial index keeps the relay's query cheap as published rows accumulate. Writing the event is one more insert on the same client:

src/orders/place-order.tsTypeScript
import type { Pool } from "pg";
import { withTransaction } from "../db/transaction";
 
export interface PlaceOrderInput {
  customerId: string;
  items: { sku: string; quantity: number }[];
  totalCents: number;
}
 
export async function placeOrder(pool: Pool, input: PlaceOrderInput): Promise<string> {
  return withTransaction(pool, async (client) => {
    const { rows } = await client.query<{ id: string }>(
      `INSERT INTO orders (customer_id, total_cents, status)
       VALUES ($1, $2, 'PENDING') RETURNING id`,
      [input.customerId, input.totalCents],
    );
    const orderId = rows[0].id;
 
    await client.query(
      `INSERT INTO outbox (aggregate_type, aggregate_id, event_type, payload)
       VALUES ('order', $1, 'OrderPlaced', $2)`,
      [orderId, JSON.stringify({ orderId, ...input })],
    );
    return orderId;
  });
}

The relay

The relay claims a batch of unpublished rows, publishes them, and marks them done:

src/outbox/relay.tsTypeScript
import { setTimeout as sleep } from "node:timers/promises";
import type { Pool } from "pg";
import { withTransaction } from "../db/transaction";
import type { Publish } from "../messaging/types";
 
interface OutboxRow {
  id: string;
  aggregate_type: string;
  aggregate_id: string;
  event_type: string;
  payload: unknown;
  headers: Record<string, string>;
}
 
export async function relayBatch(pool: Pool, publish: Publish, batchSize = 100): Promise<number> {
  return withTransaction(pool, async (client) => {
    const { rows } = await client.query<OutboxRow>(
      `SELECT id, aggregate_type, aggregate_id, event_type, payload, headers
         FROM outbox
        WHERE published_at IS NULL
        ORDER BY created_at
        LIMIT $1
        FOR UPDATE SKIP LOCKED`,
      [batchSize],
    );
 
    for (const row of rows) {
      await publish(`${row.aggregate_type}-events`, {
        id: row.id,
        type: row.event_type,
        key: row.aggregate_id,
        payload: row.payload,
        headers: row.headers,
      });
    }
 
    if (rows.length > 0) {
      await client.query("UPDATE outbox SET published_at = now() WHERE id = ANY($1::uuid[])", [
        rows.map((row) => row.id),
      ]);
    }
    return rows.length;
  });
}
 
export async function runRelay(pool: Pool, publish: Publish, signal: AbortSignal): Promise<void> {
  while (!signal.aborted) {
    const published = await relayBatch(pool, publish);
    if (published === 0) await sleep(500); // nothing to send: poll again shortly
  }
}

FOR UPDATE SKIP LOCKED makes each relay skip rows another relay has locked, so several relays can run for availability without double-publishing in normal operation. If a relay crashes after publishing but before its UPDATE commits, those rows go out again: at-least-once delivery, which is why consumers deduplicate on the outbox id.

The relay holds a transaction open while publishing, so keep batches small. Ordering is approximate, since parallel relays and out-of-order commits can reorder events; if consumers need per-aggregate order, add a version number they can check. And purge published rows on a schedule.

Change data capture as an alternative

Polling adds steady query load and up to one interval of latency. Change data capture removes the polling: Debezium reads PostgreSQL's write-ahead log via logical replication, emits changes in commit order, and ships an outbox event router for this exact pattern. The cost is more infrastructure, usually Kafka Connect, and a replication slot to monitor: an inactive slot makes PostgreSQL retain WAL until the disk fills, unless you cap it with max_slot_wal_keep_size.

Sagas for multi-service workflows

Placing an order touches orders, inventory, and payments, each with its own database. Two-phase commit across them is rarely practical: every participant must support it, it holds locks across network calls, and it re-couples their availability.

A saga replaces the single transaction with a sequence of local ones. Each step commits in its own service and triggers the next; if a step fails, compensating actions undo the completed steps in reverse order.

Compensations are semantic undos, not rollbacks: a refund offsets a charge rather than erasing it. They must be idempotent and retried until they succeed, and I order steps so the hardest to undo runs last. Releasing reserved stock is cheap and refunding a card is not, so the reservation comes first. Sagas also give up isolation, so make intermediate state explicit with statuses like PENDING.

Choreography

With choreography there is no coordinator. Each service reacts to events and publishes its own:

Text
Order      publishes OrderPlaced
Inventory  on OrderPlaced      → reserves stock → StockReserved or StockUnavailable
Payment    on StockReserved    → charges card   → PaymentCharged or PaymentDeclined
Inventory  on PaymentDeclined  → releases stock → StockReleased
Order      on PaymentCharged   → CONFIRMED
Order      on StockUnavailable → CANCELLED
Order      on StockReleased    → CANCELLED

It's loosely coupled and needs no extra component, which suits short flows. As the workflow grows, the logic scatters, and finding where order 42 is stuck means correlating logs from several services.

Orchestration

With orchestration, one component owns the workflow, usually the service responsible for the outcome. It sends commands, handles replies, and decides the next step. I model it as a state machine persisted in the database:

src/orders/order-saga.tsTypeScript
type SagaState = "RESERVING" | "CHARGING" | "RELEASING" | "CONFIRMED" | "CANCELLED";
 
type Reply =
  | "StockReserved"
  | "StockUnavailable"
  | "PaymentCharged"
  | "PaymentDeclined"
  | "StockReleased";
 
export interface OrderSaga {
  orderId: string;
  state: SagaState;
  // In one transaction: update the saga row only if it is still in its current
  // state, then insert the next command (if any) into the outbox.
  transition(next: SagaState, command?: string): Promise<void>;
}
 
export async function onReply(saga: OrderSaga, reply: Reply): Promise<void> {
  switch (`${saga.state}:${reply}`) {
    case "RESERVING:StockReserved":
      return saga.transition("CHARGING", "ChargePayment");
    case "RESERVING:StockUnavailable":
      return saga.transition("CANCELLED");
    case "CHARGING:PaymentCharged":
      return saga.transition("CONFIRMED");
    case "CHARGING:PaymentDeclined":
      return saga.transition("RELEASING", "ReleaseStock"); // compensation
    case "RELEASING:StockReleased":
      return saga.transition("CANCELLED");
    default:
      return; // duplicate or stale reply: nothing to do
  }
}

The transaction that creates the order also stores the saga in RESERVING and writes a ReserveStock command to the outbox. Replies that don't match the current state, such as duplicates, are ignored. Participants must be idempotent too: the payment service should key charges on the order ID so a redelivered ChargePayment never charges twice.

The workflow reads top to bottom in one file, which choreography can't offer; keep business rules in the participants so the orchestrator doesn't become a god service. For long-running flows, a workflow engine such as Temporal or AWS Step Functions handles persistence, timers, and retries.

Observability across service boundaries

One user action can span HTTP calls, messages, and background jobs in several processes, and you need a shared identifier to reconstruct it. The minimum is a correlation ID, created at the edge, passed on every call and message, and logged on every line. Better is distributed tracing with W3C Trace Context, which standardizes a traceparent header of four fields: version, trace ID, parent span ID, and flags.

Text
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01

OpenTelemetry auto-instrumentation propagates it on HTTP calls, and many broker-client instrumentations add it to message headers. The outbox is a gap, since the relay publishes after the request's span has ended, so capture the context when writing the row:

src/observability/trace-headers.tsTypeScript
import { context, propagation } from "@opentelemetry/api";
 
// Call while handling the request, then save the result in outbox.headers
export function traceHeaders(): Record<string, string> {
  const carrier: Record<string, string> = {};
  propagation.inject(context.active(), carrier); // adds traceparent (and tracestate)
  return carrier;
}

The consumer passes those headers to propagation.extract() and starts its span in that context, so one trace covers the request, the relay, and every consumer.

A decision guide

SituationMy choiceWhy
Caller needs the result to continueSync REST or gRPC callThe user is waiting on the answer
Hot-path reads of another service's dataLocal copy kept fresh by eventsRemoves the runtime dependency
Side effects nobody waits for (email, indexing)Event through an outboxProducer stays fast when consumers are down
One service must do a job, now or laterCommand on a queueCompeting consumers, retries, a DLQ
Many consumers, replay, or an audit trailEvent streamRetained, replayable, ordered per partition
Workflow across services that can fail midwayOrchestrated sagaCompensations instead of distributed transactions

My default recommendation

On a new system, I use synchronous calls for queries and anything a user is waiting on, each with a timeout, a circuit breaker, and retries limited to idempotent operations. State changes publish events through an outbox, handled by idempotent consumers with a watched DLQ. I avoid synchronous chains deeper than two hops, giving services local copies of data fed by events instead, and I use orchestrated sagas for workflows that need compensation.

Async is not the goal. It buys decoupling at the cost of immediate consistency and simple debugging, so I spend it where decoupling pays off.

Key takeaways

  • Decide each interaction by who needs the answer, and when, not by the tools you already run.
  • Give every synchronous call a timeout; retry only idempotent operations, with backoff and full jitter, at one layer.
  • Circuit breakers and bulkheads stop one slow dependency from becoming an outage.
  • Assume at-least-once delivery: deduplicate on producer-assigned IDs and route poison messages to a DLQ someone owns.
  • Never dual-write. A transactional outbox, drained by a relay or CDC, keeps state and events consistent.
  • Use sagas with explicit compensations across services, and propagate traceparent to follow them.

None of these patterns is exotic, and none is free. Start with timeouts, idempotent consumers, and the outbox, then add the rest as real failures demand.