My first deploy script for a Docker app was docker compose pull && docker compose up -d, and every run caused a small outage. Compose stopped the old container before starting the new one, so for a few seconds requests failed, in-flight requests were cut off, and users with sessions in process memory were logged out.

Zero-downtime deployments fix this without Kubernetes. In this post I build the pipeline for a Node.js API: graceful shutdown and health checks, a GitHub Actions workflow that pushes images to GitHub Container Registry, blue-green deployments on one VM with Docker Compose and Caddy, expand-and-contract migrations, and rollbacks in seconds. The payoff: a push to main ships while users keep clicking. The examples use Node.js 24, Docker Engine with Compose, Caddy 2, and PostgreSQL.

What actually causes downtime during a deploy

Deploy downtime is rarely a full outage; it's a burst of failed requests from four sources:

  • Killed in-flight requests. docker stop sends SIGTERM, waits 10 seconds by default, then sends SIGKILL. An app that ignores SIGTERM, or exits the instant it arrives, takes every open request with it.
  • Cold starts before readiness. A container counts as running once its process starts, while the app may still be connecting to the database, so early traffic fails.
  • Incompatible migrations. Old and new code overlap during every deploy. A release that renames a column breaks the old version the moment the migration runs.
  • Cache and session loss. In-memory sessions log everyone out on restart, and a cold local cache floods the database just as the new version starts.

The last one has the simplest fix: keep state out of the process. Store sessions in Redis or your database (JWT vs. sessions compares the options), and version cache keys when their format changes, as in caching strategies with Redis. The rest of this post handles the other three.

Graceful shutdown and health checks

Clean stops and clean starts both begin in the app.

Handling SIGTERM in a Node.js server

I use the bare node:http module; Express's app.listen() returns the same http.Server, so this carries over. The server has two health endpoints and a shutdown flag:

src/server.tsTypeScript
import http from 'node:http';
import { Pool } from 'pg';
import { routes } from './routes.js'; // your application's request handler
 
const pool = new Pool({
  connectionString: process.env.DATABASE_URL,
  connectionTimeoutMillis: 2_000, // a probe should fail fast, not hang
});
 
let shuttingDown = false;
 
const server = http.createServer(async (req, res) => {
  // While draining, tell clients and proxies not to reuse this connection.
  if (shuttingDown) res.setHeader('Connection', 'close'); 
 
  if (req.url === '/healthz') {
    // Liveness: the process is up and the event loop responds. No dependencies.
    res.writeHead(200).end('ok');
    return;
  }
 
  if (req.url === '/readyz') {
    // Readiness: should this instance receive traffic right now?
    const ready =
      !shuttingDown && (await pool.query('SELECT 1').then(() => true, () => false));
    res.writeHead(ready ? 200 : 503).end(ready ? 'ready' : 'not ready');
    return;
  }
 
  await routes(req, res);
});
 
server.listen(3000, () => console.log('listening on :3000'));

On SIGTERM, the server marks itself not ready, stops accepting connections, waits for in-flight requests, closes the database pool, and exits:

src/server.ts (continued)TypeScript
const DRAIN_TIMEOUT_MS = 20_000; // keep it below Compose's stop_grace_period
 
async function shutdown(signal: NodeJS.Signals) {
  if (shuttingDown) return;
  shuttingDown = true;
  console.log(`${signal} received, draining`);
 
  // Safety net: exit on our own terms before Docker escalates to SIGKILL.
  setTimeout(() => {
    console.error('drain timed out, exiting with requests still open');
    process.exit(1);
  }, DRAIN_TIMEOUT_MS).unref();
 
  // Stop accepting connections and close idle keep-alive sockets (Node 19+).
  // The callback runs once every in-flight request has finished.
  await new Promise<void>((resolve) => server.close(() => resolve()));
  await pool.end(); // last: in-flight requests may still have needed it
  console.log('drained, exiting');
  process.exit(0);
}
 
for (const signal of ['SIGTERM', 'SIGINT'] as const) {
  process.on(signal, () => {
    shutdown(signal).catch((err) => {
      console.error('shutdown failed', err);
      process.exit(1);
    });
  });
}

server.close() refuses new connections and, since Node.js 19, closes idle keep-alive connections. Busy connections finish their request; if a client sends another request on one during the drain, the highlighted Connection: close header closes it after the response. The pool closes last because in-flight requests may need it, and the timer guarantees an exit before Docker escalates to SIGKILL.

In my tests, a connection that was mid-request at SIGTERM stayed open until Node's keep-alive timeout closed it, about six seconds on Node.js 24, hence the generous deadline. If a load balancer routes on readiness, also pause a few seconds before server.close() so it sees the 503 first.

Getting the signal to Node: exec form, PID 1, and timeouts

A handler only helps if the signal reaches it. Three things decide that.

The CMD form. The shell form, CMD node dist/server.js, runs your command through /bin/sh -c, and the Dockerfile reference warns that the shell doesn't pass signals on, so your app may never see SIGTERM. An npm start wrapper adds a similar middleman. Use the exec form, a JSON array, so node is the main process. The health check gets its own tiny script:

src/healthcheck.tsTypeScript
// Docker runs this inside the container: exit code 0 means healthy, 1 unhealthy.
fetch('http://127.0.0.1:3000/readyz')
  .then((res) => process.exit(res.ok ? 0 : 1))
  .catch(() => process.exit(1));
DockerfileDockerfile
# syntax=docker/dockerfile:1
FROM node:24-slim AS build
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build && npm prune --omit=dev
 
FROM node:24-slim
ENV NODE_ENV=production
WORKDIR /app
COPY --from=build /app/package.json ./
COPY --from=build /app/node_modules ./node_modules
COPY --from=build /app/dist ./dist
USER node
EXPOSE 3000
 
# Healthy once /readyz answers 200. Failures during the start period don't count.
HEALTHCHECK --interval=5s --timeout=3s --start-period=30s --retries=3 \
  CMD ["node", "dist/healthcheck.js"]
 
# Exec form: node is the container's main process and receives SIGTERM itself.
CMD ["node", "dist/server.js"]

PID 1. The container's main process is PID 1, and the kernel ignores signals that PID 1 has no handler for (SIGKILL aside), so an app without a SIGTERM handler always waits out the full timeout. PID 1 must also reap orphaned child processes, which Node doesn't do. init: true in Compose, or docker run --init, runs tini as PID 1 to forward signals and reap zombies; elsewhere, install tini and use ENTRYPOINT ["tini", "--"].

The timeout. Docker waits 10 seconds after SIGTERM before sending SIGKILL. Compose's stop_grace_period raises that limit; keep it above the app's own deadline, 30 seconds against 20 here.

Liveness vs. readiness, and Docker's HEALTHCHECK

Graceful shutdown protects requests on the way out; readiness protects them on the way in.

  • Liveness (/healthz): is the process alive? It checks no dependencies, because restarting the app won't fix a database outage. Orchestrators restart instances that fail it.
  • Readiness (/readyz): should this instance get traffic right now? It fails until the server is up, then returns 503 while draining or when the database is unreachable. Load balancers stop routing to instances that fail it.

Kubernetes has a probe for each. Docker has one HEALTHCHECK per container and, outside Swarm, only reports its status without restarting anything. That makes it a natural deploy gate, so I point it at /readyz. Failures during --start-period don't count, and the first success marks the container healthy at once.

Checking the database is a trade-off: a release with a wrong DATABASE_URL never sees traffic, but a database outage marks every instance unhealthy. That's harmless here, where nothing routes on Docker's health status; behind a load balancer, I'd check only local state.

Rolling, blue-green, and canary deployments

The deployment strategy decides how traffic moves to the new version:

StrategyHow traffic movesExtra capacityRollbackGood fit
RollingReplace instances a few at a timeLittle or noneRoll out the old image againMany replicas; the Kubernetes default
Blue-greenStart a full new copy, then switch all traffic at onceA second copy during the deploySwitch back, instantly if the old copy runsOne VM or a small fleet
CanarySend a small share first, then moreA few instancesRoute the canary share backHigh traffic, good metrics

Rolling deploys mix old and new versions for the whole rollout. Canaries limit the blast radius but need weighted routing and solid metrics. Blue-green costs a second copy, on one VM just a second container, and buys an atomic cutover with an instant way back, so it's my choice for single VMs.

Building and pushing images with GitHub Actions

The workflow has three jobs: test runs the suite, build pushes an image to GitHub Container Registry (GHCR), and deploy runs a script on the server over SSH. Each job runs only if the previous one succeeded.

.github/workflows/deploy.ymlYAML
name: Deploy
 
on:
  push:
    branches: [main]
 
permissions:
  contents: read
 
env:
  IMAGE: ghcr.io/acme/shop-api # lowercase: registries reject uppercase names
 
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: actions/setup-node@v7
        with:
          node-version: 24
          cache: npm
      - run: npm ci
      - run: npm test
 
  build:
    needs: test
    runs-on: ubuntu-latest
    permissions:
      contents: read
      packages: write # lets GITHUB_TOKEN push to GHCR
    steps:
      - uses: actions/checkout@v7
      - uses: docker/setup-buildx-action@v4
      - uses: docker/login-action@v4
        with:
          registry: ghcr.io
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}
      - id: meta
        uses: docker/metadata-action@v6
        with:
          images: ${{ env.IMAGE }}
          tags: |
            type=sha,format=long
            type=ref,event=branch
      - uses: docker/build-push-action@v7
        with:
          context: .
          push: true
          tags: ${{ steps.meta.outputs.tags }}
          labels: ${{ steps.meta.outputs.labels }}
          cache-from: type=gha
          cache-to: type=gha,mode=max
 
  deploy:
    needs: build
    runs-on: ubuntu-latest
    environment: production
    concurrency: production # never two deploys against the same VM at once
    steps:
      - name: Install the deploy key and pin the server's host key
        env:
          SSH_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
          KNOWN_HOSTS: ${{ secrets.DEPLOY_KNOWN_HOSTS }}
        run: |
          install -m 700 -d ~/.ssh
          printf '%s\n' "$SSH_KEY" > ~/.ssh/id_ed25519
          chmod 600 ~/.ssh/id_ed25519
          printf '%s\n' "$KNOWN_HOSTS" > ~/.ssh/known_hosts
      - name: Blue-green deploy
        env:
          HOST: ${{ vars.DEPLOY_HOST }}
          TAG: sha-${{ github.sha }}
        run: ssh "deploy@$HOST" "/opt/shop/deploy.sh $TAG"

Key choices:

  • Immutable tags. type=sha,format=long tags each image sha- plus the full commit hash, which names exactly one build; the deploy job rebuilds that name from sha-${{ github.sha }}. The main tag is for humans, not deploys. docker/metadata-action also adds OCI labels that link the package to your repository.
  • Least privilege. Only build gets packages: write, so the built-in GITHUB_TOKEN can push to GHCR without a personal token. type=gha caches BuildKit layers between runs.
  • One deploy at a time. With concurrency, a newer deploy waits for the running one and replaces any older one still waiting. The production environment adds protection rules, like required reviewers, and scopes the SSH secrets.
  • A pinned host key. DEPLOY_KNOWN_HOSTS pins the server's host key (from ssh-keyscan, verified out of band); StrictHostKeyChecking=no would let anyone who redirects the hostname impersonate your server.

On the server, a deploy user in the docker group runs the script, and Docker's post-installation guide warns that this group grants root-level privileges, so guard that key like a root key. For a private image, log the server in to GHCR once with a personal access token (classic) limited to read:packages; GitHub's Container registry docs note that only classic tokens work.

Blue-green deployments on one VM with Docker Compose and Caddy

Two app services, blue and green, run the same image at different tags. Caddy, the only container with published ports, proxies to whichever color is live. A deploy starts the idle color, waits until it's healthy, points Caddy at it, and retires the old one.

The Compose file and Caddyfile

/opt/shop/compose.yamlYAML
name: shop
 
x-app: &app
  init: true                # tini as PID 1: forwards signals, reaps zombie processes
  restart: unless-stopped
  env_file: app.env         # DATABASE_URL and other runtime secrets
  stop_grace_period: 30s    # Docker's default is 10s; the app gives up after 20s
 
services:
  caddy:
    image: caddy:2
    restart: unless-stopped
    ports:
      - "80:80"
      - "443:443"
      - "443:443/udp"
    volumes:
      - ./caddy:/etc/caddy:ro   # the directory, so edits on the host show up inside
      - caddy_data:/data
      - caddy_config:/config
 
  blue:
    <<: *app
    image: ghcr.io/acme/shop-api:${BLUE_TAG:?set BLUE_TAG in .env}
 
  green:
    <<: *app
    image: ghcr.io/acme/shop-api:${GREEN_TAG:?set GREEN_TAG in .env}
 
volumes:
  caddy_data:
  caddy_config:

The x-app block holds what both colors share. Each tag comes from .env, which Compose reads automatically and the deploy script updates; :? makes Compose fail loudly if a tag is missing.

/opt/shop/caddy/CaddyfileText
shop.example.com {
	import /etc/caddy/upstream.caddy
}
/opt/shop/caddy/upstream.caddyText
reverse_proxy blue:3000

Caddy gets and renews the TLS certificate itself once DNS points at the VM. The upstream sits in its own file so the script rewrites one line. I mount the directory rather than single files: a single-file bind mount keeps pointing at the original file, so edits that replace it on the host, as sed -i does, never reach the container.

I chose Caddy because its /load endpoint, which caddy reload calls, blocks until the new config runs and restores the old one if anything fails. nginx can do the same with an included upstream file and nginx -s reload.

The deploy script

Terminal/opt/shop/deploy.sh
#!/usr/bin/env bash
# Blue-green deploy on a single VM.
#   ./deploy.sh sha-<commit>   start that image in the idle color, then switch to it
#   ./deploy.sh rollback       switch back to the idle color and the image it last ran
set -euo pipefail
cd "$(dirname "$0")"
 
TARGET="${1:?usage: deploy.sh <image-tag> | rollback}"
LIVE=$(grep -oE 'blue|green' caddy/upstream.caddy)  # the proxy config says who is live
if [ "$LIVE" = blue ]; then IDLE=green; else IDLE=blue; fi
TAG_VAR="${IDLE^^}_TAG"                             # BLUE_TAG or GREEN_TAG in .env
 
if [ "$TARGET" != rollback ]; then
  sed -i "/^$TAG_VAR=/d" .env && echo "$TAG_VAR=$TARGET" >> .env
  docker compose pull "$IDLE"
  # Expand-only migrations, run with the new image while the live color serves.
  docker compose run --rm --no-deps "$IDLE" node dist/migrate.js
fi
 
# 1. Start the idle color and wait until its HEALTHCHECK passes.
echo "Starting $IDLE ($(grep "^$TAG_VAR=" .env)); $LIVE keeps serving"
if ! docker compose up -d --no-deps --wait --wait-timeout 120 "$IDLE"; then
  docker compose logs --tail 50 "$IDLE"
  docker compose stop "$IDLE"
  echo "$IDLE never became healthy; $LIVE is still live" >&2
  exit 1
fi
 
# 2. Switch traffic. Caddy validates the new config, then swaps it in gracefully.
echo "reverse_proxy $IDLE:3000" > caddy/upstream.caddy
if ! docker compose exec -T caddy caddy reload --config /etc/caddy/Caddyfile; then
  echo "reverse_proxy $LIVE:3000" > caddy/upstream.caddy
  echo "Caddy rejected the switch; $LIVE is still live" >&2
  exit 1
fi
 
# 3. Retire the old color: SIGTERM, then up to stop_grace_period to drain.
#    KEEP_WARM=1 leaves it running, so a rollback only needs step 2.
if [ "${KEEP_WARM:-0}" != 1 ]; then
  docker compose stop "$LIVE"
fi
echo "$IDLE is live"

The live color comes from upstream.caddy, so the proxy config is the source of truth. A new release gets its tag recorded, its image pulled, and its migrations run while the old color serves, so those migrations must work with the live version (swap in your tool for node dist/migrate.js). docker compose up --wait returns once the HEALTHCHECK passes and fails if the container turns unhealthy, exits, or times out; the script then prints the logs and stops the new color before any user notices. If Caddy rejects the reload, the upstream file is restored. Finally, docker compose stop sends SIGTERM and the old color drains; it's stopped, not removed, which makes rollbacks fast.

First deploy and a smoke test

Copy compose.yaml, deploy.sh, caddy/Caddyfile, and an app.env with your secrets to /opt/shop, then run once as the deploy user:

Terminal
cd /opt/shop
echo "reverse_proxy blue:3000" > caddy/upstream.caddy
printf 'BLUE_TAG=none\nGREEN_TAG=none\n' > .env
chmod +x deploy.sh && chmod 600 app.env
docker login ghcr.io -u YOUR_GITHUB_USER   # password: the read:packages token
docker compose up -d caddy

Caddy starts out pointing at a blue that doesn't exist yet, so the first deploy goes to green. During a deploy, this loop prints any request that doesn't return 200:

Terminal
while true; do
  code=$(curl -s -o /dev/null -w '%{http_code}' https://shop.example.com/healthz)
  [ "$code" = 200 ] || echo "$(date +%T) got $code"
  sleep 0.2
done

Database migrations with expand and contract

Every strategy above runs two versions against one database at some point, and blue-green keeps the previous version as your rollback target, so each schema change must work with both. Expand and contract splits a breaking change, like renaming users.name to full_name, into backward-compatible steps:

StepSchema changeApplication code
1. ExpandAdd a nullable full_name columnWrites both columns, reads name
2. BackfillCopy name into full_name for existing rowsNo change
3. Switch readsNoneReads full_name, still writes both
4. Stop old writesNoneReads and writes only full_name
5. ContractDrop the name columnNo change

Each step ships as its own deploy and works with the one before it, so any step can be rolled back. Step 1's dual write:

src/users/repository.tsTypeScript
import { pool } from '../db.js';
 
export async function renameUser(id: string, fullName: string): Promise<void> {
  await pool.query(
    'UPDATE users SET name = $1, full_name = $1 WHERE id = $2', 
    [fullName, id],
  );
}

The schema steps, in PostgreSQL:

SQL
-- Step 1, expand: additive, so the version that is live right now keeps working.
SET lock_timeout = '5s'; -- give up instead of queueing behind a long transaction
ALTER TABLE users ADD COLUMN full_name text;
 
-- Step 2, backfill: small batches keep locks short. Repeat until it updates 0 rows.
UPDATE users SET full_name = name
WHERE id IN (
  SELECT id FROM users
  WHERE full_name IS NULL AND name IS NOT NULL
  LIMIT 1000
);
 
-- Step 5, contract: a later release, once nothing running or rollback-ready reads name.
SET lock_timeout = '5s';
ALTER TABLE users DROP COLUMN name;

Adding a nullable column and dropping one are quick catalog changes, but each takes a brief exclusive lock. If a long transaction holds a conflicting lock, the ALTER waits and every later query queues behind it; lock_timeout makes it give up so you can retry. Build new indexes with CREATE INDEX CONCURRENTLY so writes keep flowing, as I explain in how database indexes work.

Rolling back, and a pre-deploy checklist

Three ways back

  • Instant: keep the old color warm. With KEEP_WARM=1, the previous color keeps running, so ./deploy.sh rollback finds it healthy and only reloads Caddy. It doubles memory and database connections, and in-process scheduled jobs run twice, so I keep it for risky releases.
  • Seconds: ./deploy.sh rollback. The previous color is stopped, not removed, with its image on disk, so the script restarts it, waits for the health check, and switches back.
  • Any version: deploy an older tag. Passing an older sha- tag takes the same health-gated path. GHCR doesn't delete old images on its own; if you add a cleanup policy, keep the last few releases.

A failed deploy needs no rollback, because traffic never moved. For a hard guarantee that a tag still means the same bytes, deploy by the digest output of docker/build-push-action. And a rollback undoes code, not data: a dropped column doesn't come back, which is the real reason for expand and contract.

Pre-deploy checklist

Before I trust a service with these deploys, I check:

  • The app handles SIGTERM, with a drain deadline shorter than stop_grace_period.
  • The Dockerfile uses an exec-form CMD, and Compose sets init: true.
  • The HEALTHCHECK calls a /readyz that checks what requests need.
  • This release's migrations are expand-only.
  • Sessions and shared caches live outside the process.
  • The previous release's image tag still exists in GHCR.
  • A rollback and a deploy under load both succeed in staging.

Key takeaways

  • Downtime comes from killed requests, traffic to instances that aren't ready, schemas one version can't use, and state in process memory.
  • Drain on SIGTERM and exit before Docker's timeout; an exec-form CMD and init: true make sure the signal arrives.
  • Keep liveness and readiness separate, and gate every traffic switch on readiness.
  • On one VM, blue-green with Compose and Caddy gives an atomic cutover: start the idle color, wait for healthy, reload, retire.
  • Expand and contract keeps the schema compatible with both the live version and your rollback target.
  • Tag images by commit so a rollback is just another deploy, and rehearse it.

None of this needs Kubernetes. It's a signal handler, two endpoints, a short script, and patience with schema changes. If you do one thing this week, add the SIGTERM handler; it protects every restart, not only deploys.