For most of Java's history, a server handled each request on its own thread, and that thread sat idle whenever the code waited on a database or another service. Threads were expensive, so we pooled them, and busy services ran out of threads long before they ran out of CPU. The standard escape was reactive programming: rewrite everything as chains of Mono and Flux operators so that no thread ever waits.

Java virtual threads, final since JDK 21, offer a different deal. You keep writing plain blocking code, one thread per request, and the JDK makes each thread cheap enough to run hundreds of thousands of them. They are also easy to misuse in ways that only show up under load.

Below I show how they work with a runnable fan-out of a thousand blocking HTTP calls, then cover what bites in production and when to use something else.

Why thread-per-request hit a wall

A platform thread, the kind new Thread() has always created, wraps an operating system thread. Each one reserves a fixed-size stack up front (1 MB by default on 64-bit Linux), and the kernel schedules every one of them. A few thousand are fine; a few hundred thousand are not. So servers pool their threads: Tomcat processes requests on at most 200 threads by default.

Little's law says how much concurrency you need: requests in flight equal throughput times latency. With illustrative numbers, 1,000 requests per second at 200 ms each means 200 in flight, exactly what the pool holds. If a downstream service slows and requests take a second, you need 1,000 threads. The pool fills, requests queue, and latency climbs while the CPU idles, because those 200 threads are all waiting on sockets.

Reactive frameworks such as Project Reactor never block, so a few event-loop threads juggle thousands of requests. The price: every API returns a Mono or a Flux, stack traces stop describing your code, debugging gets harder, and blocking libraries such as JDBC drivers need special handling. Virtual threads aim for the same scalability with code that reads top to bottom.

How Java virtual threads work

A virtual thread is still a java.lang.Thread, but it isn't tied to an OS thread for its whole life. Virtual threads previewed in JDK 19 and 20 and became final in JDK 21 with JEP 444. The JDK runs them on carrier threads: a small ForkJoinPool of platform threads, one per available processor by default.

Mounting and unmounting on carrier threads

When a virtual thread mounted on a carrier calls a blocking JDK operation, such as a socket read, Thread.sleep, BlockingQueue.take, or ReentrantLock.lock, the JDK unmounts it: its stack frames move to the heap, and the carrier picks up the next ready virtual thread. When the data arrives or the timer fires, the virtual thread is rescheduled, possibly on another carrier, and continues where it left off. Underneath, network I/O uses non-blocking sockets. File I/O is the main exception, because most operating systems can't do it asynchronously: the virtual thread keeps its carrier, and the JDK temporarily adds another to compensate.

A waiting virtual thread is just a heap object sized to the stack it actually uses, so an application can have millions of them. They don't run code faster, though. In the words of JEP 444: "Virtual threads are not faster threads — they do not run code any faster than platform threads." You gain throughput, not lower latency per request.

Creating virtual threads

There are three entry points, and this program uses all of them:

CreatingVirtualThreads.javaJava
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
 
public class CreatingVirtualThreads {
 
    public static void main(String[] args) throws Exception {
        Runnable task = () -> System.out.println("Hello from " + Thread.currentThread());
 
        // 1. The builder: name it, start it, join it, like any other Thread
        Thread named = Thread.ofVirtual().name("worker-", 0).start(task);
        named.join();
 
        // 2. A shortcut when you don't need the builder
        Thread.startVirtualThread(task).join();
 
        // 3. The one you'll use most: a new virtual thread for every task
        try (ExecutorService executor = Executors.newVirtualThreadPerTaskExecutor()) {
            var isVirtual = executor.submit(() -> Thread.currentThread().isVirtual());
            System.out.println("Executor threads are virtual: " + isVirtual.get());
        } // close() waits for every submitted task to finish
    }
}

Running java CreatingVirtualThreads.java prints something like this (thread IDs vary):

Text
Hello from VirtualThread[#27,worker-0]/runnable@ForkJoinPool-1-worker-1
Hello from VirtualThread[#30]/runnable@ForkJoinPool-1-worker-1
Executor threads are virtual: true

Each line shows the virtual thread, its name if it has one, and its current carrier. I use the per-task executor most, because it plugs into code that already speaks ExecutorService. Like every ExecutorService since JDK 19, its close() waits for submitted tasks, so try-with-resources scopes a batch of work.

Fanning out blocking HTTP calls

A classic job for virtual threads is fan-out: calling another service many times for one request or batch job. The program below starts a local HTTP server that takes one second per response, standing in for a slow downstream service, then sends it 1,000 requests with java.net.http.HttpClient, each on its own virtual thread. It needs only JDK 21 or later.

Fanout.javaJava
import com.sun.net.httpserver.HttpServer;
 
import java.io.IOException;
import java.io.OutputStream;
import java.net.InetSocketAddress;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.charset.StandardCharsets;
import java.time.Duration;
import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
 
public class Fanout {
 
    static final int REQUESTS = 1_000;
 
    public static void main(String[] args) throws IOException {
        HttpServer server = startSlowServer();
        try {
            int port = server.getAddress().getPort();
            fanOut(URI.create("http://127.0.0.1:" + port + "/slow"));
        } finally {
            server.stop(0);
        }
    }
 
    static void fanOut(URI uri) {
        List<Future<Integer>> results = new ArrayList<>();
        long start = System.nanoTime();
 
        try (HttpClient client = HttpClient.newBuilder()
                     .version(HttpClient.Version.HTTP_1_1)
                     .connectTimeout(Duration.ofSeconds(5))
                     .build();
             ExecutorService executor = Executors.newVirtualThreadPerTaskExecutor()) {
            for (int i = 0; i < REQUESTS; i++) {
                results.add(executor.submit(() -> fetchStatus(client, uri)));
            }
        } // closing the executor waits until every task has finished
 
        long millis = Duration.ofNanos(System.nanoTime() - start).toMillis();
        long ok = results.stream()
                .filter(f -> f.state() == Future.State.SUCCESS && f.resultNow() == 200)
                .count();
        System.out.printf("%d of %d requests succeeded in %d ms%n",
                ok, REQUESTS, millis);
    }
 
    static int fetchStatus(HttpClient client, URI uri)
            throws IOException, InterruptedException {
        HttpRequest request = HttpRequest.newBuilder(uri)
                .timeout(Duration.ofSeconds(10))
                .build();
        // send() blocks, so the virtual thread unmounts until the response arrives
        return client.send(request, HttpResponse.BodyHandlers.ofString()).statusCode();
    }
 
    // Stand-in for a slow downstream service: every response takes one second
    static HttpServer startSlowServer() throws IOException {
        var address = new InetSocketAddress("127.0.0.1", 0); // any free port
        HttpServer server = HttpServer.create(address, 1_000); // backlog for the burst
        server.setExecutor(Executors.newVirtualThreadPerTaskExecutor());
        server.createContext("/slow", exchange -> {
            try {
                Thread.sleep(Duration.ofSeconds(1));
            } catch (InterruptedException e) {
                Thread.currentThread().interrupt();
            }
            byte[] body = "ok".getBytes(StandardCharsets.UTF_8);
            exchange.sendResponseHeaders(200, body.length);
            try (OutputStream out = exchange.getResponseBody()) {
                out.write(body);
            }
        });
        server.start();
        return server;
    }
}

Run it straight from source:

Terminal
java Fanout.java

On my 8-core Windows machine with JDK 25, a typical run prints:

Text
1000 of 1000 requests succeeded in 2183 ms

The floor is one second, since every response takes one second; the rest is the cost of opening a thousand connections at once (with 200 requests, my runs took 1.1 to 1.4 seconds). Now swap Executors.newVirtualThreadPerTaskExecutor() on line 40 for Executors.newFixedThreadPool(100). My run took 10.6 seconds, as it must: 1,000 one-second calls through 100 threads can't finish in under 10. Your numbers will differ, but not the shape.

In the highlighted lines, client.send blocks, so each virtual thread unmounts until its response arrives, and a thousand waiting requests share 8 carriers. Note also that HttpClient is AutoCloseable since JDK 21, and that the 1_000 passed to HttpServer.create is the listen backlog: with a backlog of 50, about one request in ten failed in my runs because the burst overflowed the OS accept queue.

What does not get faster

Virtual threads only help when threads wait. CPU-bound work such as hashing passwords, compressing files, or resizing images has nothing to wait for, so nothing unmounts. All virtual threads share as many carriers as you have cores: 10,000 of them hashing passwords on 8 cores still compute 8 at a time, no faster than a fixed pool of 8 platform threads.

It can be worse than a pool. JEP 444 notes that the scheduler "does not currently implement time sharing for virtual threads", so a computing virtual thread keeps its carrier until it blocks or finishes. If every carrier is busy hashing, ready I/O-bound virtual threads wait, and cheap requests get slow. I keep CPU-heavy work on a pool sized to the hardware and call it from virtual threads:

PasswordHasher.javaJava
import java.security.GeneralSecurityException;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import javax.crypto.SecretKeyFactory;
import javax.crypto.spec.PBEKeySpec;
 
public final class PasswordHasher implements AutoCloseable {
 
    // CPU-bound work: one platform thread per core is all the parallelism there is
    private final ExecutorService cpuPool =
            Executors.newFixedThreadPool(Runtime.getRuntime().availableProcessors());
 
    // Call this from a virtual thread: waiting on get() is cheap, the hashing is not
    public byte[] hash(char[] password, byte[] salt)
            throws InterruptedException, ExecutionException {
        return cpuPool.submit(() -> pbkdf2(password, salt)).get();
    }
 
    private static byte[] pbkdf2(char[] password, byte[] salt)
            throws GeneralSecurityException {
        PBEKeySpec spec = new PBEKeySpec(password, salt, 600_000, 256);
        try {
            return SecretKeyFactory.getInstance("PBKDF2WithHmacSHA256")
                    .generateSecret(spec)
                    .getEncoded();
        } finally {
            spec.clearPassword();
        }
    }
 
    @Override
    public void close() {
        cpuPool.close();
    }
}

The virtual thread parks on get() almost for free, and because pool threads are ordinary OS threads, the operating system time-slices them with the carriers, so a burst of logins can't starve everything else.

The bottleneck moves downstream

With a 200-thread pool, at most 200 requests could touch your database at once. That limit was an accident of the thread count, and virtual threads remove it. If 5,000 requests arrive together, they all reach the connection pool. Spring Boot's default pool, HikariCP, holds at most 10 connections out of the box and lets callers wait up to 30 seconds for one, so 4,990 virtual threads queue while upstream callers time out and retry. A 5,000-connection pool only moves the problem into the database: PostgreSQL allows 100 connections by default.

Limit concurrency where the scarce resource lives. For the database, that's the connection pool: size it for what the database can handle, with a short connection timeout so excess requests fail fast. For anything else, such as a partner API that tolerates 20 concurrent calls, use a Semaphore, as Oracle's virtual threads guide also recommends.

InventoryClient.javaJava
import java.io.IOException;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
import java.util.concurrent.Semaphore;
import java.util.concurrent.TimeUnit;
 
public final class InventoryClient {
 
    // The inventory team says 20 concurrent requests is what their service can take
    private final Semaphore permits = new Semaphore(20);
    private final HttpClient http;
    private final URI baseUri;
 
    public InventoryClient(HttpClient http, URI baseUri) {
        this.http = http;
        this.baseUri = baseUri;
    }
 
    public String stockLevel(String sku) throws IOException, InterruptedException {
        // Wait briefly for a permit, then fail fast instead of queueing without limit
        if (!permits.tryAcquire(500, TimeUnit.MILLISECONDS)) {
            throw new IOException("Inventory service is at capacity, try again later");
        }
        try {
            URI uri = baseUri.resolve("/stock/" + sku);
            HttpRequest request = HttpRequest.newBuilder(uri)
                    .timeout(Duration.ofSeconds(2))
                    .build();
            return http.send(request, HttpResponse.BodyHandlers.ofString()).body();
        } finally {
            permits.release();
        }
    }
}

tryAcquire with a timeout turns overload into a fast, explicit error that your handler can map to a 503 or a fallback. Release the permit in a finally block, or every exception leaks one until the client locks up.

Pinning: when a virtual thread can't let go

A virtual thread is pinned when it blocks but can't unmount, so its carrier blocks too. Pinning doesn't break correctness, but carriers are few, and the scheduler "does not compensate for pinning by expanding its parallelism" (JEP 444). Pin as many threads as you have cores on slow I/O and every other virtual thread waits; in the worst case, the application hangs.

Synchronized blocks before and after JDK 24

On JDK 21 through 23, the usual culprit was synchronized: blocking inside a synchronized block or method, such as doing I/O while holding the monitor or calling Object.wait(), pinned the carrier. The standard workaround was a ReentrantLock, which virtual threads can park on without holding the carrier:

TokenCache.javaJava
import java.io.IOException;
import java.time.Instant;
import java.util.concurrent.locks.ReentrantLock; 
 
public class TokenCache {
 
    private final ReentrantLock lock = new ReentrantLock(); 
    private final TokenService tokenService;
    private Token current;
 
    public TokenCache(TokenService tokenService) {
        this.tokenService = tokenService;
    }
 
    public synchronized Token get() throws IOException, InterruptedException { 
    public Token get() throws IOException, InterruptedException { 
        lock.lock(); 
        try { 
            if (current == null || current.expiresAt().isBefore(Instant.now())) {
                current = tokenService.fetch(); // blocking I/O inside the lock
            }
            return current;
        } finally { 
            lock.unlock(); 
        } 
    }
 
    public record Token(String value, Instant expiresAt) {}
 
    public interface TokenService {
        Token fetch() throws IOException, InterruptedException;
    }
}

JEP 491, delivered in JDK 24, fixed this in the JVM: virtual threads now unmount while blocked in synchronized code, including Object.wait(), so neither you nor your libraries need that rewrite for pinning reasons. A few cases still pin:

  • Native code, called through JNI or the Foreign Function & Memory API, that calls back into Java code that blocks.
  • Blocking inside a class initializer (a static block), while loading a class, or while waiting for another thread to initialize one.

Finding pinned threads with JFR

JDK Flight Recorder (JFR) records a jdk.VirtualThreadPinned event when a virtual thread blocks while pinned for more than 20 ms. The event is on by default, so record under realistic load and print the events:

Terminal
# Record for 60 seconds (or until the app exits) into pinning.jfr
java -XX:StartFlightRecording=duration=60s,filename=pinning.jfr -jar app.jar
 
# Print the pinned events, with enough stack frames to reach your own code
jfr print --events jdk.VirtualThreadPinned --stack-depth 20 pinning.jfr

For an application that's already running, jcmd <pid> JFR.start duration=60s filename=pinning.jfr does the same. On JDK 24 and later, each event names the reason and the carrier. Here is a trimmed one from JDK 25, from a virtual thread that slept inside a static initializer:

Text
jdk.VirtualThreadPinned {
  duration = 113 ms
  blockingOperation = "LockSupport.park"
  pinnedReason = "VM call to SlowInit.<clinit> on stack"
  carrierThread = "ForkJoinPool-1-worker-3" (javaThreadId = 36)
  ...
}

On JDK 21 through 23, the event has a stack trace but no reason. Those versions also accept -Djdk.tracePinnedThreads=full (or short), which prints a stack trace whenever a thread blocks while pinned. JEP 491 removed that property in JDK 24, and on newer JDKs setting it does nothing.

Migrating an existing service

Most existing code runs unchanged. Two areas need attention.

ThreadLocal pitfalls and scoped values

ThreadLocal works on virtual threads, so logging contexts such as SLF4J's MDC, security contexts, and transaction state carry on as before. The trouble comes from habits built for a few hundred pooled threads:

  • Caching expensive objects per thread. A ThreadLocal<SimpleDateFormat> gave each of 200 pooled threads one reusable instance. With a virtual thread per task, every task creates one and drops it. Prefer immutable, thread-safe types such as DateTimeFormatter.
  • Memory multiplied by thread count. A thread local's value exists once per live thread: cheap with 200 threads, expensive with a million. InheritableThreadLocal values are also copied into every child thread.

One upside: a forgotten value can't leak into the next request, since each virtual thread runs one task.

For passing request context down the call stack, use a ScopedValue: it is bound for the duration of a call, immutable, and gone when the call returns. It was a preview API in JDK 21 through 24 and became final in JDK 25 with JEP 506:

RequestContext.javaJava
public class RequestContext {
 
    private static final ThreadLocal<String> REQUEST_ID = new ThreadLocal<>(); 
    private static final ScopedValue<String> REQUEST_ID = ScopedValue.newInstance(); 
 
    static void handle(String requestId) {
        REQUEST_ID.set(requestId); 
        try { 
            processOrder(); 
        } finally { 
            REQUEST_ID.remove(); 
        } 
        ScopedValue.where(REQUEST_ID, requestId).run(RequestContext::processOrder); 
    }
 
    static void processOrder() {
        // Everything called from here sees the value, however deep the call stack
        System.out.println("Processing order for request " + REQUEST_ID.get());
    }
 
    public static void main(String[] args) throws InterruptedException {
        Thread.ofVirtual().start(() -> handle("req-42")).join();
    }
}

processOrder doesn't change, and there is no remove() left to forget. On JDK 21 through 24, enable preview features to run it, for example java --enable-preview --source 21 RequestContext.java on JDK 21. Subtasks forked in a StructuredTaskScope inherit scoped values, but that API is still in preview as of JDK 26 and was reworked in JDK 25, so I'd wait for it to become final.

Turning on virtual threads in Spring Boot

In Spring Boot 3.2 or later, on Java 21 or later, it's one property. I set a few related ones with it:

src/main/resources/application.propertiesProperties
spring.threads.virtual.enabled=true
# Virtual threads are daemon threads: keep the JVM alive when nothing else does
spring.main.keep-alive=true
# The connection pool is now the database's concurrency limit: size it on purpose
spring.datasource.hikari.maximum-pool-size=20
spring.datasource.hikari.connection-timeout=2000

The switch changes several things at once:

  • Tomcat and Jetty process each request on a virtual thread.
  • The applicationTaskExecutor used by @Async becomes a SimpleAsyncTaskExecutor that starts a virtual thread per task.
  • @Scheduled methods run on a SimpleAsyncTaskScheduler backed by virtual threads.
  • Integrations such as the RabbitMQ and Kafka listener containers get virtual-thread executors.
  • Thread-pool properties such as server.tomcat.threads.max stop having any effect.

To check, log Thread.currentThread() in a controller; the output should start with VirtualThread. The Spring Boot reference requires Java 21 and strongly recommends Java 24 or later.

Virtual threads, thread pools, or reactive?

Here is how I choose:

WorkloadMy defaultWhy
Handlers that mostly wait on databases and servicesVirtual threadsCheap blocking, plain code
Fanning out many I/O callsVirtual threads plus a Semaphore per downstreamCheap concurrency with explicit limits
CPU-heavy work such as hashing or image processingA fixed platform pool sized to your coresCores cap parallelism; the OS time-slices platform threads
Blocking in synchronized code you can't change, on JDK 21 to 23JDK 24 or later, or platform threads for that pathPinning can starve the carriers
Long blocking calls into native codePlatform threadsNative frames still pin on every JDK
Streams that need end-to-end backpressureReactive, such as Project ReactorBackpressure and stream operators are the point
A reactive codebase that meets its targetsKeep it reactiveA rewrite buys little

For a new blocking service on JDK 25, I default to virtual threads with an explicit limit in front of every downstream, and keep reactive code for streams.

Key takeaways

  • Virtual threads make waiting cheap, not code faster: use them for I/O-bound work and keep CPU-heavy work on a core-sized pool.
  • Create one per task with Executors.newVirtualThreadPerTaskExecutor(), never pool them, and protect scarce resources with a Semaphore or a deliberately sized connection pool.
  • Run JDK 24 or later, ideally JDK 25, so synchronized no longer pins, and use the jdk.VirtualThreadPinned JFR event to find what still does.
  • Don't cache expensive objects in ThreadLocal; carry request context in a ScopedValue, final since JDK 25.
  • In Spring Boot 3.2 or later, spring.threads.virtual.enabled=true moves request handling, @Async, and @Scheduled onto virtual threads; add spring.main.keep-alive=true and revisit your pool sizes.

Virtual threads don't remove the need to think about concurrency; they move the question from how many threads you can afford to how much load each downstream can take. Turn them on in staging with a JFR recording running, and let the pinned-thread events and your pool metrics tell you what to fix next.