← Back to blog
JavaOctober 5, 202610 min read

Java Virtual Threads in Production: How We Replaced 200-Thread Pools with One Line of Config

A hands-on account of migrating Spring Boot microservices from platform threads to Java 21 virtual threads — what broke, what improved, and the patterns we settled on after three months in production.

I've spent an embarrassing number of hours in my career staring at Grafana dashboards, watching thread pool utilization creep toward 100%, and then arguing with my team about whether we should increase maxThreads from 200 to 300 or just spin up another pod. That entire category of work — the thread pool tuning spreadsheets, the load test reports, the "optimal pool size for Service X" wiki pages — became irrelevant the month we moved to virtual threads.

This isn't a tutorial stitched together from the JEP spec. I ran this migration across multiple Spring Boot services handling real traffic in a system I describe in my experience section. Three months later, here's what actually happened — the parts that worked immediately, the parts that caught fire, and the patterns we landed on.

Why Thread Pools Were Eating Our Lunch

Picture a typical booking service. A user hits the "change flight" endpoint. That single request fans out to an airline inventory check, a fare recalculation call, a payment hold verification, and a notification dispatch. Each of those downstream calls takes anywhere from 40ms to 200ms. The thread handling that request sits idle for the vast majority of its life, doing nothing but waiting for bytes to come back over the network.

Multiply that by a thousand concurrent users and you've got a thousand OS threads burning through roughly a gigabyte of stack memory, with each thread spending 85% of its time parked on I/O. The CPU is bored. The memory is full. And you're paying for four pods when the actual compute work could fit on one.

We tried the usual fixes over the years. Bumped thread counts. Tuned connection timeouts. Introduced async patterns with CompletableFuture in the hot paths. Added circuit breakers so slow downstreams didn't hog threads forever (I wrote about that in the microservices post). Each fix helped incrementally but never addressed the root issue: one request equals one OS thread, and OS threads are heavy.

// Every service had its own version of this ritual
@Bean
public WebServerFactoryCustomizer<TomcatServletWebServerFactory> threadPoolTuning() {
    return factory -> factory.addConnectorCustomizers(connector -> {
        var handler = (AbstractProtocol<?>) connector.getProtocolHandler();
        handler.setMaxThreads(350);     // bumped from 200 after the March incident
        handler.setMinSpareThreads(50); // cargo-culted from Stack Overflow circa 2019
        handler.setConnectionTimeout(8000);
    });
}
// TODO: revisit these numbers next quarter (this TODO has been here for two years)

Eleven services. Eleven slightly different configurations. Eleven engineers who each had a different opinion about the "right" pool size.

The Migration Itself

Spring Boot 3.2 shipped with virtual thread support. The configuration is almost offensively simple:

spring:
  threads:
    virtual:
      enabled: true

That property tells Tomcat to hand each incoming request to a virtual thread instead of pulling a platform thread from the pool. Virtual threads are managed by the JVM, not the OS. They start with a few hundred bytes of stack instead of the default 1MB. When they block on I/O, the JVM parks them internally and frees the underlying carrier thread to run something else. No OS context switch. No wasted memory.

The mental model shift matters more than the config change. With platform threads, you think in terms of pool capacity — "this service can handle 200 concurrent requests before queuing." With virtual threads, you stop thinking about thread capacity entirely. The JVM will spin up as many virtual threads as you have requests. Ten thousand concurrent connections? Ten thousand virtual threads, each costing a few kilobytes. The bottleneck moves downstream — to your database connection pool, your HTTP client pool, your downstream service capacity.

What Actually Changed in Production

We didn't flip the switch everywhere at once. Started with a read-heavy service that aggregated fare data from three upstream providers. Low risk — if it degraded, users would see stale prices for a few seconds, not lose bookings.

Here's what the numbers looked like after a week of side-by-side comparison in staging, running identical traffic through both configurations:

MetricPlatform Threads (200 pool)Virtual ThreadsDelta
Requests/sec at 1K concurrent users~820~3,900+375%
p99 latency at 1K concurrent2.3 seconds210ms-91%
Heap + thread stack memory~1.7 GB~480 MB-72%
Requests/sec at 5K concurrent users~500 (severe queuing)~3,700+640%
p99 latency at 5K concurrentMost requests timed out290msNight and day

The 5K concurrent test was the eye-opener. Platform threads simply couldn't keep up — the pool was saturated, requests queued behind each other, timeouts cascaded. Virtual threads handled it with barely a latency bump because there was no pool to saturate. Every request got its own thread instantly.

Memory savings translated directly to pod count. A service that needed four pods at 2GB each (to leave headroom for thread stacks under load) now ran comfortably on two pods at 1GB. That math repeated across ten services.

Three Things That Bit Us Hard

The happy path was easy. The exceptions were instructive.

1. Synchronized Blocks Turned Into Landmines

We had a legacy rate limiter that wrapped its counter updates in synchronized. Looked harmless — it had worked fine for years. Under virtual threads, it became a chokepoint.

When a virtual thread enters a synchronized block and then does something that would normally cause it to yield (like a blocking call), it can't. The virtual thread gets pinned to its carrier thread. That carrier is now stuck until the synchronized block exits. With hundreds of virtual threads hitting the same synchronized method, the few carrier threads (typically equal to your CPU core count) all get pinned, and your service effectively freezes.

// This worked fine with platform threads. Broke everything with virtual threads.
public class RateLimiter {
    private int counter = 0;

    public synchronized boolean tryAcquire() {
        if (counter >= MAX_RATE) {
            waitForNextWindow(); // BLOCKS inside synchronized = carrier thread pinned
            counter = 0;
        }
        counter++;
        return true;
    }
}

// Fixed version — ReentrantLock cooperates with virtual thread scheduling
public class RateLimiter {
    private int counter = 0;
    private final ReentrantLock lock = new ReentrantLock();

    public boolean tryAcquire() {
        lock.lock();
        try {
            if (counter >= MAX_RATE) {
                waitForNextWindow(); // lock.lock() lets the virtual thread unmount cleanly
                counter = 0;
            }
            counter++;
            return true;
        } finally {
            lock.unlock();
        }
    }
}

The JVM gives you a detection tool: -Djdk.tracePinnedThreads=short prints a warning every time a virtual thread gets pinned. We ran every service with this flag in staging for a full week before enabling virtual threads in production. Found pinning in our own code (fixable), in a JDBC driver wrapper (fixed by upgrading), and in one payment SDK (unfixable — that service stayed on platform threads).

2. Connection Pool Starvation

This one was subtle. With platform threads, 200 threads could issue at most 200 concurrent database queries. Our HikariCP pool of 20 connections meant at most 20 queries ran simultaneously, and the other 180 threads waited their turn. The math worked out — connection wait times stayed low because demand was capped by the thread pool.

Virtual threads blew that ceiling off. Suddenly 3,000 concurrent requests could all try to grab a database connection at the same time. Twenty connections serving 3,000 requesters means long waits and potential timeouts.

The fix wasn't complicated, but it required rethinking the pool size based on what the database could handle rather than how many threads the application had:

spring:
  datasource:
    hikari:
      maximum-pool-size: 40    # sized for the database, not the thread pool
      connection-timeout: 3000 # fail fast — don't let virtual threads pile up waiting
      minimum-idle: 15

We also added a Semaphore in one service that was particularly bursty, to cap concurrent database access independently of the connection pool:

private final Semaphore dbThrottle = new Semaphore(60);

public FareResult lookupFare(String flightId) {
    dbThrottle.acquire();
    try {
        return fareRepository.findByFlightId(flightId);
    } finally {
        dbThrottle.release();
    }
}

3. ThreadLocal Initialization Costs

Platform threads live in a pool — they're created once and reused for thousands of requests. If a ThreadLocal does expensive initialization (loading certificates, building a parser, allocating a large buffer), that cost is paid once per thread and amortized across the thread's lifetime.

Virtual threads are created per-request and discarded. Every request pays the full ThreadLocal initialization cost. We found one ThreadLocal<SSLContext> that was rebuilding a TLS context on every single request — something that took ~15ms and had been invisible when it happened once per platform thread.

// Invisible cost on platform threads, 15ms per request on virtual threads
private static final ThreadLocal<SSLContext> TLS = ThreadLocal.withInitial(() -> {
    return rebuildSSLContext(); // loads keystores, builds trust managers
});

// Fix: build once, share across all threads
private static final SSLContext SHARED_TLS = rebuildSSLContext();

How We Rolled It Out

The playbook we settled on after the first few services:

Week 1 — Scan. Run every service in staging with -Djdk.tracePinnedThreads=short. Grep the logs for pinning warnings. Catalog them: our code vs. third-party, fixable vs. not.

Week 2 — Fix. Replace synchronized with ReentrantLock where we own the code. Upgrade libraries where newer versions fixed pinning. For unfixable third-party pinning, document it and mark the service as "skip for now."

Week 3 — Canary. Enable virtual threads on two low-risk services. Watch latency, error rate, memory, and — critically — HikariCP connection wait times. Compare against the previous week's baseline.

Weeks 4-6 — Rollout. One service per day. Enable, monitor for 24 hours, move to the next. Keep the old thread pool config on a feature flag so a revert is a config change, not a deployment.

Week 7 — Cleanup. Delete the custom Tomcat thread pool beans. Remove the tuning documentation. Update runbooks. Close the quarterly "review thread pool sizes" recurring ticket permanently.

Ten of eleven services migrated. The one holdout uses a payment gateway SDK with synchronized blocks buried deep in its HTTP client internals. The vendor knows about it. We check every release.

When Virtual Threads Are the Wrong Answer

If your service spends its time crunching numbers — running ML inference, processing images, computing risk scores — virtual threads buy you nothing. The bottleneck is CPU, not I/O. Platform threads with a pool sized to your core count are still correct for CPU-bound work.

Services already using reactive patterns (WebFlux with Project Reactor) also don't benefit much. Reactive code never blocks in the first place, so the problem virtual threads solve doesn't exist there. We kept WebFlux in our real-time event streaming pipeline because it's genuinely event-driven — backpressure and stream composition matter more than thread efficiency in that context.

The sweet spot is the vast majority of Spring Boot services: request comes in, call some other services, query a database, assemble a response, send it back. I/O-bound, request-response, synchronous code that blocks. That's where virtual threads turn a 200-thread ceiling into effectively unlimited concurrency.

Three Months Later

The migration was done by week seven. Here's where things stood after three months:

Pod count across the cluster dropped by about 30%. Lower memory per instance meant higher density. Some services that ran four replicas now run two. The cloud bill noticed.

Thread pool tuning as a practice is gone. No more quarterly reviews. No more debates about whether 200 or 350 is the right number. No more production incidents caused by pool exhaustion during traffic spikes. The entire category of operational toil evaporated.

Code got simpler. We removed CompletableFuture chains and @Async wrappers from several services where they existed purely to avoid blocking the thread pool. Straightforward synchronous code that reads top-to-bottom, because blocking is cheap now.

One new problem appeared. With the thread pool no longer acting as a natural throttle, downstream services occasionally saw traffic spikes they hadn't seen before. We added explicit rate limiting at the service mesh level to replace the implicit throttling the thread pool used to provide. A better solution than a thread pool — intentional rather than accidental — but something we hadn't anticipated needing.

Getting Started on Your Own Services

If you're on Java 21+ and Spring Boot 3.2+:

  1. Add spring.threads.virtual.enabled=true to one service
  2. Run staging traffic with -Djdk.tracePinnedThreads=short in JVM args
  3. Fix any pinning warnings — synchronized → ReentrantLock for the common case
  4. Review your HikariCP pool size — it's now the real concurrency bottleneck, not threads
  5. Audit ThreadLocal usage for expensive initialization
  6. Load test at 3-5x your normal concurrent user count — virtual threads can handle it, verify your downstreams can too

The migration was one of the more satisfying infrastructure changes I've done in nine years of building these systems. Modest effort, immediate payoff, and it eliminated an entire class of operational headaches. If your services spend their time waiting on I/O — and most do — virtual threads are worth your next sprint.

Have questions about the migration or running into specific issues? I've been through the sharp edges — reach out and I'll share what I can.


Related Articles

Krishna Kumar Yadav

Senior Software Engineer building distributed systems at scale. 9+ years across fintech, airline tech, and startups.

← Read more articles