Building Microservices at 130 Million Requests Per Day
How we architected the AirAsia Move platform for massive scale — circuit breakers, distributed tracing, Kafka event flows, and hard lessons from production.
When your platform processes over 130 million requests every single day, the margin for error is measured in milliseconds and the cost of a bad deployment is measured in lost bookings. I've spent the last three-plus years as a Senior Software Engineer at AirAsia's Technology Centre in Bengaluru, building and operating the microservices behind AirAsia Move — the travel super-app that handles flights, hotels, and ancillary services for millions of travelers across Southeast Asia. This is a post about what that actually looks like from the inside: the architecture decisions that worked, the ones that didn't, and the production incidents that taught me more than any design document ever could.
The Scale Problem
AirAsia Move isn't a single monolithic booking engine. It's an OTA (Online Travel Agency) platform composed of dozens of independently deployed microservices. On an average day, the system handles north of 130 million requests — roughly 1,500 requests per second sustained, with spikes well above that during flash sales and holiday booking windows.
I own services on the Manage My Booking (MMB) platform. That means flight changes, fare summaries, price-slash features, and post-booking ancillaries like seat selection and baggage add-ons. Every one of these flows involves calls to multiple upstream services — airline inventory, payment gateways, pricing engines, notification systems. A single user action like changing a flight can fan out to eight or nine downstream calls before a response comes back.
At this scale, the engineering problems aren't about whether your code compiles or whether your unit tests pass. They're about what happens when one of those eight downstream services is 200 milliseconds slower than usual. Or when an upstream dependency starts returning 503s during a deployment. Or when a Kafka consumer group rebalances at exactly the wrong moment.
Architecture Decisions That Held Up
Circuit Breakers with Resilience4j
The first pattern that proved its weight in gold was the circuit breaker. We use Resilience4j across all our Spring Boot services. The idea is straightforward: if a downstream service starts failing, stop calling it for a while instead of piling up timeouts that cascade through the entire system.
Here's a stripped-down version of how we configure a circuit breaker for an upstream pricing service:
@Bean
public CircuitBreakerConfig pricingCircuitBreakerConfig() {
return CircuitBreakerConfig.custom()
.failureRateThreshold(50)
.waitDurationInOpenState(Duration.ofSeconds(30))
.slidingWindowType(SlidingWindowType.COUNT_BASED)
.slidingWindowSize(20)
.minimumNumberOfCalls(10)
.build();
}
The numbers matter. A 50% failure threshold over a window of 20 calls means the breaker trips after 10 failures out of 20 requests. The 30-second open-state duration gives the downstream service time to recover before we start probing again. We tuned these values through load testing with JMeter — the defaults were too aggressive for our traffic pattern and tripped breakers during normal latency variance.
We pair circuit breakers with fallback methods. For the fare-summary endpoint, if the pricing service is down, we return the last-cached fare with a flag indicating staleness. The user still sees a price — it might be a few minutes old, but that's better than a blank screen or a 500 error.
@CircuitBreaker(name = "pricingService", fallbackMethod = "cachedFareFallback")
public FareSummary getFareSummary(String bookingId) {
return pricingClient.fetchFare(bookingId);
}
private FareSummary cachedFareFallback(String bookingId, Throwable t) {
log.warn("Pricing service unavailable for booking {}, using cache", bookingId);
return fareCache.getLastKnown(bookingId)
.map(fare -> fare.withStaleFlag(true))
.orElseThrow(() -> new ServiceUnavailableException("No cached fare available"));
}
Distributed Tracing with Sleuth and Zipkin
When a request crosses ten services before returning, debugging a latency spike without distributed tracing is like searching for a specific grain of sand on a beach. We use Spring Cloud Sleuth for trace propagation and Zipkin for visualization.
Every inbound request gets a trace ID. That ID propagates through every downstream HTTP call, every Kafka message, every async thread. When a user reports that their flight change took 14 seconds instead of the usual 2, I can pull up the trace ID and see exactly which service call ate the extra 12 seconds.
The configuration is minimal — Sleuth auto-instruments most of the Spring Cloud stack:
spring:
sleuth:
sampler:
probability: 0.1
zipkin:
base-url: http://zipkin-collector:9411
We sample 10% of traces in production. Sampling everything would drown Zipkin, and at 1,500 RPS, 10% still gives us 150 traces per second — more than enough to spot patterns. For specific debugging, we can force a trace by injecting a header.
The biggest win from tracing wasn't debugging individual requests — it was identifying systemic patterns. We noticed that every Tuesday between 2 AM and 4 AM UTC, latency spiked on the inventory service. Turned out a batch job was running full table scans on the same database the API was reading from. We moved the batch job to a read replica and the Tuesday spikes disappeared.
Kafka for Event-Driven Communication
Not everything needs a synchronous HTTP call. When a user completes a flight change, we need to update the booking record, send a confirmation email, recalculate loyalty points, and push a notification to the mobile app. None of those secondary actions need to block the user's response.
We use Apache Kafka as the backbone for async event flows. The booking-change service publishes a BookingChanged event, and downstream consumers — notification service, loyalty engine, analytics pipeline — each process it independently.
@KafkaListener(topics = "booking-changes", groupId = "notification-service")
public void handleBookingChange(BookingChangedEvent event) {
notificationService.sendConfirmation(event.getBookingId(), event.getPassengerEmail());
}
Kafka's partitioning model gives us parallelism for free. We partition by booking ID, which guarantees that all events for a single booking land on the same partition and get processed in order. With 12 partitions and 12 consumers in the notification group, we handle the event volume without breaking a sweat.
One hard-won lesson: always set explicit retention policies. We had a topic with the default 7-day retention that grew to 400 GB because no one noticed the consumer group was lagging by three days. That ate into disk on the Kafka brokers and started affecting other topics.
Production War Stories
The Memory Leak That Only Showed Up Under Load
Six months in, our flight-change service started getting OOMKilled by Kubernetes roughly once every 48 hours. Under normal traffic, heap usage was fine. Under sustained load, it crept up and never came back down.
I spent two days profiling with JVisualVM connected to a staging pod running production-mirrored traffic. The culprit was a ConcurrentHashMap used as an in-memory cache that had no eviction policy. Every unique booking ID that came through got cached, and the entries never expired. Under low traffic, the map stayed small. Under sustained 1,500 RPS with unique booking IDs, it grew until the JVM ran out of heap.
The fix was embarrassingly simple: swap the raw ConcurrentHashMap for a Caffeine cache with a TTL and max-size cap. Five lines of code, two days of investigation.
Cache<String, FareSnapshot> fareCache = Caffeine.newBuilder()
.maximumSize(50_000)
.expireAfterWrite(Duration.ofMinutes(10))
.build();
The Cascading Timeout
During a Diwali sale, the payment gateway started responding 300ms slower than usual — not enough to fail, just enough to be annoying. But our service had a 5-second timeout on payment calls, and the circuit breaker was configured with a count-based window. At peak traffic, the slower responses backed up the thread pool. Incoming requests started queuing. Upstream services calling us started timing out. Within 90 seconds, three services were effectively down.
The root cause wasn't the payment gateway being slow. It was our thread pool configuration. We were using a fixed thread pool of 200 threads with a bounded queue of 500. When every thread was blocked waiting on a 5-second payment call, the queue filled up and new requests got rejected.
We made three changes after that incident:
- Reduced the payment timeout from 5 seconds to 2 seconds. If payment hasn't responded in 2 seconds during peak load, it's not going to respond usefully.
- Switched to a bulkhead pattern using Resilience4j's Bulkhead, isolating the payment call to its own thread pool so a slow payment service can't starve other operations.
- Added a time-based sliding window to the circuit breaker instead of count-based, so it reacts faster during traffic spikes.
The Consumer Group Rebalance
Kafka consumer group rebalances are one of those things that sound benign in documentation and turn into emergencies in production. We had a deployment that rolled out a new consumer version. Kubernetes did a rolling restart — killing one pod, starting a new one, waiting for readiness, then moving to the next.
Every time a pod died, the consumer group rebalanced. Every rebalance paused consumption for 30-60 seconds. With 12 pods and a rolling restart, we had 12 rebalances over a 20-minute window. During those pauses, the topic lag grew. By the time the deployment finished, we had 2 million unprocessed events. The notification service was sending confirmation emails for flight changes that had happened 45 minutes earlier.
We fixed this by implementing static group membership. Each consumer gets a persistent group.instance.id tied to its pod identity. When a pod restarts, the new instance reclaims the same member ID, and Kafka skips the rebalance if it comes back within the session.timeout.ms window.
What I'd Do Differently
If I were starting this system from scratch today, three things would change.
Contract testing from day one. We caught integration bugs in staging that should have been caught by consumer-driven contract tests. When Service A changes a response field from fare_amount to fareAmount, the unit tests in Service A pass. The unit tests in Service B pass. Staging explodes. Contract tests catch this at build time.
Structured logging earlier. We adopted structured JSON logging eventually, but for the first year, half our logs were unstructured text. Searching for a specific booking ID across 10 services in Kibana was painful when some services logged bookingId=ABC123 and others logged Processing booking ABC123 for user xyz. Structured logging with consistent field names should have been a service template requirement on day one.
More aggressive load shedding. Instead of accepting every request and letting the system degrade under overload, I'd implement load shedding at the gateway level. If the system is at 90% capacity, start returning 429s to lower-priority endpoints (analytics, non-critical ancillaries) while keeping the booking and payment paths fully served.
The Takeaway
Building microservices at this scale isn't about picking the right framework or the right cloud provider. It's about understanding failure modes. Every architectural decision — circuit breakers, tracing, async messaging, bulkheads — exists because something broke in production and we needed it to not break the same way again.
The AirAsia Move platform processes 130 million requests a day not because we got the design right on the first try, but because we built the instrumentation to see what was breaking and the patterns to contain the blast radius when it did.
If you're building distributed systems and want to compare notes, or if you're looking for an engineer who's operated services at this scale, check out my experience and projects — or just reach out directly.
Related Articles
- AI-Driven Development with Claude Code — How I run autonomous agents across this same microservices codebase to accelerate feature delivery.
- From Startup to Acquisition — The early years building distributed systems on a twelve-dollar VPS before AirAsia-scale problems existed.
- OpenBanking PSD2 API Development — Building compliant banking APIs at Oracle — a different kind of scale problem where regulatory correctness outweighs throughput.