← Back to blog
AI & ToolingSeptember 28, 20269 min read

AI-Driven Development with Claude Code: Running Autonomous Agents Across Microservices

Running autonomous parallel agents across a 10+ microservice codebase. Custom permission scopes, Git workflow integration, and measurable delivery gains.

Why I Adopted AI Agents

At AirAsia, I own production microservices on the Manage My Booking platform — flight changes, fare summaries, price-slash, post-booking ancillaries. Over a dozen services, each with its own repository, its own deployment pipeline, its own quirks accumulated over years of high-traffic production use. The platform processes 130 million requests a day. There is no room for sloppy code.

The problem was never writing code. The problem was the overhead surrounding it. Context-switching between services killed my momentum. I would be deep in a fare-calculation refactor, then get pulled into fixing a circuit breaker config three repositories over, then pivot again to review a schema migration for the ancillary service. Every switch cost me 20–30 minutes of mental reload time. Multiply that across a week and you lose entire days to friction.

I started experimenting with Claude Code not because I wanted AI to write code for me, but because I needed a way to keep multiple workstreams moving in parallel without losing my mind. The difference between a chatbot and an autonomous agent is the difference between asking someone for directions and handing someone the keys. Claude Code operates directly in the terminal, reads your codebase, runs commands, writes files, and commits code. It is an agent, not an autocomplete engine.

What I did not expect was how drastically it would change my entire development workflow. Within two months, I had restructured my approach to feature delivery around a three-phase model — Inception, Construction, Operations — with Claude Code agents handling the bulk of Construction while I focused on architecture decisions and code review.

Setting Up the Workflow

Permission Scopes

The first thing I learned: you cannot give an AI agent unrestricted access to a production codebase and walk away. That is how you end up with a force-push to main at 2 AM.

Claude Code supports permission configurations that control exactly what the agent can and cannot do. I set up scoped permissions per project:

{
  "permissions": {
    "allow": [
      "Read",
      "Edit",
      "Write",
      "Bash(npm run test*)",
      "Bash(npm run lint*)",
      "Bash(mvn test*)",
      "Bash(git diff*)",
      "Bash(git status*)",
      "Bash(git log*)",
      "Bash(git add -p*)"
    ],
    "deny": [
      "Bash(git push*)",
      "Bash(git checkout main*)",
      "Bash(rm -rf*)",
      "Bash(docker*)"
    ]
  }
}

The agent can read, write, edit files, run tests and linters, view diffs, and stage changes with interactive add. It cannot push to remote, switch to main, delete directory trees, or touch Docker containers. Every push goes through me. Every merge goes through me. The agent does the work; I do the gatekeeping.

This configuration lives in each project's .claude/settings.json. Different services get different scopes. The payment-adjacent services have tighter restrictions. The internal tooling repos are looser.

Git Integration

The Git workflow was the piece that took the most iteration to get right. Early on, I let the agent commit directly, and the commit messages were fine but the staging was sloppy — it would bundle unrelated changes into a single commit.

I switched to selective staging. The agent runs git add -p to stage hunks interactively, or I configure it to stage specific files. After the agent finishes a task, I review the diff before any commit happens. This gives me an audit trail where each commit maps to a single logical change.

The branch strategy is straightforward:

  1. I create the feature branch and write a brief spec in a markdown file at the repo root
  2. The agent reads the spec, explores the codebase, and implements
  3. I review the working tree diff, adjust staging, and commit
  4. I push and open the PR

The agent never touches main. It never pushes. It never resolves merge conflicts on its own. These are hard boundaries, not suggestions.

Running Parallel Agents Across Services

This is where the real productivity gains show up. At AirAsia, a single feature often touches three to five services. A new ancillary offering might need changes in the booking service, the pricing engine, the ancillary catalog, the notification service, and the API gateway. Traditionally, I would work through these sequentially — finish one, context-switch, start the next.

With Claude Code, I run parallel terminal sessions, each with its own agent instance scoped to a specific repository. My typical setup looks like this:

  • Terminal 1: Agent working on the booking-service, implementing the new endpoint
  • Terminal 2: Agent working on the pricing-engine, adding the fare calculation logic
  • Terminal 3: Agent working on the ancillary-catalog, updating the schema and data layer
  • Me: Reviewing diffs as they come in, running integration tests, handling the architectural decisions that require cross-service awareness

Each agent operates independently in its own repo. They do not communicate with each other. I am the coordination layer. When the pricing engine agent finishes its implementation, I review it, then feed any interface changes (new request/response shapes, updated error codes) to the booking service agent as context for its next iteration.

This is not magic. It is structured delegation. The agents handle the mechanical work — writing service methods, updating DTOs, adding unit tests, wiring dependency injection — while I handle the parts that require understanding the system as a whole.

The Three-Phase Model

I structured my workflow around three phases, borrowed loosely from the Unified Process but adapted for AI-assisted development:

Inception: I do this entirely myself. Define the feature scope, draw the sequence diagrams, identify which services are affected, write the interface contracts. This phase produces a spec document that becomes the agent's input. I spend 30–60 minutes here, and the quality of this phase directly determines how useful the agent will be downstream.

Construction: This is where the agents take over. Each agent gets a spec, a repo, and a set of constraints. They implement, write tests, and stage changes. I review in rolling fashion — checking in on each terminal every 15–20 minutes, providing corrections, approving directions. A feature that used to take me three days of sequential implementation now takes a day of parallel construction plus review.

Operations: Post-merge monitoring, production validation, performance profiling. This phase is still entirely human. I watch the dashboards, check the distributed traces in Zipkin, validate that circuit breakers behave correctly under the new code paths. AI agents have no business in production operations.

Results and Productivity Gains

After six months of running this workflow, here is what I measured:

Feature delivery time dropped by roughly 40%. A cross-service feature that previously took 4–5 days from spec to PR now typically takes 2–3 days. The reduction comes almost entirely from parallel Construction — the agent handles three repos simultaneously instead of me working through them one at a time.

Unit test coverage increased. Not because I mandated higher coverage targets, but because writing tests is exactly the kind of mechanical work that agents handle well. When the agent implements a service method, it writes the corresponding tests as part of the same task. I used to skip edge-case tests when I was tired or rushing. The agent does not get tired.

Code review quality improved. This one was unexpected. When I write code myself, I review it with the same mental model I used to write it — which means I miss my own blind spots. When an agent writes code, I review it with fresh eyes. I catch more issues during review of agent-written code than I do reviewing my own work. The separation between author and reviewer is genuine, not performative.

Commit history became cleaner. Selective staging and single-purpose commits made the Git history actually useful for debugging. git bisect works when each commit is atomic. It does not work when commits are "WIP: stuff" bundles.

The numbers are not perfect. Some features do not parallelize well — tightly coupled changes where Service B literally cannot be implemented until Service A's interface is finalized. For those, the agent still helps, but the gains are smaller. Maybe 15–20% instead of 40%.

Pitfalls and Guardrails

Six months of daily use surfaced a consistent set of failure modes:

The agent optimizes locally. It will write perfectly correct code for the service it can see, while introducing an interface mismatch with a service it cannot see. This is why I am the coordination layer. Cross-service contracts must be defined by a human who holds the full system model in their head.

Generated tests can be tautological. The agent sometimes writes tests that verify the implementation rather than the behavior. A test that mocks every dependency and then asserts that mockService.method() was called exactly once is not testing anything useful. I watch for this pattern and push back when I see it.

Specification quality is the bottleneck. A vague spec produces vague code. "Add error handling to the booking endpoint" gets you generic try-catch blocks. "Return HTTP 409 with error code BOOKING_CONFLICT when the fare class is no longer available, include the current available fare classes in the response body, and emit a Kafka event to the pricing-update topic" gets you exactly what you need. I spend more time writing specs now than I did writing code before, and the tradeoff is worth it.

Do not let the agent refactor while implementing. If I ask for a new feature and the agent decides the existing code structure needs cleaning up first, I end up with a diff that mixes new functionality with refactoring. These are impossible to review properly. I set explicit boundaries: implement the feature against the existing structure, and I will open a separate refactoring PR if needed.

Human-in-the-loop checkpoints are not optional. Every 15–20 minutes, I check each agent's progress. Not because the agent will break something catastrophic — the permission scopes prevent that — but because catching a wrong direction early saves an hour of wasted work. A five-minute course correction at the 20-minute mark beats throwing away 90 minutes of implementation.

What's Next

I am currently experimenting with structured audit trails — having each agent produce a brief log of decisions made and alternatives considered during implementation. This gives me better context during review and creates documentation that survives after the PR is merged.

I am also looking at cross-agent context sharing. Right now, if the pricing engine agent discovers that an internal API returns a different response shape than documented, I manually relay that information to the booking service agent. Automating that relay — through shared context files or a coordination protocol — is the next step.

The broader pattern here is not about AI replacing developers. It is about restructuring the development workflow so that human attention goes where it matters most: architecture, system-level reasoning, production reliability. The mechanical act of translating a well-defined spec into working code with tests is exactly the kind of work that agents handle well. And at the scale we operate at AirAsia — 130 million daily requests, dozens of interconnected services — that mechanical work was consuming a disproportionate amount of my time.

If you want to see the rest of what I have built with this approach, check out my projects. And if you are running a distributed system team that could use this kind of workflow optimization, I would be glad to talk — reach out through my contact page.


Related Articles

Krishna Kumar Yadav

Senior Software Engineer building distributed systems at scale. 9+ years across fintech, airline tech, and startups.

← Read more articles