Architecting with LLMs: use plain old code when you can, LLMs when you must

generative AI   genAI development   architecting   harness engineering   idea to code   decision making   anti-patterns   pattern  

Contact me for information about consulting and training at your company.

The MEAP for Microservices Patterns 2nd edition is now available


I recently wrote about the Coding agent sandwich, which is a design pattern for a coding agent harness. The pattern consists of a tasty filling of LLM invocations sandwiched between two slices of plain old code (POC). Its practical rule is: use POC when you can, LLMs when you must.

The following diagram shows how the pattern’s worked example, the implement plan workflow, applies this rule. The workflow is decomposed into a tree of responsibilities. Each responsibility is implemented using either POC (gray) or an LLM (yellow). LLMs are used only for the responsibilities that require judgment, such as writing code and fixing a broken build. The workflow has orchestration at two levels. POC orchestrates the overall workflow, and the LLM agent that implements each task decides which action to take next within that task.

In this article, I apply that rule and the other ideas from the Coding agent sandwich pattern to architecting systems that leverage LLMs. Let’s start by looking at the timeless idea of the right tool for the right job.

The right tool for the right job

As with many new technologies, GenAI agents have attracted tremendous hype. This naturally leads to the Golden Hammer anti-pattern: the tendency to apply a shiny tool to every problem, regardless of whether it’s the best fit. As the hype settles, it becomes clearer where agents fit and where POC is the better choice.

When to use POC

POC is the right choice when you can specify how to compute the answer. This is the case when the rules are known and correctness can be precisely defined. Don’t waste your time trying to solve a problem with an LLM when a far simpler solution exists.

Beyond being the right fit for these problems, POC has the benefit of repeatability. Given the same inputs and state, POC behaves the same way every time, while an LLM is probabilistic. Repeatability also makes POC easier to develop and debug because failures are reproducible. There’s no need to experiment with different prompts and hope that the LLM does the right thing. And at runtime, POC is typically much faster and cheaper than an LLM invocation.

When to use an LLM

An LLM is a good candidate when you can specify what a good answer looks like but cannot reasonably specify how to compute it. This is the case when the input is ambiguous or unstructured, when understanding natural language is required, or when producing the answer requires judgment or interpretation. For such problems, enumerating deterministic rules is impractical.

Here are some real-world examples where an LLM is a good candidate:

  • Interpreting an ambiguous operational support request
  • Translating a user’s business intent into an enterprise transaction
  • Interpreting complex contractual or engineering requirements
  • Implementing a software change from a natural-language specification
  • Diagnosing an unfamiliar production failure

These capabilities come at a price. An LLM invocation is slow and expensive compared to executing POC. And because an LLM is probabilistic, repeated invocations with the same input can behave differently. So before you commit to an LLM, test it against representative examples to check that its error rate, cost, and latency are acceptable.

Start with functional decomposition

Designing a system that uses LLMs starts with recursive functional decomposition. You identify the responsibilities needed to solve the problem. These responsibilities include orchestration, which is the control flow that decides which responsibility to execute next.

The decomposition is hierarchical. A responsibility X can itself be decomposed into sub-responsibilities X1 and X2, together with a coordinator that orchestrates them.

For example, consider the implement plan workflow that is the worked example in the Coding agent sandwich article. Its job is to turn each task in a development plan into a Git commit. Decomposing this problem identifies seven responsibilities:

  • Getting the next task from the plan
  • Implementing the task
  • Validating the result
  • Pushing the commit and creating a pull request
  • Waiting for the CI build
  • Fixing the CI build if it fails
  • Addressing pull request review feedback

Implementing the task is itself decomposed into writing a test, writing the code, running the tests, marking the task as complete, and committing the changes. The orchestration responsibility is the loop that executes these steps for each task until no tasks remain.

For each responsibility: POC or LLM?

The next step is to decide, for each responsibility, whether to implement it using POC or an LLM. As described earlier, use POC when you can specify how to compute the answer, and use an LLM when you can only specify what a good answer looks like.

When assigning a responsibility to an LLM, keep the responsibility narrowly defined. For example, the implement plan workflow invokes a coding agent for each task in the plan, instead of using one agent run to implement the whole plan. An LLM invocation or agent run that handles many responsibilities is less reliable. The LLM might, for example, stop partway through. A large invocation or run is also wasteful because it uses an expensive LLM for work that POC could do far more cheaply.

Orchestration: POC or LLM?

When a system contains an LLM, it’s tempting to assume that the LLM should orchestrate it. But orchestration is itself a responsibility, so the same POC-or-LLM decision applies to it. There are two options:

  • POC orchestration - the right choice when you can define the steps and the branches between them in advance. The POC orchestrator decides which responsibility to execute next. When a branch requires judgment, the orchestrator can invoke an LLM to choose one of the branches that it defines, and then execute it. Making a decision is simply another narrowly defined responsibility that is implemented using an LLM.
  • LLM-based orchestration - the right choice when you cannot define the steps in advance. The LLM examines the current state and the results so far, and decides what to invoke next.

The implement plan workflow uses POC orchestration, which forms the top slice of the Coding agent sandwich.

An agent orchestrates and does the work

When an LLM decides each next action after seeing the result of the previous one, the responsibility is implemented as an agent. An agent consists of an LLM, a harness, and a set of tools. The LLM also decides when it has finished. The harness, such as Claude Code, is POC that runs the loop. It invokes the LLM, executes each tool call that the LLM requests, and passes the result back to the LLM.

The LLM in an agent does more than orchestrate. It also performs some of the responsibilities that it orchestrates and delegates the rest to tools. For example, the coding agent that implements a task writes the tests and the code itself, and it runs the tests using Gradle. The coding agent forms the filling of the sandwich.

Give LLMs POC tools

A responsibility implemented using an LLM can itself be decomposed. When it needs to act on the world, such as running tests, querying a database, or invoking an API, each action is a sub-responsibility whose rules are known. Implement each one using POC as a tool with well-defined behavior that the LLM invokes. For lessons on designing such tools, see Migrating a Claude Code skill from WebFetch to the new CircleCI CLI.

For example, the coding agent that implements a task invokes Gradle to run the tests and Git to commit the changes. The LLM decides what to do next based on the actual test results. These tools form the bottom slice of the sandwich.

LLMs often belong at the boundaries

Applying the POC-or-LLM rule to each responsibility tends to put LLMs at the boundaries of a system. At a boundary, the system receives unstructured, ambiguous input from the real world, which must be interpreted. Once the input has been interpreted, it becomes structured state that well-defined operations can act on. Those operations are implemented using POC.

For example, consider a system that handles operational support requests. Each request is written in natural language and is often ambiguous. An LLM interprets the request and converts it into a structured ticket, which specifies the category, the priority, and the affected service. A well-formed ticket can still be wrong, because the LLM might misinterpret the request. POC therefore validates the ticket, for example by checking that the affected service exists, and sends a ticket that fails validation to a person for review. POC then routes each valid ticket and updates the ticketing system. Validation can catch a ticket that names a service that doesn’t exist, but it can’t catch one that names the wrong service.

Not every LLM responsibility is at a boundary. Each LLM responsibility in the implement plan workflow interprets some input, such as a task description, the logs of a failed build, or review comments. But most of its work is writing code, which requires judgment.

Revisiting the POC-or-LLM decision

The decision to implement a responsibility using POC or an LLM can change after the initial implementation. When you first implement a responsibility, you might not understand the problem well enough to specify how to compute the answer, so you use an LLM. As your understanding improves, you might find that you can specify the computation after all. Reimplementing the responsibility using POC then makes it faster, cheaper, and repeatable.

The change can also go the other way. A responsibility that you implemented using POC might turn out to require judgment. For example, its input might be more varied than you expected, so you keep adding rules to handle new cases. In that case, reimplementing the responsibility using an LLM might be simpler than adding more rules.

In either direction, compare the two implementations before switching. Run both against the same representative inputs, and check their results, cost, and latency against the same acceptance criteria.

Summary

The key idea of this article is to design a system that uses LLMs by applying functional decomposition. You decompose the problem into responsibilities, including orchestration, and then decide how to implement each one. To avoid the Golden Hammer anti-pattern, make that decision separately for each responsibility. Use POC whenever you can specify how to compute the answer. Use an LLM only when you must, because you can only specify what a good answer looks like.

Need help with modernizing your architecture?

I help organizations modernize their architecture to enable fast flow and GenAI-powered software delivery. If you’re planning or struggling with a modernization effort, I can help.

Learn more about my modernization and architecture advisory work →


generative AI   genAI development   architecting   harness engineering   idea to code   decision making   anti-patterns   pattern  


Copyright © 2026 Chris Richardson • All rights reserved

About Microservices.io

Microservices.io is created by Chris Richardson, software architect, creator of the original CloudFoundry.com, and author of Microservices Patterns. Chris helps organizations modernize their architecture to enable fast flow and GenAI-powered software delivery.

Need help modernizing your architecture?

Avoid the trap of creating a modern legacy system — a new architecture with the same old problems.
Contact me to discuss your modernization goals.

Get Help

Microservices Patterns, 2nd edition

I am very excited to announce that the MEAP for the second edition of my book, Microservices Patterns is now available!

Learn more

ASK CHRIS

?

Got a question about microservices?

Fill in this form. If I can, I'll write a blog post that answers your question.

NEED HELP?

I help organizations improve agility and competitiveness through better software architecture.

Learn more about my consulting engagements, and training workshops.

LEARN about microservices

Chris offers numerous other resources for learning the microservice architecture.

Get the book: Microservices Patterns

Read Chris Richardson's book:

Example microservices applications

Want to see an example? Check out Chris Richardson's example applications. See code

Virtual bootcamp: Distributed data patterns in a microservice architecture

My virtual bootcamp, distributed data patterns in a microservice architecture, is now open for enrollment!

It covers the key distributed data management patterns including Saga, API Composition, and CQRS.

It consists of video lectures, code labs, and a weekly ask-me-anything video conference repeated in multiple timezones.

The regular price is $395/person but use coupon OFFEFKCW to sign up for $95 (valid until Sept 30th, 2025). There are deeper discounts for buying multiple seats.

Learn more

Learn how to create a service template and microservice chassis

Take a look at my Manning LiveProject that teaches you how to develop a service template and microservice chassis.

Signup for the newsletter


BUILD microservices

Ready to start using the microservice architecture?

Consulting services

Engage Chris to create a microservices adoption roadmap and help you define your microservice architecture,


The Eventuate platform

Use the Eventuate.io platform to tackle distributed data management challenges in your microservices architecture.

Eventuate is Chris's latest startup. It makes it easy to use the Saga pattern to manage transactions and the CQRS pattern to implement queries.


Join the microservices google group