Managing Distributed Workflows: Orchestration vs. Choreography
Below are my personal notes on Chapter 11, “Managing Distributed Workflows,” from Software Architecture: The Hard Parts. I highly recommend reading the book if you haven’t already.
A business workflow often spans multiple domains and operations that inherently depend on each other. Once these operations are distributed across different services, architects must decide how to preserve these relationships without introducing unnecessary coupling.
The challenge is not eliminating coupling, since some coupling is inherent in the problem domain, but choosing how that coupling is implemented across distributed components.
Semantic coupling is unavoidable; implementation coupling is a design choice.
Every business workflow has some inherent coupling imposed by the problem domain. If one operation depends on another, that relationship exists regardless of how the system is designed.
DDD and bounded contexts can help identify and define domain boundaries, but they cannot eliminate semantic coupling between those domains. An architect cannot reduce the semantic coupling inherent in the problem through implementation choices; they can only avoid introducing unnecessary implementation coupling or make it worse through poor design choices.
This distinction becomes especially important in distributed workflows. When a semantically coupled business process is distributed across multiple services, architects must decide how those inherent relationships should be implemented.
Three main forces contribute to this implementation coupling:
- Communication: how services exchange information.
- Consistency: how consistent their state needs to be across the workflow.
- Coordination: how the individual steps of the workflow are managed.
Managing distributed workflows is therefore largely about implementing the unavoidable semantic coupling of a business process across distributed components without introducing unnecessary implementation coupling.
Two common approaches to workflow coordination are orchestration and choreography.

Orchestration Communication Style
This communication style introduces an orchestrator/mediator to manage a distributed workflow. The orchestrator does not contain the domain behavior of the individual services it coordinates. Its responsibility is to coordinate those services, Its sole responsibility is coordinating the workflow, while domain behavior remains within the appropriate bounded contexts.
It may not seem as useful when considering only the happy path, since its usefulness becomes more apparent as semantic coupling and workflow complexity increase, particularly around error handling.
While the orchestrator/mediator is responsible for communication and coordination, each service owns its transactional state. When an error occurs, the orchestrator coordinates the necessary recovery or compensating actions using the existing communication paths, without introducing new communication paths between the participating services.
Since the orchestrator owns the workflow, it provides several advantages as workflow complexity increases: it centralizes workflow state and behavior, coordinates error handling and compensating actions, and can recover from temporary outages through mechanisms such as retries.
As workflow ownership is centralized, orchestration also introduces trade-offs: the orchestrator can become a bottleneck or single point of failure unless designed for resilience and horizontal scaling. Centralized coordination can also reduce potential parallelism and increase coupling between the orchestrator and the participating services.
Choreography Communication Style
In this communication pattern, the initiating request goes to the first service in the chain of responsibility. Each service performs its responsibility and sends a message that triggers the next participant in the workflow. Since there is no centralized coordinator, workflow state such as processing status, error states, retry counts, and compensation state may need to be tracked across the participating services.
At first glance, choreography seems straightforward. However, complexity increases when considering error cases. Unlike orchestration, handling these errors may require additional communication paths between services, causing services to interact directly with each other.
Since workflow state must be managed somehow, choreography commonly uses patterns such as Front Controller, Stateless Choreography, and Stamp Coupling.
- Front Controller: A dedicated component tracks the overall workflow state while services still coordinate through choreography. It simplifies state tracking but introduces a centralized component and additional coupling.
- Stateless Choreography: No component explicitly tracks the overall workflow state; each service reacts to incoming messages and performs its part. It’s simple but becomes difficult to manage as workflow and error-handling complexity grows.
- Stamp Coupling: Workflow state is carried along with the messages exchanged between services. This avoids centralized state storage but increases message size and couples services to the shared workflow-state structure.
Each pattern manages or propagates workflow state differently and introduces its own trade-offs in coupling, complexity, and state ownership.
Since there is no centralized coordinator, choreography provides more opportunities for parallelism, allows services to scale independently, avoids a central point of failure, and reduces coupling to an orchestrator. However, distributed coordination comes with trade-offs. Without a workflow owner, managing workflow state, error and boundary conditions becomes harder. Services need more knowledge about the overall workflow, increasing the complexity of error handling and making recovery from failures or temporary outages more difficult.
Conclusion
Neither orchestration nor choreography should always be chosen over the other. Choosing between them requires looking beyond the happy path and understanding the semantic coupling inherent in the business workflow, including its error and boundary conditions.
We cannot change the inherent coupling imposed by the business domain, but we can control the implementation coupling introduced by our architectural decisions.
As workflow complexity grows, these trade-offs become increasingly important. Our responsibility as architects is to understand the primary forces creating that complexity and choose the communication style that best fits the workflow without introducing unnecessary implementation coupling.
Originally published in Level Up Coding on Medium.