The difficult branch is usually a business rule
A workflow retries an order update after a timeout. The external system has already accepted it, but the workflow cannot tell. Adding another node does not answer whether retrying will create a duplicate. That question belongs to the business operation, regardless of which tool draws the arrows.
Our boundary for custom code is not the number of nodes on a canvas. It is whether someone can describe an operation's promises without opening the whole workflow. What does it accept? What changes? What survives a restart? Who owns the result when the caller stops waiting? Those questions reveal a service boundary much earlier than an untidy diagram does.
Keep the coordination where people can see it
Visual workflows remain useful for routing, notifications, schedules and coordination between systems. An operations colleague can inspect the intended sequence and understand where a case stopped. Replacing that visibility with a collection of small services can make maintenance worse, especially when nobody owns their deployment.
Throughput is a separate issue. n8n documents queue mode with worker processes and Redis; needing more execution capacity does not by itself prove that the workflow should be rewritten. The current queue mode documentation describes that infrastructure choice.
Consider an illustrative returns workflow. The canvas can receive the request, gather order details and route the outcome. The decision about whether the return is eligible should have a clearly named owner and a testable contract. Moving that decision out can simplify the workflow without removing its operational value.
Extract a promise, not a messy section
A useful service accepts a business command such as assessing a return, not an arbitrary bundle of workflow variables. Its response distinguishes accepted, rejected and awaiting evidence. It also defines which fields are required and which version of the business rules produced the decision.
Write tests around disagreements: an order is partly refunded, a product has been exchanged, or the customer supplies an old reference. These examples belong to the operation and should remain understandable after the workflow changes. A test suite coupled to every node label simply moves the same fragility into another repository.
The uncomfortable part is ownership. A code service needs someone responsible for releases, security updates, logs and recovery. If the team cannot name that person, extracting code may exchange a visible problem for an invisible one. A disciplined workflow with documented rules can be the better interim design.
Put durable state behind a deliberate boundary
For an operation that changes another system, assign a stable business reference before sending it. Persist the attempt and its outcome. When the result is uncertain, reconcile it with the receiving system instead of assuming a timeout means nothing happened. Design repeated requests to return the existing outcome where the operation allows it.
Keep secrets and access narrow. The coordinating workflow should not need unrestricted database access just because the extracted service does. Logging also needs a shared correlation reference, so support can follow the same case across both sides without collecting complete customer payloads everywhere.
Migration should preserve a way back. Compare decisions in a non-writing mode first, investigate differences, and switch the caller only after the intended behaviour is understood. Never let the old and new paths perform the same external action during comparison.
Make the boundary review this week
Choose the workflow that causes the most uncertainty during recovery. Write its irreversible actions, stored state, retry rules and current owner on a page. Ask an operator to explain a failed run using that page.
Extract only the rule that cannot be tested or owned cleanly in its current location. Define its input, outcomes and recovery procedure before choosing a framework. Keep the surrounding workflow visible. Success is a case that can be explained and safely resumed, not a smaller canvas or a larger codebase.
