Skip to main content

Redesign the Workflow Before You Add Another Tool

A practical operating method for e-commerce teams to simplify work, preserve essential controls and test software against real outcomes.

A hand moves a distinct orange tile through a simplified workflow above a keyboard

Most teams reach for a new tool when the real problem is the workflow around it. Consider an order split across two fulfilment locations: one parcel leaves, another misses cut-off, the dispatch message is wrong, and support must reconstruct the context across four systems.

The failure crosses inventory, fulfilment, customer communication and the promise made at checkout. Adding software may accelerate one task while leaving the end-to-end outcome unchanged.

Start with the work itself: define the outcome, trace a real case, name the controls, decide where human judgment belongs, then test any tool against that redesign.

A five-step workflow redesign method
  1. OutcomeCustomer + P&L
  2. TraceReal cases
  3. ClassifyPurpose + control
  4. GovernHuman + system
  5. DecideBuy, build or no-buy

Start with the operating outcome

“Improve operations” is too broad. Define an outcome that matters to the customer and the P&L, then establish the current baseline before looking at products.

For a delivery exception, that might mean:

Reduce the time from the first reliable failure signal to a corrected customer promise.

Reduce the share of cases reopened because the first resolution was incomplete.

Protect refund accuracy and prevent duplicate replacement orders.

Show who owns the case until the customer and the operational record agree.

The dashboard is not the outcome. Instrumentation may be necessary, but it should measure a decision already made about what good looks like. The ISO process approach starts in the same place: intended outputs, process interactions, ownership, risk and measures.

Map the real work, then test the map

Follow one recent case from the first signal to final resolution. Record the systems touched, data copied, handoffs, decisions, waits, approvals and moments when someone has to reconstruct context. This is a practical current-state trace, not a workshop imagining how the process is supposed to work.

Value-stream mapping is useful here because it makes the complete material and information flow visible and helps a team improve the whole rather than optimise one isolated step. The Lean Enterprise Institute describes the same distinction between end-to-end lead time and local process improvement.

One case is only the start. Validate the map against a normal order, a common exception, a rare high-impact failure, a peak-volume case and a recovery after an upstream system or partner fails. If owned-site, Amazon, marketplace, cross-border or multi-location orders behave differently, they need separate paths or clearly marked variants. GOV.UK experience-mapping guidance similarly recommends gathering several user experiences before consolidating a single map.

Decide what each step is for

Do not jump from “manual” to “automate.” Classify each step first:

  1. Eliminate work that has no remaining customer, commercial or control purpose.
  2. Standardise work whose rules differ only because teams have accumulated local habits.
  3. Integrate information that already exists in an authoritative source but is repeatedly re-keyed or reconciled.
  4. Automate bounded, observable work with understood failure modes.
  5. Keep human judgment where policy, customer context, material risk or irreversible action makes it necessary.

The Toyota Production System makes the sequence explicit: improve the work by hand, remove waste, define abnormalities, then build those controls into machines.

Removal still requires control judgment. An approval may protect refund authority; reconciliation may detect duplicate captures; reporting may provide statutory or audit evidence. Before removing a step, name its control objective, its owner and the evidence that proves the replacement control works.

Unnecessary work
Repeated entry, duplicate approval or reporting with no remaining customer, commercial or control purpose.
Necessary control
A deliberate check that protects money, access, compliance, customer rights or the integrity of the operational record.

Design the human-and-system boundary

“Human in the loop” is not a complete operating design. The reviewer needs the right evidence, authority, time and interface to make a real decision rather than rubber-stamp a queue.

For every automated or agent-assisted step, define the operating boundary explicitly:

What the system may decide or change.

Which source is authoritative, how fresh it must be and what happens when sources disagree.

Which confidence, value or risk thresholds require human review.

Whether an action is reversible and how it is rolled back or compensated.

What gets logged, reconciled and learned from later.

Who owns an ageing exception and where it escalates.

The safe fallback when the system, vendor or integration is unavailable.

These boundaries matter whether the automation is a conventional rules engine or an AI-enabled agent. For AI systems, the NIST AI Risk Management Framework calls for clear human roles, realistic evaluation, ongoing monitoring, third-party risk management and response and recovery plans.

Design the exception and recovery path

Happy paths prove that a workflow can complete. Exceptions show whether the operation can recover without losing money, control or customer trust.

In e-commerce that means testing partial stock, payment failure, damaged or rejected inventory, conflicting promotions, late supplier data, marketplace restrictions, unusual return patterns and carrier events that arrive out of order. The list will differ by category, channel, market and operating model.

A study of execution logs from five organisations found that exceptions, especially unexpected ones, were associated with longer throughput times. The useful lesson is not that exceptions are the whole operation. It is that known exceptions should be designed into the process, while unknown ones need monitoring and a safe response (Dijkman et al.).

An operator reviews an amber exception branching away from an orderly workflow

Buy against the redesign

Now turn the redesigned workflow into requirements. Test buy, build, integrate and no-buy options against real cases, including the awkward ones. A vendor demonstration should show failure, recovery and reconciliation, not only completion.

Measure the net operation:

Total elapsed time and queue age.

First-time-right, reopen and rework rates.

Customer outcome and avoidable contacts.

Refund, margin or inventory leakage.

Control effectiveness and auditability.

Resilience, recovery time and manual fallback.

Integration, migration, training and ongoing ownership cost.

Data access, portability and the exit path if the vendor fails.

A tool that saves ten minutes and creates a review queue may still improve the operation. The answer depends on the queue’s volume, delay, quality and risk. Local speed is evidence only when the end-to-end outcome also improves.

Before committing, challenge the redesign with people who operate comparable e-commerce workflows. A peer at the same scale and with similar channel or fulfilment complexity can expose exceptions that a vendor demonstration or an internal team accustomed to its own workarounds may miss.

The goal is neither a smaller software bill nor a larger one. It is a more capable, governable and resilient operation that keeps the customer promise when reality refuses to follow the happy path.

Bring the next decision to the Network.

Compare notes with experienced peers and contribute what you have learned.