OR68 and Gousto’s ToteSim Digital Twin

Fabrice Durier presenting ToteSim at OR68, Nottingham.

From Data to Decisions at Scale

Earlier this month, the UK and international Operational Research community descended on the University of Nottingham for OR68: From Data to Decisions.

The event featured three intense days of plenaries, poster sessions, and valuable discussions at our stand. We were honoured to present and, in particular, answer the standout question from practitioners and business leaders alike: how do you build enough stakeholder trust in an analytical model to let it directly influence physical, high-risk operations?

Decision Lab and Gousto tackled that question head-on in our joint session in the Simulation & Digital Twins stream. Presenting alongside Dr Fabrice Durier and Will Seymour from Gousto, Decision Lab’s Peter Riley unpacked the delivery of ToteSim, a data-driven simulation twin of the automated blue-tote replenishment system at Gousto’s Warrington fulfilment facility.

Earlier this year, we published a high-level overview in our case study, How Simulation Became Gousto’s Recipe for Growth. At OR68, we took the conversation a level deeper into the applied decision science: the validation mechanics, the experimental rigour, and the physical factory evidence.

Here are the key operational research insights shared during the conference.

1. When Combinatorial Variety Meets Physical Bottlenecks

As explored in our published case study, Gousto’s strategic mission to scale its menu from 50 to 200 weekly recipes created exponential combinatorial complexity behind the scenes. Fulfilling roughly 200,000 bespoke boxes and picking 5 million ingredients every week means that an ingredient pick must occur every three seconds across dynamic, non-repeating recipes.

While Gousto already maintained a mature simulation model for customer “red-box” packing lines, that model assumed blue-tote ingredient replenishment was an infinite, frictionless feed. In reality, Warrington’s automated tote network, which houses 18,000 reserve bins, 9,800 pick slots, 20 vertical lifts, and 270 automated shuttles, was tightly constrained.

Conveyor merges and shared transport spines meant that a micro-stoppage in tote replenishment could cascade across the entire facility. Yet with a 24/7 operating plant, conducting physical trial-and-error was out of the question: running live A/B experiments in a single facility carried unacceptable operational risk.

The OR challenge was to determine whether intelligent algorithmic control could unlock latent transport capacity without committing to capital-intensive physical expansions.

2. Bridging the Gap: What Makes “Simulation Twinning” Different?

Much of the academic discussion at OR68 centred on the levels, or extent, of digital twin maturity. Rather than striving for an immediate, all-singing real-time twin, ToteSim was built around pragmatic Simulation Twinning:

  • The Golden Triangle Architecture: Decoupling the physical replenishment facility into three bounded operational domains (Decant Induction, Reserve Storage, and Pick Stations) connected by the shared transport conveyor loop. This modular design allowed subsystems to be isolated or recombined across 24 project epics during an intensive five-month collaborative sprint.
  • Executing the Live Decision Stack: Crucially, ToteSim is not an abstract approximation of rules. Built in AnyLogic and fed shift-level historical states via Databricks, the model directly invokes Gousto’s live production Python decision services (including Dynamic Replenishment and Totewise). The simulation twin evaluates the exact algorithmic logic deployed in the warehouse.
  • The “Red + Blue = Purple” Modularity: Analysts can isolate the replenishment twin (Blue) with synthetic pick rates, or couple it directly with the customer order picking model (Red) to evaluate holistic “Purple” system impacts.

3. Decision-Specific Validation: Building Credibility for Capital Decisions

A core lesson shared with the OR Society delegates was how validation must be structured to build operational trust. Rather than chasing an unachievable standard of “100% global fidelity,” ToteSim’s validation strategy followed Robinson’s and Sargent’s principles of operational validity for defined decisions:

  • ~80% weighted mean accuracy across broad, multi-flow network replays against historical shift data.
  • >95% accuracy on critical, high-volume pathways (such as decant-to-reserve and reserve-to-pick flows).

This decision-specific calibration gave Gousto’s operations leadership the confidence to use ToteSim not just as a retrospective reporting dashboard, but as a proactive experimental testbed.

4. The First Factory Proof Point: “Robin Hood” Headroom Balancing

The first production deployment tested whether algorithmic optimisation could ease congestion on the shared conveyor spine feeding the pick tower.

Previously, rigid category quotas caused totes to queue at the decant line whenever their specific category limit was reached, even when spare capacity existed elsewhere on the conveyor. The team developed a drop-in controller, internally dubbed Robin Hood, based on headroom balancing:

  1. When total spine occupancy is below target, spare headroom is equalised dynamically across tote streams, clearing queues before merges freeze.
  2. When the belt reaches saturation, category limits are pulled back toward a central balance to prevent any single category from monopolising throughput.
  3. Designed as a direct drop-in replacement, it utilised the exact same input/output telemetry, requiring zero changes to physical sensors or PLC controls.

From 2,000+ Runs to Measured Production Evidence

Before factory rollout, ToteSim stress-tested the candidate algorithm across 2,000+ simulation runs, demonstrating a +4.3% improvement in cumulative spine activity area-under-the-curve under high stress ( p<0.001)p<0.001) ).

Deployed into live production at Warrington at the start of Q2 2026, a six-week post-deployment evaluation across more than 2 million live tote journeys confirmed what the simulation predicted:

  • Baseline Stability: General transit times and empty/reserve movements remained completely stable, proving that the intervention was safe and non-disruptive.
  • High-Stress Tail Reduction: The 95th-percentile queue time for constrained decant totes dropped by 24%.
  • Peak Congestion Relief: Avoidable queueing during peak operational windows reduced by 10% to 20%.

Moving from Event Models to an Enduring Decision Capability

Our headline message at OR68 was simple: a digital twin creates value long before it becomes an autonomous, fully connected IoT system.

By establishing a reliable loop—Observe  Initialise  Validate  Experiment  Deploy  Measure—ToteSim has evolved from a simulation model into an operational decision capability. Gousto is now utilising the twin to stress-test lift and shuttle failure modes, evaluate pallet induction strategies, and validate future automation logic before deploying capital.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *