NEW: The Decision Factory — a novel about decisions under uncertainty. Get it on Amazon
Decision Science · · Adam DeJans Jr.

Replay Before You Deploy

Why replay testing is one of the most important and overlooked disciplines in supply chain optimization, decision science, and analytical system adoption.

decision scienceoptimizationsupply chainsimulationproduction systems

Many optimization projects fail for a surprisingly simple reason: the team never proves that the proposed decision policy would have worked on yesterday’s business before asking people to trust it tomorrow.

The model solves. The forecast looks reasonable. The dashboard is attractive. The pilot presentation goes well. Yet when the recommendation finally reaches planners, operators, buyers, or executives, nobody wants to use it.

This hesitation is often interpreted as resistance to change.

Sometimes it is. More often it is a rational response to missing evidence.

Most operational teams have spent years living with the consequences of bad decisions. They have experienced stockouts, excess inventory, supplier failures, transportation disruptions, forecast misses, system outages, and executive initiatives that promised more than they delivered. When a new optimization system appears and claims it can improve decisions, experienced operators immediately ask a practical question:

“How do we know this would have worked last month?”

That question deserves a better answer than a PowerPoint.

The difference between a model and a policy

Many teams evaluate models when they should be evaluating policies.

A forecast model might reduce error by ten percent. A machine learning model might improve prediction quality. A MILP might generate a higher objective value than a heuristic.

Those are useful observations, but they do not necessarily tell us whether the business will improve.

Businesses do not experience forecasts. Businesses experience decisions.

Inventory gets purchased. Trucks get scheduled. Capacity gets allocated. Production gets sequenced. Customer orders get accepted or rejected.

The true unit of evaluation is not the model. It is the decision policy produced by the model.

What replay testing actually means

Replay testing means reconstructing historical decision environments and asking a simple question:

If the proposed policy had been operating at that point in time, using only information available at that moment, what would have happened?

This sounds obvious. In practice it is surprisingly difficult.

Teams frequently cheat without realizing it.

They evaluate decisions using information that was not available at decision time. They use corrected forecasts. They use cleaned datasets. They use actual lead times instead of uncertain lead times. They accidentally allow the model to see the future.

A valid replay environment must recreate the information state that existed when the decision was made.

Only then can we compare what the organization actually did against what the proposed policy would have done.

Why executives should care

Replay testing is not just a technical exercise.

It is an executive communication tool.

Most leaders do not want to debate solver parameters, branching strategies, feature engineering choices, or probabilistic calibration techniques.

They want evidence.

A replay study allows a team to say:

“Over the previous twelve months, this policy would have reduced inventory by eight percent while maintaining service levels.”

Or:

“This replenishment policy would have prevented forty percent of emergency shipments during peak season.”

Or:

“This buying cadence policy would have increased order consolidation while remaining within inventory risk tolerances.”

These statements connect analytical work to business outcomes.

Replay before simulation

Many practitioners immediately jump to simulation.

Simulation is valuable. It allows us to evaluate policies under scenarios that have never occurred.

However, replay should typically come first.

Historical replay answers the credibility question.

Simulation answers the robustness question.

Replay establishes whether a policy appears better than current practice under known operating conditions.

Simulation explores what happens when reality behaves differently.

The strongest decision systems use both.

The hidden value of replay

Most teams view replay as a validation exercise.

In reality, replay is often a discovery exercise.

It exposes missing constraints.

It reveals bad assumptions.

It uncovers data quality problems.

It highlights situations where planners consistently override recommendations.

It identifies regions where the model performs well and regions where it struggles.

Some of the most valuable findings during replay have nothing to do with optimization quality.

They reveal misunderstandings about how the business actually operates.

Replay and organizational trust

Trust is not built through technical sophistication.

Trust is built through evidence.

When planners see examples from their own categories, facilities, suppliers, and customers, the conversation changes.

The discussion moves away from whether optimization is theoretically useful.

Instead, the discussion becomes:

“Under what conditions should we trust this recommendation?”

That is a much more productive question.

A practical Optimization University principle

Before deploying any meaningful analytical decision system, ask three questions.

Can we replay it?

Can we explain it?

Can we measure it?

If the answer to any of those questions is no, the organization is probably not ready for production.

A surprising number of analytical projects spend months improving models while spending almost no effort building replay capability.

That is backwards.

The replay environment is often more valuable than the first version of the model because it creates the infrastructure needed to learn.

Final thoughts

The goal of optimization is not to produce impressive recommendations.

The goal is to improve operational decisions.

Replay testing creates the bridge between mathematical potential and operational reality.

It allows decision makers to evaluate policies using evidence rather than intuition. It creates credibility with operators. It reveals weaknesses before they become production incidents. It provides a foundation for simulation, experimentation, and continuous improvement.

The best analytical teams do not ask people to trust the future.

They start by proving they understand the past.