Planner Overrides Are Training Data
Why manual changes to analytical recommendations should be captured as structured evidence, not dismissed as noise or treated as proof that the model failed.
Most analytical systems treat an override as the end of the story.
The model recommends 1,000 units. The buyer changes it to 1,400. The system records the final order, the purchase order goes to the supplier, and everyone moves on.
That is a waste.
The difference between 1,000 and 1,400 may contain more information about the real decision problem than another month of model tuning.
A planner override can reveal a missing constraint, stale state information, a hidden business objective, a timing issue, a bad forecast, a supplier relationship, a promotional commitment, a cash concern, or a policy that is mathematically sensible but operationally impossible.
Yet many organizations capture none of this. They preserve the transaction and discard the reason.
Then six months later, the data science team retrains the model on the final order quantity as if 1,400 units were simply the correct answer.
This is how organizations teach their systems to imitate decisions they do not understand.
The override is not automatically right
There is an important distinction.
Planner overrides are evidence. They are not ground truth.
A human can know something the model does not know. A human can also panic, anchor on last year, protect a local metric, respond to an executive escalation, or simply make a bad decision.
The useful question is not:
Who was right, the model or the planner?
The useful questions are:
- What information was available to each decision maker?
- What objective was each one implicitly optimizing?
- What constraints did each one believe were real?
- What happened after the decision?
That framing turns an override from a political event into a decision-science artifact.
If the planner knew a supplier shutdown was likely and the model did not, the problem may be state estimation.
If the planner increased the order because the item was tied to a strategic launch, the problem may be objective specification.
If the planner cut the order because the warehouse was physically full, the problem may be a missing capacity constraint.
If the planner overrode the recommendation and the result was worse, the organization may have discovered a training opportunity or a misaligned incentive.
The override itself does not answer the question. It tells you where to investigate.
Most systems log outcomes, not decisions
Supply chain systems are usually good at recording what eventually happened.
They know the final purchase order. They know the shipment. They know the receipt. They know the inventory balance. They may know the forecast that was active at the time.
But a decision system should preserve the decision path.
At minimum, a serious analytical workflow should be able to reconstruct:
- The state of the world when the recommendation was generated.
- The recommendation produced by the model.
- The key assumptions behind that recommendation.
- The human adjustment, if any.
- The reason for the adjustment.
- The final action actually executed.
- The uncertainty that resolved afterward.
- The economic outcome.
Without that chain, postmortems become storytelling.
A stockout happens and someone says the forecast was bad. Excess inventory appears and someone says the buyer ordered too much. Service falls and someone says the optimizer was too conservative.
Maybe.
But if you cannot replay the state, recommendation, override, execution, and outcome, you are not debugging a decision system. You are debating memories.
Build an override taxonomy
A free-text comment box is better than nothing, but it is not enough.
The organization should define a small, practical taxonomy for why recommendations are changed.
For example:
- New information not available to the model
- Incorrect inventory or pipeline state
- Supplier constraint
- Transportation constraint
- Warehouse or labor constraint
- Promotion or event
- Strategic customer commitment
- Cash or working-capital pressure
- Risk tolerance adjustment
- Recommendation not operationally executable
- Planner judgment with no new information
- Other
The goal is not bureaucratic perfection. The goal is to make repeated failure modes visible.
If 38 percent of overrides are caused by supplier minimums that are missing from the optimization model, you do not have an adoption problem. You have a modeling problem.
If planners repeatedly override recommendations because the system ignores a promotion calendar, you do not need a workshop about trusting AI. You need the promotion calendar.
If overrides cluster around a specific region because managers are measured on local service while the optimizer maximizes network profit, you may have discovered an incentive problem.
A taxonomy turns anecdotes into a queue of engineering work.
Measure the override economically
Override rate is a weak metric by itself.
A system with a 5 percent override rate can be terrible if those overrides affect the most valuable decisions. A system with a 40 percent override rate can still be useful if most changes are small, low-value adjustments.
The more interesting measurement is economic.
For each overridden decision, preserve both candidates:
- the original analytical recommendation,
- the executed human-adjusted decision.
Then, when possible, evaluate both under the same realized history or the same simulated scenarios.
This creates a counterfactual comparison.
Suppose the model recommended ordering 1,000 units and the planner ordered 1,400. Demand later resolved at 1,150.
A simplistic analysis says the planner was wrong because 250 units remained.
A better analysis includes margin, shortage cost, markdown exposure, lead time, future demand, holding cost, supplier terms, and the value of preserving service.
A still better analysis recognizes that one realized demand path is not enough. If the original decision was made under uncertainty, compare both actions across a common set of plausible futures.
That is where simulation becomes useful.
You are not simulating to prove the model was smarter than the planner.
You are simulating to ask:
Under the information available at the time, which decision was more economically robust?
This is a much more mature question.
Use common scenarios
When comparing an original recommendation with an override, evaluate both under the same scenarios whenever possible.
If one decision is tested against one set of random demand paths and the other decision is tested against another, simulation noise can obscure the difference.
Using common random numbers gives both candidates the same futures.
The only thing changing is the decision.
This makes the comparison more informative, especially when the economic difference between two candidates is smaller than the natural variation in the system.
The same principle applies to policy tuning.
If you are comparing two replenishment policies, two allocation rules, or two buying cadences, give them the same demand, lead-time, disruption, and capacity scenarios.
Decision science is hard enough without adding unnecessary noise to the experiment.
Overrides reveal missing state
Many so-called model failures are actually information failures.
The mathematical model may be perfectly capable of making the right decision if it had the right state.
A buyer knows that a supplier representative warned of a likely delay.
A planner knows that a major customer is preparing an unannounced promotion.
A warehouse manager knows that a section of the building will be unavailable next week.
A merchant knows that a product is being repositioned and should not be replenished aggressively.
None of this appears in the official data pipeline.
The planner overrides the recommendation.
The data science team sees noncompliance.
The operations team sees common sense.
Both are looking at the wrong abstraction.
The real issue is that the decision system and the human are operating with different state variables.
Before asking people to trust the model, ask what they know that the model does not.
Then decide whether that information should become data, a formal constraint, a parameter, a scenario, or an explicit human input.
Overrides reveal missing objectives
Sometimes the model has the facts and still recommends something the business rejects.
That often means the disagreement is economic.
A network optimizer may reduce inventory at a location because the expected profit impact is small. The regional leader may protect that inventory because service level is part of their performance score.
A replenishment model may reduce an order because the expected markdown risk is high. The merchant may increase it because being out of stock during a launch is politically unacceptable.
A transportation model may consolidate shipments to reduce cost. The operations team may prefer more frequent deliveries because variability creates labor problems downstream.
These are not necessarily irrational overrides.
They may reveal objectives that were never made explicit.
The wrong response is to hide more penalties inside the objective until people stop complaining.
The better response is to surface the tradeoff.
What are we actually optimizing?
Which costs are real?
Which risks matter?
Who owns the tradeoff?
A useful model does not eliminate disagreement. It makes disagreement precise enough to resolve.
Overrides reveal incentive design
Analytical systems operate inside organizations, and organizations have local incentives.
The optimizer may be asked to maximize total network profit.
The planner may be measured on in-stock rate.
The warehouse manager may be measured on utilization.
Finance may be measured on working capital.
Procurement may be rewarded for unit cost.
Then leadership acts surprised when everyone overrides the system in a different direction.
The model is not fighting human resistance. It is colliding with the scorecard.
This is why adoption cannot be solved entirely through interface design, training, or executive sponsorship.
If the decision recommendation asks a person to hurt the metric used to evaluate their performance, override behavior is predictable.
Override telemetry can expose these conflicts.
If one team systematically adds inventory while another systematically removes it, do not immediately tune the model separately for each team. First inspect what each team is rewarded for doing.
Sometimes the most important optimization work is organizational.
Do not optimize for zero overrides
A zero-override system is not necessarily mature.
It may mean the model is excellent.
It may also mean people have stopped using it.
Or they execute the recommendation in the official system and undo it somewhere else.
Or the interface makes overriding too difficult.
Or employees have learned that disagreement is punished.
The objective should not be compliance.
The objective should be better decisions.
A healthy decision system may preserve a role for human intervention, especially when information arrives through channels the model cannot yet observe.
The important part is that intervention becomes visible and learnable.
The organization should know when humans add value, when the model adds value, and when both are repeatedly failing for the same structural reason.
The feedback loop should change the system
Collecting override reasons is useless if nothing happens afterward.
A practical review cadence might look like this:
Daily: Flag high-value overrides and obvious data-quality issues.
Weekly: Review repeated override categories, large economic deviations, and operational exceptions.
Monthly: Prioritize changes to data pipelines, model constraints, objectives, policies, and training.
Quarterly: Revisit whether the decision rights and incentives around the system still make sense.
The point is not to create another dashboard.
The point is to convert operational disagreement into system improvement.
A repeated override should eventually lead to one of a few outcomes:
- the model changes,
- the data changes,
- the process changes,
- the incentive changes,
- the planner behavior changes,
- or the organization explicitly accepts the exception.
If the same override happens every week for a year, it is no longer an exception.
It is part of the decision problem.
A production decision system should learn from disagreement
The first version of an optimization system is always incomplete.
That is normal.
The mistake is pretending the model is complete and treating every human adjustment as contamination.
The other mistake is assuming the human must be right and letting every override bypass analysis.
A stronger system does neither.
It preserves the recommendation. It preserves the override. It preserves the information available at the time. It evaluates outcomes economically. It looks for repeated patterns. Then it changes the decision architecture.
This is how a prototype becomes an operational capability.
Not because the model eventually becomes perfect.
Because the organization builds a mechanism for learning where the model, the data, the process, and the humans disagree.
The override is not the failure.
The failure is throwing away the lesson.