NEW: The Decision Factory — a novel about decisions under uncertainty. Get it on Amazon
Decision Science · · Adam DeJans Jr.

Forecast at the Grain of the Decision

Why weekly, monthly, SKU-level, and aggregate forecasts are not interchangeable once a real supply chain decision sits downstream.

forecastingoptimizationsupply-chaindecision-scienceuncertaintysimulation

Forecast at the Grain of the Decision

A forecast can be perfectly reasonable and still be wrong for the decision you are trying to make.

This happens constantly in supply chain.

A planning team forecasts 52 weeks of demand. Another system forecasts the total demand over the same 52 weeks and spreads it across time. Both forecasts may have the same annual mean. They may even have similar annual error metrics.

Then someone says, “These are basically the same forecast.”

They are not.

If the decision is made weekly, timing matters. If inventory carries from one week to the next, timing matters. If replenishment lead time is 14 weeks, timing matters. If a stockout in week 8 cannot be repaired by excess demand in week 40, timing matters.

The aggregation level of the forecast is part of the decision model whether you acknowledge it or not.

Start with the decision, not the forecast table

Before arguing about weekly versus monthly forecasts, write down the decision.

Suppose the business chooses order quantities

[ q_t \ge 0 ]

for periods (t=1,\ldots,T).

Inventory evolves as

[ I_{t+1}=I_t+q_{t-L}-D_t, ]

where (L) is lead time and (D_t) is uncertain demand in period (t).

Already, the important question is obvious: what uncertainty does the decision encounter while it is active?

If I place an order today that arrives in week 12, the distribution of demand before and after week 12 matters differently. The total demand over a year is not enough to reconstruct that exposure.

This is why forecast design should begin with questions like:

  • What action is actually being selected?
  • At what cadence can that action change?
  • What state carries between decisions?
  • How long are decisions committed before they can be revised?
  • Which constraints operate weekly, monthly, quarterly, or over the full horizon?
  • What information will become available before the next action?

If those questions are unanswered, debating MAPE at different aggregation levels is mostly theater.

Equal totals do not imply equal paths

Consider two 8-week demand paths:

[ A=(20,20,20,20,20,20,20,20) ]

and

[ B=(0,0,0,0,40,40,40,40). ]

Both total 160 units.

For an annual budget, that may be enough.

For replenishment, they are obviously different worlds.

Under path A, inventory drains steadily. Under path B, inventory sits idle for four weeks and then disappears quickly. A supplier with a six-week lead time sees a very different ordering problem. Holding cost changes. Stockout exposure changes. Capacity timing changes. The value of expediting changes.

Compressing both paths to the scalar 160 destroys information that the policy may need.

The same problem appears with uncertainty.

Suppose total 8-week demand has distribution

[ D_{1:8} \sim F. ]

Knowing (F) does not tell us how demand is distributed across weeks. Many joint weekly distributions can produce the same distribution of the total.

One may have independent weekly demand. Another may have persistent high and low regimes. Another may shift demand between adjacent weeks while keeping the total nearly fixed.

Those processes can have identical horizon totals and radically different inventory consequences.

Marginals are not paths either

There is another subtle mistake: forecasting every week separately and assuming that gives you a valid demand path model.

Suppose you estimate

[ D_t \sim F_t ]

for every week.

You still do not know the joint distribution

[ (D_1,D_2,\ldots,D_T). ]

That joint structure matters whenever decisions or states span multiple periods.

Demand can be positively correlated across weeks. Promotions can pull demand forward. A product launch can create persistent upside. A supply interruption can censor observed sales for several consecutive periods. Seasonality can make neighboring weeks move together.

If a simulator independently samples each weekly marginal, it may generate paths that look plausible one week at a time but are nonsense as trajectories.

This is a practical issue, not a statistical footnote.

A replenishment policy is exposed to trajectories.

Aggregation changes uncertainty

Aggregation usually makes forecasts look easier.

Weekly errors can partially cancel when rolled into a month. SKU errors can cancel when rolled into a category. Location errors can cancel when rolled into a region.

That is legitimate statistical pooling. It is also dangerous if the downstream decision cannot pool the same way.

Imagine two SKUs with demand errors of +100 and -100 units. Category error is zero.

If the inventory is interchangeable, perhaps that is fine.

If one SKU is stocked out while the other has 100 extra units sitting on the shelf, the category forecast being exactly correct does not help you.

The rule is simple:

Aggregation is useful only when the decision can exploit the same aggregation.

If inventory, capacity, substitution, or money is fungible across the aggregated dimension, pooling can be real. If it is not, the forecast metric may be hiding operational error.

Temporal aggregation creates the same trap

The same logic applies across time.

Suppose the business has a monthly demand forecast of 400 units.

You need weekly demand for a replenishment simulation, so you divide by four:

[ (100,100,100,100). ]

That is not a weekly forecast. It is an allocation assumption.

Maybe demand is actually

[ (50,80,110,160). ]

Maybe there is a promotion in week 2. Maybe the last week includes a holiday. Maybe customers systematically order at month-end.

The monthly total can be correct while the implied weekly path is badly wrong.

If the downstream policy is sensitive to those within-month movements, the disaggregation rule belongs in the model and needs validation like any other assumption.

The forecast horizon and decision horizon are different things

A common architecture mistake is to make the forecast horizon dictate the decision structure.

A forecast might extend 104 weeks because that is useful for long-lead-time planning. That does not mean the optimizer should make 104 equally meaningful weekly decisions.

Likewise, a policy may make only one immediate order but need a long forecast horizon to value that order correctly.

Separate three concepts:

  1. Forecast horizon: how far into the future uncertainty is represented.
  2. Decision horizon: which future actions are explicitly represented in the model.
  3. Commitment horizon: which decisions are actually executed before the system gets another chance to replan.

They can be different.

In a rolling system, you may simulate 78 weeks, optimize a sequence of future actions, execute only the first week, observe new information, and solve again.

That is normal.

Forecast coherence is useful, but it does not solve the decision problem

Hierarchical and temporal reconciliation methods can make forecasts add up properly.

Weekly forecasts can sum to monthly forecasts. SKU forecasts can sum to category forecasts. Regions can sum to the network.

This is useful. Incoherent forecasts create obvious operational problems.

But coherence is not sufficient.

A coherent forecast can still put uncertainty at the wrong grain. It can still erase correlation. It can still smooth peaks that trigger capacity or inventory constraints. It can still optimize an accuracy metric that has little relationship to the economics of the action.

The test is downstream performance.

Build the uncertainty representation around the policy

For a real replenishment problem, I usually want the uncertainty model to generate scenarios such as

[ D^{(s)}=(D_1^{(s)},D_2^{(s)},\ldots,D_T^{(s)}), ]

where (s) indexes a plausible future path.

The policy then interacts with each path.

For example, a simple order-up-to policy might be

[ q_t^{(s)}=\max{0,S_t-IP_t^{(s)}}, ]

subject to pack sizes, MOQs, capacity, supplier calendars, and whatever else is actually executable.

The simulator tracks inventory, backorders or lost sales, receipts, working capital, holding cost, and other economic consequences through time.

Now you can compare uncertainty representations by the thing that matters: the decisions they produce and the economics those decisions realize.

A useful experiment

Suppose you are deciding between two forecasting architectures:

Model A: forecast total horizon demand, then allocate it across weeks.

Model B: forecast weekly demand directly, preserving temporal structure, then aggregate when an aggregate is needed.

Do not settle this with a slide showing RMSE.

Run both through the same decision pipeline.

Use the same initial inventory, lead times, constraints, policy class, cost model, and evaluation scenarios. Then compare:

  • order quantities,
  • timing of orders,
  • inventory trajectories,
  • stockout frequency and duration,
  • excess inventory,
  • expedite usage,
  • capacity violations,
  • working capital,
  • realized contribution margin,
  • total economic regret versus a benchmark policy.

Also measure decision disagreement.

If two forecasting systems have similar accuracy but recommend materially different actions, you have found the part worth investigating.

If they produce nearly identical actions across the state space you actually encounter, the forecasting difference may not matter much operationally.

That is a far more useful conclusion than declaring one forecast 3% more accurate.

Do not evaluate on the same uncertainty model used to optimize

This matters especially when tuning policies in simulation.

If Model A generates the scenarios used to tune the policy and those same scenarios are used to evaluate it, Model A gets to define its own exam.

Use an independent evaluation process.

Historical replay is useful when point-in-time data is available. Holdout simulation is useful when history is sparse. Stress scenarios are useful for tail behavior.

At minimum, separate:

  • policy tuning scenarios,
  • policy confirmation scenarios,
  • final comparison scenarios.

If you are comparing forecast architectures, use common evaluation scenarios where possible so Monte Carlo noise does not masquerade as a model difference.

Constraints reveal when granularity matters

Forecast granularity matters most near nonlinear or discrete decision boundaries.

Consider a vendor MOQ:

[ \sum_i q_{i,t} \ge M y_t, ]

where (y_t) indicates whether an order is opened.

A small shift in weekly demand can move the economically preferred order from below the MOQ to above it. The result is not a small change in the decision. It can be an entire additional purchase order.

The same thing happens around:

  • truck capacity,
  • warehouse receiving limits,
  • production setups,
  • case packs,
  • cash budgets,
  • shelf capacity,
  • supplier blackout dates,
  • contractual order windows.

Smoothing uncertainty across those boundaries can make the model look stable precisely because you removed the variation that causes the real system to change decisions.

Metrics should follow the decision

Forecast metrics still have a place. Bias, calibration, pinball loss, CRPS, and scale-adjusted error can all diagnose useful properties.

But they are intermediate metrics.

For a decision system, I would also track:

Decision metrics

  • frequency of order changes,
  • quantity disagreement between forecast versions,
  • frequency of crossing MOQ or capacity thresholds,
  • stability of first-stage actions,
  • regret from using one forecast representation instead of another.

Economic metrics

  • margin,
  • holding cost,
  • lost-sales cost,
  • obsolescence,
  • expedite cost,
  • working capital,
  • capacity cost.

Risk metrics

  • downside quantiles,
  • severe stockout probability,
  • tail inventory exposure,
  • CVaR or another explicit downside measure when appropriate.

A forecast improvement that does not move the decision or improve these outcomes may be scientifically interesting and operationally irrelevant.

Implementation notes that save pain later

Preserve forecast vintage

Never backtest a decision using a forecast that was revised after the decision date.

Store the exact forecast distribution or scenario set that was available when the decision would have been made.

Otherwise you have leaked future information into the policy evaluation.

Keep transformations explicit

If monthly forecasts are converted to weekly values, store the transformation. If category demand is allocated to SKUs, store the allocation rule. If scenarios are reconciled, store the pre- and post-reconciliation versions.

Hidden transformations are where debugging goes to die.

Log the state with the decision

For every recommendation, save enough state to reconstruct why it happened:

  • inventory position,
  • open orders,
  • forecast vintage,
  • scenario or distribution identifier,
  • lead-time assumptions,
  • active constraints,
  • policy parameters,
  • objective components,
  • final recommended action.

When someone asks six weeks later why the system ordered 4,800 units, “that was the forecast at the time” is not enough.

Compare at multiple aggregation levels

A weekly model can still be evaluated monthly and annually. A SKU model can still be evaluated at category level.

The reverse is not generally true. Once you throw away detail, you may not be able to reconstruct the decision-relevant structure.

This is one reason I prefer preserving the finest useful representation upstream and aggregating intentionally downstream.

Common failure modes

1. Choosing granularity because the data warehouse already has it

The table is monthly, so the model is monthly.

That is an implementation convenience pretending to be problem formulation.

2. Assuming better aggregate accuracy means better decisions

Aggregation mechanically removes some noise. It can improve the metric while making the action worse.

3. Sampling weekly forecasts independently

Good weekly marginals do not automatically produce good multi-week paths.

4. Disaggregating with fixed percentages forever

A fixed 25%-per-week split may be acceptable as a temporary approximation. Treating it as truth for years is different.

5. Ignoring forecast vintage

Using today’s reconstructed history to judge yesterday’s decision is one of the easiest ways to make a policy look smarter than it was.

6. Optimizing forecast accuracy instead of policy value

The forecast is an input. The decision is the product.

What I would do in practice

If I inherited a planning system with disagreement about forecast granularity, I would not start by rebuilding the forecasting stack.

I would take a representative set of items and decisions and build a controlled comparison.

First, document the executable decision cadence, lead times, state transitions, and constraints. Second, identify exactly where each candidate forecast gains or loses information through aggregation or disaggregation. Third, feed both representations through the same policy. Fourth, evaluate the resulting actions on common out-of-sample demand paths. Fifth, inspect the cases where the decisions diverge most.

Those disagreement cases are gold.

They tell you whether the missing information is economically important. Maybe weekly seasonality matters. Maybe only the annual total matters for a long-term capacity contract. Maybe correlation matters only for a small group of long-lead-time products. Maybe the supposedly sophisticated weekly model changes almost nothing.

You do not need a philosophical answer to “What is the correct forecast granularity?”

You need an empirical answer to a narrower question:

What information must be preserved for this decision to perform well?

That is the level at which forecasting becomes decision science instead of a leaderboard exercise.