Machine Learning for Demand Forecasting in Retail

Machine learning is useful in retail demand forecasting for a fairly practical reason: retail demand gets messy at exactly the level where planners have to make decisions.
It is relatively easy to say a category will grow next month. It is much harder to decide how many units of a particular SKU should be available in a particular location, whether a promotion changed the underlying demand curve, or whether weak sales mean customers lost interest or the store simply ran out of stock.
That distinction matters.
Retailers do not make inventory decisions against a clean demand curve. They make buys, set replenishment quantities, move inventory between stores, manage WOS and decide when markdown risk has become uncomfortable. A forecast is useful when it improves those decisions.
That is where machine learning can earn its place.
Why Retail Demand Is Harder to Forecast Than a Sales Trend
Sales history looks more objective than it really is.
A POS system records what sold. It does not necessarily record what customers wanted to buy.
Consider a basic apparel example. A store receives a size run of XS through XL. Medium and Large sell quickly and are gone by Thursday. Small and XL remain available through the weekend.
The sales history now shows zero sales for M and L after Thursday. Taken literally, demand collapsed. In reality, availability collapsed.
A forecasting model trained blindly on transactions can learn the wrong lesson from that history. The same issue appears when a store receives a poor initial allocation, when an item is launched with limited history or when one product substitutes for another after a stockout.
Then add everything else that moves retail demand: price, markdowns, promotions, holidays, weather in some categories, store profile, seasonality, assortment changes, product lifecycle and local events.
At an aggregate level, much of that noise gets smoothed out. At SKU-store level, it does not.
Retail forecasting research has long highlighted the operational difficulty created by promotions, product and store hierarchies, different aggregation levels and the sheer dimensionality of retail data. The M5 forecasting competition makes the scale particularly tangible. Participants forecast 42,840 hierarchical Walmart sales series, with 30,490 at the lowest level.
That is closer to the problem planners actually face. You are not forecasting one neat sales curve. You are managing thousands of small, related and frequently noisy demand problems.
Machine learning's advantage is not that it is newer. It is that a model can learn relationships across those demand series while considering variables beyond historical sales.
That becomes useful when demand has multiple interacting causes.
What Machine Learning Actually Adds to a Retail Forecast
Traditional forecasting methods are not obsolete.
For stable products with recurring seasonality and enough clean history, relatively straightforward time-series methods can work very well. If an item sells at a predictable rate and rarely changes price or promotional status, adding model complexity for its own sake accomplishes very little.

Machine learning becomes more interesting when historical demand alone does not explain enough.
Instead of asking only, "What usually happens after this sales pattern?" the forecast can account for questions such as:
- What price is the item selling at?
- Is it on promotion?
- What happened to similar products under similar conditions?
- Is this store behaving differently from the chain?
- Is recent velocity accelerating?
- Is the product early or late in its lifecycle?
- Was inventory actually available during the historical period?
There is another useful difference. Machine learning models can learn globally across related demand series rather than treating every SKU-location combination as a completely independent forecasting problem.
That matters for sparse data.
A new SKU might have only a few weeks of sales. A low-volume SKU may sell intermittently. Neither gives an isolated model much to work with. But those products still belong to categories, share attributes with other products and sell through stores with established demand patterns.
The M5 competition showed strong performance from machine-learning approaches in a large and complex retail forecasting environment. It also showed why retailers should resist turning this into a statistical-method beauty contest. Simpler approaches remained competitive in parts of the hierarchy.
More sophisticated does not automatically mean more accurate
Retail teams should be suspicious of any forecasting project that begins with the assumption that every existing method needs to be replaced by AI.
Start with the demand problem.
Research on SKU-level promotional forecasting found that relatively simple time-series approaches performed well in normal, non-promotional periods, while regression-tree approaches using explicit promotional information delivered much stronger results when promotions changed the demand pattern.
That makes operational sense.
A basic replenishment SKU with stable weekly demand is not the same forecasting problem as a fashion item entering markdown, a seasonal SKU approaching peak demand or a product about to receive a 30% promotional discount.
Different demand behavior deserves different treatment.
Better Retail Forecasts Start With Better Demand Signals

A retailer can feed five years of transactions into a machine-learning model and still get a bad forecast.
More history is not automatically better information.
The model needs signals that help explain why demand moved. Depending on the retailer, useful inputs might include lagged sales, rolling demand, price, markdown depth, promotional status, holidays, store or channel, product attributes, lifecycle stage, inventory availability and demand behavior among related products.
This is where a lot of forecasting work becomes less glamorous and more valuable.
Promotions are a good example. If sales doubled during a promotion, the model needs to understand that the item was promoted. Otherwise, it may treat that volume as evidence of higher underlying demand.
Research using SKU-store data has shown the value of explicitly including sales and promotional features and of pooling information across demand series rather than modelling every SKU-store combination in isolation.
The operational implication is simple. Better models cannot compensate indefinitely for weak demand signals.
If promotional calendars sit in one spreadsheet, inventory availability somewhere else and sales history in the ERP, somebody eventually has to bring that information together. This is one reason spreadsheet-heavy forecasting processes become difficult to maintain as assortments and store counts grow.
Promotions are not just temporary demand spikes
Promotional forecasting gets more complicated after the promotion ends.
A retailer needs to separate baseline demand from incremental promotional demand. Discount depth matters. Timing matters. The response to previous promotions matters. Competing promotions may matter.
Then there is the stock-up effect.
Suppose customers who normally purchase an item every few weeks buy extra units during a promotion. The promotional week performs strongly, but some of that demand has been pulled forward. The week after the event can therefore run below normal.
Recent forecasting research has specifically examined this post-promotion period, an area that earlier promotion-versus-non-promotion approaches could overlook.
This has a direct inventory consequence.
A model can forecast the promotional spike reasonably well and still leave the retailer overstocked if it assumes baseline demand immediately returns afterward. WOS starts climbing just as the merchandising team is moving on to the next event.
Stockouts create the opposite distortion. If inventory was unavailable, low sales should not automatically become a low-demand training signal.
Good forecasting starts by distinguishing demand from the operational conditions that allowed demand to become a sale.
Forecast Accuracy Has to Hold Up at SKU, Store and Size Level
Retail forecasts live inside a hierarchy.
Company. Category. Product or style. SKU. Store or fulfillment location.
The higher you go, the cleaner the numbers tend to look because individual errors cancel each other out. One SKU beats plan while another misses it. One store is strong while another is weak. At category level, the total can appear surprisingly accurate.
Unfortunately, you cannot replenish a category total.
Consider a fashion retailer that correctly forecasts demand for 1,000 units of a style. On paper, that is a good forecast.
But inventory was bought against the wrong size curve.
Medium and Large run out first. XS and XL keep accumulating WOS. Total inventory may still look healthy. The customer looking for a Large does not care. Neither does the planner staring at a broken size run and realizing the remaining units are increasingly likely to require markdown.
That is why size-level forecasting matters in apparel and footwear. The unit decision happens below the style forecast.
Geography creates the same problem.
A retailer can buy roughly the correct national quantity and still create stockouts through allocation. High-velocity stores get too little. Lower-velocity stores sit on excess stock. Total inventory is technically sufficient, but it is sitting in the wrong places.
The hierarchical structure of the M5 competition is useful here because forecasts were required at very granular levels while remaining part of a broader retail hierarchy.
So when someone says a new forecasting model improved accuracy, the next question should be: where?
Better category-level forecasting can improve financial planning and OTB decisions. Better SKU-location and size-level forecasting is more directly connected to allocation, replenishment and availability.
Both matter. They solve different problems.
Measure ML by the Inventory Decisions It Improves
Forecast accuracy is an intermediate metric.
Retailers still need statistical measures. MAE, WAPE, RMSE and forecast bias can all be useful depending on the assortment and forecasting setup. They help teams compare models and understand where errors are occurring.

But none of them is the final business objective.
A planner ultimately cares whether the forecast contributes to better service levels, fewer avoidable stockouts, healthier WOS, stronger inventory turns, better sell-through and less excess stock requiring markdown.
The distinction can expose misleading improvements.
Suppose a new model improves overall accuracy across a seasonal category. Most of the improvement comes from a long tail of relatively low-value SKUs. At the same time, the model continues systematically underforecasting several products that drive a disproportionate amount of category revenue.
The accuracy dashboard improves.
The stores still stock out of the products customers actually want.
That is why bias also deserves attention. A model that is consistently wrong in one direction creates an inventory behavior. Persistent underforecasting can create stockouts and missed sales. Persistent overforecasting quietly ties up working capital and increases markdown exposure.
Backtesting should reflect this operational reality. Test models across multiple historical windows, not one convenient holdout period. Include seasonal peaks, promotions, quieter trading periods and demand shocks where the data allows it.
A forecast that performs beautifully in March and falls apart during holiday trading is not a particularly useful retail forecast.
The forecast only matters when it changes the inventory decision
The actual chain looks something like this:
Forecast → inventory target → buy or replenishment → allocation → sell-through → markdown exposure
That is the standard ML forecasting projects should ultimately be held against.
Alibaba provides a large-scale example of this broader approach. Its published retail implementation combined deep-learning demand forecasting with inventory management, promotional pricing and markdown-related optimization rather than treating forecasting as a standalone analytics exercise. The company reported annual savings of $42 million in shrinkage and inventory costs, alongside $110 million in increased sales and $13 million in increased profit.
The important part is not the size of Alibaba or the particular model architecture. It is that forecasting was connected to the decisions downstream.
For a mid-market retailer, the same principle applies on a more practical scale.
If a forecast identifies rising demand but nobody adjusts replenishment, nothing changed. If it identifies a store-level imbalance but inventory cannot be reallocated, the insight arrived without an action. If a planner gets a forecast but still has to spend hours exporting files and rebuilding Excel models to understand the inventory implication, much of the benefit has been lost.
This is where platforms such as Flagship can be useful. The value is not simply generating an AI forecast. It is bringing the forecast down to the inventory level planners work at, including size-level demand, and making changing inventory exposure visible early enough to act on it.
Retailers do not need the smartest forecasting model in the room.
They need a forecasting process that helps them carry less unnecessary stock without blindly creating more stockouts. They need to see demand changes earlier, allocate inventory more intelligently and recognize excess WOS or markdown risk before the only remaining option is a deeper discount.
Machine learning can help do that. But the standard should remain stubbornly operational.
If the forecast does not improve the inventory decision, its sophistication is mostly academic.