Retail Allocation Planning Using Reinforcement Learning Scenarios
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Supply chain planners lack effective tools for accurately forecasting and managing allocation of short life-cycle products, leading to sub-optimal allocation and resulting in lost sales and mark downs.
Innovation Solution
A system utilizing reinforced machine learning to model short life-cycle products as a semi-Markov Decision Process (SMDP) with a reward-penalty function, determining unconstrained and constrained quantities for stores and distribution centers, and integrating an allocation planner with archiving and planning systems to optimize product allocation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If generalized business rules are used for allocation, then the allocation process is simple and quick, but allocation accuracy deteriorates and results in lost sales and mark downs
Solution Approach 1:
The patent segments products into different categories based on life cycle characteristics (short life-cycle products vs. regular products). Different allocation strategies are applied to different segments: machine learning models for short life-cycle products and generalized business rules for regular products. This segmentation allows each method to operate in its optimal domain, improving overall accuracy while maintaining operational simplicity where appropriate.
Solution Approach 2:
The patent changes the fundamental parameter of allocation methodology from static business rules to dynamic machine learning models that continuously learn from historical data. The system adjusts allocation decisions based on real-time parameters such as product performance, market conditions, and inventory levels, enabling high accuracy while adapting to changing conditions without requiring complex manual processes.
2Measurement precision
If machine learning models are used for allocation, then allocation accuracy improves, but system complexity increases
Solution Approach 1:
The patent creates a universal allocation system that can handle multiple product types through a single integrated framework. The machine learning model is designed to process different product categories and life cycle characteristics using the same underlying architecture, eliminating the need for separate complex systems for each product type while maintaining high accuracy across diverse scenarios.
Solution Approach 2:
The system implements self-service through automated machine learning models that continuously train on historical allocation data and product performance information. The model automatically adjusts its predictions based on learned patterns, reducing the need for manual intervention and complex rule-based systems while maintaining high allocation accuracy through continuous self-optimization.
3Measurement precision
If detailed modeling of influencing factors is performed for each short life-cycle product, then allocation accuracy improves, but computing time and complexity increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-training machine learning models on historical allocation data and product performance information before actual allocation decisions need to be made. The model learns patterns and relationships from past data, enabling rapid predictions for new allocation scenarios without requiring time-consuming real-time analysis of all influencing factors, thus improving accuracy while reducing computing time during execution.
Data Source
AI summary
A system and method for allocation planning comprise a server comprising a processor and memory and configured to calculate a reward for a historical allocation of a product to one or more stores associated with a retailer. Embodiments include simulating what-if scenarios for the historical allocation to identify an allocation having a greater reward than the historical allocation and allocating a quantity of a product for a current allocation to the one or more stores based, at least in part, on a distance calculation of one or more independent variables for the historical allocation and the current allocation and the identified allocation having the greater reward then the historical allocation.


