Demand-driven material demand planning system based on deep reinforcement learning and optimization method

The demand-driven material requirements planning system, which utilizes deep reinforcement learning, dynamically adjusts key parameters, solving the problem of inventory mismatch in dynamic environments that traditional material requirements planning faces. This enables the system to achieve intelligent and adaptive optimization, improving the scientific nature and efficiency of planning decisions.

CN121745557APending Publication Date: 2026-03-27HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Traditional material requirements planning (MRP) struggles to adapt to market changes in dynamic environments, leading to inventory mismatches. Existing demand-driven MRP parameter settings rely on empirical rules, making it difficult to achieve intelligent and adaptive optimization.

Method used

A demand-driven material requirements planning system based on deep reinforcement learning is adopted. Through closed-loop optimization of the demand module, the demand-driven material requirements planning module, the deep reinforcement learning module, the simulation module, and the execution module, key parameters are dynamically adjusted. The agent is trained by using a deep reinforcement learning algorithm library to conduct multiple rounds of trial and error learning and optimize parameter configuration.

Benefits of technology

It significantly improves the adaptive response capability and intelligent level of planning decision-making of the demand-driven material requirements planning system, reduces the burden of manual adjustment, improves the scientific nature of parameter configuration and decision-making efficiency, and realizes continuous optimization of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745557A_ABST
    Figure CN121745557A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent manufacturing, and discloses a demand-driven material demand planning system based on deep reinforcement learning and an optimization method. The system comprises a demand module which is used for collecting, processing and outputting related demands and supplier environment data; the demand-driven material demand planning module is used for generating an inventory replenishment plan, a production plan and a material purchasing plan; the simulation module is used for constructing and operating a simulation model so as to output a simulation experiment statistical result; the deep reinforcement learning module interacts with the demand-driven material demand planning module and the simulation module, carries out training by utilizing a built-in algorithm library, and outputs an optimization parameter set of a demand-driven material demand plan; and the execution module is used for converting the plan information into an operation instruction and feeding back actual operation data and an execution result so as to make up for the deficiency of static parameter configuration of the traditional demand-driven material demand plan, realize collaborative optimization of inventory level and service rate and enhance the response capability of an enterprise to market demand change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent manufacturing technology, specifically to a demand-driven material requirements planning system and optimization method based on deep reinforcement learning. Background Technology

[0002] Material Requirements Planning (MRP), as a demand forecasting-based planning tool, relies heavily on stable and accurate forecast inputs for its effectiveness, making it more suitable for large-scale, standardized production models. However, in today's market environment with significantly increased demand uncertainty, traditional MRP can easily lead to a mismatch between inventory levels and actual demand, resulting in either excess inventory or material shortages. Demand-Driven MRP is a new planning and execution methodology designed to overcome these limitations. It integrates MRP principles, lean manufacturing, and the theory of constraints, achieving decoupling between demand and supply by setting dynamic inventory buffers at key points, thereby improving the responsiveness of the manufacturing system and enhancing the stability of internal material flow.

[0003] The core of demand-driven material requirements planning (DRP) lies in the management of dynamic buffers. Maximizing its effectiveness depends on setting several key performance indicators (SPIs), such as decoupling lead time factors, demand variability factors, peak order thresholds, and peak order horizons. The proper configuration of these SPIs has a crucial impact on the system performance of DRP.

[0004] Currently, the parameter settings for demand-driven material requirements planning (DRP) largely rely on empirical rules and industry benchmarks, making it difficult to adapt to complex and dynamic environments. When factors such as market demand, supply lead times, and product structure change, static parameter configurations often lead to a decline in system performance. Therefore, how to achieve dynamic and intelligent optimization of DRP parameters to adapt to the ever-changing environment is a key issue that urgently needs to be addressed in the current application and promotion of DRP. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a demand-driven material requirements planning (MRP) system and optimization method based on deep reinforcement learning. It aims to overcome the shortcomings of existing demand-driven MRP parameter setting methods, enabling intelligent adjustment of key parameters based on dynamically changing market and internal enterprise environments. This improves the overall operational efficiency of the system, significantly reducing inventory costs and enhancing customer satisfaction, and also provides new ideas for intelligent planning by enterprises.

[0006] To achieve the above objectives, according to one aspect of the present invention, a demand-driven material requirements planning system based on deep reinforcement learning is provided, comprising a demand module, a demand-driven material requirements planning module, a deep reinforcement learning module, a simulation module, and an execution module; The demand module is used to collect and process known demand data and supplier environment data, generate demand forecast data, and output comprehensive demand data to the demand-driven material requirements planning module and simulation module. The demand-driven material requirements planning module generates inventory replenishment plans, production plans, and material procurement plans based on comprehensive demand data obtained from the demand module, key parameters obtained from the deep reinforcement learning module, and basic replenishment strategies. It then outputs the generated plan data to the simulation module and the execution module. Simultaneously, it receives execution results from the execution module to adjust the plan formulation. The simulation module, based on the comprehensive demand data obtained from the demand module, the planned data obtained from the demand-driven material requirements planning module, the actual execution data obtained from the execution module, and the bill of materials and equipment layout information, builds and runs a simulation model, outputs the statistical results of the simulation experiment, and interacts with the deep reinforcement learning module. The deep reinforcement learning module interacts with the demand-driven material requirements planning module and the simulation module. It is trained using a deep reinforcement learning algorithm library and outputs an optimized parameter set to the demand-driven material requirements planning module as the basis for planning optimization and updates. The execution module receives planning data from the demand-driven material requirements planning module, transforms it into executable work instructions, and issues them to the actual production, logistics, and inventory processes. It also feeds back the actual operation data and execution results generated during the execution process to the simulation module and the demand-driven material requirements planning module.

[0007] Preferably, the known demand data includes historical demand data and short-term confirmed demand data; the supplier environment data includes average lead time, minimum order quantity, pricing strategy, and supply stability data; and the demand forecast data includes quantitative demand forecast data and qualitative demand forecast data.

[0008] Preferably, the generated inventory replenishment plan is determined based on average daily usage, buffer zone size, net flow location, and inventory level; the production plan includes production quantity, production time, and material consumption; and the material procurement plan includes procurement quantity and procurement date.

[0009] Preferably, the simulation model includes a production model, a logistics model, and an inventory model; the statistical results of the simulation experiment include inventory level, on-time delivery rate, and stockout quantity.

[0010] Preferably, the deep reinforcement learning algorithm library includes a dual deep Q-network algorithm, a deep deterministic policy gradient algorithm, a dominant actor-critic algorithm, and a proximal policy optimization algorithm; the optimization parameter set includes a demand variation factor and a lead time factor for calculating the buffer partition size, as well as order peak threshold and order peak horizon parameters related to net flow location calculation.

[0011] Preferably, the deep reinforcement learning module is implemented as follows: a deep reinforcement learning agent is designed and trained. The agent performs multiple rounds of trial and error learning in a simulation model constructed and run by the simulation module. At each time step, the agent selects an action based on the current state. The demand-driven material demand planning module and the simulation module receive and execute the action, and then provide feedback on the next state and reward. The agent uses the collected experience data to optimize its internal policy network or value network through a deep reinforcement learning algorithm to learn a parameter-dynamically adjusted strategy that can maximize the accumulated expected reward.

[0012] Preferably, the status information includes net flow location, inventory level, quantity of purchases in transit, and average daily usage; the action is a fixed value or relative adjustment of the parameter to be optimized; the reward is determined by a weighted combination of one or more of the three key performance indicators: average inventory, service rate, and inventory shortage cost.

[0013] Preferably, the actual operational data includes production information, logistics information, and inventory information; the execution results include production progress achievement rate, logistics plan completion status, and ending inventory status.

[0014] Preferably, the production information includes real-time production progress, equipment failure rate, and finished product yield; the logistics information includes material movement tracking information, purchase order tracking information, and logistics anomaly information; and the inventory information includes inventory consumption, inventory replenishment, and inventory warning.

[0015] According to another aspect of the present invention, the present invention provides an optimization method for demand-driven material requirements planning based on deep reinforcement learning, comprising the following steps: Step S01: When a new order arrives, the demand-driven material requirement update process is initiated based on the order characteristics and preset cycle. Step S02: Invoke the execution module to load actual operating data and the initial plan; Step S03: Call the requirements module to update the comprehensive requirements data; Step S04: Call the simulation module, run the simulation model and obtain the simulation experiment statistics based on the comprehensive requirements data, actual operation data and initial plan; Step S05: Call the deep reinforcement learning module to interact and train with the demand-driven material requirements planning module and the simulation module until the preset training termination condition is reached, and output the final optimized parameter set; Step S06: Call the demand-driven material requirements planning module to generate updated inventory replenishment plans, production plans, and material procurement plans based on the final optimized parameter set; Step S07: Execute the updated plan; Step S08: When the system status changes, restart the demand-driven material requirements planning update process.

[0016] In summary, the technical solutions conceived by this invention have the following beneficial effects compared with the prior art: 1. This invention introduces a deep reinforcement learning mechanism, which enables the demand-driven material requirements planning system to learn and dynamically adjust key parameters based on real-time changes in market demand, supplier performance, and internal enterprise operating status. This overcomes the shortcomings of traditional static parameter setting methods, such as slow response and difficulty in adapting to complex environments, and significantly improves the adaptive response capability of the demand-driven material requirements planning system to environmental changes and the level of intelligence in planning decisions.

[0017] 2. This invention significantly reduces the burden on planners to manually adjust and maintain the demand-driven material requirements planning system through automated iterative learning and optimization, reduces bias caused by subjective judgment, and improves the scientific nature of parameter configuration and decision-making efficiency.

[0018] 3. This invention, based on a learning process using simulation and actual operational data, provides a data-driven and quantifiable scientific method for configuring and continuously optimizing key parameters in a demand-driven material requirements planning (MRP) system. Through continuous monitoring of system performance and a model retraining mechanism, closed-loop continuous optimization of the application effect of the demand-driven MRP system can be achieved, making it widely applicable to different types of production and manufacturing environments. Attached Figure Description

[0019] Figure 1 This invention provides a schematic diagram of the structure of a demand-driven material requirements planning system based on deep reinforcement learning. Figure 2 This invention provides an integrated data flow diagram of a demand-driven material requirements planning system based on deep reinforcement learning. Figure 3 A flowchart for creating an inventory replenishment plan for the demand-driven material requirements planning module; Figure 4 This is a schematic diagram illustrating the interaction between the reinforcement learning agent and the simulation environment of this invention. Figure 5The present invention provides a flowchart of a demand-driven material requirements planning optimization method based on deep reinforcement learning. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0021] This invention provides a demand-driven material requirements planning (MRP) system based on deep reinforcement learning, comprising a demand module, a demand-driven MRP planning module, a deep reinforcement learning module, a simulation module, and an execution module. By constructing a closed-loop system integrating demand information processing, core logic of demand-driven MRP planning, dynamic simulation environment, intelligent optimization through deep reinforcement learning, and execution and feedback, it can achieve intelligent and adaptive dynamic adjustment of key parameters of demand-driven MRP planning.

[0022] The demand module is used to collect and process known demand data and supplier environment data to generate demand forecast data, and output comprehensive demand data to the demand-driven material demand planning module and simulation module. The known demand data includes historical demand data and short-term confirmed demand data. The supplier environment data includes average lead time, minimum order quantity, pricing strategy and supply stability data. The demand forecast data includes quantitative demand forecast data and qualitative demand forecast data. The demand-driven material requirements planning (MRP) module generates inventory replenishment plans, production plans, and material procurement plans based on comprehensive demand data obtained from the demand module, key parameters obtained from the deep reinforcement learning module, and basic replenishment strategies. It then outputs the generated plan data to the simulation and execution modules. Simultaneously, it receives execution results from the execution module to adjust the plan formulation. The generated inventory replenishment plan is determined based on average daily usage, buffer zone size, net flow location, and inventory level. The production plan includes production quantity, production time, and material consumption. The material procurement plan includes purchase quantity and purchase date. The simulation module, based on comprehensive demand data obtained from the demand module, planned data obtained from the demand-driven material requirements planning module, actual execution data obtained from the execution module, and bill of materials and equipment layout information, constructs and runs a simulation model, outputs simulation experiment statistical results, and interacts with the deep reinforcement learning module; the simulation model includes a production model, a logistics model, and an inventory model; the simulation experiment statistical results include inventory level, on-time delivery rate, and stockout quantity; The deep reinforcement learning module interacts with the demand-driven material requirements planning (MRP) module and the simulation module. It is trained using a deep reinforcement learning algorithm library and outputs an optimized parameter set to the demand-driven MRP module as the basis for plan formulation and optimization updates. The deep reinforcement learning algorithm library includes a dual deep Q-network algorithm, a deep deterministic policy gradient algorithm, a dominant actor-commentator algorithm, and a proximal policy optimization algorithm. The optimized parameter set includes a demand variation factor and a lead time factor for calculating the buffer partition size, as well as order peak thresholds and order peak horizon parameters related to net flow location calculation. The specific implementation method of the deep reinforcement learning module is as follows: a deep reinforcement learning agent is designed and trained. This agent performs multiple rounds of trial and error learning in a simulation model constructed and run by the simulation module. At each time step, the agent selects an action based on the current state. The demand-driven material requirements planning module and the simulation module receive and execute the action, thereby providing feedback on the next state and reward. The agent uses the collected experience data to optimize its internal policy network or value network through a deep reinforcement learning algorithm to learn a parameter dynamic adjustment strategy that maximizes the accumulated expected reward. The state information includes net flow location, inventory level, in-transit procurement quantity, and average daily usage. The action is a fixed value or relative adjustment of the parameter to be optimized. The reward is determined by a weighted combination of one or more of the three key performance indicators: average inventory, service rate, and inventory shortage cost. The execution module receives planning data from the demand-driven material requirements planning module, transforms it into executable work instructions, and issues them to the actual production, logistics, and inventory stages. It also feeds back the actual operational data and execution results generated during execution to the simulation module and the demand-driven material requirements planning module. The actual operational data includes production information, logistics information, and inventory information. The execution results include production progress achievement rate, logistics plan completion status, and ending inventory status. The production information includes real-time production progress, equipment failure rate, and finished product yield. The logistics information includes material movement tracking information, purchase order tracking information, and logistics anomaly information. The inventory information includes inventory consumption, inventory replenishment, and inventory warnings.

[0023] The core of demand-driven material requirements planning (DRP) lies in the management of its dynamic buffer. Maximizing its effectiveness depends on the setting of several key performance indicators (SPIs), such as the decoupling lead time factor, demand variability factor, peak order threshold, and peak order horizon. The proper configuration of these SPIs has a crucial impact on the system performance. Existing DRP systems have significant shortcomings in terms of intelligent and adaptive parameter adjustment. Parameter settings largely rely on empirical rules and industry benchmarks, making it difficult to maintain optimal performance in complex and dynamic environments, thus hindering the realization of its application potential.

[0024] Specifically, such as Figure 1 This invention provides a schematic diagram of a demand-driven material requirements planning (MRP) system based on deep reinforcement learning. The system consists of five collaborative modules: a demand module, a demand-driven MRP planning module, a simulation module, a deep reinforcement learning module, and an execution module. Specifically, the demand module, as the system's external information input, is responsible for collecting and processing various demand and environmental data, while simultaneously generating demand forecast data. The demand-driven MRP planning module formulates inventory replenishment plans, production plans, and material procurement plans based on comprehensive demand data and parameter configuration strategies. The simulation module is tightly coupled with the deep reinforcement learning module to achieve intelligent optimization of the system—the simulation module provides the environment and statistical results of simulation experiments for evaluating system performance, while the deep reinforcement learning module, through interactive learning with the demand-driven MRP planning module and the simulation module, outputs optimized demand-driven MRP parameter configuration strategies. The execution module transforms the generated plans into executable work instructions and provides feedback on actual operation data and execution results.

[0025] This system is not a linear, one-way process. The demand module, simulation module, deep reinforcement learning module, and execution module revolve around the core demand-driven material requirements planning module, forming a closed-loop optimization and execution process. This enables the system to adapt to the dynamic changes in external demand and internal operations, and continuously improve the operational efficiency of demand-driven material requirements planning.

[0026] (1) System functional modules like Figure 2 This is a schematic diagram of the integrated data flow of the demand-driven material requirements planning system based on deep reinforcement learning provided by the present invention. The following will describe the specific functional positioning, data processing, and data interaction flow of the system's five modules: the demand module, the demand-driven material requirements planning module, the simulation module, the deep reinforcement learning module, and the execution module.

[0027] (1.1) Requirements Module In this system, the demand module is positioned as the main external information input and preprocessing interface. It is used to collect and process known demand data and supplier environment data related to material demand, and generate demand forecast data, providing a structured and comprehensive demand information flow for subsequent demand-driven material demand planning and dynamic simulation.

[0028] Known demand data has a high degree of certainty and is an important basis for short-term planning. It mainly includes two types of data: First, historical demand data, such as data records covering the past 12 to 36 months extracted from the company's existing enterprise resource planning system and sales order management system, which should include information such as material codes, order dates and quantities; Second, short-term confirmed demand data, which comes directly from demand orders that have been formally confirmed and committed to by customers in the near future, and should include information such as order number, ordered material code, required quantity, delivery date, order priority and customer importance.

[0029] Supplier environment data mainly includes four types of data: First, average lead time, defined as the average time interval from the formal issuance of a purchase order to the material passing quality inspection and being ready for production, which can be obtained based on statistical analysis of historical purchase orders and delivery records; second, minimum order quantity, which refers to the minimum order quantity or minimum order amount set by the supplier for a purchase order, and purchase demands below this value need to be accumulated or combined; third, pricing strategy, including tiered pricing for different purchase quantity ranges, bulk discounts for one-time large purchases, and prices for long-term cooperation agreements; and fourth, supply stability data, which is obtained through comprehensive evaluation of multiple dimensions such as historical on-time delivery rate, delivery quantity accuracy rate, and incoming material quality pass rate.

[0030] Demand forecasting data is an important reference for medium- and long-term planning. It mainly includes two types of data: First, quantitative demand forecasting data, which is based on collected and verified historical demand data and uses built-in time series analysis methods, such as exponential smoothing, to capture the baseline level, long-term trends, and seasonal characteristics of the data to generate quantitative demand forecasting data. Second, qualitative demand forecasting data, which is generated by systematically collecting qualitative market intelligence through standardized methods, such as periodic expert questionnaires or regular cross-departmental demand review meetings, and using predefined adjustment rules based on domain knowledge to make necessary corrections to the quantitative forecasting data.

[0031] The demand module collects, processes, and generates the above demand forecast data. Finally, it structures and standardizes the demand forecast data and outputs comprehensive demand data, which is then transmitted to the demand-driven material requirements planning module and simulation module in the system.

[0032] (1.2) Demand-driven Material Requirements Planning Module The demand-driven material requirements planning (MRP) module is the core planning engine of the system. Based on the theoretical framework of demand-driven MRP and basic replenishment strategies, it combines comprehensive demand data obtained from the demand module and key parameters obtained from the deep reinforcement learning module to generate inventory replenishment plans, production plans, and material procurement plans. These plans are then output to the simulation module and the execution module. Simultaneously, it receives execution results from the execution module to adjust the plan formulation.

[0033] The inputs to the demand-driven material requirements planning (MRP) module include comprehensive demand data obtained from the demand module, key parameters obtained from the deep reinforcement learning module, plan execution status feedback from the execution module, and the theoretical framework and basic replenishment strategies of demand-driven MRP.

[0034] Based on the above input information, the demand-driven material requirements planning module performs calculations to generate inventory replenishment plans, production plans, and material procurement plans.

[0035] Inventory replenishment planning targets pre-selected strategic decoupling points for materials, such as those where multiple upstream components converge into a single common semi-finished product in the BOM structure. This plan is based on average daily usage, buffer zone size, net flow location, and inventory levels, and primarily includes, for example,... Figure 3 The process content is shown.

[0036] Calculate the average daily usage of materials: Based on the configured average daily usage calculation method, such as based on historical 12-month demand data, or using the weighted average of historical demand data and future forecast demand data, calculate the average daily usage of materials at each decoupling point.

[0037] Calculate the inventory buffer zone size: Based on the average daily usage and the demand variation factor and lead time factor obtained from the deep reinforcement learning module, calculate the size of the green, yellow, and red zones of the inventory buffer for each decoupling point material according to the following formula. The red zone includes the red baseline zone and the red safety zone.

[0038] Green Zone Max (average daily usage) Decoupling lead time Lead time factor, minimum order quantity Yellow Zone Average daily usage Decoupling lead time Red Zone Red reference area Red Safety Zone Red reference area Average daily usage Decoupling lead time Lead time factor Red Safety Zone Average daily usage Decoupling lead time Lead time factor Demand variation factor Calculate the net flow location: Based on the material inventory information and the two parameters of the order peak threshold and the order peak horizon obtained from the deep reinforcement learning module, calculate the net flow location of the decoupling point material according to the following formula.

[0039] Net flow location Inventory on hand Inventory of items for which replenishment orders have been placed but have not yet arrived Qualified order requirements In this context, qualified order demand equals the sum of three parts: previously unfulfilled orders, orders due today, and reasonable peak demand within the time window. Unfulfilled orders and orders due today should be fulfilled immediately, while the reasonable peak order demand within the time window relates to future demand that should be met. The reasonable peak order demand within the time window is determined by two parameters: the peak order threshold and the peak order horizon. Orders exceeding the peak order threshold within the peak order horizon are considered reasonable peak order demand within the time window.

[0040] Develop an inventory replenishment plan: A replenishment signal is triggered when the net flow position of the material at the decoupling point drops below the top of its yellow zone. The formula for calculating the replenishment quantity is as follows: Replenishment quantity Top of the green zone of the inventory buffer zone Current net flow position Based on this, the module will also consider the supplier's minimum order quantity or the internal minimum production batch and adjust the replenishment quantity to generate the final inventory replenishment plan. The production plan targets materials that require internal production. This plan translates inventory replenishment needs into specific production tasks, refers to the bill of materials structure, determines the consumption of materials at each level required for product production, and clarifies the planned production quantity, planned start time, and planned completion time.

[0041] The material procurement plan targets materials that require external procurement or raw materials involved in the production plan. This plan specifies the quantity of materials to be procured and the purchase order placement date, calculated backwards from the material demand time and supplier lead time. It is also adjusted based on information such as minimum order quantity and pricing strategy input from the demand module.

[0042] After completing the above-mentioned plan, the demand-driven material requirements planning module outputs it to the simulation and execution modules. In the simulation module, it serves as the basis for simulating operations and is used to evaluate the system's performance under the current parameter configuration; in the execution module, it serves as the basis for guiding actual production, logistics, and inventory operations.

[0043] Meanwhile, the demand-driven material requirements planning module also receives feedback on plan execution from the execution module. This feedback information is used to dynamically adjust the current plan, forming a closed-loop feedback between planning and execution. This allows for better adaptation to changes in actual operations and improves the accuracy of the plan.

[0044] (1.3) Simulation Module The core function of the simulation module is to construct a dynamic simulation model that closely approximates the actual operating environment of an enterprise, and to simulate the long-term operation of a demand-driven material requirements planning (MRP) system on this model. Based on comprehensive demand data obtained from the demand module, planned data obtained from the demand-driven MRP module, actual execution data obtained from the execution module, and bill of materials and equipment layout information, the model constructs and runs the simulation, outputs statistical results of the simulation experiments, and interacts with the deep reinforcement learning module.

[0045] The inputs to the simulation module include comprehensive demand data obtained from the demand module, planning data obtained from the demand-driven material requirements planning module, actual execution data, and bill of materials and equipment layout information obtained from the execution module.

[0046] The simulation module constructs a simulation model that includes three interconnected and collaborative sub-models: a production model, a logistics model, and an inventory model.

[0047] The production model is responsible for simulating the internal production and manufacturing process of an enterprise. Based on the input production plan, it can comprehensively consider the capacity constraints, processing sequence, changeover time, work-in-process buffer limits of each production unit, as well as factors such as equipment failure, quality problems leading to rework or scrap, and dynamically simulate the real-time production progress of products and the occupancy of equipment resources. It also calculates production performance indicators such as equipment failure rate and finished product yield.

[0048] The logistics model is responsible for simulating the physical flow of materials in the enterprise's production operations, including external procurement logistics and internal production logistics. For external procurement logistics, it simulates the entire process from the issuance of a purchase order to the supplier's response, preparation of goods, transportation, and final receipt and warehousing. For internal production logistics, it simulates the distribution of raw materials from the warehouse to the production line, the transfer of semi-finished products between different processes, and the flow of finished products from the production line to the finished goods warehouse. It considers factors such as the capacity of transportation tools like automated guided vehicles, batch rules, and route planning, and is also responsible for modeling and tracking various logistics anomalies.

[0049] The inventory model is responsible for tracking and managing the inventory levels of all materials in the simulation environment in real time. It can dynamically update the inventory quantity of each buffer material based on the material consumption of the production model, the material arrival of the logistics model, and order fulfillment, and generate inventory warning signals based on preset thresholds.

[0050] The simulation module drives the collaborative operation of the aforementioned sub-models through its internal simulation engine. Within a set simulation duration, it progresses according to a preset time step, simulating the performance of the demand-driven material requirements planning system under specific parameter configurations and dynamic demand inputs. During the simulation, the module continuously collects, calculates, and outputs simulation experiment statistical results as learning feedback for the deep reinforcement learning module. These output statistical results include inventory levels, on-time delivery rate, and stockout quantity.

[0051] (1.4) Deep Reinforcement Learning Module The deep reinforcement learning module is the core module for realizing intelligent and adaptive optimization of material demand planning parameters driven by demand. By interacting with the demand-driven material demand planning module and the simulation module, it uses the built-in deep reinforcement learning algorithm library for training, explores near-optimal parameter combination strategies to improve the overall performance of the system, and outputs the optimized parameter set to the demand-driven material demand planning module as the basis for planning optimization and updates.

[0052] To support optimization problems of varying complexity and different parameter spaces, the deep reinforcement learning module includes a configurable deep reinforcement learning algorithm library, which contains the following algorithms: First, there is the dual deep Q-network algorithm, which is suitable for optimization in continuous discrete spaces and alleviates the overestimation problem of traditional deep Q-networks by decoupling the selection and evaluation of the target Q-value. Second, there is the deep deterministic policy gradient algorithm, which is suitable for optimization in continuous action spaces. Third, there is the dominant actor-critic algorithm, which reduces the variance of the policy gradient by introducing a dominant function. Fourth, there is the proximal policy optimization algorithm, which improves the stability and sample efficiency of training by limiting the magnitude of policy updates.

[0053] Operators can select a matching algorithm framework based on the characteristics of the parameters to be optimized and the constraints of computing resources; agent design will be carried out in conjunction with algorithm selection, and it is necessary to define the action space, state space and reward function.

[0054] The action space defines the interventions an agent can make in a demand-driven material requirements planning (MRP) system, i.e., which parameters to adjust and how to adjust them. In this invention, the action space directly corresponds to the set of parameters to be optimized determined in the demand-driven MRP module, including four parameters: demand variability factor, lead time factor, order peak threshold, and order peak horizon. The demand variability factor and lead time factor affect the size of the green and red zones in the inventory buffer partition; the order peak threshold and order peak horizon affect the calculation of the net flow position, specifically in determining the reasonable order peak within the time window, influencing the size of qualified order demand. The agent's action can be to output new fixed values ​​for these parameters, or to output relative adjustments to the current values ​​of these parameters, such as increasing or decreasing by a certain percentage. When multiple parameters are optimized simultaneously, the action space is a multi-dimensional vector.

[0055] The state space is a quantitative representation of the current operating status of the demand-driven material requirements planning system in the simulation environment, including: the current net flow position of materials at each strategic decoupling point; the current actual inventory level; the quantity of procurement or production replenishment orders in transit; and the currently calculated average daily usage.

[0056] The reward function is calculated based on one or more key performance indicators (KPIs) obtained from the simulation module. It aims to quantify the actions performed by the agent, i.e., the contribution of parameter adjustments to the enterprise's pre-defined demand-driven material requirements planning (MRP) optimization goals. These KPIs include average inventory on hand, service rate, and inventory shortage cost. In this invention, the reward function is a configurable function that can be weighted and combined according to the enterprise's strategic priorities at different times or specific optimization goals for different product lines. For example, when the enterprise's goal is to minimize total costs while maintaining a high service level, the reward function can be designed as follows:

[0057] in, It is the reward value at the current time step; This is the actual service rate for the current period. This is the target service rate, which is used to incentivize agents to achieve service targets. It is the normalized average inventory on hand for the current period; This is the normalized inventory shortage in the current cycle; The weighting coefficients corresponding to service level, inventory cost, and stockout cost can be adjusted according to the company's objectives.

[0058] like Figure 4As shown, the training of the deep reinforcement learning agent is accomplished through interactive learning. The agent performs multiple rounds of trial-and-error learning in a simulation model built and run by the simulation module. At each time step, the agent selects an action based on its current internal policy network and the current state information obtained from the simulation module, i.e., outputting a set of demand-driven material demand planning (MDP) parameters. The MDP module and the simulation module receive and execute this action, using these new parameters to drive its internal simulation model to run one or more planning cycles. After the simulation runs, the simulation module collects relevant data and calculates immediate rewards based on performance within that cycle, feeding both pieces of information back to the agent. The agent stores the experiential data generated from this interaction in an experience replay pool. By sampling data from the experience replay pool and utilizing specific update rules of a pre-selected deep reinforcement learning algorithm, it optimizes the parameters of its internal policy network or value network. The entire interactive training process continues until the agent's performance reaches a preset convergence criterion or the specified number of training rounds is completed.

[0059] After optimizing the parameter set, the deep reinforcement learning module outputs it to the demand-driven material requirements planning module as the basis for planning optimization and updates.

[0060] (1.5) Execution Module In this optimization system, the execution module plays the role of connecting the planning layer and the actual application layer. It transforms the planning information received from the demand-driven material requirements planning module into executable work instructions and issues them to the actual production, logistics and inventory links. It also feeds back the actual operation data and execution results generated during the execution process to the simulation module and the demand-driven material requirements planning module.

[0061] The inputs to the execution module include inventory replenishment plans for materials at each strategic decoupling point, production plans for self-made parts, and material procurement plans for purchased parts. This planning information is the direct basis for the execution module to perform subsequent operations.

[0062] Upon receiving the planning information, the core task of the execution module is to transform it into executable work instructions and distribute them to the enterprise's actual production, logistics, and inventory management processes. To achieve this function, the execution module is configured to interface with the enterprise's existing information systems, such as manufacturing execution systems, warehouse management systems, and procurement management systems.

[0063] For production planning, the execution module breaks it down into specific production work orders. Each work order defines in detail the material requirements, process routes, planned start and completion times, and the production resources to be allocated. This work order information is then transmitted to the manufacturing execution system for production scheduling and execution. For material procurement planning, the execution module converts it into purchase requisitions or directly generates purchase orders, which are then transmitted to the procurement management system for subsequent supplier interaction, order confirmation, and logistics tracking. For inventory replenishment planning, the execution module generates inventory transfer instructions, which are then transmitted to the warehouse management system for in-warehouse operations.

[0064] During the execution of the plan, the execution module collects and summarizes actual operational data in real time or periodically. This data reflects the actual execution of the plan and the true state of the operating environment, and is an important foundation for achieving closed-loop optimization. Actual operational data mainly includes three categories: production information, logistics information, and inventory information.

[0065] Production information records the actual performance of the manufacturing process, including the actual start and end times of each production work order, the actual output quantity of each process, the real-time quantity of work-in-process at each workstation, the actual consumption of materials during the production process, the actual running time of equipment, the number and duration of equipment downtime due to malfunctions, and the final product yield and corresponding scrap rate of each production batch.

[0066] Logistics information tracks the physical flow of materials, including records of material movement between various storage points and workstations within the factory; actual arrival batches of purchase orders, the type and quantity of materials in each batch, and the precise arrival time; and records of various logistics anomalies that occur throughout the logistics process, such as transportation delays and damage to goods.

[0067] Inventory information reflects the dynamic changes in the inventory level of each material, including real-time records of in-stock inventory consumption; actual inventory replenishment records; and various preset inventory warning states triggered by the current inventory level, such as insufficient warnings when the inventory level is below the safety threshold or overstock warnings when the inventory level is above the upper limit threshold.

[0068] The execution module processes and integrates the collected actual operational data and phased execution results, including the overall production progress achievement rate, logistics plan completion status, and actual inventory status at the end of a planning cycle. This data is used for real-time performance monitoring and anomaly management, and also fed back to the simulation module and the demand-driven material requirements planning (MRP) module. The data fed back to the simulation module can be used to calibrate and verify the internal parameters of the simulation model, thereby continuously improving the fit between the simulation environment and the actual operating environment. The data fed back to the demand-driven MRP module can be used to evaluate the effectiveness of the current demand-driven MRP parameters and plan, identify and analyze the deviations and causes between the plan and actual execution, and dynamically adjust the plan to achieve rolling optimization.

[0069] (2) Optimization method Based on the demand-driven material requirements planning system based on deep reinforcement learning proposed in this invention, this invention also proposes a specific method for optimizing this system, such as... Figure 5 As shown, the optimization method includes the following steps: Step S01: When a new order arrives, the demand-driven material requirement update process is initiated based on the order characteristics and preset cycle. This optimization method does not run continuously; instead, it is activated by preset trigger conditions to balance optimization benefits with computational resource consumption. The process is initiated by the arrival of a new order as the basic event signal, but whether to respond immediately depends on the analysis of the order's characteristics. Order characteristic analysis is a multi-dimensional comprehensive evaluation process, including at least: the order's scale characteristics, such as whether its total amount or quantity exceeds preset thresholds; the order's priority characteristics, such as whether it comes from a strategic customer requiring priority support; and the degree of impact on the demand for strategic decoupling point materials after decomposition through the bill of materials. Based on this evaluation result, the process initiation logic is as follows: if the comprehensive characteristics of the new order are determined to have a significant impact on the existing parameter configuration, the optimization update process is immediately initiated to ensure a rapid response. Conversely, the order is included in the regular demand category, its information is cached and accumulated until the preset planned update time node is reached, at which point the routine optimization process is initiated uniformly.

[0070] Step S02: Invoke the execution module to load actual operating data and the initial plan; After triggering the update process, the initial data preparation and baseline establishment phase begins. This step involves calling the system's execution module to obtain actual operational data, including three types of data: production information, logistics information, and inventory information. Simultaneously, it loads the previous cycle's inventory replenishment plan, production plan, and material procurement plan, along with their parameter configurations, as the reference benchmark and initial conditions for this round of optimization.

[0071] Step S03: Call the requirements module to update the comprehensive requirements data; After establishing the initial operational status, the system enters the demand information update phase. This step invokes the system's demand module to execute its comprehensive data collection, processing, and generation functions. The demand module integrates all new order data and supplier environment information since the last update, and re-runs and calibrates its internal demand forecasting model to generate updated comprehensive demand data.

[0072] Subsequently, the optimization method enters its core phase—the simulation evaluation and learning optimization stage. Step S04: Call the simulation module, run the simulation model and generate simulation experiment statistical results based on the comprehensive requirements data, actual operation data and initial plan; The process involves invoking the simulation module, taking the updated comprehensive demand data, the loaded initial operating state, and the parameters from the previous cycle's demand-driven material requirements planning (DRP) as input. It then simulates the performance of the DRP system under the current unoptimized parameter configuration, outputting simulation experimental statistics that include key performance indicators such as inventory levels, on-time delivery rate, and stockout levels. These statistics establish the performance benchmark for this round of optimization and provide state information and environmental feedback for the subsequent deep reinforcement learning training process.

[0073] Step S05: Call the deep reinforcement learning module to interact and train with the demand-driven material requirements planning module and the simulation module until the preset training termination condition is reached, and output the final optimized parameter set; The process invokes the deep reinforcement learning module to perform intelligent parameter optimization. The deep reinforcement learning agent in this module obtains the current system state from the simulation module and selects an action based on its internal policy network or value network, thus outputting a new set of parameters. The demand-driven material requirements planning module and the simulation module receive these parameters and rerun the simulation model based on this configuration. They then feed back the new system state and the reward value calculated based on the simulation results to the agent. The agent uses the experiential data generated through this interaction to update its internal policy network or value network.

[0074] The iterative learning process described above will continue until the preset training termination conditions are met, such as the average cumulative reward converging and stabilizing, or the preset maximum number of training epochs being reached. After training is complete, the deep reinforcement learning module will output the final optimized parameter set.

[0075] Step S06: Call the demand-driven material requirements planning module to generate updated inventory replenishment plans, production plans, and material procurement plans based on the final optimized parameter set; After obtaining the optimized parameter set, the plan update phase begins. This step invokes the demand-driven material requirements planning module, which, based on the final optimized parameter set output by the deep reinforcement learning module and combined with the updated comprehensive demand data from the demand module, re-formulates the inventory replenishment plan, production plan, and material procurement plan according to its internal core computational logic.

[0076] Step S07: Execute the updated plan.

[0077] After the plan is updated, the step calls the execution module to convert it into executable work instructions and send them to the actual production, logistics and inventory stages.

[0078] Step S08: When the system status changes, restart the demand-driven material requirements update process.

[0079] Finally, the process enters the continuous feedback and closed-loop optimization phase. During the execution of the new plan, the execution module continuously monitors the actual operational data and status of key aspects such as production, logistics, and inventory. When a change in status that may affect overall performance is detected, the entire optimization process will restart. For example, a significant extension of the delivery cycle of a major supplier may cause a deviation from the existing plan. In this case, the system will re-analyze and simulate based on the latest data to generate new inventory replenishment plans, production plans, and material procurement plans. Conversely, if the system status remains stable for a period of time and all key performance indicators are within acceptable ranges, this round of optimization ends, and the system will continue to execute according to the current plan until the next triggering event occurs. Through continuous monitoring and feedback mechanisms, it can be ensured that the demand-driven material requirements system is always in an optimal operating state.

[0080] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A demand-driven material requirements planning system based on deep reinforcement learning, characterized in that, It includes a demand module, a demand-driven material requirements planning module, a deep reinforcement learning module, a simulation module, and an execution module; The demand module is used to collect and process known demand data and supplier environment data, generate demand forecast data, and output comprehensive demand data to the demand-driven material requirements planning module and simulation module. The demand-driven material requirements planning module generates inventory replenishment plans, production plans, and material procurement plans based on comprehensive demand data obtained from the demand module, key parameters obtained from the deep reinforcement learning module, and basic replenishment strategies. The generated plan data is then output to the simulation module and the execution module. Simultaneously, it receives execution results from the execution module to adjust the plan formulation; The simulation module, based on the comprehensive demand data obtained from the demand module, the planned data obtained from the demand-driven material requirements planning module, the actual execution data obtained from the execution module, and the bill of materials and equipment layout information, builds and runs a simulation model, outputs the statistical results of the simulation experiment, and interacts with the deep reinforcement learning module. The deep reinforcement learning module interacts with the demand-driven material requirements planning module and the simulation module. It is trained using a deep reinforcement learning algorithm library and outputs an optimized parameter set to the demand-driven material requirements planning module as the basis for planning optimization and updates. The execution module receives planning data from the demand-driven material requirements planning module, transforms it into executable work instructions, and issues them to the actual production, logistics, and inventory processes. It also feeds back the actual operation data and execution results generated during the execution process to the simulation module and the demand-driven material requirements planning module.

2. The demand-driven material requirements planning system based on deep reinforcement learning as described in claim 1, characterized in that, The known demand data includes historical demand data and short-term confirmed demand data; the supplier environment data includes average lead time, minimum order quantity, pricing strategy, and supply stability data; the demand forecast data includes quantitative demand forecast data and qualitative demand forecast data.

3. The demand-driven material requirements planning system based on deep reinforcement learning as described in claim 1, characterized in that, The generated inventory replenishment plan is determined based on average daily usage, buffer zone size, net flow location, and inventory level; the production plan includes production quantity, production time, and material consumption; the material procurement plan includes procurement quantity and procurement date.

4. The demand-driven material requirements planning system based on deep reinforcement learning as described in claim 1, characterized in that, The simulation model includes a production model, a logistics model, and an inventory model; the statistical results of the simulation experiment include inventory level, on-time delivery rate, and stockout quantity.

5. A demand-driven material requirements planning system based on deep reinforcement learning as described in claim 1, characterized in that, The deep reinforcement learning algorithm library includes the dual deep Q-network algorithm, the deep deterministic policy gradient algorithm, the dominant actor-critic algorithm, and the proximal policy optimization algorithm; the optimization parameter set includes the demand variation factor and lead time factor for calculating the buffer partition size, as well as the order peak threshold and order peak horizon parameters related to the net flow location calculation.

6. The demand-driven material requirements planning system based on deep reinforcement learning as described in claim 1, characterized in that, The specific implementation method of the deep reinforcement learning module is as follows: design and train a deep reinforcement learning agent. The agent performs multiple rounds of trial and error learning in the simulation model constructed and run by the simulation module. At each time step, the agent selects an action according to the current state. The demand-driven material demand planning module and the simulation module receive and execute the action, and then provide feedback on the next state and reward. The agent uses the collected experience data to optimize the internal policy network or value network through the deep reinforcement learning algorithm in order to learn a parameter dynamic adjustment strategy that can maximize the accumulated expected reward.

7. A demand-driven material requirements planning system based on deep reinforcement learning as described in claim 6, characterized in that, The status information includes net flow location, inventory level, quantity of purchases in transit, and average daily usage; the action is a fixed value or relative adjustment of the parameter to be optimized; the reward is determined by a weighted combination of one or more of the three key performance indicators: average inventory, service rate, and inventory shortage cost.

8. A demand-driven material requirements planning system based on deep reinforcement learning as described in claim 1, characterized in that, The actual operational data includes production information, logistics information, and inventory information; the execution results include production progress achievement rate, logistics plan completion status, and ending inventory status.

9. A demand-driven material requirements planning system based on deep reinforcement learning as described in claim 8, characterized in that, The production information includes real-time production progress, equipment failure rate, and finished product yield; the logistics information includes material movement tracking information, purchase order tracking information, and logistics anomaly information. The inventory information includes inventory consumption, inventory replenishment, and inventory warnings.

10. An optimization method for a demand-driven material requirements planning system based on any one of claims 1-9, characterized in that, Includes the following steps: Step S01: When a new order arrives, the demand-driven material requirement update process is initiated based on the order characteristics and preset cycle. Step S02: Invoke the execution module to load actual operating data and the initial plan; Step S03: Call the requirements module to update the comprehensive requirements data; Step S04: Call the simulation module, run the simulation model and obtain the simulation experiment statistics based on the comprehensive requirements data, actual operation data and initial plan; Step S05: Call the deep reinforcement learning module to interact and train with the demand-driven material requirements planning module and the simulation module until the preset training termination condition is reached, and output the final optimized parameter set; Step S06: Call the demand-driven material requirements planning module to generate updated inventory replenishment plans, production plans, and material procurement plans based on the final optimized parameter set; Step S07: Execute the updated plan; Step S08: When the system status changes, restart the demand-driven material requirements planning update process.

Citation Information

Cited By

  • Automobile assembly line dynamic periodic material distribution scheduling system

    CN122175305A