A reinforcement learning-based supply chain contingency replanning method
By combining generative models and reinforcement learning strategies, we can generate and evaluate supply chain emergency replanning schemes. This solves the problems of decision-making complexity and hard constraint violation in multi-stage coupling, achieves rapid and stable emergency replanning, and reduces execution risks and costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 四川文理学院
- Filing Date
- 2026-04-16
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies in supply chain emergency replanning have high decision-making dimensions and complex constraints that are coupled across multiple links such as procurement, production, warehousing and transportation. This makes it difficult to output high-quality solutions within emergency time limits. Reinforcement learning decision-making is prone to violating hard constraints and causing key performance indicators to fall below the bottom line. Rolling replanning is prone to causing frequent changes and fluctuations in plans.
The system constructs supply chain status data and historical replanning execution schemes, generates a limited set of candidate schemes, evaluates and ranks them through hard constraint feasibility projection and repair processing, and combines reinforcement learning strategy models. Under the control of uncertainty and key performance indicator thresholds, the final scheme is output, and the backoff solver is called to generate a backup scheme that meets the hard constraints and threshold requirements.
It enables the rapid acquisition of executable replanning solutions within emergency timeframes, reduces risk spillover, improves the feasibility and stability of the plan, and reduces manual correction and execution costs.
Smart Images

Figure REF-OBJ-1776241175464-000311 
Figure REF-OBJ-1776241175464-000312
Abstract
Description
Technical Field
[0001] This invention relates to the field of supply chain management and intelligent optimization, and in particular to a supply chain emergency replanning method based on reinforcement learning. Background Technology
[0002] Supply chain planning and scheduling are fundamental technologies in manufacturing and logistics. With the widespread adoption of information systems such as Enterprise Resource Planning (ERP), Advanced Planning and Scheduling (APS), Warehouse Management Systems (WMS), and Transportation Management Systems (TMS), companies typically develop procurement plans, production schedules, warehousing operation plans, and distribution plans based on data such as orders, inventory, capacity, and transportation resources, and continuously update these plans on a rolling basis. When the supply chain experiences sudden disruptions such as production stoppages, material delays, sudden demand changes, or transportation disruptions, the industry often employs methods such as manual experience and rule-based approaches, heuristic algorithms, mathematical programming methods (e.g., linear programming and mixed-integer programming), and simulation evaluation for emergency replanning. In recent years, intelligent decision-making methods based on machine learning and reinforcement learning have also been gradually introduced to improve response speed and global collaborative optimization capabilities.
[0003] However, existing technologies still have the following shortcomings: 1. Replanning problems that are coupled across multiple stages such as procurement, production, warehousing and transportation have high decision-making dimensions and complex constraints. Mathematical programming or heuristic algorithms are difficult to stably output high-quality solutions within emergency time limits, and the time consumed in solving or iterating increases significantly as the problem size increases.
[0004] 2. Under hard constraints such as production capacity, inventory, loading, delivery window, minimum order quantity, and contract lock-in, decision outputs based on rules or reinforcement learning are prone to producing infeasible solutions, requiring repeated manual corrections and affecting emergency response efficiency and feasibility.
[0005] 3. Strengthening learning: In actual business operations, relying on online trial and error can easily lead to key indicators such as stockout rates, penalties, and cash flow exceeding the bottom line under uncertain or abnormal data conditions. At the same time, rolling replanning can easily lead to frequent changes and fluctuations in plans, increasing execution costs such as line switching, reassignment, and rerouting.
[0006] Therefore, there is a need for a supply chain emergency replanning method that can address the shortcomings of existing technologies. Summary of the Invention
[0007] One objective of this invention is to propose a reinforcement learning-based supply chain emergency replanning method. Addressing the challenges of existing technologies where multi-stage replanning across procurement, production, warehousing, and transportation is difficult to complete within emergency timeframes under sudden disruptions, reinforcement learning decisions are prone to violating hard constraints such as capacity, inventory, loading, delivery windows, minimum order quantities, and contract lock-in, and online trial-and-error can easily lead to key performance indicators (KPIs) exceeding limits and causing frequent fluctuations in rolling plans, this invention proposes a method that constructs supply chain status data and historical replanning execution schemes during rolling replanning. A generative model generates a limited set of candidate replanning schemes, and feasible candidate sets are obtained through hard constraint feasibility projection and feasibility repair processing. A reinforcement learning strategy model is then used to evaluate and rank the feasible candidate sets to output predicted KPI values and plan change costs. Furthermore, a safety gating system is implemented based on uncertainty indicators and KPI threshold sets. If a threshold is not met, a fallback solver is invoked to generate a backup plan that meets the hard constraints and threshold requirements and is then executed. This invention offers the technical advantages of rapidly obtaining executable replanning schemes, reducing risk spillover, and balancing response speed and plan stability.
[0008] This invention provides a reinforcement learning-based supply chain emergency replanning method, comprising: S1. At each rolling replanning time, acquire supply chain business data and information on sudden disturbances to construct supply chain status data, and determine the replanning execution plan from the previous rolling replanning time as the historical replanning execution plan; S2. Based on the supply chain status data, information on sudden disturbances, and historical replanning execution plans, call the generative model to generate a set of candidate replanning plans, with the number of plans not exceeding a preset upper limit; S3. Perform feasibility processing on the set of candidate replanning plans based on a set of hard constraints to obtain a set of feasible candidate replanning plans that meet the hard constraints. Feasibility processing includes: performing feasibility projection processing and / or feasibility repair processing on candidate replanning plans that do not meet the hard constraints; S4. Combining the supply chain status data, information on sudden disturbances, historical replanning execution plans, and the set of feasible candidate replanning plans, use a reinforcement learning strategy model to process the feasible candidate replanning plans. The set of solutions is evaluated and ranked, and the optimal ranked solution is selected as the reinforcement learning recommended replanning solution. Evaluation results corresponding to the reinforcement learning recommended replanning solution are generated, including predicted values of key performance indicators (KPIs) and the cost of plan changes. S5: Based on the evaluation results, the uncertainty index of the reinforcement learning recommended replanning solution is calculated. Based on the KPI threshold set, it is determined whether the predicted KPI values meet the threshold requirements. A pass flag is generated when the uncertainty index does not exceed the preset safety threshold and the predicted KPI values meet the KPI threshold set; otherwise, a fail flag is generated. S6: If a pass flag is obtained, the reinforcement learning recommended replanning solution is determined as the final replanning execution solution; otherwise, the backoff solver is called to generate a backup replanning solution that meets both the hard constraint set and the KPI threshold set. S7: An execution instruction is generated based on the final replanning execution solution and issued for execution.
[0009] Optionally, S1 includes: At each rolling replanning time, the order demand data, inventory data, production capacity data, warehousing operation capacity data, transportation resource data, and the replanning execution plan from the previous rolling replanning time are collected, and the information on sudden disturbances is also collected. The collected data is timestamped and its consistency is verified. Missing or abnormal items are removed to form valid business data. Based on the valid business data and the sudden disturbance information, calculate status elements including available inventory, in-transit quantity, remaining capacity, remaining warehousing operation capacity, available transportation capacity, and order delivery urgency, and associate the status elements with the corresponding supply chain node identifiers and time identifiers to construct the supply chain status data; The replanning execution scheme at the previous rolling replanning moment is determined as the historical replanning execution scheme, and the supply chain status data, the sudden disturbance information, and the historical replanning execution scheme are output.
[0010] Optionally, S2 includes: Based on the supply chain status data, the sudden disturbance information, and the historical replanning execution scheme, the generated condition data is constructed. The generated condition data includes conditional elements that characterize order demand, inventory level, production capacity, warehousing operation capacity, transportation resources, and the scope of impact of sudden disturbances. Based on the generated condition data, the generated model is called multiple times to generate multiple candidate replanning schemes that do not exceed the preset upper limit of the number of candidates. The candidate replanning scheme set is formed by removing duplicate schemes from the multiple candidate replanning schemes. Each candidate replanning scheme is represented as a set of decision variables including the procurement decision variable, the production decision variable, the warehousing decision variable, and the transportation decision variable, and is associated with a corresponding time identifier.
[0011] Optionally, S3 includes: For each candidate replanning scheme in the set of candidate replanning schemes, calculate the constraint violation amount based on the set of hard constraints; When the constraint violation amount is zero, the corresponding candidate replanning scheme is marked as a feasible candidate replanning scheme; When the constraint violation amount is not zero, the feasibility projection processing or the feasibility repair processing is performed to adjust the procurement decision variables, production decision variables, warehousing decision variables and transportation decision variables in the candidate replanning scheme, and the constraint violation amount is recalculated after adjustment; When the amount of constraint violation after adjustment is zero, the adjusted candidate reprogramming scheme is marked as a feasible candidate reprogramming scheme; otherwise, the corresponding candidate reprogramming scheme is eliminated. After processing all candidate replanning schemes in the candidate replanning scheme set, all feasible candidate replanning schemes are summarized to form the feasible candidate replanning scheme set. Furthermore, the feasibility projection process includes: constructing a projection optimization problem with the objective function of minimizing the adjustment amount of the candidate replanning scheme and the constraint condition set of hard constraints, and solving the projection optimization problem to obtain the projected candidate replanning scheme. The adjustment amount is the weighted sum of the changes in the procurement decision variable, the production decision variable, the warehousing decision variable and the transportation decision variable before and after the projection.
[0012] Optionally, S4 includes: Based on the supply chain status data, the sudden disturbance information, and the historical replanning execution schemes, an evaluation object is constructed for each feasible candidate replanning scheme in the set of feasible candidate replanning schemes; A reinforcement learning strategy model is used to evaluate each evaluation object and generate predicted values of key performance indicators for each feasible candidate replanning scheme. Calculate the plan change cost corresponding to each feasible candidate replanning scheme based on the difference between each feasible candidate replanning scheme and the historical replanning execution scheme; Based on the predicted values of the key performance indicators and the cost of the plan change, an evaluation score is generated for each feasible candidate replanning scheme, and the ranking result is generated based on the evaluation score; Based on the ranking result, the optimal ranking scheme is determined from the set of feasible candidate replanning schemes as the reinforcement learning recommended replanning scheme, and an evaluation result corresponding to the reinforcement learning recommended replanning scheme is generated. The evaluation results include the predicted values of the key performance indicators and the cost of the plan changes.
[0013] Optionally, S5 includes: Based on the evaluation results, the predicted values of the key performance indicators (KPIs) corresponding to the reinforcement learning recommendation reprogramming scheme are obtained. The predicted KPIs are evaluated multiple times using an uncertainty evaluation method to obtain multiple sets of predicted KPIs. An uncertainty index is calculated based on these multiple sets of predicted KPIs, where the uncertainty index represents the dispersion of the predicted KPIs. A threshold judgment is performed on the predicted KPIs based on the KPI threshold set, where the stockout rate threshold and the penalty threshold are upper limit thresholds, and the cash flow threshold is a lower limit threshold. When the predicted stockout rate ≤ the stockout rate threshold, the predicted penalty ≤ the penalty threshold, and the predicted cash flow ≥ the cash flow threshold, the predicted KPIs are determined to satisfy the KPI threshold set. Otherwise, it is determined that the predicted value of the key performance indicator does not meet the key performance indicator threshold set; when the uncertainty indicator does not exceed the preset safety threshold and the predicted value of the key performance indicator meets the key performance indicator threshold set, the pass flag is generated; otherwise, the fail flag is generated. Furthermore, the threshold determination of the predicted value of the key performance indicator based on the key performance indicator threshold set includes: introducing a safety margin when determining the threshold, the safety margin being positively correlated with the uncertainty indicator; and determining that the predicted value of the key performance indicator satisfies the key performance indicator threshold set when the predicted value of the stockout rate, the predicted value of the penalty, and the predicted value of the cash flow respectively meet the threshold requirements after adding the safety margin.
[0014] Optionally, S6 includes: When the pass flag is true, the reinforcement learning recommended replanning scheme is determined as the final replanning execution scheme and output; When the failure flag is true, a rollback solution task is constructed based on the supply chain status data, the sudden disturbance information, and the historical replanning execution scheme. The rollback solution task includes the set of hard constraints and the set of key performance indicator thresholds. The rollback solver is invoked to solve the rollback task, generating a minimum replanning scheme that satisfies the set of hard constraints and the set of key performance indicator thresholds. The guaranteed minimum replanning scheme is determined as the final replanning execution scheme.
[0015] Optionally, the S7 includes: Based on the final replanning execution scheme, the procurement decision variables, production decision variables, warehousing decision variables, and transportation decision variables are analyzed, and procurement execution instructions corresponding to the procurement decision variables, production execution instructions corresponding to the production decision variables, warehousing execution instructions corresponding to the warehousing decision variables, and transportation execution instructions corresponding to the transportation decision variables are generated respectively. Instruction parameters are generated for the procurement execution instruction, the production execution instruction, the warehousing execution instruction, and the transportation execution instruction, respectively. The instruction parameters include an execution object identifier, a quantity parameter, and a time identifier. The procurement execution order, the production execution order, the warehousing execution order, and the transportation execution order are issued to the corresponding business systems for execution; The final replanning execution scheme is recorded as the replanning execution scheme at the current rolling replanning moment, and the replanning execution scheme at the current rolling replanning moment is updated to the historical replanning execution scheme at the next rolling replanning moment, returning to step S1.
[0016] The beneficial effects of this invention are: 1. Generate a limited number of candidate replanning schemes through a generative model, and then evaluate and rank them within the candidate set by a reinforcement learning policy model. This avoids reinforcement learning from directly outputting decisions in a high-dimensional action space, reducing the time spent on solving and learning, and enabling the coordinated replanning of multiple links such as procurement, production, warehousing and transportation to be completed within the emergency time limit under sudden disturbances.
[0017] 2. By performing feasibility projection and feasibility repair processing based on hard constraint sets on candidate solutions and eliminating unrepairable solutions, the final output solution meets supply chain hard constraints such as capacity, inventory, loading, delivery window, minimum order quantity, and contract lock-in, thereby improving the feasibility of the solution and reducing manual correction costs.
[0018] 3. A safety gating mechanism consisting of uncertainty indicators and key performance indicator thresholds is used. When a failure occurs, the rollback solver is called to generate a backup plan that meets the hard constraints and threshold requirements. This avoids key indicators such as stockout rate, penalties, and cash flow from exceeding the bottom line due to online trial and error. At the same time, the cost of plan changes is calculated by combining historical replanning execution plans to suppress rolling replanning plan oscillations and reduce the execution costs caused by frequent changes. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a supply chain emergency replanning method based on reinforcement learning proposed in this invention; Figure 2 The flowchart for step S3 illustrates the process of calculating constraint violations, projecting feasibility, repairing feasibility, and determining feasibility for candidate solutions, ultimately resulting in a set of feasible candidate solutions that satisfy the hard constraints. Detailed Implementation
[0020] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0021] refer to Figure 1 A reinforcement learning-based supply chain contingency replanning method includes: S1. At each rolling replanning time, acquire supply chain business data and information on sudden disturbances to construct supply chain status data, and determine the replanning execution plan from the previous rolling replanning time as the historical replanning execution plan; S2. Based on the supply chain status data, information on sudden disturbances, and historical replanning execution plans, call the generative model to generate a set of candidate replanning plans, with the number of plans not exceeding a preset upper limit; S3. Perform feasibility processing on the set of candidate replanning plans based on a set of hard constraints to obtain a set of feasible candidate replanning plans that meet the hard constraints. Feasibility processing includes: performing feasibility projection processing and / or feasibility repair processing on candidate replanning plans that do not meet the hard constraints; S4. Combining the supply chain status data, information on sudden disturbances, historical replanning execution plans, and the set of feasible candidate replanning plans, use a reinforcement learning strategy model to process the feasible candidate replanning plans. The set of solutions is evaluated and ranked, and the optimal ranked solution is selected as the reinforcement learning recommended replanning solution. Evaluation results corresponding to the reinforcement learning recommended replanning solution are generated, including predicted values of key performance indicators (KPIs) and the cost of plan changes. S5: Based on the evaluation results, the uncertainty index of the reinforcement learning recommended replanning solution is calculated. Based on the KPI threshold set, it is determined whether the predicted KPI values meet the threshold requirements. A pass flag is generated when the uncertainty index does not exceed the preset safety threshold and the predicted KPI values meet the KPI threshold set; otherwise, a fail flag is generated. S6: If a pass flag is obtained, the reinforcement learning recommended replanning solution is determined as the final replanning execution solution; otherwise, the backoff solver is called to generate a backup replanning solution that meets both the hard constraint set and the KPI threshold set. S7: An execution instruction is generated based on the final replanning execution solution and issued for execution.
[0022] In this specific embodiment, S1 includes: The system is based on Set the current data slice reference time and align it to the reference time. For discrete-time grids at the granularity level, data retrieval tasks are initiated from the enterprise business system to obtain order demand data, inventory data, production capacity data, warehousing operation capacity data, transportation resource data, and the replanning execution plan of the previous rolling replanning moment within the same transaction batch, while obtaining sudden disturbance information from the event management system. The order demand data uses "order line" as the smallest recording unit and includes order identifier, material identifier, required quantity, delivery deadline, and delivery node identifier. The inventory data uses "node-material" as the smallest recording unit and includes current quantity, allocated quantity, and safety stock. The production capacity data uses "production node-time period" as the smallest recording unit and includes available working hours or available output limit and locked capacity occupancy. The warehousing operation capacity data uses "warehouse node-time period" as the smallest recording unit and includes inbound capacity, outbound capacity, and locked operation occupancy. The transportation resource data uses "route or carrier resource-time period" as the smallest recording unit and includes available train or cabin limit and locked transportation occupancy. The sudden disturbance information includes disturbance type, set of affected supply chain node identifiers, start and end time of impact, and impact intensity parameters, and is used to deduct from the capacity, warehousing operation capacity, and transportation resource limit for the corresponding time period. After completing data collection, the system performs timestamp alignment and data consistency verification on the data. Timestamp alignment maps the occurrence time of all records to a unified time. The time periods are defined to ensure the uniqueness of records with the same "key field combination". Data consistency verification includes non-negativity verification, key field integrity verification, and cross-table consistency verification, and will exclude records with missing key fields, negative values, and delivery deadlines earlier than [previous dates]. Order records and abnormal records exceeding the business master data limit are identified as invalid records and removed to form valid business data; Based on valid business data and information on sudden disruptions, the system calculates and summarizes status elements according to supply chain node identifiers and time identifiers. These status elements include at least available inventory, in-transit inventory, remaining capacity, remaining warehousing capacity, available transportation capacity, and order delivery urgency. Available inventory is obtained by subtracting allocated and safety stock from current inventory. In-transit inventory is calculated from shipped inventory and... Subsequent arriving shipments are aggregated by arrival time period. Remaining capacity is calculated by subtracting the locked-in capacity from the capacity limit and adding the capacity reduction corresponding to sudden disruptions. Remaining warehousing capacity is calculated by subtracting the locked-in capacity from the warehousing capacity limit and adding the capacity reduction corresponding to sudden disruptions. Available transportation capacity is calculated by subtracting the locked-in transportation capacity from the transportation resource limit and adding the transportation capacity reduction corresponding to sudden disruptions. Order delivery urgency is calculated by the remaining time between the order delivery deadline and the current time period, combined with the standard fulfillment lead time. The calculation shows that the urgency level increases monotonically as the remaining time decreases; The system associates each status element with its corresponding supply chain node identifier. and time markers The supply chain status data is correlated and constructed, and is represented by a state vector indexed by node-time as follows: ; in Represents a node Time marker The supply chain state vector at the location, A unique code representing a supply chain node, taken from the node's master data table. Indicates a time identifier and is in the form of Discretized time period number Represents a node Time marker Available inventory at the location Represents a node Time marker The amount in transit Represents a node Time marker The remaining capacity at that location Represents a node Time marker The remaining warehousing capacity at the location Represents a node Time marker Available transportation capacity at the location Represents a node Time marker The urgency of order delivery at the location; The system identifies the replanning execution plan from the previous rolling replanning time as the historical replanning execution plan and compares it with the previous plan. It binds archives and outputs supply chain status data, information on sudden disturbances, and historical replanning and execution plans.
[0023] In this specific embodiment, S2 includes: The system receives supply chain status data, sudden disturbance information, and historical replanning execution plans, and uses the current rolling replanning time as the basis for its analysis. Build a length of starting point Data generated under specific time periods ,in It is formed by concatenating three parts in the order of fields, and the order of fields is fixed as state conditions, disturbance conditions, and historical scheme conditions; The state conditions are determined by the node set. With time set State vector on Expand to obtain and maintain with Consistent index order, and before expansion... Each state element within the system undergoes interval normalization to eliminate dimensional differences, and the normalization parameters are fixed to the upper and lower bounds obtained from offline statistics. The disturbance conditions are obtained by mapping sudden disturbance information. The mapping process converts the set of supply chain node identifiers affected by the disturbance and the start and end times of the impact into the same... The perturbation mask matrix and intensity matrix on the index are used, and the positions outside the influence range are set to zero, so that the generative model can obtain the perturbation effect range with the same granularity as the state conditions on the input side; The conditions of the historical scheme are determined by the historical replanning implementation scheme. The data is extracted from the range and expanded in the order of procurement decision variables, production decision variables, warehousing decision variables, and transportation decision variables. The quantity fields are normalized using the same normalization parameters as the state conditions to ensure consistent scaling. System call generates model pairs Perform multiple random generation attempts and set the maximum number of generation attempts to 1. At the same time, the maximum number of preset candidates is set to And the number of candidates reached Generation will terminate at this time; The generative model is a conditional variational autoencoder, whose conditional encoder... Using the number of layers The Transformer encoder structure with each hidden dimension being The number of heads of attention is The feedforward network dimension is Using drop rate Dropout is used to suppress overfitting in the decoder. Using the number of layers The network structure is a fully connected network with 512 hidden units per layer and uses the ReLU activation function. The dimension of the latent variables is set to [value missing]. Furthermore, the latent variable sampling follows a standard normal distribution to achieve diversified generation under the same conditions; The generation process satisfies: ; in Indicates the first The set of decision variables for the candidate reprogramming schemes obtained from the second generation is represented in a vectorized form. Indicates the generation of serial numbers and , Indicates the decoder network and Indicates decoder network parameters, This indicates the conditions for generating data. Indicates the first Secondary generated latent variables, Indicates a normal distribution. The dimension is The zero vector, The dimension is The identity matrix, Indicate the dimension of latent variables; The system will Map back to structured candidate replanning schemes by field and identify each time period. Generate associated procurement decision variables, production decision variables, warehousing decision variables, and transportation decision variables. During the mapping process, perform inverse normalization on quantity-type decision variables and round them to the nearest integer using the smallest unit of business measurement. Set any integer values less than zero to zero to ensure that the data type and value range of the candidate solutions meet the input requirements for subsequent feasibility processing. The system calculates a deduplication fingerprint for each structured candidate replanning scheme and performs duplicate scheme removal. The deduplication fingerprint is obtained by converting all quantitative decision variables into integers according to the smallest unit of measurement, concatenating them with all time identifiers in a fixed order to form a byte sequence, and then calculating the SHA-256 digest. Candidate replanning schemes with the same digest are judged as duplicates, and only the scheme that appears for the first time is retained. Complete at most After generating and performing deduplication, the system outputs a set of candidate replanning solutions, and the number of solutions in this set does not exceed [a certain threshold]. .
[0024] In this specific embodiment, S3 includes: The system receives a set of candidate replanning schemes and processes them one by one according to their generation number. The candidate replanning schemes are represented as vectors. Indicates and From time set The associated procurement decision variables, production decision variables, warehousing decision variables, and transportation decision variables are concatenated in a fixed field order, where... Indicates the candidate reprogramming scheme number and The value range is the sequence number of all schemes within the candidate set; The system instantiates a set of hard constraints based on supply chain status data and information on sudden disturbances, and encodes them into computable constraint checkers and projective optimization constraints. The set of hard constraints includes capacity hard constraints, inventory hard constraints, loading hard constraints, delivery window hard constraints, minimum order quantity hard constraints, and contract lock-in hard constraints. Specifically, the capacity hard constraint restricts the production decision variables of each production node at each time point from not exceeding the remaining capacity; the inventory hard constraint restricts the ending inventory of each node at each time point from being neither less than zero nor less than the safety stock level, and satisfies the inventory balance relationship between nodes; and the loading hard constraint restricts the loading of each transportation resource at each time point from exceeding the remaining capacity. The transportation decision variables do not exceed the available transportation capacity, and the loading capacity of a single vehicle or container does not exceed the maximum load or volume limit. The delivery window hard constraint restricts each order to be delivered within its delivery time window and not later than the delivery deadline. The minimum order quantity hard constraint restricts the procurement decision variable corresponding to each supplier-material-time identifier to be either zero or not less than the minimum order quantity of that supplier-material, and satisfies the constraint of being an integer multiple of the minimum unit of measurement. The contract locking hard constraint generates a locking set from the historical replanning execution scheme and fixes the decision variables in the locking set to the corresponding values in the historical replanning execution scheme. The locking set is based on the length of the locking time domain. Constructed and covering time periods from Before Procurement, production, warehousing, and transportation decisions that have been issued or committed within a given timeframe; For each candidate reprogramming scheme The system calls the constraint checker to calculate the constraint violation amount item by item and summarizes it into a total violation amount. The calculation rule for the constraint violation amount is as follows: for each inequality constraint, the "excess amount by which the left-hand value exceeds the right-hand value" is taken as the violation amount of the constraint and is taken as a non-negative value; for each equality constraint, the absolute value of the difference between the left and right ends is taken as the violation amount of the constraint and is taken as a non-negative value. Then, the violation amounts of the same type of constraint on all node identifiers and time identifiers are summed to obtain the class violation amount of the constraint, and the class violation amounts of all types of constraints are summed to obtain the total violation amount. When the total number of violations is zero, the system will... Directly mark it as a feasible candidate replanning solution and add it to the set of feasible candidate replanning solutions; When the total number of violations is not zero, the system... Perform a feasible projection process and construct it as a projection optimization problem with the objective function of minimizing the adjustment amount and the constraints of a set of hard constraints. Solve the problem using a mixed-integer linear programming solver. The projection optimization problem is written as: st ; in Indicates the candidate reprogramming scheme The projection scheme vector obtained after performing projection processing. Denotes the decision vector of the projection optimization problem and is related to Having the same field order and the same dimensions, Represents the component index of the decision vector and This represents the decision vector dimension and consists of procurement decision variables, production decision variables, warehousing decision variables, and transportation decision variables. The unfolded length within the range is determined. Represents the decision vector The One portion, Indicates candidate reprogramming schemes The One portion, This represents the absolute value operation. Indicates the first The adjustment weights for each component are fixed according to the field type: procurement component weight 1.0, production component weight 1.2, warehousing component weight 0.8, and transportation component weight 1.1. Components belonging to the locked set are assigned a value. To force the contract to remain locked, This indicates that the set of hard constraints is in The feasible domain set instantiated on the index includes the above-mentioned capacity, inventory, loading, delivery window, minimum order quantity and contract locking constraints; The solver sets a single-task solution time limit for each candidate projection task. And return the satisfactory result within the time limit The feasible solution as Or return an infeasibility flag; When a feasible solution is returned, the system uses Alternative When the constraint checker is called again to calculate the total number of violations and the total number of violations is zero, When a feasible candidate replanning scheme is added to the set of feasible replanning schemes, and the total violation is not zero or returns to the infeasibility flag, the system performs feasibility repair processing. In the repair processing, the following rule-based repair operations are performed in a fixed order: contract lock coverage, minimum unit of measurement rounding, minimum order quantity multipleing, loading integerization and capacity truncation, extension of excess capacity and warehousing operation capacity to the next period without crossing the order delivery deadline, and elimination of negative inventory by extending outbound and adjusting inbound to the next period without violating delivery window constraints. During the repair process, the constraint checker is called to update the total violation after each rule repair is completed. The repair stops when the total violation is zero and the repaired scheme is added to the set of feasible candidate replanning schemes. If the total violation is still not zero after all rule repairs are completed, the corresponding candidate replanning scheme is removed. After performing the above processing on all candidate replanning schemes in the candidate replanning scheme set, the system summarizes all feasible candidate replanning schemes and outputs the feasible candidate replanning scheme set.
[0025] In this specific embodiment, S4 includes: The system receives supply chain status data, information on sudden disturbances, and historical replanning execution plans. It also receives a set of feasible candidate replanning plans and vectorizes these historical plans into... And maintain the steps and steps Consistent field order and time set A consistent time identification order is maintained, and each feasible candidate replanning scheme in the set of feasible candidate replanning schemes is vectorized into... ,in Indicates the first The decision variable vector of 1 feasible candidate reprogramming scheme includes procurement decision variables, production decision variables, warehousing decision variables and transportation decision variables associated with time markers; The system targets each Construct evaluation objects Evaluation object It is obtained by splicing three types of data in a fixed order, namely, state data, disturbance data, and scheme data. The state data is identified by all nodes. With time markers State vector on according to Steps to obtain and use the index sequential expansion The same upper and lower bound tables are used to complete interval normalization. The disturbance data is obtained by mapping the disturbance information to the disturbance mask matrix and the intensity matrix according to... The scheme data is obtained by expanding the index in sequence. and The data is obtained by concatenating the fields after alignment, and the quantity fields are normalized using the same normalization parameters as the status data. The system will be all Input reinforcement learning policy model to generate various The corresponding predicted key performance indicators, wherein the reinforcement learning strategy model is a multi-head value assessment network trained offline and invoked online for inference. ,in The shared encoder is for the number of layers. A Transformer encoder with hidden dimensions of 1 / 2. The number of heads of attention is The feedforward network dimension is The discard rate Shared encoder After obtaining fixed-dimensional representation vectors through feature extraction, they are fed into the out-of-stock rate prediction head, penalty prediction head, and cash flow prediction head, respectively. All three prediction heads are three-layer fully connected networks with layer widths of [missing information]. The system uses the ReLU activation function to output scalar predicted values, and records the three outputs as the stockout rate predicted values. Predicted value of fines Compared with cash flow forecasts Furthermore, all three predicted values are normalized scalars to allow for in-scalar integration with the costs of plan changes; System basis and Calculation of differences between Corresponding cost of plan change ,in The calculation is performed by accumulating the absolute differences weighted by field type, with fixed weights: 1.0 for procurement decision variables, 1.2 for production decision variables, 0.8 for warehousing decision variables, and 1.1 for transportation decision variables. The accumulated result is then divided by the historical replanning execution plan. The absolute values of the corresponding fields within the range are summed to obtain the value range. The cost of changing the normalization plan; The system calculates and ranks evaluation scores based on the predicted values of key performance indicators and the costs of plan changes. The evaluation scores are calculated using the following formula: ; in Indicates the first One feasible candidate reprogramming scheme The evaluation score is calculated based on the overall performance of the solution, with a lower score indicating a better overall performance. This represents the weight of the out-of-stock rate and its value is... Indicates the first The predicted out-of-stock rate for each option. This represents the weight of the penalty and its value is... Indicates the first The projected penalty value for each option. Represents the cash flow weight and takes the value of Indicates the first The projected cash flow for each option is represented by a negative sign in the evaluation score, as a larger cash flow is considered better. This indicates a plan to change the weights and a value of [value]. Indicates the first The cost of plan changes relative to historical replanning and implementation plans; The system performs all feasible candidate reprogramming schemes according to... Sort by size from smallest to largest and in case of identical scores, use... The smaller solution is ranked as the better solution, and the solution ranked first is determined as the reinforcement learning recommendation replanning solution and its corresponding evaluation result is output. The evaluation result includes the predicted value of the out-of-stock rate, the predicted value of the penalty, the predicted value of the cash flow, and the cost of the plan change for the reinforcement learning recommendation replanning solution.
[0026] In this specific embodiment, S5 includes: The system receives the reinforcement learning-recommended reprogramming scheme and its evaluation results, and records the reinforcement learning-recommended reprogramming scheme as... ,in This represents the optimally sorted feasible candidate reprogramming scheme, and its field order corresponds to the time set. With steps and steps Consistent, and the predicted out-of-stock rate, predicted penalty, and predicted cash flow in the assessment results are respectively recorded as follows: , and ,in This represents the predicted out-of-stock rate corresponding to the reinforcement learning-recommended reprogramming scheme. This represents the predicted penalty value corresponding to the reprogramming scheme recommended by reinforcement learning. This indicates that the cash flow forecast corresponding to the reinforcement learning-recommended reprogramming scheme is equal to the sum of the three steps. Consistent with normalized scalars; The system is based on steps The same evaluation object construction rules will The evaluation criteria for this recommended solution are constructed by combining supply chain status data, information on sudden disruptions, and historical replanning and implementation schemes. and to Perform uncertainty assessment to obtain multiple sets of predicted values for key performance indicators. Uncertainty assessment adopts the following steps: Consistent reinforcement learning policy model During the inference phase, Dropout is kept on to form multiple forward evaluations of Monte Carlo Dropout, where the Dropout drop rate is maintained at a certain step. The settings in Furthermore, each forward evaluation uses an independent Dropout mask, and the system fixes the number of evaluations to a fixed value. And obtained ,in Indicates the evaluation sequence number and , Indicates the first The predicted out-of-stock rate output from this assessment. Indicates the first The predicted penalty value output by this assessment. Indicates the first The cash flow forecast output from this assessment; The system defines the dispersion of multiple assessments as an uncertainty index and uses the largest standard deviation among the predicted values of the three key performance indicators as the uncertainty index. To implement the aforementioned safety margin mechanism, the system defines the safety margin as a linear function positively correlated with the uncertainty index and applies it to the stockout rate threshold, penalty threshold, and cash flow threshold, respectively. The calculation process satisfies the following: ; in This represents an uncertainty indicator, and its value is the maximum of the standard deviations of the predicted values of the three key performance indicators. This represents the standard deviation of the predicted out-of-stock rate. This represents the standard deviation of the predicted fines. The standard deviation of the projected cash flow values. Indicates the key performance indicator type index and Indicates the number of evaluations and Indicates the evaluation sequence number. Indicates the first The output of the first evaluation Predicted values of key performance indicators (KPIs) Indicates the first The average of multiple assessments of the predicted values of key performance indicators (KPIs). Indicates the first Safety margins corresponding to key performance indicators (KPIs) This represents the safety margin factor, and its value is fixed. To ensure that the safety margin is monotonically positively correlated with the uncertainty index; The system reads the key performance indicator threshold set and safety threshold from the threshold configuration table and records the stockout rate threshold as follows: The penalty threshold is recorded as Cash flow threshold is denoted as And the preset safety threshold is recorded as ,in and upper limit threshold and The lower limit threshold; The system uses the average of multiple evaluations. and The predicted value of the key performance indicator is used in the gating judgment, and a safety margin is introduced for threshold judgment. The threshold judgment rule is fixed to simultaneously satisfying , and If the predicted value of a key performance indicator meets the threshold set of key performance indicators, then the predicted value of the key performance indicator does not meet the threshold set of key performance indicators. The system further implements uncertainty security gating and adopts rules. To determine whether the uncertainty index does not exceed the preset safety threshold, if and only if A pass flag is generated when the predicted value of the key performance indicator meets the key performance indicator threshold set; otherwise, a fail flag is generated and the pass or fail flag is linked to the threshold set. as well as The output is used to determine whether to trigger the backoff solver.
[0027] In this specific embodiment, S6 includes: The system receives pass or fail flags and also receives reinforcement learning-recommended reprogramming schemes. The evaluation results, along with the retention of supply chain status data, information on sudden disruptions, and historical replanning and implementation vectors, are also included. ,in This represents the optimal sorted feasible candidate reprogramming scheme, where the field order corresponds to the time set. Consistent This indicates the replanning execution plan at the previous rolling replanning moment. Vectorized representation within the range and the field order is the same as Consistent; When the flag is true, the system will The final replanning implementation plan is directly identified and output, and it is marked as "Reinforcement Learning Successful Plan" for subsequent auditing and backtracking. When the failure flag is true, the system constructs a backtracking solution task and solidifies this backtracking solution task into a defined mixed-integer linear programming model, the decision vector of which is denoted as . and With steps and steps The model consistently includes procurement decision variables, production decision variables, warehousing decision variables, and transportation decision variables associated with time markers. The hard constraint feasible region of the model is denoted as... and From steps The same hard constraints on capacity, inventory, loading, delivery window, minimum order quantity, and contract lock-in exist in China. Instantiated on the index and completely consistent, the key performance indicator threshold set of the model is obtained from the steps Read and keep unchanged, and record them as the out-of-stock rate thresholds respectively. Fine threshold With cash flow threshold ,in and upper limit threshold and The lower limit threshold; To ensure that a fallback replanning scheme satisfying the threshold set can still be generated when sudden disturbances cause insufficient regular resources, the system activates an emergency resource pool during the rollback solution task and injects the emergency resource pool with additional but still constrained resource entries. The emergency resource pool includes emergency procurement resources and emergency transportation resources, both of which have fixed resource caps and fixed unit cost parameters. The emergency procurement resources are modeled as a set of suppliers. The procurement decision variables within the system are constrained by a zero lead time and a total quantity limit. Emergency transportation resources are modeled as a set of carrier resources. The transportation decision variables within the scope are constrained by the upper limit of available transportation capacity, thus enabling the rollback solution task to have a feasible solution space provided that the contract lock and other hard constraints are not violated. The system invokes a rollback solver to solve the rollback task. The rollback solver is a branch-and-bound hybrid integer linear programming solver with a global solution time limit. Feasible solution priority search strategy and optimality gap threshold And within the time limit, return a feasible solution that satisfies all hard constraints and threshold constraints. The solution objective of the backtracking task is fixed as minimizing the overall cost and the cost of plan change under the conditions of satisfying the set of hard constraints and the set of key performance indicator thresholds, and is written as: , st ; in This represents the fallback replanning scheme output by the fallback solver. This represents the decision vector for the backoff solution task. This indicates that procurement costs, production costs, warehousing operation costs, transportation costs, and emergency resource surcharges are measured in a uniform currency and... The total cost obtained by summing up the costs within the scope This indicates the weight of the cost of plan changes and its value is [value]. This indicates the cost of plan changes relative to historical replanning implementation schemes and the steps involved. Consistent field type weighted absolute difference normalization calculation rules This represents the set of feasible regions with hard constraints. Indicates by Induced shortage rate and by The ratio of unmet demand to total demand within the specified range is calculated and implemented by linearizing the stockout quantity auxiliary variable. Indicates by Induced penalties and the amount of late orders and the unit price of the penalty. The results were obtained by summarizing within the range and implemented using linearized late crossover auxiliary variables. Indicates by Induced cash flow, calculated by subtracting procurement, production, warehousing, and transportation expenses from revenue collection. The time series summary within the range is obtained by using the lowest cash balance at the time point as the constraint. This indicates the out-of-stock rate threshold. Indicates the threshold for fines. Indicates the cash flow threshold; The system returns after the solver has rolled back. Then call the steps again Same constraint checker pair Perform hard constraint consistency verification and calculation and Perform a threshold consistency review. Once the review is passed, The final replanning and implementation plan is determined and output, and it is marked as a "rollback backup plan" to distinguish it from the "reinforcement learning pass plan" for management purposes.
[0028] In this specific embodiment, S7 includes: The system receives the final replanning execution plan and records it as... ,in Indicates the time set The set of decision variables within the scope, including procurement decision variables, production decision variables, warehousing decision variables, and transportation decision variables, and maintaining consistency with the steps. and steps Consistent field order and time identifier order; The system The system performs structured parsing and places the results into four types of execution detail tables. The procurement execution detail table uses supplier identifier, material identifier, and time identifier as key fields and generates procurement execution instructions. The production execution detail table uses production node identifier, work order or process identifier, and time identifier as key fields and generates production execution instructions. The warehousing execution detail table uses warehouse node identifier, operation type identifier, and time identifier as key fields and generates warehousing execution instructions. The transportation execution detail table uses carrier resource identifier or route identifier, loading unit identifier, and time identifier as key fields and generates transportation execution instructions. To ensure consistency between the issued commands and the business system, the system encapsulates each execution instruction into an instruction message with a fixed field order: ; in This represents a single instruction message. `id` represents the unique identifier of the instruction and is used for idempotent control. `cat` represents the instruction category, with a value of one of the following: procurement, production, warehousing, or transportation. `obj` represents the execution object identifier, corresponding to a combination of supplier and material identifiers in the procurement category, a combination of production node and work order identifiers in the production category, a combination of warehouse node and operation type identifiers in the warehousing category, and a combination of carrier resource identifiers or route identifiers and loading unit identifiers in the transportation category. Indicate quantity parameters and be consistent with the smallest unit of measurement and steps The range of values remains consistent after feasibility processing. Indicates time marker and ver represents the plan version number and its value corresponds to the rolling replanning time. The version number corresponds to each other and is used for concurrency control on the business system side; the hash represents the message verification digest and is used for integrity verification. After generating all instruction messages, the system calls the corresponding business system interface according to the instruction category and executes the issuance process with transaction boundaries. The procurement execution instruction is written through the purchase order interface of the enterprise resource planning system and triggers the supplier confirmation process. The production execution instruction is written through the work order issuance interface of the advanced planning and scheduling system and triggers the production line scheduling lock. The warehousing execution instruction is written through the inbound, outbound and transfer task interface of the warehouse management system and triggers the task assignment. The transportation execution instruction is written through the waybill and booking interface of the transportation management system and triggers the carrier confirmation. The system processes each... The system performs interface return code verification and writes receipts to the database. When the return is successful, the status of the corresponding execution instruction is updated to "issued" and the issuance timestamp and business system receipt identifier are recorded. When the return fails, the status of the corresponding execution instruction is updated to "issued failed" and a retry queue is triggered. The maximum number of retries is fixed at 3. If the limit is exceeded, a manual alarm is triggered. At the same time, the system maintains duplicate issuance suppression with ID as the idempotent key to avoid the business system generating duplicate orders or duplicate tasks. After all command messages have been sent, the system will Write the plan archive table with the plan version ver as the index and record it as the current rolling replanning time. The replanning execution plan is then updated to the historical replanning execution plan for the next rolling replanning time step, and the process returns to the previous step. This will allow us to enter the next round of rolling replanning and closed loop.
[0029] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0030] This invention employs a sequential emergency replanning process consisting of "candidate solution generation, hard constraint feasibility processing, reinforcement learning evaluation and ranking, uncertainty safety gating, and fallback solution." After a sudden disturbance occurs, the collaborative decisions across procurement, production, warehousing, and transportation are first limited to a controlled set of candidates, significantly compressing the reinforcement learning decision search space and shortening the emergency solution time. Subsequently, feasibility projection and feasibility repair processing are performed on the candidate solutions using a set of hard constraint conditions, ensuring that the solutions entering the reinforcement learning evaluation stage meet hard constraints such as capacity, inventory, loading, delivery window, minimum order quantity, and contract lock-in, avoiding the output of unexecutable solutions. Reinforcement learning ranks and selects feasible candidates based on predicted key performance indicators and the cost of plan changes, enabling replanning to respond quickly to disturbances while suppressing frequent changes and plan oscillations during the rolling process, reducing execution costs such as line changes, reassignments, and rerouting.
[0031] In terms of algorithm structure, this invention makes problem-oriented improvements to address the risk spillover and data uncertainty in emergency scenarios: On the one hand, it changes reinforcement learning from "directly outputting high-dimensional actions" to "evaluating and ranking feasible candidate solutions," and pre-solidifies hard constraints through projection and repair, so that the learning and reasoning process is naturally constrained by the executability boundary; on the other hand, it introduces an uncertainty index based on the degree of dispersion of multiple evaluations, and combines it with a set of key performance indicator thresholds to form a safety gating mechanism. When the recommended solution has high uncertainty or may trigger bottom-line risks such as stockout rate, penalties, and cash flow, it automatically switches to the backoff solver to generate a backup solution that meets the hard constraints and threshold requirements. This mechanism avoids the risk diffusion caused by online trial and error, and ensures that there is still a stable and feasible implementation strategy under emergency conditions.
Claims
1. A supply chain contingency replanning method based on reinforcement learning, characterized in that, include: S1. At the rolling replanning moment, acquire supply chain business data and information on sudden disturbances, construct supply chain status data, and determine the replanning execution plan at the previous rolling replanning moment as the historical replanning execution plan. S2. Based on supply chain status data, sudden disturbance information and historical replanning execution plans, call the generation model to generate a set of candidate replanning plans. The number of plans shall not exceed the preset upper limit of the number of candidates. S3. Perform feasibility processing on the candidate replanning scheme set based on the hard constraint set to obtain a set of feasible candidate replanning schemes that satisfy the hard constraint set. Feasibility processing includes: performing feasibility projection processing and / or feasibility repair processing on candidate replanning schemes that do not satisfy the hard constraint set; S4. Combining supply chain status data, sudden disturbance information, historical replanning execution schemes, and the set of feasible candidate replanning schemes, use a reinforcement learning strategy model to evaluate and rank the set of feasible candidate replanning schemes. Select the scheme with the best ranking as the reinforcement learning recommended replanning scheme, and generate the evaluation results corresponding to the reinforcement learning recommended replanning scheme, including the predicted values of key performance indicators. S5. Calculate the uncertainty index of the reinforcement learning recommended reprogramming scheme based on the evaluation results, and determine whether the predicted value of the key performance indicator meets the threshold requirements based on the key performance indicator threshold set. When the uncertainty index does not exceed the preset safety threshold and the predicted value of the key performance indicator meets the key performance indicator threshold set, a pass flag is generated; otherwise, a fail flag is generated. S6. If a pass flag is obtained, the reinforcement learning recommended reprogramming scheme is determined as the final reprogramming execution scheme; otherwise, the backoff solver is called to generate a backup reprogramming scheme that meets the hard constraint condition set and the key performance indicator threshold set. S7. Generate execution instructions based on the final reprogramming execution scheme and issue them for execution.
2. The supply chain emergency replanning method based on reinforcement learning according to claim 1, characterized in that, S1 includes: At each rolling replanning time, the order demand data, inventory data, production capacity data, warehousing operation capacity data, transportation resource data, and the replanning execution plan from the previous rolling replanning time are collected, and the information on sudden disturbances is also collected. The collected data is timestamped and its consistency is verified. Missing or abnormal items are removed to form valid business data. Based on the valid business data and the sudden disturbance information, calculate status elements including available inventory, in-transit quantity, remaining capacity, remaining warehousing operation capacity, available transportation capacity, and order delivery urgency, and associate the status elements with the corresponding supply chain node identifiers and time identifiers to construct the supply chain status data; The replanning execution scheme at the previous rolling replanning moment is determined as the historical replanning execution scheme, and the supply chain status data, the sudden disturbance information, and the historical replanning execution scheme are output.
3. The supply chain emergency replanning method based on reinforcement learning according to claim 1, characterized in that, S2 includes: Based on the supply chain status data, the sudden disturbance information, and the historical replanning execution scheme, the generated condition data is constructed. The generated condition data includes conditional elements that characterize order demand, inventory level, production capacity, warehousing operation capacity, transportation resources, and the scope of impact of sudden disturbances. Based on the generated condition data, the generated model is called multiple times to generate multiple candidate replanning schemes that do not exceed the preset upper limit of the number of candidates. The candidate replanning scheme set is formed by removing duplicate schemes from the multiple candidate replanning schemes. Each candidate replanning scheme is represented as a set of decision variables including the procurement decision variable, the production decision variable, the warehousing decision variable, and the transportation decision variable, and is associated with a corresponding time identifier.
4. The supply chain emergency replanning method based on reinforcement learning according to claim 1, characterized in that, S3 includes: For each candidate replanning scheme in the set of candidate replanning schemes, calculate the constraint violation amount based on the set of hard constraints; When the constraint violation amount is zero, the corresponding candidate replanning scheme is marked as a feasible candidate replanning scheme; When the constraint violation amount is not zero, the feasibility projection processing or the feasibility repair processing is performed to adjust the procurement decision variables, production decision variables, warehousing decision variables and transportation decision variables in the candidate replanning scheme, and the constraint violation amount is recalculated after adjustment; When the amount of constraint violation after adjustment is zero, the adjusted candidate reprogramming scheme is marked as a feasible candidate reprogramming scheme; otherwise, the corresponding candidate reprogramming scheme is eliminated. After processing all candidate replanning schemes in the candidate replanning scheme set, all feasible candidate replanning schemes are summarized to form the feasible candidate replanning scheme set.
5. The supply chain emergency replanning method based on reinforcement learning according to claim 1, characterized in that, S4 includes: Based on the supply chain status data, the sudden disturbance information, and the historical replanning execution schemes, an evaluation object is constructed for each feasible candidate replanning scheme in the set of feasible candidate replanning schemes; A reinforcement learning strategy model is used to evaluate each evaluation object and generate predicted values of key performance indicators for each feasible candidate replanning scheme. Calculate the plan change cost corresponding to each feasible candidate replanning scheme based on the difference between each feasible candidate replanning scheme and the historical replanning execution scheme; Based on the predicted values of the key performance indicators and the cost of the plan change, an evaluation score is generated for each feasible candidate replanning scheme, and the ranking result is generated based on the evaluation score; Based on the ranking result, the optimal ranking scheme is determined from the set of feasible candidate replanning schemes as the reinforcement learning recommended replanning scheme, and an evaluation result corresponding to the reinforcement learning recommended replanning scheme is generated. The evaluation results include the predicted values of the key performance indicators and the cost of the plan changes.
6. The supply chain emergency replanning method based on reinforcement learning according to claim 1, characterized in that, S5 includes: Based on the evaluation results, the predicted values of the key performance indicators (KPIs) corresponding to the reinforcement learning recommendation reprogramming scheme are obtained. The predicted KPIs are evaluated multiple times using an uncertainty evaluation method to obtain multiple sets of predicted KPIs. An uncertainty index is calculated based on these multiple sets of predicted KPIs, where the uncertainty index represents the dispersion of the predicted KPIs. A threshold judgment is performed on the predicted KPIs based on the KPI threshold set, where the stockout rate threshold and the penalty threshold are upper limit thresholds, and the cash flow threshold is a lower limit threshold. When the predicted stockout rate ≤ the stockout rate threshold, the predicted penalty ≤ the penalty threshold, and the predicted cash flow ≥ the cash flow threshold, the predicted KPIs are determined to satisfy the KPI threshold set. Otherwise, it is determined that the predicted value of the key performance indicator does not meet the key performance indicator threshold set; when the uncertainty indicator does not exceed the preset safety threshold and the predicted value of the key performance indicator meets the key performance indicator threshold set, the pass flag is generated; otherwise, the fail flag is generated.
7. A supply chain emergency replanning method based on reinforcement learning according to claim 1, characterized in that, S6 include: When the pass flag is true, the reinforcement learning recommended replanning scheme is determined as the final replanning execution scheme and output; When the failure flag is true, a rollback solution task is constructed based on the supply chain status data, the sudden disturbance information, and the historical replanning execution scheme. The rollback solution task includes the set of hard constraints and the set of key performance indicator thresholds. The rollback solver is invoked to solve the rollback task, generating a minimum replanning scheme that satisfies the set of hard constraints and the set of key performance indicator thresholds. The guaranteed minimum replanning scheme is determined as the final replanning execution scheme.
8. The supply chain emergency replanning method based on reinforcement learning according to claim 1, characterized in that, S7 includes: Based on the final replanning execution scheme, the procurement decision variables, production decision variables, warehousing decision variables, and transportation decision variables are analyzed, and procurement execution instructions corresponding to the procurement decision variables, production execution instructions corresponding to the production decision variables, warehousing execution instructions corresponding to the warehousing decision variables, and transportation execution instructions corresponding to the transportation decision variables are generated respectively. Instruction parameters are generated for the procurement execution instruction, the production execution instruction, the warehousing execution instruction, and the transportation execution instruction, respectively. The instruction parameters include an execution object identifier, a quantity parameter, and a time identifier. The procurement execution order, the production execution order, the warehousing execution order, and the transportation execution order are issued to the corresponding business systems for execution; The final replanning execution scheme is recorded as the replanning execution scheme at the current rolling replanning moment, and the replanning execution scheme at the current rolling replanning moment is updated to the historical replanning execution scheme at the next rolling replanning moment, returning to step S1.
9. A supply chain contingency replanning method based on reinforcement learning according to claim 4, characterized in that, The feasibility projection process includes: constructing a projection optimization problem with the objective function of minimizing the adjustment amount of the candidate replanning scheme and the constraint condition set of hard constraints, and solving the projection optimization problem to obtain the projected candidate replanning scheme. The adjustment amount is the weighted sum of the changes in the procurement decision variable, the production decision variable, the warehousing decision variable and the transportation decision variable before and after the projection.
10. A supply chain contingency replanning method based on reinforcement learning according to claim 6, characterized in that, The threshold determination of the predicted value of the key performance indicator based on the threshold set of key performance indicators includes: introducing a safety margin when making the threshold determination, the safety margin being positively correlated with the uncertainty indicator; and determining that the predicted value of the key performance indicator satisfies the threshold set of key performance indicators when the predicted value of the stockout rate, the predicted value of the penalty, and the predicted value of the cash flow respectively meet the threshold requirements after adding the safety margin.