Supply chain demand coordination diffusion reinforcement prediction control method

By combining a diffusion-reinforced predictive control method with robust stochastic model predictive control and diffusion model-driven offline reinforcement learning, the problems of high computational complexity, information asymmetry, and lack of disturbance modeling in traditional methods in supply chain systems are solved, achieving efficient and stable supply and demand coordination control.

CN121032154BActive Publication Date: 2026-01-27BEIHANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511563620.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-01-27
Estimated Expiration
2045-10-30

AI Technical Summary

Technical Problem

When facing dynamic environments and complex disturbances, existing supply chain systems suffer from high computational complexity and slow response speed due to traditional model predictive control methods, information asymmetry and lack of disturbance modeling in traditional distributed control methods, and insufficient representation capabilities of reinforcement learning methods. As a result, the stability and applicability of control strategies are insufficient in practical applications.

Method used

A diffusion-reinforced predictive control method is adopted, which integrates the state evolution, control input and disturbance modeling of the supply chain system through robust stochastic model predictive control and diffusion model-driven offline reinforcement learning to generate a control policy with robustness and efficient response. The diffusion model is combined to represent multimodal distribution, thereby improving the generalization ability and adaptability of the policy.

Benefits of technology

It achieves efficient and stable supply and demand coordination control in the supply chain system, improves the robustness and adaptability of the control strategy, and ensures the stability and accuracy of the system under complex disturbances and constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032154B_ABST
    Figure CN121032154B_ABST
Patent Text Reader

Abstract

The application relates to a diffusion reinforcement prediction control method for supply and demand coordination of a supply chain, belongs to the technical field of intelligent control and supply chain management, and solves the problems of insufficient modeling precision and weak generalization ability in supply chain control in the prior art, and comprises the following steps: S1: modeling of a supply chain network system, including establishing a discrete dynamic model of the supply chain network system, setting random disturbance in the discrete dynamic model, and establishing a control target to obtain a controlled network system model; S2: based on the controlled network system model, a robust stochastic model prediction control model is established to generate a control strategy for supply and demand regulation of the supply chain; S3: strategy generation and training based on a diffusion model, the robust stochastic model prediction control model is trained to obtain a robust regulation strategy model meeting real-time requirements; and S4: the trained robust regulation strategy model is used in the supply chain to generate a control strategy and is used in an execution program in the supply chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent control and supply chain management technology, specifically to a diffusion-enhanced predictive control method for supply chain supply and demand coordination. Background Technology

[0002] In the current context of continuous global economic integration, the supply chain system, as a multi-level network composed of nodes such as suppliers, manufacturers, distributors, retailers, and customers, has become a key support for enterprises to enhance competitiveness, achieve lean management, and achieve sustainable development through its efficient collaboration and dynamic response capabilities. Especially in the current environment of volatile demand and complex logistics, how to achieve intelligent collaboration and precise matching of supply and demand among supply chain nodes has become a technical challenge that urgently needs to be overcome in current research and industrial practice.

[0003] Model Predictive Control (MPC) is widely used in industrial process control due to its excellent constraint handling capabilities and rolling optimization characteristics. It models the control problem as an optimization problem, effectively handles variable constraints, and achieves system control by rolling the solution to the optimal control input. However, in research on random signal tracking, traditional MPC methods employ a decoupled "planning-control" approach: first, a reference trajectory is generated offline, and then the control input is solved online. While this two-stage optimization method has a clear structure, it requires frequent re-solving of the optimal trajectory in dynamic environments. If the reference signal is disturbed or the node state is updated, the planning must be restarted from scratch, significantly increasing computational complexity and execution latency, making it difficult to deploy on resource-constrained edge computing platforms.

[0004] On the other hand, in current supply chain systems, traditional distributed control methods generally assume that each sub-node can observe a unified reference signal. However, this premise is often difficult to meet in practical applications. In reality, only the retail nodes at the very end can grasp customer demand information in real time, while upstream nodes (such as manufacturers and suppliers) must rely on chain-like information transmission to make response decisions, forming a typical information asymmetry structure, which severely restricts the synergy and applicability of traditional control strategies. At the same time, supply chain systems are widely subject to non-Gaussian disturbances, such as transportation losses, order distortion, and demand distortion. These disturbances are nonlinear, asymmetric, and highly uncertain, far exceeding the modeling capabilities of conventional Gaussian disturbance assumptions. Existing methods lack effective mechanisms to deal with these real-world disturbances, leading to amplified tracking errors and decreased stability of the control system under disturbance conditions, making it difficult to guarantee the accuracy and reliability of supply and demand regulation.

[0005] Furthermore, the numerous constraints in supply chain systems (such as inventory capacity) lead to a highly fragmented feasible strategy space, and control strategies often exhibit multimodal distribution characteristics. Against this backdrop, Gaussian distribution policy networks, commonly used in traditional reinforcement learning methods, struggle to effectively capture the complex structure of these strategies, resulting in significant policy modeling biases and insufficient expressive power. This, in turn, limits learning effectiveness and the generalization ability of the strategies, making it difficult to meet the collaborative optimization requirements under supply chain constraints. The inability of Gaussian distribution policy networks, commonly used in traditional reinforcement learning, to effectively represent such complex policy distributions leads to insufficient modeling accuracy and weak generalization ability, thus limiting practical application effectiveness.

[0006] Several core technical challenges exist in current intelligent collaborative control of the supply chain. These challenges directly restrict the feasibility, stability, and deployment efficiency of existing methods in complex real-world scenarios. Specifically, these challenges include: poor real-time performance of the "planning-control" decoupling structure: Traditional MPC requires secondary optimization in tracking random external signals, resulting in high computational overhead and slow response speed in dynamic environments, making it difficult to deploy in multi-node, multi-constraint supply chain systems; information asymmetry and lack of disturbance modeling: Most distributed control methods assume that all nodes share a reference trajectory, ignoring the actual structure where only terminal retail nodes possess customer demand information, and failing to explicitly consider disturbance factors such as transportation delays and order information distortion, thus reducing system stability; and insufficient representational capabilities of reinforcement learning methods: In supply chain systems, control strategies often exhibit multimodal characteristics due to constraints.

[0007] Chinese invention patent application, publication number CN103810562A, entitled "A Dynamic Subject Collaboration Method for Supply Chain Systems," discloses a method where, driven by demand objectives, a leading subject establishes a project centered around the final product. Collaborating subjects are selected to form a collaborative work organization, which, under the standardized processing flow of the supply chain system and based on cooperation commitments, jointly completes the tasks set for the project. However, this method does not consider factors such as random disturbances in complex real-world scenarios and cannot provide a reliable control strategy.

[0008] Chinese invention patent application, publication number CN117522225A, entitled "Supply Chain Collaborative Management Method and System Based on Industrial Internet Identifier Resolution," discloses a method within the Industrial Internet that converts product identifier information corresponding to supply chain nodes into resolution results, converts supply chain node supply information into integrated information, and, based on demand information from the demand side and the integrated information, performs collaborative planning on supply chain nodes through a planning model to obtain collaborative planning results, which are then allocated to the supply chain nodes. However, this method also fails to consider factors such as random disturbances in complex real-world scenarios and cannot provide a reliable control strategy.

[0009] To address the aforementioned technical bottlenecks, there is an urgent need for an intelligent control method that possesses both robustness and convergence, while also being able to adapt to the complex constraints and disturbances of the supply chain system. Summary of the Invention

[0010] In view of the above problems, this invention provides a diffusion-based reinforced predictive control method for supply chain supply and demand coordination. Through a diffusion-based Robust Stochastic Model Predictive Control (DB-RSMPC) approach, the overall technical process includes the following key steps: First, robust control modeling: A robust stochastic model predictive control (RSMPC) problem is constructed, modeling the state evolution, control input, disturbances, and constraints of the supply chain system to form a stochastic optimization control problem. Second, feasibility and stability analysis: The necessary and sufficient conditions for the system to track random signals are derived, a robust control strategy with disturbance tolerance is designed, and theoretical conditions to ensure system stability under closed-loop operation are given. By solving this robust control problem, a high-quality expert control trajectory dataset satisfying the constraints is obtained. Third, diffusion strategy learning: Based on the expert data, an offline reinforcement learning algorithm driven by a diffusion model is introduced to learn the potential multimodal distribution of the strategy, thereby training a control strategy with good generalization ability and real-time response performance, ultimately achieving efficient and stable tracking control of random external signals. This invention proposes introducing a diffusion model into an offline reinforcement learning framework. By simulating the distribution of expert behavior trajectories generated by Robust Stochastic Model Predictive Control (RSMPC), it achieves the modeling and rapid generation of high-quality policies. The diffusion model can effectively characterize multimodal control behavior and has significant advantages in accuracy, policy diversity, and security compared to traditional reinforcement learning methods based on Gaussian policy modeling. Thus, it provides a more practical solution for intelligent supply and demand collaboration in the supply chain.

[0011] According to embodiments of the present invention, a diffusion-enhanced predictive control method for supply chain supply and demand coordination is provided, comprising:

[0012] S1: Modeling of the supply chain networked system, including establishing a discrete dynamics model of the supply chain networked system, setting the random disturbances in the discrete dynamics model, and establishing control objectives to obtain the controlled networked system model, wherein the supply chain networked system is a distributed system including multiple subsystems;

[0013] S2: Based on the controlled networked system model, a robust stochastic predictive control model is established to generate predictive control strategies for supply chain supply and demand regulation.

[0014] S3: Policy generation and training based on diffusion model: Train the robust stochastic model predictive control model to obtain a robust control policy model that meets real-time requirements.

[0015] S4: Apply the trained robust control strategy model to the supply chain to generate control strategies and provide them to the execution process in the supply chain;

[0016] Specifically, S2 includes:

[0017] S2.1: Based on the controlled networked system model, the reachability state of the prediction time window is introduced to establish a stochastic reachability trajectory optimization problem model to determine the optimal reachability trajectory. The optimization objective is to minimize the error between the predicted reachability output trajectory and the tracking trajectory, as well as the control energy consumption.

[0018] S2.2: For the optimization objective of the stochastic reachable trajectory optimization problem model, establish a collaborative loss function, and obtain the local cost of each subsystem in the entire prediction time window based on the predicted reachable output trajectory and tracking trajectory of each subsystem.

[0019] S2.3: Set the condition that the model of the random reachable trajectory optimization problem has a non-zero feasible solution;

[0020] S2.4: Considering random perturbations, the random reachable trajectory optimization problem model is transformed into a distribution-based stochastic predictive control model to provide predictive control strategies;

[0021] S2.5: Set the stability condition for the convergence of the tracking error distribution in the distributed system.

[0022] Optionally, S1 specifically includes:

[0023] S1.1: System modeling of the state variables, control inputs and random disturbances of each subsystem in the supply chain, modeling the networked supply chain system as a discrete dynamic model for each subsystem, where the subsystem is an independent operating unit in the supply chain, including suppliers, manufacturers and distributors;

[0024] S1.2: Establish control objectives, including establishing cooperative constraints on the state variables and control inputs of each subsystem, and establishing distributed cooperative control of the subsystems.

[0025] Optionally, in S1.1:

[0026] Each subsystem processes the state variables and control inputs at the current moment, and combines them with random disturbances to obtain the state variables at the next moment, and each subsystem processes the state variables at the current moment to obtain the output of each subsystem;

[0027] The state variables of the subsystem are the subsystem's inventory, production capacity, and demand for goods;

[0028] The control inputs of the subsystem are control inputs, including the adjustment amounts of the subsystem's output and the adjustment amounts of the delivery amount;

[0029] The random disturbance of the subsystem is a non-Gaussian disturbance, representing demand distortion. The random disturbance is set to be an independently distributed time random variable with finite second moments.

[0030] Optionally, in S2.1, an reachable state and corresponding control input are introduced as decision variables to establish a stochastic reachable trajectory optimization problem model, including:

[0031] The optimization objective of establishing a stochastic reachable trajectory optimization problem model is to minimize the deviation between the predicted reachable output trajectory and the tracking trajectory of each subsystem in the distributed system and the control energy consumption, using the reachable trajectories of each subsystem in the predicted future time window as variables. The reachable trajectory is a set including reachable states and reachable control inputs.

[0032] The constraints for establishing a stochastic reachable trajectory optimization problem model include: establishing the discrete state evolution law of each subsystem; generating the reachable state at the next moment based on the current reachable state and reachable control input within the prediction time window; obtaining the predicted reachable output trajectory of each subsystem based on the reachable state; ensuring that the reachable state of each subsystem satisfies physical and logical constraints; limiting the amplitude of the reachable control input of each subsystem; and constraining the predicted reachable output trajectory of each subsystem and its neighboring subsystems to be consistent with the cooperative reachable output trajectory.

[0033] Optionally, S2.2 specifically includes:

[0034] Based on the predicted reachable trajectories of each subsystem and the tracking trajectories received from the outside, the local cost of each subsystem is obtained throughout the prediction time window.

[0035] The local cost of each subsystem is solved to obtain the optimal reachable trajectory that minimizes the local cost, which is used as the solution for the local cost. The optimal reachable trajectory is used as the reachable trajectory at the current time and substituted into the discrete dynamics model to obtain the reachable state at the next time. The reachable state at the next time is then substituted into the stochastic reachable trajectory optimization problem model to predict the reachable trajectory at the next time.

[0036] Optionally, S2.3 sets the existence conditions for non-zero feasible solutions to the random reachable trajectory optimization problem, including:

[0037] Establish a composite control strategy for each subsystem in a distributed system, taking into account both neighbor coordination and global target tracking;

[0038] Among them, neighbor cooperation ensures that connected nodes can at least indirectly perceive the global target in the network system topology connection of each subsystem; the global target is the tracking reference trajectory signal.

[0039] Optionally, S2.4 introduces the reachable trajectory tracking error and pseudo-distribution error of the predictive control time window, as well as random disturbances, to transform the stochastic reachable trajectory optimization problem model into a distribution-based stochastic predictive control model, including:

[0040] The optimization objective is to minimize the composite cost of the subsystem, which includes stage loss, terminal loss, and cooperative loss, using the reachable trajectory and control input of the subsystem as variables.

[0041] The constraints include: for actual state variables, generating the state variables for the next time step based on the current actual state variables, control inputs, and random disturbances within the predictive control time window; predicting the actual output trajectory based on the actual state variables; ensuring the physical realizability of the ideal prediction; ensuring that the actual state variables and control inputs satisfy the operational constraints in the presence of disturbances; all prediction sequences starting from the current actual state; the composite cost function of the subsystem includes stage loss, terminal loss, and cooperative loss; generating the reachable state for the next time step based on the current reachable state and reachable control inputs within the predictive time window; each subsystem obtaining the predicted reachable output trajectory based on the reachable state; ensuring that the reachable state of each subsystem satisfies physical and logical constraints; limiting the amplitude of the reachable control inputs of each subsystem; and constraining the predicted reachable output trajectories of each subsystem and its neighboring subsystems to be consistent with the cooperative reachable output trajectories.

[0042] Specifically, by introducing the reachable trajectory tracking error and pseudo-distribution error of the predictive control time window, the stage loss and terminal loss of the subsystem are obtained.

[0043] Optionally, S2.4 also includes:

[0044] Considering the expectation of the cost function, the optimization objective of the stochastic model predictive control model is transformed into minimizing the cost of each subsystem, using the control input of each subsystem as a variable.

[0045] The input variables include state deviation variables, random disturbances, and the set of constraints for control input; the output variable is the sequence of optimal solutions, which serves as the optimal control law.

[0046] The first element of the optimal solution sequence is taken as the predictive control strategy. The predictive control strategy is used as the control input at the current moment to control the supply chain. It is then substituted into the discrete dynamics model to obtain the state variables at the next moment. The obtained state variables at the next moment are then substituted into the distribution-based stochastic predictive control model to continue updating the predictive control strategy used to solve the next moment, thereby realizing the real-time control and updating of the distributed system.

[0047] Optionally, the stability conditions in S2.5 specifically include:

[0048] To ensure the continuity of the system and its costs, the system dynamics function, stage loss function, and terminal loss function are all continuous; the stage loss function has a lower bound for all cases within the set of trajectory constraints.

[0049] The trajectory constraint set is a closed set and contains the origin; the control input set and the terminal pseudo-error set are compact sets and contain the origin.

[0050] A continuous terminal control law exists, which, when substituted into the discrete dynamics model, satisfies the trajectory constraints.

[0051] For a pair of state and control input sequences, the stage loss function has a lower bound and is continuously and strictly increasing, with a value of 0 at the origin and an infinity limit at infinity.

[0052] The perturbation input coupling matrix of the subsystem is less than or equal to the preset upper bound of the perturbation input coupling matrix;

[0053] The control parameter matrix of the terminal control strategy satisfies the condition that the system tracking reachable trajectory error converges in the distributed sense.

[0054] Compared with existing technologies, the diffusion-enhanced predictive control method for supply chain supply and demand coordination provided by this invention has at least the following beneficial effects.

[0055] (1) Ensures the solvability of the problem and improves the completeness of the method theory: This invention models the tracking problem of random external signals into an optimization problem. In the modeling process, it clearly gives the sufficient conditions for the existence of feasible solutions to the optimization problem, which theoretically ensures the feasibility of the control strategy and enhances the engineering applicability and robustness of the method.

[0056] (2) Ensure system stability: By setting sufficient conditions to ensure the stable operation of the system in a closed loop, a complete control framework is constructed, which provides support for deploying control strategies with stability guarantees in actual systems.

[0057] (3) Balancing robustness and intelligence to improve the quality of control strategy: In the stage of generating control strategy based on expert data, the diffusion model reinforcement learning method is combined, which not only retains the advantages of robust model predictive control in dealing with disturbances and constraints, but also improves the efficiency and generalization ability of strategy learning, so that the control system can show higher adaptability and intelligent decision-making ability when facing complex tasks. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the embodiments will be briefly introduced below. The features and advantages of the present invention can be more clearly understood by referring to the accompanying drawings. The accompanying drawings are schematic and should not be construed as limiting the present invention in any way. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 A flowchart of a diffusion-enhanced predictive control method for supply chain supply and demand coordination provided according to an embodiment of the present invention.

[0060] Figure 2 A schematic diagram of the coordinated regulation of the supply chain network system in the diffusion-enhanced predictive control method for supply chain supply and demand coordination provided according to an embodiment of the present invention.

[0061] Figure 3 The diagram below illustrates the intelligent robust stochastic model predictive control architecture based on a diffusion model in the diffusion-enhanced predictive control method for supply chain supply and demand coordination provided according to an embodiment of the present invention. Detailed Implementation

[0062] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.

[0063] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein. Therefore, the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0064] The following detailed description of the diffusion-enhanced predictive control method for supply chain supply and demand coordination according to an embodiment of the present invention is provided with reference to the accompanying drawings.

[0065] The supply chain supply and demand coordination diffusion-enhanced predictive control method provided by one embodiment of the present invention may include: modeling innovation, establishing a collaborative reference trajectory modeling mechanism under disturbance perception; control strategy innovation, establishing trajectory feasibility and stability conditions under robust stochastic MPC; and strategy generation innovation, learning efficient and safe control strategies based on diffusion models.

[0066] Modeling innovation, establishing a collaborative reference trajectory modeling mechanism under disturbance perception specifically includes the following processes.

[0067] First, using mathematical models System modeling is performed on the state, control, and random disturbances of each node (i.e., each subsystem) in the supply chain. This means that the state variables at the next moment can be obtained based on the current state variables, control inputs, and random disturbances of each subsystem.

[0068] Here These represent state variables, which are actual state variables, including node states such as inventory, production capacity, and demand for goods. These are control inputs, the actual control inputs, including operational quantities such as output. This refers to non-Gaussian random perturbations, including situations such as demand distortion. Represents the perturbation input matrix; This represents the state transition function of the subsystem. Represents the index of a subsystem that acts as a node in the supply chain and , For the number of subsystems, Indicates the sampling time. Indicates the variable at the sampling time The logo.

[0069] Next, a control objective is constructed for the collaborative reference trajectory model. The function of the control objective comprehensively considers the actual delivered product quantity and customer demand error of each subsystem, the consistency of the actual delivered product quantity of each node in the supply chain, and the regulation cost. By optimizing the control input, even if only the terminal node can perceive the demand tracking trajectory, other mid-to-upstream nodes can also achieve a dynamic response to the global objective based on their own and neighboring node information observed locally, combined with the estimation of random disturbances.

[0070] The effects and advantages of establishing a cooperative reference trajectory modeling mechanism under disturbance perception are significant: it can truly reflect the actual situation of information asymmetry in the supply chain; it strengthens the cooperative perception capability among multiple nodes; and it lays an analytical model foundation for subsequent robust control.

[0071] The innovation of control strategies and the establishment of trajectory feasibility and stability conditions under robust stochastic MPC include the following process.

[0072] First, reachable trajectories are introduced, and the errors between actual state variables and reachable states, as well as the errors between actual control inputs and reachable control inputs, are defined. Then, the previously set control objectives are transformed and established into three parts: balancing tracking error and regulation energy consumption by establishing a stage loss function; ensuring predictive terminal stability by establishing a terminal loss function; and achieving distributed consistency by establishing a cooperative loss function. Next, through optimization of the expected value, reachable trajectories and error control strategies are considered uniformly, and system stability conditions are given using Lyapunov theory.

[0073] The effects and advantages of establishing trajectory feasibility and stability conditions under robust stochastic MPC are: ensuring the feasibility of the control problem; and ensuring closed-loop stability.

[0074] The innovation of strategy generation, based on the diffusion model, involves the following process for efficient security control strategy learning.

[0075] First, expert data is prepared, and a trajectory set is generated using robust stochastic model predictive control (RSMPC), including the current state variables and predictive control strategies.

[0076] Next, the diffusion score model is trained and learned. , here It contains all state information. It is a regulatory strategy. This represents the score function obtained through training, and its physical meaning is the conditional probability density of the expert policy. about The logarithmic gradient of the action distribution is approximated. This approximation of the probability density gradient (score function) of the action distribution guides subsequent policy optimization. Policy training then proceeds using a combination of imitation and constraints. Finally, the policy is directly output during the inference phase, eliminating the need for sampling and reconstruction, significantly accelerating inference speed.

[0077] The effects and advantages of efficient safety control strategy learning based on diffusion models include: efficient fitting of non-Gaussian and multimodal strategies; compatibility with expert trajectories and actual disturbance information; fast inference speed and low deployment cost; and automatic satisfaction of control safety and cooperative constraints.

[0078] like Figure 1 As shown, the diffusion-enhanced predictive control method for supply chain supply and demand coordination according to another embodiment of the present invention has the following overall framework diagram: Figure 3 As shown, the specific steps include:

[0079] S1: Modeling of the supply chain networked system, including establishing a discrete dynamics model of the supply chain networked system, setting random disturbances in the discrete dynamics model, and establishing control objectives to obtain the controlled networked system model. The supply chain networked system is the supply chain system, such as... Figure 2 As shown, a networked supply chain system is a distributed system comprising multiple subsystems. Subsystems can be intelligent agents, independent operating units within the supply chain, such as suppliers, manufacturers, and distributors, and can be viewed as nodes in the networked supply chain system.

[0080] S1.1: System modeling is performed on the state variables, control inputs, and random disturbances of each subsystem in the supply chain, transforming the networked supply chain system into a discrete dynamic model for each subsystem:

[0081]

[0082] in, Indexes representing subsystems of a supply chain network system and , For the number of subsystems, Indicates the first Each subsystem at time State variables, The dimension representing the state variable. Indicates the first Each subsystem at time The control input, This indicates the dimension that controls the input. Indicates the first Each subsystem at time The output trajectory, Indicates the dimension of the output variable. Indicates the sampling time. Represents positive integers; For the first Each subsystem at time Random perturbations; Indicates the first The perturbation input matrix of each subsystem; Indicates the first The state transition function of each subsystem; Indicates the first Each subsystem at time The output matrix. The indexes of the subsystems of the above supply chain network system. Used to identify the first in the supply chain These are independent operating units, such as suppliers, manufacturers, and distributors. These subsystems can be intelligent agents and can act as nodes in the supply chain. State variables... This can represent, for example, the status of subsystems within the supply chain, such as inventory, production capacity, and demand for goods; control inputs. These are control variables, representing operational quantities such as adjustments to the output or distribution volume of a subsystem; random disturbances. Non-Gaussian perturbations can represent situations such as demand distortion. The sampling time.

[0083] Equation (1a) represents the state variables of each subsystem at the next moment, processed based on the current state variables and control input, and combined with random disturbances; Equation (1b) represents the output trajectory of each subsystem, processed based on the current state variables. Here, random disturbances are set... It is an independently distributed time random variable with finite second moments.

[0084] At sampling time closed-loop subsystem Subject to hard constraints: , ,in and It is a compact convex polyhedron, and its interior contains the origin.

[0085] In this step, a discrete dynamic model of the supply chain network system is established to dynamically describe the state variables, control inputs, and random disturbances of each subsystem of the supply chain network system at different sampling times.

[0086] S1.2: Establish control objectives, including establishing cooperative constraints on the state variables and control inputs of each subsystem, and establishing distributed cooperative control of the subsystems.

[0087] For the discrete dynamics model of the networked supply chain system established above, the core control objective is to design a distributed cooperation strategy. This ensures that each subsystem can achieve the desired output trajectory through collaborative tracking. .in, Indicates the first Collaboration strategies for individual subsystems Indicates the number of subsystems. Indicates the first Each subsystem at sampling time The predicted collaboratively reachable output trajectory considering collaborative constraints.

[0088] S1.2.1: Avoid collaborative constraints by setting state variables for each subsystem. and control input Collaboration constraints: and This ensures that the state variables and control inputs of each subsystem satisfy the constraint, thereby preventing system malfunctions due to constraint violations. This represents a pre-defined set of state variable constraints. This represents the preset set of control input constraints.

[0089] S1.2.2: Tracking performance optimization involves distributed collaborative control of subsystems, enabling each subsystem's output trajectory to quickly and accurately follow the collaboratively reachable output trajectory, and minimizing the tracking error between the collaboratively reachable output trajectory and the tracking trajectory, thereby achieving supply and demand collaborative optimization of the supply chain network system.

[0090] The above-mentioned control objectives are closely integrated with the multi-subsystem collaboration of the supply chain network system, which has the actual characteristics of random disturbances, providing a clear guide for the design of subsequent control strategies.

[0091] S2: Based on the controlled networked system model, a robust stochastic predictive control model is established to generate predictive control strategies for supply chain supply and demand regulation.

[0092] S2.1: A stochastic reachable trajectory optimization problem model is established based on the controlled networked system model to determine the optimal reachable trajectory. Its optimization objective is to minimize the cost function of tracking error and control efficiency.

[0093] Based on linearization techniques, the state transition function of the subsystem is transformed into:

[0094]

[0095] in, This represents the linearized state transition matrix. This represents the linearized control input matrix. In this implementation, the subscript of the symbol... For the identifier of the subsystem index, the symbol in This represents the parameters or variables that change with the sampling time.

[0096] To track random signals, a prediction time window is introduced. The reachable reference signal indicates the state that can be reached. and corresponding reachable control inputs Using reachable states and reachability control as decision variables, a stochastic reachable trajectory optimization problem model is established under ideal conditions without considering disturbances to determine the optimal reachable trajectory. The reachable trajectory is a set that includes reachable states and reachability control inputs.

[0097] Based on reachability states and reachability control inputs, a stochastic reachability trajectory optimization problem model is established to determine the optimal reachable trajectory:

[0098]

[0099] in, This represents the tracking signal and serves as the tracking trajectory. Indicates the first Subsystems in Time prediction The reachable trajectory, Indicates the first Subsystem Time prediction The reachable state, Indicates the first Subsystem Time prediction Reachable control input, For the reachable trajectory of the input The corresponding output trajectory, representing the predicted reachable output trajectory. For the first The cooperative achievable output trajectory of each subsystem and its neighboring subsystems Indicates the first The index of the adjacent subsystems of each subsystem and , Indicates the first The total number of neighboring subsystems of each subsystem Indicates the first Adjacent subsystems of each subsystem The predicted output trajectory can be reached; Indicates the first The local cost of each subsystem Indicates the first Subsystem Time prediction The state transition matrix, Indicates the first Subsystem Time prediction The control input matrix, Indicates the first Subsystem Time prediction The output matrix. Represents the set of constraints for a pre-defined reachable state. This represents the set of constraints for the preset reachability control input.

[0100] No. Adjacent subsystems of each subsystem The predicted output trajectory Based on the first Adjacent subsystems of each subsystem Predictable reachable state The prediction method for this reachable state can be referenced from the method for the first... A method for predicting the reachability state of a subsystem.

[0101] In this embodiment, A forecast time window for predicting the future at the current moment. Indicates the sampling time and serves as the current time. Indicates in Prediction window after time point At that moment, . Indicates the current Time in the prediction time window The identifier of the parameter or variable obtained from the time-mapping prediction; superscript It is an identifier for reachable parameters or variables. This represents the transpose of the matrix. The aforementioned distributed system is established based on the controlled networked system model from step S1.

[0102] For the first Subsystems in The entire forecast period (forecast time window) for each moment. A sequence of all reachable states. To make up for the previous round ( The optimal reachable trajectory obtained from the solution is substituted into the reachable state obtained from the solution of the discrete dynamics model. This state is then used as input to update the reachable state within the prediction time window of the current round. When substituting the optimal reachable trajectory into the discrete dynamics model, the random perturbation term is neglected, and it is treated as an ideal state without considering random perturbations. This represents the reachable state obtained from the final prediction within the prediction time window of this round.

[0103] For the first Subsystems in The entire prediction time window of the moment prediction The sequence of all reachable control inputs.

[0104] Reachable trajectories for distributed systems The term in the text represents the first term. Subsystems in Predict the entire prediction time window at any time All reachable trajectories. The optimal reachable trajectory obtained from the previous round is used as input for prediction updates within the prediction time window of this round. This is the reachable trajectory obtained from the last prediction within the prediction time window of this round.

[0105] The above establishes Equation (2a) as the optimization objective for the stochastic reachable trajectory optimization problem model. The reachable trajectories (including reachable states and reachable control inputs) of each subsystem in the predicted future time window are used as variables, and the local cost of each subsystem is minimized, that is, the deviation between the predicted reachable output trajectory and the tracking trajectory of each subsystem and the control energy consumption are minimized.

[0106] The constraints for establishing the stochastic reachable trajectory optimization problem model include: Equation (2b) describes the discrete state evolution law of each subsystem, and generates the reachable state at the next moment based on the current reachable state and reachable control input within the prediction time window; Equation (2c) indicates that each subsystem obtains the current predicted reachable output trajectory based on the current reachable state, which is used to coordinate with neighboring subsystems or track reference trajectories; Equation (2d) indicates ensuring that the reachable state of each subsystem satisfies physical and logical constraints; Equation (2e) limits the amplitude of the reachable control input of each subsystem; Equation (2f) indicates constraining the predicted reachable output trajectory of each subsystem and its neighboring subsystems to be consistent with the collaborative reachable output trajectory, so as to realize the distributed collaboration of the supply chain subsystem output.

[0107] In the above model of the random reachable trajectory optimization problem, the ideal state is not considered, and the influence of random perturbations is not taken into account.

[0108] S2.2: For the optimization objective of the stochastic reachable trajectory optimization problem model, a collaborative loss function is established, expressed as:

[0109]

[0110] in, Indicates the first The local cost of each subsystem For the first The positive definite matrix of each subsystem The global cost function of the distributed system. This represents the weighted sum of the local costs of all subsystems, used for distributed collaborative optimization. It is the reachable trajectory of the entire distributed system. Represents the reachable state sequence of the entire distributed system and , Represents the reachable control input sequence of the entire distributed system and , express Tracking trajectory in real time.

[0111] Equation (3a) above represents the local cost of each subsystem in the entire prediction time window, based on the predicted reachable output trajectory of each subsystem and the tracking trajectory received from the outside; Equation (3b) sums the local costs of all subsystems to obtain the global cost of the distributed system.

[0112] For each subsystem, the local cost problem (3a) above is solved to minimize the local cost. The solution to the local cost is denoted as... , indicating the first The predicted optimal reachable trajectory of each subsystem, including the first... The optimal reachable state of each subsystem and optimal reachable control input The error between the tracking trajectory and the optimal reachable trajectory is denoted as... Among them, subscript This represents the optimal value. Based on the first... The predicted optimal reachable trajectory of each subsystem You can also get the first Each adjacent subsystem of the subsystem The optimal reachable trajectory ,in, Including the Adjacent subsystems of each subsystem optimal reachable state and optimal reachable control input .

[0113] Based on the optimal reachable trajectories of all subsystems obtained through the solution, the optimal reachable trajectory of the distributed system as a whole that minimizes the global cost can be obtained. This serves as the solution for the global cost. Furthermore, the error between the tracking trajectory and the optimal reachable trajectory of the distributed system as a whole is... .

[0114] The sequence of optimal reachable trajectories for the entire distributed system obtained by solving the problem. This includes the optimal reachable state sequence of a distributed system. and optimal reachable control input sequence .

[0115] The obtained optimal reachable trajectory is used as the reachable trajectory at the current moment. It is substituted into the discrete dynamics model to solve for the reachable state at the next moment. Then, the reachable state at the next moment is substituted into the stochastic reachable trajectory optimization problem model to predict the reachable trajectory at the next moment.

[0116] Specifically, the local cost of each subsystem is solved to obtain the optimal reachable trajectory that minimizes the local cost, which is used as the solution for the local cost; based on the solutions for the local costs of each subsystem, the overall optimal reachable trajectory of the distributed system that minimizes the global cost is obtained, which is used as the solution for the global cost.

[0117] S2.3: Set the existence conditions for non-zero feasible solutions of the random reachable trajectory optimization problem model (2a)-(2f).

[0118] Establish the following conditions (4a)-(4c), where the origin is set to a preset set of state variable constraints. and the preset set of control input constraints If the interior point of the equation holds, then the non-zero feasible solution of the optimization problem (2a)-(2f) is represented by the following equation (4a), that is, the non-zero reachable trajectory exists.

[0119] Establish a composite control strategy for each subsystem in the distributed system, while taking into account both neighbor coordination and global target tracking:

[0120]

[0121] in, Let represent the linear subspace spanned by the column vectors of the matrix. Indicates the first Individual Subsystem and Neighboring Subsystems The relative importance weight of state differences in collaborative control. Indicates the first Can each subsystem obtain a reference trajectory signal? Indicates the first From one subsystem to its neighboring subsystem The communication link weights, Indicates the first Each subsystem at sampling time The initial state variables, Indicates the first The neighboring subsystem of each subsystem At sampling time The initial state variables, Indicates the first Subsystems in The state transition matrix; Representing the neighbor subsystem The communication topology graph nodes, To contain only the ability to The neighbor subsystem for sending information.

[0122] Furthermore, if and only if the first Subsystems in Time prediction control input matrix A full rank is achieved, and a non-zero reachable trajectory exists.

[0123] Equation (4a) above represents the composite control strategy of each subsystem in the distributed system, which takes into account both neighbor cooperation and global target tracking; the global target is the tracking reference trajectory signal. Equation (4b) represents neighbor cooperation, which means that in the network system topology connection (i.e., communication topology) of each subsystem, the connected nodes can at least indirectly perceive the global target, and each node represents each subsystem.

[0124] S2.4: Introduce random perturbations to transform the random reachable trajectory optimization problem model into a distribution-based stochastic predictive control model, which is used to provide predictive control strategies.

[0125] To better describe the problem, we introduce the reachable trajectory tracking error of the predictive control time window, i.e., the difference between the actual state variable and the reachable state. The difference: and actual control input and reachable control input The difference: ; and introduce pseudo-distribution error: ,in, It follows an expected Gaussian distribution Random variables. Among them, the actual state variables. For the first Each subsystem at sampling time Predicted time State variables, actual control inputs For the first Each subsystem at sampling time Predicted time Control input.

[0126] By introducing the aforementioned reachable trajectory tracking error and pseudo-distribution error, as well as random perturbations, the problem models (2a)-(2f) are transformed into distribution-based stochastic predictive control models:

[0127]

[0128] in, Indicates the predictive control time window and , Indicates the first The composite cost function of each subsystem Indicates the first Subsystems in Time prediction State variables, Indicates the first Subsystems in Time prediction The output trajectory, Indicates the first Subsystems in Time prediction The control input, Indicates the first Subsystem Time prediction The perturbation input coupling matrix, Indicates pseudo-distribution error exist Time prediction sampling points, Indicates the time of prediction The pseudo-distribution error, Indicates the first Subsystems in Time prediction random perturbations, Indicates the first The terminal loss function of each subsystem Indicates the first The stage loss function of each subsystem Indicates the first The measured state variables of each subsystem at the current moment. Indicates the first Random perturbations of each subsystem. The state variables obtained by substituting the predictive control strategy of the previous moment into the discrete dynamics model (1a) are solved.

[0129] , indicating the first Subsystems in Predicting the future and controlling time windows A sequence of state variables. To make up for the previous round ( The predictive control strategy obtained by solving is substituted into the state variables obtained by solving the discrete dynamics model (1a) and used as input to update the state variables within the predictive control time window of this round. This refers to the state variable obtained from the final prediction within the predictive control time window of this round.

[0130] , indicating the first Subsystems in Predicting the future and controlling time windows A sequence of state variables. For the previous round ( The obtained predictive control strategy is used as input for prediction updates within the predictive control time window of this round. This is the control input obtained from the last prediction within the predictive control time window of this round.

[0131] Equation (5a) above indicates that the optimization objective is to minimize the composite cost of each subsystem, using the reachable trajectory and control input of each subsystem as variables. The composite cost of a subsystem includes stage loss, terminal loss, and cooperative loss.

[0132] The constraints of the problem model include: Equation (5b) adds a disturbance term to the state transition law of the actual state variables. Reflecting the randomness of the actual system, the state variables for the next moment are generated based on the current actual state variables, control input, and random disturbances within the predictive control time window; Equation (5c) represents the prediction of the actual measurable output trajectory (including noise effects) based on the actual state variables; Equation (5d) indicates that the actual state variables and control input satisfy the operating constraints in the presence of disturbances; Equation (5e) indicates that all predicted sequences originate from the current actual state. Equation (5f) represents the composite cost function of the subsystem, which includes stage loss, terminal loss, and cooperative loss. Equation (5g) describes the state transition law under ideal conditions (no disturbance); it maps the ideal state (reachable state) to the output space (predicted reachable output trajectory); it guarantees the physical realizability of the ideal prediction (reachable state and reachable control input); it indicates that the ideal output (i.e., predicted reachable output trajectory) of the constrained subsystem and its neighboring subsystems tracks a common cooperative reachable output trajectory. .

[0133] Based on reachable trajectory tracking error and pseudo-distribution error, establish the stage loss function of the subsystem:

[0134]

[0135] in, For the first The positive definite matrices of each subsystem with respect to its state. Indicates the first The positive definite matrices of each subsystem with respect to the control input. Indicates the first Each subsystem at time Predicted time The pseudo-distribution error.

[0136] Based on the pseudo-distributed error, the terminal loss function of the subsystem is established:

[0137] .

[0138] Cooperative loss function of each subsystem , obtained from equation (3a).

[0139] Therefore, the composite cost of each subsystem is obtained from the stage loss, terminal loss, and collaborative loss of each subsystem.

[0140] For stochastic systems, considering the expectation of the cost function, problem (5a) is transformed into a robust stochastic model predictive control model:

[0141]

[0142] in, Represents a measure of random disturbances and describes the predictive control time window. Internal random perturbation The joint probability distribution, Represents the space of perturbation sequences. Indicates random perturbation The expected cumulative loss.

[0143] Based on the cost function, for any pseudo-distribution error The stochastic model, the robust stochastic model predictive control model (6) is transformed into:

[0144]

[0145] in, This represents the minimum expected cumulative loss. Indicates the control input of the subsystem. The set of constraints representing the control inputs of the subsystem is used as the control feasible region. The set of constraints for pseudo-distributed errors and .

[0146] In the above formula, the control input of each subsystem is used. The objective is to minimize the cost of the subsystem; the input variable is the pseudo-distribution error of the subsystem. random disturbance The set of constraints controlling the input The output variable is the optimal solution sequence, i.e., the optimal control law. Optimal solution sequence For including predictive control time windows Optimal solution elements at all times sequence.

[0147] For a given initial state, the sequence of optimal solutions is derived from a set-valued mapping. Define, where It should be noted that, It is a set-valued mapping because it is for probability measures There may be multiple solutions.

[0148] The predictive control strategy takes the first element of the optimal solution sequence, that is:

[0149]

[0150] The above formula indicates that at each moment The first element of the optimal solution sequence is selected as the predictive control strategy and updated accordingly.

[0151] Predictive control strategy based on each time step This allows for real-time control and updates. Specifically, this predictive control strategy... The current control input is used to control the supply chain, and the state variables of the next moment are obtained by substituting them into the discrete dynamics model (1a) of the subsystem. The obtained state variables of the next moment are then substituted into the distribution-based stochastic model predictive control model (5a)-(5g) to continue to update the predictive control strategy used to solve the next moment, so as to realize the real-time control and update of the distributed system.

[0152] S2.5: For stochastic predictive control models, set stability conditions to ensure the convergence of the tracking error distribution in the distributed system. The stability conditions specifically include the following conditions.

[0153] Condition 1: (Continuity of system and cost) The system dynamics function, stage loss function, and terminal loss function are all continuous; the stage loss function has a lower bound for all cases within the set of trajectory constraints.

[0154] Condition 2: (Properties of the constraint set) The trajectory constraint set is a closed set and contains the origin; the control input set and the terminal pseudo-error set are compact sets and contain the origin.

[0155] Condition 3: (Terminal control law) There exists a continuous terminal control law, which, when substituted into the discrete dynamics model, satisfies specific inclusion and inequality relationships, as well as trajectory constraints.

[0156] Condition 4: (Tracking Cost Bound) There exists a class of functions for which the stage loss function has a lower bound for the state and control input sequence pairs, and the function satisfies the following conditions: continuously and strictly increasing, with a value of 0 at the origin and a limit of infinity at infinity.

[0157] Condition 5: The perturbation coefficient matrix (perturbation input coupling matrix) of the subsystem satisfies: .in, This represents the upper bound of the preset perturbation input coupling matrix.

[0158] Condition 6: The control parameter matrix of the terminal control strategy satisfies the condition that the tracking error of the system converges in the distributional sense. This is achieved by setting the construction method of the terminal control law (Equation (9)) and by using matrix inequality constraints (Equation (10)), so that the tracking error of the system based on this terminal control law converges in the distributional sense.

[0159] For reachable control input Design terminal control strategy:

[0160]

[0161] in, Representation Subsystem exist Terminal control law at time of moment, Representation Subsystem exist Availability control input at any time Timing Subsystem exist The specific value of the pseudo-distribution error at time 1. Representation Subsystem exist The control parameter matrix at each time step. From ,for Time Prediction The reachability control input at any given time.

[0162] The following equation (10) satisfies the convergence condition of the system tracking reachable trajectory error in the sense of distribution:

[0163] (10)

[0165] in, It is a given positive constant. It is an identity matrix of appropriate dimensions. , and It is the parameter matrix to be solved.

[0166] With the assistance of control strategy (8), terminal control strategy (9), and inequality condition (10), an expert dataset is generated, including the state variables obtained in the current round. and predictive control strategies .

[0167] Therefore, by solving the optimization problem model that satisfies the constraints, the control strategy for supply chain supply and demand regulation can be obtained.

[0168] S3: Policy generation and training based on diffusion model. The robust stochastic model predictive control model is trained to obtain a robust control policy model that can meet real-time requirements.

[0169] like Figure 3 As shown, this implementation introduces a diffusion-driven offline reinforcement learning framework. Its core lies in leveraging the powerful expressive capabilities of generative diffusion models to accurately learn and simulate the potential multimodal distribution behind optimal supply and demand coordination behavior directly from high-quality expert trajectory data generated by robust stochastic model predictive control (RSMPC), thereby generating control strategies that combine high performance, high robustness, strong security, and real-time efficiency.

[0170] This offline learning framework transforms the behavior policy optimization task into a cross-entropy minimization problem, aiming to ensure that the learned policy distribution aligns with the distribution of experiential behaviors captured in expert data. When generating expert policies using a diffusion model, it abandons the traditional approach of directly regressing control actions. Instead, it first trains a score-based diffusion model to accurately learn the probability density gradient (i.e., "score") of the expert dataset. Subsequently, in the reinforcement learning policy optimization phase, the score information output by this pre-trained diffusion model is used as a guiding signal to directly drive the parameter updates of the policy network. This mechanism bypasses the cumbersome process required by traditional diffusion models in the inference phase—multiple iterations of denoising from Gaussian noise to reconstruct complex time-related action sequences—thus significantly reducing the time cost of policy generation.

[0171] To ensure safety and cooperative constraints, the Lagrange dual form of the policy optimization problem achieves a balance between imitation fidelity and constraint satisfaction, ensuring that the learned policy is both expressive and meets the requirements of safety and cooperation.

[0172] See Figure 3 To achieve the above framework, this implementation method combines value-based policy optimization with diffusion policy optimization, specifically including three stages.

[0173] Phase 1: Evaluator Training: Training two parameters as follows and The objective function is and The evaluator is a neural network. By accurately estimating the state-value function and the state-action-value function, it provides crucial gradient guidance for subsequent policy improvement, clarifying the direction of policy optimization. The evaluator training specifically includes the following steps.

[0174] Data preparation and preprocessing: Collect data on states, actions, rewards, and next states generated during environmental interactions to build an experience replay pool. Clean the data, removing outliers (such as state parameters outside the reasonable range, invalid actions, etc.), and standardize or normalize the data to ensure that data from different dimensions are of the same magnitude, facilitating neural network learning.

[0175] Neural network construction: Construct the following neural networks with parameters respectively. The objective function is A state-value function neural network, with parameters as follows: The objective function is A state-action value function (SAF) neural network takes a state as input and outputs a value estimate of that state; a state-action value function (SAF) neural network takes both a state and an action as input and outputs a value estimate of performing that action in that state. The network structure can employ a multilayer perceptron, with the number of hidden layers and neurons adjusted according to task complexity.

[0176] Network training and updates: Data batches are randomly sampled from the experience replay pool, and the target value is calculated using a temporal difference algorithm. For state-valued neural networks, parameters are updated by minimizing the mean squared error between the target value and the network output; for state-action-valued neural networks, the mean squared error is also used as the loss function, and parameters are updated based on the sampled data. Gradient descent is used to optimize the target function during training, and a target network can be introduced to improve training stability. The current network parameters are periodically copied to the target network.

[0177] Evaluation and Validation: During training, the performance of the two neural networks is periodically evaluated using a validation set. The deviation between the estimated values ​​and the true values ​​(obtainable via Monte Carlo methods, etc.) is calculated to determine if the network has achieved sufficient estimation accuracy. If the accuracy is insufficient, the network structure, training parameters, or the amount of training data needs to be adjusted, and training should be retrained.

[0178] Phase Two: Diffusion Model Training: Based on the expert dataset, optimize the parameters as follows... The objective function is A neural network. Through this process, the optimal expert policy for the corresponding state is modeled and generated.

[0179] Expert Dataset Collection and Organization: This involves collecting state-action pairs generated by domain experts executing optimal strategies in their environments, forming an expert dataset. The data in the expert dataset is then filtered to ensure its validity and representativeness, removing noisy and erroneous data. Simultaneously, the data is divided into a training set and a validation set.

[0180] Diffusion model construction: Construction parameters are The objective function is The diffusion model typically includes a forward diffusion process and a backward diffusion process. The forward diffusion process gradually introduces noise to transform the data into a Gaussian distribution, while the backward diffusion process gradually recovers the original data from the Gaussian noise. The network structure of the model can adopt architectures such as U-Net to adapt to complex data distribution modeling.

[0181] Model training process: The diffusion model is trained based on the training set data. During training, for each sample, noise is introduced at different time steps according to the forward diffusion process. Then, the backdiffusion network learns the ability to recover the data from the previous time step from the noisy data. The objective function is usually the loss function for noise prediction. By minimizing this loss function, the model parameters are updated using gradient descent. During training, the model performance is periodically evaluated on the validation set, such as the similarity between generated actions and expert actions.

[0182] Model tuning and validation: Based on the evaluation results of the validation set, the model is tuned, adjusting hyperparameters such as network depth and learning rate. After training, the model's ability to generate optimal actions in a given state is tested. By comparing the generated actions with the actions of corresponding states in the expert dataset, the model's ability to accurately model and generate optimal expert policies is verified. If the generation results are unsatisfactory, the dataset needs to be re-examined or the model structure adjusted, and training should be performed again.

[0183] Phase 3: Policy Training: Iteratively updating the neural network used to generate the control policy (parameters are...) The objective function is In this process, we not only focus on improving the performance of the strategy, but also strictly maintain constraints on deviations from the security model to ensure that the generated strategy achieves continuous performance optimization while meeting security and collaboration requirements.

[0184] Policy network initialization, construction parameters are: The objective function is A policy neural network is used to generate control policies. The network input is the environment state, and the output is the action distribution. An appropriate network structure is selected based on the task characteristics; for example, a network with Gaussian output can be used for continuous action spaces, while a network with softmax output can be used for discrete action spaces. The network parameters are then randomly initialized.

[0185] The policy update criteria are defined, clearly defining the goals of policy updates: to improve policy performance while strictly maintaining constraints on deviations from the safety model. Performance improvement can be achieved by maximizing cumulative rewards, utilizing the value estimate provided by the evaluator trained in the first stage to calculate the policy gradient and guide the policy update towards a better direction. Safety constraints can be incorporated into the objective function by introducing penalty terms, etc., to penalize the objective function when the policy may deviate from the safety model, thus limiting the direction of policy updates.

[0186] The iterative update process starts with the initial policy and proceeds through multiple rounds of updates. In each iteration, the current policy interacts with the environment to collect new empirical data, which is then stored in the experience replay pool. Next, the policy gradient is calculated by combining the value information provided by the evaluator in the first phase with the expert policy information generated by the diffusion model in the second phase. Based on the policy gradient and safety constraints, the policy network parameters are updated. Simultaneously, after each update, the policy is checked to ensure it meets safety and cooperation requirements. For example, a safety verification module verifies the policy's behavior under critical states; if not, the update magnitude is adjusted or the policy is corrected.

[0187] Performance evaluation and convergence determination: After each iteration, the performance of the current strategy is evaluated, including average cumulative reward and task completion rate. By comparing performance metrics across multiple iterations, convergence is determined. Training stops when performance metrics stabilize and reach the preset performance target, while the strategy consistently meets safety and coordination requirements; otherwise, iterative updates continue.

[0188] S4: Apply the trained robust control strategy model to the supply chain to generate control strategies for execution within the supply chain. Specifically, the trained robust control strategy model can be integrated into the supply chain's ERP (Enterprise Resource Planning) system in real time. Based on information such as order flow and inventory data in the ERP system, ordering and production scheduling strategies are output to achieve supply and demand coordination; and safety stock is automatically maintained under demand fluctuations, and production lines are quickly adjusted in the event of sudden equipment failures. This step includes the following steps.

[0189] S4.1: Real-time interface architecture between the model and the ERP system, data interface layer development: Develop standardized data interfaces to achieve bidirectional communication between the robust control strategy model and the supply chain ERP system. On one hand, by extracting core indicators such as order flow data, inventory data, and production data from the ERP system in real time, a dynamic dataset for model input is constructed; on the other hand, the output results such as ordering strategies and production scheduling instructions generated by the model are pushed to the execution module of the ERP system in a structured format to ensure that the strategy can directly drive the business process.

[0190] S4.2: Strategy generation in supply and demand coordination scenarios, dynamic generation of ordering strategies: The robust control strategy model predicts demand trends based on real-time order flow and calculates the optimal order quantity in combination with the current inventory level.

[0191] S4.3: Adaptive Control Mechanism under Dynamic Disturbances, Maintaining Safety Stock under Demand Fluctuations: The robust control strategy model incorporates a dynamic safety stock constraint module. When demand experiences a sudden surge (such as order peaks caused by promotional activities), the model triggers an inventory warning mechanism while generating an ordering strategy to ensure that inventory levels do not fall below the safety line. If demand remains weak, the model automatically reduces order quantities and digests excess inventory through promotional production scheduling, avoiding capital tied up.

[0192] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0193] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0194] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A diffusion-enhanced forecasting and control method for supply chain supply and demand coordination, characterized in that, include: S1: Modeling of the supply chain networked system, including establishing a discrete dynamics model of the supply chain networked system, setting the random disturbances in the discrete dynamics model, and establishing control objectives to obtain the controlled networked system model, wherein the supply chain networked system is a distributed system including multiple subsystems; S2: Based on the controlled networked system model, a robust stochastic predictive control model is established to generate predictive control strategies for supply chain supply and demand regulation. S3: Policy generation and training based on diffusion model: Train the robust stochastic model predictive control model to obtain a robust control policy model that meets real-time requirements. S4: Apply the trained robust control strategy model to the supply chain to generate control strategies and provide them to the execution process in the supply chain; Specifically, S2 includes: S2.1: Based on the controlled networked system model, the reachability state of the prediction time window is introduced to establish a stochastic reachability trajectory optimization problem model to determine the optimal reachability trajectory. The optimization objective is to minimize the error between the predicted reachability output trajectory and the tracking trajectory, as well as the control energy consumption. S2.2: For the optimization objective of the stochastic reachable trajectory optimization problem model, establish a collaborative loss function, and obtain the local cost of each subsystem in the entire prediction time window based on the predicted reachable output trajectory and tracking trajectory of each subsystem. S2.3: Set the condition that the model of the random reachable trajectory optimization problem has a non-zero feasible solution; S2.4: Considering random perturbations, the random reachable trajectory optimization problem model is transformed into a distribution-based stochastic predictive control model to provide predictive control strategies; S2.5: Set the stability condition for the convergence of the tracking error distribution in the distributed system.

2. The diffusion-enhanced forecasting and control method for supply chain supply and demand coordination according to claim 1, characterized in that, S1 specifically includes: S1.1: System modeling of the state variables, control inputs and random disturbances of each subsystem in the supply chain, modeling the networked supply chain system as a discrete dynamic model for each subsystem, where the subsystem is an independent operating unit in the supply chain, including suppliers, manufacturers and distributors; S1.2: Establish control objectives, including establishing cooperative constraints on the state variables and control inputs of each subsystem, and establishing distributed cooperative control of the subsystems.

3. The diffusion-enhanced forecasting and control method for supply chain supply and demand coordination according to claim 2, characterized in that, In S1.1: Each subsystem processes the state variables and control inputs at the current moment, and combines them with random disturbances to obtain the state variables at the next moment, and each subsystem processes the state variables at the current moment to obtain the output of each subsystem; The state variables of the subsystem are the subsystem's inventory, production capacity, and demand for goods; The control inputs of the subsystem are control inputs, including the adjustment amounts of the subsystem's output and the adjustment amounts of the delivery amount; The random disturbance of the subsystem is a non-Gaussian disturbance, representing demand distortion. The random disturbance is set to be an independently distributed time random variable with finite second moments.

4. The diffusion-enhanced forecasting and control method for supply chain supply and demand coordination according to claim 1, characterized in that, In S2.1, reachable states and corresponding control inputs are introduced as decision variables to establish a stochastic reachable trajectory optimization problem model, including: The optimization objective of establishing a stochastic reachable trajectory optimization problem model is to minimize the deviation between the predicted reachable output trajectory and the tracking trajectory of each subsystem in the distributed system and the control energy consumption, using the reachable trajectories of each subsystem in the predicted future time window as variables. The reachable trajectory is a set including reachable states and reachable control inputs. The constraints for establishing a stochastic reachable trajectory optimization problem model include: establishing the discrete state evolution law of each subsystem; generating the reachable state at the next moment based on the current reachable state and reachable control input within the prediction time window; obtaining the predicted reachable output trajectory of each subsystem based on the reachable state; ensuring that the reachable state of each subsystem satisfies physical and logical constraints; limiting the amplitude of the reachable control input of each subsystem; and constraining the predicted reachable output trajectory of each subsystem and its neighboring subsystems to be consistent with the cooperative reachable output trajectory.

5. The diffusion-enhanced forecasting and control method for supply chain supply and demand coordination according to claim 4, characterized in that, S2.2 specifically includes: Based on the predicted reachable trajectories of each subsystem and the tracking trajectories received from the outside, the local cost of each subsystem is obtained throughout the prediction time window. The local cost of each subsystem is solved to obtain the optimal reachable trajectory that minimizes the local cost, which is used as the solution for the local cost. The optimal reachable trajectory is used as the reachable trajectory at the current time and substituted into the discrete dynamics model to obtain the reachable state at the next time. The reachable state at the next time is then substituted into the stochastic reachable trajectory optimization problem model to predict the reachable trajectory at the next time.

6. The diffusion-enhanced forecasting and control method for supply chain supply and demand coordination according to claim 4, characterized in that, S2.3 sets the following conditions for the existence of non-zero feasible solutions to the random reachable trajectory optimization problem: Establish a composite control strategy for each subsystem in a distributed system, taking into account both neighbor coordination and global target tracking; Among them, neighbor cooperation ensures that connected nodes can at least indirectly perceive the global target in the network system topology connection of each subsystem; the global target is the tracking reference trajectory signal.

7. The diffusion-enhanced forecasting and control method for supply chain supply and demand coordination according to claim 4, characterized in that, S2.4 introduces reachable trajectory tracking error and pseudo-distribution error within the predictive control time window, as well as random disturbances, transforming the stochastic reachable trajectory optimization problem model into a distribution-based stochastic predictive control model, including: The optimization objective is to minimize the composite cost of the subsystem, which includes stage loss, terminal loss, and cooperative loss, using the reachable trajectory and control input of the subsystem as variables. The constraints include: for actual state variables, generating the state variables for the next time step based on the current actual state variables, control inputs, and random disturbances within the predictive control time window; predicting the actual output trajectory based on the actual state variables; ensuring the physical realizability of the ideal prediction; ensuring that the actual state variables and control inputs satisfy the operational constraints in the presence of disturbances; all prediction sequences starting from the current actual state; the composite cost function of the subsystem includes stage loss, terminal loss, and cooperative loss; generating the reachable state for the next time step based on the current reachable state and reachable control inputs within the predictive time window; each subsystem obtaining the predicted reachable output trajectory based on the reachable state; ensuring that the reachable state of each subsystem satisfies physical and logical constraints; limiting the amplitude of the reachable control inputs of each subsystem; and constraining the predicted reachable output trajectories of each subsystem and its neighboring subsystems to be consistent with the cooperative reachable output trajectories. Specifically, by introducing the reachable trajectory tracking error and pseudo-distribution error of the predictive control time window, the stage loss and terminal loss of the subsystem are obtained.

8. The diffusion-enhanced forecasting and control method for supply chain supply and demand coordination according to claim 7, characterized in that, S2.4 also includes: Considering the expectation of the cost function, the optimization objective of the stochastic model predictive control model is transformed into minimizing the cost of each subsystem, using the control input of each subsystem as a variable. The input variables include state deviation variables, random disturbances, and the set of constraints for control input; the output variable is the sequence of optimal solutions, which serves as the optimal control law. The first element of the optimal solution sequence is taken as the predictive control strategy. The predictive control strategy is used as the control input at the current moment to control the supply chain. It is then substituted into the discrete dynamics model to obtain the state variables at the next moment. The obtained state variables at the next moment are then substituted into the distribution-based stochastic predictive control model to continue updating the predictive control strategy used to solve the next moment, thereby realizing the real-time control and updating of the distributed system.

9. The diffusion-enhanced forecasting and control method for supply chain supply and demand coordination according to claim 1, characterized in that, The specific conditions for stability in S2.5 include: To ensure the continuity of the system and its costs, the system dynamics function, stage loss function, and terminal loss function are all continuous; the stage loss function has a lower bound for all cases within the set of trajectory constraints. The trajectory constraint set is a closed set and contains the origin; the control input set and the terminal pseudo-error set are compact sets and contain the origin. A continuous terminal control law exists, which, when substituted into the discrete dynamics model, satisfies the trajectory constraints. For a pair of state and control input sequences, the stage loss function has a lower bound and is continuously and strictly increasing, with a value of 0 at the origin and an infinity limit at infinity. The perturbation input coupling matrix of the subsystem is less than or equal to the preset upper bound of the perturbation input coupling matrix; The control parameter matrix of the terminal control strategy satisfies the condition that the system tracking reachable trajectory error converges in the distributed sense.

Citation Information

Patent Citations

  • Dynamic body synergetic method for supply chain system

    CN103810562A

  • Supply chain collaborative management method and system based on industrial internet identifier analysis

    CN117522225A

  • Multi-layer supply chain network inventory management method

    CN118278861A

  • Supply chain full-link risk early warning platform and method based on big data

    CN120278526A