An industry chain adaptive scheduling method and system fusing deep reinforcement learning
By integrating deep reinforcement learning methods, a dynamic causal graph is constructed in real time and a new policy paradigm is generated. This solves the problems of lag and poor adaptability of scheduling policies in existing technologies, realizes early identification and dynamic adaptation of the industrial chain system, improves the foresight and robustness of scheduling, and ensures the stable operation of the industrial chain.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA SHENHUA ENERGY CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-06-23
AI Technical Summary
Existing supply chain scheduling methods suffer from problems such as scheduling strategy lag and poor adaptability to deep system changes when facing highly uncertain and dynamically changing complex systems. This leads to feedback lag and the inapplicability of sticking to the original goal optimization behavior, and may even produce negative effects.
By integrating deep reinforcement learning, a dynamic causal graph is constructed in real time. Causal entropy is used to monitor system stability, predict potential risks, and generate new strategy paradigms based on instability information. Scheduling instructions are determined through simulation interaction, dynamically adapting to changes in the supply chain logic, and achieving early identification and forward-looking prediction.
It breaks through the limitations of traditional reliance on lagging performance indicators, and realizes early identification and forward prediction of causal instability within the industrial chain system. This improves the foresight, adaptability and robustness of scheduling, effectively avoids unexpected risks, and ensures the efficient and stable operation of the industrial chain.
Smart Images

Figure CN122264455A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of supply chain scheduling technology, specifically to an adaptive supply chain scheduling method and system that integrates deep reinforcement learning. Background Technology
[0002] Supply chain scheduling is a core technology area in modern manufacturing and logistics management. Its goal is to achieve specific operational objectives by coordinating all aspects of the supply chain, from raw material procurement, production and processing, inventory management to product distribution.
[0003] Existing supply chain scheduling methods are mainly based on mathematical programming or traditional simulation optimization techniques. These methods typically begin by establishing a mathematical model describing the operation of the supply chain and setting one or more fixed optimization objectives, such as minimizing total cost or minimizing order delivery cycle. The scheduling scheme is then obtained by solving the model. Some more advanced methods incorporate reinforcement learning, enabling the scheduling agent to learn strategies through interaction with the simulation environment to adapt to a certain degree of environmental dynamism.
[0004] However, the aforementioned existing technologies have inherent limitations when dealing with complex supply chain systems characterized by high uncertainty and dynamic changes. These methods typically rely indirectly on monitoring final performance indicators, such as production costs, inventory levels, or order delay rates, to assess system stability. This reliance leads to feedback lag; by the time performance indicators show significant deterioration, the internal causal logic of the system has already become unstable and persisted for some time, thus missing the opportunity for early intervention.
[0005] Furthermore, existing methods have shortcomings in their adaptive adjustment mechanisms. When a system detects a performance degradation, its adjustment behavior is typically limited to re-optimizing the parameters of the scheduling strategy within a pre-defined, fixed optimization objective function framework. For situations where the system's operating logic undergoes fundamental changes due to external shocks or internal structural evolution, existing technologies cannot identify such profound changes, let alone adjust their own optimization objectives accordingly. Therefore, when the system faces unexpected operating modes, adhering to the original objective will no longer be applicable and may even produce negative effects. Summary of the Invention
[0006] This invention provides an adaptive scheduling method and system for the industrial chain that integrates deep reinforcement learning, in order to solve the problems of lagging scheduling strategies and poor adaptability to deep changes in the system in the prior art.
[0007] In a first aspect, the present invention provides an adaptive scheduling method for the industrial chain that integrates deep reinforcement learning, the method comprising: Multiple variables within the target industry chain are acquired in real time, and a dynamic causal graph is constructed and updated in real time. The dynamic causal graph is used to characterize the causal relationship between variables, and the edges between variables are attached with causal entropy, which represents the strength of causality. Predict potential risk areas based on causal entropy, and determine whether to trigger strategy evolution based on at least one of the causal entropy and the prediction result. If policy evolution is triggered, the instability information is determined, and a new policy paradigm is generated and activated based on the instability information. The new strategy paradigm interacts with the simulation environment of the target industry chain to determine effective scheduling instructions, and then deploys these instructions to the physical execution system of the target industry chain.
[0008] The adaptive scheduling method for the industrial chain that integrates deep reinforcement learning provided by this invention constructs and updates a dynamic causal graph containing causal strength quantification edges in real time. It monitors system stability and predicts potential risks using causal entropy, triggers policy evolution, generates a new policy paradigm based on instability information, and determines and deploys scheduling instructions through simulation interaction. This method breaks through the limitations of traditional methods that rely on lagging performance indicators, achieves early identification and forward prediction of causal instability within the system, dynamically adapts to changes in the logic of the industrial chain, does not require fixed optimization targets, significantly improves the foresight, adaptability, and robustness of scheduling, effectively avoids unexpected risks, and ensures the efficient and stable operation of the industrial chain.
[0009] In an optional implementation, after deploying effective scheduling instructions to the physical execution system of the target supply chain, the method further includes: Obtain actual execution data from the physical execution system and evaluate the effectiveness of the new strategy paradigm based on the actual execution data; If the new strategy paradigm is effective, the new strategy paradigm and the corresponding instability information will be stored in the paradigm archive.
[0010] The supply chain adaptive scheduling method provided by this invention, which integrates deep reinforcement learning, evaluates the effectiveness of the new strategy paradigm by acquiring actual data from the physical execution system after deploying effective scheduling instructions. The effective paradigm and corresponding instability information are stored in the paradigm archive to form a knowledge loop. The practicality of the strategy is verified by actual data, ensuring that the scheduling instructions meet the actual operational needs of the supply chain. This provides rich historical experience for subsequent risk prediction and strategy generation, further enhancing the adaptive capability of supply chain scheduling.
[0011] In one optional implementation, multiple variables within the target industry chain are acquired in real time, and a dynamic causal graph is constructed and updated in real time, including: The data of each variable is preprocessed to obtain the standard variable data. Based on the time series data of each standard variable data, the dynamic causal discovery algorithm is used to determine the causal relationship between each variable based on a preset time window. Obtain time-series data of two variables with a causal relationship, and calculate the transfer entropy between the two variables based on the time-series data; The transfer entropy between variables with causal relationships is normalized to obtain a probability distribution, and the Shannon entropy of the probability distribution is calculated as the causal entropy of the dynamic causal graph.
[0012] The supply chain adaptive scheduling method fused with deep reinforcement learning provided by this invention achieves direct monitoring of the internal logical stability of the supply chain system by constructing a dynamic causal graph and calculating the system's causal entropy. Compared with methods that rely on lagging performance indicators such as cost and delay, it can identify potential instability risks before the system's performance indicators deteriorate significantly based on changes in the strength of causal relationships, thus providing an advance time window for strategy adjustment.
[0013] In one optional implementation, a potential risk area is predicted based on causal entropy, and a determination is made on whether to trigger strategy evolution based on at least one of the causal entropy and the prediction result, including: The causal entropy change rate is calculated based on the causal entropy at each time point, and the causal entropy is compared with the causal entropy threshold and the causal entropy change rate is compared with the change rate threshold. If the causal entropy is greater than the causal entropy threshold or the causal entropy change rate is greater than the change rate threshold, then the passive response triggers the strategy evolution. The system acquires historical evolutionary trajectory data from the paradigm archive, learns from this data to obtain meta-learning results, and then predicts the latest evolutionary trajectory data based on these results. If the prediction results meet preset conditions, the system proactively triggers policy evolution.
[0014] In one optional implementation, the historical evolutionary trajectory data includes: historical instability information and corresponding policy paradigms. Learning is performed based on the historical evolutionary trajectory data to obtain meta-learning results. The latest evolutionary trajectory data is then predicted based on the meta-learning results to obtain prediction results. If the prediction results meet preset conditions, policy evolution is triggered, including: Learning from historical evolutionary trajectory data yields a predictive model that represents the relationship between unstable information and policy paradigms, which serves as the meta-learning result. The latest evolutionary trajectory data is input into the trained prediction model to obtain the prediction vector, and the potential risk area location map is determined based on the prediction vector. The potential risk area location map is used to characterize the probability of future instability. The aggregated risk value is calculated based on the potential risk area location map. When the aggregated risk value is greater than the risk trigger threshold, the proactive predictive trigger strategy evolution is initiated.
[0015] The industry chain adaptive scheduling method provided by this invention introduces a meta-learning mechanism for historical evolutionary trajectories in the paradigm archive, enabling the scheduling system to predict potential future risks. It is not limited to passively responding to current system instability, but can actively identify areas where problems may occur in the future based on the learned evolutionary patterns. Thus, the adjustment of scheduling strategies is transformed from an event-driven passive mode to a forward-looking proactive prediction mode.
[0016] In an optional implementation, if policy evolution is triggered, instability information is determined, a new policy paradigm is generated and activated based on the instability information, including: If the evolution of the strategy is triggered by a passive response, a causal unstable region localization map is generated as an unstable information characterizing the current unstable mode. If the strategy evolution is triggered by proactive prediction, a potential risk area location map is predicted, which serves as instability information representing future instability patterns. Compare the structural similarity between the unstable information and the historical unstable information in the paradigm archive. If the structural similarity exceeds the similarity threshold, the corresponding historical unstable information is used as the matching item for the unstable information. Based on the matching terms with instability information, the policy paradigm corresponding to the matching terms is determined as the new policy paradigm.
[0017] In an optional implementation, if the structural similarity does not exceed a similarity threshold, a new strategy paradigm generation process is initiated, including: Construct a policy primitive library and encode each policy primitive in the library to obtain a primitive vector; The global state vector, instability information vector, and random noise vector of the target industry chain are input into the generator of the generative adversarial network. The instability information vector is used as the query and the primitive vectors are used as the key and value. The attention mechanism in the generator is used to generate the weights of each primitive vector. We perform weighted calculations on the primitive vectors based on their weights to obtain a composite reward function, and then encapsulate this composite reward function into a new policy paradigm.
[0018] The adaptive scheduling method for the industrial chain that integrates deep reinforcement learning provided by this invention achieves dynamic adjustment of the scheduling optimization objective by retrieving or generating a policy paradigm with a new reward function as its core. When the system faces fundamental logical changes, this method is not limited to optimizing policy parameters under the original reward function, but can generate a completely new reward function to redefine the optimization problem itself. This enables the scheduling system to adapt to unexpected system operation modes that are difficult to handle by traditional fixed-objective methods, thus expanding the scope of application of its adaptive scheduling.
[0019] Secondly, this invention provides an industry chain adaptive scheduling system that integrates deep reinforcement learning, the system comprising: The causal graph construction and update module is used to acquire multiple variables within the target industry chain in real time, construct and update a dynamic causal graph in real time. The dynamic causal graph is used to characterize the causal relationship between variables, and the edges between variables are attached with causal entropy, which characterizes the causal strength. The strategy evolution triggering module is used to predict potential risk areas based on causal entropy, and determine whether to trigger strategy evolution based on at least one of causal entropy and prediction results. The strategy paradigm management module is used to determine instability information if strategy evolution is triggered, and to generate and activate a new strategy paradigm based on the instability information. The effective scheduling instruction deployment module is used to interact with the simulation environment of the target industry chain to determine the effective scheduling instructions and deploy them to the physical execution system of the target industry chain.
[0020] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect or any corresponding embodiment thereof.
[0021] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description
[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the first process of the industry chain adaptive scheduling method integrating deep reinforcement learning according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the strategy deployment, evaluation and knowledge closure process in the industry chain adaptive scheduling method integrating deep reinforcement learning according to an embodiment of the present invention. Figure 4This is a schematic diagram of the second process of the industry chain adaptive scheduling method integrating deep reinforcement learning according to an embodiment of the present invention. Figure 5 This is a flowchart illustrating the proactive predictive evolution triggering based on meta-learning in the industry chain adaptive scheduling method integrating deep reinforcement learning according to an embodiment of the present invention. Figure 6 This is a schematic diagram of the policy paradigm retrieval and generation process in the supply chain adaptive scheduling method integrating deep reinforcement learning according to an embodiment of the present invention. Figure 7 This is a complete flowchart of a specific embodiment of the industry chain adaptive scheduling method integrating deep reinforcement learning according to an embodiment of the present invention. Figure 8 This is a structural block diagram of an industry chain adaptive scheduling system that integrates deep reinforcement learning according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0026] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0027] As an optional application scenario of this invention, such as Figure 1 As shown, the industry chain adaptive scheduling system may include at least one terminal device and at least one server. Figure 1The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.
[0028] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.
[0029] Existing supply chain scheduling methods, when faced with complex, dynamic, and highly uncertain system environments, often rely on fixed optimization objectives and lagging performance feedback. This results in insufficient adaptability to sudden and fundamental changes in system logic, making it difficult to achieve proactive risk avoidance and strategy self-evolution. To address these issues, this invention provides a supply chain adaptive scheduling method integrating deep reinforcement learning. By constructing and updating a dynamic causal graph containing quantized edges of causal strength in real time, it monitors system stability and predicts potential risks using causal entropy. After triggering strategy evolution, it generates a new strategy paradigm based on instability information. The scheduling instructions are determined and deployed through simulation interaction, achieving early identification and forward-looking prediction of internal causal instability and dynamically adapting to changes in supply chain logic.
[0030] According to an embodiment of the present invention, an embodiment of an adaptive scheduling method for supply chains integrating deep reinforcement learning is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0031] This embodiment provides an adaptive scheduling method for the supply chain that integrates deep reinforcement learning, which can be used in the aforementioned computer system. Figure 2 This is a flowchart of an industry chain adaptive scheduling method incorporating deep reinforcement learning according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Acquire multiple variables within the target industry chain in real time, construct and update a dynamic causal graph in real time. The dynamic causal graph is used to characterize the causal relationship between variables, and the edges between variables are attached with causal entropy, which characterizes the causal strength.
[0032] Specifically, multivariate time series data are continuously collected from the industrial chain simulation environment, and a dynamic causal map representing the causal relationships between key variables within the system is constructed and updated in real time using a dynamic causal discovery algorithm. The dynamic causal graph is a time-varying directed acyclic graph. Nodes represent key variables in the target industry chain, and edges represent causal relationships between variables, with each edge accompanied by a quantified causal strength. .
[0033] Step S202: Predict potential risk areas based on causal entropy, and determine whether to trigger strategy evolution based on at least one of the causal entropy and the prediction result.
[0034] Specifically, by using dynamic causal graphs to calculate causal entropy, the stability of the current causal structure of the industrial chain system can be quantitatively assessed. This can determine whether there is a risk and whether a strategy evolution process needs to be initiated. When the causal entropy or its rate of change meets the preset causal entropy instability criterion, a passive response triggers the strategy evolution.
[0035] Based on causal entropy, potential risk areas in the future can be predicted. By learning from historical experience, the causal instability risk of the industrial chain system in the future can be predicted, thereby triggering forward-looking strategy evolution. A meta-learning model analyzes the historical evolution trajectory to predict the potential risk areas of the system in the future and outputs a potential risk area location map. When the risk value of this location map exceeds a preset threshold, a proactive and predictive strategy evolution is triggered. The preset threshold is determined through statistical analysis of historical data. Specifically, the system collects the predicted probability distribution of historical operating data when no failures have occurred, and selects the 95th percentile (or 99th percentile) of this distribution as the threshold. Only when the predicted risk probability exceeds the fluctuation range under most normal conditions is it considered a risk.
[0036] In step S203, if policy evolution is triggered, the instability information is determined, and a new policy paradigm is generated and activated based on the instability information.
[0037] Specifically, when policy evolution is triggered, the causal instability information representing the current or future system instability pattern is first determined based on the trigger source (passive response or active prediction). Then, this causal instability information is used as a query to search a pre-defined paradigm archive. If a historical policy paradigm with a graph structure similarity exceeding a threshold is found, it is directly identified as the new policy paradigm; otherwise, a generative model is initiated to generate a completely new policy paradigm based on the current system state and the causal instability information.
[0038] Step S204: Interact the new strategy paradigm with the simulation environment of the target industry chain to determine the effective scheduling instructions, and deploy the effective scheduling instructions to the physical execution system of the target industry chain.
[0039] Specifically, the updated policy paradigm is activated. This new policy paradigm is a new reward function. Throughout the entire policy learning cycle, all reward signals fed back to the agent from the simulation environment will be calculated based on this new reward function. The new policy paradigm is then loaded into the simulation environment of the target industry chain. A deep reinforcement learning scheduling agent with a built-in graph neural network is used to maximize the long-term cumulative return defined by the new reward function. Through extensive interaction with the industry chain simulation environment, the agent learns and ultimately determines and outputs a set of effective scheduling instructions.
[0040] Based on the new policy paradigm, a deep reinforcement learning scheduling agent is driven to learn an executable scheduling policy that matches the paradigm.
[0041] The specific operation for activating the new policy paradigm involves loading a new reward function, its core component, into the industry chain simulation environment. Subsequently, throughout the entire policy learning cycle, all reward signals fed back to the agent from the simulation environment will be calculated based on this new reward function.
[0042] In a specific embodiment, the internal network structure of the deep reinforcement learning scheduling agent adopts a Graph Attention Network (GAT). The agent's policy network receives the current supply chain state graph as input. The node feature vectors of this state graph contain real-time business data of each node (e.g., inventory levels, output rates), while the adjacency matrix defines the physical or business connection relationships between nodes. For a manufacturing node in the supply chain, its node feature vector may specifically include: current raw material inventory, work-in-process (WIP) quantity, finished goods inventory, mean time between failures (MTBF), overall equipment efficiency (OEE), current batch defect rate, backlog of production orders, and production demand forecast for the next cycle. For a distribution center node, its feature vector may include: inventory levels of various products, inbound rate, outbound rate, average inventory turnover days, and the number of pending delivery orders. All these values are normalized before being input into the GAT network. The graph attention network, through its self-attention mechanism, assigns different attention weights to the neighboring nodes of each node, thereby aggregating neighborhood information to generate an updated representation of each node. The updated node representation is then fed into a graph readout layer and aggregated into a graph-level vector representing the state of the entire system.
[0043] The graph-level vector is input into subsequent fully connected layers, ultimately outputting an action probability distribution. The agent's action space is a multi-dimensional discrete or continuous space, where each dimension corresponds to a schedulable decision variable, such as adjusting supplier A's raw material shipment quantity to x units or setting production line B's production rate to y pieces / hour. The agent samples the probability distribution output by the policy network to generate a specific action, i.e., a set of scheduling instructions.
[0044] The scheduling and execution module employs the Proximal Policy Optimization (PPO) algorithm to train the agent. The training process is an iterative loop where the agent executes its current policy in a simulated industry chain environment, generating a series of trajectory data related to states, actions, and rewards. Then, using this trajectory data, the policy loss and value loss are calculated according to the objective function of the PPO algorithm, and the network parameters of the graph attention network and subsequent fully connected layers are updated through backpropagation. The goal of this process is to find an optimal set of network parameters that maximizes the long-term cumulative reward defined by the newly activated reward function.
[0045] The training process terminates when the agent's policy performance converges or after reaching the preset number of training iterations. At this point, the trained policy network constitutes an effective scheduling policy. The scheduling execution module will use this fixed policy network to directly determine and output the current optimal scheduling instruction through a forward propagation calculation when it receives a new real-time status of the supply chain. This set of scheduling instructions is the effective scheduling instruction and will be transmitted to the next step for deployment and evaluation.
[0046] The supply chain adaptive scheduling method fused with deep reinforcement learning provided in this embodiment constructs and updates a dynamic causal graph containing causal strength quantification edges in real time. It monitors system stability and predicts potential risks using causal entropy, triggers policy evolution, generates new policy paradigms based on instability information, and determines and deploys scheduling instructions through simulation interaction. This method breaks through the limitations of traditional methods that rely on lagging performance indicators, achieves early identification and forward prediction of causal instability within the system, dynamically adapts to changes in supply chain logic, does not require fixed optimization targets, significantly improves the foresight, adaptability, and robustness of scheduling, effectively avoids unexpected risks, and ensures the efficient and stable operation of the supply chain.
[0047] In some alternative implementations, after deploying effective scheduling instructions to the physical execution system of the target supply chain, the method further includes: Step S205: Obtain the actual execution data of the physical execution system and evaluate the effectiveness of the new strategy paradigm based on the actual execution data.
[0048] Specifically, effective scheduling instructions are transmitted to an actual physical execution system within the supply chain for deployment. Simultaneously, an evaluation unit is used to assess the practical application effectiveness of the newly activated strategy paradigm. Evaluation criteria include whether the system's causal entropy recovers to a stable level within a specified time, and whether the key performance indicators of the supply chain meet the standards.
[0049] like Figure 3 The diagram shown is a schematic of the strategy deployment, evaluation, and knowledge closure process according to this embodiment. It is mainly executed by an evaluation unit and a knowledge accumulation module 500 to complete the feedback and evolution closure of the adaptive scheduling method.
[0050] Effective scheduling instructions are transmitted to the physical execution systems in the supply chain through a system interface. This interface is responsible for translating abstract scheduling instructions (e.g., setting a production rate to a specific value or adjusting an inventory threshold) into specific operational commands that can be executed by the physical systems, such as sending an updated production plan to the manufacturing execution system or submitting new purchase order parameters to the enterprise resource planning system.
[0051] Once deployment is complete, immediately initiate the effectiveness evaluation process for the newly activated strategy paradigm. An evaluation unit will be evaluated within a pre-defined evaluation time window following the deployment of the instruction. Internally, the status of the supply chain system is continuously monitored. The criteria for successful evaluation include two parallel conditions that must be met simultaneously.
[0052] Conditions for causal stability recovery: within the evaluation time window At the end, the causal entropy of the system It must decrease to a preset steady-state threshold. Below, this threshold It is set as the average causal entropy of the system during a long historical period of stable operation.
[0053] Key performance indicator compliance conditions: within the evaluation time window Within the supply chain, a set of key performance indicators (KPIs) must be maintained within an acceptable range. For example, the on-time delivery rate of orders must be no less than 98%, and the average inventory turnover days must not exceed 30 days.
[0054] Evaluation unit in At the end, a comprehensive assessment is made to determine whether both of the above conditions have been met. If both conditions are met, the strategy evolution is evaluated as successful; otherwise, it is evaluated as a failure.
[0055] Step S206: If the new strategy paradigm is effective, store the new strategy paradigm and the corresponding instability information in the paradigm archive.
[0056] Specifically, when the effectiveness evaluation result is successful, it indicates that the new strategy paradigm is effective. The new strategy paradigm and its corresponding causal instability information that led to this evolution are stored as a new data entry in the paradigm archive used by the strategy paradigm management module 300, completing the closed-loop accumulation of knowledge for subsequent proactive predictive triggering of meta-learning and retrieval by the strategy paradigm management module 300.
[0057] The knowledge accumulation module 500 is activated only when the evaluation result is successful. This will cause the original causal unstable information of this strategy evolution (whether it comes from the location map of passive response or from the location map of active prediction) and the strategy paradigm that has just been verified as successful (whose core is the new reward function definition) to be encapsulated into a new and complete problem-solution data entry.
[0058] Subsequently, the knowledge accumulation module 500 atomically stores this newly generated data entry as an independent record in the paradigm archive used by the strategy paradigm management module 300. After this operation, the new entry becomes a permanent part of the archive, usable for future paradigm retrieval by the strategy paradigm management module 300 and meta-learning by proactive predictive triggering submodules. This process completes a full knowledge accumulation loop from problem discovery and resolution to experience consolidation. If the evaluation result is a failure, the knowledge storage operation is not performed, and the system can be configured to revert to the previous stable strategy paradigm or record the failure information for offline analysis.
[0059] When the evaluation result is a failure, the system will execute a pre-defined fault handling procedure. This procedure first rolls back the scheduling policy to the policy paradigm used before this evolution. Simultaneously, the system will generate a detailed failure log and store it in a dedicated failure case library. This failure log will contain at least the following fields: a timestamp, causal instability information that triggered this evolution, the policy paradigm judged as failed (i.e., the reward function), and the system causal entropy within the evaluation time window. The data includes the sequence of values for all core KPIs within the evaluation window. These records are available for system administrators to perform offline diagnostics and model iteration analysis.
[0060] The supply chain adaptive scheduling method integrating deep reinforcement learning provided in this embodiment evaluates the effectiveness of the new strategy paradigm by acquiring actual data from the physical execution system after deploying effective scheduling instructions. The effective paradigm and corresponding instability information are stored in the paradigm archive to form a knowledge loop. The practicality of the strategy is verified through actual data, ensuring that the scheduling instructions meet the actual operational needs of the supply chain. This provides rich historical experience for subsequent risk prediction and strategy generation, further enhancing the adaptive capability of supply chain scheduling.
[0061] This embodiment provides an adaptive scheduling method for the supply chain that integrates deep reinforcement learning, which can be used in the aforementioned computer system. Figure 4 This is a flowchart of an industry chain adaptive scheduling method incorporating deep reinforcement learning according to an embodiment of the present invention, such as... Figure 4 As shown, the process includes the following steps: Step S301: Real-time acquisition of multiple variables within the target industry chain, construction and real-time updating of a dynamic causal graph. The dynamic causal graph is used to characterize the causal relationship between variables, and the edges between variables are attached with causal entropy, which characterizes the causal strength.
[0062] Specifically, step S301 includes: Step S3011: Preprocess the data of each variable to obtain standard variable data, and use the dynamic causal discovery algorithm based on the time series data of each standard variable data to determine the causal relationship between each variable based on a preset time window.
[0063] Specifically, a set of multivariate time series data with predefined key variables is collected and received from a supply chain simulation environment or actual physical system. These variables include, but are not limited to: raw material inventory levels of suppliers at all levels, production line output rates of manufacturers, finished goods inventory levels of distribution centers, order backlogs of retailers, and transportation lead times in the logistics process. The collected raw time series data is first preprocessed. This preprocessing includes: normalizing the variable data of different dimensions using the z-score standardization method to make them conform to a standard normal distribution; and stabilizing the non-stationary time series through first-order differencing to meet the data stationarity requirements of subsequent causal discovery algorithms.
[0064] Step S3012: Obtain time series data of two variables with a causal relationship, and calculate the transfer entropy between the two variables based on the time series data.
[0065] Specifically, after preprocessing, the supply chain dynamic causal graph construction module 100 uses a constraint-based dynamic causal discovery algorithm (such as the PC-stable algorithm) to learn the causal structure between variables within a sliding time window. A time window width W and a step size S are set, and the module performs causal discovery once on the data within the time interval [t-W+1,t], generating time... The causal graph is obtained, and then this process is repeated in the next time window [t-W+S+1,t+S]. Within each time window, the specific execution steps of the PC-stable algorithm include: 1) Construct a fully connected undirected graph, where the nodes represent key variables in the industry chain; 2) For each pair of adjacent nodes in the graph ( , ), and so on, using combinations of other nodes. Given the condition, perform a conditional independence test. If it is found that... Under the conditions and If the conditions are independent, remove the node. and The edge between; 3) Repeat step 2 until there are no more borders to remove, thus obtaining the skeleton of the graph; 4) Using the v-structure and FCI direction rules, the direction of the edges in the skeleton graph is determined, and finally a directed acyclic graph is formed, namely the causal graph at time t.
[0066] In one embodiment, the conditional independence test uses the G-test (log-likelihood ratio test). For the variable , Given the condition set Z, the test statistic is calculated as follows:
[0067] in, Variables observed in a data sample Values , Values , condition set Values The frequency. By calculating the The value is compared with the critical value of the chi-square distribution at a preset significance level (e.g., To determine whether conditional independence holds.
[0068] After determining the topological structure of the causal graph, the industrial chain dynamic causal graph construction module 100 needs to quantify the causal strength of each directed edge in the graph.
[0069] In one embodiment, transfer entropy (TE) is used to measure the strength of a causal influence from one time series X to another time series Y. For two variables... and ,from arrive Transitive entropy The calculation formula is:
[0070] in, Representing variables The state at the next moment; Representing variables At any moment The historical state, The order of the historical state; Representing variables At any moment The historical state, The order of the historical state; Represents a probability distribution. This represents a conditional probability distribution.
[0071] Calculated transfer entropy value That is, it is assigned the value of a dynamic causal graph. In, from the node Pointing to node The edge ( , Quantitative causal strength Through the above steps, the dynamic causal graph construction module 100 of the industrial chain finally outputs a complete, weighted, time-varying directed acyclic graph, providing a data foundation for subsequent system causal entropy calculation and strategy evolution triggering.
[0072] Step S3013: Normalize the transfer entropy between variables with causal relationships to obtain a probability distribution, and calculate the Shannon entropy of the probability distribution as the causal entropy of the dynamic causal graph.
[0073] Specifically, the causal entropy is calculated based on the constructed dynamic causal graph and the transfer entropy. The specific process includes: (1) normalizing the causal intensity of all edges in the causal graph to obtain the probability distribution. :
[0074] in, Indicates at time Causal relationship edge ( , The normalization intensity throughout the entire dynamic causal graph; Indicates at time From variables Pointer variable The quantification of causal strength of the edges; For a moment The edge set of a dynamic causal graph; , ) is an edge set Either side of it.
[0075] (2) Calculate the Shannon entropy of the probability distribution, which is used as the causal entropy. The formula for its calculation is:
[0076] in, For a moment The system's causal entropy; For a moment The edge set of the dynamic causal graph; edge set Middle, side ( , The normalized strength of ).
[0077] The supply chain adaptive scheduling method fused with deep reinforcement learning provided in this embodiment achieves direct monitoring of the internal logical stability of the supply chain system by constructing a dynamic causal graph and calculating the system's causal entropy. Compared with methods that rely on lagging performance indicators such as cost and delay, it can identify potential instability risks before the system's performance indicators deteriorate significantly based on changes in the strength of causal relationships, thus providing an advance time window for strategy adjustment.
[0078] Step S302: Predict potential risk areas based on causal entropy, and determine whether to trigger strategy evolution based on at least one of the causal entropy and the prediction result.
[0079] Specifically, step S302 includes: Step S3021: Calculate the causal entropy change rate based on the causal entropy at each time point, and compare the causal entropy with the causal entropy threshold and the causal entropy change rate with the change rate threshold. If the causal entropy is greater than the causal entropy threshold or the causal entropy change rate is greater than the change rate threshold, then the passive response triggers the strategy evolution.
[0080] Specifically, a causal entropy instability criterion is used to determine whether to trigger policy evolution. This criterion comprises two parallel conditions, and triggering the evolution is achieved when either condition is met. The first condition is an absolute threshold criterion: causal entropy. Among them, the causal entropy threshold (i.e., the absolute threshold) This is a pre-defined constant whose value is determined by statistical analysis of the system's causal entropy values during long-term stable operation of the industrial chain system. For example, it can be set as the historical stable entropy mean plus three standard deviations. This criterion is used to identify when the overall causal structure of the system is in a persistently unstable state.
[0081] The second condition is the rate of change threshold criterion: Among them, the rate of change It can be approximated by calculating the difference between the current causal entropy value and the causal entropy value at the previous moment, while the rate of change threshold... This is also a pre-defined constant. This criterion is used to capture drastic deterioration of the system's causal structure within a short period of time, even if the absolute entropy value has not yet reached its maximum. This condition is particularly crucial for identifying situations where system stability rapidly declines due to sudden events.
[0082] Once the causal entropy instability criterion is met, triggering a passive-response policy evolution, a causal instability region location map will be immediately generated. This causal instability region location map is a subgraph. ,in, yes a subset of nodes yes A subset of edges. A node. Included In, if and only if the node is connected to at least one path in An edge in the middle. Included In, if and only if it satisfies at least one of the following conditions:
[0083] in, This is the preset minimum causal strength threshold.
[0084]
[0085] in, This is the preset threshold for the rate of change of causal intensity.
[0086] This diagram is the current dynamic causal graph. A subgraph whose generation rule is: traversal All edges in the localization graph that satisfy any of the following conditions and the nodes they connect to are included in the localization graph: 1) Quantitative Causality Strength of Edges 1) The intensity of the edge is below a preset minimum strength threshold; 2) The rate of change of the causal strength of the edge in the most recent evaluation period exceeds a preset maximum rate of change threshold.
[0087] This generated causal instability region location map will be used as causal instability information representing the current system instability mode and will be output to the strategy paradigm management module 300 for subsequent paradigm retrieval or generation.
[0088] Step S3022: Obtain historical evolutionary trajectory data from the paradigm archive, and learn based on the historical evolutionary trajectory data to obtain meta-learning results. Based on the meta-learning results, predict the latest evolutionary trajectory data to obtain prediction results. If the prediction results meet the preset conditions, actively predictively trigger strategy evolution.
[0089] In some optional implementations, the historical evolutionary trajectory data includes: historical instability information and corresponding strategy paradigms, and step S3022 above includes: Step a1: Learning is performed based on historical evolutionary trajectory data to obtain a predictive model representing the relationship between unstable information and policy paradigm as the meta-learning result.
[0090] Specifically, the paradigm archive stores historical evolutionary trajectories. The data structure of the paradigm archive is a collection of key-value pairs, where the "key" is historical causal instability information (a graph structure), and the "value" is the corresponding, validated, successful strategy paradigm. The historical evolutionary trajectory is a sequence of these key-value pairs arranged chronologically, i.e.:
[0091] In one specific embodiment, a Transformer-based sequence model is used for meta-learning. During the training phase, the data in the historical evolutionary trajectory first needs to be vectorized. For each piece of historical causal instability information (a graph), a fixed-dimensional graph neural network encoder is used to transform it into a fixed-dimensional vector. For each corresponding policy paradigm, an encoder is used to convert it into a vector. Thus, the historical evolutionary trajectory is transformed into a sequence of vector pairs: ; The training task for the Transformer model is set as follows: input a sequence of length... Historical vector pairs sequence Predict the vector corresponding to the next causal instability information that will occur. By training on a large number of historical trajectories, the model's internal self-attention mechanism learns the patterns of evolution over time among different types of causal instability patterns, serving as a meta-learning outcome.
[0092] Step a2: Input the latest evolutionary trajectory data into the trained prediction model to obtain the prediction vector, and determine the potential risk area location map based on the prediction vector. The potential risk area location map is used to characterize the probability of future instability.
[0093] Specifically, in actual operation, such as Figure 5 The diagram illustrates the process of proactive predictive evolutionary triggering based on meta-learning. The latest historical evolutionary trajectory sequence is input into a trained Transformer model. The model outputs a prediction vector, which represents the prediction for a future time step. ( This is a prediction of possible causal instability patterns (for the predicted horizon). The prediction vector is then fed into a decoder corresponding to the graph encoder, which reconstructs it into a specific graph structure—the potential risk region localization map. Each node and edge in this localization map is assigned a value between 0 and 1, representing the probability that the model predicts that node or edge will become unstable in the future.
[0094] Step a3: Calculate the aggregated risk value based on the potential risk area location map. When the aggregated risk value is greater than the risk trigger threshold, the proactive predictive strategy evolution is triggered.
[0095] Specifically, an aggregated risk value is calculated based on the potential risk area location map. The aggregated risk value can be calculated by taking the maximum value of the instability probability of all edges in the graph, or by calculating the average probability of all edges whose probabilities exceed a specific low-confidence filtering threshold. Then, this aggregated risk value is compared with a preset risk trigger threshold. A comparison is made, and the triggering condition is met. At that time, the system will trigger proactive, predictive strategy evolution. The potential risk area location map generated by the model, which contains future risk prediction information, will be output to the strategy paradigm management module 300 as causal instability information characterizing the future system instability mode.
[0096] The supply chain adaptive scheduling method integrating deep reinforcement learning provided in this embodiment introduces a meta-learning mechanism for historical evolutionary trajectories in the paradigm archive, enabling the scheduling system to predict potential future risks. It is not limited to passively responding to current system instability, but can proactively identify areas where problems may occur in the future based on the learned evolutionary patterns. This transforms the adjustment of scheduling strategies from an event-driven passive mode to a forward-looking proactive prediction mode.
[0097] In step S303, if policy evolution is triggered, the instability information is determined, and a new policy paradigm is generated and activated based on the instability information.
[0098] Specifically, step S303 includes: Step S3031: If the evolution of the strategy is triggered by a passive response, a causal unstable region location map is generated as unstable information representing the current unstable mode.
[0099] Specifically, if the policy evolution is triggered by a passive response, the current unstable region location map triggered by the passive response will be used as the instability information (i.e., causal instability information) representing the current unstable mode.
[0100] In step S3032, if the strategy evolution is triggered by proactive prediction, a potential risk area location map is predicted and used as instability information to characterize future instability modes.
[0101] Specifically, if the strategy evolution is triggered by proactive prediction, the potential risk area location map triggered by proactive prediction will be used as instability information to characterize future instability patterns.
[0102] Step S3033: Compare the structural similarity between the unstable information and the historical unstable information in the paradigm archive. If the structural similarity exceeds the similarity threshold, the corresponding historical unstable information is used as the matching item for the unstable information.
[0103] Specifically, such as Figure 6 The diagram shows the process of strategy paradigm retrieval and generation. First, paradigm retrieval is performed based on instability information. Whether it is a current unstable area location map triggered by passive response or a potential risk area location map triggered by active prediction, it is uniformly used as the input query for retrieval and generation.
[0104] During the paradigm retrieval phase, the strategy paradigm management module 300 uses the input causal instability information graph as the query graph and calculates graph structure similarity among all historical causal instability information graphs stored in the paradigm archive. The graph structure similarity calculation employs the Graph Edit Distance (GED) algorithm, which calculates the cost of the minimum number of editing operations (insertion, deletion, and replacement of nodes / edges) required to transform the query graph into a historical graph. A lower GED indicates a higher structural similarity between the two graphs. The module iterates through all historical entries in the archive, calculates the GED value between the query graph and each historical graph, and selects the historical entry with the smallest GED value as the best candidate.
[0105] Calculate the structural similarity based on the obtained minimum GED value, and then compare the structural similarity with a preset similarity threshold. Compare, if the structural similarity is greater than If a valid match is found, the historical strategy paradigm corresponding to the best candidate entry is extracted directly as the new strategy paradigm to be activated in this evolution, and then output to the scheduling and execution module. If the structural similarity is not less than... If no similar historical experience is available in the system, the retrieval phase ends and the paradigm generation phase begins immediately.
[0106] Step S3034: Based on the matching terms of the instability information, determine the policy paradigm corresponding to the matching terms as the new policy paradigm.
[0107] Specifically, based on the matching terms of instability information, an effective solution is matched for the current or predicted system instability mode, that is, a new strategy paradigm (the strategy paradigm corresponding to the matching term).
[0108] In some optional implementations, if the structural similarity does not exceed a similarity threshold, a new strategy paradigm generation process is initiated, and step S303 further includes: Step S3035: Construct a policy primitive library and encode each policy primitive in the policy primitive library to obtain a primitive vector.
[0109] Specifically, such as Figure 6 As shown, in the paradigm generation stage, a policy primitive library is first constructed, containing multiple policy primitives. This library includes a set of structured, basic reward function terms, with each policy primitive corresponding to a local optimization objective within the industry chain. For example, the library may contain the following primitives (this is just an example and not a limitation).
[0110] Inventory cost primitives: ,in, For nodes Inventory levels, This is the unit inventory cost coefficient.
[0111] Output rate primitives: ,in, To manufacture nodes The output.
[0112] Delivery cycle primitives: ,in, This represents the average order delivery time from end to end.
[0113] Order fulfillment rate primitives: ,in, To improve the on-time fulfillment rate of orders.
[0114] Step S3036: Input the global state vector, instability information vector, and random noise vector of the target industry chain into the generator of the generative adversarial network. Use the instability information vector as the query and the primitive vector as the key and value. Use the attention mechanism in the generator to generate the weights of each primitive vector.
[0115] Specifically, the policy paradigm management module 300 activates a pre-trained generative model, which in one embodiment is a Generative Adversarial Network (GAN). The GAN's generator receives three inputs: 1) A vector representing the current global state of the industrial chain system; 2) A graph vector representing the causal instability information of the current input; 3) A random noise vector. The generator's goal is to output a new reward function, which is the core of the new policy paradigm.
[0116] To make the generation process more targeted, an attention mechanism is integrated into the generator. This attention mechanism uses the graph vector of causal instability information as the query and the vector representations of all primitives in a pre-defined policy primitive library as the key and value. The attention mechanism in the generator calculates the weight distribution of each policy primitive. The weight distribution calculated by the attention mechanism reflects the degree of relevance of different policy primitives to solving the problem indicated by the current causal instability information.
[0117] Step S3037: The primitive vectors are weighted based on their respective weights to obtain a composite reward function, which is then encapsulated into a new policy paradigm.
[0118] Specifically, the generator uses an attention mechanism to weight and combine these primitives to form a composite reward function, for example... The weight The strategy primitives determined by the attention mechanism are predefined, modular basic rewards, such as minimizing inventory holding costs at a specific node, maximizing the output of a production line, or shortening the delivery cycle from a specific supplier to the manufacturer.
[0119] Based on these weights, the generator selectively combines and weights policy primitives to construct a composite, novel reward function. The GAN's discriminator then learns from the reward functions of all historically successful policy paradigms stored in the paradigm archive to determine the rationality and effectiveness of the generator's reward function. Its output loss signal is used to iteratively optimize the generator's parameters. Finally, the generator's output reward function is encapsulated into a new policy paradigm and sent to the scheduling and execution module.
[0120] The industry chain adaptive scheduling method integrating deep reinforcement learning provided in this embodiment achieves dynamic adjustment of the scheduling optimization objective by retrieving or generating a policy paradigm with a new reward function as its core. When the system faces fundamental logical changes, this method is not limited to optimizing policy parameters under the original reward function, but can generate a completely new reward function to redefine the optimization problem itself. This enables the scheduling system to adapt to unexpected system operation modes that are difficult to handle by traditional fixed-objective methods, thus expanding the scope of its adaptive scheduling application.
[0121] Step S304 involves interacting the new strategy paradigm with the simulation environment of the target industry chain to determine effective scheduling instructions, and then deploying these instructions to the physical execution system of the target industry chain. For details, please refer to [link to relevant documentation]. Figure 2 Step S204 of the illustrated embodiment will not be described again here.
[0122] To verify the applicability of the supply chain adaptive scheduling method integrating deep reinforcement learning provided in this embodiment in actual industrial scenarios, the following uses the new energy vehicle power battery supply chain as an example to illustrate the specific operation process when dealing with the risk of raw material supply interruption.
[0123] Scenario setting: Construct a simplified industry chain model, including upstream raw material suppliers (lithium mining company A, lithium mining company B), midstream cathode material manufacturers (factory C), and downstream cell manufacturers (factory D).
[0124] Key variable: The system primarily collects data on the shipment volume of lithium mining company A. Transport time for sea route H Factory C's raw material inventory and the order delivery rate of factory D .
[0125] Initial state: The system is running in normal mode. The current strategy paradigm focuses on reducing inventory holding costs and tends to maintain a low inventory level.
[0126] like Figure 7 The diagram shown illustrates the complete process of the supply chain adaptive scheduling method integrating deep reinforcement learning provided in this embodiment. The execution process specifically includes: Step S1, anomaly detection and graph update, the industrial chain dynamic causal graph construction module 100 at time... Data anomaly detected: transit time of shipping route H Significant delays occurred, and the correlation coefficient between the shipment volume of lithium mining company A and the inventory intake of factory C decreased.
[0127] Updating the dynamic causal graph using the PC-stable algorithm The results show that there is a directed edge from node lithium mining company A to node factory C. causal strength The drop from a high value of 0.85 to 0.20 indicates a blockage in the upstream supply chain.
[0128] Step S2, passive response triggered, calculate the current system causal entropy. The weakening of causal connections in the critical supply path caused the overall system entropy to rise and exceed a preset absolute threshold. .
[0129] The system determines that the current causal structure is unstable, triggers the policy evolution process, and outputs a causal instability region location map containing node A, node C and their connecting edges.
[0130] Step S3: Strategy Paradigm Generation. The strategy paradigm management module 300 uses the aforementioned positioning map as a query condition. If no matching historical strategy is found in the paradigm archive, the generative model is initiated.
[0131] Based on the current state of low inventory and supply disruption, the generator uses an attention mechanism to adjust the weight distribution of policy primitives: reducing the weight of the inventory cost primitive (allowing inventory levels to rise), and increasing the weight of the order fulfillment rate primitive and the stockout penalty primitive.
[0132] The generated policy paradigm includes a new reward function. Its guiding principle has shifted from cost minimization to prioritizing delivery assurance.
[0133] Step S4: Agent scheduling training, the scheduling execution module loads the new reward function. The deep reinforcement learning scheduling agent analyzes the state through graph neural networks and identifies a stable causal connection between backup supplier lithium mining company B and factory C.
[0134] based on To maximize returns, the agent outputs a set of scheduling instructions through the PPO algorithm: suspend new orders to lithium mining company A; initiate the procurement process to lithium mining company B, accepting a higher unit price in exchange for supply stability; and adjust the production schedule of factory C, prioritizing the production of high-priority order batches.
[0135] Step S5, Deployment and Evaluation: The scheduling instructions are transmitted to the execution system. During the evaluation period following instruction execution, monitoring data shows that the continuity of raw material warehousing at factory C has been restored, and the system's causal entropy... Once the price returns to a stable threshold range and the key performance indicator (order delivery rate) meets the preset standards, the system determines that this strategy evolution is effective.
[0136] In step S6, the knowledge accumulation module 500 stores the supply chain roadblock cause-effect diagram and the strategy paradigms for switching suppliers and ensuring delivery in the paradigm archive.
[0137] This data entry will provide training samples for the proactive predictive triggering strategy paradigm in subsequent operations, enabling it to output corresponding risk area location maps and match response strategies in advance when similar supply network risks are predicted in the future.
[0138] This embodiment also provides an adaptive scheduling system for the supply chain that integrates deep reinforcement learning. This system is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.
[0139] This embodiment provides an adaptive scheduling system for the industrial chain that integrates deep reinforcement learning, such as... Figure 8 As shown, it includes: The causal graph construction and update module 100 is used to acquire multiple variables within the target industry chain in real time, construct and update a dynamic causal graph in real time. The dynamic causal graph is used to characterize the causal relationship between variables, and the edges between variables are attached with causal entropy, which characterizes the causal strength.
[0140] The strategy evolution triggering module 200 is used to predict potential risk areas based on causal entropy, and to determine whether to trigger strategy evolution based on at least one of causal entropy and prediction results.
[0141] The strategy paradigm management module 300 is used to determine instability information if strategy evolution is triggered, and to generate and activate a new strategy paradigm based on the instability information.
[0142] The effective scheduling instruction deployment module 400 is used to interact with the simulation environment of the target industry chain to determine the effective scheduling instructions and deploy the effective scheduling instructions to the physical execution system of the target industry chain.
[0143] In some alternative implementations, the supply chain adaptive scheduling system incorporating deep reinforcement learning also includes: The knowledge accumulation module 500 is used to acquire the actual execution data of the physical execution system and evaluate the effectiveness of the new strategy paradigm based on the actual execution data; if the new strategy paradigm is effective, the new strategy paradigm and the corresponding instability information are stored in the paradigm archive.
[0144] In some optional implementations, the causal map construction and update module 100 includes: The causal relationship determination unit is used to preprocess the data of each variable to obtain standard variable data, and based on the time series data of each standard variable data, use a dynamic causal discovery algorithm to determine the causal relationship between each variable based on a preset time window.
[0145] The transfer entropy calculation unit is used to acquire time-series data of two variables with a causal relationship, and calculate the transfer entropy between the two variables based on the time-series data.
[0146] The causal entropy calculation unit is used to normalize the transfer entropy between variables with causal relationships, obtain a probability distribution, and calculate the Shannon entropy of the probability distribution as the causal entropy of the dynamic causal graph.
[0147] In some optional implementations, the strategy evolution triggering module 200 includes: The passive response unit is used to calculate the causal entropy change rate based on the causal entropy at each time point, and compare the causal entropy with the causal entropy threshold and the causal entropy change rate with the change rate threshold. If the causal entropy is greater than the causal entropy threshold, or the causal entropy change rate is greater than the change rate threshold, the passive response triggers the strategy evolution.
[0148] The proactive prediction unit is used to acquire historical evolutionary trajectory data in the paradigm archive, learn based on the historical evolutionary trajectory data to obtain meta-learning results, predict the latest evolutionary trajectory data based on the meta-learning results, and obtain prediction results. If the prediction results meet the preset conditions, the proactive prediction triggers the strategy evolution.
[0149] In some optional implementations, the strategy paradigm management module 300 includes: The first instability information determination unit is used to generate a causal instability region location map if the strategy evolution is triggered by a passive response, as instability information representing the current instability mode.
[0150] The second instability information determination unit is used to predict the location map of potential risk areas if the strategy evolution is triggered by active prediction, and to serve as instability information representing future instability modes.
[0151] The instability information matching unit is used to compare the structural similarity between instability information and historical instability information in the paradigm archive. If the structural similarity exceeds the similarity threshold, the corresponding historical instability information is used as the matching item for instability information.
[0152] The first new strategy paradigm determination unit is used to determine the strategy paradigm corresponding to the matching item based on the instability information as the new strategy paradigm.
[0153] In some optional implementations, if the structural similarity does not exceed a similarity threshold, a new strategy paradigm generation process is initiated. The strategy paradigm management module 300 further includes: The policy primitive encoding unit is used to construct a policy primitive library and encode each policy primitive in the library to obtain a primitive vector.
[0154] The weight determination unit is used to input the global state vector, instability information vector, and random noise vector of the target industry chain into the generator of the generative adversarial network. The instability information vector is used as the query, and the primitive vectors are used as the key and value. The attention mechanism in the generator is used to generate the weights of each primitive vector.
[0155] The second new policy paradigm determination unit is used to perform weighted calculations on the primitive vectors based on the weights of each primitive vector to obtain a composite reward function, and then encapsulate the composite reward function into a new policy paradigm.
[0156] The supply chain adaptive scheduling system integrating deep reinforcement learning provided in this embodiment of the invention can execute the supply chain adaptive scheduling method integrating deep reinforcement learning provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments above, and will not be repeated here.
[0157] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0158] The following is a detailed reference. Figure 9 This diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 901, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from memory 908 into random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device. The processor 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0159] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic devices to exchange data via wireless or wired communication with other devices. Although Figure 9 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0160] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a memory 908, or installed from a ROM 902. When the computer program is executed by the processor 901, it performs the functions defined in the supply chain adaptive scheduling method incorporating deep reinforcement learning according to embodiments of the present invention.
[0161] Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0162] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the supply chain adaptive scheduling method incorporating deep reinforcement learning shown in the above embodiments is implemented.
[0163] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. An adaptive scheduling method for the industrial chain integrating deep reinforcement learning, characterized in that, The method includes: Multiple variables within the target industry chain are acquired in real time, and a dynamic causal graph is constructed and updated in real time. The dynamic causal graph is used to characterize the causal relationship between variables, and the edges between variables are attached with causal entropy that characterizes the causal strength. Based on the causal entropy, potential risk areas are predicted, and based on at least one of the causal entropy and the prediction result, it is determined whether strategy evolution is triggered. If policy evolution is triggered, instability information is determined, and a new policy paradigm is generated and activated based on the instability information. The new strategy paradigm is interacted with the simulation environment of the target industry chain to determine effective scheduling instructions, and the effective scheduling instructions are deployed to the physical execution system of the target industry chain.
2. The method according to claim 1, characterized in that, After deploying the effective scheduling instructions to the physical execution system of the target industry chain, the method further includes: Obtain the actual execution data of the physical execution system, and evaluate the effectiveness of the new strategy paradigm based on the actual execution data; If the new strategy paradigm is effective, then the new strategy paradigm and the corresponding instability information are stored in the paradigm archive.
3. The method according to claim 1, characterized in that, Real-time acquisition of multiple variables within the target industry chain, construction and real-time updating of a dynamic causal graph, including: The data of each variable is preprocessed to obtain the standard variable data. Based on the time series data of each standard variable data, the dynamic causal discovery algorithm is used to determine the causal relationship between each variable based on a preset time window. Obtain time-series data of two variables that have a causal relationship, and calculate the transfer entropy between the two variables based on the time-series data; The transfer entropy between variables with causal relationships is normalized to obtain a probability distribution, and the Shannon entropy of the probability distribution is calculated as the causal entropy of the dynamic causal graph.
4. The method according to claim 1, characterized in that, Based on the causal entropy, potential risk areas are predicted, and based on at least one of the causal entropy and the prediction result, it is determined whether strategy evolution is triggered, including: The causal entropy change rate is calculated based on the causal entropy at each time point, and the causal entropy is compared with the causal entropy threshold and the causal entropy change rate is compared with the change rate threshold. If the causal entropy is greater than the causal entropy threshold or the causal entropy change rate is greater than the change rate threshold, then a passive response triggers the strategy evolution. The system acquires historical evolutionary trajectory data from the paradigm archive and learns from the historical evolutionary trajectory data to obtain meta-learning results. Based on the meta-learning results, it predicts the latest evolutionary trajectory data to obtain prediction results. If the prediction results meet preset conditions, it actively and predictively triggers strategy evolution.
5. The method according to claim 4, characterized in that, The historical evolutionary trajectory data includes: historical instability information and corresponding policy paradigms. Learning is performed based on the historical evolutionary trajectory data to obtain meta-learning results. The latest evolutionary trajectory data is then predicted based on the meta-learning results to obtain prediction results. If the prediction results meet preset conditions, policy evolution is triggered, including: Learning is performed based on the historical evolutionary trajectory data to obtain a predictive model that represents the relationship between unstable information and policy paradigms, which serves as the meta-learning result. The latest evolutionary trajectory data is input into the trained prediction model to obtain a prediction vector, and a potential risk area location map is determined based on the prediction vector. The potential risk area location map is used to characterize the probability of future instability. The aggregated risk value is calculated based on the potential risk area location map. When the aggregated risk value is greater than the risk trigger threshold, the proactive predictive strategy evolution is triggered.
6. The method according to claim 4, characterized in that, If policy evolution is triggered, instability information is determined, and a new policy paradigm is generated and activated based on the instability information, including: If the evolution of the strategy is triggered by a passive response, a causal unstable region localization map is generated as an unstable information characterizing the current unstable mode. If the strategy evolution is triggered by proactive prediction, a potential risk area location map is predicted, which serves as instability information representing future instability patterns. The structural similarity between the instability information and the historical instability information in the paradigm archive is compared. If the structural similarity exceeds the similarity threshold, the corresponding historical instability information is used as the matching item for the instability information. Based on the matching terms of the instability information, the policy paradigm corresponding to the matching terms is determined as the new policy paradigm.
7. The method according to claim 6, characterized in that, If the structural similarity does not exceed the similarity threshold, a new strategy paradigm generation process is initiated, including: Construct a policy primitive library and encode each policy primitive in the policy primitive library to obtain a primitive vector; The global state vector, instability information vector, and random noise vector of the target industry chain are input into the generator of the generative adversarial network. The instability information vector is used as the query and the primitive vectors are used as the key and value. The attention mechanism in the generator is used to generate the weights of each primitive vector. The primitive vectors are weighted based on their respective weights to obtain a composite reward function, which is then encapsulated into a new policy paradigm.
8. An adaptive scheduling system for the industrial chain integrating deep reinforcement learning, characterized in that, The system includes: The causal graph construction and update module is used to acquire multiple variables within the target industry chain in real time, construct and update a dynamic causal graph in real time. The dynamic causal graph is used to characterize the causal relationship between variables, and the edges between variables are attached with causal entropy that characterizes the causal strength. The strategy evolution triggering module is used to predict potential risk areas based on the causal entropy, and determine whether to trigger strategy evolution based on at least one of the causal entropy and the prediction result. The strategy paradigm management module is used to determine instability information if strategy evolution is triggered, and to generate and activate a new strategy paradigm based on the instability information. The effective scheduling instruction deployment module is used to interact with the simulation environment of the target industry chain to determine the effective scheduling instructions and deploy the effective scheduling instructions to the physical execution system of the target industry chain.
9. An electronic device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method of any one of claims 1 to 7.