An automated multi-agent whole-process operation scheduling platform and method

By constructing an automated multi-agent full-process operation and scheduling platform, and utilizing operational resilience knowledge graphs and generative adversarial networks for operational early warning, the problem of insufficient foresight of minor risks in traditional operation and maintenance models has been solved, achieving efficient, accurate, and economical resource scheduling.

CN121212751BActive Publication Date: 2026-02-06LIAONING ZHONGLIAN AUTOMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511769696.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-06
Estimated Expiration
2045-11-28

AI Technical Summary

Technical Problem

Traditional operation and maintenance models lack the ability to anticipate subtle early risks when dealing with complex industrial operating environments. Alarm information is difficult to automatically locate the root cause, and resource scheduling schemes are detached from actual constraints and cost-effectiveness, resulting in insufficient efficiency and accuracy of resource scheduling.

Method used

An automated multi-agent end-to-end operation scheduling platform is constructed. By acquiring the full-process operation dataset, an operation resilience knowledge graph is built. An operation early warning model is constructed by combining generative adversarial networks and graph neural networks. Operation early warning and game negotiation are carried out to generate operation control strategies and to perform operation incentive points settlement to optimize resource scheduling.

Benefits of technology

It achieves efficient and precise scheduling of operational resources, and ensures the economy and feasibility of operational control strategies through early warning of weak signals and multi-resource collaborative optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121212751B_ABST
    Figure CN121212751B_ABST
Patent Text Reader

Abstract

The application provides an automatic multi-agent whole-process operation scheduling platform and method, comprising: acquiring a whole-process operation dataset, constructing and real-time updating a whole-process operation resilience knowledge graph based on the whole-process operation dataset, constructing an operation early warning model based on a generative adversarial network and a graph neural network, and pre-training the model combined with historical whole-process operation datasets; inputting real-time whole-process operation datasets into the operation early warning model for operation early warning, generating an operation early warning dataset, conducting game negotiation based on the operation early warning dataset, obtaining game negotiation results, generating operation control strategies according to the game negotiation results, clearing operation incentive points according to the implementation effect of the operation control strategies, generating an optimization suggestion set and feeding back and optimizing the operation early warning model, and realizing efficient and accurate scheduling of operation resources.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of resource scheduling, in particular to an automatic multi-agent whole-process operation scheduling platform and method. BACKGROUND

[0002] In the current complex industrial operation environment, enterprises have generally established multi-dimensional operation and maintenance systems covering equipment, supply chain, manpower, etc., and have accumulated a large amount of real-time and historical data. The traditional operation and maintenance mode mostly adopts an alarm system based on fixed thresholds and a response process relying on human experience. This mode has certain effect in dealing with known single-point faults and provides a foundation support for ensuring the stable operation of the system.

[0003] However, with the sharp increase in system complexity and dynamics, the traditional operation and maintenance mode is facing severe challenges. On the one hand, it often lacks the ability to foresee weak early risks, and alarm information is difficult to automatically locate the root cause, resulting in ambiguous response targets. On the other hand, the resource scheduling scheme has certain problems of being divorced from actual constraints and cost-effectiveness, thereby leading to insufficient efficiency and accuracy of resource scheduling.

[0004] These limitations make the traditional technology inefficient in dealing with modern operation and maintenance scheduling, so there is an urgent need for an automatic multi-agent whole-process operation scheduling method to achieve efficient and accurate scheduling of operation resources. SUMMARY

[0005] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide an automatic multi-agent whole-process operation scheduling method, which comprises:

[0006] An operation data set is obtained, which is divided into a historical whole-process operation data set and a real-time whole-process operation data set. An operation resilience knowledge graph is constructed and updated in real time based on the operation data set;

[0007] An operation early warning model is constructed based on a generative adversarial network and a graph neural network, and the model is pre-trained in combination with the historical whole-process operation data set. The operation early warning model comprises a risk detection sub-module and a simulation early warning sub-module;

[0008] The real-time whole-process operation data set is input into the operation early warning model for operation early warning, and an operation early warning data set is generated, which comprises an operation early warning information set and operation incentive points;

[0009] Game negotiation is performed based on the operation early warning data set, a game negotiation result is obtained, and an operation control strategy is generated according to the game negotiation result;

[0010] According to the implementation effect of the operation regulation strategy, operation incentive points are cleared, an optimized suggestion set is generated, and an operation early warning model is fed back and optimized.

[0011] As a further scheme of the present application, an operation early warning model is constructed based on a generative adversarial network and a graph neural network, and is pre-trained in combination with a historical full-process operation data set, the operation early warning model comprising a risk detection sub-module and a simulation early warning sub-module, and comprising:

[0012] A digital twin environment is constructed based on a full-process operation resilience knowledge graph and a historical full-process operation data set, and is updated in real time in combination with a real-time full-process operation data set, the digital twin environment being used to simulate a full-process operation state;

[0013] The historical full-process operation data set is clustered to obtain a non-stationary operation data set, the non-stationary operation data set representing historical data in which an operation process imbalance occurs in a full-process operation process, and at least comprising operation imbalance root cause data and operation imbalance impact data;

[0014] A reinforcement learning agent is constructed, a state space corresponding to the reinforcement learning agent is defined according to the digital twin environment, an action space of the reinforcement learning agent is set as controllable noise for early risk detection, a deep neural network is used as a policy network of the reinforcement learning agent, and the policy network is pre-trained in combination with the non-stationary operation data set, and the pre-trained reinforcement learning agent is used as the risk detection sub-module;

[0015] The risk detection sub-module is used to generate controllable noise, and the controllable noise is used to be input into the simulation early warning sub-module for early operation risk identification.

[0016] As a further scheme of the present application, the method further comprises:

[0017] The controllable noise is input into the digital twin environment for simulation, to generate a simulation early warning data set, the simulation early warning data set comprising an initial state vector, controllable noise, a perturbed state vector, and an unperturbed state vector;

[0018] The initial state vector represents an initial state vector of the digital twin environment, the perturbed state vector represents a state vector sequence of the digital twin environment within a fixed time window after the controllable noise is injected, and the unperturbed state vector represents a state vector sequence of the digital twin environment within a fixed time window corresponding to no injection of the controllable noise;

[0019] The simulation early warning submodule is constructed based on a deep twin network, the simulation early warning data set is input into the simulation early warning submodule, the difference degree of the state vector after disturbance and the state vector without disturbance is obtained according to an attention mechanism network, pre-training is performed in combination with a predetermined training target, training efficiency verification is performed in combination with a training verification set constructed based on a non-stationary operation data set, and the training is ended only when a predetermined efficiency condition is reached.

[0020] As a further scheme of the present application, real-time full-process operation data set is input into the operation early warning model for operation early warning, and an operation early warning data set is generated, which includes an operation early warning information set and operation incentive points, and includes:

[0021] After the real-time full-process operation data set is input into the operation early warning model, the risk detection submodule generates corresponding controllable noise based on the real-time full-process operation data set;

[0022] The controllable noise is injected into the pre-constructed digital twin environment for simulation and generation of corresponding simulation early warning data set;

[0023] The simulation early warning submodule performs operation early warning based on the simulation early warning data set and obtains corresponding operation early warning results;

[0024] If the operation early warning result is no abnormality, the full-process operation state is continuously monitored;

[0025] If the operation early warning result is abnormality, the corresponding operation early warning information set and operation incentive points are generated;

[0026] The operation early warning information set and operation incentive points are packaged as operation early warning data set for output.

[0027] As a further scheme of the present application, the simulation early warning submodule performs operation early warning based on the simulation early warning data set and obtains corresponding operation early warning results, including:

[0028] The simulation early warning submodule analyzes the simulation early warning data set and obtains the difference degree of the corresponding digital twin environment before and after the controllable noise is injected;

[0029] The difference degree is mapped and transformed according to a preset rule to generate corresponding early risk early warning score;

[0030] If the early risk early warning score reaches a predetermined threshold, it is determined that there is an abnormality;

[0031] If the early risk early warning score does not reach the predetermined threshold, it is determined that there is no abnormality;

[0032] Based on the whole-process operation resilience knowledge graph, and using the attention mechanism to locate the root cause of the abnormal simulation early warning data set, an abnormal root node set is obtained;

[0033] Reasoning the entity nodes contained in the abnormal root node set to generate corresponding root cause reasoning information, the root cause reasoning information set at least includes abnormal root cause inducement and initial root cause inducement elimination strategy;

[0034] Encapsulating the abnormal root node set and the corresponding root cause reasoning information set into an operation early warning information set, analyzing the initial root cause inducement elimination strategy based on a preset operation point mechanism, and generating corresponding operation incentive points.

[0035] As a further scheme of the application, based on the operation early warning data set, a game negotiation is carried out to obtain a game negotiation result, and an operation control strategy is generated according to the game negotiation result, including:

[0036] Based on the operation early warning information set, an operation risk verification is carried out to obtain an operation risk verification result, and the operation early warning information set is feedback optimized according to the operation risk verification result to generate a corresponding operation risk resource set;

[0037] The operation risk resource set at least includes an operation risk root entity set, an operation risk elimination resource set and an initial root cause inducement elimination strategy;

[0038] Using multi-agent reinforcement learning and dynamic game algorithm, the operation risk resource set is combined to carry out game negotiation to generate a corresponding game negotiation result;

[0039] Based on the game negotiation result, the initial resource scheduling strategy is optimized and adjusted to generate a corresponding operation control strategy.

[0040] As a further scheme of the application, multi-agent reinforcement learning and dynamic game algorithm are used to combine the operation risk resource set to carry out game negotiation to generate a corresponding game negotiation result, including:

[0041] A corresponding game intelligent agent is configured for all operation processes of the operation whole process, and the operation whole process at least includes four operation processes of engineering construction, spare parts management, maintenance and detection service;

[0042] Based on the initial root cause inducement elimination strategy, each game intelligent agent is initialized, and each game intelligent agent synchronously combines the operation risk root entity set and the operation risk elimination resource set to analyze the rationality of the initial root cause inducement elimination strategy based on its own scheduling cost;

[0043] The initial root cause elimination strategy includes an initial operational incentive points budget for eliminating the root causes of risk. The scheduling cost is represented by the actual number of operational incentive points generated according to the preset cost conversion rules, and includes at least the actual loss cost and the waiting compensation cost.

[0044] Based on the results of the rationality analysis, each game agent negotiates with the goal of minimizing scheduling costs and generates corresponding operation and control strategies.

[0045] As a further aspect of the present invention, operational incentive points are settled based on the implementation effect of operational control strategies, an optimization suggestion set is generated, and feedback optimization is performed on the operational early warning model, including:

[0046] The settlement of operational incentive points refers to the deviation between the initial operational incentive point budget and the actual operational incentive point expenditure of the fulfilling intelligent agent of the operational control strategy, and the corresponding settlement result is generated based on the deviation value.

[0047] If the settlement result is that the budget is sufficient, the remaining operating incentive points after paying the actual operating incentive points will be settled as performance remuneration to each performing intelligent agent. Each performing intelligent agent will pay commissions in accordance with the preset rules and inject them into the risk resistance reserve pool.

[0048] If the settlement result is that the budget is sufficient but the cost exceeds the limit due to the failure of the performing agent to perform, the points compensation pool will be activated to compensate for the corresponding excess cost, and the points of the failing agent will be deducted. The points deduction includes excess cost deduction and credit deduction.

[0049] The operational incentive points in the points compensation pool are derived from the point deductions of unfulfilled intelligent agents;

[0050] If the settlement result shows that the budget seriously underestimates the actual cost, the risk-resistance reserve pool will be activated to compensate each performing intelligent agent for costs. The cost compensation includes advance payment of points costs and reward points compensation.

[0051] An optimization suggestion set is generated based on the liquidation results, and the operation early warning model is optimized based on the optimization suggestion set.

[0052] As a further aspect of the present invention, a full-process operation dataset is obtained, which is divided into a historical full-process operation dataset and a real-time full-process operation dataset. A full-process operation resilience knowledge graph is constructed and updated in real-time based on the full-process operation dataset, including:

[0053] Multi-source data collection is performed on the entire operation and scheduling process to obtain a full-process operation dataset. The full-process operation dataset includes at least engineering construction process data, spare parts management process data, maintenance and repair process data, and testing service process data.

[0054] Feature extraction and semantic analysis are performed on the full-process operation dataset to obtain the full-process operation entity set and the corresponding full-process topology association network. Taking each entity in the full-process operation entity set as a node, the full-process operation resilience knowledge graph is constructed simultaneously by combining the full-process topology association network.

[0055] Based on graph neural networks and causal reasoning, the resilience of the full-process operational resilience knowledge graph is assessed, and an operational risk score is assigned to each entity node.

[0056] The resilience assessment is expressed as a quantitative measure of the operational resilience of the current node by combining the entity node's own state health, topological importance, and resistance to risk contagion from neighboring nodes.

[0057] The state health is used to quantify the current operational status of an entity node, the topological importance is used to quantify the importance of the entity node in the entire process operation, and the neighbor node risk contagion resistance is used to quantify the ability of the current node to maintain its own state health when neighbor nodes experience operational risks.

[0058] Furthermore, embodiments of the present invention also provide an automated multi-agent end-to-end operation and scheduling platform, comprising:

[0059] The data acquisition module is used to acquire the full-process operation dataset and construct and update the full-process operation resilience knowledge graph in real time based on the full-process operation dataset.

[0060] The model building module is used to build an operational early warning model and pre-train the model using historical full-process operational datasets.

[0061] An operation early warning module is used to input real-time full-process operation datasets into the operation early warning model to generate operation early warning datasets;

[0062] The game decision-making module is used to conduct game negotiation based on the operation early warning dataset, obtain the game negotiation result, and generate operation control strategy based on the game negotiation result.

[0063] The feedback adjustment module is used to perform operational incentive points settlement and generate an optimization suggestion set based on the results of the operational incentive points settlement to optimize the operational early warning model.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] A full-process operation dataset is acquired, and a full-process operation resilience knowledge graph is constructed and updated in real time based on the full-process operation dataset. By integrating multi-source heterogeneous data of the full-process operation with the construction of the knowledge graph, the topological relationship and causal relationship of cross-link scheduling of the full-process operation are visualized, providing a data foundation for subsequent steps.

[0066] An operational early warning model is constructed based on generative adversarial networks and graph neural networks. The model is pre-trained using historical full-process operational datasets. Real-time full-process operational datasets are input into the operational early warning model to generate operational early warning datasets. By amplifying weak signals in the operational process through the operational early warning model, advanced early warning of progressive failures is achieved.

[0067] Based on the aforementioned operational early warning dataset, a game negotiation is conducted to obtain the negotiation results. Based on the negotiation results, an operational control strategy is generated, which realizes multi-resource collaborative optimization under budget constraints and ensures the economy and feasibility of the operational control strategy in actual operational scenarios.

[0068] Based on the implementation effect of the operation control strategy, the operation incentive points are settled, an optimization suggestion set is generated, and the operation early warning model is optimized by feedback, so as to achieve efficient and accurate scheduling of operation resources. Attached Figure Description

[0069] Figure 1 This is a flowchart of the steps of an automated multi-agent full-process operation scheduling method according to the present invention;

[0070] Figure 2 This is a flowchart of step S3 in the automated multi-agent full-process operation scheduling method of the present invention;

[0071] Figure 3 This is a schematic diagram of an automated multi-agent full-process operation and scheduling platform according to the present invention. Detailed Implementation

[0072] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating the steps of an automated multi-agent full-process operation scheduling method according to the present invention. Figure 2 This is a flowchart of step S3 in the automated multi-agent full-process operation scheduling method of the present invention. The following is a detailed introduction to the automated multi-agent full-process operation scheduling method.

[0073] Step S1: Obtain the full-process operation dataset, and construct and update the full-process operation resilience knowledge graph in real time based on the full-process operation dataset.

[0074] Specifically, multi-source data is collected from the entire operation and scheduling process to obtain a full-process operation dataset. This full-process operation dataset includes at least engineering construction process data, spare parts management process data, maintenance and repair process data, and testing service process data. Feature extraction and semantic analysis are performed on the full-process operation dataset to obtain a set of full-process operation entities and a corresponding full-process topology network. Using each entity in the set of full-process operation entities as a node, a full-process operation resilience knowledge graph is constructed simultaneously with the full-process topology network. Based on graph neural networks and causal reasoning, the resilience of the full-process operation resilience knowledge graph is assessed, and an operational risk score is assigned to each entity node.

[0075] Understandably, resilience assessment is expressed as quantifying the operational resilience of a current node by combining its own state health, topological importance, and resistance to risk contagion from neighboring nodes; the state health is used to quantify the current operational state of the entity node; the topological importance is used to quantify the importance of the entity node in the entire process operation; and the resistance to risk contagion from neighboring nodes is used to quantify the current node's ability to maintain its own state health when neighboring nodes experience operational risks.

[0076] In one possible embodiment, real-time and historical data are continuously collected from four core business processes: engineering construction, spare parts management, maintenance and repair, and testing services. Data preprocessing is performed on the data collected from all these business processes to generate a corresponding full-process operation data vector. Taking pump installation and commissioning as an example, pump installation and commissioning records from the engineering construction process are collected simultaneously. These records include at least alignment accuracy and initial vibration values. Spare parts management records show the inventory quantity and procurement and transportation information of the corresponding spare bearing for the pump. Maintenance and repair records show the real-time vibration sensor readings and repair work orders for the pump. The latest spectral analysis results of the lubricating oil for the pump are also collected from the testing service process. The raw data is cleaned, denoised, and missing values ​​are filled. Simultaneously, the pump number P-101, spare bearing number B-203, and repair work order E-302 in the pump installation and commissioning records are associated to generate a full-process operation data vector representing the operational resources required for the pump with pump number P-101 during installation and commissioning.

[0077] By employing graph neural network technology, a full-process operation resilience knowledge graph is constructed by simultaneously combining a full-process operation dataset composed of full-process operation data vectors. The graph nodes of the full-process operation resilience knowledge graph represent full-process operation entities, such as equipment, personnel, spare parts, and tasks. The edges represent the relationships between entities, such as dependency, supply, execution, and location. For example, reactor R-201 is an entity node. It is connected to heat exchanger E-15 through a process dependency edge, to mechanical seal S-07 in the spare parts warehouse through a spare parts supply edge, and to maintenance team T-03 through a maintenance edge. The edge connections of this reactor demonstrate that the health status of reactor R-201 directly affects heat exchanger E-15, and its maintenance depends on the availability of mechanical seal S-07 and maintenance team T-03.

[0078] An operational risk score is calculated for each node in the graph. This score quantifies the node's vulnerability and is generated by comprehensively evaluating the node's own health status, its topological importance in the network, and the potential impact of the status of its neighboring nodes. Specifically, the quantitative evaluation of the node's own health status is represented by setting corresponding monitoring indicators for each entity node, such as vibration, temperature, and inventory level, and setting a normal range [L, U] for each monitoring indicator. For numerical monitoring indicators, the indicator anomaly is set to max(0, (actual value - U) / (UL)). For Boolean indicators, the indicator anomaly is set to 0 or 1. For example, if the maintenance team is online, the indicator anomaly is set to 1. Finally, the indicator anomaly of each node is weighted and combined according to a preset weight ratio. The constant is used to generate a node health score; the topological importance is represented by a weighted fusion of the betweenness centrality, compact centrality, and eigenvector centrality of the node, generating a corresponding topological importance score; the quantification of the potential impact caused by neighboring nodes is represented by using Granger causality tests combined with historical data to calculate the impact coefficient of the state change of neighboring node j on the current node i. If there is no historical data, it is set by experts according to business logic, such as the impact intensity of spare parts inventory on equipment being higher than the impact intensity of tool use. The node health score, topological importance score, and impact coefficient are weighted and fused according to a preset weight ratio to generate a corresponding operational risk score. It should be noted that the preset weight ratio used in the calculation of the operational risk score needs to be specifically allocated based on the importance of the data involved in the actual operation process.

[0079] Taking the calculation of the operational risk score of centrifugal pump P-101 as an example, assuming its corresponding monitoring indicators are vibration and temperature, analysis of historical data of centrifugal pump P-101 reveals that vibration has a greater impact on centrifugal pump P-101 than temperature. Therefore, the weight ratio can be set to 6:4, with vibration anomaly of 0.8 and temperature anomaly of 0.4, resulting in a node health score of 0.64. Furthermore, its betweenness centrality is 0.9 with a corresponding weight of 0.5, its compact centrality is 0.7 with a corresponding weight of 0.3, and its eigenvector centrality is 0.8 with a corresponding weight of 0.2, indicating its topological importance. The health score is 0.82. The neighboring nodes of centrifugal pump P-101 include spare parts node S-51 and engineer node T-13, with health scores of 0.8 and 0.3 respectively. Assuming that historical data analysis reveals a significant increase in the probability of centrifugal pump P-101 failing within the next week whenever spare parts S-51 inventory falls below a preset safety threshold, statistical analysis and normalization determine the influence strength of spare parts S-51 to be 0.9. Similarly, historical data analysis reveals that the workload of engineer T-13 significantly impacts the failure repair time of centrifugal pump P-101. There are moderate but non-decisive influencing factors. Granger causality tests and fuzzy analysis are used to transform these factors, resulting in an influence strength of 0.6. The inverse ratio of the shortest path length between neighboring nodes and the current node is used as an attenuation factor to simulate the attenuation effect of neighboring nodes' influence within the graph. Assuming the shortest path lengths from spare part node S-51 and engineer node T-13 to centrifugal pump P-101 are 1 and 2 respectively, their corresponding attenuation factors are 1 and 0.5 respectively. Combining the influence strength, the node health score of the neighboring nodes, and the attenuation factor, the influence of spare part node S-51 and engineer node T-101 is obtained. The influence coefficient of engineer node T-13 on centrifugal pump P-101 is 0.9*0.8*1+0.6*0.3*0.5=0.81. Analysis of the historical full-process operation dataset shows that the node health score and topology importance score contribute more to the final operational risk than the influence coefficient. Therefore, assuming the weight ratio of centrifugal pump P-101 node is node health score: topology importance score: influence coefficient = 4:4:2, the corresponding operational risk score for this node is 0.64*0.4+0.82*0.4+0.81*0.2=0.746.

[0080] Step S2: Construct an operation early warning model based on generative adversarial networks and graph neural networks, and pre-train the model using historical full-process operation datasets.

[0081] Understandably, a digital twin environment is constructed based on a knowledge graph of end-to-end operational resilience and a historical end-to-end operational dataset. This digital twin environment is then updated in real time using a real-time end-to-end operational dataset. The digital twin environment is used to simulate the end-to-end operational state and contains a prediction model, such as a temporal convolutional network or a long short-term memory network. This model is pre-trained based on a large amount of health data under normal operating conditions. The prediction model can predict the state vector under future normal conditions based on the current input state vector.

[0082] Clustering is performed on the historical full-process operation dataset to obtain a non-stationary operation dataset. The non-stationary operation dataset refers to historical data on operational process imbalances that occurred during the full-process operation, and includes at least root cause data and impact data of operational imbalances. For example, data from all normal operating periods of centrifugal pump P-101 in the past year are extracted as health samples. At the same time, historical data of failures in the past year are extracted based on clustering algorithms. Assuming that centrifugal pump P-101 experienced a bearing failure six months ago, the data sequence of the 48 hours before the failure is extracted and marked as an early risk sample. The impact data caused by this failure are also obtained.

[0083] The impact data refers to the series of changes and losses caused by a failure in the operating system consisting of four stages: engineering construction, spare parts management, maintenance and repair, and testing services. For example, in the engineering construction stage, the shutdown of centrifugal pump P-101, a critical piece of equipment, causes a complete shutdown of the production line for 2 hours. This 2-hour production loss can be used as the impact data for the engineering construction stage. In the maintenance and repair and spare parts management stages, the emergency repair costs incurred by this repair, such as overtime pay, emergency spare parts allocation and transportation costs, and outsourcing service fees, are all recorded as the impact data for the maintenance and repair stage. Similarly, changes in spare parts inventory are used as the impact data for the spare parts management stage. For example, if a critical spare bearing is consumed, causing the safety stock of that model of spare parts to be exceeded, requiring emergency procurement, the change in inventory level from safe to emergency will be recorded as the impact data for this stage. In the testing service stage, assuming that the sudden shutdown of P-101 causes secondary damage to its upstream equipment, the abnormal data of these related equipment during and after the failure will be recorded as the impact data for the testing service stage.

[0084] In this embodiment, step S2 includes:

[0085] Step S2-1: Construct and train the risk detection submodule.

[0086] Specifically, a reinforcement learning agent is constructed, its corresponding state space is defined according to the digital twin environment, its action space is set to inject controllable noise for early risk detection, a deep neural network is used as its policy network, and the policy network is pre-trained simultaneously with the non-stationary operation dataset. The pre-trained reinforcement learning agent is used as a risk detection submodule, which is used to generate controllable noise, and the controllable noise is used to input into the simulation early warning submodule for early operational risk identification.

[0087] In one possible embodiment, the state space of the reinforcement learning agent is defined by combining the current state of each node in the pre-digital twin environment within the end-to-end operational resilience knowledge graph. For example, the state space can be defined by combining the corresponding states of device vibration, temperature, and load. The intensity, frequency, and target device of controllable noise within the action space are set. The policy network adopts a dual-network structure of Actor and Critic networks. The Actor network is responsible for generating action policies, taking the state vector corresponding to the state space as input and outputting the mean and variance of the generated action policies to define the specific characteristics of noise injection. For example, for the action dimension of noise intensity, the Actor network outputs a mean of 0.3 and a variance of 0.1, meaning that under the current system state, the optimal noise intensity is likely to be 0.3. Nearby, noise intensity is adjusted by combining variance, i.e., noise intensity sampling is performed according to a normal distribution of N(0.3,0.1²). The Critic network is responsible for evaluating the value of action policies. Its input is the state vector in the state space and the action vector in the action space, and it finally outputs a state-action value estimate to guide the Actor network in policy optimization. A reward function is constructed using a multi-objective weighted combination approach. The reward function includes at least a detection reward, a safety penalty, an efficiency reward, and an exploration entropy reward. The detection reward is used to reward effective stimulation of operational process anomalies. The safety penalty is used to punish excessive behavior that may lead to system crash. The efficiency reward is used to promote achieving the detection objective with minimum energy. The exploration entropy reward is used to maintain policy diversity.

[0088] The pre-training process is carried out in the digital twin environment. Taking the centrifugal pump P-101 as a specific training example, the agent's policy network generates a specific action policy based on the state vector at a certain moment. Assuming the parameters of this action policy are [type: {sine sweep frequency, probability: 0.7}, center frequency: 237Hz, bandwidth: 45Hz, gain: 0.23, duration: 2.1 seconds], after the incentive is executed in the digital twin environment, the corresponding comprehensive reward value is calculated based on the reward function, which is 0.156. The corresponding parameters of the policy network are updated through the proximal policy optimization algorithm. The entire training process adopts an iterative training method to optimize the parameters of the policy network until the preset iteration conditions are met, such as the average comprehensive reward value of the most recent 100 rounds being greater than 0.5 and the detection success rate on the validation set exceeding 85%, at which point the training ends. After training, the risk detection submodule can generate the corresponding optimal detection policy for different full-process operation states, i.e., the most suitable controllable noise. This controllable noise can be used as a detection signal to effectively and safely test the stability of the current operation process state.

[0089] Understandably, the detection reward can be represented as Tanh((actual resonance intensity - resonance intensity threshold) / resonance intensity threshold), where the resonance intensity is calculated by the simulation warning submodule and used to quantify the sensitivity to specific noise excitation. The default weight of this item is 1.0. For example, after the agent injects controllable noise, the simulation warning submodule analyzes and obtains the corresponding actual resonance intensity of 0.83. Assuming the resonance intensity threshold is 0.7, the corresponding detection reward is tanh((0.83-0.7) / 0.7)≈0.18. This detection reward means that a significant resonance was successfully triggered, and the agent receives a positive reward of +0.18. The safety penalty can be represented as -max(0,(X p -X s ) / X s ) 2 , where X p X represents the peak value of the key monitoring indicator after noise injection. s This represents the safety threshold for key monitoring indicators, with a default weight of 1.5. For example, assuming the key monitoring indicator for centrifugal pump P-101 is vibration, and the peak vibration after injecting controllable noise A1 is 12.5 mm / s, and the safety threshold is 15 mm / s, then the corresponding safety penalty term is -max(0,(12.5-15) / 15). 2 =0 means the current operation is within the safe range and there is no penalty. The vibration peak after injecting controllable noise A2 is 16.5mm / s. Therefore, the corresponding safety penalty term is -max(0,(16.5-15) / 15). 2=-0.01 means the current operation exceeds the safe limit and a penalty must be imposed; the efficiency reward can be represented as -λ*ln(1+E), where λ is a preset adjustment coefficient to ensure that this item does not dominate the reward and defaults to 1.0, E is the total energy of the injected noise, which is calculated based on the amplitude and duration of the injected noise by default, and the default weight of the efficiency reward is 0.2. For example, if the total energy of the noise injected by the agent is 0.35, then the corresponding efficiency reward is -0.1*ln(1+0.35)≈-0.03; the exploration entropy reward can be represented as η*H, where η is the fine-tuning exploration coefficient and defaults to 0. .01, where H is the entropy of the action policy in a given state. The higher the entropy value, the more uniform the probability distribution of the policy, meaning that the probability of choosing different actions is closer, and the exploration is stronger. The weight of this reward item is 1.0. For example, in a certain state, the agent's policy network outputs the mean and variance of 6 action policies. Assuming that the policy entropy H calculated based on these parameters is 2.1, the corresponding exploration entropy reward item is 0.01*2.1=0.021. Combining the above, the comprehensive reward calculated is 1.0*(0.18-0.03+0.021)+1.5*(-0.01)=0.156.

[0090] Step S2-2: Construct and train the simulation early warning submodule.

[0091] Specifically, controllable noise is input into the digital twin environment for simulation to generate a simulation early warning dataset. The simulation early warning dataset includes an initial state vector, controllable noise, a perturbed state vector, and an unperturbed state vector. A simulation early warning submodule is constructed based on a deep twin network. The simulation early warning dataset is input into the simulation early warning submodule. The difference between the perturbed state vector and the unperturbed state vector is obtained according to the attention mechanism network. Pre-training is performed in combination with a predetermined training objective. Training performance is verified in combination with a training and validation set constructed based on a non-stationary operation dataset. Training ends only when the predetermined performance condition is met.

[0092] Understandably, the initial state vector represents the initial state vector of the digital twin environment, the perturbation state vector represents the state vector sequence of the digital twin environment within a fixed time window after injecting controllable noise, and the unperturbed state vector represents the state vector sequence of the digital twin environment within a fixed time window without injecting controllable noise.

[0093] Understandably, a simulation early warning submodule is constructed based on a deep twin network architecture. This submodule is used to quantify the dynamic sensitivity of the current state to noise excitation. The deep twin network architecture consists of two subnetworks with shared weights. Each subnetwork is constructed using a temporal neural network with an embedded attention mechanism. One branch is used to receive the perturbed state vector output by the digital twin environment after injecting controllable noise perturbation. The other branch is used to receive the normal state vector sequence of the system over a period of time in the future under the same initial conditions, predicted by the prediction model embedded in the digital twin environment, i.e., the undisturbed state vector. The attention mechanism is used to guide the simulation early warning submodule to focus on the key system variables and key time points that are most significant in response to perturbation, thereby capturing abnormal patterns more accurately.

[0094] In one possible embodiment, the pre-training process uses controllable noise generated by the risk detection submodule as a disturbance source for the digital twin environment to obtain the corresponding simulation early warning dataset. Pre-training is then performed in conjunction with a preset training objective. Specifically, taking centrifugal pump P-101 as an example, assuming that during a certain training process, the corresponding unperturbed state vector sequence is {vibration: [10.0, 10.1, 10.2] mm / s, temperature: [65.0, 65.1, 65.2] ℃}, after injecting controllable noises A1 and A2, the corresponding perturbed state vector sequences are A1 {vibration: [10.1, 10.8, 12.5] mm / s, temperature: [65.1, 65.5, 66.5] ℃} and controllable noise A2 {vibration: [10.0, 10.2, 10.3] mm / s, temperature: [65.0, 10.2, 10.3] mm / s, temperature: [65.0, 10.2, 10.3] ℃}, respectively. [65.2,65.3]℃}, based on the pre-trained gated recurrent unit, features are extracted from the unperturbed state vector sequence and the perturbed state vector sequence. The pre-training dataset of the gated recurrent unit is the historical full-process operation dataset. The feature vector corresponding to the unperturbed state vector sequence is used as the benchmark to generate the benchmark feature vector [1.0,1.0]. The feature vectors corresponding to A1 and A2 are [2.8,3.2] and [1.1,1.0], respectively. The first dimension of the feature vector represents the degree to which the injected noise energy is amplified, which is used to quantify the instability of the current state. The second dimension represents the duration and coherence of the abnormal response, which is used to quantify the difficulty of recovery after deviating from the normal state (this is just an example of the transformation from the original state vector to the corresponding feature vector. The specific transformation needs to be determined in combination with the actual situation).

[0095] Calculate the Euclidean distance between the eigenvectors corresponding to A1 and A2 and the baseline eigenvector, and use the Euclidean distance as the difference degree. The difference degrees corresponding to A1 and A2 are 2.84 and 0.10, respectively. A resonance mapping function is used to generate the corresponding resonance intensity. The magnitude of the resonance intensity value directly quantifies the sensitivity of this controllable noise excitation. The higher the value, the greater the deviation between the actual response and the normal baseline, meaning that small disturbances are amplified more significantly, and the higher the risk that the current operating process is in a dynamically unstable critical state. The resonance mapping function can be expressed as 2 / [1+e^(-k*(D-D0))]-1, where k is the curvature factor with a default value of 2, used to control the mapping steepness, and D0 is the baseline difference degree with a default value of 1. If the difference and resonance intensity of A1 and A2 are 0.87 and 0.01 respectively, and the predetermined benchmark resonance intensity is 0.1, then the difference and resonance intensity of A1 both exceed the corresponding benchmark values, indicating that the controllable noise corresponding to A1 can cause drastic fluctuations in the operating state within the controllable range, and is defined as effective disturbance noise. Similarly, since the controllable noise corresponding to A2 cannot cause drastic fluctuations in the operating state within the controllable range, it is defined as ineffective disturbance noise. The training objective is to maximize the difference of effective disturbance noise and minimize the difference of ineffective disturbance noise. Iterative training is performed using a contrastive loss function, and the training performance is verified by combining a training and validation set built based on a non-stationary operation dataset. Training ends only when the predetermined performance condition is met.

[0096] Understandably, the predetermined performance conditions include: the fluctuation range of the contrastive loss function value of the training and validation sets for three consecutive times is less than a preset fluctuation threshold, such as the most recent three contrastive loss function values ​​being 0.125, 0.124, and 0.126, which are lower than the preset fluctuation threshold of 0.03; the classification accuracy of effective perturbation noise and invalid perturbation noise in the training and validation sets consistently reaches above 95%; if the contrastive loss function value of the training and validation sets does not refresh the optimal loss record for 10 consecutive times, early stopping is triggered to prevent overfitting. For example, assuming the loss value corresponding to the optimal loss record is 0.42, but the corresponding loss value for 10 consecutive training sessions does not fall below 0.42, then the early stopping mechanism is activated to prevent overfitting.

[0097] Step S3: Input the real-time full-process operation dataset into the operation early warning model to generate an operation early warning dataset.

[0098] Specifically, after the real-time full-process operation dataset is input into the operation early warning model, the risk detection submodule generates corresponding controllable noise based on the real-time full-process operation dataset; the controllable noise is injected into a pre-built digital twin environment for simulation, generating a corresponding simulation early warning dataset; the simulation early warning submodule performs operation early warning based on the simulation early warning dataset and obtains the corresponding operation early warning result; if the operation early warning result is no abnormality, the full-process operation status continues to be monitored; if the operation early warning result is abnormality, a corresponding operation early warning information set and operation incentive points are generated; the operation early warning information set and operation incentive points are encapsulated into an operation early warning dataset for output.

[0099] Understandably, the simulation early warning submodule parses the simulation early warning dataset to obtain the difference degree between the digital twin environment before and after the injection of controllable noise; it maps and transforms the difference degree according to preset rules to generate a corresponding early risk warning score; if the early risk warning score reaches a predetermined threshold, it is determined that an anomaly has occurred; if the early risk warning score does not reach the predetermined threshold, it is determined that there is no anomaly; based on the full-process operational resilience knowledge graph and using an attention mechanism, the simulation early warning dataset with anomalies is used to locate the root cause and obtain a set of anomaly root cause nodes; the entity nodes contained in the set of anomaly root cause nodes are used for root cause reasoning to generate corresponding root cause reasoning information, the root cause reasoning information set includes at least the anomaly root cause trigger and the initial root cause trigger elimination strategy; the anomaly root cause node set and the corresponding root cause reasoning information set are encapsulated into an operational early warning information set, and the initial root cause trigger elimination strategy is analyzed based on a preset operational points mechanism to generate corresponding operational incentive points.

[0100] In one possible embodiment, if the resonance intensity does not exceed a predetermined threshold, the current state is determined to be stable and no resource scheduling is required. If the resonance intensity exceeds the predetermined threshold, the gradient of the difference between the characteristics of each input node corresponding to the generated controllable noise and the final output resonance intensity is obtained based on the graph neural network inside the operation early warning model. This quantifies the contribution of each entity node to the current resonance intensity. Taking centrifugal pump P-101 as an example, its vibration data, spare bearing inventory, and the load of the engineer responsible for this scheduling task constitute the entity nodes of this operation scheduling. While calculating the resonance intensity, the operation early warning model will combine its own graph neural network to process the data corresponding to all entity nodes. Its internal self-attention mechanism will evaluate the influence relationship between each entity node. Assuming that it calculates the influence of the spare bearing inventory node on the final output... The attention weight of the output node is as high as 0.7, while the attention weight of the engineer load node is only 0.1. At the same time, by calculating the gradient, it was found that a small change in the vibration data of the centrifugal pump will cause a drastic fluctuation in the final output resonance intensity value, and its gradient value is much higher than that of other nodes. Therefore, it is determined that the spare bearing inventory and the centrifugal pump vibration are the main contributors to this resonance. The node with the highest contribution is identified as the primary excitation point of the resonance. Then, starting from the primary excitation point, probabilistic reverse reasoning is performed along the corresponding causal edge on the knowledge graph of the full-process operation resilience. Combined with the probabilistic graphical model constructed by historical fault data, the posterior probability of different root causes leading to the current resonance phenomenon is obtained, thereby determining the most likely fault root cause and its influence path. The fault root cause and its influence path are then encapsulated as an abnormal root cause inducement and output.

[0101] The root cause of the anomaly is matched with a pre-built historical case library and contingency plan library. For example, a contingency plan match is performed in the historical case library and contingency plan library in the form of {bearing wear, resonance intensity: 0.74}. Based on the matched contingency plan, the current operational resource pool is searched to find a set of candidate resources that can perform the task, such as a list of qualified engineers, inventory information of required spare parts, and availability status of special tools. Combined with relevant content such as cost efficiency, geographical location, and current load, the optimal resource combination is selected from the candidate resource set, and the approximate task loss is estimated, thereby generating the corresponding initial root cause elimination strategy. For example, senior engineer Zhang San is assigned to use spare part B-202 to replace parts, which is expected to take 4 hours.

[0102] A fuzzy mapping algorithm is used to map resonance intensity values ​​to corresponding risk levels. For example, a resonance intensity value below 0.3 is considered low risk; a resonance intensity value in the range [0.3, 0.6) is considered medium risk; a resonance intensity value in the range [0.6, 0.8) is considered high risk; and a resonance intensity value above 0.8 is considered emergency risk. Each risk level corresponds to a basic risk coefficient, such as 1.0 for low risk, 1.2 for medium risk, 1.5 for high risk, and 2.0 for emergency risk. The impact path of the current predicted risk is obtained by analyzing the number of potentially affected related equipment based on the full-process operational resilience knowledge graph, and generating corresponding risk impact coefficients. For example, the risk impact coefficient for affecting one piece of equipment is set to 1.0, the risk impact coefficient for affecting one subsystem is set to 1.3, and the risk impact coefficient for affecting the entire plant system is set to 2.0. Finally, the corresponding comprehensive risk coefficient is generated based on the formula: comprehensive risk coefficient = basic risk coefficient * risk impact coefficient.

[0103] The system analyzes all the resources required for the initial root cause elimination strategy and generates a detailed resource list. The resource list includes human resources, such as the skill level and required working hours of engineers; material resources, such as the model and quantity of spare parts; tool resources, such as the time occupied by special testing equipment; and collaboration resources, such as the working hours required for cooperation from other departments. Each resource in the resource list is priced according to its historical average cost or preset standard cost to generate a corresponding benchmark cost. For example, the cost of a senior engineer per working hour is 50 points, and the cost of a specific model of bearing is 200 points.

[0104] The initial integral budget is calculated as follows: Initial Integral Budget = Base Cost * Comprehensive Risk Coefficient * Emergency Scheduling Premium * Resource Scarcity Premium. The emergency scheduling premium is an additional premium factor, such as 1.2, if the task needs to respond during non-working hours or in a very short time. The default value is 1.0 during normal working hours or under normal response conditions. The resource scarcity premium is an additional premium factor if the required resources are currently in short supply or in high demand. Similarly, the default value is 1.0 if none of the above situations apply.

[0105] It should be noted that, in order to prevent the initial incentive budget from becoming excessively inflated, the calculated initial incentive budget will be compared with the actual cost of similar tasks in the past, and an upper limit will be set, such as not exceeding 150% of the highest historical cost. Finally, the calculated initial incentive budget will be rounded or smoothed to generate the final initial operational incentive incentive budget.

[0106] Step S4: Conduct game negotiation based on the operation early warning dataset, obtain the game negotiation result, and generate an operation control strategy based on the game negotiation result.

[0107] Specifically, operational risks are verified based on the operational early warning information set to obtain operational risk verification results. The operational early warning information set is then optimized based on these results to generate a corresponding operational risk resource set. This operational risk resource set includes at least a set of operational risk root cause entities, an operational risk elimination resource set, and an initial root cause elimination strategy. A multi-agent reinforcement learning and dynamic game theory algorithm is employed, simultaneously combined with the operational risk resource set, to conduct game negotiation and generate corresponding game negotiation results. Based on these game negotiation results, the initial resource scheduling strategy is optimized and adjusted to generate a corresponding operational control strategy.

[0108] Understandably, corresponding game-theoretic agents are configured for all operational processes throughout the entire operation process, which includes at least four operational processes: engineering construction, spare parts management, maintenance and repair, and testing services. Each game-theoretic agent is initialized based on an initial root cause elimination strategy. Each agent, based on its own scheduling cost, simultaneously analyzes the rationality of the initial root cause elimination strategy by combining the set of operational risk root cause entities and the set of operational risk elimination resources. The initial root cause elimination strategy includes an initial operational incentive point budget for eliminating risk root causes. The scheduling cost is represented by the actual number of operational incentive points generated according to a preset cost conversion rule, and includes at least the actual loss cost and the waiting compensation cost. Based on the rationality analysis results, each game-theoretic agent negotiates with the goal of minimizing the scheduling cost, generating a corresponding operational control strategy.

[0109] In one possible embodiment, corresponding agents are created for the operational resources involved based on the initial root cause elimination strategy, and an initial bid is set for each agent based on the baseline cost of each agent. A multi-round negotiation process is initiated, in which each agent adjusts its initial bid based on its own state, such as workload and resource scarcity, and in combination with the bids received from other agents, using a game theory-based negotiation strategy to generate a negotiated bid. It is checked whether the total negotiated bid exceeds the initial operational incentive budget, and the scheme that exceeds the budget is constrained and optimized. When the bid converges or the maximum number of rounds is reached, the game stops, and the optimized scheduling strategy and the corresponding bid reward are output. The scheduling strategy and the corresponding bid reward are encapsulated into an operational control strategy for output.

[0110] For example, in a certain initial root cause elimination strategy, senior engineer Zhang San is assigned to replace centrifugal pump P-101 with spare part B-202. Based on the initial root cause elimination strategy and the current state of each resource, the engineer agent is initialized as follows: {Zhang San, base cost 50 points, initial price 70 points due to high current load}, and the spare part agent is initialized as follows: {Spare part B-202, base cost 30 points, initial price 50 points due to tight inventory}. The total initial price is 70 + 50 = 120 points. Assuming the budget is only 100 points, a prompt is made that the price exceeds the budget, and each agent needs to adjust the price. Therefore, multiple rounds of price negotiation are initiated. Assuming that after multiple rounds of price negotiation, the engineer agent evaluates that if the price is not reduced, Zhang San may lose the task, but considering Zhang San's current situation... Previously under high load, the engineer was willing to lower the price to 65 points. The spare parts agent discovered an alternative spare part, B-203, with a cost of only 25 points, to replace spare part B-202. However, spare part B-203's performance was slightly inferior to spare part B-202, so the price was lowered to 45 points, bringing the total price to 110 points, which still exceeded the budget. At this point, the global coordination and optimization mechanism was activated, proposing to use spare part B-203 and asking Zhang San if he would accept the collaboration. Zhang San assessed that using spare part B-203 would increase the difficulty and risk of maintenance, and requested an additional 10 points as risk compensation. The new price was 65 + 10 + 25 = 100 points, which met the budget. The final operation scheduling strategy was output as {Engineer Zhang San: 75 points, Spare Part B-203: 25 points}.

[0111] Step S5: Based on the implementation effect of the operation control strategy, the operation incentive points are settled, an optimization suggestion set is generated, and the operation early warning model is optimized by feedback.

[0112] Specifically, the deviation between the initial operational incentive points budget and the actual operational incentive points spent by the performance agent under the operational control strategy is settled, and a corresponding settlement result is generated based on the deviation.

[0113] If the settlement result is that the budget is sufficient, the remaining operating incentive points after paying the actual operating incentive points will be settled as performance remuneration to each performing intelligent agent. Each performing intelligent agent will pay commissions in accordance with the preset rules and inject them into the risk resistance reserve pool.

[0114] If the settlement result is that the budget is sufficient but the cost exceeds the limit due to the failure of the performing agent to perform, the points compensation pool will be activated to compensate for the corresponding excess cost, and the points of the failing agent will be deducted. The points deduction includes excess cost deduction and credit deduction. The operation incentive points of the points compensation pool come from the points deduction of the failing agent.

[0115] If the settlement result shows that the budget seriously underestimates the actual cost, the risk-resistance reserve pool will be activated to compensate each performing intelligent agent for costs. The cost compensation includes advance payment of points costs and reward points compensation.

[0116] An optimization suggestion set is generated based on the liquidation results, and the operation early warning model is optimized based on the optimization suggestion set.

[0117] In one possible embodiment, the resonance intensity output of the operation early warning model corresponding to each operation scheduling, the final scheduling strategy after the game, the integral cost of actual resource consumption, and the actual execution effect of the strategy are recorded, such as the actual time spent on maintenance tasks and the recovery status of key equipment indicators after maintenance. All the above data are timestamped to form a complete causal chain record. For example, for the operation event corresponding to centrifugal pump P-101, its resonance intensity is 0.74, the initial operation incentive integral budget is 100 points, the final solution after the game is {Engineer Zhang San: 75 points, spare part B-203: 25 points}, and the execution result data is recorded as: [Actual maintenance time: {5 hours, exceeding the expected maintenance time by 4 hours}, equipment effect after maintenance: {Vibration value decreased from 12.5 mm / s to 9.8 mm / s, although it did not fully recover to the optimal 8.0 mm / s, but the risk has been eliminated}].

[0118] Based on the aforementioned causal chain record, a two-way feedback optimization is performed on the operation early warning model and the game mechanism. For the operation early warning model, the execution result data is used as the ultimate label for its prediction accuracy. That is, if after a high resonance intensity early warning, although the scheduling is successful, the maintenance effect is poor or the cost far exceeds expectations, then the sample will be marked as having a prediction bias, and a corresponding incremental training set will be generated to fine-tune the operation early warning model so that its future predictions can better match the actual resource consumption and repair difficulty. For example, if the operation event corresponding to centrifugal pump P-101 is judged to be partially effective but not optimal, then this case will be used to fine-tune the operation early warning model so that it can more accurately assess its resonance intensity and potential risks when facing similar patterns in the future.

[0119] For the game theory mechanism, the differences between the final transaction price and the initial asking price and the benchmark cost in each operation scheduling, as well as the bargaining behavior of the agent, are analyzed to generate a bargaining training set. The bargaining training set is used to optimize the bargaining strategy of the agent and dynamically update the benchmark cost of each resource in the preset resource cost library. For example, if engineer Zhang San can win the transaction at a price 50 points higher than his benchmark cost in multiple games, his benchmark cost will be gradually increased to 60 or even 65 points, making the initial budget allocation more accurate in the future. At the same time, the agent corresponding to Zhang San will also learn from the successful bargaining experience and strengthen its asking price strategy when resources are scarce.

[0120] Figure 3 This is a schematic diagram of an automated multi-agent full-process operation and scheduling platform according to the present invention.

[0121] Specifically, an automated multi-agent end-to-end operation and scheduling platform includes:

[0122] The data acquisition module is used to acquire the full-process operation dataset and construct and update the full-process operation resilience knowledge graph in real time based on the full-process operation dataset.

[0123] The model building module is used to build an operational early warning model and pre-train the model using historical full-process operational datasets.

[0124] The operation early warning module is used to input real-time full-process operation datasets into the operation early warning model to generate operation early warning datasets.

[0125] The game decision-making module is used to conduct game negotiation based on the operation early warning dataset, obtain the game negotiation result, and generate operation control strategy based on the game negotiation result.

[0126] The feedback adjustment module is used to perform operational incentive points settlement and generate an optimization suggestion set based on the results of the operational incentive points settlement to optimize the operational early warning model.

[0127] The specific usage and function of this embodiment are explained below:

[0128] First, a full-process operation dataset is acquired. Based on the full-process operation dataset, a full-process operation resilience knowledge graph is constructed and updated in real time. By integrating multi-source heterogeneous data of the full-process operation and synchronously constructing the knowledge graph, the visualization of the corresponding topological relationship when scheduling the full-process operation across links is realized, providing a data foundation for subsequent steps.

[0129] Next, an operation early warning model was constructed based on generative adversarial networks and graph neural networks. The model was pre-trained using historical full-process operation datasets. Real-time full-process operation datasets were then input into the operation early warning model to generate operation early warning datasets. By amplifying weak signals in the operation process through the operation early warning model, advanced early warning of progressive failures was achieved.

[0130] Then, based on the operational early warning dataset, a game negotiation is conducted to obtain the negotiation result. Based on the negotiation result, an operational control strategy is generated, realizing multi-resource collaborative optimization under budget constraints, thereby ensuring the economy and feasibility of the operational control strategy in actual operational scenarios.

[0131] Finally, based on the implementation effect of the operational control strategy, the operational incentive points are settled, an optimization suggestion set is generated, and the operational early warning model is optimized by feedback, thus achieving efficient and accurate scheduling of operational resources.

[0132] An electronic device, comprising:

[0133] At least one processor; and at least one memory communicatively connected to the processor; wherein the memory stores instructions executable by at least one processor, the instructions being executed by at least one processor to enable at least one processor to perform the method proposed in Embodiment 1 of the present invention.

[0134] The following is a detailed introduction to the various components of the electronic device:

[0135] In this context, the processor is the control center of the electronic device. It can be a single processor or a collective term for multiple processing elements. For example, a processor can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement Embodiment 1 of this invention, such as one or more digital signal processors (DSPs) or one or more field-programmable gate arrays (FPGAs).

[0136] The processor can perform various functions of an electronic device by running or executing software programs stored in memory and by calling data stored in memory.

[0137] The memory is used to store the software program that executes the solution of the present invention, and the execution is controlled by the processor. For specific implementation methods, please refer to the above method embodiments, which will not be repeated here.

[0138] The memory can be a real-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. The memory can be integrated with the processor or exist independently and coupled to the processor through an interface circuit of an electronic device; this embodiment of the invention does not specifically limit this.

[0139] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via limited means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0140] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0141] It should be understood that, in the embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0142] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. An automated multi-agent full-process operation scheduling method, characterized in that, Comprising the following steps: Obtain a full-process operation dataset, which is divided into a historical full-process operation dataset and a real-time full-process operation dataset, and build and update a full-process operation resilience knowledge graph based on the full-process operation dataset; Build an operation early warning model based on a generative adversarial network and a graph neural network, and pre-train the model with the historical full-process operation dataset, wherein the operation early warning model includes a risk detection submodule and a simulation early warning submodule; Input the real-time full-process operation dataset into the operation early warning model for operation early warning, and generate an operation early warning dataset, which includes an operation early warning information set and operation incentive points; The generation of the operation early warning dataset represents that when the real-time full-process operation dataset is input into the operation early warning model, the risk detection submodule generates corresponding controllable noise based on the real-time full-process operation dataset; Inject the controllable noise into the pre-built digital twin environment for simulation, and generate a corresponding simulation early warning dataset; The simulation early warning submodule performs operation early warning based on the simulation early warning dataset, and obtains corresponding operation early warning results; If the operation early warning result is normal, continue to monitor the full-process operation state; If the operation early warning result is abnormal, generate the corresponding operation early warning information set and operation incentive points; Encapsulate the operation early warning information set and operation incentive points as the operation early warning dataset for output; Based on the operation early warning dataset, game negotiation is performed to obtain game negotiation results, and operation control strategies are generated based on the game negotiation results; According to the implementation effect of the operation control strategy, the operation incentive points are cleared, an optimization suggestion set is generated, and the operation early warning model is feedback optimized. 2.The automated multi-agent full-process operation scheduling method of claim 1, wherein, Build an operation early warning model based on a generative adversarial network and a graph neural network, and pre-train the model with the historical full-process operation dataset, wherein the operation early warning model includes a risk detection submodule and a simulation early warning submodule, comprising: Build a digital twin environment based on the full-process operation resilience knowledge graph and the historical full-process operation dataset, and update the digital twin environment in real time based on the real-time full-process operation dataset, wherein the digital twin environment is used to simulate the full-process operation state; Cluster the historical full-process operation dataset to obtain a non-stationary operation dataset, which represents historical data of operation process imbalance in the full-process operation process, and at least includes operation imbalance root cause data and operation imbalance impact data; Build a reinforcement learning agent, define its corresponding state space based on the digital twin environment, set its action space as injecting controllable noise for early risk detection, use a deep neural network as its policy network, pre-train the policy network in combination with the non-stationary operation dataset, and use the pre-trained reinforcement learning agent as the risk detection submodule; The risk detection submodule is used to generate controllable noise, which is used to input into the simulation early warning submodule for early operation risk identification. 3.The automated multi-agent full-process operation scheduling method of claim 2, wherein, The method further comprises: The controllable noise is input into the digital twin environment for simulation and simulation early warning data sets are generated, which include initial state vectors, controllable noise, post-disturbance state vectors, and undisturbed state vectors; The initial state vector represents the initial state vector of the digital twin environment, the post-disturbance state vector represents the state vector sequence of the digital twin environment within a fixed time window after the controllable noise is injected, and the undisturbed state vector represents the state vector sequence of the digital twin environment within a fixed time window without the controllable noise being injected; The simulation early warning sub-module is constructed based on a deep twin network, the simulation early warning data sets are input into the simulation early warning sub-module, the difference between the post-disturbance state vector and the undisturbed state vector is obtained according to an attention mechanism network, pre-training is performed in combination with a predetermined training target, training effectiveness is verified in combination with a training and verification set constructed based on a non-stationary operation data set, and training is ended only when a predetermined effectiveness condition is met. 4.The automated multi-agent full-process operation scheduling method of claim 1, wherein, The simulation early warning sub-module performs operation early warning based on the simulation early warning data sets and obtains corresponding operation early warning results, including: The simulation early warning sub-module analyzes the simulation early warning data sets and obtains the difference between the digital twin environment before and after the controllable noise is injected; The difference is mapped and converted according to a preset rule to generate a corresponding early risk early warning score; If the early risk early warning score reaches a predetermined threshold, it is determined that an anomaly occurs; If the early risk early warning score does not reach the predetermined threshold, it is determined that there is no anomaly; Based on the full-process operation resilience knowledge graph, the simulation early warning data set with the anomaly is root located using an attention mechanism to obtain an abnormal root node set; Root reasoning is performed on entity nodes included in the abnormal root node set to generate corresponding root reasoning information, and the root reasoning information set at least includes an abnormal root cause and an initial root cause elimination strategy; The abnormal root node set and the corresponding root reasoning information set are encapsulated as an operation early warning information set, and the initial root cause elimination strategy is analyzed based on a preset operation credit mechanism to generate corresponding operation incentive credits. 5.The automated multi-agent full-process operation scheduling method of claim 1, wherein, Game negotiation is performed based on the operation early warning data set to obtain a game negotiation result, and an operation control strategy is generated according to the game negotiation result, including: Operation risk verification is performed based on the operation early warning information set to obtain an operation risk verification result, the operation early warning information set is feedback optimized according to the operation risk verification result, and a corresponding operation risk resource set is generated; The operation risk resource set at least includes an operation risk root entity set, an operation risk elimination resource set, and an initial root cause elimination strategy; Multi-agent reinforcement learning and dynamic game algorithms are used to perform game negotiation in combination with the operation risk resource set to generate corresponding game negotiation results; The initial resource scheduling strategy is optimized and adjusted based on the game negotiation result to generate a corresponding operation control strategy.

6. The automated multi-agent full-process operation scheduling method according to claim 5, characterized in that, Multi-agent reinforcement learning and dynamic game algorithms are used to perform game negotiation in combination with the operation risk resource set to generate corresponding game negotiation results, including: A corresponding game agent is configured for each operation process of an operation whole process, and the operation whole process at least includes four operation processes of engineering construction, spare part management, maintenance and detection service; Each game agent is initialized based on an initial root cause elimination strategy, and each game agent performs rationality analysis on the initial root cause elimination strategy based on a scheduling cost, in combination with an operation risk root cause entity set and an operation risk elimination resource set. The initial root cause elimination strategy contains an initial operation incentive integral budget for eliminating risk roots, and the scheduling cost is represented by an actual operation incentive integral number generated according to a preset cost conversion rule, and at least includes a real loss cost and a waiting compensation cost. Each game agent performs game negotiation to generate a corresponding operation control strategy, with the minimum scheduling cost as the game target according to the result of rationality analysis.

7. The automated multi-agent full-process operation scheduling method according to claim 1, characterized in that, According to the implementation effect of the operation control strategy, operation incentive integral clearing is performed, an optimization suggestion set is generated, and the operation early warning model is feedback optimized, including: The operation incentive integral clearing is represented by a deviation value of the initial operation incentive integral budget and the actual operation incentive integral cost number of the performance agent of the operation control strategy, and the corresponding clearing result is generated based on the deviation value; If the clearing result is budget sufficient, the remaining operation incentive integral after paying the actual operation incentive integral cost is settled as the performance remuneration to each performance agent, and each performance agent pays a commission according to a preset rule and injects it into an anti-risk reserve pool; If the clearing result is budget sufficient but the cost is over budget due to the failure of the performance agent to perform, the integral compensation pool is started to compensate for the corresponding excess cost, and the integral deduction is performed on the agent that fails to perform, including excess cost deduction and bad faith integral deduction; The operation incentive integral of the integral compensation pool comes from the integral deduction of the non-performance agent; If the clearing result is that the real cost is seriously underestimated, the anti-risk reserve pool is started to compensate the cost of each performance agent, including the cost of advance payment and the compensation of reward integral; Based on the clearing result, an optimization suggestion set is generated, and the operation early warning model is feedback optimized according to the optimization suggestion set. 8.The automated multi-agent whole-process operation scheduling method of claim 1, wherein, An operation whole process is acquired, and the operation whole process is divided into a historical operation whole process data set and a real-time operation whole process data set. Based on the operation whole process data set, an operation whole process resilience knowledge graph is constructed and updated in real time, including: Multi-source data acquisition is performed on the operation scheduling whole process to acquire the operation whole process data set, and the operation whole process data set at least includes engineering construction process data, spare part management process data, maintenance and repair process data, and detection service process data; Feature extraction and semantic analysis are performed on the operation whole process data set to acquire an operation whole process entity set and a corresponding operation whole process topology association network, and each entity in the operation whole process entity set is taken as a node to construct an operation whole process resilience knowledge graph in combination with the operation whole process topology association network; Based on the graph neural network and causal reasoning, the whole-process operation resilience knowledge graph is evaluated for resilience, and each entity node is given an operation risk score; The resilience evaluation represents the operation resilience of the current node by combining the state health degree, topological importance and adjacent node risk infection resistance of the entity node itself; Wherein, the state health degree is used to quantify the current operation state of the entity node, the topological importance is used to quantify the importance of the entity node in the whole-process operation, and the adjacent node risk infection resistance is used to quantify the ability of the current node to maintain its state health when the adjacent node has an operation risk.

9. An automated multi-agent whole-process operation scheduling platform for implementing the method of any one of claims 1 to 8, characterized in that, Comprise: A data acquisition module for acquiring a whole-process operation data set and constructing and updating a whole-process operation resilience knowledge graph based on the whole-process operation data set; A model construction module for constructing an operation early warning model and pre-training the model based on historical whole-process operation data sets; An operation early warning module for inputting real-time whole-process operation data sets into the operation early warning model for operation early warning, and generating an operation early warning data set; A game decision module for game bargaining according to the operation early warning data set, obtaining a game bargaining result, and generating an operation control strategy based on the game bargaining result; A feedback adjustment module for operation incentive points clearing, and generating an optimization suggestion set based on the result of operation incentive points clearing to feedback and optimize the operation early warning model.

Citation Information

Patent Citations

  • Project progress tracking and risk prediction system and method based on improved knowledge graph and multi-view graph neural network

    CN120181781A

  • Game theory-based offshore wind power plant target unit screening, fault repairing and generating capacity optimizing method

    CN120430451A