Automatic multi-agent full-process operation scheduling platform and method
By constructing an operational early warning model using generative adversarial networks and graph neural networks, and combining multi-agent reinforcement learning and dynamic game theory algorithms, the problem of insufficient efficiency and accuracy of resource scheduling in traditional operation and maintenance models is solved, and efficient and accurate operational resource scheduling is achieved.
Patent Information
- Application Number
- CN202511769696.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-11-28
AI Technical Summary
Traditional operation and maintenance models lack the ability to anticipate subtle early risks when dealing with complex industrial operating environments, making it difficult to automatically locate the root cause and resulting in insufficient efficiency and accuracy in resource scheduling.
An operational early warning model is constructed using generative adversarial networks and graph neural networks. The model is pre-trained using a full-process operational dataset. Resource scheduling is performed through multi-agent reinforcement learning and dynamic game theory algorithms to generate operational control strategies. Operational incentive tokens are then liquidated to optimize resource scheduling.
It enables efficient and precise scheduling of operational resources, improves the efficiency and accuracy of operational scheduling, and ensures multi-resource collaborative optimization and economy under budget constraints.
Smart Images

Figure CN121212751A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of resource scheduling, in particular to an automatic multi-agent whole-process operation scheduling platform and method. BACKGROUND
[0002] In the current complex industrial operation environment, enterprises have generally established multi-dimensional operation and maintenance systems covering equipment, supply chain, manpower, etc., and have accumulated a large amount of real-time and historical data. The traditional operation and maintenance mode mostly adopts an alarm system based on fixed thresholds and a response process relying on human experience. This mode has a certain effect in dealing with known single-point faults and provides a foundation support for ensuring the stable operation of the system.
[0003] However, with the dramatic increase in system complexity and dynamics, the traditional operation and maintenance mode is facing severe challenges. On the one hand, it often lacks the ability to foresee weak early risks, and alarm information is difficult to automatically locate the root cause, resulting in ambiguous response targets. On the other hand, the resource scheduling scheme has certain problems of being divorced from actual constraints and cost-effectiveness, thereby leading to insufficient efficiency and accuracy of resource scheduling.
[0004] These limitations make the traditional technology inefficient in dealing with modern operation and maintenance scheduling, so there is an urgent need for an automatic multi-agent whole-process operation scheduling method to achieve efficient and accurate scheduling of operation resources. SUMMARY
[0005] In view of the above-mentioned problems, in combination with the first aspect of the present application, the embodiments of the present application provide an automatic multi-agent whole-process operation scheduling method, which comprises: acquiring a whole-process operation data set, the whole-process operation data set being divided into a historical whole-process operation data set and a real-time whole-process operation data set, and constructing and updating a whole-process operation resilience knowledge graph based on the whole-process operation data set; constructing an operation early warning model based on a generative adversarial network and a graph neural network, and pre-training the model in combination with the historical whole-process operation data set, the operation early warning model comprising a risk detection sub-module and a simulation early warning sub-module; inputting the real-time whole-process operation data set into the operation early warning model for operation early warning, generating an operation early warning data set, the operation early warning data set comprising an operation early warning information set and an operation incentive token; conducting a game negotiation based on the operation early warning data set, obtaining a game negotiation result, and generating an operation control strategy according to the game negotiation result; clearing the operation incentive token according to the implementation effect of the operation control strategy, generating an optimization suggestion set, and feeding back and optimizing the operation early warning model.
[0006] As a further scheme of the present application, an operation early warning model is constructed based on a generative adversarial network and a graph neural network, and the model is pre-trained in combination with a historical full-process operation dataset, the operation early warning model comprising a risk detection submodule and a simulation early warning submodule, comprising: A digital twin environment is constructed based on a full-process operation resilience knowledge graph and a historical full-process operation dataset, and the digital twin environment is updated in real time in combination with a real-time full-process operation dataset, the digital twin environment being used to simulate a full-process operation state; The historical full-process operation dataset is clustered to obtain a non-stationary operation dataset, the non-stationary operation dataset representing historical data in which an operation process imbalance occurs in the full-process operation process and at least comprising operation imbalance root cause data and operation imbalance impact data; A reinforcement learning agent is constructed, a state space corresponding to the reinforcement learning agent is defined according to the digital twin environment, an action space of the reinforcement learning agent is set as controllable noise for early risk detection, a deep neural network is used as a policy network of the reinforcement learning agent, and the policy network is pre-trained in combination with the non-stationary operation dataset, the pre-trained reinforcement learning agent being used as the risk detection submodule; The risk detection submodule is used to generate controllable noise, and the controllable noise is used to be input into the simulation early warning submodule for early operation risk identification.
[0007] As a further scheme of the present application, the method further comprises: The controllable noise is input into the digital twin environment for simulation to generate a simulation early warning dataset, the simulation early warning dataset comprising an initial state vector, the controllable noise, a perturbed state vector, and an unperturbed state vector; The initial state vector represents an initial state vector of the digital twin environment, the perturbed state vector represents a state vector sequence of the digital twin environment within a fixed time window after the controllable noise is injected, and the unperturbed state vector represents a state vector sequence of the digital twin environment within a fixed time window corresponding to no injection of the controllable noise; The simulation early warning submodule is constructed based on a deep twin network, the simulation early warning dataset is input into the simulation early warning submodule, a difference degree of the perturbed state vector and the unperturbed state vector is obtained according to an attention mechanism network, pre-training is performed in combination with a predetermined training target, training effectiveness is verified in combination with a training and verification set constructed based on the non-stationary operation dataset, and the training is ended only when a predetermined effectiveness condition is met.
[0008] As a further scheme of the present application, characterized in that a real-time full-process operation dataset is input into the operation early warning model for operation early warning to generate an operation early warning dataset, the operation early warning dataset comprising an operation early warning information set and an operation incentive token, comprising: After the real-time full-process operation data set is input into the operation early warning model, the risk detection submodule generates corresponding controllable noise based on the real-time full-process operation data set; The controllable noise is injected into the pre-constructed digital twin environment for simulation and simulation, to generate a corresponding simulation early warning data set; The simulation early warning submodule performs operation early warning based on the simulation early warning data set to obtain corresponding operation early warning results; If the operation early warning result is no exception, continue to monitor the full-process operation state; If the operation early warning result is abnormal, generate a corresponding operation early warning information set and operation incentive token; The operation early warning information set and operation incentive token are packaged as an operation early warning data set for output.
[0009] As a further scheme of the present application, the simulation early warning submodule performs operation early warning based on the simulation early warning data set to obtain corresponding operation early warning results, including: The simulation early warning submodule analyzes the simulation early warning data set to obtain the difference degree of the digital twin environment before and after the controllable noise is injected; According to a preset rule, the difference degree is mapped and transformed to generate a corresponding early risk early warning score; If the early risk early warning score reaches a predetermined threshold, it is determined that an exception has occurred; If the early risk early warning score does not reach the predetermined threshold, it is determined that there is no exception; Based on the full-process operation resilience knowledge graph, and using the attention mechanism to locate the root cause of the simulation early warning data set that appears abnormal, an abnormal root node set is obtained; The entity nodes contained in the abnormal root node set are subjected to root cause reasoning to generate corresponding root cause reasoning information, and the root cause reasoning information set at least includes abnormal root cause inducement and initial root cause inducement elimination strategy; The abnormal root node set and the corresponding root cause reasoning information set are packaged as an operation early warning information set, and the initial root cause inducement elimination strategy is analyzed based on a preset operation token mechanism to generate a corresponding operation incentive token.
[0010] As a further scheme of the present application, based on the operation early warning data set, a game negotiation is performed to obtain a game negotiation result, and an operation control strategy is generated according to the game negotiation result, including: Based on the operation early warning information set, operation risk verification is performed to obtain an operation risk verification result, and the operation early warning information set is fed back and optimized according to the operation risk verification result to generate a corresponding operation risk resource set; The operation risk resource set at least includes an operation risk root cause entity set, an operation risk elimination resource set, and an initial root cause elimination strategy; The multi-agent reinforcement learning and the dynamic game algorithm are adopted to carry out game negotiation in combination with the operation risk resource set, and a corresponding game negotiation result is generated. Based on the game negotiation result, an initial resource scheduling strategy is optimized and adjusted, and a corresponding operation control strategy is generated.
[0011] As a further scheme of the present application, the multi-agent reinforcement learning and the dynamic game algorithm are adopted to carry out game negotiation in combination with the operation risk resource set, and a corresponding game negotiation result is generated, including: A corresponding game intelligent agent is configured for all operation processes of an operation whole process, and the operation whole process at least includes four operation processes of engineering construction, spare part management, maintenance and detection service; Based on an initial root cause elimination strategy, each game intelligent agent is initialized, and based on a scheduling cost of each game intelligent agent, the initial root cause elimination strategy is analyzed for rationality in combination with the operation risk root cause entity set and the operation risk elimination resource set. The initial root cause elimination strategy contains an initial operation incentive token budget for eliminating risk root causes, and the scheduling cost is represented as an actual operation incentive token quantity generated according to a preset cost conversion rule, and at least includes a real loss cost and a waiting compensation cost. Each game intelligent agent carries out game negotiation with the goal of minimizing the scheduling cost according to the result of the rationality analysis, and generates a corresponding operation control strategy.
[0012] As a further scheme of the present application, operation incentive token clearing is carried out according to the implementation effect of the operation control strategy, an optimization suggestion set is generated, and an operation early warning model is feedback optimized, including: The operation incentive token clearing is represented as a deviation value of clearing the initial operation incentive token budget and the actual operation incentive token consumption quantity of a performance intelligent agent of the operation control strategy, and a corresponding clearing result is generated based on the deviation value; If the clearing result is budget sufficient, the operation incentive tokens remaining after the actual operation incentive token consumption are settled as performance remuneration to each performance intelligent agent, and each performance intelligent agent pays a commission to inject into an anti-risk reserve pool according to a preset rule; If the clearing result is budget sufficient but the cost is over budget due to the failure of the performance intelligent agent to perform, a token compensation pool is started to compensate for the corresponding excess cost, and a token penalty is imposed on the intelligent agent that fails to perform, and the token penalty includes an excess cost penalty and a bad faith token penalty; The operation incentive tokens of the token compensation pool come from the token penalty of the non-performing intelligent agent; If the clearing result is that the budget seriously underestimates the real cost, a risk resistance pool is started to compensate the cost of each fulfillment agent, including the cost of token and the compensation of reward token; An optimization suggestion set is generated based on the clearing result, and the operation early warning model is feedback optimized according to the optimization suggestion set.
[0013] As a further scheme of the present application, a full-process operation dataset is acquired, the full-process operation dataset is divided into a historical full-process operation dataset and a real-time full-process operation dataset, and a full-process operation resilience knowledge graph is constructed and updated in real time based on the full-process operation dataset, including: Multi-source data acquisition is performed on an operation scheduling full process to acquire a full-process operation dataset, the full-process operation dataset at least including engineering construction process data, spare part management process data, maintenance and repair process data and detection service process data; Feature extraction and semantic analysis are performed on the full-process operation dataset to acquire a full-process operation entity set and a corresponding full-process topological correlation network, and a full-process operation resilience knowledge graph is constructed by taking each entity in the full-process operation entity set as a node and synchronously combining the full-process topological correlation network; Resilience evaluation is performed on the full-process operation resilience knowledge graph based on a graph neural network and causal reasoning, and an operation risk score is given to each entity node; The resilience evaluation represents quantification of operation resilience of the current node in combination with a state health degree, topological importance and adjacent node risk infection resistance of the entity node itself; The state health degree is used to quantify the current operation state of the entity node, the topological importance is used to quantify the importance of the entity node in the full-process operation, and the adjacent node risk infection resistance is used to quantify the ability of the current node to maintain its state health when an adjacent node has an operation risk.
[0014] In another aspect, the embodiment of the present application further provides an automatic multi-agent full-process operation scheduling platform, including: A data acquisition module is configured to acquire a full-process operation dataset and construct and update a full-process operation resilience knowledge graph based on the full-process operation dataset; A model construction module is configured to construct an operation early warning model and pre-train the model in combination with a historical full-process operation dataset; An operation early warning module is configured to input a real-time full-process operation dataset into the operation early warning model to perform operation early warning and generate an operation early warning dataset; The game decision module is used for game bargaining according to the operation early warning data set, obtaining a game bargaining result, and generating an operation control strategy based on the game bargaining result; The feedback adjustment module is used for operation incentive token clearing, and generating an optimization suggestion set based on the result of operation incentive token clearing to feedback optimize the operation early warning model.
[0015] Compared with the prior art, the present application has the following beneficial effects: The full-process operation data set is obtained, the full-process operation resilience knowledge graph is constructed and updated in real time based on the full-process operation data set, the visualization of the topological relationship and the causal relationship of cross-link scheduling in the full-process operation is realized by integrating the multi-source heterogeneous data of the full-process operation and the construction of the knowledge graph, and a data basis is provided for the subsequent steps; The operation early warning model is constructed based on the generative adversarial network and the graph neural network, the model is pre-trained in combination with the historical full-process operation data set, the real-time full-process operation data set is input into the operation early warning model for operation early warning, the operation early warning data set is generated, the weak signal in the operation process is amplified through the operation early warning model, and the early warning of the progressive failure is realized; Game bargaining is carried out based on the operation early warning data set, the game bargaining result is obtained, the operation control strategy is generated according to the game bargaining result, the multi-resource collaborative optimization under the budget constraint is realized, and the economy and the executability of the operation control strategy in the actual operation scene are ensured; According to the implementation effect of the operation control strategy, the operation incentive token clearing is carried out, the optimization suggestion set is generated, and the operation early warning model is feedback optimized, and the efficient and accurate scheduling of the operation resource is realized. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a step flow chart of an automatic multi-agent full-process operation scheduling method of the present application; Figure 2 is a step flow chart of step S3 in the automatic multi-agent full-process operation scheduling method of the present application; Figure 3 is a schematic diagram of an automatic multi-agent full-process operation scheduling platform of the present application. DETAILED DESCRIPTION
[0017] The present application will be specifically described below in conjunction with the drawings of the specification, Figure 1 is a step flow chart of an automatic multi-agent full-process operation scheduling method of the present application, Figure 2 is a step flow chart of step S3 in the automatic multi-agent full-process operation scheduling method of the present application, and the automatic multi-agent full-process operation scheduling method will be described in detail below.
[0018] Step S1, obtain a full-process operation dataset, and construct and update a full-process operation resilience knowledge graph based on the full-process operation dataset.
[0019] Specifically, multi-source data is collected for the full-process operation scheduling to obtain a full-process operation dataset, which at least includes engineering construction process data, spare part management process data, maintenance and repair process data, and detection service process data. Feature extraction and semantic analysis are performed on the full-process operation dataset to obtain a full-process operation entity set and a corresponding full-process topological correlation network. A full-process operation resilience knowledge graph is constructed by taking each entity in the full-process operation entity set as a node and synchronously combining the full-process topological correlation network. Resilience evaluation is performed on the full-process operation resilience knowledge graph based on a graph neural network and causal reasoning, and an operation risk score is assigned to each entity node.
[0020] It can be understood that resilience evaluation represents quantification of the operation resilience of a current node in combination with the state health degree, topological importance, and adjacent node risk infection resistance of the entity node itself; the state health degree is used to quantify the current operation state of the entity node; the topological importance is used to quantify the importance of the entity node in the full-process operation; and the adjacent node risk infection resistance is used to quantify the ability of the current node to maintain its state health when an adjacent node has an operation risk.
[0021] In one possible embodiment, real-time data and historical data are continuously collected from four core business links of engineering construction, spare part management, maintenance and repair, and detection service. Data preprocessing is performed on the data collected from all the above business links to generate corresponding full-process operation data vectors. Taking installation and debugging of a pump device as an example, pump device installation and debugging records from the engineering construction process are synchronously collected, which at least include centering accuracy and initial vibration value. The pump device corresponding spare bearing inventory quantity and procurement and transportation information from the spare part management process are collected. Real-time vibration sensor readings and maintenance work orders of the pump device from the maintenance and repair process are collected. The latest spectral analysis results of lubricating oil of the pump device from the detection service process are collected. The original data are cleaned, denoised, and missing value filled, and the pump number P-101, spare bearing number B-203, and maintenance work order E-302 in the pump device installation and debugging records are associated to generate a full-process operation data vector of the operation resources required by the pump device with the installation and debugging pump number P-101.
[0022] The graph neural network technology is adopted to construct a whole-process operation resilience knowledge graph in combination with a whole-process operation data set composed of whole-process operation data vectors. Graph nodes of the whole-process operation resilience knowledge graph represent whole-process operation entities, such as devices, personnel, spare parts and tasks, and edges represent relationships between entities, such as dependence, supply, execution and location. For example, a reaction kettle R-201 is an entity node, which is connected to a heat exchanger E-15 through an edge representing process dependence, to a mechanical seal S-07 in a spare parts warehouse through an edge representing spare part supply, and to a maintenance team T-03 through an edge representing maintenance. The edge connections of the reaction kettle show that the health status of the reaction kettle R-201 directly affects the heat exchanger E-15, and its maintenance depends on the availability of the mechanical seal S-07 and the maintenance team T-03.
[0023] An operation risk score is calculated for each node in the graph, which quantifies the vulnerability of the node. The operation risk score is generated by calculating a comprehensive assessment of the node's own health status, its topological importance in the network and the potential impact of its neighbor nodes on it. Specifically, the quantitative assessment of the own health status is represented as setting corresponding monitoring indicators for each entity node, such as vibration, temperature, inventory level, etc., and setting a normal range [L, U] for each monitoring indicator. For numerical monitoring indicators, the indicator abnormality is set as max(0, (actual value-U) / (U-L)). For Boolean indicators, the indicator abnormality is directly set to 0 or 1, such as the maintenance team being online, which has an indicator abnormality of 1. Finally, the indicator abnormality of each node corresponding to the monitoring indicator is combined by a pre-set weight ratio to generate a node health score. The topological importance is represented as a weighted fusion of the betweenness centrality, closeness centrality and eigenvector centrality of the node to generate a corresponding topological importance score. The potential impact of the neighbor nodes is quantified as calculating the influence coefficient of the state change of neighbor node j on the current node i using Granger causality test combined with historical data. If there is no historical data, the influence intensity is set by experts according to business logic, such as the influence intensity of spare parts inventory on devices being higher than that of tools, and the operation risk score is generated by a pre-set weight ratio of the node health score, the topological importance score and the influence coefficient. It should be noted that the pre-set weight ratio used in the calculation of the operation risk score needs to be specifically allocated according to the importance of the data involved in the actual operation process.
[0024] Taking the operation risk score of centrifugal pump P-101 as an example, assuming that the corresponding monitoring indicators are vibration and temperature, through analysis of the historical data of centrifugal pump P-101, it is found that vibration has a greater impact on centrifugal pump P-101 than temperature, so the weight ratio can be set to 6:4, and the vibration anomaly degree is 0.8, the temperature anomaly degree is 0.4, and the corresponding node health score is 0.64; the betweenness centrality is 0.9, the corresponding weight is 0.5, the closeness centrality is 0.7, the corresponding weight is 0.3, and the eigenvector centrality is 0.8, the corresponding weight is 0.2, and the corresponding topological importance score of centrifugal pump P-101 is 0.82; the corresponding neighbor nodes of centrifugal pump P-101 include spare part node S-51 and engineer node T-13, and the node health scores of the two are 0.8 and 0.3 respectively, assuming that through analysis of historical data, it is found that whenever the inventory of spare parts S-51 is lower than the preset safety line, the probability of failure of centrifugal pump P-101 within the next week shows a significant upward trend, through statistical analysis and normalization, it is determined that the influence intensity of spare parts S-51 is 0.9, and by the same token, through analysis of historical data, it is found that the work load of the responsible engineer T-13 has a moderate influence on the fault repair time of centrifugal pump P-101 but is not a decisive factor, and the influence intensity corresponding thereto is obtained by using Granger causality test and fuzzy analysis, which is 0.6, and the inverse ratio of the shortest path length between the neighbor nodes and the current node is taken as the attenuation factor to simulate the attenuation effect of the influence of the neighbor nodes in the graph, assuming that the shortest path lengths of spare part node S-51 and engineer node T-13 to centrifugal pump P-101 are 1 and 2 respectively, and the attenuation factors corresponding thereto are 1 and 0.5 respectively, combining the influence intensity, the node health score corresponding to the neighbor nodes themselves and the attenuation factor, the influence coefficient of spare part node S-51 and engineer node T-13 on centrifugal pump P-101 is 0.9*0.8*1+0.6*0.3*0.5=0.81; through analysis of the historical whole-process operation data set, it is found that the node health score and the topological importance score have a greater contribution to the final operation risk than the relative influence coefficient, so the weight ratio corresponding to centrifugal pump P-101 node is node health score:topological importance score:influence coefficient=4:4:2, and the operation risk score corresponding to the node is 0.64*0.4+0.82*0.4+0.81*0.2=0.746.
[0025] Step S2, an operation early warning model is constructed based on a generative adversarial network and a graph neural network, and model pre-training is performed in combination with a historical whole-process operation data set.
[0026] It can be understood that, based on the full-process operation resilience knowledge graph and the historical full-process operation data set, a digital twin environment is constructed, and the digital twin environment is updated in real time combined with real-time full-process operation data set, the digital twin environment is used to simulate the full-process operation state, and a prediction model is included in the digital twin environment, such as using a time series convolution network or a long short-term memory network as a prediction model, which is pre-trained based on a large amount of health data in a normal state, and the prediction model can predict the state vector in the future normal state according to the current input state vector.
[0027] The historical full-process operation data set is clustered to obtain a non-stationary operation data set, which represents historical data of operation process imbalance in the full-process operation, and at least includes operation imbalance root cause data and operation imbalance influence data, for example, extracting all normal running period data of centrifugal pump P-101 in the past year as health samples, and extracting historical data of faults occurring in the past year based on clustering algorithm, assuming that centrifugal pump P-101 occurred a bearing fault half a year ago, then extracting its data sequence 48 hours before the fault occurs, which is marked as early risk sample, and obtaining influence data caused by the fault.
[0028] The influence data is represented as a series of changes and loss data caused in the operation system composed of engineering construction, spare parts management, maintenance and repair, detection service after a fault occurs, for example, in the engineering construction link, since centrifugal pump P-101 is a key equipment, its shutdown leads to the complete shutdown of the production line for 2 hours, and the production loss of 2 hours can be used as the influence data corresponding to the engineering construction link; In the maintenance and repair link and the spare parts management link, the emergency maintenance cost generated by this maintenance, such as overtime pay, emergency spare parts allocation and transportation cost, outsourcing service cost are recorded as the influence data corresponding to the maintenance and repair link, and the change of spare parts inventory is recorded as the influence data of the spare parts management link, such as consuming a key spare bearing, which leads to the breakthrough of the safety inventory of this type of spare parts, and the emergency procurement needs to be started, and the change of inventory level from safety to emergency is recorded, which will be used as the influence data of this link; In the detection service link, assuming that due to the sudden shutdown of P-101, the upstream equipment is damaged, the abnormal data of these associated equipment during and after the fault is recorded, which will be used as the influence data corresponding to the detection service link.
[0029] In this embodiment, step S2 includes: Step S2-1, constructing and training a risk detection sub-module.
[0030] Specifically, a reinforcement learning agent is constructed, a state space corresponding to the digital twin environment is defined, controllable noise is injected for early risk detection, a deep neural network is used as the policy network, the policy network is pre-trained in combination with the non-stationary operation data set, the pre-trained reinforcement learning agent is used as a risk detection sub-module, and the controllable noise is used as input to the simulation early warning sub-module for early operation risk identification.
[0031] In one possible embodiment, the state space of the reinforcement learning agent is defined in combination with the current state of each node in the full-process operation resilience knowledge graph in the pre-digital twin environment, such as the vibration, temperature, and load of the device; the strength, frequency, and target device of the controllable noise in the action space are set; the policy network adopts a double-network structure of Actor network and Critic network, wherein the Actor network is responsible for generating the action policy, taking the state vector corresponding to the state space as input, and outputting the mean and variance of the generated action policy, which is used to define the specific characteristics of noise injection, for example, for the noise intensity dimension, the Actor network outputs a mean of 0.3 and a variance of 0.1, meaning that the optimal noise intensity is probably around 0.3 under the current system state, and the noise intensity is adjusted in combination with the variance, i.e., noise intensity is sampled according to the normal distribution N(0.3, 0.1²); the Critic network is responsible for evaluating the value of the action policy, taking the state vector of the state space and the action vector of the action space as input, and finally outputting a state-action value estimate, which is used to guide the policy optimization of the Actor network; a multi-objective weighted combination method is used to construct the reward function, which at least includes a detection reward term, a safety penalty term, an efficiency reward term, and an exploration entropy reward term, the detection reward term is used to reward the effective triggering of operation process abnormalities, the safety penalty term is used to punish overactive behaviors that may cause system collapse, the efficiency reward term is used to encourage achieving the detection goal with the minimum energy, and the exploration entropy reward term is used to maintain the diversity of the policy.
[0032] The pre-training process is performed in the digital twin environment. Taking a specific training instance corresponding to the centrifugal pump P-101 as an example, the policy network of the agent generates a specific action policy based on the state vector corresponding to a certain moment. Assuming that the parameters of the action policy correspond to [type: {sine sweep, probability: 0.7}, center frequency: 237 Hz, bandwidth: 45 Hz, gain: 0.23, duration: 2.1 seconds], after executing the excitation in the digital twin environment, the corresponding comprehensive reward value is 0.156 according to the calculation combined with the reward function. The policy network parameters are updated by the proximal policy optimization algorithm. The entire training process optimizes the parameters of the policy network in an iterative training manner until the preset iteration condition is reached, such as when the average comprehensive reward value of the last 100 rounds is greater than 0.5 and the detection success rate on the validation set exceeds 85%. After training is completed, the risk detection sub-module can generate an optimal detection strategy corresponding to different full-process operation states, that is, the most suitable controllable noise. The controllable noise can be used as a detection signal to effectively and safely test the stability of the current operation process state.
[0033] It can be understood that the detection reward item can be represented as Tanh((actual resonance intensity-resonance intensity threshold) / resonance intensity threshold), and the resonance intensity is calculated by the simulation early warning sub-module to quantify the sensitivity to a specific noise excitation. The weight of the item is 1.0 by default. For example, after the agent injects controllable noise, the actual resonance intensity is obtained as 0.83 through analysis by the simulation early warning sub-module. Assuming that the resonance intensity threshold is 0.7, the corresponding detection reward item is tanh((0.83-0.7) / 0.7)≈0.18, which means that the agent obtains a positive reward of +0.18 because the obvious resonance is successfully triggered. The safety penalty item can be represented as -max(0,(X p -X s ) / X s ) 2 , where X p represents the monitoring peak value of the key monitoring indicator after the noise is injected, and X s represents the safety threshold of the key monitoring indicator, and the weight of the item is 1.5 by default. For example, assuming that the key monitoring indicator of the centrifugal pump P-101 is vibration, the vibration peak value after injecting controllable noise A1 is 12.5 mm / s, and the safety threshold is 15 mm / s, then the corresponding safety penalty item is -max(0,(12.5-15) / 15) 2 =0, which means that the current operation is within the safety range and there is no penalty. The vibration peak value after injecting controllable noise A2 is 16.5 mm / s, and the corresponding safety penalty item is -max(0,(16.5-15) / 15) 2= -0.01, meaning that the current operation is out of the safety range, and a penalty needs to be applied; the efficiency reward term can be expressed as -λ*ln(1+E), where λ is a preset control coefficient, used to ensure that this term does not dominate the reward and is set to 1.0 by default, E is the total energy of the injected noise, which is calculated by default according to the amplitude and duration of the injected noise, and the weight of the efficiency reward term is 0.2 by default, for example, if the total energy of the noise injected by the agent is 0.35, then the corresponding efficiency reward term is -0.1*ln(1+0.35)≈-0.03; the exploration entropy reward term can be expressed as η*H, where η is a fine-tuning exploration coefficient and is set to 0.01 by default, H is the entropy of the action policy in a given state, the higher the entropy value, the more uniform the probability distribution of the policy, that is, the possibility of selecting different actions is closer, and the exploration is stronger, and the weight corresponding to this reward term is 1.0, for example, in a certain state, the agent's policy network outputs the mean and variance of 6 action strategies, assuming that the policy entropy H calculated from these parameters is 2.1, then the corresponding exploration entropy reward term is 0.01*2.1=0.021, and the comprehensive reward calculated by combining the above content is 1.0*(0.18-0.03+0.021)+1.5*(-0.01)=0.156.
[0034] Step S2-2, constructing and training the simulation early warning sub-module.
[0035] Specifically, the controllable noise is input into the digital twin environment for simulation to generate a simulation early warning data set, which includes an initial state vector, controllable noise, a disturbed state vector, and an undisturbed state vector; an analog early warning sub-module is constructed based on a deep twin network, the simulation early warning data set is input into the analog early warning sub-module, the difference between the disturbed state vector and the undisturbed state vector is obtained according to the attention mechanism network, pre-training is performed in combination with a predetermined training target, training effectiveness verification is performed in combination with a training verification set constructed based on a non-stationary operation data set, and training is ended only when a predetermined effectiveness condition is met.
[0036] It can be understood that the initial state vector represents the initial state vector of the digital twin environment, the disturbed state vector represents the state vector sequence of the digital twin environment within a fixed time window after injecting controllable noise, and the undisturbed state vector represents the state vector sequence of the digital twin environment within a fixed time window without injecting controllable noise.
[0037] It can be understood that the simulation early warning submodule is constructed based on a deep twin network architecture, the deep twin network architecture is composed of two subnetworks sharing weights, each subnetwork is constructed by using a time neural network embedded with an attention mechanism, one branch is used to receive the disturbed state vector output by the digital twin environment after injecting controllable noise disturbance, the other branch is used to receive the normal state vector sequence predicted by the prediction model embedded in the digital twin environment under the same initial condition, i.e. the undisturbed state vector; the attention mechanism is used to guide the simulation early warning submodule to focus on the key system variables and key time points with the most significant disturbance response, so as to more accurately capture abnormal patterns.
[0038] In a possible embodiment, the pre-training process takes the controllable noise generated by the risk detection submodule as the disturbance source of the digital twin environment, obtains the corresponding simulation early warning data set, and pre-trains in combination with the preset training target. Specifically, taking the centrifugal pump P-101 as an example, assuming that the corresponding undisturbed state vector sequence in a training process is {vibration: [10.0, 10.1, 10.2] mm / s, temperature: [65.0, 65.1, 65.2] °C}, and the corresponding disturbed state vector sequences after injecting controllable noises A1 and A2 are A1 {vibration: [10.1, 10.8, 12.5] mm / s, temperature: [65.1, 65.5, 66.5] °C}, and controllable noise A2 {vibration: [10.0, 10.2, 10.3] mm / s, temperature: [65.0, 65.2, 65.3] °C}, the undisturbed state vector sequence and the disturbed state vector sequence are extracted based on the pre-trained gated recurrent unit, wherein the pre-training data set of the gated recurrent unit is the historical full-process operation data set, the feature vector corresponding to the undisturbed state vector sequence is taken as the reference, the reference feature vector [1.0, 1.0] is generated, the feature vectors corresponding to A1 and A2 are [2.8, 3.2] and [1.1, 1.0] respectively, wherein the first dimension of the feature vector represents the degree of amplification of the injected noise energy, which is used to quantify the instability of the current state, and the second dimension represents the duration and coherence of the abnormal response, which is used to quantify the difficulty of recovery after deviating from the normal state (here, only the conversion from the original state vector to the corresponding feature vector is exemplified, and the specific determination needs to be combined with the actual situation).
[0039] The Euclidean distances of the eigenvectors corresponding to A1 and A2 from the reference eigenvector are calculated, and the Euclidean distances are taken as the difference degrees, so the difference degrees corresponding to A1 and A2 are 2.84 and 0.10 respectively; the corresponding resonance intensity is generated by using the resonance mapping function, the value of the resonance intensity directly quantifies the sensitivity of the controllable noise excitation this time, the higher the value, the greater the deviation of the actual response from the normal reference, that is, the smaller the disturbance is amplified more significantly, and the higher the risk of the current operation process being in a dynamic unstable critical state, the resonance mapping function can be expressed as 2 / [1+e^(-k*(D-D0))]-1, wherein k is a curvature factor and the default value is 2, used to control the steepness of the mapping, D0 is a reference difference degree and the default value is 1.5, so the resonance intensities corresponding to A1 and A2 are 0.87 and 0.01 respectively, assuming that the predetermined reference resonance intensity is 0.1, since the difference degree and the resonance intensity corresponding to A1 both exceed the corresponding reference values, it is defined as effective disturbance noise, and since the controllable noise corresponding to A2 cannot cause a dramatic fluctuation in the operation state within the controllable range, it is defined as ineffective disturbance noise, the difference degree of the effective disturbance noise is maximized and the difference degree of the ineffective disturbance noise is minimized as the training target, the contrast loss function is used for iterative training, and the training efficiency is verified by combining the training and verification set constructed based on the non-stationary operation data set, and the training is ended only when the predetermined efficiency condition is met.
[0040] It can be understood that the predetermined efficiency condition includes: the fluctuation range of the contrast loss function value of the training and verification set is less than the preset fluctuation threshold for 3 consecutive times, such as the contrast loss function values of the last 3 times are 0.125, 0.124 and 0.126, which are lower than the preset fluctuation threshold 0.03; the classification accuracy of the effective disturbance noise and the ineffective disturbance noise in the training and verification set continuously reaches more than 95%; if the contrast loss function value of the training and verification set does not refresh the optimal loss record for 10 consecutive times, the early stopping mechanism is triggered to prevent overfitting, for example, assuming that the loss value corresponding to the optimal loss record is 0.42, but the loss value corresponding to the continuous training for 10 times does not fall below 0.42, then the early stopping mechanism is started to prevent overfitting.
[0041] Step S3, inputting the real-time full-process operation data set into the operation early warning model to perform operation early warning, and generating an operation early warning data set.
[0042] Specifically, after the real-time full-process operation data set is input into the operation early warning model, the risk detection submodule generates corresponding controllable noise based on the real-time full-process operation data set; the controllable noise is injected into the pre-constructed digital twin environment for simulation and simulation to generate a corresponding simulation early warning data set; the simulation early warning submodule performs operation early warning based on the simulation early warning data set to obtain a corresponding operation early warning result; if the operation early warning result is normal, the full-process operation state is continuously monitored; if the operation early warning result is abnormal, a corresponding operation early warning information set and operation incentive token are generated; the operation early warning information set and operation incentive token are encapsulated as an operation early warning data set for output.
[0043] It can be understood that the simulation early warning submodule analyzes the simulation early warning data set to obtain the difference degree of the digital twin environment before and after the controllable noise is injected; the difference degree is mapped and transformed according to a preset rule to generate a corresponding early risk early warning score; if the early risk early warning score reaches a predetermined threshold, it is determined that an abnormality occurs; if the early risk early warning score does not reach the predetermined threshold, it is determined that there is no abnormality; based on the full-process operation resilience knowledge graph, and using the attention mechanism, the simulation early warning data set with an abnormality is subjected to root cause positioning to obtain an abnormal root node set; the entity nodes contained in the abnormal root node set are subjected to root cause reasoning to generate corresponding root cause reasoning information, and the root cause reasoning information set at least includes an abnormal root cause and an initial root cause elimination strategy; the abnormal root node set and the corresponding root cause reasoning information set are encapsulated as an operation early warning information set, and the initial root cause elimination strategy is analyzed based on a preset operation token mechanism to generate a corresponding operation incentive token.
[0044] In one possible embodiment, if the resonance intensity does not exceed the predetermined threshold, it is determined that the current state is stable and resource scheduling is not required; if the resonance intensity exceeds the predetermined threshold, the gradient of the difference between the controllable noise corresponding to the feature of each input node and the resonance intensity of the final output is obtained based on the graph neural network inside the operation early warning model, so as to quantify the contribution value of each entity node to the current resonance intensity. Taking centrifugal pump P-101 as an example, its vibration data, spare bearing inventory and the load of the engineer responsible for this scheduling task constitute the entity nodes of this operation scheduling. The operation early warning model calculates the resonance intensity while processing the data corresponding to all entity nodes by using its own graph neural network. The self-attention mechanism inside the operation early warning model evaluates the influence relationship between each entity node. It is assumed that the attention weight of the spare bearing inventory node to the final output result is as high as 0.7, while the attention weight of the engineer load node is only 0.1. At the same time, it is found through gradient calculation that a slight change in the vibration data of the centrifugal pump will cause the resonance intensity value of the final output to fluctuate dramatically, and the gradient value is much higher than that of other nodes. Therefore, it is determined that the spare bearing inventory and the centrifugal pump vibration are the main contributors to this resonance. The node with the highest contribution degree is identified as the primary excitation point of the resonance. Then, starting from the primary excitation point, probabilistic backward reasoning is performed along the corresponding causal edge on the whole-process operation resilience knowledge graph, and the posterior probability of different root causes leading to the current resonance phenomenon is obtained by combining the probabilistic graph model constructed based on historical fault data, so as to determine the most possible fault root cause and its influence path, and encapsulate the fault root cause and its influence path as an abnormal root cause inducement for output.
[0045] The abnormal root cause inducement is matched with the pre-constructed historical case library and pre-plan library. For example, in the form of {bearing wear, resonance intensity: 0.74}, the pre-plan matching is performed in the historical case library and the pre-plan library. According to the matched pre-plan, the current operation resource pool is searched to find a candidate resource set capable of executing the task, such as a list of qualified engineers, inventory information of required spare parts, and available state of special tools. The corresponding optimal resource combination is selected from the candidate resource set by considering the associated contents such as cost efficiency, geographical location and current load, and the approximate task loss is estimated, so as to generate an initial root cause inducement elimination strategy, for example, assigning senior engineer Zhang San to use spare part B-202 for part replacement, which is expected to consume 4 hours of working hours.
[0046] The resonance intensity value is mapped to the corresponding risk level by using a fuzzy mapping algorithm. For example, when the resonance intensity value is less than 0.3, it is considered as low risk; when the resonance intensity value is in the range of [0.3, 0.6), it is considered as medium risk; when the resonance intensity value is in the range of [0.6, 0.8), it is considered as high risk; and when the resonance intensity value is greater than 0.8, it is considered as urgent risk. Each risk level corresponds to a basic risk coefficient, such as 1.0 for low risk, 1.2 for medium risk, 1.5 for high risk, and 2.0 for urgent risk. The influence path of the current predicted risk is obtained, that is, the number of associated devices that may be affected is analyzed according to the whole-process operation resilience knowledge graph, and the corresponding risk influence coefficient is generated. For example, the risk influence coefficient of affecting 1 device is set to 1.0, the risk influence coefficient of affecting a subsystem is set to 1.3, and the risk influence coefficient of affecting the whole plant system is set to 2.0. Finally, the corresponding comprehensive risk coefficient is generated according to the comprehensive risk coefficient = basic risk coefficient * risk influence coefficient.
[0047] All resources required for analyzing the initial root cause elimination strategy are analyzed to generate a detailed resource list, which includes human resources such as the skill level of engineers and required working hours, material resources such as the model and quantity of spare parts, tool resources such as the occupation time of special detection equipment, and collaboration resources such as the working hours that require cooperation with other departments. Each resource in the resource list is priced according to its historical average cost or preset standard cost to generate a corresponding benchmark cost. For example, the cost of a senior engineer per hour is 50 tokens, and the cost of a specific type of bearing is 200 tokens.
[0048] The initial token budget is obtained according to the calculation method of initial token budget = benchmark cost * comprehensive risk coefficient * emergency scheduling premium * resource scarcity premium. The emergency scheduling premium is represented by an additional premium coefficient, such as 1.2 times, if the task needs to be responded to in non-working hours or extremely short time, and the default value is 1.0 in normal working hours or normal response state. The resource scarcity premium is represented by another additional premium coefficient if the required resource is currently in short supply or in high demand, and the default value is 1.0 if there is no such situation.
[0049] It should be noted that, in order to prevent the initial token budget from being excessively inflated, the calculated initial token budget is compared with the actual cost of historical similar tasks, and an upper limit is set, such as not exceeding 150% of the highest historical cost. Finally, the calculated initial token budget is rounded or smoothed to generate the final initial operation incentive token budget.
[0050] Step S4, based on the operation warning data set, game negotiation is carried out to obtain the game negotiation result, and an operation control strategy is generated according to the game negotiation result.
[0051] Specifically, based on the operation early warning information set, operation risk verification is performed to obtain an operation risk verification result, the operation early warning information set is fed back and optimized according to the operation risk verification result, and a corresponding operation risk resource set is generated, wherein the operation risk resource set at least includes an operation risk root cause entity set, an operation risk elimination resource set, and an initial root cause elimination strategy; a multi-agent reinforcement learning and dynamic game algorithm is used to combine the operation risk resource set to generate a corresponding game negotiation result; and the initial resource scheduling strategy is optimized and adjusted based on the game negotiation result to generate a corresponding operation control strategy.
[0052] It can be understood that corresponding game agents are configured for all operation processes of an operation whole process, the operation whole process at least includes four operation processes of engineering construction, spare part management, maintenance and detection service; each game agent is initialized based on an initial root cause elimination strategy, each game agent performs rationality analysis on the initial root cause elimination strategy based on its own scheduling cost, in combination with the operation risk root cause entity set and the operation risk elimination resource set, wherein the initial root cause elimination strategy contains an initial operation incentive token budget for eliminating risk root causes, the scheduling cost is represented as an actual operation incentive token quantity generated according to a preset cost conversion rule, and at least includes a real loss cost and a waiting compensation cost; each game agent performs game negotiation according to the result of rationality analysis, with the minimum scheduling cost as the game target to generate a corresponding operation control strategy.
[0053] In a possible embodiment, corresponding agents are created for the operation resources involved according to the initial root cause elimination strategy, and each agent is set with a corresponding initial offer based on its corresponding benchmark cost; a multi-round negotiation process is started, each agent adjusts the initial offer according to its own state such as work load, resource scarcity, and in combination with the received offers of other agents, generates a negotiation offer by using a negotiation strategy based on game theory; it is checked whether the comprehensive negotiation offer exceeds the initial operation incentive token budget, and the scheme exceeding the budget is optimized; when the offer converges or reaches the maximum round, the game stops, and the optimized scheduling strategy and the corresponding offer reward are output, and the scheduling strategy and the corresponding offer reward are packaged as an operation control strategy for output.
[0054] For example, the initial root cause elimination strategy is to assign senior engineer Zhang San to replace the part of centrifugal pump P-101 with spare part B-202. Based on the initial root cause elimination strategy and the current state of each resource, the engineer agent is initialized: {Zhang San, benchmark cost 50 tokens, initial asking price 70 tokens due to high current load}, and the spare part agent is initialized: {spare part B-202, benchmark cost 30 tokens, initial asking price 50 tokens due to inventory shortage}. The total initial offer is 70+50=120 tokens, which exceeds the budget of 100 tokens. The agents need to adjust the offer, so multiple rounds of negotiation are launched. After multiple rounds of negotiation, the engineer agent assesses that if the price is not reduced, Zhang San may lose the task, but considering that Zhang San is currently under high load, he is willing to reduce the offer to 65 tokens. The spare part agent finds that a replacement spare part B-203 with a cost of only 25 tokens can replace spare part B-202, but the performance of spare part B-203 is slightly inferior to that of spare part B-202, and the asking price is reduced to 45 tokens. The total offer is 110 tokens, which still exceeds the budget. At this time, the global coordination and optimization mechanism is started, and a proposal is made to use spare part B-203. Zhang San is asked whether he accepts the collaboration. Zhang San assesses that using spare part B-203 will increase the difficulty and risk of maintenance, and requires an additional 10 tokens as risk compensation. The new offer is 65+10+25=100 tokens, which meets the budget. The final operation scheduling strategy is output as {engineer Zhang San: 75 tokens, spare part B-203: 25 tokens}.
[0055] Step S5, according to the implementation effect of the operation regulation strategy, the operation incentive token clearing is carried out, the optimization suggestion set is generated, and the operation early warning model is feedback optimized.
[0056] Specifically, the initial operation incentive token budget is cleared with the deviation value of the actual operation incentive token consumption amount of the performance agent of the operation regulation strategy, and the corresponding clearing result is generated based on the deviation value.
[0057] If the clearing result is that the budget is sufficient, the remaining operation incentive tokens after the actual operation incentive token consumption are paid as performance remuneration settlement to each performance agent, and each performance agent pays the commission into the anti-risk reserve pool according to the preset rules; If the clearing result is that the budget is sufficient but the cost is over budget due to the failure of the performance agent to perform, the token compensation pool is started to compensate for the excess cost, and the token deduction is performed on the agent that fails to perform, the token deduction includes excess cost deduction and bad faith token deduction, wherein the operation incentive tokens of the token compensation pool are derived from the token deduction of the agent that fails to perform; If the clearing result is that the budget is seriously underestimated, the anti-risk reserve pool is started to compensate the cost of each performance agent, and the cost compensation includes the advance token cost and the reward token compensation; generate an optimization suggestion set based on the clearing result, and perform feedback optimization on the operation early warning model according to the optimization suggestion set.
[0058] In a possible embodiment, the resonance intensity output of the operation early warning model corresponding to each operation scheduling, the final scheduling strategy after the game, the token cost of actual resource consumption, and the actual execution effect of the strategy, such as the real time consumption of the maintenance task and the recovery of the key indicators of the equipment after maintenance, are recorded; all the above data are time-stamped and aligned to form a complete causal chain record, for example, for the operation event corresponding to the centrifugal pump P-101, the resonance intensity is 0.74, the initial operation incentive token budget is 100 tokens, the final scheme after the game is {engineer Zhang San: 75 tokens, spare part B-203: 25 tokens}, and the execution result data record is [actual maintenance time consumption: {5 hours, 4 hours more than the expected maintenance time consumption}, equipment effect after maintenance: {vibration value is reduced from 12.5 mm / s to 9.8 mm / s, although it is not completely restored to the best 8.0 mm / s, but the risk has been eliminated}].
[0059] Based on the causal chain record, the operation early warning model and the game mechanism are bidirectionally feedback optimized, for the operation early warning model, the execution result data is used as the ultimate label of the prediction accuracy, that is, if the maintenance effect is poor or the cost is much higher than expected after the high resonance intensity warning, the sample will be labeled as the prediction deviation, and the corresponding incremental training set is generated to fine-tune the operation early warning model, so that the future prediction can be more in line with the actual resource consumption and repair difficulty, for example, the operation event corresponding to the centrifugal pump P-101 is judged as partially effective but not optimal, and the case is used to fine-tune the operation early warning model, so that the operation early warning model can more accurately evaluate the resonance intensity and potential risk when facing similar patterns in the future.
[0060] For the game mechanism, the difference between the final transaction price and the initial asking price, the benchmark cost, and the bargaining behavior of the agent of each operation scheduling are analyzed, a bargaining training set is generated, the bargaining training set is used to optimize the bargaining strategy of the agent, and the benchmark cost of each resource in the preset resource cost library is dynamically updated, for example, engineer Zhang San can always transact at a price 50 tokens higher than his benchmark cost in many games, at this time, his benchmark cost will be gradually increased to 60 or even 65 tokens, so that the initial budget allocation in the future is more accurate, at the same time, the agent corresponding to Zhang San will also learn from the successful bargaining experience and strengthen his asking price strategy when resources are scarce.
[0061] Figure 3 FIG. 1 is a schematic diagram of an automatic multi-agent full-process operation scheduling platform.
[0062] Specifically, an automatic multi-agent full-process operation scheduling platform comprises: The data acquisition module is configured to acquire a full-process operation dataset and construct and update a full-process operation resilience knowledge graph based on the full-process operation dataset.
[0063] The model construction module is configured to construct an operation early warning model and pre-train the model in combination with a historical full-process operation dataset.
[0064] The operation early warning module is configured to input a real-time full-process operation dataset into the operation early warning model to perform operation early warning and generate an operation early warning dataset.
[0065] The game decision module is configured to perform game negotiation based on the operation early warning dataset, acquire a game negotiation result, and generate an operation regulation strategy based on the game negotiation result.
[0066] The feedback adjustment module is configured to perform operation incentive token clearing and generate an optimization suggestion set based on a result of the operation incentive token clearing to perform feedback optimization on the operation early warning model.
[0067] The specific use mode and role of the embodiment are described as follows: First, a full-process operation dataset is acquired, and a full-process operation resilience knowledge graph is constructed and updated in real time based on the full-process operation dataset. By integrating full-process operation multi-source heterogeneous data and synchronously constructing a knowledge graph, visualization of corresponding topological relationships during cross-link scheduling of the full-process operation is achieved, thereby providing a data basis for subsequent steps. Next, an operation early warning model is constructed based on a generative adversarial network and a graph neural network, and the model is pre-trained in combination with a historical full-process operation dataset. A real-time full-process operation dataset is input into the operation early warning model to perform operation early warning and generate an operation early warning dataset. By amplifying weak signals in the operation process through the operation early warning model, early warning of gradual faults is achieved. Then, game negotiation is performed based on the operation early warning dataset, a game negotiation result is acquired, and an operation regulation strategy is generated based on the game negotiation result. Multi-resource collaborative optimization under the condition of budget constraints is achieved, thereby ensuring the economy and executability of the operation regulation strategy in actual operation scenarios. Finally, operation incentive token clearing is performed according to an implementation effect of the operation regulation strategy, an optimization suggestion set is generated, and feedback optimization is performed on the operation early warning model. Efficient and accurate scheduling of operation resources is achieved.
[0068] An electronic device includes: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to the first embodiment of the present application.
[0069] The various components of the electronic device are described in detail as follows: The processor is the control center of the electronic device, and can be one processor or a plurality of processing elements. For example, the processor is one or more central processing units (CPUs), application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the first embodiment of the present application, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).
[0070] The processor can perform various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0071] The memory is used to store software programs for implementing the scheme of the present application, and is controlled by the processor to perform the implementation. The specific implementation manner can refer to the above method embodiments, and will not be described here.
[0072] The memory can be a real-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only (CD-ROM), or other optical disk storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and capable of being accessed by a computer, but is not limited thereto. The memory can be integrated with the processor or exist independently and coupled with the processor through the interface circuit of the electronic device, and the present embodiment is not limited in this regard.
[0073] The above-described embodiments can be implemented in whole or in part by software, hardware (e.g., circuitry), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a limited (e.g., infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid state disk.
[0074] It should be understood that the term "and / or" herein merely describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B, which can represent three cases of A alone, A and B together, and B alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it, but it can also represent an "and / or" relationship. The specific meaning can be understood according to the context before and after it.
[0075] It should be understood that in the embodiments of the present application, the size of the sequence number of each process does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0076] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. An automated multi-agent full-process operation scheduling method, characterized in that, It includes the following steps: Obtain the full-process operation dataset, which is divided into historical full-process operation dataset and real-time full-process operation dataset. Construct and update the full-process operation resilience knowledge graph in real time based on the full-process operation dataset. An operational early warning model is constructed based on generative adversarial networks and graph neural networks, and the model is pre-trained using historical full-process operational datasets. The operational early warning model includes a risk detection submodule and a simulation early warning submodule. The real-time full-process operation dataset is input into the operation early warning model to generate an operation early warning dataset, which includes a set of operation early warning information and operation incentive tokens. Based on the aforementioned operational early warning dataset, a game negotiation is conducted to obtain the negotiation results, and an operational control strategy is generated based on the negotiation results. Based on the implementation effect of the operational control strategy, operational incentive tokens are liquidated, an optimization suggestion set is generated, and feedback optimization is performed on the operational early warning model.
2. The automated multi-agent full-process operation scheduling method according to claim 1, characterized in that, An operational early warning model is constructed based on generative adversarial networks and graph neural networks, and pre-trained using historical full-process operational datasets. The operational early warning model includes a risk detection submodule and a simulation early warning submodule, comprising: A digital twin environment is constructed based on a knowledge graph of end-to-end operational resilience and a historical end-to-end operational dataset. The digital twin environment is updated in real time in conjunction with a real-time end-to-end operational dataset. The digital twin environment is used to simulate the end-to-end operational status. Cluster the historical full-process operation dataset to obtain a non-stationary operation dataset. The non-stationary operation dataset refers to historical data on operational process imbalances that occurred during the full-process operation, and includes at least root cause data and impact data of operational imbalances. A reinforcement learning agent is constructed. Its state space is defined according to the digital twin environment. Its action space is set to inject controllable noise for early risk detection. A deep neural network is used as its policy network. The policy network is pre-trained in conjunction with the non-stationary operation dataset. The pre-trained reinforcement learning agent is used as a risk detection sub-module. The risk detection submodule is used to generate controllable noise, which is then input into the simulation early warning submodule for early operational risk identification.
3. The automated multi-agent full-process operation scheduling method according to claim 2, characterized in that, The method further includes: Controllable noise is input into the digital twin environment for simulation to generate a simulation early warning dataset, which includes an initial state vector, controllable noise, a disturbed state vector, and an undisturbed state vector. The initial state vector represents the initial state vector of the digital twin environment, the perturbation state vector represents the state vector sequence of the digital twin environment within a fixed time window after injecting controllable noise, and the unperturbed state vector represents the state vector sequence of the digital twin environment within a fixed time window without injecting controllable noise. A simulated early warning submodule is constructed based on a deep twin network. The simulated early warning dataset is input into the simulated early warning submodule. The difference between the perturbed state vector and the undisturbed state vector is obtained according to the attention mechanism network. Pre-training is performed in combination with a predetermined training objective. Training performance is verified in combination with a training and validation set constructed based on a non-stationary operation dataset. Training ends only when the predetermined performance condition is met.
4. The automated multi-agent full-process operation scheduling method according to claim 1, characterized in that, The real-time, end-to-end operational dataset is input into the operational early warning model to generate an operational early warning dataset. This dataset includes a set of operational early warning information and operational incentive tokens, comprising: After the real-time full-process operation dataset is input into the operation early warning model, the risk detection submodule generates corresponding controllable noise based on the real-time full-process operation dataset; The controllable noise is injected into a pre-constructed digital twin environment for simulation, generating a corresponding simulation early warning dataset; The simulation early warning submodule performs operational early warnings based on the simulation early warning dataset and obtains the corresponding operational early warning results. If the operational alert result is no abnormality, continue to monitor the operational status of the entire process; If the operational warning result indicates an anomaly, a corresponding set of operational warning information and operational incentive tokens will be generated. The operational early warning information set and the operational incentive tokens are encapsulated into an operational early warning dataset for output.
5. The automated multi-agent full-process operation scheduling method according to claim 4, characterized in that, The simulation early warning submodule performs operational early warnings based on the simulation early warning dataset and obtains the corresponding operational early warning results, including: The simulation early warning submodule parses the simulation early warning dataset to obtain the degree of difference between the digital twin environment before and after the injection of controllable noise; The difference is mapped and transformed according to preset rules to generate a corresponding early risk warning score; If the early risk warning score reaches a predetermined threshold, it is determined that an anomaly has occurred; If the early risk warning score does not reach the predetermined threshold, it is determined that there is no abnormality; Based on the knowledge graph of operational resilience throughout the entire process, and using the attention mechanism, the root cause of the simulation early warning dataset that has anomalies is located, and the set of anomaly root cause nodes is obtained. Root cause reasoning is performed on the entity nodes contained in the set of abnormal root cause nodes to generate corresponding root cause reasoning information. The root cause reasoning information set includes at least the abnormal root cause and the initial root cause elimination strategy. The set of abnormal root cause nodes and the corresponding root cause reasoning information set are encapsulated into an operational early warning information set. The initial root cause elimination strategy is analyzed based on a preset operational token mechanism to generate corresponding operational incentive tokens.
6. The automated multi-agent full-process operation scheduling method according to claim 1, characterized in that, Based on the aforementioned operational early warning dataset, a game-theoretic negotiation is conducted to obtain the negotiation results. Based on these results, an operational control strategy is generated, including: Based on the operational early warning information set, operational risks are verified, operational risk verification results are obtained, and the operational early warning information set is optimized based on the operational risk verification results to generate a corresponding operational risk resource set. The operational risk resource set includes at least the operational risk root cause entity set, the operational risk elimination resource set, and the initial root cause elimination strategy. A multi-agent reinforcement learning and dynamic game theory algorithm is adopted, and the game negotiation is carried out simultaneously with the operational risk resource set to generate the corresponding game negotiation result. Based on the results of the game negotiation, the initial resource scheduling strategy is optimized and adjusted to generate a corresponding operation and control strategy.
7. The automated multi-agent full-process operation scheduling method according to claim 6, characterized in that, Employing multi-agent reinforcement learning and dynamic game theory algorithms, and simultaneously combining operational risk resource sets for game negotiation, corresponding game negotiation results are generated, including: Configure a corresponding game-theoretic intelligent agent for each operational process in the entire operation process, which includes at least four operational processes: engineering construction, spare parts management, maintenance and repair, and testing services. Each game agent is initialized based on the initial root cause elimination strategy. Each game agent, based on its own scheduling cost, simultaneously combines the operational risk root cause entity set and the operational risk elimination resource set to conduct a rationality analysis of the initial root cause elimination strategy. The initial root cause elimination strategy includes an initial operational incentive token budget for eliminating the root cause of risk. The scheduling cost is represented by the actual number of operational incentive tokens generated according to a preset cost conversion rule, and includes at least the actual loss cost and the waiting compensation cost. Based on the results of the rationality analysis, each game agent negotiates with the goal of minimizing scheduling costs and generates corresponding operation and control strategies.
8. The automated multi-agent full-process operation scheduling method according to claim 1, characterized in that, Based on the implementation effect of the operational control strategy, operational incentive tokens are liquidated, an optimization suggestion set is generated, and feedback optimization is applied to the operational early warning model, including: The liquidation of the operational incentive tokens is represented by the deviation between the initial operational incentive token budget and the actual operational incentive tokens spent by the fulfilling agent of the operational control strategy. The corresponding liquidation result is generated based on the deviation. If the settlement result indicates that the budget is sufficient, the remaining operating incentive tokens after paying the actual operating incentive token expenditure will be settled as performance rewards to each performing smart agent. Each performing smart agent will then pay commissions according to preset rules and inject them into the risk-resistant reserve pool. If the liquidation result is that the budget is sufficient but the cost exceeds the limit due to the failure of the performing smart agent to perform, then the token compensation pool will be activated to compensate for the corresponding excess cost, and the smart agent that failed to perform will be penalized with tokens. The token penalty includes excess cost penalty and bad faith token penalty. The operational incentive tokens for the token compensation pool are derived from the token penalties imposed on unqualified smart entities. If the liquidation result shows that the budget significantly underestimates the actual cost, then the risk-resistant reserve pool will be activated to compensate each performing smart agent for costs, including advance token costs and reward token compensation. An optimization suggestion set is generated based on the liquidation results, and the operation early warning model is optimized based on the optimization suggestion set.
9. The automated multi-agent full-process operation scheduling method according to claim 1, characterized in that, Obtain a full-process operation dataset, which is divided into historical full-process operation datasets and real-time full-process operation datasets. Construct and update a full-process operation resilience knowledge graph based on the full-process operation datasets in real time, including: Multi-source data collection is performed on the entire operation and scheduling process to obtain a full-process operation dataset. The full-process operation dataset includes at least engineering construction process data, spare parts management process data, maintenance and repair process data, and testing service process data. Feature extraction and semantic analysis are performed on the full-process operation dataset to obtain the full-process operation entity set and the corresponding full-process topology association network. Taking each entity in the full-process operation entity set as a node, the full-process operation resilience knowledge graph is constructed simultaneously by combining the full-process topology association network. Based on graph neural networks and causal reasoning, the resilience of the full-process operational resilience knowledge graph is assessed, and an operational risk score is assigned to each entity node. The resilience assessment is expressed as a quantitative measure of the operational resilience of the current node by combining the entity node's own state health, topological importance, and resistance to risk contagion from neighboring nodes. The state health is used to quantify the current operational status of an entity node, the topological importance is used to quantify the importance of the entity node in the entire process operation, and the neighbor node risk contagion resistance is used to quantify the ability of the current node to maintain its own state health when neighbor nodes experience operational risks.
10. An automated multi-agent end-to-end operation scheduling platform, used to implement the method described in any one of claims 1 to 9, characterized in that, include: The data acquisition module is used to acquire the full-process operation dataset and construct and update the full-process operation resilience knowledge graph in real time based on the full-process operation dataset. The model building module is used to build an operational early warning model and pre-train the model using historical full-process operational datasets. An operation early warning module is used to input real-time full-process operation datasets into the operation early warning model to generate operation early warning datasets; The game decision-making module is used to conduct game negotiation based on the operation early warning dataset, obtain the game negotiation result, and generate operation control strategy based on the game negotiation result. The feedback adjustment module is used to perform operational incentive token liquidation and generate an optimization suggestion set based on the results of the operational incentive token liquidation to optimize the operational early warning model.
Citation Information
Patent Citations
Multi-park low-carbon scheduling method and system based on double games
CN116862144A
E-commerce intelligent operation monitoring and collaborative decision-making method based on big data analysis
CN119539920A
Multi-level information security policy generation method based on knowledge graph
CN119728302A
Project progress tracking and risk prediction system and method based on improved knowledge graph and multi-view graph neural network
CN120181781A
Big data analysis and prediction-based strategy dynamic optimization system and method
CN120337978A
Cited By
Merchant operation model training method and device, electronic equipment and storage medium
CN121456487A
Merchant operation model training method and device, electronic equipment and storage medium
CN121456487B