Dynamic cooperative control method and system for intelligent power plant with few people on duty

By combining metacognitive optimization models and reinforcement learning, a smart power plant control method was developed, which solved the adaptive problem of thermal power plants under dynamic changes. This method enables stable, safe, and efficient operation of power plants with low human input, meeting both economic and environmental requirements.

CN121763716APending Publication Date: 2026-03-31ANHUI TONGXINYUAN TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies lack adaptability in thermal power plants, resulting in control strategies failing to remain optimal under dynamically changing unit loads, fuel characteristics, or environmental conditions, and lacking long-term optimization consistency, which affects economic efficiency and carbon emission control.

Method used

The metacognitive optimization model is used to adjust the weights of the global optimization objective function based on real-time data. The optimal control strategy is generated in the digital twin model using reinforcement learning. Distributed collaboration is achieved through a potential game mechanism. Risk assessment and human-machine decision-making are carried out in conjunction with an evidence fusion algorithm.

Benefits of technology

It enables dynamic adaptation of power plant control strategies, improves system response efficiency and safety, reduces human error, and ensures stable operation of the power plant with low manpower input, while taking into account both economic efficiency and environmental protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121763716A_ABST
    Figure CN121763716A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic cooperative control method and system for a smart power plant with few people on duty, relates to the field of power plant operation control, and solves the technical problem of insufficient adaptive ability of a control strategy in a few people on duty environment in the prior art. The method comprises the following steps: adjusting a weight coefficient of a global optimization objective function by utilizing a meta-cognitive optimization model based on historical and real-time operation data; generating an optimal control strategy in a pre-constructed digital twinborn model through a reinforcement learning algorithm by using the adjusted global optimization objective function; performing distributed negotiation on each sub-region through a potential game mechanism according to the optimal control strategy to obtain a cooperative control strategy of each sub-region; and performing simulation verification on the cooperative control strategy by using a pre-constructed digital twinborn model, and performing risk assessment and strategy decision in combination with a risk assessment model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power plant operation control, specifically a dynamic collaborative control method and system for intelligent power plants with minimal human intervention. Background Technology

[0002] With the continuous improvement of automation levels in thermal power plants, the trend towards less staffed or even unmanned operation has become increasingly prevalent in the industry. Against this backdrop, traditional control strategies face severe challenges. Existing technologies often employ multi-objective optimization methods with fixed weights. However, fixed weights cannot adapt to the dynamic changes in power plant operating conditions. When unit load, fuel characteristics, or environmental conditions change, the original weight allocation may no longer be optimal, leading to low operating efficiency. Secondly, traditional methods lack the ability to evaluate the long-term effects of control strategies, focusing only on instantaneous optimization results while neglecting the cumulative impact of historical operating performance.

[0003] Especially under the strategic goals of carbon peaking and carbon neutrality, thermal power plants need to simultaneously consider economic benefits and carbon emission control, which places higher demands on optimizing control strategies. While existing reinforcement learning methods based on instantaneous rewards can achieve online optimization, they are prone to getting trapped in local optima due to the lack of metacognitive evaluation of historical performance, and cannot maintain consistency in multi-objective collaborative optimization over long-term operation. Summary of the Invention

[0004] This application provides a dynamic collaborative control method and system for smart power plants with minimal human intervention, which solves the technical problem of insufficient adaptive capability of control strategies in environments with minimal human intervention in the prior art.

[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, a dynamic collaborative control method for smart power plants with reduced human intervention is provided, including: Real-time operation data is obtained by collecting operational data from the entire power plant and its sub-regions. Based on historical and real-time operational data, the weight coefficients of the global optimization objective function are adjusted using a metacognitive optimization model; the global optimization objective function is used to solve a multi-objective optimization problem of power plant operation economy and carbon emission intensity. By utilizing the adjusted global optimization objective function, the optimal control strategy is generated in a pre-constructed digital twin model through a reinforcement learning algorithm; Each sub-region, based on its optimal control strategy, engages in distributed negotiation through a potential game mechanism to obtain a collaborative control strategy for each sub-region. The collaborative control strategy is simulated and verified using a pre-built digital twin model, and risk assessment and strategy decision-making are performed in conjunction with a risk assessment model; the risk assessment model is a human-machine decision-making power allocation model obtained through evidence fusion algorithm.

[0006] Based on the above technical solutions, the dynamic collaborative control method for a smart power plant with reduced human intervention provided in this application combines real-time full-domain data acquisition with a metacognitive optimization model. This allows for a dynamic balance between economic efficiency and low-carbon requirements based on historical and real-time data, making the control objectives more aligned with the actual operating conditions of the power plant. Secondly, the linkage between reinforcement learning and digital twins places the control strategy generation process in a virtual simulation environment, reducing the risks and costs of actual equipment debugging while improving the optimality and adaptability of the strategy through algorithm iteration. Furthermore, the game theory mechanism replaces traditional centralized control, allowing each sub-region to autonomously negotiate and achieve collaboration, reducing the transmission delay of control commands and improving system response efficiency. Finally, the design of the evidence fusion risk assessment model and the allocation of human-machine decision-making power can accurately match the needs of scenarios with reduced human intervention. Through scientific evaluation and decision-making division of labor, human error is reduced, ensuring the stable and safe operation of the power plant with low human input.

[0007] Furthermore, the adjustment of the weight coefficients of the global optimization objective function using the metacognitive optimization model includes: The economic indicators for power plant operation are calculated based on data on power generation, fuel costs, and operation and maintenance costs. The carbon emission intensity index of the power plant operation is calculated based on carbon emission monitoring data; According to the equipment safety operation rules, system constraint violation penalty items are set, and combined with the economic indicators, carbon emission intensity indicators and system constraint violation penalty items, a global optimization objective function and a metacognitive reward function are constructed. The weight coefficients of the metacognitive reward function are consistent with the weight coefficients of the global optimization objective function. The metacognitive reward function is a reward function constructed based on the weighted integral of the global optimization objective function within the historical time window, and is used to evaluate the long-term performance of the control strategy. Based on gradient sensitivity analysis, the metacognitive reward function is optimized by time integration to obtain a global optimization objective function with adaptive weight adjustment.

[0008] Furthermore, the formula for calculating the global optimization objective function is: ;in, , representing economic indicators, , representing carbon emission intensity; , indicating a penalty for violating the constraint, w i y represents the penalty weight for exceeding the limit of the i-th parameter. i y represents the operating parameters of the i-th device. i safe δ represents the safety threshold for the i-th real-time running data. iThe safety margin for the i-th real-time running data, where α, β, and γ represent the weighting coefficients of each item; The formula for calculating the metacognitive reward function is: Where λ represents the discount factor used to adjust the decay rate of historical data, t represents the current time, τ represents the time of the integration variable, and T represents the length of the integration time window.

[0009] Furthermore, the calculation formula for the gradient sensitivity analysis is as follows: Where η represents the adjustment step size, The standard deviation of economic indicators Let f(·) represent the standard deviation of carbon emission indicators, f(·) be a decision function based on fuzzy logic, and u represent the control action.

[0010] Furthermore, the generation of the optimal control policy in the pre-built digital twin model using a reinforcement learning algorithm includes: Construct a digital twin model of the power plant; the digital twin model should be able to map at least the dynamic characteristics and coupling relationships of the boiler, steam turbine, and environmental protection equipment. The adjusted global optimization objective function is used as the reward function for the reinforcement learning agent, and policy exploration is carried out in the digital twin environment. The environment for setting up a reinforcement learning agent is defined; wherein, the action space of the environment includes, but is not limited to, continuously adjustable operation variables such as equipment speed commands, total air volume setpoints, turbine control valve opening commands, flue gas recirculation damper openings, and ammonia injection control valve openings; and the state space includes, but is not limited to, power plant operating data such as unit load, main steam pressure, main steam temperature, reheat steam temperature, furnace negative pressure, flue gas oxygen content, nitrogen oxide concentration, and carbon dioxide emission intensity. Construct an actor-critic network architecture; where the actor network takes the state space as input and outputs continuous control actions, and the critic network takes the concatenated vector of the state space and action space and outputs the Q value representing the action value function; The parameters of the actor network and the critic network are iteratively trained and validated using the deep deterministic policy gradient algorithm to obtain a converged policy network. Real-time power plant operation data is input into the converged policy network for online learning and policy generation. For each control policy generated, a Monte Carlo simulation is performed in the digital twin model to calculate the expected cumulative reward value of the control policy. The policy with the highest expected cumulative reward value is selected as the optimal control policy.

[0011] Furthermore, the distributed negotiation through a potential game mechanism to obtain the collaborative control strategy for each sub-region includes: Define a cost function for the edge controller of each sub-region: ; among which, Ji This represents the total cost of the i-th edge controller. This represents the action of the i-th controller. Let N(i) represent the set of actions of controllers other than controller i, and let N(i) represent the set of controllers adjacent to controller i. ij To characterize the potential function representing the dynamic coupling relationship between the subsystems to which controller i and controller j belong, y i Let Q represent the output control vector of the i-th controller. i R represents the weight matrix of the output tracking error. i The weight matrix representing the control action, y i ref ρ represents the control parameters issued to the i-th controller in the optimal control strategy. ij This represents the coupling strength coefficient, which is determined based on the strength of the physical association between sub-regions. Each controller randomly generates initial control actions and sets a convergence threshold and a maximum number of iterations; In each iteration, each controller fixes the control actions of the other controllers and solves for the optimal control action that minimizes the cost function; wherein, the optimization problem to be solved is to find the optimal control action that minimizes the cost function, given that the control actions of the other controllers remain unchanged. Minimize control actions u i ; When the update norm of the control actions of all controllers is less than the convergence threshold, or the number of iterations reaches the maximum number of iterations, the iteration stops, and the current combination of control actions of each controller is output as the Nash equilibrium solution of the potential game. The solution is the cooperative control strategy of each sub-region. The update norm represents the sum of the L2 norms of the changes in control actions of all controllers between two adjacent iterations.

[0012] Furthermore, the coupling potential function is determined based on the physical coupling characteristics between the control systems of the sub-regions. For the boiler control system and the turbine control system, the coupling potential function is defined as: ;in, Main steam pressure, Set the turbine inlet pressure value. Main steam temperature, Set the turbine inlet temperature setpoint. These are the weighting coefficients for the temperature coupling term; The coupling potential function between the combustion control system and the environmental protection control system is defined as: ;in, The concentration of nitrogen oxides during the power generation process. This represents the maximum denitrification capacity of the environmental control system. The concentration of sulfur dioxide during the power generation process. ξ represents the maximum desulfurization capacity of the environmental control system, and ξ is the weighting coefficient of the sulfur dioxide coupling term.

[0013] Furthermore, the simulation verification of the cooperative control strategy using a pre-built digital twin model includes: Set the initial state of the digital twin model to be consistent with the operating state of the current real power plant; The control command sequence of the aforementioned collaborative control strategy is used as input and injected into the initialized digital twin model; the control command sequence includes a complete future scheduling period. Numerical integration is performed using the variable-step-size Runge-Kutta method to advance the dynamic evolution of the digital twin model and predict the change trajectory of key operating parameters during the future scheduling period; the key operating parameters include at least unit load, main steam pressure, main steam temperature, nitrogen oxide emission concentration, carbon dioxide emission concentration, and furnace negative pressure. Based on the change trajectory of the aforementioned key operating parameters, the values ​​of safety indicators, economic indicators, environmental indicators, and stability indicators are calculated respectively; among them, The safety indicators represent the probability and maximum extent of exceeding the limits for main steam pressure, main steam temperature, and furnace negative pressure. The unit load is used to calculate the power generation revenue in the economic indicators. The environmental indicators include nitrogen oxide emission concentration and carbon emission intensity. The stability indicators represent the variance of the fluctuation of each key operating parameter. A threshold is set for each performance indicator. When the calculated value of each performance indicator is less than the corresponding threshold, the collaborative control strategy is deemed to have passed the verification; otherwise, the collaborative control strategy is deemed to have failed the verification.

[0014] Furthermore, the process of combining risk assessment models for risk assessment and strategy decision-making includes: Based on digital twin verification results, equipment health status, operation history, and environmental data, a risk assessment model based on evidence theory is constructed to obtain the risk index under the current collaborative control strategy. The calculated values ​​of each performance indicator during the simulation verification process are averaged and weighted to obtain the strategy performance score. The strategy credibility score is calculated based on the risk index and the strategy performance score; wherein, the strategy credibility score C is calculated as: C=w1×(1-risk index / preset risk threshold)+w2×(strategy performance score / benchmark score), and the benchmark score represents the average weighted sum of the threshold values ​​of each performance indicator; A multi-level decision-making mechanism is constructed based on the credibility score of the strategy to determine the human-machine decision-making ratio at different levels.

[0015] Furthermore, the calculation formula for the risk assessment model is as follows: Define the identification framework θ = low risk (L), medium risk (M), high risk (H); Construct a basic probability assignment function for each source of evidence Where k=1,2,3,4, corresponding to four data sources: digital twin verification results, device health status, operation history, and environmental data, respectively. The Dempster combination rule is used to synthesize the basic probability assignments of the four evidence sources, resulting in the synthesized basic probability assignment function m(A); where A θ; The risk index R is calculated using the formula R = 0.5 × Bel(H) + 0.5 × Bel(M); where Bel is the confidence function. .

[0016] Furthermore, the construction of a basic probability assignment function for each source of evidence includes: Using the digital twin verification result as the primary source of evidence, the basic probability allocation function for the primary source of evidence is constructed as follows: ; where N i Let be the normalized value of the i-th performance metric. The safety threshold for the i-th performance metric. Let I(·) be the warning threshold for the i-th performance metric and the indicator function. Using the device health status as a second source of evidence, the basic probability allocation function for this second source of evidence is as follows: Among them, L remain For the remaining lifespan of the equipment, L design For the equipment design life, F level The fault warning level is represented by k, which is an adjustment parameter. Using the operation history as a third source of evidence, the basic probability distribution function for the third source of evidence is constructed as follows: Among them, P success E represents the frequency of using the fully automated decision-making mode within a pre-defined historical time period. level This indicates the operator's experience level, and 'a' is an adjustment parameter. Using environmental data as a fourth source of evidence, the basic probability distribution function for this fourth source of evidence is constructed as follows: ; where ΔE j ΔE represents the rate of change of the j-th environmental parameter. j0 E represents the threshold value for the rate of change of the j-th environmental parameter. j E represents the current value of the j-th environment parameter. j0 b represents the pre-specification of the j-th environmental parameter. j c j The slope parameter is represented by the environmental parameters, which include temperature, humidity, air pressure, and wind speed.

[0017] Furthermore, the multi-level decision-making mechanism includes: When the strategy credibility score is greater than the first threshold, a fully automatic decision-making mode is adopted to automatically execute the collaborative control strategy. When the strategy credibility score is less than or equal to the first threshold and greater than the second threshold, a human-machine collaborative decision-making mode is adopted, and the collaborative control strategy is pushed to the central control platform for manual confirmation of execution or rejection of the collaborative control strategy. When the strategy credibility score is less than or equal to the second threshold, the manual decision-making mode is activated.

[0018] Secondly, this application provides a dynamic collaborative control system for smart power plants with minimal human intervention, comprising: a data acquisition module, a target optimization module, a strategy generation module, a collaborative negotiation module, and a verification and evaluation module; wherein, The data acquisition module is used to collect real-time operating data of the entire power plant and its sub-regions to form historical and real-time operating datasets. The target optimization module is used to adjust the weight coefficients of the global optimization objective function based on the data acquired by the data acquisition module and the metacognitive optimization model, so as to achieve a dynamic balance between multiple objectives such as economic efficiency and carbon emission intensity. The strategy generation module is used to generate the optimal control strategy in a pre-built digital twin model based on the adjusted global optimization objective function through a reinforcement learning algorithm. The collaborative negotiation module is used to enable each sub-region to conduct distributed negotiation based on the optimal control strategy through a potential game mechanism to obtain the collaborative control strategy for each sub-region. The verification and evaluation module is used to simulate and verify the collaborative control strategy using a digital twin model, and to conduct risk assessment and strategy decision-making using a risk assessment model constructed by evidence fusion algorithm, so as to ensure the safety and effectiveness of the control strategy.

[0019] Compared with the prior art, the beneficial effects of this application are: This application adjusts the weights of the global optimization objective function by combining economic indicators, carbon emission intensity indicators, and system constraint violation penalties through a metacognitive optimization model. This avoids the problem of insufficient adaptability to operating conditions caused by fixed objective weights in traditional control. It can also evaluate the long-term performance of the control strategy based on historical time windows, ensuring that the control direction always conforms to the actual operating needs of the power plant. Secondly, by using a digital twin model of the dynamic characteristics and coupling relationships of key equipment, combined with reinforcement learning algorithms, strategy exploration is carried out, and Monte Carlo simulation is used to select the strategy with the highest expected cumulative reward value, which reduces the risks and costs in the actual equipment commissioning process and makes the strategy more in line with the complex operating logic of the power plant. Next, a cost function containing coupling relationships is defined for the sub-region edge controller through a potential game mechanism, enabling each sub-region to quickly reach Nash equilibrium in distributed negotiation, replacing the traditional centralized control mode, reducing the transmission delay of control commands, and avoiding operational conflicts caused by independent control of sub-regions. In addition, this application constructs a reliable verification system before strategy execution. It uses a digital twin model to reproduce the current real power plant operation status, and predicts the change trajectory of key operating parameters after injecting control command sequences. It verifies the feasibility of the strategy from multiple dimensions such as safety, economy, environmental protection and stability, and prevents unqualified strategies from being put into actual operation. Finally, a risk assessment model is constructed by integrating multiple data sources, including digital twin verification results, equipment health status, operation history, and environmental data, based on evidence fusion algorithms. A multi-level human-machine decision-making mechanism is established by combining strategy performance scores. This mechanism can achieve fully automated decision-making to reduce human dependence when the strategy credibility is high, and can also initiate human-machine collaboration or manual decision-making to avoid errors when the risk is high. While reducing human input, it comprehensively ensures the safety, economy, and environmental protection of power plant operation, effectively solving the shortcomings of existing technologies in terms of dynamic adaptability, collaborative efficiency, and risk prevention and control. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A system architecture diagram of a dynamic collaborative control system for a smart power plant with minimal human intervention, provided in an embodiment of this application; Figure 2 A flowchart illustrating a dynamic collaborative control method for a smart power plant with minimal human intervention, provided in an embodiment of this application; Figure 3 A flowchart illustrating another dynamic collaborative control method for a smart power plant with minimal human intervention, provided in an embodiment of this application; Figure 4 A flowchart illustrating another dynamic collaborative control method for a smart power plant with minimal human intervention, provided in an embodiment of this application; Figure 5 This is a flowchart illustrating another dynamic collaborative control method for a smart power plant with reduced human intervention, provided in an embodiment of this application. Detailed Implementation

[0022] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0023] The dynamic collaborative control method for smart power plants with reduced human intervention provided in this application embodiment can be applied to, for example... Figure 1 In a dynamic collaborative control system for a smart power plant with minimal human intervention, as shown, Figure 1 As shown, the system includes a data acquisition module, a target optimization module, a strategy generation module, a collaborative negotiation module, and a verification and evaluation module, all connected in communication. The data acquisition module is used to collect real-time operating data of the entire power plant and its sub-regions, forming historical and real-time operating datasets; The target optimization module is used to adjust the weight coefficients of the global optimization objective function based on the data acquired by the data acquisition module and the metacognitive optimization model, so as to achieve a dynamic balance between multiple objectives such as economic efficiency and carbon emission intensity. The strategy generation module is used to generate the optimal control strategy in a pre-built digital twin model based on the adjusted global optimization objective function through a reinforcement learning algorithm. The collaborative negotiation module is used to enable each sub-region to conduct distributed negotiation based on the optimal control strategy through a potential game mechanism, so as to obtain the collaborative control strategy for each sub-region. The verification and evaluation module is used to simulate and verify the collaborative control strategy using a digital twin model, and to conduct risk assessment and strategy decision-making by combining a risk assessment model built with an evidence fusion algorithm, so as to ensure the safety and effectiveness of the control strategy.

[0024] To address the technical problems of poor dynamic adaptability, low sub-regional coordination efficiency, and insufficient control strategy security in existing smart power plant control technologies, this application provides a dynamic collaborative control method and system for smart power plants with minimal human intervention. The method includes: Real-time operation data is obtained by collecting operational data from the entire power plant and its sub-regions. Based on historical and real-time operational data, the weight coefficients of the global optimization objective function are adjusted using a metacognitive optimization model; the global optimization objective function is used to solve a multi-objective optimization problem of power plant operation economy and carbon emission intensity. By utilizing the adjusted global optimization objective function, the optimal control strategy is generated in a pre-constructed digital twin model through a reinforcement learning algorithm; Each sub-region, based on its optimal control strategy, engages in distributed negotiation through a potential game mechanism to obtain a collaborative control strategy for each sub-region. The collaborative control strategy is simulated and verified using a pre-built digital twin model, and risk assessment and strategy decision-making are carried out in combination with a risk assessment model; the risk assessment model is a human-machine decision-making power allocation model obtained by evidence fusion algorithm.

[0025] Based on this, this application can realize dynamic adaptation, efficient collaboration, safety verification and intelligent division of labor between humans and machines in power plant control. While reducing manpower input, it takes into account the economic efficiency, environmental protection and safety of operation, and meets the core needs of smart power plants with less manpower.

[0026] like Figure 2 As shown in the embodiment of this application, a dynamic collaborative control method for a smart power plant with minimal human intervention is provided, comprising: S1. Real-time acquisition of operating data for the entire power plant and its sub-regions to obtain real-time operating data.

[0027] The real-time operating data includes, but is not limited to, unit load, main steam pressure, main steam temperature, reheat steam temperature, furnace negative pressure, flue gas oxygen content, nitrogen oxide concentration, carbon dioxide emission concentration, fuel consumption, equipment operating speed, and vibration parameters.

[0028] In some implementations, sub-regions can be divided according to the core equipment systems of the power plant, such as boiler system area, steam turbine system area, environmental protection system area, and auxiliary equipment system area; they can also be divided according to physical spatial layout, such as boiler workshop area, steam turbine plant area, and desulfurization and denitrification workshop area; or they can be divided according to control function modules, such as combustion control area, steam-water control area, and emission control area; data acquisition can be achieved by relying on existing power plant data acquisition architectures such as sensors, data acquisition and monitoring control systems, and distributed control systems.

[0029] It should be noted that data preprocessing is required during the data collection process, including outlier removal, missing value completion, and data standardization, to ensure the accuracy and consistency of real-time running data.

[0030] S2. Based on historical and real-time operational data, adjust the weight coefficients of the global optimization objective function using a metacognitive optimization model.

[0031] Among them, the metacognitive optimization model is an intelligent optimization model based on metacognitive theory (including three core components: metacognitive knowledge, metacognitive monitoring, and metacognitive regulation), data-driven optimization theory, and gradient sensitivity analysis methods. It integrates the metacognitive reward function and the global optimization objective function, possessing closed-loop self-learning and adaptive dynamic adjustment capabilities. Furthermore, the metacognitive knowledge model corresponds to the accumulation of correlation patterns between "operating condition strategy performance" in historical power plant operating data, such as the balance between economic efficiency and carbon emissions under different loads; the metacognitive monitoring model corresponds to the process of evaluating the long-term performance of the control strategy in real time based on the weighted integral of the global optimization objective function within a historical time window through the metacognitive reward function; and the metacognitive regulation model corresponds to the core mechanism of iterative optimization of weight coefficients, relying on gradient sensitivity analysis to quantify the sensitivity of indicators such as economic efficiency and carbon emission intensity to control actions. The global optimization objective function is a function used to solve a multi-objective optimization problem of power plant operation economy and carbon emission intensity. Additional objectives such as equipment safety and operational stability can be incorporated according to actual needs.

[0032] In some implementations, the metacognitive optimization model can also be built based on existing optimization theories or technologies such as neural networks, fuzzy control theory, and genetic algorithms. By building a data-driven performance evaluation module, the weight coefficients of the multi-objective optimization function can be iteratively solved using methods such as gradient descent, particle swarm optimization, and ant colony optimization. The global optimization objective function can be determined by setting quantitative standards for economic indicators and carbon emission intensity indicators, combined with power plant operation objectives such as grid peak-valley scheduling requirements and environmental policy requirements, to determine the basic weight framework, which is then dynamically corrected by the metacognitive optimization model based on real-time data.

[0033] S3. Using the adjusted global optimization objective function, the optimal control strategy is generated in the pre-built digital twin model through reinforcement learning algorithm.

[0034] Among them, reinforcement learning algorithms can learn strategies in digital twin models by interacting and trying with the virtual environment based on the reward rules set by the global optimization objective function, and guide the digital twin model to simulate the power plant operation state under different control actions. The optimal control strategy refers to a set of continuous or discrete operation instructions that can make the global optimization objective function take the optimal value and at the same time satisfy the equipment operation constraints, covering control actions such as equipment start-up and shutdown, parameter adjustment (such as valve opening and speed setting).

[0035] In some implementations, the digital twin model needs to be constructed based on the geometric structure, dynamic characteristics, and coupling relationships of the power plant's physical entity. A geometric model can be built using 3D modeling software, and dynamic simulation can be achieved by combining mechanistic modeling such as boiler heat balance equations and turbine work formulas with data-driven modeling. Reinforcement learning algorithms can utilize existing algorithms such as Deep Deterministic Policy Gradient (DDPG), Proximal Policy Optimization (PPO), and Q-learning. The digital twin model is used as a virtual environment, and the calculation result of the global optimization objective function is used as the reward signal to train the agent to generate control strategies. During the generation process, multiple simulations can be run to select stable strategies that meet the target requirements under different operating conditions, which are then used as the optimal control strategy.

[0036] It should be noted that the digital twin model needs to maintain real-time data synchronization with the physical power plant. The parameter deviations of the virtual model are corrected by real-time operating data to ensure the consistency between the simulation results and the actual operating status.

[0037] For example, the DDPG algorithm can be used to construct a reinforcement learning agent. Parameters such as unit load and main steam in the digital twin model are used as state inputs, and the turbine valve opening and boiler total air volume setpoint are used as action outputs. The adjusted global optimization objective function is used as the reward. The agent simulates control actions under different loads multiple times in the digital twin model. Through iterative training, the strategy converges, and finally, the control strategy that can achieve a balance between economy and carbon emissions within the unit load range of 50%-100% is selected. This includes instructions such as the turbine valve opening range and total air volume setpoint corresponding to different loads.

[0038] S4. Each sub-region, based on its optimal control strategy, conducts distributed negotiation through a potential game mechanism to obtain a collaborative control strategy for each sub-region.

[0039] Among them, the potential game mechanism is a collaborative method based on non-cooperative game theory. It enables each sub-region to achieve global coordination through information exchange and instruction adjustment while pursuing its own control objectives (such as the steam production target of the boiler sub-region and the power generation efficiency target of the turbine sub-region), thus avoiding operational conflicts between sub-regions caused by independent control. The mechanism can set negotiation rules in combination with the physical coupling relationship between sub-regions (such as steam supply and demand, flue gas transmission) to ensure that the instructions of each sub-region are coordinated and consistent.

[0040] In some implementations, each sub-region can be configured with an edge controller. The edge controller receives the sub-region target parameters (such as the main steam pressure target of the boiler sub-region and the nitrogen oxide emission target of the environmental protection sub-region) from the optimal control strategy, and generates initial control commands based on its own real-time operating data. Information interaction between sub-regions is realized through communication methods such as industrial Ethernet and 5G industrial Internet. Each sub-region adjusts its own commands according to the initial commands of other sub-regions to reduce coordination conflicts (such as the boiler sub-region adjusting fuel supply according to the steam demand of the turbine sub-region). Convergence conditions can be set during the negotiation process (such as the change in two adjacent commands being less than a threshold). When all sub-region commands meet the convergence conditions, the final collaborative control strategy is output.

[0041] It should be noted that distributed negotiation needs to reduce data transmission latency. Edge computing technology can be used to deploy some negotiation logic locally in sub-regions to avoid instruction lag caused by centralized processing. This is especially suitable for large power plants with a large number of sub-regions and complex coupling relationships.

[0042] For example, the boiler sub-region generates an initial command for a main steam pressure of 10 MPa based on the optimal control strategy, while the turbine sub-region generates an initial command for a steam demand of 100 t / h. Through a game theory mechanism, the boiler sub-region discovers that the current steam output can only meet the demand of 90 t / h, so it adjusts the fuel supply to increase the steam output to 100 t / h, while stabilizing the main steam pressure at 10 MPa. Based on the adjustment result of the boiler sub-region, the turbine sub-region fine-tunes the valve opening from 80% to 82%, ultimately achieving coordination and obtaining a coordinated control strategy of "fuel supply increased by 5%" for the boiler sub-region and "valve opening of 82%" for the turbine sub-region.

[0043] S5. The collaborative control strategy is simulated and verified using a pre-built digital twin model, and risk assessment and strategy decision-making are carried out in conjunction with a risk assessment model.

[0044] Among them, simulation verification refers to incorporating collaborative control strategies into the digital twin model to simulate the power plant operation process over a future period (such as 1 hour or 24 hours), and monitoring the changing trends of key operating parameters and whether they meet safety thresholds; the risk assessment model is a human-machine decision-making power allocation model obtained through evidence fusion algorithms. It can integrate multi-source data such as digital twin verification results, equipment health status, and operation history records to quantify the risk level of the current control strategy and provide suggestions on the proportion of human-machine decision-making (such as fully automatic execution, execution after manual confirmation, and pure manual decision-making).

[0045] In some implementations, simulation verification can set up multiple test scenarios, including normal operating conditions, extreme operating conditions (such as extreme temperatures, sudden changes in grid load), and minor fault operating conditions (such as single sensor failure, decrease in auxiliary machine efficiency). The digital twin model is used to predict the execution effect of the collaborative control strategy in different scenarios and determine whether there are problems such as parameter exceeding limits or equipment overload. The risk assessment model can use the Dempster-Shafer evidence fusion algorithm to use the digital twin verification results (such as the probability of parameter exceeding limits), equipment health status (such as remaining lifespan, fault warning level), and operation history (such as the execution success rate of similar strategies) as evidence sources to calculate the risk index of the strategy. The strategy decision is divided into levels according to the risk index. High-risk strategies initiate manual decision-making, medium-risk strategies adopt human-machine collaborative decision-making, and low-risk strategies adopt fully automatic decision-making.

[0046] It should be noted that the simulation verification needs to cover common abnormal scenarios in power plants to avoid potential risks that are not discovered due to missing scenarios; the evidence sources of the risk assessment model need to be updated in real time to ensure the timeliness and accuracy of the risk assessment results.

[0047] Based on the above technical solutions, the dynamic collaborative control method for less-manned operation of smart power plants provided in this application ensures the reliability of the data source for control decisions through full-scenario data collection, achieves dynamic adaptation of control objectives through metacognitive optimization models, improves the optimality and security of control strategies through the combination of reinforcement learning and digital twins, enhances the collaborative efficiency of sub-regions through a potential game mechanism, and realizes intelligent division of labor between humans and machines through a risk assessment model, which can comprehensively meet the operational needs of less-manned operation of smart power plants.

[0048] In one possible implementation of this application embodiment, the above-mentioned S1 can be specifically implemented by the following S101, S102 and S103, which are described in detail below: S101. Determine the scope and core data types for real-time operation data collection.

[0049] In some implementation methods, the scope of data collection must explicitly include economic data, carbon emission data, and equipment operating parameters: economic data includes power generation, fuel cost data, and operation and maintenance cost data, used for subsequent calculation of economic indicators for power plant operation; carbon emission data includes actual carbon emissions and carbon emission quota data, used for calculating carbon emission intensity indicators; equipment operating parameters must at least cover unit load, main steam pressure, main steam temperature, reheat steam temperature, furnace negative pressure, flue gas oxygen content, nitrogen oxide concentration, and carbon dioxide emission concentration, and also include key operating parameters of boilers, turbines, and environmental protection equipment, such as the speed of desulfurization and denitrification devices, valve opening, and damper opening, used to determine whether the equipment complies with safe operation rules.

[0050] S102, Deploy data acquisition equipment and real-time transmission architecture.

[0051] In some implementation methods, suitable data acquisition equipment can be selected for different data types: power generation uses a high-precision energy meter with a precision of 0.2S level, installed in the main outgoing switchgear of the power plant; fuel consumption uses a belt scale or gas flow meter, installed at the end of the coal conveyor belt and the gas conveyor pipeline, respectively; carbon emissions and flue gas parameters (nitrogen oxides, carbon dioxide concentration, and flue gas oxygen content) use CEMS sensors that meet national environmental protection standards, installed at the boiler flue outlet and the inlet and outlet of environmental protection equipment; equipment operating parameters (pressure, temperature, speed, and valve opening) use industrial-grade sensors, installed at key monitoring points of equipment such as the boiler drum, turbine inlet, and valve actuators.

[0052] S103. Preprocess the collected raw data to obtain effective real-time running data.

[0053] The raw data may contain outliers, missing values, or inconsistent data formats. The purpose of preprocessing is to eliminate data noise, which mainly includes three steps: outlier removal, missing value completion, and data standardization. This ensures the accuracy, completeness, and consistency of real-time running data, laying the foundation for the subsequent fusion and application of historical and real-time data.

[0054] Based on the above technical solution, S1, through comprehensive data collection and processing, not only covers the core data required for subsequent calculations of economic indicators and carbon emission intensity indicators, but also includes key parameters for judging the safe operation of equipment, providing high-quality data support for subsequent operation steps.

[0055] In one possible implementation of this application embodiment, the above-mentioned S2 can be specifically implemented by the following S201, S202, S203 and S204, which are described in detail below: S201. Calculate the economic indicators of power plant operation based on historical and real-time operating data.

[0056] Among them, economic indicators It is a core parameter for quantifying the profitability of power plant operation. It needs to be calculated by combining historical and current real-time power generation, fuel costs, operation and maintenance costs, and benchmark power generation revenue to reflect the economic advantages and disadvantages of the current operating conditions relative to the benchmark level.

[0057] In some implementations, the economic indicators are calculated using the following formula: ;in, Revenue from power generation: It is obtained by multiplying the real-time power generation (unit: MWh) by the grid-approved on-grid tariff (unit: yuan / kWh), that is, revenue from power generation = power generation × 1000kWh / MWh × on-grid tariff. The power generation is taken from the real-time reading of the grid metering device, and the historical data is retrieved from the power plant's metering ledger. Fuel cost: Calculated based on real-time consumption of fuel type (coal unit: tons, gas unit: cubic meters) and purchase price (unit: yuan / ton or yuan / cubic meter), i.e. "fuel cost = fuel consumption × fuel unit price". Consumption is taken from the metering sensors of the fuel delivery system. Historical data must include additional costs such as fuel transportation and storage. Operation and maintenance costs are divided into fixed costs such as monthly equipment maintenance fees and variable costs such as consumable replacement fees and lubricant fees. Fixed costs are calculated based on the historical average monthly amortization value, while variable costs are calculated based on real-time consumable usage. Benchmark power generation revenue: The average monthly power generation revenue under normal operating conditions of the power plant over the past 12 months is taken as the normalized benchmark for economic indicators to eliminate the impact of differences in power generation at different times.

[0058] S202. Calculate the carbon emission intensity index of power plant operation based on carbon emission monitoring data.

[0059] Among them, carbon emission intensity index It is a key parameter for measuring the environmental compliance of power plants. It is calculated by the ratio of real-time actual carbon emissions to the current carbon emission quota to determine whether the current carbon emissions exceed the policy-permitted range.

[0060] In some implementations, the formula for calculating carbon emission intensity is: The actual carbon emissions can be calculated in two ways: one is to directly collect carbon dioxide concentration through online flue gas monitoring and calculate the real-time emissions based on flue gas flow rate; the other is to calculate based on fuel consumption and the corresponding fuel carbon emission coefficient, i.e., actual carbon emissions = fuel consumption × fuel carbon emission coefficient; carbon emission quotas are determined by local ecological and environmental departments based on power plant installed capacity and industry emission standards, and are divided into annual quotas and monthly decomposed quotas.

[0061] S203. Set system constraint violation penalty items and construct a global optimization objective function and a metacognitive reward function.

[0062] in, To mitigate the risk of equipment operating beyond its safety threshold, the global optimization objective function integrates the three objectives of "economy, carbon emissions, and safety," while the metacognitive reward function... The long-term performance of the control strategy is evaluated based on historical time windows, and the weight coefficients of the three factors are kept consistent, providing a basis for subsequent weight adjustments.

[0063] In some implementations, the specific steps for constructing the global optimization objective function and the metacognitive reward function are as follows: 1. Calculate system constraint violation penalties : According to the equipment safety operation rules, safety thresholds are set for each key operating parameter, such as main steam pressure, main steam temperature, and furnace negative pressure. Safety margin δ i and the weight of the penalty for exceeding the limit w i And it is calculated using the following formula: ;In the formula, y i Let max(0,·) be the real-time value of the operating parameter of the i-th device, where max(0,·) represents the value only when the parameter exceeds the limit. Exceeding the safety margin δ i The penalty value is calculated only when w. i This represents the penalty weight for exceeding the limit of the i-th parameter, which is set according to the consequences of parameter failure, such as main steam pressure wi=5, furnace negative pressure w i =3.

[0064] 2. Construct the global optimization objective function: The bureau's optimization objective function, with economic indicators as its core, combined with carbon emission intensity and safety penalties, is calculated using the following formula: ; where α, β, and γ are the weighting coefficients for economic efficiency, carbon emission intensity, and system constraint violation penalties, respectively. The initial values ​​can be set according to the power plant's operation objectives and subsequently adjusted through a metacognitive optimization model.

[0065] 3. Constructing the metacognitive reward function R meta : The metacognitive reward function is constructed based on the weighted integral of the global optimization objective function within the historical time window, and is used to evaluate the long-term performance of the control strategy. The formula is as follows: Where t is the current time, τ is the time of the integration variable, T is the length of the integration time window, usually taken as 1224 hours, λ is the discount factor, with a value range of 0.6 to 0.9, used to adjust the decay rate of historical data, and α(τ) and β(τ) are the weighting coefficients of historical time.

[0066] It should be pointed out that, In the calculation formula, the safety margin δi is usually set to 5% to 10% of the safety threshold range. For example, if the main steam pressure safety threshold is 10-12 MPa, then δi... i =0.5MPa.

[0067] S204. Optimize the metacognitive reward function based on gradient sensitivity analysis to obtain adaptively adjusted weight coefficients.

[0068] Among them, gradient sensitivity analysis is used to quantify the sensitivity of economic efficiency and carbon emission indicators to control actions. By iteratively optimizing the metacognitive reward function, the dynamic adjustment of α and β in the global optimization objective function is achieved, ensuring that the weights fit the real-time operating conditions.

[0069] In some implementations, gradient sensitivity analysis calculates the adjustment amount of α using the following formula, and the adjustment logic for β is similar: Where η represents the adjustment step size, with values ​​of 0.01 to 0.05, used to control the magnitude of weight adjustment; , These are the partial derivatives of the economic indicators and carbon emission intensity indicators with respect to the control action u, respectively, which quantify the degree of influence of the control action on the indicators. , , respectively, are the historical standard deviations of economic indicators and carbon emission intensity indicators, used to normalize partial derivatives and eliminate differences in the magnitude of the indicators; f(·) is a decision function based on fuzzy logic, with two partial derivatives as inputs and an adjustment coefficient between 0 and 1 as output.

[0070] The weight adjustment process is as follows: calculation is performed at regular intervals. and The adjustment amount is added to the current weight coefficients to obtain new α and β, which are then substituted into the global optimization objective function to complete the adaptive update.

[0071] Based on the above technical solution, S2 achieves dynamic adaptation of the weight coefficients of the global optimization objective function by quantifying core indicators, constructing multi-objective functions, and optimizing weights through gradients. This provides a scientific and dynamic objective framework for generating optimal control strategies through reinforcement learning.

[0072] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, the above S3 can be implemented through the following S301, S302, S303 and S304, which are explained in detail below: S301. Construct a digital twin model of the power plant.

[0073] Among them, the power plant digital twin model is a virtual operating environment for reinforcement learning algorithms. It needs to accurately map the core equipment characteristics and system coupling relationship of the physical power plant, provide a high-fidelity simulation basis for the exploration, training and verification of control strategies, and ensure that the generated strategies can be directly adapted to the actual operation of the power plant.

[0074] In some implementations, the process of building a digital twin model may include the following steps: First, 3D modeling software is used to construct geometric models of boilers, steam turbines, and environmental protection equipment to restore the equipment's shape, installation location, and pipeline connection relationships, ensuring that the spatial layout is consistent with the physical power plant. Then, based on the equipment mechanism equations and historical operating data, a mathematical model reflecting the dynamic characteristics of the equipment is constructed. For example, on the boiler side, a heat balance equation and a steam-water two-phase flow model can be used to simulate the changes in main steam pressure and temperature with fuel quantity and air volume; on the turbine side, a variable operating condition equation can be used to simulate the changes in power with steam parameters and valve opening; and on the environmental protection equipment side, a reaction kinetic model can be used to simulate the changes in nitrogen oxide concentration with ammonia injection quantity. Next, the physical relationships between the various devices are clarified. For example, the boiler and the steam turbine are coupled through "main steam pressure-temperature-flow rate", and the combustion system and the environmental protection system are coupled through "flue gas-composition-flow rate". The coupling relationship is then transformed into mathematical constraints and written into the model logic.

[0075] S302. Set the environment and reward function for the reinforcement learning agent.

[0076] In this context, the environment of the reinforcement learning agent defines the state perceived by the agent and the actions that can be performed, while the reward function guides the agent to explore strategies that meet the goals of economy, low carbon emissions, and safety. Together, they constitute the core interactive framework of reinforcement learning.

[0077] In some implementations, the environment settings and reward function definition can be as follows: 1. State Space: Contains key operating data of the power plant, including but not limited to: unit load P load Main steam pressure P main Main steam temperature T main Reheat steam temperature T reheat Furnace negative pressure P furnace Oxygen content (O2) and NOx concentration (C) in flue gas NOx Carbon dioxide emission intensity C CO2 The state space is represented in vector form as S=[P load ,P main ,T main ,T reheat ,P furnace O2,C NOx C CO2 All parameters can be normalized to the [0,1] interval, eliminating dimensional differences.

[0078] 2. Operation Space: Includes continuously adjustable equipment operation variables, covering combustion control, steam and water control, and environmental control, specifically including but not limited to: equipment speed command N. fan Boiler total air volume setpoint Q air Steam turbine control valve opening command μ turbineFlue gas recirculation damper opening μ recirc ammonia injection regulating valve opening degree μ ammonia The action space is represented in vector form as A=[N] fan Q air ,μ turbine ,μ recirc ,μ ammonia Each action requires physical constraints, such as μ. turbine ∈[30%,100%], to avoid unit instability caused by excessively small valve opening.

[0079] 3. Reward Function R: Directly adopt the adjusted global optimization objective function from S2 to ensure that the agent's exploration direction is consistent with the power plant's multi-objective optimization requirements. The formula is: R = When an agent performs a certain action, if it can improve... ,reduce And none If the reward value R increases, the agent is encouraged to retain that type of action; conversely, if the reward value R decreases, the agent is prompted to adjust its actions.

[0080] It should be noted that the state space needs to be synchronized with the operating data of the digital twin model in real time to ensure that the state perceived by the agent is consistent with the operating conditions of the virtual power plant; the constraints of the action space need to be strictly matched with the physical limits of the equipment to prevent the agent from exploring actions that exceed the carrying capacity of the equipment.

[0081] S303. Construct an actor-critic network architecture and train a policy network.

[0082] Among them, the Actor-Critic network architecture is the core algorithm framework for implementing reinforcement learning. The Actor network is responsible for generating control actions, and the Critic network is responsible for evaluating the value of the actions. The two work together to iterate and train, eventually obtaining a converged policy network, which ensures that the generated control actions are stable and optimal.

[0083] In some implementations, the network architecture construction and training process is as follows: First, an actor-critic network is constructed. The Actor network takes a state space vector S as input and outputs an action space vector A. It outputs a continuous control action A that conforms to physical constraints based on the current state S, i.e., A = Actor(S). The Critic network takes the concatenated vector [S,A] of the state and action spaces as input and outputs the action value function Q, reflecting the long-term cumulative reward expectation of the current state-action pair. This value is used to evaluate the value of the Actor network's output action and provides the gradient direction for updating the Actor network parameters, i.e., Q = Critic(S,A).

[0084] Next, the Deep Deterministic Policy Gradient (DDPG) algorithm is employed to improve training stability through experience replay and the target network mechanism. The core update formulas in the training process include: Target Q value calculation: ,in For the target Q value, Let R be the reward value at the current time t, γ be the discount factor with a value of 0.95, Q' be the output of the target Critic network, and A' be the output of the target Actor network. Based on the current reward and the value of the future state, the expected value of the current action is calculated to provide training labels for the Critic network.

[0085] Calculation of Critic network loss L: Where N is the experience playback batch size, The current output of the Critic network optimizes the value assessment accuracy of the Critic network by minimizing the mean square error between the current Q value and the target Q value.

[0086] Actor network gradient calculation: Where θ represents the Actor network parameters. This formula indicates that by updating the Actor network parameters along the direction of increasing value evaluation by the Critic network, the Actor can generate better actions.

[0087] Finally, the digital twin model is used as the training environment, and the agent interacts with the environment to collect experience (S). t A t ,r t ,S t+1 The data is then stored in the experience replay pool. After collecting 1000 experiences, N experiences are randomly sampled from the pool to update the network parameters. The number of training iterations is set to 10000. When the loss value of the Critic network is less than 0.01 for 100 consecutive iterations and the fluctuation of the action output by the Actor network is less than 5%, the policy network is determined to have converged, and training is stopped.

[0088] S304. Online learning generates control strategies and Monte Carlo simulation is used to select the optimal strategy.

[0089] In some implementations, the specific operation flow of step S304 is as follows: First, the real-time operating data of the power plant is input into the convergent strategy network obtained by S303. The network outputs the control action at the current moment based on the real-time state. At the same time, the real-time data and actions are fed back to the network, and the network parameters are fine-tuned through online gradient descent so that the strategy can adapt to fluctuations in operating conditions.

[0090] For each control strategy generated, it is injected into the digital twin model and subjected to N (N≥100) stochastic simulations. During each simulation, small random perturbations are added to the power plant operating parameters of the model to simulate uncertainties in actual operation. The control strategy includes control actions for one future scheduling period, such as a 24-hour continuous action sequence A. seq =[A1,A2,...,A24].

[0091] Then, the model evolution is advanced using the variable step size Long-Gekutta method, and the cumulative reward value of each simulation is recorded. Where k=1,2,...,N, and γ=0.95 is the discount factor.

[0092] Finally, calculate the expected cumulative reward value for N simulations. The strategy with the highest expected cumulative reward value is determined as the optimal control strategy. If the expected reward values ​​of multiple strategies differ by less than 1%, then strategy P is preferred. violation The strategy of setting the value to 0 prioritizes safety.

[0093] Based on the above technical solutions, S3 uses a digital twin model combined with an Actor-Critic architecture and the DDPG algorithm to ensure that the strategy can adapt to continuous action space and multi-objective requirements; and through uncertainty analysis of Monte Carlo simulation, it selects stable strategies that take into account economy, low carbon emissions and safety.

[0094] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 4 As shown, the above S4 specifically includes the following S401 to S404: S401. Define the cost function and coupling potential function for the edge controller of each sub-region.

[0095] In some implementations, the cost function is formulated as follows: The control cost of the i-th edge controller is quantified from three dimensions to guide the controller in balancing target tracking, reducing motion energy consumption, and maintaining sub-regional coordination. J i This represents the total cost of the i-th edge controller; a smaller value indicates a better control scheme. u i This indicates the action of the i-th controller, such as the total air volume command for the boiler sub-zone or the valve opening command for the turbine sub-zone. This represents the set of actions for all sub-region controllers except for the i-th controller; This indicates that the i-th controller performs action u. iThe actual output vector, such as the main steam pressure and temperature of the boiler sub-region; This represents the target parameters issued to the i-th controller in the optimal control strategy; Used to output the tracking error term, Q i To track the error weight matrix, this term penalizes the deviation between the actual output and the target. R represents the energy consumption term of the action. i This is the action weight matrix, and this item penalizes aggressive control actions; N(i) represents the set of sub-region controllers adjacent to the i-th controller; ρ ij This represents the coupling strength coefficient between the sub-regions to which the i-th and j-th controllers belong. It is set based on the strength of the physical association, such as the close association between the boiler and the steam turbine due to steam supply and demand. ij =0.8; Combustion and environmental protection are second only to flue gas transmission, ρ ij =0.5; This represents the coupling potential function, which quantifies the dynamic coupling relationship between the i-th and j-th sub-regions. The smaller the value, the better the synergy.

[0096] In some implementations, the coupling potential function is defined differently based on the physical coupling characteristics between sub-regions, transforming physical association requirements into mathematical constraints. For example, The coupling potential function between the boiler control system and the turbine control system: This is used to quantify the matching degree between boiler output steam parameters and turbine inlet requirements, avoiding turbine efficiency reduction or equipment damage caused by steam pressure / temperature deviations; among which, The actual value of the main steam pressure. Set the turbine inlet pressure value. The actual value of the main steam temperature. Set the turbine inlet temperature setpoint. This is the weighting coefficient for the temperature coupling term, which is usually set to 0.6. However, since pressure has a greater impact on turbine safety, the weighting of the pressure term is set to 1 by default.

[0097] The coupling potential function between the combustion control system and the environmental protection control system: This is used to quantify the matching degree between pollutants generated by combustion and the treatment capacity of environmental protection equipment, to avoid pollutant concentrations exceeding the treatment limit of environmental protection equipment, thus preventing emissions from exceeding standards; among which, The concentration of NOx produced by combustion. To maximize the denitrification capacity of the environmental protection system, SO2 concentration produced by combustion To maximize the desulfurization capacity of the environmental protection system, This is the weighting coefficient for the SO2 coupling term, usually set to 0.7. Because NOx emission limits are more stringent, the NOx term weight is 1 by default.

[0098] S402. Initialize control actions and set iteration parameters, including: Each sub-region edge controller, based on the optimal control strategy Based on the equipment's safe operating range, initial actions are randomly generated. ; Convergence threshold This is set as the minimum acceptable range for the change in action, typically between 0.01 and 0.05, when the update norm of all controller actions is less than... At that point, it was considered that the movement had stabilized.

[0099] S403. Iteratively solve for the optimal control action of each controller.

[0100] In some implementations, iterative solving needs to be performed according to the process of synchronous update and local optimization, as follows: From the first iteration to the kth iteration (k≥1), each controller i fixes the (k-1)th iteration action of all other controllers. Only optimize its own actions , making the cost function Minimize. The optimization problem is expressed as: ;in, This represents the minimum value of the control parameter. Indicates the maximum value of the control parameter; Since the cost function is a quadratic continuously differentiable function, the gradient descent method is used to solve the above optimization problem. The steps are as follows: Calculate the cost function for u i gradient ; According to the formula Update action; where μ is the learning rate, ranging from 0.01 to 0.05; If the updated Exceeding Then it is cut off to the boundary, such as At that time, take .

[0101] All controllers must calculate the k-th iteration action simultaneously after all actions in the (k-1)-th iteration are determined. This synchronous iteration avoids information asymmetry caused by some controllers updating prematurely.

[0102] S404, Determine if the iteration converges and output the cooperative control strategy.

[0103] In some implementations, the steps for convergence determination and instruction output may include: Define the update norm for the k-th iteration. The sum of the L2 norms of the changes in the actions of all controllers in two consecutive iterations is given by the following formula: ; The iteration is considered convergent if either of the following two conditions is met during the iteration process: Condition 1: Update norm This indicates that all controller actions have stabilized, and further iterations will not yield significant optimization; among them, The convergence threshold set for S402; Condition 2: The number of iterations k = the maximum number of iterations K max This indicates that negotiations should be terminated even if the actions are not fully stable.

[0104] When the iteration converges, output the current action combination of each controller. This combination represents the Nash equilibrium solution of the potential game—for each controller i, with the actions of other controllers remaining unchanged, individually changing... This will cause J i The value increases, therefore this solution is the optimal combination of instructions for global coordination and can be directly issued to each sub-region for execution.

[0105] Based on the above technical solutions, S4 ensures that collaboration does not conflict through cost functions and coupling potential functions, avoids information asymmetry through synchronous iteration, and the obtained Nash equilibrium solution can guarantee the optimality of instructions. Compared with traditional centralized control, it can reduce instruction transmission delay and adapt to the core requirement of autonomous collaboration without human intervention in scenarios with few people on duty, thus solving the problem of operational conflicts caused by independent control of sub-regions.

[0106] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 5 As shown, the above S5 specifically includes the following S501 to S503: S501. The collaborative control strategy is simulated and verified using a pre-built digital twin model.

[0107] The simulation and verification of digital twin models is a core step in previewing the operational effects of collaborative control strategies in actual power plants. By reproducing real operating scenarios and predicting parameter change trajectories, it is possible to identify in advance whether the strategy may have issues such as exceeding safety limits, economic losses, or exceeding environmental standards, thus avoiding equipment risks or performance losses caused by the direct implementation of the strategy.

[0108] In some implementations, the specific operations of simulation verification may include: The initial operating state of the digital twin model is set to be completely consistent with that of the real power plant, and the real-time collected power plant operating data is directly assigned to the corresponding parameters of the model to ensure that the simulation starting point is without deviation from the physical power plant. The sequence of cooperative control strategies output by S4 is used as input and injected into the initialized digital twin model in chronological order. Numerical integration using the variable-step-size Runge-Kutta method is employed to advance the dynamic evolution of the digital twin model. By discretizing the equipment mechanism equations, parameter values ​​at each time point within future timeframes are calculated step by step, ultimately yielding the trajectory of key operating parameters. These key operating parameters include at least: unit load, main steam pressure, main steam temperature, nitrogen oxide emission concentration, carbon dioxide emission concentration, and furnace negative pressure.

[0109] Based on the parameter change trajectory, safety indicators, economic indicators, environmental indicators, and stability indicators are calculated respectively. The calculation logic for each indicator is as follows: Safety indicators include the probability and maximum extent of exceeding limits for main steam pressure, main steam temperature, and furnace negative pressure. The probability of exceeding limits is calculated as: (Number of parameter exceedances / Total number of simulation time steps); the maximum extent of exceeding limits is calculated as: max(Actual parameter value - Upper limit of safety threshold, Lower limit of safety threshold - Actual parameter value).

[0110] Economic indicators: Power generation revenue is calculated based on the unit load change trajectory, combined with fuel costs and operation and maintenance costs, and can be calculated according to E in S2. economic The formula calculates that a larger value indicates better economic efficiency. Environmental indicators include NOx emission concentration and carbon emission intensity, both of which must be lower than national or local emission standards; Stability index: Calculate the variance of fluctuations for each key operating parameter, using the following formula: ;in The parameter value at time t. The variance represents the mean of the parameters over the simulation period, where n is the number of time steps. The smaller the variance, the more stable the parameters are.

[0111] A threshold is set for each performance indicator. If the calculated values ​​of all indicators are less than (or meet) the corresponding threshold, the collaborative control strategy is deemed to have passed verification. If any indicator does not meet the threshold, the verification is deemed to have failed, and the process must return to S3 to regenerate the control strategy.

[0112] S502. Construct a risk assessment model based on evidence theory and calculate the risk index of the current strategy.

[0113] Among them, the risk assessment model integrates four types of evidence sources: simulation verification results, equipment health, operation history, and environmental data, to quantify the potential risks of collaborative control strategies and provide objective risk basis for subsequent human-machine decision-making.

[0114] In some implementations, the construction of the risk assessment model and the calculation of the risk index include the following steps: First, define the identification framework and evidence sources: Define the identification framework θ = low risk (L), medium risk (M), and high risk (H).

[0115] Define evidence sources: k=1 for digital twin verification results, k=2 for device health status, k=3 for operation history, and k=4 for environmental data. Each evidence source must provide a basic probability allocation for the risk level in θ.

[0116] Next, the basic probability allocation function m for each source of evidence is constructed. k (A); among them, A θ, m k (A) represents the support of evidence k for risk level A, taking values ​​[0,1], and the sum of all mk(A) is 1: The basic probability allocation function for the first source of evidence is: Based on four performance metrics of S501, the degree to which the simulation verification results support the risk level is determined. Among them, N... i Let be the normalized value of the i-th performance metric. The safety threshold for the i-th performance metric. The warning threshold I(·) for the i-th performance metric is an indicator function, which is 1 if the condition in parentheses is met, and 0 otherwise.

[0117] The basic probability allocation function for the second source of evidence is: Based on the remaining lifespan of the equipment and the fault warning level, assess the impact of equipment operating risk on the strategy; where L remain For the remaining lifespan of the equipment, L design For the equipment design life, F level The warning level is the fault warning level, and k is an adjustment parameter, usually set to 0.8-1.2, used to control the impact of the warning level on the probability.

[0118] The basic probability allocation function for the third source of evidence is: Based on the historical success rate of fully automated decision-making and operator experience, assess the risks associated with the need for manual intervention. Among these, P... success E represents the frequency of using the fully automated decision-making mode within a pre-defined historical time period. level This indicates the operator's experience level, and 'a' is an adjustment parameter ranging from 0.5 to 1, used to control the impact of historical data on probability.

[0119] The basic probability allocation function for the fourth source of evidence is: Based on the rate of change and current values ​​of four environmental parameters—temperature, humidity, air pressure, and wind speed—the risk of environmental disturbances to strategy execution is assessed. Among these, ΔE...j ΔE represents the rate of change of the j-th environmental parameter. j0 E represents the threshold value for the rate of change of the j-th environmental parameter. j E represents the current value of the j-th environment parameter. j0 b represents the pre-specification of the j-th environmental parameter. j c j The slope parameter is represented by the environmental parameters, which include temperature, humidity, air pressure, and wind speed.

[0120] Then, the Dempster combination rule is used to synthesize evidence. For two evidence sources m1 and m2, the basic probability assignment function m(A) after synthesis is formulated as follows: Then, based on this formula, the four evidence sources are fused in sequence: first, m1 and m2 are fused; then the result is fused with m3; and finally the result is fused with m4 to obtain the final synthesized m(A).

[0121] Finally, calculate the risk index R: First, calculate the trust function Bel(A), the formula is as follows: , representing the total trust level in A, i.e., all subsets B The sum of m(B) of A; Next, calculate the risk index R to quantify the overall impact of medium to high risk: R = 0.5 × Bel(H) + 0.5 × Bel(M), and the larger the R, the higher the strategy risk.

[0122] S503. Calculate the policy credibility score and construct a multi-level decision-making mechanism.

[0123] Among them, the strategy credibility score integrates risk level and performance level to quantify the executability of the strategy; the multi-level decision-making mechanism dynamically allocates human and machine decision-making power based on the credibility score, adapting to the needs of safety priority and efficiency consideration in scenarios with few people on duty.

[0124] In some implementations, step S503 is performed as follows: (1) The calculated values ​​of the four performance indicators in S501 are averaged and weighted to obtain the strategy performance score. The formula is: ;in, This represents the actual calculated value of the i-th performance index. The maximum value of the indicator. The minimum value of the indicator. The weights for each indicator are as follows: safety is typically weighted at 0.4, while economy, environmental friendliness, and stability are each weighted at 0.2.

[0125] (2) Calculate the strategy credibility score C using the following formula: ;in, This indicates a preset risk threshold. The baseline strategy score is represented by the weighted average sum of the threshold values ​​for each performance metric.

[0126] (3) Construct a multi-level decision-making mechanism: Based on the strategy credibility score C, set two thresholds, such as the first threshold C1 and the second threshold C2, usually C1=0.8 and C2=0.5, and divide into three decision-making modes: Fully automatic decision-making mode: When C>C1, the strategy has low risk and excellent performance. No manual intervention is required. The collaborative control strategy is automatically distributed to each sub-area for execution. It is suitable for scenarios with stable operating conditions and controllable risks. Human-machine collaborative decision-making mode: When C2 < C ≤ C1, the strategy needs to be confirmed manually. The control strategy is pushed to the central control platform, and the operator can view the parameter change trajectory. After confirming that there are no problems, click "execute", or adjust the instructions after discovering risk points. It is suitable for scenarios with small fluctuations in operating conditions and medium risks. Manual decision-making mode: When C≤C2, the strategy is high-risk or has poor performance, and execution is automatically suspended, triggering an alarm on the central control platform. Operators need to re-evaluate the strategy or manually formulate control instructions. This mode is suitable for extreme working conditions.

[0127] Based on the above technical solutions, S5 constructs a security redundancy and decision-making guarantee before strategy execution through digital twin pre-verification, evidence fusion risk quantification, and multi-level human-machine decision-making processes. Among them, digital twin simulation avoids the risk of physical trial and error, evidence theory risk assessment integrates multi-source data to ensure comprehensive risk judgment, and the multi-level decision-making mechanism accurately adapts to the manpower allocation needs of scenarios with few people on duty. This reduces ineffective human intervention and eliminates the blind execution of high-risk strategies, ultimately achieving a safe, efficient, economical, and environmentally friendly collaborative control closed loop.

[0128] The above primarily describes the solutions of the embodiments of this application from the perspective of device implementation. It is understood that each device, for example, a dynamic collaborative control device for a smart power plant with minimal human intervention, includes at least one of the hardware structures and software modules corresponding to the execution of each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0129] In implementation, each step of the method provided in this embodiment can be completed by integrated logic circuits in the processor or by instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0130] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0131] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and modifications.

Claims

1. A dynamic collaborative control method for a smart power plant with less manned, characterized in that, The utility model relates to a power plant operation optimization method and system based on meta-cognition and reinforcement learning, comprising: Real-time acquisition of power plant operation data and real-time operation data; Based on historical and real-time operation data, the meta-cognition optimization model is used to adjust the weight coefficients of the global optimization objective function; The global optimization objective function is used to solve the multi-objective optimization problem of power plant operation economy and carbon emission intensity; Using the adjusted global optimization objective function, the optimal control strategy is generated in the pre-constructed digital twin model through the reinforcement learning algorithm; Each sub-region carries out distributed negotiation through potential game mechanism according to the optimal control strategy, and the collaborative control strategy of each sub-region is obtained; The collaborative control strategy is simulated and verified by using the pre-constructed digital twin model, and risk assessment and strategy decision are made in combination with the risk assessment model; the risk assessment model is a human-machine decision weight distribution model obtained by evidence fusion algorithm.

2. The dynamic collaborative control method for less manned operation of a smart power plant according to claim 1, characterized in that, The utility model relates to a power plant operation optimization method and system based on meta-cognition and reinforcement learning, comprising: According to the power generation, fuel cost and operation and maintenance cost data, the economic indicators of power plant operation are calculated; According to the carbon emission monitoring data, the carbon emission intensity indicators of power plant operation are calculated; The system constraint violation penalty term is set according to the equipment safe operation rules, and the global optimization objective function and the meta-cognition reward function are constructed in combination with the economic indicators, the carbon emission intensity indicators and the system constraint violation penalty term; the weight coefficients of the meta-cognition reward function are consistent with the weight coefficients of the global optimization objective function, and the meta-cognition reward function is a reward function constructed based on the weighted integral of the global optimization objective function in the historical time window, which is used to evaluate the long-term performance of the control strategy; Based on gradient sensitivity analysis, the meta-cognition reward function is optimized by time integral to obtain the global optimization objective function with adaptive weight adjustment.

3. The dynamic collaborative control method for less manned operation of a smart power plant according to claim 2, characterized in that, The calculation formula of the global optimization objective function Objective is: ; wherein, , represents an economic indicator, , represents carbon intensity; represents a constraint violation penalty term, w i represents the penalty weight of the i-th parameter out-of-limit, y i represents the i-th device operating parameter, represents the safety threshold of the i-th real-time operating data, δ i is the safety margin of the i-th real-time operating data, α, β, γ represent the weight coefficients of each term; The calculation formula of the metacognition reward function is: ; wherein, λ represents a discount factor, used to adjust the decay rate of historical data, t represents the current time, τ represents the integral variable time, and T represents the integral time window length.

4. The dynamic collaborative control method for less manned operation of a smart power plant according to claim 1, characterized in that, The utility model relates to a power plant operation optimization method and system based on meta-cognition and reinforcement learning, comprising: The digital twin model of power plant is constructed; the digital twin model can at least map the dynamic characteristics and coupling relationship of boiler, steam turbine and environmental protection equipment; The adjusted global optimization objective function is used as the reward function of the reinforcement learning agent to explore the strategy in the digital twin environment; The environment of the reinforcement learning agent is set; wherein, the action space of the environment includes device speed instruction, total air volume set value, steam turbine adjusting door opening degree instruction, flue gas recirculation damper opening degree, ammonia injection regulating valve opening degree, the state space includes unit load, main steam pressure, main steam temperature, reheat steam temperature, furnace negative pressure, flue gas oxygen content, nitrogen oxide concentration and carbon dioxide emission intensity; The actor-critic network architecture is constructed; wherein, the actor network takes the state space as input and outputs continuous control action, and the critic network takes the splicing vector of state space and action space as input and outputs the Q value representing the action value function; The parameters of the actor network and the critic network are iteratively trained and verified by using the deep deterministic policy gradient algorithm to obtain the converged strategy network. The real-time power plant operation data is input into the converged strategy network for online learning and strategy generation, and each time a control strategy is generated, a Monte Carlo simulation is performed in the digital twin model to calculate the expected cumulative reward value of the control strategy, and the control strategy with the highest expected cumulative reward value is selected as the optimal control strategy.

5. The dynamic collaborative control method for less manned operation of a smart power plant according to claim 1, characterized in that, The distributed negotiation through the potential game mechanism obtains the collaborative control strategy of each sub-region, including: Define a cost function for each edge controller of each sub-region: ; where J i represents the total cost of the i-th edge controller, represents the action of the i-th controller, represents the action set of other controllers except controller i, N(i) represents the set of controllers adjacent to controller i, Φ ij is a potential function representing the dynamic coupling relationship between the subsystems to which controller i and controller j belong, y i represents the output control vector of the i-th controller, Q i represents the weight matrix of the output tracking error, R i represents the weight matrix of the control action, y i ref represents the control parameter issued to the i-th controller in the optimal control strategy, ρ ij represents the coupling strength coefficient, which is determined according to the physical association strength between sub-regions; Each controller randomly generates an initial control action, and sets a convergence threshold and a maximum number of iterations; In each iteration, each controller fixes the control actions of the other controllers and solves for the optimal control action that minimizes the cost function; wherein the optimization problem solved is formulated as finding the control action u that minimizes the cost function subject to the constraints i ; When the update norm of the control action of all controllers is less than the convergence threshold, or the number of iterations reaches the maximum number of iterations, the iteration is stopped, and the current combination of control actions of each controller is output as the Nash equilibrium solution of the potential game, which is the collaborative control strategy of each sub-region; the update norm represents the sum of the two norms of the control action variation between adjacent two iterations of all controllers.

6. The dynamic collaborative control method for less manned operation of a smart power plant according to claim 1, characterized in that, The simulation verification of the collaborative control strategy by using the pre-constructed digital twin model includes: The initial state of the digital twin model is set to be consistent with the current operation state of the real power plant; The control instruction sequence of the collaborative control strategy is input into the initialized digital twin model as input; the control instruction sequence includes a complete future scheduling period; A variable step size Runge-Kutta method is used for numerical integration to advance the dynamic evolution of the digital twin model and predict the change trajectory of the key operating parameters in the future scheduling period; the key operating parameters at least include unit load, main steam pressure, main steam temperature, nitrogen oxide emission concentration, carbon dioxide emission concentration, and furnace negative pressure; Based on the change trajectory of the key operating parameters, the values of the safety index, the economy index, the environmental protection index, and the stability index are calculated respectively; wherein, The safety index represents the probability and maximum out-of-limit amplitude of the main steam pressure, the main steam temperature, and the furnace negative pressure, the unit load is used to calculate the power generation income in the economy index, the environmental protection index includes the nitrogen oxide emission concentration and the carbon emission intensity, and the stability index represents the fluctuation variance of each key operating parameter; A threshold is set for each performance index, and when the calculated values of each performance index are less than the corresponding threshold, it is determined that the collaborative control strategy passes the verification; otherwise, it is determined that the collaborative control strategy fails the verification.

7. The dynamic collaborative control method for less manned operation of a smart power plant according to claim 6, characterized in that, The risk assessment and strategy decision-making combined with the risk assessment model include: A risk assessment model based on evidence theory is constructed based on the digital twin verification result, the equipment health state, the operation history record, and the environmental data to obtain the risk index under the current collaborative control strategy; The calculated values of each performance index in the simulation verification process are averaged and weighted to obtain a strategy performance score; The strategy credibility score is calculated based on the risk index and the strategy performance score; wherein, the calculation formula of the strategy credibility score C is: C=w1×(1-risk index / preset risk threshold)+w2×(strategy performance score / benchmark score), and the benchmark score represents the average weighted sum of the set thresholds of each performance index; A multi-level decision mechanism is constructed based on the policy credibility score to determine the proportion of human-machine decision at different levels.

8. The dynamic collaborative control method for less manned operation of a smart power plant according to claim 7, characterized in that, The calculation formula of the risk assessment model is: The identification framework is defined as θ = low risk, medium risk, and high risk; Constructing basic probability assignment functions for each evidence source ; wherein k = 1, 2, 3, 4, respectively corresponding to the digital twin verification result, the equipment health status, the operation history record, and the environmental data four data sources; The basic probability distribution of four evidence sources is synthesized by using Dempster combination rule to obtain the synthesized basic probability distribution function m(A); wherein A θ; A risk index R is calculated according to the formula R = 0.5 x Bel(H) + 0.5 x Bel(M); where Bel is a belief function, H represents high risk in the identification framework, and M represents medium risk in the identification framework.

9. The dynamic collaborative control method for less manned operation of a smart power plant according to claim 7, characterized in that, The multi-level decision mechanism includes: When the policy credibility score is greater than a first threshold, an automatic decision mode is adopted to automatically execute the collaborative control strategy; When the policy credibility score is less than or equal to the first threshold and greater than a second threshold, a human-machine collaborative decision mode is adopted to push the collaborative control strategy to a central control platform for manual confirmation of execution or denial of execution of the collaborative control strategy; When the policy credibility score is less than or equal to the second threshold, a manual decision mode is started.

10. A dynamic collaborative control system for less manned operation of a smart power plant, characterized in that, It includes: a data acquisition module, a target optimization module, a strategy generation module, a collaborative negotiation module, and a verification and evaluation module, wherein The data acquisition module is configured to acquire real-time operation data of the entire power plant and each sub-region to form a historical and real-time operation data set; The target optimization module is configured to adjust the weight coefficient of a global optimization objective function using a meta-cognition optimization model based on the data acquired by the data acquisition module to achieve dynamic balance of economic efficiency and carbon emission intensity multi-objectives; The strategy generation module is configured to generate an optimal control strategy in a pre-constructed digital twin model based on the adjusted global optimization objective function through a reinforcement learning algorithm; The collaborative negotiation module is configured to enable each sub-region to perform distributed negotiation through a potential game mechanism based on the optimal control strategy to obtain a collaborative control strategy of each sub-region; The verification and evaluation module is configured to simulate and verify the collaborative control strategy using the digital twin model and perform risk assessment and strategy decision-making in combination with a risk assessment model constructed using an evidence fusion algorithm.

Citation Information

Patent Citations

  • Reinforced learning control method and control system suitable for combined power and heat supply type virtual power plant

    CN116859739A

  • Intelligent thermal power plant layered optimization control method based on multiple agents

    CN120044789A

  • Industrial park multi-target collaborative optimization scheduling system and method based on artificial intelligence

    CN120146482A

  • Virtual power plant green energy consumption cooperation method and system fusing digital twinning and reinforcement learning

    CN120409220A

  • Confrontation game cooperative control method and system for heterogeneous unmanned ship cluster

    CN120447543A

Cited By

  • Multi-point linkage intelligent control method and system for coal-fired boiler

    CN122041169A