An artificial intelligence-based operation control optimization method and system for a thermal power plant

Through multi-agent game reinforcement learning and dynamic benchmarking modeling, the operation control system of thermal power plants has achieved synergistic optimization of energy consumption, carbon emissions and economic efficiency, solving the problems of high energy consumption, lagging carbon emissions and difficulty in balancing economic efficiency in traditional control systems, and improving operation efficiency and stability.

CN121300048BActive Publication Date: 2026-06-02BEIJING GREEN POWER FRESH TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING GREEN POWER FRESH TECHNOLOGY CO LTD
Filing Date
2025-11-04
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Thermal power plant operation and control systems struggle to dynamically adapt to optimal control strategies based on load changes, coal type fluctuations, and environmental disturbances. This results in high energy consumption, lagging carbon emission control, and difficulty in balancing economic efficiency. The lack of full-process modeling capabilities based on multi-source operation data and multi-dimensional evaluation mechanisms prevents the achievement of the synergistic optimization goals of improving energy efficiency, reducing carbon emissions, and balancing economic performance.

Method used

By integrating multi-agent game reinforcement learning and dynamic benchmarking, a benchmark database is constructed by collecting operating data of thermal power plant units, generating real-time operating condition deviation vectors. Multi-agent game reinforcement learning is used to construct an intelligent optimization network, generating comprehensive operation adjustment strategies to achieve coordinated optimization of energy consumption, carbon emissions, and economic efficiency.

Benefits of technology

It has achieved multi-objective synergistic optimization of energy consumption control, carbon emission regulation and economic improvement of thermal power units, improved the responsiveness and stability of operation strategies, reduced energy consumption and carbon emissions, and improved economic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121300048B_ABST
    Figure CN121300048B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on artificial intelligence's thermal power plant operation control optimization method and system, including following steps: the operation data of thermal power plant unit is collected, and operation dataset is constructed;The operation dataset is preprocessed;Performance calculation is carried out, and index set is generated;According to index set, sample is filtered from operation dataset, benchmark value set is constructed, and dynamic benchmark value database is established;Comparison analysis is carried out, and the multi-dimensional deviation between current working condition and optimal working condition is calculated;Through the intelligent optimization network constructed based on multi-agent game reinforcement learning, strategy game update is executed, and the optimization instruction set is obtained through game balance process;Dynamic benchmark value database is revised, and the multi-objective operation control optimization closed-loop management of thermal power unit is realized.The application fuses multi-agent game reinforcement learning and dynamic benchmark modeling, realizes thermal power unit energy consumption, carbon emission and economic synergy optimization, with the advantages of strong intelligence, high energy efficiency and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of operation control optimization, and in particular to an artificial intelligence-based method and system for optimizing the operation control of thermal power plants. Background Technology

[0002] Currently, the operation and control systems of thermal power plants mainly rely on fixed threshold strategies or empirical rule bases to guide unit regulation. In actual operation, it is difficult to dynamically adapt to the optimal control strategy according to load changes, coal type fluctuations and environmental disturbances, resulting in high energy consumption, lagging carbon emission control and difficulty in balancing economic efficiency. Although some methods based on expert systems or fuzzy control have introduced a certain degree of intelligent regulation capability, they still lack the ability to deeply model and dynamically feedback complex operating data, making it difficult to support fine-grained adjustments under multi-objective constraints.

[0003] To address the aforementioned issues, existing technologies generally lack the ability to model the entire process based on multi-source operational data, lack a multi-dimensional evaluation mechanism that combines energy consumption, carbon emissions, and economic indicators, and have not effectively introduced reinforcement learning and game theory mechanisms to achieve dynamic optimization and adaptive game-theoretic updates of operational strategies. Consequently, they are unable to achieve the synergistic optimization goals of improving energy efficiency, reducing carbon emissions, and achieving economic balance. Summary of the Invention

[0004] One objective of this invention is to propose an artificial intelligence-based method and system for optimizing the operation control of thermal power plants. This invention integrates multi-agent game-theoretic reinforcement learning and dynamic benchmarking modeling to achieve coordinated optimization of energy consumption, carbon emissions, and economic efficiency of thermal power units, and has the advantages of strong intelligence, high energy efficiency, and excellent stability.

[0005] An artificial intelligence-based method for optimizing the operation control of a thermal power plant, according to an embodiment of the present invention, includes the following steps:

[0006] Collect operational data of thermal power plant units under different loads, coal types, and environmental conditions to construct an operational dataset;

[0007] The runtime dataset is preprocessed to generate a preprocessed runtime dataset;

[0008] The preprocessed running dataset is used to perform performance calculations to obtain a set of equipment performance indicators, and then a set of energy consumption indicators, a set of carbon emission indicators, and a set of economic indicators are generated based on the set of equipment performance indicators.

[0009] Based on the sets of energy consumption indicators, carbon emission indicators, and economic indicators, samples are selected from the operational dataset, corresponding operational data are extracted, a set of benchmark values ​​is constructed, and a dynamic benchmark value database is established.

[0010] The real-time operating data of the unit is compared and analyzed with the set of benchmark values ​​in the dynamic benchmark value database to calculate the multi-dimensional deviation between the current operating condition and the optimal operating condition, and generate a real-time operating condition deviation vector.

[0011] The real-time operating condition deviation vector is input into an intelligent optimization network constructed based on multi-agent game reinforcement learning, and the strategy game update is performed. Through the game equilibrium process, a comprehensive operation adjustment strategy vector is generated to obtain the optimized instruction set.

[0012] Collect feedback data after operational adjustments, correct the dynamic benchmark value database, and realize multi-objective operation control optimization closed-loop management of thermal power units.

[0013] Optionally, the operating data includes boiler system operating data, steam turbine system operating data, air-cooled system operating data, desulfurization and denitrification system operating data, and heating system operating data.

[0014] Optionally, the preprocessing includes time series alignment, noise filtering, outlier removal, missing value imputation, and data normalization.

[0015] Optionally, the set of equipment performance indicators includes combustion thermal efficiency, turbine internal efficiency, cooling efficiency, desulfurization efficiency, and heat transfer efficiency, which correspond to the operating status of the boiler system, turbine system, air-cooled system, desulfurization and denitrification system, and heating system, respectively.

[0016] Optionally, the construction of the dynamic benchmark database specifically includes:

[0017] Based on the energy consumption index set, carbon emission index set, and economic index set, the corresponding energy efficiency index, carbon emission index, and economic index are extracted for each sample in the operational dataset. Joint constraint screening is performed based on the preset energy consumption threshold, carbon emission threshold, and economic constraint limit. When the energy efficiency index of a sample is less than or equal to the energy consumption threshold, the carbon emission index is less than or equal to the carbon emission threshold, and the economic index is not higher than the economic constraint limit, it is determined that the sample meets all three constraint conditions at the same time, and the sample that meets the conditions is selected into the candidate sample set.

[0018] Normalization is performed on the energy efficiency indicators, carbon emission indicators and economic indicators in the candidate sample set to map the indicators with different dimensions to a unified numerical range, and a comprehensive score is calculated based on the weights of the three types of indicators: energy consumption, carbon emission and economy.

[0019] The candidate sample set is sorted from high to low according to the comprehensive score. When the comprehensive score of a sample is within the sample interval of the previous preset proportion, the sample is determined to be the optimal interval sample. All optimal interval samples are combined into a preferred sample set, and the corresponding running data is extracted from the preferred sample set.

[0020] A benchmark value set is constructed based on the operational data in the preferred sample set;

[0021] The benchmark value set is classified and stored according to unit load, coal type and environmental conditions to establish a dynamic benchmark value database.

[0022] Optionally, the generation of the real-time operating condition deviation vector specifically includes:

[0023] Real-time energy consumption indicators, real-time carbon emission indicators, and real-time economic indicators are obtained from real-time operating data to form a set of real-time indicators for the unit.

[0024] The system calls upon the dynamic benchmark database to retrieve a set of benchmark values ​​that match the current unit load, coal type, and environmental conditions.

[0025] The set of real-time indicators of the unit is compared with the set of benchmark values, and the deviation between each indicator is calculated.

[0026] Energy consumption deviation, carbon emission deviation, and economic deviation are normalized and mapped to a unified scale space. A real-time operating condition deviation vector is constructed by splicing these parameters.

[0027] Optionally, the real-time operating condition deviation vector consists of an energy consumption deviation component, a carbon emission deviation component, and an economic deviation component.

[0028] Optionally, the generation of the optimized instruction set specifically includes:

[0029] The real-time operating condition deviation vector is input into an intelligent optimization network constructed based on multi-agent game reinforcement learning. The intelligent optimization network includes an energy consumption agent, a carbon emission agent, and an economic agent. Each agent corresponds to a set of nodes in the network structure. Each node in the set of nodes represents the operating parameters corresponding to the agent's objective. The connection weights between the nodes are calculated, and an adjacency matrix is ​​established based on the connection weights.

[0030] At each decision point, a state dependency matrix is ​​constructed based on the rate of change of energy consumption deviation, carbon emission deviation, and economic deviation, and the adjacency matrix is ​​updated.

[0031] The real-time operating condition deviation vector is used as the initial input to the intelligent optimization network;

[0032] Each agent performs graph convolution operations based on the adjacency matrix to update the policy latent vector;

[0033] The instantaneous reward values ​​of the energy-consuming agent, carbon-emission agent, and economic agent are calculated based on the energy consumption deviation, carbon emission deviation, and economic efficiency deviation, respectively.

[0034] To perform multi-agent policy updates, the agent policy parameter set from the previous decision time step is invoked. The policy gradient is calculated based on the immediate reward values ​​of each agent and the connection weights between neighboring agents. The updated policy parameter set is then performed in the policy parameter space along the reward gradient direction according to the learning rate, resulting in a new agent policy parameter set.

[0035] ;

[0036] in, Represents a node At any moment The strategy parameters, Represents a node At any moment The strategy parameters, Represents the mathematical expectation. Represents the policy function. This indicates calculating the gradient. Indicates the learning rate. Represents a node Instant rewards Represents a node Instant rewards;

[0037] During the strategy update process, a two-layer optimization mechanism of structural evolution and game equilibrium is implemented. Based on the Nash equilibrium condition, the equilibrium point of the multi-agent strategy is determined, and a comprehensive operational adjustment strategy vector is generated.

[0038] An optimized instruction set is generated based on the comprehensive operation adjustment strategy vector.

[0039] Optionally, the optimized instruction set includes fuel distribution adjustment, air flow regulation, ammonia injection rate correction, and heating load distribution.

[0040] An artificial intelligence-based thermal power plant operation control optimization system according to an embodiment of the present invention includes:

[0041] The operation data acquisition module is used to collect operating data from thermal power plant units and build an operating dataset.

[0042] The data preprocessing module is used to preprocess the running dataset;

[0043] The performance calculation module is used to form a set of equipment performance indicators and generate a set of energy consumption indicators, a set of carbon emission indicators, and a set of economic indicators based on the set of equipment performance indicators.

[0044] The benchmark value construction module is used to filter samples from the running dataset, construct a set of benchmark values, and establish a dynamic benchmark value database.

[0045] The working condition deviation calculation module is used to call the set of benchmark values ​​in the dynamic benchmark value database, perform comparative analysis, calculate the multi-dimensional deviation between the current working condition and the optimal working condition, and generate a real-time working condition deviation vector.

[0046] The intelligent optimization module is used to input the real-time operating condition deviation vector into the intelligent optimization network constructed based on multi-agent game reinforcement learning, perform strategy game update, generate a comprehensive operation adjustment strategy vector through the game equilibrium process, and output the optimization instruction set.

[0047] The feedback optimization module is used to collect feedback data after the operation is adjusted and to correct the dynamic benchmark value database.

[0048] The beneficial effects of this invention are:

[0049] This invention proposes an artificial intelligence-based optimization method for the operation control of thermal power plants. It breaks through the technical bottleneck of traditional operation control systems that rely on fixed rules and single optimization objectives, and realizes multi-objective collaborative optimization of energy consumption control, carbon emission regulation, and economic improvement. By collecting historical operation data of thermal power plant units under different loads, coal types, and environmental conditions, a high-quality operation dataset is constructed. On this basis, standardization processing and performance index extraction are performed, which effectively solves the problems of coarse operation status assessment and single control strategy in the prior art. At the same time, by establishing a dynamic benchmark database containing energy consumption, carbon emission, and economic indicators, the unit operation status can be continuously compared with the optimal sample, realizing the quantitative assessment of operating condition deviation.

[0050] Furthermore, this invention introduces a multi-agent game-theoretic reinforcement learning framework to construct an intelligent optimization network. It builds energy consumption agents, carbon emission agents, and economic agents respectively, and executes reinforcement learning strategy updates based on heterogeneous objectives. Through structural evolution and strategy game mechanisms, it achieves Nash equilibrium on a dynamic adjacency graph, generating a comprehensive operational adjustment strategy with global balancing capabilities. This effectively overcomes the problem of uncoordinated objective conflicts in traditional methods. In actual operation, this method can dynamically generate an optimization instruction set based on real-time operating condition deviation vectors, driving each control unit of the unit to perform intelligent adjustments. It also continuously corrects the dynamic benchmark database based on operational feedback data, forming an adaptive learning and closed-loop optimization mechanism, thereby improving the responsiveness and stability of the operational strategy. Attached Figure Description

[0051] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0052] Fig. 1 This is a flowchart of an artificial intelligence-based optimization method for the operation control of thermal power plants proposed in this invention;

[0053] Fig. 2 This is a diagram of the intelligent optimization network structure based on multi-agent game reinforcement learning for the artificial intelligence-based thermal power plant operation control optimization method proposed in this invention. Detailed Implementation

[0054] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0055] refer to Figs. 1-2 An artificial intelligence-based optimization method for the operation control of thermal power plants includes the following steps:

[0056] Collect operational data of thermal power plant units under different loads, coal types, and environmental conditions to construct an operational dataset;

[0057] The runtime dataset is preprocessed to generate a preprocessed runtime dataset;

[0058] The preprocessed running dataset is used to perform performance calculations to obtain a set of equipment performance indicators, and then a set of energy consumption indicators, a set of carbon emission indicators, and a set of economic indicators are generated based on the set of equipment performance indicators.

[0059] Based on the sets of energy consumption indicators, carbon emission indicators, and economic indicators, samples are selected from the operational dataset, corresponding operational data are extracted, a set of benchmark values ​​is constructed, and a dynamic benchmark value database is established.

[0060] The real-time operating data of the unit is compared and analyzed with the set of benchmark values ​​in the dynamic benchmark value database to calculate the multi-dimensional deviation between the current operating condition and the optimal operating condition, and generate a real-time operating condition deviation vector.

[0061] The real-time operating condition deviation vector is input into an intelligent optimization network constructed based on multi-agent game reinforcement learning, and the strategy game update is performed. Through the game equilibrium process, a comprehensive operation adjustment strategy vector is generated to obtain the optimized instruction set.

[0062] Collect feedback data after operational adjustments, correct the dynamic benchmark value database, and realize multi-objective operation control optimization closed-loop management of thermal power units.

[0063] In this embodiment, the operating data includes boiler system operating data, turbine system operating data, air-cooling system operating data, desulfurization and denitrification system operating data, and heating system operating data. It refers to a set of multi-source operating parameters acquired and recorded in real time during the actual operation of the thermal power plant unit, which is used to reflect the overall thermodynamic state, energy transfer process, and pollutant control of the unit. The boiler system operating data includes main steam temperature, main steam pressure, fuel input, and feedwater flow rate. The turbine system operating data includes main steam flow rate, shaft power, and cooling water flow rate. The air-cooling system operating data includes fan speed, inlet air temperature, and outlet air temperature. The desulfurization and denitrification system operating data includes inlet flue gas concentration, outlet flue gas concentration, lime slurry flow rate, and ammonia injection rate. The heating system operating data includes supply and return flow rates and supply and return temperature difference.

[0064] In this embodiment, the preprocessing includes time series alignment, noise filtering, outlier removal, missing value imputation, and data normalization.

[0065] In this embodiment, the set of equipment performance indicators includes combustion thermal efficiency, turbine internal efficiency, cooling efficiency, removal efficiency, and heat transfer efficiency, which correspond to the operating status of the boiler system, turbine system, air-cooled system, desulfurization and denitrification system, and heating system, respectively. The combustion thermal efficiency is used to reflect the conversion level of fuel energy into steam thermal energy, the turbine internal efficiency is used to reflect the conversion capability of steam thermal energy into mechanical work, the cooling efficiency is used to reflect the heat dissipation capability of the air-cooled system, the removal efficiency is used to reflect the removal capability of pollutants, and the heat transfer efficiency is used to reflect the utilization degree of heat energy output.

[0066] The generation of the energy consumption index set specifically includes: calculating the unit energy consumption value based on the ratio of unit output power to fuel input, according to the combustion thermal efficiency of the boiler system, the internal efficiency of the turbine system, the cooling efficiency of the air-cooled system, and the heat transfer efficiency of the heating system, to characterize the fuel energy input required for the unit to generate a unit of electricity; calculating the energy consumption deviation value based on the difference in the time change rate between the combustion thermal efficiency and the cooling efficiency of the air-cooled system, to reflect the degree of deviation of the unit's current energy consumption level from its historical best energy efficiency state; calculating the load response coefficient value based on the product of the heat transfer efficiency of the heating system and the unit load factor, which describes the energy transfer stability of the unit under load fluctuation conditions; and weighting the unit energy consumption value, energy consumption deviation value, and load response coefficient value according to preset weights to obtain the energy consumption index set, which comprehensively characterizes the overall energy efficiency level of the unit. The unit load factor characterizes the ratio of the unit's average power generation load to its rated load within a specified time period, and describes the unit's load utilization rate and load fluctuation characteristics under different operating conditions.

[0067] The generation of the carbon emission index set specifically includes: calculating the carbon emission value per unit of electricity generated based on the product of fuel input and fuel carbon emission coefficient, used to characterize the unit's carbon emission level under current fuel combustion conditions; calculating the carbon emission correction value based on the ratio of the time change rate of the desulfurization and denitrification system's removal efficiency to the load fluctuation amplitude within a preset time window, used to describe the emission response capability of the desulfurization and denitrification system under load variation conditions; calculating the comprehensive pollutant emission value based on the product relationship between ammonia injection rate and fuel input, used to reflect the overall synergistic treatment effect of the unit in the multi-pollutant emission control process; and weighting and summing the carbon emission value, carbon emission correction value, and comprehensive pollutant emission value according to preset weights to obtain the carbon emission index set, used to comprehensively characterize the unit's carbon emission control level during the operating cycle.

[0068] The generation of the economic indicator set specifically includes: calculating the operating cost per unit of power generation based on the weighted sum of fuel input, water consumption, maintenance operation volume, and ammonia injection rate, to reflect the comprehensive operating expenditure of the unit under current load conditions; calculating the economic deviation value based on the rate of change of the ratio of the internal efficiency of the turbine system to the combustion thermal efficiency of the boiler system, to characterize the degree of deviation of the unit's operating economy from the target economic state; calculating the comprehensive economic value based on the weighted sum of fuel cost, water consumption cost, maintenance cost, and ammonia injection cost within a preset time window, to describe the comprehensive economic performance of the unit under energy consumption and emission constraints; and weighting the operating cost value, economic deviation value, and comprehensive economic value according to preset weights to obtain the economic indicator set, which is used to comprehensively characterize the economic operating level of the unit under multiple objective constraints.

[0069] In all the calculation and fusion steps involving the sets of energy consumption indicators, carbon emission indicators, and economic indicators, the values ​​of each sub-indicator are normalized to map data of different dimensions to a unified numerical range, so as to ensure that the multidimensional indicators have consistent scale comparability and calculation stability.

[0070] In this embodiment, the construction of the dynamic benchmark database specifically includes:

[0071] Based on the energy consumption index set, carbon emission index set, and economic index set, the corresponding energy efficiency index, carbon emission index, and economic index are extracted for each sample in the operational dataset. Joint constraint screening is performed based on the preset energy consumption threshold, carbon emission threshold, and economic constraint limit. When the energy efficiency index of a sample is less than or equal to the energy consumption threshold, the carbon emission index is less than or equal to the carbon emission threshold, and the economic index is not higher than the economic constraint limit, it is determined that the sample meets all three constraint conditions at the same time, and the sample that meets the conditions is selected into the candidate sample set.

[0072] Normalization is performed on the energy efficiency indicators, carbon emission indicators and economic indicators in the candidate sample set to map the indicators of different dimensions to a unified numerical range, and a comprehensive score is calculated based on the weights of the three types of indicators: energy consumption, carbon emission and economy. The comprehensive score is used to measure the overall operating level of the sample under multi-objective constraints.

[0073] The candidate sample set is sorted from high to low according to the comprehensive score. When the comprehensive score of a sample is within the sample interval of the previous preset proportion, the sample is determined to be the optimal interval sample. All optimal interval samples are combined into a preferred sample set, and the corresponding running data is extracted from the preferred sample set.

[0074] A benchmark value set is constructed based on the operational data in the preferred sample set;

[0075] The benchmark value set is classified and stored according to unit load, coal type and environmental conditions to establish a dynamic benchmark value database. The dynamic benchmark value database has a periodic update mechanism, which can correct the benchmark value set based on new operating samples, so that the benchmark values ​​maintain timeliness and representativeness under different operating conditions, and provide dynamic benchmark information for the comparative analysis of the real-time operating status of the unit and the optimization of operation control.

[0076] In this embodiment, the generation of the real-time operating condition deviation vector specifically includes:

[0077] Real-time energy consumption indicators, real-time carbon emission indicators, and real-time economic indicators are obtained from real-time operating data to form a set of real-time indicators for the unit.

[0078] The system calls upon the dynamic benchmark database to retrieve a set of benchmark values ​​that match the current unit load, coal type, and environmental conditions.

[0079] The specific steps of retrieving a set of benchmark values ​​that match the current unit load, coal type, and environmental conditions include: After calling the dynamic benchmark value database, performing category filtering based on the current unit's coal type, extracting a set of benchmark samples with the same coal type from the dynamic benchmark value database to obtain a coal type matching sample set; then, in the coal type matching sample set, calculating the load difference of each benchmark sample record based on the current unit load value, filtering records with load differences within a preset threshold range to generate a load matching sample set; next, in the load matching sample set, calculating the environmental difference based on the current ambient temperature and humidity, and filtering several sample records with small environmental differences to generate an environmental matching sample set; extracting the corresponding indicator values ​​for the benchmark energy consumption index, benchmark carbon emission index, and benchmark economic index in the environmental matching sample set to obtain a set of candidate benchmark indicators; performing a weighted average calculation on each indicator value in the candidate benchmark indicator set according to a preset weight coefficient to obtain the benchmark energy consumption index value, benchmark carbon emission index value, and benchmark economic index value under the current unit operating state, and combining the three indicator values ​​to form a benchmark value set;

[0080] The set of real-time indicators of the unit is compared with the set of benchmark values, and the deviation between each indicator is calculated. The deviation includes energy consumption deviation, carbon emission deviation and economic deviation, which respectively represent the multi-dimensional deviation of the current operating condition from the benchmark operating condition.

[0081] Energy consumption deviation, carbon emission deviation, and economic deviation are normalized and mapped to a unified scale space. A real-time operating condition deviation vector is constructed by splicing these parameters.

[0082] In this embodiment, the real-time operating condition deviation vector consists of an energy consumption deviation component, a carbon emission deviation component, and an economic deviation component, and is used to characterize the comprehensive deviation characteristics of the unit's current operating condition in terms of energy consumption, carbon emission, and economic dimensions.

[0083] In this embodiment, the generation of the optimized instruction set specifically includes:

[0084] The real-time operating condition deviation vector is input into an intelligent optimization network constructed based on multi-agent game-theoretic reinforcement learning. This network includes energy consumption agents, carbon emission agents, and economic agents. Each agent corresponds to a set of nodes in the network structure. Each node in the set represents an operating parameter corresponding to the agent's objective. The energy consumption agent's node set includes nodes for fuel input, air flow, main steam temperature, and boiler combustion thermal efficiency. The carbon emission agent's node set includes nodes for flue gas inlet concentration, flue gas outlet concentration, removal efficiency, and ammonia injection rate. The economic agent's node set includes nodes for fuel unit price, water cost, ammonia injection cost, and maintenance time. Connection weights between nodes are calculated. These weights characterize the dynamic coupling relationship between energy consumption, carbon emission, and economic objectives. An adjacency matrix is ​​then established based on these connection weights.

[0085] ;

[0086] in, Represents a node and nodes Connection weights between them Represents a node Energy consumption deviation, Represents a node Energy consumption deviation, Represents a node carbon emission deviation, Represents a node carbon emission deviation, Represents a node Economic deviation, Represents a node Economic deviation, This represents the energy consumption coupling weighting coefficient. This represents the carbon emission coupling weighting coefficient. This represents the economic coupling weight coefficient;

[0087] At each decision point, a state dependency matrix is ​​constructed based on the rates of change of energy consumption deviation, carbon emission deviation, and economic deviation, and the adjacency matrix is ​​updated:

[0088] ;

[0089] in, Indicates time The adjacency matrix, Indicates connection weights. Indicates time The state dependency matrix, This represents element-wise multiplication, used to dynamically adjust edge weights along the dimension of the rate of change of deviation.

[0090] The real-time operating condition deviation vector is used as the initial input of the intelligent optimization network and as the state input benchmark during the graph convolution propagation process, so as to realize the joint strategy evolution of the agent under the three-dimensional target dimensions of energy consumption, carbon emission and economic efficiency.

[0091] Each agent performs graph convolution operation based on the adjacency matrix to update the policy latent vector, where the initial value of the policy latent representation is the real-time operating condition deviation vector:

[0092] ;

[0093] in, Indicates time The strategy implicit vector, This represents the policy latent vector from the previous time step. Represents the training weight matrix. This represents a non-linear activation function used to extract high-order interaction features among multiple agents;

[0094] The strategy latent vector is used to characterize the joint state characteristics of the three-dimensional objectives of energy consumption, carbon emission and economy under the current working conditions. The strategy latent vector is input into the multi-agent game learning process and participates in the strategy parameter update and game balance process as a state input. It is used to drive the strategy update of the three types of agents of energy consumption, carbon emission and economy and the game optimization among multiple agents to achieve dynamic balance at the system level.

[0095] The instant reward values ​​of the energy consumption agent, carbon emission agent, and economic agent are calculated based on the energy consumption deviation, carbon emission deviation, and economic deviation, respectively. The instant reward value of the energy consumption agent is calculated by multiplying the energy consumption deviation value by the energy consumption penalty coefficient and taking the negative value. It is used to reflect the degree of penalty when energy consumption exceeds the benchmark value. The instant reward value of the carbon emission agent is calculated by multiplying the carbon emission deviation value by the carbon emission penalty coefficient and taking the negative value. It is used to reflect the degree of penalty when carbon emission exceeds the standard. The instant reward value of the economic agent is calculated by subtracting the weighted result of the energy consumption penalty item and the carbon emission penalty item from the economic deviation value. It is used to characterize the balance between comprehensive economic benefits and energy consumption and carbon emission constraints.

[0096] The multi-agent policy update process involves calling the agent policy parameter set from the previous decision time, calculating the policy gradient value based on the immediate reward value of each agent and the connection weights between neighboring agents, and updating the policy parameter space along the reward gradient direction according to the learning rate to obtain a new agent policy parameter set. This ensures that the policy update result can achieve a game equilibrium between energy consumption targets, carbon emission targets, and economic targets.

[0097] ;

[0098] in, Represents a node At any moment The strategy parameters, Represents a node At any moment The strategy parameters, Represents the mathematical expectation. Represents the policy function. This indicates calculating the gradient. Indicates the learning rate. Represents a node Instant rewards Represents a node Instant rewards;

[0099] During the strategy update process, a two-layer optimization mechanism of structural evolution and game equilibrium is implemented. Based on the Nash equilibrium condition, the equilibrium point of the multi-agent strategy is determined, and a comprehensive operational adjustment strategy vector is generated.

[0100] The aforementioned two-layer optimization mechanism for structural evolution and game equilibrium specifically includes: calling the policy parameter sets and instant reward sequences of the energy-consuming agent, carbon-emission agent, and economic agent at the current decision moment; calculating the policy return change rate based on the reward difference between two consecutive decision moments to reflect the direction and magnitude of the return adjustments made by each agent; performing minimum-maximum normalization on the policy return change rate, mapping its numerical range to the interval between 0 and 1 to unify the return dimensions of different dimensions; in the structural evolution stage, setting a return sensitivity coefficient based on the average return fluctuation amplitude during the historical training process of each agent, proportionally allocating the normalized return change rate, and updating the step size according to the learning rate control parameter to generate the structure. An evolutionary update matrix is ​​used to characterize the degree of policy dependence among agents. During the game equilibrium phase, the structural evolutionary update matrix and the immediate rewards of each agent are called to calculate the payoff response value one by one. It is determined whether the payoff of the agent increases when the policies of other agents remain unchanged. If the payoff is still increasing, the current update direction is maintained; otherwise, the policy parameters are fine-tuned in the opposite direction. If the payoff change of all agents is less than the threshold of 1% in five consecutive iterations, the policy equilibrium state is considered to have been reached. The equilibrium policy parameters of the three types of agents are extracted and weighted and fused to form a comprehensive operation adjustment policy vector, which serves as the optimal guidance information for the unit operation control system, realizing dynamic Pareto optimal control under the three-dimensional objectives of energy consumption, carbon emissions, and economy.

[0101] An optimized instruction set is generated based on the comprehensive operation adjustment strategy vector.

[0102] In this embodiment, the optimized instruction set includes fuel distribution adjustment, air flow regulation, ammonia injection rate correction, and heating load distribution, which are used to guide the coordinated operation and regulation of the boiler system, turbine system, air cooling system, and desulfurization and denitrification system.

[0103] An artificial intelligence-based operation control optimization system for thermal power plants includes:

[0104] The operation data acquisition module is used to collect operating data from thermal power plant units and build an operating dataset.

[0105] The data preprocessing module is used to preprocess the running dataset;

[0106] The performance calculation module is used to form a set of equipment performance indicators and generate a set of energy consumption indicators, a set of carbon emission indicators, and a set of economic indicators based on the set of equipment performance indicators.

[0107] The benchmark value construction module is used to filter samples from the running dataset, construct a set of benchmark values, and establish a dynamic benchmark value database.

[0108] The working condition deviation calculation module is used to call the set of benchmark values ​​in the dynamic benchmark value database, perform comparative analysis, calculate the multi-dimensional deviation between the current working condition and the optimal working condition, and generate a real-time working condition deviation vector.

[0109] The intelligent optimization module is used to input the real-time operating condition deviation vector into the intelligent optimization network constructed based on multi-agent game reinforcement learning, perform strategy game update, generate a comprehensive operation adjustment strategy vector through the game equilibrium process, and output the optimization instruction set.

[0110] The feedback optimization module is used to collect feedback data after the operation is adjusted and to correct the dynamic benchmark value database.

[0111] Example 1:

[0112] To verify the feasibility of this invention in practice, it was applied to the daily operation control optimization of a 300MW coal-fired unit in a large thermal power plant. In actual operation, this unit has long faced problems such as large fluctuations in energy consumption, high carbon emission intensity, and lagging economic dispatch response. Especially under the influence of frequent load adjustments, drastic fluctuations in coal quality, and hot and humid climate, the traditional operation control strategy based on a single indicator can hardly balance energy conservation and emission reduction with economy, resulting in low operating efficiency and imbalance of control objectives.

[0113] In this application scenario, comprehensive data collection is performed on the unit's operating data across different load segments, coal types, and multiple days. Through structured cleaning and time alignment, an operating dataset containing operating parameters, energy efficiency indicators, carbon emission records, and economic output is established. Subsequently, the system extracts performance indicators and models multi-objective indicators according to the method described in this invention, constructing a multi-dimensional operating evaluation space that includes three types of indicators: energy consumption, carbon emission, and economics. A dynamic benchmark database is also constructed by combining time-series distribution. By comparing the data with the unit's current real-time data, the system can automatically generate an operating deviation vector and input this vector into an intelligent optimization network constructed based on a multi-agent game reinforcement learning mechanism.

[0114] The intelligent optimization network comprises three collaborative sub-agents: an energy consumption agent, a carbon emission agent, and an economic agent. Each agent represents its optimization goal orientation through a policy implicit vector and updates its implicit state through a graph convolution mechanism. This process further executes structural evolution and policy game theory. During policy evolution, agents form policy conflicts and coordination relationships based on state sharing and information interaction. Through reinforcement learning training, the system gradually approaches the multi-objective game equilibrium and outputs a comprehensive operational adjustment policy vector, forming an optimized instruction set that can directly guide operational parameters such as combustion regulation, heat supply distribution, and coal feeding control.

[0115] After the unit operators make on-site adjustments according to the optimization instructions, the system will continuously collect the adjusted operation feedback data and automatically correct the benchmark value database and strategy parameters, realizing a closed-loop management process from strategy generation to parameter verification and benchmark value update. The results of several weeks of trial operation show that this method can effectively cope with multi-objective conflicts in the actual operating environment and achieve stable and efficient operation control and regulation.

[0116] To verify the performance of the present invention, it was compared with the traditional method. The comparison results are shown in Table 1.

[0117] Table 1. Performance Comparison of the Invention and Traditional Methods

[0118]

[0119] In terms of energy consumption, the method of this invention reduces the standard coal consumption per unit of power generation from 296g / kWh under the traditional strategy to 287g / kWh, achieving an energy saving of 3.04%. This optimization benefits from the dynamic updating capability of the multi-agent game mechanism for the operating strategy, making the adjustment of combustion efficiency more precise. Especially when the load changes frequently, the dynamic response of the graph convolution structure and the strategy implicit vector can capture the operating condition deviation and output matching instructions more quickly, thereby reducing unnecessary energy waste.

[0120] In terms of responsiveness, the adjustment response time of the coal feeding system is reduced from 17.5 seconds to 9.8 seconds, improving efficiency by 44%. Traditional methods rely heavily on empirical rules and static adjustment formulas, resulting in lag in response. In contrast, this invention continuously approaches the local optimal control strategy through reinforcement learning, achieving second-level dynamic control and improving the sensitivity and accuracy of the response.

[0121] In terms of carbon emissions, the carbon emission factor decreased by 4.18%, indicating a significant reduction in carbon emissions per unit of electricity generated. At the same time, the method of this invention reduced the number of times the denitrification system exceeded the standard, from an average of 4.2 times per month to 1.3 times, a reduction of more than 69%. This is attributed to the system's ability to monitor and quickly correct carbon emission-related operating conditions in real time, thereby maintaining the stable operation of the equipment within the environmental compliance range.

[0122] In terms of economics, the variable cost per unit of electricity generation was reduced by 5.12%, indicating that the optimization strategy not only saves energy and reduces emissions, but also brings actual benefits in terms of fuel, operation and maintenance costs. In addition, the loss ratio during load fluctuations also decreased from 11.7% to 7.4%, reflecting the improvement of operational stability brought about by the invention and reducing economic losses caused by frequent start-ups or load swings.

[0123] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An artificial intelligence-based optimization method for the operation control of thermal power plants, characterized in that, Includes the following steps: Collect operational data of thermal power plant units under different loads, coal types, and environmental conditions to construct an operational dataset; The runtime dataset is preprocessed to generate a preprocessed runtime dataset; The preprocessed running dataset is used to perform performance calculations to obtain a set of equipment performance indicators, and then a set of energy consumption indicators, a set of carbon emission indicators, and a set of economic indicators are generated based on the set of equipment performance indicators. Based on the sets of energy consumption indicators, carbon emission indicators, and economic indicators, samples are selected from the operational dataset, corresponding operational data are extracted, a set of benchmark values ​​is constructed, and a dynamic benchmark value database is established. By comparing and analyzing real-time operating data with the benchmark value set in the dynamic benchmark value database, the multi-dimensional deviation between the current operating condition and the optimal operating condition is calculated, and a real-time operating condition deviation vector is generated. The real-time operating condition deviation vector is input into an intelligent optimization network constructed based on multi-agent game reinforcement learning, and the strategy game update is performed. Through the game equilibrium process, a comprehensive operation adjustment strategy vector is generated to obtain the optimized instruction set. Collect feedback data after operational adjustments, correct the dynamic benchmark value database, and realize multi-objective operation control optimization closed-loop management of thermal power units; The construction of the dynamic benchmark database specifically includes: Based on the energy consumption index set, carbon emission index set, and economic index set, the corresponding energy efficiency index, carbon emission index, and economic index are extracted for each sample in the operational dataset. Joint constraint screening is performed based on the preset energy consumption threshold, carbon emission threshold, and economic constraint limit. When the energy efficiency index of a sample is less than or equal to the energy consumption threshold, the carbon emission index is less than or equal to the carbon emission threshold, and the economic index is not higher than the economic constraint limit, it is determined that the sample meets all three constraint conditions at the same time, and the sample that meets the conditions is selected into the candidate sample set. Normalization is performed on the energy efficiency indicators, carbon emission indicators and economic indicators in the candidate sample set to map the indicators with different dimensions to a unified numerical range, and a comprehensive score is calculated based on the weights of the three types of indicators: energy consumption, carbon emission and economy. The candidate sample set is sorted from high to low according to the comprehensive score. When the comprehensive score of a sample is within the sample interval of the previous preset proportion, the sample is determined to be the optimal interval sample. All optimal interval samples are combined into a preferred sample set, and the corresponding running data is extracted from the preferred sample set. A benchmark value set is constructed based on the operational data in the preferred sample set; The benchmark value set is classified and stored according to unit load, coal type and environmental conditions to establish a dynamic benchmark value database; The generation of the real-time operating condition deviation vector specifically includes: Real-time energy consumption indicators, real-time carbon emission indicators, and real-time economic indicators are obtained from real-time operating data to form a set of real-time indicators for the unit. The system calls upon the dynamic benchmark database to retrieve a set of benchmark values ​​that match the current unit load, coal type, and environmental conditions. The set of real-time indicators of the unit is compared with the set of benchmark values, and the deviation between each indicator is calculated. Normalization is performed on the energy consumption deviation, carbon emission deviation and economic deviation, and they are mapped to a unified scale space. A real-time operating condition deviation vector is constructed by splicing the components. The real-time operating condition deviation vector consists of energy consumption deviation components, carbon emission deviation components and economic deviation components. The generation of the optimized instruction set specifically includes: The real-time operating condition deviation vector is input into an intelligent optimization network constructed based on multi-agent game reinforcement learning. The intelligent optimization network includes an energy consumption agent, a carbon emission agent, and an economic agent. Each agent corresponds to a set of nodes in the network structure. Each node in the set of nodes represents the operating parameters corresponding to the agent's objective. The connection weights between the nodes are calculated, and an adjacency matrix is ​​established based on the connection weights. At each decision point, a state dependency matrix is ​​constructed based on the rate of change of energy consumption deviation, carbon emission deviation, and economic deviation, and the adjacency matrix is ​​updated. The real-time operating condition deviation vector is used as the initial input to the intelligent optimization network; Each agent performs graph convolution operation based on the adjacency matrix to update the policy latent vector, where the initial value of the policy latent representation is the real-time operating condition deviation vector: ; in, Indicates time The strategy implicit vector, This represents the policy latent vector from the previous time step. Represents the training weight matrix. This represents a non-linear activation function used to extract high-order interaction features among multiple agents. Indicates time The adjacency matrix; The instantaneous reward values ​​of the energy-consuming agent, carbon-emission agent, and economic agent are calculated based on the energy consumption deviation, carbon emission deviation, and economic efficiency deviation, respectively. To perform multi-agent policy updates, the agent policy parameter set from the previous decision time step is invoked. The policy gradient is calculated based on the immediate reward values ​​of each agent and the connection weights between neighboring agents. The updated policy parameter set is then performed in the policy parameter space along the reward gradient direction according to the learning rate, resulting in a new agent policy parameter set. ; in, Represents a node At any moment The strategy parameters, Represents a node At any moment The strategy parameters, Represents the mathematical expectation. Represents the policy function. This indicates calculating the gradient. Indicates the learning rate. Represents a node Instant rewards Represents a node Instant rewards Represents a node and nodes Connection weights between them; The reward gradient direction is the logarithmic policy function. Joint Award Items The overall composition relative to the strategy parameters The direction of parameter update obtained after calculating the gradient; During the strategy update process, a two-layer optimization mechanism of structural evolution and game equilibrium is implemented. Based on the Nash equilibrium condition, the equilibrium point of the multi-agent strategy is determined, and a comprehensive operational adjustment strategy vector is generated. An optimized instruction set is generated based on the comprehensive operation adjustment strategy vector.

2. The method for optimizing the operation control of a thermal power plant based on artificial intelligence according to claim 1, characterized in that, The operational data includes boiler system operational data, steam turbine system operational data, air-cooled system operational data, desulfurization and denitrification system operational data, and heating system operational data.

3. The method for optimizing the operation control of a thermal power plant based on artificial intelligence according to claim 1, characterized in that, The preprocessing includes time series alignment, noise filtering, outlier removal, missing value imputation, and data normalization.

4. The method for optimizing the operation control of a thermal power plant based on artificial intelligence according to claim 1, characterized in that, The set of equipment performance indicators includes combustion thermal efficiency, turbine internal efficiency, cooling efficiency, desulfurization efficiency, and heat transfer efficiency, which correspond to the operating status of the boiler system, turbine system, air-cooled system, desulfurization and denitrification system, and heating system, respectively.

5. The method for optimizing the operation control of a thermal power plant based on artificial intelligence according to claim 1, characterized in that, The optimized instruction set includes fuel distribution adjustment, air flow regulation, ammonia injection rate correction, and heating load distribution.

6. An artificial intelligence-based thermal power plant operation control optimization system, executing the artificial intelligence-based thermal power plant operation control optimization method according to any one of claims 1 to 5, characterized in that, include: The operation data acquisition module is used to collect operating data from thermal power plant units and build an operating dataset. The data preprocessing module is used to preprocess the running dataset; The performance calculation module is used to form a set of equipment performance indicators and generate a set of energy consumption indicators, a set of carbon emission indicators, and a set of economic indicators based on the set of equipment performance indicators. The benchmark value construction module is used to filter samples from the running dataset, construct a set of benchmark values, and establish a dynamic benchmark value database. The working condition deviation calculation module is used to call the set of benchmark values ​​in the dynamic benchmark value database, perform comparative analysis, calculate the multi-dimensional deviation between the current working condition and the optimal working condition, and generate a real-time working condition deviation vector. The intelligent optimization module is used to input the real-time operating condition deviation vector into the intelligent optimization network constructed based on multi-agent game reinforcement learning, perform strategy game update, generate a comprehensive operation adjustment strategy vector through the game equilibrium process, and output the optimization instruction set. The feedback optimization module is used to collect feedback data after operational adjustments and to correct the dynamic benchmark value database. The construction of the dynamic benchmark database specifically includes: Based on the energy consumption index set, carbon emission index set, and economic index set, the corresponding energy efficiency index, carbon emission index, and economic index are extracted for each sample in the operational dataset. Joint constraint screening is performed based on the preset energy consumption threshold, carbon emission threshold, and economic constraint limit. When the energy efficiency index of a sample is less than or equal to the energy consumption threshold, the carbon emission index is less than or equal to the carbon emission threshold, and the economic index is not higher than the economic constraint limit, it is determined that the sample meets all three constraint conditions at the same time, and the sample that meets the conditions is selected into the candidate sample set. Normalization is performed on the energy efficiency indicators, carbon emission indicators and economic indicators in the candidate sample set to map the indicators with different dimensions to a unified numerical range, and a comprehensive score is calculated based on the weights of the three types of indicators: energy consumption, carbon emission and economy. The candidate sample set is sorted from high to low according to the comprehensive score. When the comprehensive score of a sample is within the sample interval of the previous preset proportion, the sample is determined to be the optimal interval sample. All optimal interval samples are combined into a preferred sample set, and the corresponding running data is extracted from the preferred sample set. A benchmark value set is constructed based on the operational data in the preferred sample set; The benchmark value set is classified and stored according to unit load, coal type and environmental conditions to establish a dynamic benchmark value database; The generation of the real-time operating condition deviation vector specifically includes: Real-time energy consumption indicators, real-time carbon emission indicators, and real-time economic indicators are obtained from real-time operating data to form a set of real-time indicators for the unit. The system calls upon the dynamic benchmark database to retrieve a set of benchmark values ​​that match the current unit load, coal type, and environmental conditions. The set of real-time indicators of the unit is compared with the set of benchmark values, and the deviation between each indicator is calculated. Normalization is performed on the energy consumption deviation, carbon emission deviation and economic deviation, and they are mapped to a unified scale space. A real-time operating condition deviation vector is constructed by splicing the components. The real-time operating condition deviation vector consists of energy consumption deviation components, carbon emission deviation components and economic deviation components. The generation of the optimized instruction set specifically includes: The real-time operating condition deviation vector is input into an intelligent optimization network constructed based on multi-agent game reinforcement learning. The intelligent optimization network includes an energy consumption agent, a carbon emission agent, and an economic agent. Each agent corresponds to a set of nodes in the network structure. Each node in the set of nodes represents the operating parameters corresponding to the agent's objective. The connection weights between the nodes are calculated, and an adjacency matrix is ​​established based on the connection weights. At each decision point, a state dependency matrix is ​​constructed based on the rate of change of energy consumption deviation, carbon emission deviation, and economic deviation, and the adjacency matrix is ​​updated. The real-time operating condition deviation vector is used as the initial input to the intelligent optimization network; Each agent performs graph convolution operation based on the adjacency matrix to update the policy latent vector, where the initial value of the policy latent representation is the real-time operating condition deviation vector: ; in, Indicates time The strategy implicit vector, This represents the policy latent vector from the previous time step. Represents the training weight matrix. This represents a non-linear activation function used to extract high-order interaction features among multiple agents. Indicates time The adjacency matrix; The instantaneous reward values ​​of the energy-consuming agent, carbon-emission agent, and economic agent are calculated based on the energy consumption deviation, carbon emission deviation, and economic efficiency deviation, respectively. To perform multi-agent policy updates, the agent policy parameter set from the previous decision time step is invoked. The policy gradient is calculated based on the immediate reward values ​​of each agent and the connection weights between neighboring agents. The updated policy parameter set is then performed in the policy parameter space along the reward gradient direction according to the learning rate, resulting in a new agent policy parameter set. ; in, Represents a node At any moment The strategy parameters, Represents a node At any moment The strategy parameters, Represents the mathematical expectation. Represents the policy function. This indicates calculating the gradient. Indicates the learning rate. Represents a node Instant rewards Represents a node Instant rewards Represents a node and nodes Connection weights between them; The reward gradient direction is the logarithmic policy function. Joint Award Items The overall composition relative to the strategy parameters The direction of parameter update obtained after calculating the gradient; During the strategy update process, a two-layer optimization mechanism of structural evolution and game equilibrium is implemented. Based on the Nash equilibrium condition, the equilibrium point of the multi-agent strategy is determined, and a comprehensive operational adjustment strategy vector is generated. An optimized instruction set is generated based on the comprehensive operation adjustment strategy vector.