A hierarchical optimization control method for smart thermal power plants based on multi-agent systems
By employing a multi-agent intelligent thermal power plant hierarchical optimization control method, the complexity of scheduling and equipment control in thermal power plants under the uncertainty of new energy sources is solved, and the collaborative cooperation of multiple agents and efficient and economical operation are realized.
Patent Information
- Application Number
- CN202510178744.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The scheduling and equipment control of thermal power plants are complex under the uncertainty of new energy output and environmental variability. The coupling and correlation of various units make it difficult to achieve effective coordination, which affects economic benefits.
A multi-agent intelligent thermal power plant hierarchical optimization control method is adopted. By setting up intelligent agents for thermal power units, wind, solar and energy storage, simulation, plant-level scheduling decision and equipment-level control, a multi-agent simulation model is constructed. Multi-agent reinforcement learning algorithm is used to predict electric and heat load and optimize equipment control, so as to realize the collaborative cooperation of multiple agents.
It improves the adaptability and collaborative efficiency of thermal power plants in complex environments, enables economical and efficient scheduling and equipment control, and enhances the reliability of algorithmic decision-making.
Smart Images

Figure CN120044789B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of smart thermal power plant technology, specifically relating to a hierarchical optimization control method for smart thermal power plants based on multi-agent systems. Background Technology
[0002] In recent years, under the national policy background of vigorously developing clean energy, although the installed capacity of new energy sources such as photovoltaic and wind power has continued to grow, coal-fired power units are still the main equipment for energy supply in my country. For a long time to come, thermal power units will remain the main force in the production of electricity and heat. To build an efficient, clean, and sustainable thermal power plant energy supply industry, in addition to strictly enforcing energy efficiency standards for new units and implementing energy-saving and emission-reduction upgrades and renovations of existing thermal power units, it is also necessary to rationally allocate the output power of thermal power units, reduce the total operating costs and fuel consumption of thermal power plants, optimize resources, improve unit load rate and operating quality, and achieve the best economic and social benefits.
[0003] Currently, thermal power plants are gradually adopting a combined energy supply model of new energy units and thermal power units. The main problems they face are: due to the uncertainty of new energy output and the variability of the external environment, the scheduling and equipment control of thermal power plant units are more complex. In addition, the operating output of each thermal power unit and new energy unit in a thermal power plant is coupled and related. Each unit must not only consider its own operating status, but also the operating status of other units in order to achieve effective synergy and improve the economic benefits of the thermal power plant.
[0004] Based on the above technical problems, a new hierarchical optimization control method for smart thermal power plants based on multi-agent systems needs to be designed. Summary of the Invention
[0005] The technical problem to be solved by this invention is to overcome the shortcomings of the prior art and provide a hierarchical optimization control method for smart thermal power plants based on multi-agent systems. This method can treat each piece of equipment in the thermal power plant as an independent decision-making agent and, based on the current state and behavioral interaction information of other agents, effectively realize multi-agent collaborative cooperation, improve the adaptability and cooperation efficiency of multi-agent systems in complex thermal power plant environments, achieve economical and efficient scheduling, and at the same time improve the reliability of algorithm decision-making through continuous training by utilizing corresponding multi-agent reinforcement learning algorithms.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0007] This invention provides a hierarchical optimization control method for smart thermal power plants based on multi-agent systems, comprising:
[0008] Step S1: Set up a multi-agent system for a smart thermal power plant, including at least the intelligent agents of each thermal power unit, wind, solar and energy storage intelligent agents, simulation intelligent agents, upper-level plant-level scheduling and decision-making intelligent agents, and lower-level equipment-level control intelligent agents, and build a multi-agent simulation model that can reproduce the actual operation of the smart thermal power plant.
[0009] Step S2: The inferential agent uses a multi-agent simulation model to obtain historical and real-time operating data of the thermal power plant, combines external dynamic parameters to predict the thermal power plant's electrical and thermal load, wind and solar power generation, and energy storage, and infers the future operating status of the thermal power plant based on the prediction results.
[0010] Step S3: The upper-level plant-level scheduling decision-making agent acquires the observation data of each thermal power unit agent, wind, solar and storage agent, and simulation agent, and calls the preset multi-agent reinforcement learning module to model the thermal power plant's electrical and thermal load optimization allocation problem, and performs plant-level multi-agent action space design, state space design, and multi-objective reward function design to solve and obtain the optimal load allocation results of each thermal power unit agent and wind, solar and storage agent.
[0011] Step S4: The lower-level equipment-level control agent, based on the optimal load allocation results of each thermal power unit's agent, defines the steam turbine, boiler body, boiler auxiliary equipment, and heating steam extraction device in each thermal power unit as corresponding equipment agents. It integrates the interaction information between each equipment agent, equipment operation data, and equipment parameter control mechanisms to achieve the goal of minimizing load allocation results and operating costs. It performs equipment-level multi-agent optimization control modeling and solves the optimal control commands for each equipment agent.
[0012] Furthermore, in step S1, constructing a multi-agent simulation model capable of reproducing the actual operation of a smart thermal power plant includes:
[0013] The actual operating data of the thermal power plant is read into each thermal power unit intelligent agent, wind, solar and energy storage intelligent agent, simulation intelligent agent, upper-level plant-level scheduling and decision-making intelligent agent, and lower-level equipment-level control intelligent agent. Combined with the operating mechanism of each intelligent agent, a simulation model of each intelligent agent in the thermal power plant is established. The operating process rules and basic functions of each intelligent agent are reproduced and transformed into rule knowledge base, static knowledge base and dynamic knowledge base of each intelligent agent.
[0014] The system integrates the operating data, rule knowledge base, interactive information, and scheduling control parameters of each thermal power unit intelligent agent, wind, solar and energy storage intelligent agent, inference intelligent agent, upper-level plant-level scheduling decision intelligent agent, and lower-level equipment-level control intelligent agent using a pre-set modeling master intelligent agent. At the same time, it uses a pre-set environmental intelligent agent to simulate the external environmental conditions of the thermal power plant, forming a multi-agent simulation model of the thermal power plant with the cooperation of each intelligent agent.
[0015] The accuracy of the simulation model is judged based on the error between the simulation data output by the multi-agent simulation model of the thermal power plant and the actual data. If it does not meet expectations, the model parameters are modified until the error is within a reasonable range, and finally a multi-agent simulation model that can reproduce the actual operation of the smart thermal power plant is obtained.
[0016] Furthermore, step S2 includes:
[0017] The inferential agent uses a multi-agent simulation model to obtain historical and real-time operating data of each agent in the thermal power plant. Combined with external factors such as temperature, humidity, wind speed, solar radiation intensity, electricity price information, and peak and valley load information, the inferential agent uses these data as environmental status information.
[0018] Based on environmental state information, the inferential agent selects different feature extraction algorithms and prediction models as action spaces to predict the thermal load of thermal power plants, wind and solar power generation, and energy storage. The prediction accuracy of the prediction model is used as the reward function to select the optimal prediction model. The corresponding prediction results and external dynamic parameters for future periods are input into the multi-agent simulation model to infer the future operating state of thermal power plants.
[0019] Furthermore, step S3 includes:
[0020] The upper-level plant-level scheduling decision-making intelligent agent acquires observation data from each thermal power unit intelligent agent, wind-solar-storage intelligent agent, and inference intelligent agent, including at least the power generation and future predicted power generation of each thermal power unit, the current power generation and future predicted wind-solar power generation, energy storage, the operating coal consumption characteristics of each thermal power unit, the electricity and heat load demand, weather conditions, and electricity price information.
[0021] The pre-defined multi-agent reinforcement learning module for scheduling models the optimal allocation of electrical and thermal loads in thermal power plants, represented as follows:
[0022] M =<G,U,r,O,n,π,γ> ;
[0023] G represents the global state of the plant-level scheduling decision; U represents the action set, including the actions chosen by each agent; r represents the reward function; each agent has its own observation value o∈O, and the agent has a policy π; n represents the number of agents; γ represents the discount factor γ∈[0,1]; where the plant-level multi-agent state space includes the observation data of each agent, the action space includes the operating state and operating output of each agent, and the multi-objective reward function includes minimizing plant-level operating costs, minimizing carbon emissions, and maximizing the absorption capacity of wind and solar new energy.
[0024] In the upper-level plant-level scheduling decision-making intelligent agent, each intelligent agent uses neural network units to store observation data and historical actions;
[0025] Initialize the environment, network training parameters, and experience replay pool for optimal allocation of electrical and thermal loads in thermal power plants;
[0026] Using the observation data obtained by the i-th agent at time t Each action is worth Q. i Select Action And generate messages To communicate;
[0027] The trajectory sets of each agent are stored in the experience replay pool for sampling and training.
[0028] Obtain the action value Q of each agent at time t and the global reward Q. total Training is performed, and the overall objective function loss is calculated to update the network parameters of the multi-agent reinforcement learning module;
[0029] The multi-agent reinforcement learning module, which has been trained and updated, outputs the optimal load allocation results for each thermal power unit agent and the wind, solar and energy storage agent.
[0030] Furthermore, the multi-agent reinforcement learning module includes a neural network unit, a joint network, and an information interaction unit; it processes the observation data of each agent. Interaction information and actions of other intelligent agents The input is fed into the neural network unit, which selects actions and outputs the Q-values of each agent. i (τ i ,u i,t ), For the action observation trajectory of the i-th agent, the Q-values selected by each agent and the global state are input into the joint network. Based on the input global state, each local value function and global value function are trained to maximize the global Q-value: maxQ total =g({Q i} i∈n g(·) represents the relationship between the global value function and the local value function, realizing the global Q-value Q. total The optimization process involves encoding the historical observation data and actions of each agent using the information interaction unit and sending messages to other agents, represented as follows: M i,t-1 For the message generated by the i-th agent at time t-1, h i,t-1 Let φ be the hidden state of the LSTM network at time t-1, including the historical information of the i-th agent, and let φ be the parameters of the joint network.
[0031] Furthermore, in step S4, device-level multi-agent optimization control modeling is performed, including:
[0032] Each thermal power unit defines its steam turbine, boiler body, boiler auxiliary equipment, and heating steam extraction device as a corresponding intelligent entity. The steam turbine intelligent entity is used for the operation control of the steam turbine, including the adjustment of speed, power, and temperature parameters. The boiler body intelligent entity is used for the control of the steam-water system and combustion system, including boiler water level, main steam pressure, steam-water ratio, fuel supply, air volume adjustment, and furnace temperature. The boiler auxiliary equipment intelligent entity is used for the control of ventilation equipment, coal conveying equipment, pulverizing equipment, water supply equipment, and ash removal equipment. The heating steam extraction device intelligent entity is used for the control of heating steam extraction, adjusting the extraction steam volume and pressure according to the heating load demand.
[0033] Multi-agent reinforcement learning is applied to equipment optimization control. It integrates equipment operation data, inter-device coupling and interaction information, and equipment parameter control mechanisms from lower-level device-level control agents. Each device agent makes autonomous decisions based on its current state and collaboration with other device agents. The equipment optimization control problem is modeled as a Markov decision process, expressed as:
[0034] W =<M,S,A,P,O′,r′,γ′> ;
[0035] M = {1, 2, ..., m} is a set of m device agents; S is the state space observed by all device agents; A is the action space for each device agent; P is the transition probability function from any state to any state after taking a joint action; O′ is the observation space of the device agents; r′ is the reward function, set according to the goal of minimizing load allocation and operating costs; γ′ is the discount factor; where each device agent j selects an action a. j ∈A, forming a joint action vector a={a1,a2,…,a n}∈A M .
[0036] Furthermore, in step S4, when solving for the optimal control instructions for each device agent, an improved multi-agent deep deterministic policy gradient algorithm is adopted: a value function decomposition method is introduced to decompose the value function and quantify the contribution of each device agent.
[0037] Furthermore, in the improved multi-agent deep deterministic policy gradient algorithm, the Actor network is responsible for generating actions, and the Critic network is responsible for calculating the value function of the environment state and the current action selected by the Actor network. τ j For the historical observation information of the j-th device intelligent agent; φ j Value function Network parameters; μ jFor deterministic policies, and by decomposing the value function through a hybrid network, each device agent is allowed to make independent decisions based on its own observations and historical information, achieving a balance between individual autonomous learning and global optimization.
[0038] Furthermore, the device agents share a centralized network of critics, the coefficients of which are represented as follows:
[0039]
[0040] φ is the value function of the joint action. The parameters are: τ is the set of historical observations of all device agents; a is the set of actions of all device agents; g ψ It is a hybrid network and is a nonlinear monotonic function;
[0041] The critic network is trained by minimizing a loss function, as follows:
[0042]
[0043] θ - φ - ψ - θ represents the parameters of the Actor network, Critic network, and hybrid network, respectively; μ(τ′; θ) - ) is the strategy given to the Actor network.
[0044] Furthermore, the Actor network includes a hidden layer and a gated recurrent unit (GRU) layer; the Critic network includes a hidden layer and integrates a hybrid network.
[0045] The beneficial effects of this invention are:
[0046] This invention establishes a multi-agent system for smart thermal power plants, builds a multi-agent simulation model, and introduces a multi-agent reinforcement learning algorithm for upper-level plant-level scheduling decisions and lower-level equipment-level optimization control. It enables each piece of equipment in the thermal power plant to act as an independent decision-making agent, effectively achieving multi-agent collaboration based on its current state and the behavioral interaction information of other agents. This improves the adaptability and collaborative efficiency of the multi-agent system in complex thermal power plant environments, achieving economical and efficient scheduling. Furthermore, by continuously training the corresponding multi-agent reinforcement learning algorithm, the reliability of the algorithm's decisions is improved.
[0047] Other features and advantages will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained through the structures particularly pointed out in the description and the drawings.
[0048] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0049] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0050] Figure 1 This is a flowchart of a hierarchical optimization control method for smart thermal power plants based on multi-agent systems, according to the present invention.
[0051] Figure 2 This is a schematic diagram of the multi-agent architecture of the smart thermal power plant of the present invention. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] Example 1
[0054] like Figure 1 , Figure 2 As shown in the figure, this embodiment 1 provides a hierarchical optimization control method for smart thermal power plants based on multi-agent systems, which includes:
[0055] Step S1: Set up a multi-agent system for a smart thermal power plant, including at least the intelligent agents of each thermal power unit, wind, solar and energy storage intelligent agents, simulation intelligent agents, upper-level plant-level scheduling and decision-making intelligent agents, and lower-level equipment-level control intelligent agents, and build a multi-agent simulation model that can reproduce the actual operation of the smart thermal power plant.
[0056] Step S2: The inferential agent uses a multi-agent simulation model to obtain historical and real-time operating data of the thermal power plant, combines external dynamic parameters to predict the thermal power plant's electrical and thermal load, wind and solar power generation, and energy storage, and infers the future operating status of the thermal power plant based on the prediction results.
[0057] Step S3: The upper-level plant-level scheduling decision-making agent acquires the observation data of each thermal power unit agent, wind, solar and storage agent, and simulation agent, and calls the preset multi-agent reinforcement learning module to model the thermal power plant's electrical and thermal load optimization allocation problem, and performs plant-level multi-agent action space design, state space design, and multi-objective reward function design to solve and obtain the optimal load allocation results of each thermal power unit agent and wind, solar and storage agent.
[0058] Step S4: The lower-level equipment-level control agent, based on the optimal load allocation results of each thermal power unit's agent, defines the steam turbine, boiler body, boiler auxiliary equipment, and heating steam extraction device in each thermal power unit as corresponding equipment agents. It integrates the interaction information between each equipment agent, equipment operation data, and equipment parameter control mechanisms to achieve the goal of minimizing load allocation results and operating costs. It performs equipment-level multi-agent optimization control modeling and solves the optimal control commands for each equipment agent.
[0059] In this embodiment, step S1 involves constructing a multi-agent simulation model capable of reproducing the actual operation of a smart thermal power plant, including:
[0060] The actual operating data of the thermal power plant is read into each thermal power unit intelligent agent, wind, solar and energy storage intelligent agent, simulation intelligent agent, upper-level plant-level scheduling and decision-making intelligent agent, and lower-level equipment-level control intelligent agent. Combined with the operating mechanism of each intelligent agent, a simulation model of each intelligent agent in the thermal power plant is established. The operating process rules and basic functions of each intelligent agent are reproduced and transformed into rule knowledge base, static knowledge base and dynamic knowledge base of each intelligent agent.
[0061] The system integrates the operating data, rule knowledge base, interactive information, and scheduling control parameters of each thermal power unit intelligent agent, wind, solar and energy storage intelligent agent, inference intelligent agent, upper-level plant-level scheduling decision intelligent agent, and lower-level equipment-level control intelligent agent using a pre-set modeling master intelligent agent. At the same time, it uses a pre-set environmental intelligent agent to simulate the external environmental conditions of the thermal power plant, forming a multi-agent simulation model of the thermal power plant with the cooperation of each intelligent agent.
[0062] The accuracy of the simulation model is judged based on the error between the simulation data output by the multi-agent simulation model of the thermal power plant and the actual data. If it does not meet expectations, the model parameters are modified until the error is within a reasonable range, and finally a multi-agent simulation model that can reproduce the actual operation of the smart thermal power plant is obtained.
[0063] It should be noted that the static knowledge base includes the basic parameters of each thermal power unit and wind, solar, and energy storage equipment; the dynamic knowledge base includes the operating data, scheduling strategies, and equipment control parameters of each thermal power unit and wind, solar, and energy storage equipment. The main intelligent agent for modeling includes a demonstration animation module and a simulation of the entire thermal power plant's operating scenario.
[0064] In this embodiment, step S2 includes:
[0065] The inferential agent uses a multi-agent simulation model to obtain historical and real-time operating data of each agent in the thermal power plant. Combined with external factors such as temperature, humidity, wind speed, solar radiation intensity, electricity price information, and peak and valley load information, the inferential agent uses these data as environmental status information.
[0066] Based on environmental state information, the inferential agent selects different feature extraction algorithms and prediction models as action spaces to predict the thermal load of thermal power plants, wind and solar power generation, and energy storage. The prediction accuracy of the prediction model is used as the reward function to select the optimal prediction model. The corresponding prediction results and external dynamic parameters for future periods are input into the multi-agent simulation model to infer the future operating state of thermal power plants.
[0067] In this embodiment, step S3 includes:
[0068] The upper-level plant-level scheduling decision-making intelligent agent acquires observation data from each thermal power unit intelligent agent, wind-solar-storage intelligent agent, and inference intelligent agent, including at least the power generation and future predicted power generation of each thermal power unit, the current power generation and future predicted wind-solar power generation, energy storage, the operating coal consumption characteristics of each thermal power unit, the electricity and heat load demand, weather conditions, and electricity price information.
[0069] The pre-defined multi-agent reinforcement learning module for scheduling models the optimal allocation of electrical and thermal loads in thermal power plants, represented as follows:
[0070] M =<G,U,r,O,n,π,γ> ;
[0071] G represents the global state of the plant-level scheduling decision; U represents the action set, including the actions chosen by each agent; r represents the reward function; each agent has its own observation value o∈O, and the agent has a policy π; n represents the number of agents; γ represents the discount factor γ∈[0,1]; where the plant-level multi-agent state space includes the observation data of each agent, the action space includes the operating state and operating output of each agent, and the multi-objective reward function includes minimizing plant-level operating costs, minimizing carbon emissions, and maximizing the absorption capacity of wind and solar new energy.
[0072] In the upper-level plant-level scheduling decision-making intelligent agent, each intelligent agent uses neural network units to store observation data and historical actions;
[0073] Initialize the environment, network training parameters, and experience replay pool for optimal allocation of electrical and thermal loads in thermal power plants;
[0074] Using the observation data obtained by the i-th agent at time t Each action is worth Q. i Select Action And generate messages To communicate;
[0075] The trajectory sets of each agent are stored in the experience replay pool for sampling and training.
[0076] Obtain the action value Q of each agent at time t and the global reward Q. total Training is performed, and the overall objective function loss is calculated to update the network parameters of the multi-agent reinforcement learning module;
[0077] The multi-agent reinforcement learning module, which has been trained and updated, outputs the optimal load allocation results for each thermal power unit agent and the wind, solar and energy storage agent.
[0078] In practical applications, multi-agent reinforcement learning involves simultaneous learning and decision-making within the same environment. Each agent makes decisions based on its own observations, while also considering the behaviors and strategies of other agents. These agents may be cooperative, competitive, or a combination of both. The environment's feedback depends not only on the actions of individual agents but also on the actions of others; that is, an agent's optimal strategy may change as the strategies of other agents change. The multi-agent interaction process with the environment is as follows: Each agent selects an action based on its state, forming an action set that interacts with the environment. The environment then provides each agent with a set of states and a joint reward. Each agent then selects its next action based on the total reward and its own state.
[0079] It should be noted that the minimum plant-level operating cost is expressed as:
[0080]
[0081] N g C represents the number of thermal power units. i,t,g Let C be the operating cost of the i-th thermal power unit at time t; t,w C represents the operation and maintenance cost of the wind turbine generator at time t. t,pv C represents the operation and maintenance cost of the photovoltaic power generation unit at time t; t,es C represents the operation and maintenance cost of the energy storage device at time t. t,w,aban Let C be the cost of wind curtailment at time t; t,pv,aban Let be the cost of wasting light at time t;
[0082] The minimum carbon emissions are expressed as:
[0083]
[0084] e i,t,g Let be the carbon emissions of the i-th thermal power unit at time t;
[0085] The strongest capacity for wind and solar energy to absorb renewable energy is represented by:
[0086]
[0087] P t,w,abon P t,pv,abon These represent the amount of wind and solar power curtailed at time t, respectively.
[0088] In this embodiment, the multi-agent reinforcement learning module includes a neural network unit, a joint network, and an information interaction unit; it processes the observation data of each agent. Interaction information and actions of other intelligent agents The input is fed into the neural network unit, which selects actions and outputs the Q-values of each agent. i (τ i ,u i,t ), For the action observation trajectory of the i-th agent, the Q-values selected by each agent and the global state are input into the joint network. Based on the input global state, each local value function and global value function are trained to maximize the global Q-value: maxQ total =g({Q i} i∈n g(·) represents the relationship between the global value function and the local value function, realizing the global Q-value Q. total The optimization process involves encoding the historical observation data and actions of each agent using the information interaction unit and sending messages to other agents, represented as follows: M i,t-1 For the message generated by the i-th agent at time t-1, h i,t-1 Let φ be the hidden state of the LSTM network at time t-1, including the historical information of the i-th agent, and let φ be the parameters of the joint network.
[0089] In practical applications, neural network units filter out unreasonable actions based on current observations, calculate the corresponding Q-values, select appropriate actions for each agent, process the actions of each agent into an action set, and the joint network provides global control for each agent, calculating the global reward Q. total This approach considers the decisions of all agents, reflecting the overall collaborative situation and enhancing collaborative learning and autonomous decision-making capabilities. Through interaction between multiple agents and the power plant environment, agent collaboration is learned. A joint network is used to estimate the joint action value, which is then treated as a non-linear combination of each agent's values. This ensures that the joint action value is monotonic with respect to each agent's value, maximizing the joint action value during the learning process and achieving consistency between training and evaluation strategies.
[0090]
[0091] Q totalThe output of the joint network is generated based on the Q-value function estimation of agent i. Through the joint network, the strategies of each agent can be comprehensively considered to achieve effective cooperation in complex environments.
[0092] In the joint network, in order to enhance the cooperative performance of the agents, an experience replay mechanism is integrated in combination with the deep Q-learning algorithm. During training, the agent's observations, global state, actions, rewards and action sets are stored in the experience replay pool.
[0093] In the information interaction unit, agent i obtains historical information from other agents to supplement its own observations, represented as:
[0094]
[0095] For agent i, the final observation at time t includes local observations and received messages; E contains all messages received from agent i's neighboring agents at time t-1; i Let i be the set of agents adjacent to agent i; concat is the concatenation of the agent's local observations and received messages to form a combined observation vector.
[0096] In this embodiment, step S4, which involves performing device-level multi-agent optimization control modeling, includes:
[0097] Each thermal power unit defines its steam turbine, boiler body, boiler auxiliary equipment, and heating steam extraction device as a corresponding intelligent entity. The steam turbine intelligent entity is used for the operation control of the steam turbine, including the adjustment of speed, power, and temperature parameters. The boiler body intelligent entity is used for the control of the steam-water system and combustion system, including boiler water level, main steam pressure, steam-water ratio, fuel supply, air volume adjustment, and furnace temperature. The boiler auxiliary equipment intelligent entity is used for the control of ventilation equipment, coal conveying equipment, pulverizing equipment, water supply equipment, and ash removal equipment. The heating steam extraction device intelligent entity is used for the control of heating steam extraction, adjusting the extraction steam volume and pressure according to the heating load demand.
[0098] Multi-agent reinforcement learning is applied to equipment optimization control. It integrates equipment operation data, inter-device coupling and interaction information, and equipment parameter control mechanisms from lower-level device-level control agents. Each device agent makes autonomous decisions based on its current state and collaboration with other device agents. The equipment optimization control problem is modeled as a Markov decision process, expressed as:
[0099] W =<M,S,A,P,O′,r′,γ′> ;
[0100] M = {1, 2, ..., m} is a set of m device agents; S is the state space observed by all device agents; A is the action space for each device agent; P is the transition probability function from any state to any state after taking a joint action; O′ is the observation space of the device agents; r′ is the reward function, set according to the goal of minimizing load allocation and operating costs; γ′ is the discount factor; where each device agent j selects an action a. j ∈A, forming a joint action vector a={a1,a2,…,a n}∈A M .
[0101] In this embodiment, in step S4, when solving for the optimal control instructions for each device agent, an improved multi-agent deep deterministic policy gradient algorithm is adopted: a value function decomposition method is introduced to decompose the value function and quantify the contribution of each device agent.
[0102] In this embodiment, in the improved multi-agent deep deterministic policy gradient algorithm, the Actor network is responsible for generating actions, and the Critic network is responsible for calculating the value function of the environment state and the current action selected by the Actor network. τ j For the historical observation information of the j-th device intelligent agent; φ j Value function Network parameters; μ j For deterministic policies, and by decomposing the value function through a hybrid network, each device agent is allowed to make independent decisions based on its own observations and historical information, achieving a balance between individual autonomous learning and global optimization.
[0103] In this embodiment, the device agents share a centralized network of critics, the coefficients of which are represented as follows:
[0104]
[0105] φ is the value function of the joint action. The parameters are: τ is the set of historical observations of all device agents; a is the set of actions of all device agents; g ψ It is a hybrid network and is a nonlinear monotonic function;
[0106] The critic network is trained by minimizing a loss function, as follows:
[0107]
[0108] θ - φ - ψ -θ represents the parameters of the Actor network, Critic network, and hybrid network, respectively; μ(τ′; θ) - ) is the strategy given to the Actor network.
[0109] In this embodiment, the Actor network includes a hidden layer and a gated recurrent unit (GRU) layer; the Critic network includes a hidden layer and integrates a hybrid network.
[0110] It should be noted that in the Actor network, the hidden layers are used to process input information and extract features, while the GRU layer is used to enhance the processing capability of time series data, thereby improving the understanding of environmental dynamics; the Critic network, through the introduction of a hybrid network, can effectively integrate the value estimates of multiple agents, and achieve accurate estimation of the global value function in complex multi-agent environments.
[0111] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0112] Furthermore, the functional modules in the various embodiments of this invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. If the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the steps of the methods in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0113] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A hierarchical optimization control method for smart thermal power plants based on multi-agent systems, characterized in that, It includes: Step S1: Set up a multi-agent system for a smart thermal power plant, including at least the intelligent agents of each thermal power unit, wind, solar and energy storage intelligent agents, simulation intelligent agents, upper-level plant-level scheduling and decision-making intelligent agents, and lower-level equipment-level control intelligent agents, and build a multi-agent simulation model that can reproduce the actual operation of the smart thermal power plant. Step S2: The inference agent uses a multi-agent simulation model to obtain historical and real-time operating data of the thermal power plant. Combined with external temperature, humidity, wind speed, solar radiation intensity, electricity price information, and peak and valley load information, the inference agent uses these as environmental state information to predict the thermal power plant's electrical and thermal load, wind and solar power generation, and energy storage. Based on the prediction results, the inference agent then extrapolates the future operating status of the thermal power plant. Step S3: The upper-level plant-level scheduling decision-making agent acquires the observation data of each thermal power unit agent, wind-solar-storage agent, and inference agent, and calls the preset multi-agent reinforcement learning module to model the thermal power plant's electrical and thermal load optimization allocation problem. It then performs plant-level multi-agent action space design, state space design, and multi-objective reward function design to solve for the optimal load allocation results of each thermal power unit agent and wind-solar-storage agent. The multi-objective reward function includes minimizing plant-level operating costs, minimizing carbon emissions, and maximizing the absorption capacity of wind and solar new energy sources. Step S4: The lower-level equipment-level control agent, based on the optimal load allocation results of each thermal power unit's agent, defines the steam turbine, boiler body, boiler auxiliary equipment, and heating steam extraction device in each thermal power unit as the corresponding equipment agent. It integrates the interaction information between the equipment agents, equipment operation data, and equipment parameter control mechanisms. With the goal of minimizing the load allocation results and operating costs, it uses Markov decision process to perform equipment-level multi-agent optimization control modeling and solves the optimal control commands for each equipment agent. In step S4, when solving for the optimal control instructions for each device agent, an improved multi-agent deep deterministic policy gradient algorithm is used: a value function decomposition method is introduced to decompose the value function and quantify the contribution of each device agent. In the improved multi-agent deep deterministic policy gradient algorithm, the Actor network is responsible for generating actions, and the Critic network is responsible for calculating the value function of the environment state and the current action selected by the Actor network. , For the first Historical observation information of each device's intelligent agent; For each device intelligent agent Select an action; Value function Network parameters; For deterministic strategies, and by decomposing the value function through a hybrid network, each device agent is allowed to make independent decisions based on its own observations and historical information, achieving a balance between individual autonomous learning and global optimization; The device agents share a centralized network of commenters, whose coefficients are represented as follows: ; Value function of joint actions Parameters; The historical observation set of all device agents; The set of actions for all device agents; It is a hybrid network and is a nonlinear monotonic function; The number of intelligent agents in the device; The commentator network is trained by minimizing a loss function, as follows: ; ; For the reward function; Discount factor; , , These are the parameters for the Actor network, Critic network, and hybrid network, respectively. The strategy given for the Actor network.
2. The hierarchical optimization control method for intelligent thermal power plants according to claim 1, characterized in that, In step S1, a multi-agent simulation model capable of reproducing the actual operation of a smart thermal power plant is constructed, including: The actual operating data of the thermal power plant is read into each thermal power unit intelligent agent, wind, solar and energy storage intelligent agent, simulation intelligent agent, upper-level plant-level scheduling and decision-making intelligent agent, and lower-level equipment-level control intelligent agent. Combined with the operating mechanism of each intelligent agent, a simulation model of each intelligent agent in the thermal power plant is established. The operating process rules and basic functions of each intelligent agent are reproduced and transformed into rule knowledge base, static knowledge base and dynamic knowledge base of each intelligent agent. The system integrates the operating data, rule knowledge base, interactive information, and scheduling control parameters of each thermal power unit intelligent agent, wind, solar and energy storage intelligent agent, inference intelligent agent, upper-level plant-level scheduling decision intelligent agent, and lower-level equipment-level control intelligent agent using a pre-set modeling master intelligent agent. At the same time, it uses a pre-set environmental intelligent agent to simulate the external environmental conditions of the thermal power plant, forming a multi-agent simulation model of the thermal power plant with the cooperation of each intelligent agent. The accuracy of the simulation model is judged based on the error between the simulation data output by the multi-agent simulation model of the thermal power plant and the actual data. If it does not meet expectations, the model parameters are modified until the error is within a reasonable range, and finally a multi-agent simulation model that can reproduce the actual operation of the smart thermal power plant is obtained.
3. The hierarchical optimization control method for intelligent thermal power plants according to claim 1, characterized in that, Step S2 includes: Based on environmental state information, the inferential agent selects different feature extraction algorithms and prediction models as action spaces to predict the thermal load of thermal power plants, wind and solar power generation, and energy storage. The prediction accuracy of the prediction model is used as the reward function to select the optimal prediction model. The corresponding prediction results and external dynamic parameters for future periods are input into the multi-agent simulation model to infer the future operating state of thermal power plants.
4. The hierarchical optimization control method for intelligent thermal power plants according to claim 1, characterized in that, Step S3 includes: The upper-level plant-level scheduling decision-making intelligent agent acquires observation data from each thermal power unit intelligent agent, wind-solar-storage intelligent agent, and inference intelligent agent, including at least the power generation and future predicted power generation of each thermal power unit, the current power generation and future predicted wind-solar power generation, energy storage, the operating coal consumption characteristics of each thermal power unit, the electricity and heat load demand, weather conditions, and electricity price information. The pre-defined multi-agent reinforcement learning module for scheduling models the optimal allocation of electrical and thermal loads in thermal power plants, represented as follows: ; This represents the overall state for plant-level scheduling decisions; The action set includes the actions chosen by each agent; each agent has its own observations. Intelligent agents possess strategies ; The number of agents; Discount factor Among them, the plant-level multi-agent state space includes the observation data of each agent, the action space includes the running state and running output of each agent, and the reward function is a multi-objective reward function; In the upper-level plant-level scheduling decision-making intelligent agent, each intelligent agent uses neural network units to store observation data and historical actions; Initialize the environment, network training parameters, and experience replay pool for optimal allocation of electrical and thermal loads in thermal power plants; Using the first The observation data obtained by the agent at time t Value of each action Select Action and generate messages To communicate; The trajectory sets of each agent are stored in the experience replay pool for sampling and training. Obtain the action value of each agent at time t Value and global reward Training is performed, and the overall objective function loss is calculated to update the network parameters of the multi-agent reinforcement learning module; The multi-agent reinforcement learning module, which has been trained and updated, outputs the optimal load allocation results for each thermal power unit agent and the wind, solar and energy storage agent.
5. The hierarchical optimization control method for intelligent thermal power plants according to claim 4, characterized in that, The multi-agent reinforcement learning module includes a neural network unit, a joint network, and an information interaction unit; it processes the observation data of each agent. Interaction information and actions of other intelligent agents The input is fed into the neural network unit, which selects actions and outputs the results for each agent. value , For the first The action observation trajectory of each agent, and the selected agents The local and global values are input into a joint network. Based on the input global state, each local and global value function is trained to maximize the global value. value: , To establish the relationship between global value functions and local value functions, implement global... value The optimization process involves encoding the historical observation data and actions of each agent using the information interaction unit and sending messages to other agents, represented as follows: , For the first The message generated by the agent at time t-1 Let be the hidden state of the LSTM network at time t-1, including the t-1... Historical information of each intelligent agent These are the parameters of the joint network.
6. The hierarchical optimization control method for intelligent thermal power plants according to claim 1, characterized in that, In step S4, device-level multi-agent optimization control modeling is performed, including: Each thermal power unit defines its steam turbine, boiler body, boiler auxiliary equipment, and heating steam extraction device as a corresponding intelligent entity. The steam turbine intelligent entity is used for the operation control of the steam turbine, including the adjustment of speed, power, and temperature parameters. The boiler body intelligent entity is used for the control of the steam-water system and combustion system, including boiler water level, main steam pressure, steam-water ratio, fuel supply, air volume adjustment, and furnace temperature. The boiler auxiliary equipment intelligent entity is used for the control of ventilation equipment, coal conveying equipment, pulverizing equipment, water supply equipment, and ash removal equipment. The heating steam extraction device intelligent entity is used for the control of heating steam extraction, adjusting the extraction steam volume and pressure according to the heating load demand. Multi-agent reinforcement learning is applied to equipment optimization control. It integrates equipment operation data, inter-device coupling and interaction information, and equipment parameter control mechanisms from lower-level device-level control agents. Each device agent makes autonomous decisions based on its current state and collaboration with other device agents. The equipment optimization control problem is modeled as a Markov decision process, expressed as: ; for A collection of intelligent devices; The state space observed by all device agents; The action space for each device agent; Let be the transition probability function from any state to any state after taking joint action; For the observation space of the device's intelligent agent; The reward function is set based on satisfying the load allocation results and the goal of minimizing operating costs; among which, the action... To form a joint action vector .
7. The hierarchical optimization control method for intelligent thermal power plants according to claim 1, characterized in that, The Actor network includes a hidden layer and a gated recurrent unit (GRU) layer; the Critic network includes a hidden layer and integrates a hybrid network.
Citation Information
Patent Citations
Virtual power plant intelligent control method and system based on multiple agents
CN120474103A