Intelligent thermal power plant layered optimization control method based on multiple agents
By setting up multiple agents in thermal power plants and using multi-agent simulation models and reinforcement learning algorithms, collaborative cooperation and optimization control of each agent are achieved, and complex scheduling and control problems of thermal power plants in the joint mode of new energy and thermal power units are solved, and economic benefits and decision-making reliability are improved.
Patent Information
- Application Number
- CN202510178744.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The thermal power plants face complex unit scheduling and equipment control under the joint model of new energy and thermal power units, and the operation output of each unit is coupled and correlated, making it difficult to achieve effective coordinated coupling to improve economic benefits.
The layered optimization control method of smart thermal power plants based on multi-agents is adopted, and the multi-agents include thermal power unit agents, wind and light storage agents, deduction agents and upper and lower-level scheduling decision-making agents are set. Through the multi-agent simulation model and multi-agent reinforcement learning algorithm, the collaborative cooperation and optimization control of each agent are realized.
It improves the adaptability and collaboration efficiency of multiple agents in complex thermal power plant environments, realizes cost-effective scheduling, and enhances the reliability of algorithmic decision-making.
Smart Images

Figure CN120044789A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent thermal power plants, and particularly relates to a hierarchical optimization control method for intelligent thermal power plants based on multi - agents. Background Technique
[0002] In recent years, under the policy background of the country's strong development of clean energy, although the installed capacity of new energy sources such as photovoltaic and wind energy has been continuously increasing, at present, coal - fired units in China are still the main equipment for energy supply. Thermal power units will still be the main force in the production of electric energy and heat energy for a long time. To build an efficient, clean and sustainable development energy supply industry for thermal power plants, in addition to strictly setting the energy efficiency access threshold for newly built units and implementing energy conservation and emission reduction upgrades and transformations for thermal power unit equipment in existing thermal power plants, etc., it is also necessary to reasonably allocate the output power of thermal power units, reduce the total operating cost and fuel consumption of thermal power plants, so as to reasonably optimize resources, improve the unit load rate and operating quality, and achieve the best economic and social benefits.
[0003] At present, thermal power plants are gradually adopting the combined mode of new energy units and thermal power units for energy supply. The main problems faced are: due to the uncertainty of new energy output and the variability of the external environment, etc., the unit scheduling and equipment control of thermal power plants become more complex. In addition, the operating outputs of each thermal power unit and new energy unit in a thermal power plant have coupled correlations. Each unit not only has to consider its own operating conditions, but also has to comprehensively consider the operating conditions of other units in order to achieve effective collaborative coupling of the units and improve the economic benefits of the thermal power plant.
[0004] Based on the above - mentioned technical problems, it is necessary to design a new hierarchical optimization control method for intelligent thermal power plants based on multi - agents. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a hierarchical optimization control method for intelligent thermal power plants based on multi - agents, which can regard each device of the thermal power plant as an independent decision - making agent, and based on the current state and the behavior interaction information of other agents, effectively realize the collaborative cooperation of multi - agents, improve the adaptability and cooperation efficiency of multi - agents in a complex thermal power plant environment, achieve economic and efficient scheduling, and at the same time use the corresponding multi - agent reinforcement learning algorithm. Through continuous training, the reliability of algorithm decision - making is improved.
[0006] To solve the above - mentioned technical problems, the technical solution of the present invention is:
[0007] The present invention provides a hierarchical optimization control method for intelligent thermal power plants based on multi - agents, which includes:
[0008] Step S1: Set up the multi - agent system for the intelligent thermal power plant, including at least intelligent agents for each thermal power unit, intelligent agents for wind - solar - energy storage, a deduction intelligent agent, an upper - level plant - level scheduling decision - making intelligent agent, and a lower - level equipment - level control intelligent agent, and construct a multi - agent simulation model that can reproduce the actual operation of the intelligent thermal power plant;
[0009] Step S2: The deduction intelligent agent uses the multi - agent simulation model to obtain the historical operation data and real - time operation data of the thermal power plant, combines external dynamic parameters to predict the thermal - electric load, wind - solar power generation, and energy storage of the thermal power plant, and deduces the future operation state of the thermal power plant according to the prediction results;
[0010] Step S3: The upper - level plant - level scheduling decision - making intelligent agent obtains the observation data of each thermal power unit intelligent agent, wind - solar - energy storage intelligent agent, and deduction intelligent agent, and calls the preset multi - agent reinforcement learning module to model the thermal - electric load optimal distribution problem of the thermal power plant, design the plant - level multi - agent action space, state space, and multi - objective reward function, and solve to obtain the optimal load distribution results of each thermal power unit intelligent agent and wind - solar - energy storage intelligent agent;
[0011] Step S4: The lower - level equipment - level control intelligent agent is based on the optimal load distribution results of each thermal power unit intelligent agent, defines the steam turbine, boiler body, boiler auxiliary equipment, and heat - supply extraction device in each thermal power unit as corresponding equipment intelligent agents, integrates the interaction information, equipment operation data, and equipment parameter control mechanisms among the equipment intelligent agents, and takes the goal of meeting the load distribution results and minimizing the operation cost to conduct equipment - level multi - agent optimal control modeling, and solve to obtain the optimal control instructions of each equipment intelligent agent.
[0012] Furthermore, in step S1, constructing a multi - agent simulation model that can reproduce the actual operation of the intelligent thermal power plant includes:
[0013] Read the actual operation data of the thermal power plant into each thermal power unit intelligent agent, wind - solar - energy storage intelligent agent, deduction intelligent agent, upper - level plant - level scheduling decision - making intelligent agent, and lower - level equipment - level control intelligent agent, and combine the operation mechanisms of each intelligent agent to establish simulation models of each intelligent agent in the thermal power plant, reproduce the operation process rules and basic functions of each intelligent agent, and transform them into rule knowledge bases, static knowledge bases, and dynamic knowledge bases of each intelligent agent;
[0014] Use the preset modeling main intelligent agent to integrate the operation data, rule knowledge bases, interaction information, and scheduling control parameters of each thermal power unit intelligent agent, wind - solar - energy storage intelligent agent, deduction intelligent agent, upper - level plant - level scheduling decision - making intelligent agent, and lower - level equipment - level control intelligent agent. At the same time, use the preset environment intelligent agent to simulate the external environmental conditions of the thermal power plant to form a multi - agent simulation model of the thermal power plant in which each intelligent agent collaborates;
[0015] The accuracy of the simulation model is judged according to the error between the simulation data and the actual data output by the multi-agent simulation model of the thermal power plant. If it does not meet the expectation, the model parameters are modified until the error is within a reasonable range, and finally a multi-agent simulation model that can reproduce the actual operation of the intelligent thermal power plant is obtained.
[0016] Further, the step S2 includes:
[0017] The inference agent uses the multi-agent simulation model to obtain the historical operation data and real-time operation data of each agent in the thermal power plant, and combines the external temperature, humidity, wind speed, solar radiation intensity, electricity price information, and load peak-valley period information as the environmental state information of the inference agent.
[0018] Based on the environmental state information, the inference agent selects different feature extraction algorithms and prediction models as the action space to predict the thermal and electric loads, wind and solar power generation, and energy storage capacity of the thermal power plant, and takes the optimal prediction accuracy of the prediction model as the reward function to select the optimal prediction model, and inputs the corresponding prediction results and external dynamic parameters in the future period into the multi-agent simulation model to deduce the future operation state of the thermal power plant.
[0019] Further, the step S3 includes:
[0020] The upper-level plant-level scheduling decision-making agent obtains the observation data of each thermal power unit agent, wind-solar-storage agent, and inference agent, including at least the power generation and future predicted power generation of each thermal power unit, the current power generation and future predicted wind-solar power generation of wind and solar, the energy storage capacity, the operating coal consumption characteristics of each thermal power unit, the thermal and electric load demand, weather conditions, and electricity price information.
[0021] The pre-set multi-agent reinforcement learning module for scheduling models the optimal allocation of the thermal and electric loads of the thermal power plant as:
[0022] M = <G, U, r, O, n, π, γ>;
[0023] G is the global state of the plant-level scheduling decision; U is the action set, including the actions selected by each agent; r is the reward function; each agent has its own observation value o ∈ O, and the agent has a policy π; n is the number of agents; γ is the discount factor γ ∈ [0, 1]; among them, the plant-level multi-agent state space includes the observation data of each agent, the action space includes the operation state and operation output of each agent, and the multi-objective reward function includes the minimum plant-level operation cost, the minimum carbon emission, and the strongest wind-solar new energy consumption capacity.
[0024] In the upper-level plant-level scheduling decision-making agent, each agent uses a neural network unit to store the observation data and historical actions.
[0025] Initialize the environment, network training parameters, and experience replay pool for the optimal distribution of thermal and electrical loads in a thermal power plant;
[0026] Use the observation data obtained by the i-th agent at time t Each action value Q i , select an action And generate a message Conduct communication;
[0027] Store the trajectory sets of each agent in the experience replay pool for sampling and training;
[0028] Obtain the action value Q value and global reward Q of each agent at time t total , conduct training, and calculate the loss of the overall objective function to update the network parameters of the multi-agent reinforcement learning module;
[0029] Use the multi-agent reinforcement learning module after training and updating to output the optimal load distribution results of each thermal power generation unit agent and wind-solar-storage agent.
[0030] Furthermore, the multi-agent reinforcement learning module includes a neural network unit, a joint network, and an information interaction unit; input the observation data of each agent itself The interaction information and actions of other agents Into the neural network unit, select an action and output the Q value Q of each agent i (τ i , u i,t ), Is the action observation trajectory of the i-th agent, input the Q values selected by each agent and the global state into the joint network, and train each local value function and global value function according to the input global state, so that its goal is to maximize the global Q value: maxQ total = g({Q i} i∈n ), g(·) is the relationship between the global value function and the local value function, realize the optimization of the global Q value Q total , and at the same time use the information interaction unit to encode the historical observation data and actions of each agent and send messages to other agents, expressed as: M i,t-1 Is the message generated by the i-th agent at time t-1, h i,t-1 Is the hidden state of the LSTM network at time t-1, including the historical information of the i-th agent, and φ is the parameter of the joint network.
[0031] Furthermore, in step S4, device-level multi-agent optimization control modeling is performed, including:
[0032] Define the steam turbine, boiler body, boiler auxiliary equipment, and heat extraction steam device in each thermal power unit as the corresponding equipment agents; the steam turbine agent is used to control the operation of the steam turbine, including regulating rotational speed, power, and temperature parameters; the boiler body agent is used to control the steam-water system and combustion system, including boiler water level, main steam pressure, steam-water ratio, fuel supply, air volume regulation, and furnace temperature; the boiler auxiliary equipment agent is used to control ventilation equipment, coal conveying equipment, coal pulverizing equipment, feed water equipment, and ash removal equipment; the heat extraction steam device agent is used to control heat extraction steam, and adjusts the extraction steam volume and pressure according to the heat supply load demand;
[0033] Apply multi-agent reinforcement learning to equipment optimization control. Use the lower-layer device-level control agent to integrate the equipment operation data, coupling interaction information between devices, and device parameter control mechanisms of each equipment agent. And each equipment agent makes autonomous decisions based on the current state and cooperation with other equipment agents. Model the equipment optimization control problem as a Markov decision process, expressed as:
[0034] W = <M, S, A, P, O′, r′, γ′>;
[0035] M = {1, 2, …, m} is the set of m equipment agents; S is the state space observed by all equipment agents; A is the action space of each equipment agent; P is the transition probability function from any state to any state after taking joint actions; O′ is the observation space of the equipment agent; r′ is the reward function, which is set according to the goal of meeting the load distribution result and minimizing the operation cost; γ′ is the discount factor; where, each equipment agent j selects an action a j ∈A, forming a joint action vector a = {a 1 , a 2 , …, a n} ∈ A M .
[0036] Furthermore, in step S4, when solving to obtain the optimal control instructions of each equipment agent, use an improved multi-agent deep deterministic policy gradient algorithm: introduce the value function decomposition method to decompose the value function and quantitatively evaluate the contribution of each equipment agent respectively.
[0037] Furthermore, in the improved multi-agent deep deterministic policy gradient algorithm, the Actor network is responsible for generating actions, and the Critic network is responsible for calculating the value function of the environmental state and the current action selected by the Actor network τ j is the historical observation information of the j-th equipment agent; φ j is the network parameter of the value function μj For the deterministic policy and the decomposition of the value function through the hybrid network, each device agent is allowed to make independent decisions based on its own observations and historical information, achieving a balance between individual autonomous learning and global optimization.
[0038] Furthermore, the device agents share a centralized critic network, and its coefficients are expressed as:
[0039]
[0040] φ is the parameter of the joint action value function ; τ is the set of historical observations of all device agents; a is the set of actions of all device agents; g ψ is the hybrid network, which is a non - linear monotonic function;
[0041] The critic network is trained by minimizing the loss function, which is expressed as:
[0042]
[0043] θ - , φ - , ψ - are the parameters of the actor network, the critic network, and the hybrid network respectively; μ(τ′; θ - ) is the policy given by the actor network.
[0044] Furthermore, the actor network includes a hidden layer and a gated recurrent unit (GRU) layer; the critic network includes a hidden layer and integrates the hybrid network.
[0045] The beneficial effects of the present invention are:
[0046] By setting up multi - agents in the intelligent thermal power plant, establishing a multi - agent simulation model, and introducing a multi - agent reinforcement learning algorithm for upper - layer plant - level scheduling decisions and lower - layer device - level optimization control, the present invention can take each device of the thermal power plant as an independent decision - making agent, and based on the current state and the behavior interaction information of other agents, effectively achieve multi - agent collaborative cooperation, improve the adaptability and cooperation efficiency of multi - agents in the complex thermal power plant environment, realize economic and efficient scheduling, and at the same time, by using the corresponding multi - agent reinforcement learning algorithm and continuous training, improve the reliability of algorithm decision - making.
[0047] Other features and advantages will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification and the drawings.
[0048] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. Description of the Drawings
[0049] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 Flowchart of a hierarchical optimization control method for a smart thermal power plant based on multi-agent of the present invention;
[0051] Figure 2 Schematic diagram of the multi-agent architecture of the smart thermal power plant of the present invention. Specific Embodiments
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0053] Embodiment 1
[0054] As Figure 1 、 Figure 2 shown, Embodiment 1 of the present invention provides a hierarchical optimization control method for a smart thermal power plant based on multi-agent, which includes:
[0055] Step S1: Set up multi-agents for the smart thermal power plant, including at least smart agents for each thermal power unit, smart agents for wind-solar-storage, a deduction agent, an upper-level plant-level scheduling decision-making agent, and a lower-level equipment-level control agent, and construct a multi-agent simulation model that can reproduce the actual operation of the smart thermal power plant;
[0056] Step S2: The deduction agent uses the multi-agent simulation model to obtain the historical operation data and real-time operation data of the thermal power plant, combines external dynamic parameters to predict the thermal and electrical loads of the thermal power plant, the wind-solar power generation, and the energy storage capacity, and deduces the future operation state of the thermal power plant according to the prediction results;
[0057] Step S3: The upper-level plant-level scheduling decision-making agent obtains the observation data of each thermal power unit agent, wind-solar-storage agent, and deduction agent, and invokes the preset multi-agent reinforcement learning module to model the problem of optimizing the distribution of thermal and electric loads in the thermal power plant. It designs the action space, state space, and multi-objective reward function of the plant-level multi-agent, and solves to obtain the optimal load distribution results of each thermal power unit agent and wind-solar-storage agent.
[0058] Step S4: The lower-level equipment-level control agent, based on the optimal load distribution results of each thermal power unit agent, defines the steam turbine, boiler body, boiler auxiliary equipment, and heating extraction steam device in each thermal power unit as corresponding equipment agents, integrates the interaction information, equipment operation data, and equipment parameter control mechanisms among the equipment agents, and takes minimizing the load distribution result and operation cost as the goal to conduct equipment-level multi-agent optimization control modeling, and solves to obtain the optimal control instructions of each equipment agent.
[0059] In this embodiment, in step S1, a multi-agent simulation model capable of reproducing the actual operation of the intelligent thermal power plant is constructed, including:
[0060] Read the actual operation data of the thermal power plant into each thermal power unit agent, wind-solar-storage agent, deduction agent, upper-level plant-level scheduling decision-making agent, and lower-level equipment-level control agent, and combine the operation mechanisms of each agent to establish the simulation models of each agent in the thermal power plant, reproduce the operation process rules and basic functions of each agent, and convert them into the rule knowledge bases, static knowledge bases, and dynamic knowledge bases of each agent.
[0061] Use the preset modeling main agent to integrate the operation data, rule knowledge bases, interaction information, and scheduling control parameters of each thermal power unit agent, wind-solar-storage agent, deduction agent, upper-level plant-level scheduling decision-making agent, and lower-level equipment-level control agent. At the same time, use the preset environment agent to simulate the external environmental conditions of the thermal power plant to form a multi-agent simulation model of the thermal power plant in which each agent collaborates.
[0062] Judge the accuracy of the simulation model according to the error between the simulation data and the actual data output by the multi-agent simulation model of the thermal power plant. If it does not meet the expectation, modify the model parameters until the error is within a reasonable range, and finally obtain a multi-agent simulation model capable of reproducing the actual operation of the intelligent thermal power plant.
[0063] It should be noted that the static knowledge base includes the basic parameters of each thermal power unit and wind-solar-storage equipment; the dynamic knowledge base includes the operation data, scheduling strategies, and equipment control parameters of each thermal power unit and wind-solar-storage equipment. There is a demonstration animation module inside the modeling main agent and it simulates the operation scenario of the entire thermal power plant.
[0064] In this embodiment, step S2 includes:
[0065] The inference agent uses the multi-agent simulation model to obtain the historical operation data and real-time operation data of each agent in the thermal power plant, and combines the external temperature, humidity, wind speed, solar radiation intensity, electricity price information, and load peak-valley period information as the environmental state information of the inference agent;
[0066] Based on the environmental state information, the inference agent selects different feature extraction algorithms and prediction models as the action space to predict the thermal and electrical loads, wind and solar power generation, and energy storage capacity of the thermal power plant. Taking the optimal prediction accuracy of the prediction model as the reward function, the optimal prediction model is selected, and the corresponding prediction results and external dynamic parameters in the future time period are input into the multi-agent simulation model to deduce the future operation state of the thermal power plant.
[0067] In this embodiment, the step S3 includes:
[0068] The upper-level plant-level dispatching decision-making agent obtains the observation data of each thermal power unit agent, wind-solar-storage agent, and inference agent, including at least the power generation and future predicted power generation of each thermal power unit, the current power generation and future predicted wind-solar power generation of wind and solar, energy storage capacity, the operating coal consumption characteristics of each thermal power unit, thermal and electrical load demands, weather conditions, and electricity price information;
[0069] The pre-set multi-agent reinforcement learning module for dispatching models the optimal allocation of the thermal and electrical loads of the thermal power plant as:
[0070] M = <G, U, r, O, n, π, γ>;
[0071] G is the global state of the plant-level dispatching decision; U is the set of actions, including the actions selected by each agent; r is the reward function; each agent has its own observation value o ∈ O, and the agent has a policy π; n is the number of agents; γ is the discount factor γ ∈ [0, 1]; among them, the plant-level multi-agent state space includes the observation data of each agent, the action space includes the operation state and operation output of each agent, and the multi-objective reward function includes the minimum plant-level operation cost, the minimum carbon emission, and the strongest wind-solar new energy consumption capacity;
[0072] In the upper-level plant-level dispatching decision-making agent, each agent uses a neural network unit to save the observation data and historical actions;
[0073] Initialize the environment, network training parameters, and experience replay pool for the optimal allocation of the thermal and electrical loads of the thermal power plant;
[0074] Using the observation data obtained by the i-th agent at time t Each action value Q i , select the action and generate a message Communicate;
[0075] Store the trajectory sets of each agent into the experience replay pool for sampling and training;
[0076] Obtain the action-value Q-values and the global reward Q of each agent at time t total , conduct training, and calculate the loss of the overall objective function to update the network parameters of the multi-agent reinforcement learning module;
[0077] Use the multi-agent reinforcement learning module after training update to output the optimal load distribution results of each thermal power unit agent and wind-solar-storage agent.
[0078] In practical applications, multi-agent reinforcement learning is a scenario where learning and decision-making are carried out simultaneously in the same environment. Each agent makes decisions based on its own observations and needs to consider the behaviors and strategies of other agents. These agents may be in a cooperative relationship, a competitive relationship, or both. The feedback of the environment depends not only on the actions of a single agent but also on the actions of other agents. That is, the optimal strategy of one agent may change with the changes in the strategies of other agents. The interaction process between multi-agents and the environment: Each agent selects an action according to the state, forms an action set to interact with the environment, and then the environment will feedback a state set and a joint reward to each agent. The agent then selects the next action based on the total reward and the state of each agent.
[0079] It should be noted that the minimum plant-level operating cost is expressed as:
[0080]
[0081] N g is the number of thermal power units; C i,t,g is the operating cost of the i-th thermal power unit at time t; C t,w is the operation and maintenance cost of the wind turbine generator at time t; C t,pv is the operation and maintenance cost of the photovoltaic generator at time t; C t,es is the operation and maintenance cost of the energy storage device at time t; C t,w,aban is the curtailment cost of wind power at time t; C t,pv,aban is the curtailment cost of photovoltaic power at time t;
[0082] The minimum carbon emissions are expressed as:
[0083]
[0084] e i,t,g is the carbon emissions of the i-th thermal power unit at time t;
[0085] The strongest consumption capacity of wind-solar new energy is expressed as:
[0086]
[0087] P t,w,abon 、P t,pv,abon are the wind curtailment volume and the photovoltaic curtailment volume at time t, respectively.
[0088] In this embodiment, the multi-agent reinforcement learning module includes a neural network unit, a joint network, and an information interaction unit; the observation data of each agent itself the interaction information and actions of other agents are input into the neural network unit to select actions and output the Q-value Q i (τ i , u i,t ) of the i-th agent. The Q-values selected by each agent and the global state are input into the joint network. According to the input global state, each local value function and the global value function are trained to maximize the global Q-value: maxQ total = g({Q i} i∈n ), where g(·) is the relationship between the global value function and the local value function, to optimize the global Q-value Q total . At the same time, the information interaction unit encodes the historical observation data and actions of each agent and sends messages to other agents, which is expressed as: M i,t-1 is the message generated by the i-th agent at time t-1, h i,t-1 is the hidden state of the LSTM network at time t-1, including the historical information of the i-th agent, and φ is the parameter of the joint network.
[0089] In actual applications, the neural network unit filters out unreasonable actions according to the current observation value, calculates the corresponding Q-value, selects the appropriate next action for each agent, and then processes the actions of each agent into an action set. The joint network provides global control for each agent, calculates the global reward Q total , considers the decisions of all agents, reflects the global cooperation situation, and enhances the collaborative learning and autonomous decision-making capabilities. Through the interactive learning of multi-agents and the thermal power plant environment, the agent cooperation is carried out. The joint network estimates the joint action value, which is used as a non-linear combination of each agent value, to ensure that the joint action value is monotonic with respect to each agent value, and maximizes the joint action value during the learning process to achieve the consistency between the training and evaluation strategies:
[0090]
[0091] Q total It is the output result of the joint network, which is generated based on the Q-value function estimation of agent i. Through the joint network, the strategies of each agent can be comprehensively considered to achieve effective cooperation in complex environments.
[0092] In the joint network, to enhance the collaborative performance of agents, the experience replay mechanism is integrated in combination with the deep Q-learning algorithm. During training, the observations, global states, actions, rewards, and action sets of agents are stored in the experience replay pool.
[0093] In the information interaction unit, agent i obtains historical information from other agents to supplement its own observations, which is expressed as:
[0094]
[0095] is the final observation of agent i at time t, including local observations and received messages; is the set of all messages received from the neighboring agents of agent i at time t - 1; E i is the set of neighboring agents of agent i; concat is to concatenate the local observations and received messages of the agent to form a comprehensive observation vector.
[0096] In this embodiment, in step S4, device-level multi-agent optimization control modeling is performed, including:
[0097] The steam turbines, boiler bodies, boiler auxiliary equipment, and heat extraction and supply devices in each thermal power unit are defined as corresponding device agents; the steam turbine agent is used to control the operation of the steam turbine, including regulating the rotational speed, power, and temperature parameters; the boiler body agent is used to control the steam-water system and combustion system, including boiler water level, main steam pressure, steam-water ratio, fuel supply, air volume regulation, and furnace temperature; the boiler auxiliary equipment agent is used to control ventilation equipment, coal conveying equipment, coal pulverizing equipment, water supply equipment, and ash removal equipment; the heat extraction and supply device agent is used to control heat extraction and supply, and adjust the extraction steam volume and pressure according to the heat supply load demand;
[0098] Applying multi-agent reinforcement learning to device optimization control, using the lower-layer device-level control agents to integrate the device operation data, coupling interaction information between devices, and device parameter control mechanisms of each device agent, and each device agent makes autonomous decisions based on the current state and cooperation with other device agents, and models the device optimization control problem as a Markov decision process, which is expressed as:
[0099] W = <M, S, A, P, O′, r′, γ′>;
[0100] Let \(M = \{1, 2, \ldots, m\}\) be a set of \(m\) device agents; \(S\) is the state space observed by all device agents; \(A\) is the action space of each device agent; \(P\) is the transition probability function from any state to any state after taking a joint action; \(O'\) is the observation space of the device agent; \(r'\) is the reward function, which is set according to the goal of satisfying the load distribution result and minimizing the operating cost; \(\gamma'\) is the discount factor; where each device agent \(j\) selects an action \(a\) j \(\in A\), forming a joint action vector \(a=\{a\) 1 , a\) 2 , \ldots, a\) n \}\in A M .
[0101] In this embodiment, in step S4, when solving for the optimal control instructions of each device agent, an improved multi-agent deep deterministic policy gradient algorithm is adopted: introducing a value function decomposition method to decompose the value function and separately quantify and evaluate the contributions of each device agent.
[0102] In this embodiment, in the improved multi-agent deep deterministic policy gradient algorithm, the Actor network is responsible for generating actions, and the Critic network is responsible for calculating the value function of the environmental state and the current action selected by the Actor network \(\tau\) j is the historical observation information of the \(j\)-th device agent; \(\varphi\) j is the network parameter of the value function \(\mu\) j is the deterministic policy, and by decomposing the value function through a hybrid network, each device agent is allowed to make independent decisions based on its own observations and historical information, achieving a balance between individual autonomous learning and global optimization.
[0103] In this embodiment, the device agents share a centralized Critic network, and its coefficients are expressed as:
[0104]
[0105] \(\varphi\) is the parameter of the joint action value function ; \(\tau\) is the set of historical observations of all device agents; \(a\) is the set of actions of all device agents; \(g\) ψ is a hybrid network, which is a non-linear monotonic function;
[0106] The Critic network is trained by minimizing the loss function, which is expressed as:
[0107]
[0108] \(\theta\) - , \(\varphi\) -, ψ - are the parameters of the Actor network, the Critic network, and the hybrid network respectively; μ(τ′; θ - ) is the policy given by the Actor network.
[0109] In this embodiment, the Actor network includes a hidden layer and a gated recurrent unit (GRU) layer; the Critic network includes a hidden layer and integrates the hybrid network.
[0110] It should be noted that the hidden layer in the Actor network is used to process input information and extract features, and the GRU layer is used to enhance the processing ability of time series data, thereby improving the understanding of the environmental dynamics; by introducing the hybrid network, the Critic network can effectively fuse the value estimates of multiple agents and achieve an accurate estimate of the global value function in a complex multi-agent environment.
[0111] In several embodiments provided by this application, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0112] In addition, each functional module in various embodiments of the present invention may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes.
[0113] Inspired by the ideal embodiments of the present invention described above, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A multi-agent-based intelligent thermal power plant hierarchical optimization control method, characterized in that: It includes: Step S1, setting up a multi-agent of a smart thermal power plant, including at least the agents of each thermal power unit, the wind, solar and storage agents, the deduction agent, the upper-level plant-level dispatching decision agent, and the lower-level equipment-level control agent, and building a multi-agent simulation model that can reproduce the actual operation of the smart thermal power plant; Step S2, the deduction agent uses the multi-agent simulation model to obtain the historical operation data and real-time operation data of the thermal power plant, combines the external dynamic parameters to predict the thermal power plant's electric and thermal load, wind and solar power generation, and storage capacity, and deduces the future operation status of the thermal power plant based on the prediction results; Step S3, the upper-level plant-level dispatch decision agent obtains the observation data of each thermal power unit agent, wind, solar and storage agent and deduction agent, and calls the preset multi-agent reinforcement learning module to model the thermal power plant electric and thermal load optimization allocation problem, and performs plant-level multi-agent action space design, state space design and multi-objective reward function design to solve and obtain the optimal load allocation results of each thermal power unit agent and wind, solar and storage agent; Step S4, the lower-level equipment-level control agent is based on the optimal load distribution results of each thermal power unit agent, and defines the steam turbine, boiler body, boiler auxiliary equipment, and heating steam extraction device in each thermal power unit as the corresponding equipment agent, integrates the interactive information and equipment operation data and equipment parameter control mechanism between each equipment agent, and performs equipment-level multi-agent optimization control modeling to achieve the goal of satisfying the load distribution results and minimizing the operating cost, and solves and obtains the optimal control instructions for each equipment agent.
2. The hierarchical optimization control method for a smart thermal power plant according to claim 1 is characterized in that: In step S1, a multi-agent simulation model capable of reproducing the actual operation of a smart thermal power plant is constructed, including: Read the actual operation data of the thermal power plant into each thermal power unit intelligent agent, wind, solar and storage intelligent agent, deduction intelligent agent, upper-level plant-level dispatch decision intelligent agent, and lower-level equipment-level control intelligent agent, and combine the operation mechanism of each intelligent agent to establish a simulation model of each intelligent agent in the thermal power plant, reproduce the operation process rules and basic functions of each intelligent agent, and convert them into rule knowledge bases, static knowledge bases and dynamic knowledge bases of each intelligent agent; The preset modeling main agent is used to integrate the operation data, rule knowledge base, interaction information and dispatch control parameters of each thermal power unit agent, wind, solar and storage agent, deduction agent, upper-level plant-level dispatch decision agent, and lower-level equipment-level control agent. At the same time, the preset environmental agent is used to simulate the external environmental conditions of the thermal power plant to form a multi-agent simulation model of a thermal power plant in which each agent collaborates. The accuracy of the simulation model is judged based on the error between the simulation data output by the multi-agent simulation model of the thermal power plant and the actual data. If it does not meet expectations, the model parameters are modified until the error is within a reasonable range. Finally, a multi-agent simulation model that can reproduce the actual operation of the smart thermal power plant is obtained.
3. The hierarchical optimization control method for a smart thermal power plant according to claim 1 is characterized in that: The step S2 comprises: The deduction agent uses the multi-agent simulation model to obtain the historical and real-time operation data of each agent in the thermal power plant, and combines the external temperature, humidity, wind speed, solar radiation intensity, electricity price information, and load peak and valley period information as the environmental status information of the deduction agent; The deduction agent selects different feature extraction algorithms and prediction models as the action space according to the environmental state information, and performs thermal load forecasting, wind and solar power generation forecasting and energy storage forecasting of thermal power plants. The optimal prediction model is selected with the optimal prediction accuracy as the reward function, and the corresponding prediction results and external dynamic parameters of the future time period are input into the multi-agent simulation model to deduce the future operating status of the thermal power plant.
4. The hierarchical optimization control method for a smart thermal power plant according to claim 1 is characterized in that: The step S3 comprises: The upper-level plant-level dispatch decision-making agent obtains the observation data of each thermal power unit agent, wind, solar and storage agent and deduction agent, including at least the power generation of each thermal power unit and the future predicted power generation, the current power generation of wind and solar and the future predicted power generation of wind and solar, storage capacity, the operating coal consumption characteristics of each thermal power unit, electric and thermal load demand, weather conditions, and electricity price information; The multi-agent reinforcement learning module preset by the dispatcher models the optimal distribution of electric and thermal loads in the thermal power plant, which is expressed as: M=<G,U,r,O,n,π,γ> ; G is the global state of the plant-level scheduling decision; U is the action set, including the actions selected by each agent; r is the reward function; each agent has its own observation value o∈O, and the agent has a strategy π; n is the number of agents; γ is the discount factor γ∈[0,1]; among them, the plant-level multi-agent state space includes the observation data of each agent, the action space includes the operating status and operating output of each agent, and the multi-objective reward function includes the minimum plant-level operating cost, the minimum carbon emissions and the strongest wind and solar energy absorption capacity; The upper-level plant-level dispatch decision agent uses neural network units to save observation data and historical actions. Initialize the environment, network training parameters and experience playback pool for optimal distribution of thermal power plant electric and thermal loads; Using the observation data obtained by the i-th agent at time t Each action has a value of Q i , select action and generate the message to communicate; The trajectory set of each agent is stored in the experience replay pool for sampling and training; Get the action value Q value and global reward Q of each agent at time t total , conduct training, and calculate the overall objective function loss to update the network parameters of the multi-agent reinforcement learning module; The multi-agent reinforcement learning module that has completed training and updating is used to output the optimal load distribution results for each thermal power unit agent and wind, solar and storage agent.
5. The hierarchical optimization control method for a smart thermal power plant according to claim 4 is characterized in that: The multi-agent reinforcement learning module includes a neural network unit, a joint network and an information interaction unit; each agent's own observation data Interaction information and actions of other agents Input into the neural network unit, select actions and output the Q value Q of each agent i (τ i ,u i,t ), For the action observation trajectory of the i-th agent, the Q value selected by each agent and the global state are input into the joint network. According to the input global state, each local value function and the global value function are trained to maximize the global Q value: max Q total = g({Q i } i∈n ), g(·) is the relationship between the global value function and the local value function, realizing the global Q value Q total Optimize, and use the information interaction unit to encode the historical observation data and actions of each agent, and send messages to other agents, which is expressed as: M i,t-1 is the message generated by the i-th agent at time t-1, h i,t-1 is the hidden state of the LSTM network at time t-1, including the historical information of the i-th agent, and φ is the parameter of the joint network.
6. The hierarchical optimization control method for a smart thermal power plant according to claim 1 is characterized in that: In step S4, device-level multi-agent optimization control modeling is performed, including: The steam turbine, boiler body, boiler auxiliary equipment, and heat extraction device in each thermal power unit are defined as corresponding equipment intelligent bodies; the steam turbine intelligent body is used to control the operation of the steam turbine, including the adjustment of speed, power, and temperature parameters; the boiler body intelligent body is used to control the steam-water system and the combustion system, including boiler water level, main steam pressure, steam-water ratio, fuel supply, air volume adjustment, and furnace temperature; the boiler auxiliary equipment intelligent body is used to control ventilation equipment, coal transportation equipment, pulverizing equipment, water supply equipment, and ash removal equipment; the heat extraction device intelligent body is used to control heat extraction, and adjust the extraction volume and pressure according to the heat load demand; Multi-agent reinforcement learning is applied to equipment optimization control. The lower-level equipment-level control agent is used to integrate the equipment operation data of each equipment agent, the coupling interaction information between equipment, and the equipment parameter control mechanism. Each equipment agent makes autonomous decisions based on the current state and in collaboration with other equipment agents. The equipment optimization control problem is modeled as a Markov decision process, which can be expressed as: W=<M,S,A,P,O′,r′,γ′> ; M = {1, 2, ..., m} is a set of m device agents; S is the state space observed by all device agents; A is the action space of each device agent; P is the transition probability function from any state to any state after taking joint actions; O' is the observation space of the device agent; r' is the reward function, which is set according to the load distribution result and the minimum operating cost target; γ' is the discount factor; where each device agent j selects an action a j ∈A, forming a joint action vector a={a1,a2,…,a n }∈A M .
7. The hierarchical optimization control method for a smart thermal power plant according to claim 1 is characterized in that: In step S4, when solving for the optimal control instructions of each device agent, an improved multi-agent deep deterministic policy gradient algorithm is used: a value function decomposition method is introduced to decompose the value function and quantify and evaluate the contribution of each device agent.
8. The hierarchical optimization control method for a smart thermal power plant according to claim 7 is characterized in that: In the improved multi-agent deep deterministic policy gradient algorithm, the actor network is responsible for generating actions, and the critic network is responsible for calculating the value function of the environment state and the current action selected by the actor network. τ j is the historical observation information of the jth device agent; φ j is the value function The network parameters of j It is a deterministic strategy and decomposes the value function through a hybrid network, allowing each device agent to make independent decisions based on its own observations and historical information, achieving a balance between individual autonomous learning and global optimization.
9. The hierarchical optimization control method for a smart thermal power plant according to claim 8 is characterized in that: The device agents share a centralized critic network, whose coefficients are expressed as: φ is the joint action value function Parameters; τ is the historical observation set of all device agents; a is the action set of all device agents; g ψ It is a hybrid network and is a nonlinear monotonic function; The critic network is trained by minimizing the loss function, expressed as: θ - ,φ - , - are the parameters of the actor network, critic network and hybrid network respectively; μ(τ′; θ - ) is the strategy given by the actor network.
10. The hierarchical optimization control method for a smart thermal power plant according to claim 8, characterized in that: The actor network includes a hidden layer and a gated recurrent unit (GRU) layer; the critic network includes a hidden layer and an integrated hybrid network.
Citation Information
Patent Citations
Virtual power plant adjustable resource accurate control method considering demand response
CN113904380A
Method and device for acquiring energy scheduling strategy
CN117808259A
Layered optimization scheduling method for deep peak regulation of thermal power generating unit
CN118214088A
Plant-level multi-unit scheduling and control method containing digital twinning and boundary condition quantification
CN118889386A
Virtual power plant intelligent control method and system based on multiple agents
CN120474103A
Cited By
Boiler scheduling and control system based on multi-agent collaborative decision-making
CN121232747A
Boiler scheduling and control system based on multi-agent collaborative decision-making
CN121232747B
Thermal power plant operation control optimization method and system based on artificial intelligence
CN121300048A
Optimization control method and system based on lexicographical order multi-objective reinforcement learning
CN121704187A
Dynamic cooperative control method and system for intelligent power plant with few people on duty
CN121763716A