Energy storage auxiliary thermal power unit deep reinforcement learning load frequency control method

By optimizing the output strategy of the energy storage system through deep reinforcement learning, the problems of long response time and high control strategy complexity of the energy storage system in grid frequency regulation are solved, and the coordinated operation of energy storage and thermal power units and stable regulation of grid frequency are realized.

CN119051070BActive Publication Date: 2026-06-05NANJING UNIV OF POSTS & TELECOMM

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2024-07-15
Publication Date
2026-06-05

Smart Images

  • Figure CN119051070B_ABST
    Figure CN119051070B_ABST
Patent Text Reader

Abstract

The application discloses a kind of energy storage auxiliary thermal power unit deep reinforcement learning load frequency control methods, including establishing the deep reinforcement learning environment of power system load frequency control problem, to simulate the characteristics and control demand of each component in actual power system;For lithium ion battery energy storage system in power system component, design the energy storage system output control strategy considering power system frequency deviation and energy storage system state of charge;The power system load frequency control problem is modeled as Markov decision model, combined with Markov decision model and energy storage system output control law, build the deep reinforcement learning load frequency control framework based on Actor-Critic architecture;Combined with deep deterministic policy gradient (DDPG) algorithm and random network distillation technology, design the load frequency controller of double output, including agent state, action and multi-objective dynamic adaptive reward function;Design agent solving algorithm based on dynamic weight adjustment.The load frequency control method provided by the application does not depend on specific model, has strong environmental adaptability, compared with the traditional energy storage output constant control strategy, the complementary coordination control of energy storage system and thermal power unit can be realized, and the frequency modulation effect of power system is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of load frequency control, specifically relating to a deep reinforcement learning load frequency control method for energy storage-assisted thermal power units. Background Technology

[0002] Automatic generation control (AGC) is a crucial technology for achieving active power balance and system frequency stability in power grids. Currently, traditional AGC systems primarily utilize thermal power units, which suffer from drawbacks such as long response time delays, low ramp-up rates, and insufficient AGC command tracking capabilities. Furthermore, with the large-scale integration of renewable energy into the grid, its weak inertia, volatility, and uncertainty place higher demands on the grid's frequency regulation speed. Therefore, introducing higher-quality frequency regulation resources to address the strong random disturbances caused by large-scale renewable energy grid integration has significant academic and engineering application value.

[0003] In recent years, energy storage technologies, represented by electrochemical batteries, have developed rapidly, and their large-scale application in power systems is increasing rapidly. Currently, a large body of literature has conducted a series of studies on the application of energy storage batteries in secondary frequency regulation of power systems. The paper "Secondary frequency control strategy for BESS considering their degree of participation" (Energy Reports, 2020, 594-602) proposes a participation-based energy storage output control strategy to dynamically adjust the output of energy storage participating in secondary frequency regulation. The paper "Fuzzy logic control and battery energy storage system based power system secondary frequency regulation" (2023 7th International Conference on Green Energy and Applications, 2023) comprehensively considers regional control deviation (ACE) and energy storage SOC signal, proposing a fuzzy control strategy for energy storage output based on ACE and SOC to smooth the output of the energy storage system participating in secondary frequency regulation. However, most existing literature uses fuzzy control and piecewise functions to represent the energy storage output function. While this alleviates the overcharging and over-discharging problems of energy storage to some extent, it also introduces new issues. First, it weakens the rapid response characteristics of energy storage; second, it easily causes secondary disturbances to the system frequency at the critical point of output.

[0004] Therefore, designing a reasonable energy storage output function is crucial. Based on existing research, this invention selects the sigmoid function as the fundamental curve for energy storage output, using SOC as the independent variable and energy storage output power as the dependent variable to construct the charge / discharge output function curve of the energy storage system. This function ensures both the smoothness of energy storage output and takes into account the rapid response characteristics of energy storage.

[0005] With the grid integration of a high proportion of distributed renewable energy sources, the control performance of traditional controllers in these power grids needs further improvement when facing complex operating conditions such as numerous random disturbances, system parameter and structural changes. In recent years, reinforcement learning algorithms have been widely used in the field of load frequency control due to their advantages such as strong decision-making ability, high environmental adaptability, and the ability to continuously explore and learn from experience to obtain the optimal control strategy. The design of load frequency controllers based on reinforcement learning has become a hot topic in AGC research. The literature "A frequency control strategy of large grid with energy storage based on multi-agent algorithm" (2023 IEEE 3rd International Conference on Power, Electronics and Computer Applications, 2023, 181-184) proposes a load frequency control technology for energy storage grids based on the DDPG algorithm for load frequency control systems that include thermal power, renewable energy and energy storage batteries, so as to achieve the purpose of energy storage assisting grid frequency regulation. The paper "A safe policy learning-based method for decentralized and economic frequency control in isolated networked-microgrid systems" (IEEE Transactions on Sustainable Energy, 2022, 1982-1993) proposes a data-driven decentralized economic frequency control method for isolated microgrid systems. This method integrates energy storage output strategy as part of the controller design into a reinforcement learning controller, and limits the energy storage output under different scenarios through constraints. Based on existing literature, current research on energy storage participation in secondary frequency regulation still has the following shortcomings:

[0006] (1) The control strategy does not take into account the SOC state of the energy storage system. If the SOC is too high or too low, it will reduce the bidirectional frequency regulation capability of the energy storage battery, and there is no SOC self-recovery function of the energy storage system.

[0007] (2) The energy storage output control strategy is based on the unilateral frequency regulation demand of the power grid, without fully considering the frequency regulation capability of energy storage and the frequency regulation demand of the power grid, and fails to make full use of the complementary advantages between energy storage and the power grid.

[0008] (3) Existing methods use the output of energy storage as a constraint condition for controller design, which increases the complexity of controller design. In addition, the existing methods consider the energy storage output constraints in a relatively simple way and do not take into account the output characteristics of different energy storage SOC ranges.

[0009] (4) Existing methods only use a simple linear weighting method to solve the multi-objective reward function, which makes it difficult to determine reasonable multi-objective weights and thus cannot solve for the optimal strategy. Summary of the Invention

[0010] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a deep reinforcement learning load frequency control method for energy storage-assisted thermal power units. This method comprehensively considers the grid frequency regulation requirements and energy storage frequency regulation capabilities when providing frequency regulation actions, achieving complementary and coordinated operation of energy storage and conventional units. Simultaneously, it solves the coordination problem of multiple optimization objectives through dynamic weight adaptive adjustment.

[0011] To achieve the above objectives, the present invention provides a deep reinforcement learning load frequency control method for energy storage-assisted thermal power units, comprising the following steps:

[0012] S1. Establish a deep reinforcement learning environment for the load frequency control problem of power systems to simulate the characteristics and control requirements of various components in actual power systems;

[0013] S2. For lithium-ion battery energy storage systems in power system components, design an energy storage system output control strategy that takes into account both power system frequency deviation and energy storage system state of charge (SOC);

[0014] S3. The load frequency control problem of the power system is modeled as a Markov decision model, and a deep reinforcement learning load frequency control framework based on the Actor-Critic architecture is built by combining the Markov decision model and the output control law of the energy storage system.

[0015] S4. Combining Deep Deterministic Policy Gradient (DDPG) and stochastic network distillation techniques, a dual-output load frequency controller (agent) is designed, including the agent's input state, output action, and multi-objective dynamic adaptive reward function. The input state is the frequency deviation Δf between the energy storage system's real-time SOC and the power system's frequency. The output action is A = [a1, a2], where a1 serves as the frequency regulation control command for the thermal power unit, and a2 serves as the charging and discharging control command for the energy storage system. The adaptive reward function is generated by external rewards. and internal rewards composition;

[0016] S5. Based on the external and internal rewards in the multi-objective dynamic adaptive reward function, design a solution algorithm for a dual-output load frequency controller (agent);

[0017] Furthermore, a general deep reinforcement learning environment for load frequency control models that can accurately simulate the characteristics of real power frequency regulation systems was constructed; the model components of the load frequency control model include thermal power units, lithium-ion battery energy storage systems, photovoltaic power generation, wind power generation, and loads.

[0018] Furthermore, the output control strategy of the energy storage system includes two modes: normal adaptive frequency regulation and abnormal energy storage self-recovery.

[0019] 1) Normal adaptive frequency regulation mode: When the system frequency deviation exceeds the frequency dead zone, it indicates that the system has a frequency regulation requirement. The charging and discharging pattern P of the energy storage should be determined based on the energy storage operation characteristics and considering the real-time SOC state of the energy storage. c and P d ,

[0020]

[0021] Among them, P r Rated charge / discharge power for energy storage; SOC l SOC h SOC0 and SOC1 represent the different zones of the energy storage state of charge, which are identified by parameters based on the operating characteristics of the energy storage system; n is an adjustable factor, determined based on the operating characteristics of the energy storage battery.

[0022] 2) Abnormal Energy Storage Self-Recovery Mode: When the system frequency deviation is in the frequency regulation dead zone, and the energy storage SOC state is poor, the primary output target of the energy storage is to restore the SOC. The controller utilizes the remaining frequency regulation reserve capacity of the conventional units, while taking into account the limitations of the grid operation state, to restore the SOC state of the energy storage. The following SOC self-recovery charging and discharging law P is designed. c1 and P d1 :

[0023]

[0024] Among them, n1 and n2 are adjustable factors that can be determined according to the operating characteristics of the energy storage battery.

[0025] Furthermore, the steps for modeling the power system load frequency control problem as a Markov decision model and combining the Markov decision model with the energy storage system output control law to build a deep reinforcement learning load frequency control framework based on the Actor-Critic architecture are as follows:

[0026] S31. A Markov decision process consists of a quadruple.<S,A,P,R> Let S represent the set of states, A represent the set of actions, P represent the state transition function, and R represent the reward function. In each interaction between the agent and the environment, the power system is considered as the environment interacting with the agent. First, the agent, based on the observed system state s... t Based on the current strategy, give action a t Then, the system executes action a. t And update the system's running status to s t+1 At the same time, it returns the reward R to the agent for the action performed. t Finally, the agent collects all interaction information, completes the iteration and update of its own policy, until the policy converges to the optimal value.

[0027] S32. In this framework, the load frequency control uses a power system model as the environment and the system frequency deviation signal and energy storage system SOC signal provided by the environment as the controller input. The agent outputs two frequency regulation actions, a1 and a2, based on the input information, which are applied to the thermal power unit and the energy storage system, respectively. The a1 signal is superimposed with the adaptive factor and applied to the thermal power unit model, while the a2 signal is applied to the energy storage system output strategy module to generate active power output from the energy storage system to balance load changes and achieve load frequency control of the power system. The energy storage output strategy module can independently control the active power output and SOC self-recovery of the energy storage system, avoiding the use of energy storage output as a constraint condition for controller design in existing methods, which greatly simplifies the controller design. In addition, the output constraints considered in existing methods are relatively simple, only including the upper and lower bounds of SOC, and do not consider the output characteristics of different energy storage SOC ranges.

[0028] Furthermore, the designed dual-output load frequency controller includes an agent input state, an agent output action, and a multi-objective dynamic adaptive reward function, wherein...

[0029] State S: In a power system, the state observed by the agent includes the real-time SOC of the energy storage system and the frequency deviation Δf;

[0030] Action A: For the intelligent agent, its task is to control the electrical power output of the energy storage system and the thermal power unit. The action set is defined as follows:

[0031] A = [a1, a2]

[0032] a1 applies to the thermal power unit, and a2 applies to the energy storage system;

[0033] Multi-objective dynamic adaptive reward function: The primary objective of the load frequency control system is to quickly suppress system frequency oscillations and ensure the safe and stable operation of the system. Secondly, it fully utilizes the rapid response characteristics of energy storage to participate in system frequency regulation, while ensuring that the SOC of the energy storage system is within a suitable range. Finally, based on rapid frequency regulation, the agent should fully explore the unknown states of the power system environment to improve the controller's environmental adaptability. Therefore, the following multi-objective dynamic adaptive reward function is designed:

[0034]

[0035] in, This represents the external reward value designed based on real-time power grid status feedback; it is a true reflection of the system frequency deviation. The internal reward value represents the evaluation of the novelty of the current state by the random network distillation algorithm, and λ1 and λ2 represent the two objectives, respectively. and The weighting coefficients are positive. Defined as:

[0036]

[0037] Where α1, α2, and α3 are the partition values ​​of frequency deviation; β1 > 0, β2 > 0, and β3 > 0 are the weights corresponding to the reward function of each control region; C is the round end flag, and γ > 0 is the weight coefficient;

[0038] To ensure the energy storage's state of charge (SOC) remains at a good level and to guarantee the effectiveness of training rounds, a round-end penalty γC is introduced into the reward function. This penalty represents a large negative reward value given to the agent when the system frequency deviation is too large or the SOC of the energy storage system is low during a training round, ending the training round and proceeding to the next round. The round-end flag is defined as follows:

[0039] C=1,|Δf|>α4 or SOC<α5

[0040] Where α4 and α5 are the values ​​of the constraints;

[0041] The weight coefficients λ1 and λ2 of the two objective functions are expressed as follows:

[0042] Average system frequency deviation Δf ave This is a direct reflection of the frequency modulation effect of the control strategy. If there is a large frequency error, then selecting appropriate control measures to reduce it becomes particularly important. On the other hand, if there is a small deviation, it indicates that the system frequency error is already within the steady-state range, so correcting this indicator is not so important, and other objectives, such as external reward objectives, can be prioritized. The weighting coefficients are determined by the following formula:

[0043]

[0044] Where gain1 represents the dynamic gain of the weights. This represents a normal weight, indicating the relative priority of the corresponding objective. gain1 is determined by the following formula:

[0045]

[0046] Where, Δf max Δf represents the maximum acceptable frequency deviation under steady-state conditions of the power system. ave h1 represents the system average frequency deviation at the end of the DDPG agent round training, and h1 represents the rate of change of gain.

[0047] The error between the predicted and target values ​​of a stochastic network distillation model is an objective evaluation of the network's prediction performance. When the error is large, control measures should be taken to improve the network's prediction performance; when the error is small, other objectives, such as internal reward objectives, should be prioritized. The weighting coefficient is defined as follows:

[0048]

[0049] Where gain2 represents the dynamic gain of the weights. This represents a normal weight, indicating the relative priority of the corresponding objective. gain2 is determined by the following formula:

[0050]

[0051] in, This represents the maximum acceptable error between the predicted and target values ​​of the random network distillation model. h1 represents the average error between the predicted and target values ​​of the stochastic network distillation model at the end of the DDPG agent round training, and h2 represents the gain change rate.

[0052] Furthermore, the internal reward of the multi-objective dynamic adaptive reward function Obtained by the random network distillation algorithm, including the following steps:

[0053] S41, Environmental state feature extraction based on random network distillation:

[0054] Assume the state of the power system at time t is s. t The state at time t+1 is s t+1 When state s t+1 The input is fed into a random network, and the prediction network provides state prediction features. The target network provides the state target features f(s) t+1The random network consists of a prediction network and a target network. Each network comprises two fully connected layers and two ReLU layers, with parameters as follows: and θ f ;

[0055] S42, Solving for internal rewards based on environmental state characteristics:

[0056] State prediction features For state s t+1 Novelty prediction value, state target feature f(s) t+1 ) represents state s t+1 The novelty target value is used to predict the output of the target network through training. The mean squared error between the predicted value and the target value represents the internal reward term.

[0057]

[0058] Furthermore, the design of the multi-objective deep reinforcement learning solution algorithm based on dynamic weight adjustment, based on the external and internal rewards in the multi-objective dynamic adaptive reward function, includes:

[0059] Initialize constant weights for the multi-objective dynamic adaptive reward function Deep reinforcement learning network parameters, random distillation network parameters θ f Select an appropriate steady-state maximum acceptable frequency deviation Δf max Maximum acceptable error of random network distillation model

[0060] The agent interacts with the environment; the action network outputs control actions a1 and a2 to the power system; and the evaluation network evaluates the reward function. and Action value q e and q i At the same time, the system provides the reward value R for the current action. e and R i ;

[0061] Based on the evaluation action value, a multi-objective weighted loss function is calculated. The evaluation network parameters are updated by minimizing the weighted loss function, the action network parameters are updated using gradient descent, and the stochastic distillation network is updated by minimizing the mean squared error between the predicted and target values. The multi-objective weighted loss function is expressed as follows:

[0062]

[0063] in denoted as the evaluation network's target evaluation value for the action and state at the next time step, and N represents the number of randomly sampled samples.

[0064] The agent repeatedly interacts with the environment until the end of the round, and calculates the average frequency deviation Δf within that round. ave The average error between the predicted and target values ​​of the random network distillation model And determine Δf ave and If the error is large within the specified interval, adjust the increase rate of dynamic gains h1 and h2, calculate and update the values ​​of weighting coefficients λ1 and λ2. If the error is small, fine-tune the weighting coefficients to obtain the best frequency modulation performance.

[0065] The agent continues its interaction with the power system environment in the next round until the multi-objective dynamic adaptive reward function R is reached. t When the cumulative reward converges to 0, it means that the agent has solved for the optimal control strategy.

[0066] This invention provides a computer device, comprising: a processor, a computer-readable storage medium, and a storage for storing a computer program, wherein, when executing the computer program, the processor is configured to implement the steps of the deep reinforcement learning load frequency control method for energy storage-assisted thermal power units; and when executing the computer program, the processor performs the following functions:

[0067] The data acquisition function is used to collect real-time operating environment information of the power system, including the active power output of each unit, system frequency deviation, and SOC of the energy storage system.

[0068] The computational function takes environmental information as input, calculates the action value function under different states, and further obtains the loss function for neural network updates;

[0069] The output function outputs control actions based on environmental input and action value function estimation.

[0070] Storage function, used to store computer programs and information about trained neural networks.

[0071] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements the steps of the method.

[0072] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0073] (1) When giving the output rules of the energy storage system, this invention considers the energy storage SOC and the self-recovery of the energy storage SOC. The energy storage output control strategy takes into account both the grid frequency regulation requirements and the energy storage frequency regulation capability.

[0074] (2) The energy storage output strategy of the present invention is independent of the load frequency controller, which simplifies the controller design and reduces the complexity of the controller strategy solution. At the same time, the design of the energy storage output strategy can take into account more complex constraints, so as to achieve precise output control of the energy storage system.

[0075] (3) This invention introduces random network distillation technology to generate new evaluation indexes for reward function design, which incentivizes agents to explore the unknown state space extensively, thereby solving for the optimal control strategy.

[0076] (4) The present invention can better coordinate the proportion of external rewards and internal rewards through the dynamic weight adaptive adjustment method, so as to achieve the best frequency regulation effect and the optimal state of charge level of the energy storage system. Attached Figure Description

[0077] Figure 1 This is an overall flowchart of the load frequency control method according to an embodiment of the present invention;

[0078] Figure 2 This is a load frequency control framework diagram according to an embodiment of the present invention;

[0079] Figure 3 This is the dynamic gain gain1 variation curve of an embodiment of the present invention;

[0080] Figure 4 Here is the dynamic gain gain2 variation curve of this embodiment of the invention:

[0081] Figure 5 This is a load change diagram according to an embodiment of the present invention;

[0082] Figure 6 This is a comparison curve of frequency deviations in embodiments of the present invention;

[0083] In this invention, curve 1 represents the method of the present invention, and curve 2 represents the PI method with a fixed ratio.

[0084] Figure 7 This is a comparison curve of the output power of thermal power units according to embodiments of the present invention;

[0085] In this invention, curve 1 represents the method of the present invention, and curve 2 represents the PI method with a fixed ratio.

[0086] Figure 8 This is a comparison curve of the output power of the energy storage system according to an embodiment of the present invention;

[0087] In this invention, curve 1 represents the method of the present invention, and curve 2 represents the PI method with a fixed ratio.

[0088] Figure 9 This is a comparison curve of the SOC of the energy storage system according to an embodiment of the present invention;

[0089] In this invention, curve 1 represents the method of the present invention, and curve 2 represents the PI method with a fixed ratio. Detailed Implementation

[0090] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0091] Example 1

[0092] like Figures 1 to 2 As shown, a deep reinforcement learning load frequency control method for energy storage-assisted thermal power units includes the following steps:

[0093] S1. Establish a deep reinforcement learning environment for the load frequency control problem of power systems to simulate the characteristics and control requirements of various components in actual power systems;

[0094] S2. For lithium-ion battery energy storage systems in power system components, design an energy storage system output control strategy that takes into account both power system frequency deviation and energy storage system state of charge (SOC);

[0095] S3. The load frequency control problem of the power system is modeled as a Markov decision model, and a deep reinforcement learning load frequency control framework based on the Actor-Critic architecture is built by combining the Markov decision model and the output control law of the energy storage system.

[0096] S4. Combining Deep Deterministic Policy Gradient (DDPG) and stochastic network distillation techniques, a dual-output load frequency controller (agent) is designed, including the agent's input state, output action, and multi-objective dynamic adaptive reward function. The input state is the frequency deviation Δf between the energy storage system's real-time SOC and the power system's frequency. The output action is A = [a1, a2], where a1 serves as the frequency regulation control command for the thermal power unit, and a2 serves as the charging and discharging control command for the energy storage system. The adaptive reward function is generated by external rewards. and internal rewards composition;

[0097] S5. Based on the external and internal rewards in the multi-objective dynamic adaptive reward function, design a solution algorithm for a dual-output load frequency controller (agent);

[0098] Furthermore, in step S1, the transfer function model of the load frequency control system is constructed using the MATLAB toolbox. In this example, thermal power units, energy storage systems, photovoltaic power generation and wind power generation systems are introduced to simulate the random uncontrollability brought about by the grid connection of new energy sources.

[0099] S11: The transfer function model for thermal power units is established as follows:

[0100]

[0101] Among them, T g T is the time constant of the governor of the thermal power unit. CH T RH F HP These are the turbine time constant, reheater time constant, and reheater gain, respectively.

[0102] S12: The transfer function model of the energy storage system is expressed as:

[0103]

[0104] Among them, T e The energy storage output response time constant;

[0105] S13: Establish the transfer function model of the photovoltaic system, the formula is as follows:

[0106]

[0107] Among them, K PVi and T PVi These represent the gain and time constant of the photovoltaic system, respectively.

[0108] S14: Establish the transfer function model of the wind power generation system, as shown in the following formula:

[0109]

[0110] Among them, K Wi and T Wi These represent the gain and time constant of the wind power generation system, respectively.

[0111] S15: Establish the transfer function model between the generator and the load, as shown in the following formula:

[0112]

[0113] Where M and D are the system's rotational inertia and damping coefficient, respectively;

[0114] Furthermore, the method for constructing the output control law of the energy storage system is as follows: First, determine the position of the system frequency deviation. When the system frequency deviation crosses the frequency dead zone, it indicates that the system has a frequency regulation requirement. At this time, the energy storage system should work in the normal adaptive frequency regulation mode. Based on the magnitude of the frequency deviation and considering the real-time SOC state of the energy storage, determine the charging and discharging law of the energy storage at this time. When the system frequency deviation is in the frequency regulation dead zone and the SOC state of the energy storage is poor, the energy storage system should work in the non-normal energy storage self-recovery working mode. The energy storage takes the recovery of SOC as the primary output target, makes full use of the remaining frequency regulation reserve capacity of conventional units, takes into account the grid operation status limitations, and restores the SOC state of the energy storage.

[0115] Furthermore, the deep reinforcement learning load frequency control framework based on the Actor-Critic architecture is as follows: Figure 2 As shown. The designed control framework adopts an artificial intelligence controller obtained through deep reinforcement learning. The control module uses a load frequency controller designed with the DDPG algorithm that integrates random network distillation technology. This controller collects the grid frequency deviation signal and the energy storage battery SOC signal in real time. According to the reward function, the controller action network is guided to output action signals a1 and a2 that act on the conventional unit and the energy storage battery, respectively, to realize the load frequency control of the grid.

[0116] Furthermore, the dual-output load frequency controller comprises three parts: agent input state, agent output action, and multi-objective dynamic adaptive reward function.

[0117] State S: In a power system, the state observed by the agent includes the real-time SOC of the energy storage system and the frequency deviation Δf;

[0118] Action A: For the intelligent agent, its task is to control the electrical power output of the energy storage system and the thermal power unit. The action set is defined as follows:

[0119] A = [a1, a2]

[0120] a1 applies to the thermal power unit, and a2 applies to the energy storage system;

[0121] Multi-objective dynamic adaptive reward function: The primary objective of the load frequency control system is to quickly suppress system frequency oscillations and ensure the safe and stable operation of the system. Secondly, it fully utilizes the rapid response characteristics of energy storage to participate in system frequency regulation, while ensuring that the SOC of the energy storage system is within a suitable range. Finally, based on rapid frequency regulation, the agent should fully explore the unknown states of the power system environment to improve the controller's environmental adaptability. Therefore, the following multi-objective dynamic adaptive reward function is designed:

[0122]

[0123] in, This represents the external reward value designed based on real-time power grid status feedback; it is a true reflection of the system frequency deviation. The internal reward value represents the evaluation of the novelty of the current state by the random network distillation algorithm, and λ1 and λ2 represent the two objectives, respectively. and The weighting coefficients are positive. Defined as:

[0124]

[0125] Where α1, α2, and α3 are the partition values ​​of frequency deviation; β1 > 0, β2 > 0, and β3 > 0 are the weights corresponding to the reward function of each control region; C is the round end flag, and γ > 0 is the weight coefficient;

[0126] To ensure the energy storage's state of charge (SOC) remains at a good level and to guarantee the effectiveness of training rounds, a round-end penalty γC is introduced into the reward function. This penalty represents a large negative reward value given to the agent when the system frequency deviation is too large or the SOC of the energy storage system is low during a training round, ending the training round and proceeding to the next round. The round-end flag is defined as follows:

[0127] C=1,|Δf|>α4 or SOC<α5

[0128] Where α4 and α5 are the values ​​of the constraints;

[0129] The weight coefficients λ1 and λ2 of the two objective functions are expressed as follows:

[0130] Average system frequency deviation Δf ave This is a direct reflection of the frequency modulation effect of the control strategy. If there is a large frequency error, then selecting appropriate control measures to reduce it becomes particularly important. On the other hand, if there is a small deviation, it indicates that the system frequency error is already within the steady-state range, so correcting this indicator is not so important, and other objectives, such as external reward objectives, can be prioritized. The weighting coefficients are determined by the following formula:

[0131]

[0132] Where gain1 represents the dynamic gain of the weights. This represents a normal weight, indicating the relative priority of the corresponding objective. gain1 is determined by the following formula:

[0133]

[0134] Where, Δf maxΔf represents the maximum acceptable frequency deviation under steady-state conditions of the power system. ave h1 represents the system average frequency deviation at the end of the DDPG agent round training, and h1 represents the rate of change of gain; the curve of gain1 is as follows: Figure 3 As shown.

[0135] The error between the predicted and target values ​​of a stochastic network distillation model is an objective evaluation of the network's prediction performance. When the error is large, control measures should be taken to improve the network's prediction performance; when the error is small, other objectives, such as internal reward objectives, should be prioritized. The weighting coefficient is defined as follows:

[0136]

[0137] Where gain2 represents the dynamic gain of the weights. This represents a normal weight, indicating the relative priority of the corresponding objective. gain2 is determined by the following formula:

[0138]

[0139] in, This represents the maximum acceptable error between the predicted and target values ​​of the random network distillation model. h1 represents the average error between the predicted and target values ​​of the stochastic network distillation model at the end of the DDPG agent training rounds, and h2 represents the rate of change of gain. The gain2 curve is shown below. Figure 4 As shown.

[0140] Furthermore, the internal reward of the multi-objective dynamic adaptive reward function Obtained by the random network distillation algorithm, including the following steps:

[0141] S41, Environmental state feature extraction based on random network distillation:

[0142] Assume the state of the power system at time t is s. t The state at time t+1 is s t+1 When state s t+1 The input is fed into a random network, and the prediction network provides state prediction features. The target network provides the state target features f(s) t+1 The random network consists of a prediction network and a target network. Each network comprises two fully connected layers and two ReLU layers, with parameters as follows: and θ f ;

[0143] S42, Solving for internal rewards based on environmental state characteristics:

[0144] State prediction features For state s t+1 Novelty prediction value, state target feature f(s) t+1 ) represents state s t+1 The novelty target value is used to predict the output of the target network through training. The mean squared error between the predicted value and the target value represents the internal reward term.

[0145]

[0146] Furthermore, the design of a multi-objective deep reinforcement learning solution algorithm based on dynamic weight adjustment, based on the external and internal rewards in the multi-objective dynamic adaptive reward function, includes:

[0147] Initialize constant weights for the multi-objective dynamic adaptive reward function Deep reinforcement learning network parameters, random distillation network parameters θ f Select an appropriate steady-state maximum acceptable frequency deviation Δf max Maximum acceptable error of random network distillation model

[0148] The agent interacts with the environment; the action network outputs control actions a1 and a2 to the power system; and the evaluation network evaluates the reward function. and Action value q e and q i At the same time, the system provides the reward value R for the current action. e and R i ;

[0149] Based on the evaluation action value, a multi-objective weighted loss function is calculated. The evaluation network parameters are updated by minimizing the weighted loss function, the action network parameters are updated using gradient descent, and the stochastic distillation network is updated by minimizing the mean squared error between the predicted and target values. The multi-objective weighted loss function is expressed as follows:

[0150]

[0151] in denoted as the evaluation network's target evaluation value for the action and state at the next time step, and N represents the number of randomly sampled samples.

[0152] The agent repeatedly interacts with the environment until the end of the round, and calculates the average frequency deviation Δf within that round. ave The average error between the predicted and target values ​​of the random network distillation model And determine Δf ave and If the error is large within the specified interval, adjust the increase rate of dynamic gains h1 and h2, calculate and update the values ​​of weighting coefficients λ1 and λ2. If the error is small, fine-tune the weighting coefficients to obtain the best frequency modulation performance.

[0153] The agent continues its interaction with the power system environment in the next round until the multi-objective dynamic adaptive reward function R is reached. t When the cumulative reward converges to 0, it means that the agent has solved the optimal control policy.

[0154] To verify the beneficial effects of the proposed deep reinforcement learning load frequency control method for energy storage-assisted thermal power units, the following simulation experiments were conducted. The reinforcement learning training environment for this invention was built using Matlab / Simulink. The rated power of the thermal power unit was 1000MW, the frequency regulation reserve capacity was -50 to 50MW, the ramp rate was 30MW / min, and the rated power and capacity of the energy storage were 12.5MW / 2.5. Perimeter normalization was performed based on a rated frequency of 50Hz and the maximum rated capacity of the unit. The system simulation parameters are shown in Table 1.

[0155] Table 1 Simulation Parameters

[0156] parameter value parameter value parameter value parameter value <![CDATA[T g / s]]> 0.08 <![CDATA[K W1 ]]> 1 <![CDATA[SOC1]]> 0.20 <![CDATA[α5]]> 0.2 <![CDATA[T RH / s]]> 10 <![CDATA[K W2 ]]> 1.4 <![CDATA[SOC0]]> 0.45 <![CDATA[β1]]> 1 <![CDATA[T CH / s]]> 0.3 <![CDATA[T W1 / s]]> 1.25 <![CDATA[SOC1]]> 0.55 <![CDATA[β2]]> 5 <![CDATA[F HP / s]]> 0.5 <![CDATA[T W2 / s]]> 0.041 <![CDATA[SOC h ]]> 0.75 <![CDATA[β3]]> 20 <![CDATA[T e / s]]> 0.04 <![CDATA[K PV1 ]]> -18 <![CDATA[α1]]> 0.003 γ -40 <![CDATA[D / (pu·Hz -1 )]]> 1 <![CDATA[K PV2 ]]> 900 <![CDATA[α2]]> 0.005 n 10 <![CDATA[M / (pu·Hz -1 )]]> 10 <![CDATA[T PV1 / s]]> 100 <![CDATA[α3]]> 0.01 <![CDATA[n1]]> 20 <![CDATA[P r / (MW·h)]]> 2.5 <![CDATA[T PV2 / s]]> 50 <![CDATA[α4]]> 0.2 <![CDATA[n2]]> 10

[0157] according to Figure 3 The load disturbance is used to simulate the load changes of a real power system, and the simulation results are as follows: Figures 5-9 As shown in the figure, the deep reinforcement learning load frequency control method for energy storage-assisted thermal power units proposed in this invention, compared with traditional fixed-proportional PI control frequency regulation, can dynamically adjust the active power output of the energy storage system and the thermal power unit, achieving safe and stable operation of the power system with only a smaller energy storage output. Regarding energy storage SOC maintenance, this invention has better SOC maintenance performance, better reserves capacity for subsequent frequency regulation tasks, and extends the frequency regulation time of the energy storage system.

[0158] In summary, as demonstrated by the simulation results, the proposed deep reinforcement learning load frequency control method for energy storage-assisted thermal power units can dynamically adjust the outputs of the energy storage system and the thermal power unit, rapidly suppressing power system frequency oscillations and ensuring the safe and stable operation of the system. Furthermore, the dynamic weight adjustment algorithm achieves optimal coordination among multiple objectives. This verifies the effectiveness of the proposed deep reinforcement learning load frequency control method for energy storage-assisted thermal power units.

[0159] This invention provides a computer device, comprising: a processor, a computer-readable storage medium, and a storage for storing a computer program, wherein, when executing the computer program, the processor is configured to implement the steps of the deep reinforcement learning load frequency control method for energy storage-assisted thermal power units; and when executing the computer program, the processor performs the following functions:

[0160] The data acquisition function is used to collect real-time operating environment information of the power system, including the active power output of each unit, system frequency deviation, and SOC of the energy storage system.

[0161] The computational function takes environmental information as input, calculates the action value function under different states, and further obtains the loss function for neural network updates;

[0162] The output function outputs control actions based on environmental input and action value function estimation.

[0163] Storage function, used to store computer programs and information about trained neural networks.

[0164] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements the steps of the method.

[0165] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A deep reinforcement learning load frequency control method for energy storage-assisted thermal power units, characterized in that, Includes the following steps: S1. Establish a deep reinforcement learning environment for the load frequency control problem of power systems to simulate the characteristics and control requirements of various components in actual power systems; S2. For lithium-ion battery energy storage systems in power system components, design an energy storage system output control strategy that takes into account both power system frequency deviation and energy storage system state of charge (SOC); S3. Model the power system load frequency control problem as a Markov decision model, and combine the Markov decision model with the output control law of the energy storage system to build an Actor-based system. - A deep reinforcement learning load frequency control framework based on the Critic architecture. S4. Combining Deep Deterministic Policy Gradient (DDPG) and stochastic network distillation techniques, a dual-output load frequency controller is designed, including agent input state, agent output action, and multi-objective dynamic adaptive reward function. The input state is the frequency deviation Δf between the energy storage system's real-time SOC and the power system frequency. The output action is A = [a1, a2], where a1 serves as the frequency regulation control command for the thermal power unit, and a2 serves as the charging and discharging control command for the energy storage system. The adaptive reward function is generated by external rewards. and internal rewards Composition; The multi-objective dynamic adaptive reward function is designed as follows: in, This represents the external reward value designed based on real-time power grid status feedback; it is a true reflection of the system frequency deviation. The internal reward value represents the evaluation of the novelty of the current state by the random network distillation algorithm, and λ1 and λ2 represent the two objectives, respectively. and The weighting coefficients are positive. Defined as: Where α1, α2, and α3 are the partition values ​​of the frequency deviation; β1 > 0, β2 > 0, and β3 > 0 are the weights corresponding to the reward function of each control region; C is the round end flag, and γ > 0 is the weight coefficient; the round end flag is defined as: C = 1, |Δf| > α4 or SOC < α5 Where α4 and α5 are the values ​​of the constraints; S5. Based on the external and internal rewards in the multi-objective dynamic adaptive reward function, design a solution algorithm for a dual-output load frequency controller.

2. The method for deep reinforcement learning load frequency control of energy storage-assisted thermal power units according to claim 1, characterized in that, In step S1, a deep reinforcement learning environment is constructed for a general load frequency control model that can accurately simulate the characteristics of a real power frequency regulation system. The model components of the load frequency control model include thermal power units, lithium-ion battery energy storage systems, photovoltaic power generation, wind power generation, and loads.

3. The deep reinforcement learning load frequency control method for energy storage-assisted thermal power units according to claim 1, characterized in that, In step S2, the output control strategy of the energy storage system includes two modes: normal adaptive frequency regulation and abnormal energy storage self-recovery. 1) Normal adaptive frequency regulation mode: When the system frequency deviation exceeds the frequency dead zone, it indicates that the system has a frequency regulation requirement. Based on the energy storage operation characteristics and considering the real-time SOC state of the energy storage, the charging and discharging pattern P of the energy storage at this time is determined. c and P d : Among them, P r Rated charge / discharge power for energy storage; SOC l SOC h SOC0 and SOC1 represent the different zones of the energy storage state of charge, which are identified by parameters based on the operating characteristics of the energy storage system; n is an adjustable factor, determined based on the operating characteristics of the energy storage battery. 2) Abnormal Energy Storage Self-Recovery Mode: When the system frequency deviation is in the frequency regulation dead zone, and the energy storage SOC state is poor, the primary output target of the energy storage is to restore the SOC. The controller utilizes the remaining frequency regulation reserve capacity of the conventional units, while taking into account the limitations of the grid operation state, to restore the SOC state of the energy storage. The following SOC self-recovery charging and discharging law P is designed. c1 and P d1 : Among them, n1 and n2 are adjustable factors that can be determined according to the operating characteristics of the energy storage battery.

4. The deep reinforcement learning load frequency control method for energy storage-assisted thermal power units according to claim 1, characterized in that, In step S3, the power system load frequency control problem is modeled as a Markov decision model. Combining the Markov decision model with the energy storage system output control law, an Actor-based model is constructed. - The steps of the Critic architecture's deep reinforcement learning load frequency control framework are as follows: S31. A Markov decision process consists of a quadruple.<S,A,P,R> Let S represent the set of states, A represent the set of actions, P represent the state transition function, and R represent the reward function. In each interaction between the agent and the environment, the power system is considered as the environment interacting with the agent. First, the agent, based on the observed system state s... t Based on the current strategy, give action a t Then, the system executes action a. t And update the system's running status to s t+1 At the same time, it returns the reward R to the agent for the action performed. t Finally, the agent collects all interaction information, completes the iteration and update of its own policy, until the policy converges to the optimal value. S32. The load frequency control framework uses the power system model as the environment and the system frequency deviation signal and energy storage system SOC signal given by the environment as the controller input. The intelligent agent outputs two frequency regulation actions a1 and a2 according to the input information, which are applied to the thermal power unit and the energy storage system respectively. The a1 signal is superimposed with the adaptive factor and applied to the thermal power unit model. The a2 signal is applied to the energy storage system output strategy module to generate energy storage active power output to balance load changes and realize the load frequency control of the power system.

5. The method for deep reinforcement learning load frequency control of energy storage-assisted thermal power units according to claim 1, characterized in that, In step S4, a dual-output load frequency controller is designed, including agent input state, agent output action, and multi-objective dynamic adaptive reward function, wherein... State S: In a power system, the state observed by the agent includes the real-time SOC of the energy storage system and the frequency deviation Δf; Action A: For the intelligent agent, its task is to control the electrical power output of the energy storage system and the thermal power unit. The action set is defined as follows: A = [a1, a2] a1 applies to the thermal power unit, and a2 applies to the energy storage system; The weighting coefficients λ1 and λ2 are designed as follows: Where gain1 represents the dynamic gain of the weights. Let gain1 be a normal weight, representing the relative priority of the corresponding objective; gain1 is determined by the following formula: Where, Δf max Δf represents the maximum acceptable frequency deviation under steady-state conditions of the power system. ave h1 represents the system average frequency deviation at the end of the DDPG agent round training, and h1 represents the rate of change of gain. Where gain2 represents the dynamic gain of the weights. Let represent a normal weight, which indicates the relative priority of the corresponding objective; gain2 is determined by the following formula: in, This represents the maximum acceptable error between the predicted and target values ​​of the random network distillation model. h1 represents the average error between the predicted and target values ​​of the stochastic network distillation model at the end of the DDPG agent round training, and h2 represents the gain change rate.

6. The deep reinforcement learning load frequency control method for energy storage-assisted thermal power units according to claim 1, characterized in that, In step S4, the internal reward of the multi-objective dynamic adaptive reward function Obtained by the random network distillation algorithm, including the following steps: S41. Environmental state feature extraction based on random network distillation: Assume the state of the power system at time t is s. t The state at time t+1 is s t+1 When state s t+1 The input is fed into a random network, and the prediction network provides state prediction features. The target network provides the state target features f(s) t+1 The random network consists of a prediction network and a target network. Each network comprises two fully connected layers and two ReLU layers, with parameters as follows: and θ f ; S42. Solving for internal rewards based on environmental state characteristics: State prediction features For state s t+1 Novelty prediction value, state target feature f(s) t+1 ) represents state s t+1 The novelty target value is used to predict the output of the target network through training. The mean squared error between the predicted value and the target value represents the internal reward term.

7. The method for deep reinforcement learning load frequency control of energy storage-assisted thermal power units according to claim 1, characterized in that, In step S5, based on the external and internal rewards in the multi-objective dynamic adaptive reward function, a multi-objective deep reinforcement learning solution algorithm based on dynamic weight adjustment is designed, including: Initialize constant weights for the multi-objective dynamic adaptive reward function Deep reinforcement learning network parameters, random distillation network parameters θ f Select an appropriate steady-state maximum acceptable frequency deviation Δf max Maximum acceptable error of random network distillation model The agent interacts with the environment; the action network outputs control actions a1 and a2 to the power system; and the evaluation network evaluates the reward function. and Action value q e and q i At the same time, the system provides the reward value R for the current action. e and R i ; Based on the evaluation action value, a multi-objective weighted loss function is calculated. The evaluation network parameters are updated by minimizing the weighted loss function, the action network parameters are updated using gradient descent, and the stochastic distillation network is updated by minimizing the mean squared error between the predicted and target values. The multi-objective weighted loss function is expressed as follows: in, denoted as the evaluation network's target evaluation value for the action and state at the next time step, where N represents the number of randomly sampled samples; The agent repeatedly interacts with the environment until the end of the round, and calculates the average frequency deviation Δf within that round. ave The average error between the predicted and target values ​​of the random network distillation model And determine Δf ave and If the error is large within the specified interval, adjust the increase rate of dynamic gains h1 and h2, calculate and update the values ​​of weighting coefficients λ1 and λ2; if the error is small, fine-tune the weighting coefficients to obtain the best frequency modulation performance. The agent continues its interaction with the power system environment in the next round until the multi-objective dynamic adaptive reward function R is reached. t When the cumulative reward converges to 0, it means that the agent has solved the optimal control policy.

8. A computer device, characterized in that, include: A processor, a computer-readable storage medium, and a means for storing a computer program, wherein the processor, when executing the computer program, is configured to implement the steps of the deep reinforcement learning load frequency control method for energy storage-assisted thermal power units according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-7.