AGC collaborative optimization method based on time sequence deep reinforcement learning PID control
By combining LSTM and TD3 neural network to improve the PID controller and optimize the AGC control parameters, the problem of poor frequency adjustment effect in high proportion renewable energy access grid is solved, and more efficient frequency stability and adaptability are achieved.
Patent Information
- Application Number
- CN202510446854.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional PID control algorithms are difficult to effectively deal with the frequency fluctuations and uncertainties of the power system after a high proportion of renewable energy is connected to the power grid, resulting in poor frequency adjustment effect.
Combined with the timing deep reinforcement learning algorithm, the PID controller is improved through the design of LSTM network and TD3 neural network, optimize the AGC control parameters of the multi-source power system, and enhance the robustness and adaptability of frequency adjustment.
The frequency stability and regulation capabilities of multi-source interconnected power systems are improved, and the random fluctuations of renewable energy can be better cope with, and the adaptability bottleneck of traditional control algorithms can be reduced.
Smart Images

Figure CN120300832A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power systems, and specifically relates to a multi-source collaborative optimization application scenario of PID control improved based on a time-series deep reinforcement learning algorithm in an AGC scenario. Background Technique
[0002] In the context of the increasingly prominent global energy crisis and environmental problems, renewable clean energies such as wind power, photovoltaic power, and hydropower have become the focus of attention in the international energy field. However, new energy has characteristics of weak inertia and strong random fluctuations, and its large-scale access will inevitably affect the real-time balance of power generation and load in the power system and threaten the system frequency quality. Automatic generation control is an important technical means for the power grid to maintain frequency quality, and it keeps the power system frequency and the tie-line power deviation between regions within a predetermined value by balancing the power generation and load consumption. With the large-scale access of renewable energies such as wind and light to the power grid, traditional load frequency control methods are difficult to guarantee the control effect in the face of renewable energies with strong randomness and volatility.
[0003] To ensure the frequency quality of a high-proportion new energy power system, the participation of new energy in frequency regulation has attracted much attention. The application of smart grid technology makes it possible to monitor and control renewable energies in real time. To fully exploit the frequency regulation potential of renewable resources, it is of great engineering significance to establish a frequency response model of an interconnected power system including hydropower, thermal power, wind power, and energy storage. Based on the dynamic correction mechanism of area control error, PID control realizes the automatic generation control (AGC) of the power grid by adjusting the active power output of the generator set in real time. This classic control strategy has been widely deployed in power system dispatching due to its engineering practicability.
[0004] With the development of intelligent control theory, the current research focus is on the collaborative innovation of modern optimization control technology and traditional PID architecture. Since a high proportion of power electronic devices and a high proportion of renewable energies introduce additional complexity and nonlinearity to the system. For the uncertainties and changes in a multi-source interconnected power system, traditional control algorithms are often powerless. To break through the adaptability bottleneck of traditional algorithms, machine learning frameworks have achieved breakthrough applications in the field of complex power grid optimization control in recent years. In this context, deep reinforcement learning has gradually developed into a frontier technical path for coping with multiple uncertainties in the power system due to its autonomous decision-making ability and dynamic environment adaptation mechanism.
[0005] Therefore, combining with the advanced deep reinforcement learning algorithm and optimizing the AGC control instructions are crucial for the frequency stability of multi-source interconnected power systems. By using the modeling ability of the temporal neural network to capture the temporal correlation of action decisions, enhancing the state representation and policy generalization ability of the reinforcement learning agent, and adaptively optimizing the AGC parameters by making full use of the historical information of the power system, the robustness and adaptability of frequency regulation can be improved. This has important practical significance for coping with the frequency regulation pressure faced by multi-source interconnected power systems considering thermal, hydro, wind, solar, and energy storage, and it is also a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] The object of the present invention is to propose an AGC collaborative optimization method based on temporal deep reinforcement learning PID control, which combines the temporal modeling ability of the Long Short-Term Memory (LSTM) network with the neural network design of the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to design a PID controller based on the L-TD3 algorithm. The AGC control parameters of the multi-source power system are optimized in real time to achieve the coordinated operation of various energy sources, thereby ensuring the frequency stability of the power system.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] An AGC collaborative optimization method based on temporal deep reinforcement learning PID control, comprising the following steps:
[0009] (1) Establish a frequency response control model for a multi-source interconnected power system;
[0010] a. The frequency of each regional system is modeled as:
[0011]
[0012] In the formula, M i is the inertia coefficient of the power system in region i; D i is the load damping coefficient of the system; ΔP Gi is the sum of the outputs of all power sources in the system; ΔP Di is the sum of the fluctuating loads in the system.
[0013] b. The power exchange of each regional tie line is modeled according to the actual situation. Control area 1 exchanges power with control area 2 and control area 3 respectively:
[0014]
[0015] In the formula: T ijIt is the synchronization coefficient of the tie line between regions i and j.
[0016] c. The power generation resources of each regional power system are modeled according to the regional resource characteristics, and renewable energy is incorporated into the secondary frequency regulation loop for AGC real-time scheduling.
[0017] The expression of the frequency response model of thermal power units:
[0018]
[0019] In the formula: T g is the governor time constant; T t is the steam turbine time constant; T r is the reheat time constant; K r is the reheat coefficient.
[0020] The equivalent model of the frequency response of hydropower units:
[0021]
[0022] In the formula, T w is the inertia time constant of the water turbine; T R , T I are the time constants of the water turbine regulating valve respectively.
[0023] The equivalent model of the wind farm:
[0024]
[0025] In the formula: K wind is the frequency modulation coefficient for wind power to participate in frequency modulation; T wind is the pitch response time constant.
[0026] The equivalent model of the photovoltaic power station:
[0027]
[0028] In the formula: K pv is the frequency modulation coefficient for the photovoltaic power station to participate in frequency modulation; T pv is the response time constant of the photovoltaic power station.
[0029] The equivalent model of the energy storage power station:
[0030]
[0031] In the formula: T bess is the time constant of the battery energy storage link; K bess is the energy storage gain coefficient.
[0032] (2) Design of the AGC control architecture based on deep reinforcement learning;
[0033] The control areas 1 and 2 adopt the TBC control mode and are responsible for regulating the tie-line deviation and frequency deviation; the control area 3 adopts the FTC control mode and is responsible for regulating the tie-line deviation. The AGC is controlled in the PID manner. The calculation formula for the total area regulation requirement (ARR) of the control areas 1, 2, and 3 at time t is as follows:
[0034]
[0035] In the formula, K Pi is the proportional gain coefficient of control area i, and P ACEi (t) is the area control deviation value of control area i at time t; K Ii is the integral gain coefficient of control area i, and K Di is the differential gain coefficient of control area i.
[0036] The AGC controller for each area is a deep reinforcement learning agent. The agent interacts with the environment to obtain a joint state set transformed from the joint action set of multiple agents and the observations of each agent. At the same time, the environment gives feedback according to this high-dimensional set, and then obtains the next joint action state. Through the continuous interaction between multiple agents and the environment, the cycle of action-state-reward process, the optimal optimization strategy that maximizes the cumulative reward of multiple agents can be finally obtained.
[0037] a. Reinforcement learning state space: Select the environmental state variables monitored in real time by the regional power grid, such as area control deviation, area frequency deviation, area frequency change rate, area tie-line power, etc. as the observable states of the agent. The agent masters the global information and conducts learning optimization.
[0038] b. Reinforcement learning action space: The agents in each area act independently, and the PID controller parameters are defined as control actions. The constraint conditions are:
[0039]
[0040] In the formula, K Pi , K Ii and K Di are the proportional, integral, and differential control parameters of the regional PID control respectively, and K Pmax , K Imax and K Dmax are the upper limits of the adjustment of the proportional, integral, and differential control parameters of the regional PID control respectively.
[0041] c. Design of Reinforcement Learning Reward Function: The design of the agent's reward function incorporates the dynamic characteristics of each region. Taking the absolute value of the area control error |ACE| as the objective function can seek the optimal solution of the control strategy that maximizes the long-term benefits of the CPS and suppresses large fluctuations in power. The magnitude of the absolute value of the frequency deviation |Δf| can also directly reflect the quality of the system control performance. Therefore, a reinforcement learning reward function is designed using the linear weighting of the regional frequency deviation and the area control error in a multi-source interconnected power system:
[0042] R i = -ω f |Δf(i)| 2 - ω a |ACE(i)| 2 (10)
[0043] where ω f and ω a are the weight coefficients of |Δf(i)| and |ACE(i)| respectively.
[0044] (3) Design of the Neural Network of the Deep Reinforcement Learning Algorithm with Stacked LSTM.
[0045] Introduce LSTM into the actor network and critic network of TD3 to enhance the system state modeling ability, enabling the control strategy to better adapt to the dynamic environment.
[0046] LSTM enables the actor network to generate a smoother AGC adjustment strategy based on historical information, improving frequency stability. Output the hidden state to capture the dynamic time series dependence of the system.
[0047] h t = LSTM(s t , h t-1 ) (11)
[0048] where h t and s t are the system information and the set of environmental states at the current time t, and h t-1 is the system historical information at the past time. The agent selects the action a t based on the policy π combined with the historical information:
[0049] a t = π θ (s t , h t-1 ) (12)
[0050] The critic network also considers the time correlation for Q-value update:
[0051] h t = LSTM([st , a t , h t-1 ) (13)
[0052]
[0053] where r t is the environmental reward and γ is the discount factor.
[0054] The AGC collaborative optimization method based on temporal deep reinforcement learning PID control provided by the present invention provides a design solution for the AGC collaborative optimization of a multi-source interconnected power system based on a temporal neural network and a TD3 deep reinforcement learning algorithm. The frequency modulation loop incorporating renewable energy can provide a more flexible response and enhance the resilience of the power system. The LSTM further enhances the TD3 deep reinforcement learning algorithm's ability to model temporal states, enabling the control strategy to fully utilize historical information and improve the robustness and adaptability of frequency regulation. This method can adaptively optimize AGC parameters, improve the regulation ability of interconnected power grids, and provide an efficient reinforcement learning solution for smart grid frequency control. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0056] Figure 1 Frequency response model of a multi-source interconnected power system considering thermal, hydro, wind, and photovoltaic power generation resources provided for the embodiment;
[0057] Figure 2 L-TD3 deep reinforcement learning architecture for a multi-source interconnected power system provided for the embodiment;
[0058] Figure 3 Actor-critic neural network structure combined with an LSTM network provided for the embodiment;
[0059] Figure 4 Comparison of reward training curves based on different deep reinforcement learning algorithms;
[0060] Figure 5 Comparison of three-area frequency deviation curves based on different deep reinforcement learning algorithms;
[0061] Figure 6 Comparison of three-area frequency change rate curves based on different deep reinforcement learning algorithms;
[0062] Figure 7 Comparison of three - area area control error curves based on different deep reinforcement learning algorithms Specific implementation mode
[0063] The embodiments of the present invention are implemented on the premise of the technical solution, and detailed implementation modes are given, but the protection scope of the present invention is not limited to the following embodiments
[0064] The comprehensive inertia control parameter design method for the temperature - controlled load frequency response proposed by the present invention includes the following steps
[0065] 1. As shown in Figure 1 , taking the AGC of a certain interconnected power grid as the research background, the power grid is divided into 3 control areas, and each area has different resource characteristics and control modes. Control areas 1 and 2 adopt the TBC control mode, which is responsible for regulating the tie - line deviation and frequency deviation; control area 3 adopts the FTC control mode, which is responsible for regulating the tie - line deviation. A frequency response control model of a multi - source interconnected power system is established
[0066] (1) The frequency of each regional system is modeled as
[0067]
[0068] where M i is the inertia coefficient of the power system in region i; D i is the load damping coefficient of the system; ΔP Gi is the sum of the outputs of all power sources in the system; ΔP Di is the sum of the fluctuating loads in the system
[0069] (2) The power exchange of each regional tie - line is modeled according to the actual situation. Control area 1 exchanges power with control area 2 and control area 3 respectively
[0070]
[0071] where: T ij is the tie - line synchronization coefficient between regions i and j
[0072] (3) The power generation resources of each regional power system are modeled according to the regional resource characteristics, and renewable energy is incorporated into the secondary frequency modulation loop for AGC real - time scheduling. The frequency response model expression of thermal power units
[0073]
[0074] where: T g is the governor time constant; T t is the steam turbine time constant; T r is the reheat time constant; K ris the reheat coefficient.
[0075] Equivalent model of the frequency response of hydropower units:
[0076]
[0077] In the formula, T w is the inertia time constant of the water turbine; T R , T I are the time constants of the water turbine regulating valve respectively. Figure 1 In R r and R h are the primary frequency regulation droop control coefficients of thermal power units and hydropower units respectively.
[0078] Equivalent model of wind farm:
[0079]
[0080] In the formula: K wind is the frequency modulation coefficient for wind power to participate in frequency modulation; T wind is the pitch response time constant.
[0081] Equivalent model of photovoltaic power station:
[0082]
[0083] In the formula: K pv is the frequency modulation coefficient for the photovoltaic power station to participate in frequency modulation; T pv is the response time constant of the photovoltaic power station.
[0084] Equivalent model of energy storage power station:
[0085]
[0086] In the formula: T bess is the time constant of the battery energy storage link; K bess is the energy storage gain coefficient.
[0087] Figure 1 In r hi , r ri , r ni are the power response proportionality coefficients of hydropower, thermal power and renewable energy (wind power, photovoltaic power and energy storage) respectively. By reasonably configuring renewable energy to participate in frequency modulation, the frequent start and stop of traditional generating units can be reduced, and the operating cost can be reduced. In addition, in the event of an emergency (such as a sudden increase in load or a power generation failure), the frequency modulation loop incorporating renewable energy can provide a more flexible response and enhance the resilience of the power system.
[0088] 2. Design of AGC control architecture based on deep reinforcement learning;
[0089] The AGC controller for each area is a deep reinforcement learning agent. The agent interacts with the environment to obtain a set of joint states transformed from the joint action set of multiple agents and the observations of each agent. At the same time, the environment gives feedback based on this high-dimensional set, and then obtains the next joint action state. This process can be modeled as a Markov process (S, A, P, R, γ), where S is the state, A is the action, P is the state transition model, R is the reward function, and γ ∈ [0, 1] is the discount factor. Based on the policy π, during the change process of the joint action state, all agents will obtain their respective rewards at the corresponding moments and calculate the cumulative return. Through the continuous interaction between multiple agents and the environment, cycling through the action-state-reward process, the optimal optimization strategy that maximizes the cumulative return of multiple agents can be finally obtained.
[0090] (1) Reinforcement learning state space: The AGC system comprehensively perceives the grid operation information, selects the environmental state variables of real-time monitoring of the regional grid, such as area control error, area frequency deviation, area frequency change rate, area tie-line power, etc. as the observable states of the agent. The agent masters the global information and conducts learning optimization.
[0091] (2) Reinforcement learning action space: The agents in each area act independently, and define the PID controller parameters as the control actions. The constraint conditions are:
[0092]
[0093] In the formula, K Pi 、K Ii and K Di are the proportional, integral, and differential control parameters of the regional PID control respectively, and K Pmax 、K Imax and K Dmax are the upper limits of the adjustment of the proportional, integral, and differential control parameters of the regional PID control respectively.
[0094] (3) Design of the reinforcement learning reward function: The design of the agent's reward function includes the dynamic characteristics of each area. Taking the absolute value of the area control error |ACE| as the objective function can seek the optimal solution of the control strategy that satisfies the long-term maximization of CPS benefits and suppresses large power fluctuations. The magnitude of the absolute value of the frequency deviation |Δf| can also directly reflect the quality of the system control performance. Therefore, a reinforcement learning reward function is designed using the linear weighting of the regional frequency deviation and the regional control error of the multi-source interconnected power system:
[0095] R i =-ω f |Δf(i)| 2 -ω a |ACE(i)| 2 (23)
[0096] Where ω f and ω a are the weight coefficients of |Δf(i)| and |ACE(i)| respectively.
[0097] 3. Neural network design of the L-TD3 deep reinforcement learning algorithm with stacked LSTM.
[0098] (1) TD3 algorithm design:
[0099] The TD3 algorithm is improved based on DDPG. By introducing a double Q-network to reduce overestimation bias, delaying policy updates to improve stability, and introducing target policy noise to enhance exploration ability, the robustness and performance of training are improved. The interaction process between the L-TD3 algorithm and the environment is as Figure 2 shown.
[0100] Optimizing the AGC controller parameters requires finding the optimal solution in a complex non-linear dynamic system. TD3 alleviates the problem of Q-value overestimation through a double Q-network, making the optimization result more reliable, avoiding overly aggressive control strategies, and improving frequency stability. TD3 trains two independent Q-networks Q θ1 (s,a) and Q θ2 (s,a), and takes the minimum value to evaluate the actions generated by the actor network.
[0101]
[0102] Where s t and a t are the environmental state and action set at time t respectively, r(s t ,a t ) is the environmental reward, θ' i is the parameter of the target Q-network, π θ' is the target policy network, and γ is the discount factor.
[0103] The loss function of the critic network is:
[0104]
[0105] The target policy smoothing mechanism of TD3 adds Gaussian noise to the target action, avoiding over-reliance on local optima of the policy, thereby reducing drastic changes in control parameters, improving the smoothness of control signals, and ensuring the executability of regulation.
[0106]
[0107] And truncate the noise to prevent invalid actions due to excessive noise.
[0108] ε = clip(ε, -c, c) (27)
[0109] TD3 updates through a delayed strategy, reducing the dependence of the policy network on unstable Q-values, making the training more stable and helping to find the global optimal control parameters. The policy network is updated once every d steps of the Q-network update. The parameters of the target network are softly updated.
[0110]
[0111] In the formula, τ is the soft update coefficient, which is much less than 1.
[0112] (2) Neural network design:
[0113] Traditional TD3 only makes decisions based on the current state s t and does not fully utilize historical information, resulting in poor generalization ability of the policy. Therefore, this method introduces LSTM into the actor network and critic network of TD3 to improve the system state modeling ability, enabling the control policy to better adapt to the dynamic environment. The improved actor-critic network structure in this paper is as Figure 3 shown.
[0114] LSTM enables the actor network to generate a smoother AGC adjustment strategy based on historical information, improving frequency stability. The output hidden state captures the dynamic time series dependencies of the system.
[0115] h t = LSTM(s t , h t-1 ) (29)
[0116] In the formula, h t and s t are the system information and environmental state set at the current time t, and h t-1 is the system historical information at past times. The agent selects the action a t based on the policy π combined with historical information:
[0117] a t = π θ (s t , h t-1 ) (30)
[0118] The critic network also considers the time correlation for Q-value update:
[0119] h t = LSTM([s t , a t , h t-1 ) (31)
[0120]
[0121] where r t is the environmental reward, and γ is the discount factor.
[0122] To verify the AGC collaborative optimization dynamic performance of the PID control strategy based on the L-TD3 algorithm in a certain interconnected power grid interconnection system, the frequency response scenario of a multi-source three-region interconnected power system is taken as an example for analysis.
[0123] At t = 1s, the system control area 1 is subjected to a 0.1 p.u. step power disturbance, and the performances of three reinforcement learning algorithms, namely L-TD3, TD3, and DDPG, in frequency stability control are compared and analyzed. The experiment is carried out from two dimensions of training convergence and frequency response control index, and the significant advantages of the L-TD3 algorithm in dealing with time-varying disturbances are verified.
[0124] Analysis of the convergence of the training process:
[0125] Figure 4 (a-c) respectively show the comparison of the training reward curves of the L-TD3, TD3, and DDPG algorithms. The experimental results show that L-TD3 is significantly superior to the traditional algorithms in terms of convergence speed, reward value optimization, and policy exploration efficiency.
[0126] L-TD3 shows a lower reward variance during the training process, and the average reward value is much higher than those of the TD3 and DDPG algorithms. The training curve of L-TD3 shows that the LSTM layer in its network structure effectively suppresses the policy oscillation caused by environmental randomness by memorizing historical state-action pairs, thus realizing more stable gradient updates.
[0127] Moreover, for the exploration strategy, it is significantly superior to other algorithms. L-TD3 locates the high-reward policy area through the time-series correlation exploration mechanism in the early iteration. At the 18th iteration, a high-performance policy with a reward value of -0.0151 is explored, as Figure 4 (a) is marked by an arrow. Its reward value is increased by 77.96% and 98.99% compared with TD3 (-0.685) and DDPG (-1.477) in the same stage respectively, and the optimization efficiency is greatly improved. Due to the time-series modeling ability of LSTM, its hidden state h t dynamically screens key state features through the gating mechanism, such as the correlation between the frequency deviation change rate and the regional power imbalance, thus accelerating the directional search for high-value policies. The reward values of the TD3 algorithm and the DDPG algorithm converge to -0.111 and -0.137 at the 50th iteration.
[0128] This difference stems from the dual optimization mechanism of L-TD3: Delayed update of the policy network: By separating the update frequencies of the policy network and the value network, the policy overshoot problem of traditional DDPG-like algorithms is alleviated; Temporal credit assignment: The algorithm based on LSTM for backpropagation along time optimizes the cross-period reward assignment, enabling more accurate gradient feedback for early control actions. Therefore, the intelligent controller based on L-TD3 is more suitable for the real-time requirements of the AGC control scenario in a multi-area interconnected power system.
[0129] Comparison of frequency response characteristics:
[0130] Figures 5 - 7 The frequency response results of the interconnected power system under different algorithms are shown as follows. Figure 5 As shown, the PID control based on the L-TD3 algorithm exhibits the best dynamic characteristics in the frequency deviation control of the three areas:
[0131] In control area 1, the L-TD3 algorithm quickly suppresses the frequency deviation Δf to the lowest value of -0.0094 Hz at t = 4.0 s after the disturbance. Compared with TD3 reaching the lowest value of -0.020 Hz at t = 5.7 s and DDPG reaching the lowest value of -0.020 Hz at t = 6.0 s, the frequency modulation effect is significantly improved by 53%. In the steady state stage, the steady state value of the frequency deviation converges to near zero deviation, while the steady state deviation of TD3 is -0.0028 Hz and that of DDPG is -0.0052 Hz. The steady state errors are reduced by 89.3% and 94.2% respectively. Similarly, in control areas 2 and 3, compared with the two algorithms, the frequency modulation effects are improved by 41.5% and 54.3% respectively, and the frequency steady state accuracy is significantly improved.
[0132] Figure 6 shows the change of the frequency change rate in the three areas. Since control area 1 is the disturbance source area, the maximum value of the frequency change rate is limited by the system inertia. However, under the optimization of the L-TD3 algorithm in control area 2, the peak value of the frequency change rate is suppressed to f ratemax = -0.0098 Hz / s. Compared with f ratemax = -0.011 Hz / s of TD3 and f ratemax = -0.014 Hz / s of DDPG, the control effects are significantly improved by 10.91% and 30% respectively. The optimization effect in control area 3 is more significant, and the reduction rates of the peak values of the frequency change rate reach 19.23% and 38.24% respectively.
[0133] Figure 7 Under the optimization of the L-TD3 algorithm in, the control effects of the area control deviations in the three control areas all show superiority. The maximum value of the area control deviation in control area 1 is P ACEmax = -0.0019 p.u. Compared with P ACEmax = -0.0022 p.u. of TD3 and P ACEmax= -0.0025 p.u., and the control effects are significantly improved by 13.64% and 24% respectively. Similarly, in Control Areas 2 and 3, comparing the two algorithms, the control effects are improved by 19.12% and 20.29% respectively. It shows that the PID control based on the L-TD3 algorithm has significantly better collaborative regulation ability for cross-region power imbalance than the traditional method.
[0134] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for AGC collaborative optimization based on time-series deep reinforcement learning PID control, comprising the following steps: (1) Establish a frequency response control model for a multi-source interconnected power system; a. Model the frequency of each regional system as: Where, M i is the inertia coefficient of the power system in area i; D i is the load damping coefficient of the system; ΔP Gi is the sum of the outputs of all power sources in the system; ΔP Di is the sum of the fluctuating loads in the system. b. Model the power exchange of each regional tie line according to the actual situation. Control area 1 exchanges power with control area 2 and control area 3 respectively: Where: T ij is the tie-line synchronization coefficient between regions i and j. c. Model the power generation resources of each regional power system according to the regional resource characteristics, and incorporate renewable energy into the secondary frequency regulation loop for AGC real-time scheduling. Expression of the frequency response model of thermal power units: Where: T g is the governor time constant; T t is the steam turbine time constant; T r is the reheat time constant; K r is the reheat coefficient. Equivalent model of the frequency response of hydropower units: Where, T w is the inertia time constant of the water turbine; T R , T I are respectively the time constants of the water turbine regulating valve. Equivalent model of wind farms: Where: K wind is the frequency modulation coefficient for wind power to participate in frequency modulation; T wind is the pitch response time constant. Equivalent model of photovoltaic power stations: Where: K pv is the frequency modulation coefficient for the PV power station to participate in frequency modulation; T pv is the response time constant of the PV power station. Equivalent model of energy storage power stations: Where: T bess is the time constant of the battery energy storage link; K bess is the energy storage gain coefficient. (2) Design of an AGC control architecture based on deep reinforcement learning; Control areas 1 and 2 adopt the TBC control mode, which is responsible for regulating the tie line deviation and frequency deviation; control area 3 adopts the FTC control mode, which is responsible for regulating the tie line deviation. AGC is controlled in a PID manner. The calculation formula for the total regional regulation demand (Area Regulation Requirement, ARR) of control areas 1, 2, and 3 at time t is: where, K Pi is the proportional gain coefficient of control zone i, and P ACEi (t) is the area control deviation value of control zone i at time t; K Ii is the integral gain coefficient of control zone i, and K Di is the differential gain coefficient of control zone i. The AGC controller for each region is a deep reinforcement learning agent. The agent interacts with the environment to obtain a joint state set transformed from the joint action set of multiple agents and the observations of each agent. At the same time, the environment gives feedback based on this high-dimensional set, and then obtains the next joint action state. Through the continuous interaction between multiple agents and the environment, cycling through the action-state-reward process, the best optimization strategy that maximizes the cumulative return of multiple agents can be finally obtained. a. Reinforcement learning state space: Select environmental state variables monitored in real time by the regional power grid, such as regional control deviation, regional frequency deviation, regional frequency change rate, regional tie line power, etc. as the observable states of the agent. The agent masters the global information and conducts learning and optimization. b. Reinforcement learning action space: The agents in each region act independently, and define the PID controller parameters as control actions. The constraint conditions are: where K Pi , K Ii and K Di are respectively the proportional, integral, and derivative control parameters of the regional PID control, and K Pmax , K Imax and K Dmax are respectively the upper limits for adjusting the proportional, integral, and derivative control parameters of the regional PID control. c. Design of the reinforcement learning reward function: The design of the agent's reward function includes the dynamic characteristics of each region. Taking the absolute value of the regional control deviation |ACE| as the objective function can seek the optimal solution of the control strategy that satisfies the long-term maximum benefit of CPS and suppresses large fluctuations in power. The magnitude of the absolute value of the frequency deviation |Δf| can also directly reflect the quality of the system control performance. Therefore, use the linear weighting of the regional frequency deviation and regional control deviation of the multi-source interconnected power system to design the reinforcement learning reward function: R i = -ω f |Δf(i)| 2 -ω a |ACE(i)| 2 (10) where ω f and ω a are the weight coefficients of |Δf(i)| and |ACE(i)| respectively.
2. The method for AGC collaborative optimization based on time-series deep reinforcement learning PID control according to claim 1, wherein: The frequency response of wind power, photovoltaic and energy storage power generation resources is equivalent, a frequency response model is established and incorporated into the secondary frequency regulation loop, enhancing the resilience and frequency response flexibility of the interconnected power system. For a multi-region interconnected power system considering thermal, hydro, wind, photovoltaic and energy storage power generation resources, a PID controller based on time-series deep reinforcement learning is designed, improving the frequency stability of the power system.
3. The PID control method based on temporal deep reinforcement learning includes the following steps: (1) Design of the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm: TD3 trains two independent Q networks Q θ1 (s, a) and Q θ2 (s, a), and takes the minimum value to evaluate the actions generated by the actor network. where s t and a t are the environmental state and action set at time t respectively, r(s t , a t ) is the environmental reward, θ' i are the parameters of the target Q-network, π θ' is the target policy network, and γ is the discount factor. The loss function of the critic network is: The target policy smoothing mechanism of TD3 adds Gaussian noise to the target action to avoid over-reliance of the policy on local optima, thereby reducing drastic changes in control parameters, improving the smoothness of control signals, and ensuring the executability of regulation. And the noise is truncated to prevent invalid actions caused by excessive noise. ε = clip(ε, -c, c) (14) TD3 updates the policy with a delay, reducing the dependence of the policy network on unstable Q-values, making the training more stable and helping to find the global optimal control parameters. The policy network is updated once every d steps of Q-network update. Soft update is performed on the parameters of the target network. In the formula, τ is the soft update coefficient, which is much less than 1. (2) Neural network design of the deep reinforcement learning algorithm with a stacked Long Short-Term Memory (LSTM) network. LSTM is introduced into the actor network and critic network of TD3 to enhance the system state modeling ability, enabling the control policy to better adapt to the dynamic environment. LSTM enables the actor network to generate a smoother AGC regulation policy based on historical information, improving frequency stability. The output hidden state captures the dynamic temporal dependencies of the system. h t = LSTM(s t , h t-1 )(16) where h t and s t are the system information and environmental state set at the current time t, and h t-1 is the system historical information at past times. The agent selects an action a based on the policy π and combines historical information t : a t = π θ (s t , h t-1 )(17) The critic network also updates the Q-value considering temporal correlation: h t = LSTM([s t , a t , h t-1 ) (18) where r t is the environmental reward and γ is the discount factor.
4. The PID control method based on temporal deep reinforcement learning according to claim 3, characterized in that: TD3 improves the robustness and performance of training by introducing a dual Q-network to reduce overestimation bias, delaying policy updates to improve stability, and introducing target policy noise to enhance exploration ability; LSTM captures the temporal correlation of action decisions, thereby enhancing the state representation and policy generalization ability of the agent in partially observable scenarios; The PID controller based on temporal deep reinforcement learning can adaptively optimize AGC parameters, improve the regulation ability of the multi-source interconnected power grid, and provide an efficient reinforcement learning solution for intelligent grid frequency control.
Citation Information
Cited By
Network source load frequency control simulation modeling method
CN120784856A
A web source load frequency control simulation modeling method
CN120784856B
Novel power system primary frequency modulation optimization method based on deep learning
CN120896185A
Optimal primary frequency control method and system based on reinforcement learning
CN121566489A
Rotary power flow controller parameter adaptive adjustment method based on reinforcement learning
CN121602401A