Ultra-low frequency oscillation damping joint parameter optimization method for grid-connected hydroelectric system based on deep reinforcement learning model identification

By optimizing the parameters of the governor and additional damping controller of the hydropower system using a deep reinforcement learning model and a chaotic particle swarm optimization algorithm, the problem of poor suppression of ultra-low frequency oscillations in traditional methods is solved, and the stability and robustness of the system are improved.

CN121172718APending Publication Date: 2025-12-19SANXIA JINSHAJIANG YUNCHUAN HYDROPOWER DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511215960.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively suppress ultra-low frequency oscillations in hydropower systems. Traditional methods, such as optimizing governor parameters individually or adding additional damping controllers, suffer from interactive effects and implementation difficulties, resulting in poor damping characteristic optimization.

Method used

A deep reinforcement learning model is used to identify the hydropower system model. The governor and additional damping controller are combined as a whole, and their parameters are optimized by chaotic particle swarm optimization algorithm. A multi-loop additional damping controller is constructed to provide positive damping and improve the system damping characteristics.

Benefits of technology

The joint optimization of governor parameters and additional damping controller significantly improved the damping characteristics of the hydropower system, suppressed ultra-low frequency oscillations, and enhanced the robustness and stability of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121172718A_ABST
    Figure CN121172718A_ABST
Patent Text Reader

Abstract

The invention provides a grid-connected hydroelectric system ultralow frequency oscillation damping joint parameter optimization method based on deep reinforcement learning model identification. The method comprises the following steps: constructing a hydroelectric generating set speed regulation model; defining other models of the power system except the hydroelectric generating set in the hydroelectric system as external system models, and identifying a transfer function model of the external system models by adopting a deep reinforcement learning algorithm, so as to obtain a relationship between the guide vane opening degree of the hydroelectric generating set in the hydroelectric system and the system frequency response; based on a hydroelectric generating set speed regulation model and an external system model, an additional damping controller is used as one of control links of a speed regulator, the additional damping controller and the speed regulator are regarded as a whole, and parameters of the speed regulator and parameters of the additional damping controller are optimized based on a chaotic particle swarm algorithm. According to the method, the phase compensation of the mechanical torque of the hydroelectric generating set is realized on the whole through parameter joint optimization, and the damping characteristic of the set is improved, so that the ultra-low frequency oscillation is effectively inhibited.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of hydroelectric unit control, in particular to a grid-connected hydroelectric system super-low frequency oscillation damping joint parameter optimization method based on deep reinforcement learning model identification. BACKGROUND

[0002] With the deepening construction of the "double high" characteristic new power system, the asynchronous interconnection of the alternating current power grid through the direct current is one of the key technical paths to realize the safe and stable operation of the system. After the asynchronous interconnection of the alternating current power grid and the main grid, the power balance and stability characteristics of the power grid will change profoundly, and the main stability problem will change from the power angle stability between regions to the frequency stability problem within the region. The current frequency stability analysis mainly focuses on two problems, namely the large-scale frequency deviation problem under active power disturbance and the super-low frequency oscillation problem related to the water turbine. Actual engineering and scientific research have shown that in the power system with a high proportion of hydropower, super-low frequency oscillation phenomenon will occur. This phenomenon will cause the water turbine guide vane to adjust frequently and the excitation system to be overloaded, threatening the safety of equipment. Further, it may cause system instability, virtual synchronous machine resonance and DC transmission lockout risks. Therefore, it is urgent to have a strong power system damping characteristic optimization method to suppress the occurrence of super-low frequency oscillation.

[0003] After the asynchronous interconnection of the regional power grid with a high proportion of hydropower, the inertia of the regional power grid decreases, and the governor parameters of the water turbine units in the grid are generally set to be large, and the governor action characteristics are very sensitive. This makes the high-hydropower regional power grid operating asynchronously prone to super-low frequency oscillation phenomenon with a frequency lower than 0.1Hz when the system load disturbance or short-circuit fault occurs. When super-low frequency oscillation occurs, all the generator speeds in the grid oscillate in unison, causing the electromagnetic power and mechanical power of the units to also oscillate, which may have a serious impact on the safe and stable operation of the power grid. Therefore, with the further promotion of asynchronous interconnection projects, the research on super-low frequency oscillation problem is of great significance to maintain the safe and stable operation of the power grid.

[0004] It is generally believed that the cause of the ultra-low frequency oscillation is mainly the negative damping torque caused by the unreasonable setting of the hydroelectric generator speed regulation system and the water hammer effect of the prime mover system. To suppress ultra-low frequency oscillation, the existing technology mainly optimizes the speed regulator parameters and adds an additional damping controller. In terms of speed regulator parameter optimization, the main technology focuses on the influence of speed regulator PID parameters, water hammer effect coefficient and regulation coefficient on ultra-low frequency oscillation. However, parameter optimization for the speed regulator will greatly weaken its primary frequency modulation capability, which is not conducive to the rapid recovery of frequency after system failure. In addition, this method cannot completely suppress the ultra-low frequency oscillation phenomenon in regional power grids with a high proportion of hydropower. Therefore, many technologies also add an additional damping controller to directly inject a damping power component opposite in phase to the oscillation frequency to improve system damping. The water turbine has high inertia characteristics and dynamic nonlinear perturbation characteristics, and its multiple parameters present time-varying coupling relationship, making controller design and implementation difficult.

[0005] An accurate simulation system model is the basis for performance verification of ultra-low frequency oscillation suppression methods. However, the speed regulation system of the water turbine involves complex mechanical, fluid and electromagnetic coupling problems, resulting in complex nonlinear dynamic characteristics of the power system containing hydroelectric generating units. Traditional system identification methods such as Prony identification method and TLS-ESPRIT identification algorithm are greatly affected in high noise and multiple signal environments, making it difficult to accurately identify the power system model containing hydroelectric generating units.

[0006] In addition, the existing damping characteristic optimization methods mainly optimize the reduction of negative damping and the increase of positive damping separately, which has the following shortcomings, resulting in poor ultra-low frequency oscillation suppression effect.

[0007] (1) In terms of reducing negative damping, the main method is to design speed regulator parameter optimization, including PID parameters, water hammer effect coefficient and regulation coefficient, etc. However, parameter optimization for the speed regulator will greatly weaken its primary frequency modulation capability, which is not conducive to the rapid recovery of frequency after system failure.

[0008] (2) In terms of increasing positive damping, the existing improvement method is to add an additional damping controller. However, the water turbine has high inertia characteristics and dynamic nonlinear perturbation characteristics, and its multiple parameters present time-varying coupling relationship, making controller design and implementation difficult.

[0009] Both of the above two control schemes are designed separately, and due to the complex nonlinear dynamic coupling effect in the speed regulation system of the hydroelectric generating set, the suppression of the ultra-low frequency oscillation needs to comprehensively consider the optimization of the main control parameters of the speed regulator and the synergistic effect of the additional damping controller. The main control loop parameters of the speed regulator directly affect the inertia and steady-state regulation performance of the hydraulic turbine, and the additional damping controller suppresses the oscillation by introducing a damping signal of a specific frequency band, but there is an indiscernible interaction between the two in the frequency domain. The integral action of the main control loop may amplify the low-frequency phase lag, and if the compensation signal of the additional controller does not match the inherent dynamics of the speed regulator, the gain superposition may cause the actuator to be saturated or the phase cancellation may weaken the damping effect. SUMMARY

[0010] The present application aims to at least solve one of the above technical problems in the prior art.

[0011] To this end, the present application provides a method for jointly optimizing parameters of ultra-low frequency oscillation damping of a grid-connected hydroelectric system based on deep reinforcement learning model identification.

[0012] The present application provides a method for jointly optimizing parameters of ultra-low frequency oscillation damping of a grid-connected hydroelectric system based on deep reinforcement learning model identification, comprising: constructing a hydroelectric generating set speed regulation model, the hydroelectric generating set speed regulation model comprising a hydraulic system model, a hydraulic turbine model and a speed regulation system model; defining other models of the power system in the hydroelectric system except the hydroelectric generating set as an external system model, identifying the transfer function model of the external system model using a deep reinforcement learning algorithm, and then obtaining the relationship between the guide vane opening of the hydroelectric generating set and the system frequency response in the hydroelectric system; based on the hydroelectric generating set speed regulation model and the external system model, taking the additional damping controller as one of the control links of the speed regulator, taking the additional damping controller and the speed regulator as a whole, and optimizing the parameters of the speed regulator and the additional damping controller based on a chaotic particle swarm algorithm.

[0013] The method for jointly optimizing parameters of ultra-low frequency oscillation damping of a grid-connected hydroelectric system based on deep reinforcement learning model identification according to the technical solution of the present application can also have the following additional technical features: In the above technical solution, constructing the hydraulic system model comprises: based on the control equation of the typical water hammer theory, ignoring the friction coefficient of the water conduit, obtaining:

[0014] assuming that the water flow is incompressible and the loss of the water flow in the water conduit is negligible, the rigid water hammer equation is:

[0015] Considering the elasticity of the water conduit, the elastic water hammer equation under the compressible characteristics of water flow is:

[0016] wherein, represents a water hammer transfer function; represents a Laplace transform of a downstream water head; represents a Laplace transform of a downstream water flow rate; represents a water hammer effect time constant; represents an elastic time of the water conduit; and s represents a complex frequency in the Laplace transform.

[0017] In the above technical solution, the water turbine model is constructed by: Assuming that the water conduit wall is rigid and the water flow is incompressible, based on the rigid water hammer equation and ignoring the influence of the unit speed, the water turbine prime mover model is represented as:

[0018] wherein, represents a transfer function of the water turbine prime mover model; represents a transfer coefficient of the torque to the water turbine guide vane opening; represents a transfer coefficient of the flow rate to the water turbine guide vane opening; represents a transfer coefficient of the torque to the water head; represents a transfer coefficient of the flow rate to the water head.

[0019] In the above technical solution, the speed regulation system includes a PID regulator and a mechanical hydraulic system, and the speed regulation system model is constructed by: The transfer function of the PID regulator is represented as:

[0020] wherein, represents a transfer function of the PID regulator model; represents an output of the PID controller, i.e., a control input of the mechanical power; represents an angular frequency deviation; represents a deviation gain coefficient; represents a proportional gain of the regulator; represents an integral gain of the regulator; represents a differential gain of the regulator; represents a steady-state slip coefficient; The transfer function of the mechanical hydraulic system is represented as:

[0021] wherein, represents a transfer function of the mechanical hydraulic system; represents an opening deviation of the prime mover; represents a proportional gain of the mechanical hydraulic system; represents an integral gain of the mechanical hydraulic system; represents a differential gain of the mechanical hydraulic system; represents an opening time constant of the oil motor; represents a time constant of a feedback link of the oil motor.

[0022] In the technical solution, the deep reinforcement learning algorithm adopts a double-delay deep deterministic policy gradient algorithm; The deep reinforcement learning algorithm is used to identify the transfer function model of the external system model, and includes: The environment is defined to include a state space and a reward; in the state space, parameters of each order of the transfer function of the external system model are selected as state variables, an action is defined as adjusting the model parameters, and the reward is designed as an error measurement of the system output and the model predicted output quality inspection; The DRL agent is initialized, and the DRL agent includes an action network and an evaluation network, and is used to learn and evaluate the effect of the action; The DRL agent selects an action according to the current state, that is, adjusts the parameters of the transfer function; After the DRL agent performs the action, the environment corrects the parameters of the external system model according to the action, and provides a new state and a reward; The DRL agent evaluates the effect of the action through the comment network, and learns how to better select the action to minimize the error according to the reward signal; The agent continuously interacts with the environment to improve the identification accuracy of the external system model, and corrects the transfer function model parameters of the external system model according to the learning result of the agent, so as to improve the prediction performance of the model.

[0023] In the double-delay deep deterministic policy gradient algorithm, the calculation method of the target value for optimizing the comment network is:

[0024] wherein, represents a target value; represents an immediate reward, represents a discount factor, and are outputs of two target comment networks, represents a next state, represents a target action generated by the action network; The calculation method of the target action includes: ​

[0025] wherein, represents an action generated by the target action network; represents truncated Gaussian noise, which is a random variable sampled from a truncated normal distribution , and c represents a truncation threshold; Based on the target value, the review network is updated by minimizing the loss function, and the expression of the loss function is:

[0026] wherein, represents a loss function; represents an expected value operator; represents a value function estimated by the th review network, represents the parameters of the th review network, s represents the current state, represents the current action; The policy of the action network is optimized by maximizing the value function estimated by the review network, and the gradient calculation method of the action network includes:

[0027] wherein, represents the gradient of the parameters of the policy , represents the performance indicator of the policy , i.e. the objective function; represents the gradient of the action , which represents the sensitivity of the value function to the action; represents the action generated by the policy with parameters in state s; represents the gradient of the parameters of the policy .

[0028] In the above technical solution, the additional damping controller is taken as one of the control links of the speed regulator, and the additional damping controller and the speed regulator are regarded as a whole, which includes: On the basis of the speed regulator model, a multi-loop additional damping controller designed based on the phase compensation principle is introduced, and the multi-loop additional damping controller includes a low-pass filter link, a direct-current isolation link and a three-stage series phase compensation link connected in series; The additional mechanical torque component proportional to the speed deviation signal is generated in the unit dynamic process through the multi-loop additional damping controller, thereby providing positive damping; The cut-off frequency of the low-pass filter link is set to the upper limit of the frequency of the ultra-low frequency oscillation mode. The cut-off frequency of the DC blocking link is set to the lower limit of the frequency of the ultra-low frequency oscillation mode. The phase compensation link is configured with a three-stage first-order lag compensation network to achieve optimal phase matching between the control signal and the system dynamic response.

[0029] In the above technical solution, the expression of the first-order lag compensation network is:

[0030] wherein, represents the transfer function of the first-order phase compensation link; represents a gain coefficient for adjusting the degree of phase compensation; T represents a time constant; with an angular frequency When the input signal is an angular frequency, the lead phase compensation angle provided by the first-order lag compensation network is:

[0031] wherein, represents the lead phase compensation angle.

[0032] In the above technical solution, the multi-loop additional damping controller further includes a final-stage gain adjustment unit located at the lower stage of the three-stage series phase compensation link; The final-stage gain adjustment unit ensures that the additional torque signal maintains optimal damping strength under different operating conditions of the system through amplitude-frequency characteristic optimization.

[0033] In the above technical solution, the parameters of the speed regulator and the additional damping controller are optimized based on an improved chaotic particle swarm algorithm, including: The speed and position of the chaotic initialized particles are calculated; The fitness value of the particles is calculated; wherein the damping ratio of the ultra-low frequency oscillation mode is set as the fitness function value; It is judged whether the current particle fitness value is better than the individual and group extreme values; if yes, the individual extreme value and the group extreme value are updated and the next step is entered; if not, the next step is directly entered; The inertia weight and acceleration coefficient are updated non-linearly; The speed and position of the particles are updated until the termination condition is reached, and the optimal solution is output; The objective function and constraint condition of the improved chaotic particle swarm algorithm are represented as:

[0034] wherein, represents the particle fitness value; represents the time constant of the additional damping controller; Indicates the gain coefficient; This represents the proportional gain of the PID controller in the speed governor. This represents the integral coefficient of the PID control of the speed governor.

[0035] In summary, due to the adoption of the above-mentioned technical features, the beneficial effects of the present invention are: This patent proposes a joint parameter optimization method for hydropower unit governors and additional damping control based on a deep reinforcement learning model. It utilizes an accurate model of a high-proportion hydropower system identified by a deep reinforcement learning algorithm to jointly optimize governor parameters and additional frequency control parameters. By comprehensively considering the interaction between the governor's main control loop parameters and the additional damping controller parameters, the method effectively improves the damping characteristics of the hydropower system, thereby avoiding ultra-low frequency oscillations. Specifically, this method treats the governor itself and the additional damping controller as a whole, comprehensively considering the optimization of governor parameters and additional frequency control parameters. This joint optimization breaks through the limitations of traditional independent optimization, achieving a comprehensive improvement in multi-objective dynamic performance. Simultaneously, by enhancing broadband oscillation suppression and virtual inertia compensation capabilities, it significantly improves the robustness of high-proportion renewable energy power grids, supporting the safe and stable operation of "dual-high" power systems, combining technological advancement with economic practicality.

[0036] Additional aspects and advantages of the invention will become apparent in the following description or may be learned by practice of the invention. Attached Figure Description

[0037] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a flowchart of an embodiment of the present invention of a method for optimizing the joint parameters of ultra-low frequency oscillation damping in a grid-connected hydropower system based on deep reinforcement learning model identification; Figure 2 This is a schematic diagram of the hydropower unit structure in the method for optimizing the joint parameters of ultra-low frequency oscillation damping in a grid-connected hydropower system based on deep reinforcement learning model identification, according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the external system model identification method framework in the joint parameter optimization method for ultra-low frequency oscillation damping of grid-connected hydropower system based on deep reinforcement learning model identification, according to an embodiment of the present invention. Figure 4 This is a diagram of the TD3 algorithm architecture in an embodiment of the present invention, which is a method for optimizing the joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification. Figure 5A parameter joint optimization framework diagram in a grid-connected hydropower system ultra-low frequency oscillation damping joint parameter optimization method based on deep reinforcement learning model identification according to an embodiment of the present application is shown in the figure. Figure 6 An improved hydropower generator governor model schematic diagram considering additional damping control in a grid-connected hydropower system ultra-low frequency oscillation damping joint parameter optimization method based on deep reinforcement learning model identification according to an embodiment of the present application is shown in the figure. Figure 7 An improved chaotic particle swarm optimization algorithm flowchart in a grid-connected hydropower system ultra-low frequency oscillation damping joint parameter optimization method based on deep reinforcement learning model identification according to an embodiment of the present application is shown in the figure. Figure 8 An ultra-low frequency oscillation schematic diagram under load disturbance using original parameters according to an embodiment of the present application is shown in the figure. Figure 9 A comparison schematic diagram of the suppression effect of different methods on ultra-low frequency oscillation under load disturbance according to an embodiment of the present application is shown in the figure. Figure 10 An ultra-low frequency oscillation schematic diagram under short-circuit fault using original parameters according to an embodiment of the present application is shown in the figure. Figure 11 A comparison schematic diagram of the suppression effect of different methods on ultra-low frequency oscillation under short-circuit fault according to an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0038] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0039] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, however, the present application can also be implemented in other different ways from those described herein, therefore, the protection scope of the present application is not limited by the specific embodiments disclosed below.

[0040] The grid-connected hydropower system ultra-low frequency oscillation damping joint parameter optimization method based on deep reinforcement learning model identification according to some embodiments of the present application will be described below with reference to Figures 1 to 11 .

[0041] Some embodiments of the present application provide a grid-connected hydropower system ultra-low frequency oscillation damping joint parameter optimization method based on deep reinforcement learning model identification.

[0042] As shown in Figure 1 , the first embodiment of the present application proposes a grid-connected hydropower system ultra-low frequency oscillation damping joint parameter optimization method based on deep reinforcement learning model identification, including the following steps S1-S3.

[0043] S1, a water turbine generator speed regulation model is constructed, the water turbine generator speed regulation model comprising a hydraulic system model, a water turbine model and a speed regulation system model.

[0044] Specifically, the water turbine generator speed regulation system comprises a hydraulic system, a water turbine and a speed regulator, and the structure is as shown in Figure 2 The hydraulic system mainly comprises a pressure pipeline, a water diversion culvert and the like structures, and is characterized by a differential equation or an algebraic equation. The water turbine adopts a linear modeling method, and the speed regulation system model is jointly constituted by a PID controller and a mechanical hydraulic system.

[0045] In some embodiments, constructing the hydraulic system model comprises: Based on the typical water hammer theory control equation after Laplace transform, the friction coefficient of the water diversion pipeline is ignored, and the following is obtained:

[0046] Assuming that the water flow is incompressible and the loss of the water flow in the water diversion pipeline is negligible, the rigid water hammer equation is:

[0047] Considering the elasticity of the water diversion pipeline, the elastic water hammer equation under the compressibility of the water flow is:

[0048] wherein, represents a water hammer transfer function; represents the Laplace transform of the downstream water head; represents the Laplace transform of the downstream water flow rate; represents a water hammer effect time constant; represents an elastic time of the water diversion pipeline; and s represents a complex frequency in the Laplace transform.

[0049] In some embodiments, constructing the water turbine model comprises: Assuming that the wall of the water diversion pipeline is rigid and the water flow is incompressible, based on the rigid water hammer equation and ignoring the influence of the unit speed, the water turbine prime mover model is represented as:

[0050] wherein, represents a transfer function of the water turbine prime mover model; represents a transfer coefficient of the torque to the water turbine guide vane opening degree; represents a transfer coefficient of the flow rate to the water turbine guide vane opening degree; represents a transfer coefficient of the torque to the water head; represents a transfer coefficient of the flow rate to the water head.

[0051] A hydroelectric generator set speed regulator maintains the system frequency within the rated range by adjusting the mechanical power of the generator. The structure of a typical hydroelectric generator set PID speed regulator is analyzed. The hydroelectric generator set PID speed regulator is composed of a PID regulator and a mechanical hydraulic system. In some embodiments, constructing the speed regulation system model includes: The transfer function of the PID regulator is represented as:

[0052] wherein, represents the transfer function of the PID regulator model; represents the output of the PID controller, i.e., the control input of the mechanical power; represents the angular frequency deviation; represents the deviation gain coefficient; represents the regulator proportional gain; represents the regulator integral gain; represents the regulator derivative gain; represents the steady-state slip coefficient; The transfer function of the mechanical hydraulic system is represented as:

[0053] wherein, represents the transfer function of the mechanical hydraulic system; represents the opening deviation of the prime mover; represents the mechanical hydraulic system proportional gain; represents the mechanical hydraulic system integral gain; represents the mechanical hydraulic system derivative gain; represents the oil motor opening time constant; represents the time constant of the oil motor feedback link.

[0054] S2, define the other model of the power system in the hydroelectric system except the hydroelectric generator set as an external system model, use a deep reinforcement learning algorithm to identify the transfer function model of the external system model, and then obtain the relationship between the guide vane opening of the hydroelectric generator set and the system frequency response in the hydroelectric system.

[0055] Specifically, in order to obtain the relationship between the guide vane opening of the hydroelectric generator set and the system frequency response in the hydroelectric system, the transfer function model of the other model of the power system except the hydroelectric generator set is identified by deep reinforcement learning (Deep Reinforcement Learning, DRL), and the identification method framework is as shown in Figure 3 In deep reinforcement learning, the agent perceives the high-dimensional state space through a neural network, generates actions, and dynamically optimizes policy parameters based on the immediate reward feedback from the environment, realizing end-to-end mapping from state to action.

[0056] In some embodiments, the deep reinforcement learning algorithm employs a Twin Delayed Deep Deterministic Policy Gradient (TD3). As an efficient deep reinforcement learning method based on the Actor-Critic architecture, TD3 has been widely applied to optimization problems in continuous action space. By introducing mechanisms such as double critic networks, delayed policy updates, and target policy smoothing, TD3 effectively alleviates the overestimation bias and training instability problems prevalent in traditional deep deterministic policy gradient algorithms, thereby significantly improving stability and robustness and enabling effective identification tasks for high-dimensional and complex systems.

[0057] Specifically, in the model identification task of a hydroelectric system, the TD3 algorithm learns and identifies the dynamic characteristics of the hydroelectric system through real-time interaction with the system. Specifically, the state (State) can be composed of key variables of the hydroelectric system, and the parameters of each order of the system transfer function are selected as state variables; the action (Action) represents the adjustment amount of the model parameters to optimize the accuracy of the system model; and the reward (Reward) is designed as the error measure between the system output and the model predicted output, usually using mean square error (MSE) to evaluate the identification effect. The framework of the agent identifying the system model through interaction with the environment is shown in Figure 3 Through this interactive learning mechanism, the agent continuously interacts with the system, accurately captures the nonlinear and complex characteristics of the hydroelectric system according to the system response, and thus achieves precise system identification.

[0058] In one specific embodiment, the deep reinforcement learning algorithm is employed to identify the transfer function model of the external system model, comprising: defining an environment including a state space and a reward; in the state space, selecting the parameters of each order of the transfer function of the external system model as state variables, and defining the action as adjusting the model parameters; and designing the reward as the error measure between the system output and the model predicted output; initializing a DRL agent, which includes an action network and an evaluation network, for learning and evaluating the effect of the action; the DRL agent selects an action according to the current state, i.e., adjusts the parameters of the transfer function; After the DRL agent performs the action, the environment modifies the parameters of the external system model according to the action and provides a new state and a reward; the DRL agent evaluates the effect of the action through the evaluation network, and learns how to better select the action to minimize the error according to the reward signal; The agent continuously interacts with the environment, improves the recognition accuracy of the external system model, and corrects the transfer function model parameters of the external system model according to the learning result of the agent, so as to improve the prediction performance of the model.

[0059] The overall architecture of the TD3 algorithm is shown in Figure 4 It should be noted that the overall architecture of the TD3 algorithm is known to those skilled in the art and will not be described here.

[0060] It should be noted that in the double-delay deep deterministic policy gradient algorithm: The calculation method of the target value for optimizing the comment network is:

[0061] Among them, represents the target value; represents the immediate reward, represents the discount factor, and are the outputs of two target comment networks, represents the next state, represents the target action generated by the action network; In order to smooth the target action and introduce randomness, the calculation method of the target action includes:

[0062] Among them, represents the action generated by the target action network; represents truncated Gaussian noise, which is a random variable sampled from the truncated normal distribution , and c represents the truncation threshold; Based on the target value, the comment network is updated by minimizing the loss function, and the expression of the loss function is:

[0063] Among them, represents the loss function; represents the expected value operator; represents the value function estimated by the evaluation network, represents the parameters of the evaluation network, s represents the current state, represents the current action; The strategy of the action network is optimized by maximizing the value function estimated by the evaluation network, and the gradient calculation method of the action network includes:

[0064] where, denotes the gradient of the policy with respect to its parameters , denotes the performance metric of the policy , i.e. the objective function; denotes the gradient of the value function with respect to the action , denotes the action generated by the policy with parameters in state s; denotes the gradient of the policy with respect to its parameters .

[0065] By optimizing the gradient, the action network generates actions that maximize the Critic network estimate, thereby continuously improving the policy.

[0066] S3, based on the speed regulation model of the hydroelectric generating set and the external system model, the additional damping controller is taken as one of the control links of the speed regulator, the additional damping controller and the speed regulator are regarded as a whole, and the chaos particle swarm optimization (PSO) is used to optimize the parameters of the speed regulator and the additional damping controller.

[0067] Specifically, in the joint parameter optimization method of the present disclosure, based on the above hydroelectric generating set and external system model, the additional damping controller is taken as one of the control links of the speed regulator, the additional damping controller and the speed regulator are regarded as a whole, and a joint parameter optimization method considering the additional damping controller and the speed regulator of the hydroelectric generating set is proposed. The process of the joint parameter optimization method is introduced as follows, and the overall framework is shown in Figure 5 .

[0068] In some embodiments, the additional damping controller is taken as one of the control links of the speed regulator, and the additional damping controller and the speed regulator are regarded as a whole, including: A multi-loop additional damping controller based on the phase compensation principle is introduced on the basis of the speed regulator model, and the multi-loop additional damping controller includes a low-pass filter link, a direct-current isolation link and a three-stage series phase compensation link in series. The additional mechanical torque component proportional to the speed deviation signal is generated in the dynamic process of the unit through the multi-loop additional damping controller, thereby providing positive damping.

[0069] The speed regulator model of the hydroelectric generating set after adding the additional damping controller is shown in Figure 6 .

[0070] Specifically, the ultra-low frequency oscillation mode is mainly concentrated in 0.01Hz to 0.1Hz, and according to this characteristic, the additional damping controller is designed as follows: The cut-off frequency of the low-pass filter link is set to the upper limit of the ultra-low frequency oscillation mode, i.e. 0.1 Hz, which is accurately matched with the upper limit of the oscillation frequency, effectively attenuates harmonic components higher than the threshold, and particularly suppresses subsynchronous oscillation and high-frequency interference signals, avoiding false responses of the controller when the power grid appears subsynchronous / ultrasynchronous resonance.

[0071] The cut-off frequency of the DC blocking link is set to the lower limit of the ultra-low frequency oscillation mode, i.e. 0.01 Hz, and the steady-state deviation component in the speed signal is eliminated by constructing an inertial differential link.

[0072] The phase compensation link is configured with a three-stage first-order lag compensation network to achieve optimal phase matching between the control signal and the system dynamic response.

[0073] In some embodiments, the multi-loop additional damping controller further comprises a final-stage gain adjustment unit located at the lower stage of the three-stage series phase compensation link; the final-stage gain adjustment unit ensures that the additional torque signal maintains optimal damping strength under different operating conditions of the system through amplitude-frequency characteristic optimization. The output of the final-stage gain adjustment unit is superimposed with the output of the governor part and then input to the mechanical hydraulic system, thereby adjusting the mechanical power and further adjusting the guide vane opening of the hydraulic turbine, so as to change the output power of the hydraulic turbine to adapt to the changes of the power grid load.

[0074] In some embodiments, the expression of the first-order lag compensation network is:

[0075] wherein, G(s) represents the transfer function of the first-order phase compensation link; K represents a gain coefficient for adjusting the degree of phase compensation; T represents a time constant; When the input signal is an angular frequency , the lead phase compensation angle provided by the first-order lag compensation network is:

[0076] wherein, G(s) represents the lead phase compensation angle.

[0077] Taking the derivative of the above formula, when the derivative is zero, the maximum lead compensation angle of the first-order phase compensation link is:

[0078] When , At this time, the amplitude-frequency characteristic of the phase compensation link is close to 0, which does not meet the requirements. The optimal compensation angle of the single-stage phase compensation unit is 60 degrees, and the signal characteristics of the system after compensation still meet the requirements when the system vibrates at an extremely low frequency. To ensure the stability of the system under extreme conditions, a three-stage series compensation structure is adopted, which can accurately adjust the direction of the mechanical torque output by the system to a specific area (the third quadrant), effectively meeting the needs of various complex scenarios in actual operation.

[0079] In some embodiments, the improved chaotic particle swarm optimization (WPSO) is used to optimize the parameters of the speed regulator and the additional damping controller, and the optimization process is as shown in Figure 7 , which includes: S31, initializing the speed and position of the particles; S32, calculating the fitness value of the particles; wherein the damping ratio of the ultra-low frequency oscillation mode is set as the fitness function value; S33, determining whether the current particle fitness value is better than the individual and group extreme values; if yes, updating the individual and group extreme values and then proceeding to the next step; if not, directly proceeding to the next step; S34, updating the inertia weight and acceleration coefficient in a nonlinear manner; S35, updating the speed and position of the particles and returning to S33 until the termination condition is reached, and then outputting the optimal solution.

[0080] It should be noted that the improved chaotic particle swarm optimization (WPSO) optimization process has two main improvements compared to the standard particle swarm optimization (PSO): (1) The initial population of the standard PSO algorithm is randomly generated, and the distribution of the initial population is not representative. Chaos is a state of irregular motion with obvious nonlinear characteristics. Using the chaotic method to initialize the position and speed of the particle swarm can improve the diversity of the initial population.

[0081] (2) The inertia weight coefficient w of the standard PSO algorithm is replaced by a nonlinear decreasing inertia constant setting method, which can effectively improve the global and local search ability and convergence speed of the particle swarm algorithm.

[0082] Specifically, to optimize the parameters of the additional damping controller and improve the ultra-low frequency oscillation damping of the system, the fitness function value J is defined as:

[0083] wherein ξ represents the damping ratio of the ultra-low frequency oscillation mode, which is the fitness function value J.

[0084] The objective function and constraint conditions of the improved chaotic particle swarm optimization are represented as:

[0085] wherein, represents the particle fitness value; represents the additional damping controller time constant; represents the gain coefficient; represents the proportional coefficient of the governor; represents the integral coefficient of the governor. , The ranges of a, b, and c are set to [1, 10], [0.1, 2], and the additional damping controller time constant The typical range of K is [0.5, 50], and the gain coefficient K typically ranges from [0.1, 1].

[0086] In one specific embodiment, based on the method of the present disclosure, improvements are made on the basis of the standard four-machine two-area model to obtain a simulation test system. The main improvement is to replace the G2 and G4 generators with hydroelectric generators and simulate the system ultra-low frequency oscillation phenomenon under a step load signal. The improved four-machine two-area system contains two alternating current systems and is interconnected through a direct current transmission line. Both the sending and receiving ends contain two generators, one of which is a hydroelectric generator. The rated capacity of the generators is 650 MVA. Based on this model, the performance of the model identification method based on deep reinforcement learning, the hydroelectric generator speed regulation system parameter optimization method, and the ultra-low frequency oscillation suppression effect under load disturbance and short-circuit fault conditions are verified. The main steps are as follows: (1) Build a simulation model: combine the two hydroelectric generator speed regulators, the additional damping controller, and the identified system transfer function to establish a closed-loop simulation model; (2) Deep reinforcement learning for model identification: set the order of the external system model, and identify it using deep reinforcement learning. In this process, the ranges of state variables and action variables, hyperparameters, etc. of deep reinforcement learning need to be set. The performance of the deep reinforcement learning-based model identification method is evaluated in terms of the stabilization process of the evaluation reward curve and whether the finally identified parameters meet the requirements.

[0087] (3) Parameter joint optimization: and based on the improved chaotic particle swarm optimization algorithm, the speed regulator parameters and the additional damping controller parameters are optimized. In this process, key parameters such as inertia weight, learning rate, particle size, and iteration number of the particle swarm optimization algorithm need to be set. The performance of the improved chaotic particle swarm optimization algorithm-based parameter joint optimization method is evaluated in terms of the stabilization process of the fitness curve and whether the finally optimized parameters meet the requirements.

[0088] (4) Verification of suppression effect under load disturbance: at a certain time, simulate the loss of system load. Compare the ultra-low frequency oscillation suppression effect of the generator under three conditions: only with the original speed regulator, with the additional fixed parameter damping controller, and with the joint optimization of the additional damping controller and the speed regulator parameters.

[0089] (5) Verification of suppression effect under short-circuit fault: At a specific moment, a three-phase short-circuit ground fault occurs on the AC bus of the simulation system, and the fault is cleared after a period of time. The suppression effect of ultra-low frequency oscillation of the generator under three conditions are compared: the original parameters of the governor, the addition of a fixed parameter additional damping controller, and the joint optimization of the additional damping controller and governor parameters.

[0090] The verification results show that when identifying the transfer function of the external system of the hydropower unit, DRL can autonomously explore the temporal characteristics of the input-output signals under power grid disturbance scenarios, and simultaneously optimize the accuracy of parameter estimation and dynamic response fitting. Meanwhile, its end-to-end learning framework can effectively integrate noise and uncertainty in the operating data, avoiding identification bias caused by model mismatch in traditional methods, and providing more robust dynamic boundary conditions for subsequent collaborative optimization of hydropower unit parameters. Through joint optimization of governor parameters and additional damping control parameters, the damping characteristics of the system are improved, effectively suppressing the ultra-low frequency oscillation phenomenon of the hydropower system. Under load disturbance and short-circuit fault conditions, the joint parameter optimization method performs better than both methods with and without an additional damping controller.

[0091] Specifically, before optimization using the method disclosed herein, the ultra-low frequency oscillation under load disturbance using the original parameters is as follows: Figure 8 As shown, the suppression effects of different methods on ultra-low frequency oscillations under load disturbances are compared as follows: Figure 9 As shown, the suppression effect of the method disclosed in this paper is significantly better than that of the fixed-parameter additional damping controller method. Table 1 shows a comparison of the damping ratios of each method in the main oscillation modes of the system under load disturbance.

[0092] Table 1 Comparison of damping ratios for the main oscillation modes of the system under load disturbance.

[0093] Before optimization using the method disclosed herein, the ultra-low frequency oscillation under short-circuit faults using the original parameters is as follows: Figure 10 As shown, the suppression effects of different methods on ultra-low frequency oscillations under short-circuit faults are compared as follows: Figure 11 As shown, the suppression effect of the method disclosed in this paper is significantly better than that of the fixed-parameter additional damping controller method. Table 2 shows a comparison of the damping ratios of each method in the main oscillation modes of the system during short-circuit faults.

[0094] Table 2 Comparison of damping ratios of the main oscillation modes of the system during short-circuit faults.

[0095] Joint parameter optimization helps to achieve phase compensation of the mechanical torque of the hydropower unit as a whole, improve the unit's damping characteristics, and thus effectively suppress ultra-low frequency oscillations.

[0096] In this specification, illustrative statements about the terminology used do not necessarily limit the scope of the embodiments or examples to a given embodiment or example. Moreover, descriptive terminology such as first, second, etc. is not necessarily used to describe a particular embodiment or example, rather, such terminology is used in accordance with its ordinary meaning to distinguish between two or more instances of an element.

[0097] Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of the present application.

Claims

1. A method for optimizing the joint parameters of ultra-low frequency oscillation damping in a grid-connected hydropower system based on deep reinforcement learning model identification, characterized in that, include: A hydropower unit speed regulation model is constructed, which includes a hydraulic system model, a turbine model, and a speed regulation system model. The power system models other than hydropower units in the hydropower system are defined as external system models. Deep reinforcement learning algorithms are used to identify the transfer function models of the external system models, thereby obtaining the relationship between the guide vane opening of the hydropower units and the system frequency response in the hydropower system. Based on the hydropower unit speed regulation model and external system model, the additional damping controller is regarded as one of the control links of the governor. The additional damping controller and the governor are regarded as a whole, and the governor parameters and additional damping controller parameters are optimized based on the chaotic particle swarm algorithm.

2. The method for optimizing joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification as described in claim 1, characterized in that, Constructing the hydraulic system model includes: Based on the typical water hammer theory governing equations, after Laplace transformation, and neglecting the friction coefficient of the water intake pipe, we obtain: Assuming the water flow is incompressible and the losses in the water pipe are negligible, the rigid water hammer equation is as follows: Considering the elasticity of the water pipe, the elastic water hammer equation under the compressible characteristics of the water flow is: in, Represents the water hammer transfer function; This represents the Laplace transform of the downstream head; The Laplace transform of the downstream water flow rate; Indicates the time constant of the water hammer effect; represents the elastic time of the water pipe; s represents the complex frequency in the Laplace transform.

3. The method for optimizing joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification as described in claim 2, characterized in that, Constructing the turbine model includes: Assuming the water intake pipe wall is rigid and the water flow is incompressible; based on the rigid water hammer equation and neglecting the influence of unit speed, the turbine prime mover model is expressed as: in, The transfer function representing the prime mover model of the water turbine; This represents the torque transmission coefficient to the turbine guide vane opening. This represents the transfer coefficient of flow rate to the turbine guide vane opening. This represents the torque transfer coefficient to the water head. This represents the transfer coefficient of flow rate to head.

4. The method for optimizing joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification as described in claim 3, characterized in that, The speed control system includes a PID controller and a mechanical-hydraulic system. The model of the speed control system includes: The transfer function of a PID controller is expressed as: in, The transfer function representing the PID controller model; This represents the output of the PID controller, i.e., the control input for mechanical power. Indicates angular frequency deviation; Indicates the bias gain coefficient; Indicates the proportional gain of the regulator; Indicates the integral gain of the regulator; Indicates the differential gain of the regulator; Expressed as the permanent slip coefficient; The transfer function of a mechanical hydraulic system is expressed as: in, Represents the transfer function of a mechanical hydraulic system; This indicates the opening deviation of the prime mover; Indicates the proportional gain of a mechanical hydraulic system; Indicates the integral gain of the mechanical hydraulic system; This represents the differential gain of a mechanical hydraulic system. Indicates the hydrator start-up time constant; This represents the time constant of the feedback loop in the hydraulic actuator.

5. The method for optimizing joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification as described in claim 4, characterized in that, The deep reinforcement learning algorithm employs a dual-delay deep deterministic strategy gradient algorithm. The process of identifying the transfer function model of the external system model using a deep reinforcement learning algorithm includes: The environment is defined as a state space and a reward. In the state space, the parameters of each order of the external system model transfer function are selected as state variables, and the action is defined as adjusting the model parameters. The reward is designed as an error metric for quality control between the system output and the model prediction output. Initialize the DRL agent, which includes an action network and an evaluation network for learning and evaluating the effects of actions; The DRL agent selects an action based on the current state, which is to adjust the parameters of the transfer function. After the DRL agent performs an action, the environment corrects the parameters of the external system model based on the action and provides new states and rewards. The DRL agent evaluates the effectiveness of actions through a comment network and learns how to better select actions to minimize errors based on reward signals; The agent continuously interacts with the environment to improve the accuracy of identifying external system models. Based on the learning results of the agent, the transfer function model parameters of the external system model are corrected to improve the predictive performance of the model.

6. The method for optimizing joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification as described in claim 5, characterized in that, In the dual-delay deep deterministic policy gradient algorithm: The method for calculating the target value used to optimize the comment network is as follows: in, Indicates the target value; Indicates an immediate reward. Indicates the discount factor. and These are the outputs of two target comment networks. Indicates the next state. This represents the target action generated by the action network; The methods for calculating the target action include: in, This represents the action generated by the target action network; This represents truncated Gaussian noise, derived from a truncated normal distribution. The random variable sampled in the middle, where c represents the cutoff threshold; Based on the target value, the comment network is updated by minimizing the loss function, which is expressed as: in, Represents the loss function; This represents the expected value operator; Indicates by the first The value function estimated by the evaluation network Indicates the first The parameters of an evaluation network, s Indicates the current state. Indicates the current action; The gradient calculation methods for the action network include: optimizing the action network's policy by maximizing the value function estimated by the evaluation network; and further methods for calculating the gradient of the action network. in, Indicates the strategy parameters Find the gradient. Representation Strategy The performance metrics, i.e., the objective function; Indicates the action beg The gradient of represents the sensitivity of the value function to the action; This indicates that in state s, the parameter is... strategy The generated action; Indicates the strategy Its parameters Find the gradient.

7. The method for optimizing joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification as described in claim 1, characterized in that, The addition of the damping controller as one of the control elements of the speed governor, and the consideration of the additional damping controller and the speed governor as a whole, includes: Based on the governor model, a multi-loop additional damping controller designed based on the phase compensation principle is introduced. The multi-loop additional damping controller includes a low-pass filter, a DC blocking link, and a three-stage series phase compensation link connected in series. The multi-loop additional damping controller generates an additional mechanical torque component that is proportional to the speed deviation signal during the dynamic process of the unit, thereby providing positive damping. The cutoff frequency of the low-pass filter is set to the upper frequency limit of the ultra-low frequency oscillation mode. The cutoff frequency of the DC blocking element is set to the lower limit of the frequency of the ultra-low frequency oscillation mode; The phase compensation stage is equipped with a three-level first-order lag compensation network to achieve optimal phase matching between the control signal and the system dynamic response.

8. The method for optimizing joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification as described in claim 7, characterized in that, The expression for the first-order lag compensation network is: in, The transfer function of the first-order phase compensation circuit is represented. This represents the gain coefficient, used to adjust the degree of phase compensation; T represents the time constant. In terms of angular frequency When the input signal is a first-order hysteresis compensation network, the leading phase compensation angle provided by the network is: in, This indicates the leading phase compensation angle.

9. The method for optimizing joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification as described in claim 7, characterized in that, The multi-loop additional damping controller also includes a final-stage gain adjustment unit located below the three-stage series phase compensation stage; The final-stage gain adjustment unit ensures that the additional torque signal maintains optimal damping strength under different system operating conditions through amplitude-frequency characteristic optimization.

10. The method for optimizing joint parameters of ultra-low frequency oscillation damping in grid-connected hydropower systems based on deep reinforcement learning model identification as described in claim 7, characterized in that, The governor parameters and additional damping controller parameters are optimized based on an improved chaotic particle swarm optimization algorithm, including: Chaos initializes the velocity and position of particles; Calculate the fitness value of the particles; where the damping ratio of the ultra-low frequency oscillation mode is set as the fitness function value; Determine if the current particle fitness value is better than the individual and population extreme values; if so, update the individual and population extreme values ​​and proceed to the next step; otherwise, proceed directly to the next step. The inertia weight and acceleration coefficient are updated nonlinearly; Update the particle's velocity and position until the termination condition is met, then output the optimal solution; The objective function and constraints of the improved chaotic particle swarm optimization algorithm are expressed as follows: in, This represents the particle fitness value; Indicates the time constant of the additional damping controller; Indicates the gain coefficient; Indicates the proportional gain of the speed controller; This represents the integral coefficient of the speed governor.