A network configuration type wind power grid connection subsynchronous oscillation suppression method based on improved second-order active disturbance rejection

CN122659902APending Publication Date: 2026-08-28YUNNAN ELECTRIC POWER TESTING & RES INST (GRP) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610787698.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0004]进一步分析表明,当风电经串联补偿装置并网系统受到扰动后,构网型网侧换流器电流内环的比例-积分控制器会在特定频段引入负阻尼效应,不仅不能有效抑制扰动电流,反而会对其产生放大作用,从而诱发并加剧次同步振荡

Benefits of technology

[0021] The beneficial effects of this invention are as follows: By replacing the original voltage and current dual-loop PI control of the grid-side converter with second-order linear active disturbance rejection control, the problem of weak disturbance rejection capability and easy introduction of negative damping to induce subsynchronous oscillations in traditional PI control is effectively solved. At the same time, by adaptively optimizing the core parameters of the active disturbance rejection controller through an intelligent agent trained based on a deep deterministic strategy gradient algorithm, the defects of manually tuned parameters that are difficult to adapt to complex and variable power grid conditions are overcome. The control parameters can be dynamically adjusted according to the real-time operating status of the system, realizing adaptive suppression of subsynchronous oscillations in the grid-connected wind turbine system including a series compensation device, and significantly improving the operational stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122659902A_ABST
    Figure CN122659902A_ABST
Patent Text Reader

Abstract

The application discloses a network-constructing type wind power grid-connected subsynchronous oscillation suppression method based on improved second-order active disturbance rejection, and the oscillation suppression method comprises the following steps: constructing a parameterized model, wherein the parameterized model comprises a network-constructing type grid-side converter; adopting a second-order linear active disturbance rejection control to replace voltage and current double-loop control in the network-constructing type grid-side converter, so as to obtain a network-constructing type grid-side converter control model based on the second-order active disturbance rejection; constructing and training an intelligent agent comprising a deep deterministic policy gradient algorithm, so that the intelligent agent can output optimal active disturbance rejection parameters of the network-constructing type grid-side converter control model; collecting system states in real time and inputting the system states into the trained intelligent agent, so as to obtain the optimal active disturbance rejection parameters; and applying the optimal active disturbance rejection parameters to the network-constructing type grid-side converter control model, generating a modulation signal to drive the converter, and realizing adaptive suppression of subsynchronous oscillation of a network-constructing type wind turbine grid-connected system comprising a series compensation device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid oscillation suppression technology, and in particular to a method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system. Background Technology

[0002] The large-scale integration of renewable energy sources, such as wind and solar power, and numerous power electronic converters into the power system is increasingly characterized by "low inertia and weak damping," posing a severe challenge to frequency and voltage stability. Virtual synchronous machine control, by simulating the external characteristics of a synchronous generator through active-frequency and reactive-voltage equations, can provide virtual inertia and voltage support for the system, making it a key technology for enhancing the active support capability of renewable energy grid connection.

[0003] However, the introduction of virtual synchronous machine control has further complicated the subsynchronous oscillation problem in grid-connected wind power systems via series compensation capacitors. When subsynchronous oscillations occur in such systems, they may induce serious accidents such as wind turbine shutdowns and equipment damage, endangering the safe operation of the system.

[0004] Further analysis reveals that when a wind power grid-connected system via a series compensation device is disturbed, the proportional-integral controller in the inner current loop of the grid-connected grid-side converter introduces a negative damping effect in a specific frequency band. This not only fails to effectively suppress the disturbance current but also amplifies it, thereby inducing and exacerbating subsynchronous oscillations. Therefore, with the integration of numerous new energy sources and power electronic devices into the grid, the dynamic interaction of the system becomes increasingly complex. The existing control strategies for wind power grid-connected grid-side converters, which rely solely on a linear combination of the proportional and integral values ​​of the error in the inner current loop, are no longer sufficient to effectively suppress subsynchronous oscillations and may even become a contributing factor to worsening these oscillations. A new control scheme is urgently needed to address this problem. Summary of the Invention

[0005] In view of the above-mentioned prior art, the present invention provides a method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system, which mainly solves the technical problems existing in the above-mentioned background art.

[0006] To achieve the above objectives, the technical solution of this invention is implemented as follows: This invention discloses a method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system. The oscillation suppression method includes: A parameterized model of a grid-connected wind turbine system with a series compensation device is constructed. The parameterized model includes a grid-connected grid-side converter, and the control strategy of the grid-connected grid-side converter includes voltage and current dual-loop control. The voltage and current dual-loop control in the grid-side converter is replaced by second-order linear active disturbance rejection control, resulting in a control model for the grid-side converter based on second-order active disturbance rejection. Using the parameterized model as the interaction environment, an agent containing a deep deterministic policy gradient algorithm is constructed and trained, so that the agent can output the optimal active disturbance rejection parameters of the grid-side converter control model. The system status is collected in real time and input into the trained agent to obtain the optimal active disturbance rejection parameters; The optimal active disturbance rejection parameters are applied to the control model of the grid-type grid-side converter to generate a modulation signal to drive the converter, thereby achieving adaptive suppression of subsynchronous oscillations in the grid-type wind turbine grid-connected system containing a series compensation device.

[0007] Optionally, a second-order active disturbance rejection controller with state error feedback and total disturbance feedforward compensation can be constructed based on a linear extended state observer and a linear tracking differentiator.

[0008] Optionally, the expression for the linearly extended state observer is:

[0009] In the formula, To output the tracking signal for y, The tracking signal is used to output the first derivative of y. The tracking signal for the total system disturbance. , , The time derivative of the corresponding state variable. , , For the gain of the linearly extended state observer, This represents the gain coefficient of the control channel in the original second-order system model. This is the voltage control command for the d-axis port of the grid-side converter.

[0010] Optionally, the linear tracking differentiator uses a simplified control law to generate the basic control quantity, the expression of which is:

[0011] In the formula, This is the reference value for the grid connection point voltage on the d-axis. For linear tracking differentiator output, , This is the gain of the linear tracking differentiator.

[0012] Optionally, based on the output of the tracking differentiator and the tracking signal of the extended state observer, a state error feedback quantity is generated, and the total disturbance feedforward compensation quantity is superimposed to obtain the output of the second-order active disturbance rejection controller.

[0013] Optionally, using the parameterized model as the interaction environment, an agent with an Actor-Critic architecture is constructed, and the agent is trained using a deep deterministic policy gradient algorithm, specifically including: The parameterized model is used as the environment for agent interaction. The observation state, action variables and reward function of the agent are defined. The observation state is the grid connection point voltage and grid current in the synchronous rotating coordinate system. The action variables are the gain parameters of the linear extended state observer in the second-order linear active disturbance rejection control. Construct an agent comprising a Critic network and an Actor network. The Actor network takes the observed state as input and outputs an action variable for adjusting the active disturbance rejection control parameters. The Critic network takes the observed state and the action variable for adjusting the active disturbance rejection control parameters as input and outputs a state-action evaluation function. The agent is trained using a deep deterministic policy gradient algorithm, enabling it to output the optimal self-disturbance rejection core parameters.

[0014] Optionally, the expression for the state-action evaluation function is:

[0015] In the formula, o represents the observed state of the environment. For the action variable of active disturbance rejection control parameter adjustment, For the expectation, , For the new state and new action variables in the next moment, The main evaluation network parameters, As a discount factor, For the reward function, The new state-action evaluation function for the next moment.

[0016] Optionally, the Critic network updates its parameters by minimizing a loss function, the expression of which is:

[0017]

[0018] In the formula, For storing experience data generated by the interaction between intelligent agents and the environment Experience replay buffer, For the target Actor network parameters, r This is the reward value.

[0019] Optionally, the reward function is the ratio and integral of the error between the instantaneous active power output by the virtual synchronizer and the active power reference value. The expression for the reward function is as follows:

[0020] In the formula, The instantaneous active power output by the virtual synchronous machine. This is a reference value for active power. , These are the weighting coefficients.

[0021] The beneficial effects of this invention are as follows: By replacing the original voltage and current dual-loop PI control of the grid-side converter with second-order linear active disturbance rejection control, the problem of weak disturbance rejection capability and easy introduction of negative damping to induce subsynchronous oscillations in traditional PI control is effectively solved. At the same time, by adaptively optimizing the core parameters of the active disturbance rejection controller through an intelligent agent trained based on a deep deterministic strategy gradient algorithm, the defects of manually tuned parameters that are difficult to adapt to complex and variable power grid conditions are overcome. The control parameters can be dynamically adjusted according to the real-time operating status of the system, realizing adaptive suppression of subsynchronous oscillations in the grid-connected wind turbine system including a series compensation device, and significantly improving the operational stability and reliability of the system. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only preferred embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart illustrating the subsynchronous oscillation suppression method for grid-connected wind power based on an improved second-order active disturbance rejection in the embodiments of this application. Figure 2 Structure and control block diagram of a grid-connected system for direct-drive wind turbines via series compensation device; Figure 3 To improve the second-order active disturbance rejection control block diagram; Figure 4 The control block diagram is shown for the subsynchronous oscillation suppression architecture based on the improved second-order active disturbance rejection strategy. Figure 5 A schematic diagram of the reward curve for improving the second-order active disturbance rejection strategy; Figure 6 A schematic diagram of the output active power curves for voltage and current loops using PI and an improved second-order active disturbance rejection strategy. Figure 7 A schematic diagram of the output current spectrum analysis for the voltage-current loop using an improved second-order active disturbance rejection strategy. Figure 8 This is a schematic diagram of the voltage-current loop using PI output current spectrum analysis. Detailed Implementation

[0024] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. In the following description, the expression "some embodiments" refers to a subset of all possible embodiments; however, it should be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0025] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention can be practiced without one or more of these details. In other instances, certain technical features well-known in the art have not been described in order to avoid obscuring the invention.

[0026] It should be understood that the present invention can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein. Rather, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the invention to those skilled in the art. Furthermore, the terminology used herein is intended only to describe particular embodiments and is not intended to limit the invention. When used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “compose” and / or “comprising,” when used in this specification, identify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups. When used herein, the term “and / or” includes any and all combinations of the associated listed items.

[0027] It should also be noted that when an element is referred to as being "fixed to" another element, it can be directly attached to the other element or there may be an intervening element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or there may be an intervening element. The terms "vertical," "horizontal," "inner," "outer," "left," "right," and similar expressions used herein are for illustrative purposes only and do not represent the only possible implementation.

[0028] To fully understand this invention, a detailed structure will be presented in the following description to illustrate the technical solution proposed by this invention. Optional embodiments of the invention are described in detail below; however, in addition to these detailed descriptions, the invention may have other embodiments.

[0029] Please refer to the attached document. Figure 1 The first aspect of this application provides a method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system. The oscillation suppression method includes: S1. Construct a parameterized model of a grid-connected wind turbine system containing a series compensation device. The parameterized model includes a grid-connected grid-side converter, and the control strategy of the grid-connected grid-side converter includes voltage and current dual-loop control. See Figure 2 The established parameterized model includes the grid-type direct-drive wind turbine generator body, the turbine-side converter, the grid-type grid-side converter, the DC bus, the transmission line, the series compensation device, and the power grid. The turbine-side output terminal of the direct-drive wind turbine generator is connected to the AC side of the turbine-side converter. The DC side of the turbine-side converter is connected to the DC side of the grid-side converter through the DC bus. The AC side of the grid-side converter is connected to the power grid through the transmission line and the series compensation device, forming a complete back-to-back converter grid-connected topology.

[0030] Furthermore, the grid-type grid-side converter is equipped with corresponding control strategies, including virtual synchronous machine control, voltage loop, current loop, coordinate transformation, and PWM modulation. To facilitate understanding of the technical solution of this invention, the control strategies of the grid-type grid-side converter are described in detail below.

[0031] 1) Virtual Synchronous Machine Control Virtual synchronous machines simulate the inertia and damping characteristics of synchronous machines, giving them the characteristics of inertia and primary frequency regulation. Virtual synchronous machine control includes instantaneous power calculation, active power control, and reactive power control.

[0032] The instantaneous output power of virtual synchronous control is calculated based on the grid connection point voltage and grid current according to instantaneous power theory, as shown in the following formula:

[0033] In the formula, P e , Q e These represent the instantaneous active power and instantaneous reactive power outputs, respectively. u pd , u pq They are respectively dq The grid connection point voltage in the coordinate system is determined by the grid connection point. abc Three-phase voltage u pabc Obtained through coordinate transformation; i gd , i gq They are respectively dq In the coordinate system, the grid current is composed of the three-phase current of the grid.i gabc It is obtained through coordinate transformation.

[0034] The active power control and reactive power control of the virtual synchronous machine are obtained from the active power-frequency regulation equation and the reactive power-voltage equation, respectively, as shown below:

[0035] in,

[0036]

[0037]

[0038] In the formula, T set Given a reference torque; T e Electromagnetic torque; Output angular frequency for the virtual synchronizer; Electromagnetic torque; D p This is the active damping coefficient; D q This is the reactive power damping coefficient; P set This is a reference value for active power. Q set This is a reference value for reactive power. K The inertia coefficient; The phase of the internal potential of the virtual synchronous machine; E m This is the effective value of the internal potential; V nom This is the effective value of the rated voltage; V This is the effective value of the output voltage.

[0039] 2) Voltage and current loop Based on the structure of the grid-side converter of the grid-type direct-drive wind turbine, it can be obtained that... dq Nodal voltage equations in coordinate system.

[0040]

[0041] In the formula, u cd , u cq They are respectively dq Converter port voltage in coordinate system; L f , R f These are the filter resistor and the filter inductor, respectively.

[0042] Depend on dq The mathematical equations for current loop control can be derived from the nodal voltage equations in the coordinate system, as shown below.

[0043]

[0044] In the formula, i dref , i qref They are respectively dq Reference value of current in coordinate system; k cp , k cd These are the proportional and integral parameters of the current loop, respectively. G c = k cp + k ci / s .

[0045] The current reference value for current loop control is provided by the voltage loop, and the mathematical expression is shown below.

[0046]

[0047] In the formula, u dref , u qref They are respectively dq Voltage reference value in coordinate system; k vp , k vd These are the proportional and integral parameters of the voltage loop, respectively. G v = k vp + k vi / s The voltage reference value is provided by a virtual synchronous machine. Based on the grid voltage-oriented vector control method, the technical solution of this invention adopts... d Axis orientation, therefore let u dref = E m , u qref =0.

[0048] S2. Replace the voltage and current dual-loop control in the grid-side converter with second-order linear active disturbance rejection control to obtain a control model for the grid-side converter based on second-order active disturbance rejection.

[0049] See Figure 3-5In grid-side converters of grid-type direct-drive wind turbines, the voltage and current loops typically employ proportional-integral (PI) control strategies. However, PI controllers are relatively simple in structure and struggle to effectively suppress system disturbances. Furthermore, the current loop may amplify disturbance currents during dynamic response, potentially inducing new subsynchronous oscillations and impacting system stability and reliability.

[0050] To address the aforementioned problems, Active Disturbance Rejection Control (ADRC) is a nonlinear control method developed based on PID control and utilizing an error feedback mechanism. ADRC mainly comprises four parts: a tracking differentiator, an extended state observer, nonlinear state error feedback, and disturbance estimation compensation. The extended state observer is the core component of ADRC, enhancing the system's anti-interference capability through real-time disturbance estimation and compensation. The technical solution of this invention is based on a linear extended state observer and a linear tracking differentiator, constructing a second-order ADRC controller with state error feedback and total disturbance feedforward compensation. Furthermore, the second-order linear ADRC replaces the PI control with voltage-current loop control, thereby suppressing subsynchronous oscillations.

[0051] Specifically, when constructing a second-order active disturbance rejection controller, the second-order system is obtained based on the aforementioned mathematical expression of the voltage and current loop of the grid-side converter. Taking the d-axis as an example, the expression of its second-order system is as follows:

[0052] In the formula, for d The second-order derivative of the converter port voltage without coupling terms; , They are respectively d The second and first derivatives of the grid connection point voltage.

[0053] Let the converter port voltage without coupling terms be For the system input and grid connection point voltage For the system output, the second-order system can be rearranged to obtain the following formula:

[0054] After further simplification, the following equivalent substitution can be made:

[0055]

[0056] in, Let the total disturbance of the system be represented by the following state variables:

[0057] The second-order system can be described in the form of state equations, as shown below:

[0058] A third-order linearly extended state observer can be obtained from the system's state equations. The expression for the linearly extended state observer is as follows:

[0059] In the formula, To output the tracking signal for y, The tracking signal is used to output the first derivative of y. The tracking signal for the total system disturbance. , , The time derivative of the corresponding state variable. , , For the gain of the linearly extended state observer, This represents the gain coefficient of the control channel in the original second-order system model. This is the voltage control command for the d-axis port of the grid-side converter.

[0060] Furthermore, the linear tracking differentiator uses a simplified control law to generate the basic control quantity, the expression of which is:

[0061] In the formula, This is the reference value for the grid connection point voltage on the d-axis. For linear tracking differentiator output, , This is the gain of the linear tracking differentiator.

[0062] Furthermore, a total disturbance feedforward compensation mechanism is introduced, combining the basic control quantity output by the linear tracking differentiator with the total disturbance estimate output by the linear extended state observer to obtain a second-order active disturbance rejection controller, the structure of which is as follows: Figure 3 As shown, the final output of the second-order active disturbance rejection controller is:

[0063] S3. Using the parameterized model as the interaction environment, construct and train an agent containing a deep deterministic policy gradient algorithm, so that the agent can output the optimal active disturbance rejection parameters of the grid-side converter control model.

[0064] The Deep Deterministic Policy Gradient (DDPG) algorithm is an offline policy-based deep reinforcement learning algorithm that combines the advantages of Deep Q-Network (DQN) and Deterministic Policy Gradient (DPG) algorithms. It is specifically designed to solve sequential decision-making problems in continuous action spaces. Traditional DQN algorithms can only handle discrete action spaces and cannot be directly applied to continuous action scenarios such as parameter tuning; while traditional DPG algorithms suffer from training instability and divergence. The DDPG algorithm effectively solves these problems by introducing an Actor-Critic dual-network architecture, an experience replay mechanism, and a soft update mechanism for the target network, enabling it to learn the optimal deterministic policy in high-dimensional continuous states and action spaces.

[0065] The core architecture of the DDPG algorithm comprises two types of neural networks: a policy network (Actor network) and a value network (Critic network). Each network has two independent parameter systems: a main network and a target network, forming a dual-network structure. The Actor network outputs deterministic action variables based on the current observed state, with its core objective being to maximize long-term cumulative reward. The Critic network evaluates the long-term value of performing the corresponding action in the current state, with its core objective being to accurately predict the state-action value function and provide gradient guidance for the Actor network's parameter updates. An experience replay mechanism stores all experience samples generated during the agent's interaction with the environment. During training, small batches of samples are randomly drawn from the buffer for learning, effectively breaking the temporal correlation between samples and avoiding gradient oscillations during network training. The target network soft update mechanism allows the target network's parameters to slowly follow the changes in the main network, rather than directly copying the main network's parameters. This ensures the stability of the target value function during training and solves the training divergence problem caused by drastic fluctuations in target values ​​in traditional reinforcement learning algorithms. In some embodiments of the present invention, the parameterized model is used as the interaction environment for the intelligent agent. This model can accurately simulate the dynamic response characteristics of the system under different operating conditions and disturbances, providing a real and reliable interaction platform for the training of the intelligent agent.

[0066] In some embodiments of the present invention, when constructing an agent containing a deep deterministic policy gradient algorithm, the agent's observation state must first be defined as the grid connection point voltage and grid current in a synchronously rotating dq coordinate system, the expression of which is: , The d-axis component of the grid connection point voltage. The q-axis component of the grid connection point voltage. The d-axis component of the grid current. This represents the q-axis component of the grid current.

[0067] The action variable of the agent is defined as the gain parameter of the linear extended state observer in second-order linear active disturbance rejection control, and its expression is: These are the three gain coefficients of the third-order linear extended state observer, which are core parameters determining the disturbance estimation accuracy and control performance of the active disturbance rejection controller. A reward function is also defined, using the ratio and integral of the error between the instantaneous active power output of the virtual synchronizer and the active power reference value as the reward function, and its expression is:

[0068] In the formula, The instantaneous active power output by the virtual synchronous machine. This is a reference value for active power. , The weighting coefficients are used to guide the agent to optimize parameters in order to reduce active power tracking error and improve system stability.

[0069] An agent is constructed comprising a main Actor network, a target Actor network, a main Critic network, and a target Critic network. The Actor network takes the observed state o as input and outputs the action variable a adjusted by the active disturbance rejection control parameters. The Critic network takes the observed state o and the action variable a adjusted by the active disturbance rejection control parameters as input and outputs a state-action evaluation function, which measures the expected value of the long-term cumulative reward that can be obtained by performing the corresponding action in the current state.

[0070] After the agent is built, it is trained. During the training process, the interaction samples are stored using an experience replay mechanism, the parameters of the main Critic network are updated by minimizing the mean square error loss function, the parameters of the main Actor network are updated by the policy gradient ascent method, and the parameters of the target network are updated by a soft update mechanism.

[0071] Specifically, before training begins, the weight parameters of the main Actor network, target Actor network, main Critic network, and target Critic network are randomly initialized. The initial parameters of the target Actor network are set to be the same as those of the main Actor network, and the initial parameters of the target Critic network are set to be the same as those of the main Critic network. At the same time, an experience replay buffer of a preset capacity is initialized to store the state, action, reward, and next state sample data generated during the interaction between the agent and the environment.

[0072] In each training round, the parameterized model is first reset to its initial running state to obtain the initial observation state. The agent then acts on the environment based on the action variables output by the main Actor network, superimposed with Gaussian exploration noise, to ensure that the agent has sufficient exploration capabilities in the early stages of training to traverse different parameter spaces. After receiving the action variables, the environment updates its running state, generates the observation state for the next time step, and calculates the immediate reward value for the current time step according to the aforementioned reward function. The four-tuple experience sample consisting of the current observation state, action variables, immediate reward value, and the next observation state is stored in the experience replay buffer.

[0073] Once the number of samples in the experience replay buffer reaches a preset minimum batch threshold, experience samples of a preset batch size are randomly selected from the experience replay buffer for training. Random sampling breaks the temporal correlation between samples, preventing gradient oscillations during network training. For each selected experience sample, the target Actor network generates a target action based on the next observation state, and then the target Critic network calculates the target state-action value based on the next observation state and the target action. The expression for the state-action evaluation function is as follows: The expression is:

[0074] In the formula, o represents the observed state of the environment. For the action variable of active disturbance rejection control parameter adjustment, For the expectation, , For the new state and new action variables in the next moment, The main evaluation network parameters, As a discount factor, For the reward function, The new state-action evaluation function for the next moment.

[0075] Based on the calculated target Q-value and the current state-action value output by the main Critic network, a mean squared error loss function is constructed. This loss function is used to measure the state-action evaluation function assessed by the evaluation network. Compared with expected cumulative rewards y m To address the differences between the two, gradient descent is used to update the parameters of the main Critic network, gradually approximating the true long-term cumulative reward expectation. The loss function expression is as follows:

[0076]

[0077] In the formula, For storing experience data generated by the interaction between intelligent agents and the environment Experience replay buffer, For the target Actor network parameters, r This is the reward value.

[0078] After the parameters of the main Critic network are updated, the parameters of the main Actor network are updated using the policy gradient ascent method, with the goal of maximizing the expected value of the state-action value function output by the main Critic network. The objective is to maximize the expected value of the state-action evaluation function so that the main Actor network can output the optimal action that yields the maximum long-term reward in the current state. The policy gradient expression is as follows:

[0079] In the formula, Action strategy Expected rewards; The gradient of the policy probability with respect to the parameters measures the impact of changes in network parameters on the policy.

[0080] After each update of the main network parameters, a soft update mechanism is used to synchronize the parameters of the target Actor network and the target Critic network. The soft update rules are as follows:

[0081]

[0082] In the formula, The learning rate represents the speed at which the network parameters are updated. The soft update mechanism allows the parameters of the target network to slowly follow the changes of the main network, ensuring the stability of the target value function and avoiding algorithm divergence caused by drastic fluctuations in the target value during training.

[0083] The process of agent-environment interaction, experience sample collection, small-batch sample training, and network parameter updates is repeated until the number of training rounds reaches the preset maximum number of iterations, or the average reward value fluctuation of multiple consecutive training rounds is less than a preset threshold and no longer significantly increases. At this point, the agent training is considered converged. After training, the agent can output the optimal linear extended state observer gain parameters based on the real-time operating state of the system, enabling the second-order active disturbance rejection controller to maintain excellent disturbance suppression capability and subsynchronous oscillation suppression effect under different power grid operating conditions.

[0084] Furthermore, after training, the action network parameters of each agent are retained. Based on the current environmental observation state, the agent generates adjustment actions for the extended state observer gain, thereby achieving adaptive suppression of subsynchronous oscillations in the grid-connected wind power system via series compensation device. The parameters are shown in the table below.

[0085] Table 1 System Simulation Parameters

[0086] S4. Collect the system status in real time and input it into the trained agent to obtain the optimal active disturbance rejection parameters; Specifically, during operation, the voltage and current transformers configured at the grid connection point collect the three-phase grid connection point voltage and three-phase grid current signals in real time. These signals are then converted into grid connection point voltage components and grid current components in the synchronous rotating dq coordinate system through coordinate transformation, thus forming the real-time observation status of the intelligent agent.

[0087] The real-time observed states are input into the trained agent's main Actor network. Based on the learned parameter tuning strategy, the main Actor network outputs the optimal linear extended state observer gain parameters corresponding to the current system operating state. During this process, the agent no longer adds exploratory noise and directly outputs deterministic optimal parameter values, ensuring the stability and reliability of the control process. S5. Apply the optimal active disturbance rejection parameters to the grid-side converter control model to generate a modulation signal to drive the converter, thereby achieving adaptive suppression of subsynchronous oscillations in the grid-connected wind turbine system containing a series compensation device.

[0088] Specifically, the gain parameters of the optimal linear extended state observer output by the agent in real time are dynamically updated to the third-order linear extended state observer of the second-order active disturbance rejection controller. Based on the updated parameters, the second-order active disturbance rejection controller combines the real-time collected grid connection point voltage signal and the voltage reference value output by the virtual synchronous machine to obtain the grid-side converter port voltage control command in the dq coordinate system.

[0089] The aforementioned port voltage control commands are converted into three-phase voltage commands in an abc three-phase stationary coordinate system via coordinate transformation. These commands are then sent to the pulse width modulation module to generate corresponding drive pulse signals, which drive the insulated gate bipolar transistor (IGBT) switching devices of the grid-side converter to turn on and off according to preset logic, thereby achieving precise control of the output voltage and current of the grid-side converter. When system operating conditions change or grid disturbances occur, the intelligent agent can quickly adjust the core parameters of the active disturbance rejection controller (ADRC) based on real-time system status data. This ensures that the ADRC maintains optimal disturbance estimation accuracy and compensation effect, effectively suppressing potential subsynchronous oscillations in the system and significantly improving the operational stability and reliability of grid-connected wind turbine systems including series compensation devices. The effect is as follows: Figure 6-8 As shown.

[0090] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system, characterized in that, Oscillation suppression methods include: A parameterized model of a grid-connected wind turbine system with a series compensation device is constructed. The parameterized model includes a grid-connected grid-side converter, and the control strategy of the grid-connected grid-side converter includes voltage and current dual-loop control. The voltage and current dual-loop control in the grid-side converter is replaced by second-order linear active disturbance rejection control, resulting in a control model for the grid-side converter based on second-order active disturbance rejection. Using the parameterized model as the interaction environment, an agent containing a deep deterministic policy gradient algorithm is constructed and trained, so that the agent can output the optimal active disturbance rejection parameters of the grid-side converter control model. The system status is collected in real time and input into the trained agent to obtain the optimal active disturbance rejection parameters; The optimal active disturbance rejection parameters are applied to the control model of the grid-type grid-side converter to generate a modulation signal to drive the converter, thereby achieving adaptive suppression of subsynchronous oscillations in the grid-type wind turbine grid-connected system containing a series compensation device.

2. The method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system according to claim 1, characterized in that, A second-order active disturbance rejection controller (ADRC) with state error feedback and total disturbance feedforward compensation is constructed based on a linear extended state observer and a linear tracking differentiator.

3. The method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system according to claim 1, characterized in that, The expression for the linearly extended state observer is: In the formula, To output the tracking signal for y, To provide a tracking signal for the first derivative of the output y, The tracking signal for the total system disturbance. , , Let be the time derivative of the corresponding state variable. , , For the gain of the linearly extended state observer, This represents the gain coefficient of the control channel in the original second-order system model. This is the voltage control command for the d-axis port of the grid-side converter.

4. The method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system according to claim 3, characterized in that, The linear tracking differentiator uses a simplified control law to generate the basic control quantity, the expression of which is: In the formula, This is the reference value for the grid connection point voltage on the d-axis. For linear tracking differentiator output, , This is the gain of the linear tracking differentiator.

5. The method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system according to claim 4, characterized in that, Based on the output of the tracking differentiator and the tracking signal of the extended state observer, a state error feedback quantity is generated, and the total disturbance feedforward compensation quantity is superimposed to obtain the output of the second-order active disturbance rejection controller.

6. The method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system according to claim 5, characterized in that, Using the parameterized model as the interaction environment, an Actor-Critic architecture agent is constructed, and the agent is trained using a deep deterministic policy gradient algorithm, specifically including: The parameterized model is used as the environment for agent interaction. The observation state, action variables and reward function of the agent are defined. The observation state is the grid connection point voltage and grid current in the synchronous rotating coordinate system. The action variables are the gain parameters of the linear extended state observer in the second-order linear active disturbance rejection control. Construct an agent comprising a Critic network and an Actor network. The Actor network takes the observed state as input and outputs an action variable for adjusting the active disturbance rejection control parameters. The Critic network takes the observed state and the action variable for adjusting the active disturbance rejection control parameters as input and outputs a state-action evaluation function. The agent is trained using a deep deterministic policy gradient algorithm, enabling it to output the optimal self-disturbance rejection core parameters.

7. The method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system according to claim 6, characterized in that, The expression for the state-action evaluation function is: In the formula, o represents the observed state of the environment. For the action variable of active disturbance rejection control parameter adjustment, As expected, , For the new state and new action variables in the next moment, The main evaluation network parameters, As a discount factor, For the reward function, The new state-action evaluation function for the next moment.

8. The method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system according to claim 7, characterized in that, The Critic network updates its parameters by minimizing a loss function, the expression of which is: In the formula, For storing experience data generated by the interaction between intelligent agents and the environment Experience replay buffer, For the target Actor network parameters, r This is the reward value.

9. A method for suppressing subsynchronous oscillations in grid-connected wind power based on an improved second-order active disturbance rejection system, as described in claim 8, is characterized in that... The reward function is defined as the ratio and integral of the error between the instantaneous active power output by the virtual synchronizer and the active power reference value. The expression for the reward function is as follows: In the formula, The instantaneous active power output by the virtual synchronous machine. This is a reference value for active power. , These are the weighting coefficients.