Power system transient stability control method and device based on active low voltage ride through of wind turbine generator
By constructing an active low-voltage ride-through control system for wind turbines based on a reinforcement learning network with a proximity optimization strategy, the transient stability problem caused by the reduction of power system inertia due to wind turbine connection was solved. This achieved fast response and optimized transient stability control, improving the safety and stability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEAST DIANLI UNIVERSITY
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-08
AI Technical Summary
The integration of wind turbines reduces the inertia of the power system, threatening the system's transient stability. Traditional low-voltage ride-through technology has a slow response speed and poor flexibility, making it difficult to effectively cope with the transient instability of the power system.
An active low-voltage ride-through control system for wind turbines is constructed using a reinforcement learning network based on a proximity optimization strategy. This system consists of an active low-voltage ride-through module, an active and reactive power optimization control module, an initialization module, an interaction module, a data processing module, a decision-making module, and a learning module. The reinforcement learning network is then used to optimize the control strategy of the wind turbines and enable them to respond quickly to power system disturbances.
It improves the flexibility of low voltage ride-through of wind turbine units and the transient stability of the system, and can quickly provide transient stability control measures when the wind turbine grid-connected system is disturbed, thereby improving the safety and stability of the system.
Smart Images

Figure CN122000895A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system transient stability control technology, and more specifically, to a power system transient stability control system and method based on active low voltage ride-through of wind turbine generators. Background Technology
[0002] In recent years, my country has actively and steadily promoted carbon peaking and carbon neutrality, and steadily advanced the green and low-carbon transformation of energy, placing greater emphasis on promoting the development of new and clean energy. In traditional power systems dominated by centralized thermal and hydropower, the transient stability of the power grid mainly relies on the rotor inertia and regulation capabilities of synchronous generators. However, the integration of wind turbines has altered the dynamic characteristics of the power system. In particular, the introduction of doubly-fed induction generators and full-power wind turbines has significantly reduced system inertia, thus posing a new threat to system transient stability. Therefore, transient stability control of systems containing wind turbines is of great significance.
[0003] The low-voltage ride-through capability of wind turbines can change the power balance of the system, thus serving as a means of transient stability control. However, traditional low-voltage ride-through is affected by the degree of voltage drop and the extent of rotor current exceeding the limit, resulting in slow response speed and poor flexibility. Therefore, actively controlling the low-voltage ride-through of wind turbines based on the transient instability of the system and utilizing the wind turbine's own fault protection behavior for power system transient stability control can provide more control means for the transient stability control of power systems connected to the grid.
[0004] In recent years, many wind farms, both domestically and internationally, have successfully addressed system instability issues caused by wind turbine disconnection through improved low-voltage ride-through (LVRT) technology in practical engineering projects. Therefore, in-depth research into the LVRT transient stability of wind turbine-integrated systems not only helps improve the safety and stability of the existing power grid but also lays a solid foundation for the further development of new energy power generation. This research is crucial in practical engineering applications, directly impacting the reliability and sustainability of future large-scale wind power grid integration.
[0005] Application content
[0006] To address the aforementioned problems, this invention proposes a power system transient stability control device based on active low-voltage ride-through of wind turbines, comprising:
[0007] The system includes an active low-voltage ride-through module, an active and reactive power optimization control module, an initialization module, an interaction module, a data processing module, a decision-making module, and a learning module.
[0008] Furthermore, the active low-voltage ride-through module is used to control the active switching of the wind turbine crowbar circuit and the active locking of the rotor converter.
[0009] The active and reactive power optimization control module is used to optimize the low voltage ride-through process, actively control the active and reactive power of the wind turbine, control the active power output of the system, and adjust the system voltage level affected by disturbances.
[0010] The initialization module is used to configure the network parameters of the reinforcement learning network based on the proximity optimization strategy algorithm, set the maximum number of interactions in each cycle between interactions and the number of cycles to be trained, and read the preset set of low voltage ride-through termination instructions and the set of characteristic electrical quantities of the power system.
[0011] The interaction module is used to perform the following operations: after the power system runs for one interaction interval step, it reads the characteristic electrical quantity data once; if the reinforcement learning network outputs an instruction to end low voltage ride-through, it transmits the instruction to the power system.
[0012] The data processing module is used to classify the characteristic electrical quantities read by the interaction module into the following three categories: decision-making measure electrical quantities, control effect electrical quantities, and safety constraint electrical quantities;
[0013] The evaluation module is used to obtain a reward value based on the voltage amplitude and phase angle of each node in the system using a reward function;
[0014] The decision module is used to input the electrical quantity of the decision measure into the reinforcement learning network as input data, and to output the active low voltage ride-through end control command of the wind turbine as the reinforcement learning network, so that the wind turbine can provide corresponding active low voltage ride-through measures when the power system is disturbed, so that the system transient stability can be restored to a safe range.
[0015] The learning module is used to judge the system transient stability recovery effect based on the control effect electrical quantity and to judge whether the power system safety constraint is triggered based on the safety constraint electrical quantity. On this basis, it updates the network parameters of the reinforcement learning network in combination with the reward value obtained by the evaluation module.
[0016] Furthermore, the reinforcement learning network based on the proximity optimization strategy algorithm comprises two neural networks: a policy neural network and a value neural network. The input of the policy neural network is the real-time node voltage amplitude and phase angle, and the output is the active low-voltage ride-through termination control command for the wind turbine. The input of the value neural network is the system stability index sBTTC value determined after executing the policy, and the output is the neural network weights used to update the policy neural network and the value neural network. The characteristic electrical quantities include the node voltage amplitude and phase angle at several time points within the most recent interaction interval, the synchronous machine power angle difference, the active power and reactive power of each unit, the system bus voltage threshold, and the low-voltage ride-through time threshold. The network parameters of the reinforcement learning network based on the proximity optimization strategy algorithm include the learning rate, batch size, gradient pruning size, and discount factor size. The set of active low-voltage ride-through termination control commands for the wind turbine is constructed from the crowbar circuit disconnection command and the rotor converter reconnection command.
[0017] Furthermore, the interaction interval is the time interval at which the reinforcement learning network interacts with the power system, and each interaction interval is set to 1 second;
[0018] Furthermore, the decision-making electrical quantities include real-time node voltage amplitude and phase angle, which are used as input values for the reinforcement learning network; the control effect electrical quantities include the synchronous machine power angle difference, and the control effect of the transient stability control command given by the reinforcement learning network in the previous interaction interval is judged by the degree of reduction of the power angle difference; the safety constraint electrical quantities, namely the low-voltage ride-through time threshold, include the active power, reactive power, system bus voltage threshold, and low-voltage ride-through time threshold of each unit.
[0019] Furthermore, the input to the reinforcement learning network in the decision module is the electrical quantity of the decision measure, and the output is the active low-voltage ride-through termination control command of the wind turbine. The active low-voltage ride-through termination command includes the active disconnection of the crowbar circuit and the active reconnection of the rotor converter.
[0020] Furthermore, the reward function is set as follows:
[0021] If the sBTTC index value is greater than 0.6, the reward function is set as follows: BTTC i ×1000;
[0022] If the sBTTC index value is in the range of [0.2, 0.6], the reward is set as a linear function:
[0023] (BTTC i -0.2)×1000;
[0024] If the sBTTC index value is < 0.2, the reward function is set as follows: BTTC i ×(-1000).
[0025] To address the above problems, this invention also proposes a power system transient stability control method based on active low-voltage ride-through of wind turbines, comprising:
[0026] Step 1: Construct a power system model including active low-voltage ride-through wind turbines;
[0027] The active low-voltage ride-through capability of the wind turbine is achieved through two controls: hardware control (active switching of the crowbar circuit and active interlocking of the rotor converter) and software control (adjusting the active and reactive currents of the wind turbine current loop to optimize the low-voltage ride-through process). The hardware and software protection work together to complete the active low-voltage ride-through control of the wind turbine and integrate the wind turbine model into the power system.
[0028] Step 2: Extract the characteristic electrical quantities of multiple historical moments of the power system node model as state values in reinforcement learning to construct observation data. The characteristic electrical quantities include the voltage amplitude and phase angle of each node at several time points within one or more recent interaction intervals, the power angle difference of the synchronizing machine, the active power and reactive power of each unit, the system bus voltage threshold, and the low voltage ride-through time threshold.
[0029] Step 3: Construct a reinforcement learning network based on the nearest neighbor optimization algorithm.
[0030] The reinforcement learning network based on the nearest neighbor optimization strategy algorithm includes two neural networks: a policy neural network and a value neural network. The input of the policy neural network is the real-time node voltage amplitude and phase angle, and the output is the active low voltage ride-through termination control command of the wind turbine. The input of the value neural network is the system stability index sBTTC value after executing the policy, and the output is the neural network weights used to update the policy neural network and the value neural network.
[0031] Step 4: Analyze and classify the extracted characteristic electrical quantities;
[0032] The characteristic electrical quantities are divided into decision-making electrical quantities, control effect electrical quantities, and safety constraint electrical quantities. The decision-making electrical quantities include real-time node voltage amplitude and phase angle, which are used as input values for the reinforcement learning network. The control effect electrical quantities include the synchronous machine power angle difference, and the degree of reduction in the power angle difference is used to judge the control effect of the wind turbine active low-voltage ride-through termination command given by the reinforcement learning network in the previous interaction interval. The safety constraint electrical quantities include the active power, reactive power, system bus voltage threshold, and low-voltage ride-through time threshold of each unit.
[0033] Step 5: Use the constructed reinforcement learning network based on the nearest neighbor optimization strategy algorithm to optimize and train various feature electrical quantity data to generate a strategy model;
[0034] Step 6: Extract the real-time electrical quantities of the power system for decision-making measures, use them as input to the strategy model, and output the active low-voltage cutoff command for the wind turbine. Send the output command to the power system for execution, so that the grid-connected system including the wind turbine can provide corresponding transient stability control measures online for different operating conditions.
[0035] Furthermore, in step five, the strategy model refers to the strategy neural network structure and parameters in the trained reinforcement learning network based on the proximity optimization strategy algorithm. The input layer is the electrical quantity of the decision measure, and the output is the wind turbine active low voltage ride-through termination command.
[0036] Through the above design scheme, this invention can bring the following beneficial effects: This invention uses a reinforcement learning model based on the nearest neighbor policy optimization algorithm as the decision-making subject, and the actual power system as the environment. It extracts characteristic electrical quantities from multiple historical moments of the power system node model as state values in reinforcement learning to construct observation data; it constructs a reinforcement learning network based on the nearest neighbor policy optimization algorithm; it analyzes and classifies the electrical quantities, using a portion as input to the reinforcement learning network and another portion as output to update the network parameters. The transient stability control measures are then used as the output of the reinforcement learning network, and optimization training is conducted through reinforcement learning to generate a policy model. By extracting the policy model, the active low-voltage control system for wind turbines can quickly provide transient stability control measures when the grid-connected system containing wind turbines is disturbed, ensuring that the system's transient stability reaches a safe range. This invention improves the transient stability of current grid-connected power systems containing wind turbines under complex operating conditions.
[0037] Beneficial effects
[0038] This invention improves the flexibility of low-voltage ride-through of wind turbines and the transient stability of the wind turbine grid-connected power system when disturbances occur. It uses reinforcement learning to determine the duration of active low-voltage ride-through that best improves the transient stability of the system, and obtains the optimal control strategy for improving the transient stability of the system based on active low-voltage ride-through of wind turbines, thereby improving the transient stability of the system with wind turbines under large disturbances.
[0039] Those skilled in the art should understand that the above embodiments are merely illustrative of the specific content of this disclosure and do not limit its scope. System capacity, voltage, line parameters, etc., will vary depending on the specific circumstances of the power electronic grid-connected generator set and its grid connection. Based on this disclosure, those skilled in the art can make other changes or adjustments, and these changes still fall within the scope of this disclosure. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings involved in the embodiments or the prior art are briefly described below. Obviously, these drawings illustrate several embodiments of the present invention, and those skilled in the art can derive other possible drawings based on these drawings without creative effort. The purpose of the drawings is limited to illustrating specific embodiments and does not limit the scope of the present invention.
[0041] Figure 1 This is a flowchart of the power system transient stability control method based on active low voltage ride-through of wind turbines proposed in this invention;
[0042] Figure 2 This is a schematic diagram of the power system transient stability control device based on active low voltage ride-through of wind turbines proposed in this invention;
[0043] Figure 3 This is a computational example of an improved 3-machine 9-node system used for demonstration purposes;
[0044] Figure 4 This is a block diagram of the inner loop current control of a doubly fed wind turbine with an active and reactive power optimization control module. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0046] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, this invention is not limited to the following embodiments, and specific implementation methods can be determined according to the technical solutions of this invention and actual circumstances. To avoid obscuring the essence of this invention, well-known methods, processes, flows, components, and circuits are not described in detail.
[0047] The present invention proposes a power system transient stability control method based on active low voltage ride-through of wind turbines, comprising:
[0048] 1. Construct an improved three-machine nine-node model of a doubly-fed wind turbine with active low-voltage ride-through capability, and integrate the doubly-fed wind turbine capable of active low-voltage ride-through into the three-machine nine-node system;
[0049] 2. Extract the characteristic electrical quantities of multiple historical moments of the power system node model as state values in reinforcement learning to construct observation data. The characteristic electrical quantities include the voltage amplitude and phase angle of each node at several time points within one or more recent interaction intervals, the power angle difference of the synchronizing machine, the active power and reactive power of each unit, the system bus voltage, and the low voltage ride-through time threshold.
[0050] 3. Construct a reinforcement learning network to output a set of instructions based on the equipment in the power system that participates in the active low-voltage ride-through transient stability control of wind turbine units;
[0051] 4. Construct a reinforcement learning network based on the nearest neighbor optimization policy algorithm. The network structure of the reinforcement learning network consists of two neural networks: a policy network and a value network.
[0052] 5. Analyze the extracted characteristic electrical quantities and calculate the stability index sBTTCi value of the key branch. The formula is as follows:
[0053]
[0054] In the formula: i is the number of critical branches; m and n are the nodes at both ends of branch i; The voltage factor for branch i;
[0055] For the i-th branch phase factor; Let Δθ be the voltage across node i of branch i. i The phase angle difference of branch i;
[0056] 6. Calculate the reward value based on the key branch stability indicator sBTTC value. If the sBTTC value > 0.6, the reward function is set to BTTC. i ×1000; If the sBTTC index value is in the range of [0.2, 0.6], the reward is set as a linear function (BTTC). i -0.2)×1000; If the sBTTC index value < 0.2, the reward function is set to: BTTC i ×(-1000);
[0057] The formula for calculating the reward value is as follows:
[0058]
[0059] Different reward functions are set for different transient stability improvement effects, so that the larger the sBTTC index value, the larger the reward function and the better the transient stability control.
[0060] 7. Strategy
[0061] A policy is a mapping from state to action, which refers to a distribution on the action set given a state, that is, assigning an action probability to each state s.
[0062] 8. At the beginning, the power system is in a certain initial state. The reinforcement learning network of the dispatch center system issues actions to the power grid according to the policy distribution, interacts with the environment, changes the environmental state, and feeds it back to the dispatch center system as the state for the next decision stage. The reward is calculated and this process is repeated until the last decision stage.
[0063] The above process is solved using a deep reinforcement learning algorithm to obtain the optimal system transient stability control strategy based on active low voltage ride-through of wind turbines.
[0064] 9. The deep reinforcement learning algorithm used is the nearest neighbor optimization algorithm, which includes a policy neural network and a value neural network;
[0065] The input to the policy neural network is state s t The output is the mean and standard deviation of the normal distribution of actions, which is also the policy distribution. Then, action a is obtained by sampling. t The total reward function is:
[0066]
[0067] Where t represents the t-th interaction between reinforcement learning and the power system; θ represents the parameters of the policy neural network; The parameters of the policy neural network before the update are: T = r, ... t (θ) represents the state s in the old and new strategies. t Next action a t The probability ratio of being selected; λ represents the degree of guidance in early learning values, implying a trade-off between more bias and more variance; Q(s) t ,a t ) represents the actual sampling discount reward, indicating the reward in state s. t Next, execute action a. t Value; V(s) t ) represents the fitted discount reward, indicating state s t The value of can also be expressed in state s. t The average value of performing all actions; V(s) T ) represents the discounted reward for one cycle; is the discount factor, used in reinforcement learning to adjust the near-term and long-term effects, ranging from [0,1]; is the potential function, representing the advantage of the current action compared to the average action; r tε is the reward value at time t; ε is the gradient clipping degree, usually taken as 0.2. clip means that the KL divergence before and after the policy update is controlled between 1-ε and 1+ε, and gradients outside this range are directly ignored.
[0068] The input to the value neural network is state s t The output is the neural network weights used to update the policy neural network and the value neural network;
[0069] The loss function L(θ) of the network is evaluated as follows:
[0070]
[0071] z = r t +σV(s t+1 )
[0072] Where δ is the TD error, and the evaluation network updates its parameters by minimizing the TD error; z is the cumulative reward after discount; E represents the expected value; V(s) t ) represents the fitting discount reward.
[0073] 10. The transient stability control process for active low-voltage ride-through of wind turbines based on the nearest neighbor optimization strategy algorithm includes the following steps:
[0074] Step 1: Initialize neural network weights and biases; initialize parameters such as learning rate, batch size, gradient clipping size ε, discount factor size γ, environment initialization, and set the number of training interaction steps to 0;
[0075] Step 2: Read the observed state s at time t t This includes: the voltage magnitude and phase angle of each node at several time points within one or more recent interaction intervals;
[0076] Step 3: Input the observation data into the policy neural network. The policy neural network outputs the corresponding policy, that is, the action distribution. Sample the distribution to obtain the active low-altitude crossing end command.
[0077] Step 4: Apply the active low-voltage crossing termination command to the real-time power system from time t to t+1. After the action interacts with the environment, the environment is updated, and the observed state s at time t+1 is obtained. t+1 The instant reward r is calculated according to formulas (1)-(2). t ;
[0078] Step 5: Store s t a t r t Update state observations s t =s t +1;
[0079] Step 6: Update time t = t + 1, and repeat steps 2 to 5 until the specified number of interaction steps is reached;
[0080] Step 7: Set the observation state s t+1 The input is fed into a value neural network, which outputs the fitted discounted reward V(s). t According to the reward r stored in step 5 t According to Q(s) in formula (5) t ,a t )) Calculate the cumulative discount reward for each time point;
[0081] Step 8: Store the state s of each interaction. t Action a t Discount Rewards Q(s) t ,a t This data is used to form a batch, and the strategy neural network and value neural network are updated using this batch of data. Update steps:
[0082] ① Calculate the advantage function. Consider the states s within the batch. t Input is fed into a value neural network, and the value neural network...
[0083] Output the V(s) of this batch. t According to formula (5) and the batch Q(s) t ,a t ), calculate the dominance function for each state within a batch;
[0084] ② Update the policy neural network: According to formula (3) and the batch data state s t Action a t The policy neural network needs to minimize the loss function, so the negative of the objective function is used as the loss function and passed back to update the parameters of the policy neural network.
[0085] ③ Update the value neural network: According to formula (6) and the batch data state s t Discount Rewards Q(s) t ,a t Calculate the loss function L(θ) and backpropagate it to update the parameters of the value neural network;
[0086] Step 9: Increment the number of interactions by one, and repeat steps 2 to 8 until the specified number of interactions is reached, then stop training;
[0087] Step 10: Save the strategy and value neural network model, and perform tests, saving the test data.
[0088] 11. When a disturbance occurs in the power system, the nodal voltage magnitude and phase angle at several time points within one or more recent interaction intervals are extracted as model inputs. This allows for the output of the optimal transient stability strategy for the wind turbine active low-voltage ride-through control system under that state. For example... Figure 1 As shown, the specific implementation process of a power system transient stability control method and device based on active low-voltage ride-through of wind turbines is as follows:
[0089] S1. Execute the initialization module, configure the reinforcement learning parameters based on the proximity policy optimization algorithm; set the maximum number of interactions in each cycle between interactions and the number of training cycles, and read the preset wind turbine active low voltage ride-through end time instruction set and the set of electrical quantity data that need to be observed.
[0090] S2. Execute the interaction module. After the power system runs for one interaction interval step, read the electrical quantity data once. If the reinforcement learning network outputs the wind turbine active low voltage ride-through end control command, then the command is transmitted to the power system.
[0091] S3. Execute the data processing module, and use reinforcement learning to classify the electrical quantity data obtained in step S2 into control effect electrical quantities, decision-making measure electrical quantities, and safety constraint electrical quantities.
[0092] S4. Execute the evaluation module to obtain the reward value through formula (1)-(2);
[0093] S5, the execution decision module, takes the voltage amplitude and phase angle of each node as the input of the reinforcement learning network and gets the active low-voltage cutoff command of the wind turbine as the output.
[0094] S6. Execute the learning module. The reinforcement learning network updates its own neural network parameters according to the reward value obtained in step S4 and the formulas (3)-(7). If the reward value is high, the probability of the wind turbine active low-voltage termination command given in step S5 under the power system operating condition in step S2 will be increased, and vice versa.
[0095] S7. Determine if the maximum number of interactions for a single loop has been reached. If not, repeat steps S2 to S6; otherwise, end the current loop.
[0096] S8. Determine whether the required number of iterations has been reached. If not, return to step S1; otherwise, automatically save the trained model and exit the program.
[0097] S9. The application phase involves calling the trained reinforcement learning network and repeating steps S1 to S5 to ensure the transient stability of the power system.
[0098] Example demonstration:
[0099] To demonstrate the effectiveness of the invention, a structure was constructed as follows: Figure 2 The improved 3-machine 9-node power system shown has the following invention-related settings:
[0100] Active low voltage ride-through equipment: Doubly fed wind turbine
[0101] Number of active low-voltage ride-through fans: 30 units;
[0102] Training fault scenario: A short circuit occurs on bus 8, the wind turbine actively drives through the low-voltage circuit, the short circuit fault is cleared after 0.15 seconds;
[0103] Time of failure: 0.15 seconds;
[0104] Maximum number of interactions per loop: 60;
[0105] Interaction interval: 1 second;
[0106] Observed electrical quantities: voltage magnitude and phase angle at each node;
[0107] Training loop count: 20000;
[0108] Reinforcement learning parameters: learning rate 0.000636, batch size 256, gradient clipping size 0.2, discount factor 0.9.
[0109] The reward function is set as follows:
[0110]
[0111] In the formula: sBTTCi is the stability index value of the key branch. Different reward functions are set for different transient stability improvement effects, so that the larger the sBTTC index value, the larger the reward function, and the better the transient stability control.
[0112] By utilizing reinforcement learning to make transient stability control decisions for active low-voltage ride-through of wind turbines when power systems experience disturbances, this invention demonstrates a power system based on active low-voltage ride-through of wind turbines.
[0113] The effectiveness of transient stability control methods and devices for transient stability control of grid-connected systems with wind turbines.
[0114] Those skilled in the art should understand that the above embodiments are merely illustrative of the content of this disclosure and do not limit its scope. The system capacity, voltage, line parameters, etc., shown may vary depending on the specific circumstances of the power electronic grid-connected generator set and its grid connection. Based on this disclosure, those skilled in the art can make other changes or adjustments, and these changes still fall within the scope of this disclosure.
Claims
1. A power system transient stability control device based on active low-voltage ride-through of wind turbine generators, characterized in that, include: The system includes an active low-voltage ride-through module, an active and reactive power optimization control module, an initialization module, an interaction module, a data processing module, a decision-making module, and a learning module.
2. The power system transient stability control device based on active low-voltage ride-through of wind turbine generators according to claim 1, characterized in that: The active low-voltage ride-through module is used to control the active switching of the wind turbine crowbar circuit and the active locking of the rotor converter. The active and reactive power optimization control module is used to optimize the low voltage ride-through process, actively control the active and reactive power of the wind turbine, control the active power output of the system, and adjust the system voltage level affected by disturbances. The initialization module is used to configure the network parameters of the reinforcement learning network based on the proximity optimization strategy algorithm, set the maximum number of interactions in each cycle between interactions and the number of cycles to be trained, and read the preset set of low voltage ride-through termination instructions and the set of characteristic electrical quantities of the power system. The interaction module is used to perform the following operations: after the power system runs for one interaction interval step, it reads the characteristic electrical quantity data once; if the reinforcement learning network outputs an instruction to end low voltage ride-through, it transmits the instruction to the power system. The data processing module is used to classify the characteristic electrical quantities read by the interaction module into the following three categories: decision-making measure electrical quantities, control effect electrical quantities, and safety constraint electrical quantities; The evaluation module is used to obtain a reward value based on the voltage amplitude and phase angle of each node in the system using a reward function; The decision module is used to input the electrical quantity of the decision measure into the reinforcement learning network as input data, and to output the active low voltage ride-through end control command of the wind turbine as the reinforcement learning network, so that the wind turbine can provide corresponding active low voltage ride-through measures when the power system is disturbed, so that the system transient stability can be restored to a safe range. The learning module is used to judge the system transient stability recovery effect based on the control effect electrical quantity and to judge whether the power system safety constraint is triggered based on the safety constraint electrical quantity. On this basis, it updates the network parameters of the reinforcement learning network in combination with the reward value obtained by the evaluation module.
3. The power system transient stability control device based on active low-voltage ride-through of wind turbines according to claim 2, characterized in that: The reinforcement learning network based on the proximity optimization strategy algorithm comprises two neural networks: a policy neural network and a value neural network. The input of the policy neural network is the real-time node voltage amplitude and phase angle, and the output is the active low-voltage ride-through termination control command for the wind turbine. The input of the value neural network is the system stability index sBTTC value determined after executing the policy, and the output is the neural network weights used to update the policy neural network and the value neural network. The characteristic electrical quantities include the node voltage amplitude and phase angle at several time points within the most recent interaction interval, the synchronous machine power angle difference, the active power and reactive power of each unit, the system bus voltage threshold, and the low-voltage ride-through time threshold. The network parameters of the reinforcement learning network based on the proximity optimization strategy algorithm include the learning rate, batch size, gradient pruning size, and discount factor size. The set of active low-voltage ride-through termination control commands for the wind turbine is constructed from the crowbar circuit disconnection command and the rotor converter reconnection command.
4. The power system transient stability control device based on active low-voltage ride-through of wind turbine generators according to claim 3, characterized in that: The interaction interval is the time at which the reinforcement learning network interacts with the power system, and each interaction interval is set to 1 second.
5. The power system transient stability control device based on active low-voltage ride-through of wind turbine generators according to claim 2, characterized in that: The decision-making electrical quantities include real-time node voltage amplitude and phase angle, which are used as input values for the reinforcement learning network. The control effect electrical quantities include the synchronous machine power angle difference, and the control effect of the transient stability control command given by the reinforcement learning network in the previous interaction interval is judged by the degree of reduction of the power angle difference. The safety constraint electrical quantities, namely the low-voltage ride-through threshold, include the active power, reactive power, system bus voltage threshold, and low-voltage ride-through time threshold of each unit.
6. The power system transient stability control device based on active low-voltage ride-through of wind turbine generators according to claim 2, characterized in that: The input to the reinforcement learning network in the decision module is the electrical quantity of the decision measure, and the output is the active low voltage ride-through termination control command of the wind turbine. The active low voltage ride-through termination command includes the active disconnection of the crowbar circuit and the active reconnection of the rotor converter.
7. The power system transient stability control device based on active low-voltage ride-through of wind turbine generators according to claim 2, characterized in that: The reward function is set as follows: If the sBTTC index value is greater than 0.6, the reward function is set as follows: BTTC i ×1000; If the sBTTC index value is in the range of [0.2, 0.6], the reward is set as a linear function: (BTTC i -0.2)×1000; If the sBTTC index value is < 0.2, the reward function is set as follows: BTTC i ×(-1000).
8. A method for using the power system transient stability control device based on active low-voltage ride-through of wind turbines according to any one of claims 1 to 7, characterized in that, include: Step 1: Build a power system model with active low-voltage ride-through wind turbines. The active low-voltage ride-through capability of the wind turbines is achieved through hardware control and software control. The hardware control includes active switching of the crowbar circuit and active locking of the rotor converter. The software control includes optimizing the active and reactive currents for the low-voltage ride-through process. Step 2: Extract the characteristic electrical quantities of multiple historical moments of the power system node model as state values in reinforcement learning to construct observation data. The characteristic electrical quantities of multiple historical moments of the power system node model include node voltage amplitude and phase angle at several time points within the interaction interval, synchronous machine power angle difference, unit active power, reactive power, system bus voltage threshold, and low voltage ride-through time threshold. Step 3: Construct a reinforcement learning network based on the proximity policy optimization algorithm. The reinforcement learning network based on the proximity policy optimization algorithm includes a policy neural network and a value neural network. The input of the policy neural network is the real-time node voltage amplitude and phase angle, and the output is the system transient stability control command after the wind turbine active low voltage ride-through ends. The input of the value neural network is the system stability index sBTTC value judged after the policy is executed, and the output is the neural network weights used to update the policy neural network and the value neural network. Step four involves analyzing and classifying the characteristic electrical quantities. These quantities are categorized into decision-making electrical quantities and control-effect electrical quantities, including power angle difference reduction and safety constraint electrical quantities. The decision-making electrical quantities include real-time node voltage amplitude and phase angle, which are used as input values for the reinforcement learning network. The control-effect electrical quantities include the synchronous machine power angle difference; the degree of power angle difference reduction is used to determine the control effect of the transient stability control commands given by the reinforcement learning network within the interaction interval. The safety constraint electrical quantities include low-voltage ride-through time thresholds, including the active power, reactive power, system bus voltage threshold, and low-voltage ride-through time threshold for each unit. Step 5: Use the reinforcement learning network based on the nearest neighbor optimization strategy algorithm to optimize and train the feature electrical quantity data to generate a strategy model; Step 6: Extract the real-time electrical quantities of the power system for decision-making measures, use them as input to the strategy model, and output the active low-voltage ride-through termination control command for the wind turbine. Send the output command to the power system for execution.
9. A power system transient stability control method based on active low-voltage ride-through of wind turbine generators according to claim 8, characterized in that: In step five, the strategy model refers to the strategy neural network structure and parameters in the trained reinforcement learning network based on the nearest neighbor optimization strategy algorithm. The input layer is the electrical quantity of the decision measure, and the output is the active low voltage ride-through termination control command of the wind turbine.