New energy power generation system stability improvement method based on deep reinforcement learning

Through a method based on deep reinforcement learning, the agent is trained to learn control parameter adjustment strategies from real-time running data, solving the stability problem of new energy power generation systems during the grid connection process, and achieving system stability improvement and rapid response.

CN120127626AActive Publication Date: 2025-06-10SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510174957.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

During the grid connection process, new energy power generation systems are prone to cause stability problems such as sub-synchronous oscillation or over-synchronous oscillation. The existing stability control methods are difficult to adapt to complex and changeable operating conditions and large-scale time-varying scenarios.

Method used

Using a method based on deep reinforcement learning, we use real-time operation data of the new energy power generation system, design reward functions, offline training of deep reinforcement learning agents, obtain control parameters, and adaptively adjust control parameters online to improve system stability.

Benefits of technology

It realizes online adaptive adjustment of controller parameters according to changes in the system operating conditions, improves the stability of the new energy power generation system, avoids the dependence on oscillating component detection links and mechanism cognitiveness, and has fast response and strong adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120127626A_ABST
    Figure CN120127626A_ABST
Patent Text Reader

Abstract

The invention provides a new energy power generation system stability improvement method based on deep reinforcement learning. The method comprises the steps of obtaining real-time operation data of a new energy power generation system; a reward function is designed, the input of the reward function is the real-time operation data and the stability judgment basis of the new energy power generation system, and the output of the reward function is a reward value; training a deep reinforcement learning agent offline by adopting the real-time operation data and the reward value of the new energy power generation system; inputting the real-time operation data into the trained deep reinforcement learning agent to obtain new energy power generation system control parameters; and with the control parameters of the new energy power generation system as the standard, the control parameters of the new energy power generation system are adaptively adjusted on line. According to the invention, the controller parameters can be adaptively adjusted on line according to the system operation condition change, and the stability of the new energy power generation system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of new energy power generation, and in particular, to a method for improving the stability of a new energy power generation system based on deep reinforcement learning. Background Art

[0002] With the continuous growth of the global demand for clean energy and the transformation of the energy structure, renewable energy power generation technology has developed rapidly. However, during the grid connection process of a new energy power generation system, since a power electronic converter is used as the grid connection interface device, stability problems such as sub-synchronous oscillation or super-synchronous oscillation are likely to occur, which seriously affect the effective utilization of new energy and the safe and stable operation of the power grid. Currently, for the stability problems of new energy power generation grid connection, many scholars at home and abroad have proposed various stability control methods, such as active damping control. However, most of these methods are based on fixed damping control structures and parameters, mainly for suppressing oscillations in specific working conditions and specific frequency bands, so it is difficult to adapt to scenarios with complex and variable operating conditions and large-range time-varying oscillation amplitude-frequency. In addition, in recent years, some scholars have also proposed an adaptive oscillation suppression strategy, which suppresses oscillations by online adjusting damping parameters, and these parameters are adjusted according to the detected oscillation components. However, on the one hand, this strategy is limited by the rapid detection time of oscillation components after oscillation occurs, which to a certain extent limits the effect of oscillation suppression; on the other hand, the online adjustment principle of damping parameters is based on existing mechanism cognition, so it is difficult to be applied to "black / gray box" systems. Summary of the Invention

[0003] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method for improving the stability of a new energy power generation system based on deep reinforcement learning.

[0004] According to one aspect of the present invention, there is provided a method for improving the stability of a new energy power generation system based on deep reinforcement learning, including:

[0005] Obtaining real-time operation data of the new energy power generation system;

[0006] Designing a reward function, the input of which is the real-time operation data of the new energy power generation system and the stability evaluation basis, and the output is a reward value;

[0007] Using the real-time operation data of the new energy power generation system and the reward value to offline train a deep reinforcement learning agent;

[0008] Inputting the real-time operation data into the trained deep reinforcement learning agent to obtain control parameters of the new energy power generation system;

[0009] Taking the control parameters of the new energy power generation system as the standard, adaptively adjust the control parameters of the new energy power generation system online.

[0010] Preferably, the new energy power generation system includes a photovoltaic power generation system, a doubly-fed wind power generation system or a direct-drive wind power generation system.

[0011] Preferably, the real-time operation data of the new energy power generation system are data under multiple working conditions, and the multiple working conditions refer to multiple short-circuit ratios and power variations; the real-time operation data include the DC bus voltage, the three-phase voltage and current on the AC side of the converter, the output power, and the errors between the above physical quantities and the given values.

[0012] Preferably, the stability evaluation criterion refers to a parameter that can quantify the stability of the system using numerical indicators, and this parameter refers to the system phase margin;

[0013] The process of obtaining the system phase margin is as follows:

[0014] Establish a broadband impedance model of the new energy grid-connected system;

[0015] Based on the broadband impedance model, use the impedance stability analysis method to calculate the system phase margin under different operating conditions and controller parameters, where the operating conditions include the operating point and the grid strength.

[0016] Preferably, the reward value includes two parts: instantaneous error and stability; among them, the instantaneous error is calculated according to the operation data of the new energy power generation system, and the stability is calculated according to the quantization index obtained from the stability criterion;

[0017] The deep reinforcement learning agent learns by maximizing the reward value.

[0018] Preferably, the deep reinforcement learning agent is a neural network including a Critic value network and an Actor policy network.

[0019] Preferably, using the real-time operation data and the reward value of the new energy power generation system to offline train the deep reinforcement learning agent includes:

[0020] Obtain the real-time operation data, including the DC bus voltage, the three-phase voltage and current on the AC side of the converter, the output active and reactive power and their error values, and input them into the agent to be trained;

[0021] Input the real-time operation data and the phase margin of the current system state into the reward function, calculate the corresponding reward value, and input the reward value into the agent to be trained;

[0022] Training is carried out in an offline environment through a simulation system or an actual system, using a real-time simulation environment and historical experience replay for training. Through multiple rounds of simulation, the neural network parameters of the agent to be trained are continuously corrected to maximize the expected reward value;

[0023] Train until the reward value can reach the set standard.

[0024] Preferably, the training is carried out in an offline environment through a simulation system or an actual system, using a real-time simulation environment and historical experience replay for training. Through multiple rounds of simulation, the neural network parameters of the agent to be trained are continuously corrected to maximize the expected reward value, including:

[0025] Define the loss for each step;

[0026] After defining the loss for each step, the neural network parameters are updated using the gradient descent method or the gradient ascent method, and soft update of the target network is used:

[0027]

[0028] θ′←τθ+(1-τ)θ′

[0029] In the formula, θ is the parameter of the main network of the value network, θ’ is the parameter of the target network of the value network, α is the learning rate, τ is the soft update coefficient, and L(θ) is the loss function, which is defined using the Bellman error as:

[0030]

[0031] In the formula, r is the reward, s is the state, a is the output action, γ is the discount factor; Q is the total return value obtained by network evaluation, indicating the cumulative return that the agent can obtain in the future after taking a certain action a in the given state s, and s', a represent the state and action at the next moment.

[0032] Preferably, a regularization factor λ can also be added to the loss function, specifically:

[0033]

[0034] Preferably, real-time operation data is continuously collected in an online environment for fine-tuning and optimizing the neural network parameters.

[0035] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects:

[0036] The embodiments of the present invention can adaptively adjust the controller parameters online according to the changes in the system operation conditions, improve the stability of the new energy power generation system. This method does not require an oscillation component detection link, and the online parameter adjustment does not depend on mechanism cognition, having the advantages of fast response speed and strong adaptability to working condition changes. Description of the Drawings

[0037] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non - limiting embodiments with reference to the accompanying drawings:

[0038] Figure 1 It is a flowchart of a method for improving the stability of a new - energy power generation system based on deep reinforcement learning in a preferred embodiment of the present invention.

[0039] Figure 2 It is a schematic diagram of the framework structure of a method for improving the stability of a new - energy power generation system based on deep reinforcement learning in a preferred embodiment of the present invention.

[0040] Figure 3 It is a schematic diagram of the structure of a typical wind power generation system in a preferred embodiment of the present invention.

[0041] Figure 4 It is a flowchart of deep reinforcement learning in a preferred embodiment of the present invention.

[0042] Figure 5 It is a schematic diagram of the control structure of a new - energy power generation system combined with a deep - reinforcement - learning agent in a preferred embodiment of the present invention.

[0043] Figure 6 It is a simulation waveform diagram of a new - energy power generation system with conventional control under a step - change of active power in a preferred embodiment of the present invention.

[0044] Figure 7 It is a simulation waveform diagram of a new - energy power generation system based on deep reinforcement learning under a step - change of active power in a preferred embodiment of the present invention. Detailed Description of the Preferred Embodiments

[0045] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0046] In an embodiment of the present invention, a method for improving the stability of a new - energy power generation system based on deep reinforcement learning is provided. Refer to Figure 1 and Figure 2 , and it can adopt the following steps:

[0047] Step 1: Obtain the real - time operation data of the new - energy power generation system;

[0048] Step 2: Design a reward function. The input of this reward function is the real-time operation data of the new energy power generation system and the stability evaluation criterion, and the output is the reward value.

[0049] Step 3: Use the real-time operation data of the new energy power generation system obtained in Step 1 and the reward value obtained in Step 2 to offline train a deep reinforcement learning agent.

[0050] Step 4: Input the real-time operation data into the deep reinforcement learning agent trained in Step 3 to obtain the control parameters of the new energy power generation system.

[0051] Step 5: Take the control parameters of the new energy power generation system obtained in Step 4 as the standard to online adaptively adjust the control parameters of the new energy power generation system.

[0052] In the above embodiment, by training a deep reinforcement learning agent, the agent learns the parameter adjustment strategy that can improve stability, especially in the case of a weak power grid and high-power output, thereby enhancing the stability of the new energy grid-connected system.

[0053] In a preferred embodiment, the real-time operation data of the new energy power generation system is data under various working conditions, that is, data under various short-circuit ratios and power changes. The real-time operation data includes the DC bus voltage, the three-phase voltage and current on the AC side of the converter, the output power, and the errors between the above physical quantities and the given values.

[0054] In a preferred embodiment, the stability evaluation criterion can quantify the stability of the system as a numerical index. The acquisition process is specifically as follows:

[0055] Establish a broadband impedance model of the new energy grid-connected system;

[0056] Use the impedance stability analysis method to calculate the system phase margin under different operating conditions (including different operating points and grid strengths) and controller parameters.

[0057] In the above embodiment, the real-time data used is the operation data under different short-circuit ratios and power changes, and the criterion used is the system phase margin under different operating conditions, so that the trained agent can adapt to the usage requirements of different operating conditions and learn the parameter adjustment strategies to improve stability in various situations.

[0058] In a preferred embodiment, the designed reward function includes two parts: instantaneous error and stability. The instantaneous error is calculated based on the relevant operation data of the new energy power generation system, and the stability is calculated based on the quantified index obtained from the stability criterion. The overall goal is that the smaller the instantaneous error and the closer the phase margin is to the preset value, the higher the reward. The reward is used as an important basis during offline training, and the agent learns by maximizing the obtained reward; while during online application, no reward input is required.

[0059] Further, in a specific embodiment, the reward function adopted is specifically:

[0060] r = r 1 + r 2

[0061] r 1 represents the instantaneous error, and r 2 represents stability.

[0062]

[0063] r 2 = -ω 5 ·|PM - 30°|

[0064] + ω 6 ·S(20° < PM < 60°)

[0065] - ω 6 ·S(PM < 0° or PM > 90°)

[0066] - 2ω 6 ·S(PM < -20° or PM > 110°)

[0067] Among them, represents the instantaneous error of the DC voltage, e Q represents the instantaneous error of the reactive power, U dc,ref represents the DC voltage reference value, P ref represents the active power reference value. Each ω is the weight of each item, PM is the phase margin, and the definition of the S function is:

[0068]

[0069] In order to obtain a more accurate agent, in a preferred embodiment of the present invention, the training method of the agent is a deep reinforcement learning method, including DQN (Deep Q-Network), DDPG (deep deterministic policy gradient), TD3 (Twin Delayed Deep Deterministic policy gradient), PPO (Proximal Policy Optimization) method, etc.; the agent is composed of a neural network, and the training process is to adjust the network parameters based on the training operation data.

[0070] Further, in a preferred embodiment, the training process of the agent specifically includes:

[0071] 1) Obtain real-time operating data, including DC bus voltage, three-phase voltage and current on the AC side of the converter, output active and reactive power and their error values, and input them into the agent to be trained. The real-time operating parameters need to be used as the observations of the agent. By observing these values, the agent can evaluate the current operating conditions, which helps to make corresponding decisions to improve stability.

[0072] 2) Evaluate the stability of the current system state through a pre-designed reward function. The reward function combines stability evaluation criteria such as phase margin, calculates the corresponding reward value, and inputs it into the agent to be trained.

[0073] 3) Training can be carried out in an offline environment through a simulation system or an actual system. Use a real-time simulation environment and historical experience replay for training, and continuously correct the network parameters through multiple rounds of simulation, so as to maximize the expected reward;

[0074] 4) Train until the agent can adaptively adjust the control parameters under different operating conditions to improve the stability of the system. Specifically:

[0075] The reward function reflects the stability of the system, so the goal of training is to maximize the reward function. During training, random initial operating conditions are given in each episode. If the rewards obtained in consecutive multiple training episodes can reach a certain threshold, it means that the adaptive stability improvement effect of the agent under different operating conditions is good enough, and training can be stopped.

[0076] Subsequently, data can be continuously collected in an online environment for fine-tuning and optimization.

[0077] Furthermore, in a preferred embodiment, step 3) is further described. Specifically:

[0078] After defining the loss for each step, the network parameters are updated using the gradient descent method or the gradient ascent method, and soft update of the target network is used. The formula is as follows:

[0079]

[0080] θ′←τθ+(1-τ)θ′

[0081] In the formula, θ is the parameter of the main network of the value network, θ’ is the parameter of the target network of the value network, α is the learning rate, τ is the soft update coefficient, and L(θ) is the loss function. The loss function generally uses the Bellman error definition:

[0082]

[0083] Where r is the reward, s is the state, a is the output action, and γ is the discount factor. The Q-value is the total return value evaluated by the network, similar to the reward function, indicating the cumulative return that the agent can obtain in the future (weighted by the discount factor) after taking a certain action a in the given state s. s', a’ represent the state and action at the next moment. That is, the DDPG method attempts to make the current Q-value (the last term Q) approach a more stable target value (the sum of the first two terms r + γQ) to obtain more stable training.

[0084] For the above loss function, a regularization factor λ can also be additionally added, such as:

[0085]

[0086] Among them, the target network is a copy of the main network and is used to assist in learning. The actor network and the critic network each have a main network and a target network. The target structure is the same as that of the main network, but the parameter update methods are different: the parameters of the main network are directly updated by the gradient descent method. The parameters of the target network are updated from the main network at a slower rate, usually using the "soft update" method in the document. Using the target network can provide a relatively stable reference value for the update of the main network, avoiding situations such as oscillations or divergence (gradient explosion) during the learning process, and is an auxiliary means for training.

[0087] In an embodiment of the present invention, the collected operation data is input into the trained agent. The Actor network of the agent adopts a deep neural network architecture, including multiple fully connected hidden layers and non-linear activation functions (such as ReLU and tanh), aiming to extract state features and generate appropriate control strategies. The input of this network is the real-time operation data of the current system. Linear mapping is performed on the original input through multiple fully connected layers, and non-linear activation functions are used for feature extraction to capture the complex relationships between state variables and ensure that the control strategy can adapt to various operating states. After a series of linear and non-linear mapping processes, the optimized control strategy parameters are finally output through the output layer.

[0088] To verify the feasibility and effectiveness of the new energy power generation system stability improvement method in the above embodiment, a specific embodiment of the present invention takes a typical wind power generation system as an example for verification. As Figure 3 shown, the typical wind power generation system provided by this embodiment may include: a wind turbine, a generator, a converter, and a filter. The converter in the wind power generation system is a back-to-back converter, which consists of a machine-side converter and a grid-side converter. The machine-side converter is responsible for controlling the speed or torque of the generator to achieve maximum wind energy capture; the grid-side converter is responsible for controlling the DC bus voltage and achieving synchronization with the power grid.

[0089] Figure 4 As shown, the flowchart of deep reinforcement learning provided by this embodiment may include the following operations:

[0090] S1. Parameter predefined: Roughly set the initial parameter values according to experience, specifically the main circuit parameters and controller parameters. The main circuit parameters include the rated value of the DC side voltage, the rated value of the AC side voltage, the rated output power, the value of the DC bus capacitor, the value of the filter inductor, etc. The controller parameters include the PI control parameters of the current loop, the DC voltage loop, the phase-locked loop, and the reactive power loop.

[0091] S2. Configure the environment and associate the simulation model:

[0092] First, manually bind the agent variable to the RL Agent module to enable the training function to recognize the intelligent agent. Subsequently, start setting the dimensions and intervals of the observations and actions (i.e., the inputs and outputs of the intelligent agent).

[0093] Second, determine the dimensions of the observations and actions. To avoid overly unreasonable action outputs, set an interval for the actions.

[0094] Furthermore, determine the interval reference values. After selecting the reference values, considering giving each parameter a certain exploration space, set the upper and lower bounds to a certain multiple of the reference values. Considering the respective influences of the two parameters and their combined effects on the bandwidth of each control link, set the interval of the proportional coefficient to 1 / 4 times to 4 times, and set the interval of the integral coefficient to 1 / 8 times to 8 times. Within this parameter selection range, it is basically possible to ensure that the bandwidth of each link fluctuates around 1 / 4 times to 4 times of the reference value, leaving enough exploration space without completely unreasonable action outputs.

[0095] In addition, the environment must include an initialization function to restore the environment to a random initial state before the start of each round of training. In this function, in order to simulate the change of the working point in practice, set the short-circuit ratio to vary between 2 and 6, the grid voltage to vary between 0.95 and 1.05 times the rated voltage, and the output power to vary between 0.2 and 1.0 times the rated power.

[0096] In addition, set various short-circuit ratios and power changes so that the intelligent agent can learn more different working conditions.

[0097] S3. Create and set up the Critic and Actor networks: First, design the structure of the neural network. The network structure is generally determined by two aspects. One is the number of layers of the network, and the other is the type of each layer and the number of neurons it contains. A relatively simple network may not be able to effectively capture the details in the information, resulting in a poor fitting effect of the network; an overly complex network may have excessive data analysis capabilities, so that it can fit the noise in the data, thus causing the problem of overfitting and insufficient generalization ability. Therefore, the network structure needs to be continuously adjusted according to the actual training effect to achieve the optimal. In this embodiment, the Critic network needs to receive information from both observations and actions at the same time. Among them, the information from observations has a higher dimension and richer details. Therefore, a layer of ReLU function is used for processing, and then it is merged with the information from actions. After the two are merged, two layers of ReLU functions and two fully connected layers with 64 neurons are used for processing. Since the Actor network needs to analyze information from observations and derive strategies, it also requires a certain degree of complexity. However, since it finally needs to be deployed in the actual scenario, considering the computational amount and real-time calculation time, it should not be too complex. Therefore, the structure of the Critic network in this embodiment is: Observation input layer - Fully connected layer - ReLU - Fully connected layer - Merging layer (merged with the input from the Actor network) - ReLU - Fully connected - ReLU - Fully connected output layer. The structure of the Actor network is Input layer - Fully connected - ReLU - Fully connected - ReLU - Fully connected - tanh - Output layer. Among them, different network structures (such as different numbers of layers, different arrangements between layers, the number of neurons, the selection of non-linear activation functions, etc.) only affect the training effect and actual performance, but all belong to the critic-actor framework.

[0098] S4. Set the agent options and create the agent: Before training, in order to specify the behavior of the agent during training, certain hyperparameters need to be manually configured and continuously adjusted according to the training results. In this embodiment, the agent hyperparameters include the agent sampling period, discount factor, experience pool length, experience pool sample size, network learning rate, learning gradient threshold, L2 regularization factor, training noise, minimum noise, noise decay rate, target policy smoothing noise, etc. Among them: The discount factor can make the value function converge more easily; the experience pool length needs to be large enough to achieve stable value network updates; the learning gradient threshold can ensure that each learning step size is within an appropriate range to avoid negative impacts such as gradient explosion; the exploration noise needs to be moderate to ensure that the agent learns an effective and not overly aggressive strategy; the L2 regularization factor can limit the coefficients of high-degree polynomials, thus avoiding overfitting of the entire network.

[0099] S5. Set training options: Before training, it is also necessary to set the parameters related to the entire training process. In this embodiment, the training process parameters include the maximum number of episodes per single training, the number of steps per episode, the average reward calculation window length, the stop training criterion, the save agent criterion, whether to output the training results to the console, whether to plot the training results, etc.

[0100] S6. Start training.

[0101] As Figure 5 shown, in the new energy power generation system control structure combined with a deep reinforcement learning agent provided in this embodiment, it includes a deep reinforcement learning agent. Using the operation-related data of the new energy power generation system as input, it outputs the control parameters of the new energy power generation system and updates them to the control system of the new energy power generation system to improve the stability of the system. The agent obtained through deep reinforcement learning observes the environmental state State S i (DC bus voltage, three-phase voltage and current on the AC side of the converter, output active and reactive power and their error values, etc.) and then needs to interact with the environment by selecting actions through a policy. In this embodiment, the agent is implemented by changing the PI parameters (K pdc , K idc , K pi , K ii , K ppll , K ipll ) of the DC voltage loop, current loop, and phase-locked loop, so as to track the preset environmental state and improve stability. Specifically: The input of the Critic network includes the current operation data and the control policy output by the Actor network. Its role is to evaluate the quality of the actions generated by the current actor network in combination with the environmental state and calculate the Q value as the basis for optimizing and training the actor network. This network consists of multiple fully connected layers, and the ReLU activation function is used for nonlinear transformation between the fully connected layers, and finally outputs a scalar Q value. The input of the Actor network is the real-time collected data, and the output is the control policy parameters (including the PI parameters of the DC voltage loop, the PI parameters of the current loop, the PI parameters of the PLL, etc.). This network consists of multiple fully connected layers, which include ReLU and tanh activation functions to extract state features and generate reasonable control policies. Finally, the optimized control policy parameters are output to adjust the performance of the control system in real time.

[0102] As Figure 6As shown, the conventional control provided by this embodiment is the reactive power / direct current voltage outer loop and current inner loop control with fixed parameters. During the simulation process, the initial short-circuit ratio is 4, and the initial output power is 0.2 p.u. Subsequently, at 0.5 s, it is switched to a weak power grid with a short-circuit ratio of 2.7. Subsequently, at 1 s, 2 s, and 3 s respectively, the active power output is increased to 0.6 p.u., 0.8 p.u., and 1.0 p.u. As can be seen from the figure, as the active power increases, the system gradually oscillates and finally diverges and becomes unstable.

[0103] As Figure 7 shown, during the simulation process of the new energy power generation system based on the deep reinforcement learning agent provided by this embodiment, the initial short-circuit ratio is 4, and the initial output power is 0.2 p.u. Subsequently, at 0.5 s, it is switched to a weak power grid with a short-circuit ratio of 2.7. Subsequently, at 1 s, 2 s, and 3 s respectively, the active power output is increased to 0.6 p.u., 0.8 p.u., and 1.0 p.u. As can be seen from the figure, at different active power output levels, the system can maintain stability, and the overshoot is small, which reflects the advantages of the method for improving the stability of the new energy power generation system based on deep reinforcement learning provided by the present invention.

[0104] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention. The above preferred features can be used in any combination without conflict.

Claims

1. A method for improving the stability of a new energy power generation system based on deep reinforcement learning, characterized in that: include: Obtain real-time operation data of new energy power generation systems; Designing a reward function, the input of which is the real-time operation data and stability evaluation basis of the new energy power generation system, and the output of which is a reward value; Using the real-time operation data and the reward value of the new energy power generation system to train a deep reinforcement learning agent offline; Inputting the real-time operation data into the trained deep reinforcement learning agent to obtain the control parameters of the new energy power generation system; Taking the control parameters of the renewable energy power generation system as a standard, the control parameters of the renewable energy power generation system are adjusted online and adaptively.

2. A method for improving the stability of a new energy power generation system according to claim 1, characterized in that: The new energy power generation system includes a photovoltaic power generation system, a double-fed wind power generation system or a direct-drive wind power generation system.

3. A method for improving the stability of a new energy power generation system according to claim 1, characterized in that: The real-time operating data of the new energy power generation system is data under multiple operating conditions, and the multiple operating conditions refer to multiple short-circuit ratios and power changes; the real-time operating data includes the DC bus voltage, the three-phase voltage and current on the AC side of the converter, the output power, and the errors between the above-mentioned physical quantities and the given values.

4. A method for improving the stability of a new energy power generation system according to claim 1, characterized in that: The stability evaluation basis refers to a parameter that can quantify the stability of the system using a numerical indicator, and the parameter refers to the system phase margin; The process of obtaining the system phase margin is as follows: Establish a broadband impedance model for renewable energy grid-connected systems; Based on the broadband impedance model, an impedance stability analysis method is used to calculate the system phase margin under different operating conditions and controller parameters, wherein the operating conditions include operating points and grid strength.

5. A method for improving the stability of a new energy power generation system according to claim 4, characterized in that: The reward value includes two parts: instantaneous error and stability; wherein the instantaneous error is calculated according to the operating data of the new energy power generation system, and the stability is calculated according to the quantitative index obtained by the stability criterion; The deep reinforcement learning agent learns by maximizing the reward value.

6. A method for improving the stability of a new energy power generation system according to claim 1, characterized in that: The deep reinforcement learning agent is a neural network including a Critic value network and an Actor strategy network.

7. A method for improving the stability of a new energy power generation system according to claim 6, characterized in that: The method of using the real-time operation data and the reward value of the new energy power generation system to offline train a deep reinforcement learning agent includes: Acquire the real-time operation data, including DC bus voltage, converter AC three-phase voltage and current, output active and reactive power and their error values, and input them into the intelligent agent to be trained; Inputting the real-time operation data and the phase margin of the current system state into the reward function, calculating a corresponding reward value, and inputting the reward value into the intelligent agent to be trained; Training is performed in an offline environment through a simulation system or an actual system, using a real-time simulation environment and historical experience playback for training. The neural network parameters of the agent to be trained are continuously corrected through multiple rounds of simulation to maximize the expected reward value. Train until the reward value reaches the set standard.

8. A method for improving the stability of a new energy power generation system according to claim 7, characterized in that: The training is performed in an offline environment through a simulation system or an actual system, using a real-time simulation environment and historical experience playback for training, and continuously correcting the neural network parameters of the agent to be trained through multiple rounds of simulation to maximize the expected reward value, including: Define the loss at each step; The neural network parameters are updated using the gradient descent or gradient ascent method after defining the loss at each step, and the target network is soft updated: θ′←τθ+(1-τ)θ′ Where θ is the parameter of the value network main network, θ' is the parameter of the value network target network, α is the learning rate, τ is the soft update coefficient, and L(θ) is the loss function, which is defined using Bellman error as: In the formula, r is the reward, s is the state, a is the output action, and γ is the discount factor; Q is the total return value obtained by network evaluation, which means the cumulative return that the agent can obtain in the future after taking an action a in a given state s, and s' and a' represent the state and action at the next moment.

9. A method for improving the stability of a new energy power generation system according to claim 8, characterized in that: The loss function further adds a regularization factor λ, specifically:

10. A method for improving the stability of a new energy power generation system according to claim 9, characterized in that: Real-time operation data continues to be collected in the online environment to fine-tune and optimize the neural network parameters.

Citation Information

Patent Citations

  • Power distribution network multi-region cooperative reactive power optimization method based on multi-agent reinforcement learning

    CN115483703A

  • Power distribution network operation strategy intelligent generation method and device based on reinforcement learning

    CN116523327A

  • Flexible DC power grid fault intelligent recovery method based on data-physical fusion

    CN117117824A

  • Spectral programmable optical frequency comb generation method based on deep reinforcement learning

    CN117192865A

  • Grid-connected converter dynamic simulation method based on time-varying impedance characteristics of new energy station

    CN118199146A

Cited By

  • Motor driving method and device, computer equipment and storage medium

    CN121193154A