A Fast Frequency Control Method and System for New Energy Power Systems
Through the improved dual-layer control method combined with the voltage source converter, the frequency stability problem in the new energy power system is solved, fast response and high-precision frequency control are achieved, and complex dynamic operating conditions are adapted to the stability and reliability of the system are improved.
Patent Information
- Application Number
- CN202510361315.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The reduction in inertia and damping levels in new energy power systems leads to frequency stability problems. Traditional frequency control methods are slow to respond and have low control accuracy, making it difficult to adapt to complex dynamic working conditions.
The improved TD3 reinforcement learning algorithm is used to combine with the voltage source converter (VSC) for dual-layer control, and the set output power is generated through outer layer supervision control, and the active and reactive control is carried out in combination with sag control and power strategy. The modulated voltage reference signal is generated to control VSC, achieving rapid frequency adjustment.
It improves the response speed and accuracy of frequency control, can quickly adapt to the complex dynamic operating conditions of new energy power systems, and improves the stability and reliability of frequency control.
Smart Images

Figure CN119891269B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and specifically relates to a fast frequency control method and system for a new energy power system. Background Art
[0002] Since new energy power generation (such as wind energy, photovoltaic, etc.) is usually connected to the power grid through power electronic devices, this transformation has led to a significant reduction in system inertia and damping. Traditional frequency control strategies that rely on the inertial characteristics of synchronous generators show obvious limitations in such a low-inertia power grid, including slow response speed, low control accuracy, and inability to adapt to fast dynamic disturbances. These limitations pose a major challenge to the grid frequency stability. In the context of grid frequency stability, frequency control strategies are mainly divided into primary frequency regulation, secondary frequency regulation, and ancillary service control. However, traditional primary and secondary frequency regulation methods have a slow response speed and are difficult to effectively maintain frequency stability in a power grid with a high new energy penetration rate. Especially when the power grid is disturbed, the frequency response of traditional control methods often has significant overshoot and a long recovery time.
[0003] In recent years, the frequency control technology based on the Voltage Source Converter (VSC) has received extensive attention. The voltage source converter can provide frequency support for a low-inertia power grid through its fast dynamic power regulation ability. However, how to optimize the control strategy of the voltage source converter to fully utilize its response speed and at the same time adapt to complex dynamic working conditions is still an urgent problem to be solved. Reinforcement Learning (RL) does not depend on the accurate model of the system, can learn in an unknown environment, and dynamically adjust the strategy according to the environmental changes. It is suitable for complex and high-dimensional state spaces. Reinforcement learning provides new ideas and methods for dealing with the control of complex power systems. For example, existing deep reinforcement learning algorithms include frequency control methods based on Deep Deterministic Policy Gradient (DDPG algorithm), and its main steps are as follows: Design of policy network and value network: The DDPG algorithm uses two independent networks, which are used to generate control actions (policy network) and evaluate action values (value network) respectively. Experience replay mechanism: By storing and reusing past experience samples, the training efficiency and algorithm stability are improved. Optimization of continuous action space: The deterministic policy gradient is used to directly optimize continuous control actions, so as to achieve frequency control. However, grid frequency control mainly relies on traditional frequency regulation methods based on droop control and frequency control methods based on Model Predictive Control (MPC). Droop control adjusts the output power of generators or energy storage devices through the linear relationship between frequency and power to respond to frequency deviations. The traditional droop control is essentially a proportional control strategy, which requires accurate setting of droop coefficients, is difficult to quickly respond to large-scale disturbances, and is difficult to adapt to the nonlinear complex new energy power system with variable operating conditions. Model predictive control requires an accurate system model, has a high computational complexity, and poor real-time performance, which limits its application in fast response scenarios. Therefore, the DDPG algorithm has the following deficiencies: Overestimation problem of action value function: Due to the evaluation error of a single value network, the frequency control strategy falls into a local optimum. Insufficient sample efficiency: The random experience replay mechanism cannot effectively utilize samples with high learning value, which affects the training efficiency of the algorithm. Summary of the Invention
[0004] The technical problem to be solved by the present invention: Aiming at the above problems of the prior art, a fast frequency control method and system for a new energy power system are provided. The present invention aims to solve the frequency stability problem caused by the reduction of inertia and damping levels in the current power system, especially in the context of the increasing penetration rate of new energy, the traditional frequency control methods have problems such as slow response speed, low control accuracy, and difficulty in adapting to complex dynamic working conditions. It quickly responds to large-scale disturbances in the new energy power system, adapts to the operation requirements of the complex nonlinear new energy power system, and provides an efficient and accurate frequency control solution for the power grid.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0006] A fast frequency control method for a new energy power system includes the following steps: In outer layer control, a supervised control method based on reinforcement learning is adopted to generate a set output power ; the active set output power is obtained through droop control and power strategy and the reactive set output power ; combining the set output power of the voltage source converter VSC, the active power , the given active power set balance point and the frequency set balance point to perform active power control to generate a control frequency ; combining the reactive power , the given reactive power set balance point and the voltage set point to perform reactive power control to generate a control voltage ; according to the control voltage and the control frequency generate a modulation voltage reference signal through inner layer control ; the modulation voltage reference signal generates a control signal modulation ratio through PWM modulation to control the voltage source converter VSC.
[0007] Optionally, when the supervised control method based on reinforcement learning is adopted to generate the set output power in the outer layer control, the supervised control method based on reinforcement learning adopted is the improved TD3 algorithm. The improved TD3 algorithm improves the way of obtaining experience from the experience replay pool for the TD3 algorithm to the prioritized experience replay mechanism, and uses the current generator frequency and the rate of change of frequency RoCoF of the new energy power system as the state, and the set output power as the corresponding action, and uses a preset reward function to iteratively obtain the optimal set output power . The prioritized experience replay mechanism refers to extracting the temporal difference error in the TD3 algorithm as the priority of the experience in the experience replay pool, and selecting the size of the priority to sample experience from the experience replay pool to replace the random sampling experience of the TD3 algorithm from the experience replay pool.
[0008] Optionally, the function expression of the reward function is:
[0009] ,
[0010] ,
[0011] ,
[0012] wherein, is the reward, is the cost coefficient for controlling the action; is the action, and the action is the set output power ; is the deviation cost term, is the deviation penalty term, and are both weight factors, is the state, and the state is the current generator frequency and the rate of change of frequency RoCoF, is the ideal state.
[0013] Optionally, the function expression for obtaining the active set output power and the reactive set output power through droop control and power strategy is:
[0014] ,
[0015] ,
[0016] wherein, represents the filter voltage, represents the transformer current, is a 90° rotation matrix, is the transpose operation.
[0017] Optionally, the function expression for performing active power control to generate the control frequency by combining the set output power of the voltage source converter VSC, the active power , the given active power set balance point and the frequency set balance point is:
[0018] ,
[0019] ,
[0020] wherein, is the given frequency set balance point, is the active power droop coefficient, is the given active power set balance point, is the low-pass filtered active power, is the time, is with respect to time The differential of is the cut-off frequency of the low-pass filter, and
[0021] Optionally, the combined reactive power , the given reactive power setting balance point and the voltage set point are used for reactive power control to generate a control voltage , and the functional expression is:
[0022] ,
[0023] ,
[0024] wherein, is the given voltage set point, is the reactive power droop coefficient, is the given reactive power setting balance point, is the reactive power after low-pass filtering, is time, is the differential of with respect to time, is the cut-off frequency of the low-pass filter, and
[0025] Optionally, the modulation voltage reference signal is generated by inner-layer control based on the control voltage and the control frequency , and the functional expression is:
[0026] ,
[0027] wherein, and are proportional gains, is the control voltage, is the reference voltage, is the control frequency, and
[0028] In addition, the present invention further provides a fast frequency control system for a new energy power system, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the fast frequency control method for the new energy power system.
[0029] In addition, the present invention further provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the fast frequency control method for the new energy power system through a processor.
[0030] In addition, the present invention also provides a computer program product, including a computer program or instruction, which is programmed or configured to execute the fast frequency control method for the new energy power system through a processor.
[0031] Compared with the prior art, the present invention mainly has the following advantages: the present invention includes generating a set output power by adopting a supervised control method based on reinforcement learning in the outer layer control ; obtaining the active set output power through droop control and power strategy and the reactive set output power ; combining the set output power , active power , the given active power set balance point and the frequency set balance point to perform active power control to generate a control frequency ; combining reactive power , the given reactive power set balance point and the voltage set point to perform reactive power control to generate a control voltage ; generating a modulation voltage reference signal through inner layer control according to the control voltage and the control frequency ; generating a control signal modulation ratio by PWM modulation of the modulation voltage reference signal to control the voltage source converter VSC. Through the double-layer control strategy of outer layer control and inner layer control, combining the generation of set output power by adopting a supervised control method based on reinforcement learning in the outer layer control , it can solve the frequency stability problem caused by the reduction of inertia and damping level in the current power system. Especially in the context of the increasing new energy penetration rate, the traditional frequency control method has problems such as slow response speed, low control accuracy, and difficulty in adapting to complex dynamic working conditions. It gives full play to the fast response characteristics of the voltage source converter, quickly adjusts the power output when disturbed, reduces the frequency deviation and accelerates the system recovery, and improves the stability and reliability of frequency control.
[0032] BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic diagram of the basic process of the method of the embodiment of the present invention.
[0033] Figure 2 It is a schematic diagram of the training process of the improved TD3 algorithm in the embodiment of the present invention.
[0034] Figure 3Schematic diagram of the control process for applying the improved TD3 algorithm to frequency control in the embodiments of the present invention.
[0035] Figure 4 Schematic diagram of the simulation structure of the IEEE node system in the embodiments of the present invention.
[0036] Figure 5 Schematic diagram of the frequency control result of the improved TD3 algorithm in the embodiments of the present invention.
[0037] Figure 6 Comparison diagram of the average frequency curves of the improved TD3 algorithm in the embodiments of the present invention. Detailed implementation manners
[0038] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.
[0039] As Figure 1 shown, the fast frequency control method for a new energy power system in this embodiment includes the following steps: in the outer layer control, a supervision control method based on reinforcement learning is adopted to generate a set output power ; the active set output power and the reactive set output power are obtained through droop control and power strategy ; the active power control is combined with the set output power of the voltage source converter VSC, the given active power set balance point and the frequency set balance point to generate a control frequency ; the reactive power control is combined with the reactive power , the given reactive power set balance point and the voltage set point to generate a control voltage ; according to the control voltage and the control frequency , a modulation voltage reference signal is generated through the inner layer control ; the modulation voltage reference signal is used to generate a control signal modulation ratio
[0040] through PWM modulation to control the voltage source converter VSC. Figure 1 As shown, the method of this embodiment includes an improved TD3 algorithm and a fast frequency control model. The fast frequency control model includes a DC side circuit and a lossless switching unit connected to the power grid, and through an RLC filter Connected to the power grid, its mathematical model is based on the dq coordinate system and is represented in per-unit form. The DC side circuit includes a DC link capacitor for energy storage , a constant current source representing the input current of renewable energy generation and a controlled DC current source representing the current related to the energy flexibility of the system . The DC control maintains the DC link voltage at a set reference value and utilizes the flexibility of the DC side energy storage to achieve control. The adopted VSC control is based on a two-level scheme including internal and external control loops. In the outer loop control of the VSC, the droop control and the power measurement system output and and adjusts the set values of frequency and voltage control to control the output voltage amplitude and frequency. The active power P and reactive power Q controllers are designed to control the output voltage amplitude and frequency .
[0041] The fast frequency control method for the new energy power system in this embodiment can solve the frequency stability problem in the current power system due to the reduction of inertia and damping levels. Especially in the context of the increasing penetration of new energy, the traditional frequency control methods have problems such as slow response speed, low control accuracy, and difficulty in adapting to complex dynamic working conditions. It gives full play to the fast response characteristics of the voltage source converter, quickly adjusts the power output when disturbed, reduces the frequency deviation and accelerates the system recovery, and improves the stability and reliability of frequency control.
[0042] In the outer layer control, when using the reinforcement learning-based supervisory control method to generate the set output power , the reinforcement learning-based supervisory control method can adopt the DDPG (Deep Deterministic Policy Gradient) algorithm and the TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm. However, the DDPG algorithm uses the evaluation error of a single value network and has an overestimation problem; moreover, the random experience replay mechanism of the DDPG algorithm and the TD3 algorithm cannot effectively utilize samples with high learning value, which affects the training efficiency of the algorithm. On this basis, as an alternative implementation, when using the reinforcement learning-based supervisory control method to generate the set output power in the outer layer control of this embodiment, the adopted reinforcement learning-based supervisory control method is the improved TD3 algorithm, such as Figure 2As shown, the improved TD3 algorithm improves the way of obtaining experiences from the experience replay pool for the TD3 algorithm to a prioritized experience replay mechanism, and uses the current generator frequency and rate of change of frequency RoCoF of the new energy power system as the state, and sets the output power as the corresponding action, and uses a preset reward function to iteratively obtain the optimal set output power . The prioritized experience replay mechanism refers to extracting the temporal difference error in the TD3 algorithm (a key indicator in reinforcement learning, used to measure the deviation between the currently estimated action value and the target action value) as the priority of the experiences in the experience replay pool, and selecting the priority size to sample experiences from the experience replay pool to replace the random sampling experiences in the experience replay pool of the TD3 algorithm. On the one hand, the improved TD3 algorithm in this embodiment can effectively avoid the overestimation problem of a single value network based on the inherent double Critic network of the TD3 algorithm, improving the stability and accuracy of the algorithm; on the other hand, through the prioritized experience replay mechanism, according to the learning value of the experience samples, it preferentially replays experiences with extremely high success rates or extremely high failure rates, thereby accelerating the learning process. Combining these technical improvements, the improved TD3 algorithm can quickly respond to large-scale disturbances, adapt to the operation requirements of complex nonlinear new energy power systems, and provide an efficient and accurate frequency control solution for the power grid.
[0043] As Figure 2 shown, in this embodiment, an experience sampled from the experience replay pool is represented as:
[0044] ,
[0045] wherein, is the state at the current moment, is the action at the current moment, is the reward at the current moment, is the state at the next moment. The TD3 algorithm includes a policy network (Actor network ), and two value networks: Critic network 1 and Critic network 2, as well as a policy target network and two value target networks: Critic target network 1 and Critic target network 2. The network parameters of the Actor network are denoted as , the network parameters of Critic network 1 are denoted as and the network parameters of Critic network 2 are denoted as ; the network parameters of the Actor target network are denoted as , the network parameters of Critic target network 1 are denoted as , and the network parameters of Critic target network 2 are denoted as ; The network structures of the policy network and the policy target network are the same, the network structures of the Critic network 1 and the Critic target network 1 are the same, and the network structures of the Critic network 2 and the Critic target network 2 are the same. The policy network (Actor network ) is used to generate an action based on the current state by adding noise:
[0046] ,
[0047] where is the action after adding noise, and the action is the set output power ; is the action generated by the Actor network based on the current state , the state is the current generator frequency and the rate of change of frequency RoCoF, is the noise, and it satisfies the distribution , where is the standard deviation of the system state s.
[0048] In this embodiment, the calculation function expression of the temporal difference error is:
[0049] ,
[0050] where is the temporal difference error of the j-th experience, is the reward, is the discount factor, is the Q value obtained by the value target network based on the state at the next moment and the action ; is the Q value obtained by the value network based on the state at the current moment and the action . As an optional implementation manner, when extracting the temporal difference error in the TD3 algorithm as the priority of the experience in the experience replay pool in this embodiment, it includes sampling k specified experiences from the experience replay pool according to the importance sampling weights of each experience, and calculating and updating their importance sampling weights according to the above formula; and the importance sampling weight of any j-th experience is calculated by the function expression:
[0051] ,
[0052] where is the total number of experiences in the experience replay pool, is the normalized probability of the j-th experience, and there is:
[0053] ,
[0054] wherein, and are the priority values of the j-th experience and the i-th experience respectively, is the priority adjustment parameter. When , the relative gap between high-priority samples and low-priority samples will decrease, and the influence of low-priority samples may increase. When , the power of the priority will make the value of high-priority samples grow faster and the value of low-priority samples grow slower, thereby enhancing the probability of high-priority samples being selected. The initial value of the priority value can be a random value, and through the update operation, its value becomes the corresponding importance sampling weight.
[0055] In this embodiment, the functional expression of the reward function is:
[0056] ,
[0057] ,
[0058] ,
[0059] wherein, is the reward, is the cost coefficient of the control action; is the action, and the action is the set output power ; is the deviation cost term, is the deviation penalty term, and are both weight factors, is the state, and the state is the current generator frequency and the rate of change of frequency RoCoF, is the ideal state.
[0060] In this embodiment, the functional expression for updating the network parameters of Critic network 1 and the network parameters of Critic network 2 is:
[0061] ,
[0062] In the above formula, represents the network parameters of Critic network 1 or the network parameters of Critic network 2, is the target value, is the Q value obtained by Critic network 1 or Critic network 2 based on the state s and the action a, and N is the number of samples;
[0063] ,
[0064] In the above formula, is the reward obtained based on the state s and action a at the current moment, is the discount factor, is the Q value obtained by the Critic target network 1 or the Critic target network 2 based on the state at the next moment and the action . In addition, in this embodiment, the target network delay update is adopted for updating the policy target network and the two value target networks to reduce the fluctuation during policy update, and further improve the convergence efficiency and stability of the algorithm. The function expressions for updating the policy target network and the two value target networks are as follows:
[0065] ,
[0066] In the above formula, represents the network parameters of the Critic target network 1 or the network parameters of the Critic target network 2 , is the soft update coefficient, represents the network parameters of the Critic network 1 or the network parameters of the Critic network 2 , is the soft update coefficient, are the network parameters of the Actor target network, are the network parameters of the Actor network.
[0067] After the agent training of the improved TD3 algorithm in this embodiment is completed, it can be applied to an actual system or a simulation environment to perform frequency control such as Figure 3As shown in the figure, the process includes: Loading the trained agent: Loading the agent parameters saved during the training process. Initializing the environment and control parameters: Initializing the state of the simulation or real environment, and setting the initial frequency, load status, and disturbance conditions according to the system model. Setting the execution mode of the agent. If in the real environment, disable the exploration strategy and directly execute deterministic actions. Frequency control loop: During the operation of the system, the agent adjusts the frequency in real time. At each time step, obtain the action (power adjustment amount) output by the agent according to the current state. Continuously monitor the frequency deviation to ensure that the control effect meets the requirements. Recording the control effect: Recording data such as the frequency value, power output, and frequency deviation at each time step. Saving the data in the form of a file or chart for subsequent analysis. Drawing a curve of frequency vs. time to evaluate the dynamic performance and stability of the control. Evaluating the control performance: Comparing the control effect of the agent with traditional methods (such as droop control, DDPG control, TD3 control) to demonstrate the superiority of the method.
[0068] In this embodiment, the droop control and power strategy are used to obtain the active set output power and the reactive set output power The function expressions are:
[0069] ,
[0070] ,
[0071] where represents the filter voltage, represents the transformer current, is a 90° rotation matrix, is the transpose operation.
[0072] In this embodiment, combining the set output power of the voltage source converter VSC, the active power , the given active power set balance point and the frequency set balance point to perform active control to generate the control frequency The function expression is:
[0073] ,
[0074] ,
[0075] where is the given frequency set balance point, is the active power droop coefficient, is the given active power set balance point, is the active power after low-pass filtering, is time, is the differential with respect to time and is the cut-off frequency of the low-pass filter, is the active power.
[0076] In this embodiment, combined with the reactive power , the given reactive power set balance point and the voltage set point reactive power control is performed to generate the control voltage The functional expression is:
[0077] ,
[0078] ,
[0079] where is the given voltage set point, is the reactive power droop coefficient, is the given reactive power set balance point, is the reactive power after low-pass filtering, is time, is the differential with respect to time and is the cut-off frequency of the low-pass filter, is the reactive power.
[0080] In this embodiment, according to the control voltage and the control frequency the modulation voltage reference signal is generated through inner-layer control. The functional expression is:
[0081] ,
[0082] where and are the proportional gains, is the control voltage, is the reference voltage, is the control frequency, is the reference frequency.
[0083] To verify the performance of the fast frequency control method of this embodiment for new energy power systems in fast frequency control tasks, in this embodiment, at Figure 4The IEEE 14-node test system shown is used for simulation tests. In this embodiment, based on the standard IEEE 14-bus system, two VSCs are introduced into this system, which are respectively connected to Bus 1 and Bus 2. The converter-based device is equipped with a power of 850 MW. The disturbances for performance evaluation control are generator loss and load loss, both of which simulate step changes in active power injection on the network buses of interest. It is assumed that when the frequency deviation exceeds ±0.5 Hz or the rate of change of frequency RoCoF exceeds ±1 Hz / s, load shedding (a power system protection measure) will be triggered. To test and verify the performance, an 800-MW disturbance is applied to Bus 14, and the obtained results are as Figure 5 shown. In addition, in this embodiment, the performances of other existing frequency control methods are also compared, including droop control, DDPG algorithm, TD3 algorithm, and improved TD3 algorithm. By analyzing the frequency response of the system under load mutations, the system average frequency control effects of each method are obtained in this embodiment method, as Figure 6 shown. It can be seen from Figure 6 that, as a traditional baseline method, droop control exhibits obvious frequency deviations under load disturbances, with a delay in the frequency recovery speed and a large steady-state error. Droop control cannot effectively meet the frequency control requirements in complex environments. The DDPG algorithm improves the shortcomings of droop control and has a relatively small frequency drop and the ability to enter the recovery stage faster. However, the DDPG algorithm introduces the overestimation problem of a single neural network, resulting in the frequency recovery tending to fall into a local optimum. The TD3 algorithm eliminates the overestimation bias problem in DDPG by introducing a dual evaluation network and a delayed update mechanism, significantly improving the frequency control performance. The improved TD3 algorithm introduces a priority experience replay mechanism based on TD3, which improves the learning efficiency of the control strategy and reduces the oscillation in the control strategy through the reinforcement learning of the main samples. The experimental results show that the improved TD3 algorithm has the best frequency control performance, achieving a small frequency deviation and fast stability.
[0084] In summary, the fast frequency control method for the new energy power system in this embodiment integrates the reinforcement learning control technology and the VSC technology into the new energy power system, and constructs a fast frequency control model for the new energy power system based on reinforcement learning. The improved TD3 algorithm is used to interactively train the improved power system. By introducing a dual Critic network and a delayed update mechanism, the problem of overestimation of Q value is solved, and the accuracy of control decisions is improved; at the same time, combined with the prioritized experience replay mechanism, the priorities of learning samples are dynamically adjusted according to the TD error, so that high-value experiences are fully utilized, thereby accelerating the algorithm convergence speed and obtaining the optimal fast frequency control strategy.
[0085] In addition, this embodiment also provides a fast frequency control system for a new energy power system, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the fast frequency control method for the new energy power system.
[0086] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the fast frequency control method for the new energy power system through a processor.
[0087] In addition, this embodiment also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the fast frequency control method for the new energy power system through a processor.
[0088] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for realizing the functions in the processFigure 1 One process or multiple processes and / or boxes Figure 1 Steps of the functions specified in one box or multiple boxes.
[0089] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. A rapid frequency control method for a new energy power system, characterized in that It includes the following steps: generating a set output power by adopting a supervised control method based on reinforcement learning in the outer layer control ; obtaining active power and reactive power through droop control and power strategy ; combining the set output power of the voltage source converter VSC, active power , a given active power set balance point and a frequency set balance point to perform active power control to generate a control frequency ; combining reactive power , a given reactive power set balance point and a voltage set point to perform reactive power control to generate a control voltage ; generating a modulation voltage reference signal through inner layer control according to the control voltage ; Generate a modulation voltage reference signal Generate a control signal modulation ratio through PWM modulation To control the voltage source converter VSC; The above-mentioned outer layer control uses a supervised control method based on reinforcement learning to generate a set output power When the supervised control method based on reinforcement learning used is the improved TD3 algorithm, the improved TD3 algorithm improves the way of obtaining experiences from the experience replay pool for the TD3 algorithm to a prioritized experience replay mechanism, and uses the current generator frequency and rate of change of frequency RoCoF of the new energy power system as the state, and the set output power as the corresponding action, and uses a preset reward function to iteratively obtain the optimal set output power The prioritized experience replay mechanism refers to extracting the temporal difference error in the TD3 algorithm as the priority of the experiences in the experience replay pool, and selecting the size of the priority to sample experiences from the experience replay pool to replace the random sampling experiences of the TD3 algorithm from the experience replay pool; Said according to the control voltage and the control frequency generate a modulation voltage reference signal through inner layer control The functional expression is: , Among them, and are proportional gains, is the control voltage, is the reference voltage, is the control frequency, is the reference frequency.
2. The rapid frequency control method for a new energy power system according to claim 1, characterized in that The functional expression of the said reward function is as follows: , , , wherein, is the reward, is the cost coefficient of the control action; is the action, which is used to adjust the set output power ; is the deviation cost term, is the deviation penalty term, and are both weight factors, is the state, and the state is the current generator frequency and the rate of change of frequency RoCoF, is the ideal state.
3. The rapid frequency control method for a new energy power system according to claim 1, characterized in that The active power obtained through droop control and power strategy and reactive power The functional expression is as follows: , , Among them, represents the filter voltage, represents the transformer current, is a 90° rotation matrix, is a transpose operation.
4. The rapid frequency control method for a new energy power system according to claim 1, characterized in that, The set output power combined with the voltage source converter VSC, active power, the given active power set balance point and the frequency set balance point are used for active power control to generate the function expression of the control frequency , , Among them, is the given frequency set balance point, is the active power droop coefficient, is the given active power set balance point, is the active power after low-pass filtering, is time, is the differential with respect to time and is the cut-off frequency of the low-pass filter, is the active power.
5. The rapid frequency control method for a new energy power system according to claim 1, wherein The combined reactive power , the given reactive power setting balance point and the voltage set point are used for reactive power control to generate a control voltage The functional expression of which is: , , wherein, is the given voltage set point, is the reactive power droop coefficient, is the given reactive power set balance point, is the reactive power after low-pass filtering, is time, is the differential of with respect to time is the cut-off frequency of the low-pass filter, is the reactive power.
6. A fast frequency control system for a new energy power system, comprising a microprocessor and a memory connected to each other, characterized in that, The said microprocessor is programmed or configured to execute the fast frequency control method for new energy power systems described in any one of claims 1 to 5.
7. A computer-readable storage medium storing a computer program or instructions, characterized in that, The said computer program or instruction is programmed or configured to execute, through a processor, the fast frequency control method for new energy power systems described in any one of claims 1 to 5.
8. A computer program product, comprising a computer program or instructions, characterized in that, The said computer program or instruction is programmed or configured to execute, through a processor, the fast frequency control method for new energy power systems described in any one of claims 1 to 5.
Citation Information
Patent Citations
VSC-MTDC interconnection system frequency control method based on virtual synchronous machine
CN116054236A
Shared energy storage frequency adjusting method and system based on deep reinforcement learning
CN119298090A