A method and apparatus for optimizing the configuration of generator sets in hybrid new energy power stations

By combining graph attention networks and a deep reinforcement learning model with a near-end policy optimization algorithm, the problem of optimal configuration of generator units in hybrid renewable energy power plants was solved, achieving the effects of improving stability under small disturbances and reducing operating costs.

CN119834370BActive Publication Date: 2025-11-14WUHAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411875191.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-11-14
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

How to optimize the capacity ratio and location distribution of grid-connected and grid-connected generator units in hybrid renewable energy power stations to improve stability under minor disturbances, reduce the number of grid-connected generator units, and lower operating costs.

Method used

A deep reinforcement learning (DRL) model combining graph attention network (GAT) and proximal policy optimization (PPO) algorithm is used to establish an optimal configuration model. An online optimal configuration strategy is generated through offline training to adapt to changes in external power grid strength and internal topology.

Benefits of technology

This improves the small-interference stability of hybrid new energy power stations and reduces the number of grid-connected generator units required, resulting in better economic benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119834370B_ABST
    Figure CN119834370B_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for optimizing the configuration of generator sets in a hybrid renewable energy power station, relating to the field of renewable energy power station technology. The method includes: establishing an optimized configuration model for generator sets in a hybrid renewable energy power station, where the configuration of generator sets and energy storage devices are used as optimization variables, aiming to improve small-disturbance stability margin and reduce the number of GFM generator sets; integrating GAT with a PPO-based DRL framework to obtain a GAT-PPO-based DRL model; offline training the GAT-PPO-based DRL model using the optimized configuration model; and online outputting an optimized configuration strategy based on the external grid strength and the topology information of the hybrid renewable energy power station using the trained GAT-PPO-based DRL model. This invention can adapt to changes in external grid strength and internal topology in HRPS (High-Speed ​​Power Supply System), and the generated optimized configuration strategy not only improves the small-disturbance stability of the hybrid renewable energy power station but also allows for the configuration of fewer GFM generator sets to achieve better economic benefits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of new energy power station technology, specifically to a method and apparatus for optimizing the configuration of generator sets in a hybrid new energy power station. Background Technology

[0002] The "dual-carbon" strategic goal has promoted the development of new energy power generation in my country. However, the stability issues caused by the large-scale centralized grid connection of new energy power sources have become increasingly prominent. This is mainly manifested in the complex broadband oscillations that occur when grid-following (GFL) controlled generators are connected to weak grids. In contrast, grid-forming (GFM) controlled generators possess active grid support capabilities and exhibit stronger stability under weak grid conditions. However, under strong grid conditions, GFM generators may increase the risk of low-frequency oscillations, affecting small-disturbance stability. Furthermore, configuring too many GFM generators in a power station leads to higher operating costs. Therefore, optimizing the capacity ratio and location distribution of GFM generators in hybrid renewable power stations (HRPS) is currently a research hotspot. Summary of the Invention

[0003] The purpose of this invention is to provide a method and apparatus for optimizing the configuration of generator sets in hybrid new energy power stations, which solves the problems of how to optimize the capacity ratio and location distribution of GFM generator sets in HRPS in the prior art. It can adapt to the internal topology changes of HRPS, generate optimized configuration strategies online, not only improve the stability of HRPS under small disturbances, but also configure fewer GFM generator sets to achieve better economic benefits.

[0004] To achieve the above objectives, in a first aspect, the present invention provides a method for optimizing the configuration of generator sets in hybrid new energy power stations, comprising:

[0005] Step 1: Establish an optimization configuration model for hybrid new energy power station generator sets. The optimization configuration model uses the configuration of generator sets and energy storage devices as optimization variables, with the goal of improving small disturbance stability margin and reducing the number of GFM generator sets.

[0006] Step 2: Integrate GAT with the PPO-based DRL framework to obtain the GAT-PPO-based DRL model;

[0007] Step 3: Use the optimization configuration model to train the DRL model based on GAT-PPO offline. After the training is completed, the DRL model based on GAT-PPO will output the optimization configuration strategy online based on the external grid strength and the topology information of the hybrid new energy power station.

[0008] According to the present invention, a method for optimizing the configuration of generator sets in a hybrid new energy power station includes an optimization strategy that includes the capacity ratio and location distribution of GFM generator sets and GFL generator sets in the hybrid new energy power station, as well as the control type of the energy storage device.

[0009] According to the method for optimizing the configuration of generator sets in a hybrid new energy power station provided by the present invention, step 2 specifically includes:

[0010] The state variables are obtained based on the current topology and state data of the hybrid new energy power station;

[0011] The internal configuration actions are generated based on the state variables and applied to the hybrid new energy power station. The internal configuration actions are evaluated based on the changes in the state of the hybrid new energy power station.

[0012] The reward function is determined based on the relationship between the state variables and the preset constraints.

[0013] According to the present invention, a method for optimizing the configuration of generator units in a hybrid new energy power station is provided, wherein the expression for the state variables is:

[0014] w1(X)=min{ξ L1 ,ξ L2 ,...,ξ Lk -0.05

[0015]

[0016]

[0017] S t =[S topo ,S h ]

[0018] In the formula, S t For state variables; s topo S h These represent the topology and status data of the current hybrid renewable energy power station; w1, w2, and w3 represent the low-frequency band small-interference stability index, the sub / supersynchronous band small-interference stability index, and the operating cost index, respectively; ξ Lk ξ represents the damping ratio corresponding to the k-th low-frequency closed-loop pole; Sn and These represent the damping ratio and critical damping ratio corresponding to the closed-loop pole of the nth sub- / supersynchronous frequency band, respectively; N GFM X represents the number of GFM generator sets; X is the optimization variable, representing the configuration vector of generator sets and energy storage devices within the hybrid new energy power station.

[0019] According to the method for optimizing the configuration of generator units in a hybrid new energy power station provided by the present invention, the expression of the reward function is as follows:

[0020]

[0021] In the formula, r is the reward value.

[0022] According to the present invention, a method for optimizing the configuration of generator sets in a hybrid renewable energy power station is provided. The DRL model based on GAT-PPO includes GAT, a policy network, and a value network. GAT is used to obtain the node feature set based on the current state variables of the hybrid renewable energy power station. The policy network is used to generate configuration strategies for GFM generator sets, GFL generator sets, and energy storage devices. The value network is used to evaluate the merits of the configuration strategies for GFM generator sets, GFL generator sets, and energy storage devices.

[0023] According to the method for optimizing the configuration of generator units in a hybrid new energy power station provided by the present invention, step 3 specifically includes: characterizing the node characteristics of the hybrid new energy power station through GAT, updating the policy network and value network, and applying the trained DRL model based on GAT-PPO online.

[0024] According to the present invention, a method for optimizing the configuration of generator units in a hybrid new energy power station, the method for updating the strategy network includes: inputting a node feature set represented by GAT into the strategy network to obtain the distribution of the new strategy and the old strategy; calculating the probability of selecting each action under the new strategy and the old strategy respectively based on the distribution of the new strategy and the old strategy; dividing the probability under the new strategy by the probability under the old strategy to obtain the probability ratio; calculating the objective function value of the DRL model based on GAT-PPO using the dominance function and the probability ratio; using the negative value corresponding to the objective function value as the loss function of the strategy network; and updating the parameters of the strategy network through backpropagation using the loss function until a new strategy that meets the shearing requirements is obtained.

[0025] According to the present invention, a method for optimizing the configuration of generator sets in a hybrid new energy power station is provided, which updates the value network by: inputting the node feature set represented by GAT into the value network, calculating the value function under the current state through a forward propagation mechanism, and calculating the loss function to update the value network by gradient.

[0026] Secondly, the present invention provides a hybrid new energy power station generator set optimization configuration device, comprising:

[0027] Establish a unit to build an optimal configuration model for generator sets in hybrid new energy power stations. The optimal configuration model uses the configuration of generator sets and energy storage devices as optimization variables, with the goal of improving small disturbance stability margin and reducing the number of GFM generator sets.

[0028] The fusion unit is used to fuse GAT with the PPO-based DRL framework to obtain a GAT-PPO-based DRL model.

[0029] The training output unit is used to train the GAT-PPO-based DRL model offline using the optimized configuration model. After the training is completed, the GAT-PPO-based DRL model outputs the optimized configuration strategy online based on the external grid strength and the topology information of the hybrid renewable energy power station.

[0030] The technical solution of the present invention has at least the following technical effects:

[0031] This invention provides a method and apparatus for optimizing the configuration of generator units in a hybrid renewable energy power station. The method includes: establishing an optimization configuration model for the generator units in the hybrid renewable energy power station, where the configuration of generator units and energy storage devices are used as optimization variables, aiming to improve small-disturbance stability margin and reduce the number of GFM generator units; integrating GAT with a PPO-based DRL framework to obtain a GAT-PPO-based DRL model; offline training the GAT-PPO-based DRL model using the optimization configuration model; and online outputting an optimized configuration strategy based on the external grid strength and the topology information of the hybrid renewable energy power station using the trained GAT-PPO-based DRL model. This invention can adapt to changes in external grid strength and internal topology in HRPS (High-Speed ​​Rail System), and the generated optimized configuration strategy not only improves the small-disturbance stability of the hybrid renewable energy power station but also allows for the configuration of fewer GFM generator units to achieve better economic benefits. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0033] In the attached diagram:

[0034] Figure 1 This is a schematic diagram illustrating the training process of the DRL model based on GAT-PPO of this invention.

[0035] Figure 2 This is a schematic diagram of the structure of the present invention, which includes an energy storage device and a grid-connected HRPS.

[0036] Figure 3 This is a comparison chart of the training reward curves of the DRL models based on GAT-PPO and PPO in scenario 1 of this invention.

[0037] Figure 4 This is a diagram showing the optimized configuration of GFM in HRPS based on GAT-PPO under three scenarios of this invention;

[0038] Figure 5 This is a flowchart of the method for optimizing the configuration of generator sets in hybrid new energy power stations according to the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0040] The following detailed description of some embodiments of the present invention will be provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0041] This invention's in-depth research reveals that the small-disturbance stability of HRPS is significantly influenced by factors such as external grid strength (short-circuit ratio), the capacity ratio of GFM / GFL generator sets, and the specific location of GFM generator sets within the HRPS. Therefore, by using the configuration of generator sets and energy storage devices as optimization variables, and aiming to improve the small-disturbance stability margin and reduce the number of GFM generator sets, an optimized configuration model for the capacity ratio and location distribution of GFM generator sets in the HRPS can be established. Solving this optimized configuration model faces challenges due to the time-varying nature of the external power system grid strength affecting the small-disturbance stability of the HRPS. Model-driven methods often suffer from low efficiency and heavy computational burden, making online application difficult. However, the deep reinforcement learning (DRL) method, employing offline training and online application, allows for rapid online solution of the optimized configuration model after training.

[0042] Currently, online optimization techniques for small-disturbance stability based on deep reinforcement learning focus on controller parameter optimization, with few applications in the optimal configuration of GFM / GFL generator sets in HRPS. Furthermore, since the topology of HRPS affects its small-disturbance stability, topology changes need to be considered. Therefore, this invention proposes an optimal configuration method for HRPS generator sets, which is an online solution for the optimal configuration of GFM / GFL generator sets in HRPS based on the GAT-PPO DRL model.

[0043] The HRPS generator set optimization configuration method provided by this invention can adapt to internal topology changes within the HRPS. The online generated optimization configuration strategy not only improves the small-disturbance stability of hybrid renewable energy power plants but also allows for the configuration of fewer GFM generator sets to achieve better economic efficiency. This method integrates a graph attention network (GAT), which has stronger generalization ability to different topologies, with a deep reinforcement learning (DRL) framework based on the proximal policy optimization (PPO) algorithm to obtain a GAT-PPO-based DRL model. The optimization configuration model is then trained offline. After training, the GAT-PPO-based DRL model outputs an online optimization configuration strategy, including the capacity ratio and location distribution of GFM / GFL generator sets in the HRPS and the control type of energy storage devices, based on the external grid strength (short-circuit ratio) and HRPS topology information.

[0044] Please see Figure 5 This invention provides a method for optimizing the configuration of generator sets in hybrid new energy power stations, including:

[0045] Step 1: Establish an optimal configuration model for hybrid new energy power station generator sets. This optimal configuration model uses the configuration of generator sets and energy storage devices as optimization variables, with the goal of improving small disturbance stability margin and reducing the number of GFM generator sets.

[0046] Step 2: Integrate GAT with the PPO-based DRL framework to obtain the GAT-PPO-based DRL model;

[0047] Step 3: The GAT-PPO-based DRL model is trained offline using an optimized configuration model. After training, the GAT-PPO-based DRL model outputs an optimized configuration strategy online based on the external grid strength (short-circuit ratio) and the topology information of the hybrid renewable energy power station. This optimized configuration strategy includes the capacity ratio and location distribution of GFM and GFL generators in the hybrid renewable energy power station, as well as the control type of the energy storage device.

[0048] Specifically, the design process of the DRL model based on GAT-PPO is as follows, that is, step 2 specifically includes:

[0049] 1) Design of the state space

[0050] State variable S t Based on the current HRPS topology data S topo and state data S h Composition. State data S hThis includes oscillation modes, damping thresholds, small-disturbance stability indices (w1, w2) in the low-frequency band and sub / supersynchronous frequency band, and operating cost indices (w3). Specifically, to study the interaction between the HRPS and the power grid, dq impedance modeling is used to calculate eigenvalues ​​and damping ratios, quantitatively evaluating the small-disturbance stability and stability margin of the grid-connected system. Configuring too many GFM generator units in the HRPS leads to higher operating costs.

[0051] w1(X)=min{ξ L1 ,ξ L2 ,...,ξ Lk -0.05 (1)

[0052]

[0053] S t =[S topo ,S h (4)

[0054] In the formula, ξ Lk ξ represents the damping ratio corresponding to the k-th low-frequency closed-loop pole; Sn and These represent the damping ratio and critical damping ratio corresponding to the closed-loop pole of the nth sub- / supersynchronous frequency band, respectively; N GFM X represents the number of GFM generator sets; X is the optimization variable, which in this invention represents the configuration vector of generator sets and energy storage devices within a hybrid new energy power station.

[0055] 2) Design of motion space

[0056] The action space is based on the current state variable S of the HRPS. t Generate internal configuration action A t And it acts on HRPS, evaluating internal configuration actions A based on changes in HRPS status. t Specifically, taking the hybrid renewable energy power station as the environment, the configuration vector X of the generator sets and energy storage devices inside the HRPS at the current moment is defined. t Consider it as internal configuration action A t .

[0057] A t = [ X t (5)

[0058] 3) Design of the reward function

[0059] The reward function is based on the state variable S t The relationship with the preset constraints is determined by partitioning according to the constraints and the objective function. The objective function is considered only if the constraints are met; otherwise, a penalty value is returned. The current HRPS state variable S... tWhen constraints are not met, a larger penalty value is returned directly, aiming to ensure that the agent in the DRL makes decisions within the constraint domain. When S t A reward is given when the constraints are met, aiming to enable the agent in the DRL to find the optimal decision based on feasible options. Considering both the small-disturbance stability margin and the number of GFM generators, the expression for the reward function is as follows:

[0060]

[0061] In the formula, r is the reward value.

[0062] Specifically, the DRL model based on GAT-PPO proposed in this invention mainly includes three parts: GAT, policy network, and value network. GAT is used to obtain the node feature set based on the current state variables of the hybrid renewable energy power station. The policy network is used to generate configuration strategies for GFM generator sets, GFL generator sets, and energy storage devices. The value network is used to evaluate the merits of the configuration strategies for GFM generator sets, GFL generator sets, and energy storage devices.

[0063] The application process of the DRL model based on GAT-PPO proposed in this invention mainly includes three steps, with step 3 specifically including: 1. Characterizing the node characteristics of hybrid renewable energy power stations using GAT; 2. Updating the policy network and value network; 3. Applying the trained GAT-PPO-based DRL model online. Unlike the traditional PPO algorithm, in order to perceive changes in the HRPS topology, the state information of GAT-PPO needs to be used to characterize the node features (topology data S) of the HRPS through GAT. topo , Generator set control type and status data S h The current HRPS state variable S is first represented during training. t Input into GAT, output node feature set S gat Then, the node feature set S gat The inputs are fed into the policy network and the value network for training, and the outputs are the internal configuration actions A of the generator sets and energy storage devices within the HRPS. t and the action A used to evaluate the generation of internal configuration t The value function of superiority and inferiority.

[0064] The update process for the policy network and value network is as follows:

[0065] The node feature set S represented by GAT gatThe input is fed into the policy network, resulting in two policy distributions: old and new. Based on these distributions, the probability of choosing each action under each policy is calculated. Then, the probability under the new policy is divided by the probability under the old policy to obtain the probability ratio. The objective function of the GAT-PPO-based DRL model is calculated using the dominance function and the probability ratio. The negative value of the objective function is then used as the loss function for the policy network. The parameters of the policy network are updated through backpropagation using the loss function until a new policy that meets the pruning requirements is obtained. The node feature set S, represented by GAT, is then... gat The input is fed into the value network, and the value function for the current state is calculated through the forward propagation mechanism. Then, the loss function is calculated to update the gradient of the value network.

[0066] Understandably, in the value network, the reward discount factor γ directly affects the expected long-term return assessment. As the reward discount factor γ increases, GAT-PPO pays more attention to long-term rewards, thus impacting policy updates. When the loss function values ​​of both the policy network and the value network are stable and close to a small value, and the moving average reward is positive and tends to stabilize, it indicates that the model training has converged.

[0067] Finally, the trained GAT-PPO-based DRL model outputs online optimization configuration strategies based on the external short-circuit ratio and internal topology information of the HRPS, including the capacity ratio and location distribution of GFM / GFL generator sets in the HRPS and the control type of energy storage devices.

[0068] Based on the same inventive concept, another embodiment of the present invention provides a hybrid new energy power station generator set optimization configuration device, which corresponds to the method of the aforementioned embodiment, and the device includes:

[0069] Establish a unit to build an optimal configuration model for generator sets in hybrid new energy power stations. The optimal configuration model uses the configuration of generator sets and energy storage devices as optimization variables, with the goal of improving small disturbance stability margin and reducing the number of GFM generator sets.

[0070] The fusion unit is used to fuse GAT with the PPO-based DRL framework to obtain a GAT-PPO-based DRL model.

[0071] The training output unit is used to train the GAT-PPO-based DRL model offline using the optimized configuration model. After the training is completed, the GAT-PPO-based DRL model outputs the optimized configuration strategy online based on the external grid strength and the topology information of the hybrid renewable energy power station.

[0072] The following is a specific application example of the present invention.

[0073] To verify the superiority of the online solution method for the optimal configuration strategy based on the GAT-PPO DRL model proposed in this invention, Figure 2 An optimal configuration model is established for the grid-connected / grid-connected HRPS (High-Speed ​​Power Supply System) with energy storage devices. All generators within the HRPS are direct-drive wind turbines, categorized into grid-connected control and grid-connected control types based on the converter's control mode. The HRPS employs a chain structure with multiple feeders, consisting of five feeders, each connecting 18 wind turbines, totaling 90 wind turbines with a capacity of 2MW each. Wind turbines on the same feeder are connected via a 0.43km busbar, with inductance and resistance per unit length of 1.23mH / km and 0.17Ω / km, respectively. A 20MW energy storage device is connected to the turbine side at the PCC (Point of Common Coupling). Furthermore, this invention considers the interconnection topology of wind turbines at different feeder ends within a hybrid renewable energy power station. To explore the optimal configuration of GFM / GFL generators in the HRPS under different grid strengths, an external grid short-circuit ratio K is designed. SCR There are three scenarios, namely 1, 2, and 3.

[0074] The optimal configuration model for GFM / GFL generator sets in hybrid renewable energy power plants is shown below:

[0075] 1) Optimize variables

[0076] Vector X = [x1, x2, ..., x 91 The configuration of the energy storage devices in the GFM / GFL generator set and HRPS is characterized, therefore the number of optimization variables is 91. Among them, x k (k = 1, 2, ..., 90) is a 0 / 1 variable; a value of 1 indicates that the k-th unit is a GFM generator set, and a value of 0 indicates that the k-th unit is a GFL generator set. 91 This indicates the control type of the configured energy storage device. A value of 1 indicates that the energy storage device is under GFM control, and a value of 0 indicates that the energy storage device is under GFL control.

[0077] 2) Objective function

[0078] The optimization objective is to maximize the small disturbance stability margin in the low-frequency band and the sub / supersynchronous band while minimizing the operating cost. The small disturbance stability indices for the low-frequency band and the sub / supersynchronous band are constructed based on their respective small disturbance stability margins, as shown in formulas (1) and (2). Too many GFM generator sets in the HRPS will lead to higher operating costs, as shown in formula (3). The optimization objective function is as follows:

[0079] {min{-w1(X)}; min{-w2(X)}; min{w3(X)}} (7)

[0080] In the formula, w1, w2, and w3 are the low-frequency band small-interference stability index, the sub / supersynchronous band small-interference stability index, and the operating cost index, respectively; X is the optimization variable.

[0081] 3) Constraints

[0082] The following constraints must be met: In the low-frequency band, the damping ratio must be greater than the stability threshold (5%); in the sub- / supersynchronous frequency band, the damping ratio must be greater than the critical damping; the number of GFM generator sets must be no less than 10 and no more than 50. These constraints can be expressed as:

[0083]

[0084] In the formula, f Lk , k, f St t and t represent the frequency and number of low-frequency oscillation modes, and the frequency and number of sub / supersynchronous oscillation modes, respectively.

[0085] A DRL model based on GAT-PPO was used to solve the optimization configuration models under three different scenarios. Based on parameter comparison experiments, the following parameters of the GAT-PPO method were determined: a two-layer GAT was used to represent the topology of HRPS; in the PPO structure, three fully connected layers were selected; the total number of time steps was set to 100,000; the number of rounds was set to 2048; and the reward discount factor γ was set to 0.99. Under different scenarios, the capacity allocation and location distribution strategies of GFM / GFL generator sets obtained from the optimization configuration model of GFM / GFL generator sets in HRPS were solved using both the GAT-PPO-based DRL model and the PPO-based DRL model, respectively. The corresponding objective function values ​​are shown in Table 1. The training reward curves of the GAT-PPO-based DRL model and the PPO-based DRL model are similar in the three scenarios. Taking scenario 1 as an example, the model training reward curves are compared as follows: Figure 3 As shown. The optimal GFM configuration structure in the HRPS of the DRL model based on GAT-PPO in three scenarios is as follows. Figure 4 As shown.

[0086] Table 1 Optimal GFM configuration strategies in HRPS under three scenarios

[0087]

[0088] As shown in Table 1, in Scenario 1, although the number of GFM generator sets configured by the PPO-based DRL model and the GAT-PPO-based DRL model is the same, the stability margin of the GAT-PPO-based DRL model is 0.0601, which is 10.58% higher than the 0.0519 of the PPO-based DRL model. In Scenarios 2 and 3, the number of GFM generator sets configured by the GAT-PPO-based DRL model is less than that of the PPO-based DRL model, and the stability margin of the GAT-PPO-based DRL model is 13.15% and 28.13% higher than that of the PPO-based DRL model, respectively. It can be seen that the optimized configuration strategy generated by the GAT-PPO-based DRL model not only reduces the number of GFM generator sets but also improves the stability margin under small disturbances.

[0089] Taking scenario 1 as an example, as training progresses, Figure 3 The reward values ​​of the GAT-PPO-based and PPO-based DRL models continuously increase in each training round, gradually converging after approximately 400 training rounds. Compared to the PPO-based DRL model, the GAT-PPO-based DRL model not only converges faster but also achieves a higher final convergence value for the reward function. This indicates that GAT can better extract key information using attention mechanisms, flexibly guiding the agent in DRL to make optimal decisions.

[0090] from Figure 4 As can be seen, as the external grid short-circuit ratio increases from 1 to 3, the number of GFM generator sets in the HRPS decreases from 32 to 17, and the location of the GFM generator sets gradually transitions from the beginning to the end of the feeder. For example, in scenario 1, most of the GFM generator sets are configured at the beginning of each feeder, while in scenarios 2 and 3, most of the GFM generator sets are configured at the end of each feeder.

[0091] In summary, the proposed method and apparatus for optimizing the configuration of generator units in hybrid renewable energy power plants utilizes a GAT-PPO-based DRL model for online solution. Compared to other conventional DRL models, the PPO algorithm, through probability ratio shearing technology, limits the strategy update step size, improving the stability and efficiency of the learning process, and possesses advantages such as fast convergence speed and wide adaptability. Meanwhile, GAT effectively characterizes the topology information of the HRPS and the control type of the generator units, enabling the final optimized configuration strategy to better adapt to the topology changes of the HRPS. Therefore, this invention can adapt to both the external grid strength and internal topology changes of the HRPS, and the generated optimized configuration strategy not only improves the small-disturbance stability of hybrid renewable energy power plants but also allows for the configuration of fewer GFM generator units to achieve better economic benefits.

[0092] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. It should be understood that the invention is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A method for optimizing the configuration of generator sets in a hybrid new energy power station, characterized in that, include: Step 1: Establish an optimal configuration model for hybrid new energy power station generator sets. The optimal configuration model uses the configuration of generator sets and energy storage devices as optimization variables, with the goal of improving small disturbance stability margin and reducing the number of GFM generator sets. Step 2: Integrate GAT with the PPO-based DRL framework to obtain the GAT-PPO-based DRL model; Step 3: The optimized configuration model is used to train the DRL model based on GAT-PPO offline. After training, the DRL model based on GAT-PPO outputs the optimized configuration strategy online based on the external power grid strength and the topology information of the hybrid new energy power station. Step 2 specifically includes: The state variables are obtained based on the current topology and state data of the hybrid new energy power station; An internal configuration action is generated based on the state variables and applied to the hybrid new energy power station. The internal configuration action is evaluated based on the changes in the state variables of the hybrid new energy power station. The reward function is determined based on the relationship between the state variables and the preset constraints. The GAT is used to obtain the node feature set based on the current state variables of the hybrid renewable energy power station.

2. The method for optimizing the configuration of generator sets in hybrid new energy power stations according to claim 1, characterized in that, The optimized configuration strategy includes the capacity ratio and location distribution of GFM generator sets and GFL generator sets in hybrid new energy power stations, as well as the control type of the energy storage device.

3. The method for optimizing the configuration of generator sets in hybrid new energy power stations according to claim 1, characterized in that, The expression for the state variable is: In the formula, For state variables; , These are the topology data and status data of the current hybrid renewable energy power station, respectively. w 1. w 2 and w 3 are the low-frequency band small-interference stability index, the sub / supersynchronous frequency band small-interference stability index, and the operating cost index, respectively; Indicates the first k The damping ratio corresponding to each low-frequency closed-loop pole; and They represent the first n Damping ratio and critical damping ratio corresponding to the closed-loop poles of each sub / supersynchronous frequency band; N GFM The number of GFM generator sets; X To optimize variables, we represent the configuration vector of generator sets and energy storage devices within a hybrid new energy power station.

4. The method for optimizing the configuration of generator sets in hybrid new energy power stations according to claim 3, characterized in that, The expression for the reward function is: In the formula, This is the reward value.

5. The method for optimizing the configuration of generator sets in hybrid new energy power stations according to claim 1, characterized in that, The DRL model based on GAT-PPO includes GAT, a policy network, and a value network; the policy network is used to generate configuration strategies for GFM generator sets, GFL generator sets, and energy storage devices, and the value network is used to evaluate the merits of the configuration strategies for GFM generator sets, GFL generator sets, and energy storage devices.

6. The method for optimizing the configuration of generator sets in hybrid new energy power stations according to claim 5, characterized in that, Step 3 specifically includes: characterizing the node features of the hybrid new energy power station through GAT, updating the policy network and value network, and applying the trained DRL model based on GAT-PPO online.

7. The method for optimizing the configuration of generator sets in hybrid new energy power stations according to claim 6, characterized in that, The updated policy network includes: inputting the node feature set represented by GAT into the policy network to obtain the distribution of the new policy and the old policy; calculating the probability of selecting each action under the new policy and the old policy respectively according to the distribution of the new policy and the old policy; dividing the probability under the new policy by the probability under the old policy to obtain the probability ratio; calculating the objective function value of the DRL model based on GAT-PPO using the dominance function and the probability ratio; using the negative value corresponding to the objective function value as the loss function of the policy network; and updating the parameters of the policy network through backpropagation using the loss function until a new policy that meets the pruning requirement is obtained.

8. The method for optimizing the configuration of generator sets in hybrid new energy power stations according to claim 6, characterized in that, The updated value network includes: inputting the node feature set represented by GAT into the value network, calculating the value function in the current state through a forward propagation mechanism, and calculating the loss function to update the value network using gradients.

9. A hybrid new energy power station generator set optimization configuration device, characterized in that, include: A unit is established to build an optimal configuration model for generator sets in hybrid new energy power stations. The optimal configuration model uses the configuration of generator sets and energy storage devices as optimization variables, with the goal of improving the small disturbance stability margin and reducing the number of GFM generator sets. The fusion unit is used to fuse GAT with the PPO-based DRL framework to obtain a GAT-PPO-based DRL model. The training output unit is used to train the GAT-PPO-based DRL model offline using the optimized configuration model, and to output the optimized configuration strategy online based on the external power grid strength and the topology information of the hybrid new energy power station after the training is completed. The fusion unit is specifically used for: The state variables are obtained based on the current topology and state data of the hybrid new energy power station; An internal configuration action is generated based on the state variables and applied to the hybrid new energy power station. The internal configuration action is evaluated based on the changes in the state variables of the hybrid new energy power station. The reward function is determined based on the relationship between the state variables and the preset constraints. The GAT is used to obtain the node feature set based on the current state variables of the hybrid renewable energy power station.