Multi-agent reinforcement learning-based multi-machine wide-area damping cooperative control system and method

By installing wide-area damping controllers in wind turbines and synchronous generator units, and using multi-agent reinforcement learning to train agents to generate optimal control strategies, the problem of poor performance of traditional controllers under different states is solved, thereby improving the stability and damping of the power system.

CN122026540APending Publication Date: 2026-05-12NORTHEAST DIANLI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHEAST DIANLI UNIVERSITY
Filing Date
2025-12-30
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Traditional wide-area damping controllers cannot guarantee the control effect of the power system under different operating conditions. The interaction and influence between wind farms and synchronous units are increased, leading to low-frequency oscillations and stability problems in the power grid. Traditional optimization algorithms cannot coordinate the optimization of multiple controllers.

Method used

A multi-machine wide-area damping cooperative control system based on a multi-agent deep deterministic policy gradient algorithm is adopted. By installing wide-area damping controllers in wind turbines and synchronous turbines, multi-agent reinforcement learning is used to train agents to generate optimal control strategies and optimize controller parameters to improve system damping.

Benefits of technology

It improves the system damping of the power system during low-frequency oscillations, reduces the damage caused by oscillations, and enhances the stability and transmission capacity of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122026540A_ABST
    Figure CN122026540A_ABST
Patent Text Reader

Abstract

The invention relates to the field of dynamic control of a power system, discloses a multi-agent reinforcement learning-based multi-machine wide-area damping cooperative control system and method, and aims to solve the problems that the low-frequency oscillation risk of a power grid is aggravated under the high proportion of new energy, the cooperative optimization of a traditional controller is insufficient and the like. An installation site and a feedback signal are determined in combination with participation factor analysis and observability analysis, and a multi-agent depth deterministic strategy gradient algorithm is adopted to optimize parameters of a doubly-fed fan and a synchronous machine controller; through training of the steps of initialization, interaction, judgment, decision making, learning and the like, the intelligent agent can quickly generate an optimal control strategy. The method can effectively improve system damping, rapidly suppress power fluctuation and guarantee stable operation of a power grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic control and stability technology of power systems, and more specifically, to a multi-machine wide-area damping cooperative control system and method based on multi-agent reinforcement learning. Background Technology

[0002] With the increasing interconnection of power grids, the complex network structure of power systems can lead to low-frequency oscillations at any time, which greatly limits the transmission efficiency of power system interconnection lines. Wind power, due to its high output and random fluctuations, and its connection to the grid via power electronic devices, exacerbates the risk of oscillations in power systems containing wind power. As large-capacity wind farms are planned and implemented, they will replace the power generation of traditional power sources, leading to a decrease in the damping effect of power system stabilizers on traditional synchronous machines. Traditional wide-area damping controllers, with fixed parameters, cannot guarantee the control effect of the system under different operating conditions. The interaction and influence between wind farms or wind farm clusters and the connected power system also increase. Among these, the low-frequency oscillations and stability of the power grid that large-scale wind power integration systems may face are key concerns. In particular, low-frequency oscillations in power systems containing large-capacity wind farms may lead to power system instability and affect the safe operation of wind turbines in large-capacity wind farms.

[0003] To alleviate the problem of insufficient damping in the inter-mode oscillation of interconnected power systems, the rapid power modulation characteristics of wind turbines can be utilized to add wide-area damping control. Considering the randomness and intermittency of wind turbine output, wide-area damping control can also be added to synchronous turbines. To better enable wide-area damping coordinated control of wind turbines, especially the widely used doubly-fed induction generators (DFIGs), and synchronous turbines, the coordinated tuning of the parameters of the two types of damping controllers has been a critical challenge. Traditional wide-area signal selection and installation location determination only address a single mode and cannot achieve ideal results under multiple operating modes. The controller input is single-input, which can lead to controller failure in the event of a system fault. Traditional mathematical programming methods are somewhat insufficient for finding the optimal solution in nonlinear and multi-extremum optimization problems, and traditional optimization algorithms can only optimize a single controller, failing to perform coordinated optimization when multiple controllers exist in the system. Therefore, a multi-agent deep deterministic policy gradient algorithm is introduced to solve the multi-controller coordinated optimization problem.

[0004] Wide-area damping control greatly helps improve the dynamic stability and transmission capacity of a system. With the increasing grid-connected capacity of wind turbines, achieving wide-area damping coordinated control between wind turbines and synchronous generators in the power system is of great significance for ensuring the safe and stable operation of large-scale interconnected power systems.

[0005] Application content

[0006] To address the aforementioned problems, this invention proposes a multi-agent wide-area damping cooperative control system and method based on multi-agent reinforcement learning. The trained agents possess the ability to rapidly generate optimal multi-agent wide-area damping cooperative control strategies when the system experiences low-frequency oscillations. This improves system damping during low-frequency oscillations in the power system, thereby reducing the damage caused by oscillations and enhancing the stability of the power system.

[0007] To address the aforementioned problems, this invention proposes a multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning, comprising:

[0008] Build a wide-area damping controller based on wide-area measurement signals.

[0009] Determine the installation location of the wide-area damping controller and the selection of the feedback signal.

[0010] Based on the multi-agent deep deterministic policy gradient algorithm, the observations of the agents and the controller parameters to be optimized are determined for training. The trained agents have the ability to quickly generate the optimal multi-machine wide-area damped cooperative control strategy when the system experiences low-frequency oscillations.

[0011] The multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning includes: drawing on the design method of existing PSS, selecting the wide-area damping control loop based on the comprehensive geometric index of the system's observability / controllability, and installing the wide-area damping controller based on the wide-area measurement signal in the reactive power control link of the rotor-side converter of the doubly fed wind turbine and the excitation link of the synchronous machine.

[0012] To avoid the problem of poor damping effect that may occur when a wide-area damping controller uses local signal input, this invention uses a wide-area measurement signal as input for the wide-area damping controller.

[0013] The multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning includes: using participation factor analysis of power system simulation examples to select the installation location of the wide-area damping controller. Taking the two-zone four-machine system example, the installation location of the wide-area damping controller is determined to be the doubly fed wind turbine DFIG and the synchronous machine G3.

[0014] Using the system observability analysis of power system simulation examples, the feedback signal of the wide-area damping controller is selected. Taking the two-zone four-machine system as an example, the feedback signal of the wide-area damping controller is determined to be the synchronous machine speed difference.

[0015] The multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning includes:

[0016] The controller parameters to be optimized for the wide-area damping controller are determined to be gain K and lead / lag parameters T1 and T3. To distinguish the two wide-area damping controllers to be optimized mentioned in S11, the two sets of controller parameters are divided as follows: the parameters to be optimized installed in the synchronous machine G3 are: K-G3, T1-G3, T3-G3; the parameters to be optimized installed in the doubly fed wind turbine DFIG are: K-DFIG, T1-DFIG, T3-DFIG. The value range of K is (0-100], the value range of T1 and T3 is (0-1], and the remaining controller parameters Tw is 10, T2 and T4 are both 0.5.

[0017] The observed variables in the controller parameter optimization process are determined, and the damping ratio and frequency of the dominant mode that best reflects the system stability are selected in the power system modal simulation analysis.

[0018] Based on the relevant content mentioned in S11 and S31, within the framework of the multi-agent deep deterministic policy gradient algorithm, two agents are selected to optimize two controllers respectively. Agent 1 optimizes the parameters of the wide-area damping controller attached to the synchronous machine G3, and Agent 2 optimizes the parameters of the wide-area damping controller attached to the doubly fed induction generator (DFIG). According to the value range of each parameter mentioned in S31, within this algorithm framework, the two agents learn and continuously optimize together, enabling better collaboration between different controllers. The trained agents have the ability to quickly generate the optimal multi-machine wide-area damping collaborative control strategy when the power system experiences low-frequency oscillations.

[0019] The reinforcement learning network constructed for the multi-machine wide-area damped cooperative control system based on multi-agent reinforcement learning consists of the following components:

[0020] Initialization module: Used to configure the parameters of the reinforcement learning network based on the multi-agent deep deterministic policy gradient algorithm, set the maximum number of interactions per cycle to 32 and the required number of training cycles to 50,000, set the initial values ​​of the controller parameters, and read the damping ratio and frequency of the dominant mode under the initial values;

[0021] Interactive module: After the power system simulation example runs, the dominant mode is selected based on the damping ratio and frequency. After each simulation, the data of the dominant mode is read. Before each simulation, the agent generates controller parameters and passes these controller parameters to the corresponding wide-area damping controller in the power system simulation example, and then the simulation is performed.

[0022] Evaluation module: Based on the damping ratio of the selected dominant modes, the reward value is obtained using the reward function;

[0023] Decision module: Used to input the damping ratio and frequency of the dominant mode into the reinforcement learning network, and use the combination of the two sets of controllers as the output of the reinforcement learning network, so that when the system oscillates at low frequency, the optimal combination of the two sets of wide-area damping controller parameters can be generated quickly.

[0024] Learning module: Determines the damping effect of controller parameters based on the damping ratio of the dominant mode, and updates the network parameters of the reinforcement learning network based on the reward value obtained from the evaluation module.

[0025] The reinforcement learning network based on the multi-agent deep deterministic policy gradient algorithm comprises two agents, each containing two neural networks: a policy neural network and a value neural network. The policy neural network takes as input the damping ratio and frequency of the dominant oscillation mode selected after each simulation and outputs a combination of control parameters. The value neural network takes as input the damping ratio and frequency of the dominant oscillation mode selected after each simulation and the control parameters, and outputs the neural network weights used to update the policy neural network and the value neural network.

[0026] The reward function for a multi-machine wide-area damped cooperative control system based on multi-agent reinforcement learning is as follows:

[0027] If the damping ratio after simulation is between 5% and 10%, the bonus value is 1000.

[0028] If the damping ratio after simulation is between 10% and 20%, the bonus value is 10.

[0029] If the damping ratio after simulation is greater than 20%, the bonus value is -50.

[0030] If the damping ratio after simulation is between 0% and 5%, the bonus value is the corresponding damping ratio multiplied by -10.

[0031] If the damping ratio after simulation is less than 0%, i.e., negative damping occurs, the bonus value is -10000.

[0032] To address the aforementioned issues, this invention also proposes a multi-machine wide-area damped cooperative control method based on multi-agent reinforcement learning, with the specific steps as follows:

[0033] Step 1: Build a wide-area damping controller based on the wide-area measurement system, determine the location of the controller installation using participation factor analysis, determine the controller feedback signal based on system observability, and determine the controller parameters to be optimized subsequently.

[0034] Step 2: Construct the network of the reinforcement learning algorithm based on the deep deterministic policy gradient of multi-agent agents.

[0035] Step 3: Determine the input to the reinforcement learning network as the damping ratio and frequency of the selected dominant modes, and determine the number of interactions with the power system simulation examples during the training process.

[0036] Step 4: Use the constructed reinforcement learning network based on the multi-agent deep deterministic policy gradient algorithm to optimize and train the electrical quantities input into the network to generate a policy model;

[0037] Step 5: Input the damping ratio and frequency of the power system into the above strategy model. The output is the optimal combination of two sets of controller parameters. Transmit the output controller parameter combination to the power system for execution, so that when the system oscillates at low frequency, it can quickly generate the two sets of optimal wide-area damping controller parameters and thus improve the damping of the system.

[0038] Beneficial effects

[0039] This invention uses a multi-agent deep deterministic policy gradient algorithm model as the decision-making body and a simulation example in power system simulation software as the environment. First, a multi-machine wide-area damping control model is built in a two-zone, four-machine system within the power system simulation software, including wide-area damping controllers attached to the wind turbines and synchronous machines. The damping ratio and frequency of the dominant modes selected through modal analysis are used as input to a reinforcement learning network to update its parameters. The output consists of two sets of controller parameters generated by the reinforcement learning network based on the input. Before simulation, the generated controller parameters are applied to the corresponding controller parameters. The agents are trained within the multi-agent deep deterministic policy gradient algorithm framework. The trained agents possess the ability to quickly generate optimal multi-machine wide-area damping cooperative control strategies when the system experiences low-frequency oscillations. This improves the damping under oscillating modes in the power system. This invention provides a new solution for multi-machine wide-area damping cooperative control in power systems, improving the operational stability of the power system. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings involved in the embodiments or the prior art are briefly described below. Obviously, these drawings illustrate several embodiments of the present invention, and those skilled in the art can derive other possible drawings based on these drawings without creative effort. The purpose of the drawings is limited to illustrating specific embodiments and does not limit the scope of the present invention.

[0041] Figure 1 This is a flowchart of the multi-machine wide-area damping cooperative control method based on multi-agent reinforcement learning proposed in this invention;

[0042] Figure 2 This is a two-zone, four-machine system example used for demonstration purposes;

[0043] Figure 3 This is a block diagram of the wide-area damping controller used in the example demonstration;

[0044] Figure 4 This is a block diagram of the additional damping control strategy used in the example demonstration;

[0045] Figure 5 This example demonstrates the active power fluctuation curve during a three-phase short circuit in a tie line under the multi-machine wide-area damping cooperative control method based on multi-agent reinforcement learning proposed in this invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this invention clearer and easier to understand, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the embodiments described in this invention are only one of all embodiments, and this invention is not limited to any particular embodiment; a corresponding embodiment can be selected according to specific circumstances. To avoid obscuring the essence of this invention, well-known methods, processes, flows, components, and circuits are not described in detail.

[0047] The present invention proposes a multi-agent wide-area damping cooperative control method based on multi-agent reinforcement learning, comprising: constructing a wide-area damping controller based on wide-area measurement signals; determining the installation location of the wide-area damping controller and the selection of feedback signals; and training the agent based on the multi-agent deep deterministic policy gradient algorithm to determine the agent's observations and the controller parameters to be optimized. The trained agent has the ability to quickly generate the optimal multi-agent wide-area damping cooperative control strategy when the system experiences low-frequency oscillations.

[0048] A multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning includes: drawing on the design method of existing PSS, selecting the wide-area damping control loop based on the comprehensive geometric index of the system's observability / controllability, and installing the wide-area damping controller based on the wide-area measurement signal in the reactive power control link of the rotor-side converter of the doubly fed wind turbine and the excitation link of the synchronous machine.

[0049] To avoid the problem of poor damping effect that may occur when a wide-area damping controller uses local signal input, this invention uses a wide-area measurement signal as input for the wide-area damping controller.

[0050] A multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning includes: using participation factor analysis of power system simulation examples to select the installation location of the wide-area damping controller. Taking the two-zone four-machine system example, the installation location of the wide-area damping controller is determined to be the doubly fed wind turbine DFIG and the synchronous machine G3.

[0051] Using the system observability analysis of power system simulation examples, the feedback signal of the wide-area damping controller is selected. Taking the two-zone four-machine system as an example, the feedback signal of the wide-area damping controller is determined to be the synchronous machine speed difference.

[0052] A multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning includes:

[0053] The controller parameters to be optimized for the wide-area damping controller are determined to be gain K and lead / lag parameters T1 and T3. To distinguish the two wide-area damping controllers to be optimized mentioned in S11, the two sets of controller parameters are divided as follows: the parameters to be optimized installed in the synchronous machine G3 are: K-G3, T1-G3, T3-G3; the parameters to be optimized installed in the doubly fed wind turbine DFIG are: K-DFIG, T1-DFIG, T3-DFIG. The value range of K is (0-100], the value range of T1 and T3 is (0-1], and the remaining controller parameters Tw is 10, T2 and T4 are both 0.5.

[0054] The observed variables in the controller parameter optimization process are determined, and the damping ratio and frequency of the dominant mode that best reflects the system stability are selected in the power system modal simulation analysis.

[0055] Based on the relevant content mentioned in S11 and S31, within the framework of the multi-agent deep deterministic policy gradient algorithm, two agents are selected to optimize two controllers respectively. Agent 1 optimizes the parameters of the wide-area damping controller attached to the synchronous machine G3, and Agent 2 optimizes the parameters of the wide-area damping controller attached to the doubly fed induction generator (DFIG). According to the value range of each parameter mentioned in S31, within this algorithm framework, the two agents learn and continuously optimize together, enabling better collaboration between different controllers. The trained agents have the ability to quickly generate the optimal multi-machine wide-area damping collaborative control strategy when the power system experiences low-frequency oscillations.

[0056] The reinforcement learning network constructed for the multi-machine wide-area damped cooperative control system based on multi-agent reinforcement learning consists of the following components:

[0057] Initialization module: Used to configure the parameters of the reinforcement learning network based on the multi-agent deep deterministic policy gradient algorithm, set the maximum number of interactions per cycle to 32 and the required number of training cycles to 50,000, set the initial values ​​of the controller parameters, and read the damping ratio and frequency of the dominant mode under the initial values;

[0058] Interactive module: After the power system simulation example runs, the dominant mode is selected based on the damping ratio and frequency. After each simulation, the data of the dominant mode is read. Before each simulation, the agent generates controller parameters and passes these controller parameters to the corresponding wide-area damping controller in the power system simulation example, and then the simulation is performed.

[0059] Evaluation module: Based on the damping ratio of the selected dominant modes, the reward value is obtained using the reward function;

[0060] Decision module: Used to input the damping ratio and frequency of the dominant mode into the reinforcement learning network, and use the combination of the two sets of controllers as the output of the reinforcement learning network, so that when the system oscillates at low frequency, the optimal combination of the two sets of wide-area damping controller parameters can be generated quickly.

[0061] Learning module: Determines the damping effect of controller parameters based on the damping ratio of the dominant mode, and updates the network parameters of the reinforcement learning network based on the reward value obtained from the evaluation module.

[0062] The reinforcement learning network based on the multi-agent deep deterministic policy gradient algorithm comprises two agents, each containing two neural networks: a policy neural network and a value neural network. The policy neural network takes as input the damping ratio and frequency of the dominant oscillation mode selected after each simulation and outputs a combination of control parameters. The value neural network takes as input the damping ratio and frequency of the dominant oscillation mode selected after each simulation and the control parameters, and outputs the neural network weights used to update the policy neural network and the value neural network.

[0063] The reward function for a multi-machine wide-area damped cooperative control system based on multi-agent reinforcement learning is as follows:

[0064] If the damping ratio after simulation is between 5% and 10%, the bonus value is 1000.

[0065] If the damping ratio after simulation is between 10% and 20%, the bonus value is 10.

[0066] If the damping ratio after simulation is greater than 20%, the bonus value is -50.

[0067] If the damping ratio after simulation is between 0% and 5%, the bonus value is the corresponding damping ratio multiplied by -10.

[0068] If the damping ratio after simulation is less than 0%, i.e., negative damping occurs, the bonus value is -10000.

[0069] To address the aforementioned issues, this invention also proposes a multi-machine wide-area damped cooperative control method based on multi-agent reinforcement learning, with the specific steps as follows:

[0070] Step 1: Build a wide-area damping controller based on the wide-area measurement system, determine the location of the controller installation using participation factor analysis, determine the controller feedback signal based on system observability, and determine the controller parameters to be optimized subsequently.

[0071] Step 2: Construct the network of the reinforcement learning algorithm based on the deep deterministic policy gradient of multi-agent agents.

[0072] Step 3: Determine the input to the reinforcement learning network as the damping ratio and frequency of the selected dominant modes, and determine the number of interactions with the power system simulation examples during the training process.

[0073] Step 4: Use the constructed reinforcement learning network based on the multi-agent deep deterministic policy gradient algorithm to optimize and train the electrical quantities input into the network to generate a policy model;

[0074] Step 5: Input the damping ratio and frequency of the power system into the above strategy model. The output is the optimal combination of two sets of controller parameters. Transmit the output controller parameter combination to the power system for execution, so that when the system oscillates at low frequency, it can quickly generate the two sets of optimal wide-area damping controller parameters and thus improve the damping of the system.

[0075] In this embodiment, the specific steps are as follows:

[0076] 1. Construct a wide-area damping controller based on a wide-area measurement system, determine the location for installing the controller using participation factor analysis, determine the controller's feedback signal based on system observability, and determine the controller parameters to be optimized subsequently.

[0077] 2. Determine the observation space and action space of the reinforcement learning network. The observation space consists of the frequency and damping ratio of the dominant mode selected after modal simulation, and the action space consists of the combination of two sets of controller parameters generated by the two agents taking values ​​within the range of each controller parameter.

[0078] 3. Determine the number of interactions with the power system simulation examples during the training process. Before each simulation, the controller parameters generated by the agent in the reinforcement learning network are first applied to the corresponding controller, and then the simulation is performed. The damping ratio and frequency of the dominant mode obtained after the simulation are used to update the reinforcement learning network, so that the generated control strategy is more in line with the system.

[0079] 4. Construct a reinforcement learning network based on a multi-agent deep deterministic policy gradient algorithm. The two agents in the reinforcement learning network have the same network structure, which contains two neural networks, namely a policy network and a value network.

[0080] The input layer of the policy neural network takes the frequency and damping ratio of the dominant mode selected after each simulation as input, and the output layer outputs a combination of controller parameters.

[0081] The input layer of the value neural network takes as input the frequency and damping ratio of the dominant mode selected after the real simulation and the combination of controller parameters, and the output layer outputs the neural network weights used to update the policy neural network and the value neural network.

[0082] 5. Develop a reward function based on the damping ratio of the dominant mode selected after simulation:

[0083] If the damping ratio after simulation is between 5% and 10%, the bonus value is 1000.

[0084] If the damping ratio after simulation is between 10% and 20%, the bonus value is 10.

[0085] If the damping ratio after simulation is greater than 20%, the bonus value is -50.

[0086] If the damping ratio after simulation is between 0% and 5%, the bonus value is the corresponding damping ratio multiplied by -10.

[0087] If the damping ratio after simulation is less than 0%, i.e., negative damping occurs, then the bonus value is -10000.

[0088] 6. Strategy

[0089] A policy is a mapping from state to action, which refers to a distribution on the action set given a state, that is, assigning an action probability to each state s.

[0090] 7. Before the simulation interaction, the power system is in an initial state s0. The agent issues an action a0 to the power system according to the policy distribution π, determines the damping ratio and frequency of the dominant oscillation mode in the next stage, and transmits the action command to the simulation and environment interaction. The environmental state changes and is fed back to the agent as the state s1 of the next decision stage. The reward r0 is calculated, and this process is repeated until the last decision stage. In this process, the two agents share the state and execute in a distributed manner, and learn together to continuously optimize.

[0091] The above process is solved using a multi-agent deep deterministic strategy gradient algorithm to obtain the optimal combination of parameters for the multi-machine wide-area damping controller.

[0092] like Figure 1 As shown in the flowchart, the multi-machine wide-area damped cooperative control method based on multi-agent reinforcement learning includes the following steps:

[0093] S1. Execute the initialization module, configure the reinforcement learning parameters based on multi-agent deep deterministic policy gradient; set the initial parameters of the controller, set the maximum number of interactions per loop to 32, and the number of training loops to 50,000.

[0094] S2. Execute the interaction module. In the first interaction, the initial controller parameters are transmitted to the power system to perform modal simulation and read the damping ratio and frequency of the dominant mode. After that, the controller parameters generated by the decision module are transmitted to the controller to perform simulation until the training is completed.

[0095] S3. The execution evaluation module reads the damping ratio of the dominant mode in the simulation after each application of the controller parameters, and determines the reward based on the reward function to determine the quality of this action.

[0096] S4, the execution decision module, takes the damping ratio of the dominant mode read after each simulation as the input of the agent, and outputs a combination of two controller parameters. This combination of controller parameters is then transmitted to the interaction module and sent to the power system to perform the simulation.

[0097] S5, Execution learning module: The two agents update the parameters of their own neural networks based on the reward value obtained in S5. The two agents learn together and work together until they learn a strategy that meets the current working conditions.

[0098] S6. Determine if the maximum number of interactions for a single loop has been reached. If not, repeat steps S2 to S5; otherwise, end the current loop.

[0099] S7. Determine whether the required number of iterations has been reached. If not, return to step S1; otherwise, automatically save the trained model and exit the program.

[0100] S8. The application stage involves calling the trained reinforcement learning network to repeat steps S1 to S4 to improve the system damping under low-frequency oscillations of the power system.

[0101] Example demonstration:

[0102] To demonstrate the effect, a structure such as Figure 2 The simulation example of the two-zone four-machine system shown here has the same parameters as the basic example parameters, which will not be elaborated on here.

[0103] Figure 3 The diagram shows the structure of a wide-area damping controller. The input is the speed difference of the synchronous machine. The controller parameters to be optimized are K, T1, and T3, while the other parameters are constants.

[0104] Figure 4 The diagram shown is a block diagram of the additional damping control strategy. Figure 3 The wide-area damping controller is attached to the reactive power circuit of the rotor-side converter of the doubly fed wind turbine and the excitation circuit of the synchronous machine G3.

[0105] To verify the correctness of the proposed method, a three-phase short-circuit fault was set on the transmission line of the four-machine system in Zone 2 at 1 second, lasting for 0.1 seconds, with a simulation duration of 60 seconds, and the change in active power of the transmission line was observed.

[0106] Figure 5The comparison shows the control effect of the controller parameters. The blue line represents the controller without wide-area damping, the green line represents the controller parameters obtained using the particle swarm optimization algorithm, and the red line represents the controller parameters obtained using this method. It can be clearly seen that the controller parameters optimized by this method can suppress transmission line power fluctuations more quickly and the effect is more obvious, which can illustrate the effectiveness and correctness of this method.

[0107] Those skilled in the art should understand that the above embodiments are merely illustrative of the content of this disclosure and do not limit its scope. The system capacity, voltage, line parameters, etc., shown may vary depending on the specific circumstances of the power electronic grid-connected generator set and its grid connection. Based on this disclosure, those skilled in the art can make other changes or adjustments, and these changes still fall within the scope of this disclosure.

Claims

1. A multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning, characterized in that, include: Step S1: Build a wide-area damping controller based on wide-area measurement signals; Step S2: Determine the installation location of the wide-area damping controller and the selection of the feedback signal; Step S3: Based on the multi-agent deep deterministic policy gradient algorithm, determine the agent's observations and the controller parameters to be optimized for training. The trained agent has the ability to generate a multi-machine wide-area damping cooperative control strategy when the system experiences low-frequency oscillations.

2. The multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning according to claim 1, characterized in that, Step S1 specifically includes: Step S11: Select a wide-area damping control loop based on the comprehensive geometric index of the observability and controllability of the system. Install the wide-area damping controller based on the wide-area measurement signal in the reactive power control link of the rotor-side converter of the doubly fed wind turbine and the excitation link of the synchronous machine. Step S12: The present invention uses a wide-area measurement signal as input for a wide-area damping controller.

3. The multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning according to claim 1, characterized in that, Step S2 specifically includes: Step S21: Using the participation factor analysis of the power system simulation example, select the installation location of the wide-area damping controller. Taking the two-zone four-machine system example, the installation location of the wide-area damping controller is determined to be the doubly fed wind turbine DFIG and the synchronous machine G3. Step S22: Using the system observability analysis of the power system simulation example, select the feedback signal of the wide-area damping controller. Taking the two-zone four-machine system example, the feedback signal of the wide-area damping controller is determined to be the synchronous machine speed difference.

4. The multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning according to claim 1, characterized in that, Step S3 specifically includes: Step S31: Determine the controller parameters to be optimized for the wide-area damping controller as gain K and lead / lag parameters T1 and T3. To distinguish the two wide-area damping controllers to be optimized mentioned in S11, the two sets of controller parameters are divided as follows: the parameters to be optimized installed in the synchronous machine G3 are: K-G3, T1-G3, T3-G3; the parameters to be optimized installed in the doubly fed wind turbine DFIG are: K-DFIG, T1-DFIG, T3-DFIG. The value range of K is (0-100], the value range of T1 and T3 is (0-1], and the remaining controller parameters Tw is 10, T2 and T4 are both 0.

5. Step S32: Determine the observed variables in the controller parameter optimization process, and select the damping ratio and frequency of the dominant mode that best reflects the system stability in the power system modal simulation analysis; Step S32: Based on the relevant content mentioned in S11 and S31, within the framework of the multi-agent deep deterministic policy gradient algorithm, two agents are selected to optimize two controllers respectively. Agent 1 optimizes the parameters of the wide-area damping controller attached to the synchronous machine G3, and Agent 2 optimizes the parameters of the wide-area damping controller attached to the doubly fed wind turbine DFIG. According to the value range of each parameter mentioned in S31, within this algorithm framework, the two agents learn together and continuously optimize, enabling better collaboration between different controllers. The trained agents have the ability to quickly generate the optimal multi-machine wide-area damping collaborative control strategy when the power system experiences low-frequency oscillations.

5. The multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning according to claim 1, characterized in that, The constructed reinforcement learning network consists of the following components: Initialization module: Used to configure the parameters of the reinforcement learning network based on the multi-agent deep deterministic policy gradient algorithm, set the maximum number of interactions per cycle and the number of training cycles required, set the initial values ​​of the controller parameters, and read the damping ratio and frequency of the dominant mode under the initial values; Interactive module: After the power system simulation example runs, the dominant mode is selected based on the damping ratio and frequency. After each simulation, the data of the dominant mode is read. Before each simulation, the agent generates controller parameters and passes these controller parameters to the corresponding wide-area damping controller in the power system simulation example, and then performs the simulation. Evaluation module: Based on the damping ratio of the selected dominant modes, the reward value is obtained using the reward function; Decision module: Used to input the damping ratio and frequency of the dominant mode into the reinforcement learning network, and use the combination of the two sets of controllers as the output of the reinforcement learning network, so that when the system oscillates at low frequency, the optimal combination of the two sets of wide-area damping controller parameters can be generated quickly. Learning module: Determines the damping effect of controller parameters based on the damping ratio of the dominant mode, and updates the network parameters of the reinforcement learning network based on the reward value obtained from the evaluation module.

6. The multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning according to claim 1, characterized in that, The reinforcement learning network based on the multi-agent deep deterministic policy gradient algorithm comprises two neural networks: a policy neural network and a value neural network. The input of the policy neural network is the damping ratio and frequency of the dominant oscillation mode selected after each simulation, and the output is a combination of control parameters. The input of the value neural network is the damping ratio and frequency of the dominant oscillation mode selected after each simulation and the combination of the control parameters, and the output is the neural network weights used to update the policy neural network and the value neural network.

7. The multi-machine wide-area damping cooperative control system based on multi-agent reinforcement learning according to claim 5, characterized in that: The reward function is set as follows: If the damping ratio after simulation is between 5% and 10%, the bonus value is 1000. If the damping ratio after simulation is between 10% and 20%, the bonus value is 10. If the damping ratio after simulation is greater than 20%, the bonus value is -50. If the damping ratio after simulation is between 0% and 5%, the bonus value is the corresponding damping ratio multiplied by -10. If the damping ratio after simulation is less than 0%, i.e., negative damping occurs, then the bonus value is -10000.

8. The method for a wide-area damped cooperative control system based on multi-agent reinforcement learning according to any one of claims 1 to 7, characterized in that, include: Step 1: Build a wide-area damping controller based on the wide-area measurement system, determine the location of the controller installation using the participation factor analysis method, determine the controller feedback signal based on the system observability, and determine the controller parameters to be optimized in the future. Step 2: Construct a network based on a strong learning algorithm using multi-agent deep deterministic policy gradients; Step 3: Determine the input to the reinforcement learning network as the damping ratio and frequency of the selected dominant modes, and determine the number of interactions with the power system simulation examples during the training process; Step 4: Use the constructed reinforcement learning network based on the multi-agent deep deterministic policy gradient algorithm to optimize and train the electrical quantities input into the network to generate a policy model; Step 5: Input the damping ratio and frequency of the power system into the above strategy model. The output is the optimal combination of two sets of controller parameters. Transmit the output controller parameter combination to the power system for execution, so that when the system oscillates at low frequency, it can quickly generate the two sets of optimal wide-area damping controller parameters and thus improve the damping of the system.