An Oscillation Coordination Suppression Method for a New Energy Grid-Connected System Based on the MA-DDPG Algorithm
The multi-agent deep reinforcement learning model constructed through the multi-agent depth deterministic strategy gradient (MA-DDPG) algorithm coordinates the new energy to send multiple inverters in the system through flexible direct transmission, solving the wide-band oscillation problem and achieving stable operation and efficient control of the system.
Patent Information
- Application Number
- CN202411589511.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-11-08
AI Technical Summary
The dynamic interaction between multiple inverters in the flexible direct transmission system is prone to trigger broadband oscillation, and existing control methods are difficult to coordinate the oscillation suppression strategies of multiple inverters, resulting in unstable system operation.
The multi-agent depth deterministic strategy gradient (MA-DDPG) algorithm is used to construct a multi-agent deep reinforcement learning model. By adjusting the parameters of the additional damping controller online, an optimal control strategy for the new energy through flexible direct delivery system is generated to achieve coordinated control between multiple inverters.
It effectively suppresses the wide-frequency oscillation in the flexible direct transmission system of the new energy system, improves the stability and adaptability of the system, and can maintain the stable operation of the system under various complex operating conditions.
Smart Images

Figure CN119543204B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer application technologies, and particularly to a method for suppressing oscillation coordination in a new energy grid-connected system based on the MA-DDPG algorithm. Background Art
[0002] With the growing global demand for clean energy, the access ratio of renewable energy sources such as wind energy and solar energy has been increasing year by year. Due to its advantages in long-distance and large-scale power transmission, flexible DC transmission technology has gradually become an important technology for delivering renewable energy to load centers. However, the interaction of multiple converters in the new energy transmission system via flexible DC is likely to trigger broadband oscillations, which cover the low-frequency to high-frequency range and may seriously affect the operation stability of the system.
[0003] Currently, the oscillation suppression methods in the prior art mainly include the following several solutions:
[0004] 1. Single converter control: Most oscillation suppression methods are optimized and designed for a single converter and cannot effectively cope with the dynamic interaction and oscillation between multiple converters.
[0005] 2. Fixed parameter control: Traditional supplementary damping control methods rely on pre-set fixed parameters, which cannot be automatically adjusted when the system operating conditions change, resulting in poor control effects.
[0006] 3. Lack of coordinated control between converters: Due to the dynamic interaction of multiple converters in the new energy transmission system via flexible DC, adjusting the control parameters of a certain converter alone may cause new oscillations in other converters, and it is difficult for existing control methods to coordinate the oscillation suppression strategies of multiple converters.
[0007] With the increasing penetration rate of renewable energy in the power system, the system operating conditions become more and more complex, and the existing single-agent and fixed parameter control methods are difficult to meet the dynamically changing requirements. Therefore, an intelligent control method that can coordinate multiple converters is needed to ensure the stable operation of the system under various working conditions. Summary of the Invention
[0008] An embodiment of the present invention provides a method for suppressing oscillation coordination in a new energy grid-connected system based on the multi-agent deep deterministic policy gradient (MA-DDPG) algorithm.
[0009] To achieve the above object, the present invention adopts the following technical solutions.
[0010] A method for suppressing oscillation coordination in a new energy grid-connected system based on the MA-DDPG algorithm includes:
[0011] Obtain the system parameters of the new energy transmission system via flexible DC;
[0012] Construct a multi-agent deep reinforcement learning model based on the system parameters of the new energy transmission system via flexible DC, and define the state space set, action space set and reward function of the agents;
[0013] Use the MA-DDPG algorithm to train the multi-agent deep reinforcement learning model until convergence, and obtain the trained multi-agent deep reinforcement learning model;
[0014] Online adjust the parameters of the additional damping controller through the trained multi-agent deep reinforcement learning model, generate the optimal control strategy for the new energy transmission system via flexible DC, and use the optimal control strategy to suppress system oscillations.
[0015] Preferably, the obtaining of the system parameters of the new energy transmission system via flexible DC includes:
[0016] Set the new energy transmission system via flexible DC to include: a new energy rectifier VSC1, a flexible DC sending-end inverter VSC2 and a receiving-end converter VSC3. VSC1 adopts VdcIq control, VSC2 adopts Vf control, and VSC3 adopts VdcIq control. Install additional damping controllers on VSC1 and VSC2, and obtain the system parameters of the new energy transmission system via flexible DC. The system parameters include dynamic parameters and electrical quantity data, and the electrical quantity data includes DC-side voltage and current and AC-side d-axis voltage data.
[0017] Preferably, the constructing of the multi-agent deep reinforcement learning model based on the system parameters of the new energy transmission system via flexible DC and defining the state space set, action space set and reward function of the agents includes:
[0018] Construct a multi-agent deep reinforcement learning model based on the system parameters of the new energy transmission system via flexible DC. The multi-agent deep reinforcement learning model includes two agents, Agent1 and Agent2. Agent1 is connected to VSC1 and serves as the additional damping controller of VSC1, and Agent2 is connected to VSC2 and serves as the additional damping controller of VSC2;
[0019] The state space of each agent includes: control reference value, deviation quantity and system operation state. The state space of agent Agent1 is defined as:
[0020]
[0021] where v dc1 is the DC voltage of VSC1, Δv dc1 (t) is the deviation of v dc1 from its reference value, and i dc1 (t) is the input DC current of the new energy output;
[0022] The state space of Agent2 is defined as:
[0023]
[0024] where v Ld2 is the d-axis output voltage of VSC2, and Δv Ld2 is the deviation of v Ld2 from its reference value, and Ld2 is the reference value of v
[0025] The action space of each agent is the parameters of the additional damping controller;
[0026] The action space of Agent VSC1 is:
[0027] The action space of Agent VSC2 is:
[0028] where K is the gain, and T 1 and T 2 are time constants;
[0029] Design a data-driven reward function:
[0030] r t = -|Δv dc1 (t)| - |Δv Ld2 (t)|
[0031] where Δv dc1 (t) is the deviation of the DC-side voltage of VSC1 from its reference value, and Δv Ld2 (t) is the deviation of the d-axis output voltage of VSC2 from its reference value.
[0032] Preferably, the multi-agent deep reinforcement learning model is trained using the MA-DDPG algorithm until convergence to obtain a trained multi-agent deep reinforcement learning model, including:
[0033] In the collaborative training of the multi-agent deep reinforcement learning model, the MA-DDPG algorithm is used to achieve the coordinated control between the converters. The inputs of the training process include the state space information of each agent: for Agent1 (VSC1), it includes the DC voltage v dc1 and its deviation Δv dc1 , the input DC current i dc1 and the output active power P 1 ; for Agent2 (VSC2), it includes the d-axis output voltage v Ld2 and its deviation Δv Ld2and the output active power P 2 , the output of the training is the action space of each agent, that is, the parameters of the additional damping controller, including the parameter K of the VSC1 controller VSC1 (t), T 1 VSC1 (t), and the parameter K of the VSC2 controller VSC2 (t), T 1 VSC2 (t),
[0034] The MA-DDPG algorithm is trained using a centralized training and decentralized execution framework. During the centralized training phase, the two agents share the same experience replay pool for storing interactive data samples (s t , a t , r t , s t+1 ), and also share a reward function r based on electrical quantity measurement values. Each agent contains an Actor network and a Critic network. The Actor network is responsible for outputting control actions based on the current state, and the Critic network is responsible for evaluating the value of the current state-action pair;
[0035] During the training process, the Actor and Critic networks of the two agents update the network parameters through interaction. When updating, the Actor network will consider the action influence of the other agent, and when evaluating, the Critic network takes the states and actions of the two agents as inputs;
[0036] A shared data-driven reward function is used to coordinately train multiple converters,
[0037] During the training process of the multi-agent deep reinforcement learning model, the agents of VSC1 and VSC2 share the same data-driven reward function r t , and this reward function is constructed based on real-time measurement data, and the calculation formula is:
[0038] r t = -|Δv dc1 (t)| - |Δv Ld2 (t)|
[0039] where, Δv dc1 (t) represents the DC voltage deviation of VSC1 at time t, and Δv Ld2 (t) represents the d-axis voltage deviation of VSC2 at time t. The negative sign indicates that the smaller the deviation, the larger the reward value;
[0040] The training of the multi-agent deep reinforcement learning model needs to meet the following conditions simultaneously: First, the cumulative reward values of both agents reach convergence, that is, the change range of the cumulative reward values in consecutive training rounds is less than a preset threshold; Second, the oscillation suppression effect of the system under different operating conditions meets the standard; Third, the control parameters output by both agents remain relatively stable in consecutive training rounds, and the change range of the parameters is less than the set threshold;
[0041] When the above conditions are met simultaneously, the training of the multi-agent deep reinforcement learning model ends, and a trained multi-agent deep reinforcement learning model is obtained. Preferably, the method of online adjusting the parameters of the additional damping controller by the trained multi-agent deep reinforcement learning model to generate the optimal control strategy of the new energy transmission system via flexible DC transmission, and using the optimal control strategy for system oscillation suppression includes:
[0042] The input data of the trained multi-agent deep reinforcement learning model includes the state space information of each agent: For Agent1 (VSC1), it includes the DC voltage v dc1 and its deviation Δv dc1 , the input DC current i dc1 and the output active power P 1 ; For Agent2 (VSC2), it includes the d-axis output voltage v Ld2 and its deviation Δv Ld2 and the output active power P 2 ;
[0043] The multi-agent deep reinforcement learning model preprocesses the input data by normalization to make the input data meet the numerical range set during training, and then inputs the input data into the corresponding Actor networks of VSC1 and VSC2 respectively. Based on the current state, the Actor network obtains its respective optimal control strategy through forward calculation. The optimal control strategy includes the parameters K VSC1 , T 1 VSC1 , of the VSC1 controller, and the parameters K VSC2 , T 1 VSC2 ,
[0044] It can be seen from the technical solutions provided by the embodiments of the present invention above that the new energy transmission system via flexible DC transmission in the present invention can coordinate the operations of multiple converters to achieve global optimal control of the system. Compared with the existing single control method, the present invention has stronger adaptability and better oscillation suppression effect, and can maintain the stable operation of the system under various complex working conditions.
[0045] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent from the following description, or will be learned through the practice of the present invention. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 It is a processing flowchart of a method for adaptively adjusting additional damping control parameters provided for an embodiment of the present invention;
[0048] Figure 2 It is a system topology diagram of a system for transmitting new energy through a flexible direct current provided for an embodiment of the present invention;
[0049] Figure 3 It is a network structure diagram of a TD3 algorithm provided for an embodiment of the present invention;
[0050] Figure 4 It is a comparison diagram of oscillation suppression effects provided for an embodiment of the present invention. Detailed Embodiments
[0051] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0052] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used here may include wireless connection or coupling. The term "and / or" used here includes any unit and all combinations of one or more related listed items.
[0053] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless defined as here.
[0054] For the convenience of understanding the embodiments of the present invention, the following will further explain with several specific embodiments in conjunction with the drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.
[0055] The present invention proposes a method for suppressing oscillation coordination in a new energy grid-connected system based on the MA-DDPG algorithm, a coordinated control method based on multi-agent deep reinforcement learning, which is used to suppress the broadband oscillation during the transmission of renewable energy through a flexible DC transmission system. This method adopts a centralized training / decentralized execution framework to achieve coordinated control among multiple converters and improve the flexibility of oscillation suppression.
[0056] The processing flow of a method for suppressing oscillation coordination in a new energy grid-connected system based on the MA-DDPG algorithm provided by an embodiment of the present invention is as Figure 1 shown, and includes the following processing steps:
[0057] Step S10: Obtain the system parameters of the new energy transmitted through the flexible DC system.
[0058] Step S20: Based on the system parameters of the new energy transmitted through the flexible DC system, construct a multi-agent deep reinforcement learning model, and define the state space set, action space set and reward function of the agent.
[0059] Step S30: Use the MA-DDPG algorithm to train the multi-agent deep reinforcement learning model until convergence.
[0060] Step S40: Online adjust the parameters of the additional damping controller through the trained multi-agent deep reinforcement learning model, generate the optimal control strategy for the new energy transmitted through the flexible DC system, and use the optimal control strategy to suppress system oscillation.
[0061] The above step S10 includes: The structure of a new energy transmitted through the flexible DC system provided by an embodiment of the present invention is as Figure 2 shown. This system includes: a new energy rectifier VSC1, a flexible DC sending-end inverter VSC2 and a receiving-end converter VSC3. VSC1 adopts VdcIq control, VSC2 adopts Vf control, and VSC3 adopts VdcIq control. An additional damping controller is installed on VSC1 and VSC2 to suppress the broadband oscillation of the system.
[0062] Obtain the system parameters of the above new energy flexible DC transmission system. The system parameters include dynamic parameters and electrical quantity data. The above electrical quantity data includes the state space information of each agent: for Agent1 (VSC1), it includes the DC voltage v dc1 and its deviation Δv dc1 , the input DC current i dc1 and the output active power P 1 ; for Agent2 (VSC2), it includes the d-axis output voltage v Ld2 and its deviation Δv Ld2 and the output active power P 2 .
[0063] The above step S20 includes:
[0064] Based on the system parameters of the above new energy flexible DC transmission system, construct a multi-agent deep reinforcement learning model. The multi-agent deep reinforcement learning model includes two agents (Agent1 and Agent2). Agent1 is connected to VSC1 and serves as an additional damping controller for VSC1. Agent2 is connected to VSC2 and serves as an additional damping controller for VSC2. The two agents (Agent1 and Agent2) are trained using the same set of reward functions to achieve coordinated control.
[0065] The state space of each agent includes: control reference value, deviation, and system operating state. Among them, the state space of agent Agent1 is defined as:
[0066] s t VSC1 ={v dc1 (t), Δv dc1 (t), i dc1 (t)}
[0067] where v dc1 is the DC voltage of VSC1, Δv dc1 (t) is the deviation of v dc1 from its reference value, and i dc1 (t) is the input DC current of the new energy output.
[0068] The state space of agent Agent2 is defined as:
[0069]
[0070] where v Ld2 is the d-axis output voltage of VSC2, Δv Ld2 is the deviation of v Ld2 from its reference value, is vLd2 Reference value
[0071] The action space of each agent is the parameter of the additional damping controller
[0072] The action space of agent VSC1 is:
[0073] The action space of agent VSC2 is:
[0074] where K is the gain, T 1 and T 2 are time constants
[0075] To achieve the coordinated control of the two agents, the present invention designs a data-driven reward function:
[0076] r t =-|Δv dc1 (t)|-|Δv Ld2 (t)|
[0077] where Δv dc1 (t) is the deviation between the DC-side voltage of VSC1 and its reference value, and Δv Ld2 (t) is the deviation between the d-axis output voltage of VSC2 and its reference value. This reward function is based on measured data and belongs to a data-driven reward function
[0078] The above step S30 includes: The present invention uses the multi-agent deep deterministic policy gradient (MA-DDPG) algorithm to train the multi-agent deep reinforcement learning model. The above-shared data-driven reward function is used to perform coordinated training on multiple converters
[0079] Before describing the algorithm steps, first explain the mathematical symbols used:
[0080] - The actor network of the i-th agent maps the state to an action
[0081] - The critic network of the i-th agent evaluates the value of the state-action pair
[0082] - The parameters of the actor network of the i-th agent
[0083] - The parameters of the critic network of the i-th agent
[0084] -N t : Exploration noise, usually generated using the Ornstein-Uhlenbeck process
[0085] The training process of the above multi-agent deep reinforcement learning model includes:
[0086] 1. Initialization:
[0087] - Initialize the actor network for each agent i and the critic network
[0088] - Initialize the target networks μ i ′ and Q i ′;
[0089] - Initialize the experience replay buffer D;
[0090] 2. For each episode:
[0091] - Initialize the random process N for action exploration;
[0092] - Receive the initial observation state s 1 ;
[0093] For each time step t:
[0094] - For each agent i, select an action a according to the current policy i,t = μ i (s t ) + N t ;
[0095] - Execute the action a t = (a 1,t ,..., a N,t ), observe the reward r t and the new state s t+1 s_{t+1};
[0096] - Store the transition (s t , a t , r t , s t+1 ) into D;
[0097] - Randomly sample a batch of transitions (s, a, r, s') from D;
[0098] - For each agent i:
[0099] - Calculate the target value
[0100] where r i is the immediate reward, γ is the discount factor, Q i ′ is the target critic network, μ 1 ′(s′),..., μ′ N(s′) is the target action for all agents;
[0101] - Update the critic network to minimize the loss:
[0102] Use MSE (mean - square error) as the loss function, and the goal is to minimize the gap between the predicted Q - value and the target Q - value.
[0103] - Update the actor network using policy gradient:
[0104]
[0105] - Soft - update the target network of each agent:
[0106] where τ is the mixing parameter for soft - update, usually taking a relatively small value.
[0107] 3. End the loop to obtain the trained multi - agent deep reinforcement learning model. The training of the multi - agent deep reinforcement learning model terminates when the following conditions are simultaneously met: First, the cumulative reward values of both agents reach convergence, that is, the change range of the cumulative reward values in consecutive multiple training rounds is less than a preset threshold; Second, the oscillation suppression effect of the system under different operating conditions meets the standard, specifically manifested as the fluctuation amplitude of the DC voltage v dc1 of VSC1 and the d - axis voltage v Ld2 of VSC2 are both less than 2% of the rated value and can return to stability within 0.5 seconds; Third, the control parameters output by both agents remain relatively stable in consecutive multiple training rounds, and the change range of the parameters is less than the set threshold. When the above conditions are simultaneously met, it can be considered that the multi - agent deep reinforcement learning model is trained. At this time, save the trained model parameters for subsequent online control.
[0108] The above step S40 includes: online application.
[0109] After training, the two agents can output adaptive additional damping control parameters for VSC1 and VSC2 respectively according to different operating conditions, achieving decentralized execution.
[0110] 7. Control effect
[0111] As verified by simulation as Figure 4 shown, the multi - agent deep reinforcement learning method proposed by the present invention can more effectively suppress system oscillation and improve the stability and adaptability of the system compared with the traditional fixed - parameter method and the DQN method under different frequencies and different operating conditions.
[0112] In summary, the present invention realizes coordinated control based on multi-agent deep reinforcement learning, effectively suppresses broadband oscillations in the new energy transmission system via VSC-HVDC, and improves the stability of the system.
[0113] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0114] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present invention, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0115] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0116] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for coordinated suppression of oscillation of a new energy grid-connected system based on the MA-DDPG algorithm, characterized in that: include: Obtain system parameters of the new energy flexible direct current transmission system; Based on the system parameters of the new energy flexible direct current transmission system, a multi-agent deep reinforcement learning model is constructed to define the state space set, action space set and reward function of the agent; Use the MA-DDPG algorithm to train the multi-agent deep reinforcement learning model until convergence to obtain a trained multi-agent deep reinforcement learning model; The parameters of the additional damping controller are adjusted online through the trained multi-agent deep reinforcement learning model to generate the optimal control strategy for the new energy flexible direct current transmission system, and the optimal control strategy is used to suppress system oscillations. The system parameters of the flexible direct current transmission system for new energy are obtained, including: A new energy flexible direct current transmission system is provided, including: a new energy rectifier VSC1, a flexible direct current sending-end inverter VSC2 and a receiving-end converter VSC3. VSC1 adopts VdcIq control, VSC2 adopts Vf control, and VSC3 adopts VdcIq control. Additional damping controllers are installed on VSC1 and VSC2 to obtain system parameters of the new energy flexible direct current transmission system. The system parameters include dynamic parameters and electrical quantity data. The electrical quantity data include DC side voltage and current and AC side d-axis voltage data.
2. The method according to claim 1, characterized in that The multi-agent deep reinforcement learning model is constructed based on the system parameters of the new energy flexible direct current transmission system, and the state space set, action space set and reward function of the agent are defined, including: A multi-agent deep reinforcement learning model is constructed based on the system parameters of the new energy flexible direct current transmission system. The multi-agent deep reinforcement learning model includes two agents, Agent 1 and Agent 2. Agent 1 is connected to VSC 1 and serves as an additional damping controller of VSC 1. Agent 2 is connected to VSC 2 and serves as an additional damping controller of VSC 2. The state space of each agent includes: control reference value, deviation and system operation status. The state space of agent Agent1 is defined as: s t VSC1 ={v dc1 (t),Δv dc1 (t),i dc1 (t)} where v dc1 is the DC voltage of VSC1, Δv dc1 (t) is v dc1 Deviation from its reference value, i dc1 (t) Input DC current for renewable energy output; The state space of Agent2 is defined as: where v Ld2 is the d-axis output voltage of VSC2, Δv Ld2 v Ld2 Deviation from its reference value, v Ld2 Reference value of The action space of each agent is the parameters of the additional damping controller; The action space of agent VSC1 is: The action space of agent VSC2 is: Where K is the gain, T1 and T2 are time constants; Design a data-driven reward function: r t =-|Δv dc1 (t)|-|Δv Ld2 (t)| Where Δv dc1 (t) is the deviation of the DC voltage of VSC1 from its reference value, Δv Ld2 (t) is the deviation of the d-axis output voltage of VSC2 from its reference value.
3. The method according to claim 2, characterized in that The multi-agent deep reinforcement learning model is trained using the MA-DDPG algorithm until convergence to obtain a trained multi-agent deep reinforcement learning model, including: In the collaborative training of the multi-agent deep reinforcement learning model, the MA-DDPG algorithm is used to achieve coordinated control between converters. The input of the training process includes the state space information of each agent: for agent Agent1 (VSC1), including the DC voltage v dc1 and its deviation Δv dc1 、Input DC current i dc1 and output active power P1; for agent Agent2 (VSC2), including d-axis output voltage v Ld2 and its deviation Δv Ld2 and output active power P2. The output of the training is the action space of each agent, which includes the parameters K of the VSC1 controller VSC1 (t),T1 VSC1 (t), And the parameter K of the VSC2 controller VSC2 (t),T1 VSC2 (t), The MA-DDPG algorithm adopts a centralized training and decentralized execution framework for training. In the centralized training phase, the two agents share the same experience replay pool to store interaction data samples (s t ,a t ,r t ,s t+1 ), and share a reward function r based on the electrical quantity measurement value. Each agent contains an Actor network and a Critic network, where the Actor network is responsible for outputting control actions based on the current state, and the Critic network is responsible for evaluating the value of the current state-action pair; During the training process, the Actor and Critic networks of the two agents interact to update network parameters. The Actor network considers the influence of the other agent's actions when updating, and the Critic network takes the states and actions of the two agents as input when evaluating. Coordinated training of multiple inverters using a shared data-driven reward function. During the training of the multi-agent deep reinforcement learning model, the agents of VSC1 and VSC2 share the same data-driven reward function r t , the reward function is built based on real-time measurement data, and the calculation formula is: r t =-|Δv dc1 (t)|-|Δv Ld2 (t)| Where Δv dc1 (t) represents the DC voltage deviation of VSC1 at time t, Δv Ld2 (t) represents the d-axis voltage deviation of VSC2 at time t, and the negative sign indicates that the smaller the deviation, the greater the reward value; The training termination of the multi-agent deep reinforcement learning model requires the following conditions to be met at the same time: first, the cumulative reward value change of the two agents in multiple consecutive training rounds is less than the preset threshold; second, the oscillation suppression effect of the system under different operating conditions meets the standard; third, the control parameters output by the two agents remain relatively stable in multiple consecutive training rounds, and the change range of the parameters is less than the set threshold; When the above conditions are met at the same time, the training of the multi-agent deep reinforcement learning model is completed, and a trained multi-agent deep reinforcement learning model is obtained.
4. The method according to claim 3, characterized in that The method of adjusting the parameters of the additional damping controller online through the trained multi-agent deep reinforcement learning model to generate the optimal control strategy for the new energy flexible direct current transmission system, and using the optimal control strategy to suppress system oscillations, includes: The input data of the trained multi-agent deep reinforcement learning model includes the state space information of each agent: for agent Agent1 (VSC1), it includes the DC voltage v dc1 and its deviation Δv dc1 、Input DC current i dc1 and output active power P1; for agent Agent2 (VSC2), including d-axis output voltage v Ld2 and its deviation Δv Ld2 And output active power P2; The multi-agent deep reinforcement learning model normalizes the input data to make it meet the numerical range set during training, and then inputs the input data into the Actor networks corresponding to VSC1 and VSC2 respectively. Based on the current state, the Actor networks obtain their respective optimal control strategies through forward calculation. The optimal control strategy includes the parameter K of the VSC1 controller. VSC1 ,T1 VSC1 , And the parameter K of the VSC2 controller VSC2 ,T1 VSC2 ,
Citation Information
Patent Citations
Multi-VSG micro-grid coordination control method based on deep reinforcement learning
CN118263889A
Broadband oscillation suppression method for new energy access power grid and electronic equipment
CN118713111A