A method for suppressing oscillations in a new energy transmission system via flexible DC based on the double-delay deep deterministic policy gradient algorithm
By applying the deep reinforcement learning model of the TD3 algorithm in the new energy trans-flexible direct delivery system, adaptively adjusting the additional damping control parameters, the problem of difficulty in effectively suppressing system oscillations in the existing technology is solved, and higher stability and control accuracy are achieved.
Patent Information
- Application Number
- CN202411589509.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-08
AI Technical Summary
The prior art is difficult to adjust the control parameters adaptively, and it is difficult to effectively suppress the oscillation of the new energy supply system through flexible direct transmission.
A deep reinforcement learning model based on TD3 algorithm is adopted. By obtaining the dynamic parameters and electrical quantity data of the system, state space, action space and reward functions are constructed, and the deep reinforcement learning model is trained to output additional damping control parameters, which is used to adjust the parameters of the additional damping controller online.
Adaptive oscillation suppression of the new energy through flexible direct delivery system is realized, and control parameters can be adjusted according to real-time operating conditions to improve the stability and control accuracy of the system.
Smart Images

Figure CN119518834B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system control, and particularly to an oscillation suppression method for a new energy transmission system via flexible DC transmission based on the TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm. Background Art
[0002] With the large-scale access of renewable energy, flexible DC transmission has become an important way for renewable energy transmission. However, dynamic interactions in the power transmission system may lead to wide-band oscillations, seriously threatening the safe and stable operation of the power transmission system.
[0003] Currently, the oscillation suppression methods in the existing power transmission systems mainly include parameter optimization, replacement of non-linear controllers, and additional damping control, etc. However, these methods often adopt a set of fixed parameters and are difficult to adapt to the changes in the system operating conditions. Therefore, there is an urgent need for an oscillation suppression method that can adaptively adjust control parameters. Summary of the Invention
[0004] Embodiments of the present invention provide an oscillation suppression method for a new energy transmission system via flexible DC transmission based on the TD3 algorithm to effectively suppress oscillations in the new energy transmission system via flexible DC transmission.
[0005] To achieve the above object, the present invention adopts the following technical solutions.
[0006] An oscillation suppression method for a new energy transmission system via flexible DC transmission based on the Twin Delayed Deep Deterministic Policy Gradient TD3 algorithm, comprising:
[0007] Obtaining the dynamic parameters and electrical quantity data of the new energy transmission system via flexible DC transmission, where the electrical quantity data includes DC side voltage and current and AC side d-axis voltage data;
[0008] Constructing a deep reinforcement learning model based on the dynamic parameters and electrical quantity data of the new energy transmission system via flexible DC transmission, defining the state space set, action space set, and reward function of the deep reinforcement learning model; training the deep reinforcement learning model using the TD3 algorithm to obtain a trained deep reinforcement learning model;
[0009] Inputting the current state space data of the new energy transmission system via flexible DC transmission into the trained deep reinforcement learning model, the trained deep reinforcement learning model outputs additional damping control parameters, applying the additional damping control parameters to an additional damping controller, and online adjusting the parameters of the additional damping controller to generate the optimal control strategy for the new energy transmission system via flexible DC transmission.
[0010] Preferably, constructing a deep reinforcement learning model based on the dynamic parameters and electrical quantity data of the new energy flexible DC transmission system includes:
[0011] Constructing a deep reinforcement learning model based on a deep neural network structure including an input layer, a hidden layer, and an output layer. The deep reinforcement learning model includes a state space, an action space, and a reward function. The state space includes the voltage, current, and power electrical quantities of the system. The action space includes the control parameters of the additional damping controller. The reward function is designed based on the system performance index;
[0012] Taking the deep reinforcement learning model as an adaptive optimizer for the additional damping controller. The deep reinforcement learning model extracts system features through a multi-layer deep neural network structure and outputs the additional damping control parameters.
[0013] Preferably, defining the state space set of the deep reinforcement learning model is as follows:
[0014] (1) For the agent of VSC1, its state space is defined as follows:
[0015]
[0016] Where:
[0017] -v dc1 Is the DC voltage of VSC1
[0018] -Δv dc1 (t) is the deviation of v dc1 From its reference value
[0019] -i dc1 (t) is the input DC current of the new energy output
[0020] (2) For the agent of VSC2, its state space is defined as follows:
[0021] Where:
[0022] -v Ld2 Is the d-axis output voltage of VSC2
[0023] -Δv Ld2 Is the deviation of v Ld2 From its reference value
[0024] - Is the reference value of v Ld2
[0025] Defining the action space of the deep reinforcement learning model is as follows:
[0026] The action space of VSC1:
[0027] Action space of VSC2:
[0028] where K is the gain, and T1 and T2 are the time constants.
[0029] Define the reward function of the deep reinforcement learning model as follows:
[0030] The damping ratio based on the dominant oscillation mode of the system is defined as follows:
[0031] r t VSC1 = min[ξ VSC1 (t) - ξ * , 0]
[0032] r t VSC2 = min[ξ VSC2 (t) - ξ * , 0]
[0033] where ξ VSC1 is the system damping ratio when implementing the additional damping control of VSC1, and ξ * is the target damping ratio. The reward function r of VSC2 t VSC2 is defined in a similar way.
[0034] Reward calculation based on the damping ratio of the dominant oscillation mode of the system
[0035]
[0036] where ξ(t) is the system damping ratio at the current moment, and ξ target is the target damping ratio.
[0037] When the system damping ratio ξ VSC1 is greater than the target value, the reward is 0
[0038] When the system damping ratio ξ VSC1 is less than the target value, a negative reward is given.
[0039] Preferably, training the deep reinforcement learning model using the TD3 algorithm to obtain a trained deep reinforcement learning model includes:
[0040] Initializing the actor network parameters, critic network parameters, target network parameters, and experience replay buffer of the TD3 algorithm. For each training episode, select an action based on the current policy and add exploration noise, execute the action and observe the reward and the next state, store the transition in the experience replay buffer, sample a batch of transitions from the buffer for training, update the critic network and the actor network, and softly update the target network;
[0041] The TD3 algorithm uses two critic networks to estimate the Q value, forms the target value using the minimum Q value, adds noise to the target policy action and clips it, delays the update of the policy network and the target network until the deep reinforcement learning model converges, and obtains the trained deep reinforcement learning model, obtains the trained deep reinforcement learning model.
[0042] Preferably, inputting the current state space data of the new energy transmitted through the flexible DC transmission system into the trained deep reinforcement learning model, the trained deep reinforcement learning model outputs the additional damping control parameter, applying the additional damping control parameter to the additional damping controller, and online adjusting the parameters of the additional damping controller to generate the optimal control strategy of the new energy transmitted through the flexible DC transmission system, including:
[0043] Obtain the current state space data s of the new energy transmitted through the flexible DC transmission system, and this state space data s includes the equivalent DC current of the new energy, the DC side voltage of the VSC, and the AC side voltage;
[0044] Input the state space data s into the trained deep reinforcement learning model, and the trained deep reinforcement learning model calculates the additional damping control parameter a←π φ (s), and this additional damping control parameter a includes the gain and the time constant;
[0045] Apply the additional damping control parameter a to the additional damping controller, and the output signal of the additional damping controller acts on the DC voltage control loop or the AC voltage control loop of the VSC, online adjust the additional damping control parameter a of the additional damping controller, and through the real-time evaluation and decision-making of the system state by the deep reinforcement learning model, continuously adjust through online optimization to achieve the optimal operating state of the system and generate the optimal control strategy of the new energy transmitted through the flexible DC transmission system.
[0046] It can be seen from the technical solutions provided by the embodiments of the present invention described above that the present invention provides an oscillation suppression control method for a new energy transmitted through a flexible DC transmission system, and designs an additional damping controller through the TD3 algorithm to achieve the adaptive suppression of system oscillation.
[0047] Additional aspects and advantages of the present invention will be given in part in the following description, and these will become apparent from the following description, or can be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0049] Figure 1 It is a system topology diagram of a flexible DC transmission system for new energy provided by an embodiment of the present invention;
[0050] Figure 2 It is a processing flowchart of a method for adaptively adjusting additional damping control parameters provided by an embodiment of the present invention;
[0051] Figure 3 It is a network structure diagram of a TD3 algorithm provided by an embodiment of the present invention;
[0052] Figure 4 It is a comparison diagram of oscillation suppression effects provided by an embodiment of the present invention. Detailed Embodiments
[0053] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0054] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any and all combinations of one or more of the associated listed items.
[0055] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as herein.
[0056] For the convenience of understanding the embodiments of the present invention, the following will further explain with several specific embodiments in conjunction with the drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.
[0057] The research object of the present invention is a flexible DC transmission system for renewable energy. The structure of the flexible DC transmission system for renewable energy is as Figure 1 shown, including: a renewable energy converter VSC1, a sending-end converter VSC2 and a receiving-end converter VSC3 of a VSC-HVDC (Voltage Source Converter-High Voltage Direct Current). VSC1 adopts VdcIq (DC Voltage and Q-axis Current Control) control, VSC2 adopts Vf (voltage-frequency control) control, and VSC3 adopts VdcIq (DC voltage and q-axis current control) control. The present invention installs additional damping controllers on VSC1 and VSC2. The adaptive additional damping control can adaptively adjust its own internal parameters according to the changes of real-time working conditions to effectively suppress oscillations.
[0058] The processing flow of a method for suppressing oscillations in a flexible DC transmission system for renewable energy based on the TD3 algorithm provided by the embodiments of the present invention is as Figure 2 shown, including the following processing steps:
[0059] Step S10: Obtain the dynamic parameters and electrical quantity data of the flexible DC transmission system for renewable energy. The above electrical quantity data includes data such as DC side voltage and current and AC side d-axis voltage.
[0060] Step S20: Construct a deep reinforcement learning model based on the dynamic parameters and electrical quantity data of the flexible DC transmission system for renewable energy, and define the state space set, action space set and reward function of the deep reinforcement learning model.
[0061] Step S30: Use the TD3 algorithm to train the deep reinforcement learning model until convergence.
[0062] Step S40: The TD3 algorithm is used to design an adaptive additional damping controller. The parameters of the additional damping controller are adjusted online through the trained deep reinforcement learning model to generate the optimal control strategy for the new energy transmission system via VSC-HVDC to suppress system oscillations.
[0063] The above-mentioned step S10 includes: In practical applications, a PMU (phasor measurement unit) can be used to obtain the dynamic parameters and electrical quantity data of the new energy transmission system via VSC-HVDC. The PMU is widely used in the data acquisition of the new energy transmission system via VSC-HVDC and is usually installed at positions such as the AC bus of the new energy power station, the AC side of the VSC converter station, the key nodes of the system, and the grid connection points. It can synchronously collect key electrical quantity data such as voltage / current phasors (including amplitude and phase angle), frequency and rate of change of frequency, active / reactive power, and sequence components. The sampling frequency is usually ≥50 times per second, and time synchronization is ensured through the satellite positioning system with a microsecond-level synchronization accuracy.
[0064] The above-mentioned step S20 includes: The deep reinforcement learning model has a deep neural network structure and consists of three core elements: the state space, the action space, and the reward function. Among them, the state space includes key electrical quantities such as the voltage, current, and power of the system; the action space includes the control parameters of the additional damping controller; the reward function is designed based on system performance indicators such as stability. The deep reinforcement learning model uses a deep neural network as a value function approximator or a policy function to extract system features through a multi-layer network structure and output the control strategy.
[0065] The deep reinforcement learning model of the embodiment of the present invention can be used as an adaptive optimizer for the additional damping controller. It continuously optimizes the control parameters by learning the dynamic characteristics of the system in real time, enabling the damping controller to adapt to the changes in the system operating state. The output of the deep reinforcement learning model can be directly used as the control signal of the damping controller or used to dynamically adjust the parameters of the damping controller, such as the gain coefficient, time constant, etc.
[0066] The construction process of the above-mentioned deep reinforcement learning model includes: First, preprocessing operations such as cleaning, normalization, and time alignment need to be performed on the raw data collected by devices such as PMUs. Then, key features (such as voltage, current, power, etc.) are extracted from the preprocessed data as the input of the state space, and based on this, a deep reinforcement learning model with a deep neural network structure including an input layer, a hidden layer, and an output layer is designed. After the model is constructed, historical operation data is used for offline training to establish an initial policy model. After the model is deployed, the control strategy is continuously optimized and updated through real-time data. At the same time, the control effect is continuously monitored and the model parameters are dynamically adjusted according to the system response, finally forming a closed-loop adaptive learning mechanism to continuously improve the control performance.
[0067] The present invention uses the TD3 algorithm to design an adaptive additional damping controller, and the system operating condition parameters in the state space include: the equivalent DC current of new energy, the DC side voltage of the VSC, and the AC side voltage.
[0068] Define the state space set of the deep reinforcement learning model as follows:
[0069] (1) For the agent of VSC1, its state space is defined as follows:
[0070] s t VSC1 ={v dc1 (t), Δv dc1 (t), i dc1 (t)}
[0071] Where:
[0072] -v dc1 is the DC voltage of VSC1
[0073] -Δv dc1 (t) is the deviation of v dc1 from its reference value
[0074] -i dc1 (t) is the input DC current of the new energy output
[0075] (2) For the agent of VSC2, its state space is defined as follows:
[0076] Where:
[0077] -v Ld2 is the d-axis output voltage of VSC2
[0078] -Δv Ld2 is the deviation of v Ld2 from its reference value
[0079] - is the reference value of v Ld2
[0080] Define the action space of the deep reinforcement learning model as follows:
[0081] The action space of VSC1:
[0082] The action space of VSC2:
[0083] Where K is the gain, and T1 and T2 are the time constants.
[0084] Define the reward function of the deep reinforcement learning model as follows:
[0085] The damping ratio based on the system-dominated oscillation mode is defined as follows:
[0086] r t VSC1 = min[ξ VSC1 (t) - ξ * , 0]
[0087] r t VSC2 = min[ξ VSC2 (t) - ξ * , 0]
[0088] where ξ VSC1 is the system damping ratio when implementing the additional damping control of VSC1, and ξ * is the target damping ratio. The reward function r of VSC2 t VSC2 is defined in a similar way.
[0089] Reward calculation is performed based on the damping ratio of the system-dominated oscillation mode
[0090]
[0091] where ξ(t) is the system damping ratio at the current moment, and ξ target is the target damping ratio.
[0092] When the system damping ratio ξ VSC1 is greater than the target value, the reward is 0
[0093] When the system damping ratio ξ VSC1 is less than the target value, a negative reward is given.
[0094] The system eigenvalues are calculated through the frequency-domain impedance network modeling method and are used for the calculation of the reward function.
[0095] The above step S30 includes:
[0096] Figure 3 This is the network structure diagram of a TD3 algorithm provided by an embodiment of the present invention. The implementation steps of the TD3 algorithm include: The key of the TD3 algorithm lies in maintaining two critic networks to estimate the Q value and using the minimum value of these two networks to form the target, thereby reducing the overestimation bias. The specific steps are as follows:
[0097] (1) Initialization:
[0098] - Initialize the actor network parameters θ, the critic network parameters φ1 and φ2
[0099] - Initialize the target network parameters θ', φ1' and φ2'
[0100] - Initialize the experience replay buffer R
[0101] (2) For each training episode: Select an action based on the current policy and add exploration noise, execute the action and observe the reward and the next state, store the transition in the experience replay buffer, sample a batch of transitions from the buffer for training, update the critic network and the actor network, and softly update the target network.
[0102] a) Initialize the power system
[0103] b) For each time step t:
[0104] - Select an action with exploration noise: a ← π φ (s) + ε,
[0105] - Execute the action and observe the reward r and the next state s'
[0106] - Store a state transition process (s, a, r, s′) in R
[0107] - Sample N state transition processes (s i , a i , r i , s i ′)
[0108] - Calculate the target Q value:
[0109]
[0110] where ε is the noise for target policy smoothing
[0111] - Update the critic network:
[0112]
[0113] - Update the actor network every d steps:
[0114]
[0115] - Softly update the target network:
[0116] φ′ ← τφ + (1 - τ)φ′
[0117] θ i ′ ← τθ i + (1 - τ)θ i ′ for i = 1, 2
[0118] The process of training a deep reinforcement learning model using the TD3 algorithm includes: establishing two critic networks to estimate the Q value, using the minimum Q value to form the target value, reducing the overestimation bias, adding noise to the target policy action and performing clipping, delaying the update of the policy network and the target network until the deep reinforcement learning model converges, and obtaining the trained deep reinforcement learning model. During the training phase, data under different operating conditions and disturbance levels are used to improve the adaptability of the controller.
[0119] The convergence of the deep reinforcement learning model is mainly judged by monitoring the change trend of the reward function value, the loss function value of the policy network, and the system key performance indicators (such as transition time, overshoot, steady-state error, etc.); when the fluctuation amplitude of these indicators is lower than the preset threshold within multiple consecutive training cycles and the dynamic response characteristics of the system meet the control requirements, it can be considered that the model has converged. At the same time, it is also necessary to verify the generalization ability of the model through cross-validation and testing under different operating conditions to ensure that the model can maintain a stable control effect under various operating conditions. In practical applications, multiple convergence criteria can be set, and the model is considered to be fully converged only when all criteria are met.
[0120] The above step S40 includes:
[0121] After the deep reinforcement learning model is trained, the additional damping controller can operate online according to the real-time measurement data:
[0122] The input data of the trained deep reinforcement learning model is the data corresponding to the previously defined state space measured online, that is, the equivalent DC current of the new energy, the DC side voltage of the VSC, and the AC side voltage. The output data is the data corresponding to the previously defined action space, that is, the control parameters of the additional damping controller, including the gain and the time constant.
[0123] (1) Obtain the current state space data s of the system, and the state space data s includes the equivalent DC current of the new energy, the DC side voltage of the VSC, and the AC side voltage.
[0124] (2) Input the above state space data s into the trained deep reinforcement learning model, and the trained deep reinforcement learning model calculates the additional damping control parameter a←π φ (s), and the additional damping control parameter a includes the gain and the time constant.
[0125] (3) Apply the additional damping control parameter a to the additional damping controller, and apply the output signal of the additional damping controller to the DC voltage control loop or the AC voltage control loop of the VSC. Adjust the additional damping control parameter a of the additional damping controller online. Through the real-time evaluation and decision-making of the system state by the deep reinforcement learning model, overall consider the dynamic characteristics, stability requirements and operation constraints of the system, and continuously adjust through online optimization to achieve the optimal operation state of the system, generate the optimal control strategy for the new energy transmission system via VSC-HVDC to achieve system oscillation suppression.
[0126] Verification of control effect Figure 4 The figure below is a comparison chart of oscillation suppression effects provided by the embodiments of the present invention. The method of the present invention has been verified under three different working conditions:
[0127] Condition 1: i dc1 = 1000 A, R Load = 0.1 Ω → 0.05 Ω,
[0128] Condition 2: i dc1 = 900 A, R Load = 0.1 Ω → 0.05 Ω,
[0129] Condition 3: i dc1 = 1100 A, R Load = 0.1 Ω → 0.15 Ω,
[0130] As Figure 4 shown, the simulation results show that compared with the traditional fixed parameter method and the DQN (Deep Q-Network) method, the method of the present invention has better oscillation suppression effect, manifested as smaller overshoot and faster adjustment time. This verifies that the method of the present invention has good self-adaptability to changes in working conditions.
[0131] The method of the present invention can be applied to multiple VSC converters in the system simultaneously to achieve coordinated control. The method of the present invention is applicable to suppressing low-frequency oscillation, subsynchronous oscillation and high-frequency oscillation.
[0132] In summary, the present invention provides an oscillation suppression control method for a new energy transmission system via VSC-HVDC, which realizes the adaptive adjustment of the additional damping control parameter through the TD3 algorithm and achieves the adaptive suppression of system oscillation.
[0133] The present invention effectively solves the problem of overestimation of Q-values by adopting a dual-critic network structure, improving the training effect; based on the model-driven design of the reward function, making full use of the system damping characteristic information to accelerate the training process; the actual application results show that compared with the traditional fixed-parameter method, the method of the present invention has a better oscillation suppression effect.
[0134] The method of the present invention can effectively guide the intelligent agent to adjust the controller parameters to achieve adaptive suppression of oscillations under different working conditions by designing a model-based reward function and combining the calculation of system eigenvalues. By using a two-stage training mechanism, in the initial stage, a single intelligent agent is trained in a model-driven manner, and then multi-agent coordinated training is carried out through a data-driven reward function to improve the overall control effect. Compared with the existing methods, the present invention has significant advantages in terms of the response speed, self-adaptability, and control accuracy of oscillation suppression.
[0135] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0136] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0137] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and the relevant parts can be referred to the partial description of the method embodiments. The device and system embodiments described above are only illustrative, and the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0138] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for suppressing oscillation of a new energy flexible direct current transmission system based on a double-delayed deep deterministic policy gradient algorithm, characterized in that: include: Obtain dynamic parameters and electrical quantity data of the new energy flexible direct current transmission system, wherein the electrical quantity data includes the voltage and current on the direct current side and the d-axis voltage data on the alternating current side; Based on the dynamic parameters and electrical quantity data of the new energy flexible direct current transmission system, a deep reinforcement learning model is constructed, and the state space set, action space set and reward function of the deep reinforcement learning model are defined; the deep reinforcement learning model is trained using the TD3 algorithm to obtain a trained deep reinforcement learning model; Inputting the current state space data of the new energy through flexible direct current transmission system into the trained deep reinforcement learning model, the trained deep reinforcement learning model outputs additional damping control parameters, applying the additional damping control parameters to the additional damping controller, adjusting the parameters of the additional damping controller online, and generating the optimal control strategy of the new energy through flexible direct current transmission system; The deep reinforcement learning model constructed based on the dynamic parameters and electrical quantity data of the new energy flexible direct current transmission system includes: Constructing a deep reinforcement learning model based on a deep neural network structure including an input layer, a hidden layer, and an output layer, wherein the deep reinforcement learning model includes a state space, an action space, and a reward function, wherein the state space includes voltage, current, and power electrical quantities of the system, the action space includes control parameters of an additional damping controller, and the reward function is designed based on system performance indicators; The deep reinforcement learning model is used as an adaptive optimizer of the additional damping controller. The deep reinforcement learning model extracts system features through a multi-layer deep neural network structure and outputs additional damping control parameters.
2. The method according to claim 1, characterized in that: The state space set of the deep reinforcement learning model is defined as follows: (1) For the VSC1 agent, its state space is defined as follows: s t VSC1 ={v dc1 (t),Δv dc1 (t),i dc1 (t)} in: -v dc1 is the DC voltage of VSC1 -Δv dc1 (t) is v dc1 Deviation from its reference value -i dc1 (t) is the input DC current of the renewable energy output (2) For the VSC2 agent, its state space is defined as follows: in: -v Ld2 is the d-axis output voltage of VSC2 -Δv Ld2 v Ld2 Deviation from its reference value - v Ld2 Reference value The action space of the deep reinforcement learning model is defined as follows: Action space of VSC1: Action space of VSC2: Where K is the gain, T1 and T2 are time constants; The reward function of the deep reinforcement learning model is defined as follows: The damping ratio based on the dominant oscillation mode of the system is defined as follows: r t VSC1 =min[ξ VSC1 (t)-ξ * ,0] r t VSC2 =min[ξ VSC2 (t)-ξ * ,0] where ξ VSC1 is the system damping ratio when VSC1 additional damping control is implemented, ξ * is the target damping ratio, and the reward function r of VSC2 t VSC2 The definition is similar; Reward calculation based on the damping ratio of the dominant oscillation mode of the system Among them, ξ(t) is the system damping ratio at the current moment, ξ target is the target damping ratio When the system damping ratio ξ VSC1 When it is greater than the target value, the reward is 0 When the system damping ratio ξ VSC1 When it is less than the target value, a negative reward is given.
3. The method according to claim 2, characterized in that The method of using the TD3 algorithm to train the deep reinforcement learning model to obtain a trained deep reinforcement learning model includes: Initialize the actor network parameters, critic network parameters, target network parameters and experience replay buffer of the TD3 algorithm. For each training round, select actions based on the current strategy and add exploration noise, execute actions and observe rewards and next states, store transfers in the experience replay buffer, sample batch transfers from the buffer for training, update the critic network and actor network, and soft-update the target network. The TD3 algorithm uses two critic networks to estimate Q values, uses the minimum Q value to form the target value, adds noise to the target policy action and prunes it, delays updating the policy network and the target network until the deep reinforcement learning model converges, and obtains a trained deep reinforcement learning model.
4. The method according to claim 3, characterized in that The method of inputting the current state space data of the new energy through flexible direct current transmission system into the trained deep reinforcement learning model, the trained deep reinforcement learning model outputting additional damping control parameters, applying the additional damping control parameters to the additional damping controller, adjusting the parameters of the additional damping controller online, and generating the optimal control strategy of the new energy through flexible direct current transmission system includes: Obtain the current state space data s of the new energy flexible direct current transmission system, where the state space data s includes the new energy equivalent direct current, the VSC direct current side voltage and the alternating current side voltage; The state space data s is input into the trained deep reinforcement learning model, and the trained deep reinforcement learning model calculates the additional damping control parameter a←π φ (s), the additional damping control parameter a includes a gain and a time constant; The additional damping control parameter a is applied to the additional damping controller, the output signal of the additional damping controller is applied to the DC voltage control loop or the AC voltage control loop of the VSC, the additional damping control parameter a of the additional damping controller is adjusted online, and the system state is evaluated and decided in real time by a deep reinforcement learning model. Through online optimization, continuous adjustments are made to achieve the optimal operating state of the system, and the optimal control strategy for the new energy flexible direct current transmission system is generated.
Citation Information
Patent Citations
Power system AC optimal power flow decision-making method, device, equipment and medium
CN117335414A
Subsynchronous oscillation suppression method, device and equipment and storage medium
CN118889391A