Flexible DC system oscillation emergency control method and system based on deep reinforcement learning

Through the flexible DC system oscillation emergency control method based on deep reinforcement learning, a state space model is constructed and converted into a Markov decision process to generate control instructions, which solves the problem that traditional methods are difficult to suppress flexible DC system oscillations and achieves fast and accurate oscillation suppression.

CN120710077APending Publication Date: 2025-09-26STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510980775.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional offline oscillation analysis and control strategies are unable to meet the oscillation suppression needs of flexible DC systems under the complex structure of large power grids. Commonly used control methods require the reconstruction of the original structure, making it difficult for dispatchers to suppress oscillations quickly and effectively.

Method used

An emergency oscillation control method for a flexible DC system based on deep reinforcement learning is adopted. The three-phase voltage and current signals are obtained for preprocessing, a state space model is constructed, and the model is converted into a Markov decision process. The control instructions are generated using a proximal strategy optimization algorithm to achieve emergency oscillation control of the flexible DC system.

Benefits of technology

Real-time suppression of oscillations in flexible DC systems can be achieved without additional electrical equipment. The strategy network and system status are interactively updated, improving the accuracy and adaptability of oscillation suppression and eliminating dependence on internal parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120710077A_ABST
    Figure CN120710077A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power system control, in particular to a flexible direct current system oscillation emergency control method and system based on deep reinforcement learning, the oscillation emergency control problem is converted into a Markov decision process, a state space in deep reinforcement learning is formed by combining the current system damping strength and a controllable variable, and the system oscillation emergency control method and system based on deep reinforcement learning are obtained. The power increment is used as an action space of the flexible power transmission system; and a near-end strategy optimization algorithm is adopted to solve the specific output of the system power for guaranteeing the stable operation of the flexible direct-current power transmission system in real time. A flexible direct-current power transmission system regulation and control strategy is formed through machine learning, and real-time oscillation suppression can be realized through the control potential of the flexible direct-current power transmission system without additional electrical equipment. And the constructed strategy network continuously performs parameter interaction with the system operation state, continuously updates network parameters according to result feedback while outputting the regulation and control strategy, continuously enhances the value evaluation and action selection performance of the network, and ensures the accuracy of decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power system control technology, and in particular to a method and system for emergency control of oscillations in a flexible direct current system based on deep reinforcement learning. Background Art

[0002] Flexible DC transmission technology has been widely used in renewable energy grid integration and large-scale grid interconnection due to its advantages, including good controllability, low low-order harmonic content, easy boosting and expansion, and adaptability to offshore wind power integration. However, the large-scale commissioning of flexible DC transmission projects has led to an increase in the penetration of power electronics in power systems. This has led to interactions between heterogeneous power electronics devices, such as renewable energy converters and energy storage devices, and between power grids due to the high-frequency nonlinear links introduced by power electronics control. This has led to numerous stability issues in the dynamic processes of the power grid, with oscillation instability being a particularly prominent problem. Oscillation issues not only affect power quality and equipment safety but can also induce system instability and divergence, seriously endangering the safe and stable operation of the power system.

[0003] As the scale of power grids continues to grow and their structures become increasingly complex, the system's dynamic characteristics are becoming highly nonlinear. Traditional offline oscillation analysis and control strategies can no longer meet application requirements. Commonly used suppression methods based on additional control devices require the original control structure to be reshaped, and oscillation suppression may not be achieved after major changes in the power system's grid structure and operating mode. Therefore, it is difficult for dispatchers to quickly and effectively suppress oscillations at the operational level.

[0004] The information disclosed in this background technology section is only intended to deepen the understanding of the overall background technology of the present disclosure and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art known to those skilled in the art. Summary of the Invention

[0005] The present invention provides a flexible DC system oscillation emergency control method based on deep reinforcement learning, which can effectively solve the problems in the background technology.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is:

[0007] A method for emergency control of oscillations of a flexible DC system based on deep reinforcement learning, comprising the following steps:

[0008] Obtain three-phase voltage and three-phase current signals at the access point of the flexible DC transmission system;

[0009] Preprocess the electrical signal, filter out high-frequency noise, and use the improved Z-score method to detect and replace abnormal data points;

[0010] Based on the preprocessed data, the state space model of the HVDC Flexible system is constructed using the vector fitting method.

[0011] Extract the real part of the state characteristic root σ through eigenvalue analysis t , and determine whether to trigger oscillation control;

[0012] The oscillation emergency control problem of HVDC flexible system is transformed into a Markov decision process (MDP) model;

[0013] The proximal strategy optimization algorithm is used to solve the MDP model, generate control instructions and send them to the control center.

[0014] Furthermore, the pre-processing of the electrical signal includes:

[0015] According to the oscillation problem of the flexible DC transmission system, a low-pass filter with a cutoff frequency of 10kHz is used to filter out high-frequency noise;

[0016] The implementation of the improved Z-score method includes the following steps:

[0017] Construct a sliding window with the current data point x as the center, including 3 data points before and after;

[0018] Calculate the mean x of the data in the window m and standard deviation x σ ;

[0019] If|xx m | / x σ >3, then replace x with x m .

[0020] Furthermore, the constructing of the state space model of the flexible DC transmission system by using the vector fitting method includes the following steps:

[0021] Construct the fitting transfer function, which is specifically expressed as:

[0022]

[0023] Among them, N is the initial order, r j is the residue, a j h fit (s) is the pole, d and k are rational numbers, and s is the complex frequency;

[0024] Optimize the parameters by the least squares method so that h fit (s) approximates h(s) actual system transfer function

[0025] The state space model of the HVDC Flexible system is specifically expressed as follows:

[0026]

[0027] Among them, X, U, and Y are state variables, input variables, and output variables respectively; A, B, C, and D are state matrices, input matrices, output matrices, and direct transfer matrices respectively.

[0028] Furthermore, the real part of the state characteristic root σ is extracted by eigenvalue analysis. t , and judging whether to trigger oscillation control, including the following steps:

[0029] Calculate the characteristic root λ of the state matrix A of the flexible DC transmission system t =σ t +jω t , extract the real part of the characteristic root σ t As a basis for judging the system oscillation damping characteristics;

[0030] When σ t When <-0.7, the system is judged to be in a strong damping state;

[0031] When σ t When ≥-0.7, the system is judged to be in a weak damping state.

[0032] Furthermore, a Markov decision process (MDP) model is constructed, including the definition:

[0033] State space S t =(σ t ,Q t ,P t ,V t ), where σ t is the real part of the system characteristic root, Q t is the reactive power output of the system at time t, P t is the active power output of the system at time t, V t is the voltage of the node connected to VSC at time t;

[0034] Action space A t is the controllable variable of the system at time t;

[0035] Reward function R t The design is a constraint priority mechanism, specifically expressed as:

[0036]

[0037] Among them, α constrain is a set of system constraints, including node voltage constraints, active power output constraints, and reactive power output constraints; R t,p is the penalty for not meeting the constraint, R t,r The reward when the constraint is satisfied.

[0038] Furthermore, the proximal policy optimization algorithm is based on the Actor-Critic framework, where the policy (Actor) network outputs the action probability distribution, the value (Critic) network calculates the state value function, and updates the neural network parameters through gradient descent;

[0039] Actor network parameter update uses the clipping objective function, which is specifically expressed as:

[0040]

[0041] The loss function of the Critic network is, specifically expressed as:

[0042]

[0043] Among them, θ represents the Actor network parameters, Represents the loss function of the Actor network The gradient of the parameter θ, η θ represents the dynamic learning rate of the Actor network, E represents the expectation, w represents the Critic network parameter, V w (s t ) represents the current value function; Represents the target value function used to evaluate the accuracy of the critic network output.

[0044] Furthermore, the advantage function Introducing the Actor network, specifically expressed as:

[0045]

[0046] Q w (s,a) means executing action a according to strategy π in the current state t The reward expectation, V w (s) represents the expected reward for executing all control schemes according to the strategy π in the current state, which is specifically expressed as:

[0047] Q w (s, a) = E(R t |s t =s,a t =a;π);

[0048] V w (s)=E(R t |s t =s;π).

[0049] Further, Based on the temporal difference algorithm, it is specifically expressed as:

[0050]

[0051] After constructing the loss function, the critic network parameters are updated through gradient calculation, which is specifically expressed as:

[0052]

[0053] in, is the loss function L V (w) Gradient with respect to parameter w; η w is the Critic network learning rate.

[0054] Furthermore, the dynamic learning rate of the Actor network is calculated by the ratio of the sampling probabilities of the new and old strategies, which is specifically expressed as:

[0055]

[0056] Among them, η θ,base Represents the baseline learning rate of the Actor network; π θ (a t ,s t ) is the sampling probability of the new strategy, is the sampling probability of the old strategy; ε is the hyperparameter that limits the clipping interval; the CLIP function controls the change between the old and new strategies within the interval [1-ε,1+ε].

[0057] A flexible DC system oscillation emergency control system based on deep reinforcement learning, which adopts the above-mentioned flexible DC system oscillation emergency control method based on deep reinforcement learning, comprises:

[0058] A signal acquisition module is used to obtain three-phase voltage and three-phase current signals at the access point of the flexible DC transmission system;

[0059] The preprocessing module uses a low-pass filter with a cutoff frequency of 10 kHz to filter out high-frequency noise in the signal and uses the improved Z-score method to detect and replace abnormal data points;

[0060] a modeling module, connected to the preprocessing module, and constructing a state space model of the flexible DC transmission system using a vector fitting method;

[0061] A stability judgment module, connected to the modeling module, calculates the characteristic roots of the state matrix A and extracts the real part of the characteristic roots to determine whether to trigger oscillation control;

[0062] A reinforcement learning decision module is connected to the stability judgment module, models the oscillation control problem as a Markov decision process, uses a proximal policy optimization algorithm to solve the MDP, and generates control instructions;

[0063] The command output module is connected to the reinforcement learning decision module and is used to send control commands to the flexible direct current transmission control center.

[0064] The beneficial effects of the present invention are:

[0065] The present invention proposes a method for emergency control of oscillations in a flexible direct current (HVDC) system based on deep reinforcement learning. By forming a control strategy for the flexible HVDC transmission system through machine learning, real-time oscillation suppression can be achieved through the control potential of the flexible HVDC transmission system itself without the need for additional electrical equipment.

[0066] The constructed policy network continuously interacts with the system's operating status through parameter interaction. While outputting control strategies, it continuously updates network parameters based on feedback, continuously enhancing the network's value assessment and action selection performance to ensure decision-making accuracy.

[0067] After collecting and training the corresponding data sets to build a neural network, only the system's online operation data needs to be input to output the decision results for suppressing oscillations. Compared with traditional additional damping control devices, this device gets rid of the dependence on the internal parameters of the flexible DC converter and the power grid, and has stronger adaptability for suppressing oscillations in highly flexible DC transmission systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0069] Figure 1 This is a flow chart of the method for emergency control of oscillations of a flexible DC system based on deep reinforcement learning in the present invention;

[0070] Figure 2 The topology structure of the flexible direct current transmission system in the present invention;

[0071] Figure 3 Schematic diagram of the oscillation emergency control method of a flexible DC system based on deep reinforcement learning in the present invention;

[0072] Figure 4 This is the reward function convergence curve calculated by the flexible DC system oscillation emergency control method based on deep reinforcement learning in the present invention;

[0073] Figure 5 This is the power waveform at the PCC point where the flexible DC system is connected to the AC system in the present invention. DETAILED DESCRIPTION

[0074] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0075] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly attached to the other element or there may be an intermediate element. When an element is referred to as being "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only and do not represent the only implementation methods.

[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0077] Taking into account the powerful ability of machine learning to handle nonlinear problems, the present invention introduces it into the field of oscillation control. By collecting and training data sets to construct the logic from system modes to control variables, the system increment value that ensures system stability is quickly and accurately solved, forming a new solution for oscillation emergency control.

[0078] By transforming the oscillation emergency control problem into a Markov decision process (MDP), the state space in deep reinforcement learning is constructed by combining the current system damping strength and controllable variables (such as system power output), and the power increment is used as the action space of the flexible transmission system; on this basis, the proximal strategy optimization algorithm is used to solve the specific system power output that ensures the stable operation of the flexible direct current transmission system in real time, thereby effectively realizing the oscillation emergency control of the flexible direct current transmission system.

[0079] Specifically, such as Figure 1 The method for emergency control of oscillations of a flexible DC system based on deep reinforcement learning is shown, comprising the following steps:

[0080] The three-phase voltage and three-phase current signals at the access point of the flexible DC transmission system are obtained; the electrical signals are preprocessed to filter out high-frequency noise and an improved Z-score method is used to detect and replace abnormal data points; based on the preprocessed data, a vector fitting method is used to construct a state space model of the flexible DC transmission system; and the real part of the state characteristic root σ is extracted through eigenvalue analysis. t, and determine whether to trigger oscillation control; transform the oscillation emergency control problem of the flexible DC transmission system into a Markov decision process (MDP) model; use the proximal strategy optimization algorithm to solve the MDP model, generate control instructions and send them to the control center for reference by operators.

[0081] In this embodiment, electrical signals are preprocessed to remove noise and abnormal data from the initial voltage and current signals. First, the collected voltage and current data are fed into a low-pass filter to remove high-frequency noise from the original signals. Furthermore, the abnormal point data is further processed to minimize the amount of interference introduced by the sampling process, thereby preventing any impact on the oscillation emergency control method.

[0082] First, a suitable low-pass filter cutoff frequency must be selected. The low-pass filter must ensure that the fundamental frequency and oscillation frequency signals are retained while filtering out high-frequency noise. Due to the oscillation issues in the flexible DC transmission system, a low-pass filter with a cutoff frequency of 10kHz is used to filter out high-frequency noise.

[0083] Secondly, detect abnormal data points. For the collected data points, it is necessary to determine whether there are outliers. The present invention uses the improved Z score method to detect outliers. The implementation of the improved Z score method includes the following steps:

[0084] Construct a sliding window with the current data point x as the center, including 3 data points before and after it; calculate the mean x of the data in the window m and standard deviation x σ ; If | xx m | / x σ >3, then x is considered an outlier.

[0085] Finally, it is necessary to deal with abnormal data points. Common methods for dealing with outliers include deletion, correction, and replacement. Based on the abnormal data judgment rule of the improved Z score method, the abnormal point data x is replaced by x m .

[0086] In this embodiment, a vector fitting method is used to construct a state space model of a flexible DC transmission system, including the following steps:

[0087] Assuming that the original transfer function of the system is h(s), construct the fractional cumulative form fitting function h fit (s) Construct the fitting transfer function, which is specifically expressed as:

[0088]

[0089] Among them, N is the initial order, r j is the residue, a j h fit (s) is the pole, d and k are rational numbers, and s is the complex frequency;

[0090] Optimize the parameters by the least squares method so that h fit (s) approximates h(s) actual system transfer function

[0091] The state space model of the HVDC Flexible system is specifically expressed as follows:

[0092]

[0093] Among them, X, U, and Y are state variables, input variables, and output variables respectively; A, B, C, and D are state matrices, input matrices, output matrices, and direct transfer matrices respectively.

[0094] Furthermore, the real part of the state characteristic root σ is extracted by eigenvalue analysis t As a basis for judging the system oscillation damping characteristics and determining whether to trigger oscillation control, the following steps are included:

[0095] Calculate the characteristic root λ of the state matrix A of the flexible DC transmission system t =σ t +jω t , extract the real part of the characteristic root σ t As a basis for judging the system oscillation damping characteristics;

[0096] When σ t When <-0.7, the system is judged to be in a strong damping state;

[0097] When σ t When ≥-0.7, the system is judged to be in a weak damping state.

[0098] The obtained damping evaluation results are input into the neural network for designing the reward mechanism.

[0099] In this embodiment, a Markov decision process (MDP) model is constructed, including defining:

[0100] State space S t =(σ t ,Q t ,P t ,V t ), where S t is the state set of the system at time t, σ t is the real part of the system characteristic root, Q t is the reactive power output of the system at time t, P t is the active power output of the system at time t, V t is the voltage of the node connected to VSC at time t;

[0101] Action space A t is the controllable variable of the system at time t, such as reactive output, active output, etc.

[0102] Reward function R t Designed as a constraint priority mechanism, R t To obtain rewards after executing an action, the constraint target partitioning-objective function method is used to design a reward mechanism. Specifically, the reward calculation of the objective function is considered only when the constraint conditions are met. Specifically, it is expressed as:

[0103]

[0104] Among them, α constrain is a set of system constraints, including node voltage constraints, active power output constraints, and reactive power output constraints; R t,p is the penalty for not meeting the constraint, R t,r The goal is to determine the optimal control solution for the reward when the constraints are met.

[0105] Reward function R t The specific settings are as follows:

[0106]

[0107] In this embodiment, the proximal policy optimization algorithm is based on the Actor-Critic framework, where the policy (Actor) network outputs the action probability distribution, the value (Critic) network calculates the state value function, and updates the neural network parameters through gradient descent;

[0108] The input of the Actor network is the current state vector s t , the output is the mean and variance of the controllable variables of the system. The control strategy π is formed by constructing the multivariate normal probability distribution function of the action vector. The corresponding control scheme can be obtained by sampling the normal probability distribution function. The control scheme is then passed to the system model to calculate the characteristic roots of the system after control and generate the state vector s at the next moment. t+1 , then the training sample sequence (s t ,a t ,r t ,s t+1 ), output samples to the experience buffer pool.

[0109] Furthermore, the input of the Critic network is the current state vector s t The output is the value function V used to evaluate the quality of the control scheme w (s).

[0110] Furthermore, the Critic network updates the gradient by constructing a loss function. The loss function of the Critic network is, which is specifically expressed as:

[0111]

[0112] Among them, E represents expectation, w represents the critic network parameter, V w (s t ) represents the current value function; Represents the target value function used to evaluate the accuracy of the critic network output.

[0113] Based on the temporal difference algorithm, it is specifically expressed as:

[0114]

[0115] After constructing the loss function, the critic network parameters are updated through gradient calculation, which is specifically expressed as:

[0116]

[0117] in, is the loss function L V (w) Gradient with respect to parameter w; η w is the Critic network learning rate.

[0118] Furthermore, the Actor network is updated by the clipped alternative objective function, which Introducing the Actor network, specifically expressed as:

[0119]

[0120] Q w (s,a) means executing action a according to strategy π in the current state t The reward expectation, V w (s) represents the expected reward for executing all control schemes according to the strategy π in the current state, which is specifically expressed as:

[0121] Q w (s, a) = E(R t |s t =s,a t =a;π);

[0122] V w (s)=E(R t |s t =s;π).

[0123] Furthermore, the Actor network parameter update adopts the clipping objective function, which is specifically expressed as:

[0124]

[0125] Among them, θ represents the Actor network parameters, Represents the loss function of the Actor network The gradient of the parameter θ, η θ Represents the dynamic learning rate of the Actor network;

[0126] The dynamic learning rate of the actor network is calculated by the ratio of the sampling probabilities of the new and old strategies, which is specifically expressed as:

[0127]

[0128] Among them, η θ,base Represents the baseline learning rate of the Actor network; π θ (a t ,s t ) is the sampling probability of the new strategy, is the sampling probability of the old strategy; ε is the hyperparameter that limits the clipping interval; the CLIP function controls the change between the old and new strategies within the interval [1-ε,1+ε] to improve the stability of the algorithm.

[0129] The proximal policy optimization algorithm generates samples through continuous interaction with the model and performs batch training through gradient descent until the maximum iteration cycle is reached and the reward function converges. The Actor network can then be put into online application to generate the optimal control solution.

[0130] Specific implementation: This invention takes the flexible direct current transmission system as an example to further illustrate the flexible direct current system oscillation emergency control method based on deep reinforcement learning proposed in this invention. The flexible direct current transmission system topology is as follows: Figure 2 As shown, the flexible DC converter is connected to the 500kV AC grid through a 500 / 220kV and 220 / 110kV two-stage transformer, and the flexible DC converter adopts voltage and current dual closed-loop vector control.

[0131] On this basis, the schematic diagram of the designed flexible DC system oscillation emergency control method based on deep reinforcement learning is shown in the figure below. Figure 3 As shown in the figure. Based on the Markov reward process of the flexible DC transmission system scheduling model, the present invention uses a fully connected layer to build the strategy and value networks. Both networks contain two hidden layers, with the number of neurons in each layer being 64 and 32 respectively. The hidden layer uses the ReLU activation function. The learning rates of the strategy and value networks are μ o =0.0001, μ π =0.00001, the discount factor is 0.99, the clipping hyperparameter ε is 0.25, and the Adam optimizer is used to update the network. The training results are as follows Figure 4 As shown. Figure 4It can be seen that the designed reinforcement learning algorithm has converged and can be used for emergency control of oscillations in flexible transmission systems. In order to verify the effectiveness and accuracy of the above analysis, the flexible DC transmission power is set to 0.96pu, the grid strength SCR is 2.2, and a disturbance is introduced at t = 1s. Under the action of the disturbance, the power waveform at the PCC point where the flexible DC system is connected to the AC system is as follows: Figure 5 shown.

[0132] Furthermore, when the system does not incorporate oscillation decisions, the power at the PCC gradually diverges under the influence of disturbances, and the system gradually becomes unstable. Calculation of the oscillation mode shows that the real part of the oscillation mode, σ, of the flexible DC transmission system is 1.11 > -0.7. If, at t = 1.2s, the reinforcement learning decision is implemented (increasing the reactive power by 0.2 pu), the system oscillation converges after a brief transient transition. At this point, the real part of the oscillation mode, σ, decreases to -1.9, which is less than the designed threshold of -0.7, indicating that the system can operate reliably and stably.

[0133] The present invention also discloses a flexible DC system oscillation emergency control system based on deep reinforcement learning, comprising:

[0134] A signal acquisition module is used to obtain three-phase voltage and three-phase current signals at the access point of the flexible DC transmission system;

[0135] The preprocessing module uses a low-pass filter with a cutoff frequency of 10 kHz to filter out high-frequency noise in the signal and uses the improved Z-score method to detect and replace abnormal data points;

[0136] a modeling module, connected to the preprocessing module, and constructing a state space model of the flexible DC transmission system using a vector fitting method;

[0137] A stability judgment module, connected to the modeling module, calculates the characteristic roots of the state matrix A and extracts the real part of the characteristic roots to determine whether to trigger oscillation control;

[0138] A reinforcement learning decision module is connected to the stability judgment module, models the oscillation control problem as a Markov decision process, uses a proximal policy optimization algorithm to solve the MDP, and generates control instructions;

[0139] The command output module is connected to the reinforcement learning decision module and is used to send control commands to the flexible direct current transmission control center.

[0140] Those skilled in the art will appreciate that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for emergency control of oscillations of a flexible DC system based on deep reinforcement learning, characterized in that: The following steps are involved: Obtain three-phase voltage and three-phase current signals at the access point of the flexible DC transmission system; Preprocess the electrical signal, filter out high-frequency noise, and use the improved Z-score method to detect and replace abnormal data points; Based on the preprocessed data, the state space model of the HVDC Flexible system is constructed using the vector fitting method. Perform eigenvalue analysis on the state matrix A of the state space model of the flexible DC transmission system and extract the real part of the state characteristic root σ t , through σ t The value of determines whether to trigger the oscillation emergency control of the flexible DC transmission system; When the characteristic analysis result triggers the emergency control of the HVDC flexible system oscillation, it is converted into a Markov decision process MDP model; The proximal strategy optimization algorithm is used to solve the Markov decision process MDP model, generate control instructions and send them to the control center to perform emergency control of the oscillation of the flexible DC transmission system.

2. The method for emergency control of oscillation of a flexible DC system based on deep reinforcement learning according to claim 1, characterized in that: The preprocessing of the electrical signal comprises: According to the oscillation problem of the flexible DC transmission system, a low-pass filter with a cutoff frequency of 10kHz is used to filter out high-frequency noise; The implementation of the improved Z-score method includes the following steps: Construct a sliding window with the current data point x as the center, including 3 data points before and after; Calculate the mean x of the data in the window m and standard deviation x σ ; If|xx m | / x σ >3, then replace x with x m .

3. The method for emergency control of oscillation of a flexible DC system based on deep reinforcement learning according to claim 1, characterized in that: The method of constructing a state space model of a flexible DC transmission system using a vector fitting method includes the following steps: Construct the fitting transfer function, which is specifically expressed as: Among them, N is the initial order, r j is the residue, a j h fit (s) is the pole, d and k are rational numbers, and s is the complex frequency; Optimize the parameters by the least squares method so that h fit (s) approximates h(s) actual system transfer function, h fit (s) is the fitting expression of the transfer function h(s) of the HVDC Flexible system; The state space model of the HVDC Flexible system is specifically expressed as follows: in, represents the time derivative of the state variable, X, U, and Y are the state variable, input variable, and output variable respectively; A, B, C, and D are the state matrix, input matrix, output matrix, and direct transfer matrix respectively.

4. The method for emergency control of oscillation of a flexible DC system based on deep reinforcement learning according to claim 3, characterized in that: The state characteristic root real part σ is extracted by eigenvalue analysis t , and judging whether to trigger oscillation control, including the following steps: Calculate the characteristic root λ of the state matrix A of the flexible DC transmission system t =σ t +jω t , extract the real part of the characteristic root σ t As the basis for judging the system oscillation damping characteristics, ω t represents the oscillation frequency; When σ t When <-0.7, the system is judged to be in a strong damping state; When σ t When ≥-0.7, the system is judged to be in a weak damping state.

5. The method for emergency control of oscillation of a flexible DC system based on deep reinforcement learning according to claim 1, characterized in that: Construct a Markov decision process MDP model, including the definition: State space s t =(σ t ,Q t ,P t ,V t ), where σ t is the real part of the system characteristic root, Q t is the reactive power output of the system at time t, P t is the active power output of the system at time t, V t is the voltage of the node connected to VSC at time t; Action space A t is the controllable variable of the system at time t; Reward function R t The design is a constraint priority mechanism, specifically expressed as: Among them, α constrain is a set of system constraints, including node voltage constraints, active power output constraints, and reactive power output constraints; R t,p is the penalty for not meeting the constraint, R t,r The reward when the constraint is satisfied.

6. The method for emergency control of oscillation of a flexible DC system based on deep reinforcement learning according to claim 1, characterized in that: The proximal policy optimization algorithm is based on the Actor-Critic framework, where the policy actor network outputs the action probability distribution, the value critic network calculates the state value function, and updates the neural network parameters through gradient descent; Actor network parameter update uses the clipping objective function, which is specifically expressed as: The loss function of the Critic network is L V (w), specifically expressed as: Among them, θ represents the Actor network parameters, Represents the loss function of the Actor network The gradient of the parameter θ, η θ represents the dynamic learning rate of the Actor network, E() represents the expected operation, w represents the Critic network parameter, V w (s t ) represents the current value function; Represents the target value function used to evaluate the accuracy of the critic network output.

7. The method for emergency control of oscillation of a flexible DC system based on deep reinforcement learning according to claim 6, characterized in that: The advantage function Introducing the Actor network, specifically expressed as: Q w (s,a) represents the expected reward for executing action a according to strategy π in the current state s, V w (s) represents the expected reward for executing all control schemes according to the strategy π in the current state s, which is specifically expressed as: Q w (s,a)=E(R t |s t =s,a t =a;π); V w (s)=E(R t |s t =s;π); Among them, R t is the reward function, and E() represents the expected operation.

8. The method for emergency control of oscillation of a flexible DC system based on deep reinforcement learning according to claim 6, characterized in that: Based on the temporal difference algorithm, it is specifically expressed as: After constructing the loss function, the critic network parameters are updated through gradient calculation, which is specifically expressed as: in, is the loss function L V (w) Gradient of the weight parameter w of the Critic network; η w is the Critic network learning rate.

9. The method for emergency control of oscillation of a flexible DC system based on deep reinforcement learning according to claim 6, characterized in that: Dynamic learning rate η of the Actor network θ The calculation is performed by the ratio of the sampling probabilities of the new and old strategies, which can be expressed as: Among them, η θ,base Represents the baseline learning rate of the Actor network; π θ (a t ,s t ) is the sampling probability of the new strategy, is the sampling probability of the old strategy; ε is the hyperparameter that limits the clipping interval; the CLIP function controls the change between the old and new strategies within the interval [1-ε,1+ε].

10. A flexible DC system oscillation emergency control system based on deep reinforcement learning, characterized in that: The method for emergency control of oscillations of a flexible DC system based on deep reinforcement learning according to any one of claims 1 to 9 comprises: A signal acquisition module is used to obtain three-phase voltage and three-phase current signals at the access point of the flexible DC transmission system; The preprocessing module uses a low-pass filter with a cutoff frequency of 10 kHz to filter out high-frequency noise in the signal and uses the improved Z-score method to detect and replace abnormal data points; a modeling module, connected to the preprocessing module, and constructing a state space model of the flexible DC transmission system using a vector fitting method; A stability judgment module, connected to the modeling module, calculates the characteristic roots of the state matrix A and extracts the real part of the characteristic roots to determine whether to trigger oscillation control; A reinforcement learning decision module is connected to the stability judgment module, models the oscillation control problem as a Markov decision process, adopts a proximal strategy optimization algorithm to solve the MDP, and generates control instructions; an instruction output module is connected to the reinforcement learning decision module and is used to send the control instructions to the flexible direct current transmission control center.