A virtual inertia regulation method for DC transmission based on safety reinforcement learning
Through safety reinforcement learning and multi-agent collaborative decision-making, the virtual inertia adjustment of DC transmission is optimized, and the problems of inaccurate inertia evaluation and slow training speed in traditional methods are solved, and the stability and frequency response of the power grid are optimized, which avoids power over limits and improves the disturbance resistance of the power system.
Patent Information
- Application Number
- CN202411945867.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-12-27
AI Technical Summary
In power systems with high proportion of renewable energy, traditional virtual inertia evaluation methods are difficult to accurately reflect the system inertia level, resulting in DC transmission equipment being unable to effectively provide frequency support when frequency disturbances, and the existing reinforcement learning training speed is slow, which can easily lead to power overlimiting problems.
Using a method based on security reinforcement learning, through collaborative decision-making by multi-agents and combined with deep neural network models, virtual inertia adjustment of DC transmission is optimized, safety constraints are introduced, and virtual inertia parameters are dynamically adjusted within the safety boundary, and frequency response is optimized.
It improves the stability and disturbance resistance of the power grid, optimizes the virtual inertia response, avoids the problem of power overruns, and improves the training speed and decision-making accuracy of traditional reinforcement learning.
Smart Images

Figure CN119813330B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power systems, and in particular to a method for regulating virtual inertia of direct current transmission based on safety reinforcement learning. Background Art
[0002] In new power systems, inertia is a key metric for assessing grid stability. Traditional power systems rely on the rotating mechanical inertia of synchronous generators to mitigate frequency fluctuations. However, as power systems transition to a high proportion of renewable energy, the level of inertia in the grid is gradually decreasing, and the problem of insufficient system inertia is becoming increasingly prominent. For synchronous generators, mechanical inertia can be directly calculated from the generator's rotational mass and operating speed. This inertia supports the stability of the power system during frequency disturbances. For renewable energy and other power electronic devices, virtual inertia is an inertial response artificially simulated through control strategies. Research has been conducted on the call and evaluation of virtual inertia for devices such as DC link capacitors, supercapacitors, and energy storage batteries. However, existing virtual inertia assessments often rely on empirical or linear models, which struggle to accurately reflect the system's inertia level under complex dynamic conditions. Furthermore, the capacity limitations and response speed of energy storage devices also affect the effectiveness of virtual inertia.
[0003] LCC-HVDC (grid-commutated converter high-voltage direct current) and VSC-HVDC (voltage-source converter high-voltage direct current), important HVDC transmission technologies, have limited inertia provision capabilities. LCC-HVDC's dynamic response is slow when the grid is disturbed, making it unable to provide effective frequency support. VSC-HVDC has strong voltage and power control capabilities, but its low inertia, particularly in weak grids, makes it difficult to provide effective frequency support. Previous research has demonstrated how VSC (voltage-source converter) emulates synchronous machine inertia, modulating power according to a frequency reference command and schedule to achieve inertia response, primary and secondary frequency modulation, and ultimately frequency support. Whether it is an LCC-HVDC or VSC-HVDC system, the research and implementation of virtual inertia control is aimed at improving grid stability and dynamic response capabilities. The setting of virtual inertia parameters plays a key role in the strength of DC's support for frequency stability. Traditional virtual inertia parameter settings are set in advance based on DC capacity and predicted system operating status, which is difficult to cope with complex and changeable real-time operating conditions. Real-time data feedback and prediction algorithms or intelligent control algorithms are needed to achieve more accurate virtual inertia simulation.
[0004] Safety reinforcement learning (SRL) builds upon standard RL by incorporating safety constraints and mechanisms. This allows the agent to not only pursue optimal control objectives when making decisions, but also ensure that the system remains within safety boundaries during exploration. Previous studies have applied SRL to three typical power system problems: frequency regulation, voltage control, and energy management. In HVDC systems, RL can dynamically adjust virtual inertia based on grid status and disturbances, optimizing the system's frequency response. However, improper inertia settings can easily lead to power over-limit issues in DC systems. This problem is caused by multiple factors, including the DC's own frequency response, the magnitude of grid frequency or power fluctuations, and the characteristics of the grid's frequency response. The introduction of complex constraints slows down traditional RL training and reduces decision-making accuracy. Summary of the Invention
[0005] In response to the above problems, the present invention provides a virtual inertia adjustment method for direct current transmission based on safety reinforcement learning, which helps to improve the stability and anti-disturbance capability of the power grid.
[0006] The technical solution of the present invention is: a method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning, comprising the following steps:
[0007] (1) According to the control logic of LCC-HVDC and VSC-HVDC with adjustable inertia in the system, determine the virtual inertia parameters in the control logic;
[0008] (2) Evaluate the inertia of the synchronous machine, asynchronous motor, new energy and energy storage in the system to obtain the current system inertia;
[0009] (3) Based on steps (1) and (2), a constrained Markov decision process for multi-infeed DC inertia control of regional power grid is proposed, including action, state, benefit, environment and decision constraints;
[0010] (4) Based on step (3), a multi-agent method is used for training, and each agent competes with each other until the optimal decision-making agent is obtained for online decision-making.
[0011] In step (1),
[0012] For LCC-HVDC, by introducing the differential links of series and parallel connection, the specific mathematical expression is as follows:
[0013]
[0014] Where, ΔI dcref Indicates the DC current reference correction value, K f1 and K f2Represent the coefficients of the frequency differential term and the proportional term, f and f respectively N They represent the current frequency and rated frequency of the system respectively, and dt represents the time derivative;
[0015] For VSC-HVDC, the specific mathematical expression is as follows:
[0016] ΔP ref =(h1s+d1)(ff N )
[0017] Where ΔP ref represents the active reference correction value, h1 and d1 represent the virtual inertia and primary frequency modulation coefficient of the VSC, respectively, and s is a complex variable in the complex frequency domain.
[0018] In step (2),
[0019] 1) Synchronous motor inertia H syn Expressed as:
[0020]
[0021] Where, E syn is the rotor energy storage at the rated speed of the synchronous machine, S syn is the rated capacity, J is the moment of inertia of the synchronous machine rotor, ω 0m is the rated speed of the synchronous machine;
[0022] 2) Inertia of asynchronous motor H asy Expressed as:
[0023]
[0024] Where, K1, K2, K3 are the asynchronous motor control parameters, H am is the moment of inertia of the asynchronous motor rotor;
[0025] 3) The inertia of new energy and energy storage is expressed as:
[0026]
[0027] Among them, E re is the equivalent rotor kinetic energy of the new energy inertia response, H i is half of the inertia control coefficient, L is the set of new energy units; S re,i is the capacity of the i-th new energy source in the system;
[0028] Current system inertia H sys Expressed as:
[0029]
[0030] Among them, E kis the sum of the kinetic energy of all units, S total is the sum of all unit capacities, G is the set of synchronous machines, S asy is the sum of the asynchronous machine capacities, H syn,i and S syn,i are the inertia and capacity of the i-th synchronous machine in the system respectively; H re,i is the inertia of the i-th new energy device in the system; H asy,i is the inertia of the i-th asynchronous machine in the system; S re is the total capacity of the new energy units in the system, K i is the load ratio of the asynchronous motor on bus i, L i is the total load of busbar i, and α is the correlation coefficient of asynchronous motor energy release.
[0031] In step (3), the action of the constrained Markov decision process is the virtual inertia parameter of the LCC-HVDC and VSC-HVDC control logic. The specific action formula of the constrained Markov decision process is:
[0032] a t =x t ,x t ∈[0,1]
[0033] Among them, a t To constrain the actions of the Markov decision process, x t is the corresponding inertia parameter.
[0034] In step (3), the state of the constrained Markov decision process includes the following variables: the output upper limit vector p of each controllable DC in the region upper , the output lower limit vector p of each controllable DC in the region lower , the current output vector p of each controllable DC in the area con 、Current system inertia H sys 、The current system DC output vector p dc ; The state formula of the specific constrained Markov decision process is:
[0035] s t =(p upper ,p lower , p con ,H sys ,p dc )
[0036] Among them, s t is the state of the constrained Markov decision process.
[0037] In step (3), the benefit of the constrained Markov decision process includes the system frequency response improvement benefit r f , DC inertia support cost c ine; The specific constraint Markov decision process profit formula is:
[0038] r t =r f -c ine
[0039]
[0040] Among them, r t To constrain the payoff of the Markov decision process, r f For the system frequency response improvement benefit, c ine is the DC inertia support cost; M is the number of frequency observation points, N is the number of controllable DC, k rf 、k i,1 、k i,2 is the profit calculation coefficient, Δf j , Δf m is the absolute value of the frequency change rate between the j-th and m-th observation points in the frequency response.
[0041] In step (3), the environment of the constrained Markov decision process is the regional power grid frequency response N-1 scanning tool. After inputting the parameters of each controllable DC virtual inertia in the environment, the tool automatically simulates the power grid frequency response after each DC blocking and extracts the calculated state s t With the return r t Required variables, output calculated state s t With the return r t .
[0042] In step (3), in the constrained Markov decision process, the decision constraints include the output constraints of each line and the frequency change rate constraints of each observation point. The constraint formula is as follows:
[0043]
[0044] Δf j ≤Δf upper
[0045] in, and are the upper and lower limits of the output of the i-th DC line, is the maximum output of the ith DC line in the frequency response, Δf upper It is the absolute value of the maximum frequency change rate allowed by the system.
[0046] In step (4), each agent is responsible for different control objectives, among which agent A is responsible for optimizing the system frequency response to improve the benefit r f , Agent B is responsible for optimizing the DC inertia support cost c ine , achieving collaborative decision-making through a centrally trained decentralized execution framework.
[0047] In step (4), a deep neural network model is used as the security layer of security reinforcement learning and integrated into multi-agent security reinforcement learning, including security layer verification, agent execution action, environment state update, experience replay and sampling, model evaluation, optimal model judgment, and model preservation; wherein, the neural network input during security layer training is the state of the constrained Markov decision process and the action taken by the current agent;
[0048] Specifically including: the output upper limit vector p of each controllable DC in the area upper , the output lower limit vector p of each controllable DC in the region lower , the current output vector p of each controllable DC in the area con 、Current system inertia H sys 、The current system DC output vector p dc , the virtual inertia parameter vector x of each LCC-HVDC and VSC-HVDC control logic; the safety layer input neural network is as follows:
[0049] input=(p upper ,p lower , p con ,H sys ,p dc ,x)
[0050] The output of the neural network during safety layer training is whether the output of each line exceeds the limit and whether the frequency change rate of each observation point exceeds the allowable value;
[0051] O s =(P state ,Δf state )
[0052] Among them, P istate ,Δf state Indicates whether each DC exceeds the limit and whether the frequency change of each observation point exceeds the limit. s is the output of the security layer;
[0053] Before integrating the trained deep neural network model into the deep reinforcement learning action, the two intelligent agents first select an action and input the current state space and action into the trained deep neural network model. If the action satisfies the line output constraints and the frequency change rate constraints of each observation point, the action is executed to continue training. If the constraints are not met, the action is resampled until the constraints are met.
[0054] During operation, the present invention adjusts the virtual inertia parameters of multiple DC systems online based on measured grid information, enabling it to cope with the complex and ever-changing real-time operating conditions of the grid and optimize the system's frequency response characteristics. To address the problem of DC power over-limits caused by improper inertia settings, the present invention uses a built-in safety layer to learn the relationship between decision variables and constraint over-limits. During training, the present invention analyzes various factors that lead to power over-limits, including the DC's own frequency response, grid frequency or power fluctuations, and grid frequency response characteristics. This improves the training speed of traditional reinforcement learning and optimizes the power system's virtual inertia response while ensuring safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a flow chart of the method of the present invention;
[0056] Figure 2 This is the LCC-HVDC frequency support control logic diagram;
[0057] Figure 3 It is the VSC-HVDC frequency support control logic diagram;
[0058] Figure 4 is a schematic diagram of the IEEE 14-node standard test example in the present invention,
[0059] Figure 5 is a schematic diagram of system inertia distribution before optimization;
[0060] Figure 6 It is a schematic diagram of the inertia distribution of the optimized system;
[0061] Figure 7 It is a virtual inertia decision-making agent training process based on safety reinforcement learning. DETAILED DESCRIPTION
[0062] like Figure 1-7 As shown, the present invention proposes a method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning, comprising the following steps:
[0063] (1) According to the control logic of the LCC-HVDC and VSC-HVDC with adjustable inertia in the system, the virtual inertia parameter in the control logic is determined. Specifically, the coefficient of the control link that adjusts the active output by the frequency change rate is the virtual inertia parameter; the virtual inertia parameter is the adjustment object of the subsequent algorithm.
[0064] (2) Evaluate the inertia of other devices in the system except the inertia-adjustable DC, which is used to constitute the state variables in the Markov decision process and input into the intelligent agent to support the virtual inertia decision process, specifically including the inertia of synchronous machines, asynchronous motors, new energy and energy storage;
[0065] (3) A constrained Markov decision process for multi-input DC inertia control in regional power grids is proposed to realize the interaction between the intelligent agent and the training environment in the subsequent training process. In order to support the safety reinforcement learning training process, in addition to the four elements of action, state, benefit, and environment, it also includes decision constraints for safety layer training.
[0066] Among them, the action is the control parameter of the inertia-adjustable DC, the state contains the information required for DC inertia decision-making, the benefit includes the system frequency response improvement benefit and various control costs, the environment is the system frequency response N-1 scanning program, and the constraints include frequency safety constraints and DC output constraints.
[0067] (4) A multi-agent method is used for training, where each agent competes with each other until the optimal decision-making agent is obtained, that is, a DC virtual inertia decision-making strategy is constructed for online decision-making. At the same time, to ensure that the agent's actions meet the safety constraints in the system, a deep neural network model is trained as a safety layer. The safety layer is introduced during the reinforcement learning training process. Before each action is executed, it is judged whether the action meets the safety constraints. If not, the action is resampled and replaced until the constraints are met.
[0068] In the application, after the optimal intelligent agent is trained, the state elements of the constrained Markov decision process in step (3) are collected from the actual system and input into the intelligent agent. The intelligent agent will automatically output the optimal virtual inertia of each DC to realize the online decision process.
[0069] The reinforcement learning-based online inertia adjustment method proposed in this paper adjusts the virtual inertia parameters of multiple DC systems online based on measured grid information. This method can address the complex and ever-changing real-time operating conditions of the power grid and optimize the system's frequency response characteristics. To address the problem of DC power over-limits caused by improper inertia settings, a built-in safety layer is used to learn the relationship between decision variables and constraint over-limits. During training, the method analyzes various factors that lead to power over-limits, such as the DC's own frequency response, grid frequency or power fluctuations, and grid frequency response characteristics. This method improves the training speed of traditional reinforcement learning and optimizes the power system's virtual inertia response while ensuring safety.
[0070] In step (1), the specific control logics of LCC-HVDC and VSC-HVDC are as follows: Figure 2 、 Figure 3 As shown, the following are the specific introductions:
[0071] For LCC-HVDC, by introducing series and parallel differential links on the basis of traditional proportional control, LCC-HVDC can realize the inertia support function of the AC system, thereby effectively suppressing the AC frequency change rate under disturbance conditions. Its control block diagram is shown in the figure below. In this control mode, LCC-HVDC simulates the inertial response mechanism of the synchronous generator, thereby providing inertia support for the power system. The specific mathematical expression is as follows:
[0072]
[0073] Where, ΔI dcref Indicates the DC current reference correction value, K f1 and K f2 Represent the coefficients of the frequency differential term and the proportional term, f and f respectively N Respectively represent the current frequency and rated frequency of the system, d represents the derivative, and dt represents the derivative with respect to time. The inertia of these two controls is determined by the coefficient K of the frequency differential term. f1 reflect.
[0074] In formula (1), the above formula is parallel differential, corresponding to Figure 2 Figure (a); the following formula is a series differential, corresponding to Figure 2 Middle (b) Figure;
[0075] As can be seen from the above formula, the core principle of LCC-HVDC to achieve inertia support is to detect the frequency change rate of the sending or receiving system on the rectifier side and linearly adjust the amplitude of the DC current, thereby achieving rapid modulation of the DC power and slowing down the frequency change rate.
[0076] For VSC-HVDC, by adding a frequency correction link before the power control loop, the active output and the derivative component of the frequency can be correlated, so that the VSC has the ability to support the system inertia. The specific mathematical expression is as follows:
[0077] ΔP ref =(h1s+d1)(ff N ) (2)
[0078] Where, ΔP ref represents the active reference correction value, h1 and d1 represent the virtual inertia and primary frequency modulation coefficient of the VSC, respectively, and s is a complex variable in the complex frequency domain. The inertia provided by the VSC to the system is h1.
[0079] In this invention, real-time control of system inertia is the action of the intelligent agent. Simultaneously, its active output regulation range is constrained by the sending-end converter. Because the active output of LCCs and VSCs has upper and lower limits, the inertia support provided by these devices to the system also has upper and lower limits.
[0080] In step (2), the system inertia evaluation method is as follows:
[0081] System inertia is a crucial parameter for maintaining power system stability during external disturbances, directly impacting the magnitude and speed of system frequency fluctuations. The intelligent agent assesses the adequacy of the current system's inertia based on real-time monitoring. If the system inertia is insufficient, the intelligent agent makes decisions based on pre-set control strategies, dynamically adjusting the system's active power output to ensure the grid maintains stability and security despite insufficient inertia, effectively improving the smart grid's dynamic response and disturbance tolerance.
[0082] The current inertia measurement can adopt an estimation method based on the equivalent aggregation model to statistically sum different inertia sources. First, it is necessary to aggregate the models of various resources to obtain the inertia contributions from different resources.
[0083] 1) Synchronous motor inertia
[0084] Inertia is an inherent property of objects, manifesting as resistance to changes in motion. Power system inertia manifests as resistance to frequency changes caused by external disturbances, and is a key factor in ensuring system frequency stability. Generators and prime movers rely on unbalanced torque to accelerate or decelerate. Taking into account the damping coefficient, for a single synchronous machine, this can be written as:
[0085]
[0086] Where H syn is the inertia of the synchronous machine, T m -T e is the difference between the mechanical torque and the electromagnetic torque, D is the attenuation coefficient, and Δω is the deviation of the rotor speed from the rated speed. dω / dt is the derivative of ω.
[0087] The fundamental reason for the existence of inertia in a power system is that the generator speed and the system maintain a certain relationship. Inertia is the resistance of the synchronous generator. For a single synchronous generator, the inertia constant of the i-th synchronous generator, or in this case, the inertia of the synchronous generator, is usually defined as:
[0088]
[0089] E syn The energy stored in the rotor of the synchronous machine at rated speed, usually in MWs, S syn is the rated capacity in MVA, J is the moment of inertia of the synchronous machine rotor, ω 0m is the rated speed of the synchronous machine. For hydropower units, H syn The value of is usually 2-4s, and that of thermal power plants is usually 3-9s.
[0090] 2) Inertia of asynchronous motor Hasy Expressed as:
[0091]
[0092] K1, K2, K3 are asynchronous motor control parameters, H am is the moment of inertia of the asynchronous motor rotor.
[0093] Therefore, measuring the inertia response of an asynchronous motor when the frequency changes is a challenge. To this end, the present invention proposes the following measurement method. The perturbation measurement method shows that the total rotor kinetic energy and the inertia constant of the asynchronous motor are linearly related, and this relationship can be approximated using a linear expression.
[0094]
[0095] Where K i is the load ratio of the asynchronous motor on bus i, L i is the total load of busbar i, H asy,i is the inertia of the i-th asynchronous motor in the system. By setting the time window to 0.1s-0.4s, setting the system power disturbance and the least squares linearization curve, the slope can be measured and α = 0.9868 is obtained; M is the set of asynchronous motors; E asy is the total kinetic energy of all asynchronous motors in the system.
[0096] 3) Inertia of new energy and energy storage
[0097] The impact of new energy and energy storage can be expressed as:
[0098]
[0099] Among them, E re is the equivalent rotor kinetic energy of the new energy inertia response, H i is half of the inertia control coefficient, L is the set of new energy units; S re,i is the capacity of the i-th new energy source in the system.
[0100] Therefore, the inertia response capability of the entire system can be measured, which can be achieved by summing the equivalent rotational kinetic energy. When the system has a large amount of out-of-zone DC, in order to be able to measure the current inertia in different scenarios, the inertia constant is written as follows.
[0101]
[0102] Among them, E k is the sum of the kinetic energy of all units, S total is the sum of all unit capacities, G is the set of synchronous machines, S asyis the sum of the asynchronous machine capacities, and the variable subscript i represents the unit number. syn,i and S syn,i is the inertia and capacity of the i-th synchronous machine in the system; H re,i is the inertia of the i-th new energy device in the system; H asy,i is the inertia of the i-th asynchronous machine in the system; S re is the total capacity of new energy units in the system,
[0103] In step (3), the action of the constrained Markov decision process is the virtual inertia parameter x of the LCC-HVDC and VSC-HVDC control logic i , where i is the DC number; this parameter is normalized data, with a value range of [0,1]. i ∈[0,1], the corresponding DC virtual inertia parameter is set to H i,max , where H i,max is the maximum inertia of the ith DC line under the response speed limit, corresponding to the control parameter K in formulas (1) and (2) respectively. f1 , K f1 , h1. The specific action formula of the constrained Markov decision process is:
[0104] a t =x t ,x t ∈[0,1] (10)
[0105] Among them, a t is the action of the constrained Markov decision process. t is the corresponding inertia parameter.
[0106] In step (3), the state of the constrained Markov decision process includes the following variables: the output upper limit vector p of each controllable DC in the region upper , the output lower limit vector p of each controllable DC in the region lower , the current output vector p of each controllable DC in the area con 、Current system inertia H sys , the current system DC output vector (that is, only considering the large power fluctuations that may occur in the system under N-1 conditions) p dc The specific state formula of the constrained Markov decision process is:
[0107] s t =(p upper ,p lower , p con ,H sys ,p dc ) (11)
[0108] Among them, s tis the state of the constrained Markov decision process
[0109] In step (3), the benefit of the constrained Markov decision process includes the system frequency response improvement benefit r f , DC inertia support cost c ine The payoff formula for the constrained Markov decision process is:
[0110] r t =r f -c ine (12)
[0111]
[0112] Among them, r t To constrain the payoff of the Markov decision process, r f For the system frequency response improvement benefit, c ine is the DC inertia support cost; M is the number of frequency observation points, N is the number of controllable DC, k rf 、k i,1 、k i,2 is the profit calculation coefficient, Δf j , Δf m is the absolute value of the frequency change rate between the j-th and m-th observation points in the frequency response.
[0113] In step (3), the environment of the constrained Markov decision process is the regional power grid frequency response N-1 scanning tool. After inputting the parameters of each controllable DC virtual inertia in the environment, the tool automatically simulates the power grid frequency response after each DC blocking and extracts the calculated state s t With the return r t Required variables, output calculated state s t With the return r t .
[0114] In step (3), in the constrained Markov decision process, in addition to the four elements of action, state, benefit, and environment, decision constraints are also included for safety layer training. Specific decision constraints include the output constraints of each line and the frequency change rate constraints of each observation point. The constraint formula is as follows:
[0115]
[0116] Δf j ≤Δf upper (16)
[0117] in, and are the upper and lower limits of the output of the i-th DC line, Δf is the maximum output of the i-th DC in the frequency response. upperIt is the absolute value of the maximum frequency change rate allowed by the system.
[0118] Based on the above constraints, two constraint vectors are formed, namely, the frequency change rate exceeding limit vector Δf of each observation point of the system after N-1 scans over , after N-1 scans, the controllable DC power over-limit situation vector Δp of the system con,over , the constraint vector will be used for training the safety layer in the subsequent reinforcement learning algorithm.
[0119] In step (4), a multi-agent method is used for training. Each agent is responsible for different control objectives. Agents compete with each other to optimize the performance of the entire system in collaboration. Among them, agent A is responsible for optimizing the system frequency response to improve the benefit r f , Agent B is responsible for optimizing the DC inertia support cost c ine , this collaborative decision-making is achieved through the Central Training Decentralized Execution (CTDE) framework. The dual-agent collaborative decision-making process is shown in Figure 7 .
[0120] In step (4), a deep neural network model is used as the security layer of safety reinforcement learning and integrated into multi-agent safety reinforcement learning, including safety layer verification (i.e., whether the DC power does not exceed the limit and the frequency change does not exceed the limit. If so, the next action is executed. If not, the action that meets the safety constraint conditions is selected and the action is repeated), agent execution action, environment state update, experience playback and sampling, model evaluation, optimal model judgment, and model saving steps, such as Figure 7 As shown in Figure 2. The neural network input during safety layer training is the state of the constrained Markov decision process and the action taken by the current agent. Specifically, it includes: the output upper limit vector p of each controllable DC in the region upper , the output lower limit vector p of each controllable DC in the region lower , the current output vector p of each controllable DC in the area con 、Current system inertia H sys , the current system DC output vector (that is, only considering the large power fluctuations that may occur in the system under N-1 conditions) p dc , the virtual inertia parameter vector x of each LCC-HVDC and VSC-HVDC control logic. The safety layer input neural network is as follows:
[0121] input=(p upper ,p lower , p con ,H sys ,p dc ,x) (19)
[0122] The output of the neural network during safety layer training is whether the output of each line exceeds the limit and whether the frequency change rate of each observation point exceeds the allowable value.
[0123] O s =(P state ,Δf state ) (20)
[0124] Among them, P istate ,Δf state Indicates whether each DC exceeds the limit and whether the frequency change of each observation point exceeds the limit. s This is the output of the security layer.
[0125] Before integrating the trained deep neural network model into the deep reinforcement learning action, the two intelligent agents first select an action and input the current state space and action into the trained deep neural network model. If the action satisfies the line output constraints and the frequency change rate constraints of each observation point, the action is executed to continue training. If the constraints are not met, the action is resampled until the constraints are met.
[0126] Safety reinforcement learning technology learns the relationship between decision variables and constraint violations through a built-in safety layer. It has good performance in dealing with decision problems with complex constraints and can optimize the virtual inertia response of the power system while ensuring safety.
[0127] In specific applications,
[0128] In order to verify the regulation effect of the deep reinforcement learning method proposed in this invention on the virtual inertia of DC transmission, the IEEE 14-node standard test case is used for simulation. The simulation is carried out based on PSCAD / EMTDC. The node system grid structure is as follows: Figure 4 As shown. In this test example, node 4 is connected to the synchronous thermal power generator set, represented by G4; nodes 1, 2, and 3 are all connected to the grid-following VSC, represented by G1, G2, and G3 respectively, and node 5 is connected to the grid-forming VSC, represented by G5. In this example, node 4 and node 2 are the points with the shortest and longest electrical distances from the synchronous units, respectively. According to the negative correlation between node frequency and node inertia, the inertia distribution of the system can be reflected by measuring the bus frequency of node 4 and node 2. In order to reflect the frequency change under the node load disturbance state, this example is set to increase the load at node 13 at 4s. Before and after optimization using the deep reinforcement learning method, the measured frequencies of bus 4 and bus 2 are as follows: Figure 5 and Figure 6 After optimization, the frequency fluctuations of bus 2 in the initial, stable, and disturbed states are significantly reduced, and the frequency deviation from bus 4 is significantly reduced. This shows that the deep reinforcement learning method designed in this invention can effectively improve the system inertia distribution and make the inertia distribution of each node more uniform.
[0129] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be covered by the scope of the claims of the present invention.
Claims
1. A method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning, characterized in that: The following steps are involved: (1) According to the control logic of LCC-HVDC and VSC-HVDC with adjustable inertia in the system, determine the virtual inertia parameters in the control logic; (2) Evaluate the inertia of the synchronous machine, asynchronous motor, new energy and energy storage in the system to obtain the current system inertia; (3) Based on steps (1) and (2), a constrained Markov decision process for multi-infeed DC inertia control of regional power grid is proposed, including action, state, benefit, environment and decision constraints; (4) Based on step (3), a multi-agent method is used for training, and each agent competes with each other until the optimal decision agent is obtained for online decision-making; In step (2), 1) Synchronous motor inertia H syn Expressed as: Where, E syn is the rotor energy storage at the rated speed of the synchronous machine, S syn is the rated capacity, J is the moment of inertia of the synchronous machine rotor, ω 0m is the rated speed of the synchronous machine; 2) Inertia of asynchronous motor H asy Expressed as: Where, K1, K2, K3 are the asynchronous motor control parameters, H am is the moment of inertia of the asynchronous motor rotor; 3) The inertia of new energy and energy storage is expressed as: Among them, E re is the equivalent rotor kinetic energy of the new energy inertia response, H i is half of the inertia control coefficient, L is the set of new energy units; S re,i is the capacity of the i-th new energy source in the system; Current system inertia H sys Expressed as: Among them, E k is the sum of the kinetic energy of all units, S total is the sum of all unit capacities, G is the set of synchronous machines, S asy is the sum of the asynchronous machine capacities, H syn,i and S syn,i are the inertia and capacity of the i-th synchronous machine in the system respectively; H re,i is the inertia of the i-th new energy device in the system; H asy,i is the inertia of the i-th asynchronous machine in the system; S re is the total capacity of the new energy units in the system, K i is the load ratio of the asynchronous motor on bus i, L i is the total load of busbar i, and α is the correlation coefficient of asynchronous motor energy release.
2. The method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning according to claim 1, characterized in that: In step (1), For LCC-HVDC, by introducing the differential links of series and parallel connection, the specific mathematical expression is as follows: Where, ΔI dcref Indicates the DC current reference correction value, K f1 and K f2 Represent the coefficients of the frequency differential term and the proportional term, f and f respectively N They represent the current frequency and rated frequency of the system respectively, and dt represents the time derivative; For VSC-HVDC, the specific mathematical expression is as follows: ΔP ref =(h1s+d1)(ff N ) Where ΔP ref represents the active reference correction value, h1 and d1 represent the virtual inertia and primary frequency modulation coefficient of the VSC, respectively, and s is a complex variable in the complex frequency domain.
3. The method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning according to claim 1, characterized in that: In step (3), the action of the constrained Markov decision process is the virtual inertia parameter of the LCC-HVDC and VSC-HVDC control logic. The specific action formula of the constrained Markov decision process is: a t =x t ,x t ∈[0,1] Among them, a t To constrain the actions of the Markov decision process, x t is the corresponding inertia parameter.
4. The method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning according to claim 1, characterized in that: In step (3), the state of the constrained Markov decision process includes the following variables: the output upper limit vector p of each controllable DC in the region upper , the output lower limit vector p of each controllable DC in the region lower , the current output vector p of each controllable DC in the area con 、Current system inertia H sys 、The current system DC output vector p dc ; The state formula of the specific constrained Markov decision process is: s t =(p upper ,p lower ,p con ,H sys ,p dc ) Among them, s t is the state of the constrained Markov decision process.
5. The method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning according to claim 1, characterized in that: In step (3), the benefit of the constrained Markov decision process includes the system frequency response improvement benefit r f , DC inertia support cost c ine ; The specific constraint Markov decision process profit formula is: r t =r f -c ine Among them, r t To constrain the payoff of the Markov decision process, r f For the system frequency response improvement benefit, c ine is the DC inertia support cost; M is the number of frequency observation points, N is the number of controllable DC, k rf 、k i,1 、k i,2 is the profit calculation coefficient, Δf j , Δf m is the absolute value of the frequency change rate between the j-th and m-th observation points in the frequency response.
6. The method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning according to claim 1, characterized in that: In step (3), the environment of the constrained Markov decision process is the regional power grid frequency response N-1 scanning tool. After inputting the parameters of each controllable DC virtual inertia in the environment, the tool automatically simulates the power grid frequency response after each DC blocking and extracts the calculated state s t With the return r t Required variables, output calculated state s t With the return r t .
7. The method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning according to claim 1, characterized in that: In step (3), in the constrained Markov decision process, the decision constraints include the output constraints of each line and the frequency change rate constraints of each observation point. The constraint formula is as follows: Δf j ≤Δf upper in, and are the upper and lower limits of the output of the i-th DC line, is the maximum output of the ith DC line in the frequency response, Δf upper It is the absolute value of the maximum frequency change rate allowed by the system.
8. The method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning according to claim 1, characterized in that: In step (4), each agent is responsible for different control objectives, among which agent A is responsible for optimizing the system frequency response to improve the benefit r f , Agent B is responsible for optimizing the DC inertia support cost c ine , achieving collaborative decision-making through a centrally trained decentralized execution framework.
9. The method for adjusting virtual inertia of direct current transmission based on safety reinforcement learning according to claim 8, characterized in that: In step (4), a deep neural network model is used as the security layer of security reinforcement learning and integrated into multi-agent security reinforcement learning, including security layer verification, agent execution action, environment state update, experience replay and sampling, model evaluation, optimal model judgment, and model preservation; wherein, the neural network input during security layer training is the state of the constrained Markov decision process and the action taken by the current agent; Specifically including: the output upper limit vector p of each controllable DC in the area upper , the output lower limit vector p of each controllable DC in the region lower , the current output vector p of each controllable DC in the area con 、Current system inertia H sys 、The current system DC output vector p dc , the virtual inertia parameter vector x of each LCC-HVDC and VSC-HVDC control logic; the safety layer input neural network is as follows: input=(p upper ,p lower ,p con ,H sys ,p dc ,x) The output of the neural network during safety layer training is whether the output of each line exceeds the limit and whether the frequency change rate of each observation point exceeds the allowable value; O s =(P state ,Δf state ) Among them, P istate ,Δf state Indicates whether each DC exceeds the limit and whether the frequency change of each observation point exceeds the limit. s is the output of the security layer; Before integrating the trained deep neural network model into the deep reinforcement learning action, the two intelligent agents first select an action and input the current state space and action into the trained deep neural network model. If the action satisfies the line output constraints and the frequency change rate constraints of each observation point, the action is executed to continue training. If the constraints are not met, the action is resampled until the constraints are met.
Citation Information
Patent Citations
Virtual inertia target generation method of virtual synchronous generator
CN117411064A
Method for increasing power of power transmission corridor containing embedded direct current
CN118646099A