Multi-section transmission limit calculation method based on multi-objective transferable reinforcement learning

By employing a multi-objective transferable reinforcement learning method, an inverse search model, and a decomposition strategy, the complexity of multi-section power transmission limit calculation is solved, achieving fast and accurate Pareto front output and improving computational efficiency and accuracy.

CN119691331BActive Publication Date: 2025-10-24SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411705909.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-10-24
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately calculate the transmission limit for multiple sections, which increases the complexity of system operation and the difficulty of scheduling. Moreover, the calculation results are often conservative or inaccurate, wasting potential and hindering the consumption of new energy sources.

Method used

We employ a multi-objective transferable reinforcement learning approach, combining an inverse search model, a decomposition strategy, and a Markov decision process with a neighborhood parameter transfer strategy and an improved DDPG algorithm to achieve an end-to-end Pareto front for multi-section coupling limits.

Benefits of technology

It enables rapid and accurate calculation of the Pareto front of multi-section coupling limits, saving 81.2% of the solution time, improving the supervolume index by 8.9%, and providing an automated multi-section transmission limit setting scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691331B_ABST
    Figure CN119691331B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of electrical engineering, and particularly discloses a multi-section power transmission limit calculation method based on multi-target migratory reinforcement learning, which comprises the following steps: considering transient power angle and voltage stability constraints, establishing a reverse search model for power transmission limit calculation; adopting a decomposition strategy to decompose the multi-section coupled limit calculation problem of the reverse search model into sub-problem models; converting different sub-problem models into Markov decision processes, and combining a neighborhood-based parameter migration strategy and an improved DDPG algorithm to solve the problems; and realizing end-to-end output of the Pareto front of the multi-section coupled limit through training of the model. The method has better convergence and diversity compared with traditional multi-target meta-heuristic algorithms.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electrical engineering, in particular to a multi-section power transmission limit calculation method based on multi-objective transferable reinforcement learning. BACKGROUND

[0002] With the continuous increase of new energy penetration and the substantial growth of load, the uncertainty of the power system is increasingly strong, and the system operation mode is more complex, making it difficult for the dispatching department to accurately identify extreme scenarios and accurately grasp the section limit. In order to reserve a certain safety margin, the power transmission limit is often intentionally lowered in engineering, which further wastes the potential of section power transmission and hinders the accommodation of new energy.

[0003] The calculation of section power transmission limit belongs to the solution of high-dimensional non-convex nonlinear model, which makes the calculation of power transmission limit in actual engineering often degenerate to the traditional engineering idea, relying on intensive manpower and operation experience to carry out limited "trial and error" checking of extreme modes, and its accuracy is difficult to verify. At present, there are few methods to calculate the section power transmission limit, mainly including expert experience method and computational geometry method. The rules made by the expert experience method may be too subjective to find the optimal solution from the mechanism to plan the whole, and the matching of the section limit in the expert database considers the conservative threshold and the rough division of typical operation modes, which reduces the section transmission capacity. The computational geometry method often loses part of the characteristics associated with the operation mode, and the calculation result of the limit is also conservative, and it has limitations when facing three-dimensional or more coupled sections.

[0004] The multi-section power transmission limit setting problem evolves into a non-convex nonlinear multi-objective optimization problem with differential equations. The solution of this problem is not limited to a single optimal solution, but a Pareto front composed of multiple optimal solutions, which can be considered as a stable convex boundary in the section power flow cut set space. The non-convex nonlinear model may have multiple local optimal solutions, difficulty in ensuring convergence, high computational complexity, and other problems, and is affected by multiple operation mode adjustments and operating state constraints. The actual Pareto boundary of this problem is often not a convex boundary in the ordinary sense, which further increases the difficulty of solving the model. SUMMARY

[0005] The purpose of the present application is to overcome the shortcomings of the prior art and provide a multi-section power transmission limit calculation method based on multi-objective transferable reinforcement learning.

[0006] The purpose of the present application is achieved by the following technical solution: a multi-section power transmission limit calculation method based on multi-objective transferable reinforcement learning, which comprises the following steps:

[0007] Step S1, considering transient power angle and voltage stability constraints, an inverse search model for power transmission limit calculation is established;

[0008] Step S2, using a decomposition strategy to decompose the multi-section coupling limit calculation problem of the reverse search model into sub-problem models;

[0009] Step S3: Convert different sub-problem models into Markov decision processes and solve them by combining the neighborhood-based parameter migration strategy with the improved DDPG algorithm;

[0010] Step S4: Implement the Pareto frontier (stable convex boundary) of the end-to-end output multi-section coupling limit by training the sub-problem model.

[0011] Specifically, the reverse search model for transmission quota calculation takes minimizing the external power transmission due to section instability as the objective function, and its expression is:

[0012]

[0013] Among them, P ij For the contact line I ij The transmission power along the specified direction; Z is the set of tie lines;

[0014] Specifically, the system power flow equation of the reverse search model for the transmission limit calculation is:

[0015]

[0016] Among them, i∈N, N is the set of all nodes; P Gi , Q Gi are the active power and reactive power output by the generator respectively; Q Ci Inject power into the reactive power compensation device; P Li , Q Li are the node active and reactive loads respectively; U i ,θ i is the node voltage amplitude and phase angle; Y ij , α ij is the magnitude and phase angle of the node admittance matrix;

[0017] The synchronous machine rotor motion equation of the reverse search model for the transmission limit calculation is:

[0018]

[0019] Where i∈S G , S G for all generators aggregated; and are the rotor speed at time τ and the rated speed of the rotor respectively; is the rotor angle; E′ di and E′ qiare the d-axis and q-axis components of the generator voltage, respectively; M i is the inertia constant; T Mi is the output mechanical power; D i is the damping torque coefficient; I di and I qi are the d-axis and q-axis components of the generator current, respectively; X di and X qi is the synchronous reactance; X' di and X' qi is the transient reactance; T' d0i and T' q0i are the open-circuit time constants of the d-axis and q-axis, respectively; E fdi is the excitation output voltage. Specifically, the constraint conditions of the inverse search model of the power transmission limit calculation include a steady-state operation constraint and a transient instability constraint, and the steady-state operation constraint is expressed as:

[0020]

[0021] where P Gi and Q Gi are the active power and reactive power output by the generator, respectively; U i is the node voltage amplitude; I ij is the current of the line l ij after a fault occurs; is the upper limit of the current of the line I ij ; S G represents a set of generators; N represents a set of nodes; and Z represents a set of tie lines.

[0022] Specifically, the transient instability constraint is:

[0023]

[0024] where x(τ) is a state variable; y(τ) is an algebraic variable; u is a control variable; x(τ0), y(τ0) represent an initial operating mode; τ is a simulation time step; (τ0, τ end ] represents a transient process time period; a transient power angle instability constraint and a transient voltage instability constraint;

[0025] a transient power angle stability evaluation index f TSI :

[0026]

[0027] where f TSI <0, the system is transiently unstable, and the larger f TSI , the more stable the system is.

[0028] a voltage stability evaluation index ηU,i :

[0029]

[0030] The voltage is stable when the bus voltage at the central point of the system is lower than 0.75pu and the duration is less than 1s, and the voltage is critically stable when the duration is within 0.8~1s; and the voltage is unstable when the bus voltage at the central point is less than 0.75pu and the duration is greater than 1s; where i∈N; U i,min The node voltage U in the transient process after the fault is cleared i Minimum value; U i,cr is the voltage threshold of node i, which is 0.75 pu; T cr To allow V i,min The duration of <0.75pu is 1s; T b The node voltage is lower than U i,cr The total duration of τ cr,j Indicates the jth U i,min =U i,cr time step; U i,sta is the voltage when node i returns to normal; k τ is the critical voltage shift factor, which is 0.75; when η U,i <0, it is determined that the busbar transient voltage is unstable;

[0031] The transient instability constraint satisfies the transient power angle instability constraint or the voltage instability constraint. The transient instability constraint is processed using the Big-M method. The model is easy to solve by rewriting equation (5) into the following equation:

[0032]

[0033] Among them, M1 and M2 are sufficiently small numbers; δ i is a 0-1 decision variable, when δ I =1, the i-th constraint is satisfied.

[0034] Specifically, when calculating the multi-section coupled transmission limit, multiple sections are optimized simultaneously, and the specific definition is as follows:

[0035]

[0036] in, Contains the state, algebra, and control variables of the power system; The objective function consists of m different independent section limit calculations; represent the system equality constraints and inequality constraints respectively, Including stable operation constraints and transient instability constraints.

[0037] Specifically, the multi-surface coupling limit calculation problem in step S2 is converted into n scalar optimization sub-problems by weighted sum, and the scalar objective function after conversion is represented as:

[0038]

[0039] wherein, represents the objective function of the i-th sub-problem, is a weight vector, and m is the number of solved sections. Specifically, the Markov decision process is composed of a five-tuple <S, A, r, P, γ>, which is specifically described as follows:

[0040] State space S:

[0041] S = [P G ,U G ,P L ,Q L ,P f ,f TSI ,η U ] (11)

[0042] wherein, P G is the active power of the generator; U G is the node voltage; P L is the active power of the load; Q L is the reactive power of the load; P f is the sum of the transmission power flow of all tie lines in the selected section; f TSI is a transient power angle stability evaluation index, and η U is a transient voltage stability evaluation index.

[0043] The action space A includes the generator action space a gen and the total load action space a load , a gen is the active power adjustment amount ΔP g of the generator, and the adjustment range is N respectively represent the maximum active power of the generator and the training step number of RL per round; a load is the active power adjustment amount ΔP Load_all of the load, and the adjustment range is represents the maximum active power of the total load fluctuation.

[0044] The reward function r is represented as:

[0045] r con = -[f TSI + M1(1-δ1)] - [η U + M2(1-δ2)] + (δ1+δ2-1) (12)

[0046]

[0047] where M1 and M2 are sufficiently small numbers; δ i is a 0-1 decision variable, when δ i = 1, the ith constraint is satisfied; r con is the reward value of the agent exploring the operation mode to satisfy the transient instability constraint, and is a continuous reward; g is the reward function calculation condition; r i- is the final reward value of the sub-problem i; r i-obj is the reward value of the scalar objective function of the sub-problem i after normalization; ω s , s = 1, 2 are weight coefficients of the reward items; K is a penalty given when the power flow is not converged, the line is overloaded, the power flow is reversed, or the generator is out of limit;

[0048] The discount factor γ, γ ∈ [0, 1] represents the discount of future rewards.

[0049] The state transition matrix P is determined by the interactive environment.

[0050] Specifically, in step S3, the MORL method is used to model the limit calculation sub-problem as a neural network by combining the decomposition method of the multi-objective calculation problem and the parameter migration strategy, so as to represent the network model parameters of the i-1th sub-problem, define [ω, b] as the network parameters that have not been optimized, and [ω * , b * ] as the optimized neural network parameters. It is assumed that the i-1th limit calculation sub-problem has been solved, and its network parameters have reached or are close to the optimal value. Therefore, when training the i th sub-problem, the experience knowledge of the latter can be used for fast solving, and the optimal network parameters of the latter are used as the starting parameters of the neural network of the former.

[0051] Specifically, the improved deep deterministic policy gradient method is used to train the MDP model of the sub-problem, so that the agent fully interacts with the environment, increases the random exploration ability, and considers adding random noise to the action:

[0052] a t = μ(s t | θ μ ) + ξ t (16)

[0053] where μ(s t | θ μ ) is the Actor network, and ξ t represents random noise.

[0054] After performing the action a t , the immediate reward r t+1 and the next state S t+1, forming a quadruple <S t , a t , r t+1 , S t+1 > is stored in the experience replay pool R, a Batch_size quadruples are randomly selected from R as the input of the current actor network and critic network, and the true value Q i is calculated.

[0055]

[0056] Wherein, gamma is a discount factor.

[0057] The critic network parameter theta Q and the actor network parameter theta μ are updated by minimizing the loss function and calculating the policy gradient respectively:

[0058]

[0059] Wherein, L is the loss value when the critic network is updated, and the Q value of the critic network is closer to the target Q value by minimizing the loss function; N is the number of samples sampled from the experience replay pool each time; Is the gradient of the actor network, which is calculated based on the Q value of the critic network to update the weight of the actor network; Is the gradient of the critic network Q value to the action, which reflects the influence of action selection on the Q value.

[0060] Finally, the obtained (theta Q , theta μ ) is updated to the target network parameter through a relatively stable soft update:

[0061]

[0062] Wherein, tau is a divergence factor, and the network parameters are updated by repeatedly iterating formula (16) to formula (20), and the algorithm finally converges.

[0063] The present application has the following advantages:

[0064] The application follows the solution idea of a multi-objective optimization problem, and proposes a multi-section power transmission limit calculation method based on multi-objective reinforcement learning, realizes the end-to-end mode to quickly output the accurate Pareto front of multi-section coupling limit, and the front is the stable convex boundary that needs to be found in engineering. Compared with the traditional multi-objective meta-heuristic algorithm, the application saves 81.2% of the solution time while grasping good convergence and diversity, and the hypervolume index is improved by at least 8.9%. In addition, the introduced neighbor-based parameter migration strategy enables the current sub-problem to learn from the model parameters of the previous sub-problem, which improves the convergence of the algorithm while greatly improving the training speed. The multi-objective transferable reinforcement learning method used in the application provides a new automatic solution for the multi-section power transmission limit setting which is difficult to set in engineering. BRIEF DESCRIPTION OF DRAWINGS

[0065] Figure 1 It is a flowchart of the calculation method of the application;

[0066] Figure 2 It is a sub-problem decomposition schematic diagram of the application;

[0067] Figure 3 It is a general framework schematic diagram of the MORL algorithm calculation of the application;

[0068] Figure 4 It is a neighbor-based parameter migration strategy process schematic diagram of the application;

[0069] Figure 5 It is a system structure and section division schematic diagram of the application;

[0070] Figure 6 It is a reward value curve schematic diagram in the training process of the first sub-problem of the application;

[0071] Figure 7 It is a section 1 and section 2 power flow change schematic diagram after the reward value curve is flattened of the application;

[0072] Figure 8 It is a Pareto front schematic diagram of the MORL in the two-section power transmission limit problem of the application;

[0073] Figure 9 It is a stable and unstable working condition schematic diagram in the two-section power flow space of the application;

[0074] Figure 10 It is a section power transmission limit search trajectory schematic diagram of the application;

[0075] Figure 11 It is a Pareto front comparison schematic diagram of the MORL and the NSGA-II algorithm of the application;

[0076] Figure 12 A schematic diagram contrasting the Pareto frontiers obtained by the MORL and MOEA / D algorithms of the present application;

[0077] Figure 13 A training performance comparison graph (I) of the present application using the parameter migration strategy and the original strategy;

[0078] Figure 14 A training performance comparison graph (II) of the present application using the parameter migration strategy and the original strategy. DETAILED DESCRIPTION

[0079] In order to make the purpose of the present application clearer, the technical solutions and advantages are described in further detail below in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application, i.e., the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations.

[0080] Therefore, the detailed description of the embodiments of the present application provided below in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without making creative efforts fall within the scope of the present application.

[0081] It should be noted that the relational terms such as "first" and "second" and the like are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus including a series of elements includes not only those elements, but also other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus including the element.

[0082] The present application will be further described below in combination with the accompanying drawings, but the scope of protection of the present application is not limited to the following description.

[0083] As shown in the following figure, the multi-section power transmission limit calculation method based on multi-objective migratory reinforcement learning includes the following steps: Figures 1 to 14

[0084] ​Step S1: Considering the transient power angle and voltage stability constraints, establish a reverse search model for transmission limit calculation;

[0085] Step S2, using a decomposition strategy to decompose the multi-section coupling limit calculation problem of the reverse search model into sub-problem models;

[0086] Step S3: Convert different sub-problem models into Markov decision processes and solve them by combining the neighborhood-based parameter migration strategy with the improved DDPG algorithm;

[0087] Step S4: Implement the Pareto frontier (stable convex boundary) of the end-to-end output multi-section coupling limit by training the sub-problem model.

[0088] The traditional method for calculating the transmission quota is to first determine the local operating boundary and then select a limited operating mode to perform a traversal calculation of the TTC, and take its upper (lower) limit to obtain the quota value. This method is greatly affected by the operating mode and has low calculation efficiency. The accuracy of the calculation result also needs to be considered. In order to maximize the calculation efficiency of the transmission quota, the present invention proposes a reverse search model for the section transmission quota. The section transmission quota can be obtained by reverse searching from the instability space. The reverse search model for the transmission quota calculation takes minimizing the external power transmission of the section instability as the objective function, and its expression is:

[0089]

[0090] Among them, P ij For the contact line I ij The transmission power along the specified direction; Z is the set of tie lines;

[0091] Furthermore, the system power flow equation of the reverse search model for the transmission limit calculation is:

[0092]

[0093] Among them, i∈N, N is the set of all nodes; P Gi , Q Gi are the active power and reactive power output by the generator respectively; Q Ci Inject power into the reactive power compensation device; Li , Q Li are the node active and reactive loads respectively; U i ,θ i is the node voltage amplitude and phase angle; Y ij , α ij is the magnitude and phase angle of the node admittance matrix;

[0094] The synchronous machine rotor motion equation of the reverse search model for the transmission limit calculation is:

[0095]

[0096] Where i∈S G , S G for all generators aggregated; and are the rotor speed at time τ and the rated speed of the rotor respectively; is the rotor angle; E′ di and E′ qi are the d-axis and q-axis components of the generator voltage respectively; M i is the inertia constant; T Mi is the output mechanical power; D i is the damping torque coefficient; I di and I qi are the d-axis and q-axis components of the generator current respectively; X di With X qi is the synchronous reactance; X′ di and X′ qi is the transient reactance; T′ d0i and T′ q0i are the open-circuit time constants of the d-axis and q-axis respectively; E fdi is the excitation output voltage. Furthermore, the constraints of the reverse search model for the transmission limit calculation include steady-state operation constraints and transient instability constraints. The steady-state operation constraints are expressed as:

[0097]

[0098] Among them, P Gi , Q Gi are the active power and reactive power output by the generator respectively; U i is the node voltage amplitude; I ij For line l ij Current after a fault occurs; For line I ij Current upper limit; S G represents the generator set; N represents the node set; Z represents the tie line set; Equation (4) represents the upper and lower limit constraints of power source active output, upper and lower limit constraints of power source reactive power, upper and lower limit constraints of node voltage, and line thermal stability constraints respectively;

[0099] The transient instability constraint is:

[0100]

[0101] Among them, x(τ) is the state variable; y(τ) is the algebraic variable; u is the control variable; x(τ0) and y(τ0) represent the initial operation mode; τ is the simulation time step; (τ0,τ end ] represents the transient process time period; transient angle instability constraint and transient voltage instability constraint;

[0102] transient angle stability evaluation index f TSI :

[0103]

[0104] where f TSI <0, the system is transient unstable, and f TSI , the larger the system is more stable;

[0105] voltage stability evaluation index η U,i :

[0106]

[0107] The system is determined to be voltage stable when the voltage of the pivotal bus is lower than 0.75pu and the duration is within 1s, and the duration is within the range of 0.8-1s is critical voltage stable; while the voltage of the pivotal bus is lower than 0.75pu and the duration is greater than 1s is voltage unstable; where i∈N; U i,min is the minimum value of the node voltage U i after the transient process after the fault is cleared; U i,cr is the voltage threshold of node i, which is 0.75pu; T cr is the duration allowed for V i,min <0.75pu, which is 1s; T b is the total duration of the node voltage being lower than U i,cr ; τ cr,j represents the time step of the jth U i,min =U i,cr ; U i,sta is the voltage of node i when it returns to normal; k τ is the critical voltage offset factor, which is 0.75; when η U,i <0, it is determined that the busbar transient voltage is unstable; hereinafter, η U represents the voltage stability margin of the system, η U is the minimum value of η U,i .

[0108] Further, the transient instability constraint adopted by the section transmission limit calculation model is an "or" form constraint, that is, the reverse search model of the transmission limit calculation satisfies the transient angle instability constraint or the voltage instability constraint, the Big-M method is adopted to process the transient instability constraint, and the model is easy to solve by rewriting formula (5) into the following formula:

[0109]

[0110] where M1 and M2 are sufficiently small numbers; δi is 0-1 decision variable, when δ I i-th constraint is satisfied, i.e. The method can be used for the design of RL reward function.

[0111] Further, the complex power system often needs to consider the influence caused by multiple sections coupling with each other. Therefore, the multi-section coupling power transmission limit needs to be calculated, and the problem belongs to a multi-objective optimization problem. When the multi-section coupling power transmission limit is calculated, multiple sections are simultaneously optimized, and the specific definition is as follows.

[0112]

[0113] Wherein, contains the state, algebra and control variables of the power system; comprises m different independent section limit calculation objective functions; respectively represent system equation constraints and inequality constraints, including stable operation constraints and transient instability constraints. The objective of the function is to find a set of optimal solutions in the trade-off between different objectives, i.e. Pareto optimal solution. Further, the present application intends to introduce a decomposition strategy to model the multi-section coupling limit calculation problem, which is decomposed into a set of scalar optimization sub-problems. The optimal solutions corresponding to different sub-problems are obtained by solving the sub-problems in a collaborative optimization manner, and these optimal solutions constitute the Pareto front of the original problem. The present application adopts the well-known weighted sum (Weighted sum, WS) method to decompose the multi-section limit calculation problem. In step S2, the multi-section coupling limit calculation problem is converted into n scalar optimization sub-problems by weighted sum, and the converted scalar objective function is represented as:

[0114]

[0115] Wherein, represents the objective function of the i-th sub-problem, is a weight vector, and m is the number of solved sections; the Pareto front of the multi-section power transmission limit problem can be described by the n optimal solutions of the above n sub-problems. The larger the value of n is, the more complete the decomposition is, and the more sufficient the description is. The process can be intuitively explained by Figure 2

[0116] ​Further, the multi-section limit calculation problem can be decomposed into multiple sub-problems with different weight vectors, and the optimal solutions of all sub-problems constitute the Pareto front of the original problem. Therefore, the solution of the sub-problem is crucial, and the objective function of the model is unique. First, the section limit calculation model is reconstructed into the form of Markov Decision Process (MDP), which is composed of five tuples <S, A, r, P, γ>, and is described as follows:

[0117] State space S:

[0118]

[0119] where P G is the active power of the generator; U G is the node voltage; P L is the active load; Q L is the reactive load; P f is the sum of transmission power flow of all tie lines in the selected section; f TSI is the transient angle stability evaluation index, and η U is the transient voltage stability evaluation index; S t is used to represent the state vector at time t;

[0120] The action space A includes the generator action space a gen and the total load action space a load , a gen is the active power adjustment amount ΔP g of the generator, and the adjustment range is N respectively represent the maximum active power of the generator and the training step number of RL per round; a load is the active power adjustment amount ΔP Load_all of the load, and the adjustment range is , which represents the maximum active power of the total load; in order to improve the stability of the training, the power output of the generator except the balancing machine is constrained in the range of , and the total active power of the load is constrained in the range of . Hereinafter, a i = [a gen , a load] represents the action vector reward function r at time t, which is a hard index to evaluate the quality of the agent's action and plays a decisive role in the solving process of MDP. The reward function is designed by combining discrete rewards and continuous rewards. First, the constraint conditions are processed. When the power flow is not convergent, the line is overloaded, the power flow is reversed, and other situations occur, a large penalty is given to improve the training speed; for transient instability constraints, the transient power angle stability index and voltage stability margin are selected as the key targets, and the Big-M processing method is used to design the continuous reward value. Then, the scalar objective function of the sub-problem is normalized and added to the total reward function. For sub-problem i, the reward function can be expressed by the following formula:

[0121] r con = -[f TSI + M1(1 - δ1)] - [η U + M2(1 - δ2)] + (δ1 + δ2 - 1) (12)

[0122]

[0123] where M1 and M2 are sufficiently small numbers; δ i is a 0-1 decision variable, when δ i = 1, the i-th constraint is satisfied; r con is the reward value of the agent exploring the operating mode to satisfy the transient instability constraint, which is a continuous reward; g is the reward function calculation condition; r i- total is the final reward value of sub-problem i; r i-obj is the reward value of the normalized scalar objective function of sub-problem i; ω s , s = 1, 2 are the weight coefficients of the reward items; K is the penalty given when the power flow is not convergent / line is overloaded / power flow is reversed / generator is out of limit;

[0124] The discount factor γ, γ ∈ [0, 1] represents the discount of future rewards.

[0125] The state transition matrix P is determined by the interactive environment.

[0126] The MORL calculation framework constructed by the MDP model and the coordinated optimization of multiple sub-problems is shown in Figure 3 .

[0127] When the number of sub-problems is large, solving the above multiple MDP model requires a large amount of time cost. Because for different sub-problems, the agent needs to explore the optimal strategy again, and does not make good use of the self-learning and knowledge transfer ability of RL. According to formula (10), the two adjacent sub-problems have similar weight vectors, so their optimal solutions are most likely to be similar. Based on this, a sub-problem can be solved by learning from its adjacent sub-problems.

[0128] The decomposition method combined with the multi-objective calculation problem and the parameter migration strategy in step S3 model the quota calculation sub-problems into neural networks by the MORL method, so as to The network model parameters representing the i-1th sub-problem are defined as [ω, b], and [ω * ,b * ] are the optimized neural network parameters. It is assumed that the i-1th quota calculation sub-problem has been solved, and the network parameters have reached the optimal or near-optimal value. When training the i-th sub-problem, the experience knowledge of the latter can be used to quickly solve it, and the optimal network parameters of the latter are used as the starting parameters of the former neural network. In this way, by migrating the model parameters of the previous adjacent sub-problem to the current sub-problem as the starting parameters, the training time cost of the model is effectively reduced. Simply put, the network model parameters are sequentially deployed to adjacent sub-problems in the direction of weight reduction (increase), and the mode is as shown in Figure 4 The strategy provides the possibility for MORL to quickly solve multi-section quota problems. The multi-section power transmission quota calculation process based on MORL combined with the decomposition method of the multi-objective calculation problem and the parameter migration strategy is shown in Table 1:

[0129] Table 1 Pseudo code of overall calculation process

[0130]

[0131]

[0132] Further, the improved deep deterministic policy gradient method is used to train the MDP model of the sub-problem. DDPG draws on the DQN experience replay library idea to break the correlation between data, improve the stability of the training, and realize the rapid convergence of the algorithm through the Actor-Critic network. The original DDPG contains four neural networks, namely the Actor network μ(s|θ μ ), the Critic network Q(s, a|θ Q ), and the copied target network μ tar (s|θ μ ) and Q tar (s, a|θ Q ), which enables the agent to fully interact with the environment, increases the random exploration ability, and considers adding random noise to the action:

[0133] a t =μ(s t |θ μ )+ξ t (16)

[0134] where μ(s t |θ μ ) is the Actor network, ξ t represents random noise;

[0135] After performing action a t , the immediate reward r t+1 and the next state S t+1 are obtained, forming a four-tuple <S t , a t , r t+1 , S t+1 > which is stored in the experience replay pool R. A Batch_size number of four-tuples are randomly sampled from R as the input of the current Actor network and Critic network, and the true value Q i is calculated:

[0136]

[0137] where γ is the discount factor.

[0138] The Critic network parameter θ Q and the Actor network parameter θ μ are updated by minimizing the loss function and calculating the policy gradient, respectively:

[0139]

[0140] where L is the loss value when the Critic network is updated, and by minimizing the loss function, the Q value of the Critic network is closer to the target Q value; N is the number of samples from the experience replay pool in each training; is the gradient of the Actor network, which is calculated based on the Q value of the Critic network to update the weights of the Actor network; is the gradient of the Critic network Q value with respect to the action, which reflects the influence of action selection on the Q value.

[0141] Finally, the obtained (θ Q , θ μ ) is used to update the target network parameters through a relatively stable soft update:

[0142]

[0143] where τ is the divergence factor. By repeatedly iterating equations (16) to (20), the network parameters are constantly updated, and the algorithm eventually converges.

[0144] The improved method of DDPG adopted by the application, the original DDPG algorithm adopts the method of experience replay for training, which fully considers the time correlation between samples and reuses the samples, however, this method ignores the importance of different samples, and randomly extracting a batch of samples may cause a large number of low reward value samples to be trained, and then cause DDPG to converge slowly. The application adopts a classification experience replay method based on immediate reward, and the samples with immediate reward less than the average reward value are put into the experience buffer pool 1, and the samples greater than the average reward value are put into the experience buffer pool 2. When a batch of samples are extracted in proportion, more samples are extracted from the buffer pool 2 each time, and a small amount of samples are extracted from the buffer pool 1, the purpose is to maintain the sample diversity while improving the convergence performance.

[0145] The above is the improved DDPG algorithm training process for single target problem, and the training process of the sub-problem model decomposed for the multi-section transmission limit calculation problem is shown in Table 2:

[0146] Table 2 Pseudo code of sub-problem training algorithm

[0147]

[0148]

[0149] Example system:

[0150] The application adopts IEEE 39-node system as a test example, the coupling sections investigated are transmission section 1 composed of AC tie lines 1-39, 2-3 and 18-3, and transmission section 2 composed of AC tie lines 15-16, and the system structure and section division are shown in Figure 5 The thermal stability constraint, transient power angle stability constraint and transient voltage stability constraint are considered, the key fault of the section is set as three-phase short-circuit fault of the tie line at 0.1s, and the fault line is cut off at 0.15s; the transient stability simulation time is 5s. The system reference power is 100MW, the generator adopts a classical model, the generator excitation control is considered, and the load adopts a constant impedance model.

[0151] The continuous action space of the generator and total load trained by the MORL algorithm is set as: a_bound=[-0.6,0.6], p L_all _bound=[-6,6]. The system adopts a three-layer BP neural network to construct the Actor network and the Critic network. In the training process, the classification buffer pool is adopted in the playback mechanism, the extraction ratio of the two types of buffer pool Batch_size is 0.7, random noise is added to the action selection, and the remaining network parameter settings are shown in Table 3:

[0152] Table 3 Parameter settings of network model

[0153]

[0154]

[0155] Effectiveness analysis of multi-section coupling limit calculation method:

[0156] The simulation starts from a randomly selected system instability mode, trains the initialized sub-problems first, and then transmits the model weights of subsequent sub-problems using the domain parameter migration strategy. The training step number of the first sub-problem is 60000 steps, and data statistics are performed in 20 steps as a round. The data in the training process is normalized. The reward value curve in the training process of the first sub-problem is as shown in Figure 6 .

[0157] As can be seen from Figure 6 , the reward value curve has a large amplitude fluctuation before about 500 rounds, and has basically reached convergence after 1000 rounds of training. The reason for the oscillation is that the decision-making agent explores the optimal strategy through continuous "trial and error", but this does not affect the convergence process. Before 500 rounds, the reward value curve oscillates violently, and at this time the system is often in a stable state or power flow divergence state, and cannot reflect the process of the intelligent agent searching for the transmission limit boundary from the instability operating mode (the power flow when the system is stable is often small). After the curve tends to be stable, according to the definition of the reward function, the operating mode of the system at this time is basically unstable, which meets the constraint condition of the limit calculation. Figure 7 The power flow changes of section 1 and section 2 after the reward value curve is flat are given, and the section power flow corresponding to each round is the cumulative average value of 20 update steps. It can be seen that within 1000-2000 training rounds, both target values have undergone huge changes, among which the power flow of section 1 has dropped from 800MW to about 600MW; the power flow of section 2 has repeatedly risen and fallen to 350MW and stabilized. The above shows that the decision result of the MORL method is reasonable, and for a sub-problem under a certain weight, the section transmission limit can be well searched, Figure 7 The finally converged section power flow value is the coupling section transmission limit value of the first sub-problem. However, in actual dispatching, the section transmission limit value under a certain weight is not concerned, and it is also unrealistic to apply a bias to the section. The present application hopes to find a convex boundary domain of mutually restrained section transmission limits to guide the optimization control of multiple sections.

[0158] For the trained first sub-problem model, the neighborhood-based parameter migration strategy transmits the model parameters to the adjacent sub-problems, trains a small number of rounds, and updates the step number to 20 steps. 50 groups of weight vectors are generated, that is, 50 scalar optimization sub-problems, and the Pareto frontier of the multi-objective migratory reinforcement learning algorithm on the two-section transmission limit problem is obtained, and the visualization result is as shown inFigure 8 It can be seen that the Pareto front on the section transmission limit problem is different from the conventional two-objective optimization Pareto front (the conventional PF is "slide" type), which is similar to the "S" type curve, which is due to the selection of the two sections by the present application, which are mutually clamped, not simply the same increase and decrease relationship, which is determined by the non-convex nonlinear model of the section limit calculation, and the Pareto front is the safety convex boundary of the two section limits. By rough division, it can be determined that the system is generally stable in the lower left corner of the power flow space, and the upper right corner of the power flow space is generally unstable.

[0159] In order to verify the effectiveness of the MORL method for searching the Pareto front, the present application checks the operating mode in the two-section power flow space by time domain simulation method to determine the stability boundary. The generator output and the total load are randomly fluctuated within their upper and lower limit ranges, and the terminal voltage is fluctuated within the value range. Latin hypercube sampling is used to generate 12000 operating conditions, the three-phase short circuit fault is set as the contingency set, and the transient stability is checked. Figure 9 The stability and instability condition diagrams in the two-section power flow space are given. In the figure, the circle represents the stable mode, and the cross represents the unstable mode. Since the stable and unstable operating modes are mixed, the unstable mode is covered on the stable mode to better distinguish the safety boundary. By comparing Figure 8 With Figure 9 It can be seen that the convex boundary of the section limit searched by MORL is basically consistent with the stable boundary obtained by sampling, which verifies the effectiveness of the method proposed in the present application. In addition, the boundary searched by MORL is more conservative than the sampling boundary, which may be due to the insufficient number of samples, which further illustrates that the search precision of the method in the present application is high.

[0160] In order to intuitively describe the relationship between the training data and the result, Figure 10 The search trajectory of the agent for the section transmission limit operating point is given. As can be seen from the figure, different training weights will obtain different transmission limit boundary points under the initial condition of point A, which is due to the selection of different weight vectors will change the reward function in MORL, and then lead to the search of the transmission limit with certain differences. In addition, when the agent searches the transmission limit similar to that obtained at the initial point A under the initial condition of point B, the weight vectors used by the two are basically consistent, which proves that the method has good adaptability to the change of the initial operating mode.

[0161] Comparison results with metaheuristic algorithms:

[0162] To better demonstrate the superiority of the MORL algorithm, this paper used the most widely used multi-objective metaheuristic algorithm as a comparison algorithm. We also selected the NSGA-II and MOEA / D algorithms, which have shown the best performance in this field, for testing. To ensure the accuracy of the test results, each case was repeated 20 times, and the best result was used as the comparison data. Figure 11 and 12 The Pareto fronts of section transmission limits obtained by different iteration numbers of NSGA-II algorithm and MOEA / D algorithm are compared with the search results of MORL.

[0163] Depend on Figure 11 It can be seen that for NSGA-II that performs a normal number of iterations (100 times), the Pareto front obtained is quite different from the actual result, and it is easy to fall into the local optimum. As the number of iterations gradually increases, the convergence level of NSGA-II gradually improves. When it increases to 8 times the initial number, the Pareto front of the section transmission limit searched by NSGA-II is close to that of the MORL algorithm, but it is slightly lower than the MORL algorithm in terms of convergence and diversity. Figure 12 In the example, when the number of iterations increases to 800, the local Pareto front found by the MOEA / D algorithm is very similar to the MORL search results, and even performs better locally. However, the MOEA / D algorithm is significantly inferior to MORL and NSGA-II in terms of diversity, and the non-dominated solutions found are all concentrated in a small area, which is unacceptable for constructing a complete convex boundary for transmission quota stability. As can be seen from the above, the MORL algorithm proposed in this paper outperforms the NSGA-II and MOEA / D algorithms in both convergence and diversity, verifying that the MORL algorithm has stronger global optimal search performance.

[0164] To further intuitively show the specific performance of each algorithm, Table 4 gives the hypervolume (HV) index and the calculation time consumption of different algorithms (including different iteration numbers of the same algorithm). It can be seen that the HV index obtained by MORL is the largest among all algorithms, that is, MORL has the best convergence and diversity. In terms of calculation time consumption, the average running time of MORL is 117.93s, which is much lower than that of NSGA-II and MOEA / D algorithms with 800 iterations for better convergence. Although the NSGA-II algorithm with 100 iterations takes less time to solve, its convergence is very poor, while MORL can search for the best Pareto front in all aspects with only about 40s more time, showing the high efficiency of MORL algorithm in solving. In general, the NSGA-II algorithm has good diversity but poor convergence; the MOEA / D algorithm has good convergence but poor diversity and needs to consume a lot of solving time. MORL algorithm improves the HV index by 8.9% and saves 81.2% of the solving time compared with the above optimal multi-objective heuristic algorithm.

[0165] Table 4 Hypervolume index and calculation time consumption of algorithm for solving transmission limit problem

[0166]

[0167] It is worth mentioning that the MORL algorithm of the present application takes 13.69h to train the transmission limit stability boundary, which is caused by the need for the agent to continuously change the generator output action for power flow calculation and transient stability analysis and the simultaneous solution of multiple sub-problems. The trained model can be directly used to output the Pareto front.

[0168] Effectiveness analysis of parameter transfer strategy:

[0169] The present application first compares the model training of the original strategy and the transfer strategy. The original strategy trains each sub-problem for 3000 rounds, with 20 update steps per round. Figure 13 The Pareto frontiers output by the two training strategies are given. It can be found that when trained with the same number of rounds, the model of the original strategy has very poor convergence and diversity, because the agent cannot use the knowledge learned in the past when training the sub-problems from scratch each time, and the key information of the neighborhood is lost.

[0170] In addition, in order to eliminate the result error caused by the number of training times, the present application further expands the training rounds of the original strategy to 12000, with 20 update steps per round, and shows the results after training in Figure 14From the figure, if the parameter migration of the neighborhood is not performed, even if the training times are increased by 3 times, the performance of the original strategy is still not good enough. Therefore, the application of the parameter migration strategy based on the neighborhood to the MORL algorithm is very practical and efficient, which greatly reduces the training time while improving the convergence.

[0171] The application follows the solution idea of a multi-objective optimization problem, and proposes a multi-section power transmission limit calculation method based on multi-objective reinforcement learning, realizes the fast output of the accurate Pareto front of multi-section coupling limits in an end-to-end manner, and the front is the stable convex boundary that needs to be found in engineering. Compared with the traditional multi-objective meta-heuristic algorithm, the method saves 81.2% of the solution time while maintaining good convergence and diversity, and the hypervolume index is improved by at least 8.9%. In addition, the introduced parameter migration strategy based on the neighborhood enables the current sub-problem to learn from the model parameters of the previous sub-problem, which greatly improves the training speed while improving the convergence of the algorithm. The multi-objective transferable reinforcement learning method used in the application provides a new automatic solution for the multi-section power transmission limit setting which is difficult to set in engineering.

[0172] The above description is only a preferred embodiment of the application, and does not limit the application in any form. Any skilled person in the art can make many possible changes, modifications and equivalent embodiments to the technical solution of the application without departing from the scope of the technical solution of the application. Therefore, any modification, equivalent change and modification of the above embodiments within the scope of the technical solution of the application are all within the protection scope of the technical solution.

Claims

1. A multi-contour power transmission limit calculation method based on multi-objective transferable reinforcement learning, characterized in that: The method comprises the following steps: Step S1, considering transient power angle and voltage stability constraints, a reverse search model for transmission limit calculation is established; Step S2, a decomposition strategy is used to decompose the multi-section coupled limit calculation problem of the reverse search model into sub-problem models; Step S3, different sub-problem models are converted into Markov decision processes, and an improved DDPG algorithm combined with a neighborhood-based parameter migration strategy is used for solving; Step S4, by training the sub-problem models, a Pareto front of the multi-section coupled limit is output end-to-end; The reverse search model for transmission limit calculation takes the minimum out-of-delivery power of section instability as the objective function, and its expression is: (1) wherein, is a tie line the transport power in the specified direction; is a set of tie lines; The system power flow equation of the reverse search model for transmission limit calculation is: (2) wherein, , is the set of all nodes; , are the active and reactive power outputs of the generators, respectively; is the injected power of the reactive power compensation device; , are the active and reactive loads at the nodes, respectively; , are the voltage magnitude and phase angle at the nodes, respectively; , are the admittance matrix magnitude and phase angle at the nodes, respectively. The synchronous machine rotor motion equation of the reverse search model for transmission limit calculation is: (3) wherein , is the set of all generators; with are the rotor speed and the rotor rated speed, respectively; is the rotor angle; and are the d- and q-axis components of the generator voltage, respectively; are the d- and q-axis components of the generator voltage, respectively; is the inertia constant; is the output mechanical power; is the damping torque coefficient; and are the d- and q-axis components of the generator current, respectively; and are the synchronous reactances; and are the transient reactances; and are the open-circuit time constants of the d- and q-axis, respectively; is the field output voltage;​​​​​​ The constraint conditions of the reverse search model for transmission limit calculation include steady-state operation constraints and transient instability constraints, and the steady-state operation constraints are expressed as: (4) wherein, , are the generator output active and reactive power, respectively; is the node voltage magnitude; is the line current after a fault; is the line current upper limit; denotes the set of generators; denotes the set of nodes; denotes the set of tie lines; The transient instability constraints are: (5) wherein, is a state variable; is an algebraic variable; is a control variable; , denotes an initial operating mode; is a simulation time step; denotes a transient process time period; , transient power angle instability constraints and transient voltage instability constraints; Transient power angle stability assessment index : (6) wherein, The system is transiently unstable when, and The larger the system is, the more stable it is; Voltage stability assessment index : (7) The system is determined to be voltage stable when the center point bus voltage is less than 0.75pu and the duration is within 1s, and the duration is within the range of 0.8-1s, which is critical voltage stability; and the center point bus voltage is less than 0.75pu and the duration is greater than 1s, which is voltage instability; wherein, ; The node voltage is cleared after the transient process of the fault The minimum value; The voltage threshold of the node is 0.75pu; The duration allowed is 1s; The total duration of the node voltage being less than ; The time step of the first ; The time step of the first ; The voltage of the node when it returns to normal; The critical voltage offset factor is 0.75; when , it is determined that the bus transient voltage is unstable; The transient instability constraints satisfy transient power angle instability constraints or voltage instability constraints, the Big-M method is used to process the transient instability constraints, and the model is easy to solve by rewriting formula (5) into the following formula: (8) in, and is a real number; is a 0-1 decision variable, when When, The constraints are satisfied; When the multi-section coupled transmission limit is calculated, multiple sections are optimized at the same time, and the specific definition is as follows: (9) wherein, comprising a state, an algebra, and control variables of the power system; including a target function composed of , representing system equality constraints and inequality constraints, respectively, including steady-state operation constraints and transient instability constraints; In step S2, the multi-section coupled limit calculation problem is converted into scalar optimization sub-problems by weighted sum, and the converted scalar objective function is expressed as: (10) wherein, represents the objective function of the th sub-problem, is a weight vector, is the number of solutions of the section; The Markov decision process is composed of a five tuple and is specified as follows: State space : (11) wherein, Pgen is the generator active power; V is the node voltage; Pload is the load active power; Qload is the load reactive power; Ptrans is the sum of transmission power flow of all tie lines in the selected section; Pangle is the transient angle stability assessment index, Pvoltage is the transient voltage stability assessment index; Action space Including generator action space And total load action space , The active power adjustment amount of the generator , the adjustment range is , , Respectively represent the maximum active power of the generator and the training step number of RL per round; The active power adjustment amount of the load , the adjustment range is , Indicates the maximum active power of the total load fluctuation reward function r is represented as: (12) (13) (14) (15) where, is the reward value for the agent to explore the operating mode that satisfies the transient instability constraint, which is a continuous reward; is the reward function calculation condition, and is a real number; is a 0-1 decision variable, when the first constraint is satisfied; is the final reward value of the sub-problem ; is the reward value of the scalar objective function normalization of the sub-problem ; , is the weight coefficient of the reward item; is the penalty given when the power flow is not converged / line overload / power flow reversal / generator out-of-limit. discount factor , denotes the discounting of future rewards; The state transition matrix is determined by the interaction environment. The decomposition method combined with the parameter migration strategy in step S3 models the quota calculation sub-problems into neural networks by MORL method, so as to represent the network model parameters of the th sub-problem, define as the network parameters that have not been optimized, as the optimized neural network parameters, assume that the th quota calculation sub-problem has been solved, and the network parameters thereof have reached or are close to the optimum, then when training the th sub-problem, the experience knowledge of the sub-problem is used for fast solution, and the optimal network parameters of the latter are taken as the starting parameters of the neural network of the former. The improved deep deterministic policy gradient method is used for training the MDP model of the sub-problem. The original DDPG contains four neural networks, namely, an actor network , a critic network , a copied target network and ; in order to enable the agent to fully interact with the environment and increase the random exploration ability, random noise is added to the action: (16) wherein represents random noise; Execute an action Get instant rewards and the next state , forming a quadruple Stored in the experience replay pool In, from Randomly extract Batch_size quadruples as the input of the current Actor network and Critic network to calculate the true value : (17) wherein is a discount factor; updating Critic network parameters by minimizing the loss function and computing policy gradient and Actor network parameters : (18) (19) Wherein, L is the loss value of the Critic network during updating, and the Q value of the Critic network is closer to the target Q value by minimizing the loss function; W is the number of samples sampled from the experience replay pool each time; is the gradient of the Actor network, which is calculated based on the Q value of the Critic network, and the weight of the Actor network is updated; is the gradient of the Q value of the Critic network with respect to the action, which reflects the influence of action selection on the Q value; Finally, the obtained The target network parameters are updated in a relatively smooth soft update manner: (20) wherein, is the divergence factor, and the network parameters are updated by repeatedly iterating equations (16) through (20), and the algorithm eventually converges.

Citation Information

Patent Citations

  • Multi-objective reinforcement learning method and device based on Pareto optimization

    CN114742231A

  • Multi-target safety optimization method of micro-grid energy system based on deep reinforcement learning

    CN114897266A