Power Grid Flow Convergence Adjustment Method Based on PPO-HELM

By combining the HELM and PPO algorithms, load nodes are screened and capacitors are switched on and off, which solves the convergence problem of traditional power flow calculation under load fluctuations and renewable energy, and achieves rapid stability adjustment of the power grid and efficient power flow convergence.

CN119891215BActive Publication Date: 2025-10-14HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411799414.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-10-14
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

Traditional power flow calculation methods have difficulty effectively solving convergence problems when faced with uncertainties such as load fluctuations and renewable energy, which may lead to overload or other potential problems in the power grid system. Existing algorithms also have poor real-time response and efficiency.

Method used

Combining the Holomorphic Embedded Power Flow Method (HELM) with the Proximal Policy Optimization (PPO) algorithm, the voltage stability index is introduced to screen load nodes and perform capacitor switching operations. The deep reinforcement learning algorithm is used to optimize power flow adjustment, and the reward mechanism of the PPO algorithm is used to accelerate power flow convergence.

Benefits of technology

It improves the adaptive capability and convergence speed of power grid flow adjustment, shortens the active power and reactive power adjustment time, and ensures the continuous and stable operation of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119891215B_ABST
    Figure CN119891215B_ABST
Patent Text Reader

Abstract

The application discloses a power grid power flow convergence adjustment method based on PPO-HELM, which firstly introduces voltage stability judgment indexes and voltage collapse point indexes by using a power flow calculation method of HELM.Secondly, after generating a power flow non-convergence state by adjusting a power grid load level, the voltage collapse point index values in the state are calculated by using HELM, and are sorted from large to small, and according to the sorting result, the first m corresponding load nodes are selected to participate in power flow adjustment.Then, according to the voltage stability judgment indexes, the voltage collapse point indexes and the selected load nodes, a PPO algorithm environment based on HELM is established.Finally, a target function is constructed, a PPO algorithm model is built, power flow convergence adjustment actions are output, and training and testing are performed.The application reduces the number of load nodes needing capacitor switching operation, shortens the active power and reactive power regulation time, and improves the adjustment efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of electric power information technology and relates to a power grid power flow convergence adjustment method based on PPO-HELM. BACKGROUND

[0002] Convergence of power flow calculation is a key factor in analyzing the safety and stability of a power system. If the power flow calculation cannot converge, it may indicate that the system is overloading or has other potential problems. Power flow calculation is not only applicable to normal operating conditions, but also to evaluating the overall performance of the system. If the power flow calculation cannot converge, it means that there may be problems in the system design, which needs to be optimized to ensure the stability and safety of the system.

[0003] The problem of non-convergence of power flow may be caused by two reasons: one is that the system itself has no solution, and the other is that the algorithm cannot effectively find a solution. Traditional power flow calculation methods have difficulty in effectively solving the convergence problem in the face of load fluctuations and uncertainties such as renewable energy. The holomorphic embedding power flow calculation method can avoid iteration, reduce computational complexity, and improve reliability. Traditional methods and machine learning algorithms such as support vector regression (SVR) and ridge regression (RR) have poor efficiency and real-time response in power flow adjustment, making it difficult to adapt to dynamic changes in the power grid. Therefore, advanced power flow calculation methods and deep reinforcement learning algorithms need to be introduced to improve the adaptive ability and convergence speed of power flow adjustment. SUMMARY

[0004] In response to the above problems, a power grid flow convergence adjustment method based on HELM-PPO is proposed. This method combines the Holomorphic Embedding Load Flow Method (HELM) and the Proximal Policy Optimization (PPO) algorithm for the convergence adjustment of power grid flows. The HELM flow calculation method is adopted and a voltage stability evaluation index is introduced. According to the voltage collapse point index, the load nodes participating in the flow adjustment are screened out, and capacitor switching operations are performed on these nodes. In addition, the average value of the voltage stability index of each node and the maximum value of the voltage collapse point index of each node are used to set the reward mechanism of the PPO algorithm. The IEEE 14-node system and the IEEE 57-node system are used as examples to verify the effectiveness of the proposed method. The present invention proposes to combine the Holomorphic Embedding Method (HELM) with the Proximal Policy Optimization (PPO) deep reinforcement learning algorithm to deal with the power grid flow convergence problem. This method uses the HELM power flow calculation method to calculate the voltage collapse point index and voltage stability index, which can be used to assess whether the system's power flow solution is stable and guide the power flow of the power grid toward stability. The PPO algorithm fully utilizes the advantages of deep learning and reinforcement learning algorithms. It can not only use neural networks to learn the complex nonlinear characteristics of the power grid, but also effectively deal with high-dimensional problems.

[0005] Specifically, the HELM method is used to calculate the power flow, and the voltage stability index L is introduced based on the power flow solution. ij , and voltage stability index VSI based on sensitivity analysis HELM1i and Voltage Collapse Indicator VSI HELM2i . Calculate the VSI of each node in the power grid HELM1i , and calculate the average VSI of these nodes ave . Calculate the VSI of each node in the power grid HELM2i value, and find the maximum VSI among them max Secondly, calculate the voltage stability index VSI of each node HELM2i , and based on this index, select the load nodes participating in the power flow adjustment, and select the generators whose initial active power is less than a certain value to join the power flow adjustment process, thereby reducing the number of nodes that need to participate in the adjustment, shortening the adjustment time of active power and reactive power, and improving the stability of the power grid. Finally, construct the PPO model and combine it with VSI ave and VSI maxThe trained model can monitor power flow data in real time, make intelligent adjustments based on real-time changes, and quickly respond to the impact of power flow fluctuations to ensure the continued stable operation of the power grid.

[0006] The method of the present invention specifically is:

[0007] (1) Using HELM's power flow calculation method to introduce the voltage stability index VSI HELM1i and Voltage Collapse Indicator VSI HELM2i .

[0008] The voltage stability index VSI is introduced by calculating the sensitivity of each node voltage based on HELM. HELM1i and Voltage Collapse Indicator VSI HELM2i .

[0009] According to HELM, the sensitivity of each node voltage can be calculated. First, an embedded pure virtual function is constructed as shown in formula (1).

[0010]

[0011] Where V i [n] represents the nth voltage component of node i in HELM power flow calculation, s n Represents the nth-order term of the frequency domain operator s.

[0012] Nonlinear sensitivity of voltage to injected power and As shown in formula (2):

[0013]

[0014] Through theoretical derivation and simulation analysis, a stability judgment index VSI is proposed HELM1i :

[0015]

[0016] Calculate the VSI of each node HELM1i value, and calculate the VSI of these nodes HELM1i Average VSI value ave :

[0017]

[0018] Where: VSI HELM1i is the voltage stability index, node i is the bus number of the distribution network (i∈2~T), node j is the bus number of the reactive power injection, and real refers to the real part of the corresponding complex number. j represents the reactive power of node j, P jVSI i represents the n-th coefficient term of voltage at node i. T represents the total number of nodes. represents the partial derivative of the n-th coefficient term of voltage at node i with respect to the reactive power of node j. represents the partial derivative of the n-th coefficient term of voltage at node i with respect to the active power of node j. VSI HELM1i The smaller the value of VSI , the better the voltage stability of the system.

[0019] In order to determine the voltage stability of the system as a whole, an index VSI HELM2i is added:

[0020]

[0021] When VSI HELM2i <0, the power grid system is stable, otherwise, the voltage of the power distribution network system has reached the voltage collapse point.

[0022] The VSI HELM2i values of each node are calculated, and the maximum value VSI max is found:

[0023]

[0024] The VSI HELM1i and VSI HELM2i indices in step (1) output node data, therefore, an L ij index is added to view the stability state of each transmission line. The calculation formula of L ij is as follows:

[0025]

[0026] In the formula: X ij represents the reactance between node i and node j, R ij represents the resistance between node i and node j, U i represents the voltage amplitude of node i; when L ij <1, the power grid system is in a voltage stable state, the smaller the L ij index, the higher the voltage stability; when L ij index approaches 1, it means that the voltage stability of the system is declining, and there is a risk of voltage collapse; when L ij ≥1, the voltage of the power grid system has lost stability and entered a voltage collapse state.

[0027] The L ij values of each line in the power grid are calculated, and the maximum value is selected as L max :

[0028]

[0029] (2)Screening nodes participating in the flow adjustment;

[0030] After generating the flow divergence state by adjusting the power grid load level, the VSI HELM2i value in this state is calculated by HELM, and the VSI HELM2i values are sorted from large to small, and the load nodes i corresponding to the first m VSI HELM2i values are selected to participate in the flow adjustment. The capacitors on these load nodes are switched to adjust the reactive power, which can effectively adjust the node voltage level and optimize the voltage distribution of the power grid. At the same time, the generators with small initial output and large adjustment margin are used for active power adjustment, which makes them have flexible positive and negative power adjustment capability, and realizes more effective flow distribution.

[0031] (3)According to the voltage stability index VSI HELM1i , the voltage collapse point index VSI HELM2i and the selected load nodes, the PPO algorithm environment based on HELM is established.

[0032] The deep reinforcement learning model of Proximal Policy Optimization (PPO) includes an environment and an agent. The environment provides the state s t , the agent takes action a t and feeds back the result to the environment, and the environment generates a new state s t+1 and a reward R, and stores the experience tuple (s t , a t , s t+1 , R) in the experience pool , then samples the experience pool to train the agent. The agent learns the optimal strategy through interaction with the environment. This interaction process includes two key networks: the Actor network and the Critic network. The goal of the Actor network is to maximize the cumulative return of the policy π, and the goal of the Critic network is to optimize the value network by minimizing the loss function. The Actor network includes a current policy network π θ (a t |s t ) and an old policy network π θold (a t |s t ), while the Critic network includes a value network PPO algorithm combines the advantages of deep learning and reinforcement learning, and uses deep neural networks to approximate the value function and policy function, so that the agent can more accurately estimate these functions, thereby improving the learning efficiency and performance.

[0033] First, we need to build the environment. Gi , reactive power Q of load S Li , VSI of each node HELM1i The mean of the values ​​and the VSI in each node HELM2i The maximum value of state s t :

[0034]

[0035] Where, Represents the generator G i The active power, Indicates load L i Reactive power, VSI ave Indicates the VSI of each node HELM1i Mean value, VSI max Indicates the VSI in each node HELM2i Set the upper limit of the generator active power Gs max and lower limit Gs min And the upper limit of load reactive power Ls max and lower limit Ls min :

[0036]

[0037] Where, Represents the generator G i The upper limit of active power, Represents the generator G i The lower limit of active power, Indicates load L i The upper limit of reactive power, Indicates load L i Lower limit of reactive power.

[0038] According to the nodes selected for power flow adjustment in step (2), power flow adjustment is achieved by adjusting the generator output or capacitor switching operation of these nodes. Action a t as follows:

[0039]

[0040] Where, Represents the generator G i The change in active power output, Indicates load L i The change in reactive power.

[0041] a tis the continuous space, and the upper and lower limits of the action value space are set as follows:

[0042]

[0043] wherein, represents the maximum value of the single action generator output, represents the minimum value of the single action generator output; represents the maximum value of the single action capacitor switching, represents the minimum value of the single action capacitor switching, ΔP minG is the lower limit of the generator active power, ΔP maxG is the upper limit of the generator active power; ΔQ minL is the minimum value of the load reactive power, ΔQ maxL is the maximum value of the load reactive power.

[0044] The next state after the action is made is:

[0045] s t+1 ~ f(s t+1 |s t , a t ) (15)

[0046] wherein: f is the state transition function.

[0047] Reward mechanism setting: when the state is within the effective range, that is, does not exceed the preset upper and lower limits, the reward function is set according to VSI max and VSI ave .

[0048] The reward function is set as follows:

[0049] 1. When VSI max is positive, the reward function is set according to the initial value VSI max of VSI ini . Specifically as follows:

[0050] a. If 0 < VSI max ≤ (1.5 * VSI ini ), when VSI max is in this range, a continuous reward R = (-4 * VSI max ) / (3 * VSI ini ) is set.

[0051] b. When VSI max > (1.5 * VSI ini ), when VSI max is in this range, a penalty R = -5 is given.

[0052] 2. VSI max When it is a non-positive number:

[0053] a. In VSI ave ≥ω1*VSI o In the case of o ) / VSI ave , where VSI o For VSI ave The initial value of , ω1 is the coefficient.

[0054] b. In VSI ave <ω1*VSI o In this case, reward R=5 is given.

[0055] (4) Construct the objective function, build the PPO algorithm model, and output the flow convergence adjustment action.

[0056] After the environment is created, the intelligent agent is built. The initial objective function of the PPO algorithm is As shown below:

[0057]

[0058] Where ξ is the ratio of new and old strategies; is the advantage function estimate; θ is the parameter of the Actor network; E is the expectation.

[0059]

[0060] Where, π θ is the current policy, π θold is the old strategy, θ is the parameter of the actor network, π θ (a t |s t ) is the current strategy in state s t Next take action a t The probability of π θold (a t |s t ) is the old policy in state s t Next take action a t The probability of; t is the time step; s t 、a t They are the states s observed at the tth time step. t and the action a t .

[0061]

[0062] In the formula The target value of the state value network; V(s t) is in state s t The state value function under r t is the reward obtained at the tth time step; γ is the discount factor; V(s t+1 ) is in state s t+1 The state value function under .

[0063] In order to limit the amplitude of policy updates, PPO introduces a clipping objective function as follows:

[0064]

[0065] Where clip(ξ,1-ε,1+ε) is the clipping function, which limits the size of the ratio of the new and old strategies ξ to the range of [1-ε,1+ε] to prevent drastic changes in the strategy.

[0066] Furthermore, in order to improve the performance of the strategy while maintaining the stability of the training algorithm, PPO combines the optimization of the value function and the entropy term. The final objective function of PPO consists of three parts:

[0067]

[0068] Where c1 is the weight coefficient of the value function objective term, and c2 is the weight coefficient of the entropy objective term.

[0069] Value function objective term and entropy target term The calculation formula is as follows:

[0070]

[0071] H(π(·|s t ))=-π θ (a|s)logπ θ (a|s) (24)

[0072] In the formula Represents the parameters of the Critic network, For state s t The state value function under represents the target value of the state value network, H(π(·|s t )) is the entropy of the strategy, π θ (a t |s t ) is the current strategy in state s t Next take action a t The probability of , log represents the logarithmic operation.

[0073] Loss function of the PPO algorithm The calculation formula is as follows:

[0074]

[0075] Use gradient descent algorithm to minimize the loss function Update the policy network parameters θ and value network parameters The calculation process is shown in the formula:

[0076]

[0077] Where α is the learning rate and β is the entropy regularization coefficient.

[0078] (5) The PPO algorithm model is trained and tested to obtain the results of the tidal current convergence adjustment.

[0079] When HELM is used for power flow calculation, a non-convergent power flow state can be generated by adjusting the load level in the power grid, which serves as a training sample for the PPO algorithm model.

[0080] Beneficial effects of the present invention: HELM is used to calculate the voltage stability index VSI of each node HELM2i , and according to VSI HELM2i The load nodes that participate in the power flow adjustment are selected by the value, thereby reducing the number of load nodes that need capacitor switching operations. At the same time, generators with initial active power less than a certain value are selected to participate in the adjustment. These two screening processes help shorten the active power and reactive power adjustment time and improve the adjustment efficiency. In addition, by using VSI max 、VSI ave Using this indicator to set the reward function of the PPO algorithm allows for more comprehensive monitoring of the grid state, timely identification of non-convergent power flows, and accelerated convergence of reward accumulation, further improving the algorithm's regulation accuracy and responsiveness. A trained model can adjust the grid power flow from a non-convergent state to a convergent state, further regulating the grid to a more stable level. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figure 1 This is the PPO schematic diagram;

[0082] Figure 2 This is the PPO network structure diagram;

[0083] Figure 3 This is the training flow chart;

[0084] Figure 4 The cumulative reward of the PPO algorithm for the IEEE14-node system;

[0085] Figure 5(a) shows the IEEE 14-node system VSI HELM2i Comparison chart before and after value adjustment;

[0086] Figure 5(b) is a VSI of the IEEE 14-node system HELM1i before and after value adjustment.

[0087] Figure 5(c) is a L of the IEEE 14-node system ij before and after value adjustment. DETAILED DESCRIPTION

[0088] First, based on the PPO environment of HELM, the node power grid model is established, and the power flow data of the power grid is initialized. The network parameters are initialized ψ, φ and θ. In the process of the PPO algorithm, at the beginning of training, first use the HELM method to perform power flow calculation on the power grid model in the current state s t , the power flow data of the power grid in the current state can be obtained. Calculate the stability index, and the action a t is determined to participate in the power flow adjustment of the generator active power and the load reactive power, and a t is obtained s t+1 and reward R, and then (s t , a t , s t+1 , R) is stored in the experience pool for subsequent training. With the continuous update of the data in the experience pool, the network parameters will also be continuously adjusted and updated. The training process will continue until the stop condition is met.

[0089] a, reference Figure 1 , the present application is as follows:

[0090] b, step (1): use the power flow calculation method based on holomorphic embedding to construct the IEEE 14 power grid model. By using HELM and pandapower to perform power flow calculation on the 14-node power grid model, the power flow data of each node and each branch obtained by the two methods is compared. The results show that the power flow calculation results of the two methods are consistent.

[0091] c, step (2): calculate the VSI HELM2i index, and select the load nodes participating in the power flow adjustment.

[0092] For the current power flow non-convergent state, calculate the VSI HEIM2i index of each node, and output the VSI HELM2i value of all nodes, and the calculation results are as shown in Table 1:

[0093] Table 1 VSI HELM2i value of each branch of the IEEE 14 power grid model

[0094]

[0095]

[0096] According to the calculation results, the adjustment is made by adjusting the reactive power injection of the loads corresponding to nodes 4 and 5. At the same time, all generators in the system are selected to participate in the power flow adjustment.

[0097] d. Step (3): Establish a PPO algorithm environment based on HELM.

[0098] First, based on the calculation results, the reactive power of the loads on nodes 4 and 5 and the active power output of all generators are determined as adjustable variables. Then, the active power of the generators is Reactive power of load VSI of each node HELM1i The mean of the values ​​and the VSI in each node HELM2i The maximum value is the current state s t , get the state s at the next moment t+1 Utilize VSI-based max and VSI ave The reward R is determined by the calculation formula, and the size of the reward is used to evaluate the performance of the model.

[0099] e. Step (4): Build the PPO algorithm model.

[0100] refer to Figure 2 , the input of the Actor network is the state, and the output is the action probability distribution parameter. The input of the Critic network is the state, and the output is the target value of the state value network and state value function Network update process reference for PPO deep reinforcement learning Figure 3 In the Actor network, the Actor is given the current state s by the environment. t Execution Strategy π θ , calculate the new and old strategy ratio ξ and the clipping objective function And according to the strategy π θ Calculate entropy objective

[0101] In the Critic network, the target value based on the state value network and the current state value function Calculating the advantage function Sum value function objective term By the clipping objective function Value function objective term and entropy target term The loss function that makes up PPO By adjusting the parameters θ and Perform gradient descent training updates to achieve the policy network πθ and value network training and optimization.

[0102] f. Step (5): Train and test the model to obtain the final adjustment results.

[0103] The power flow convergence adjustment process is as follows:

[0104]

[0105]

[0106] This example uses the Pytorch framework to build a neural network and uses the Adam optimizer for training. Adam is a commonly used adaptive gradient descent algorithm that can usually converge to good results in a relatively short period of time when training neural networks. When using the PPO algorithm based on the Pytorch framework for training, the PPO algorithm parameter settings are shown in Table 2:

[0107] Table 2 PPO algorithm parameter settings

[0108]

[0109] Two experimental schemes are set up in the IEEE14 power grid model. In scheme 1, VSI is used max and VSI ave Indicators are used to set the reward mechanism of the algorithm. Then, in solution 2, L max The reward mechanism of the indicator reset algorithm is set as follows:

[0110] 1.L max When <1:

[0111] A reward of R=5 is given.

[0112] 2.L max ≥1, according to L max The initial value L ini The size sets the reward function. The details are as follows:

[0113] a.1 <L max ≤(1.3*L ini ), when L max When in this range, reward R=-2L is given max / (1.3*L ini ).

[0114] b. If L max >(1.3*L ini ), when L max When in this range, a reward R=-5 is given.

[0115] Scheme 3: Based on the IEEE 14-node system of pandapower, the convergence flag of power flow calculation can be judged by the return value of the runpp function. When the power flow converges, R = 5, and when the power flow does not converge, R = 0.

[0116] In the IEEE 14-node system, the total number of training steps k total is 35000, p total is 100, and the number of training rounds is 350. The cumulative reward of the two experimental schemes is referred to Figure 4 , and the curve of scheme 1-PPO-VSI represents the cumulative reward curve obtained by using the PPO algorithm and introducing VSI max and VSI ave index in scheme 1. max The curve of scheme 2-PPO-L represents the cumulative reward curve obtained by using the PPO algorithm and introducing L max index in scheme 2. The curve of scheme 3-PPO represents the cumulative reward curve obtained by using the PPO algorithm in scheme 3. According to the analysis of the cumulative reward curve, Figure 4 in scheme 1, the reward gradually increases in the first 160 training rounds, and after 160 rounds, it tends to be stable and the cumulative reward remains at a relatively high level. In contrast, in scheme 2, it is observed from the curve that the reward curve fluctuates greatly in the first 60 rounds, and after 60 rounds, it tends to be stable, but the reward curve still remains at a relatively low level, which indicates that there is a certain limitation in the reward setting in scheme 1. In scheme 3, the cumulative reward always remains at the 0 scale line, which indicates that the power flow has never converged during the power flow adjustment process.

[0117] The above trained model is tested in the initial environment. The change curves of each index before and after power flow adjustment are respectively referred to Figures 5(a)-5(c) . From Figures 5(a)-5(c) and Table 3, it can be seen that in scheme 1, before adjustment, VSI HELM2i > 0, and after adjustment, VSI max < 0. This indicates that the power flow of the power grid has successfully converged after adjustment. And from Figures 5(a)-5(c) , it can be seen that after adjustment, the value of VSI HELM1i of each node has decreased. According to the test results of Figures 5(a)-5(c) and Table 3, in scheme 2, before adjustment, L max > 1, and after power flow adjustment, L max has decreased but is still greater than 1. This indicates that the power flow still does not converge after adjustment in scheme 2, and the adjustment fails. The above results show that the power flow adjustment result after introducing VSI max and VSI ave is always better than introducing Lmax the results after.

[0118] Table 3 Values of each index before and after adjustment of IEEE 14-node system

[0119]

Claims

1. A power grid power flow convergence adjustment method based on PPO-HELM, characterized in that: The following steps are involved: Step 1: Use HELM's power flow calculation method to introduce the voltage stability index VSI HELM1i and Voltage Collapse Indicator VSI HELM2i The specific implementation process is as follows: According to HELM, the sensitivity of each order of grid node voltage is calculated and the embedded pure virtual function V is constructed. i (s): Where V i [n] represents the nth voltage component of node i in HELM power flow calculation, s n represents the nth order term of the frequency domain operator s; Nonlinear sensitivity of voltage to injected power and As shown in formula (2): Proposed stability judgment index VSI HELM1i : Calculate the VSI of each node HELM1i value, and calculate the VSI of these nodes HELM1i Average VSI value ave ; Where node i is the bus number of the distribution network, node j is the bus number of the reactive power injection, and real refers to the real part of the corresponding complex number; Q j represents the reactive power of node j, P j represents the active power of node j, V i [n] represents the nth coefficient term of the voltage at node i; T represents the total number of nodes; The nth coefficient term representing the voltage at node i is used to find the partial derivative of the reactive power at node j; The nth coefficient term representing the voltage at node i is calculated as the partial derivative of the active power at node j; Add VSI indicator HELM2i : Step 2: After adjusting the grid load level to generate a non-convergent state, use HELM to calculate the VSI under this state. HELM2i value, and set VSI HELM2i Sort the values ​​from largest to smallest and select the first m VSIs based on the sorting results. HELM2i The load node i corresponding to the value participates in the power flow adjustment; Step 3: Based on the voltage stability index VSI HELM1i , Voltage Collapse Index VSI HELM2i As well as the selected load nodes, establish the HELM-based PPO algorithm environment; Step 4: Construct the objective function, build the PPO algorithm model, and output the power flow convergence adjustment action; Step 5: Train and test the PPO algorithm model to obtain the results of power flow convergence adjustment.

2. The power grid power flow convergence adjustment method based on PPO-HELM according to claim 1, characterized in that: The step 1 also includes: when the VSI HELM2i <0, the power grid system is stable, otherwise the distribution network system voltage has reached the voltage collapse point; Calculate the VSI of each node HELM2i Values, and find the maximum value VSI max ; VSI HELM1i and VSI HELM2i The data output by the indicator is node data, increase L ij Indicators to check the stability of each transmission line, L ij The calculation formula is as follows: Where: X ij represents the reactance between node i and node j, R ij represents the resistance between node i and node j, U i The voltage amplitude of node i; when L ij <1, the power grid system is in a voltage stable state; when L ij When ≥1, the power grid system enters the voltage collapse state; Calculate the L of each line in the power grid ij value, and select the maximum value as L max .

3. The power grid power flow convergence adjustment method based on PPO-HELM according to claim 2, characterized in that: The HELM-based PPO algorithm environment is as follows: The deep reinforcement learning model of proximal policy optimization PPO consists of an environment and an agent. The environment gives the state s t , the agent takes action a t And feed the results back to the environment, and the environment generates a new state s t+1 and reward R, and the experience tuple (s t , a t ,s t+1 ,R) deposited into the experience pool Then, the experience pool Perform sampling training on the agent. The agent learns the optimal strategy through interaction with the environment. The interaction process includes: Actor network and Critic network. The Actor network contains a current strategy network π θ (a t |s t ) and the old policy network π θold (a t |s t ), while the Critic network includes the value network The active power P of the generator Gi , reactive power Q of load S Li , VSI of each node HELM1i The mean of the values ​​and the VSI in each node HELM2i The maximum value of state s t : Where, For the generator G i The active power, is the load L i Reactive power; Set the upper limit of the generator active power Gs max and lower limit Gs min And the upper limit of load reactive power Ls max and lower limit Ls min : Where, Represents the generator G i The upper limit of active power, Represents the generator G i The lower limit of active power, Indicates load L i The upper limit of reactive power, Indicates load L i Lower limit of reactive power; According to the nodes selected for power flow adjustment in step 2, power flow adjustment is achieved by adjusting the generator output or capacitor switching operation of these nodes. Action a t as follows: Where, For the generator G i The change in active power output, is the load L i Change in reactive power; a t It is a continuous space, and upper and lower limits must be set for the action value space: Where, Indicates the maximum output of the generator in a single action. Indicates the minimum output of the generator for a single action; Indicates the maximum value of a single-action capacitor switching. Indicates the minimum value of single-action capacitor switching, ΔP minG is the lower limit of the generator’s active output, ΔP maxG The upper limit of the active output of the generator; ΔQ minL is the minimum value of the load reactive power, ΔQ maxL is the maximum value of the load reactive power; The next state after taking the action is: s t+1 ~f(s t+1 |s t ,a t ) (12) Where: f is the state transfer function; Reward mechanism setting: When the status is within the valid range, according to VSI max and VSI ave Set the reward function.

4. The power grid power flow convergence adjustment method based on PPO-HELM according to claim 3, characterized in that: The reward function is specifically set as follows: VSI max is a positive number, according to VSI max Initial value VSI ini Set up the reward function; specifically: a. If 0 <VSI max ≤(1.5*VSI ini ), when VSI max In this range, set the continuous reward R = (-4*VSI max ) / (3*VSI ini ); VSI max >(1.5*VSI ini ), when VSI max If it is within this range, a penalty R=-5 is given; VSI max is a non-positive number: In VSI ave ≥ω1*VSI o In the case of o ) / VSI ave ; Among them, VSI o For VSI ave The initial value of , ω1 is the coefficient; In VSI ave <ω1*VSI o In this case, reward R=5 is given.

5. The power grid power flow convergence adjustment method based on PPO-HELM according to claim 4, characterized in that: The specific process of step 4 is as follows: After the environment is created, the intelligent agent is built; the initial objective function of the PPO algorithm is As shown below: ξ is the ratio of new and old strategies; is the advantage function estimate; θ is the parameter of the Actor network; E is the expectation; PPO introduces a clipping objective function as follows: Where clip(ξ,1-ε,1+ε) is the clipping function, which limits the size of ξ to the range of [1-ε,1+ε]; The final PPO objective function consists of three parts: Where c1 is the weight coefficient of the value function objective term, and c2 is the weight coefficient of the entropy objective term; Value function objective term and entropy target term The calculation formula is as follows: H(π(·|s t ))=-π θ (a|s)logπ θ (a|s) (18) In the formula Represents the parameters of the Critic network, For state s t The state value function under represents the target value of the state value network, H(π(·|s t )) is the entropy of the strategy, π θ (a t |s t ) is the current strategy in state s t Next take action a t The probability of logarithmic operation; the loss function of the PPO algorithm is 6. The power grid power flow convergence adjustment method based on PPO-HELM according to claim 5, characterized in that: In the step 5, it is also included that, when using HELM to perform power flow calculation, a non-convergent power flow state is generated by adjusting the load level in the power grid, which serves as a training sample for the PPO algorithm model.

Citation Information

Patent Citations

  • HELM load flow calculation method considering load static characteristics

    CN114036751A

  • Power transmission section power adjustment method based on SAC deep reinforcement learning

    CN118137457A