Reinforcement learning-based multi-objective dispatch optimization control method and system for power system

By constructing a power system model and training an intelligent agent using a multi-objective scheduling optimization method based on deep reinforcement learning, the coordination problem between distributed power sources and grid control was solved, achieving efficient and stable operation and resource optimization of the power system.

CN119695928BActive Publication Date: 2025-12-12STATE GRID JIANGSU ECONOMIC RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411483206.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-12-12
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively coordinate the control of distributed power sources and the power grid, leading to problems such as voltage exceeding limits, increased line losses, and three-phase imbalance. Furthermore, existing algorithms are ill-suited to complex new power systems and are prone to getting trapped in local optima.

Method used

A multi-objective scheduling optimization method based on deep reinforcement learning is adopted. By constructing photovoltaic models, wind turbine models and composite load models, the sensitivity of voltage nodes is analyzed using the Monte Carlo method, and a near-end policy optimization agent is trained to achieve real-time optimization control of the power system.

Benefits of technology

It enables efficient scheduling and optimization of complex power systems, alleviates voltage fluctuations and over-limit issues, and improves grid stability and resource utilization for new energy integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119695928B_ABST
    Figure CN119695928B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-objective scheduling optimization methods based on reinforcement learning, comprising the following steps: obtaining the historical voltage data of each node of power system;Randomly generate power and load disturbance, change each node voltage value by solving through power flow calculation, record each branch current of system, total network loss and total reactive power compensation capacity;The sensitivity vector of each voltage node to branch current, system total network loss, total reactive power compensation is calculated using Monte Carlo method, and the total weighted sensitivity vector is obtained by weighting addition normalization as the "weight" of each node;Define the product of node "weight" and its time sequence operating voltage deviation value as node weighted voltage deviation value, and the global accumulation is system weighted voltage deviation value;Introduce system weighted voltage deviation value to train near-end strategy optimization reinforcement learning intelligent agent;From this, power system voltage real-time optimization control strategy.The application can effectively realize multi-objective control of power system, and is conducive to safe, reliable and economic operation of power system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-objective scheduling optimization of power systems, and particularly relates to a multi-objective scheduling optimization method and system considering time sequence operation voltage weighted offset value of power grids based on reinforcement learning. BACKGROUND

[0002] In the process of new power system construction, the penetration rate of distributed renewable energy is constantly increasing. However, with the growth of distributed power capacity, its many adverse effects on power quality and security and stability of the power system gradually emerge.

[0003] The output of distributed power is constantly changing due to the intermittency and volatility of renewable energy, and the power output curve is difficult to synchronize with the load power. Therefore, after a large number of distributed power is connected to the new power system, the problems of voltage out-of-limit and line loss faced by the power grid operation are becoming more and more serious, and the three-phase imbalance characteristics are more prominent. Problems such as increase of harmonic content in distribution network and protection misoperation also frequently occur. Unbalanced operation of the system will cause a sharp increase in voltage negative sequence components, causing abnormal operation of equipment and increased loss, and in severe cases, it will threaten the safe and stable operation of the power system, and also cause the problem of islanding effect. In actual operation, distributed power will also have adverse effects on voltage, frequency, protection, harmonics, and rotor angle stability of the power system.

[0004] With the increasing of new energy installed capacity, more and more adjustable power equipment appears in the new distribution network. The control performance of controllable inverter, SVG and other equipment is getting better and better. However, if an effective coordinated control method cannot be established between the compensation device and the distributed power, the utilization rate of resources and the operation efficiency of the power system cannot be effectively improved, and problems such as overvoltage and network loss increase cannot be alleviated. In the field of coordinated optimization control between power system components and power electronic equipment, scholars at home and abroad have carried out research. At present, the common control means of power electronic equipment includes: virtual synchronous generator technology (VSG), droop control, PQ control, etc. These technologies are mature and widely used, simple and reliable, and have good local control effect, but lack the ability of overall coordinated optimization. Therefore, some algorithms need to be combined to realize overall coordinated control optimization, such as: linear and nonlinear programming method, dynamic programming method, multi-objective particle swarm optimization algorithm, genetic algorithm, etc. However, these algorithms are difficult to represent the complex model of new active distribution network and are easy to fall into local optimum. The rapid development of artificial intelligence, big data analysis, Internet of Things and other technologies provides an effective method for active distribution network optimization control. Deep reinforcement learning can handle high-dimensional, large-scale and uncertain control problems and has strong decision-making ability. At present, the application of power system based on deep reinforcement learning has attracted widespread attention. However, the actual power system operation has strong complexity and uncertainty, and is closely related to climate conditions, economic environment and human activities, which is difficult to represent with deterministic mathematical laws. In order to play a good role in actual control, further research is needed. Many scholars apply genetic algorithm, fuzzy chance constrained programming, multi-objective particle swarm optimization algorithm and other methods to active distribution network optimization scheduling. These algorithms have high requirements for problem modeling and are not suitable for complex new power systems and are easy to fall into local optimum. SUMMARY

[0005] The purpose of the present application is to solve the defects in the prior art, and to provide a multi-objective scheduling optimization method and system considering the time sequence operation voltage weighted offset value of the power grid based on reinforcement learning, which realizes the cooperative optimization of the composite control target based on day-ahead optimization real-time control and deep reinforcement learning, and is beneficial to building a more safe, reliable and economic new system.

[0006] In order to achieve the above purpose, the present application is realized by the following technical scheme:

[0007] The multi-objective scheduling optimization control method of the power system based on reinforcement learning comprises:

[0008] Obtain the topological structure of the power system and the historical data of the voltage of each node under multiple scenarios, and construct a photovoltaic model, a wind turbine model and a composite load model;

[0009] Randomly generate power system power supply node voltage and load disturbance, solve through power flow calculation, record the changed voltage amplitude of each node, record the current value of each branch current, system real-time total network loss, system real-time total reactive compensation under this state;

[0010] The sensitivity matrix of each voltage node to each current branch, the sensitivity vector of the system real-time total network loss, and the sensitivity vector of the system real-time total reactive compensation are obtained by using the Monte Carlo method, and the total weighted sensitivity vector is obtained by weighting and adding normalization, which is defined as the "weight" of each node;

[0011] The node "weight" is multiplied by the time sequence running voltage offset value to define the node weighted voltage offset value, and the global accumulation is the system weighted voltage offset value under the current state of the system;

[0012] The system weighted voltage offset value is introduced to train the proximal policy optimization reinforcement learning intelligent agent; and the trained intelligent optimization control model is used for real-time optimization control of the power system voltage.

[0013] The reinforcement learning algorithm is used for power system control, and the proximal policy optimization (PPO) is used to train the intelligent agent, so that multi-objective scheduling optimization is realized, and the intelligent optimization control model obtained based on the global weighted voltage offset value of the system can control different voltage nodes with different accuracy according to the "weight" corresponding to the voltage node.

[0014] Further, the sensitivity matrix of each voltage node to each current branch is as follows:

[0015]

[0016] In the formula, α mn is the sensitivity of the nth node to the mth line in the line;

[0017] The sensitivity vector of the voltage node to the system real-time total network loss is as follows:

[0018] [β s1 ,β s2 ,...,β sn ]

[0019] In the formula, β sn is the sensitivity of the nth node to the system real-time total network loss;

[0020] The sensitivity vector of the voltage node to the system real-time total reactive compensation is as follows:

[0021] [γ s1 ,γ s2 ,...,γ sn ]

[0022] In the formula, γsn The real-time total reactive power compensation sensitivity of the system for the nth node pair.

[0023] The weighted voltage deviation value of the above system is obtained by the following steps:

[0024] The weighted voltage deviation value of the system with respect to the current δU α The weighted voltage deviation value of the system with respect to the loss δU β The weighted voltage deviation value of the system with respect to the reactive power compensation δUγ:

[0025] δU α = α1δu1+ α2δu2+ α3δu3+... + α n δu n

[0026] δU β = β1δu1+ β2δu2+ β3δu3+... + β n δu n

[0027] δU γ = γ1δu1+ γ2δu2+ γ3δu3+... + γ n δu n

[0028] In the formula, δu n is the voltage deviation value of the nth node,

[0029] α n is the "weight" of the branch current of the nth node after processing in the weighted voltage deviation value,

[0030] β n is the "weight" of the circuit loss of the nth node after processing in the weighted voltage deviation value,

[0031] γ n is the "weight" of the system reactive power compensation of the nth node after processing in the weighted voltage deviation value;

[0032] The comprehensive control target priority is set, and the weighted voltage coefficients of each part are set:

[0033] A + B + C = 1

[0034] In the formula, A, B, and C respectively correspond to the weighted voltage coefficients of the branch current, the system loss, and the reactive power compensation three parts;

[0035] The total weighted voltage deviation value δU of the system is calculated:

[0036]

[0037] The state space of the reinforcement learning agent includes the state space of each phase voltage of each node, the switching state of controllable equipment, and the gear position, as detailed below:

[0038] S = [Sta] u Sta C Sta O Sta G Sta S ]

[0039] In the formula, S is the complete set of the state space, which can be decomposed into 5 subsets containing node voltage information and the states of various controllable devices.

[0040]

[0041]

[0042] In the formula, Sta u This represents the state-space voltage subset formed by the voltages of each phase. Indicates the first node on node k The node voltage of the phase, where It consists of three phases, a, b, and c, with each phase voltage differing by 120 degrees.

[0043] Sta C =[Sta c1 Sta c2 ,...,Sta ci ]

[0044] Sta O =[Sta o1 Sta o2 ,…,Sta oi ]

[0045] Sta G =[Sta g1 Sta g2 ,...,Sta gi ]

[0046] Sta S =[Sta s1 Sta s2 ,...,Sta si ]

[0047] In the formula, Sta C Sta represents the state-space capacitance subset consisting of all capacitor switching states. ci Indicates the switching status of each capacitor; Sta O Sta represents a state-space transformer subset consisting of the tap positions of all on-load tap-changing transformers. oirepresents the tap position state of each on-load tap-changing transformer; Sta G represents the state space generator subset composed of all generator output states, Sta G represents the real-time active power output of each generator unit; Sta S represents the state space breaker subset composed of all breaker opening and closing states, Sta si represents the opening and closing state of each breaker.

[0048] wherein the action space Act of the reinforcement learning agent includes Act c , Act o , Act G , Act s four sub-action sets:

[0049] Act = [Act c , Act o , Act G , Act s ]

[0050] Act C = [Act c1 , Act c2 ,…, Act ci ]

[0051] C i = [c i1 , c i2 ]

[0052] wherein Act C represents the capacitor action space subset composed of all capacitor switching states, Act ci is the switching state of the capacitor with index i, C i represents the action set that the i-th capacitor can take, c i1 is the action of switching off the capacitor, c i2 is the action of switching on the capacitor.

[0053] Act O = [Act o1 , Act o2 ,…, Act oi ]

[0054] O i = [o i1 , o i2 , o i3 ,..., o ix ]

[0055] wherein Act Orepresents a subset of the transformer action space consisting of all OLTC tap positions, Act oi represents the tap position of the i-th OLTC, O i represents the action set available to the i-th transformer, o ix represents the different different connections.

[0056] Act G = [Act g1 , Act g2 , …, Act gi ]

[0057] G i = [g i1 , g i2 ]

[0058] wherein, Act G represents a subset of the generator action space consisting of all generator output states, Act gi represents the action of the i-th generator, G i represents the action set available to the i-th generator, g i1 gives a step increase in the active output of the generator, g i2 gives a step decrease in the active output of the generator.

[0059] Act S = [Act s1 , Act s2 , …, Act si ]

[0060] S i = [s i1 , s i2 ]

[0061] wherein, Act S represents a subset of the generator action space consisting of all circuit breaker opening and closing states, Act si represents the action of the i-th circuit breaker, S i represents the action set available to the i-th circuit breaker, s i1 s gives an open state of the circuit breaker, s i2 s gives a closed state of the circuit breaker.

[0062] The reward function of the reinforcement learning agent is constructed based on the system global weighted voltage deviation value, specifically as follows:

[0063] reward 1t = δU t_1 - δU t_2

[0064] That is, at time t, the reward is the reduction of the weighted voltage offset value before and after the agent action; in the formula, δU t_1 , δU t_2 represents the voltage of each node before and after the agent action at time t;

[0065] And the required control voltage of the model is constrained as follows:

[0066]

[0067] In the formula, U is the voltage of each node;

[0068] When the voltage of a certain node exceeds the safe range, the reward received by the agent at this moment is 2t As follows:

[0069]

[0070] In the formula, M is 1000; U k is the voltage of any node;

[0071] And the generator set is set to have an output limit penalty reward 3t , as follows:

[0072]

[0073] In the formula, N is 5, P gen i is the real-time output value of any generator set, P gen j_min and P gen j_max are the upper and lower limits of the output of the generator set, respectively;

[0074] The global reward function of the reinforcement learning agent is as follows:

[0075] reward = reward 1t + reward 2t + reward 3t .

[0076] The training method of the above reinforcement learning agent is as follows:

[0077] During training, first, the parameters of the "actor_net" and "critic_net" networks in the proximal policy optimization deep reinforcement learning algorithm need to be initialized and set, and a buffer area is set, and the initial state of the load power curve, the distributed power system power curve, the capacitor, the line switch and the on-load tap changer, and the initial output of the generator set are input;

[0078] For any time, when the agent transits from time t k-1 to time t kWhen the power flow calculation is performed, the voltage composition of each node at the moment is obtained, and the initial state before the action of the agent is input into the two networks of "actor_net" and "critic_net", and the "actor_net" makes a judgment to obtain the strategy; when the action sequence group is generated, the default load power consumption and photovoltaic system output power do not change, the agent updates the movable device state, and the power flow operation is performed again to obtain a new state, which is also returned to the agent, and the two states are input into the "critic_net" to obtain the corresponding reward value;

[0079] vector (Sta k- , Sta k+ , R k , Act k ) will be placed in the cache buffer, wherein Sta k- represents the state before the action of the agent at time t k , Sta k+ represents the state after the action of the agent at time t k , R k represents the reward value obtained by the agent at time t k due to the action, and Act k is the action taken by the agent at this time, after each test, the agent will randomly extract L groups of vectors at this time point in the previous training process from the cache buffer, and input L groups (Sta k- , Sta k+ ) into "critic_net", the network will update the parameters, and determine the advantage function output by the network according to the original output value, L groups (Sta k- , Act k ) will be input into "actor_net", the probability ratio of taking the action under the new and old strategies is calculated, the network parameters are updated, and a new strategy function is generated;

[0080] Within fifteen minutes from t k to t k+1 , the system will maintain a steady state S k , until the time t k+1 , the load power consumption and the output power of the distributed power system change, the power flow calculation is performed again to obtain the new state Sta k before the action of the agent Act k+1- , a new action group A k+1 is taken, and then the decision training process of the agent is consistent with the time t k .

[0081] The application also provides a multi-objective scheduling optimization control system of a power system based on reinforcement learning.

[0082] an acquisition module, configured to acquire power system topology and historical voltage data of each node under multiple scenarios;

[0083] an optimization control module, configured to input the voltage data into the trained intelligent optimization control model to obtain a real-time optimization control scheme for the voltage of the power system;

[0084] The intelligent optimization control model is obtained by introducing a system weighted voltage offset value and training a proximal policy optimization reinforcement learning intelligent agent.

[0085] The system weighted voltage offset value is calculated as follows: the sensitivity matrix of each voltage node to each current branch, the sensitivity vector of each voltage node to the real-time total network loss of the system, and the sensitivity vector of each voltage node to the real-time total reactive power compensation amount of the system are obtained by using the Monte Carlo method, the total weighted sensitivity vector is obtained by weighting and adding and normalizing, and is defined as the "weight" of each node; the node "weight" is multiplied by the time sequence running voltage offset value to define the node weighted voltage offset value, and the global accumulation is obtained.

[0086] The present application has the following advantages and effects compared with the prior art:

[0087] (1) The present application applies a deep learning algorithm method to voltage control of a power system, trains an intelligent agent by using a PPO algorithm, can simultaneously complete control of different power equipment, the model can better realize scheduling optimization for a composite control target, and improves the control effect of the model; meanwhile, the model has efficient scheduling calculation capability and is suitable for composite target control of a large-scale power system.

[0088] (2) The present application constructs a deep learning intelligent agent reward function based on a system global weighted voltage offset value, the model can implement control with different degrees of accuracy according to the differences in the "weights" corresponding to different voltage nodes. The strategy generation speed is faster, the optimal solution can be found in a short time, the intelligent agent strategy updating process is accelerated, and a better optimization effect can be obtained under the guarantee of a constraint condition.

[0089] (3) The present application trains an intelligent optimization control model, the model is based on historical power operation data, and is obtained based on a Markov decision and a reinforcement learning process, exploration of an offline intelligent agent based on a behavior strategy on an environment, and training of network parameters by using a policy gradient algorithm, therefore, the voltage control scheme obtained based on the control module is a more excellent voltage control strategy, can effectively improve power flow distribution, relieve voltage out-of-limit, suppress voltage fluctuation, and thus can realize real-time optimization control of the voltage of a distribution network under high-proportion new energy access. BRIEF DESCRIPTION OF DRAWINGS

[0090] Figure 1The weighted voltage offset value generation method flowchart in the embodiment 1 of the application;

[0091] Figure 2 The multi-target optimization scheduling deep reinforcement learning agent training flowchart in the embodiment 2 of the application;

[0092] Figure 3 The bus voltage of each time point before optimization in the embodiment 3 of the application

[0093] Figure 4 The bus voltage of each time point after optimization in the embodiment 3 of the application;

[0094] Figure 5 The network loss comparison chart before and after optimization in the embodiment 3 of the application. DETAILED DESCRIPTION

[0095] The application will be described in detail below with reference to the drawings and specific embodiments.

[0096] Embodiment one

[0097] Acquisition of system global reinforcement voltage offset value

[0098] The application constructs the concept of weighted voltage offset value and completes the calculation of the weighted voltage offset value parameters of the power system model, inputs the weighted voltage offset value as a reward into the reinforcement learning distribution network model, and finally sets the model under the reward to analyze the node voltage control degree under different "weight" positions. As shown in the figure, the specific steps are as follows: Figure 1

[0099] Obtain the new power system topology structure and the historical data of each node voltage under multiple scenarios, construct the photovoltaic model, wind turbine model and composite load model;

[0100] Randomly change the source node voltage and load to calculate the voltage amplitude of each node under different states, record the branch current value, system real-time total network loss and system real-time total reactive power compensation under this state:

[0101] I=[I1,I2,...,I m ],P loss ,Q total

[0102] I is defined as a one-dimensional current matrix containing m branches in the system, P loss is defined as the total network loss of the system Q total is defined as the system real-time total reactive power compensation;

[0103] ​The sensitivity matrix of each voltage node to each current branch, the sensitivity vector of the real-time total network loss to the system, and the sensitivity vector of the real-time total reactive compensation to the system are obtained by using the Monte Carlo method:

[0104] The sensitivity matrix of each voltage node to each current branch is as follows:

[0105]

[0106] In the formula, α mn is the sensitivity of the nth node Node to the mth line in the line;

[0107] The sensitivity vector of the real-time total network loss to the system is as follows:

[0108] [β s1 ,β s2 ,...,β sn ]

[0109] In the formula, β sn is the sensitivity of the nth node Node to the real-time total network loss of the system;

[0110] The sensitivity vector of the real-time total reactive compensation to the system is as follows:

[0111] [γ s1 ,γ s2 ,...,γ sn ]

[0112] In the formula, γ sn is the sensitivity of the nth node Node to the real-time total reactive compensation of the system;

[0113] The principal component analysis method is used to reduce the current sensitivity matrix to a one-dimensional vector, and a composite sensitivity vector of voltage to all line currents is obtained:

[0114] The PCA function is called to obtain the composite sensitivity of voltage to all line currents:

[0115] [α 01 ,α 02 ,...,α 0n ]

[0116] In the formula, α 0n is the composite sensitivity of Node n to all branch currents in the system;

[0117] The three node voltage sensitivity vectors are normalized, so that the sum of the internal elements of the vector is 1:

[0118] α1+α2+α3+...+α n =1

[0119] β1+β2+β3+...+βn = 1

[0120] γ1+γ2+γ3+...+γ n = 1

[0121] In the formula, α n , β n , γ n are the branch current of the nth node, the real-time total network loss of the system, the final "weight" position parameter of the real-time total reactive power based on sensitivity of the system, respectively;

[0122] The weighted voltage offset value δU of the system about current is calculated α , the weighted voltage offset value δU about loss is calculated β , and the weighted voltage offset value δUγ about system reactive power compensation is calculated:

[0123] δUα=α1δu1+α2δu2+α3δu3+...+αnδun

[0124] δUβ=β1δu1+β2δu2+β3δu3+...+βnδun

[0125] δUγ=γ1δu1+γ2δu2+γ3δu3+...+γnδun

[0126] δu n is the voltage offset value of Node n,

[0127] α n is the "weight" of the branch current of Node n in the weighted voltage offset value after processing,

[0128] β n is the "weight" of the circuit loss of Node n in the weighted voltage offset value after processing,

[0129] γ n is the "weight" of the system reactive power compensation of Node n in the weighted voltage offset value after processing;

[0130] The weighted voltage coefficients of each part are set according to the priority of the comprehensive control target:

[0131] A+B+C=1

[0132] In the formula, A, B, and C are the weighted voltage coefficients of the branch current, the system loss, and the reactive power compensation, respectively;

[0133] The total weighted voltage offset value δU of the system is calculated:

[0134] δU=AδU α +BδU β +CδU γ

[0135] Since each voltage offset value is obtained by multiplying the node voltage by the weight owned by the node, δU can also be expressed as the sum of each node voltage multiplied by different coefficients:

[0136]

[0137] Embodiment Two

[0138] Multi-objective comprehensive control method

[0139] The power system intelligent control constraint set considering the power flow power constraint, the generator output constraint and the node voltage constraint is constructed:

[0140]

[0141]

[0142] S cource-toal +S var-toal =S load-total +S line-total

[0143] In the formula, m is the number of centralized power sources, n is the number of centralized load, S source-total is the total power output apparent power, S source-i is the apparent power of the i-th generator set, S load-total is the apparent power sum of all loads, S load-i is the apparent power of the i-th centralized load group, S var-total is the apparent power sum of all reactive power compensation devices, S line-total is the apparent power sum consumed by the system line.

[0144] 0.95U Kn ≤U k ≤1.05U kn

[0145] In the formula, U k is the real-time voltage of any node, U Kn is the rated voltage of the node.

[0146] P gen j_min ≤P gen j ≤P gen j_max

[0147] In the formula, P gen j is the real-time output of any generator, P gen j_min and P gen j_max are the upper and lower limits of the output of the generator, respectively.

[0148] The deep reinforcement intelligent agent state space containing the voltage of each node, the controlled device switching and the gear state is constructed as follows:

[0149] S = [Sta u , Sta C , Sta O , Sta G , Sta S ]

[0150] where S is the state space universe, which can be decomposed into five subsets containing node voltage information and the state of each controllable device.

[0151]

[0152]

[0153] where Sta u represents the state space voltage subset composed of each phase voltage, represents the node voltage of the kth phase at node k, where k = 1, 2, …, K, and is composed of three phases a, b, and c, with a phase difference of 120 degrees between each phase voltage;

[0154] Sta C = [Sta c1 , Sta c2 , …, Sta ci ]

[0155] Sta O = [Sta o1 , Sta o2 , …, Sta oi ]

[0156] Sta G = [Sta g1 , Sta g2 , …, Sta gi ]

[0157] Sta S = [Sta s1 , Sta s2 , …, Sta si ]

[0158] where Sta C represents the state space capacitor subset composed of all capacitor switching states, Sta ci represents the switching state of each capacitor; Sta O represents the state space transformer subset composed of all on-load tap-changing transformer tap states, Sta oi represents the tap state of each on-load tap-changing transformer; and Sta GSta G represents the real-time active power output of each generator unit; Sta S represents the state space breaker subset composed of all breaker opening and closing states, Sta si represents the opening and closing states of each breaker.

[0159] The discrete intelligent agent action space containing complete controllable devices is constructed, and the action space Act of the reinforcement learning intelligent agent includes Act c , Act o , Act G , Act s Four sub-action sets:

[0160] Act = [Act c , Act o , Act G , Act s ]

[0161] Act C = [Act c1 , Act c2 ,…, Act ci ]

[0162] C i = [c i1 ,c i2 ]

[0163] In the formula, Act C represents the capacitor action space subset composed of all capacitor switching states, Act ci is the switching state of the capacitor with index i, C i represents the action set that the i-th capacitor can take, c i1 is to remove the capacitor, c i2 is to put in the capacitor.

[0164] Act O = [Act o1 , Act o2 ,…, Act oi ]

[0165] O i = [o i1 ,o i2 ,o i3 ,..., o ix ]

[0166] In the formula, Act O represents the transformer action space subset composed of all on-load tap changer position states, Actoi represents the gear state of the i-th OLTC, O i represents the action set that the i-th transformer can take, o ix represents its different different connections.

[0167] Act G = [Act g1 , Act g2 , …, Act gi ]

[0168] G i = [g i1 , g i2 ]

[0169] wherein Act G represents the generator action space subset consisting of all generator output states, Act gi represents the action of the i-th generator group, G i represents the action set that the i-th generator group can take, g i1 is to increase the active output of the generator by one gradient, g i2 is to decrease the active output of the generator by one gradient.

[0170] Act S = [Act s1 , Act s2 , …, Act si ]

[0171] S i = [s i1 , s i2 ]

[0172] wherein Act S represents the generator action space subset consisting of all circuit breaker opening and closing states, Act si represents the action of the i-th circuit breaker, S i represents the action set that the i-th circuit breaker can take, s i1 represents opening the circuit breaker, s i2 represents closing the circuit breaker.

[0173] The reward function of the deep learning agent is constructed based on the system global weighted voltage deviation value:

[0174] reward 1t = δU t_1 - δU t_2

[0175] That is, at time t, the reward is the reduction amount of the weighted voltage deviation value before and after the action of the agent; wherein δU t_1 , δUt_2 represent the voltage of each Node before and after the agent acts at time t.

[0176] In addition, in order to ensure the voltage stability of the power system, the required control voltage of the model is constrained as follows:

[0177]

[0178] wherein is the voltage of each Node.

[0179] When the voltage of a certain Node exceeds the safe range, it will have a huge impact on the safety of the power system, so the agent will be greatly punished M in this state, and the reward received by the agent at this moment is 2t as follows:

[0180]

[0181] M is set to 1000 in the model, wherein U k is the unit value of the voltage of any Node.

[0182] In addition, the generator set is set with an output limit penalty reward 3t , which is as follows:

[0183]

[0184] N is set to 5 in the model, wherein P gen i is the real-time output value of any generator set.

[0185] reward = reward 1t + reward 2t + reward 3t

[0186] The reward is the global reward function of the agent action;

[0187] The agent is trained multiple times based on a large number of scenarios, as shown in the following formula: Figure 2 During training, the parameters of the two networks "actor_net" and "critic_net" in the proximal policy optimization (PPO) deep reinforcement learning algorithm need to be initialized and set, and a buffer area is set, and the initial state of the load power curve, the distributed power supply system power curve, the capacitor, the line switch and the on-load voltage regulator tap, and the initial output of the generator set are input.

[0188] For any time, when the agent transits from time t k-1 to time t kAt this time, the real-time power of the photovoltaic system and the real-time power consumption of the load in the power distribution network system will change, so the power of the photovoltaic system and the power of the load corresponding to the system introduction time are needed, and the power flow calculation is performed to obtain the initial state of each Node voltage before the action of the intelligent agent at the time. The intelligent agent needs to be input into two networks, "actor_net" to make a judgment to obtain a strategy, and here the intelligent agent is allowed to complete the action step by step, and the maximum step length of the action is set to 10, that is, the intelligent agent has ten times to select the action device and decide the opportunity to complete the cut or the opportunity to complete the cut. When the action sequence group is generated, the default load power consumption and photovoltaic system output power do not change, the intelligent agent updates the actionable device state, and performs power flow calculation again to obtain a new state, which will also be returned to the intelligent agent. The two states are input into "critic_net" to obtain the corresponding reward value. The vector (Sta k- , Sta k+ , R k , Act k ) will be placed in the cache area, where Sta k- represents the state of the intelligent agent before taking action at time t k , Sta k+ represents the state of the intelligent agent after taking action at time t k , R k represents the reward value obtained by the intelligent agent at time t k due to taking action, and Act k is the action taken by the intelligent agent at this time. After each test, the intelligent agent will randomly extract L groups of vectors at this time point from the cache area during the previous training process, and input L groups (Sta k- , Sta k+ ) into "critic_net". The network will update the parameters, and according to the original output value of the network, the advantage function output by the network will be determined again. L groups (Sta k- , Act k ) will be input into "actor_net" to calculate the probability ratio of taking this action under the new and old strategies, update the network parameters, and generate a new strategy function.

[0189] Within fifteen minutes from t k to t k+1 , the system will remain in a steady state S k until time t k+1 , the load power consumption and the output power of the distributed power system change, and the power flow calculation is performed again to obtain the new state Sta k of the intelligent agent before the action Act k+1- , and continue to take a new action group A k+1 , and then the decision training process of the intelligent agent at time tk The situation of one day will be divided into 96 points for decision-making, and the distributed power system power and load power curve of the day will be trained multiple times.

[0190] Effect Example 3

[0191] The voltage of each bus at each time point before and after optimization is shown in Figure 3 , Figure 4 respectively. Before optimization, due to the large amount of new energy access to the power grid, more bus voltages have already exceeded the safety range of the national standard, especially at noon and during the time with large electricity consumption, the maximum voltage value is up to 1.16, and the minimum is as low as 0.78, the out-of-limit situation is very serious. After optimization, the voltage out-of-limit situation is successfully improved, and the voltage value of each bus is stabilized between 0.95-1.05.

[0192] The comparison of network loss before and after optimization is shown in Figure 5 . Before adopting the control strategy, the maximum network loss and the average network loss are reduced by 46.67% and 25.45% respectively.

Claims

1. A power system multi-objective dispatch optimization control method based on reinforcement learning, characterized in that, The power system multi-objective scheduling optimization method comprises: acquire the power system topology structure and the historical data of each node voltage under multiple scenarios, construct a photovoltaic model, a wind turbine model and a composite load model; randomly generate power system source node voltage and load disturbance, and obtain the changed voltage amplitude of each node through power flow calculation, and record the current value of each branch, the real-time total network loss of the system and the real-time total reactive power compensation of the system under this state; use the Monte Carlo method to analyze and obtain the sensitivity matrix of each voltage node to each current branch, the sensitivity vector of each voltage node to the real-time total network loss of the system and the sensitivity vector of each voltage node to the real-time total reactive power compensation of the system, add the weights and normalize to obtain the total weighted sensitivity vector, which is defined as the "weight" of each node; multiply the node "weight" by the time sequence running voltage offset value to define the node weighted voltage offset value, and globally accumulate the node weighted voltage offset value to obtain the system weighted voltage offset value under the current state of the system; introduce the system weighted voltage offset value to train the proximal policy optimization reinforcement learning agent; and perform real-time optimization control of the power system voltage according to the trained intelligent optimization control model.

2. The power system multi-objective dispatch optimization control method of claim 1, wherein, The sensitivity matrix of the voltage node to each current branch is in the following form: wherein α mn is the sensitivity of the mth line in the pair of nodes to the nth node; The sensitivity vector of the voltage node to the real-time total network loss of the system is in the following form: [β s1 ,β s2 ,...,β sn ] where β sn is the sensitivity of the system real-time total network loss to the nth node pair; The sensitivity vector of the voltage node to the real-time total reactive power compensation of the system is in the following form: [gamma s1 , gamma s2 ,..., gamma sn ] where γ sn is the real-time total reactive power compensation sensitivity of the system to the nth node pair.

3. The power system multi-objective dispatch optimization control method of claim 2, wherein, The system weighted voltage offset value is obtained by the following steps: The computing system is configured to determine a weighted voltage shift value δU for the current α a weighted voltage shift value δU for the losses β a weighted voltage shift value δUγ for the system reactive power compensation δu α = α1δu1+ α2δu2+ α3δu3+... + α n δu n δu β = β1δu1+ β2δu2+ β3δu3+... + β n δu n δu γ = γ1δu1+ γ2δu2+ γ3δu3+... + γ n δu n where δu n is the voltage offset value of the nth node, a n The "weight" of the branch current of the nth node after processing in the weighted voltage offset value, beta n to handle the "weight" of the circuit loss in the weighted voltage offset value for the nth node pair after processing, gamma n "weight" of the n-th node after treatment in the weighted voltage deviation value for reactive power compensation of the system; set the weighted voltage coefficients of each part according to the priority of the comprehensive control target: A+B+C=1 In the formula, A, B and C respectively correspond to the weighted voltage coefficients of the branch current, the system loss and the reactive power compensation. Calculate the total weighted voltage offset value δU of the system:

4. The power system multi-objective dispatch optimization control method of claim 3, wherein, The state space of the reinforcement learning agent includes the voltage of each node, the state of the controllable device switching or gear position, and is specifically as follows: S = [Sta u , Sta C , Sta O , Sta G , Sta S ] In the formula, S is the full set of the state space, which is decomposed into five subsets containing node voltage information and the state of various controllable devices. Sta C = [Sta c1 , Sta c2 ,..., Sta ci ] Sta O = [Sta o1 , Sta o2 ,..., Sta oi ] Sra G = [Sta g1 , Sta g2 ,..., Sta gi ] Sta S = [Sta s1 , Sta s2 ,..., Sta si ] where Sta C represents the state space capacitor subset consisting of all capacitor switching states, Sta ci represents the switching state of each capacitor; Sta O represents the state space transformer subset consisting of all OLTC tap states, Sta oi represents the tap state of each OLTC; Sta G represents the state space generator subset consisting of all generator output states, Sta gi represents the real-time active power output of each generator unit; Sta S represents the state space breaker subset consisting of all breaker opening states, Sta si represents the opening state of each breaker.

5. The power system multi-objective dispatch optimization control method of claim 4, wherein, The action space Act of the reinforcement learning agent comprises Act c , Act o , Act G , Act s four sets of sub-actions, denoted as follows: Act = [Act c , Act o , Act G , Act s ] Act C = [Act c1 , Act c2 , …, Act ci ] C i = [c i1 , c i2 ] wherein Act C represents a subset of the capacitor action space consisting of all off states, Act ci is the off state of capacitor labeled i, C i represents the set of actions that the ith capacitor can take, c i1 is the off state of the capacitor, c i2 is the on state of the capacitor; Act O = [Act o1 , Act o2 , …, Act oi ] O i = [o i1 , o i2 , o i3 ,..., o ix ] wherein Act O denotes a subset of the transformer action space consisting of all on-load tap changer states, Act oi represents the on-load tap changer state of the i-th transformer, O i denotes the set of actions the i-th transformer can take, o ix denotes its different terminals; Act G = [Act g1 , Act g2 , …, Act gi ] G i = [g i1 , g i2 ] where Act G represents a subset of the generator action space consisting of all generator output states, Act gi represents the action of the i-th generator group, G i represents the set of actions that the i-th generator group can take, g i1 to increase the active power output of the generator by one step, g i2 to decrease the active power output of the generator by one step; Act S = [Act s1 , Act s2 , …, Act si ] S i = [s i1 ,s i2 ] wherein Act S denotes a subset of the generator action space consisting of all breaker open-close states, Act si represents the action of the i-th breaker, S i denotes the set of actions that the i-th breaker can take, s i1 denotes opening the breaker, s i2 denotes closing the breaker.

6. The power system multi-objective dispatch optimization control method of claim 5, wherein, The reward function of the reinforcement learning agent is constructed based on the global weighted voltage offset value of the system, and is specifically as follows: reward 1t = δU t_1 - δU t_2 That is, at time t, the reward is the reduction of the weighted voltage offset value before and after the action of the agent; In the formula, δU t_1 , δU t_2 represents the voltage of each node before and after the action of the agent at time t The control voltage required by the model is subjected to the following constraints: In the formula V is the voltage of each node; The reward received by the agent at the moment when the voltage of a certain node exceeds the safe range 2t As follows: where M is 1000; U k is an arbitrary node voltage unit And set the generator set output limit punishment reward 3t As follows: where N is 5, P gen i is the real-time output value of any generator set, P gen j_min and P gen j_max are the upper and lower limits of the output of the generator set, respectively. The global reward function of the reinforcement learning agent is as follows: reward = reward 1t + reward 2t + reward 3t .

7. The power system multi-objective dispatch optimization control method of claim 6, wherein, The training method of the reinforcement learning agent is as follows: During training, first, the parameters of "actor_net" and "critic_net" in the proximal policy optimization deep reinforcement learning algorithm are initialized and set, and a buffer area is set, and the initial state of the load power curve, the distributed power system power curve, the capacitor, the line switch and the on-load tap changer and the initial output of the generator set are input. For any time, when the agent transits from time t k-1 to time t k , the power flow calculation is performed to obtain the initial state of each node voltage before the action of the agent at the time, which is input into the two networks "actor_net" and "critic_net" respectively, and the "actor_net" makes a judgment to obtain the strategy; after the action sequence group is generated, the default load power consumption and the photovoltaic system output power do not change, the agent updates the actionable device state, and performs power flow calculation again based on the state to obtain a new state, which will also be returned to the agent, and the two states are input into the "critic_net" to obtain the corresponding reward value; Vector (S k- , S k+ , R k , A k ) will be placed in the cache buffer, where S k- represents the state space set of the agent before action at time t k , S k+ represents the state space set of the agent after action at time t k , R k is the reward value obtained by the agent at time t k , and A k is the action taken by the agent at time t k . After each test, the agent will randomly extract L groups of vectors at this time point from the cache buffer in the previous training process, and put L groups (S k- , S k+ ) into "critic_net". The network will update the parameters, and determine the advantage function output according to the original output value of the network. L groups (S k- , A k ) will be put into "actor_net" to calculate the probability ratio of taking the action under the new and old strategies, update the network parameters, and generate a new strategy function. In the fifteen minutes from t k to t k+1 , the system will maintain the steady state S k , until the time t k+1 , the load power consumption and the distributed power system output power of the system change, the power flow calculation to the new state S k+1- , that is, the initial state before the action of the intelligent agent at time t k+1 , continue to take new action group A k+1 , and then the decision-making training process of the intelligent agent is consistent with time t k .

8. A reinforcement learning based multi-objective dispatch optimization control system for power systems, characterized in that, The power system multi-objective scheduling optimization control system comprises: an acquisition module configured to acquire the power system topology structure and the historical data of each node voltage under multiple scenarios; an optimization control module configured to input the voltage data into the trained intelligent optimization control model to obtain a real-time optimization control scheme of the power system voltage; The intelligent optimization control model is obtained by introducing the system weighted voltage offset value to train the proximal policy optimization reinforcement learning agent. The system weighted voltage offset value calculation method is as follows: the sensitivity matrix of each voltage node to each current branch, the sensitivity vector of the real-time total network loss of the system, and the sensitivity vector of the real-time total reactive power compensation of the system are obtained by using the Monte Carlo method analysis respectively, the weighted sensitivity vector is obtained by weighting and adding normalization, which is defined as the "weight" of each node; the node "weight" is multiplied by the time sequence running voltage offset value to define the node weighted voltage offset value, and the global accumulation is obtained.

9. The power system multi-objective dispatch optimization control system of claim 8, wherein, The system weighted voltage offset value is obtained by the following steps: The computing system about the weighted voltage shift value δU of the current α , about the weighted voltage shift value δU of the loss β , about the weighted voltage shift value δUγ of the system reactive compensation δu α = α1δu1+ α2δ2+ α3δu3+... + α n δu n δu β = β1δu1+ β2δu2+ β3δu3+... + β n δu n δu γ = γ1δu1+ γ2δu2+ γ3δu3+... + γ n δu n where δu n is the voltage offset value of the nth node, a n The "weight" of the branch current of the nth node after processing in the weighted voltage offset value, β n to handle the "weight" of the circuit loss in the weighted voltage offset value for the nth node pair after processing, gamma n "weight" of the (n-1)th node in the weighted voltage deviation value for handling reactive power compensation of the nth node The priority of the comprehensive control target is set, and the weighted voltage coefficients of each part are set: A+B+C=1 In the formula, A, B and C respectively correspond to the weighted voltage coefficients of the branch current, the system loss and the reactive power compensation; The system total weighted voltage offset value δU is calculated:

10. The power system multi-objective dispatch optimization control system of claim 9, wherein, The state space of the reinforcement learning agent includes the state space of each node, each phase voltage, controllable device switching and gear state, and is specifically as follows: S=[Stau,StaC,StaO,StaG,StaS] In the formula, S is the whole set of state space, which is decomposed into five subsets containing node voltage information and states of various controllable devices: where Stau represents a state space voltage subset composed of phase voltages, represents the kth node voltage of the phase of the node, where consists of three phases a, b, c, and the phase difference between each phase voltage is 120 degrees. StaC=[Stac1,Stac2,...,Staci] StaO=[Stao1,Stao2,…,Staoi] StaG=[Stag1,Stag2,...,Stagi StaS=[Stas1,Stas2,...,Stasi] In the formula, StaC represents the state space capacitor subset composed of the switching states of all capacitors, and Staci represents the switching state of each capacitor; StaO represents the state space transformer subset composed of the tap position states of all on-load regulation transformers, and Staoi represents the tap position state of each on-load regulation transformer; StaG represents the state space generator subset composed of the states of all generators, and StaG represents the real-time active power output of each generator group; StaS represents the state space breaker subset composed of the switching states of all breakers, and Stasi represents the switching state of each breaker; The action space Act of the reinforcement learning agent includes four sub-action sets Actc, Acto, ActG and Acts, which are respectively represented as follows: Act=[Actc,Acto,ActG,Acts] ActC=[Actc1,Actc2,…,Actci] Ci=[ci1,ci2] In the formula, ActC represents the capacitor action space subset composed of the switching states of all capacitors, Actci is the switching state of the capacitor with the label i, Ci represents the action set that can be taken by the i-th capacitor, and ci1 is the switching-off of the capacitor and ci2 is the switching-on of the capacitor; Acto=[Acto1,Acto2,…,Actoi] Oi=[oi1,oi2,oi3,...,oix] In the formula, ActO represents a transformer action space subset composed of all on-load tap changer position states, Actoi represents the position state of the i-th on-load tap changer, Oi represents the action set that the i-th transformer can take, and oix represents different connections thereof; ActG = [Actg1, Actg2, …, Actgi] Gi = [g i1, g i2] In the formula, ActG represents a generator action space subset composed of all generator output states, Actgi represents the action of the i-th generator unit, Gi represents the action set that the i-th generator unit can take, g i1 is an increase of the active power output of the generator by one gradient, and g i2 is a decrease of the active power output of the generator by one gradient; ActS = [Acts1, Acts2, …, Actsi] Si = [s i1, s i2] In the formula, ActS represents a generator action space subset composed of all circuit breaker opening and closing states, Actsi represents the action of the i-th circuit breaker, Si represents the action set that the i-th circuit breaker can take, s i1 represents opening of the circuit breaker, and s i2 represents closing of the circuit breaker; The reward function of the reinforcement learning agent is constructed based on a system global weighted voltage deviation value, and is specifically as follows: reward 1t = δUt _1 - δUt _2 That is, at time t, the reward is the reduction of the weighted voltage offset value before and after the action of the agent; in the formula, δUt _1 , δUt _2 represents the voltage of each node before and after the action of the agent at time t And the control voltage required by the model is constrained as follows: In the formula V is the voltage of each node; When a certain node voltage exceeds a safe range, the agent is rewarded at this moment 2t As follows: In the formula, M is 1000; and Uk is a voltage unit value of any node; And set the generator set output limit punishment reward 3t As follows: In the formula, N is 5, Pgeni is the real-time output value of any generator set, Pgen j _min and Pgen j _max are the upper and lower limits of the output of the generator set, respectively. The global reward function of the reinforcement learning agent is as follows: reward = reward 1t + reward 2t + reward 3t .

Citation Information

Patent Citations

  • Power system voltage control method based on near-end strategy optimization algorithm

    CN115409650A

  • Power system low-carbon optimization scheduling method based on deep reinforcement learning

    CN116562464A