Multi-constraint multi-target power distribution network optimization control decision-making method based on state feedback

By applying a multi-constrained multi-objective optimization control method based on state feedback and a self-attention multi-agent deep reinforcement learning algorithm in the power system, the management problems of intermittent and volatility of renewable energy in the power grid are solved, and the independent optimization control and efficient and stable operation of the power grid are achieved.

CN120127633APending Publication Date: 2025-06-10HUZHOU ELECTRIC POWER SUPPLY CO OF STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197911.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

How to effectively integrate intermittent and volatile renewable energy in the power system to ensure the safe and reliable operation of the power grid, especially when the power market structure is complex.

Method used

The multi-constrained multi-objective distribution network optimization control decision-making method is adopted based on state feedback, and the multi-agent deep reinforcement learning method (AT-MADRL) is modeled and optimized reactive voltage control problems to achieve independent optimization control of the distribution network.

Benefits of technology

The dual goals of voltage control and loss minimization are achieved, ensuring efficient and stable power grid operation, adapting to changes in power grid operation status, and meeting diverse user needs and complex market structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120127633A_ABST
    Figure CN120127633A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-constraint multi-target power distribution network optimization control decision-making method based on state feedback. The method comprises the following steps: firstly, establishing a multi-objective function and constraint conditions of a reactive voltage control problem for a power distribution network; the reactive voltage control problem is modeled as a Markov decision process De-POMDP; and solving a reactive voltage control problem by utilizing AT-MADRL, and searching an optimal joint determination strategy. According to the method, the self-service optimization control decision network of the power distribution network is constructed, the dual goals of voltage control and loss minimization are achieved by monitoring the running state of the power distribution network in real time and dynamically adjusting the reactive power output of the PV and the SVC, and it is ensured that the decision process can be adjusted according to the real-time state of the power distribution network. According to the method, a state feedback mechanism and a multi-constraint multi-target optimization control method are combined, a decision network model for distribution network autonomous optimization control is achieved, the method can adapt to changes of the operation state of a power grid, and efficient and stable operation of the power grid is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power system automation, and relates to a multi-constraint and multi-objective distribution network optimal control decision-making method based on state feedback. Background Art

[0002] With the acceleration of the urbanization process and the improvement of the industrialization level, the power consumption shows an increasing trend. The global attention to environmental protection and sustainable development has promoted the rapid development of renewable energy. The wide access of clean energy such as solar energy and wind energy provides a green and low-carbon energy option for the power system. However, these renewable energies usually have intermittency and volatility, bringing new challenges to the stable operation of the power grid. Therefore, how to effectively integrate these new energies to ensure the safe and reliable operation of the power grid has become an urgent issue to be solved in the power industry. In addition, the power market is developing towards complexity, and the continuous deepening of the market-oriented reform has made power transactions more diversified. The different types of power generation enterprises and diverse user demands have led to the complexity of the market structure, which poses higher requirements for the management of the power distribution network.

[0003] Therefore, developing a technology that can achieve autonomous optimal control of the distribution network is crucial for improving the stability and efficiency of the power grid. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the above-mentioned prior art, and provide a multi-constraint and multi-objective distribution network optimal control decision-making method based on state feedback.

[0005] The present invention is implemented as follows. A multi-constraint and multi-objective distribution network optimal control decision-making method based on state feedback, the method comprising:

[0006] Establish a multi-objective function and constraint conditions for the reactive power voltage control problem of the distribution network;

[0007] Based on the established multi-objective function and constraint conditions, model the reactive power voltage control problem as a distributed partially observable Markov decision process Dec-POMDP;

[0008] Use the self-attention and temporal memory multi-agent deep reinforcement learning method AT-MADRL to solve the reactive power voltage control problem, and find the optimal joint deterministic strategy to maximize the cumulative discounted return, that is, the optimal distribution network optimal control scheme;

[0009] Wherein, the self-attention and temporal memory multi-agent deep reinforcement learning method AT-MADRL is a multi-agent deep reinforcement learning method introducing a self-optimizing control decision network for the distribution network;

[0010] The self-optimizing control decision network of the distribution network includes a feature extraction network, an auxiliary training network, an improved policy network for each agent, and an improved value network shared by all agents.

[0011] Preferably, the multi-objective function of the reactive power voltage control problem is expressed as follows:

[0012]

[0013] In the formula: F U (U i,t ) is the voltage deviation; T is the total number of scheduling cycles; t is the time index; is the set of all nodes in the distribution network except the slack node; i is the node index; U i,t is the voltage magnitude; P loss,t is the network loss; α is the coordination factor; P PV,i,t , Q PV,i,t , Q SVC,i,t , P L,i,t respectively represent the PV active power output, PV reactive power output, SVC reactive power output, and load active power demand; U 0 , U max , U min respectively represent the voltage reference value, the safety upper limit, and the safety lower limit; is a Gaussian distribution function with a mean of U 0 and a standard deviation of σ u ; β 1 , β 2 , β 3 , β 4 are shape parameters; P 0,t is the injected active power of the slack node.

[0014] Preferably, the constraint conditions of the reactive power voltage control problem include the following:

[0015] P PV,i,t -P L,i,t = ∑ i′∈I U i,t U i′,t (G ii′ cosθ ii′,t +B ii′ sinθ ii′,t )

[0016]

[0017] 0 ≤ P PV,i,t ≤ P PV,i,max

[0018] P PV,i,t 2 +QPV,i,t 2 ≤S PV,i,max 2

[0019] Q SVC,i,min ≤Q SVC,i,t ≤Q SVC,i,max

[0020] Where: θ ii′,t =θ i,t -θ i′,t is the phase angle θ of node i i,t and the phase angle θ of node i′ i′,t The difference between ii′ and B ii′ Respectively represent the real and imaginary parts of the node admittance matrix elements; Q L,i,t is the reactive power demand of load; P PV,i,max , S PV,i,max Respectively represent the maximum active power and rated apparent power of PV; Q SVC,i,max , Q SVC,i,min They are the maximum and minimum values ​​of SVC reactive output respectively.

[0021] Preferably, the reactive power voltage control problem is modeled as a distributed partially observable Markov decision process Dec-POMDP, and the specific implementation process is as follows:

[0022] Each PV and SVC in the distribution network is considered as a separate intelligent agent j;

[0023] The set of all node states in the distribution network at time t is recorded as where ε i,t =[P L,i,t ,Q L,i,t ,P PV,i,t ,U i,t-1 ,θ i,t-1 ,Q PV,i,t-1 ,Q SVC,i,t-1 ] is the state vector of node i at time t, including the load active demand, load reactive demand, PV active output at the current moment, as well as the voltage amplitude and phase angle, PV reactive output and SVC reactive output at the previous moment; if there is no load, PV or SVC at the node, the active power P and reactive power Q are both 0;

[0024] The joint observation of all agents at time t is recorded as in, is the observation sequence of agent j at time t, is the node set in the area where agent j is located;

[0025] The joint action of all agents at time t is recorded as where a j,t = [Q PV,j,t , Q SVC,j,t is the action of agent j at time t, which is the power constraint of the reactive power output of PV and SVC at the current time;

[0026] The probability of transferring from state s t and action a t to the next-time state s t+1 is denoted as which includes power flow distribution, as well as random variations of load and PV;

[0027] The multi-objective function of the reactive power voltage control problem is transformed into a single-step reward r t , r t ∈ r, and its expression is

[0028] Furthermore, a distributed partially observable Markov decision process Dec-POMDP is composed of the tuple where: is the set of agents, and j represents the index of the agent; is the state space; is the joint observation space, represents the observation space of agent j; is the joint action space, represents the action space of agent j; is the state transition probability function, satisfying the Markov property; is the joint deterministic policy, represents the deterministic policy of agent j; is the global reward, represents the value of the reward, which is a real number reflecting the immediate return obtained by the system when transferring to the next state S after taking a certain action A in a specific state S; γ ∈ [0, 1] is the discount rate, indicating the attention to future returns;

[0029] After each agent j observes o j,t at time t, it makes an action a j according to the policy μ j,t . After all agents take actions, the environment generates a new state s according to the state transition probability function t+1 and returns a single-step reward r t , thus completing a Markov decision process.

[0030] Preferably, the feature extraction network includes a linear transformation layer and a self-attention encoder, specifically through the state vectors ε of all nodes within the area i,t, forming an observation sequence o j,t ; each state vector ε i,t respectively undergoes linear transformation processing by a linear transformation layer and then is input into the self-attention encoder to obtain an observation feature λ j,t ; wherein, the self-attention encoder includes L A layers of masked self-attention encoding layers connected in series in sequence and one layer of unmasked self-attention encoding layer; L A ≥1; in each layer of masked self-attention encoding layer, the topological mask D j is the adjacency matrix corresponding to the region , D j is a symmetric matrix, the corresponding matrix elements of connected nodes are 0, and those of unconnected nodes are -∞.

[0031] Preferably, the auxiliary training network includes a linear projection layer, a layer normalization layer, a ReLU activation function, and a linear transformation layer. Specifically, the observation feature λ j,t output by the feature extraction network undergoes a linear projection layer, a layer normalization layer, a ReLU activation layer, and a linear transformation layer to obtain a deviation prediction value υ j,t of the regional voltage; then, according to the deviation prediction value υ j,t of the regional voltage, it helps the feature extraction network update parameters and mine characteristics related to voltage control;

[0032] The process of helping the feature extraction network update parameters and mine characteristics related to voltage control according to the deviation prediction value υ j,t of the regional voltage is specifically as follows:

[0033] The original observation sequence o j,t contains the voltage values of each node in the region at the previous moment According to the voltage deviation of each node, a self-supervised learning label ψ j,t is obtained, as shown in the following formula:

[0034]

[0035] where ρ is a correction coefficient used to control the influence degree of the deviation prediction value υ j,t on the label correction.

[0036] Preferably, the improved policy network adopts temporal memory neurons. Specifically, each neuron receives the hidden state h j,t-1 at the previous moment and the memory cell c j,t-1 to obtain the hidden state h j,t and the memory cell c j,t at the current moment, and sends them to the next neuron until the last neuron linearly transforms the output hidden state h j,t to obtain the action a of the agent jj,t ; The input of the first neuron is the observed feature λ output by the feature extraction network j,t The feature after linear projection, layer normalization, and ReLU activation in sequence.

[0037] Preferably, the improved value network includes a first linear projection layer, L A layers of self-attention encoders without masks, a linear projection layer, a layer normalization layer, a ReLU activation function, a temporal memory neuron, and a second linear projection layer. Specifically, the agent action a obtained by the improved policy network j,t After linear projection, it is input into L A layers of self-attention encoders without masks, and then the output sequence is concatenated into a vector; then it goes through linear projection, layer normalization, and ReLU activation in sequence and is input into the temporal memory neuron to obtain the hidden state h t which contains all historical information before time t; finally, h t is linearly projected to the global value Q glo,t .

[0038] Preferably, the self-attention and temporal memory multi-agent deep reinforcement learning method AT-MADRL is used to solve the reactive power voltage control problem, which is divided into two processes: centralized training and distributed execution; among them, the centralized training process is completed in the centralized control master station of the distribution network, enabling agents to share observations o j,t and actions a j,t , so as to learn a globally coordinated control strategy; the distributed execution process is implemented by PV and SVC distributed devices, and by collecting the state variables ε i,t in the local area where they are located, the reactive power compensation action is calculated to achieve online control decision-making.

[0039] Preferably, the centralized training process is divided into two parts: collecting experience and updating network parameters;

[0040] When collecting experience, agent j first collects the observation sequence o from the area j,t , together with the topology mask D j and sends them into the feature extraction network δ j to obtain the feature vector λ j,t , then the improved policy network generates the action a j,t according to the behavior policy; the actions a j,t of all agents form the joint action a t , and the distribution network environment executes a t to obtain the corresponding reward r t , and then transfers to the next moment state s t+1; Such a Markov decision process is carried out at each time step, resulting in an interaction experience, that is, a quadruple (s t , a t , r t , s t+1 ), which contains the current control strategy; The interaction experience is incorporated into the sample library as labeled data, that is, the experience replay array; After collecting a sufficient number of experiences, a small batch of samples are randomly drawn from the experience replay array to update the neural network parameters;

[0041] When updating the network parameters, an auxiliary training network and an improved value network are adopted to help update the network parameters of the feature extraction network and the improved policy network; After the training is completed, the network parameters of the feature extraction network and the improved policy network are sent to the edge computing devices of PV and SVC.

[0042] The present invention has the following beneficial effects:

[0043] 1. The present invention constructs a self-optimizing control decision network for a distribution network that can collect and feedback the operation state data of the distribution network in real time, aiming to dynamically adjust the reactive power output of photovoltaic devices and SVCs by monitoring the operation state of the distribution network in real time, so as to achieve the dual goals of voltage control and loss minimization, and ensure that the decision-making process can be adjusted according to the real-time state of the power grid.

[0044] 2. The present invention studies and develops a multi-constraint multi-objective optimization control method, which considers various constraints and objectives of the power grid operation. Through algorithm optimization, an optimal or approximate optimal control strategy is found on the premise of meeting all constraint conditions to achieve all set objectives.

[0045] 3. The present invention combines the state feedback mechanism with the multi-constraint multi-objective optimization control method to realize a decision network model for autonomous optimization control of the distribution network, which can adapt to the changes in the operation state of the power grid and ensure the high efficiency and stability of the power grid operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 is a flowchart of the distribution network optimization control decision method provided by the embodiment of the present invention.

[0048] Figure 2It is the parameter update process among the feature extraction network, auxiliary training network, and improved policy network for allocating each agent in the self-optimizing control decision network of the distribution network provided by the embodiments of the present invention.

[0049] Figure 3 It is the globally coordinated control strategy learned among the agents in the self-optimizing control decision network of the distribution network provided by the embodiments of the present invention. Specific embodiments

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0051] See the attached Figure 1 , Figure 2 and Figure 3 , this embodiment provides a multi-constraint and multi-objective distribution network optimization control decision method based on state feedback. The method includes:

[0052] Step S1: Establish a multi-objective function and constraint conditions for the reactive power voltage control (Volt-Var control, VVC) problem of the distribution network;

[0053] Specifically, the multi-objective function of the reactive power voltage control problem includes voltage deviation and network loss. In order to more effectively distinguish the severity of voltage inside and outside the safe range, the voltage deviation is described by an improved U-shaped function, using the Gaussian distribution function inside the safe range and the Laplace distribution function outside the safe range. The expression is as follows:

[0054]

[0055] In the formula: F U (U i,t ) is the voltage deviation; T is the total scheduling period number; t is the time index; is the set of all nodes in the distribution network except the balance node; i is the node index; U i,t is the voltage amplitude; P loss,t is the network loss; α is the coordination factor used to balance the voltage deviation and network loss; P PV,i,t , Q PV,i,t , Q SVC,i,t , P L,i,t respectively represent the active power output of distributed photovoltaic (PV), the reactive power output of PV, the reactive power output of static var compensator (SVC), and the active power demand of the load; U 0 , U max , U min respectively represent the voltage reference value, the safety upper limit, and the safety lower limit; is a Gaussian distribution function with a mean of U 0 and a standard deviation of σ u ; β 1 , β 2 , β 3 , β 4 are shape parameters used to adjust the smoothness of the function; P 0,t is the active power injected by the balanced node;

[0056] The constraint conditions of the reactive power - voltage control problem include power flow equation constraints, power constraints of PV and SVC; specifically:

[0057] P PV,i,t -P L,i,t =∑ i′∈I U i,t U i′,t (G ii′ cosθ ii′,t +B ii′ sinθ ii′,t )

[0058]

[0059] 0 ≤ P PV,i,t ≤ P PV,i,max

[0060] P PV,i,t 2 +Q PV,i,t 2 ≤ S PV,i,max 2

[0061] Q SVC,i,min ≤ Q SVC,i,t ≤ Q SVC,i,max

[0062] where: θ ii′,t =θ i,t -θ i′,t is the difference between the phase angle θ i,t of node i and the phase angle θ i′,t of node i'; G ii′ and B ii′ represent the real part and the imaginary part of the elements of the node admittance matrix respectively; Q L,i,t is the reactive power demand of the load; P PV,i,max , S PV,i,max represent the maximum active power of PV and the rated apparent power respectively; Q SVC,i,max , Q SVC,i,min are the maximum value and the minimum value of the reactive power output of SVC respectively.

[0063] Step S2: There are multiple distributed reactive power devices installed in the distribution network. Due to issues such as communication delay and privacy protection, each device based on the state feedback mechanism can only observe the state variables in a local area. Therefore, based on the established multi-objective function and constraints, the reactive power voltage control problem is modeled as a decentralized partially observable Markov decision process (Dec-POMDP), specifically as follows:

[0064] Each PV and SVC in the distribution network is regarded as a separate agent j;

[0065] The set of all node states in the distribution network at time t is denoted as where ε i,t =[P L,i,t , Q L,i,t , P PV,i,t , U i,t-1 , θ i,t-1 , Q PV,i,t-1 , Q SVC,i,t-1 is the state vector of node i at time t, including the active power demand of the load, the reactive power demand of the load, the active power output of the PV, as well as the voltage amplitude and phase angle, the reactive power output of the PV, and the reactive power output of the SVC at the previous moment. If there is no load, PV, or SVC at this node, both the active power (P) and the reactive power (Q) are 0.

[0066] The joint observation of all agents at time t is denoted as where, is the observation sequence of agent j at time t, is the set of nodes in the area where agent j is located. Multiple agents j and j' in the same area can share observations, that is, o j,t = o j′,t .

[0067] The joint action of all agents at time t is denoted as where a j,t =[Q PV,j,t , Q SVC,j,t is the action of agent j at time t, that is, the power constraint of the reactive power output of the PV and SVC at the current moment.

[0068] The probability of transferring from state s t and action a t to the next moment state s t+1 is denoted as which includes power flow distribution, as well as the random changes of load and PV.

[0069] Convert the multi-objective function of the reactive power voltage control problem into a single-step reward r t , r t ∈r, and its expression is

[0070] Furthermore, from the tuple constitute a distributed partially observable Markov decision process Dec-POMDP, where: is the set of agents, and j represents the index of the agent; is the state space; is the joint observation space, represents the observation space of agent j; is the joint action space, represents the action space of agent j; is the state transition probability function, which satisfies the Markov property; is the joint deterministic policy, represents the deterministic policy of agent j; is the global reward, represents the value of the reward, which is a real number and reflects the immediate return obtained when the system transfers to the next state S after taking a certain action A in a specific state S; γ∈[0,1] is the discount rate, indicating the attention to future returns.

[0071] After each agent j observes o j,t at time t, according to the policy μ j makes the action a j,t , after all agents take actions, the environment generates a new state s according to the state transition probability function t+1 , and returns the single-step reward r t , thus completing a Markov decision process.

[0072] The reward value can be calculated based on observable quantities such as node voltage, load power, and PV output. VVC is a fully cooperative problem, so agents share rewards. The goal of the Dec-POMDP model is to find the optimal joint control strategy, and the distributed reactive power devices coordinate with each other to minimize the voltage deviation and network loss during the entire scheduling period.

[0073] Step S3, use the multi-agent deep reinforcement learning method with self-attention and temporal memory AT-MADRL to solve the reactive power voltage control problem and find the optimal joint deterministic policy μ * , so that the cumulative discounted return is maximized, that is, the optimal distribution network optimization control scheme.

[0074]

[0075] Among them, Eμ denotes the expected value under policy μ, and argmax represents the function to take the maximum value;

[0076] Among them, the self-attention and temporal-memory multi-agent deep reinforcement learning method (AT-MADRL) is a multi-agent deep reinforcement learning method introduced into the self-optimizing control decision network of the distribution network;

[0077] The self-optimizing control decision network of the distribution network is constructed by using a self-attention encoder and considering the network topology. It includes a feature extraction network, an auxiliary training network, an improved policy network for each agent, and an improved value network shared by all agents.

[0078] The feature extraction network is constructed by using a self-attention encoder and considering the network topology. Using a self-attention encoder, the feature extraction network can identify the positions of key nodes, master the topological structure of the region, and obtain compact high-dimensional features. The number of parameters of the self-attention encoder is not affected by the length of the input sequence and will not increase the training cost as the number of nodes increases. It is suitable for large-scale distribution systems. Exemplarily, the feature extraction network includes a linear transformation layer and a self-attention encoder, specifically through the region where agent j is located the state vectors ε of all nodes within i,t to form the observation sequence o j,t ; each state vector ε i,t is respectively processed by linear transformation of the linear transformation layer and then input into the self-attention encoder to obtain the observation feature λ j,t ;

[0079] The self-attention encoder includes L A layers of masked self-attention encoding layers connected in series and one layer of unmasked self-attention encoding layer; L A ≥1;

[0080] In each layer of masked self-attention encoding layer, the topological mask D j is the adjacency matrix corresponding to the region , D j is a symmetric matrix, the corresponding matrix elements of connected nodes are 0, and those of unconnected nodes are -∞.

[0081] The feature extraction network can extract observation features from the perspective of the agent. Let δ represent the feature extraction network, as shown in the following formula:

[0082]

[0083] Exemplarily, the auxiliary training network includes a linear projection layer, a layer normalization layer, a ReLU activation function, and a linear transformation layer. Specifically, it is the observed feature λ output by the feature extraction network j,t After passing through the linear projection layer, the layer normalization layer, the ReLU activation layer, and the linear transformation layer, a deviation prediction value υ of the regional voltage is obtained j,t , which is an important indicator related to voltage control.

[0084] Then, based on the deviation prediction value υ of the regional voltage j,t help the feature extraction network update its parameters and mine the characteristics related to voltage control.

[0085] The process of helping the feature extraction network update its parameters and mine the characteristics related to voltage control based on the deviation prediction value υ of the regional voltage j,t is specifically implemented as follows:

[0086] The original observation sequence o j,t contains the voltage values of each node in the region at the previous moment According to the voltage deviation of each node, a self-supervised learning label ψ is obtained j,t , as shown in the following formula:

[0087]

[0088] where ρ is a correction coefficient used to control the influence degree of the deviation prediction value υ j,t on the label correction.

[0089] In this embodiment, self-supervised learning is adopted to construct an auxiliary training network for regional voltage deviation prediction, so as to help the upstream feature extraction network update its parameters and mine the characteristics related to voltage control. Introducing the auxiliary training network can improve the sample efficiency, accelerate and stabilize the training process, and enhance the robustness of the agent. Let ω represent the auxiliary training network, as shown in the following formula:

[0090]

[0091] Exemplarily, the role of the policy network is to generate corresponding actions according to the input observed features, so as to control the output of the agent. However, in the existing MADRL algorithm, the structure of the policy network adopts a simple feedforward neural network (FNN), and its learning ability is limited, and the action performance is poor. Since during the entire scheduling period, state variables such as PV and load show strong temporal characteristics, remembering the information of historical moments helps the agent understand the global state, so as to make the optimal action selection within the limited local observation and achieve better full-cycle control effects. Therefore, the improved policy network adopts temporal memory neurons. Specifically, each neuron receives the hidden state h of the previous moment j,t-1 and the memory cell cj,t-1 , the hidden state h at the current moment is obtained j,t and the memory cell c j,t , and they are sent to the next neuron until the output hidden state h of the last neuron j,t is linearly transformed to obtain the action a of the agent j j,t ; the input of the first neuron is the observed feature λ output by the feature extraction network j,t The feature after linear projection, layer normalization, and ReLU activation in sequence.

[0092] Exemplarily, the improved value network includes a first linear projection layer, an L A layer self-attention encoder without mask, a linear projection layer, a layer normalization layer, a ReLU activation function, a temporal memory neuron, and a second linear projection layer. Specifically, the action a of the agent obtained by the improved policy network j,t after linear projection, is input into the L A layer self-attention encoder without mask, and then the output sequence is concatenated into a vector. Then, it goes through linear projection, layer normalization, and ReLU activation in sequence, and is input into the temporal memory neuron to obtain the hidden state h t which contains all historical information before time t. Finally, h t is linearly projected to the global value Q glo,t . Let represent the improved value network, as shown in the following formula:

[0093]

[0094] The improved value network can effectively establish its correlation, improve the accuracy of action evaluation, and enable each reactive power device to be effectively coordinated. The importance of the actions of different devices for the overall reactive power voltage control of the system varies, so it is necessary to distinguish the weight sizes of each device for the global value. The traditional value network composed of FNN cannot effectively identify the weight relationship, while the improved value network with the introduction of the self-attention encoder can, according to the features and actions of all agents at the current moment, through vector concatenation, linear projection, scaled dot product calculation, and softmax normalization in sequence, obtain the weight sizes of each device, thereby depicting the influence degree of each agent on the overall voltage and power flow distribution of the system. For different working conditions, the improved value network will automatically adjust the weight sizes according to the changes of input features and actions, without manual participation, showing strong intelligence, and at the same time enhancing the interpretability of the DRL algorithm.

[0095] Specifically, the self-attention and temporal memory multi-agent deep reinforcement learning method AT-MADRL is used to solve the reactive power voltage control problem, which is divided into two processes: centralized training and distributed execution. Among them, the centralized training process is completed in the centralized control master station of the distribution network, enabling agents to share observations o j,t and actions a j,t , so as to learn a globally coordinated control strategy. The distributed execution process is implemented by PV and SVC distributed devices. By collecting the state variables ε i,t in the local area where they are located, the reactive power compensation actions are calculated to achieve online control decisions. The auxiliary training network and the improved value network are only used in the training process to help update the network parameters. After the training is completed, only the parameters of the feature extraction network and the improved policy network are sent to the edge computing devices of PV and SVC. When the reactive power equipment makes a decision, only the feed-forward operation of the neural network needs to be performed, and the calculation time is only in milliseconds, with extremely high timeliness, so as to cope with the rapid fluctuations of new energy output.

[0096] The centralized training process is divided into two parts: collecting experience and updating network parameters;

[0097] When collecting experience, agent j first collects the observation sequence o from the area j,t , together with the topology mask D j , and sends them into the feature extraction network δ j to obtain the feature vector λ j,t . Then, the improved policy network generates the action a j,t according to the behavior policy. When constructing the objective function, more distribution network control objectives can be considered, such as reducing energy consumption, ensuring power supply reliability, optimizing resource allocation, etc. Only the multi-objective function related to the VVC problem is given below. Its behavior policy is different from the target policy, as follows:

[0098]

[0099] In the formula: ξ beh is the behavior noise, randomly drawn from a Gaussian distribution beh with a mean of 0 and a standard deviation of σ .

[0100] The joint action a j,t is composed of the actions a t of all agents. The distribution network environment executes a t to obtain the corresponding reward r t , and then transfers to the next moment state s t+1 according to the random changes of PV and load.。Such a Markov decision process is carried out at each time step, resulting in an interaction experience, that is, a quadruple (s t , a t , r t , s t+1 ), which contains the current control strategy.

[0101] The interaction experience is incorporated into the sample library as labeled data, that is, the experience replay array. Using experience replay can, on the one hand, break the correlation of sequences, make the extracted samples independent of each other, and ensure the stability of the training process. On the other hand, it can reuse the collected experience, thereby improving the sample efficiency.

[0102] After collecting a sufficient number of experiences, a small batch of samples are randomly drawn from the experience replay array to update the neural network parameters. Suppose a total of L samples are drawn, and the l-th quadruple is (s l , a l , r l , s l '), where The superscript' represents the variable at the next moment.

[0103] When updating the network parameters, an auxiliary training network and an improved value network are adopted to help update the network parameters of the feature extraction network and the improved policy network. After the training is completed, the network parameters of the feature extraction network and the improved policy network are sent to the edge computing devices of PV and SVC; specifically:

[0104] First, the observation sequence o j,l and the topology mask D j are sent into the feature extraction network δ j and the auxiliary training network ω j in turn to generate the feature vector λ j,l and the deviation prediction υ j,l . Then, according to the node voltages included in o j,l , the actual value ψ j,l of the voltage deviation is calculated and used as the label for self-supervised learning. Calculate the mean square error of the L samples to obtain the voltage deviation loss lossυ j , and then update the parameters of the feature extraction network δ j and the auxiliary training network ω j by gradient descent as follows: as follows:

[0105]

[0106] In the formula: η δ , η ω is the learning rate of the corresponding network.

[0107] An embodiment of the present invention provides an electronic device. Specifically, the electronic device includes a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the multi-constraint multi-objective distribution network optimization control decision-making method according to any one of the implementation manners is implemented.

[0108] Among them, the memory may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0109] The bus may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0110] Among them, the memory is used to store a program. After receiving an execution instruction, the processor executes the program. The method executed by the device defined by the flow process disclosed in any embodiment of the foregoing embodiments of the present invention can be applied to the processor or implemented by the processor.

[0111] A processor may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0112] The computer program product of the readable storage medium provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the multi-constraint multi-objective distribution network optimization control decision-making method described in the foregoing method embodiments. For the specific implementation, reference can be made to the foregoing method embodiments, and details are not described herein again.

[0113] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0114] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications are also regarded as the protection scope of the present invention.

Claims

1. A multi-constraint and multi-objective distribution network optimization control decision method based on state feedback, characterized in that: The method comprises: Establish multi-objective functions and constraints for reactive power and voltage control problems in distribution networks; Based on the established multi-objective functions and constraints, the reactive power and voltage control problem is modeled as a distributed partially observable Markov decision process Dec-POMDP; The self-attention and temporal memory multi-agent deep reinforcement learning method AT-MADRL is used to solve the reactive voltage control problem and find the optimal joint determination strategy to maximize the cumulative discounted return, that is, the optimal distribution network optimization control scheme. Among them, the self-attention and temporal memory multi-agent deep reinforcement learning method AT-MADRL is a multi-agent deep reinforcement learning method that introduces a self-service optimization control decision network for distribution networks; The distribution network self-service optimization control decision network includes a feature extraction network allocated to each intelligent agent, an auxiliary training network, an improved strategy network, and an improved value network shared by all intelligent agents.

2. The method according to claim 1, characterized in that: The multi-objective function of the reactive power voltage control problem is expressed as follows: Where: F U (U i,t ) is the voltage deviation; T is the total number of scheduling cycles; t is the time index; is the set of all nodes in the distribution network except the balancing node; i is the node index; U i,t is the voltage amplitude; P loss,t is the network loss; α is the coordination factor; P PV,i,t , Q PV,i,t , Q SVC,i,t , P L,i,t Respectively represent PV active output, PV reactive output, SVC reactive output and load active demand; U0, U max ,U min Respectively represent voltage reference value, upper safety limit and lower safety limit; The mean is U0 and the standard deviation is σ u Gaussian distribution function; β1, β2, β3, β4 are shape parameters; P 0,t It is the injected active power of the balancing node.

3. The method according to claim 1 or 2, characterized in that: The constraints of the reactive voltage control problem include the following: P PV,i,t -P L,i,t =∑ i′∈I U i,t U i′,t (G ii′ cosθ ii′,t +B ii′ sinθ ii′,t ) 0≤P PV,i,t ≤P PV,i,max P PV,i,t 2 +Q PV,i,t 2 ≤S PV,i,max 2 Q SVC,i,min ≤Q SVC,i,t ≤Q SVC,i,max Where: θ ii′,t =θ i,t -θ i′,t is the phase angle θ of node i i,t and the phase angle θ of node i′ i′,t The difference between ii′ and B ii′ Respectively represent the real and imaginary parts of the node admittance matrix elements; Q L,i,t is the reactive power demand of load; P PV,i,max , S PV,i,max Respectively represent the maximum active power and rated apparent power of PV; Q SVC,i,max , Q SVC,i,min They are the maximum and minimum values ​​of SVC reactive output respectively.

4. The method according to claim 1, characterized in that: The specific implementation process of modeling the reactive power voltage control problem as a distributed partially observable Markov decision process Dec-POMDP is as follows: Each PV and SVC in the distribution network is considered as a separate intelligent agent j; The set of all node states in the distribution network at time t is recorded as where ε i,t =[P L,i,t ,Q L,i,t ,P PV,i,t ,U i,t-1 ,θ i,t-1 ,Q PV,i,t-1 ,Q SVC,i,t-1 ] is the state vector of node i at time t, including the load active demand, load reactive demand, PV active output at the current moment, as well as the voltage amplitude and phase angle, PV reactive output and SVC reactive output at the previous moment; if there is no load, PV or SVC at the node, the active power P and reactive power Q are both 0; The joint observation of all agents at time t is recorded as in, is the observation sequence of agent j at time t, is the node set in the area where agent j is located; The joint action of all agents at time t is recorded as where a j,t =[Q PV,j,t ,Q SVC,j,t ] is the action of agent j at time t, that is, the power constraint of reactive power output of PV and SVC at the current moment; By state t and action a t Transfer to the next state s t+1 The probability of It includes power flow distribution, as well as random variations of load and PV; The multi-objective function of the reactive power and voltage control problem is transformed into a single-step reward r t , r t ∈r, its expression is Then by the tuple It constitutes a distributed partially observable Markov decision process Dec-POMDP, where: is the set of agents, j represents the index of the agent; is the state space; is the joint observation space, represents the observation space of agent j; is the joint action space, represents the action space of agent j; is the state transition probability function, satisfying the Markov property; For the joint deterministic strategy, represents the deterministic strategy of agent j; For global rewards, represents the value of the reward, which is a real number that reflects the immediate return obtained by the system after taking an action A in a specific state S and transferring to the next state S; γ∈[0,1] is the discount rate, which indicates the attention paid to future returns; Each agent j observes o at time t j,t After that, according to the strategy μ j Make an action j,t , the environment after all agents act, according to the state transition probability function Generate a new state s t+1 , and returns the single-step reward r t , thus completing a Markov decision process.

5. The method according to claim 1, characterized in that: The feature extraction network includes a linear transformation layer and a self-attention encoder, specifically, the region where the agent j is located The state vector ε of all nodes in i,t , forming the observation sequence o j,t ; Each state vector ε i,t After being processed by the linear transformation layer, they are input into the self-attention encoder to obtain the observed feature λ j,t ; Wherein, the self-attention encoder includes L A Layer with masked self-attention encoding layer, layer without masked self-attention encoding layer; L A ≥1; topological mask D in each masked self-attention encoding layer j For Region The corresponding adjacency matrix, D j It is a symmetric matrix, the corresponding matrix elements of connected nodes are 0, and those of unconnected nodes are -∞.

6. The method according to claim 1, characterized in that: The auxiliary training network includes a linear projection layer, a layer normalization layer, a ReLU activation function and a linear transformation layer, specifically, the observed feature λ output by the feature extraction network j,t After the linear projection layer, layer normalization layer, ReLU activation layer and linear transformation layer, the deviation prediction value υ of the regional voltage is obtained. j,t ; Then, according to the deviation prediction value of the regional voltage, j,t Help the feature extraction network update parameters and mine characteristics related to voltage control; The deviation prediction value υ based on the regional voltage j,t Help the feature extraction network update parameters and mine the characteristics related to voltage control. The specific implementation process is: The original observation sequence o j,t Contains the voltage value of each node in the region at the previous moment According to the voltage deviation of each node, the label ψ of self-supervised learning is obtained j,t , as shown below: Among them, ρ is a correction coefficient used to control the deviation prediction value υ j,t The degree of impact on label correction.

7. The method according to claim 1, characterized in that: The improved strategy network adopts temporal memory neurons, specifically, each neuron receives the hidden state h of the previous moment. j,t-1 and memory element c j,t-1 , get the hidden state h at the current moment j,t and memory element c j,t , and sent to the next neuron until the last neuron outputs the hidden state h j,t After linear transformation, we get the action a of agent j j,t ; The input of the first neuron is the observed feature λ output by the feature extraction network j,t Features after linear projection, layer normalization, and ReLU activation.

8. The method according to claim 1, characterized in that: The improved value network includes a first linear projection layer, L A The first layer is a self-attention encoder without mask, a linear projection layer, a normalization layer, a ReLU activation function, a temporal memory neuron, and a second linear projection layer. Specifically, the agent action a obtained by the improved strategy network is j,t After linear projection, input to L A The output sequence is concatenated into a vector in the self-attention encoder without mask. After that, it is linearly projected, layer normalized, and ReLU activated, and then input into the temporal memory neuron to obtain the hidden state h t contains all the historical information before time t; finally, h t Linear projection to global value Q glo,t .

9. The method according to claim 1, characterized in that: The self-attention and temporal memory multi-agent deep reinforcement learning method AT-MADRL is used to solve the reactive voltage control problem. It is divided into two processes: centralized training and distributed execution. The centralized training process is completed in the centralized control master station of the distribution network, so that the agents can share observations. j,t and action a j,t , thereby learning the global coordinated control strategy; the distributed execution process is implemented by PV and SVC distributed devices, and the state quantity ε in the local area is collected i,t , the reactive power compensation action is calculated and the online control decision is realized.

10. The method according to claim 9, characterized in that: The centralized training process is divided into two parts: collecting experience and updating network parameters; When collecting experience, agent j first collects Collect observation sequence o j,t , together with the topological mask D j Feed it into the feature extraction network δ j In the equation, we get the eigenvector λ j,t , and then improve the strategy network to generate action a according to the behavior strategy j,t ; The actions a of all agents j,t Composition of joint actions t , the distribution network environment executes a t , and get the corresponding reward r t , and then transfer to the next state s according to the random changes of PV and load t+1 ; Such a Markov decision process is carried out at each time step, which generates an interactive experience, that is, a four-tuple (s t ,a t ,r t ,s t+1 ), which contains the current control strategy; the interaction experience is included in the sample library as label data, that is, the experience replay array; when a sufficient amount of experience is collected, a small batch of samples is randomly extracted from the experience replay array to update the neural network parameters; When updating network parameters, auxiliary training networks and improved value networks are used to help update the network parameters of feature extraction networks and improved strategy networks; after the training is completed, the network parameters of feature extraction networks and improved strategy networks are sent to the edge computing devices of PV and SVC.