Power distribution network emergency recovery method, system and device and storage medium

By combining spatiotemporal graph attention capsule networks with long short-term memory networks and optimizing with dual deep Q networks, the problems of spatiotemporal feature decoupling and topological mutation in distribution network emergency recovery are solved, enabling efficient and safe generation of fault recovery strategies and ensuring the rapid response and reliability of the distribution network.

CN120855264APending Publication Date: 2025-10-28GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510726284.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing emergency recovery methods for distribution networks cannot accurately decouple and represent spatiotemporal characteristics, have high operational risks, and cannot adapt to topological changes caused by switching actions, resulting in distorted node embedding.

Method used

By combining spatiotemporal graph attention capsule network with long short-term memory network, and constructing a power distribution network emergency recovery model through graph structured representation and finite Markov decision process, and combining dual deep Q network optimization decision, dynamic feature fusion and optimal strategy search under security constraints are achieved.

Benefits of technology

It achieves high-density spatiotemporal embedding expression, reduces the risk of dangerous operations, ensures the reliability and rapid fault recovery capability of the distribution network, and meets the minute-level timeliness requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120855264A_ABST
    Figure CN120855264A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution network emergency recovery method, system and device and a storage medium, and the method comprises the steps: obtaining the operation environment information of an active power distribution network system, constructing a feature extraction model, carrying out the graph structured representation of the active power distribution network system, and extracting the emergency recovery decision features of the power distribution network; constructing a power distribution network emergency recovery model based on the power distribution network emergency recovery decision features; and converting the emergency recovery problem of the power distribution network into a finite Markov decision process, and calculating and updating operation state parameters of the power distribution network to obtain an optimal emergency recovery strategy of the power distribution network. According to the invention, through collaborative optimization of a space-time diagram attention mechanism and a double-Q network, multi-dimensional feature fusion under topology dynamic change is realized; the capacity of capturing implicit topological association is enhanced, and the pertinence and robustness of feature fusion are improved; through parameter soft update and a dynamic punishment mechanism, optimal strategy search under security constraints is realized, voltage out-of-limit and line overload risks are reduced, and physical feasibility of a recovery strategy is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of emergency restoration of power distribution networks, and in particular to a method, system, equipment and storage medium for emergency restoration of power distribution networks. Background Technology

[0002] Traditional methods employ graph convolutional networks and long short-term memory networks to independently model spatial topology and temporal series, resulting in a lack of spatiotemporal feature interaction. Specifically, in the spatial dimension, static graph convolution only aggregates single-hop neighborhood information, failing to capture dynamically formed cross-regional power supply paths during fault recovery (such as multi-layer topological dependencies caused by tie switch closures), leading to missed detection of critical line overload risks. In the temporal dimension, independent long short-term memory network modules ignore the impact of network reconstruction actions on temporal evolution. Traditional deep Q-networks suffer from two major drawbacks in switch combination decisions: Q-value overestimation and ineffective action exploration. Evaluating the value of actions with a single network leads to systematic overestimation bias, causing frequent switching operations and equipment overload; while the ε-greedy strategy generates many illegal actions under radial topology constraints, requiring repeated rollbacks to reduce exploration efficiency. Static graph neural networks assume a fixed adjacency matrix and cannot adapt to topological changes caused by switching actions, resulting in distorted node embeddings.

[0003] Existing graph convolutional network methods assume a fixed distribution network topology and update embedded features only through node attributes. However, fault recovery processes involve dynamic switching of switch states, resulting in the time-varying nature of the adjacency matrix A not being modeled. Existing graph attention mechanisms aggregate neighbor information through scalar attention coefficients but cannot characterize multidimensional relationships between nodes (such as voltage phase difference and power transmission direction). The attitude matrix of capsule networks could encode such high-order features, but current research has not yet combined it with spatiotemporal graph models, leading to low-dimensional embedding distortion of key electrical parameters. Summary of the Invention

[0004] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides a distribution network emergency recovery method, system, device, and storage medium to solve the problems of existing distribution network emergency recovery methods, such as inaccurate decoupling and characterization of the spatiotemporal characteristics of the distribution network, high operational risks, and inability to adapt to topological abrupt changes caused by switching actions, resulting in node embedding distortion.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0006] In a first aspect, embodiments of the present invention provide a method for emergency restoration of a power distribution network, comprising:

[0007] Obtain information about the operating environment of the active distribution network system;

[0008] A feature extraction model is constructed to perform graph-structured representation of the active distribution network system and extract the emergency recovery decision features of the distribution network.

[0009] A distribution network emergency recovery model is constructed based on the aforementioned distribution network emergency recovery decision characteristics;

[0010] The emergency recovery problem of the distribution network is transformed into a finite Markov decision process. The operating state parameters of the distribution network are calculated and updated to obtain the optimal emergency recovery strategy for the distribution network.

[0011] As a preferred embodiment of the distribution network emergency recovery method of the present invention, the acquisition of active distribution network system operating environment information includes: node current, node voltage, load active power and load reactive power.

[0012] As a preferred embodiment of the distribution network emergency recovery method described in this invention, the following is included: constructing a feature extraction model to perform graph-structured representation of the active distribution network system:

[0013] The feature extraction model learns graph topology knowledge by introducing a long short-term memory network-graph attention capsule network to perceive and extract information about irregular environments in active distribution networks.

[0014] The topology of the distribution network is represented by an adjacency matrix and a feature matrix. The adjacency matrix describes the connection relationship between nodes, and the feature matrix contains the voltage, current, active power and reactive power of each node at different time points. The branches with connection relationships and the nodes with power information in the distribution network system are mapped to the top set and edge set in a graph structure.

[0015] As a preferred embodiment of the distribution network emergency recovery method of the present invention, the extraction of distribution network emergency recovery decision features includes: modeling time series data using a long short-term memory network to extract dynamic features from the time series data; the structure of the long short-term memory network includes a forget gate, an input gate, and an output gate; the input format is a time series vector in the form of a sliding window, and the encoder long short-term memory network outputs a hidden state, which contains the dynamic features of the time series.

[0016] The features output by the Long Short-Term Memory Network are weighted, the weight of each feature is calculated and normalized, and spatial features are extracted using a graph attention capsule network.

[0017] Node attributes are mapped to a high-dimensional space, graph capsule layers are used for graph convolution operations, and intermediate features are calculated using graph Laplacian matrix and multinomial filter. The output of each layer is a concatenation of multiple hop count neighborhood information to represent the node state. Node embedding and graph embedding are obtained through linear transformation, and the graph embedding is used to obtain the final distribution network emergency recovery decision features through a multilayer perceptron.

[0018] The beneficial effects of this preferred technical solution are as follows: by coupling dynamic graph convolution and long short-term memory network into a spatiotemporal graph attention capsule network for temporal modeling, it breaks through the limitation of independent processing of spatial topology and temporal data in traditional methods; multi-order graph convolution accurately captures the correlation of cross-regional power supply paths, and the long short-term memory network gating mechanism synchronously updates the topology evolution and the fluctuation characteristics of new energy output, realizing high-density spatiotemporal embedding expression, providing global-local collaborative perception capability for decision-making, and ensuring the reliability and security of power supply.

[0019] As a preferred embodiment of the distribution network emergency recovery method of the present invention, the construction of a distribution network emergency recovery model based on the distribution network emergency recovery decision characteristics includes: the emergency recovery objective function of the distribution network emergency recovery model under the distribution network fault scenario is expressed as:

[0020] F U = (f1+f2+f3)

[0021]

[0022] Where f1, f2, and f3 represent minimizing network loss, minimizing the number of switching actions, and maximizing load recovery, respectively. l R is the current in branch l. l Let N be the resistance of branch l. z Let N be the set of branches. d For a set of segmented switches, N l For the set of interconnecting switches, x z Indicates the state of the switch, λ h x is the weighting coefficient for load h. h Given the state of load h, P h N is the active power of load h. f This represents the total number of loads to be restored.

[0023] Constraints include power flow constraints, safety constraints, and radial operation constraints.

[0024] As a preferred embodiment of the distribution network emergency restoration method described in this invention, the method involves transforming the distribution network emergency restoration problem into a finite Markov decision process, including:

[0025] The distribution network system environment is perceived through a feature extraction model, and the emergency recovery decision features of the distribution network are obtained from the node adjacency matrix and the feature matrix as the state space;

[0026] The sequence of switching operations for each step in the fault recovery process is used as the action space.

[0027] When the fault is recovered, the system load recovery rate reaches 100% and all operating constraints are met, a sparse reward value is obtained.

[0028] Given a state, the action-value function is calculated by weighting the cumulative reward of the target time step by a discount factor, and the optimal strategy is found so that performing the corresponding action in any state can maximize the long-term return.

[0029] As a preferred embodiment of the distribution network emergency recovery method described in this invention, the calculation and updating of distribution network operating status parameters to obtain the optimal distribution network emergency recovery strategy includes:

[0030] The current Q-network and the target Q-network are responsible for action selection and value evaluation, respectively. The target value is calculated by the target Q-network, and the current Q-network updates its parameters based on the difference between the predicted value and the target value.

[0031] The mean squared error loss function is used to measure the difference between the current Q value and the target Q value. Gradient descent is used to optimize until iterative convergence, and action decisions are output to determine the state of each switch in the distribution network.

[0032] After executing an action, the state at the next moment is obtained, and the corresponding reward value is obtained. The experience is stored in the experience replay buffer. Small batches of experience are randomly sampled from the buffer, and the network parameters are updated using the loss function of the dual deep Q-learning network. The parameters of the target network are adjusted using a soft update method. If the set number of rounds or the average reward converges, the training ends; otherwise, the next training cycle continues until the optimal distribution network emergency recovery strategy is obtained.

[0033] The beneficial effects of this preferred technical solution are that it achieves dual optimization through the collaborative design of a dual-depth Q-network architecture and a dynamic action masking mechanism. The dual-depth Q-network decouples action selection from value assessment and suppresses Q-value overestimation bias through target network soft updates, effectively reducing the risk of dangerous operations.

[0034] Secondly, the present invention provides a power distribution network emergency recovery system, comprising:

[0035] The environmental sensing module is used to acquire information about the operating environment of the active distribution network system.

[0036] The feature extraction module is used to construct a feature extraction model to perform graph-structured representation of the active distribution network system and extract the emergency recovery decision features of the distribution network.

[0037] The model building module is used to build a distribution network emergency recovery model based on the distribution network emergency recovery decision characteristics.

[0038] The optimization decision module is used to transform the emergency recovery problem of the distribution network into a finite Markov decision process, calculate and update the operating state parameters of the distribution network, and obtain the optimal emergency recovery strategy for the distribution network.

[0039] Thirdly, the present invention provides an electronic device, comprising:

[0040] Memory and processor;

[0041] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the power distribution network emergency restoration method.

[0042] Fourthly, the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the power distribution network emergency recovery method.

[0043] Compared with existing technologies, the advantages of this invention are as follows: This invention breaks through the limitation of independent processing of spatial topology and temporal data in traditional methods by coupling dynamic graph convolution and long short-term memory network temporal modeling through spatiotemporal graph attention capsule network; multi-order graph convolution accurately captures cross-regional power supply path correlation, and the long short-term memory network gating mechanism synchronously updates the topology evolution and new energy output fluctuation characteristics, realizing high-density spatiotemporal embedding expression and providing global-local collaborative perception capability for decision-making; dual optimization is achieved through the collaborative design of dual-depth Q network architecture and dynamic action masking mechanism. The dual-depth Q network decouples action selection and value assessment, and suppresses Q-value overestimation bias through target network soft update, effectively reducing the risk of dangerous operation; based on the adjacency matrix real-time update and context feature fusion mechanism (power supply, over-limit indicators), it adaptively perceives topological changes caused by switching actions. The strategy generation time is lower than that of traditional mathematical optimization methods, meeting the stringent timeliness requirements of minute-level fault recovery, ensuring the reliability and response speed of the system, and enabling rapid power restoration after a fault occurs, ensuring the normal operation of important loads. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0045] Figure 1 This is a schematic diagram of a method flow for an emergency power distribution network restoration method, system, equipment, and storage medium according to an embodiment of the present invention;

[0046] Figure 2 This is a framework diagram of a power distribution network emergency recovery strategy according to an embodiment of the present invention, which includes a power distribution network emergency recovery method, system, equipment, and storage medium.

[0047] Figure 3 This is a schematic diagram of a power distribution network emergency recovery strategy flow according to an embodiment of the present invention, which describes a power distribution network emergency recovery method, system, equipment and storage medium.

[0048] Figure 4 This is an improved IEEE-33 node topology diagram of a power distribution network emergency recovery method, system, equipment and storage medium according to an embodiment of the present invention;

[0049] Figure 5 This is a schematic diagram of the reward value obtained during training according to an embodiment of the distribution network emergency recovery method, system, equipment, and storage medium of the present invention.

[0050] Figure 6 This is a node voltage distribution diagram of a power distribution network according to an embodiment of the present invention, which describes a power distribution network emergency recovery method, system, equipment, and storage medium. Detailed Implementation

[0051] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0052] Example 1, referring to Figures 1-3 As one embodiment of the present invention, this embodiment provides a power distribution network emergency restoration method, including:

[0053] S100: Obtain information on the operating environment of the active distribution network system;

[0054] S200: Construct a feature extraction model to perform graph-structured representation of the active distribution network system and extract the emergency recovery decision features of the distribution network;

[0055] S300: Constructing an emergency recovery model for distribution networks based on the emergency recovery decision-making characteristics of distribution networks;

[0056] S400: The emergency recovery problem of the distribution network is transformed into a finite Markov decision process, which calculates and updates the operating state parameters of the distribution network to obtain the optimal emergency recovery strategy for the distribution network.

[0057] It should be noted that traditional methods use graph convolutional networks and long short-term memory networks to independently model spatial topology and temporal series, resulting in a lack of spatiotemporal feature interaction. In the spatial dimension, static graph convolution only aggregates single-hop neighborhood information and cannot capture dynamically formed cross-regional power supply paths during fault recovery, leading to missed detection of critical line overload risks. In the temporal dimension, independent long short-term memory network modules ignore the impact of network reconstruction actions on temporal evolution. Traditional deep Q-networks suffer from two major drawbacks in switch combination decisions: Q-value overestimation and ineffective action exploration. Evaluating the value of actions by a single network can lead to systematic overestimation bias, causing frequent switching operations and equipment overload; while the ε-greedy strategy generates many illegal actions under radial topology constraints, requiring repeated rollbacks to reduce exploration efficiency. Static graph neural networks assume a fixed adjacency matrix and cannot adapt to topological changes caused by switching actions, resulting in distorted node embedding. This application presents a distribution network emergency recovery framework based on spatiotemporal graph attention capsule networks and double deep Q-learning. Through the collaborative optimization of spatiotemporal graph attention mechanisms and double Q-networks, it achieves multi-dimensional feature fusion under dynamic topology changes. A feature extraction model based on spatiotemporal graph attention capsule networks is established to enhance the ability to capture implicit topological associations, effectively distinguish key nodes from redundant information, and improve the targeting and robustness of feature fusion. A distribution network emergency recovery solution model based on double deep Q-networks (DDQN) is designed. Through parameter soft updates and dynamic penalty mechanisms, it achieves optimal strategy search under safety constraints, reduces the risk of voltage exceedance and line overload, and ensures the physical feasibility of the recovery strategy.

[0058] In this embodiment of the application, obtaining the operating environment information of the active distribution network system in step S100 includes: node current, node voltage, load active power and load reactive power.

[0059] In this embodiment of the application, step S200, which involves constructing a feature extraction model to perform graph-structured representation of the active distribution network system, includes:

[0060] The feature extraction model learns graph topology knowledge by introducing a Long Short-Term Memory Network-Graph Attention Capsule Network (LGACN) to perceive and extract information about irregular environments in the active distribution network.

[0061] The topology of the distribution network is represented by an adjacency matrix and a feature matrix. The adjacency matrix describes the connection relationship between nodes, and the feature matrix contains the voltage, current, active power and reactive power of each node at different time points. The branches with connection relationships and the nodes with power information in the distribution network system are mapped to the top set and edge set in a graph structure.

[0062] For example, for a distribution network system G, its graph structured form is represented as follows:

[0063] G = A, X

[0064] Here, A is called the vertex set, and X is called the edge set;

[0065] Furthermore, the top set represents the adjacency matrix formed by the connection relationships of nodes G in the distribution network system. If the node v of the distribution network system G... i With v j If there is a connection between them, then matrix element A ij =1, otherwise A ij =0; The topology connection relationship of the distribution network system G is represented as follows:

[0066]

[0067] Edge sets are used to store the voltage of distribution network nodes. With line current Active power at each node and reactive power The characteristic matrix formed is represented as follows:

[0068]

[0069] X t ={h t,1 ,h t,2 ,...,h t,N},h t,i ∈R d

[0070] Among them, h t,i For node v i The input feature vector at time t, where N is the number of nodes, d is the number of node features, and T is the time length.

[0071] In this embodiment of the application, step S200 of extracting the emergency recovery decision features of the power distribution network includes: modeling the time series data using a long short-term memory network and extracting the dynamic features in the time series data; the structure of the long short-term memory network includes a forget gate, an input gate, and an output gate; the input format is a time series vector in the form of a sliding window, and the encoder long short-term memory network outputs a hidden state, which contains the dynamic features of the time series.

[0072] Specifically, Long Short-Term Memory (LSTM) networks are used to extract dynamic features X from sequences. t LSTM utilizes a specially designed forgetting gate f t Input gate i t and output gate ot To alleviate the problems of vanishing and exploding gradients; f t The part that was decided will be forgotten, while i t The determined part will be added to c t In the middle, output gate o t According to h t-1 and x t Calculate the output gate value. The structure of the LSTM unit is shown below:

[0073] f t =σ(W f h t-1 +V f x t +b f )

[0074] i t =σ(W i h t-1 +V i x t +b i )

[0075] o t =σ(W o h t-1 +V o x t +b o )

[0076]

[0077] h t =o t ·tanh(c t )

[0078] Where σ is the Sigmoid function, tanh is the hyperbolic tangent function, and c t Represents the state of a time unit, used for memory storage, c t Based on the previous unit state c t-1 and candidate memory states Update its parameters, W f W i W o W a W y V f V i V o V a ,b f ,b i ,b o ,b a ,b yThese are the parameters updated during training, given the input vector x = (x1,...,x...). t ,...,x L ), where x t ∈R n Let t represent an n-dimensional vector, where t∈[1,L] represents the time step, and L represents the sliding window length.

[0079] encoder Take vector x as input and output hidden state h. m =(h m,1 ,h m,2 ,...,h m,L ):

[0080] (h m,1 ,...,h m,L )=f enc (x1,x2,...,x L ;θ enc )

[0081] Where m∈[1,M], M represents the number of layers in the encoder LSTM, k represents the number of hidden units in the encoder LSTM, and θ enc =[W enc U enc ,b enc ] represents the learnable parameters of the encoder LSTM.

[0082] In this embodiment of the application, step S200 of extracting the emergency recovery decision features of the distribution network includes: weighting the features output by the long short-term memory network, calculating the weight of each feature and normalizing it, and using a graph attention capsule network to extract spatial features.

[0083] Node attributes are mapped to a high-dimensional space, graph capsule layers are used for graph convolution operations, and intermediate features are calculated using graph Laplacian matrix and multinomial filter. The output of each layer is a concatenation of multiple hop count neighborhood information to represent the node state. Node embedding and graph embedding are obtained through linear transformation, and the graph embedding is used to obtain the final distribution network emergency recovery decision features through a multilayer perceptron.

[0084] Specifically, to improve the model's ability to capture long dependencies and contextual information in time series prediction, an attention enhancement mechanism is introduced to assign weights to each part of the input data, allowing the model to focus on key features.

[0085] Furthermore, an attention enhancement mechanism is introduced to dynamically adjust the LSTM output h. m The feature weights are expressed as:

[0086] e t =tanh(W ah m +b a )

[0087]

[0088] Among them, e t For feature h m The weight value, W a With b a For the weights and biases of the attention layer, e t,i For e t The i-th element, z t,i N represents the normalized weight values. e For e t The number of elements.

[0089] Furthermore, a linear transformation is used to transform the node attribute z t,i Projecting onto a higher-dimensional space, it can be represented as:

[0090] F i init =W init ×z t,i +b init

[0091] in, And b init These are the learnable weights and biases, respectively.

[0092] The cardinality of a vector or set is denoted by |·|, h0 represents the projected length, and F... init For a matrix Indicates all

[0093] Each feature vector F i init (i∈N) will pass through a series of graph capsule layers, which use polynomial-form graph convolutional filters to compute a matrix. Defined as:

[0094]

[0095] in, Let p represent the graph Laplacian matrix, where p is the order of the statistical moments and K is the order of the convolution filter. This represents the output of the (l-1)th layer. express Perform p element-wise multiplications. variable It is a matrix where each row is the intermediate feature vector for each node i∈N, and for a given value of p, it incorporates features from... The node information of skip neighbors. The output of the l-th layer is obtained by concatenating all the nodes in the skip neighborhood. The result is as follows:

[0096]

[0097] It should be noted that a larger h l The p-value allows for the computation of more detailed and comprehensive node state representations in the final stage and intermediate steps. Similarly, a larger p-value helps each intermediate feature vector better encode the intermediate states. Compared to the scalar embedding of traditional graph convolutions, this structured embedding can more efficiently capture network characteristics, and excessively high h-values... l Both p and will incur additional training costs. The final node embedding is achieved through... The calculation is performed using a linear transformation:

[0098]

[0099] Among them, W F It is a size of The learnable weight matrix.

[0100] Furthermore, graph embedding is achieved by embedding nodes into a matrix F. Nodes The transmission is calculated by averaging through a series of linear layers:

[0101] F graph =Mean(W g2 ×(W g1 ×F Nodes ))

[0102] in, and For ease of explanation, the bias term has been omitted.

[0103] It should also be noted that some key state variables, such as power supply E, are... supp Voltage exceeding limit V viol and branch power flow l E This information cannot be directly mapped to graph node information. Voltage exceeding limits directly affects grid safety and requires stability; branch power flow reflects line status (on / off / fault), and power supply is strongly correlated with network switching behavior. To incorporate this information, a background feature vector is constructed:

[0104] F context =Feedforward(Concat([E supp V viol ,l E ]))

[0105] Furthermore, state embedding F final F graph and Fcontext Add them together and pass them through a multilayer perceptron for computation:

[0106] F final =MLP(F graph +F context )

[0107] It should be noted that, in constructing the emergency recovery target under the fault scenario of the power distribution network, this application embodiment should restore as much of the remaining load as possible while ensuring the power supply to important loads, and at the same time reduce the number of switching operations and network losses.

[0108] In this embodiment of the application, step S300, which involves constructing a distribution network emergency recovery model based on the distribution network emergency recovery decision characteristics, includes: the emergency recovery objective function of the distribution network emergency recovery model under distribution network fault scenarios is expressed as:

[0109] F U = (f1+f2+f3)

[0110]

[0111] Where f1, f2, and f3 represent minimizing network loss, minimizing the number of switching actions, and maximizing load recovery, respectively. l R is the current in branch l. l Let N be the resistance of branch l. z Let N be the set of branches. d For a set of segmented switches, N l For the set of interconnecting switches, x z Indicates the state of the switch, λ h x is the weighting coefficient for load h. h Given the state of load h, P h N is the active power of load h. f This represents the total number of loads to be restored.

[0112] Furthermore, a 1 indicates a closed switch, and a 0 indicates an open switch; a sectionalizing switch is usually in the closed state, therefore when x z When x is open, it means that the sectionalizing switch has been activated once; the tie switch is usually open, therefore when x is closed... z When the switch is closed, it means that the contact switch has been activated once. The sum of these values ​​represents the number of switch operations during the recovery process.

[0113] In this embodiment of the application, the constraints of constructing the distribution network emergency recovery model based on the distribution network emergency recovery decision characteristics in step S300 include power flow constraints, security constraints, and radial operation constraints.

[0114] It should be noted that, based on power flow constraints, security constraints, and radial operation constraints, restrictions are imposed on the emergency recovery of the distribution network from different aspects to ensure the feasibility and rationality of the emergency plan; specifically:

[0115] Current flow constraints are represented as:

[0116]

[0117] Safety constraints are represented as:

[0118]

[0119] Among them, P Gi and Q Gi P represents the active and reactive power injected into node i, respectively. Li and Q Li G represents the active and reactive power injected into node i, respectively. ij and B ij θ represents the conductance and susceptance of the branch between nodes i and j, respectively. ij For voltage U i and U j The phase angle difference, P ij Let be the transmission power of line ij. The maximum power allowed by the line. and These are the upper and lower voltage limits for node n, respectively.

[0120] Radial running constraints are represented as follows:

[0121] g k ∈G k

[0122] Among them, g k For the distribution network structure after power restoration, G k It is a collection of radial structures of the power distribution network.

[0123] In this embodiment of the application, step S400, which transforms the emergency restoration problem of the distribution network into a finite Markov decision process, includes:

[0124] The distribution network system environment is perceived through a feature extraction model, and the emergency recovery decision features of the distribution network are obtained from the node adjacency matrix and the feature matrix as the state space;

[0125] The sequence of switching operations for each step in the fault recovery process is used as the action space.

[0126] When the fault is recovered, the system load recovery rate reaches 100% and all operating constraints are met, a sparse reward value is obtained.

[0127] Given a state, the action-value function is calculated by weighting the cumulative reward of the target time step by a discount factor, and the optimal strategy is found so that performing the corresponding action in any state can maximize the long-term return.

[0128] Furthermore, the state space represents the external environment information of the agent, and the constructed state space needs to accurately describe the irregular environmental information of the power distribution network system. The power distribution network system environment is perceived through the LGACN model. A new feature matrix F is obtained from the node adjacency matrix A and the feature matrix X. final The state space is represented as follows:

[0129] s t =F final

[0130] Furthermore, the action space refers to the actions that the intelligent agent will take after perceiving its external environment. In this embodiment, the specific switching operation sequence of each action in the fault recovery process is represented as follows:

[0131]

[0132] Among them, l i To change the switching state of the i-th line in the system, N l Let N be the number of branches in the system, and Nj be the set of branches that have been operated on in the j-th round, to avoid invalid actions.

[0133] It should be noted that changing the switch state of the i-th line in the system can avoid invalid action selection; if the i-th line in the current system is in an open state, its line switch is closed to reconnect the line; if the i-th line in the current system is in a closed state, its line switch is opened to disconnect the line and exit operation.

[0134] Furthermore, the reward function is the feedback value obtained by the agent after perceiving the external environment and taking action. In this embodiment, the reward is derived from the reward mechanism defined by minimizing the system objective function, power balance constraints, safety constraints, and device operation constraints, and is expressed as:

[0135] r t =-F U -F f -F p,c

[0136]

[0137] Among them, F f Assuming a sparse reward value, after the current action is completed, the system load recovery rate is 100%, and all operational constraints are met, a larger sparse reward value is assigned to strengthen the guidance of the agent's learning direction.

[0138] Furthermore, the penalty component for the action includes voltage over-limit penalty, current over-limit penalty, and distribution network radial topology constraint penalty. The penalty for the c-th action in the h-th round is:

[0139] F p,c =-P V,c -P I,c -P Loop,c

[0140]

[0141] Among them, F p,c P is the penalty portion of the reward function for the c-th action in the current round; V,c P I,c and P Loop,c These are the voltage over-limit penalty, current over-limit penalty, and distribution network radial topology constraint penalty for the c-th operation, respectively. K U K is the penalty value set when a voltage over-limit occurs. I K is the penalty value set when a current over-limit occurs. Loop This is a penalty for the radial topology constraints of the distribution network.

[0142] Furthermore, the action-value function: given state s t Below, energy scheduling a t The quality can be represented by the expected sum of future rewards over K time steps, i.e., the action value function Q. π (s t ,a t To evaluate, it is represented as:

[0143]

[0144] Here, π represents the policy that maps the system state to energy scheduling, and γ∈[0,1] represents the discount rate of future rewards relative to current rewards. The goal of this scheduling problem is to find the optimal policy π. * To obtain the optimal action-value function Q * (s,a):

[0145]

[0146] In this embodiment of the application, the step S400 of calculating and updating the distribution network operating status parameters to obtain the optimal distribution network emergency recovery strategy includes:

[0147] The current Q-network and the target Q-network are responsible for action selection and value evaluation, respectively. The target value is calculated by the target Q-network, and the current Q-network updates its parameters based on the difference between the predicted value and the target value.

[0148] The mean squared error loss function is used to measure the difference between the current Q value and the target Q value. Gradient descent is used to optimize until iterative convergence, and action decisions are output to determine the state of each switch in the distribution network.

[0149] After executing an action, the state at the next moment is obtained, and the corresponding reward value is obtained. The experience is stored in the experience replay buffer. Small batches of experience are randomly sampled from the buffer, and the network parameters are updated using the loss function of the dual deep Q-learning network. The parameters of the target network are adjusted using a soft update method. If the set number of rounds or the average reward converges, the training ends; otherwise, the next training cycle continues until the optimal distribution network emergency recovery strategy is obtained.

[0150] It should be noted that Dual Deep Q-Learning Network (DDQN) is an improvement on the traditional deep Q-network algorithm, aiming to address the overestimation problem in Q-learning. DDQN reduces bias in Q-learning by decoupling action selection from Q-value evaluation. DDQN introduces two independent Q-value functions Q0. A and Q B Each has a different parameter θ. A and θ B The update rules are as follows:

[0151] Q A (s t ,a t )←Q A (s t ,a t )+α[r t +γQ B (s t+1 argmax a′ Q A (s t+1 ,a′;θ B ))-Q A (s t ,a t )]

[0152] Where α is the learning rate, γ is the discount factor, and a′ is the next action.

[0153] Furthermore, in DDQN, action selection and Q-value evaluation are performed separately through two independent neural networks, and the target value of DDQN is... The calculation is as follows:

[0154]

[0155] Furthermore, the current Q-network θ is used to select the next action a′, while the target Q-network θ - To evaluate the value of this action, the loss function L(θ) of DDQN is:

[0156]

[0157] Furthermore, in each cycle of refactoring and updating the parameters of the distribution network, the feature matrix X, which represents the node power information of the distribution network, and the adjacency matrix A, which represents the topological connectivity, are used as inputs to LGACN.

[0158] By leveraging LSTM and GACN to extract knowledge encompassed by the distribution network topology, a new feature matrix F will be output, embedding distribution network node information and connection relationships. final This serves as the environment state space for DDQN.

[0159] The agent introduces random noise Ω t Select the corresponding distribution network restoration strategy and execute output plan a. t Afterwards, the agent receives a corresponding reward value r. t And observe the operating status of the distribution network at the next moment. t+1 .

[0160] The agent determines the network update direction based on the policy gradient through an experience replay mechanism and updates the parameters of the target network using a soft update method. By continuously iterating and optimizing the network parameters, the policy converges when the end condition of the training round is met, and the final output is an operation sequence that can directly guide fault recovery.

[0161] After one cycle ends, the agent determines whether the termination conditions, including the cycle limit and the average reward limit, have been met. Otherwise, the agent moves to the next cycle and repeats the above process.

[0162] In an optional implementation, the number of rounds is limited to 1000, which is suitable for improving the IEEE-33 node system and can balance computational resources and convergence efficiency. The average reward limit is set to meet the requirements of a load recovery rate ≥95% and a reward fluctuation ≤3% over 100 consecutive rounds, ensuring that the strategy remains stable in a high recovery efficiency range without significant performance fluctuations.

[0163] It should be noted that this application overcomes the limitations of traditional methods in independently processing spatial topology and temporal data by coupling dynamic graph convolution and long short-term memory network temporal modeling through spatiotemporal graph attention capsule network. Multi-order graph convolution accurately captures cross-regional power supply path correlations, and the long short-term memory network gating mechanism synchronously updates topology evolution and new energy output fluctuation characteristics, realizing high-density spatiotemporal embedding expression and providing global-local collaborative perception capabilities for decision-making. Through the collaborative design of dual-depth Q network architecture and dynamic action masking mechanism, dual optimization is achieved. The dual-depth Q network decouples action selection and value assessment, and suppresses Q-value overestimation bias through target network soft update, effectively reducing the risk of dangerous operations. Based on the adjacency matrix real-time update and context feature fusion mechanism (power supply, over-limit indicators), it adaptively perceives topological changes caused by switching actions. The strategy generation time is lower than that of traditional mathematical optimization methods, meeting the stringent timeliness requirements of minute-level fault recovery, ensuring the reliability and response speed of the system, enabling rapid power restoration after a fault occurs, and ensuring the normal operation of important loads.

[0164] Example 2: The above example is an illustrative scheme of a distribution network emergency recovery method. It should be noted that the technical solution of this distribution network emergency recovery system belongs to the same concept as the technical solution of the above-described distribution network emergency recovery method. Details not described in detail in this example can be found in the description of the above-described distribution network emergency recovery method.

[0165] This embodiment provides a power distribution network emergency recovery system, comprising:

[0166] The environmental sensing module is used to acquire information about the operating environment of the active distribution network system.

[0167] The feature extraction module is used to build a feature extraction model to perform graph-structured representation of the active distribution network system and extract the emergency recovery decision features of the distribution network.

[0168] The model building module is used to build an emergency recovery model for the distribution network based on the emergency recovery decision characteristics of the distribution network.

[0169] The optimization decision module is used to transform the emergency recovery problem of the distribution network into a finite Markov decision process, calculate and update the operating state parameters of the distribution network, and obtain the optimal emergency recovery strategy for the distribution network.

[0170] This embodiment also provides an electronic device applicable to power distribution network emergency restoration methods, including:

[0171] The system includes a memory and a processor. The memory stores computer-executable instructions, and the processor executes these instructions to implement the power distribution network emergency restoration method proposed in the above embodiments.

[0172] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the distribution network emergency recovery method proposed in the above embodiments.

[0173] The storage medium proposed in this embodiment and the method for implementing emergency restoration of power distribution networks proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0174] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0175] Example 3, referring to Figures 4 to 6 This is one embodiment of the present invention. In this embodiment, an improved IEEE 33-node system is used to simulate a real power distribution network system environment for simulation testing to verify the beneficial effects of the present invention.

[0176] The improved IEEE-33 node topology in this embodiment is as follows: Figure 4 As shown, the improved reference voltage is 12.66kV. Distributed photovoltaic systems are connected at nodes 6, 18, 21, and 24, with rated capacities of 2.5MW, 1.5MW, 1.5MW, and 2MW, respectively; energy storage systems are connected at nodes 27 and 32, each with a rated capacity of 3MW and an adjustable power of [-500, 500]kW. In this embodiment, all simulations were verified in a server environment configured with an AMD Ryzen 9 5900X, an NVIDIA RTX 4080, 1GB of RAM, and the simulation software Matlab2022b.

[0177] Figure 5 The reward value per training round of the proposed algorithm is given, and the training time for 1000 rounds is 6.56 hours. Figure 5It can be seen that the agent adapts to the environment and obtains more rewards through continuous trial and error. In the initial stage, the agent is encouraged to explore more feasible solutions according to a random policy, accumulating certain "experience" through a large number of iterative learning cycles. Therefore, the reward value is low and exhibits relatively obvious oscillations in this stage. Subsequently, based on the experience accumulated in the early stage, the agent makes more comprehensive use of environmental information to obtain the optimal decision solution. In the middle and late stages, the reward value gradually increases, and finally enters the convergence range after 790 rounds, with the reward value stabilizing at approximately -144834.

[0178] Load recovery rate refers to the ratio of current online load to initial total load, with lines identified by their start and end numbers. Assume a fault occurs on lines 3-4, and lines 5-6 are close to the main grid power source, with most tie lines located downstream. Three fault times are defined: 09:00, distributed photovoltaic (PV) output is partially utilized, and the load is in its rising phase; 12:00, distributed PV output reaches its peak; 17:00, the load reaches its maximum for the day, and sunset causes no PV output. Power can only be restored by reconfiguring the network topology through line switches, reconnecting the de-energized load to the main grid power source. The emergency recovery strategies provided by the trained agent for some faults are shown in Table 1.

[0179] Table 1 Results of Emergency Recovery Strategies

[0180]

[0181]

[0182] The method proposed in this invention can flexibly adjust the recovery strategy according to different fault times and power supply conditions. At 9:00, when distributed photovoltaic output is partially restored and load increases, and line 3-4 experiences a fault, power supply is restored by constructing an effective grid topology through reasonable line switching. At 12:00, when photovoltaic output reaches its peak, the strategy is streamlined and efficient. At 17:00, when photovoltaic output is low and load is at its maximum, although the operating conditions are complex, the switching can still be controlled in an orderly manner. When line 5-6 experiences a fault at 18:00 and photovoltaic power is unsupported, the method can also address issues such as voltage exceeding limits and loop networks sequentially. The load recovery rate performs excellently under different scenarios, consistently exceeding 90%. When line 3-4 experiences a fault, the recovery rate is 98.74% at 9:00, 100% at 12:00, and 96.42% at 17:00. This demonstrates that the method of this invention can effectively restore power supply under various operating conditions, ensuring the normal operation of a high proportion of load, and strongly proving its effectiveness and reliability in distribution network fault recovery.

[0183] Figure 6 This illustrates the node voltage distribution after emergency recovery. (By...) Figure 5It can be seen that no node voltage exceeded the limit. After fault recovery, the node voltages were all within the safe operating range of the power grid (0.95-1.05). The voltage deviation decreased by 20.6%. Therefore, the fault recovery method of this invention ensures that the grid voltage remains within the safe range, improving the power supply quality of the power grid.

[0184] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for emergency restoration of a power distribution network, characterized in that, include: Obtain information about the operating environment of the active distribution network system; A feature extraction model is constructed to perform graph-structured representation of the active distribution network system and extract the emergency recovery decision features of the distribution network. A distribution network emergency recovery model is constructed based on the aforementioned distribution network emergency recovery decision characteristics; The emergency recovery problem of the distribution network is transformed into a finite Markov decision process. The operating state parameters of the distribution network are calculated and updated to obtain the optimal emergency recovery strategy for the distribution network.

2. The distribution network emergency restoration method as described in claim 1, characterized in that, The information obtained about the operating environment of the active distribution network system includes: node current, node voltage, load active power, and load reactive power.

3. The distribution network emergency restoration method as described in claim 1 or 2, characterized in that, Constructing a feature extraction model to perform graph-structured representation of the active distribution network system includes: The feature extraction model learns graph topology knowledge by introducing a long short-term memory network-graph attention capsule network to perceive and extract information about irregular environments in active distribution networks. The topology of the distribution network is represented by an adjacency matrix and a feature matrix. The adjacency matrix describes the connection relationship between nodes, and the feature matrix contains the voltage, current, active power and reactive power of each node at different time points. The branches with connection relationships and the nodes with power information in the distribution network system are mapped to the top set and edge set in a graph structure.

4. The distribution network emergency restoration method as described in claim 3, characterized in that, Extracting the decision features for emergency recovery of the power distribution network includes: using a long short-term memory network to model time series data and extracting dynamic features from the time series data; the structure of the long short-term memory network includes a forget gate, an input gate, and an output gate; the input format is a time series vector in the form of a sliding window, which is output as a hidden state by the encoder long short-term memory network, and the hidden state contains the dynamic features of the time series. The features output by the Long Short-Term Memory Network are weighted, the weight of each feature is calculated and normalized, and spatial features are extracted using a graph attention capsule network. Node attributes are mapped to a high-dimensional space, graph capsule layers are used for graph convolution operations, and intermediate features are calculated using graph Laplacian matrix and multinomial filter. The output of each layer is a concatenation of multiple hop count neighborhood information to represent the node state. Node embedding and graph embedding are obtained through linear transformation, and the graph embedding is used to obtain the final distribution network emergency recovery decision features through a multilayer perceptron.

5. The distribution network emergency restoration method as described in claim 4, characterized in that, The distribution network emergency recovery model constructed based on the aforementioned distribution network emergency recovery decision characteristics includes: the emergency recovery objective function of the distribution network emergency recovery model under distribution network fault scenarios is expressed as: F U (f1+f2+f3) Where f1, f2, and f3 represent minimizing network loss, minimizing the number of switching actions, and maximizing load recovery, respectively. l R is the current in branch l. l Let N be the resistance of branch l. z Let N be the set of branches. d For a set of segmented switches, N l For the set of interconnecting switches, x z Indicates the state of the switch, λ h x is the weighting coefficient for load h. h Given the state of load h, P h N is the active power of load h. f This represents the total number of loads to be restored. Constraints include power flow constraints, safety constraints, and radial operation constraints.

6. The distribution network emergency restoration method as described in claim 5, characterized in that, Transforming the emergency recovery problem of the power distribution network into a finite Markov decision process includes: The distribution network system environment is perceived through a feature extraction model, and the emergency recovery decision features of the distribution network are obtained from the node adjacency matrix and the feature matrix as the state space; The sequence of switching operations for each step in the fault recovery process is used as the action space. When the fault is recovered, the system load recovery rate reaches 100% and all operating constraints are met, a sparse reward value is obtained. Given a state, the action-value function is calculated by weighting the cumulative reward of the target time step by a discount factor, and the optimal strategy is found so that performing the corresponding action in any state can maximize the long-term return.

7. The emergency restoration method for power distribution networks as described in claim 6, characterized in that, The calculation and updating of distribution network operating status parameters, and the resulting optimal distribution network emergency recovery strategy, include: The current Q-network and the target Q-network are responsible for action selection and value evaluation, respectively. The target value is calculated by the target Q-network, and the current Q-network updates its parameters based on the difference between the predicted value and the target value. The mean squared error loss function is used to measure the difference between the current Q value and the target Q value. Gradient descent is used to optimize until iterative convergence, and action decisions are output to determine the state of each switch in the distribution network. After executing an action, the state at the next moment is obtained, and the corresponding reward value is obtained. The experience is stored in the experience replay buffer. Small batches of experience are randomly sampled from the buffer, and the network parameters are updated using the loss function of the dual deep Q-learning network. The parameters of the target network are adjusted using a soft update method. If the set number of rounds or the average reward converges, the training ends; otherwise, the next training cycle continues until the optimal distribution network emergency recovery strategy is obtained.

8. A power distribution network emergency recovery system, applied to the method described in any one of claims 1-7, characterized in that, include: The environmental sensing module is used to acquire information about the operating environment of the active distribution network system. The feature extraction module is used to construct a feature extraction model to perform graph-structured representation of the active distribution network system and extract the emergency recovery decision features of the distribution network. The model building module is used to build a distribution network emergency recovery model based on the distribution network emergency recovery decision characteristics. The optimization decision module is used to transform the emergency recovery problem of the distribution network into a finite Markov decision process, calculate and update the operating state parameters of the distribution network, and obtain the optimal emergency recovery strategy for the distribution network.

9. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the power distribution network emergency restoration method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores computer-executable instructions that, when executed by a processor, implement the steps of the power distribution network emergency restoration method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Load recovery method and system for fault area of power distribution network

    CN121097679A

  • Load restoration methods and systems for distribution network fault areas

    CN121097679B

  • Mobile emergency DC power supply intelligent scheduling system

    CN121146574A