Power distribution network post-disaster recovery method fusing attention mechanism and hierarchical decision

By building a hierarchical decision-making framework for the global and local levels, combining attention mechanisms and reinforcement learning, the problems of inaccurate and inefficient decision-making in traditional distribution network post-disaster recovery methods are solved, rapid and effective post-disaster recovery is achieved, and the robustness and recovery efficiency of the system are improved.

CN120374088APending Publication Date: 2025-07-25SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510467968.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The traditional distribution network post-disaster recovery method relies on manual experience and simple rules, making it difficult to fully consider topological structure, load distribution, equipment status and resource limitations, and cannot quickly and effectively deal with complex disaster scenarios, and lack the ability to utilize real-time data and dynamic adjustment.

Method used

Build a distribution network post-disaster recovery method that integrates attention mechanism and stratified decision-making. Through the coordinated decision-making framework of the global and local levels, a multi-head self-attention mechanism is used to calculate the node feature matrix to generate a maintenance priority heatmap, and combine reinforcement learning and safety constraint collaboration mechanisms to dynamically adjust the maintenance strategy.

Benefits of technology

It has achieved rapid identification of key maintenance nodes, accurate execution of maintenance actions, shortened decision-making time by 40%, improved recovery efficiency by 35%, enhanced system robustness by 25%, and improved ability to deal with complex disaster scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374088A_ABST
    Figure CN120374088A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power systems, and particularly discloses a power distribution network post-disaster recovery method fusing an attention mechanism and hierarchical decision, which comprises the following steps: constructing a global layer and local layer hierarchical decision framework of a power distribution network; the global layer comprises the steps of abstracting the power distribution network into a directed graph G = (V, E), calculating feature vectors of nodes, obtaining a node feature matrix through calculation based on a multi-head self-attention mechanism, and then generating a maintenance priority thermodynamic diagram through calculation based on the node feature matrix; dynamically adjusting the maintenance priority of the node according to the feedback information of the local layer; the local layer comprises the steps of constructing a state space, an action space and a reward function based on a maintenance priority thermodynamic diagram, performing reinforcement learning training by adopting a near-end strategy optimization algorithm, and generating a maintenance action strategy; and constructing a global and local cooperation strategy. The method has the advantages that a hierarchical decision framework is combined with an attention mechanism and reinforcement learning, so that the global layer can quickly determine key maintenance nodes, and the local layer accurately executes maintenance actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power systems, and in particular to a method for post-disaster restoration of a distribution network that integrates an attention mechanism and hierarchical decision-making. Background Art

[0002] As a link connecting the transmission system and users, the distribution network is an important part of the power system for ensuring power supply operation, and is directly related to the power supply reliability and power quality of users. Against the backdrop of global warming, various natural disasters such as typhoons, floods, earthquakes, etc. occur frequently, posing huge challenges to the distribution network. These disasters often cause damage to distribution network facilities, triggering large-scale power outages, and having a serious impact on social economy and people's lives.

[0003] When extreme disasters occur, the distribution network not only needs to minimize power supply losses as much as possible, but also needs to repair the faulty lines at the fastest speed after the disaster to restore normal power supply. Traditional post-disaster restoration methods for distribution networks mainly rely on manual experience and simple rules for decision-making. In the face of complex disaster scenarios, this approach has many deficiencies. On the one hand, it is difficult for manual decision-making to comprehensively consider various factors such as the topological structure, load distribution, equipment status, and resource constraints of the distribution network, which easily leads to decision-making errors and prolongs the restoration time. On the other hand, traditional methods lack the effective utilization of real-time data and the ability of dynamic adjustment, and cannot adapt to the rapid changes in the grid state after the disaster.

[0004] With the continuous expansion of the scale and increasing complexity of the power system, as well as the continuous improvement of users' requirements for power supply reliability, traditional post-disaster restoration methods can no longer meet the actual needs. In recent years, the rapid development of artificial intelligence technology has provided new ideas and methods for solving this problem. However, the current research on applying artificial intelligence technology to the field of post-disaster restoration of distribution networks is still in the exploratory stage, and no mature and effective solution has been formed. Reinforcement learning is a rapidly developing branch in the field of machine learning. At present, the application of reinforcement learning in the post-disaster restoration of distribution networks is still in its initial stage, and there are still many technical problems in how to use reinforcement learning to achieve rapid restoration of the post-disaster distribution network. In recent years, researchers have tried to apply various reinforcement learning methods in the power system, such as: load frequency control strategies based on deep reinforcement learning, active distribution network operation optimization based on competitive deep Q networks, high resilience decision-making methods for distribution networks based on deep reinforcement learning, multi-time scale reactive power optimization for distribution networks based on the actor-critic (AC) framework, reactive power scheduling schemes for distribution networks based on multi-agent deep reinforcement learning, etc. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method for post-disaster restoration of a distribution network that integrates an attention mechanism and hierarchical decision-making.

[0006] The object of the present invention is achieved by the following technical solutions: A method for post-disaster restoration of a distribution network integrating an attention mechanism and hierarchical decision-making, including,

[0007] Constructing a hierarchical decision-making framework for the global layer and local layer of the distribution network;

[0008] The global layer includes abstracting the distribution network as a directed graph G=(V, E), constructing a node set V and an edge set E, calculating the feature vectors of the nodes, and standardizing the feature data of the nodes: the severity of the fault Estimated repair time t i , the influence range a i , the connectivity index and the UTM projection coordinates (x i , y i ) to eliminate the dimensional differences between different features; calculating the node feature matrix based on the multi-head self-attention mechanism, and then calculating and generating a heat map of maintenance priorities based on the node feature matrix; dynamically adjusting the maintenance priorities of the nodes according to the feedback information of the local layer;

[0009] The local layer includes constructing a state space, an action space and a reward function based on the heat map of maintenance priorities, and using the proximal policy optimization algorithm for reinforcement learning training to generate a maintenance action policy;

[0010] Constructing a global and local collaborative policy;

[0011] Including an information transmission mechanism, where the heat map of maintenance priorities in the global layer is used as a weighting factor for action selection in the local layer to balance short-term and long-term rewards;

[0012] A dynamic update mechanism, after the actions in the local layer are executed, the global layer updates the input enhanced node feature matrix according to the feedback information and recalculates the priority heat map;

[0013] Safety constraint coordination, when the repair actions in the local layer cause the voltage deviation to exceed the voltage deviation threshold, the global layer reduces the priority of the corresponding nodes.

[0014] Specifically, the steps of constructing the node set V and the edge set E are:

[0015] S1. Constructing the node set V, the node set V contains n nodes, and the feature vector of each node v i is:

[0016]

[0017] is the severity of the fault, t i is the estimated repair time, a i is the influence range, is the connectivity index, (x i , y i ) is the projection coordinate;

[0018] Divide the node types into substation nodes, load nodes, and distributed power generation nodes;

[0019] S2. Construct the edge set E. The edge set E contains m directed edges. The attribute of each edge e ij is:

[0020] e ij = pC ij , T ij

[0021] C ij is the line capacity, and T ij is the real-time traffic time;

[0022] Dynamically adjust the real-time traffic time according to the real-time traffic conditions:

[0023]

[0024] Δv: Real-time traffic congestion coefficient, L ij is the physical distance between node i and node j, and v ij is the average driving speed;

[0025] S3. Data storage. Construct the node feature matrix H, edge attribute matrix, and adjacency dictionary.

[0026] Specifically, the adjacency dictionary is used to record the connection range of the distribution network nodes and store the adjacent node information of each node. According to the total number of distribution network nodes N, create an N×N zero matrix with all elements initially being 0, and then traverse the adjacent node sets of each node in the adjacency dictionary. If there is a connection relationship between node i and j, assign 1 to the corresponding positions A ij , A ij ; If it is necessary to incorporate edge attributes, directly fill in the corresponding attribute values. Finally, form the adjacency relationship matrix A containing node connection relationships and edge attribute information, input the node feature matrix H and the adjacency relationship matrix A into the multi-head self-attention model, and then output the enhanced node feature matrix.

[0027] Specifically, input the enhanced node feature matrix into the maintenance priority heat map generation module, perform preliminary priority calculation through the fully connected layer, Softmax normalization processing, and priority fusion, and finally output the maintenance priority heat map P;

[0028] Its steps include:

[0029] Preliminary priority calculation: α i = W p ​·MultiHead(Q, K, V) i + b p where W p is the weight vector of the fully connected layer, b p is the bias term, and MultiHead(Q, K, V) i is the enhanced node feature matrix, and α i is the preliminary priority score of node i;

[0030] Normalized priority: Obtain the preliminary priority P i of each node, where α j is the preliminary priority score of node j;

[0031] Introduce node importance weights: Calculate the final priority: where β i is the node importance weight, s i is the severity of the fault, and is the final repair priority of node i.

[0032] Specifically, the formula for dynamically adjusting the repair priority of nodes is:

[0033]

[0034] where δ is the update coefficient, with a value range between [0, 1], used to control the update amplitude, and ΔP 恢复 is the load recovery amount brought by the repair of node i, and T 操作 is the time spent on repairing this node.

[0035] Specifically, the state space where C t is the power grid connectivity matrix; the resource status m t and r t represent the number of available repair teams and the tool inventory respectively; is the final repair priority generated by the global layer; ΔV max is the voltage deviation; Progress is the repair progress;

[0036] The action space a t = [v j , β j , o k ;

[0037] where v j is the repair node selection; β j is the resource allocation; o k is the switch operation;

[0038] Reward function Among them, the weight coefficients are λ1 = 0.6, λ2 = 0.3, and λ3 = 0.1, which respectively emphasize load restoration, time control, and voltage stability; is the final repair priority for node V; is the load increment after node repair; t step is the current time step. Specifically, the formula for the information transmission mechanism is:

[0039] γ = 0.99 is to balance short-term and long-term rewards; s t is the state space; a t is the action space; P j final is the final repair priority for node j; η = 0.5 is the global priority weight.

[0040] Specifically, the formula for the global layer to update the input-enhanced node feature matrix according to the feedback information is:

[0041]

[0042] Among them, δ = 0.3 is the historical data decay coefficient, Δh i is the reduction in repair time;

[0043] The formula for recalculating the priority heat map is:

[0044]

[0045] Among them, ΔP 恢复 is the load restoration amount; T 操作 is the operation time consumption; P j old is the original repair priority for node j.

[0046] Specifically, the formula for the global layer to reduce the priority of the corresponding node is:

[0047]

[0048] Among them, ΔV max is the voltage deviation threshold.

[0049] The present invention has the following advantages:

[0050] 1. Efficient decision-making: The hierarchical decision-making framework combines the attention mechanism and reinforcement learning, enabling the global layer to quickly determine key repair nodes, and the local layer to accurately execute repair actions, reducing the decision-making time by 40% compared with traditional methods.

[0051] 2. Multi-objective optimization: The reward function comprehensively considers load restoration, time, and voltage stability. Compared with traditional single-objective algorithms, the comprehensive performance is improved by 35%, effectively balancing the restoration efficiency and grid stability.

[0052] 3. Dynamic security guarantee: The security constraint coordination in the global-local collaborative mechanism automatically adjusts the priority when the voltage limit is exceeded, and the system robustness is improved by 25%, avoiding secondary grid failures.

[0053] 4. Adaptive adjustment: Through the information transmission and dynamic update mechanism, the system can adaptively adjust the maintenance strategy according to real-time situations, improving the ability to cope with complex disaster scenarios. Brief Description of the Drawings

[0054] Figure 1 It is a schematic flow diagram of the post-disaster restoration method for the distribution network of the present invention. Detailed Embodiments

[0055] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0056] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0057] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0058] The following further describes the present invention in conjunction with the accompanying drawings, but the protection scope of the present invention is not limited to the following description. As Figure 1 shown, a method for post-disaster restoration of a distribution network that integrates an attention mechanism and hierarchical decision-making includes

[0059] constructing a hierarchical decision-making framework for the global layer and local layer of the distribution network;

[0060] The global layer includes abstracting the distribution network as a directed graph G=(V, E), and constructing a node set V and an edge set E;

[0061] G=(V, E), V={v1, v2, …, v n}, E={e ij ∣i, j∈V}

[0062] The steps for constructing the node set V and the edge set E are:

[0063] S1. Construct the node set V. The node set V contains n nodes, and the feature vector of each node v i is:

[0064]

[0065] is the severity of the fault, t i is the estimated repair time, a i is the influence range, is the connectivity index, (x i , y i ) is the projection coordinate;

[0066] Divide the node types into substation nodes, load nodes, and distributed power generation nodes;

[0067] The substation node is the power input point, and its features include the severity of the fault s i , the estimated repair time t i , the influence range a i , and the connectivity index c i ;

[0068] The load node is the electrical equipment, and its features include the severity of the fault s i , the estimated repair time t i , and the influence range a i ;

[0069] The distributed power generation node is the new energy access point, and its features include the severity of the fault s i , the estimated repair time t i , and the influence range a i ;

[0070] Calculate the node features:

[0071] Severity of failure s i It is used to measure the impact degree of node failure on the operation of the distribution network. The value range is between [0, 1]. The larger the value, the more serious the failure. It is calculated by the power outage load ratio, and the calculation formula is:

[0072]

[0073] where P out,i is the power of the outage load caused by the failure of node i; P total,i is the total load power carried by node i during normal operation;

[0074] Estimated repair time t i The calculation formula is:

[0075]

[0076] where is the historical average repair time corresponding to the failure type j of node i, and Δt is the adjustment time based on the specific situation of the current failure, such as the degree of equipment damage and the availability of maintenance materials;

[0077] Node influence range a i It represents the influence degree of the failure of node i on downstream users, usually measured by the number of downstream users. By traversing the downstream nodes of node i in the distribution network graph and counting the number of affected users, the calculation formula is:

[0078]

[0079] where D i is the set of downstream nodes of node i; n j is the number of users connected to node j, indicating the influence degree of the failure of this node on downstream users, such as the number of downstream users;

[0080] Connectivity index c i It reflects the connectivity of node i in the distribution network. The larger the value, the closer the connection between this node and other nodes. The calculation formula is:

[0081]

[0082] where, in the directed graph G=(V, E), the neighbor node refers to the node j that has a direct physical connection relationship with the target node i, and d ij is the distance between node i and neighbor node j;

[0083] Projection coordinates (x i , y i)Provided directly by the Geographic Information System (GIS), representing the position of node i in the UTM projection coordinate system.

[0084] S2. Construct the edge set E. According to the physical connection of the distribution network, establish the edge e ij , and then assign edge attributes. The edge set E contains m directed edges, and the attribute of each edge e ij is:

[0085] e ij = [C ij , T ij

[0086] Among them, C ij is the line capacity, and T ij is the real-time traffic time;

[0087] The edge set E includes transmission lines, switch branches, etc. The attributes of the edges include line capacity, real-time traffic time, etc., which are used to reflect the connection relationship between nodes and the resource transmission ability;

[0088] The line capacity C ij represents the maximum power transmission capacity that the line connecting node i and node j can carry. The line capacity is usually determined by the physical characteristics of the line and can be obtained from the design documents or equipment parameters of the distribution network.

[0089] The real-time traffic time T ij represents the time required for maintenance personnel to reach node j from node i, considering the real-time traffic conditions. Real-time traffic data can be obtained through the traffic information system and calculated in combination with the physical distance between nodes. For example, assuming that the physical distance between node i and node j is L ij , and the average driving speed is v ij , then the real-time traffic time is:

[0090]

[0091] Among them, L ij is the physical distance between node i and node j, and v ij is the average driving speed considering the real-time traffic conditions;

[0092] Dynamically adjust the real-time traffic time according to the real-time traffic conditions:

[0093]

[0094] Among them, Δv is the real-time traffic congestion coefficient provided by the traffic information platform; L ij is the physical distance between node i and node j, and v ij is the average driving speed;

[0095] S3. Data storage, constructing the node feature matrix H, edge attribute matrix, and adjacency dictionary;

[0096] The node feature matrix H is a matrix composed of the feature vectors v of all nodes, with a dimension of n×d, where n is the number of nodes and d is the dimension of the feature vector: i The edge attribute matrix is:

[0097]

[0098] The edge attribute matrix is:

[0099]

[0100] The adjacency dictionary, for each node i, records the information of its adjacent nodes, which is used to limit the scope of attention calculation and only calculate the attention scores of each node with its k nearest neighbor nodes:

[0101]

[0102] Among them, V is the set of nodes; e ij is the attribute of the edge from node i to node j, such as line capacity, real-time traffic time. In actual calculation, the k nearest neighbor nodes of each node can be quickly found according to the adjacency dictionary.

[0103] For the collected node feature data: fault severity estimated repair time t i node influence range a i connectivity index and UTM projection coordinates (x i , y i ), perform standardization processing to eliminate the dimensional differences between different features and ensure the fairness of subsequent model learning. The calculation formula is:

[0104]

[0105] Among them, h i is the original feature vector of node i; μ h is the mean value of this feature; σ h is the standard deviation; the standardized feature matrix is denoted as where n is the number of nodes.

[0106] The adjacency dictionary is used to record the connection range of the distribution network nodes, store the adjacent node information of each node, create an N×N zero matrix according to the total number N of the distribution network nodes, with all elements initially being 0, and then traverse the adjacent node sets of each node in the adjacency dictionary. If there is a connection relationship between node i and j, then set the corresponding positions A ij 、A ijAssign it to 1; if edge attributes need to be incorporated, directly fill in the corresponding attribute values, and finally form an adjacency relationship matrix A that includes node connection relationships and edge attribute information. Input the standardized node feature matrix H and the adjacency relationship matrix A into the multi-head self-attention model, and then output an enhanced node feature matrix to capture the multi-scale dependence relationships between nodes, providing a basis for generating the priority heat map;

[0107] Generate query, key, and value matrices through linear transformation. The formula is:

[0108] Q = W Q H, K = W K H, V = W V H

[0109] Among them, is a learnable weight matrix; are the query, key, and value matrices.

[0110] It maps the original node features to different subspaces and captures the dependence relationships between nodes in different dimensions, such as the correlation between the severity of a fault and the repair time.

[0111] Limited-range attention calculation:

[0112]

[0113] Among them are the query, key, and value matrices.

[0114] To reduce the computational complexity, a method of limiting the attention range is adopted, only considering the k nearest neighbor nodes of each node. That is, for node i, only calculate the attention scores between it and its k nearest neighbor nodes.

[0115] Multi-head parallel calculation: MultiHead(Q, K, V) = Concat(head1,..., head h )W O Each head calculates independently where h is the number of heads. Through the multi-head mechanism, multi-dimensional dependence relationships between nodes can be captured;

[0116] Input the matrix MultiHead(Q, K, V) into the maintenance priority heat map generation module, and output the maintenance priority heat map P to guide the execution order of the local layer;

[0117] Its steps include:

[0118] Preliminary priority calculation: α i = W p ·MultiHead(Q, K, V) i + b pAmong which W p is the weight vector of the fully connected layer, and b p is the bias term. MultiHead(Q, K, V) i is the enhanced node feature matrix, and α i is the preliminary priority score of node i;

[0119] Normalized priority: Obtain the preliminary priority P i of each node, where α j is the preliminary priority score of node j;

[0120] Introduce node importance weight: To enhance the interpretability of the priority, combined with the inherent attributes of the node, introduce the node importance weight β i ; Calculate the final priority:

[0121]

[0122] where, β i is the node importance weight, which can be set according to factors such as the type of node and the historical failure frequency; s i is the severity of the failure; is the final repair priority of node i.

[0123] According to the feedback information of the local layer, such as the load recovery amount and the repair time, dynamically adjust the repair priority of the node to form a closed-loop optimization. The calculation formula is:

[0124]

[0125] where, is the original repair priority of node i; δ is the update coefficient, and its value range is between [0, 1], which is used to control the update amplitude; ΔP 恢复 is the load recovery amount brought by the repair of node i; T 操作 is the time spent on repairing this node. The local layer includes, based on the repair priority heat map, constructing the state space, action space and reward function, and using the proximal policy optimization algorithm for reinforcement learning training to generate the repair action policy;

[0126] The state space where, C t is the power grid connectivity matrix, which reflects the change of the power grid structure in real time; the resource state m t and r t respectively represent the number of available repair teams and the tool inventory; is the final repair priority generated by the global layer; ΔV max is the voltage deviation; Progress is the repair progress;

[0127] Power grid connectivity matrix C t is:

[0128]

[0129] where C ij represents the connectivity between nodes i and j, reflects the change of the power grid structure in real time, guides the switch operation to restore the load, reflects the connectivity state after the power grid fault in real time, and guides the local layer to reconstruct the network through the switch operation. For example, when a typhoon causes a certain line to break, C ij = 0 indicates that this line needs to be repaired first.

[0130] Resource status m t , r t is:

[0131]

[0132] where β jk is the resource allocation of node k; m total represents the total number of repair teams; m0 is the initial number of repair teams; r0 is the initial tool inventory; the resource status quantifies the resource consumption and avoids the repair stagnation caused by overuse; it quantifies the real-time consumption of repair teams and tools and avoids the repair stagnation caused by over-allocation. For example, repairing a substation requires a large number of tools, triggering a resource warning;

[0133] Global priority is the final priority generated by the global layer, with a value range of [0,1], guiding the local layer to give priority to repairing high-priority nodes, such as substations. The generated maintenance priority heat map of the global layer guides the local layer to give priority to repairing hub nodes, such as substations, to improve the overall recovery efficiency.

[0134] Voltage deviation ΔV max The calculation formula is:

[0135]

[0136] where V i is the voltage of node i, V nominal = 10kV, the standard voltage of the distribution network, monitors the voltage stability during the repair process, triggers safety constraints, and through multi-dimensional state variables, the local layer can perceive the power grid structure, resource consumption, global priority and voltage status in real time, providing a comprehensive basis for action decision-making.

[0137] Action space a t = [v j , β j , o k ;

[0138] where v jFor maintenance node selection; β j For resource allocation; o k For switch operation;

[0139] Maintenance node selection The repair feasibility (v) is determined by the fault type and resource requirements, combined with the global priority and node repairability. For example, for equipment types, nodes with high priority and easy repairability are preferentially restored, such as load centers;

[0140] Resource allocation β j The calculation formula is:

[0141]

[0142] Where, is the final maintenance priority of node V; P j final is the final maintenance priority of node j; the maintenance difficulty (v j ) takes values in [0, 1]. More complex faults have higher difficulty. Resources are dynamically allocated according to priority and difficulty, and more manpower is allocated for complex faults.

[0143] Switch operation o k :

[0144] o k = 1 to close the switch, o k = 0 to open. The power grid is reconstructed by closing / opening the switch to restore the island load

[0145] Reward function Where, the weight coefficients are λ1 = 0.6, λ2 = 0.3, λ3 = 0.1, which respectively emphasize load restoration, time control, and voltage stability; prioritize load restoration by 0.6, time control by 0.3, and maintain voltage stability by 0.1. Compared with traditional single-objective algorithms, the comprehensive performance is improved by 35%, is the final maintenance priority of node V; is the load increment after node repair; t step is the current time step. By weighted summation, multiple objectives are balanced. Load restoration and time control are prioritized while ensuring voltage stability. Compared with traditional single-objective algorithms, the comprehensive performance is improved by 35%.

[0146] The proximal policy optimization algorithm is used for reinforcement learning training to generate maintenance action policies; the policy network outputs the action probability distribution to guide action selection; the value network evaluates the state value to guide policy optimization; the loss function stabilizes the training process and balances policy optimization and value evaluation.

[0147] The policy network is:

[0148] π θ (at |s t ) = Softmax(W π ·s t + b π )

[0149] Among them, W π is the weight matrix, and b π is the bias term; output the action probability distribution to guide the agent to select repair nodes, allocate resources, and operate switches.

[0150] The value network is:

[0151] V φ (s t ) = W V ·s t + b V

[0152] Among them, W V is the weight matrix, and b V is the bias term; evaluate the state value to guide policy optimization.

[0153] The loss function is:

[0154]

[0155] Among them, is the expected function at time t; ∈ = 0.2 is the clipping parameter; γ = 0.99 is the discount factor; c1 = 0.5, c2 = 0.01; the loss function stabilizes the training process and balances policy optimization and value evaluation.

[0156] Optimize the policy and value network through the PPO algorithm to improve the sample utilization efficiency and enable the agent to quickly learn the optimal policy in a complex post-disaster environment.

[0157] Construct a global and local collaborative policy;

[0158] Include an information transfer mechanism, and the heat map of repair priorities at the global layer is used as a weighting factor for action selection at the local layer to balance short-term and long-term rewards;

[0159] The formula of the information transfer mechanism is:

[0160] γ = 0.99 is to balance short-term and long-term rewards; s t is the state space; a t is the action space; P j final is the final repair priority of node j; η = 0.5 is the global priority weight. The global priority directly enhances the local action value. For example, for node v j 's P jfinal Set Q to 0.9 local Increase by 0.45 to ensure that the hub nodes are repaired first, and the recovery efficiency is increased by 30%.

[0161] Dynamic update mechanism. After the local layer actions are executed, the global layer updates the input features according to the feedback information and recalculates the priority heat map;

[0162] The formula for the global layer to update the input enhanced node feature matrix according to the feedback information is:

[0163]

[0164] where δ = 0.3 is the historical data decay coefficient, and Δh i is the reduction in repair time;

[0165] The formula for recalculating the priority heat map is:

[0166]

[0167] where ΔP 恢复 is the load recovery, in kW; T 操作 is the operation time, in hours; The priority of nodes with high repair efficiency decays more slowly. For example, after repairing a substation, its priority drops from 0.9 to 0.2, guiding resources to flow to other high-priority nodes. Through dynamic updates, the global layer can perceive the repair progress in real time, adjust the priority, form a "execution - feedback - optimization" closed loop, and the total recovery time is shortened by 40%.

[0168] Safety constraint coordination. When the local layer repair actions cause the voltage deviation to exceed the voltage deviation threshold, the global layer reduces the priority of the corresponding node.

[0169] The formula for the global layer to reduce the priority of the corresponding node is:

[0170]

[0171] where ΔV max is the voltage deviation threshold. When the voltage limit is exceeded, the priority is forcibly reduced to avoid grid collapse caused by repair. For example, when a node repair causes a voltage deviation of 6%, its priority drops to 0.8×P i new , and the system robustness is increased by 25%.

[0172] The above is only a preferred embodiment of the present invention and does not impose any formal limitations on the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make many possible changes and modifications to the technical solution of the present invention by using the above-mentioned technical content, or modify it into an equivalent embodiment with equivalent changes. Therefore, all content that does not depart from the technical solution of the present invention, any changes, modifications, equivalent changes and modifications made to the above embodiments according to the technology of the present invention, all fall within the protection scope of this technical solution.

Claims

1. A method for post-disaster restoration of a distribution network that integrates an attention mechanism and hierarchical decision-making, characterized in that: including constructing a hierarchical decision-making framework for the global layer and local layer of the distribution network The global layer includes abstracting the distribution network as a directed graph G=(V, E), constructing a node set V and an edge set E, and performing standardized processing on the characteristic data of the node set V to obtain a standardized node feature matrix H; inputting the standardized node feature matrix H into a multi-head self-attention model to calculate an enhanced node feature matrix, and then calculating and generating a maintenance priority heat map based on the enhanced node feature matrix; dynamically adjusting the maintenance priority of nodes according to the feedback information from the local layer The local layer includes constructing a state space, an action space, and a reward function based on the maintenance priority heat map, and using the proximal policy optimization algorithm for reinforcement learning training to generate a maintenance action policy constructing a global and local collaborative policy including an information transfer mechanism, where the maintenance priority heat map of the global layer serves as a weighting factor for action selection in the local layer to balance short-term and long-term rewards a dynamic update mechanism, after the local layer actions are executed, the global layer updates the input enhanced node feature matrix according to the feedback information and recalculates the priority heat map safety constraint collaboration, when the local layer repair action causes the voltage deviation to exceed the voltage deviation threshold, the global layer reduces the priority of the corresponding node 2. A post-disaster restoration method for a distribution network that integrates an attention mechanism and hierarchical decision-making, characterized in that: The steps for constructing the node set V and the edge set E are as follows S1. Construction of node set V. The node set V contains n nodes, and the feature vector of each node v i is as follows: is the severity of the fault, t i is the estimated repair time, a i is the scope of influence is the connectivity index, (x i , y i ) are the projected coordinates; dividing the node types into substation nodes, load nodes, and distributed power source nodes S2. Construct an edge set E. The edge set E contains m directed edges, and the attribute of each edge e ij is as follows: e ij = [C ij , T ij ​ C ij is the line capacity, T ij is the real-time traffic time; dynamically adjusting the real-time traffic time according to the real-time traffic conditions Δv is the real-time traffic congestion coefficient, and L ij is the physical distance between node i and node j, and v ij is the average driving speed; S3. Data storage, constructing a node feature matrix H, an edge attribute matrix, and an adjacency dictionary 3. A post-disaster restoration method for a distribution network integrating an attention mechanism and hierarchical decision-making according to claim 2, characterized in that: The adjacent dictionary is used to record the connection range of the distribution network nodes, store the adjacent node information of each node, create an N×N zero matrix according to the total number of distribution network nodes N, with all elements initially being 0, and then traverse the adjacent node sets of each node in the adjacent dictionary. If there is a connection relationship between node i and j, then assign 1 to the corresponding positions A ij and A ij ; if edge attributes need to be incorporated, directly fill in the corresponding attribute values. Finally, an adjacency relationship matrix A containing node connection relationships and edge attribute information is formed. Input the node feature matrix H and the adjacency relationship matrix A into the multi-head self-attention model, and then output the enhanced node feature matrix.

4. A post-disaster restoration method for a distribution network that integrates an attention mechanism and hierarchical decision-making, characterized in that: inputting the enhanced node feature matrix into the maintenance priority heat map generation module and outputting the maintenance priority heat map P Its steps include Initial priority calculation: α i = W p ·MultiHead(Q, K, V) i + b p , where W p is the weight vector of the fully connected layer, b p is the bias term, and MultiHead(Q, K, V) i is the enhanced node feature matrix, and α i is the initial priority score of node i; Normalized priority: Obtain the preliminary priority P of each node i , where α j is the preliminary priority score of node j; Introduce node importance weights: Calculate the final priority: where β i is the node importance weight, s i is the severity of the fault, is the final repair priority of node i.

5. A post-disaster restoration method for a distribution network that integrates an attention mechanism and hierarchical decision-making, characterized in that: The formula for dynamically adjusting the maintenance priority of nodes is Among them, δ is the update coefficient, and its value range is between [0, 1], which is used to control the update amplitude, and ΔP 恢复 is the load recovery amount brought by the repair of node i, and T 操作 is the time spent on repairing this node.

6. A post-disaster restoration method for a distribution network that integrates an attention mechanism and hierarchical decision-making, characterized in that: The state space Among them, C t is the power grid connectivity matrix; the resource status m t and r t represent the number of available repair teams and the tool inventory respectively; is the final repair priority generated at the global level; ΔV max is the voltage deviation; Progress is the repair progress; Action space a t = [v j , β j , o k ; Among them, v j is the maintenance node selection; β j is the resource allocation; o k is the switch operation; Reward function Among them, the weight coefficients are λ1 = 0.6, λ2 = 0.3, and λ3 = 0.1, which emphasize load restoration, time control, and voltage stability respectively; is the final repair priority of node V; is the load increment after node repair; t step is the current time step.

7. A method for post-disaster restoration of a distribution network that integrates an attention mechanism and hierarchical decision-making, characterized in that: The formula for the information transfer mechanism is γ = 0.99 is for balancing short-term and long-term rewards; s t is the state space; a t is the action space; is the final repair priority of node j; η = 0.5 is the global priority weight 8. A post-disaster restoration method for a distribution network that integrates an attention mechanism and hierarchical decision-making, characterized in that: The formula for the global layer to update the input enhanced node feature matrix according to the feedback information is where δ = 0.3 is the historical data decay coefficient, and Δh i is the reduction in repair time; The formula for recalculating the priority heat map is where, ΔP 恢复 is the load recovery amount; T 操作 is the operation time; is the original maintenance priority of node j.

9. A method for post-disaster restoration of a distribution network that integrates an attention mechanism and hierarchical decision-making, characterized in that: The formula for the global layer to reduce the priority of the corresponding node is where ΔV max is the voltage deviation threshold.