A method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning.
By constructing a dynamic model based on graph neural network reinforcement learning, a scheduling strategy for mobile emergency energy storage vehicles is generated. This solves the problem of improving the resilience of the power distribution network caused by the uncertainty of fault repair time under extreme weather conditions, and realizes the optimized scheduling of mobile emergency energy storage vehicles, thereby improving the resilience and post-disaster recovery capability of the power distribution network.
Patent Information
- Application Number
- CN202310064061.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-13
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2043-01-13
AI Technical Summary
Most existing studies assume that the repair time for faults in transportation networks and power systems is a fixed value, which fails to effectively account for the uncertainty of repair time under extreme weather conditions. This makes it difficult to formulate effective mobile emergency energy storage vehicle dispatch strategies to improve the resilience of the distribution network.
A dynamic model of the transportation network and power system is constructed using a graph neural network reinforcement learning method to generate a scheduling strategy for mobile emergency energy storage vehicles. The behavior strategy of the emergency vehicles is generated by the ε-Greedy algorithm and the graph neural network reinforcement learning algorithm. The reward function is calculated by combining the distribution network reconfiguration and optimal power flow model to achieve the optimal scheduling of mobile emergency energy storage vehicles.
It effectively solves the problem of uncertainty in fault repair time. By optimizing the scheduling of mobile emergency energy storage vehicles, it enhances the resilience of the power distribution network, reduces the risk of power outages caused by extreme disasters, and accelerates the recovery of critical loads.
Smart Images

Figure CN116151562B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power distribution network resilience enhancement technology, specifically involving a method for dispatching mobile emergency vehicles and enhancing power distribution network resilience based on graph neural network reinforcement learning. Background Technology
[0002] In recent years, with global climate change, extreme weather disasters have become increasingly frequent, causing enormous economic losses and social impacts. As a crucial component of the power system and directly connected to users, the distribution network's ability to withstand disasters and accidents, effectively ensuring power supply for people's production and daily life and reducing the economic and social losses caused by power outages, has received widespread attention. Damage caused by extreme weather is often Nk faults, and power grids operating on a reliability-based basis are powerless under such severe incidents. Therefore, it is necessary to study strategies for enhancing the resilience of urban distribution networks in response to extreme disasters such as typhoons.
[0003] Currently, with the increasing prevalence of electrified transportation, the mobile energy storage flowing through transportation networks is constantly expanding, forming a large-scale, diverse, and topologically flexible discrete energy network, providing a material and energy foundation for enhancing the resilience of the power distribution network. Dedicated mobile emergency energy storage vehicles are the most widely used emergency power supply resource for the power grid. The power supply capacity of an emergency energy storage vehicle depends on its onboard battery, typically ranging from 200 to 500 kWh. Flexible deployment of emergency power supply vehicles can optimize disaster prevention measures, reduce the risk of power outages caused by extreme disasters, and accelerate the recovery of critical loads after disasters, significantly improving the resilience of the power distribution network. However, most existing studies assume that the repair time for potential transportation network and power system line faults is a fixed time, and some even believe that these faults exist continuously within the study period. In reality, due to extreme weather changes, varying degrees of faults in transportation networks and power system lines, and differences in the professional skills of emergency repair personnel, the fault repair time is not a fixed time, but rather follows a normal distribution with a mean and variance over a certain period. Therefore, how to formulate distribution network recovery strategies by dispatching mobile emergency energy storage vehicles to improve the resilience of distribution networks, taking into account the uncertainty of repair time for potential transportation network and power system line faults, is an urgent problem to be solved. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning. This method can formulate power distribution network recovery strategies by dispatching mobile emergency energy storage vehicles under the uncertainty of repair time for potential traffic network and power system line faults, thereby improving the resilience of power distribution networks.
[0005] The main steps are as follows: Step S1: Initialize the distribution network-transportation network model; Step S2: Randomly restore the faults of the transportation network and power system lines; Step S3: Construct the state of the graph neural network reinforcement learning algorithm based on the information of the distribution network and transportation system; Step S4: Generate the scheduling behavior strategy of the mobile emergency energy storage vehicle based on the ε-Greedy algorithm and the graph neural network reinforcement learning algorithm; Step S5: Execute the scheduling strategy of the mobile emergency energy storage vehicle and judge and update the state of the mobile energy storage vehicle; Step S6: Calculate the distribution network reconstruction strategy and calculate the reward function of the mobile emergency energy storage vehicle based on the reconstruction and optimization of the distribution network; Step S7: Update the state of the graph neural network reinforcement learning algorithm; Step S8: Store the information of the current step in the memory unit; Step S9: Determine whether the predetermined time has been reached; if not, execute (2) to (8); if yes, output the graph neural network reinforcement learning algorithm parameters and the corresponding optimized scheduling results.
[0006] The technical solution adopted by this invention to solve its technical problem is:
[0007] A method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning, characterized by the following steps:
[0008] Step S1: Initialize the power distribution network-transportation network model;
[0009] Step S2: Random recovery of faults in transportation network and power system lines;
[0010] Step S3: Based on the information from the power distribution network and transportation system, construct the state x for the graph neural network reinforcement learning algorithm. i,t ;
[0011] Step S4: Generate a dispatching behavior strategy for the mobile emergency energy storage vehicle based on the ε-Greedy algorithm and the graph neural network reinforcement learning algorithm. i,t ;
[0012] Step S5: Execute the dispatch strategy for the mobile emergency energy storage vehicle. i,t It also judges and updates the status of the mobile energy storage vehicle;
[0013] Step S6: Calculate the distribution network reconfiguration strategy, and calculate the reward function r for the mobile emergency energy storage vehicle based on the reconfiguration and optimization of the distribution network. i,t ;
[0014] Step S7: The state x' of the graph neural network reinforcement learning algorithm i,t renew;
[0015] Step S8: Transfer the information from the current step (x) i,t ,a i,t ,r i,t ,x'i,t The weights are stored in memory unit D and updated using the stochastic gradient descent method for the graph neural network reinforcement learning algorithm.
[0016] Step S9: Determine whether the predetermined time T has been reached. end If not, proceed to steps S2 to S8; if yes, output the graph neural network reinforcement learning algorithm parameters and the corresponding optimization scheduling results.
[0017] Furthermore, step S1 specifically includes the following steps:
[0018] Step S11: Initialize the distribution network model, including at least the following: upper and lower limit voltages of distribution network system nodes, line and transformer parameters, upper and lower constraints on voltage and output of distributed generators, output of renewable energy, distribution network load rate, charging station capacity and location, location of faulty lines and estimated recovery time.
[0019] Step S12: Initialize the traffic network model, including at least: traffic network nodes, road length, road capacity and free-flow speed, the location of charging stations in the traffic network, the location of faulty traffic nodes and lines and the estimated recovery time.
[0020] Step S13: Initialize neural network parameters, including at least: initialization of neural network weights W and bias term b, and initialization of hyperparameters such as learning rate α, discount factor γ, batch size B and memory unit capacity D;
[0021] Step S14: Initialize time t=0.
[0022] Furthermore, in step S2, we consider that extreme events can cause weak lines in the power distribution network to break, which can also damage roads. As the weather improves and emergency personnel carry out repairs, the faulty power distribution lines and damaged roads in the transportation network will gradually be repaired. However, the repair time exhibits a certain degree of randomness due to differences in the extent of damage and the varying levels of expertise among emergency repair personnel. Therefore, we assume that the recovery time of faults in the transportation network and power distribution lines follows a certain normal distribution. in Indicates faulty component L i Average repair time; Indicates faulty component L i Time variance of repair.
[0023] Further, in step S3, the state x of the graph neural network reinforcement learning algorithm i,t Including the state of the i-th electric vehicle at time t (EV) i,t Neighboring electric vehicle status Ne i,t Line status That is, whether the line is disconnected and the power load demand. and renewable energy output Right now:
[0024]
[0025] Wherein, formula (2) represents the state EV of the i-th electric vehicle. i,t Including the next node when an electric vehicle travels to a charging station Road number Electric vehicle speed v i,t and the remaining SOC of the mobile energy storage vehicle i,t ;
[0026] Formula (3) represents the neighboring electric vehicle state Ne. i,t This includes the states of each neighboring electric vehicle k, such as the next node of the kth electric vehicle that is adjacent to the i-th electric vehicle. Then the road number it is located on Electric vehicle speed v i,k,t and remaining power SOC i,k,t .
[0027] Further, step S4 includes the following steps:
[0028] Step S41: Generate the scheduling behavior strategy of the mobile energy storage vehicle in a random manner with probability ε, i.e.
[0029] a i,t = np.random.randint(||A action ||) (4)
[0030] In the formula, ||A action || represents the number of actions of the mobile emergency energy storage vehicle; when a i,t =0 indicates that the mobile emergency energy storage vehicle is in a charging or discharging state at this location; when a i,t If ≠0, it means the mobile emergency energy storage vehicle follows dispatch behavior strategy a. i,t The vehicle is driven to the next destination for charging or discharging, thereby reducing the load on the distribution network and improving its resilience.
[0031] Step S42: Generate a dispatching behavior strategy a for the mobile emergency energy storage vehicle with a probability of 1-ε based on the experience of the graph neural network reinforcement learning algorithm. i,t ,Right now:
[0032]
[0033] In the formula, argmax(·) represents the parameter corresponding to the maximum value; Q(x i,t ,a;θ t) indicates that the i-th mobile emergency energy storage vehicle is in state x i,t θ is the adjacency matrix A and the action-value function under scheduling behavior policy a. t This represents the neural network parameters of the graph neural network reinforcement learning algorithm at time t;
[0034] The neural network structure of the graph neural network reinforcement learning algorithm includes a single input layer, whose input includes the state set x of all mobile emergency energy storage vehicles. t ={x 1,t ,x 2,t ,...,x N,t} and the relation matrix consisting of mobile emergency energy storage vehicles, i.e., the adjacency matrix A; and a fully connected layer is used to process the input state x. t Perform feature extraction x' t Then, a two-layer graph attention neural network is connected to process the proposed feature x'. t The adjacency matrix A is processed; finally, a fully connected layer is connected to output the mobile emergency energy storage vehicle in state x. t The action value functions corresponding to the adjacency matrix A are obtained below; the mobile emergency energy storage vehicle agent selects the scheduling behavior strategy based on the action value functions.
[0035] Furthermore, in step S6, the reward function r of the mobile emergency energy storage vehicle is calculated based on the distribution network reconfiguration and optimal power flow optimization. i,t Specifically, the following steps are included:
[0036] Step S61: Update the distribution network load factor and calculate the charging requirements of the mobile energy storage vehicle. and maximum discharge power
[0037] Step S62: Establish distribution network reconfiguration and optimal power flow models:
[0038]
[0039] -Mα mn,t ≤V mn,t ≤Mα mn,t (9)
[0040]
[0041]
[0042] In the formula, S b S represents the set of all nodes in a power distribution network; r This indicates that the first and last nodes of the connected substations, distributed power generation nodes, and faulty lines constitute the potential root node set; N b α represents the number of nodes in the distribution network. mn,tThis represents the line state at time t, and is a binary variable; 1 indicates a closed line, and 0 indicates an open line. n,t This is a binary variable; if the potential root node n is the root node of the distribution network during time period t, its value is 1; otherwise, it is 0. M is a sufficiently large constant. V mn,t The virtual current flows through branch mn at time t;
[0043] Equation (6) indicates that the distribution network established during the network reconfiguration process must satisfy a radial topology, that is, the number of closed lines is equal to the number of network nodes minus the number of subgraphs.
[0044] Equations (7) to (9) are the constraints for the reconfiguration of the distribution network, which require that each sub-graph is internally connected; pf represents the set of faulty lines at time t. mn,t and qf mn,t This represents the active and reactive power flow of the distribution network branch mn at time t; pd m,t and qd m,t pg represents the restored active and reactive loads of the distribution network branch at node m at time t; m,t and qg m,t This represents the active and reactive power output of the distributed power source m. and This represents the active and reactive power of the mobile emergency energy storage vehicle at the m-th node. and This represents the active and reactive power outputs of the wind turbine connected to the m-th node; and These represent the original active power load demand and reactive power demand of the m-th node, respectively. and These represent the maximum charging and discharging power that the mobile emergency energy storage vehicle can provide at the m-th node; and These represent the rated active power and reactive power of branch mn, respectively;
[0045] Equations (11) to (20) represent the optimal power flow model of the distribution network;
[0046] Based on the requirements for establishing distribution network reconfiguration and optimal power flow in the distribution network, the objective function is constructed as follows:
[0047]
[0048] In the formula, This represents the unit value of the load at the i-th node; This represents the load loss at the i-th node at time t, i.e. Δt represents the time scale of the calculation; c iS represents the cost per unit load generated by a distributed generator; g Represents a set of distributed generators; pg i,t This represents the output active power of the distributed generator;
[0049] Step S63: Solve the model using the Gurobi solver to obtain the distribution network reconfiguration scheme, optimal power flow distribution, and power of the emergency mobile energy storage vehicle.
[0050] Step S64: Calculate the reward function r of the mobile emergency energy storage vehicle based on the optimization obtained above. i,t :
[0051]
[0052] In the formula, Equation (22) indicates that the reward function of the mobile emergency energy storage vehicle is mainly determined by the amount of load restored, so as to supply power to the load of the distribution network as quickly as possible and achieve the goal of improving the resilience of the distribution network.
[0053] Furthermore, in step S7, the weights of the graph neural network reinforcement learning algorithm are updated based on the stochastic gradient descent method, specifically including:
[0054] Step S71: Randomly select a certain number of samples from memory unit D;
[0055] Step S72: Construct the loss function as shown in Equation (23), and update the weights of the graph neural network reinforcement learning algorithm according to the stochastic gradient descent method under the extracted sample as shown in Equation (24);
[0056]
[0057] In the formula, x, a, x', and a' represent the current state, action, and the state and action at the next time step, respectively; r represents the immediate reward for reinforcement learning in a graph neural network; θ t γ represents the parameters of the graph neural network reinforcement learning algorithm at the current time t; 0≤γ≤1 represents the discount factor, which reflects the influence of the future Q value on the current action; In the target graph neural network reinforcement learning algorithm, the parameter θ represents... t 'State-Action Values'
[0058]
[0059] In the formula, θ t This represents the parameters of the graph neural network reinforcement learning algorithm at the current time t; Indicates the relationship with θ tPerform differentiation; α represents the learning rate;
[0060] Step S73: After each given number of steps, adjust the current graph neural network reinforcement learning parameters θ. t Reinforcement learning parameters θ of the target graph neural network t Update.
[0061] Compared with the prior art, the present invention and its preferred embodiments have the following beneficial effects:
[0062] This method treats mobile energy storage vehicles (MEVs) in the study area as agents and abstracts their dynamic relationships as edges, thus transforming their collaborative relationships into a dynamic network graph structure. The graph neural network reinforcement learning method combines the powerful dynamic relationship processing capabilities of graph neural networks with the powerful sequential stochastic optimization decision-making capabilities of reinforcement learning, effectively solving the dynamic interaction problem of multi-agent systems under uncertain environments. The proposed graph neural network reinforcement learning method formulates mobile emergency vehicle scheduling strategies, effectively addressing the resilience improvement problem considering the uncertainty of repair time for potential traffic network and power system line faults. Based on an active distribution network considering renewable energy output, a distribution network reconfiguration and optimal power flow model is constructed to formulate distribution network resilience improvement strategies, and the reward function for mobile emergency energy storage vehicles is calculated based on this to achieve optimal scheduling of mobile emergency energy storage vehicles. In summary, the method proposed in this invention can effectively consider the uncertainty of repair time for traffic network and power system line faults, and achieve distribution network resilience improvement by formulating distribution network recovery strategies through the scheduling of mobile emergency energy storage vehicles. Attached Figure Description
[0063] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0064] Figure 1 This is a flowchart of a preferred embodiment of the present invention. Detailed Implementation
[0065] To make the features and advantages of this patent more apparent and understandable, specific embodiments are provided below for detailed explanation:
[0066] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0067] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0068] like Figure 1 The diagram shows a method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning, provided in this embodiment. The method includes the following steps:
[0069] S11: Initialization of the power distribution network-transportation network model;
[0070] S12: Random recovery from faults in transportation network and power system lines;
[0071] S13: Based on the information from the power distribution network and transportation system, construct the state x for the graph neural network reinforcement learning algorithm. i,t ;
[0072] S14: Generate a dispatching behavior strategy for mobile emergency energy storage vehicles based on the ε-Greedy algorithm and graph neural network reinforcement learning algorithm. i,t ;
[0073] S15: Implement the dispatch strategy for mobile emergency energy storage vehicles. i,t It also judges and updates the status of the mobile energy storage vehicle;
[0074] S16: Calculate the distribution network reconfiguration strategy, and calculate the reward function r for the mobile emergency energy storage vehicle based on the reconfiguration and optimization of the distribution network. i,t ;
[0075] S17: The state x' of the graph neural network reinforcement learning algorithm i,t renew;
[0076] S18: Transfer the information from the current step (x) i,t ,a i,t ,r i,t ,x' i,t The weights are stored in memory unit D and updated using the stochastic gradient descent method for the graph neural network reinforcement learning algorithm.
[0077] S19: Determine if the predetermined time T has been reached. end If not, execute steps (S12) to (S17); if yes, output the graph neural network reinforcement learning algorithm parameters and the corresponding optimization scheduling results.
[0078] The following is a detailed description of the solution in this embodiment:
[0079] I. Distribution Network-Transportation Network Model Initialization. The main steps include: 1) Distribution network model initialization, including the initialization of upper and lower limit voltages of distribution network system nodes, line and transformer parameters, upper and lower constraints of distributed generator voltage and output, output of renewable energy, distribution network load factor, charging station capacity and location, location of faulty lines and estimated recovery time, etc.; 2) Transportation network model initialization, including the initialization of transportation network nodes, road length, road capacity and free-flow speed, location of charging stations in the transportation network, location of faulty traffic nodes and lines and estimated recovery time, etc.; 3) Neural network parameter initialization, including the initialization of neural network weights W and bias term b, learning rate α, discount factor γ, batch size B and memory unit capacity D, etc.; 4) Time t=0 initialization.
[0080] The mobile energy storage vehicles in the study area are treated as agents, with each vehicle considered as a node n∈N. The connections between the mobile energy storage vehicles are considered as edges e∈E, thus forming a graph network structure G=(N,E). For each electric vehicle i in the current state x... i,t Initialize the adjacency matrix A.
[0081] II. Faults in transportation networks and power systems will be restored randomly.
[0082] Extreme events can cause weak lines in the power distribution network to break, and also damage roads. As the weather improves and emergency personnel work to repair them, the faulty power distribution lines and damaged roads will gradually be repaired. However, the repair time varies depending on the extent of the damage and the skill level of the emergency repair personnel, exhibiting a degree of randomness. Therefore, this embodiment assumes that the recovery time of the road network and power distribution network faults follows a certain normal distribution. in Indicates faulty component L i Average repair time; Indicates faulty component L i Time variance of repair.
[0083] III. Based on information from the power distribution network and transportation system, construct the state x of the graph neural network reinforcement learning algorithm. i,t .
[0084] The state x of the graph neural network reinforcement learning algorithm i,t Including the state of the i-th electric vehicle at time t (EV) i,t Neighboring electric vehicle status Ne i,t Line status (Is the line disconnected?) Power load demand and renewable energy output Composition, that is
[0085]
[0086] Wherein, formula (48) represents the state EV of the i-th electric vehicle. i,t Including the next node when an electric vehicle travels to a charging station Road number Electric vehicle speed v i,t and the remaining SOC of the mobile energy storage vehicle i,t Formula (49) represents the state Ne of the nearest electric vehicle. i,t This includes the states of each neighboring electric vehicle k, such as the next node of the kth electric vehicle that is adjacent to the i-th electric vehicle. The road number it is located on Electric vehicle speed v i,k,t and remaining power SOC i,k,t .
[0087] IV. Generating a dispatching behavior strategy for mobile emergency energy storage vehicles based on the ε-Greedy algorithm and graph neural network reinforcement learning algorithm. i,t Includes the following steps:
[0088] Step 41: Generate the scheduling strategy for the mobile energy storage vehicle in a random manner with probability ε, i.e.
[0089] a i,t = np.random.randint(||A action ||) (50)
[0090] In the formula, ||A action || represents the number of actions of the mobile emergency energy storage vehicle; when a i,t =0 indicates that the mobile emergency energy storage vehicle is in a charging or discharging state at this location; when a i,t If ≠0, it means the mobile emergency energy storage vehicle follows dispatch behavior strategy a. i,t The vehicle is driven to the next destination for charging or discharging, thereby reducing the load on the distribution network and improving its resilience.
[0091] Step 42: Generate a dispatching behavior strategy a for the mobile emergency energy storage vehicle with a probability of 1-ε based on the experience of the graph neural network reinforcement learning algorithm. i,t ,Right now
[0092]
[0093] In the formula, argmax(·) represents the parameter corresponding to the maximum value; Q(x i,t ,a;θ t ) indicates that the i-th mobile emergency energy storage vehicle is in state x i,tθ is the adjacency matrix A and the action-value function under scheduling behavior policy a. t This represents the neural network parameters of the graph neural network reinforcement learning algorithm at time t.
[0094] The graph neural network reinforcement learning algorithm in this embodiment has a neural network structure including a one-layer input layer, whose input includes the state set x of all mobile emergency energy storage vehicles. t ={x 1,t ,x 2,t ,...,x N,t The state x is then processed by a fully connected layer, which represents the relationship matrix consisting of mobile emergency energy storage vehicles, i.e., the adjacency matrix A. t Perform feature extraction x' t Secondly, the two-layer graph attention neural network is used to process the proposed feature x'. t The adjacency matrix A is processed; finally, a fully connected layer is connected to output the mobile emergency energy storage vehicle in state x. t The action value functions corresponding to the adjacency matrix A are obtained below. The mobile emergency energy storage vehicle agent can select the scheduling behavior strategy based on these action value functions.
[0095] V. Implementation of the dispatch strategy for mobile emergency energy storage vehicles a i,t It also assesses and updates the status of the mobile energy storage vehicle.
[0096] VI. Calculate the distribution network reconfiguration strategy, and calculate the reward function r for the mobile emergency energy storage vehicle based on the reconfiguration and optimization of the distribution network. i,t It includes the following steps:
[0097] Step S61: Update the distribution network load factor and calculate the charging requirements of the mobile energy storage vehicle. and maximum discharge power
[0098] Step S62: Establish distribution network reconfiguration and optimal power flow models:
[0099]
[0100] -Mα mn,t ≤V mn,t ≤Mα mn,t (55)
[0101]
[0102]
[0103] In the formula, S b S represents the set of all nodes in a power distribution network; rThis indicates that the first and last nodes of the substation, distributed power source (generator) nodes, and faulty lines constitute the potential root node set; N b α represents the number of nodes in the distribution network. mn,t This represents the line state at time t, and is a binary variable; 1 indicates a closed line, and 0 indicates an open line. n,t This is a binary variable; if the potential root node n is the root node of the distribution network during time period t, its value is 1; otherwise, it is 0. M is a sufficiently large constant; V mn,t Let t be the virtual current flow through branch mn at time t.
[0104] Equation (52) indicates that the distribution network established during the network reconfiguration process must satisfy a radial topology, that is, the number of closed lines is equal to the number of network nodes minus the number of subgraphs.
[0105] Equations (53) to (55) are the constraints for the reconfiguration of the distribution network, which require that each sub-graph is internally connected. pf represents the set of faulty lines at time t. mn,t and qf mn,t This represents the active and reactive power flow of the distribution network branch mn at time t; pd m,t and qd m,t pg represents the restored active and reactive loads of the distribution network branch at node m at time t; m,t and qg m,t This represents the active and reactive power output of the distributed power source m. and This represents the active and reactive power of the mobile emergency energy storage vehicle at the m-th node. and This represents the active and reactive power output of the wind turbine connected to the m-th node. and These represent the original active power load demand and reactive power demand of the m-th node, respectively. and These represent the maximum charging and discharging power that the mobile emergency energy storage vehicle can provide at the m-th node; and These represent the rated active power and reactive power of branch mn, respectively.
[0106] Equations (57) to (66) represent the optimal power flow model of the distribution network.
[0107] Based on the requirements for establishing distribution network reconfiguration and optimal power flow in the distribution network, the objective function is constructed as follows:
[0108]
[0109] In the formula, This represents the unit value of the load at the i-th node; This represents the load loss at the i-th node at time t, i.e. Δt represents the time scale of the calculation; c i S represents the cost per unit load generated by a distributed generator; g Represents a set of distributed generators; pg i,t This represents the output active power of a distributed generator.
[0110] Step S63: Solve the above model using the Gurobi solver to obtain the distribution network reconfiguration scheme, optimal power flow distribution, and power of the emergency mobile energy storage vehicle.
[0111] Step S64: Calculate the reward function r of the mobile emergency energy storage vehicle based on the optimization obtained above. i,t :
[0112]
[0113] In the formula, The value of the load unit restored by the mobile energy storage vehicle is represented by Equation (68). Equation (68) shows that the reward function of the mobile emergency energy storage vehicle is mainly determined by the amount of load restored, so as to supply power to the load of the distribution network as quickly as possible and achieve the purpose of improving the resilience of the distribution network.
[0114] VII. The State x' of the Graph Neural Network Reinforcement Learning Algorithm i,t Updates include updating the electric vehicle status (EV). i,t Neighboring electric vehicle status Ne i,t .
[0115] 8. Transfer the information from the current step (x) i,t ,a i,t ,r i,t ,x' i,t The weights are stored in memory unit D and updated using stochastic gradient descent for the graph neural network reinforcement learning algorithm. This mainly includes the following steps:
[0116] Step 71: Randomly select a certain number of samples from memory unit D;
[0117] Step 72: Construct the loss function as shown in Equation (69), and update the weights of the graph neural network reinforcement learning algorithm according to the stochastic gradient descent method under the extracted sample as shown in Equation (70);
[0118]
[0119] In the formula, x, a, x', and a' represent the current state, action, and the state and action at the next moment, respectively; θ tγ represents the parameters of the graph neural network reinforcement learning algorithm at the current time t; 0≤γ≤1 represents the discount factor, which reflects the influence of the future Q value on the current action; In the target graph neural network reinforcement learning algorithm, the parameter θ represents... t 'State-action value'.
[0120]
[0121] In the formula, θ t This represents the parameters of the graph neural network reinforcement learning algorithm at the current time t; Indicates the relationship with θ t Perform the differentiation operation; α represents the learning rate.
[0122] Step 73: After a certain number of steps, adjust the current graph neural network reinforcement learning parameters θ. t Reinforcement learning parameters θ of the target graph neural network t Update.
[0123] 9. Determine whether the predetermined time T has been reached. end If not, proceed to steps two through eight; if yes, output the graph neural network reinforcement learning algorithm parameters and corresponding output results.
[0124] This invention provides a method for dispatching mobile emergency vehicles and improving the resilience of distribution networks based on graph neural network reinforcement learning. This method treats mobile energy storage vehicles in the study area as agents and abstracts their dynamic relationships as edges, thus transforming the collaborative relationships of mobile energy storage vehicles into a dynamic network graph structure. The graph neural network reinforcement learning method combines the powerful dynamic relationship processing capabilities of graph neural networks with the powerful sequential stochastic optimization decision-making capabilities of reinforcement learning, effectively solving the dynamic interaction problem of multi-agents under uncertain environments. Therefore, this embodiment proposes a graph neural network reinforcement learning method to formulate a mobile emergency vehicle dispatch strategy, effectively addressing the resilience improvement problem considering the uncertainty of repair time for potential traffic network and power system line faults. Based on an active distribution network considering renewable energy output, a distribution network reconfiguration and optimal power flow model is constructed to formulate the distribution network resilience improvement strategy, and the reward function of the mobile emergency energy storage vehicle is calculated based on this to achieve optimal dispatching of the mobile emergency energy storage vehicle. In summary, the method proposed in this embodiment can effectively consider the uncertainty of repair time for traffic network and power system line faults, and improve the resilience of the distribution network by dispatching mobile emergency energy storage vehicles to formulate a distribution network recovery strategy.
[0125] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0127] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0128] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0129] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
[0130] This patent is not limited to the above-described preferred embodiments. Anyone can derive other forms of mobile emergency vehicle dispatching and power grid resilience enhancement methods based on graph neural network reinforcement learning under the guidance of this patent. All equivalent changes and modifications made within the scope of this patent application shall fall within the scope of this patent.
Claims
1. A method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning, characterized in that, Includes the following steps: Step S1: Initialize the power distribution network-transportation network model; Step S2: Random recovery of faults in transportation network and power system lines; Step S3: Based on the information from the power distribution network and transportation system, construct the state x for the graph neural network reinforcement learning algorithm. i,t ; Step S4: Generate a dispatching behavior strategy for the mobile emergency energy storage vehicle based on the ε-Greedy algorithm and the graph neural network reinforcement learning algorithm. i,t ; Step S5: Execute the dispatching strategy for the mobile emergency energy storage vehicle. i,t It also judges and updates the status of the mobile energy storage vehicle; Step S6: Calculate the distribution network reconfiguration strategy, and calculate the reward function r for the mobile emergency energy storage vehicle based on the reconfiguration and optimization of the distribution network. i,t ; Step S7: The state x' of the graph neural network reinforcement learning algorithm i,t renew; Step S8: Transfer the information from the current step (x) i,t ,a i,t ,r i,t ,x' i,t The weights are stored in memory unit D and updated using the stochastic gradient descent method for the graph neural network reinforcement learning algorithm. Step S9: Determine whether the predetermined time T has been reached. end If not, proceed to steps S2 to S8; if yes, output the graph neural network reinforcement learning algorithm parameters and the corresponding optimization scheduling results. In step S6, the reward function r of the mobile emergency energy storage vehicle is calculated based on the distribution network reconfiguration and optimal power flow optimization. i,t Specifically, the following steps are included: Step S61: Update the distribution network load factor and calculate the maximum charging power of the mobile energy storage vehicle. and maximum discharge power Step S62: Establish distribution network reconfiguration and optimal power flow models: -But mn,t ≤V mn,t ≤Mα mn,t (9) In the formula, S b S represents the set of all nodes in a power distribution network; r This indicates that the first and last nodes of the connected substations, distributed power generation nodes, and faulty lines constitute the potential root node set; N b α represents the number of nodes in the distribution network. mn,t This represents the line state at time t, and is a binary variable; 1 indicates a closed line, and 0 indicates an open line. n,t This is a binary variable; if the potential root node n is the root node of the distribution network during time period t, its value is 1; otherwise, it is 0. M is a sufficiently large constant. V mn,t The virtual current flows through branch mn at time t; Equation (6) indicates that the distribution network established during the network reconfiguration process must satisfy a radial topology, that is, the number of closed lines is equal to the number of network nodes minus the number of subgraphs; Equations (7) to (9) are the constraints for the reconfiguration of the distribution network, which require that each sub-graph is internally connected; pf represents the set of faulty lines at time t. mn,t and qf mn,t This represents the active and reactive power flow of the distribution network branch mn at time t; pd m,t and qd m,t pg represents the restored active and reactive loads of the distribution network branch at node m at time t; m,t and qg m,t This represents the active and reactive power output of node m corresponding to the distributed power source. and This represents the active and reactive power of the mobile emergency energy storage vehicle at the m-th node. and This represents the active and reactive power outputs of the wind turbine connected to the m-th node; and These represent the original active power load demand and reactive power demand of the m-th node, respectively. and These represent the maximum charging and discharging power that the mobile emergency energy storage vehicle can provide at the m-th node; and These represent the rated active power and reactive power of branch mn, respectively; Equations (11) to (20) represent the optimal power flow model of the distribution network; Based on the requirements for establishing distribution network reconfiguration and optimal power flow in the distribution network, the objective function is constructed as follows: In the formula, This represents the unit value of the load at the m-th node; This represents the load loss at the i-th node at time m, i.e. Δt represents the time scale of the calculation; c m S represents the cost per unit load generated by a distributed generator; g Represents a set of distributed generators; pg m,t This represents the output active power of the distributed generator; Step S63: Solve the model using the Gurobi solver to obtain the distribution network reconfiguration scheme, optimal power flow distribution, and power of the emergency mobile energy storage vehicle. Step S64: Calculate the reward function r of the mobile emergency energy storage vehicle based on the optimization obtained above. i,t : Equation (22) shows that the reward function of the mobile emergency energy storage vehicle is mainly determined by the amount of restored load, so as to supply power to the load of the distribution network as quickly as possible and achieve the purpose of improving the resilience of the distribution network.
2. The method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning according to claim 1, characterized in that, Step S1 specifically includes the following steps: Step S11: Initialize the distribution network model, including at least the following: upper and lower limit voltages of distribution network system nodes, line and transformer parameters, upper and lower constraints on voltage and output of distributed generators, output of renewable energy, distribution network load rate, charging station capacity and location, location of faulty lines and estimated recovery time. Step S12: Initialize the traffic network model, including at least: traffic network nodes, road length, road capacity and free-flow speed, the location of charging stations in the traffic network, the location of faulty traffic nodes and lines and the estimated recovery time. Step S13: Initialize neural network parameters, including at least: initialization of neural network weights W and bias term b, and initialization of hyperparameters such as learning rate α, discount factor γ, batch size B and memory unit capacity D; Step S14: Initialize time t=0.
3. The method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning according to claim 1, characterized in that, In step S2, we consider extreme events causing weak lines in the power distribution network to break, which also damages roads. As the weather improves and emergency personnel work to repair them, the faulty power distribution lines and damaged roads will gradually be repaired. However, the repair time exhibits a degree of randomness due to the varying degrees of damage and the different levels of expertise among emergency repair personnel. Therefore, we assume that the recovery time of the road network and power distribution network faults follows a certain normal distribution. in Indicates faulty component L i Mean repair time; Indicates faulty component L i Time variance of repair.
4. The method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning according to claim 1, characterized in that, In step S3, the state x of the graph neural network reinforcement learning algorithm i,t Including the state of the i-th electric vehicle at time t (EV) i,t Neighboring electric vehicle status Ne i,t Line status That is, whether the line is disconnected and the power load demand. and renewable energy output Right now: Wherein, formula (2) represents the state EV of the i-th electric vehicle. i,t Including the next node when an electric vehicle travels to a charging station Road number Electric vehicle speed v i,t and the remaining SOC of the mobile energy storage vehicle i,t ; Formula (3) represents the state Ne of the nearest electric vehicle. i,t This includes the states of each neighboring electric vehicle k, including the next node of the kth electric vehicle that is adjacent to the i-th electric vehicle. Then the road number it is located on Electric vehicle speed v i,k,t and remaining power SOC i,k,t .
5. The method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Generate the scheduling behavior strategy of the mobile energy storage vehicle in a random manner with probability ε, i.e. the i,t =np.random.randint(||A action ||) (4) In the formula, ||A action || represents the number of actions of the mobile emergency energy storage vehicle; when a i,t =0 indicates that the mobile emergency energy storage vehicle is in a charging or discharging state at its current location; when a i,t If ≠0, it means the mobile emergency energy storage vehicle follows dispatch behavior strategy a. i,t The vehicle is driven to the next destination for charging or discharging, thereby reducing the load on the distribution network and improving its resilience. Step S42: Generate a dispatching behavior strategy a for the mobile emergency energy storage vehicle with a probability of 1-ε based on the experience of the graph neural network reinforcement learning algorithm. i,t ,Right now: In the formula, argmax(·) represents the parameter corresponding to the maximum value; Q(x i,t ,a;θ t ) indicates that the i-th mobile emergency energy storage vehicle is in state x i,t θ is the adjacency matrix A and the action-value function under scheduling behavior policy a. t This represents the neural network parameters of the graph neural network reinforcement learning algorithm at time t; The neural network structure of the graph neural network reinforcement learning algorithm includes a single input layer, whose input includes the state set x of all mobile emergency energy storage vehicles. t ={x 1,t ,x 2,t ,...,x N,t } and the relation matrix consisting of mobile emergency energy storage vehicles, i.e., the adjacency matrix A; and a fully connected layer is used to process the input state x. t Perform feature extraction x' t Then, a two-layer graph attention neural network is connected to process the proposed feature x'. t The adjacency matrix A is processed; finally, a fully connected layer is connected to output the mobile emergency energy storage vehicle in state x. t The action value functions corresponding to the adjacency matrix A are obtained below; the mobile emergency energy storage vehicle agent selects the scheduling behavior strategy based on the action value functions.
6. The method for dispatching mobile emergency vehicles and improving the resilience of power distribution networks based on graph neural network reinforcement learning according to claim 1, characterized in that, In step S7, the weights of the graph neural network reinforcement learning algorithm are updated based on the stochastic gradient descent method, specifically including: Step S71: Randomly select a certain number of samples from memory unit D; Step S72: Construct the loss function as shown in Equation (23), and update the weights of the graph neural network reinforcement learning algorithm according to the stochastic gradient descent method under the extracted sample as shown in Equation (24); In the formula, x, a, x', and a' represent the current state, action, and the state and action at the next time step, respectively; r represents the immediate reward for reinforcement learning in a graph neural network; θ t γ represents the parameters of the graph neural network reinforcement learning algorithm at the current time t; 0≤γ≤1 represents the discount factor, which reflects the influence of the future Q value on the current action; In the target graph neural network reinforcement learning algorithm, the parameter θ represents... t 'State-Action Values' In the formula, θ t This represents the parameters of the graph neural network reinforcement learning algorithm at the current time t; Indicates the relationship with θ t Perform differentiation; α represents the learning rate; Step S73: After each given number of steps, adjust the current graph neural network reinforcement learning parameters θ. t Reinforcement learning parameters θ of the target graph neural network t Update.