Power-traffic coupling system toughness improvement method based on multi-type mobile emergency resource optimization scheduling
Through the graph diffusion attention network reinforcement learning algorithm and the improved semi-dynamic user equilibrium model, mobile energy storage vehicles, line repair teams and road repair teams are coordinated to dispatch, which solves the problem of poor resilience improvement of the power-transportation coupling system in the existing technology and achieves more efficient emergency resource scheduling and system resilience improvement.
Patent Information
- Application Number
- CN202510837109.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
AI Technical Summary
Existing research has failed to effectively consider the impact of the dynamic changes in power line faults and traffic road fault states on the dispatching strategies of mobile energy storage vehicles, line repair teams, and road repair teams in the coordinated dispatching of multiple types of mobile emergency resources, resulting in poor results in improving the resilience of the power-transportation coupling system. In addition, the existing model solving algorithm is inefficient and difficult to cope with complex post-disaster scenarios.
The graph diffuse attention network reinforcement learning algorithm (GDATD3QN) is adopted, combined with the improved semi-dynamic user equilibrium model and continuous averaging algorithm, to construct a multi-task neural network structure to coordinate the dispatch of mobile energy storage vehicles, line repair teams and road repair teams, optimize the resilience improvement strategy of the power-transportation coupling system, and consider the influence of multiple uncertain factors.
The resilience of the power-transportation coupling system after extreme events has been improved. Through improved models and algorithms, the efficiency of strategy solving and information dissemination capabilities have been improved, the changes in urban traffic travel routes have been accurately described, and the scheduling strategy of emergency resources has been optimized.
Smart Images

Figure CN120707112A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of power-transportation coupling, and specifically relates to a method for improving the resilience of a power-transportation coupling system based on optimal scheduling of multiple types of mobile emergency resources. Background Art
[0002] With the continued expansion of new power systems and the advancement of urban transportation electrification, the interaction between power and transportation systems is becoming increasingly close, exhibiting highly interconnected, multi-layered interactions. However, while this increased coupling between power and transportation systems improves energy allocation and travel efficiency, it also makes local faults more likely to propagate between the two networks, increasing the risk of power outages and road paralysis, and reducing the resilience of both systems to extreme events such as extreme weather disasters and man-made attacks.
[0003] Mobile emergency resources, with their flexible spatial and temporal transfer capabilities, enable the spatiotemporal transmission of energy and dynamic fault repair. Combined with distributed power generation and power system network reconfiguration, they can effectively enhance the resilience of the coupled power-transportation system. Mobile energy storage vehicles are vehicles that store and flexibly transport energy to various locations. During natural disasters or power outages, mobile energy storage vehicles can quickly connect to charging stations to provide temporary power to critical facilities and equipment. They are the most widely used emergency power supply resource for power grids. However, dispatching mobile energy storage vehicles to support loads in isolated power grids without distributed resource access is only the beginning of a post-disaster emergency response. Faulty lines can cause persistent load shedding on the power supply side, especially during the process of mobile energy storage vehicles replenishing power and maneuvering. Therefore, it is necessary to dispatch line repair teams to repair power line faults and coordinate with mobile energy storage vehicles to restore power system loads after disasters. Furthermore, extreme events can directly or indirectly impact transportation networks, causing road damage and traffic disruptions. This not only limits the mobility of mobile energy storage vehicles and line repair teams within the transportation system, but also affects the normal travel of urban transportation users, causing changes in the distribution of road traffic in the transportation system, increasing the travel time costs of transportation users and the dispatching costs of mobile emergency resources. Therefore, it is necessary to dispatch road repair teams to repair faulty roads after disasters to improve the resilience of the transportation system.
[0004] However, the mathematical models and solution algorithms constructed in existing research in the process of coordinated scheduling of multiple types of mobile emergency resources have certain defects. First, the existing research results only study the emergency support of power loads in the short period after the disaster, and the resilience improvement mathematical models constructed ignore the impact of the dynamic changes of power line faults on the scheduling strategy of mobile energy storage vehicles and the power system network reconstruction strategy. Secondly, some resilience improvement mathematical models do not consider the impact of road faults on the routing strategy of mobile emergency resources, and at the same time ignore the restrictions of road fault status changes on the travel path selection of urban traffic users, and roughly make deterministic assumptions about the distribution of urban traffic traffic. Finally, when faced with larger-scale power-traffic coupling systems and the limited number of mobile emergency resources, which leads to a decrease in effective information and an increase in interference information in the adjacent information, the existing model solution algorithm has the problems of low strategy solution efficiency and insufficient information exploration capabilities. Based on the above analysis, it is urgent to construct a mathematical model for the optimal scheduling of multiple types of mobile emergency resources that can comprehensively consider the influence of uncertain factors such as the fault status of power system lines, the fault status of traffic system roads, the distribution of road traffic volume, the output of renewable energy, and the dynamic changes in the adjacent relationship of mobile emergency resources; at the same time, it is necessary to develop a model solving algorithm to efficiently and accurately coordinate the optimal scheduling strategy of multiple types of mobile emergency resources, so as to effectively improve the resilience of the power-transportation coupling system after extreme events. Summary of the Invention
[0005] The present invention aims to address the problems of existing mathematical models and solution algorithms by providing a method for improving the resilience of the power-transportation coupling system based on the optimal scheduling of multiple types of mobile emergency resources. This method collaboratively dispatches mobile energy storage vehicles, line repair teams, and road repair teams to carry out power load support, line fault repair, and road fault maintenance tasks. It also proposes an urban traffic allocation model based on an improved semi-dynamic user equilibrium to analyze the impact of changes in road traffic flow distribution caused by changes in urban traffic road fault status after a disaster on the optimal scheduling strategy for mobile emergency resources. Furthermore, it comprehensively considers multiple uncertainties, such as the power system line fault status, the transportation system road fault status, road traffic flow distribution, renewable energy output, and the adjacency of mobile emergency resources, to construct a mathematical model for improving the resilience of the power-transportation coupling system that takes urban traffic allocation into account. To efficiently and accurately solve this mathematical model, this strategy proposes a Graph Diffusion Attention Network Dueling Double Deep QNetwork (GDATD3QN) reinforcement learning algorithm. This algorithm improves on the traditional graph attention network by innovatively integrating the graph diffusion mechanism with the dynamic attention mechanism to construct the graph diffusion attention network. This network uses the generalized graph diffusion operation of Personalized PageRank theory to sparsify the original first-order neighborhood into a new, diffuse neighborhood. It then extracts the generalized adjacency relationships within this neighborhood and constructs a graph diffusion matrix. This effectively improves the model's information dissemination capabilities, captures more distant node interactions, and mitigates the attention collapse problem inherent in traditional graph attention networks. Furthermore, the algorithm employs a single-input, multi-output multi-task neural network architecture to simultaneously solve optimal scheduling policies for mobile energy storage vehicles, line repair teams, and road repair teams based on collected complex graph data and dynamic adjacency information for mobile emergency resources, effectively improving policy solution efficiency. The algorithm also incorporates the Double DQN algorithm, Dueling DQN algorithm, and prioritized experience replay to improve sampling efficiency and training effectiveness. Urban travelers, however, often experience significant route perception biases due to factors such as incomplete traffic information and personal travel habits. To account for this bounded rational decision-making behavior, this strategy proposes a continuous averaging algorithm with random utility to solve the urban travel assignment model and obtain road traffic flow distribution.
[0006] To achieve the above objectives, the technical solution of the present invention is: a method for improving the resilience of a power-transportation coupled system based on the optimal scheduling of multiple types of mobile emergency resources, comprising the following steps:
[0007] Step S1: Initialize the power-transportation coupling system model, the mobile emergency resource multi-task graph reinforcement learning model, the power-transportation coupling system failure scenario, and the travel demand of urban transportation users;
[0008] Step S2: Initialize the graph diffusion attention network reinforcement learning algorithm environment, the power-transportation coupling system fault scenario, and the urban transportation user travel demand;
[0009] Step S3: Based on the power-traffic coupling system information and the state information of multiple types of mobile emergency resources, the mobile energy storage vehicle state information of the graph diffusion attention network reinforcement learning algorithm is constructed. Line repair team status information and road repair crew status information Get the state matrix Construct the adjacency matrix A based on the adjacency relationship of mobile emergency resources t ;
[0010] Step S4: Calculate the mobile energy storage vehicle routing strategy based on the ε-Greedy strategy and the graph diffusion attention network reinforcement learning algorithm Line repair team routing strategy and road repair team routing strategies
[0011] Step S5: Execute the mobile energy storage vehicle routing strategy separately Line repair team routing strategy and road repair team routing strategies and update the status of mobile energy storage vehicles, line repair teams, and road repair teams;
[0012] Step S6: Initialize the number of iterations of the continuous average algorithm considering random utility n = 1, solve the urban traffic travel allocation model based on the improved semi-dynamic user equilibrium based on the current traffic system road fault state, and obtain the road traffic flow distribution and stranded traffic flow r t rs and user travel time costs
[0013] Step S7: Convert the mobile energy storage vehicle scheduling model, the line repair team scheduling model, the road repair team scheduling model, the power system reconstruction model, and the power system optimal power flow model into a mixed integer second-order cone programming model MISOCP, solve the mobile energy storage vehicle power scheduling strategy, the line repair team maintenance scheduling strategy, and the road repair team maintenance scheduling strategy, and calculate the reward function r of each mobile emergency resource t M 、r t L and r t R ;
[0014] Step S8: Update the mobile energy storage vehicle status information of the graph diffusion attention network reinforcement learning algorithm Line repair team status information and road repair crew status information Get the next state matrix o t+1 , update the adjacency matrix A t+1 ;
[0015] Step S9: The experience sample and initial priority Stored in the mobile energy storage vehicle memory unit D M , the experience sample and initial priority Stored in the line repair team memory unit D L , the experience sample and initial priority Stored in the road repair team memory unit D R ;
[0016] Step S10: Prioritized Experience Replay (PER) strategy is used to sample the importance of the memory unit, and the weights of the neural network of the multi-task structured graph diffusion attention network reinforcement learning algorithm are updated based on the stochastic gradient descent method;
[0017] Step S11: Determine whether the end time T has been reached End If not, execute steps S3 to S10;
[0018] Step S12: Determine whether the training end number E has been reached End If not, execute steps S3 to S11; if so, output the weight parameters of the neural network of the graph diffusion attention network reinforcement learning algorithm and the routing and scheduling strategies of the mobile energy storage vehicle, line repair team and road repair team.
[0019] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, any of the method steps described above can be implemented.
[0020] Compared with the existing technology, the present invention has the following advantages: The present invention is used to solve the problem of improving the resilience of the power-transportation coupling system based on the optimal scheduling of multiple types of mobile emergency resources under the influence of multiple uncertain factors. Compared with the existing technology, the specific advantages of the present invention are as follows:
[0021] First, in response to the deficiency of existing mathematical models that ignore the dynamic changes in the status of power line faults and traffic road faults, the present invention coordinates the dispatch of three mobile emergency resources, namely mobile energy storage vehicles, line repair teams and road repair teams, to repair post-disaster power system load reduction, line faults and traffic system road faults, and is committed to restoring the performance of the power-transportation coupling system to normal levels. It avoids the problems of existing technologies that ignore the dynamic changes in the status of power-side line faults and traffic-side road faults, and only conducts research on resilience improvement strategies in the short period after the disaster. Therefore, the present invention can achieve better resilience improvement effects on the power-transportation coupling system than the existing technologies.
[0022] Secondly, to address the problem that existing mathematical models ignore the dynamic changes in urban traffic flow distribution, this paper constructs an urban traffic trip allocation model with dynamic changes in feasible paths based on the improved semi-dynamic user equilibrium theory. This model explores the impact of dynamic changes in road faults on the distribution of urban traffic flow, travel time costs, and the optimal scheduling strategy for mobile emergency resources. Furthermore, considering that in real-world situations, urban traffic users may experience path perception bias due to factors such as incomplete traffic information and personal travel habits, this paper uses a continuous averaging algorithm with random utility to solve the urban traffic trip allocation model based on the improved semi-dynamic user equilibrium to describe this bounded rational decision-making behavior.
[0023] Finally, to address the inefficient policy solving and insufficient information exploration capabilities of existing model-solving algorithms, this paper proposes a Graph Diffusion Attention Network (GDATD3QN) reinforcement learning algorithm to solve a mathematical model for improving the resilience of a power-transportation coupled system that takes into account urban travel allocation. This algorithm improves upon traditional graph attention networks by innovatively integrating graph diffusion and dynamic attention mechanisms to construct a graph diffusion attention network. This network uses the generalized graph diffusion operation of the Personalized PageRank theory to sparsify the original first-order neighborhood into a new, diffuse neighborhood. It then extracts the generalized adjacency relationships within this neighborhood and constructs a graph diffusion matrix, effectively improving the model's information dissemination capabilities, capturing more distant node interactions, and addressing the attention collapse problem inherent in traditional graph attention networks. Furthermore, the algorithm employs a single-input, multi-output multi-task neural network architecture to simultaneously solve optimized scheduling strategies for mobile energy storage vehicles, line repair teams, and road repair teams based on collected complex graph structure data and dynamic adjacency information for mobile emergency resources, effectively improving the efficiency of policy solving. The algorithm also combines the Double DQN algorithm, Dueling DQN algorithm and priority experience replay strategies to improve the algorithm's sampling efficiency and training effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 Flow chart of the method of the present invention.
[0025] Figure 2 The neural network structure of the graph diffusion attention network reinforcement learning algorithm. DETAILED DESCRIPTION
[0026] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0027] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art to which this application belongs.
[0028] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0029] like Figure 1 As shown in FIG, a method for improving the resilience of a power-transportation coupling system based on optimal scheduling of multiple types of mobile emergency resources according to the present invention includes the following steps:
[0030] S1: Initialize the power-transportation coupling system model, the multi-task graph reinforcement learning model for mobile emergency resources, the power-transportation coupling system failure scenario, and the travel needs of urban transportation users;
[0031] S2: Initialization of the graph diffusion attention network reinforcement learning algorithm environment, initialization of the power-transportation coupling system fault scenario, and initialization of urban transportation user travel demand;
[0032] S3: Based on the power-traffic coupling system information and the status information of multiple types of mobile emergency resources, the mobile energy storage vehicle status information of the graph diffusion attention network reinforcement learning algorithm is constructed Line repair team status information and road repair crew status information Get the state matrix Construct the adjacency matrix A based on the adjacency relationship of mobile emergency resources t ;
[0033] S4: Calculate the routing strategy of mobile energy storage vehicles based on the ε-Greedy strategy and the graph diffusion attention network reinforcement learning algorithm Line repair team routing strategy and road repair team routing strategies
[0034] S5: Execute the routing strategy of mobile energy storage vehicles separately Line repair team routing strategy and road repair team routing strategies and update the status of mobile energy storage vehicles, line repair teams, and road repair teams;
[0035] S6: Initialize the number of iterations of the continuous average algorithm considering random utility n = 1, solve the urban traffic travel allocation model based on the improved semi-dynamic user equilibrium based on the current traffic system road fault state, and obtain the road traffic flow distribution and stranded traffic flow r t rs and user travel time costs
[0036] S7: Convert the mobile energy storage vehicle scheduling model, line repair team scheduling model, road repair team scheduling model, power system reconstruction model, and power system optimal power flow model into a mixed integer second-order cone programming model MISOCP, solve the mobile energy storage vehicle power scheduling strategy, line repair team maintenance scheduling strategy, and road repair team maintenance scheduling strategy, and calculate the reward function r of each mobile emergency resource tM 、r t L and r t R ;
[0037] S8: Updating the state of the mobile energy storage vehicle using a graph diffusion attention network reinforcement learning algorithm Line repair team status and the status of the road repair team Get the next state matrix o t+1 , update the adjacency matrix A t+1 ;
[0038] S9: Experience Sample and initial priority Stored in the mobile energy storage vehicle memory unit D M , the experience sample and initial priority Stored in the line repair team memory unit D L , the experience sample and initial priority Stored in the road repair team memory unit D R ;
[0039] S10: Prioritized Experience Replay (PER) strategy is used to sample importance of memory units, and the weights of the graph diffusion attention network based on the multi-task structure are updated based on the stochastic gradient descent method;
[0040] S11: Determine whether the end time T has been reached End If not, execute steps S3 to S10;
[0041] S12: Determine whether the training end number E has been reached End If not, execute steps S2 to S11; if so, output the weight parameters of the neural network of the graph diffusion attention network reinforcement learning algorithm and the routing and scheduling strategies of the mobile energy storage vehicle, line repair team and road repair team.
[0042] The following is a detailed introduction to this embodiment:
[0043] 1. Initialize the power-transportation coupling system model, the mobile emergency resource multi-task graph reinforcement learning model, the power-transportation coupling system failure scenario, and the travel needs of urban transportation users.
[0044] (1) Initialize the number of training scenes to e = 1;
[0045] (2) Initialization of the power-transportation coupling system model, including: power system node voltage limits, line and transformer parameter settings, distributed generator voltage and output upper and lower constraints, renewable energy output, power system load rate, charging station access location in the power system; traffic system nodes, road length, road capacity and free travel speed, and charging station location in the traffic system; initial charging station location settings for mobile energy storage vehicles, and warehouse location settings for line repair teams and road repair teams;
[0046] (3) Constructing a multi-layer dynamic graph model of power-transportation-mobile emergency resources, including: First, based on graph theory, the nodes in the power system and the lines connecting the nodes are regarded as points and edges, which is the basis of the power system graph model. The model can be expressed as Among them, N DN , L DN and Represent the power node set, power line set and power side charging station node set respectively. Secondly, the traffic intersections and traffic roads in the traffic system are regarded as the traffic system graph model based on points and edges; the model can be expressed as Among them, N TN , L TN 、 and They represent the traffic node set, traffic road set, traffic side charging station node set, line repair team warehouse node set, and road repair team warehouse node set. Then, all mobile energy storage vehicles, line repair teams, and road repair teams are transformed into intelligent agents, each agent is regarded as a node, and its adjacency relationship is regarded as an edge, to construct a dynamic graph model of mobile emergency resources. Where M is the set of mobile energy storage vehicles, R l Assemble the line repair team, R r Assemble for the road repair team, is the adjacency matrix composed of the adjacency relationship of mobile emergency resources. Finally, the power system graph model Traffic system graph model and dynamic graph model of mobile emergency resources A multi-layer dynamic graph model of power, transportation and mobile emergency resources is constructed to analyze the coupling relationship between the power system and the transportation system, as well as the dynamic interaction process between mobile emergency resources and the power-transportation coupling system.
[0047] (4) Initialization of the multi-task graph reinforcement learning model for mobile emergency resources, including: graph diffusion attention neural network layer number and neuron setting, weight parameter initialization; learning rate α M , α L and α R , sampling batch size N s and memory unit DM 、D L and D R Initialization of hyperparameters such as capacity;
[0048] (5) Initialization of power-transport coupling system fault scenarios, including: simulating power-transport coupling network fault scenarios after extreme events: randomly setting some line interruptions in the power system to construct a set of damaged lines Simulate the interruption of weak power lines in real life after extreme events; and randomly set some road damage in the traffic system to construct a damaged road set Simulate the situation in real life where some roads in the traffic system are blocked by obstacles caused by extreme events.
[0049] (6) Initialization of urban traffic user travel demand, including: loading all urban traffic travel demand sets RS, setting the total urban traffic user travel demand and the proportion coefficient of each travel demand rs to the total travel demand, and obtaining the basic travel demand q of each travel demand rs at different times rs,t .
[0050] 2. Initialize the graph diffusion attention network reinforcement learning algorithm environment, the power-transportation coupling system failure scenario, and the travel needs of urban transportation users.
[0051] (1) The simulation time is initialized to t = 1;
[0052] (2) Initialization of various variables in the environment of the graph diffusion attention network reinforcement learning algorithm, including: the load rate of each node in the power system, the output of renewable energy, the traffic flow and speed of each road in the transportation system; the initial position of the mobile energy storage vehicle (set at the charging station node), the initial SOC value of the mobile energy storage vehicle, and the initial speed of the mobile energy storage vehicle; the initial position of the line repair team (set in the line repair team warehouse), the repair materials of the line repair team, and the initial speed of the line repair team; the initial position of the road repair team (set in the road repair team warehouse), the repair materials of the road repair team, and the initial speed of the road repair team;
[0053] (3) Initialization of power-transportation coupling system fault scenarios and urban transportation user travel demand, including: initialization of the fault line set constructed above Fault Road Collection and the basic travel demand q of each city's travel demand rs rs,t .
[0054] 3. Based on the power-traffic coupling system information and the status information of multiple types of mobile emergency resources, the mobile energy storage vehicle status information of the graph diffusion attention network reinforcement learning algorithm is constructed. Line repair team status information and road repair crew status information Get the state matrix Construct the adjacency matrix A based on the adjacency relationship of mobile emergency resources t .
[0055] (1) Construct the status information of the mobile energy storage vehicle based on the power-traffic coupling system information and the status information of the mobile emergency resource itself Including the mobile energy storage vehicle's own status information Adjacent mobile emergency resource status information Traffic system road information at your location Power system node load data P t ED , wind turbine output P t WT , specifically expressed as:
[0056]
[0057] Where, M, R l 、R r , T and N r are the mobile energy storage vehicle set, line repair team set, road repair team set, time set, and adjacent mobile emergency resource set; m and b are the numbers of the mobile energy storage vehicle, l, r, n, and t are the line repair team number, road repair team number, adjacent mobile emergency resource number, and simulation time period, respectively; Represents the location information of the mobile energy storage vehicle in the traffic system at time t; Indicates the moving speed of the mobile energy storage vehicle in the traffic system; S m,t Indicates the remaining SOC value of the mobile energy storage vehicle; and are the information of the mobile energy storage vehicle, line repair team and road repair team adjacent to the mth mobile energy storage vehicle at time t; N r is the number of adjacent emergency resources of the mobile energy storage vehicle; J is the number of adjacent roads of the mobile energy storage vehicle; Indicates the jth road number adjacent to the mth mobile energy storage vehicle at time t, and Indicates road The first and last node numbers of and Respectively represent roads road status, road length, road traffic volume and road travel time; Indicates the current traffic system node location number, Indicates the position number of the next node to be selected, Indicates the position number of the previous node, Indicates the road number to be entered and Indicates the previous road number; and S b,n,t Respectively represent the location information, moving speed and remaining SOC value of the nth mobile energy storage vehicle in the traffic system; RT l,n,t and RT l,n,t They represent the remaining number of repair materials of the nth line repair team and road repair team with mobile energy storage vehicles respectively.
[0058] (2) Construct the status information of the line repair team based on the power-transportation coupling network information and the status information of the mobile emergency resources themselves Including the status information of the line repair team itself Adjacent mobile emergency resource status information Traffic system road information at your location Information on faulty lines repaired by the line repair team and line maintenance time Specifically expressed as:
[0059]
[0060] Where l and c are the numbers of the line repair teams.
[0061] (3) Construct the status information of the road repair team based on the power-traffic coupling network information and the status information of the mobile emergency resources themselves Including the road repair team's own status information Adjacent mobile emergency resource status information Traffic system road information at your location Information on faulty roads repaired by the road repair team and road maintenance time Specifically expressed as:
[0062]
[0063]
[0064] Where r and e are the numbers of the road repair teams.
[0065] (4) Construct a state matrix based on the status information of the mobile energy storage vehicle, line repair team, and road repair team
[0066] (5) Construct the adjacency matrix A based on the adjacency relationship of all mobile emergency resources t , specifically expressed as:
[0067]
[0068] Where, N is the adjacency relationship formed by all mobile emergency resources at time t; v The number of mobile emergency resources can be obtained by N v =|M|+|R l |+|R r |Calculate; is a binary variable, indicating the adjacency relationship between mobile emergency resources i and j. If they are adjacent, it is 1, otherwise it is 0. The adjacency matrix A t It will be input into the graph diffusion attention network reinforcement learning algorithm to participate in the calculation of the action value function.
[0069] 4. Calculate the routing strategy of mobile energy storage vehicles based on the ε-Greedy strategy and the graph diffusion attention network reinforcement learning algorithm Line repair team routing strategy and road repair team routing strategies
[0070] (1) First, calculate the random number p at time t t , whose value range is [0,1), and the ε-Greedy strategy is used to compare the current greedy ε t value and random number p t , if there is p t <ε t , then a random selection method (i.e., randomly select) is used to generate the actions of all mobile energy storage vehicles from the routing action set And form the routing behavior strategy of mobile energy storage vehicles All line repair teams' actions Line repair team routing behavior strategy All road repair crew actions Road repair team routing behavior strategy Specifically, as shown in formula (132):
[0071]
[0072] Where, Represents the routing action set of mobile energy storage vehicle routing, line repair team and road repair team; when the selected action a i,t = 0 means stopping at the original location, and also means that the mobile emergency resource route reaches the destination location (the mobile energy storage vehicle is the charging station or the initial charging station node, the line repair team is the fault line or warehouse node, and the road repair team is the fault road or warehouse node), and is ready to accept the subsequent model scheduling decision; when the action a is selected i,t≠0 means that the mobile energy storage vehicle, line repair team, and road repair team choose to enter the next traffic road to their destination according to the routing strategy;
[0073] (2) If p t ≥ε t , the routing behavior strategy of the mobile energy storage vehicle is generated by the graph diffusion attention network reinforcement learning algorithm Line repair team routing behavior strategy and the routing behavior strategy of the road repair team Specifically, as shown in formulas (133)-(135):
[0074]
[0075] In the formula, argmax() represents the parameter corresponding to the maximum value; and The multi-task structure of the neural network reinforcement learning algorithm for the graph diffusion attention network is based on the mobile emergency resource state matrix o t , mobile emergency resource adjacency matrix A t and the current network parameters θ t At the same time, the routing actions calculated for the mobile energy storage vehicle, line repair team and road repair team The value function of .
[0076] (3) The neural network of the diffusion attention network reinforcement learning algorithm is a single-input multi-output structure of a multi-task neural network. This structure is based on the mobile emergency resource state matrix o t and the adjacency matrix A t The action-value functions of mobile energy storage vehicles, line repair teams, and road repair teams are simultaneously calculated, allowing a single neural network to simultaneously control all mobile emergency resources, improving the resilience of the power-transportation coupled system. Unlike building independent neural networks for similar mobile emergency resources, the multi-task neural network with a single-input, multi-output structure employed in this patent effectively improves the efficiency of policy solving.
[0077] The neural network structure consists of a fully connected input layer, two graph diffusion attention layers, and three parallel output layers using the Dueling DQN strategy. First, the fully connected input layer receives the input mobile emergency resource state set o t Perform preliminary feature extraction and dimension conversion, output Secondly, the graph diffusion attention network layer uses the generalized graph diffusion operation of PersonalizedPageRank theory based on the graph diffusion mechanism to move the original first-order neighborhood to the emergency resource adjacency matrix A t Sparse, and then extract the generalized adjacency relationship of mobile emergency resources to construct a graph diffusion matrix The specific calculation formula is:
[0078]
[0079] T sym =D -1 / 2 A t D -1 / 2 (137)
[0080] S PPR =α ppr (I-(1-α ppr )T sym ) -1 (138)
[0081]
[0082] Where D is the degree matrix, N v is the number of mobile emergency resources; T sym is a symmetric normalized transition matrix, the sum of each row is 1; S PPR and are the PPR diffusion matrix and the diffusion matrix after pruning operation respectively; α ppr The PPR parameter is also called the restart probability, which is used to control the range of information diffusion and is adjusted according to the task's dependence on the balance between global information and local information. Its value range is (0,1); ppr Approaching 1, diffusion tends to retain the original node information and reduce the influence of other nodes. ppr As it approaches 0, the diffusion tends to be more inclined to the global information propagation of the entire graph, and I is the unit matrix; ε spa is the pruning threshold, the diffusion matrix S PPR Less than ε spa The elements of are set to 0 to sparse the matrix; ε spa Set to [0.01, 0.1]; is the diffusion matrix The degree matrix of GDC is the final normalized diffusion matrix; Equations (30) and (31) calculate the transition matrix T sym ; Equations (32) and (33) calculate the diffusion matrix S PPR And perform pruning operation on it; Equations (34) and (35) calculate the normalized diffusion matrix A GDC , and is used in the subsequent feature extraction of graph neural networks.
[0083] In addition, the graph diffusion attention layer uses a dynamic attention mechanism to the state matrix Sum diffusion matrix A GDC Calculate the attention coefficient and output feature information based on the graph diffusion attention mechanism The specific calculation formula is:
[0084]
[0085] Where σ() is the ReLU nonlinear activation function; A collection of adjacent mobile emergency resources; is the weight vector; Represents the weight matrix for linear transformation; || represents the merge operation; represents the normalized graph diffusion attention coefficient of mobile emergency resource j to mobile emergency resource i. Formula (144) calls K groups of independent attention mechanism layers and averages the calculated K groups of attention coefficients to obtain the final feature information output Learning process of diffuse attention networks with stable graphs.
[0086] Finally, the neural network constructs three parallel output layers to calculate the action Q value for the mobile energy storage vehicle, line repair team, and road repair team. Each output layer adopts the strategy of Dueling DQN algorithm to split the feature information into two fully connected branches: the first branch outputs the scalar value of the state function The second branch outputs the action advantage value function vector Therefore, the calculation of the action Q value can be expressed as:
[0087]
[0088] Where, ψ t for The fully connected neural network parameters of the branch; for The fully connected neural network parameters of the branch. The mobile energy storage vehicle, line repair team and road repair team select routing actions according to the action value function calculated by their respective neural networks. and Figure diffusion attention network reinforcement learning algorithm neural network structure Figure 2 shown.
[0089] 5. Execute the mobile energy storage vehicle routing strategy separately Line repair team routing strategy and road repair team routing strategies And update the status of mobile energy storage vehicles, line repair teams and road repair teams.
[0090] VI. Initialize the number of iterations of the continuous average algorithm considering random utility n = 1, solve the urban traffic travel allocation model based on the improved semi-dynamic user equilibrium based on the current traffic system road fault state, and obtain the road traffic flow distribution and stranded traffic flow r t rs and user travel time costs
[0091] (1) Construct an urban transportation travel allocation model based on improved semi-dynamic user equilibrium, as shown below:
[0092]
[0093] Where, L TN and They represent the traffic system roads and fault road sets respectively; RS represents the user travel start-end pair set, Where r is the starting point of the user's trip, s is the end point of the trip; K rs,t represents the set of feasible paths for the travel demand rs at time t; k represents a feasible path in the set of feasible paths; α tn and are road delay coefficients, usually taking values of 0.15 and 4; M is an integer much larger than the basic road travel time; is the path traffic flow of feasible path k; is a binary variable. If the traffic road l tn Set to 1 when it is on a feasible path k, otherwise it is 0; represents the travel time of feasible path k; r t rs and They are respectively represented as the stranded vehicle flow of travel demand rs at time t and t-1; q rs,t and They represent the travel demand rs at time t and the revised travel demand respectively; Represents the travel demand rs in the current feasible path set K rs,t The minimum travel time cost under the condition of faulty roads; as the faulty roads are repaired, the feasible path set K of travel demand rs rs,t will also change, and a new set of feasible paths with a lower travel time cost than the previous one will be obtained. Formula (146) indicates that the travel time of the faulty road is a number that is much larger than the basic travel time and can be considered to be infinite, and traffic users cannot pass through this road; Formula (147) calculates the traffic flow of the traffic road by searching all feasible paths under the user travel demand set; Formula (148) calculates the travel time of feasible path k; Formula (149) calculates the travel demand rs in the feasible path set K rs,t The stranded vehicle flow under the condition of rs,tThe modified travel demand of travel demand rs at time t is calculated by using the traffic flow at time t-1 and the traffic flow at time t-1; Formula (151) indicates that the sum of the traffic flow on all feasible paths of travel demand rs is equal to the corresponding modified travel demand; Formula (152) ensures that the traffic flow on each feasible path is non-negative and the actual travel time will not be less than K under the current feasible path set. rs,t The minimum travel time cost.
[0094] (2) Initialize the number of iterations of the continuous averaging algorithm considering random utility n = 1;
[0095] (3) The continuous average algorithm considering random utility is used to solve the urban traffic travel allocation model based on the improved semi-dynamic user equilibrium to obtain the road traffic flow distribution and the stranded traffic flow r t rs and user travel time costs The iterative solution process is as follows:
[0096] (a) At time t, the road fault status of the traffic system at the current moment is updated based on the road repair decision made by the road repair team at time t-1;
[0097] (b) Based on the current road fault status, the Dijkstra shortest path method is used to solve the feasible path set of all user travel demands rs;
[0098] (c) performing the nth iteration of the continuous averaging algorithm taking into account random utilities;
[0099] (d) All the travel demands q of traffic users are rs,t Assigned to the feasible path set K rs,t In the calculation, the traffic system road flow and the flow of each feasible path are obtained Update all road travel times;
[0100] (e) Solve for the traffic flow rate r t rs , and based on the stranded traffic flow at time t-1 Calculate the modified travel demand of rs;
[0101] (f) Calculate the selection probability of feasible path k for travel demand rs based on the Logit random utility model Specifically, it is shown in the following formula (153):
[0102]
[0103] Where δ is a scale parameter that represents the user's familiarity with traffic network information and is set to 0.3; is the feasible path set K of user travel demand rs at time t rs,tThe time cost of path k in .
[0104] (g) Based on the current iteration number n, update the path flow of the feasible path k of the travel demand rs through formula (154)
[0105]
[0106] (h) Determine whether the convergence condition is met by using equation (155). If so, proceed to step (i); otherwise, return to step (c):
[0107]
[0108] Where, ε logit Indicates the convergence accuracy, which is a fixed value; The parameter n represents the current number of iterations. Formula (155) represents the number of feasible paths assigned to the set K. rs,t When the difference between the proportion of traffic flow of any feasible path k in the total travel demand and the path selection probability calculated by the Logit random utility model is small enough, the traffic flow distribution of the entire model reaches a balanced state.
[0109] (i) The traffic assignment model is solved and the output is the traffic flow distribution, road travel time, and stranded traffic flow r of the traffic system at time t. t rs and user travel time costs
[0110] 7. Convert the mobile energy storage vehicle dispatch model, line repair team dispatch model, road repair team dispatch model, power system reconstruction model and power system optimal power flow model into a mixed integer second-order cone programming model MISOCP, solve the mobile energy storage vehicle power dispatch strategy, line repair team maintenance dispatch strategy and road repair team maintenance dispatch strategy, and calculate the reward function r of each mobile emergency resource t M 、r t L and r t R .
[0111] (1) Construct a mobile energy storage vehicle dispatch model, a power system reconstruction model, and a power system optimal power flow model, as shown below:
[0112]
[0113]
[0114] Where, Equations (156)-(162) are the mobile energy storage vehicle scheduling models; Represents the set of charging station nodes in the transportation system; is a binary variable, indicating whether the mobile energy storage vehicle is located at the traffic system node n tn , if it is at the node, it is 1, otherwise it is 0; and are binary variables, representing the charging and discharging status of the mth mobile energy storage vehicle at time t; and They represent the charging active power, discharging active power and reactive power at time t respectively; and Respectively represent the upper limit of active power and reactive power; S m,t represents the SOC at time t; and are the charging efficiency coefficient and the discharging efficiency coefficient respectively; S m and are the upper and lower limits of the SOC value respectively; Formulas (156)-(158) limit the charging and discharging power and reactive power of the mobile energy storage vehicle; Formula (159) limits the charging and discharging behavior of the mobile energy storage vehicle, that is, the mobile energy storage vehicle can only be in the charging state or the discharging state. At the same time, if the current location of the mobile energy storage vehicle is not located at the charging station node of the transportation system, it cannot be charged or discharged; Formulas (160)-(161) constrain the SOC value of the mobile energy storage vehicle; Formula (162) indicates that after the partial fault line of the power system is repaired, the power system load can be fully restored, and at this time there is no need to dispatch the mobile energy storage vehicle to perform the load support task;
[0115] Where, Equations (163)-(168) are the power system reconstruction models; L DN 、 L DN,sw and denote the set of power system lines, the set of damaged lines, the set of lines equipped with tie switches, and the set of power system source nodes, respectively; ρ(·) and φ(·) denote the set of parent nodes and the set of child nodes, respectively; α ij,t Indicates the connection status of the power system line (i, j) at time t, which is 1 if connected and 0 if disconnected; N DN represents the number of nodes in the power system; γ j,t is a binary variable, which is 1 if the power system node j is the source node at time t, otherwise it is 0; f ij,trepresents the virtual power flow into the power system node j at time t; M is a large constant. Constraint (163) sets the line state constraint based on the power system line fault in step S2; Constraint (164) is the power system radiation constraint, which describes the relationship between the power system line state (normal line, fault line and line equipped with remote switch) and the power system source node; Constraints (165)-(168) describe the virtual power flow balance constraints of the power system;
[0116] Where, Equations (169)-(179) are the optimal power flow models of the power system; and They represent the active power and reactive power flowing through the power system line (i, j) at time t, respectively; and Respectively represent the active and reactive outputs of the wind turbine WT; and They represent the active power exchange amount and reactive power exchange amount between the mobile energy storage vehicle and node j respectively; and Both represent the load loss of node j; and They represent the load demand of node j respectively; and They represent the maximum active power and maximum reactive power of the power system line (i, j) respectively; represents the upper capacity limit of the power system line (i, j); represents the square value of the voltage at the power system node j at time t; r ij and x ij Respectively represent the line resistance and reactance values of the power system line (i, j); and Respectively represent the upper and lower limits of the square value of the voltage of the power system node j at time t. Constraints (169) and (170) describe the active power balance and reactive power balance of the power system line; constraints (171) and (172) describe the active power and reactive power of the mobile energy storage vehicle charging (or discharging) at the charging station node at time t; constraint (173) limits the active power and reactive power of the mobile energy storage vehicle when it is not at the charging station node to 0; constraints (174) and (175) limit the active power and reactive power flowing through the power system line at time t; constraint (176) limits the thermal capacity of the line (i, j); constraints (177)-(179) describe the voltage relationship between adjacent nodes of the power system;
[0117] (2) Construct a dispatching model for the line repair team, as shown below:
[0118]
[0119]
[0120] Where, For the line repair team to repair the power system fault line at time t The repair decision variable is 1 if the line repair team chooses to repair the line, otherwise it is 0; Indicates the location of the traffic node where the fault line is located n dn ; The time required for the line repair team to repair the faulty line; The flag of the line repair team dispatch at time t. If it is 0, the dispatch continues. Constraint (180) ensures that the line repair team must arrive at the traffic node where the fault line is located before repairing it; constraint (181) restricts the fault line to only be repaired by one line repair team at the same time; constraint (182) describes that the line repair team has repaired T rep The faulty line at time t can resume normal operation at time t+1; constraint (183) restricts the faulty line that has been repaired from failing again in the future; constraint (184) indicates that when all faulty lines are repaired and the line repair team arrives at the line repair team warehouse at time t, the dispatch flag position is 1, and there is no need to dispatch the line repair team to perform the fault repair task.
[0121] (3) Construct a road repair team dispatch model, as shown below:
[0122]
[0123] Where, and Both are binary variables. If the road repair team is located at the traffic system node n tn Or road l tn If yes, it is 1, otherwise it is 0; It is a binary variable, indicating the road repair team's response to the faulty road. The repair decision is 1 if it is repaired, otherwise it is 0; Is a binary variable, indicating the road at time t Status; To repair damaged roads Time required; A binary variable used to determine whether to continue dispatching the road repair team to the warehouse location. The road repair team scheduling model is similar to the line repair team scheduling model and will not be described in detail here.
[0124] (4) This patent constructs a mathematical model for improving the resilience of the power-transportation coupling system taking into account urban traffic distribution, with the minimum total loss cost of the power-transportation coupling system as the objective function. The specific model is as follows:
[0125]
[0126] Where c dn 、c tn and c mer They are the unit load loss cost of the power system, the travel time cost of traffic users, and the unit time dispatch cost of mobile emergency resources. Calculated using load reduction at time t; Traffic system loss cost Calculated by formula (192), the travel cost of traffic users is and mobile emergency resource dispatch costs The model consists of two parts. Equation (193) calculates the travel cost of traffic users by multiplying the travel demand rs at time t by the incremental travel cost. In other words, minimizing the travel cost of traffic users in the model actually minimizes the incremental travel cost of all users. Equation (194) calculates the dispatching costs of mobile energy storage vehicles, line repair teams, and road repair teams at time t.
[0127] (5) Constraints (171) and (172) are linearized using the large M method, and constraint (176) is linearized using the second-order cone relaxation method; constraints (181) and (182) as well as constraints (186) and (187) of the line repair team and road repair team scheduling models are linearized using the McCormick envelope method;
[0128] (6) Using equations (190)-(194) as the objective function and equations (156)-(189) as the constraints, a mixed integer second-order cone programming model MISOCP is constructed and solved using the Gurobi commercial solver. The node load of the power system restored by the mobile energy storage vehicle is obtained, and the reward function of the mobile energy storage vehicle is calculated using equations (195) and (196).
[0129]
[0130] Where, represents the set of power system nodes whose loads are supported by mobile energy storage vehicles at time t; Power system node The load recovery amount; is the load level weight coefficient of the power system node.
[0131] By solving MISOCP, we can obtain the line maintenance strategy or scheduling strategy of the line repair team, and calculate the reward function of the road repair team through equations (197)-(201)
[0132]
[0133] Where, Represents the set of faulty lines repaired by the line repair team; It represents the travel time from the starting node b to the target node p of the traffic system; The contribution reward value of the line repair team for repairing the faulty line at time t; Rewards for the line repair team to complete the task, including: Reward for completing fault line repair, and The rewards for returning to the warehouse are all constant; is the penalty value of the line repair team at time t; It is a binary variable, which means that it is 1 when the road selected at time t is the same as that at time t-1, and 0 otherwise; and are all constants; and All are weight coefficients; Indicates the current flowing through the fault line repaired by the line repair team at time t The active power of the line repair team. Formula (198) calculates the total active power flowing through the fault line repaired by the line repair team at time t to calculate the contribution reward value of the line repair team for repairing the fault line; Formula (200) calculates the reward of the line repair team after repairing the fault line and returning to the warehouse, which is a fixed value; Formula (201) calculates the penalty value of the line repair team at time t, which includes the penalty for going back and the penalty for the line repair team repairing all fault lines but failing to return to the warehouse. And the penalty for the line repair team failing to repair all faulty lines Three parts.
[0134] By solving MISOCP, we can obtain the road maintenance strategy or scheduling strategy of the road repair team, and calculate the reward function of the road repair team through equations (202)-(206)
[0135]
[0136]
[0137] Where, The road repaired by the road repair team Contribution reward value; Formula (203) is calculated by the ratio of the sum of traffic flows on the faulty roads repaired by the road repair team at time t to the sum of traffic flows on all roads in the traffic system.
[0138] 8. Update the state of the mobile energy storage vehicle using the graph diffusion attention network reinforcement learning algorithm Line repair team status and the status of the road repair team Get the next state matrix o t+1 , update the adjacency matrix A t+1 .
[0139] 9. Experience Sample and initial priority Stored in the mobile energy storage vehicle memory unit D M , the experience sample and initial priority Stored in the line repair team memory unit D L , the experience sample and initial priority Stored in the road repair team memory unit D R .
[0140] 10. The prioritized experience replay (PER) strategy is used to sample the importance of memory units, and the weights of the graph diffusion attention network based on the multi-task structure are updated based on the stochastic gradient descent method.
[0141] (1) The input layer parameters of the multi-task neural network of the graph diffusion attention network reinforcement learning algorithm are recorded as θ In , the parameters of the first and second layer diffusion attention layers are denoted as θ G1 and θ G2 , the three parallel output layer parameters are recorded as θ m ,θ l and θ r ; The neural network parameters for the forward propagation and reverse update of the Q value of the mobile energy storage vehicle agent are recorded as θ In +θ G1 +θ G2 +θ m →θ M ; The neural network parameters for the line repair team and the road repair team are respectively denoted as θ In +θ G1 +θ G2 +θ l →θ L and θ In +θ G1 +θ G2 +θ r →θ R ;
[0142] (2) In the mobile energy storage vehicle memory unit D M In the importance sampling method, samples N are collected respectively s , and calculate the importance weight of each sample by formula (207):
[0143]
[0144] Where N D Represents memory unit D M The total number of samples in ; Indicates the priority of the sample; β is a hyperparameter.
[0145] (3) Combined with the strategy of Double DQN algorithm, the current network θ is adopted t To select the optimal action of the agent at the next moment, the target network To evaluate the value of an action, we make full use of the two neural networks of the DQN algorithm to separate action selection and strategy evaluation to reduce the risk of overestimating the Q value. We use equations (208) and (209) to calculate the loss function:
[0146]
[0147] Where θ t and are the current network parameters and target network parameters of the graph diffusion attention network reinforcement learning algorithm respectively; γ is the discount factor reflecting the impact of the future action value Q value on the current action, and its value is [0,1];
[0148] (4) Parameters of the mobile energy storage vehicle neural network based on the stochastic gradient descent method and sample importance weight Update, specifically:
[0149]
[0150] Where, α M is the learning rate of the gradient descent algorithm of the current network of the mobile energy storage vehicle. And after a fixed number of steps N up Then, the parameters of the target network of the mobile energy storage vehicle are adjusted in the following way. To update:
[0151]
[0152] (5) TD-error value calculated according to formula (208) Calculate the new priority for each sample Where ∈ is a small constant to prevent the priority from being zero. M Update the priority of each sample in
[0153] (6) Refer to steps (2) to (5) to calculate the neural network parameters of the line repair team Make updates;
[0154] (7) Refer to steps (2) to (5) to calculate the neural network parameters of the road repair team to update.
[0155] 11. Determine whether the end time T has been reached End If not, execute steps 3 to 10.
[0156] 12. Determine whether the training end number E has been reached End If not, execute steps 3 to 11; if so, output the weight parameters of the neural network of the graph diffusion attention network reinforcement learning algorithm and the routing and scheduling strategies of the mobile energy storage vehicle, line repair team and road repair team.
[0157] This invention provides a strategy for optimizing the scheduling of multiple mobile emergency resources and enhancing the resilience of the power-transportation coupling system based on graph diffusion attention network reinforcement learning. This strategy simultaneously coordinates the scheduling of three mobile emergency resources: mobile energy storage vehicles, line repair teams, and road repair teams, to enhance the resilience of the post-disaster power-transportation coupling system. It also constructs an urban traffic allocation model based on an improved semi-dynamic user equilibrium and employs a continuous averaging algorithm with random utility to solve the road traffic flow distribution and travel time cost. This strategy proposes a graph diffusion attention network reinforcement learning algorithm to solve the proposed mathematical model for enhancing the resilience of the power-transportation coupling system, which takes into account urban traffic allocation. This algorithm innovatively integrates graph diffusion mechanisms, dynamic attention mechanisms, and a multi-task structure to construct a graph diffusion attention network, enhancing the model's information dissemination capabilities and strategy solution speed, while addressing the attention collapse problem inherent in traditional graph attention networks. Furthermore, this algorithm combines the Double DQN algorithm, the Dueling DQN algorithm, and a prioritized experience replay strategy to improve sampling efficiency and training effectiveness. This algorithm can solve the optimal scheduling strategy for multiple mobile emergency resources, effectively enhancing the resilience of the post-disaster power-transportation coupling system.
[0158] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0159] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0160] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0162] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.
[0163] This patent is not limited to the above-mentioned optimal implementation method. Anyone can derive various other forms of mobile emergency vehicle dispatching and power system resilience improvement methods based on graph neural network reinforcement learning under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of this invention should be covered by this patent.
Claims
1. A method for improving the resilience of a power-transportation coupling system based on optimal scheduling of multi-type mobile emergency resources, characterized in that: The steps include: Step S1: Initialize the power-transportation coupling system model, the mobile emergency resource multi-task graph reinforcement learning model, the power-transportation coupling system failure scenario, and the travel demand of urban transportation users; Step S2: Initialize the graph diffusion attention network reinforcement learning algorithm environment, the power-transportation coupling system fault scenario, and the urban transportation user travel demand; Step S3: Based on the power-traffic coupling system information and the state information of multiple types of mobile emergency resources, the mobile energy storage vehicle state information of the graph diffusion attention network reinforcement learning algorithm is constructed. Line repair team status information and road repair crew status information Get the state matrix Construct the adjacency matrix A based on the adjacency relationship of mobile emergency resources t ; Step S4: Calculate the mobile energy storage vehicle routing strategy based on the ε-Greedy strategy and the graph diffusion attention network reinforcement learning algorithm Line repair team routing strategy and road repair team routing strategies Step S5: Execute the mobile energy storage vehicle routing strategy separately Line repair team routing strategy and road repair team routing strategies and update the status of mobile energy storage vehicles, line repair teams, and road repair teams; Step S6: Initialize the number of iterations of the continuous average algorithm considering random utility n = 1, solve the urban traffic travel allocation model based on the improved semi-dynamic user equilibrium based on the current traffic system road fault state, and obtain the road traffic flow distribution and stranded traffic flow and user travel time costs Step S7: Convert the mobile energy storage vehicle scheduling model, the line repair team scheduling model, the road repair team scheduling model, the power system reconstruction model, and the power system optimal power flow model into a mixed integer second-order cone programming model MISOCP, solve the mobile energy storage vehicle power scheduling strategy, the line repair team maintenance scheduling strategy, and the road repair team maintenance scheduling strategy, and calculate the reward function r of each mobile emergency resource t M 、r t L and r t R ; Step S8: Update the mobile energy storage vehicle status information of the graph diffusion attention network reinforcement learning algorithm Line repair team status information and road repair crew status information Get the next state matrix o t+1 , update the adjacency matrix A t+1 ; Step S9: The experience sample and initial priority Stored in the mobile energy storage vehicle memory unit D M , the experience sample and initial priority Stored in the line repair team memory unit D L , the experience sample and initial priority Stored in the road repair team memory unit D R ; Step S10: Prioritized Experience Replay (PER) strategy is used to sample the importance of the memory unit, and the weights of the neural network of the multi-task structured graph diffusion attention network reinforcement learning algorithm are updated based on the stochastic gradient descent method; Step S11: Determine whether the end time T has been reached End If not, execute steps S3 to S10; Step S12: Determine whether the training end number E has been reached End If not, execute steps S3 to S11; if so, output the weight parameters of the neural network of the graph diffusion attention network reinforcement learning algorithm and the routing and scheduling strategies of the mobile energy storage vehicle, line repair team and road repair team.
2. The method for improving the resilience of the power-transportation coupling system based on the optimal scheduling of multi-type mobile emergency resources according to claim 1 is characterized in that: Step S1 is specifically implemented as follows: Step S11: Initialization of the power-transportation coupling system model, including: power system node voltage limits, line and transformer parameter settings, distributed generator voltage and output upper and lower constraints, renewable energy output, power system load rate, charging station access location in the power system; transportation system nodes, road length, road capacity and free travel speed, and charging station location in the transportation system; initial charging station location settings for mobile energy storage vehicles, and warehouse locations for line and road repair teams; Step S12: Initialization of the multi-task graph reinforcement learning model for mobile emergency resources, including: setting the number of neural network layers and neurons of the graph diffusion attention network reinforcement learning algorithm, initialization of weight parameters; learning rate α M , α L and α R , sampling batch size N s and memory unit D M 、D L and D R Initialization of capacity hyperparameters; Step S13: Initialize the power-transport coupling system fault scenario, including: simulating the power-transport coupling network fault scenario after an extreme event occurs: randomly setting some lines to be interrupted in the power system to construct a set of damaged lines Simulate the interruption of weak power lines in real life after extreme events; and randomly set some road damage in the traffic system to construct a damaged road set Simulate the situation in real life where some roads in the traffic system are blocked by obstacles caused by extreme events; Step S14: Initialize the travel demand of urban traffic users, including: loading all urban traffic travel demand sets RS, setting the total travel demand of urban traffic users and the proportion coefficient of each travel demand rs to the total travel demand, and obtaining the basic travel demand q of each travel demand rs at different times rs,t ; Step S15: Initialize the number of training episodes e=1.
3. The method for improving the resilience of the power-transportation coupling system based on the optimal scheduling of multi-type mobile emergency resources according to claim 2 is characterized in that: Step S2 is specifically implemented as follows: Step S21: Initialize training time t=1; Step S22: Initialize various variables in the graph diffusion attention network reinforcement learning algorithm environment, including: load rate of each node in the power system, renewable energy output, traffic flow and speed of each road in the transportation system; initial position of the mobile energy storage vehicle, initial battery state of charge (SOC) of the mobile energy storage vehicle, initial speed of the mobile energy storage vehicle; initial position of the line repair team, repair materials of the line repair team, initial speed of the line repair team; initial position of the road repair team, repair materials of the road repair team, initial speed of the road repair team; Step S23: Initialize the fault scenario of the power-transportation coupling system, including: initializing the fault line set constructed in step S1 and fault road collection Step S24: Initialize the travel demand of urban traffic users, including: initializing the basic travel demand q of each urban traffic travel demand rs in step S1 rs,t .
4. The method for improving the resilience of the power-transportation coupling system based on the optimal scheduling of multi-type mobile emergency resources according to claim 1 is characterized in that: Step S3 is specifically implemented as follows: Step S31: Construct the status information of the mobile energy storage vehicle based on the power-traffic coupling system information and the mobile emergency resource status information. Including the mobile energy storage vehicle's own status information Adjacent mobile emergency resource status information Traffic system road information at your location Power system node load data P t ED , wind turbine output P t WT , specifically expressed as: Where, M, R l 、R r , T and N r are the mobile energy storage vehicle set, line repair team set, road repair team set, time set, and adjacent mobile emergency resource set; m and b are the numbers of the mobile energy storage vehicle, l, r, n, and t are the line repair team number, road repair team number, adjacent mobile emergency resource number, and simulation time period, respectively; Represents the location information of the mobile energy storage vehicle in the traffic system at time t; Indicates the moving speed of the mobile energy storage vehicle in the traffic system; S m,t Indicates the remaining SOC value of the mobile energy storage vehicle; and are the information of the mobile energy storage vehicle, line repair team and road repair team adjacent to the mth mobile energy storage vehicle at time t; N r is the number of adjacent emergency resources of the mobile energy storage vehicle; J is the number of adjacent roads of the mobile energy storage vehicle; Indicates the jth road number adjacent to the mth mobile energy storage vehicle at time t, and Indicates road The first and last node numbers of and Respectively represent roads road status, road length, road traffic volume and road travel time; Indicates the current traffic system node location number, Indicates the position number of the next node to be selected, Indicates the position number of the previous node, Indicates the road number to be entered and Indicates the previous road number; and S b,n,t They respectively represent the position information, moving speed and remaining SOC value of the nth mobile energy storage vehicle in the traffic system; RT l,n,t and RT l,n,t They represent the remaining number of repair materials of the nth line repair team and road repair team with mobile energy storage vehicle respectively; Step S32: Construct the status information of the line repair team based on the power-transportation coupling network information and the mobile emergency resource status information Including the status information of the line repair team itself Adjacent mobile emergency resource status information Traffic system road information at your location Information on faulty lines repaired by the line repair team and line maintenance time Specifically expressed as: Where, l and c are the numbers of the line repair teams; Step S33: Construct the status information of the road repair team based on the power-traffic coupling network information and the status information of the mobile emergency resources themselves Including the road repair team's own status information Adjacent mobile emergency resource status information Traffic system road information at your location Information on faulty roads repaired by the road repair team and road maintenance time Specifically expressed as: Where r and e are the numbers of the road repair teams; Step S34: Constructing a state matrix based on the state information of the mobile energy storage vehicle, the line repair team, and the road repair team Step 35: Construct the adjacency matrix A based on the adjacency relationships of all mobile emergency resources t , specifically expressed as: Where, N is the adjacency relationship formed by all mobile emergency resources at time t; v is the number of mobile emergency resources, through N v =|M|+|R l |+|R r |Calculate; is a binary variable, indicating the adjacency relationship between mobile emergency resources i and j. If they are adjacent, it is 1, otherwise it is 0. Adjacency matrix A t It will be input into the graph diffusion attention network reinforcement learning algorithm to participate in the calculation of the action value function.
5. The method for improving the resilience of the power-transportation coupling system based on the optimal scheduling of multi-type mobile emergency resources according to claim 1 is characterized in that: Step S4 is specifically implemented as follows: Step S41: First, calculate the random number p at time t t , whose value range is [0,1), and the ε-Greedy strategy is used to compare the current greedy ε t value and random number p t , if there is p t <ε t , then a random selection method is used to generate the actions of all mobile energy storage vehicles from the routing action set And form the routing behavior strategy of mobile energy storage vehicles All line repair teams' actions Line repair team routing behavior strategy All road repair crew actions Road repair team routing behavior strategy Specifically, as shown in formula (26): Where, Represents the routing action set of mobile energy storage vehicle routing, line repair team and road repair team; when the selected action a i,t = 0 means stopping at the original location, and also means that the mobile emergency resource route reaches the destination location, the mobile energy storage vehicle is the charging station or the initial charging station node, the line repair team is the fault line or warehouse node, and the road repair team is the fault road or warehouse node, ready to accept the subsequent model scheduling decision; when the action a is selected i,t ≠0 means that the mobile energy storage vehicle, line repair team, and road repair team choose to enter the next traffic road to their destination according to the routing strategy; Step S42: If there is p t ≥ε t , the routing behavior strategy of the mobile energy storage vehicle is generated by the graph diffusion attention network reinforcement learning algorithm Line repair team routing behavior strategy and the routing behavior strategy of the road repair team Specifically, as shown in formulas (27)-(29): In the formula, argmax() represents the parameter corresponding to the maximum value; and The multi-task structure of the neural network reinforcement learning algorithm for the graph diffusion attention network is based on the mobile emergency resource state matrix o t , mobile emergency resource adjacency matrix A t and the current network parameters θ t At the same time, the routing actions calculated for the mobile energy storage vehicle, line repair team and road repair team The value function of .
6. The method for improving the resilience of the power-transportation coupling system based on the optimal scheduling of multi-type mobile emergency resources according to claim 1 is characterized in that: The neural network of the diffusion attention network reinforcement learning algorithm is a single-input multi-output structure of a multi-task neural network. This structure is based on the mobile emergency resource state matrix o t and the adjacency matrix A t The action value functions of mobile energy storage vehicles, line repair teams, and road repair teams are simultaneously calculated, meaning that only one neural network is used to simultaneously control all mobile emergency resources to improve the resilience of the power-transportation coupling system. The neural network structure of the graph diffusion attention network reinforcement learning algorithm consists of a fully connected input layer, two graph diffusion attention layers, and three parallel output layers using the Dueling DQN strategy; first, the fully connected input layer inputs the mobile emergency resource state set o t Perform preliminary feature extraction and dimension conversion, output Secondly, the graph diffusion attention layer uses the generalized graph diffusion operation of PersonalizedPageRank theory based on the graph diffusion mechanism to move the original first-order neighborhood to the emergency resource adjacency matrix A t Sparse, and then extract the generalized adjacency relationship of mobile emergency resources to construct a graph diffusion matrix The specific calculation formula is: T sym =D -1 / 2 A t D -1 / 2 (31) S PPR =a ppr (I-(1-a) ppr )T sym ) -1 (32) Where D is the degree matrix, N v is the number of mobile emergency resources; T sym is a symmetric normalized transition matrix, the sum of each row is 1; S PPR and are the PPR diffusion matrix and the diffusion matrix after pruning operation respectively; α ppr The PPR parameter is also called the restart probability, which is used to control the range of information diffusion and is adjusted according to the task's dependence on the balance between global information and local information. Its value range is (0,1); ppr Approaching 1, diffusion tends to retain the original node information and reduce the influence of other nodes. ppr As it approaches 0, the diffusion tends to be more inclined to the global information propagation of the entire graph, and I is the unit matrix; ε spa is the pruning threshold, the diffusion matrix S PPR Less than ε spa The elements of are set to 0 to sparse the matrix; ε spa Set to [0.01, 0.1]; is the diffusion matrix The degree matrix of GDC is the final normalized diffusion matrix; Equations (30) and (31) calculate the transition matrix T sym ; Equations (32) and (33) calculate the diffusion matrix S PPR And perform pruning operation on it; Equations (34) and (35) calculate the normalized diffusion matrix A GDC , and used in the subsequent feature extraction of graph neural networks; In addition, the graph diffusion attention layer uses a dynamic attention mechanism to the state matrix Sum diffusion matrix A GDC Calculate the attention coefficient and output feature information based on the graph diffusion attention mechanism The specific calculation formula is: Where σ() is the ReLU nonlinear activation function; A collection of adjacent mobile emergency resources; is the weight vector; Represents the weight matrix for linear transformation; || represents the merge operation; represents the normalized graph diffusion attention coefficient of mobile emergency resource j to mobile emergency resource i; Formula (38) calls K groups of independent attention mechanism layers and averages the calculated K groups of attention coefficients to obtain the final feature information output The learning process of diffuse attention networks with stable graphs; Finally, the neural network constructs three parallel output layers to calculate the action Q value for the mobile energy storage vehicle, line repair team, and road repair team respectively; each output layer adopts the strategy of Dueling DQN algorithm to divert the feature information to two fully connected branches: the first branch outputs the scalar value of the state function The second branch outputs the action advantage value function vector Therefore, the calculation of the action Q value is expressed as: Where, ψ t for The fully connected neural network parameters of the branch; θ t for The fully connected neural network parameters of the branch; the mobile energy storage vehicle, line repair team and road repair team select routing actions according to the action value function calculated by their respective neural networks and 7. The method for improving the resilience of the power-transportation coupling system based on the optimal scheduling of multi-type mobile emergency resources according to claim 1 is characterized in that: Step S6 is specifically implemented as follows: Step S61: Constructing an urban traffic travel allocation model based on improved semi-dynamic user equilibrium, as shown below: Where, L TN and They represent the traffic system roads and fault road sets respectively; RS represents the user travel start-end pair set, Where r is the starting point of the user's trip, s is the end point of the trip; K rs,t represents the set of feasible paths for the travel demand rs at time t; k represents a feasible path in the set of feasible paths; α tn and are all road delay coefficients; M is an integer much larger than the basic road travel time; is the path traffic flow of feasible path k; is a binary variable. If the traffic road l tn Set to 1 when it is on a feasible path k, otherwise it is 0; represents the travel time of feasible path k; r t rs and They are respectively represented as the stranded vehicle flow of travel demand rs at time t and t-1; q rs,t and They represent the travel demand rs at time t and the revised travel demand respectively; Represents the travel demand rs in the current feasible path set K rs,t The minimum travel time cost under the condition of faulty roads; as the faulty roads are repaired, the feasible path set K of travel demand rs rs,t It will also change accordingly, and a new feasible path set with a smaller travel time cost than the previous one will be obtained; Formula (40) indicates that the travel time of the faulty road is a number that is much larger than the basic travel time and tends to infinity, and traffic users cannot pass through this road; Formula (41) calculates the traffic flow of the traffic road by searching all feasible paths under the user travel demand set; Formula (42) calculates the travel time of the feasible path k; Formula (43) calculates the travel demand rs in the feasible path set K rs,t The stranded vehicle flow under the condition of rs,t The modified travel demand of travel demand rs at time t is calculated by the traffic flow at time t-1 and time t. Formula (45) shows that the sum of the traffic flow on all feasible paths of travel demand rs is equal to the corresponding modified travel demand. Formula (46) ensures that the traffic flow on each feasible path is non-negative and the actual travel time will not be less than K under the current feasible path set. rs,t The minimum travel time cost; Step S62: Initialize the number of iterations of the continuous averaging algorithm considering random utility n=1; Step S63: Using the continuous average algorithm considering random utility to solve the urban traffic travel allocation model based on the improved semi-dynamic user equilibrium, obtain the road traffic flow distribution and the stranded traffic flow r t rs and user travel time costs The iterative solution process is as follows: (1) At time t, the road fault status of the current traffic system is updated based on the road repair decision made by the road repair team at time t-1; (2) Based on the current road fault status, the Dijkstra shortest path method is used to solve the feasible path set of all user travel demands rs; (3) performing the nth iteration of the continuous averaging algorithm taking into account random utilities; (4) All the travel demands of traffic users q rs,t Assigned to the feasible path set K rs,t In the calculation, the traffic system road flow and the flow of each feasible path are obtained Update all road travel times; (5) Solve the traffic flow rate r t rs , and based on the stranded traffic flow at time t-1 Calculate the modified travel demand of rs; (6) Calculate the selection probability of the feasible path k for travel demand rs based on the Logit random utility model Specifically, it is shown in the following formula (47): Where δ is a scale parameter, which represents the user's familiarity with traffic network information; is the feasible path set K of user travel demand rs at time t rs,t The time cost of path k; (7) Based on the current iteration number n, the path flow of the feasible path k of the travel demand rs is updated through formula (48) (8) Determine whether the convergence condition is met by equation (49); if so, proceed to step (9); otherwise, return to step (3): Where, ε logit Indicates the convergence accuracy, which is a fixed value; The parameter n represents the current number of iterations; Formula (49) represents the number of feasible paths assigned to the set K. rs,t When the difference between the proportion of traffic flow of any feasible path k in the total travel demand and the path selection probability calculated by the Logit random utility model is small enough, the traffic flow distribution of the entire model reaches equilibrium; (9) The traffic distribution model is solved and the output is the traffic flow distribution, road travel time, and stranded traffic flow r of the traffic system at time t. t rs and user travel time costs 8. The method for improving the resilience of the power-transportation coupling system based on the optimal scheduling of multi-type mobile emergency resources according to claim 1 is characterized in that: Step S7 is specifically implemented as follows: Step S71: Constructing a mobile energy storage vehicle dispatching model, a power system reconstruction model, and a power system optimal power flow model, as shown below: Where, Equations (50)-(56) are the mobile energy storage vehicle scheduling models; Represents the set of charging station nodes in the transportation system; is a binary variable, indicating whether the mobile energy storage vehicle is located at the traffic system node n tn , if it is at the node, it is 1, otherwise it is 0; and are binary variables, representing the charging and discharging status of the mth mobile energy storage vehicle at time t; and They represent the charging active power, discharging active power and reactive power at time t respectively; and Respectively represent the upper limit of active power and reactive power; S m,t represents the SOC at time t; and are the charging efficiency coefficient and the discharging efficiency coefficient respectively; S m and are the upper and lower limits of the SOC value respectively; Formulas (50)-(52) limit the charging and discharging power and reactive power of the mobile energy storage vehicle; Formula (53) limits the charging and discharging behavior of the mobile energy storage vehicle, that is, the mobile energy storage vehicle can only be in the charging state or the discharging state. At the same time, if the current location of the mobile energy storage vehicle is not located at the charging station node of the transportation system, it cannot be charged or discharged; Formulas (54)-(55) constrain the SOC value of the mobile energy storage vehicle; Formula (56) indicates that after the partial fault line of the power system is repaired, the power system load can be fully restored, and at this time there is no need to dispatch the mobile energy storage vehicle to perform the load support task; Where, Equations (57)-(62) are the power system reconstruction models; L DN 、 L DN,sw and denote the set of power system lines, the set of damaged lines, the set of lines equipped with tie switches, and the set of power system source nodes, respectively; ρ(·) and φ(·) denote the set of parent nodes and the set of child nodes, respectively; α ij,t Indicates the connection status of the power system line (i, j) at time t, which is 1 if connected and 0 if disconnected; N DN represents the number of nodes in the power system; γ j,t is a binary variable, which is 1 if the power system node j is the source node at time t, otherwise it is 0; f ij,t represents the virtual power flow into the power system node j at time t; M is a constant with a large value; constraint (57) sets the line state constraint based on the power system line fault in step S2; constraint (58) is the power system radiation constraint, which describes the relationship between the power system line state and the power system source node; constraints (59)-(62) describe the virtual power flow balance constraint of the power system; Where, Equations (63)-(73) are the optimal power flow models of the power system; and They represent the active power and reactive power flowing through the power system line (i, j) at time t, respectively; and Respectively represent the active and reactive outputs of the wind turbine WT; and They represent the active power exchange amount and reactive power exchange amount between the mobile energy storage vehicle and node j respectively; and Both represent the load loss of node j; and They represent the load demand of node j respectively; and They represent the maximum active power and maximum reactive power of the power system line (i, j) respectively; represents the upper capacity limit of the power system line (i, j); represents the square value of the voltage at the power system node j at time t; r ij and x ij Respectively represent the line resistance and reactance values of the power system line (i, j); and They represent the upper and lower limits of the square value of the voltage of the power system node j at time t respectively; constraints (63) and (64) describe the active power balance and reactive power balance of the power system line; constraints (65) and (66) describe the active power and reactive power of the mobile energy storage vehicle charging or discharging at the charging station node at time t; constraint (67) limits the active power and reactive power of the mobile energy storage vehicle charging and discharging to 0 when it is not at the charging station node; constraints (68) and (69) limit the active power and reactive power flowing through the power system line at time t; constraint (70) limits the thermal capacity of the line (i, j); constraints (71)-(73) describe the voltage relationship between adjacent nodes of the power system; Step S72: Construct a line repair team dispatch model, as shown below: Where, For the line repair team to repair the power system fault line at time t The repair decision variable is 1 if the line repair team chooses to repair the line, otherwise it is 0; Indicates the location of the traffic node where the fault line is located n dn ; The time required for the line repair team to repair the faulty line; is the dispatch flag of the line repair team at time t. If it is 0, the dispatch continues. Constraint (74) ensures that the line repair team must arrive at the traffic node where the faulty line is located before repairing it. Constraint (75) restricts the fault line to be repaired by only one line repair team at the same time; Constraint (76) describes the number of fault lines that have been repaired by the line repair team at time t. rep The faulty line at time t can be restored to normal operation at time t+1; constraint (77) restricts the faulty line that has been repaired from failing again in the future; Constraint (78) indicates that when all faulty lines are repaired and the line repair team arrives at the line repair team warehouse at time t, the dispatch flag is set to 1, and there is no need to dispatch the line repair team to perform the fault repair task. Step S73: Construct a road repair team dispatch model, as shown below: Where, and Both are binary variables. If the road repair team is located at the traffic system node n tn Or road l tn If yes, it is 1, otherwise it is 0; It is a binary variable, indicating the road repair team's response to the faulty road. The repair decision is 1 if it is repaired, otherwise it is 0; Is a binary variable, indicating the road at time t Status; To repair damaged roads Time required; A binary variable used to determine whether to continue dispatching the road repair team arriving at the warehouse. The road repair team dispatch model is similar to the line repair team dispatch model. Step S74: Taking the minimization of the total loss cost of the power-transportation coupling system as the objective function, a mathematical model for improving the resilience of the power-transportation coupling system taking into account urban traffic distribution is constructed, as shown below: Where c dn 、c tn and c mer They are the unit load loss cost of the power system, the travel time cost of traffic users and the unit time dispatch cost of mobile emergency resources; the power system loss cost Calculated using load reduction at time t; Traffic system loss cost Calculated by formula (86), the travel cost of traffic users is and mobile emergency resource dispatch costs It consists of two parts; among them, formula (87) calculates the travel cost of traffic users by multiplying the travel demand of rs at time t by the travel cost increment. That is, minimizing the travel cost of traffic users in the model is actually minimizing the travel cost increment of all users; formula (88) calculates the dispatching cost of mobile energy storage vehicles, line repair teams and road repair teams at time t; Step S75: Constraints (65) and (66) are linearized using the large M method, and constraint (70) is linearized using the second-order cone relaxation method; constraints (75) and (76) as well as constraints (80) and (81) of the line repair team and road repair team scheduling models are linearized using the McCormick envelope method; Step S76: Using equations (84)-(88) as the objective function and equations (50)-(83) as the constraints, a mixed integer second-order cone programming model MISOCP is constructed and solved using the Gurobi commercial solver; the node load of the power system restored by the mobile energy storage vehicle is obtained, and the mobile energy storage vehicle reward function is calculated using equations (89) and (90). Where, represents the set of power system nodes whose loads are supported by mobile energy storage vehicles at time t; Power system node The load recovery amount; is the load level weight coefficient of the power system node; By solving MISOCP, we can obtain the line maintenance strategy or scheduling strategy of the line repair team, and calculate the reward function of the road repair team through equations (91)-(95) Where, Represents the set of faulty lines repaired by the line repair team; It represents the travel time from the starting node b to the target node p of the traffic system; The contribution reward value of the line repair team for repairing the faulty line at time t; Rewards for the line repair team to complete the task, including: Reward for completing fault line repair, and The rewards for returning to the warehouse are all constant; is the penalty value of the line repair team at time t; It is a binary variable, which means that it is 1 when the road selected at time t is the same as that at time t-1, and 0 otherwise; and are all constants; and All are weight coefficients; Indicates the current flowing through the fault line repaired by the line repair team at time t The active power of the line repair team at time t is calculated by the sum of the total active power flowing through the fault line repaired by the line repair team. The contribution reward value of the line repair team for repairing the fault line is calculated by formula (92). The reward value of the line repair team after repairing the fault line and returning to the warehouse is a fixed value. The penalty value of the line repair team at time t is calculated by formula (93). Specifically, it includes the penalty for going back and the penalty for the line repair team repairing all fault lines but failing to return to the warehouse. And the penalty for the line repair team failing to repair all faulty lines Three parts; By solving MISOCP, we can obtain the road maintenance strategy or scheduling strategy of the road repair team, and calculate the reward function of the road repair team through equations (96)-(100): Where, The road repaired by the road repair team Contribution reward value; Formula (97) is calculated by the ratio of the sum of traffic flows on the faulty roads repaired by the road repair team at time t to the sum of traffic flows on all roads in the traffic system.
9. The method for improving the resilience of the power-transportation coupling system based on the optimal scheduling of multi-type mobile emergency resources according to claim 1 is characterized in that: Step S10 is specifically implemented as follows: Step S101: The input layer parameter of the graph diffusion attention network reinforcement learning algorithm neural network is recorded as θ In , the parameters of the first and second layer diffusion attention layers are denoted as θ G1 and θ G2 , the three parallel output layer parameters are recorded as θ m ,θ l and θ r ; The neural network parameters for the forward propagation and reverse update of the Q value of the mobile energy storage vehicle agent are recorded as θ In +θ G1 +θ G2 +θ m →θ M ; The neural network parameters for the line repair team and the road repair team are respectively denoted as θ In +θ G1 +θ G2 +θ l →θ L and θ In +θ G1 +θ G2 +θ r →θ R ; Step S102: In the mobile energy storage vehicle memory unit D M In the importance sampling method, samples N are collected respectively s , and calculate the importance weight of each sample by formula (101): Where N D Represents memory unit D M The total number of samples in ; Indicates the priority of the sample; β is a hyperparameter; Step S103: Combine the Double DQN algorithm strategy and adopt the current network θ t To select the optimal action of the agent at the next moment, the target network To evaluate the value of an action, we make full use of the two neural networks of the DQN algorithm to separate action selection and strategy evaluation to reduce the risk of overestimating the Q value. We use equations (102) and (103) to calculate the loss function: Where θ t and are the current network parameters and target network parameters of the graph diffusion attention network reinforcement learning algorithm respectively; γ is the discount factor reflecting the impact of the future action value Q value on the current action, and its value is [0,1]; Step S104: Optimize the neural network parameters θ of the mobile energy storage vehicle based on the stochastic gradient descent method and the sample importance weight. t M Update, specifically: Where, α M is the learning rate of the gradient descent algorithm of the current network of the mobile energy storage vehicle; and after a fixed number of steps N up Then, the parameters of the target network of the mobile energy storage vehicle are adjusted in the following way. To update: Step S105: TD-error value calculated according to formula (102) Calculate the new priority for each sample Among them, ∈ is a small constant to prevent the priority from being zero; in the mobile energy storage vehicle, the reinforcement learning memory unit D M Update the priority of each sample in Step S106: Refer to steps S102 to S105 to calculate the neural network parameters of the line repair team Make updates; Step S107: Refer to steps S102 to S105 to calculate the neural network parameters of the road repair team. to update.
10. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 9 can be implemented.
Citation Information
Cited By
Power-traffic coupling network scene generation method based on topology embedded dynamics
CN121882838A