Pre-disaster deployment method of multi-region electricity-hydrogen integrated energy system mobile power supply based on deep reinforcement learning
The deep reinforcement learning framework optimizes mobile power source deployment in multi-regional electric-hydrogen energy systems, addressing inefficiencies in existing methods by integrating distributed power sources and network reconstruction to enhance system resilience during disasters.
Patent Information
- Application Number
- CN202510166524.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-15
AI Technical Summary
In the pre-disaster deployment of multi-regional electricity-hydrogen integrated energy systems, traditional methods such as stochastic planning and robust optimization cannot effectively utilize the failure probability information predicted before disasters, resulting in too conservative decision-making or difficulty in fully considering the failure uncertainty under extreme disasters while ensuring solution efficiency. The application of deep reinforcement learning in this field has not been studied.
Using a method based on deep reinforcement learning, a pre-disaster deployment framework for mobile power supply in multi-regional electricity-hydrogen integrated energy system is established, and an in-disaster emergency response model with coordinated scheduling of distributed power supply and network reconstruction is built. Using deep reinforcement learning network to solve the objective functions and constraints, and optimizing the deployment decisions of mobile power supply and hydrogen energy systems.
On the premise of ensuring solution efficiency, the cross-regional support potential of mobile power supply and the uncertainty of multi-regional failures are fully considered, and the elasticity and post-disaster recovery capabilities of multi-regional electricity-hydrogen integrated energy systems are improved.
Smart Images

Figure CN120317698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of integrated energy system applications, and specifically to a pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning. Background Technique
[0002] As a secondary energy source with the most promising development prospects in the 21st century, hydrogen energy, due to its unique cross-time-and-space flexible adjustment characteristics, can not only enhance the flexibility of the distribution network, but also enhance the resilience of the distribution network through hydrogen fuel cell vehicles and hydrogen tube trailers in extreme disasters. In addition, the highly redundant road network provides a more reliable option for hydrogen energy scheduling and timely migration between different nodes and regions. Hydrogen energy shows great potential in enhancing the resilience of the distribution network. Currently, no research has focused on the pre-disaster deployment strategy of mobile power sources in a multi-region electric-hydrogen integrated energy system, and the pre-disaster cross-region scheduling mechanism of the multi-region electric-hydrogen integrated energy system is not clear. In addition, the pre-disaster deployment of the multi-region electric-hydrogen integrated energy system involves the deployment of mobile power sources within multiple regions, including a large amount of uncertainty of faulty lines. Traditional methods mainly use stochastic programming or robust optimization to solve the above-mentioned pre-disaster deployment problem of mobile power sources. However, robust optimization only optimizes for the worst-case scenario and cannot effectively utilize the fault probability information obtained from pre-disaster prediction, resulting in overly conservative decisions. Stochastic programming optimizes the expected optimal results under multiple typical scenarios by generating and aggregating scenarios, which represent uncertainties, but this only represents the optimal under typical scenarios and does not mean that the optimal can be achieved in actual situations. In addition, for each additional scenario considered in stochastic programming, the problem scale doubles, and the quality of its solution highly depends on the representativeness of the aggregated scenarios. Therefore, it is difficult to fully consider the impact of multi-region fault uncertainties on the scheduling results under extreme disasters while ensuring the solution efficiency.
[0003] Deep reinforcement learning is a data-driven method that can learn the optimal strategy through repeated interactions with the environment without any prior knowledge. Deep reinforcement learning can effectively utilize the growing amount of data in the environment to capture various uncertainties. In addition, a well-trained strategy can be directly deployed into the actual test process within a few milliseconds. Recently, deep reinforcement learning has been successfully applied to many power system operation problems, such as voltage regulation and economic dispatch. However, there is currently no research applying deep reinforcement learning to solve the pre-disaster deployment problem of mobile power sources. Summary of the Invention
[0004] The object of the present invention is to provide a pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, including the following steps:
[0005] 1) Consider the cross-regional support potential of mobile power sources and establish a pre-disaster deployment framework for mobile power sources in a multi-region electric-hydrogen integrated energy system considering regional mutual assistance;
[0006] 2) Based on the pre-disaster deployment framework of mobile power sources in a multi-region electric-hydrogen integrated energy system, construct a disaster response model considering the coordinated scheduling of distributed power sources and network reconfiguration;
[0007] 3) Use a deep reinforcement learning network to solve the disaster response model considering the coordinated scheduling of distributed power sources and network reconfiguration, and obtain the pre-disaster deployment decision of the multi-region electric-hydrogen integrated energy system.
[0008] Furthermore, the pre-disaster deployment framework of mobile power sources in the multi-region electric-hydrogen integrated energy system includes mobile power sources and distribution networks;
[0009] The mobile power sources are deployed at the distribution network access points within each region.
[0010] Furthermore, the objective function of the disaster response model considering the coordinated scheduling of distributed power sources and network reconfiguration is as follows:
[0011] min(C lose,E +C lose,H +C gen,E +C gen,H ) (1)
[0012]
[0013] In the formula, C lose,E is the total penalty for the cut-off electric load; C lose,H is the total penalty for the cut-off hydrogen load; C gen,E is the total operating cost of mobile electrical energy storage; C gen,H is the total operating cost of hydrogen fuel cell vehicles and hydrogen fuel cells in the hydrogen energy system; N E , N T1 represent the distribution network nodes and the set of scheduling periods within the region respectively; c lose,Ei is the unit cut-off electric load penalty at node i; P lose,Ei,t is the cut-off electric load at node i at time t; c lose,H is the unit cut-off hydrogen load penalty; M lose,Ht is the cut-off hydrogen load at time t; N EV / N HEV / N RC represent the sets composed of mobile electrical energy storage / hydrogen fuel cell vehicles / maintenance personnel respectively; N M represents the set of grid access points; c gen,E is the unit operating cost of mobile electrical energy storage, and P EV,outi,t,n is the discharge power of the nth mobile electrical energy storage located at node i at time t; c gen,His the unit operating cost of the hydrogen fuel cell, P HEV,outi,t,n is the discharge power of the nth hydrogen fuel cell vehicle at node i and time t; P FCt is the discharge power of the hydrogen fuel cell at time t in the hydrogen energy system.
[0014] Furthermore, the constraint conditions of the disaster response model considering the coordinated scheduling of distributed power sources and network reconfiguration include mobile power operation constraints, hydrogen energy system operation constraints, network reconfiguration constraints, load shedding balance constraints, and distribution network power flow constraints.
[0015] Furthermore, the mobile power operation constraints include the uniqueness constraint of the charging / discharging state of mobile electrical energy storage, the charging / discharging power constraint of mobile electrical energy storage, the energy balance constraint of mobile electrical energy storage, the upper / lower limit constraint of the energy storage of mobile electrical energy storage, the uniqueness constraint of the charging / discharging state of hydrogen fuel cell vehicles, the charging / discharging power constraint of hydrogen fuel cell vehicles, the energy balance constraint of hydrogen fuel cell vehicles, the upper / lower limit constraint of the energy storage of hydrogen fuel cell vehicles, and the maximum number of mobile power sources connected to the grid at the grid connection point constraint;
[0016] Among them, the uniqueness constraint of the charging / discharging state of mobile electrical energy storage is as follows:
[0017]
[0018] The charging / discharging power constraint of mobile electrical energy storage is as follows:
[0019]
[0020] The energy balance constraint of mobile electrical energy storage is as follows:
[0021]
[0022] The upper / lower limit constraint of the energy storage of mobile electrical energy storage is as follows:
[0023]
[0024] The uniqueness constraint of the charging / discharging state of hydrogen fuel cell vehicles is as follows:
[0025]
[0026] The charging / discharging power constraint of hydrogen fuel cell vehicles is as follows:
[0027]
[0028] The energy balance constraint of hydrogen fuel cell vehicles is as follows:
[0029]
[0030] The energy storage upper / lower limits of the hydrogen fuel power generation vehicle are as follows:
[0031]
[0032] The constraint on the maximum number of mobile power sources connected to the power grid at the grid connection point is as follows:
[0033]
[0034] In the formula, u EV,ini,t,n / u EV,outi,t,n are the state variables of the charging / discharging energy of the nth mobile energy storage at time t at node i. If it is 1, it means that the nth mobile energy storage is in the charging / discharging state at node i at time t; P EV,ini,t,n / P EV,outi,t,n are the charging / discharging powers of the nth mobile energy storage at time t at node i; P EV,in,max / P EV,out,max are the upper limits of the charging / discharging powers of the mobile power source; S EVt,n is the energy storage level of the nth mobile energy storage at time t; η EV,in / η EV,out are the charging / discharging efficiencies of the mobile energy storage; S EV,max / S EV,min are the upper / lower limits of the energy storage of the mobile energy storage; u HEV,ini,t,n / u HEV,outi,t,n are the state variables of the charging / discharging energy of the nth hydrogen fuel power generation vehicle at time t at node i. If it is 1, it means that the nth hydrogen fuel power generation vehicle is in the charging / discharging state at node i at time t; M HEV,ini,t,n / P HEV,outi,t,n are the charging mass / discharging power of the nth hydrogen fuel power generation vehicle at time t at node i; M HEV,in,max / P HEV,out,max are the upper limits of the charging mass / discharging power of the hydrogen fuel power generation vehicle; S HEVt,n is the energy storage level of the nth hydrogen fuel power generation vehicle at time t; η HEV is the power generation efficiency of the hydrogen fuel power generation vehicle; S HEV,max / S HEV,min are the upper / lower limits of the energy storage of the hydrogen fuel power generation vehicle.
[0035] Furthermore, the operation constraints of the hydrogen energy system include the uniqueness constraints of the operation states of the electrolyzer and the hydrogen fuel cell, the operation power constraint of the electrolyzer, the operation power constraint of the hydrogen fuel cell, the upper / lower limits of the energy storage of the hydrogen storage tank, and the energy balance constraint of the hydrogen storage tank;
[0036] The uniqueness constraints of the operation states of the electrolyzer and the hydrogen fuel cell are as follows:
[0037]
[0038] The operating power constraint of the electrolyzer is as follows:
[0039]
[0040] The operating power constraint of the hydrogen fuel cell is as follows:
[0041]
[0042] The upper / lower limit constraints of the energy storage of the hydrogen storage tank are as follows:
[0043]
[0044] The energy balance constraint of the hydrogen storage tank is as follows:
[0045]
[0046] Where, u EDt / u FCt are the operating state variables of the electrolyzer and the hydrogen fuel cell at time t, respectively. If it is 1, it means that the electrolyzer / hydrogen fuel cell is in the working state at time t; are the operating powers of the electrolyzer / hydrogen fuel cell at time t, respectively; P ED,max / P FC,max are the maximum operating powers of the electrolyzer / hydrogen fuel cell, respectively; η ED / η FC are the energy conversion efficiencies of the electrolyzer / hydrogen fuel cell, respectively; LHV H is the lower heating value of hydrogen; is the hydrogen storage capacity of the hydrogen storage tank at time t; M load,Ht is the hydrogen load at time t; S HS,min / S HS,max are the minimum / maximum hydrogen storage masses of the hydrogen storage tank, respectively; η HS,in / η HS,out are the hydrogen charging / discharging efficiencies of the hydrogen storage tank, respectively; If there is only a hydrogen refueling station deployed in the area, then u EDt and u FCt are both 0.
[0047] Furthermore, the network reconstruction constraint is as follows:
[0048]
[0049] Where, δ Fij,t is the closing state variable of line ij in the virtual power flow at time t; δ Fi,t is the state variable indicating whether the potential root node i serves as the root node of the microgrid at time t; M is a sufficiently large constant; P Fij,t is the virtual power flow through branch ij at time t.
[0050] Furthermore, the load shedding balance constraint is as follows:
[0051]
[0052] where P load,Ei,t is the demand electrical load at node i at time t; P re,Ei,t is the actual supply electrical load at node i at time t; M load,Ht is the demand hydrogen load at time t; M re,Ht is the actual supply hydrogen load at time t.
[0053] Furthermore, the distribution network power flow constraint is as follows:
[0054]
[0055] where P ij,t , Q ij,t are the active power and reactive power of line ij at time t, respectively. Q maxij are the upper limits of the active power and reactive power of line ij, respectively; r ij , x ij are the resistance and reactance of line ij, respectively. is the total active and reactive power injected by distributed power sources at node j at time t; is the square of the voltage at node i at time t.
[0056] Furthermore, the deep reinforcement learning network includes an action network and an evaluation network;
[0057] The state space of the action network is as follows:
[0058] o = [U branch,all , P load,all , P H,all , Q H,all , P EV,all , Q EV,all , P HEV,all , Q HEV,all (29)
[0059] where o represents the state space of the agent, which includes the topological state information U branch,all of all regions, the load prediction information P load,all , the maximum power generation information P H,all of the hydrogen energy system, the hydrogen storage capacity information Q H,all , and the maximum power generation power information P EV,all of all mobile electrical energy storage, the mobile electrical energy storage capacity information Q EV,all , the maximum power generation power information P HEV,all of the hydrogen fuel power generation vehicle, and the hydrogen fuel power generation vehicle capacity information QHEV,all 。
[0060] The action space of the action network is as follows:
[0061]
[0062] In the formula, a is the agent's action space, which includes all deployment results a of mobile electric energy storage EVj and all deployment results a of emergency hydrogen fuel power generation vehicles HEVj , discrete action a u,EVj ∈
[0063] {1, 2, …, N M} Selecting 1 means deploying the jth mobile electric energy storage to the first grid connection
[0064] point, and the same applies to emergency hydrogen fuel power generation vehicles;
[0065] The reward function of the evaluation network is as follows:
[0066]
[0067] In the formula, the agent reward function r up is the sum of all regional load shedding losses, r downi is the reward function of the ith lower-level agent, which is the load shedding loss of region i, C lose,Ei is for region i
[0068] total penalty for cutting electric load; C lose,Hi is the total penalty for cutting hydrogen load in region i;
[0069] The loss function of the evaluation network is as follows:
[0070]
[0071]
[0072] In the formula, x n,j is the preference value corresponding to the jth action dimension in the nth action output by the Actor network; p n,j is the discrete action probability; D experience pool, α H is the temperature coefficient, which is used to control the weight of the entropy regularization term in the loss function; is the entropy regularization term; A is the total number of agent actions, K n are the total action latitudes of the nth action respectively; H tar is the target entropy of the agent;
[0073] During the training process of the deep reinforcement learning network, the parameters are updated as follows:
[0074]
[0075] In the formula, α represents the learning rate of the agent; represents the gradient; θ m represents the parameters of the policy network and the value network.
[0076] The technical effect of the present invention is beyond doubt. The present invention can fully consider the cross-regional support potential of mobile power sources, and at the same time fully characterize the uncertainty of multi-regional faults on the premise of ensuring the solution efficiency, improving the resilience of the multi-regional electricity-hydrogen integrated energy system. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 is the pre-disaster deployment framework of mobile power sources for a multi-regional electricity-hydrogen integrated energy system considering regional mutual assistance;
[0078] Figure 2 is the network structure of the SAC algorithm;
[0079] Figure 3 is the multi-regional electricity-hydrogen-transportation network coupling system;
[0080] Figure 4 are the electrical load curves of each region;
[0081] Figure 5 are the hydrogen load curves of each region;
[0082] Figure 6 is the convergence situation of the SAC deep reinforcement learning method under 3000 generations of training;
[0083] Figure 7 are the load shedding losses at all levels in different regions under two cases;
[0084] Figure 8 is the line fault situation and mobile power source deployment result of Scenario 1;
[0085] Figure 9 is the line fault situation and mobile power source deployment result of Scenario 2. DETAILED DESCRIPTION OF THE INVENTION
[0086] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above-mentioned subject scope of the present invention is limited to the following embodiments. Without departing from the above-mentioned technical idea of the present invention, various substitutions and changes made according to ordinary technical knowledge and customary means in the art shall be included within the protection scope of the present invention.
[0087] Embodiment 1:
[0088] See Figures 1 to 9, A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, comprising the following steps:
[0089] 1) Considering the cross-region support potential of mobile power sources, establish a pre-disaster deployment framework for mobile power sources in a multi-region electric-hydrogen integrated energy system considering regional mutual assistance;
[0090] 2) Based on the pre-disaster deployment framework for mobile power sources in a multi-region electric-hydrogen integrated energy system, construct a disaster response model considering the coordinated scheduling of distributed power sources and network reconfiguration;
[0091] 3) Use a deep reinforcement learning network to solve the disaster response model considering the coordinated scheduling of distributed power sources and network reconfiguration, and obtain the pre-disaster deployment decision for the multi-region electric-hydrogen integrated energy system.
[0092] Example 2:
[0093] A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, with the technical content the same as in Example 1. Further, the pre-disaster deployment framework for mobile power sources in the multi-region electric-hydrogen integrated energy system includes mobile power sources and distribution networks;
[0094] The mobile power sources are deployed to the distribution network access points within each region.
[0095] Example 3:
[0096] A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, with the technical content the same as in any one of Examples 1-2. Further, the objective function of the disaster response model considering the coordinated scheduling of distributed power sources and network reconfiguration is as follows:
[0097] min(C lose,E +C lose,H +C gen,E +C gen,H ) (1)
[0098]
[0099] In the formula, C lose,E is the total penalty for the cut-off electrical load; C lose,H is the total penalty for the cut-off hydrogen load; C gen,E is the total operating cost of the mobile energy storage; C gen,H is the total operating cost of the hydrogen fuel cell vehicle and the hydrogen fuel cell in the hydrogen energy system; N E , N T1 respectively represent the distribution network nodes and the set of scheduling time periods within the region; c lose,Ei is the unit cut-off electrical load penalty at node i; P lose,Ei,t is the cut-off electrical load at node i at time t; close,H is the unit hydrogen load penalty; M lose,Ht is the hydrogen load at time t; N EV / N HEV / N RC respectively represent the sets composed of mobile electric energy storage / hydrogen fuel power generation vehicles / maintenance personnel; N M represents the set of grid access points; c gen,E is the unit operating cost of mobile electric energy storage, P EV,outi,t,n is the discharge power of the nth mobile electric energy storage at node i at time t; c gen,H is the unit operating cost of hydrogen fuel cells, P HEV,outi,t,n is the discharge power of the nth hydrogen fuel power generation vehicle at node i at time t; P FCt is the discharge power of the hydrogen fuel cells in the hydrogen energy system at time t.
[0100] Example 4:
[0101] A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, the technical content is the same as any one of Examples 1-3. Further, the constraint conditions of the in-disaster emergency response model considering the coordinated scheduling of distributed power sources and network reconstruction include mobile power source operation constraints, hydrogen energy system operation constraints, network reconstruction constraints, load shedding balance constraints, and distribution network power flow constraints.
[0102] Example 5:
[0103] A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, the technical content is the same as any one of Examples 1-4. Further, the mobile power source operation constraints include the uniqueness constraint of the charge / discharge energy state of mobile electric energy storage, the charge / discharge power constraint of mobile electric energy storage, the energy balance constraint of mobile electric energy storage, the upper / lower limit constraint of the energy storage of mobile electric energy storage, the uniqueness constraint of the charge / discharge energy state of hydrogen fuel power generation vehicles, the charge / discharge power constraint of hydrogen fuel power generation vehicles, the energy balance constraint of hydrogen fuel power generation vehicles, the upper / lower limit constraint of the energy storage of hydrogen fuel power generation vehicles, and the maximum number of mobile power sources connected to the grid at the grid access point limit constraint;
[0104] Among them, the uniqueness constraint of the charge / discharge energy state of mobile electric energy storage is as follows:
[0105]
[0106] The charge / discharge power constraint of mobile electric energy storage is as follows:
[0107]
[0108] The energy balance constraint of mobile electric energy storage is as follows:
[0109]
[0110] The energy storage upper / lower limit constraints for mobile electrical energy storage are as follows:
[0111]
[0112] The uniqueness constraint for the charging / discharging energy state of a hydrogen fuel cell electric vehicle is as follows:
[0113]
[0114] The charging / discharging power constraints for a hydrogen fuel cell electric vehicle are as follows:
[0115]
[0116] The energy balance constraint for a hydrogen fuel cell electric vehicle is as follows:
[0117]
[0118] The energy storage upper / lower limit constraints for a hydrogen fuel cell electric vehicle are as follows:
[0119]
[0120] The maximum number of mobile power sources connected to the grid at the grid connection point is restricted as follows:
[0121]
[0122] where u EV,ini,t,n / u EV,outi,t,n are the state variables of the nth mobile electrical energy storage for charging / discharging at node i at time t. If it is 1, it means the nth mobile electrical energy storage is in the charging / discharging state at node i at time t; P EV,ini,t,n / P EV,outi,t,n are the charging / discharging powers of the nth mobile electrical energy storage at node i at time t; P EV,in,max / P EV,out,max are the upper limits of the charging / discharging powers of the mobile power source; S EVt,n is the energy storage level of the nth mobile electrical energy storage at time t; η EV,in / η EV,out are the charging / discharging efficiencies of the mobile electrical energy storage; S EV,max / S EV,min are the upper / lower limits of the energy storage of the mobile electrical energy storage; u HEV,ini,t,n / u HEV,outi,t,n are the state variables of the nth hydrogen fuel cell electric vehicle for charging / discharging at node i at time t. If it is 1, it means the nth hydrogen fuel cell electric vehicle is in the charging / discharging state at node i at time t; M HEV,ini,t,n / P HEV,outi,t,n are the charging mass / discharging power of the nth hydrogen fuel cell electric vehicle at node i at time t; MHEV,in,max / P HEV,out,max are the upper limits of the charging mass / discharging power of the hydrogen fuel cell vehicle; S HEVt,n is the energy storage level of the nth hydrogen fuel cell vehicle at time t; η HEV is the power generation efficiency of the hydrogen fuel cell vehicle; S HEV,max / S HEV,min are the upper / lower limits of the energy storage of the hydrogen fuel cell vehicle.
[0123] Example 6:
[0124] A method for pre-disaster deployment of mobile power sources in a multi-region electricity-hydrogen integrated energy system based on deep reinforcement learning, the technical content is the same as any one of Examples 1-5. Further, the operation constraints of the hydrogen energy system include the uniqueness constraints of the operation states of the electrolyzer and the hydrogen fuel cell, the operation power constraint of the electrolyzer, the operation power constraint of the hydrogen fuel cell, the upper / lower limits of the energy storage of the hydrogen storage tank, and the energy balance constraint of the hydrogen storage tank;
[0125] The uniqueness constraints of the operation states of the electrolyzer and the hydrogen fuel cell are as follows:
[0126]
[0127] The operation power constraint of the electrolyzer is as follows:
[0128]
[0129] The operation power constraint of the hydrogen fuel cell is as follows:
[0130]
[0131] The upper / lower limits of the energy storage of the hydrogen storage tank are as follows:
[0132]
[0133] The energy balance constraint of the hydrogen storage tank is as follows:
[0134]
[0135] In the formula, u EDt / u FCt are the operation state variables of the electrolyzer and the hydrogen fuel cell at time t, and if it is 1, it means that the electrolyzer / hydrogen fuel cell is in the working state at time t; are the operation powers of the electrolyzer / hydrogen fuel cell at time t; P ED,max / P FC,max are the maximum operation powers of the electrolyzer / hydrogen fuel cell; η ED / η FC are the energy conversion efficiencies of the electrolyzer / hydrogen fuel cell; LHV His the lower heating value of hydrogen; is the hydrogen storage capacity of the hydrogen storage tank at time t; M load,Ht is the hydrogen load at time t; S HS,min / S HS,max are the minimum / maximum hydrogen storage masses of the hydrogen storage tank respectively; η HS,in / η HS,out are the hydrogen charging / discharging efficiencies of the hydrogen storage tank respectively; if only a hydrogen refueling station is deployed in the area, then u EDt and u FCt are both 0.
[0136] Example 7:
[0137] A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, the technical content is the same as any one of Examples 1-6. Further, the network reconstruction constraints are as follows:
[0138]
[0139] In the formula, δ Fij,t is the closing state variable of line ij in the virtual power flow at time t; δ Fi,t is the state variable indicating whether the potential root node i serves as the root node of the microgrid at time t; M is a sufficiently large constant; P Fij,t is the virtual power flow through branch ij at time t.
[0140] Example 8:
[0141] A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, the technical content is the same as any one of Examples 1-7. Further, the load shedding balance constraints are as follows:
[0142]
[0143] In the formula, P load,Ei,t is the required electrical load of node i at time t; P re,Ei,t is the actual supplied electrical load of node i at time t; M load,Ht is the required hydrogen load at time t; M re,Ht is the actual supplied hydrogen load at time t.
[0144] Example 9:
[0145] A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, the technical content is the same as any one of Examples 1-8. Further, the distribution network power flow constraints are as follows:
[0146]
[0147] In the formula, P ij,t, Q ij,t are the active power and reactive power of line ij at time t, respectively. Q maxij are the upper limits of the active power and reactive power of line ij, respectively. r ij , x ij are the resistance and reactance of line ij, respectively. is the total active and reactive power injected by distributed power sources at node j at time t; is the square of the voltage at node i at time t.
[0148] Embodiment 10:
[0149] A method for pre-disaster deployment of mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, the technical content is the same as any one of Embodiments 1-9. Further, the deep reinforcement learning network includes an action network and an evaluation network;
[0150] The state space of the action network is as follows:
[0151] o = [U branch,all , P load,all , P H,all , Q H,all , P EV,all , Q EV,all , P HEV,all , Q HEV,all (29)
[0152] In the formula, o represents the state space of the agent, which includes the topological state information U of all regions branch,all , load prediction information P load,all , maximum power generation information P of the hydrogen energy system H,all , hydrogen storage capacity information Q H,all , and the maximum power generation power information P of all mobile electrical energy storage EV,all , mobile electrical energy storage capacity information Q EV,all , maximum power generation power information P of hydrogen fuel power generation vehicles HEV,all , hydrogen fuel power generation vehicle capacity information Q HEV,all .
[0153] The action space of the action network is as follows:
[0154]
[0155] In the formula, a is the action space of the agent, which includes the deployment results a of all mobile electrical energy storage EVj and the deployment results a of all emergency hydrogen fuel power generation vehicles HEVj , the discrete action a u,EVj ∈
[0156] {1, 2,..., NM}, selecting 1 means deploying the jth mobile electric energy storage to the first grid connection
[0157] point, and the same applies to the emergency hydrogen fuel power generation vehicle;
[0158] The reward function of the evaluation network is as follows:
[0159]
[0160] In the formula, the agent reward function r up is the sum of the load shedding losses in all regions, and r downi is the reward function of the ith lower-level agent, which is the load shedding loss in region i, and C lose,Ei is for region i
[0161] The total penalty for the cut-off electrical load; C lose,Hi is the total penalty for the cut-off hydrogen load in region i;
[0162] The loss function of the evaluation network is as follows:
[0163]
[0164] In the formula, x n,j is the preference value corresponding to the jth action dimension in the nth action output by the Actor network; p n,j is the discrete action probability; D is the experience pool, and α H is the temperature coefficient, which is used to control the weight of the entropy regularization term in the loss function; is the entropy regularization term; A is the total number of actions of the agent, and K n are the total action latitudes of the nth action respectively; H tar is the target entropy of the agent;
[0165] During the training process of the deep reinforcement learning network, the parameters are updated as follows:
[0166]
[0167] In the formula, α represents the learning rate of the agent; represents the gradient; θ m represents the parameters of the policy network and the value network.
[0168] Example 11:
[0169] A method for pre-disaster deployment of mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning is as follows:
[0170] Considering the cross - time - space flexibility of the hydrogen energy system and the cross - regional support ability of mobile emergency resources, a pre - disaster deployment strategy for mobile power sources in a multi - regional electricity - hydrogen integrated energy system based on deep reinforcement learning is established. The results show that this model can fully consider the cross - time - space flexibility of the hydrogen energy system and the cross - regional support ability of mobile emergency resources, and improve the resilience in the post - disaster recovery of the multi - regional electricity - hydrogen integrated energy system. A pre - disaster deployment strategy for mobile power sources in a multi - regional electricity - hydrogen integrated energy system based on deep reinforcement learning specifically includes the following steps:
[0171] (1) Fully considering the cross - regional support potential of mobile power sources, a pre - disaster deployment framework for mobile power sources in a multi - regional electricity - hydrogen integrated energy system considering regional mutual assistance is established.
[0172] The pre - disaster deployment framework for mobile power sources in a multi - regional electricity - hydrogen integrated energy system considering regional mutual assistance is as Figure 1 shown. Before a disaster occurs, the load in the distribution network is powered by the upper - level main grid. After the disaster occurs, the distribution network loses the main grid power supply and several line faults occur. Therefore, before an extreme disaster is about to come, the joint disaster - resistance center will fully consider the uncertainty of the fault lines in each region, combine the operating status, capacity and distribution of distributed power sources in each region, and formulate a pre - disaster deployment strategy for mobile power sources based on the deep reinforcement learning algorithm. Specifically, the mobile power sources are deployed in advance to the grid access points in each region to ensure that the mobile power sources can be quickly put into use and provide necessary support when a disaster occurs. After the disaster occurs, according to the actual line fault situation, the distribution network operators in each region use the mobile power sources pre - deployed by the joint disaster - resistance center and the hydrogen energy system in the region, and combine network reconstruction technology to quickly transfer the power supply to important loads, so as to ensure the supply of important loads to the greatest extent.
[0173] (2) Fully considering the energy supply potential of distributed power sources and network reconstruction in the emergency response during the disaster, the method for establishing an emergency response model for the electricity - hydrogen integrated energy system considering the coordinated scheduling of distributed power sources and network reconstruction is as follows:
[0174] The emergency response model for the electricity - hydrogen integrated energy system considering the coordinated scheduling of distributed power sources and network reconstruction aims to minimize the expected total operating cost of the electricity - hydrogen integrated energy system in the region, including the penalty for the cut - off electricity load, the penalty for the cut - off hydrogen load, and the operating cost of distributed power sources.
[0175] Specifically, it is as follows:
[0176] min(C lose,E +C lose,H +C gen,E +C gen,H ) (1)
[0177] In the formula, C lose,E is the total penalty for the cut - off electricity load; Close,H is the total penalty for the hydrogen curtailment load; C gen,E is the total operating cost of the mobile electrical energy storage; C gen,H is the total operating cost of the hydrogen fuel cell vehicle and the hydrogen fuel cell in the hydrogen energy system, and the specific composition is as follows:
[0178]
[0179] In the formula, N E , N T1 respectively represent the distribution network node and the set of dispatching periods within the region; c lose,Ei is the penalty for the unit electrical load curtailment at node i; P lose,Ei,t is the amount of electrical load curtailment at node i at time t; c lose,H is the penalty for the unit hydrogen curtailment load; M lose,Ht is the amount of hydrogen curtailment load at time t; N EV / N HEV / N RC respectively represent the sets composed of mobile electrical energy storage / hydrogen fuel cell vehicle / maintenance personnel; N M represents the set of grid connection points; c gen,E is the unit operating cost of the mobile electrical energy storage, P EV,outi,t,n is the discharge power of the nth mobile electrical energy storage located at node i at time t; c gen,H is the unit operating cost of the hydrogen fuel cell, P HEV,outi,t,n is the discharge power of the nth hydrogen fuel cell vehicle located at node i at time t; P FCt is the discharge power of the hydrogen fuel cell in the hydrogen energy system at time t.
[0180] 1) Operating constraints of the mobile power source
[0181] The mobile power source can access the grid through the grid connection node to provide power support to the grid, thereby improving the system resilience. The operating constraints of the mobile power source include the operating constraints of the mobile electrical energy storage and the hydrogen fuel cell vehicle. Equation (6) is the uniqueness constraint of the charging / discharging energy state of the mobile electrical energy storage;
[0182] Equations (7)-(8) are the charging / discharging power constraints of the mobile electrical energy storage; Equation (9) is the energy balance constraint of the mobile electrical energy storage; Equation (10) is the upper / lower limit constraint of the energy storage of the mobile electrical energy storage; Equation (11)
[0183] is the uniqueness constraint of the charging / discharging energy state of the hydrogen fuel cell vehicle; Equations (12)-(13) are the charging / discharging power constraints of the hydrogen fuel cell vehicle; Equation (14) is the energy balance constraint of the hydrogen fuel cell vehicle;
[0184] Equation (15) is the upper / lower limit constraint of the energy storage of the hydrogen fuel cell vehicle; Equation (16) is the constraint on the maximum number of mobile power sources accessing the grid at the grid connection point.
[0185]
[0186] Wherein, u EV,ini,t,n / u EV,outi,t,n are the state variables of the charging / discharging energy of the nth mobile electric energy storage at the t-th moment at node i respectively. If it is 1, it means that the nth mobile electric energy storage is in the charging / discharging energy state at the t-th moment at node i; P EV,ini,t,n / P EV,outi,t,n are the charging / discharging power of the nth mobile electric energy storage at the t-th moment at node i respectively; P EV,in,max / P EV,out,max are the upper limits of the charging / discharging power of the mobile power supply respectively; S EVt,n is the energy storage level of the nth mobile electric energy storage at the t-th moment; η EV,in / η EV,out are the charging / discharging efficiencies of the mobile electric energy storage respectively; S EV,max / S EV,min are the upper / lower limits of the energy storage of the mobile electric energy storage; u HEV,ini,t,n / u HEV,outi,t,n are the state variables of the charging / discharging energy of the nth hydrogen fuel cell vehicle at the t-th moment at node i respectively. If it is 1, it means that the nth hydrogen fuel cell vehicle is in the charging / discharging energy state at the t-th moment at node i; M HEV,ini,t,n / P HEV,outi,t,n are the charging mass / discharging power of the nth hydrogen fuel cell vehicle at the t-th moment at node i respectively; M HEV,in,max / P HEV,out,max are the upper limits of the charging mass / discharging power of the hydrogen fuel cell vehicle respectively; S HEVt,n is the energy storage level of the nth hydrogen fuel cell vehicle at the t-th moment; η HEV is the power generation efficiency of the hydrogen fuel cell vehicle; S HEV,max / S HEV,min are the upper / lower limits of the energy storage of the hydrogen fuel cell vehicle.
[0187] 3) Operation constraints of the hydrogen energy system
[0188] The hydrogen energy system can play the flexibility of hydrogen energy on the time scale through the coordinated cooperation of hydrogen fuel cells, electrolyzers and hydrogen storage tanks. The operation model of the hydrogen energy system can be expressed as shown in Eqs. (17)-(21). Eq. (17) is the uniqueness constraint of the operation states of the electrolyzer and the hydrogen fuel cell; Eq. (18) is the operation power constraint of the electrolyzer; Eq. (19) is the operation power constraint of the hydrogen fuel cell; Eq. (20) is the upper / lower limit constraint of the energy storage of the hydrogen storage tank; Eq. (21) is the energy balance constraint of the hydrogen storage tank.
[0189]
[0190] Wherein, u EDt / u FCtThey are the operating state variables of the electrolyzer and the hydrogen fuel cell at time t. If it is 1, it means that the electrolyzer / hydrogen fuel cell is in the working state at time t; They are the operating powers of the electrolyzer / hydrogen fuel cell at time t; P ED,max / P FC,max They are the maximum operating powers of the electrolyzer / hydrogen fuel cell; η ED / η FC They are the energy conversion efficiencies of the electrolyzer / hydrogen fuel cell; LHV H is the lower heating value of hydrogen; is the hydrogen storage capacity of the hydrogen storage tank at time t; M load,Ht is the hydrogen load at time t; S HS,min / S HS,max They are the minimum / maximum hydrogen storage masses of the hydrogen storage tank; η HS,in / η HS,out They are the hydrogen charging / discharging efficiencies of the hydrogen storage tank respectively; If only hydrogen refueling stations are deployed in the area, then u EDt and u FCt are both 0.
[0191] 4) Network reconfiguration constraints
[0192] After a fault occurs, the distribution network can be reconfigured into multiple microgrids by adjusting the operating states of the tie switches and distributed power sources. Let the nodes connected to the substation and distributed power sources form the potential root node set N G , and the set composed of all the lines in the distribution network is N BR . The microgrid established during the network reconfiguration process needs to meet the radial topological structure, and its necessary and sufficient conditions are: ① The number of closed lines is equal to the number of network nodes minus the number of subgraphs; ② Each subgraph is internally connected, as shown in Eqs. (22)-(23).
[0193]
[0194]
[0195] In the formula, δ Fij,t is the closed state variable of line ij in the virtual power flow at time t; δ Fi,t is the state variable indicating whether the potential root node i serves as the root node of the microgrid at time t; M is a sufficiently large constant; P Fij,t is the virtual power flow through branch ij at time t.
[0196] 5) Load shedding balance constraints
[0197] Eqs. (24)-(27) represent the relationship that the actual supply load, load shedding and demand load need to satisfy.
[0198]
[0199] Wherein, P load,Ei,t is the demand electric load at node i at time t; P re,Ei,t is the actual supply electric load at node i at time t; M load,Ht is the demand hydrogen load at time t; M re,Ht is the actual supply hydrogen load at time t.
[0200] 6) Power flow constraints of the distribution network
[0201] Based on the improved LinDistFlow linear power flow model, this paper simulates the power flow distribution of the distribution network during post-disaster recovery, as shown in Equation (28):
[0202]
[0203] Wherein, P ij,t , Q ij,t are the active power and reactive power of line ij at time t, respectively, Q maxij are the upper limits of the active power and reactive power of line ij, respectively, r ij , x ij are the resistance and reactance of line ij, respectively, is the total active and reactive power injected by distributed power sources at node j at time t; is the square of the voltage at node i at time t.
[0204] (3) Pre-disaster deployment of the multi-region electric-hydrogen integrated energy system based on deep reinforcement learning
[0205] Modeling and solving of the decision-making process.
[0206] 1) Modeling of the pre-disaster deployment decision-making process of the multi-region electric-hydrogen integrated energy system
[0207] 1.1) State space
[0208] The agent can obtain global observation variables, including global topology fault probability information, global agent power and capacity information, etc.
[0209] o = [U branch,all , P load,all , P H,all , Q H,all , P EV,all , Q EV,all , P HEV,all , Q HEV,all (29)
[0210] Wherein, o represents the state space of the agent, which includes the topology state information U branch,all of all regions, load prediction information P load,all, the maximum power generation information P of the hydrogen energy system H,all , the hydrogen energy storage capacity information Q H,all , and the maximum power generation information P of all mobile electrical energy storage EV,all , the mobile electrical energy storage capacity information Q EV,all , the maximum power generation information P of the hydrogen fuel power generation vehicle HEV,all , the capacity information Q of the hydrogen fuel power generation vehicle HEV,all .
[0211] 1.2) Action space
[0212] The agent is deployed through
[0213]
[0214] where a is the action space of the agent, which includes the deployment results a of all mobile electrical energy storage EVj and the deployment results a of all emergency hydrogen fuel power generation vehicles HEVj , the discrete action a u,EVj ∈
[0215] {1, 2, …, N M}; choosing 1 means deploying the jth mobile electrical energy storage to the first grid connection point, and the same applies to the emergency hydrogen fuel power generation vehicle.
[0216] 1.3) Environment
[0217] The environment mainly consists of an integrated electricity-hydrogen energy system disaster response model considering network reconstruction and coordinated dispatching of distributed power sources. That is, after the agent gives the deployment decision of mobile power sources, the deployed mobile power sources and the fixed power sources in the network are used, and combined with network reconstruction technology, the weighted load is restored as much as possible, so as to evaluate the resilience level of the integrated electricity-hydrogen energy system after the agent's decision.
[0218] The integrated electricity-hydrogen energy system disaster response model aims to minimize the weighted load curtailment, and comprehensively considers the operation constraints of mobile power sources, the operation constraints of the hydrogen energy system, the network reconstruction constraints, the load balance constraints, and the power flow constraints. The specific model is as follows.
[0219] Objective function: Equation (1), Constraint conditions: Equations (6)-(28)
[0220] 1.4 Reward function
[0221] After obtaining the deployment structures of all mobile electrical energy storage and emergency hydrogen fuel power generation vehicles, the load shedding loss under this deployment result can be obtained by solving the integrated electricity-hydrogen energy system disaster response model described in Section 2.3.2. The problem studied aims to maximize the multi-region weighted load restoration, so the reward function is set as follows.
[0222]
[0223] In the formula, the agent reward function r up is the sum of the load shedding losses of all regions, and r downi is the reward function of the i-th lower-level agent, which is the load shedding loss of region i. C lose,Ei is the total penalty for the electrical load shedding in region i; C lose,Hi is the total penalty for the hydrogen load shedding in region i, which can be calculated based on the emergency response model for the integrated electricity-hydrogen energy system during a disaster proposed in Section 2.3.2.
[0224] 2) Pre-disaster deployment strategy for the multi-region integrated electricity-hydrogen energy system based on the discrete SAC algorithm
[0225] SAC (Soft Actor-Critic) is a reinforcement learning algorithm based on policy gradients, belonging to the off-policy algorithm, and introducing entropy regularization, aiming to improve the stability and efficiency of training by maximizing the expected return and exploring the policy. The network structure of the SAC algorithm is as Figure 2 shown. The policy network is used to generate the pre-disaster deployment locations of mobile power sources for the multi-region integrated electricity-hydrogen energy system, and the value network is used to evaluate the quality of the mobile power source deployment plan. The π and θ parameters are the parameters of the policy network and the value network respectively, including the weight vector W and the bias vector b of the neural network.
[0226] SAC is mainly applied to continuous action control, while the pre-disaster deployment actions of the multi-region integrated electricity-hydrogen energy system are all discrete actions. To model this action feature, first, the action network generates a softmax distribution to output the corresponding probabilities of all possible discrete actions, and then sampling from this distribution can obtain the discrete actions. Generally, the formula for the softmax distribution is as follows.
[0227]
[0228] In the formula, x n,j is the preference value corresponding to the j-th action dimension in the n-th action output by the Actor network, which reflects the relative selection tendency of the n-th action to select the j-th action dimension, but does not directly represent the probability of the action. It can be expressed as the selection probability p of the action through Equation (33). n,j . It should be noted that in this paper, it is necessary to output the deployment locations of multiple mobile electrical energy storage and emergency fuel power generation vehicles simultaneously. Therefore, multiple discrete actions need to be output simultaneously, and a softmax distribution needs to be constructed for each output of a discrete action.
[0229] For the pre-disaster deployment problem of the multi-region electricity-hydrogen integrated energy system studied in this paper, the action space is relatively large, and traditional reinforcement learning algorithms often struggle to effectively explore the entire action space. Therefore, this paper adopts the SAC algorithm as the core optimization algorithm. By introducing the strategy optimization idea of maximum entropy, it can promote more randomness and exploration while maximizing the reward, thus avoiding falling into local optimal solutions. The loss function calculation formula of its policy network is shown in Equation (34).
[0230]
[0231] In the formula, D is the experience pool, and α H is the temperature coefficient, which is used to control the weight of the entropy regularization term in the loss function; is the entropy regularization term, which reflects the randomness of the current policy. A higher entropy value means that the policy distribution of the agent is more dispersed, and it may assign larger selection probabilities to multiple actions, reflecting strong exploration behavior. When the entropy value is lower, it means that the policy distribution of the agent is more concentrated, and the agent is more inclined to select certain specific actions. The calculation formula of the entropy regularization term for discrete actions is as shown in the formula:
[0232]
[0233] In the formula, A is the total number of actions of the agent, and K n are the total action latitudes of the nth action respectively. In addition, to reduce the problem of reduced learning efficiency caused by overestimation of the Q value, this paper introduces two Q value functions with the same structure and takes the smaller value to participate in the calculation of the loss function of the policy network.
[0234]
[0235] In the conventional SAC algorithm, the Q value update depends on the immediate reward at the current moment, the reward and state value at the next time step. However, for the pre-disaster deployment of the mobile power supply studied in this paper, it is a one-time single-step decision, and only the immediate reward obtained from the current decision needs to be concerned. Therefore, the calculation formula of the Q value network loss function is simplified as shown in Equation (37):
[0236]
[0237] To ensure that the agent selects the best strategy after fully exploring the environment, the SAC algorithm adjusts the temperature coefficient, and the loss function calculation of the temperature coefficient is as shown in Equation (4-46).
[0238]
[0239] In the formula, H tarIt is the target entropy of the agent. By setting the target entropy, the SAC algorithm can automatically adjust the temperature coefficient to keep the entropy value near the target entropy, so that it can maintain a certain degree of exploration while effectively using the learned strategy to obtain higher rewards.
[0240] Based on the above loss function calculation formula, an action network, two evaluation networks, and the temperature coefficient can be updated as follows:
[0241]
[0242] In the formula, α represents the learning rate of the agent.
[0243] 3) Training process
[0244] During the training process of the SAC algorithm, the agent continuously interacts with the environment, collects state, action, and reward information, and stores these experiences in the experience pool. Then, a small batch of data is randomly sampled from the experience pool for updating the evaluation network and the action network. The evaluation network evaluates the value of the action by minimizing the error, and the action network optimizes the decision by maximizing the reward and entropy. By continuously repeating this process, the agent can gradually learn the optimal strategy in a highly uncertain environment. The training framework of the proposed method is shown in Algorithm 1.
[0245]
[0246] 4) Testing process
[0247] During the testing process, we first collect the parameters of the action network of the agent trained by Algorithm 1. The evaluation network is no longer needed during the testing process. For each test scenario, the action network is used to obtain the mobile power supply deployment decision, which is then substituted into the environment to verify the deployment. Generally speaking, the SAC testing process proposed in this paper is shown in Algorithm 3-2.
[0248]
[0249]
[0250] In this paper, a multi-region electricity-hydrogen-transportation network coupling system containing 4 electricity-hydrogen integrated energy systems and their corresponding transportation networks is constructed and simulated for verification. The multi-region electricity-hydrogen integrated energy system consists of 4 improved IEEE 33-node distribution networks. Among them, hydrogen refueling stations (including hydrogen storage tanks and hydrogen dispensers) are equipped in regions 1-2, and hydrogen energy systems (including electrolyzers, hydrogen fuel cells, hydrogen storage tanks and hydrogen dispensers) are equipped in regions 3-4, both of which are configured at node 10. In addition, each electricity-hydrogen integrated energy system contains 4 mobile power grid access points, which are respectively configured at nodes 4, 13, 26, and 32, as Figure 3As shown. The electrical loads are divided into three levels according to their importance, and the weight coefficients are 1000, 100, and 5 respectively, as shown in Table 2. The line damage risks in each region are shown in Table 3. Each region is equipped with two mobile electrical energy storages and two hydrogen fuel power generation vehicles. The relevant operating parameters of the mobile power supply and hydrogen energy equipment are shown in Table 1. The time scale of emergency dispatching during the disaster is 6 hours. In this paper, Gurobi is used to model and solve the optimization model. The SAC algorithm in this paper is written based on Python 3.8, and the neural network framework is built by tensorflow 2.6.0.
[0251] Table 1
[0252]
[0253] Table 2
[0254]
[0255]
[0256] Table 3
[0257]
[0258] To verify the effectiveness of the method proposed in this paper, the following cases are set up for comparative analysis of the test system.
[0259] Case 1: The method proposed in this paper, based on the SAC algorithm, conducts pre-disaster collaborative dispatching of the multi-region electricity-hydrogen integrated energy system.
[0260] Case 2: Using the stochastic programming method, based on the electricity-hydrogen integrated energy system emergency response model proposed in Section 2.3 of this paper, conducts pre-disaster collaborative dispatching of the multi-region electricity-hydrogen integrated energy system.
[0261] Case 3: Using the stochastic programming method, based on the electricity-hydrogen integrated energy system emergency response model proposed in Section 2.3 of this paper, conducts pre-disaster collaborative dispatching of the multi-region electricity-hydrogen integrated energy system, and limits the maximum solution time to 12,000 seconds to make it more in line with the actual usage requirements.
[0262] Figure 4 Shows the convergence of the SAC deep reinforcement learning method under 3000 generations of training, where the solid line depicts the moving average reward over 100 epochs. From Figure 4 It can be seen that due to the differences in topological failures within each training epoch, there are certain fluctuations in the reward function values within the training epoch, but it does not affect the overall convergence trend.
[0263] Table 4
[0264] Method Case1 Case2 Case3 <![CDATA[Load shedding loss (×10 5 )]]> 300029 303565 331069 Solution time (seconds) 0.029 37455 12000
[0265] Table 4 shows the test results of each case under 100 typical scenarios. Among them, the method proposed in this paper can obtain the highest reward value. The stochastic programming algorithm with time limit and the stochastic programming method to obtain the optimal solution are 1.18% and 10.34% lower than the method proposed in this paper respectively. The method proposed in this paper is superior to the stochastic programming method in both the characterization of uncertainty and the solution time. This is because the stochastic programming method extracts the uncertainty of topological faults and loads through typical scenarios, and its adaptability to uncertainty depends on the selection of typical scenarios. To ensure the solution efficiency, only a small number of scenarios can be used, making it difficult to fully characterize the uncertainty of multiple regions. While deep reinforcement learning is not restricted by this, and it can directly learn the optimal strategy from a large number of scenarios covering various uncertainties through interaction with the environment. It should be noted that the load shedding loss during the training of the SAC algorithm is higher than that in the test results. This is because, by introducing the maximum entropy objective, the SAC algorithm not only maximizes the expected return when optimizing the policy, but also encourages the exploration of the policy. At the same time, the agent samples actions based on the policy probability distribution, making each action have a certain probability of being selected, thus promoting a wider exploration. While in the test stage, the agent no longer explores, and the policy changes to a "deterministic" mode, directly selecting the action with the maximum expected return. Therefore, the load shedding loss during the test is better than that during the training.
[0266] To study the impact of cross-regional resource sharing on the pre-disaster collaborative scheduling of multi-region electric-hydrogen integrated energy systems and verify the effectiveness of the method proposed in this paper, the following cases are set in this paper for comparative analysis of the test system.
[0267] Case1: The method proposed in this paper, using the SAC algorithm to collaboratively perform pre-disaster collaborative scheduling.
[0268] Case4: Using the SAC algorithm, each region independently conducts pre-disaster to in-disaster collaborative deployment.
[0269] 1) Analysis of the impact of pre-disaster cross-regional resource sharing on the system load shedding loss
[0270] The operating results of each region of the system under different cases are shown in Table 5. In Case4, when the mobile resource allocation in Region 3 and Region 4 is the same, their load shedding losses are much lower than those in Region 1 and Region 2. This is because in extreme disasters, the hydrogen energy system can generate electricity through hydrogen fuel cells using a large amount of stored hydrogen and feed it back to the power grid, providing strong resilience support for the power grid and thus greatly reducing the load shedding loss. After considering cross-regional resource sharing, the total load shedding loss of the multi-region electric-hydrogen integrated energy system has decreased by 47.6%. Among them, the load shedding losses of Region 3 and Region 4 have increased by 21.5×10 3 yuan and 24.4×10 3yuan, and the load shedding losses in Region 1 and Region 2 decreased by 172.9×10 3 yuan and 158.6×10 3 yuan respectively. It shows that the method in this paper can reasonably allocate mobile power sources according to the disaster conditions in each region and the resilience resources within the region, minimizing the total load shedding loss of the multi-region electric-hydrogen integrated energy system. In addition, after considering cross-regional resource sharing, a small amount of hydrogen load shedding loss occurred in Region 3 because the gap of important electric loads in Region 3 was large, and the hydrogen energy system would feedback part of the hydrogen used to supply hydrogen loads in the hydrogen storage tank to the power grid through a hydrogen fuel cell to ensure the supply of important electric loads as much as possible.
[0271] Table 5
[0272]
[0273] Figure 7 The load shedding losses at all levels in different regions under two cases were compared. It can be seen from the figure that after the pre-disaster deployment, all primary loads in each region can be supplied during the emergency response stage during the disaster, indicating the significance of pre-disaster mobile power source deployment for ensuring the supply of important loads. In addition, compared with Case 4, the secondary and tertiary load shedding losses in Region 3 and Region 4 in Case 3 increased by 28.2×10 3 yuan and 17.8×10 3 yuan respectively, while the secondary and tertiary load shedding losses in Region 1 and Region 2 decreased by 325.6×10 3 yuan and 4.0×10 3 yuan respectively. That is, the mobile power sources originally mainly used to supply tertiary loads in Region 3 and Region 4 were allocated to Region 1 and Region 2 to make full use of the existing resilience resources to restore important loads. To sum up, the method in this paper can analyze the potential disaster conditions in each region and the existing resilience resources within the region to reasonably allocate MRR, and give priority to ensuring the supply of important loads under disasters.
[0274] 2) Specific scenario analysis
[0275] In this section, Scenario 1 in the test scenarios is selected to specifically analyze the pre-disaster deployment of mobile power sources. The line fault conditions and the deployment results of mobile power sources in Scenario 1 are as Figure 8 shown. Among them, to ensure the continuous power supply of the distribution network load, the mobile power source does not move after reaching the access point.
[0276] From Figure 8It can be seen that although regions 1 and 3 are subject to the N-6 fault, network reconstruction through tie lines enables internal connectivity within the system. Although it is still disconnected from the main grid at this time, internal connectivity within the system can significantly increase the power supply scope of distributed power sources. For example, in region 3, only the hydrogen energy system and a mobile energy storage can supply all the primary loads and most of the secondary loads within the region, and the amount of secondary load shedding is much smaller than that in region 4 which has a hydrogen energy system and three emergency hydrogen fuel power generation vehicles deployed. Therefore, considering tie lines and network reconstruction during the in-disaster emergency response simulation can more accurately evaluate the resilience level of the system during emergency response, thereby making more accurate pre-disaster deployment decisions for the multi-region electricity-hydrogen integrated energy system.
[0277] For regions 2 and 4, the damage is more severe, resulting in the N-9 fault. Although the distributed power sources cooperate with network reconstruction and act immediately during the disaster, 5 non-connected microgrids are still formed. Among them, microgrid 5 formed by nodes 16, 17, 18, and 33 does not have resilience resources such as mobile power access points and hydrogen energy systems. Therefore, during the in-disaster emergency response stage, it cannot obtain power supply and can only resume power supply after the damaged lines are repaired by maintenance personnel during the post-disaster recovery stage. It is worth noting that microgrid 3 in region 4 has a grid access point inside, but no mobile power is deployed there, while region 2 with the same structure has emergency hydrogen fuel power generation vehicles deployed. This is because microgrid 3 is connected to microgrid 4 through two lines, and there is a hydrogen energy system in microgrid 4 that can provide a large amount of high-power support. Although in the current scenario, due to the faults of both line 8-9 and line 22-12, the hydrogen energy system cannot provide power support for microgrid 3, as long as one of line 8-9 and line 22-12 is operating normally, the hydrogen energy system can provide power support for microgrid 3 as Figure 9 shown. Correspondingly, there is no hydrogen energy system in region 2 that can provide such a large amount of power support, so emergency hydrogen fuel power generation vehicles must be connected to microgrid 3 in region 2 for power supply.
[0278] The proposed pre-disaster deployment strategy for mobile power sources in a multi-region electricity-hydrogen integrated energy system based on deep reinforcement learning includes establishing a pre-disaster deployment framework for mobile power sources in a multi-region electricity-hydrogen integrated energy system considering regional mutual assistance; establishing an in-disaster emergency response model for an electricity-hydrogen integrated energy system considering the coordinated scheduling of distributed power sources and network reconstruction; establishing a decision-making process and solution method for the pre-disaster deployment of mobile power sources in a multi-region electricity-hydrogen integrated energy system based on deep reinforcement learning. The method proposed in this paper can fully consider uncertainties such as line faults and existing resilience resources within the region, reasonably allocate and deploy mobile power sources to each region, prioritize the supply of important loads, and improve the resilience level of the multi-region electricity-hydrogen integrated energy system.
Claims
1. A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning, characterized in that It includes the following steps: 1) Considering the cross-region support potential of mobile power sources, establish a pre-disaster deployment framework for mobile power sources in a multi-region electric-hydrogen integrated energy system considering regional mutual assistance. 2) Based on the pre-disaster deployment framework for mobile power sources in a multi-region electric-hydrogen integrated energy system, construct a disaster response model during the disaster considering the coordinated scheduling of distributed power sources and network reconfiguration; 3) Use a deep reinforcement learning network to solve the disaster response model during the disaster considering the coordinated scheduling of distributed power sources and network reconfiguration, and obtain the pre-disaster deployment decision for the multi-region electric-hydrogen integrated energy system.
2. The pre-disaster deployment strategy of the mobile power source for the multi-region electric-hydrogen integrated energy system based on deep reinforcement learning according to claim 1, wherein, The pre-disaster deployment framework for mobile power sources in the multi-region electric-hydrogen integrated energy system includes mobile power sources and distribution networks; The mobile power sources are deployed to the distribution network access points in each region.
3. The pre-disaster deployment strategy of the mobile power source for the multi-region electric-hydrogen integrated energy system based on deep reinforcement learning according to claim 1, characterized in that, The objective function of the disaster response model during the disaster considering the coordinated scheduling of distributed power sources and network reconfiguration is as follows: min(C lose,E +C lose,H +C gen,E +C gen,H ) (1) Wherein, C lose,E is the total penalty for the cut power load; C lose,H is the total penalty for the cut hydrogen load; C gen,E is the total operating cost of the mobile electrical energy storage; C gen,H is the total operating cost of the hydrogen fuel cell vehicle and the hydrogen fuel cells in the hydrogen energy system; N E and N T1 respectively represent the set of distribution network nodes and the set of dispatching periods within the region; c lose,Ei is the penalty for the unit cut power load of node i; P lose,Ei,t is the cut power load of node i at time t; c lose,H is the penalty for the unit cut hydrogen load; M lose,Ht is the cut hydrogen load at time t; N EV / N HEV / N RC respectively represent the sets composed of mobile electrical energy storage / hydrogen fuel cell vehicle / maintenance personnel; N M represents the set of grid access points; c gen,E is the unit operating cost of the mobile electrical energy storage, P EV,outi,t,n is the discharge power of the nth mobile electrical energy storage located at node i at time t; c gen ,H is the unit operating cost of the hydrogen fuel cell, P HEV,outi,t,n is the discharge power of the nth hydrogen fuel cell vehicle located at node i at time t; P FCt is the discharge power of the hydrogen fuel cells in the hydrogen energy system at time t.
4. The pre-disaster deployment strategy of the mobile power source for the multi-region electric-hydrogen integrated energy system based on deep reinforcement learning according to claim 1, wherein, The constraint conditions of the disaster response model during the disaster considering the coordinated scheduling of distributed power sources and network reconfiguration include mobile power source operation constraints, hydrogen energy system operation constraints, network reconfiguration constraints, load shedding balance constraints, and distribution network power flow constraints.
5. The pre-disaster deployment strategy of a mobile power source for a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning according to claim 4, characterized in that, The mobile power source operation constraints include the uniqueness constraint of the charging / discharging state of mobile electrical energy storage, the charging / discharging power constraint of mobile electrical energy storage, the energy balance constraint of mobile electrical energy storage, the upper / lower limit constraint of the energy storage of mobile electrical energy storage, the uniqueness constraint of the charging / discharging state of hydrogen fuel cell vehicles, the charging / discharging power constraint of hydrogen fuel cell vehicles, the energy balance constraint of hydrogen fuel cell vehicles, the upper / lower limit constraint of the energy storage of hydrogen fuel cell vehicles, and the maximum number of mobile power sources connected to the grid at the grid access point limit constraint; Among them, the uniqueness constraint of the charging / discharging state of mobile electrical energy storage is as follows: The charging / discharging power constraint of mobile electrical energy storage is as follows: The energy balance constraint of mobile electrical energy storage is as follows: The upper / lower limit constraint of the energy storage of mobile electrical energy storage is as follows: The uniqueness constraint of the charging / discharging state of hydrogen fuel cell vehicles is as follows: The charging / discharging power constraint of hydrogen fuel cell vehicles is as follows: The energy balance constraint of hydrogen fuel cell vehicles is as follows: The upper / lower limit constraint of the energy storage of hydrogen fuel cell vehicles is as follows: The maximum number of mobile power sources connected to the grid at the grid access point limit constraint is as follows: where u EV,ini,t,n and u EV,outi,t,n are the state variables of the charging / discharging energy of the nth mobile electric energy storage at time t at node i. If it is 1, it means that the nth mobile electric energy storage is in the charging / discharging state at node i at time t; P EV,ini,t,n and P EV,outi,t,n are the charging / discharging powers of the nth mobile electric energy storage at time t at node i respectively; P EV,in,max and P EV,out,max are the upper limits of the charging / discharging powers of the mobile power source respectively; S EVt,n is the energy storage level of the nth mobile electric energy storage at time t; η EV,in and η EV,out are the charging / discharging efficiencies of the mobile electric energy storage respectively; S EV,max and S EV,min are the upper / lower limits of the energy storage of the mobile electric energy storage; u HEV,ini,t,n and u HEV,outi,t,n are the state variables of the charging / discharging energy of the nth hydrogen fuel cell vehicle at time t at node i. If it is 1, it means that the nth hydrogen fuel cell vehicle is in the charging / discharging state at node i at time t; M HEV,ini,t,n and P HEV,outi,t,n are the charging mass / discharging power of the nth hydrogen fuel cell vehicle at time t at node i respectively; M HEV,in,max and P HEV,out,max are the upper limits of the charging mass / discharging power of the hydrogen fuel cell vehicle respectively; S HEVt,n is the energy storage level of the nth hydrogen fuel cell vehicle at time t; η HEV is the power generation efficiency of the hydrogen fuel cell vehicle; S HEV ,max and S HEV,min are the upper / lower limits of the energy storage of the hydrogen fuel cell vehicle.
6. The pre-disaster deployment strategy of the mobile power supply for the multi-region electric-hydrogen integrated energy system based on deep reinforcement learning according to claim 4, wherein, The hydrogen energy system operation constraints include the uniqueness constraint of the operation states of electrolyzers and hydrogen fuel cells, the operation power constraint of electrolyzers, the operation power constraint of hydrogen fuel cells, the upper / lower limit constraint of the energy storage of hydrogen storage tanks, and the energy balance constraint of hydrogen storage tanks; The uniqueness constraint of the operation states of electrolyzers and hydrogen fuel cells is as follows: The operation power constraint of electrolyzers is as follows: The operation power constraint of hydrogen fuel cells is as follows: The upper / lower limit constraint of the energy storage of hydrogen storage tanks is as follows: The energy balance constraint of hydrogen storage tanks is as follows: where, u EDt , u FCt are the operating state variables of the electrolyzer and the hydrogen fuel cell at time t, respectively. If it is 1, it means that the electrolyzer / hydrogen fuel cell is in the working state at time t; P FCt are the operating powers of the electrolyzer / hydrogen fuel cell at time t, respectively; P ED,max , P FC,max are the maximum operating powers of the electrolyzer / hydrogen fuel cell, respectively; η ED , η FC are the energy conversion efficiencies of the electrolyzer / hydrogen fuel cell, respectively; LHV H is the lower heating value of hydrogen; is the hydrogen storage capacity of the hydrogen storage tank at time t; M load,Ht is the hydrogen load at time t; S HS,min , S HS,max are the minimum / maximum hydrogen storage masses of the hydrogen storage tank, respectively; η HS,in , η HS,out are the hydrogen charging / discharging efficiencies of the hydrogen storage tank, respectively; If only a hydrogen refueling station is deployed in the area, then u EDt and u FCt are both 0.
7. A pre-disaster deployment strategy for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning according to claim 4, characterized in that, The network reconfiguration constraints are as follows: where, δ Fij,t is the closing state variable of line ij in the virtual power flow at time period t; δ Fi,t is the state variable indicating whether the potential root node i serves as the root node of the microgrid in period t; M is a sufficiently large constant; P Fij,t is the virtual power flow through branch ij at time t.
8. A pre-disaster deployment strategy for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning according to claim 4, characterized in that The load shedding balance constraints are as follows: Where, P load,Ei,t is the required electricity load of node i at time t; P re,Ei,t is the actual supplied electricity load of node i at time t; M load,Ht is the required hydrogen load at time t; M re,Ht is the actual supplied hydrogen load at time t.
9. The pre-disaster deployment strategy of a mobile power source for a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning according to claim 4, characterized in that, The distribution network power flow constraints are as follows: Wherein, P ij,t and Q ij,t are the active power and reactive power of line ij at time t respectively, and Q maxij are the upper limits of the active power and reactive power of line ij respectively, r ij and x ij are the resistance and reactance of line ij respectively, is the total active and reactive power injected by the distributed power source at node j at time t; is the square of the voltage at node i at time t.
10. A pre-disaster deployment method for mobile power sources in a multi-region electric-hydrogen integrated energy system based on deep reinforcement learning according to claim 1, characterized in that, The deep reinforcement learning network includes an action network and an evaluation network; The state space of the action network is as follows: o = [U branch,all , P load,all , P H,all , Q H,all , P EV,all , Q EV,all , P HEV,all , Q HEV,all (29) Wherein, o represents the state space of the agent, which includes the topological state information U of all regions branch,all , the load forecasting information P load,all , the maximum power generation information P of the hydrogen energy system H,all , the hydrogen energy storage capacity information Q H,all , and the maximum power generation information P of all mobile electrical energy storage EV,all , the mobile electrical energy storage capacity information Q EV,all , the maximum power generation information P of the hydrogen fuel power generation vehicle HEV,all , the hydrogen fuel power generation vehicle capacity information Q HEV,all . The action space of the action network is as follows: where a is the agent action space, which includes all deployment results a of mobile electric energy storage EVj and all deployment results a of emergency hydrogen fuel power generation vehicles HEVj , discrete action a u,EVj ∈ {1, 2, …, N M}, selecting 1 means deploying the j-th mobile electric energy storage to the first grid connection point, and the same applies to the emergency hydrogen fuel power generation vehicle; The reward function of the evaluation network is as follows: where the agent reward function r up is the sum of the load shedding losses in all regions, and r downi is the reward function of the i-th lower-level agent, which is the load shedding loss in region i, and C lose,Ei is the total penalty for the electrical load shedding in region i; C lose,Hi is the total penalty for the hydrogen load shedding in region i; The loss function of the evaluation network is as follows: where x n,j is the preference value corresponding to the j-th action dimension in the n-th action output by the Actor network; p n,j is the discrete action probability; D is the experience pool, and α H is the temperature coefficient used to control the weight of the entropy regularization term in the loss function; is the entropy regularization term; A is the total number of actions of the agent, and K n is the total action dimension of the n-th action; H tar is the target entropy of the agent; L(θ), L(α H ) is the loss function; During the training process of the deep reinforcement learning network, the parameter update is as follows: Where α represents the learning rate of the agent; represents the gradient; θ m represents the parameters of the policy network and the value network.
Citation Information
Cited By
Double-four-foot series combination control method, system, medium and equipment
CN121043154A