Method and device for generating reconstruction plan during disaster period of power distribution network based on reinforcement learning
By constructing a disaster-affected power distribution network reconstruction model based on reinforcement learning, the problem of low efficiency of traditional optimization algorithms in disaster response is solved, and a safe and feasible reconstruction plan is generated quickly, thereby improving the disaster response capability of the power distribution network.
Patent Information
- Application Number
- CN202510993735.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
When faced with strong uncertainties in disasters, existing technologies and traditional optimization algorithms struggle to develop effective reconfiguration schemes for distribution networks within a limited timeframe, resulting in insufficient disaster response capabilities.
A reinforcement learning-based approach is used to construct a reconfiguration model for the power distribution network during disasters and transform it into a Markov decision process model. The model is then trained offline using a reinforcement learning agent to generate a reconfiguration plan before the disaster occurs.
It improves the resilience of the distribution network during disasters, reduces load loss, meets engineering and physical constraints, and ensures the safety and feasibility of reconfiguration plans.
Smart Images

Figure CN120875385A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power system disaster emergency technology, and specifically relates to a method and apparatus for generating reconfiguration plans for distribution networks during disasters based on reinforcement learning. Background Technology
[0002] With the increasing frequency of extreme disasters such as typhoons in recent years, power distribution networks have experienced faults such as tower collapses and line breaks, affecting users' normal electricity demand and causing significant economic losses. Developing a reasonable power distribution network reconstruction plan before a disaster strikes can effectively improve the network's resilience in the face of disasters and ensure the normal power supply to the load to the greatest extent possible.
[0003] Existing methods for reconfiguring distribution networks primarily model the problem as an optimization problem with 0-1 integer variables, solving it using traditional optimization algorithms. However, due to the strong uncertainty of disasters, the number of fault scenarios to be considered when formulating a reconfiguration plan is large. Traditional optimization algorithms often consume a lot of time when facing complex scenarios, and may even fail to obtain the optimal solution within a finite time. As a result, traditional optimization algorithms cannot formulate reasonable and effective reconfiguration schemes to improve the resilience of the distribution network in the face of disasters. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a method and apparatus for generating reconfiguration plans for distribution networks during disasters based on reinforcement learning. This invention can generate reconfiguration plans based on the probability of branch faults in the distribution network before a disaster occurs, effectively improving the resilience of the distribution network in responding to disasters and reducing load losses caused by disasters.
[0005] A first aspect of this invention proposes a method for generating a reconfiguration plan for a distribution network during a disaster based on reinforcement learning, comprising:
[0006] Establish a reconfiguration model for the power distribution network during disasters;
[0007] The disaster-affected power distribution network reconstruction model is transformed into a Markov decision process model.
[0008] Construct a reinforcement learning agent corresponding to the power distribution network; based on the Markov decision process model, generate training samples for offline training of the reinforcement learning agent, and finally obtain the trained reinforcement learning agent;
[0009] The trained reinforcement learning agent is used to generate online reconstruction plans for the power distribution network during disasters.
[0010] In one specific embodiment of the present invention, establishing a power distribution network reconstruction model during a disaster includes:
[0011] 1) Establish the objective function of the power distribution network reconfiguration model during a disaster, as expressed below:
[0012]
[0013] Where T is the total duration of the reconfiguration plan; N is the set of load nodes in the distribution network; and L is the set of branch switches in the distribution network. Let be the expected economic cost of node i due to load loss at time t; The additional economic cost incurred by branch l at time t due to the switching action;
[0014] in,
[0015]
[0016] in, Let N be the load cost coefficient for node i; s be the fault scenario number; N is the load cost coefficient for node i. l The number of branches that may fail; Pro situ,s (t) represents the probability of fault scenario s occurring at time t; P i (t) represents the predicted load of node i at time t; p i,s (t) represents the actual load satisfied by node i at time t under scenario s;
[0017]
[0018] in, The cost of performing a single switching operation for branch l; a l (t) is a 0-1 binary variable representing whether the switch of branch l is operated at time t, a l (t) = 0 means that the switch of branch l was not operated at time t, a l (t) = 1 indicates that the switch of branch l is operated at time t;
[0019] 2) Establish constraints for the reconfiguration model of the distribution network during a disaster, including:
[0020] Distribution network topology constraints, expressed as follows:
[0021]
[0022] Where, α ij,t β ij,t u l,t Both are binary variables of 0 and 1; α ij,t α represents the connection status between node i and node j at time t. ij,t =1 indicates that there is a branch connection between node i and node j at time t, α ij,t=0 means that there is no branch connection between node i and node j at time t, or the branch switch is in the open state; β ij,t β represents whether node j is the parent node of node i at time t. ij,t =1 means that at time t, node j is the parent node of node i, β ij,t =0 means that at time t, node j is not the parent node of node i; u l,t u represents the switch status of branch l at time t. l,t =1 indicates that the switch of branch l is in the closed state at time t, u l,t =0 means that the branch l switch is in the open state at time t; Λ is the set of all branches, S is the set of all power supply nodes, Ω is the set of all nodes, and Ω / S is the set of all non-power supply nodes;
[0023] The power balance constraint at distribution network nodes is expressed as follows:
[0024]
[0025] Among them, P ij,s (t) represents the active power flowing from node i to node j in scenario s at time t, P ki,s (t) represents the active power flowing from node k to node i in scenario s at time t; This represents the active power output of the main power supply located at node i in scenario s at time t; This represents the active power output of the distributed power source located at node i in scenario s at time t.
[0026] The power range constraint for distribution network nodes is expressed as follows:
[0027] 0≤p i,s (t)≤P i (t) (12)
[0028]
[0029] in, and These represent the upper limits of active power output of the main power source and the distributed power source located at node i, respectively.
[0030] The power range constraint for distribution network branches is expressed as follows:
[0031]
[0032] y l,t,s =u l,t *S l,t,s (16)
[0033] Among them, y l,t,sy represents the connection status of branch l in scenario s at time t, and is a binary variable of 0-1. l,t,s =1 indicates that branch l is in a connected state in scenario s at time t, y l,t,s =0 means that branch l is in an unconnected state under scenario s at time t;
[0034] S represents the maximum transmission power of the branch; l,s,t Let S be the fault condition of branch l in scenario s at time t, and be a binary variable of 0-1. l,s,t =1 means that branch l is not faulty in scenario s at time t, S l,s,t =0 indicates that branch l is faulty in scenario s at time t;
[0035] The constraints on the number of distribution network reconfigurations and the number of switching operations are expressed as follows:
[0036]
[0037]
[0038] ∑ t ∑ l a l,t ≤max_action*once_max (21)
[0039] Among them, a l,t This represents whether branch l is switched at time t, and is a binary variable of 0 and 1; a l,t =1 indicates that branch l is switched at time t; a l,t =0 means that branch l was not switched at time t;
[0040] on o ff l,t This represents the operation performed by the switch in branch l at time t, with values of 1, 0, and -1; on o ff l,t =1, indicating that the switch of branch l is closed at time t; on o ff l,t =0, indicating that the switch of branch l was not operated at time t; on o ff l,t =-1, indicating that the switch of branch l is disconnected at time t;
[0041] `max_action` and `once_max` represent the maximum number of refactoring actions in a single refactoring plan and the maximum number of on / off actions during each refactoring, respectively.
[0042] In a specific embodiment of the present invention, the step of transforming the distribution network reconstruction model during a disaster into a Markov decision process model includes:
[0043] 1) Construct the state variables for the reconfiguration operation of the distribution network during the disaster at time t:
[0044]
[0045] Among them, Pro break This is a fault probability vector for all branches within the distribution network at all times, with the dimension being the total number of branches. Each element in this vector represents the fault probability of the corresponding branch. Let t be the vector of switch closure status of all branches within the distribution network at time t. The dimension is the total number of branches. Each element in this vector is a binary variable of 0-1, representing the switch state of the corresponding branch. Let be the predicted active load vector of all nodes within the distribution network at time t, with dimension equal to the total number of nodes. Each element in this vector represents the predicted active load value of the corresponding node. Let t be the identifier vector indicating whether each action can be executed under all constraints at time t. Its dimension is the total number of branches + 1, corresponding to the dimension of the action space. Each element of this vector is a binary variable of 0-1, representing whether the corresponding action can be executed.
[0046] in, The process of determining the value of each element in the algorithm is as follows:
[0047] After each acquisition of the distribution network status, pre-simulation is performed on all actions sequentially, and actions that exceed the constraints of the reconstructed model during the disaster period are excluded. Set the corresponding element in the middle to 0;
[0048] 2) Construct variables for the reconfiguration operations of the distribution network during the disaster at time t:
[0049] a RL,t =(a t (23)
[0050] Among them, a t The vector represents whether all branches within the distribution network have performed switching operations at time t. Its dimension is the total number of branches. Each element in this vector is a binary variable of 0 and 1, representing whether the corresponding branch has performed a switching operation. keep is a binary variable of 0 and 1, indicating whether the distribution network will no longer perform switching operations at time t. keep = 0 means continue to perform switching operations; keep = 1 means stop performing switching operations.
[0051] in:
[0052] ∑a t +keep=1 (24)
[0053] Where, ∑a t For a tThe sum of all elements in the expression represents the number of switching operations performed at the current moment.
[0054] 3) Construct the reward function for the reconfiguration operation of the distribution network during the disaster at time t:
[0055]
[0056] Where, r t The reward at time t; β is the expected increase in load provided by the distribution network topology at time t relative to the original topology. load and β switch The weights for load loss cost and switching operation cost are respectively, satisfying:
[0057] β load ,β switch ∈(0,1)
[0058] β load +β switch =1.
[0059] In one specific embodiment of the present invention, the construction of the reinforcement learning agent corresponding to the power distribution network includes:
[0060] 1) Constructing a policy neural network π for the reinforcement learning agent corresponding to the distribution network. θ Randomly initialize π θ The parameters θ; π θ The input is the state variable s RL,t The output is the action variable a. RL,t The probability distribution;
[0061] For this policy neural network:
[0062] Using softmax to generate action variable a RL,t Before the probability distribution, for each Elements that are equal to 0 in a are placed in a. RL,t If the value of the corresponding element is set to -∞, the probability of this action being selected is set to 0 after the softmax operation.
[0063] 2) Construct the value function neural network for the reinforcement learning agent corresponding to the distribution network. Random initialization parameters The input is the state variable s RL,t The output is the agent's expected cumulative discount reward. The estimated value, where γ is the discount factor;
[0064] 3) Construct a neural network for the objective value function of the reinforcement learning agent corresponding to the power distribution network. Structure and same, parameters initial value and parameters Their initial values are the same.
[0065] In one specific embodiment of the present invention, the offline training of the reinforcement learning agent includes:
[0066] 1) Construct an initial empty experience pool for the reinforcement learning agent to store experience samples (s) RL,t ,a RL,t ,r t ,s RL,t+1 );
[0067] 2) Randomly generate a set of scenarios, including: the fault probability of all branches in the distribution network at each time, and the predicted active load of all nodes in the distribution network at each time.
[0068] 3) Based on the Markov decision process model, the reinforcement learning agent is made to perform a Markov interaction process with the power distribution network environment, and the experience sample data generated in the process is stored in the experience pool.
[0069] Then determine:
[0070] If the number of samples in the experience pool reaches the preset sample number threshold, proceed to step 4); otherwise, return to step 2).
[0071] 4) Randomly select a set of samples from the experience pool of the reinforcement learning agent. The number of samples is Computational Value Function Neural Network Loss function:
[0072]
[0073] Among them, y t Approximate for time t The target value is calculated using the following expression:
[0074]
[0075] Calculated Then, update using gradient descent. Parameters;
[0076] 5) Randomly select a set of samples from the experience pool of the reinforcement learning agent. The number of samples is Computational value function neural network πθ Loss function:
[0077]
[0078] Calculated Then, π is updated using gradient descent. θ Parameters;
[0079] 6) Utilize the updated Parameter update Parameters:
[0080]
[0081] Where δ is the update rate of the target value function neural network;
[0082] 7) Repeat steps 2)-6) until the upper limit of the set training rounds of the reinforcement learning agent is reached, and the reinforcement learning agent that has been trained offline is obtained.
[0083] In one specific embodiment of the present invention, the step of generating an online reconstruction plan for a power distribution network during a disaster using the trained reinforcement learning agent includes:
[0084] 1) At time t, construct the state variables of the Markov decision process using the measured values of the distribution network.
[0085] 2) The state variable s obtained in step 4-1) RL,t Input a reinforcement learning agent that has been trained offline, and the reinforcement learning agent outputs the reconstruction operation action variable a at time t. RL,t =(a t (keep);
[0086] 3) Determine the result of step 2):
[0087] If a RL,t If keep = 0, then let s RL,t =s RL,t+1 Then return to step 2) to continue generating the next switching operation at time t; if a RL,t If keep = 1, then the switching operation at time t is completed. Let t = t + 1, and then return to step 1.
[0088] Once all time-based switching operations have been generated, the distribution network reconfiguration plan generation is complete.
[0089] In one specific embodiment of the present invention, it further includes:
[0090] Reinforcement learning agents generate new samples (s) during online phases.RL,t ,a RL,t ,r t ,s RL,t+1 The data is then stored in the experience pool of the reinforcement learning agent; subsequently, the reinforcement learning agent is trained again using samples from the updated experience pool.
[0091] A second aspect of the present invention provides a reinforcement learning-based device for generating reconfiguration plans for power distribution networks during disasters, comprising:
[0092] The reconstructed model generation module is used to establish a reconstructed model of the power distribution network during a disaster.
[0093] The Markov decision process model conversion module is used to convert the disaster-affected power distribution network reconstruction model into a Markov decision process model.
[0094] The reinforcement learning agent training module is used to construct a reinforcement learning agent corresponding to the power distribution network; based on the Markov decision process model, training samples are generated for offline training of the reinforcement learning agent, and finally the trained reinforcement learning agent is obtained.
[0095] The reconstruction plan generation module is used to generate reconstruction plans for the power distribution network during disasters online using the trained reinforcement learning agent.
[0096] A third aspect of the present invention provides an electronic device comprising:
[0097] At least one processor; and a memory communicatively connected to said at least one processor;
[0098] The memory stores instructions that can be executed by the at least one processor, and the instructions are configured to execute the above-described reinforcement learning-based method for generating a reconfiguration plan for a power distribution network during a disaster.
[0099] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions for causing the computer to execute a method for generating a reconfiguration plan for a power distribution network during a disaster based on reinforcement learning.
[0100] The features and beneficial effects of this invention are as follows:
[0101] 1) This invention proposes a method and apparatus for generating reconfiguration plans for distribution networks during disasters based on reinforcement learning. It utilizes a reinforcement learning agent improved with a policy neural network to efficiently output reconfiguration plans during disasters, based on the predicted fault probability of the distribution network during the disaster period. Furthermore, it satisfies both engineering and physical constraints, improving the resilience of the distribution network to natural disasters and reducing load losses caused by disasters.
[0102] 2) In the offline stage, the present invention randomly generates data such as the probability of faults in distribution network branches and the load of nodes, and uses a reinforcement learning agent to interact with the environment to obtain sample data to train the reinforcement learning agent. Then, in the online stage, the reconstruction plan is quickly obtained by using the information from the interaction with the actual distribution network.
[0103] 3) When the intelligent agent outputs a switch, the present invention restricts actions that exceed the constraints, so that the reconfiguration plan can meet the engineering and physical constraints and ensure the plan is safe and feasible. Attached Figure Description
[0104] Figure 1 This is an overall flowchart of a method for generating a reconfiguration plan for a power distribution network during a disaster, based on reinforcement learning, according to an embodiment of the present invention. Detailed Implementation
[0105] This invention proposes a method and apparatus for generating reconfiguration plans for power distribution networks during disasters based on reinforcement learning. The following is a detailed description in conjunction with the accompanying drawings and specific embodiments.
[0106] A first aspect of this invention proposes a method for generating a reconfiguration plan for a distribution network during a disaster based on reinforcement learning, comprising:
[0107] Establish a reconfiguration model for the power distribution network during disasters;
[0108] The disaster-affected power distribution network reconstruction model is transformed into a Markov decision process model.
[0109] Construct a reinforcement learning agent corresponding to the power distribution network; based on the Markov decision process model, generate training samples for offline training of the reinforcement learning agent, and finally obtain the trained reinforcement learning agent;
[0110] The trained reinforcement learning agent is used to generate online reconstruction plans for the power distribution network during disasters.
[0111] In a specific embodiment of the present invention, the overall process of the reinforcement learning-based method for generating a reconfiguration plan for a power distribution network during a disaster is as follows: Figure 1 As shown, it includes the following steps:
[0112] 1) Establish a reconstruction model for the power distribution network during disasters.
[0113] In this embodiment, the reconfiguration model of the power distribution network during a disaster consists of an objective function and constraints.
[0114] The objective function is to minimize the expected value of the additional economic costs caused by disasters during the distribution network's disaster period. The constraints include: distribution network topology constraints, distribution network node power balance constraints, distribution network node power range constraints, distribution network branch power range constraints, and constraints on the number of distribution network reconfigurations and switching operations. The specific steps are as follows:
[0115] 1-1) Establish the objective function of the reconfiguration model for the power distribution network during a disaster.
[0116] In a specific embodiment of the present invention, the objective function expression is:
[0117]
[0118] Where T is the total duration of the reconfiguration plan; N is the set of load nodes in the distribution network; and L is the set of branch switches in the distribution network. Let be the expected economic cost of node i due to load loss at time t; The additional economic cost incurred by branch l at time t due to the switching action;
[0119] in,
[0120]
[0121] in, Let N be the load cost coefficient for node i; s be the fault scenario number; N is the load cost coefficient for node i. l The number of possible faulty branches determines the number of fault scenarios. In this embodiment, since each branch has both faulty and non-faulty states, the total number of fault scenarios is: Let's take 10 branches, numbered 1 to 10, that might fail as an example. The scenario where all branches are functioning correctly is represented by the ten-bit binary number 0000000000, indicating s = 0. The scenario where branch 1 fails while the other branches are functioning correctly is represented by the ten-bit binary number 0000000001, indicating s = 1. The scenario where all branches fail is represented by the ten-bit binary number 1111111111, indicating s = 2. 10 -1. Pro situ,s (t) represents the probability of fault scenario s occurring at time t; P i (t) represents the predicted load of node i at time t; p i,s (t) represents the load actually satisfied by node i at time t under scenario s.
[0122]
[0123] in, The cost of performing a single switching operation for branch l; a l(t) represents whether the switch of branch l was operated at time t, and is a binary variable of 0-1. l (t) = 0 means that the switch of branch l was not operated at time t, a l (t) = 1 indicates that the switch of branch l is operated at time t.
[0124] 1-2) Establish the constraints for the reconfiguration model of the distribution network during a disaster, including:
[0125] Distribution network topology constraints, expressed as follows:
[0126]
[0127] Where, α ij,t β ij,t u l,t All are binary variables of 0 and 1. α ij,t α represents the connection status between node i and node j at time t. ij,t =1 indicates that there is a branch connection between node i and node j at time t, α ij,t =0 indicates that there is no branch connection between node i and node j at time t, or the branch switch is in the open state. β ij,t β represents whether node j is the parent node of node i at time t. ij,t =1 means that at time t, node j is the parent node of node i, β ij,t =0 means that at time t, node j is not the parent node of node i. l,t u represents the switch status of branch l at time t. l,t =1 indicates that the switch of branch l is in the closed state at time t, u l,t =0 indicates that the branch switch l is in the open state at time t. Λ is the set of all branches, S is the set of all power supply nodes, Ω is the set of all nodes, and Ω / S is the set of all non-power supply nodes.
[0128] The power balance constraint at distribution network nodes is expressed as follows:
[0129]
[0130] Among them, P ij,s (t) represents the active power flowing from node i to node j in scenario s at time t, P ki,s (t) represents the active power flowing from node k to node i in scenario s at time t; This represents the active power output of the main power supply located at node i in scenario s at time t; This represents the active power output of the distributed power source located at node i in scenario s at time t.
[0131] The power range constraint for distribution network nodes is expressed as follows:
[0132] 0≤p i,s (t)≤P i (t) (12)
[0133]
[0134]
[0135] in, and These represent the upper limits of active power output of the main power source and the distributed power source located at node i, respectively.
[0136] The power range constraint for distribution network branches is expressed as follows:
[0137]
[0138] y l,t,s =u l,t *S l,t,s (16)
[0139] Among them, y l,t,s y represents the connection status of branch l in scenario s at time t, and is a binary variable of 0-1. l,t,s =1 indicates that branch l is in a connected state in scenario s at time t, y l,t,s =0 means that branch l is in an unconnected state under scenario s at time t. S represents the maximum transmission power of the branch; l,s,t Let S be the fault condition of branch l in scenario s at time t, and be a binary variable of 0-1. l,s,t =1 means that branch l is not faulty in scenario s at time t, S l,s,t =0 represents a fault in branch l under scenario s at time t, which is a known parameter after the fault scenario is determined.
[0140] The constraints on the number of distribution network reconfigurations and the number of switching operations are expressed as follows:
[0141]
[0142] ∑ t ∑ l a l,t ≤max_action*once_max (21)
[0143] Among them, a l,t This represents whether branch l is switched at time t, and is a binary variable of 0 and 1; a l,t =1 indicates that branch l is switched at time t; a l,t =0 means that branch l was not switched at time t.
[0144] ono ff l,t This represents the operation performed by the switch in branch l at time t, with values of 1, 0, and -1; on o ff l,t =1, indicating that the switch of branch l is closed at time t; on o ff l,t =0, indicating that the switch of branch l was not operated at time t; on o ff l,t =-1, which means that the switch of branch l is disconnected at time t.
[0145] `max_action` and `once_max` represent the maximum number of refactoring actions in a single refactoring plan and the maximum number of on / off actions during each refactoring, respectively.
[0146] 2) Transform the distribution network reconfiguration model obtained in step 1) during the disaster period into a Markov decision process model; the specific steps are as follows:
[0147] 2-1) Construct the reconfiguration operation state variables of the distribution network at time t during the disaster period:
[0148]
[0149] Among them, Pro break This is a fault probability vector for all branches within the distribution network at all times. Its dimension is the total number of branches. Each element in this vector represents the fault probability of the corresponding labeled branch. For example, the first element represents the fault probability of the labeled branch 1. Let t be the vector of the switching status of all branches in the distribution network at time t. The dimension is the total number of branches. Each element in this vector is a binary variable of 0-1, representing the switching status of the corresponding labeled branch. For example, the first element being equal to 0 means that the labeled branch 1 is in the open state. Let t be the predicted active load vector of all nodes in the distribution network at time t. The dimension is the total number of nodes. Each element in this vector represents the predicted active load value of the corresponding labeled node. For example, the first element represents the predicted active load value of the labeled node 1. Let t be a vector indicating whether each action can be executed under all constraints at time t. Its dimension is the total number of branches + 1, corresponding to the dimension of the action space. Each element in this vector is a binary variable of 0-1, representing whether the corresponding action can be executed. For example, if the first element is equal to 1, it means that the action of the label 1 can be executed, that is, the switch of the branch labeled 1 is allowed to be operated.
[0150] In this embodiment, The method for determining the value of each element in the algorithm is as follows:
[0151] After each acquisition of the distribution network status, pre-simulation is performed on all actions sequentially, and actions that exceed the constraints of the reconstructed model during the distribution network disaster period are excluded. The corresponding element in the model is set to 0. For example, after performing a pre-simulation of the first action, i.e., after executing the operation on the first branch switch, if the model constraints are exceeded, then... The first element is set to 0; the rest... Setting the option to 1 indicates that the remaining actions can be selected.
[0152] 2-2) Construct the variables for the reconfiguration operation during the disaster period of the distribution network at time t:
[0153] a RL,t =(a t (23)
[0154] Among them, a t The vector represents whether all branches within the distribution network have performed switching operations at time t. Its dimension is the total number of branches. Each element in this vector is a binary variable (0-1), representing whether the switching operation for the corresponding labeled branch has been performed. For example, the first element equal to 1 indicates that the switching operation was performed on branch number 1. `keep` is a binary variable (0-1) indicating whether the distribution network will no longer perform switching operations at time t; `keep = 0` means continuing to perform switching operations; `keep = 1` means stopping the switching operations.
[0155] In this embodiment, for a RL,t The following constraints apply:
[0156] ∑a t +keep=1 (24)
[0157] Where, ∑a t For a t The summation of all elements in the expression represents the number of switching operations performed at the current moment.
[0158] Equation (24) shows that when the agent outputs an action, it either does not perform a switching operation or performs a switching operation on at most one branch at a time; that is, it decomposes the reconstruction operation at a moment into multiple switching operation steps.
[0159] 2-3) Construct the reward function for the reconfiguration operation of the distribution network during the disaster at time t:
[0160]
[0161] Where, r t The reward at time t; β is the expected increase in load provided by the distribution network topology at time t relative to the original topology. load and β switchThe weights for load loss cost and switching operation cost are respectively determined, subject to the following constraints:
[0162] β load ,β switch ∈(0,1)
[0163] β load +β switch =1
[0164] 3) Construct a reinforcement learning agent corresponding to the distribution network. Based on the Markov decision process model obtained in step 2), generate training samples for offline training of the reinforcement learning agent, and finally obtain the trained reinforcement learning agent; the specific steps are as follows:
[0165] 3-1) Constructing the policy neural network π for the reinforcement learning agent corresponding to the distribution network θ Randomly initialize π θ The parameters θ; π θ The input is the state variable s RL,t The output is the action variable a. RL,t The probability distribution; the structure of the network includes an input layer, a hidden layer and an output layer connected in sequence. The number of neurons in the input layer is the dimension of the state variables, and the number of neurons in the output layer is the dimension of the action variables; in this embodiment, there is one hidden layer with 512 neurons.
[0166] In this embodiment, the policy neural network is adjusted as follows:
[0167] Using softmax to generate action variable a RL,t Before the probability distribution, for each Elements that are equal to 0 in a are placed in a. RL,t The value of the corresponding element is set to -∞, for example If the first element in the array is 0, then a will be... RL,t If the action value corresponding to branch number 1 is set to -∞, then the probability of this action being selected will also be set to 0 after the softmax operation.
[0168] 3-2) Constructing a value function neural network for the reinforcement learning agent corresponding to the distribution network. Random initialization parameters The input is the state variable s RL,t The output is the agent's expected cumulative discount reward. The estimated value is given by γ, where γ is the discount factor, ranging from 0 to 1. The network structure consists of an input layer, a hidden layer, and an output layer connected in sequence. The number of neurons in the input layer is equal to the dimension of the state variables, and the number of neurons in the output layer is equal to the dimension of the action variables. In this embodiment, there is one hidden layer with 256 neurons.
[0169] 3-3) Constructing a neural network for the objective value function of the reinforcement learning agent corresponding to the distribution network. Structure and Completely identical parameters initial value and parameters Their initial values are the same.
[0170] 3-4) Construct an initially empty experience pool for the reinforcement learning agent. This experience pool is used to store experience samples (s). RL,t ,a RL,t ,r t ,s RL,t+1 In one specific embodiment of the present invention, the capacity of the experience pool is 2000.
[0171] 3-5) Randomly generate a set of scenarios, including: the fault probability of all branches in the distribution network at each time, and the predicted active load of all nodes in the distribution network at each time.
[0172] 3-6) Based on the Markov decision process model obtained in step 2), let the reinforcement learning agent perform a Markov interaction process with the power distribution network environment, and store the experience sample data generated in the process into the experience pool.
[0173] Then determine:
[0174] If the number of samples in the experience pool reaches the preset sample number threshold (500 in one specific embodiment of the present invention), then proceed to step 3-7; otherwise, return to step 3-5.
[0175] 3-7) Randomly select a set of samples from the experience pool of the reinforcement learning agent. The number of samples is In a specific embodiment of the present invention Computational Value Function Neural Network Loss function:
[0176]
[0177] Among them, y t Approximate for time t The target value is calculated using the following expression:
[0178]
[0179] Calculated Then, update using gradient descent. The parameters.
[0180] 3-8) Randomly select a set of samples from the experience pool of the reinforcement learning agent. The number of samples is In this embodiment (in and The values of π can be different, but are usually the same. The neural network calculates the value function π. θ Loss function:
[0181]
[0182] Calculated Then, π is updated using gradient descent. θ The parameters.
[0183] 3-9) Utilize the updated Parameter update Parameters:
[0184]
[0185] Wherein, δ is the update rate of the target value function neural network, and in a specific embodiment of the present invention, δ = 0.005.
[0186] 3-10) Repeat steps 3-5)-3-9) until the upper limit of the set training rounds for the reinforcement learning agent is reached, and the offline-trained reinforcement learning agent is obtained. In this embodiment, the upper limit of the training rounds is usually between 5000 and 50000. In a specific embodiment of the present invention, the upper limit of the training rounds for the reinforcement learning agent is 5000.
[0187] 4) Utilize the trained reinforcement learning agent to output the switching operations of each branch in the distribution network online, thereby generating a reconfiguration plan for the distribution network during a disaster; the specific steps are as follows:
[0188] 4-1) At time t, construct the state variables of the Markov decision process using the measured values of the distribution network. As input to the reinforcement learning agent.
[0189] 4-2) The state variable s obtained in step 4-1) RL,t Input a reinforcement learning agent that has been trained offline, and output the reconstruction operation action variable a at time t. RL,t =(a t (keep).
[0190] 4-3) Determine the result of step 4-2):
[0191] If a RL,t If keep = 0, then let s RL,t =s RL,t+1 Then return to step 4-2) to continue generating the next switching operation at time t; if a RL,t If keep = 1, then the switching operation at time t is completed. Let t = t + 1, and then return to step 4-1.
[0192] Once all time-based switching operations have been generated, the distribution network reconfiguration plan generation is complete.
[0193] Furthermore, the method described in this embodiment also includes:
[0194] 5) Reinforcement learning agents generate new samples (s) during the online phase. RL,t ,a RL,t ,r t ,S RL,t+1 And store it in the experience pool of the reinforcement learning agent.
[0195] Subsequently, the reinforcement learning agent is trained again using samples from the updated experience pool. The number of new experience samples required to start a new round of training depends on the time required for online communication and training. In one specific embodiment of the invention, a new round of training is performed on the reinforcement learning agent each time a switching action is generated. During each round of training, the loss function calculation method and network update method are the same as in the offline phase, allowing the agent's policy to be continuously optimized and approach the optimal policy.
[0196] In this embodiment, when the total number of samples in the experience pool of the reinforcement learning agent reaches the preset capacity limit, the sample that has been stored in the experience pool for the longest time will be removed for each new sample added.
[0197] To implement the above embodiments, a second aspect of the present invention proposes a reinforcement learning-based device for generating a reconfiguration plan for a power distribution network during a disaster, comprising:
[0198] The reconstructed model generation module is used to establish a reconstructed model of the power distribution network during a disaster.
[0199] The Markov decision process model conversion module is used to convert the disaster-affected power distribution network reconstruction model into a Markov decision process model.
[0200] The reinforcement learning agent training module is used to construct a reinforcement learning agent corresponding to the power distribution network; based on the Markov decision process model, training samples are generated for offline training of the reinforcement learning agent, and finally the trained reinforcement learning agent is obtained.
[0201] The reconstruction plan generation module is used to generate reconstruction plans for the power distribution network during disasters online using the trained reinforcement learning agent.
[0202] It should be noted that the foregoing explanation of an embodiment of a method for generating a reconfiguration plan for a distribution network during a disaster based on reinforcement learning also applies to a device for generating a reconfiguration plan for a distribution network during a disaster based on reinforcement learning in this embodiment, and will not be repeated here. According to the embodiment of the present invention, a device for generating a reconfiguration plan for a distribution network during a disaster based on reinforcement learning establishes a reconfiguration model for the distribution network during a disaster; transforms the reconfiguration model into a Markov decision process model; constructs a reinforcement learning agent corresponding to the distribution network; generates training samples based on the Markov decision process model for offline training of the reinforcement learning agent, ultimately obtaining a trained reinforcement learning agent; and uses the trained reinforcement learning agent to generate a reconfiguration plan for the distribution network during a disaster online. This enables the generation of a reconfiguration plan based on the probability of branch faults in the distribution network before a disaster occurs, effectively improving the resilience of the distribution network in responding to disasters and reducing load losses caused by disasters.
[0203] In one specific embodiment of the present invention, establishing a power distribution network reconstruction model during a disaster includes:
[0204] 1) Establish the objective function of the power distribution network reconfiguration model during a disaster, as expressed below:
[0205]
[0206] Where T is the total duration of the reconfiguration plan; N is the set of load nodes in the distribution network; and L is the set of branch switches in the distribution network. Let be the expected economic cost of node i due to load loss at time t; The additional economic cost incurred by branch l at time t due to the switching action;
[0207] in,
[0208]
[0209] in, Let N be the load cost coefficient for node i; s be the fault scenario number; N is the load cost coefficient for node i. l The number of branches that may fail; Pro situ,s (t) represents the probability of fault scenario s occurring at time t; Pi (t) represents the predicted load of node i at time t; p i,s (t) represents the actual load satisfied by node i at time t under scenario s;
[0210]
[0211] in, The cost of performing a single switching operation for branch l; a l (t) is a 0-1 binary variable representing whether the switch of branch l is operated at time t, a l (t) = 0 means that the switch of branch l was not operated at time t, a l (t) = 1 indicates that the switch of branch l is operated at time t;
[0212] 2) Establish constraints for the reconfiguration model of the distribution network during a disaster, including:
[0213] Distribution network topology constraints, expressed as follows:
[0214]
[0215] Where, α ij,t β ij,t u l,t Both are binary variables of 0 and 1; α ij,t α represents the connection status between node i and node j at time t. ij,t =1 indicates that there is a branch connection between node i and node j at time t, α ij,t =0 means that there is no branch connection between node i and node j at time t, or the branch switch is in the open state; β ij,t β represents whether node j is the parent node of node i at time t. ij,t =1 means that at time t, node j is the parent node of node i, β ij,t =0 means that at time t, node j is not the parent node of node i; u l,t u represents the switch status of branch l at time t. l,t =1 indicates that the switch of branch l is in the closed state at time t, u l,t =0 means that the branch l switch is in the open state at time t; Λ is the set of all branches, S is the set of all power supply nodes, Ω is the set of all nodes, and Ω / S is the set of all non-power supply nodes;
[0216] The power balance constraint at distribution network nodes is expressed as follows:
[0217]
[0218] Among them, P ij,s (t) represents the active power flowing from node i to node j in scenario s at time t, Pki,s (t) represents the active power flowing from node k to node i in scenario s at time t; This represents the active power output of the main power supply located at node i in scenario s at time t; This represents the active power output of the distributed power source located at node i in scenario s at time t.
[0219] The power range constraint for distribution network nodes is expressed as follows:
[0220] 0≤p i,s (t)≤P i (t) (12)
[0221]
[0222] in, and These represent the upper limits of active power output of the main power source and the distributed power source located at node i, respectively.
[0223] The power range constraint for distribution network branches is expressed as follows:
[0224]
[0225] y l,t,s =u l,t *S l,t,s (16)
[0226] Among them, y l,t,s y represents the connection status of branch l in scenario s at time t, and is a binary variable of 0-1. l,t,s =1 indicates that branch l is in a connected state in scenario s at time t, y l,t,s =0 means that branch l is in an unconnected state under scenario s at time t;
[0227] S represents the maximum transmission power of the branch; l,s,t Let S be the fault condition of branch l in scenario s at time t, and be a binary variable of 0-1. l,s,t =1 means that branch l is not faulty in scenario s at time t, S l,s,t =0 indicates that branch l is faulty in scenario s at time t;
[0228] The constraints on the number of distribution network reconfigurations and the number of switching operations are expressed as follows:
[0229]
[0230]
[0231] ∑ t ∑ l a l,t≤max_action*once_max (21)
[0232] Among them, a l,t This represents whether branch l is switched at time t, and is a binary variable of 0 and 1; a l,t =1 indicates that branch l is switched at time t; a l,t =0 means that branch l was not switched at time t;
[0233] on o ff l,t This represents the operation performed by the switch in branch l at time t, with values of 1, 0, and -1; on o ff l,t =1, indicating that the switch of branch l is closed at time t; on o ff l,t =0, indicating that the switch of branch l was not operated at time t; on o ff l,t =-1, indicating that the switch of branch l is disconnected at time t;
[0234] `max_action` and `once_max` represent the maximum number of refactoring actions in a single refactoring plan and the maximum number of on / off actions during each refactoring, respectively.
[0235] In a specific embodiment of the present invention, the step of transforming the distribution network reconstruction model during a disaster into a Markov decision process model includes:
[0236] 1) Construct the state variables for the reconfiguration operation of the distribution network during the disaster at time t:
[0237]
[0238] Among them, Pro break This is a fault probability vector for all branches within the distribution network at all times, with the dimension being the total number of branches. Each element in this vector represents the fault probability of the corresponding branch. Let t be the vector of switch closure status of all branches within the distribution network at time t. The dimension is the total number of branches. Each element in this vector is a binary variable of 0-1, representing the switch state of the corresponding branch. Let be the predicted active load vector of all nodes within the distribution network at time t, with dimension equal to the total number of nodes. Each element in this vector represents the predicted active load value of the corresponding node. Let t be the identifier vector indicating whether each action can be executed under all constraints at time t. Its dimension is the total number of branches + 1, corresponding to the dimension of the action space. Each element of this vector is a binary variable of 0-1, representing whether the corresponding action can be executed.
[0239] in, The process of determining the value of each element in the algorithm is as follows:
[0240] After each acquisition of the distribution network status, pre-simulation is performed on all actions sequentially, and actions that exceed the constraints of the reconstructed model during the disaster period are excluded. Set the corresponding element in the middle to 0;
[0241] 2) Construct variables for the reconfiguration operations of the distribution network during the disaster at time t:
[0242] a RL,t =(a t (23)
[0243] Among them, a t The vector represents whether all branches within the distribution network have performed switching operations at time t. Its dimension is the total number of branches. Each element in this vector is a binary variable of 0 and 1, representing whether the corresponding branch has performed a switching operation. keep is a binary variable of 0 and 1, indicating whether the distribution network will no longer perform switching operations at time t. keep = 0 means continue to perform switching operations; keep = 1 means stop performing switching operations.
[0244] in:
[0245] ∑a t +keep=1 (24)
[0246] Where, ∑a t For a t The sum of all elements in the expression represents the number of switching operations performed at the current moment.
[0247] 3) Construct the reward function for the reconfiguration operation of the distribution network during the disaster at time t:
[0248]
[0249] Where, r t The reward at time t; β is the expected increase in load provided by the distribution network topology at time t relative to the original topology. load and β switch The weights for load loss cost and switching operation cost are respectively, satisfying:
[0250] β load ,β switch ∈(0,1)
[0251] β load +β switch =1.
[0252] In one specific embodiment of the present invention, the construction of the reinforcement learning agent corresponding to the power distribution network includes:
[0253] 1) Constructing a policy neural network π for the reinforcement learning agent corresponding to the distribution network. θ Randomly initialize π θ The parameters θ; π θ The input is the state variable s RL,t The output is the action variable a. RL,t The probability distribution;
[0254] For this policy neural network:
[0255] Using softmax to generate action variable a RL,t Before the probability distribution, for each Elements that are equal to 0 in a are placed in a. RL,t If the value of the corresponding element is set to -∞, the probability of this action being selected is set to 0 after the softmax operation.
[0256] 2) Construct the value function neural network for the reinforcement learning agent corresponding to the distribution network. Random initialization parameters The input is the state variable s RL,t The output is the agent's expected cumulative discount reward. The estimated value, where γ is the discount factor;
[0257] 3) Construct a neural network for the objective value function of the reinforcement learning agent corresponding to the power distribution network. Structure and same, parameters initial value and parameters Their initial values are the same.
[0258] In one specific embodiment of the present invention, the offline training of the reinforcement learning agent includes:
[0259] 1) Construct an initial empty experience pool for the reinforcement learning agent to store experience samples (s) RL,t ,a RL,t ,r t ,s RL,t+1 );
[0260] 2) Randomly generate a set of scenarios, including: the fault probability of all branches in the distribution network at each time, and the predicted active load of all nodes in the distribution network at each time.
[0261] 3) Based on the Markov decision process model, the reinforcement learning agent is made to perform a Markov interaction process with the power distribution network environment, and the experience sample data generated in the process is stored in the experience pool.
[0262] Then determine:
[0263] If the number of samples in the experience pool reaches the preset sample number threshold, proceed to step 4); otherwise, return to step 2).
[0264] 4) Randomly select a set of samples from the experience pool of the reinforcement learning agent. The number of samples is Computational Value Function Neural Network Loss function:
[0265]
[0266] Among them, y t Approximate for time t The target value is calculated using the following expression:
[0267]
[0268] Calculated Then, update using gradient descent. Parameters;
[0269] 5) Randomly select a set of samples from the experience pool of the reinforcement learning agent. The number of samples is Computational value function neural network π θ Loss function:
[0270]
[0271] Calculated Then, π is updated using gradient descent. θ Parameters;
[0272] 6) Utilize the updated Parameter update Parameters:
[0273]
[0274] Where δ is the update rate of the target value function neural network;
[0275] 7) Repeat steps 2)-6) until the upper limit of the set training rounds of the reinforcement learning agent is reached, and the reinforcement learning agent that has been trained offline is obtained.
[0276] In one specific embodiment of the present invention, the step of generating an online reconstruction plan for a power distribution network during a disaster using the trained reinforcement learning agent includes:
[0277] 1) At time t, construct the state variables of the Markov decision process using the measured values of the distribution network.
[0278] 2) The state variable s obtained in step 4-1) RL,t Input a reinforcement learning agent that has been trained offline, and the reinforcement learning agent outputs the reconstruction operation action variable a at time t. RL,t =(a t (keep);
[0279] 3) Determine the result of step 2):
[0280] If a RL,t If keep = 0, then let s RL,t =s RL,t+1 Then return to step 2) to continue generating the next switching operation at time t; if a RL,t If keep = 1, then the switching operation at time t is completed. Let t = t + 1, and then return to step 1.
[0281] Once all time-based switching operations have been generated, the distribution network reconfiguration plan generation is complete.
[0282] In one specific embodiment of the present invention, it further includes:
[0283] Reinforcement learning agents generate new samples (s) during online phases. RL,t ,a RL,t ,r t ,s RL,t+1 The data is then stored in the experience pool of the reinforcement learning agent; subsequently, the reinforcement learning agent is trained again using samples from the updated experience pool.
[0284] To implement the above embodiments, a third aspect of the present invention provides an electronic device, comprising:
[0285] At least one processor; and a memory communicatively connected to said at least one processor;
[0286] The memory stores instructions that can be executed by the at least one processor, and the instructions are configured to execute the above-described reinforcement learning-based method for generating a reconfiguration plan for a power distribution network during a disaster.
[0287] To implement the above embodiments, a fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions for causing the computer to execute a reinforcement learning-based method for generating a reconfiguration plan for a power distribution network during a disaster.
[0288] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0289] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform a reinforcement learning-based method for generating a reconfiguration plan for a power distribution network during a disaster, as described in the above embodiments.
[0290] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0291] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0292] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0293] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.
[0294] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0295] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0296] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0297] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0298] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for generating reconfiguration plans for power distribution networks during disasters based on reinforcement learning, characterized in that, include: Establish a reconfiguration model for the power distribution network during disasters; The disaster-affected power distribution network reconstruction model is transformed into a Markov decision process model. Construct reinforcement learning agents for power distribution networks; Based on the Markov decision process model, training samples are generated for offline training of the reinforcement learning agent, and finally the trained reinforcement learning agent is obtained. The trained reinforcement learning agent is used to generate online reconstruction plans for the power distribution network during disasters.
2. The method according to claim 1, characterized in that, The establishment of the power distribution network reconstruction model during disasters includes: 1) Establish the objective function of the power distribution network reconfiguration model during a disaster, as expressed below: Where T is the total duration of the reconfiguration plan; N is the set of load nodes in the distribution network; and L is the set of branch switches in the distribution network. Let be the expected economic cost of node i due to load loss at time t; The additional economic cost incurred by branch l at time t due to the switching action; in, in, Let N be the load cost coefficient for node i; s be the fault scenario number; N is the load cost coefficient for node i. l The number of branches that may fail; Pro situ,s (t) represents the probability of fault scenario s occurring at time t; P i (t) represents the predicted load of node i at time t; p i,s (t) represents the actual load satisfied by node i at time t under scenario s; in, The cost of performing a single switching operation for branch l; a l (t) is a 0-1 binary variable representing whether the switch of branch l is operated at time t, a l (t) = 0 means that the switch of branch l was not operated at time t, a l (t) = 1 indicates that the switch of branch l is operated at time t; 2) Establish constraints for the reconfiguration model of the distribution network during a disaster, including: Distribution network topology constraints, expressed as follows: Where, α ij,t β ij,t u l,t Both are binary variables of 0 and 1; α ij,t α represents the connection status between node i and node j at time t. ij,t =1 indicates that there is a branch connection between node i and node j at time t, α ij,t =0 means that there is no branch connection between node i and node j at time t, or the branch switch is in the open state; β ij,t β represents whether node j is the parent node of node i at time t. ij,t =1 means that at time t, node j is the parent node of node i, β ij,t =0 means that at time t, node j is not the parent node of node i; u l,t u represents the switch status of branch l at time t. l,t =1 indicates that the switch of branch l is in the closed state at time t, u l,t =0 means that the branch l switch is in the open state at time t; Λ is the set of all branches, S is the set of all power supply nodes, Ω is the set of all nodes, and Ω / S is the set of all non-power supply nodes; The power balance constraint at distribution network nodes is expressed as follows: Among them, P ij,s (t) represents the active power flowing from node i to node j in scenario s at time t, P ki,s (t) represents the active power flowing from node k to node i in scenario s at time t; This represents the active power output of the main power supply located at node i in scenario s at time t; This represents the active power output of the distributed power source located at node i in scenario s at time t. The power range constraint for distribution network nodes is expressed as follows: 0≤p i,s (t)≤P i (t) (12) in, and These represent the upper limits of active power output of the main power source and the distributed power source located at node i, respectively. The power range constraint for distribution network branches is expressed as follows: y l,t,s =u l,t *S l,t,s (16) Among them, y l,t,s y represents the connection status of branch l in scenario s at time t, and is a binary variable of 0-1. l,t,s =1 indicates that branch l is in a connected state in scenario s at time t, y l,t,s =0 means that branch l is in an unconnected state under scenario s at time t; S represents the maximum transmission power of the branch; l,s,t Let S be the fault condition of branch l in scenario s at time t, and be a binary variable of 0-1. l,s,t =1 means that branch l is not faulty in scenario s at time t, S l,s,t =0 indicates that branch l is faulty in scenario s at time t; The constraints on the number of distribution network reconfigurations and the number of switching operations are expressed as follows: ∑ t ∑ l a l,t ≤max_action*once_max (21) Among them, a l,t This represents whether branch l is switched at time t, and is a binary variable of 0 and 1; a l,t =1 indicates that branch l is switched at time t; a l,t =0 means that branch l was not switched at time t; on o ff l,t This represents the operation performed by the switch in branch l at time t, with values of 1, 0, and -1; on o ff l,t =1, indicating that the switch of branch l is closed at time t; on o ff l,t =0, indicating that the switch of branch l was not operated at time t; on o ff l,t =-1, indicating that the switch of branch l is disconnected at time t; `max_action` and `once_max` represent the maximum number of refactoring actions in a single refactoring plan and the maximum number of on / off actions during each refactoring, respectively.
3. The method according to claim 2, characterized in that, The process of transforming the disaster-affected power distribution network reconstruction model into a Markov decision process model includes: 1) Construct the state variables for the reconfiguration operation of the distribution network during the disaster at time t: Among them, Pro break This is a fault probability vector for all branches within the distribution network at all times, with the dimension being the total number of branches. Each element in this vector represents the fault probability of the corresponding branch. Let t be the vector of switch closure status of all branches within the distribution network at time t. The dimension is the total number of branches. Each element in this vector is a binary variable of 0-1, representing the switch state of the corresponding branch. Let be the predicted active load vector of all nodes within the distribution network at time t, with dimension equal to the total number of nodes. Each element in this vector represents the predicted active load value of the corresponding node. Let t be the identifier vector indicating whether each action can be executed under all constraints at time t. Its dimension is the total number of branches + 1, corresponding to the dimension of the action space. Each element of this vector is a binary variable of 0-1, representing whether the corresponding action can be executed. in, The process of determining the value of each element in the algorithm is as follows: After each acquisition of the distribution network status, pre-simulation is performed on all actions sequentially, and actions that exceed the constraints of the reconstructed model during the disaster period are excluded. Set the corresponding element in the middle to 0; 2) Construct variables for the reconfiguration operations of the distribution network during the disaster at time t: a RL,t =(a t ,keep) (23) Among them, a t The vector represents whether all branches within the distribution network have performed switching operations at time t. Its dimension is the total number of branches. Each element in this vector is a binary variable of 0 and 1, representing whether the corresponding branch has performed a switching operation. keep is a binary variable of 0 and 1, indicating whether the distribution network will no longer perform switching operations at time t. keep = 0 means continue to perform switching operations; keep = 1 means stop performing switching operations. in: ∑a t +keep=1 (24) Where, ∑a t For a t The sum of all elements in the expression represents the number of switching operations performed at the current moment. 3) Construct the reward function for the reconfiguration operation of the distribution network during the disaster at time t: Where, r t The reward at time t; β is the expected increase in load provided by the distribution network topology at time t relative to the original topology. load and β switch The weights for load loss cost and switching operation cost are respectively, satisfying: β load ,β switch ∈(0,1) β load +β switch =1。 4. The method according to claim 3, characterized in that, The construction of the reinforcement learning agent corresponding to the power distribution network includes: 1) Constructing a policy neural network π for the reinforcement learning agent corresponding to the distribution network. θ Randomly initialize π θ The parameters θ; π θ The input is the state variable s RL,t The output is the action variable a. RL,t The probability distribution; For this policy neural network: Using softmax to generate action variable a RL,t Before the probability distribution, for each Elements that are equal to 0 in a are placed in a. RL,t If the value of the corresponding element is set to -∞, the probability of this action being selected is set to 0 after the softmax operation. 2) Construct the value function neural network for the reinforcement learning agent corresponding to the distribution network. Random initialization parameters The input is the state variable s RL,t The output is the agent's expected cumulative discount reward. The estimated value, where γ is the discount factor; 3) Construct a neural network for the objective value function of the reinforcement learning agent corresponding to the power distribution network. Structure and same, parameters initial value and parameters Their initial values are the same.
5. The method according to claim 4, characterized in that, The offline training of the reinforcement learning agent includes: 1) Construct an initial empty experience pool for the reinforcement learning agent to store experience samples (s) RL,t ,a RL,t ,r t ,s RL,t+1 ); 2) Randomly generate a set of scenarios, including: the fault probability of all branches in the distribution network at each time, and the predicted active load of all nodes in the distribution network at each time. 3) Based on the Markov decision process model, the reinforcement learning agent is made to perform a Markov interaction process with the power distribution network environment, and the experience sample data generated in the process is stored in the experience pool. Then determine: If the number of samples in the experience pool reaches the preset sample number threshold, proceed to step 4); otherwise, return to step 2). 4) Randomly select a set of samples from the experience pool of the reinforcement learning agent. The number of samples is Computational Value Function Neural Network Loss function: Among them, y t Approximate for time t The target value is calculated using the following expression: Calculated Then, update using gradient descent. Parameters; 5) Randomly select a set of samples from the experience pool of the reinforcement learning agent. The number of samples is Computational value function neural network π θ Loss function: Calculated Then, π is updated using gradient descent. θ Parameters; 6) Utilize the updated Parameter update Parameters: Where δ is the update rate of the target value function neural network; 7) Repeat steps 2)-6) until the upper limit of the set training rounds of the reinforcement learning agent is reached, and the reinforcement learning agent that has been trained offline is obtained.
6. The method according to claim 5, characterized in that, The process of generating an online reconstruction plan for the power distribution network during a disaster using the trained reinforcement learning agent includes: 1) At time t, construct the state variables of the Markov decision process using the measured values of the distribution network. 2) The state variable s obtained in step 4-1) RL,t Input a reinforcement learning agent that has been trained offline, and the reinforcement learning agent outputs the reconstruction operation action variable a at time t. RL,t =(a t (keep); 3) Determine the result of step 2): If a RL,t If keep = 0, then let s RL,t =s RL,t+1 Then return to step 2) to continue generating the next switching operation at time t; if a RL,t If keep = 1, then the switching operation at time t is completed. Let t = t + 1, and then return to step 1. Once all time-based switching operations have been generated, the distribution network reconfiguration plan generation is complete.
7. The method according to claim 6, characterized in that, Also includes: Reinforcement learning agents generate new samples (s) during online phases. RL,t ,a RL,t ,r t ,s RL,t+1 The data is then stored in the experience pool of the reinforcement learning agent; subsequently, the reinforcement learning agent is trained again using samples from the updated experience pool.
8. A device for generating reconfiguration plans for a power distribution network during a disaster based on reinforcement learning, characterized in that, include: The reconstructed model generation module is used to establish a reconstructed model of the power distribution network during a disaster. The Markov decision process model conversion module is used to convert the disaster-affected power distribution network reconstruction model into a Markov decision process model. The reinforcement learning agent training module is used to construct reinforcement learning agents corresponding to the power distribution network. Based on the Markov decision process model, training samples are generated for offline training of the reinforcement learning agent, and finally the trained reinforcement learning agent is obtained. The reconstruction plan generation module is used to generate reconstruction plans for the power distribution network during disasters online using the trained reinforcement learning agent.
9. An electronic device, characterized in that, include: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being configured to perform the method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method according to any one of claims 1-7.