Power distribution network dynamic maintenance decision-making method and system considering restoration time uncertainty
By establishing a deep reinforcement learning framework for Markov decision-making process model and DQN algorithm in post-disaster maintenance of distribution networks, the instability of power system caused by repair time uncertainty is solved, and economic losses are minimized and system stability is improved.
Patent Information
- Application Number
- CN202510335611.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The existing distribution network post-disaster maintenance decision model fails to effectively consider the uncertainty of repair time, resulting in increased power system stability and economic losses.
Establish a Markov decision-making process model, combine the deep reinforcement learning framework of the DQN algorithm, and train the neural network using the actual repair time scenario of the fault evaluation link to generate the optimal policy network, and dynamically generate maintenance decisions.
By considering the uncertainty of repair time, repair decisions are dynamically generated to reduce economic losses and improve the stability of the power system.
Smart Images

Figure CN120258766A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distribution network maintenance, and particularly relates to a dynamic maintenance decision-making method and system for a distribution network considering the uncertainty of repair time. Background Technique
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] Extreme weather events can cause a large number of substation and transmission line failures, resulting in power outages of users to varying degrees and causing significant economic losses. If the maintenance tasks of each key faulty facility cannot be allocated under limited resources and their maintenance order cannot be determined, it is very likely to delay the process of restoring power supply to the system. Therefore, formulating a reasonable maintenance strategy can prevent concurrent failures of other infrastructure facilities and prevent the further expansion of the accident scope, which has important practical significance for improving the resilience of the distribution network.
[0004] With the development of machine learning (ML), more and more work has focused on using ML methods to solve combinatorial optimization problems, such as the traveling salesman problem. Reinforcement learning (RL) is a method in the field of machine learning, which mainly focuses on how to learn and improve decision-making strategies based on environmental feedback. Deep reinforcement learning (DRL) is the combination of reinforcement learning and deep learning. Deep learning provides powerful representation learning capabilities, enabling the agent to better understand and represent the environment. In deep reinforcement learning, the agent uses a deep neural network to approximate the value function or policy, which enables the agent to handle more complex and high-dimensional state and action spaces.
[0005] The problem of post-disaster emergency repair of the distribution network is a problem of dispatching each repair team to each fault point in the distribution network for emergency repair until all faults are repaired. For this problem, it is of great significance to quickly formulate scientific and effective maintenance decisions. However, the existing maintenance decision-making model assumes that the repair time is a determined parameter before the implementation of the maintenance plan, without considering the uncertainty of the repair time. In actual engineering applications, the repair of faulty components is often affected by emergencies and human factors, and it is necessary for dispatchers to estimate the repair time based on on-site inspection results and combined with expert experience. Due to the urgency of post-disaster maintenance, there is a large deviation between the actual repair time and the predicted repair time, which challenges the stability of the power system, reduces the possibility of power supply, and causes economic losses, etc. Therefore, it is necessary to consider the uncertainty of the repair time when formulating the distribution network maintenance decision. Summary of the Invention
[0006] To overcome the deficiencies of the above-mentioned existing technologies, the present invention provides a method and system for dynamic maintenance decision-making of a distribution network considering the uncertainty of repair time, which fully considers the impact of the uncertainty of fault repair time on maintenance decision-making, dynamically generates maintenance decisions for the distribution network, reduces economic losses, and improves the stability of the power system.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] In the first aspect, the present invention provides a method for dynamic maintenance decision-making of a distribution network considering the uncertainty of repair time, including:
[0009] Establish the dynamic maintenance problem of faulty components in the distribution network as a Markov decision process model;
[0010] Build a deep reinforcement learning framework based on the DQN algorithm for the Markov decision process model, and use the actual fault repair time scenarios obtained in the fault assessment link to train the neural network to obtain an optimal policy network;
[0011] In the actual emergency repair stage, obtain the current fault status and input the current fault status into the optimal policy network to obtain the maintenance tasks to be executed;
[0012] Among them, the specific content of the Markov decision process model includes:
[0013] Construct decision points for the time points when maintenance tasks are assigned to maintenance teams;
[0014] Construct an action set assigned to the current idle maintenance team and the state of the currently observable maintenance sequence;
[0015] Design a reward function and a cost function according to the cumulative load supply deficit between decision points and the state transition path;
[0016] Build a Markov decision process model based on the decision points, actions, states, reward function, and the cost function.
[0017] In the second aspect, the present invention provides a system for dynamic maintenance decision-making of a distribution network considering the uncertainty of repair time, including:
[0018] A model establishment module, which is configured to: establish the dynamic maintenance problem of faulty components in the distribution network as a Markov decision process model;
[0019] Among them, the specific content of the Markov decision process model includes:
[0020] Construct decision points for the time points when maintenance tasks are assigned to maintenance teams;
[0021] Construct a set of actions assigned to the current idle maintenance team and the state of the currently observable maintenance sequence;
[0022] Design a reward function and a cost function based on the cumulative load supply deficit between decision points and the state transition path;
[0023] Construct a Markov decision process model based on the decision points, actions, states, reward function, and the cost function;
[0024] A training module, which is configured to: build a deep reinforcement learning framework based on the DQN algorithm for the Markov decision process model, and use the actual fault repair time scenarios obtained in the fault assessment session to train the neural network to obtain an optimal policy network;
[0025] A maintenance decision module, which is configured to: in the actual emergency repair stage, obtain the current fault state and input the current fault state into the optimal policy network to obtain the maintenance tasks to be executed.
[0026] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0027] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.
[0028] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.
[0029] The above one or more technical solutions have the following beneficial effects:
[0030] In the present invention, a Markov decision process model for the dynamic maintenance problem of faulty components in a distribution network is established. Based on the actual fault repair time scenarios, the neural network in the DQN algorithm is trained to obtain an optimal policy network, and then an online dynamic maintenance decision is obtained according to the optimal policy network. The present invention fully considers the impact of the uncertainty of fault repair time on maintenance decisions, dynamically generates maintenance decisions for the distribution network, minimizes economic losses, and improves the stability of the power system.
[0031] The advantages of the additional aspects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not constitute an improper limitation to the invention.
[0033] Figure 1 It is a schematic diagram of the relationship between the maintenance sequence and the load supply level curve in the first embodiment of the present invention;
[0034] Figure 2 It is a neural network architecture diagram of the DQN algorithm in the first embodiment of the present invention;
[0035] Figure 3 It is a flowchart of the operation of the DQN algorithm in the first embodiment of the present invention;
[0036] Figure 4 It is a schematic diagram of the overall process of dynamic maintenance decision-making in the first embodiment of the present invention. Detailed implementation manners
[0037] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0038] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners of the present invention.
[0039] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0040] Embodiment 1
[0041] This embodiment discloses a method for dynamic maintenance decision-making of a distribution network considering the uncertainty of repair time, including:
[0042] Step 1: Collect the distribution network and its fault information under extreme disasters to obtain the actual fault repair time scenario.
[0043] At the key nodes and frequently fault-occurring areas of the distribution network, sensors and Internet of Things devices are deployed to collect and monitor the state of the distribution network in real time. After the disaster occurs, information about the fault points, the damage conditions of the fault points, etc. in the post-disaster distribution network are obtained through the big data platform. After the data is collected, the fault information is cleaned and integrated, invalid data and error data are removed, and the data from different sources are integrated together for subsequent analysis and prediction.
[0044] (1) In the fault assessment link, obtain the probability distribution of the maintenance duration of the fault and generate the actual fault repair time scenario. At the same time, obtain the location of the base station and the information of the maintenance team from the power department.
[0045] (2) During the actual maintenance process, the status information of the faulty component is returned in real time, interacts with the DQN neural network, and obtains the action in this state.
[0046] Step 2: Establish a Markov decision process model for the dynamic maintenance problem of faulty components in the distribution network.
[0047] Specifically, the dynamic maintenance decision-making process exhibits the Markov property, that is, the future loss value only depends on the maintenance sequence observed at the current decision point and the newly arranged maintenance tasks.
[0048] The following part specifically describes the dynamic maintenance decision problem, and defines elements such as decision points, actions, states, and rewards.
[0049] (1) Decision points under dynamic decision-making.
[0050] The decision point is the time point when the dispatcher assigns a new maintenance task to an idle maintenance team. Assigning a new task to an idle maintenance team will not interrupt the ongoing maintenance, so multiple maintenance teams are allowed to perform maintenance simultaneously. For example, in Figure 1 the maintenance sequence, when assigning the maintenance task of faulty component C5 to maintenance team 2 at decision point DP4, maintenance team 1 will continue to maintain C2 without being affected. The following two situations will generate decision points:
[0051] Situation 1: A certain maintenance team just finishes the maintenance of a certain faulty component, and at this time the maintenance team changes from a busy state to an idle state.
[0052] Situation 2: A certain maintenance team starts to be schedulable at a certain moment. At this time, the maintenance team has not performed any maintenance tasks and needs to be assigned a new task.
[0053] Figure 1 There are a total of 6 decision points in . Among them, DP3, DP4, DP5, and DP6 correspond to Situation 1 and are generated after the maintenance of C1, C4, C2, and C5 respectively. DP1 and DP2 correspond to Situation 2. Whenever a decision point is generated, the dynamic maintenance decision-making process will enter a new decision-making stage, and each stage corresponds to an accumulated load supply deficit. For example, the generation of DP4 marks the entry into stage 4, which is represented by DS4 in the figure. The cumulative system load supply deficit in this stage is Ξ 4→5 . The moments of two different decision points may coincide. For example, the decision points DP1 and DP2 generated by Situation 2 are at the same moment. Among them, the number of decision points and the number of stages are equal to the number of faulty components. Therefore, the serial numbers of decision points and stages are set as τ = 1, 2,..., |Ω a |. Among them, Ω a represents the set of all faulty components, and |Ω a | is Ωa The number of elements in
[0054] (2) Actions under dynamic decision-making.
[0055] In the dynamic maintenance decision-making model, an action refers to a maintenance task assigned to the currently idle maintenance team. The action set describes the feasible domain of the actions. The action sets at each decision point are different, as shown in the following formula:
[0056]
[0057] In the formula, A τ represents the action set at decision point τ; U τ represents the set of components that have not been scheduled for maintenance at decision point τ; t τ represents the time at the current decision point; r(τ) represents the idle maintenance team that needs to be assigned tasks at decision point τ; H r(τ) represents the set of repairable components of r(τ); the schedulable time window of r(τ) is [T r(τ),min , T r(τ),max . and respectively represent the lead time and duration of maintenance task a τ , where is a random variable.
[0058] Equation (1) indicates that to assign maintenance task a τ at decision point τ, a τ must not have been assigned to any maintenance team before, and the maintenance team r(τ) has the ability to repair a τ . Another case is that the set may contain the element a τ = Null, indicating that when the action set is empty or all faulty components have been repaired, no new tasks are assigned to the maintenance team.
[0059] Meanwhile, Equation (1) also needs to satisfy two constraint conditions. The first constraint: The time at the decision point must be greater than the lower limit T r(τ),min of the callable time window of the maintenance team r(τ), that is, r(τ) is in a schedulable state; The second constraint: The probability that r(τ) finishes a τ later than the upper limit T r(τ),max of its schedulable time window is not greater than ε. The smaller ε is, the smaller the tolerance for the completion time exceeding the working period. There may be multiple intervals for the callable time window of the maintenance team. For example, if the maintenance process is interrupted due to unexpected reasons, at this time, the time window [T r(τ),min , T r(τ),max will be cut into multiple sub-windows, and the scheduler needs to adjust the action set according to the state after the maintenance interruption.
[0060] (3) State under dynamic decision-making.
[0061] The state in dynamic maintenance decision-making contains all the information of the currently observable maintenance sequence. The features included in the state are as follows:
[0062] S τ ={C τ ,O τ ,o τ ,s τ ,p τ ,t τ}
[0063] Among them, S τ represents the state of decision point τ; C τ and O τ represent the set of components that have been repaired and the set of components being repaired at decision point τ respectively; o τ is used to save the elapsed repair time of all maintenance tasks in O τ ; s τ saves the start times of all maintenance tasks that have been assigned (i.e., the elements in C τ and O τ ) before time t τ ; p τ saves the last maintenance task executed by all maintenance teams before decision point τ; finally, the time t τ at which the decision point is located is also one of the state features.
[0064] When the decision point is generated by case one, before the maintenance team r(τ) accepts a new order, it needs to feedback the following information to the dispatcher:
[0065]
[0066] Among them, e τ represents the maintenance task completed by r(τ) at the decision point; and represent the observed maintenance duration and the end time of the maintenance of e τ , which can be regarded as a sample instance of random variables and . If the current decision point is generated by case two, then let e τ be an empty set, and be 0.
[0067] (4) Reward under dynamic decision-making.
[0068] The reward in dynamic maintenance decision-making represents the state S τ of the current decision point after selecting action a τ+1 until the state S of the next decision pointτ+1 The obtainable benefit, i.e., Ξ τ→τ+1 represents the cumulative load supply deficit between adjacent decision points τ and τ + 1. Ξ τ→τ+1 The specific calculation formula of Ξ is:
[0069]
[0070] In the formula, represents the load supply value of node b, i.e., the economic loss per unit of lost load; and respectively represent the proportion of the node load supply deficit and the ideal node load supply level; v i,t is the operable state of the faulty component i at time t. Ξ τ→τ+1 Ξ can be expressed as the integral of the load supply deficit over the time interval [t τ , t τ+1 .
[0071] Inside the integral of Equation (2), the first and second constraints indicate that the repaired components must be in the operable state during the time period [t τ , t τ , t τ+1 , while the un-repaired faulty components are in the inoperable state. In the third constraint, x and Λ represent the operation-related variables and the set of operation constraints respectively.
[0072] Specifically, Λ contains the following constraints:
[0073] Assuming that the component is put into operation immediately after repair, the operation states of each component are:
[0074]
[0075] In the formula, Ω b represents the set of node substations; Ω l represents the set of transmission lines; u b,t / u l,t represents the actual operation state of the faulty substation / line at time t; v b,t represents the operable state of the faulty substation at time t; s(l) and e(l) represent the start and end nodes of line l; v l,t / v s(l),t / v e(l),t represent the operable states of the faulty line / start node of the line / end node of the line at time t.
[0076] Node active and reactive power balance equations:
[0077]
[0078] In the formula, and respectively represent the active / reactive power output of the generator unit and the line power flow; V b,t represents the magnitude of the node voltage; and respectively represent the shunt conductance and susceptance of the node;
[0079] If the transmission line is in operation, the power flow of this line needs to satisfy the power flow equation:
[0080]
[0081] where V s(l),t / V e(l),t and θ s(l),t / θ e(l),t respectively represent the voltage magnitude / voltage phase angle at the head / tail end of the faulty line; and respectively represent the branch conductance and susceptance; M is an infinite value.
[0082] Constraints on the active and reactive power output range of the generator unit. If the substation where the generator unit is located is not in operation, the output of the generator unit is 0:
[0083]
[0084] where and respectively represent the maximum / minimum active / reactive power output of the generator unit.
[0085] Limit the power flow on each line:
[0086]
[0087] where represents the maximum active power of the line power flow;
[0088] Specify the range of the node voltage magnitude and phase angle:
[0089]
[0090] where V b,max / θ b,max and V b,min / θ b,min respectively represent the maximum / minimum values of the voltage magnitude / phase angle.
[0091] The load supply at each node cannot exceed its ideal load supply level:
[0092]
[0093] (5) State transition path under dynamic decision-making.
[0094] Combined with the current decision point state S τ , the feedback information E of the next decision point τ+1 , and the action a of the current decision point τ , the state S of the next decision point can be obtained τ+1 = Tr(S τ , E τ+1 , a τ ), which is specifically as follows:
[0095] The set C of repaired components at decision point τ + 1 τ+1 should be C τ union with the task e completed at this decision point τ+1 , that is, C τ+1 = C τ ∪ e τ+1 .
[0096] The set O of components under repair τ+1 should be the union of O τ and the task a arranged in stage τ τ , while deleting e τ+1 , that is, O τ+1 = (O τ ∪ a τ ) / e τ+1 .
[0097] For the repair tasks in O τ+1 , the increment of the elapsed repair time is the difference between the times of decision points τ and τ + 1, which is described as:
[0098]
[0099] Update the start time of the task a assigned at decision point τ τ to t τ :
[0100]
[0101] Update the last task currently being executed by r(τ):
[0102] p r(τ+1),τ+1 = e τ+1
[0103] Finally, set the time of decision point τ + 1 to the end time of e τ+1 :
[0104]
[0105] In particular, the state of the initial decision point (τ + 1) is set as follows:
[0106]
[0107] At the initial decision point, no failed components have been repaired or are under repair. Therefore, C τ , O τ , o τ and p τ are all empty sets. The current time and the repair task times of each component are both defaulted to 0. When the repair process is interrupted, the dispatcher starts dynamic decision-making from the initial decision point again, and the state of the initial decision point needs to be set according to the actually observed repair sequence.
[0108] The termination decision point τe is defined as:
[0109]
[0110] In the termination decision point, all failed components must be repaired, that is, C τ =Ω a and O τ and o τ are all empty sets. The time of this decision point is set to T, and generally a relatively large value is set as the maximum time span of the post-disaster recovery stage. s τe and p τe are affected by the uncertainty of the repair time and the actions of the decision point, so their termination conditions are not restricted. The termination decision point τe does not need to issue a scheduling instruction and is only used to judge whether the repair process is over.
[0111] The status of the repair sequence starts from S1 and changes each time a new task is assigned to the repair team at a decision point until it becomes S τe . During this process, the changing process of the status {S1→S2→…S τe} is called the state transition path. Hereinafter, the state transition path between the decision point τ and τ′ (τ′>τ) is briefly written as S τ →S τ′ .
[0112] (6) The objective function under dynamic decision-making.
[0113] The Markov decision process model takes the cost function J τ (S τ ) as the objective function. The cost function J τ (S τ ) is defined as the expected cumulative load supply deficit of all possible state transition paths S τ →S τe between the decision point τ and τe:
[0114]
[0115] In the formula, Denote the policy function that maps each state to the action taken in that state, i.e., a τ = μ τ (S τ ); Ξ τ→τe Denote the state transition path S τ → S τe The cumulative load deficit between decision points τ and τe under the state; Ξ τ→τe Further divided by the decision point into the cumulative load supply deficit between adjacent decision points. γ is the value decay function. When γ < 1, it means that the dispatcher is more sensitive to the current load loss. Here, it is assumed that the value of the load supply deficit in each stage is equal, i.e., γ = 1.
[0116] According to Bellman's optimal theorem, the essence of solving the above cost function is to find an optimal policy that satisfies:
[0117]
[0118] where Denote the optimal cost function value corresponding to state S τ .
[0119] Use Q(S τ , a τ ) to measure the expected value of the reward that can be obtained after taking action a τ starting from state S τ . Its relationship with the cost function is:
[0120]
[0121] For the optimal cost function and the optimal action value function Q * (S τ , a τ ), the relationship can be expressed by the following formula:
[0122]
[0123] Equation (13) means that the optimal cost function value of state S τ is equal to the minimum of the optimal Q values of all possible actions in that state, i.e., the Q value obtained by taking the optimal action under state S τ is the optimal cost function value of that state
[0124] Step 3: Build a deep reinforcement learning framework based on the DQN algorithm for the Markov decision process model, and use the actual fault repair time scenario obtained from the fault assessment link to train the neural network to obtain the optimal policy network.
[0125] The essence of the DQN algorithm is to solve the Bellman optimal equation of an action-value function:
[0126]
[0127] That is, for the current state S τ taking action a τ the optimal action value Q * (S τ , a τ ) is equal to the immediately obtained reward Ξ τ→τ+1 plus the optimal action value Q * (S τ+1 , a τ+1 ) obtained by choosing the optimal action in the next state. The discount factor γ is used to adjust the balance between current and future rewards.
[0128] There is a recursive relationship in the above formula. In theory, starting from the termination decision point, the optimal actions of all states can be obtained by backward recursion. However, in reality, there are countless states just for the termination decision point. Therefore, backward recursion based on state enumeration cannot be used for the above Markov decision process model.
[0129] Therefore, in this embodiment, the DQN algorithm is adopted, and by training the neural network therein, the optimal behavior is continuously approximated to determine the optimal action in each state.
[0130] From the above model, the characteristics of state S τ can be analyzed. C τ and O τ are both high-order discrete variables, while o τ , s τ and t τ are continuous variables; for the action set A τ which are the faulty components to be repaired and belong to the discrete action set. The DQN algorithm is a reinforcement learning algorithm based on Q-learning, used to solve the problem of value function approximation in reinforcement learning. In the DQN algorithm, the agent selects actions according to the current state and receives feedback from the environment (i.e., rewards and the next state). For the problem of discrete actions in continuous states, the DQN algorithm approximates the Q-value function by introducing function approximation, such as neural networks, so that the agent can make decisions in the continuous state space. The architecture of DQN is as Figure 2 shown. This neural network consists of an input layer, two convolutional layers (CNN) and one fully connected layer (output layer).
[0131] Based on the above MDP model, the state at the current decision point is input into the neural network to estimate the Q value. Then, through continuous iterative updates, the Q value gradually approaches the true optimal value, thereby approaching the optimal policy and obtaining the optimal action in the current state. Specifically:
[0132] Step 3-1: According to the Temporal Difference (TD) algorithm, each update of the Q value in each step approaches in the direction of the sum of the estimated optimal Q value at the next moment and the immediate reward. The iterative update form of the Bellman equation is as follows:
[0133]
[0134] In the formula, α represents the learning factor.
[0135] The above formula uses the temporal difference learning objective to incrementally update Q(S τ ,a τ ), that is, move Q(S τ ,a τ ) closer to the TD target .
[0136] To make Q(S τ ,a τ ) converge to the true TD target, α needs to satisfy the following conditions:
[0137] Σ t α t (x)=∞
[0138]
[0139] Step 3-2: Input the current state into the Q-function neural network to obtain the Q values of each action, and use the ε-greedy greedy algorithm to sample and execute actions. Specifically
[0140] During the interaction between the agent and the environment, the agent takes the action corresponding to the optimal Q value estimated by the Q-function neural network at each step to ensure that when the Q value estimated by the model iterates to the optimal value, the corresponding policy of the agent is also the optimal policy. Just to avoid falling into local optima, when the agent takes an action, it will still randomly select an action from among many actions with a small probability.
[0141] Let the probability of selecting a random action be ε, and sample p ∈ [0,1] according to a uniform distribution. Then the action a τ selected at the decision point τ satisfies:
[0142]
[0143] where rand(·) represents random sampling according to a uniform distribution from the action set, and A τ represents the action set at the decision point τ.
[0144] Step 3-3, Experience replay. After executing the current action in the current state, the agent immediately receives the reward and the next state returned by the environment. Then, collect the quadruple composed of the current state, the current action, the current reward, and the next state and put it into the experience replay pool. Specifically,
[0145] In general supervised learning, it is assumed that the training data is independently and identically distributed. Each time the neural network is trained, one or several data are randomly sampled from the training data for gradient descent. As learning progresses, each training data will be used multiple times. Therefore, in order to better combine Q-learning and deep neural networks, the DQN algorithm adopts the experience replay method. The specific approach is to set up an experience replay pool, store the quadruple data (state, action, reward, next state) sampled from the environment each time into the experience replay pool, and then randomly sample several data from the experience replay pool for training when training the Q network. This can achieve the following two functions:
[0146] (1) Make the samples satisfy the independent assumption. The data sampled by interacting in the MDP itself does not satisfy the independent assumption because the state at this moment is related to the state at the previous moment. Non-independent and identically distributed data has a great impact on training the neural network and will cause the neural network to fit to the recently trained data. Using experience replay can break the correlation between samples and make them satisfy the independent assumption.
[0147] (2) Improve the sample efficiency. Each sample can be used multiple times, which is very suitable for the gradient learning of deep neural networks.
[0148] Step 3-4, When the number of quadruples in the experience replay pool is large enough, batch sample quadruples from the experience replay pool to calculate the Q-value loss function and update the neural network parameters ω. If there are still faulty components not repaired, return to Step 3-2; otherwise, end the training of one round of the episode. Specifically,
[0149] When there is enough data in the experience replay pool, randomly sample a small batch of data for training. Then, for a set of data (S τi , a τi , Ξ τi→τi+1 , S τi+1 ), the loss function of the Q network can be constructed in the form of mean squared error:
[0150]
[0151] Where ω are the neural network parameters.
[0152] During the training process, since the target itself contains the output of the neural network, the target keeps changing while updating the network parameters, which causes instability in the neural network training. Therefore, two sets of Q networks are needed. One is the original training network Q ω (S τ ,a τ ), which is used to calculate the Q ω (S τ ,a τ ) term in the original loss function, and the normal gradient descent method is used for updating. The other is the target network Q ω′ (S τ ,a τ ), which is used to calculate the term in the loss function, where ω′ represents the parameters in the target network.
[0153] Meanwhile, to make the update target more stable, the target network is not updated at every step. The target network uses a set of older parameters from the training network. The training network Q ω (S τ ,a τ ) is updated at every step during training, while the parameters of the target network are synchronized with the training network only every n time steps, i.e., ω′←ω. This makes the target network more stable than the training network.
[0154] In summary, the specific process of the DQN algorithm is as follows:
[0155] Initialize the network Q ω (S τ ,a τ ) with random network parameters ω;
[0156] Initialize the target network Q ω′ (S τ ,a τ ) by copying the same network parameters ω′←ω;
[0157] Initialize the experience replay pool R;
[0158] for episode e = 1 → E do
[0159] Obtain the environmental state S1;
[0160] for time step τ = 1 → τe do
[0161] Select an action according to the current network Q ω (S τ ,a τ ) with the ε-greedy greedy policy;
[0162] Execute action a τ , and obtain a reward Ξ τ→τ+1 , and the environmental state changes to S τ+1 ;
[0163] Store (S τ , a τ , Ξ τ→τ+1 , S τ+1 ) into the experience replay pool R;
[0164] If there is enough data in R, sample N data {(S τi , a τi , Ξ τi→τi+1 , S τi+1 )} i=1,...,N ;
[0165] For each data, calculate with the target network
[0166] Minimize the target loss And update the current network Q ω (S τ , a τ );
[0167] Synchronize the target network and the training network every n time steps, and update the target network accordingly;
[0168] end for
[0169] end for
[0170] The high-level workflow of the DQN algorithm is as Figure 3 shown.
[0171] After training for a sufficient number of episodes, end the training. For subsequent work, the trained Q-function neural network can be directly used to determine the actions in each state and execute them.
[0172] Step 4, actual emergency repair phase, obtain the current fault state, and input the current fault state into the optimal policy network to obtain the maintenance tasks to be executed.
[0173] Step 4-1, based on the above Step 1, obtain the probability distribution of the maintenance duration of the fault from the fault assessment link, and generate the actual fault repair time scenario.
[0174] Step 4-2, based on the above Step 2, establish a Markov decision process model. For the above actual fault repair time scenario, based on the above Step 3, train the Q-function neural network.
[0175] Step 4-3. In the actual emergency repair process, input the current fault status into the trained neural network to obtain the Q function corresponding to each action, and select the action with the minimum Q value. Assign the maintenance task corresponding to this action to the maintenance team idle at the current decision point.
[0176] Step 4-4. Wait for the maintenance team to feedback new information. According to the feedback information E τ+1 and the current action Based on the state transition path, jump to the next decision point and update the state.
[0177] Loop and execute Step 4-3 to Step 4-4 until the maintenance tasks of all faulty components have been assigned to the maintenance team, then the entire decision-making process can be completed.
[0178] Figure 4 summarizes the overall calculation process of the dynamic maintenance decision-making model. Among them, the blue box and the red box represent the offline training and online decision-making processes respectively.
[0179] Embodiment 2
[0180] The purpose of this embodiment is to provide a distribution network dynamic maintenance decision-making system considering the uncertainty of repair time, including:
[0181] A model establishment module, which is configured to: establish the dynamic maintenance problem of faulty components in the distribution network as a Markov decision process model;
[0182] Among them, the specific content of the Markov decision process model includes:
[0183] Construct decision points for the time points when maintenance tasks are assigned to maintenance teams;
[0184] Construct a set of actions assigned to the currently idle maintenance team and the state of the currently observable maintenance sequence;
[0185] Design a reward function and a cost function according to the cumulative load supply deficit between decision points and the state transition path;
[0186] Based on the decision points, actions, states, reward function and the cost function, construct a Markov decision process model;
[0187] A training module, which is configured to: build a deep reinforcement learning framework based on the DQN algorithm for the Markov decision process model, and use the actual fault repair time scenarios obtained in the fault assessment link to train the neural network to obtain an optimal policy network;
[0188] A maintenance decision-making module, which is configured to: in the actual emergency repair stage, obtain the current fault status and input the current fault status into the optimal policy network to obtain the maintenance tasks to be executed.
[0189] In more embodiments, the following is also provided:
[0190] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0191] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0192] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0193] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.
[0194] The method in Embodiment 1 can be directly embodied as being executed and completed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0195] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented and completed.
[0196] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the processes / methods described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided among program modules as needed. The machine-executable instructions for program modules can be executed within a local or distributed device. In a distributed device, program modules can be located in local and remote storage media.
[0197] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0198] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that a device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.
[0199] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0200] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A dynamic maintenance decision-making method for a distribution network considering the uncertainty of repair time, characterized in that, Including: Establish the dynamic maintenance problem of faulty components in the distribution network as a Markov decision process model; Build a deep reinforcement learning framework based on the DQN algorithm for the Markov decision process model, and use the actual fault repair time scenario obtained in the fault assessment link to train the neural network to obtain the optimal policy network; In the actual emergency repair stage, obtain the current fault status and input the current fault status into the optimal policy network to obtain the maintenance tasks to be executed; Among them, the specific content of the Markov decision process model includes: Construct decision points for the time points when maintenance tasks are assigned to maintenance teams; Construct the action set assigned to the current idle maintenance team and the state of the currently observable maintenance sequence; Design a reward function and a cost function according to the cumulative load supply deficit, state transition path between decision points, and state transition path; Construct a Markov decision process model based on the decision points, actions, states, reward function, and cost function.
2. The dynamic maintenance decision-making method for a distribution network considering the uncertainty of repair time according to claim 1, wherein, The action sets under each decision point are represented as: Among them, A τ represents the set of actions at decision point τ; U τ represents the set of components that have not been scheduled for maintenance at decision point τ; t τ represents the time at the current decision point; r(τ) represents the idle maintenance team that needs to allocate tasks at decision point τ; H r(τ) represents the set of repairable components of r(τ); the schedulable time window of r(τ) is [T r(τ),min , T r(τ),max , and respectively represent the lead time and duration of maintenance task a τ , where is a random variable; The state under each decision point contains all the information of the currently observable maintenance sequence, specifically: the time t at which the decision point itself is located τ , the set of components that have been repaired at the decision point and the set of components being repaired; the elapsed repair time of all maintenance tasks in the set of components being repaired at the decision point; all the start times of the maintenance tasks that have been assigned before time t τ ; the last maintenance task performed by all maintenance teams before the decision point.
3. A dynamic maintenance decision-making method for a distribution network considering the uncertainty of repair time as described in claim 1, characterized in that Design a reward function according to the cumulative load supply deficit between decision points, specifically: Among them, represents the load supply value of node b, that is, the economic loss per unit of lost load; and respectively represent the proportion of the node load supply deficit and the ideal node load supply level; v i,t is the operable state of the faulty component i at time t; x and Λ respectively represent the operation-related variables and the set of operation constraints.
4. A dynamic maintenance decision-making method for a distribution network considering the uncertainty of repair time as described in claim 1, characterized in that, Establishing a Markov decision process model for the dynamic maintenance problem of faulty components in the distribution network also includes: Construct the state transition path S τ+1 = Tr(S τ , E τ+1 , a τ ), where S τ+1 is the state of the next decision point, S τ is the state of the current decision point, E τ+1 is the feedback information of the next decision point, a τ is the action of the current decision point; Based on the state transition path, calculate the cost function for dynamic decision-making starting from the current decision point; among them, the cost function is the expected cumulative load supply deficit of all possible state transition paths between the current decision point and the next decision point; According to the cost function, determine the optimal cost function and the optimal action value function.
5. A dynamic maintenance decision-making method for a distribution network considering the uncertainty of repair time as claimed in claim 4, characterized in that Build a deep reinforcement learning framework based on the DQN algorithm for the Markov decision process model, and use the actual fault repair time scenario obtained in the fault assessment link to train the neural network to obtain the optimal policy network, specifically: Initialize the training network and the target network with random network parameters; Initialize the experience replay pool; Input the current state into the training network, select an action using the ε-greedy greedy algorithm and execute it to obtain the corresponding reward; After executing the current action in the current state, collect the current state, current action, current reward, and next state and form a quadruple, and put the quadruple into the experience recovery pool; Batch sample quadruples from the experience recovery pool to calculate the Q-value loss of the target network, and update the parameters of the training network with the goal of minimizing the target loss.
6. The dynamic maintenance decision-making method for a distribution network considering the uncertainty of repair time according to claim 5, wherein, For the trained Q-function neural network, formulate a dynamic maintenance decision method, determine the actions in each state and execute them to achieve online dynamic maintenance decision-making for the distribution network after a disaster, specifically: In the actual emergency repair link, input the current fault status into the trained neural network to obtain the Q function corresponding to each action, and select the action that minimizes the Q value. Assign the maintenance task corresponding to this action to the maintenance team idle at the current decision point; Wait for the maintenance team to provide new feedback. Based on the feedback information E τ+1 and the current action a τ * , based on the state transition path, jump to the next decision point and update the state; Loop until the maintenance tasks of all faulty components have been assigned to the maintenance team, and the entire decision-making process can be completed.
7. A distribution network dynamic maintenance decision-making system considering the uncertainty of repair time, characterized in that Including: A model establishment module, which is configured to: establish the dynamic maintenance problem of faulty components in the distribution network as a Markov decision process model; Among them, the specific content of the Markov decision process model includes: Construct decision points for the time points when maintenance tasks are assigned to maintenance teams; Construct a set of actions assigned to the currently idle maintenance teams, as well as the state of the currently observable maintenance sequence; Design a reward function and a cost function based on the cumulative load supply deficit between decision points and the state transition path; Construct a Markov decision process model based on the decision points, actions, states, reward function, and the cost function; A training module, which is configured to: build a deep reinforcement learning framework based on the DQN algorithm for the Markov decision process model, and use the actual fault repair time scenarios obtained in the fault assessment link to train the neural network to obtain an optimal policy network; A maintenance decision module, which is configured to: in the actual emergency repair stage, obtain the current fault state and input the current fault state into the optimal policy network to obtain the maintenance tasks to be executed.
8. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method according to any one of claims 1-6 is completed.
9. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the method according to any one of claims 1-6 is completed.
10. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-6 is implemented.