A method for constructing a repair decision model, a system, and a power grid system repair method

By adopting a dual competitive deep Q network and DC optimal current model in the power grid system, the messy and time-consuming problems of the power grid recovery process are solved, efficient and accurate repair decisions are achieved, and the rapid recovery ability and flexibility of the power grid system are improved.

CN116306298BActive Publication Date: 2025-07-25GUANGDONG POWER GRID CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310291421.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2025-07-25
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

The existing technology has the power grid recovery process that is chaotic and time-consuming in the event of large-scale power outages, making it difficult to respond instantly to the repair time of the faulty components, resulting in long-term failure of the power grid system, and the existing deep reinforcement learning methods cannot accurately reflect the elastic changes of the power grid and the insufficient learning ability of large-scale action space.

Method used

A dual competitive deep Q network is adopted, combining the node constraints and distribution line constraints of the power grid system, a single-stage DC optimal current model is established, and the repair time problem is expressed as the Markov dynamic decision-making process. The repair decision model is trained through the gradient descent method to simplify the sampling steps to achieve efficient repair decision-making.

Benefits of technology

It improves the output efficiency and accuracy of the repair decision model, ensures that the power grid system recovers quickly after failure, improves the system's elasticity and simplifies repair decisions under large-scale state action space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116306298B_ABST
    Figure CN116306298B_ABST
Patent Text Reader

Abstract

The present invention discloses a construction method of a repair decision model, a system and a power grid system repair method, including: taking the maximization of system resilience as the objective function, and establishing a single-stage DC optimal power flow model according to the node constraints and distribution line constraints of the power grid system; obtaining the repair time problems of all distribution lines in the power grid system, and based on the single-stage DC optimal power flow model, formulating the repair time problems as a Markov decision process; using a double-competitive deep Q-network to solve the Markov decision process to obtain the corresponding loss function, and training the original deep Q-network based on the loss function to obtain a repair decision model. The present invention takes the maximization of system resilience as the objective function, establishes the repair time problems, fully considers the changes in system resilience, and solves the decision process corresponding to the repair time problems through a double-competitive deep Q-network, simplifies the sampling steps, and realizes the efficient output of repair decisions under a large-scale state-action space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power grid system fault repair, and particularly to a method for constructing a repair decision model, a system, and a power grid system repair method. Background Art

[0002] The increasing dependence of modern society on electric energy has prompted the power grid to respond to known threats in a reliable manner and ensure high-quality power supply. The power system often shows high vulnerability in the face of natural disasters, so elastic strategies are needed to improve the ability of the power grid to absorb shocks and quickly recover from destructive events. Frequent and severe natural disasters have been constantly plaguing power grid managers, which highlights the importance of power grid resilience research. Timely repair of faulty components such as power towers and distribution lines after extreme outages can effectively reduce load shedding in the power system.

[0003] Currently, the method of optimization modeling is usually adopted to obtain the best repair sequence for restoring the function of the power grid system. However, in the event of a large-scale power outage, the restoration of the power grid is usually chaotic and time-consuming. It is very difficult to solve complex optimization models and determine the best repair order for large-scale systems, and the solution time is too long. Especially when considering the repair time of uncertain faulty components, it is impossible to respond immediately to emergency outages. In addition, the deep reinforcement learning method can solve the combinatorial optimization problem of power grid resilience by automatically learning decision-making experience offline, and has recently been widely applied to the decision-making of uncertain emergencies in the power grid. However, the existing methods cannot accurately reflect the resilience changes of the power grid, and at the same time, the learning ability for considering the huge action space in an uncertain repair environment is still poor. Summary of the Invention

[0004] The present invention provides a method for constructing a repair decision model, a system, and a power grid system repair method, which realizes the maximization of system resilience, simplifies the sampling steps of the deep Q-network, improves the efficiency and accuracy of the repair decision model outputting repair decisions, and prevents the power grid system from being in a faulty state for a long time.

[0005] To solve the above technical problems, an embodiment of the present invention provides a method for constructing a repair decision model, including:

[0006] Taking the maximization of system resilience as the objective function, and establishing a single-stage DC optimal power flow model according to the node constraints and distribution line constraints of the power grid system;

[0007] Obtaining the repair time problem of all distribution lines in the power grid system, and formulating the repair time problem as a Markov decision process based on the single-stage DC optimal power flow model;

[0008] The double-competitive deep Q-network is used to solve the Markov dynamic decision-making process to obtain the loss function of the original deep Q-network in the double-competitive deep Q-network. Then, based on the loss function, the original deep Q-network is trained, and the trained original deep Q-network is used as the repair decision-making model;

[0009] Among them, the double-competitive deep Q-network includes the original deep Q-network, the target deep Q-network, and the competitive network structure.

[0010] Implementing the embodiments of the present invention, with the maximization of system resilience as the objective function and using the node constraints and distribution line constraints of the power grid system as the constraint functions, a single-stage DC optimal power flow model is established to determine the optimal energy distribution and repair time problem through the single-stage DC optimal power flow model. Then, the repair time problem is formulated as a Markov dynamic decision-making process to adapt to the solution mode of the double-competitive deep Q-network, resulting in the trained repair decision-making model not only being able to fully consider system resilience and maximize system resilience but also being able to simplify the sampling steps through the competitive network structure in the double-competitive deep Q-network, thereby achieving efficient output of repair decisions under a large-scale state-action space to assist the staff in carrying out maintenance work.

[0011] As a preferred solution, the double-competitive deep Q-network is used to solve the Markov dynamic decision-making process to obtain the loss function of the original deep Q-network in the double-competitive deep Q-network. Then, based on the loss function, the original deep Q-network is trained, and the trained original deep Q-network is used as the repair decision-making model, specifically:

[0012] The double-competitive deep Q-network is used to solve the Markov dynamic decision-making process to obtain the corresponding target value and estimated value, and based on the target value and the estimated value, the loss function of the original deep Q-network is calculated;

[0013] The gradient descent method is used to iteratively optimize the original deep Q-network in the double-competitive deep Q-network. During each iterative optimization, the network parameters of the original deep Q-network are adjusted, and the Markov dynamic decision-making process is solved based on the adjusted double-competitive deep Q-network to update the loss function until the current loss function is minimized. Then, the training of the original deep Q-network is completed, and the current original deep Q-network is used as the repair decision-making model.

[0014] Implementing the preferred solution of the embodiments of the present invention, when solving the Markov dynamic decision-making process using the dual-competitive deep Q-network, obtain the loss function of the current original deep Q-network, and use the gradient descent method with the minimization of the loss function of the original deep Q-network as the optimization objective and the network parameters of the original deep Q-network as the optimization variables to iteratively train the original deep Q-network, update the network parameters of the original deep Q-network, and then use the original deep Q-network that achieves the optimization objective as the repair decision model to further improve the accuracy of the repair decision model.

[0015] As a preferred solution, when using the dual-competitive deep Q-network to solve the Markov dynamic decision-making process to obtain the corresponding target value and estimated value, specifically:

[0016] According to the preset double Q-network algorithm, use the target deep Q-network in the dual-competitive deep Q-network to evaluate the optimal repair strategy output by the original deep Q-network to obtain the corresponding target value; where the double Q-network algorithm is specifically:

[0017]

[0018] In the formula, represents the target value corresponding to the optimal repair strategy of all distribution lines in the power grid system at the (s + 1) stage, represents the immediate benefit based on system resilience at the (s + 1) stage, γ represents the future benefit discount factor, represents the state space composed of the operating states of all distribution lines in the power grid system at the (s + 1) stage, A s+1 represents the action space composed of the repair actions of all distribution lines in the power grid system at the (s + 1) stage, θ p represents the network parameters of the original deep Q-network, θ w represents the network parameters of the target deep Q-network, Q p represents the output result of the original deep Q-network, Q w represents the output result of the target deep Q-network;

[0019] According to the preset competitive network algorithm, use the competitive network structure in the dual-competitive deep Q-network to solve the Markov dynamic decision-making process to obtain the corresponding estimated value; where the double Q-network algorithm is specifically:

[0020]

[0021] In the formula, represents the estimated value of the power grid system based on the state space and the action space A s of, Denote the state space composed of the operating states of all distribution lines in the power grid system at stage s, A s Denote the action space composed of the repair actions of all distribution lines in the power grid system at stage s, V(S s |θ p ) represents the state space The corresponding scalar, A s ′ represents the selected action space A s The corresponding possible alternative action space, Denote in the state space The advantage function of the action space A s under this condition, Denote in the state space The advantage function of the possible alternative action space A s ′ under this condition.

[0022] Implement the preferred solution of the embodiment of the present invention, adopt a double-competitive deep Q-network to solve the Markov dynamic decision-making process, so that the original deep Q-network selects the best repair action and the target deep Q-network evaluates this best repair action, preventing the problem that the original deep Q-network overestimates the Q-value of the repair action and leads to too large a target value, so as to improve the accuracy and stability of the repair decision-making model, enabling the repair decision-making model to output a more efficient power grid system repair strategy. In addition, through the competitive network structure in the double-competitive deep Q-network, estimate the operating state and repair action respectively to improve the accuracy of the estimated value - the Q-value, and avoid unnecessary estimation of each repair action.

[0023] As a preferred solution, the method for constructing a repair decision-making model further includes:

[0024] Establish the node constraints and the distribution line constraints of the power grid system;

[0025] Traverse each repair stage, take the maximization of system resilience as the objective function, and use the node constraints and the distribution line constraints as the limiting conditions of the decision variables, and adjust each decision variable to obtain the system resilience corresponding to each repair stage; among them, the objective function is specifically:

[0026]

[0027] In the formula, the decision variable g ns represents the power generation of node n in the power grid system at stage s, the decision variable d ns represents the amount of demand met by node n in the power grid system at stage s, the decision variable f ls represents the magnitude of the current passing through distribution line l in the power grid system at stage s, the decision variable θ nsDenote the phase angle of node n in the s stage of the power grid system, R s Denote the system resilience in the s stage of the restoration phase.

[0028] Implement the preferred solution of the embodiment of the present invention. Based on the power saving constraint and distribution line constraint of the power grid system, adjust the power generation amount of each node in the power grid system, the amount of demand met by each node, the magnitude of the current passing through each distribution line, and the phase angle of each node, so as to accurately reflect the resilience change of the power grid system.

[0029] As a preferred solution, the node constraints include: node power generation constraint, node demand satisfaction constraint, node flow balance constraint, and node phase angle constraint;

[0030] Among them, the node power generation constraint is specifically:

[0031]

[0032] In the formula, g ns Denote the power generation amount of node n in the s stage, Denote, Denote the set of nodes in the power grid system, S denotes the set of stages in the power grid system;

[0033] The node demand satisfaction constraint is specifically:

[0034]

[0035] In the formula, d ns Denote the amount of demand met by node n in the s stage, Denote;

[0036] The node flow balance constraint is specifically:

[0037]

[0038] In the formula, n and n ′ Respectively denote the two end nodes of the distribution line l;

[0039] The node phase angle constraint is specifically:

[0040]

[0041] In the formula, θ ns Denote the phase angle of node n in the s stage, θ Denote the lower limit constraint of the phase angle, Denote the upper limit constraint of the phase angle.

[0042] Implement the preferred solution of the embodiment of the present invention, and set constraint functions for the power generation of each node, the amount of demand met by each node, the flow balance of each node, and the phase angle of each node, so as to ensure the node balance of the distribution lines in the power grid system while meeting the goal of maximizing system resilience.

[0043] As a preferred solution, the distribution line constraints include: distribution line energy limit constraints and distribution line impedance and phase angle constraints;

[0044] Among them, the distribution line energy limit constraint is specifically:

[0045]

[0046] In the formula, f ls represents the magnitude of the current passing through distribution line l in stage s, z ls represents the operating state of distribution line l in stage s, f l max represents the maximum value of the magnitude of the current passing through distribution line l, represents the set of distribution lines in the power grid system;

[0047] The distribution line impedance and phase angle constraint is specifically:

[0048]

[0049]

[0050] In the formula, X R represents the impedance of distribution line l, represents the phase angle of one end node n1 of distribution line l, represents the phase angle of the other end node n2 of distribution line l, and M represents a very large natural number used to ensure the effectiveness of the constraint.

[0051] Implement the preferred solution of the embodiment of the present invention, and set corresponding constraint functions for the magnitude of the current passing through each distribution line and the impedance of each distribution line, so as to always ensure the load balance of the distribution lines in the power grid system during the process of achieving the goal of maximizing system resilience.

[0052] As a preferred solution, the problem of obtaining the repair time of all distribution lines in the power grid system is specifically:

[0053] According to the preset repair time algorithm, calculate the determined repair time of each distribution line based on the length of each distribution line in the power grid system, and add the determined repair time of each distribution line to the uncertainty of each distribution line respectively to obtain the repair time of each distribution line;

[0054] Express the repair times of all the distribution lines as a stochastic optimization sequence problem to obtain the repair time problem of all the distribution lines in the power grid system.

[0055] Implement the preferred solution of the embodiments of the present invention. Considering that the repair times of the damaged components in the power grid system may be affected by many uncertain factors, such as the fault location, the degree of component aging, logistics and technical delays, etc., on the basis of the determined repair times of each distribution line, add the uncertain quantities of each distribution line to obtain the repair times of each distribution line, improve the accuracy of the repair times, and thus optimize the performance of the repair decision-making model.

[0056] As a preferred solution, based on the single-stage DC optimal power flow model, express the repair time problem as a Markov decision process, specifically:

[0057] Through the single-stage DC optimal power flow model, obtain the optimal system resilience of the power grid system, and establish a reward function according to the optimal system resilience of the power grid system and the repair time problem;

[0058] According to the state space, the action space, the state transition probability function, the reward function, and the discount factor, establish a five-tuple to express the Markov decision process of the repair time problem.

[0059] Implement the preferred solution of the embodiments of the present invention. Based on the optimal system resilience obtained from the single-stage DC optimal power flow model and the time repair problem, establish a corresponding reward function, and according to the state space, the action space, the state transition probability function, the reward function, and the discount factor, establish a five-tuple, and use this five-tuple to express the Markov decision process of the repair time problem, so that the Markov decision process can be solved in reinforcement learning, thereby maximizing the future benefits of the repair decision-making model.

[0060] To solve the same technical problem, the embodiments of the present invention also provide a power grid system repair method based on deep reinforcement learning, including:

[0061] Obtain the operating states of all the distribution lines in the power grid system, and input all the operating states into the repair decision-making model so that the repair decision-making model outputs the corresponding optimal repair strategy; wherein, the repair decision-making model is constructed by using the construction method of a repair decision-making model as described above;

[0062] Repair all the distribution lines in the power grid system in sequence according to the optimal repair strategy.

[0063] To solve the same technical problem, an embodiment of the present invention further provides a system for constructing a repair decision model, including:

[0064] A model establishment module, which takes maximizing system elasticity as the objective function and establishes a single-stage DC optimal power flow model according to the node constraints and distribution line constraints of the power grid system;

[0065] A representation conversion module, which is used to obtain the repair time problems of all distribution lines in the power grid system, and based on the single-stage DC optimal power flow model, express the repair time problems as a Markov decision process;

[0066] A model training module, which uses a double-competitive deep Q network to solve the Markov decision process to obtain the loss function of the original deep Q network in the double-competitive deep Q network, and then trains the original deep Q network based on the loss function, and uses the trained original deep Q network as a repair decision model; wherein, the double-competitive deep Q network includes the original deep Q network, the target deep Q network and the competitive network structure. Description of the Drawings

[0067] Figure 1 : It is a schematic flowchart of a method for constructing a repair decision model provided in Embodiment 1 of the present invention;

[0068] Figure 2 : It is a schematic flowchart of a method for repairing a power grid system based on deep reinforcement learning provided in Embodiment 2 of the present invention;

[0069] Figure 3 : It is a schematic structural diagram of a system for constructing a repair decision model provided in Embodiment 3 of the present invention. Detailed Embodiments

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0071] Embodiment 1:

[0072] Please refer to Figure 1 , which is a method for constructing a repair decision model provided in an embodiment of the present invention. The method includes steps S1 to S3, and the specific steps are as follows:

[0073] Step S1: Establish a single - stage DC optimal power flow model with the maximization of system resilience as the objective function and according to the node constraints and distribution line constraints of the power grid system.

[0074] As an optimized solution, step S1 includes steps S11 to S14, and the specific steps are as follows:

[0075] Step S11: To reflect the performance of the power grid system, refer to Equation (1). Define the ratio of the demand met to the target demand at stage s as the system resilience index at stage s.

[0076]

[0077] In the formula, d ns represents the demand met at node n at stage s, represents the target demand at node n at stage s, t s represents the duration of stage s. Among them, nodes include supply nodes and demand nodes.

[0078] Step S12: Traverse each repair stage. Refer to Equation (2). With the maximization of system resilience as the objective function and taking the node constraints and distribution line constraints of the power grid system as the limiting conditions of decision variables, adjust each decision variable to obtain the system resilience R corresponding to each repair stage s ; Among them, taking the maximization of system resilience as the objective function is specifically:

[0079]

[0080] In the formula, the decision variable g ns represents the power generation at node n in the power grid system at stage s, the decision variable d ns represents the amount of demand met at node n in the power grid system at stage s, the decision variable f ls represents the magnitude of the current passing through distribution line l in the power grid system at stage s, the decision variable θ ns represents the phase angle of node n in the power grid system at stage s, R s represents the system resilience at stage s of the repair stage, which is used to describe the improvement of the resilience of the power grid system between two states.

[0081] In this embodiment, refer to Equations (3)(4)(5)(6)(7)(8)(9). The node constraints of the power grid system include node power generation constraints, node demand - met constraints, node flow - balance constraints, and node phase - angle constraints. The distribution line constraints of the power grid system include distribution line energy - limit constraints and distribution line impedance and phase - angle constraints.

[0082] Among them, the node power generation constraint is used to set the maximum power generation of supply nodes and demand nodes, specifically:

[0083]

[0084] In the formula, g ns represents the power generation of node n in stage s, represents, represents the set of nodes in the power grid system, and S represents the set of stages in the power grid system.

[0085] Among them, the node satisfies the demand constraint, which is used to set the satisfaction of the supply node and the demand node for the demand, specifically:

[0086]

[0087] In the formula, d ns represents the amount of demand satisfied by node n in stage s, represents the target demand of node n in stage s.

[0088] Among them, the distribution line energy limit constraint is used to determine the energy limit of each distribution line. The faulty line cannot transmit power before it is repaired, specifically:

[0089]

[0090] In the formula, f ls represents the magnitude of the current passing through distribution line l in stage s, z ls represents the operating state of distribution line l in stage s, f l max represents the maximum value of the magnitude of the current passing through distribution line l, represents the set of distribution lines in the power grid system.

[0091] Among them, the node flow balance constraint is used to limit the flow balance of each node, specifically:

[0092]

[0093] In the formula, n and n′ respectively represent the two end nodes of distribution line l.

[0094] Among them, the distribution line impedance and phase angle constraint is used for the DC power flow equation composed of the reactance on the line and the phase angle at the node, and limits the impedance of the distribution line and the phase angle of the distribution line, specifically:

[0095]

[0096]

[0097] In the formula, X R represents the impedance of distribution line l, Represents the phase angle of the end node n1 of the distribution line l, Represents the phase angle of the other end node n2 of the distribution line l, and M represents a very large natural number used to ensure the effectiveness of the constraint.

[0098] Among them, the node phase angle constraint is used to limit the feasible range of the phase angle of the node, specifically:

[0099]

[0100] In the formula, θ ns Represents the phase angle of node n at stage s, θ Represents the lower limit constraint of the phase angle, Represents the upper limit constraint of the phase angle.

[0101] Step S2: Obtain the repair time problems of all distribution lines in the power grid system, and based on the single-stage DC optimal power flow model, formulate the repair time problems as a Markov dynamic decision process.

[0102] As an optimal solution, step S2 includes steps S21 to S24, and the specific steps are as follows:

[0103] Step S21: Refer to Equation (10). According to the preset repair time algorithm, calculate the determined repair time of each distribution line in the power grid system based on the length L l of each distribution line, and refer to Equation (11). Add the determined repair time of each distribution line to the uncertainty δ l of each distribution line to obtain the repair time T l of each distribution line.

[0104]

[0105] In the formula, L l is the length of the distribution line l, N l represents the daily repair length of a maintenance personnel, and c represents the number of maintenance personnel.

[0106]

[0107] In the formula, δ l is the uncertainty of the repair time of the distribution line l. In practice, the repair time of damaged components may be affected by many uncertain factors, such as the fault location, component aging degree, logistics and technical delays, etc. Therefore, the present invention assumes that the uncertainty δ l of the repair time follows a lognormal distribution, specifically refer to Equation (12).

[0108]

[0109] In the formula, μ l and σ l are respectively the mean value and the standard deviation of the logarithmic variable in the lognormal distribution.

[0110] Step S22. Referring to Equations (13) - (26), express the repair time of all distribution lines as a stochastic optimal sequence problem to obtain the repair time problem of all distribution lines in the power grid system.

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125] In the formula, a ls = 1 indicates that the damaged distribution line l in the power grid system has been selected for repair in the s stage, and z ls represents the operating state of the distribution line in the s stage, and t sRepresents the repair time in stage s. The objective function (13) represents maximizing the expected resilience of the power grid system under uncertain repair times. The constraints (14) and (15) stipulate that only one distribution line can be repaired in each repair stage, and the repaired distribution line cannot be selected again. The constraint (16) sets the initial state of all distribution lines. The constraint (17) indicates that the distribution line is available after restoration. The duration of each stage is determined by the constraint (18). The constraints (19) to (25) are the same as the constraints (3) to (9). The constraint (26) defines the action value and state value of the distribution line.

[0126] Step S23: Obtain the optimal system resilience of the power grid system through a single-stage DC optimal power flow model, and establish a reward function based on the optimal system resilience of the power grid system and the repair time problem.

[0127] As an optimal solution, step S23 includes steps S231 to S234, and the specific steps are as follows:

[0128] Step S231: Establish the state space of the power grid system to describe the state of the power grid system. A distribution line has only two states, normal or faulty. In this embodiment, it is assumed that all distribution lines are damaged by natural disasters in the initial state. When a damaged distribution line is repaired in stage s, its state returns to normal (z ls = 1). Refer to Equation (27), the state space of the power grid system consists of the operating states of all distribution lines.

[0129]

[0130] Step S232: Establish the action space of the power grid system to describe the actions of the agent. When dealing with the repair problem, the action of the agent in each stage is to select a suitable line for repair. If distribution line l is selected for repair in stage s, then a ls = 1. Refer to Equation (28), define the action space of the agent to stage s. The number of action values in the action space A s increases with the increase of the stage. In reinforcement learning, the agent, that is, the decision maker of the power grid system, interacts with the environment (such as the power grid system) and learns experience (i.e., the reward after the action) to maximize future rewards.

[0131]

[0132] Step S233: Establish the reward function of the power grid system to describe the reward or benefit after the agent executes the action. Refer to Equation (29), calculate the average resilience change of the power grid system between stage (s - 1) and stage s. Used to measure the average change in system resilience during maintenance, which can reflect the efficiency of maintenance personnel. Due to the uncertainty of the environment during the repair process, in this embodiment, referring to Equation (30), the expected value of system resilience is used as the final reward for each state to establish a reward function

[0133]

[0134]

[0135] In the formula, represents the state space of the DCOPF model at stage s and the action space A s The optimal solution obtained under. δ is the uncertain quantity during the repair process and follows a lognormal distribution.

[0136] Step S24, referring to Equation (31), combine the five elements of the state space, action space, state transition probability function, reward function, and discount factor to establish a five-tuple to represent the Markov dynamic decision-making process of the repair time problem, so as to adapt to the dual-competitive deep Q-network for solution.

[0137]

[0138] In the formula, represents the state space; A represents the action space; P is the transition probability between different operating states of the power grid system, and P = 1 for all stages of this problem; R represents the benefit or reward function of the agent for each action; γ ∈ [0, 1] is the given future benefit discount factor.

[0139] Step S3, adopt the dual-competitive deep Q-network to solve the Markov dynamic decision-making process to obtain the loss function of the original deep Q-network in the dual-competitive deep Q-network, and then based on the loss function, train the original deep Q-network and use the trained original deep Q-network as the repair decision model; among them, the dual-competitive deep Q-network includes the original deep Q-network, the target deep Q-network, and the competitive network structure.

[0140] As a preferred solution, Step S3 includes Step S31 to Step S32, and the specific steps are as follows:

[0141] Step S31, use the dual-competitive deep Q-network to solve the Markov dynamic decision-making process to obtain the corresponding target value and estimated value, and calculate the loss function of the original deep Q-network according to the target value and estimated value.

[0142] As a preferred solution, Step S31 includes Step S311 to Step S312, and the specific steps are as follows:

[0143] Step S311, derive the optimal Q value of the deep Q-network (DQN) in the state space and the action space A s See Equation (32), which is composed of the sum of the immediate reward and the expected value of the next state reward. Then, see Equation (33), and evaluate the optimal repair strategy output by the original deep Q-network using the target deep Q-network in the double competitive deep Q-network according to the preset double Q-network algorithm to obtain the corresponding target value

[0144]

[0145] In the formula represents the immediate reward based on system resilience at stage s, and γ represents the given future reward discount factor.

[0146]

[0147] In the formula represents the target value corresponding to the optimal repair strategy of all distribution lines in the power grid system at stage (s + 1), represents the immediate reward based on system resilience at stage (s + 1), γ represents the future reward discount factor, represents the state space composed of the operating states of all distribution lines in the power grid system at stage (s + 1), A s+1 represents the action space composed of the repair actions of all distribution lines in the power grid system at stage (s + 1), θ p represents the network parameters of the original deep Q-network, θ w represents the network parameters of the target deep Q-network, Q p represents the output result of the original deep Q-network, Q w represents the output result of the target deep Q-network.

[0148] Step S312, see Equation (34). To avoid unnecessary estimation of each action value, solve the Markov decision-making process using the competitive network structure in the double competitive deep Q-network according to the preset competitive network algorithm to obtain the corresponding estimated value to improve the accuracy of the finally output Q value.

[0149]

[0150] In the formula represents the estimated value of the power grid system based on the state space and the action space A s ​​Denote the state space composed of the operating states of all distribution lines in the power grid system at stage s, A s Denote the action space composed of the repair actions of all distribution lines in the power grid system at stage s, V(S s | p ) represents the state space corresponding scalar, A s ′ represents the selected action space A s corresponding possible alternative action space, represents in the state space under the action space A s advantage function, represents in the state space under the possible alternative action space A s ′ advantage function.

[0151] Step S313, please refer to Equation (35), calculate the loss function L(θ of the original deep Q-network based on the target value and the estimated value p ).

[0152]

[0153] Step S32, adopt the gradient descent method to iteratively optimize the original deep Q-network in the double-competitive deep Q-network. Adjust the network parameters θ of the original deep Q-network during each iterative optimization p , and solve the Markov dynamic decision-making process according to the adjusted double-competitive deep Q-network to update the loss function L(θ p ). When the current loss function is minimized, the training of the original deep Q-network is completed, and the current original deep Q-network is used as the repair decision model.

[0154] In this embodiment, when a set of data is given with the network parameters θ of the original deep Q-network p as the optimization variable, with the minimization of the loss function L(θ of the original deep Q-network p ) as the objective function, adopt the gradient descent method to iteratively train the original deep Q-network in the double-competitive deep Q-network, and use the original deep Q-network that satisfies the objective function as the repair decision model.

[0155] Embodiment 2:

[0156] Please refer to Figure 2 , which is a schematic flowchart of a power grid system repair method based on deep reinforcement learning provided by an embodiment of the present invention. This method includes steps S4 to step S5, and the specific steps are as follows:

[0157] Step S4: When a disaster occurs, obtain the operating status of all distribution lines in the power grid system to form a status set. And input the status set into the repair decision-making model until all elements in the status set are equal to 1, which indicates that all distribution lines in the system have returned to normal. At this time, the repair decision-making model outputs the corresponding optimal repair strategy A. d Among them, the repair decision-making model is constructed by using the construction method of a repair decision-making model described in Embodiment 1.

[0158] Step S5: Repair all distribution lines in the power grid system in sequence according to the optimal repair strategy.

[0159] In this embodiment, for the IEEE 123-node system and the IEEE 300-node system, the repair strategy acquisition method based on Stochastic Optimization (SO) and the power grid system repair method based on Deep Reinforcement Learning (DRL) described in Embodiment 2 are respectively used to generate the optimal repair strategies corresponding to 1000 uncertain repair time scenarios, and calculate the average system resilience corresponding to the samples composed of 1000 uncertain repair time scenarios to demonstrate the performance of the dual competitive Q network. For the specific calculation performance comparison, please refer to Table 1.

[0160] Among them, since the repair actions before all demands are met still have the opportunity to supply energy to the load nodes, these repair actions are defined as critical repair actions, and the system resilience and the number of critical repair actions are used as the criteria for evaluating the repair strategy. The former reflects the performance of the system, and the latter describes the repair efficiency. As can be seen from Table 1, the power grid system repair method based on DRL can generate better repair strategies for the power grid system in a shorter time, while it is difficult for the repair strategy acquisition method based on SO to obtain the optimal result within the given time.

[0161] Table 1 Comparison table of calculation performance of two methods under different systems

[0162]

[0163] Embodiment 3:

[0164] Please refer to Figure 3 , which is a schematic structural diagram of a construction system of a repair decision-making model provided by an embodiment of the present invention. The system includes a model establishment module M1, a representation conversion module M2, and a model training module M3. The specific functions of each module are as follows:

[0165] The model establishment module M1 is used to establish a single - stage DC optimal power flow model with the maximization of system resilience as the objective function and according to the node constraints and distribution line constraints of the power grid system.

[0166] The expression conversion module M2 is used to obtain the repair time problems of all distribution lines in the power grid system and, based on the single - stage DC optimal power flow model, express the repair time problems as a Markov decision - making process.

[0167] The model training module M3 is used to solve the Markov decision - making process by using a double - competitive deep Q - network to obtain the loss function of the original deep Q - network in the double - competitive deep Q - network, and then train the original deep Q - network based on the loss function and use the trained original deep Q - network as the repair decision model. Among them, the double - competitive deep Q - network includes an original deep Q - network, a target deep Q - network, and a competitive network structure.

[0168] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above - described system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here.

[0169] Compared with the prior art, all embodiments of the present invention have the following beneficial effects:

[0170] The present invention provides a method for constructing a repair decision model, a system, and a power grid system repair method. With the maximization of system resilience as the objective function and using the node constraints and distribution line constraints of the power grid system as constraint functions, a single - stage DC optimal power flow model is established to determine the optimal energy distribution and repair time problems through the single - stage DC optimal power flow model, and then the repair time problems are expressed as a Markov decision - making process to adapt to the solution mode of the double - competitive deep Q - network. As a result, the trained repair decision model can not only fully consider system resilience to maximize system resilience, but also simplify the sampling steps through the competitive network structure in the double - competitive deep Q - network, thereby achieving efficient output of repair decisions in a large - scale state - action space. Additionally, by using the double - competitive deep Q - network to solve the Markov decision - making process, the original deep Q - network selects the best repair action and the target deep Q - network evaluates this best repair action, preventing the problem that the original deep Q - network overestimates the Q - value of the repair action and leads to too large a target value. At the same time, the original deep Q - network is trained using the loss function of the original deep Q - network, improving the accuracy of the repair decision model from multiple aspects and enabling the repair decision model to output a more efficient power grid system repair strategy.

[0171] Furthermore, constraint functions are respectively set for the power generation of each node, the amount of demand met by each node, the magnitude of the current passing through the distribution line, the flow balance of each node, the impedance of each distribution line, and the phase angle of each node, so as to ensure the node balance and load balance of the distribution lines in the power grid system while meeting the goal of maximizing system resilience.

[0172] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for constructing a repair decision model, characterized in that, Including: Taking the maximization of system resilience as the objective function, and establishing a single-stage DC optimal power flow model according to the node constraints and distribution line constraints of the power grid system; Obtaining the repair time problem of all distribution lines in the power grid system, and formulating the repair time problem as a Markov decision process based on the single-stage DC optimal power flow model; Among them, the obtaining of the repair time problem of all distribution lines in the power grid system is specifically: According to the preset repair time algorithm, calculate the determined repair time of each distribution line according to the length of each distribution line in the power grid system, and add the determined repair time of each distribution line to the uncertainty of each distribution line respectively to obtain the repair time of each distribution line; Express the repair times of all the distribution lines as a stochastic optimal sequence problem to obtain the repair time problem of all the distribution lines in the power grid system; Using a double-competitive deep Q-network to solve the Markov decision process to obtain the loss function of the original deep Q-network in the double-competitive deep Q-network, then training the original deep Q-network based on the loss function, and using the trained original deep Q-network as a repair decision model; Among them, the double-competitive deep Q-network includes the original deep Q-network, the target deep Q-network and the competitive network structure; Among them, the using a double-competitive deep Q-network to solve the Markov decision process to obtain the loss function of the original deep Q-network in the double-competitive deep Q-network, then training the original deep Q-network based on the loss function, and using the trained original deep Q-network as a repair decision model is specifically: Using the double-competitive deep Q-network to solve the Markov decision process to obtain the corresponding target value and estimated value, and calculating the loss function of the original deep Q-network according to the target value and the estimated value; Using the gradient descent method to iteratively optimize the original deep Q-network in the double-competitive deep Q-network, adjusting the network parameters of the original deep Q-network during each iterative optimization, and solving the Markov decision process according to the adjusted double-competitive deep Q-network to update the loss function until the current loss function is minimized, then the training of the original deep Q-network is completed, and the current original deep Q-network is used as the repair decision model.

2. The construction method of a repair decision model according to claim 1, characterized in that The using the double-competitive deep Q-network to solve the Markov decision process to obtain the corresponding target value and estimated value is specifically: According to the preset double Q-network algorithm, using the target deep Q-network in the double-competitive deep Q-network to evaluate the best repair strategy output by the original deep Q-network to obtain the corresponding target value; Among them, the double Q-network algorithm is specifically: In the formula, represents the objective value corresponding to the optimal repair strategy of all distribution lines in the power grid system at the (s + 1) stage, represents the immediate benefit based on system resilience at the (s + 1) stage, and γ represents the future benefit discount factor, represents the state space composed of the operating states of all distribution lines in the power grid system at the (s + 1) stage, A s+1 represents the action space composed of the repair actions of all distribution lines in the power grid system at the (s + 1) stage, θ p represents the network parameters of the original deep Q network, θ w represents the network parameters of the target deep Q network, Q p represents the output result of the original deep Q network, Q w represents the output result of the target deep Q network; According to the preset competitive network algorithm, the competitive network structure in the double competitive deep Q network is used to solve the Markov dynamic decision-making process to obtain the corresponding estimated value; wherein, the double Q network algorithm is specifically as follows: In the formula, represents the estimated value of the power grid system based on the state space and the action space A s . represents the state space composed of the operating states of all distribution lines of the power grid system at stage s, and A s represents the action space composed of the repair actions of all distribution lines of the power grid system at stage s. V(S s |θ p ) represents the scalar corresponding to the state space , and A s ' represents the possible alternative action space corresponding to the selected action space A s . represents the advantage function of the action space A under the state space s , and represents the advantage function of the possible alternative action space A ' under the state space s .

3. The construction method of a repair decision model according to claim 1, characterized in that It further includes: Establish the node constraints and distribution line constraints of the power grid system; Traverse each repair stage, take the maximization of system resilience as the objective function, and use the node constraints and distribution line constraints as the limiting conditions of decision variables to adjust each decision variable to obtain the system resilience corresponding to each repair stage; wherein, the objective function is specifically as follows: In the formula, the decision variable g ns represents the power generation of node n in the power grid system at stage s. The decision variable d ns represents the quantity of demand satisfied by node n in the power grid system at stage s. The decision variable f ls represents the magnitude of the current passing through distribution line l in the power grid system at stage s. The decision variable θ ns represents the phase angle of node n in the power grid system at stage s. R s represents the system resilience at stage s during the restoration phase.

4. The construction method of a repair decision-making model according to claim 3, characterized in that The node constraints include: node power generation constraint, node demand satisfaction constraint, node flow balance constraint, and node phase angle constraint; Among them, the node power generation constraint is specifically: where, g ns represents the power generation of node n in stage s, represents the target power generation of node n in stage s, represents the set of nodes in the power grid system, and S represents the set of stages in the power grid system; The node demand satisfaction constraint is specifically: where d ns represents the quantity of node n meeting the demand in stage s, and represents the target demand of node n in stage s. The node flow balance constraint is specifically: In the formula, n and n' respectively represent the two end nodes of the distribution line l; The node phase angle constraint is specifically: where θ ns represents the phase angle of node n at stage s, θ represents the lower bound constraint of the phase angle, and θ represents the upper bound constraint of the phase angle.

5. The construction method of a repair decision model according to claim 3, characterized in that The distribution line constraints include: distribution line energy limit constraint, and distribution line impedance and phase angle constraint; Among them, the distribution line energy limit constraint is specifically: Where, f ls represents the magnitude of the current passing through the distribution line l at stage s, and z ls represents the operating state of the distribution line l at stage s, and f l max represents the maximum value of the magnitude of the current passing through the distribution line l, represents the set of distribution lines in the power grid system, and S represents the set of stages in the power grid system; The distribution line impedance and phase angle constraint is specifically: where X R represents the impedance of the distribution line l, represents the phase angle of one end node n1 of the distribution line l, represents the phase angle of the other end node n2 of the distribution line l, and M represents a very large natural number used to ensure the effectiveness of the constraint.

6. The construction method of a repair decision model according to claim 1, characterized in that The problem of repair time is formulated as a Markov dynamic decision-making process based on the single-stage DC optimal power flow model, specifically: Through the single-stage DC optimal power flow model, the optimal system resilience of the power grid system is obtained, and a reward function is established according to the optimal system resilience of the power grid system and the repair time problem; According to the state space, action space, state transition probability function, the reward function, and the discount factor, a five-tuple is established to represent the Markov dynamic decision-making process of the repair time problem.

7. A power grid system repair method based on deep reinforcement learning, characterized in that, It includes: Obtain the operating states of all distribution lines in the power grid system and input all the operating states into the repair decision model so that the repair decision model outputs the corresponding optimal repair strategy; wherein, the repair decision model is constructed by using the construction method of a repair decision model according to any one of claims 1 to 6; Repair all the distribution lines in the power grid system in sequence according to the optimal repair strategy.

8. A construction system for a repair decision model, characterized in that, It includes: A model establishment module for taking the maximization of system resilience as the objective function and establishing a single-stage DC optimal power flow model according to the node constraints and distribution line constraints of the power grid system; A formulation conversion module for obtaining the repair time problem of all distribution lines in the power grid system and formulating the repair time problem as a Markov dynamic decision-making process based on the single-stage DC optimal power flow model; Among them, obtaining the repair time problem of all distribution lines in the power grid system is specifically: According to the preset repair time algorithm, and based on the lengths of the distribution lines in the power grid system, the determined repair times of the distribution lines are calculated, and the determined repair times of the distribution lines are respectively added to the uncertainties of the distribution lines to obtain the repair times of the distribution lines; the repair times of all the distribution lines are expressed as a stochastic optimal sequence problem to obtain the repair time problem of all the distribution lines in the power grid system; A model training module, which is used to solve the Markov decision process by using a double-competitive deep Q network to obtain the loss function of the original deep Q network in the double-competitive deep Q network, and then based on the loss function, train the original deep Q network, and use the trained original deep Q network as a repair decision model; wherein, the double-competitive deep Q network includes the original deep Q network, a target deep Q network and a competitive network structure; Wherein, the step of using a double-competitive deep Q network to solve the Markov decision process to obtain the loss function of the original deep Q network in the double-competitive deep Q network, and then based on the loss function, training the original deep Q network and using the trained original deep Q network as a repair decision model is specifically as follows: Use the double-competitive deep Q network to solve the Markov decision process to obtain the corresponding target value and estimated value, and calculate the loss function of the original deep Q network according to the target value and the estimated value; use the gradient descent method to perform iterative optimization on the original deep Q network in the double-competitive deep Q network, adjust the network parameters of the original deep Q network during each iterative optimization, and solve the Markov decision process according to the adjusted double-competitive deep Q network to update the loss function until the current loss function is minimized, then complete the training of the original deep Q network, and use the current original deep Q network as the repair decision model.

Citation Information

Patent Citations

  • Transformer substation fault handling method suitable for digital transfer scene

    CN114266487A

  • Method and computer program product for drug discovery using weighted grand canonical metropolis Monte Carlo sampling

    US20040267456A1