Power distribution system dynamic reconfiguration method based on deep reinforcement learning and computer device

By modeling the distribution system recovery process as a Markov decision process and integrating physical information neural networks, the problems of low efficiency and insufficient robustness of distribution system reconstruction under extreme weather conditions in traditional methods are solved, a dynamic reconstruction method based on deep reinforcement learning is implemented, and the system's resilience and recovery efficiency are improved.

CN119696072BActive Publication Date: 2025-10-21THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411867272.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-10-21
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Existing technologies have difficulty achieving fast and effective dynamic reconstruction of distribution systems under extreme weather conditions. Traditional optimization algorithms are computationally complex and rely on complete data, and artificial intelligence algorithms lack physical information models, resulting in insufficient strategy learning and an inability to adapt to actual resilience improvement issues.

Method used

A method based on deep reinforcement learning is used to model the distribution system recovery process as a Markov decision process. Cold load impact constraints, power flow constraints, radial topology constraints, and frequency response constraints are added. The physical information neural network is integrated for parameter update optimization, and training and dynamic reconstruction are performed through a deep reinforcement learning model.

Benefits of technology

It enables the real-time decision-making capability of the distribution system under extreme weather conditions, improves the system's resilience and recovery efficiency, enhances the robustness and reliability of the recovery strategy, and adapts to actual recovery scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119696072B_ABST
    Figure CN119696072B_ABST
Patent Text Reader

Abstract

The application relates to the field of power distribution system recovery, in particular to a power distribution system dynamic reconstruction method based on deep reinforcement learning and a computer device, which realizes that a power distribution system operator can make real-time decisions according to the current state of the power distribution system, and greatly improves the flexibility of the power distribution system. The method comprises the following steps: modeling the entire recovery process of the power distribution system as a Markov decision process, the decision process is solved by using a deep reinforcement learning model; adding cold load impact constraints, power flow constraints, radial topology constraints and frequency response constraints in the entire recovery process of the power distribution system; optimizing the entire recovery process of the power distribution system; adopting a deep reinforcement learning model fused with a physical information neural network to optimize parameter updating; training the deep reinforcement learning model; and performing dynamic reconstruction and sequential recovery of the power distribution system according to the trained deep reinforcement learning model. The application is suitable for power distribution system recovery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power distribution system restoration, and in particular to a power distribution system dynamic reconstruction method and computer device based on deep reinforcement learning. Background Art

[0002] Today, distribution systems are becoming more flexible, efficient, and reliable, thanks to the integration of intelligent devices (such as distributed generators, smart meters, and remote-controlled switchgear). However, frequent extreme weather events pose a significant threat to the safe and reliable operation of modern distribution systems, such as those in green building clusters. These systems face the risk of large-scale power outages, which can cause severe infrastructure damage and lead to enormous economic losses. Over the past decade, due to the frequent occurrence of extreme weather events caused by global climate change and the widespread deployment of renewable energy generation, which is highly dependent on weather, enhancing the resilience of power systems to extreme weather conditions has become a major issue emphasized by energy systems worldwide. Furthermore, with the integration of massive amounts of rapidly changing renewable energy generation, power electronics, and information devices, the physical and information characteristics of next-generation power systems have become increasingly complex. Furthermore, given the diverse nature of extreme weather conditions, the traditional power system approach of tailoring resilience enhancement strategies to specific extreme weather scenarios based on the characteristics of traditional components and fully known information fails to reflect the inherent response characteristics of the system structure. Therefore, it is urgent to develop new, systematic theories and methods for assessing and enhancing the resilience of power system structures, grounded in the essential characteristics of next-generation power systems. Smart facilities can ensure power supply to critical loads when the distribution system is forced to disconnect from the main grid.

[0003] Most current research on distribution system network reconfiguration focuses on considering various uncertainties, and most of these solutions rely on traditional optimization algorithms. However, these algorithms require complete system data and parameters before calculation. On the one hand, collecting complete system data and obtaining accurate network parameters after a disaster occurs is impractical. On the other hand, traditional optimization algorithms are extremely time-consuming, especially when considering various system uncertainties, where the computational complexity often increases exponentially.

[0004] Given the difficulty AI-based algorithms face in handling complex physical constraints, existing research typically terminates training if the agent violates any operational constraints. This results in insufficient or even incomplete learning of the policy. Furthermore, limited research has focused on improving AI algorithms to address real-world resilience optimization problems. Furthermore, models lacking physical information make it difficult to apply the results to real-world power systems. Therefore, incorporating physical information into AI algorithms is crucial for applying AI to distribution system resilience. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a dynamic reconstruction method and computer device for a distribution system based on deep reinforcement learning, so that the distribution system operator can make real-time decisions based on the current state of the distribution system, greatly improving the flexibility of the distribution system.

[0006] The present invention adopts the following technical solutions to achieve the above-mentioned objectives. In a first aspect, the present invention provides a method for dynamic reconstruction of a power distribution system based on deep reinforcement learning, comprising:

[0007] S1. Model the entire distribution system recovery process as a Markov decision process, which consists of a state set, an action set, a transition probability, a reward function, and a discount factor. The decision maker selects an action, and the environment transitions to the next state based on the action and receives a corresponding reward based on the transition result. By solving the Markov decision process, the optimal strategy for distribution system recovery under the current state is found.

[0008] S2. Add cooling load impact constraints, power flow constraints, radial topology constraints, and frequency response constraints to the entire restoration process of the distribution system;

[0009] S3. Optimize the entire recovery process of the power distribution system;

[0010] S4, using a deep reinforcement learning model that integrates physical information neural networks to perform parameter update optimization;

[0011] S5. Train the deep reinforcement learning model;

[0012] S6. Dynamically reconstruct and sequentially restore the power distribution system based on the trained deep reinforcement learning model.

[0013] Furthermore, in step S1, the state set of the Markov decision process includes the following elements:

[0014] System topology status at time t, line switch status at time t, load recovery status at time t, active / reactive power of restored loads at time t, and voltage at each system node at time t;

[0015] The action set of a Markov decision process consists of the following elements:

[0016] Circuit breaker behavior, whether it is closed at the current moment;

[0017] Load recovery decision behavior, whether to restore at the current moment;

[0018] The reward function considers multiple aspects of the given behavior. When the given behavior violates the physical operation constraints, a negative reward will be obtained. At the same time, a positive reward will be obtained based on the importance of restoring the load at the current moment. The reward function considers the following elements:

[0019] Priority of restored loads, whether diesel generators are overloaded, whether radial constraints are violated, and whether frequency response constraints are violated;

[0020] The Markov decision process is solved using a deep reinforcement learning algorithm. This state transition probability (matrix) can be implemented using a neural network. The neural network learns the state transition process by updating parameters. The deep learning algorithm uses a policy neural network to map states to actions. The mapping process is as follows:

[0021] The agent observes the state of the environment and models it as a vector;

[0022] Convert the vector into a tensor, input the tensor into the neural network, encode the information, and output the action vector;

[0023] The agent executes the action vector.

[0024] Furthermore, in step S2, the cooling load impact constraint is as follows:

[0025]

[0026] in and They are the maximum deviation coefficient and normal deviation coefficient of the constant temperature load under the influence of the cold load shock, N represents the total number of recovery operation periods, D represents the total number of loads, and a d,t is the decay rate of the cooling load shock phenomenon, D d,t The time period from the start of load charging to the acquisition of diversity; C d,t is the time period from when the load starts to acquire diversity to time t, expressed as:

[0027] C d,t =(t-1)Δt-D d,t ;

[0028] The active power demand and reactive power demand of the load when it is restored are:

[0029]

[0030] in, is a binary variable, indicating the charging state of the load, ΔP d (k) and ΔQ d (k) According to the formula L d(t) is calculated to represent the difference between active power and reactive power at time t and time t-1, respectively. and They represent the rated active power and reactive power of load d respectively.

[0031] The power flow constraints are as follows:

[0032]

[0033] in, They are active power flow, reactive power flow, active generator output, reactive generator output, remote control switch control state, and node i voltage state; and They represent the maximum active and reactive power flows allowed to pass through the line, δ ij represents the connectivity status of the line, ε represents the set of connected edges, r ij with x ij They represent the resistance and reactance of the line respectively.

[0034] The radial topology constraints are as follows:

[0035] σ∈Ω

[0036]

[0037] Among them, σ ij,t Describes the set of all possible tree topologies of the system, δ swi,t Indicates the open and closed state of the switch. Indicates the number of nodes in the system, Indicates the number of root nodes in the tree structure, represents the flow rate of the virtual flow, σ ij,t Indicates whether the line has a switch that can be opened and closed, M is a real number.

[0038] The frequency response constraints are as follows:

[0039]

[0040] The frequency response constraint means that the load restored by the distribution system at each step is limited by the frequency response constraints of all units, γ FRR Represents the system frequency response control coefficient.

[0041] Furthermore, in step S3, the objective function of the optimization is to maximize the load recovery amount. The objective function of the optimization problem is:

[0042] where ω d Indicates the weight of the load, o d,t Indicates the recovery state of the load, is the load demand of load d at time t.

[0043] Furthermore, step S4 specifically includes:

[0044] Load information and grid topology information are input in the form of a matrix, and a convolutional neural network is used to encode the system state information. After the state information is input, it is processed by the policy neural network and the initial operation action of the intelligent agent is output. The action includes the load position to be restored in the current state and the line switch action;

[0045] At the same time, based on the current state information, a feasible solution to the restoration strategy under the current state is obtained by solving a single-period optimization problem. The objective function of the optimization problem is to maximize the number of loads restored. After linearizing the nonlinear constraints, a mixed integer programming solution is solved using a commercial solver.

[0046] The optimization problem is integrated into the parameter update of the neural network as knowledge, the behavior vector of the feasible solution obtained by the optimization problem is calculated, and the cosine similarity between it and the initial action vector given by the intelligent agent is calculated. The cosine similarity is then combined with the original policy gradient function for optimization.

[0047] Furthermore, step S5 specifically includes:

[0048] The policy network of the deep reinforcement learning model makes decisions based on the observed environmental state, and the value network of the deep reinforcement learning model gives the evaluation value of the decision. At each time step, the agent outputs an action from the policy network based on the current state and interacts with the environment to obtain rewards and the next state. Then, the experience replay mechanism is used to store the state, action, reward and next state in a buffer. Then, these experiences are randomly sampled to update the policy network and value network. The value network is optimized by minimizing the mean square error, and the deterministic policy gradient is used to update the policy network. By continuously iterating this process, the agent learns the optimal strategy.

[0049] Furthermore, in step S6, performing sequential recovery specifically includes:

[0050] When the distribution system detects a fault, it collects data including the load recovery status, the power demand of the restored load, the line connectivity status, the system voltage parameters, and the system flow parameters. The collected information is then encoded and input into the deep reinforcement learning model. The policy network of the deep reinforcement learning model will derive the optimal recovery decision for the system at the current moment based on the obtained system status information. This recovery decision will be transmitted to each sensor through communication equipment and executed by the distribution system operator. Afterwards, the system status information will be updated. Through this iterative behavior, the distribution system will execute sequential recovery decisions step by step.

[0051] In a second aspect, the present invention provides a computer device comprising a memory storing program instructions, wherein when the program instructions are executed, the above-mentioned method for dynamic reconstruction of a power distribution system based on deep reinforcement learning is executed.

[0052] The beneficial effects of the present invention are:

[0053] The present invention is based on a deep reinforcement learning algorithm model for the sequential restoration of distribution systems and the formation of dynamic microgrids, and takes into account the uncertainty of distributed low-voltage user conditions by utilizing topological flexibility. The load uncertainty inherited from the distributed low-voltage user conditions during the restoration process is resolved through a Markov decision process, and a deep deterministic policy gradient algorithm framework integrating physical information neural networks is developed, through which the agent can better learn the physical information of operational constraints. The method of the present invention can be trained through offline and online learning. In offline training, OpenDSS is used to simulate the system simulation environment. In particular, online learning can improve the adaptability of the strategy to actual restoration scenarios. In addition, the restoration strategy also fully considers the cold load shock phenomenon of building constant temperature loads under the uncertainty of environmental factors.

[0054] This invention integrates physical information neural networks into reinforcement learning algorithms for the first time, incorporating physical information into the learning strategy of the intelligent agent, further improving the robustness and reliability of the recovery strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 This is a flow chart of a method for dynamic reconfiguration of a power distribution system based on deep reinforcement learning provided by an embodiment of the present invention;

[0056] Figure 2 2 is a schematic diagram of the topology of a power distribution system provided by an embodiment of the present invention. The left side of the dotted line is a schematic diagram of a sequential restoration process without considering dynamic reconstruction, and the right side of the dotted line is a schematic diagram of a sequential restoration process with dynamic reconstruction.

[0057] Figure 3 Schematic diagram of a sequential recovery framework provided by an embodiment of the present invention;

[0058] Figure 4 This is a schematic diagram of cooling load shock provided by an embodiment of the present invention;

[0059] Figure 5 This is a schematic diagram of actual load power fluctuations provided by an embodiment of the present invention;

[0060] Figure 6 This is a schematic diagram of a physical information neural network framework provided by an embodiment of the present invention;

[0061] Figure 7is a schematic diagram of a microgrid after dynamic reconstruction provided by an embodiment of the present invention;

[0062] Figure 8 This is a flowchart of a sequential recovery process provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0063] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0064] Common approaches to sequential restoration of power systems can be divided into two categories: model-based methods and artificial intelligence-based methods. Artificial intelligence-based methods enable distribution system operators to make real-time decisions based on the current state of the system, thus providing greater potential for improving system resilience. Although the above-mentioned model-free methods have shown great potential in sequential restoration decision-making, there are still some challenges that need to be addressed. This work focuses on deep reinforcement learning algorithms. On the one hand, the training environment should contain more real-world scenario elements (e.g., uncertainties in distributed low-voltage users) to make the strategy practical. On the other hand, more physical information, such as physical operating constraints, should be integrated into such model-free methods. Given that most deep reinforcement learning methods use neural networks as state transition functions, integrating physical information can significantly improve interpretability and enable the agent to learn strategies more effectively.

[0065] Based on this, the present invention utilizes a deep reinforcement learning algorithm model for sequential restoration of distribution systems and the formation of dynamic microgrids, leveraging topological flexibility to account for the uncertainty of distributed low-voltage user conditions. A Markov decision process is used to address the load uncertainty inherited from distributed low-voltage user conditions during the restoration process. Furthermore, a deep reinforcement learning model is developed that incorporates a physical information neural network. Through this deep reinforcement learning algorithm model, the agent can better learn the physical information of the operational constraints.

[0066] When the main grid faces a security threat from the distribution system, the distribution system will be disconnected from the main grid and continue to supply power using its own backup energy. Therefore, efficient microgrid reconstruction technology and sequential restoration are of great significance to improving the resilience of the distribution system. Considering that when the load experiences a long power outage, a cold load shock phenomenon will occur, that is, some constant temperature loads, such as air conditioning systems or refrigeration equipment, when these devices are started, the sudden increase in load demand causes the instantaneous load of the equipment to increase. This phenomenon may cause the transformer to overload, thereby hindering system recovery. At this time, making full use of the topological flexibility of the system can improve the recovery efficiency, such as Figure 2 shown.

[0067] The left side shows a sequential recovery process without considering dynamic reconfiguration, that is, without leveraging topological flexibility. It can be seen that the generators are not fully utilized. The right side shows a sequential load recovery process that fully utilizes the topological flexibility of the network. It can be seen that the recovery strategy on the right has a significant advantage in terms of recovery time.

[0068] In addition, restoration will face multiple uncertainties, such as environmental factors, load demand, etc. And in extreme scenarios, it is very difficult to obtain grid data, and it is difficult to use traditional optimization algorithms to solve. Therefore, the present invention uses artificial intelligence algorithms to solve the sequential restoration problem and considers the uncertainty of load demand. The proposed algorithm can be combined with the fault management system in the distribution system. The proposed framework is as follows Figure 3 When a system fault is detected, system data is uploaded to the control center. The control center uses algorithms to determine the optimal system scheduling decision for that state and remotely control the corresponding equipment.

[0069] Specifically, the present invention provides a method for dynamic reconstruction of a power distribution system based on deep reinforcement learning, such as Figure 1 As shown, specifically including:

[0070] S1. The entire restoration process of the distribution system is modeled as a Markov decision process, which is solved using a deep reinforcement learning model.

[0071] S2. Add cooling load impact constraints, power flow constraints, radial topology constraints, and frequency response constraints to the entire restoration process of the distribution system;

[0072] S3. Optimize the entire recovery process of the power distribution system;

[0073] S4, using a deep reinforcement learning model that integrates physical information neural networks to perform parameter update optimization;

[0074] S5. Train the deep reinforcement learning model;

[0075] S6. Dynamically reconstruct and sequentially restore the power distribution system based on the trained deep reinforcement learning model.

[0076] Specifically, the Markov decision process consists of a state set, an action set, a transition probability, a reward function, and a discount factor. The decision maker selects an action, and the environment transitions to the next state based on that action. The decision maker receives a corresponding reward based on the transition result. By solving the Markov decision process, the optimal strategy for restoring the distribution system under the current state is found.

[0077] The state set of a Markov decision process contains the following elements:

[0078] System topology status at time t, line switch status at time t, load recovery status at time t, active / reactive power of restored loads at time t, and voltage at each system node at time t;

[0079] The action set of a Markov decision process consists of the following elements:

[0080] Circuit breaker behavior, whether it is closed at the current moment;

[0081] Load recovery decision behavior, whether to restore at the current moment;

[0082] The reward function considers multiple aspects of the given behavior. When the given behavior violates the physical operation constraints, a negative reward will be obtained. At the same time, a positive reward will be obtained based on the importance of restoring the load at the current moment. The reward function considers the following elements:

[0083] Priority of restored loads, whether diesel generators are overloaded, whether radial constraints are violated, and whether frequency response constraints are violated;

[0084] Due to the multiple uncertainties inherent in the sequential restoration of distribution systems, state transition probabilities are difficult to determine ex ante. A Markov decision process is solved using a deep reinforcement learning algorithm. This state transition probability matrix can be implemented using a neural network. The neural network learns the state transition process through parameter updates. The deep learning algorithm uses a policy neural network to map states to actions. The mapping process is as follows:

[0085] The agent observes the state of the environment and models it as a vector;

[0086] Convert the vector into a tensor, input the tensor into the neural network, encode the information, and output the action vector;

[0087] The agent executes the action vector.

[0088] In addition, a discount factor is considered during the training process of the deep reinforcement learning algorithm. The discount factor is a hyperparameter that can be fine-tuned based on the performance of the model.

[0089] Specifically, the present invention considers the cold load shock phenomenon that occurs when the load is restored. That is, when the load faces a long power outage, it will lose its diversity. At this time, if the load is powered on, its power demand will far exceed the power required during the power outage. This phenomenon will affect the power demand when the load is restored, that is, the load power demand will dynamically change according to the characteristics of this phenomenon. Cold load shock is as follows: Figure 4 As shown, the present invention models the cooling load impact process as the following mathematical expression, and considers these constraints when updating the state to update the power demand of the load.

[0090] The cooling load impulse constraints are as follows:

[0091]

[0092] in, and They are the maximum deviation coefficient and normal deviation coefficient of constant temperature load under the influence of cold load shock, N represents the total number of recovery operation periods, D represents the total number of loads, and a d,t is the decay rate of the cooling load shock phenomenon, D d,t The time period from the start of load charging to the acquisition of diversity; C d,t is the time period from when the load starts to acquire diversity to time t, expressed as:

[0093] C d,t =(t-1)Δt-D d,t ;

[0094] The active power demand and reactive power demand of the load when it is restored are:

[0095]

[0096] in, is a binary variable, indicating the charging state of the load; ΔP d (k) and ΔQ d (k) According to the formula L d (t) is calculated to represent the difference between active power and reactive power at time t and time t-1, respectively. and They represent the rated active power and reactive power of load d respectively.

[0097] Specifically, the sequential restoration problem needs to consider the following constraints: power flow constraints, cooling load impact constraints, radial topology constraints, and frequency response constraints. The optimization objective function is to maximize the load restoration amount. The objective function of the optimization problem is:

[0098] where ω d Indicates the weight of the load, o d,t Indicates the recovery status of the load, is the load demand of load d at time t.

[0099] The power flow constraints are as follows:

[0100]

[0101] in, They are active power flow, reactive power flow, active generator output, reactive generator output, remote control switch control state, and node i voltage state. and They represent the maximum active and reactive power flows allowed to pass through the line, δ ij represents the connectivity status of the line, ε represents the set of connected edges, r ij with x ij The present invention takes into account the uncertainty of the load during the recovery process. Therefore, the real load demand curve is Figure 4 The curve shown is the baseline, which fluctuates randomly up and down by 10%. Figure 5 shown.

[0102] The radial topology constraints are as follows:

[0103] σ∈Ω

[0104]

[0105] Among them, σ ij,t Describes the set of all possible tree topologies of the system, δ swi,t Indicates the open and closed state of the switch. Indicates the number of nodes in the system, Indicates the number of root nodes in the tree structure, represents the flow rate of the virtual flow, σ ij,t Indicates whether the line has a switch that can be opened and closed, M is a sufficiently large real number;

[0106] The frequency response constraints are as follows:

[0107]

[0108] Frequency response constraints indicate that the load restored by the distribution system at each step is limited by the frequency response constraints of all units. FRR Represents the system frequency response control coefficient.

[0109] Specifically, the present invention adopts a physical information neural network to improve the learning effect of the intelligent agent on physical constraints. The present invention integrates the deep reinforcement learning model parameter update framework of the physical information neural network as follows: Figure 6 As shown, the load information includes: load recovery status and active / reactive power demand of the restored load, and the grid topology includes the closing status of the line switch and the number of microgrids in the system.

[0110] Load information and grid topology information are input in the form of matrices, undergoing further tensor transformations. A convolutional neural network is then used to encode system state information. After this state information is input, it is processed by the policy neural network of the deep reinforcement learning model, which outputs the agent's initial action. This action includes the load position to be restored and the line switching action to be performed in the current state.

[0111] At the same time, based on the current state information, a feasible solution to the restoration strategy under the current state is obtained by solving a single-period optimization problem. The objective function of the optimization problem is to maximize the number of loads restored. After linearizing the nonlinear constraints, a mixed integer programming solution is solved using a commercial solver.

[0112] The optimization problem is integrated into the parameter update of the neural network as knowledge, the behavior vector of the feasible solution obtained by the optimization problem is calculated, and the cosine similarity between it and the initial action vector given by the intelligent agent is calculated. The cosine similarity is then combined with the original policy gradient function for optimization.

[0113] Specifically, this paper employs a deep reinforcement learning algorithm model to solve the proposed sequential recovery problem. Power flow constraints are systematically simulated using OpenDSS software to simulate realistic system states. The deep reinforcement learning algorithm code is written in Python. Through an interface, the deep reinforcement learning algorithm interacts with the simulation environment built by OpenDSS to enable agent strategy learning.

[0114] During the training phase of a deep reinforcement learning algorithm model, the model's policy network first makes decisions based on the observed environment state, while the value network evaluates these decisions. At each time step, the agent outputs an action from the policy network based on its current state and interacts with the environment to obtain a reward and the next state. Then, using an experience replay mechanism, the state, action, reward, and next state are stored in a buffer. These experiences are then randomly sampled to update the policy and value networks. The value network is optimized by minimizing the mean squared error, and the policy network is updated using deterministic policy gradients. Through repeated iterations of this process, the agent learns the optimal policy.

[0115] During the deep reinforcement learning algorithm model testing phase, the agent will directly use the policy network to take actions based on the current system state without the need for gradient updates.

[0116] This paper uses the DQN (deep Q-network) algorithm results as a reference to compare the network reconstruction results. The network topology obtained by the DQN algorithm is as follows: Figure 7 shown.

[0117] The present invention provides sequential recovery results, as shown in Table 1.

[0118] Table 1 Comparison of sequential recovery results

[0119] Model number Restore system power Recovery load quantity Load recovery rate 1 6124.02KW 128 100% 2 5702.91KW 117 91.41% 3 5908.77KW 125 97.66% 4 5409.35KW 127 99.22%

[0120] In Table 1, model number 1 uses OU-Noise for exploration action, and the model is combined with physical information neural network, model number 2 uses Gaussian noise for exploration action, and the model is combined with physical information neural network, model number 3 uses OU-Noise for exploration action, and the model is not combined with physical information neural network, and model number 4 uses Deep Q-learning algorithm.

[0121] It can be seen from the data in Table 1 that the recovery strategy provided by the deep reinforcement learning algorithm model of the present invention has a better load recovery rate than the compared methods.

[0122] The process of sequential recovery of the present invention is as follows Figure 8 As shown, first the distribution system fault detection is performed, and after the fault is detected, the fault is isolated. Then the distribution system collects data, and then sends the system data to the central control end (control center). The strategy is derived based on the strategy network, and then the strategy is executed to update the system status and finally complete the recovery.

[0123] Specifically, when the distribution system detects a fault, it collects data. The collected data includes the load recovery status, the power demand of the restored load, the line connectivity status, the system voltage parameters, and the system flow parameters. The collected information is then encoded and input into the deep reinforcement learning model. The strategy network of the deep reinforcement learning model will derive the optimal recovery decision for the system at the current moment based on the obtained system status information. The recovery decision will be transmitted to each sensor through communication equipment and executed by the distribution system operator. Afterwards, the system status information will be updated. Through this iterative behavior, the distribution system will execute sequential recovery decisions step by step.

[0124] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.

Claims

1. A dynamic reconstruction method for power distribution system based on deep reinforcement learning, characterized by: include: S1. The entire restoration process of the distribution system is modeled as a Markov decision process, which is solved using a deep reinforcement learning model. S2. Add cooling load impact constraints, power flow constraints, radial topology constraints, and frequency response constraints to the entire restoration process of the distribution system; The cooling load impulse constraints are as follows: , ; in and They are the maximum deviation coefficient and normal deviation coefficient of constant temperature load under the influence of cooling load shock, Indicates the total number of recovery periods. Indicates the total load, is the decay rate of the cooling load shock phenomenon, The time period from the start of charging the load to the achievement of diversity; Start getting diversity for your load The time period of a moment is expressed as: ; The active power demand and reactive power demand of the load when it is restored are: , ; , ; in, is a binary variable, indicating the charging state of the load, and According to the formula Calculate and represent the difference between active power and reactive power at time t and time t-1, respectively. and Respectively represent the rated active power and reactive power of load d; The radial topology constraints are as follows: in, Describes the set of all possible tree topologies of the distribution system, Indicates the open and closed state of the switch. represents the number of nodes in the power distribution system, Indicates the number of root nodes in the tree structure, represents the flow of virtual flow from node i to node j, represents the flow of the virtual flow from node j to node i, M is a real number, Represents the set of connected edges; The frequency response constraints are as follows: Frequency response constraints indicate that the load restored by the distribution system at each step is limited by the frequency response constraints of all units. represents the system frequency response control coefficient, Indicates the output of active generator g at time t; S3. Optimize the entire recovery process of the power distribution system; S4, using a deep reinforcement learning model that integrates physical information neural networks to perform parameter update optimization; S5. Train the deep reinforcement learning model; S6. Dynamically reconstruct and sequentially restore the power distribution system based on the trained deep reinforcement learning model.

2. The method for dynamic reconstruction of a power distribution system based on deep reinforcement learning according to claim 1, characterized in that: In step S1, the Markov decision process consists of a state set, an action set, a transition probability, a reward function, and a discount factor. The decision maker selects an action, and the environment transitions to the next state based on the action and receives a corresponding reward based on the transition result. By solving the Markov decision process, the optimal strategy for restoring the distribution system under the current state is found. The state set of a Markov decision process contains the following elements: System topology status at time t, line switch status at time t, load recovery status at time t, active / reactive power of restored loads at time t, and voltage at each system node at time t; The action set of a Markov decision process consists of the following elements: Circuit breaker behavior, whether it is closed at the current moment; Load recovery decision behavior, whether to restore at the current moment; The reward function considers multiple angles based on the given behavior. If the given behavior violates the physical operation constraints, a negative reward will be obtained. At the same time, a corresponding positive reward will be given based on the importance of restoring the load at the current moment. The elements considered in the reward function include: Priority of restored loads, whether diesel generators are overloaded, whether radial constraints are violated, and whether frequency response constraints are violated; The Markov decision process is solved using a deep reinforcement learning algorithm. The state transition probability can be implemented using a neural network. The neural network learns the state transition process by updating parameters. The deep learning algorithm uses a policy neural network to map states to actions. The mapping process is as follows: The agent observes the state of the environment and models it as a vector; Convert the vector into a tensor, input the tensor into the neural network, encode the information, and output the action vector; The agent executes the action vector.

3. The method for dynamic reconstruction of a power distribution system based on deep reinforcement learning according to claim 1, characterized in that: In step S2, the power flow constraints are as follows: in, They are active power flow, reactive power flow, the output of active generator g at time t, the output of reactive generator g at time t, the control state of remote control switch, the voltage state of node i, and They represent the maximum active and reactive power flows allowed to pass through the line, Indicates the connection status of the line. represents the set of connected edges, and are the resistance and reactance of the line respectively, Indicates the voltage state of node j.

4. The method for dynamic reconstruction of a power distribution system based on deep reinforcement learning according to claim 1, characterized in that: In step S3, the objective function of optimization is to maximize the load recovery amount. The objective function of the optimization problem is: ,in represents the weight of the load, Indicates the recovery status of the load.

5. The method for dynamic reconstruction of a power distribution system based on deep reinforcement learning according to claim 1, characterized in that: Step S4 specifically includes: Load information and grid topology information are input in the form of a matrix, and a convolutional neural network is used to encode the system state information. After the state information is input, it is processed by the policy neural network and the initial operation action of the intelligent agent is output. The action includes the load position to be restored in the current state and the line switch action; At the same time, based on the current state information, a feasible solution to the restoration strategy under the current state is obtained by solving a single-period optimization problem. The objective function of the optimization problem is to maximize the number of loads restored. After linearizing the nonlinear constraints, a mixed integer programming solution is solved using a commercial solver. The optimization problem is integrated into the parameter update of the neural network as knowledge, the behavior vector of the feasible solution obtained by the optimization problem is calculated, and the cosine similarity between it and the initial action vector given by the intelligent agent is calculated. The cosine similarity is then combined with the original policy gradient function for optimization.

6. The method for dynamic reconstruction of a power distribution system based on deep reinforcement learning according to claim 1, characterized in that: Step S5 specifically includes: The policy network of the deep reinforcement learning model makes decisions based on the observed environmental state, and the value network of the deep reinforcement learning model gives the evaluation value of the decision. At each time step, the agent outputs an action from the policy network based on the current state and interacts with the environment to obtain rewards and the next state. Then, the experience replay mechanism is used to store the state, action, reward and next state in a buffer. Then, these experiences are randomly sampled to update the policy network and value network. The value network is optimized by minimizing the mean square error, and the deterministic policy gradient is used to update the policy network. By continuously iterating this process, the agent learns the optimal strategy.

7. The method for dynamic reconstruction of a power distribution system based on deep reinforcement learning according to claim 1, characterized in that: In step S6, performing sequential recovery specifically includes: When a distribution system detects a fault, it collects data. The collected data includes the load recovery status, the power demand of the restored load, the line connectivity status, the system voltage parameters, and the system power flow parameters. The collected information is then encoded and input into the deep reinforcement learning model. The policy network of the deep reinforcement learning model will use the obtained system status information to derive the optimal recovery decision for the system at the current moment. This recovery decision will be transmitted to each sensor through communication equipment and executed by the distribution system operator. After that, the system status information will be updated. Through iterative update behavior, the distribution system will execute sequential recovery decisions step by step.

8. A computer device comprising a memory storing program instructions, characterized in that: When the program instructions are executed, the method for dynamic reconstruction of a power distribution system based on deep reinforcement learning as described in any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Active power distribution network fault recovery method based on reinforcement learning method

    CN113872198A

  • Online dynamic decision-making method and system for unit restoration

    EP3780307A1