Power distribution network dynamic network reconfiguration method based on constraint reinforcement safety reinforcement learning

By constructing a constrained Markov decision process and a spatiotemporal awareness neural network, the energy loss and voltage/current problems caused by uncertainties in high-penetration distribution networks of photovoltaic and wind power generation were solved, achieving efficient dynamic network reconfiguration under safety constraints and improving the operation performance of the distribution network.

CN119578564BActive Publication Date: 2026-04-17STATE GRID JIANGSU ECONOMIC RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID JIANGSU ECONOMIC RES INST
Filing Date
2024-11-13
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively address issues such as increased energy loss and voltage/current exceeding limits caused by uncertainties in distribution networks with high penetration of photovoltaic and wind power generation. Furthermore, existing deep reinforcement learning methods suffer from high computational burden and unstable performance when satisfying security constraints.

Method used

A constraint-enhanced security reinforcement learning approach is adopted to model the dynamic network reconfiguration task of the distribution network as a constrained Markov decision process. Combining the interior point policy optimization algorithm and the spatiotemporal awareness neural network, a security layer and a mask of feasible actions are designed to satisfy hard constraints. Graph convolutional networks and gated recurrent unit networks are used to enhance the spatiotemporal awareness of the agent.

Benefits of technology

It achieves reduced energy loss, optimized voltage distribution, and improved new energy carrying capacity in high-penetration distribution networks of photovoltaic and wind power generation, and can perform real-time online calculations to meet safety constraints. Compared with existing methods, it is simpler to calculate and has better performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119578564B_ABST
    Figure CN119578564B_ABST
Patent Text Reader

Abstract

The application proposes a power distribution network dynamic network reconstruction method based on constraint reinforcement security reinforcement learning, describes the DDNR task of the power distribution network as a constraint Markov decision process (CMDP), and then respectively adopts an interior point policy optimization (IPO) algorithm and a method of adding a security layer based on branch exchange to process the soft constraint (node voltage and line current constraint) and the hard constraint (radial topology constraint) in the DDNR process, and meanwhile, the application proposes a space-time perception neural network model to increase the dynamic space-time change characteristics of the power distribution network power flow in the DDNR process. The method has an important role in reducing the energy loss of the power distribution network, optimizing the voltage distribution, and increasing the new energy carrying capacity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of intelligent reinforcement learning technology and novel power system technology, specifically relating to a dynamic network reconfiguration method for distribution networks based on constraint-enhanced security reinforcement learning. Background Technology

[0002] Distributed renewable energy sources (DREs) such as photovoltaics (PV) and wind power (WP) exhibit strong intermittency and volatility. Their high integration rate increases the uncertainty of power flow state transitions in distribution networks (PDNs) and the complexity of optimization problems. Network reconfiguration technology can achieve spatial energy transfer by changing the states of sectionalizing switches and tie switches in the PDN. With the increasing penetration of remote-controlled switches and DREs, dynamic distribution network reconfiguration (DDNR) plays an important role in improving the carrying capacity of DREs, reducing PDN energy losses, and improving voltage distribution.

[0003] DDNR can be described as a mixed-integer nonlinear programming problem, where integer variables represent the states of remotely controlled switches. Current research on the DDNR problem faces two main challenges. First, it requires constructing a stochastic probability distribution that accurately describes the uncertainties in the state transitions of DRE power and PDN power flow. Second, DDNR is a typical sequential decision problem, requiring distribution operators (DSOs) to solve for switch states across multiple decision steps while predicting future uncertainties, significantly increasing the computational burden.

[0004] Currently, researchers are focusing on three types of methods: mathematical programming methods, heuristic methods, and learning-based methods. To handle the uncertainty of distribution network reconfiguration (DRE), stochastic programming methods first construct the uncertainty scenario and then solve the DRE / DNR problem as a mixed-integer second-order cone programming model. Robust optimization methods achieve minimum network loss reconfiguration of DRE-intensive distribution network (PDN) by constructing an uncertainty set. Heuristic methods use a simple stepwise search process to analyze possible solutions and logically select high-quality solutions. Branch switching is a common heuristic method, which uses branch switching strategies to find the optimal switch pairs with the goal of minimizing the active power loss of loops. Based on this, research has proposed a heuristic method for switch opening and switching to solve multi-time-step stochastic distribution network reconfiguration. Meanwhile, metaheuristic intelligent optimization algorithms such as ant colony optimization, genetic algorithms, and particle swarm optimization are also used to solve the complex DRE / DNR problem.

[0005] However, as the uncertainty of PDN increases, the scenario size of stochastic programming methods grows exponentially, significantly increasing the computational burden. Furthermore, robust optimization methods typically produce conservative solutions, all of which degrade the performance of mathematical programming methods. While heuristic methods alleviate the computational burden to some extent, they cannot guarantee optimality and cannot handle the real-time uncertainty of power flow states during PDN operation.

[0006] Deep reinforcement learning (DRL) methods can capture the uncertainty of DRE from historical data and continuously learn the dynamic characteristics of high-dimensional nonlinear systems during interaction with the environment, demonstrating good performance in solving the sequential decision problem of DDNR. Some studies have used Deep Q-learning algorithms to solve the optimal network topology of high-penetration PDNs in new energy sources, or proposed novel algorithms such as Deep Q-Network and SAC to solve the optimal network topology. However, with the increase of DRE in PDNs, existing DRL-based schemes exhibit two new challenges. First, the DDNR problem has a large number of security constraints, including radial topology constraints (hard constraints) and node voltage and line current constraints (soft constraints). Existing DRL methods satisfy these constraints by simplifying the PDN structure and adding penalty terms to the reward function; however, this approach requires a cumbersome process of designing PDN reconstruction rules and penalty factors, and inappropriate design may reduce the performance of the DRL algorithm. Secondly, in order to capture the evolution characteristics of the PDN power flow state, the DRL agent needs to perceive the timing characteristics of source load power generation and consumption and the dynamic spatial characteristics of PND topology changes when performing DDNR, rather than simply inputting the state variables into the agent's neural network in the form of a one-dimensional vector. Summary of the Invention

[0007] This invention proposes a distribution network dynamic network reconfiguration (DDNR) method based on constraint-enhanced security reinforcement learning. The DDNR task of the distribution network is described as a constrained Markov decision process (CMDP). Then, the soft constraints (node ​​voltage and line current constraints) and hard constraints (radial topology constraints) in the DDNR process are handled by the interior point policy optimization (IPO) algorithm and the addition of a security layer, respectively. At the same time, this invention proposes a spatiotemporal awareness neural network model to enhance the reinforcement learning agent's awareness of the dynamic spatiotemporal changes in the distribution network power flow during the DDNR process.

[0008] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0009] Therefore, this invention provides a dynamic network reconfiguration method for distribution networks based on constraint-enhanced security reinforcement learning, which plays an important role in reducing distribution network (PDN) energy loss, optimizing voltage distribution, and increasing the carrying capacity of new energy sources. To solve the above technical problems, this invention provides the following technical solution: a dynamic network reconfiguration method for distribution networks based on constraint-enhanced security reinforcement learning, comprising:

[0010] Step S1: Treat the DDNR task of the high-penetration active distribution network of photovoltaic and wind power generation as a CMDP model, and construct a secure reinforcement learning solution framework based on the IPO algorithm. The logarithmic barrier function of IPO enables the agent to satisfy the soft constraints of the DDNR task when learning the policy.

[0011] Step S2: Design a security layer based on a feasible action mask in conjunction with the branch switching mechanism of the distribution network, and apply the security layer to the output layer of the IPO policy network so that the policy output by the agent always meets the hard constraints of the DDNR task.

[0012] Step S3: Construct a spatiotemporal awareness neural network model and apply the spatiotemporal awareness neural network to the input layer of the IPO policy network and the value network, so that the agent can better perceive the spatiotemporal change characteristics of the power flow in the distribution network during the DDNR process;

[0013] Step S4: Design a specific CMDP model for the DDNR task based on the improved IPO solution framework, including the specific forms of the state space, action space, reward function, constraint function, and state transition function;

[0014] Step S5: Use the improved IPO algorithm constructed in steps S1, S2, and S3 to solve the CMDP model in step S4, design offline training rules to train the agent to learn the optimal DDNR policy, save the trained agent's policy, and apply the policy to the online execution process of the distribution network.

[0015] In step S3, the spatiotemporal awareness neural network model includes a GCN network and a GRU network. The GCN network performs convolution operations on graph structure data by defining an invariant convolution kernel. The GRU network regulates the transmission and memory of information through a gating mechanism. The structure of the GRU network includes a reset gate and an update gate.

[0016] In step S4, the CMDP model includes:

[0017] State Space: The state space contains spatial topology information and electrical characteristics representing multiple time periods, let the PDN state at the current time t be... ,in, Pt This refers to the net injection of active power into the node. Qt This refers to the net reactive power injection at the node. VtThe node voltage amplitude, It The amplitude of the branch current. Because of the active power loss of the line, therefore The Xt moving window, composed of past w steps, is used to represent st;

[0018] Action space: The action of DDNR is set as a switch pair SP. If the closed switch and the open switch are the same, it is considered that PDN is not performing topology reconstruction at the current moment and does not incur switching costs.

[0019] Reward Function: The reward function is the network loss cost and switching action cost of the PDN. The reward function at time t is... ,in For network loss costs, For the cost of switching operation, branch at time t ij Active power loss;

[0020] Constraint function: The node voltage Vi is constrained as follows: The branch current Iij is limited to: Vmin is the lower voltage limit, Vmax is the upper voltage limit, N is the number of distribution network node voltages, Imax is the upper limit of branch current, and ε is the number of distribution network branches.

[0021] State transition probability distribution: State transition probability distribution The DRL agent takes action. at Time state from st Transferred to st The probability of +1, where, This represents the action of the agent at time t. For a set of actions, Indicates the states at times t and t+1. For a set of states, This represents the power flow calculation process.

[0022] More specifically, in step S1, the CMDP includes 6 elements, containing a state space. Action space reward function constraint functions State transition probability distribution Discount factor The DRL agent, based on the current state st Select Action at By a strategy The goal of the DRL agent is to find a strategy that maximizes the cumulative reward of the discount while satisfying the constraints. The expression is as follows:

[0023]

[0024] In the formula, Representative strategy Relative to the cumulative discount cost of the constraint, Representation Strategy feasible domain, As a discount factor, Let be the reward function at time t, and d be the constraint. Represents the expectation of a random source. The control horizon of the agent is represented; the goal of CMDP is to select a policy that satisfies the constraints. To maximize cumulative discount returns .

[0025] More specifically, in step S1, the policy gradient algorithm uses policy gradient theory to optimize the objective. As shown in the following formula:

[0026]

[0027] In the above formula, For the parameters of the policy network, This represents the current policy function based on... s t Select Action a t The probability, where N is the length of a control horizon, and the state-action value function. Q ( s t , a t Use one parameter as The neural network representation of is calculated as follows:

[0028]

[0029] In the above formula, In the state Select action below The rewards received It is a state-value function;

[0030] The policy gradient algorithm is a near-end policy optimization algorithm, and its objective function is expressed as:

[0031]

[0032] in:

[0033]

[0034] In the above formula, , Let be the expectation, representing the empirical average over a finite sample; Indicating in strategy Down t The advantage function estimate for the decision step; Indicates the ratio of the old to the new strategies; This is the old strategy before the update; For strategy parameters; This is a truncation function that controls the changes between the old and new strategies within a certain range. Inside, The truncation constant is used to set the range of policy updates.

[0035] More specifically, in the IPO algorithm of step S1, there is an index function for each constraint-satisfied problem. Satisfy the following expression:

[0036] ,

[0037] That is, when the constraints are satisfied, the problem is solved as an unconstrained policy optimization problem that only considers the reward; when any constraint is violated, the policy needs to be adjusted first to satisfy the constraint.

[0038] The logarithmic barrier function is defined as follows:

[0039]

[0040] Where k is a hyperparameter;

[0041] At this point, by expanding the objective function using the logarithmic barrier function, the final objective function of the IPO algorithm becomes:

[0042] .

[0043] More specifically, in step S2, the principle of distribution network reconfiguration based on branch switching mechanism is: close a tie switch i to form a loop, and then open a sectionalizing switch j in the loop to restore the distribution network to a radial shape. This action is regarded as a switch pair SP(i, j).

[0044] Store all feasible switch pairs SP(i, j) at time t in a binary matrix at time t. Mt If the switch is feasible for SP, then Mt (i, j) = 1; otherwise 0; select from the binary mask matrix M at each decision step. Mt Reconstruct the swap pairs (i, j)=1;

[0045] The output of the IPO strategy network is set to a two-dimensional array h.ij (s t ), its dimensions and M Similarly, in order to output feasible switch pairs SP, a security layer based on feasible action masks is added to the output layer of the policy network, represented as:

[0046]

[0047] In the formula, It is a 2D array representing the probability distribution of SP(i, j). ; Let represent the state value of switch pair (i, j) at time t, and k and l represent the total number of switches that can be closed and opened. This represents the hidden layer state values ​​of all actions output by the action network at time t. This represents the state value of all switch pairs at time t.

[0048] More specifically, the update process of the spatiotemporal awareness neural network is represented by the following equation:

[0049]

[0050] In the above formula, The feature matrix of PDN, Let PDN be the adjacency matrix. This represents the adjacency matrix plus self-loops. Let be the degree matrix of the adjacency matrix. It is a non-linear activation function. The weight matrix is ​​a learnable matrix. For the new feature representation after GCN, To reset the door, To update the door, and These are the hidden states at the current time and the previous time, respectively. , and These are the weight matrices, , and For bias terms, Let X be a candidate hidden state in a GRU network. t Let t represent the characteristic vector of the distribution network at time t.

[0051] More specifically, the process is divided into exploration, training, and execution phases;

[0052] During the exploration phase, the IPO algorithm is based on the current strategy. Interact with the PDN environment;

[0053] During the training phase, the parameters of the actor and critic networks of the IPO agent are adjusted based on the current empirical trajectory. and The IPO agent performs updates while capturing the spatiotemporal dynamics of the uncertainty of the PDN environment and the DDNR problem from the experience trajectory.

[0054] During the execution phase, the trained policy network is saved, the spatiotemporal features of the environment are extracted through the spatiotemporal graph convolutional network, and the spatiotemporal features are input into the policy network with action masks to generate a secure DDNR policy.

[0055] The present invention also provides a computer device, the computer device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-described method for dynamic network reconfiguration of distribution networks based on constraint-enhanced security reinforcement learning.

[0056] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described method for dynamic network reconfiguration of a distribution network based on constraint-enhanced security reinforcement learning.

[0057] The beneficial effects of this invention are as follows: This invention achieves dynamic reconfiguration of active power distribution networks with high penetration of photovoltaic (PV) and wind power (WP) generation through a secure reinforcement learning method, solving problems such as increased active power loss and voltage / current exceeding limits caused by the integration of new energy sources. Furthermore, the proposed method can perform real-time online calculation and execution in the face of uncertainties in PV and WP, ​​and the obtained strategy possesses security guarantees. Compared with existing technologies, the method proposed in this invention has the following three advantages:

[0058] The IPO algorithm used introduces a logarithmic barrier function. This design allows the algorithm to filter out policies that do not meet the constraints during the training process, thus taking into account the soft constraints in the DDNR process. Compared with existing reinforcement learning methods, IPO does not require specified penalty clauses or adjustment of cumbersome penalty factors, and its calculation is simpler.

[0059] When performing DDNR, it is necessary to always satisfy the hard constraints of the radial topology of the distribution network (i.e., open-loop operation). Therefore, this invention combines the idea of ​​branch switching method and designs an action mask based on the state of all switches in the distribution network at each decision step. This mask is then applied to the output layer of the IPO strategy network, which ensures both the hard constraints of the radial topology of the distribution network and reduces the dimension of the action space.

[0060] The spatiotemporal awareness neural network model proposed in this invention integrates graph convolutional network (GCN) and gated recurrent unit neural network (GRU). GCN is used to obtain the spatial topology and power flow distribution characteristics of the DDNR process, while GRU is used to capture the timing characteristics of power generation and consumption of PV, WP, and base load. Therefore, with the support of the spatiotemporal awareness neural network model, the reinforcement learning agent can better perceive the strong spatiotemporal uncertainty of PDN in the DDNR process. Attached Figure Description

[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0062] Figure 1 This is a schematic diagram of a dynamic network reconfiguration method for distribution networks based on constraint-enhanced security reinforcement learning, provided as an embodiment of the present invention.

[0063] Figure 2 This is a schematic diagram of an IEEE 14-node distribution network provided for one embodiment of the present invention.

[0064] Figure 3 This is a schematic diagram of a dynamic network reconfiguration framework for distribution networks based on constraint-enhanced security reinforcement learning, provided as an embodiment of the present invention.

[0065] Figure 4 This is a schematic diagram of an IEEE 33-node distribution network provided for one embodiment of the present invention.

[0066] Figure 5 The training results of the proposed method provided in one embodiment of the present invention on two test systems.

[0067] Figure 6 This is a comparison chart of PDN operating costs on a typical test day under different schemes provided in one embodiment of the present invention. Detailed Implementation

[0068] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0069] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0070] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0071] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0072] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0073] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0074] Reference Figure 1-3 This is the first embodiment of the present invention, which provides a method for dynamic network reconfiguration of distribution networks based on constraint-enhanced security reinforcement learning, including:

[0075] Step S1: Treat the DDNR task of the high-penetration active distribution network of photovoltaic and wind power generation as a CMDP model, and construct a secure reinforcement learning solution framework based on the IPO algorithm. The logarithmic barrier function of IPO can enable the agent to satisfy the soft constraints of the DDNR task when learning the policy.

[0076] A general CMDP consists of 6 elements, including the state space. Action space reward function constraint functions State transition probability distribution Discount factor The DRL agent, based on the current state... st Select Action at By a strategy Decision, strategy It can be represented by a parameterized neural network. The goal of the DRL agent is to find a strategy that maximizes the cumulative reward of the discount while satisfying the constraints. The expression is as follows:

[0077] (1)

[0078] In the formula, Representative strategy Relative to the cumulative discount cost of the constraint, Representation Strategy feasible domain, As a discount factor, Let be the reward function at time t, and d be the constraint. Represents the expectation of a random source. This represents the control horizon of the agent. The goal of CMDP is to select a policy that satisfies the constraints. To maximize cumulative discount returns Common policy gradient algorithms use policy gradient theory to optimize the objective. As shown in the following formula:

[0079] (2)

[0080] In the above formula, For the parameters of the policy network, This represents the current policy function based on... s t Select Action a t The probability, where N is the length of a control horizon, and the state-action value function. Q ( s t , a t Use one parameter as The neural network representation of is calculated as follows:

[0081] (3)

[0082] In the above formula, In the state Select action below The rewards received Given the state-value function, the Proximal Policy Optimization (PPO) algorithm, as a common policy gradient algorithm, employs a truncation function to make the gradient update in equation (2) more stable and efficient. Its objective function can be expressed as:

[0083] (4)

[0084] in:

[0085] (5)

[0086] In the above formula, , Let be the expectation, representing the empirical average over a finite sample; Indicating in strategy Down t The advantage function estimate for the decision step; Indicates the ratio of the old to the new strategies; This is the old strategy before the update; For strategy parameters; This is a truncation function that controls the changes between the old and new strategies within a certain range. Inside, The truncation constant is used to set the range of policy updates.

[0087] Based on the framework of the PPO algorithm described above, this invention further proposes the IPO algorithm to solve the CMDP problem. The IPO algorithm inherits the pruning objective of the PPO algorithm and expands the constraints through a logarithmic barrier function. For each problem satisfying the constraints, if there is an index function... satisfy:

[0088] (6)

[0089] That is, when the constraints are satisfied, the problem is solved as an unconstrained policy optimization problem considering only the reward; when any constraint is violated, the policy needs to be adjusted to satisfy the constraint first. Thus, by expanding the objective function, the original CMDP problem can be simplified into an unconstrained optimization problem. The logarithmic barrier function is a differentiable approximation of the objective function, defined as:

[0090] (7)

[0091] Where k is a hyperparameter. Now, by expanding the objective function using a logarithmic barrier function, the final objective function of the IPO algorithm is:

[0092] (8)

[0093] Step S2: Design a security layer based on a feasible action mask in conjunction with the branch switching mechanism of the distribution network, and apply the security layer to the output layer of the IPO policy network so that the policy output by the agent always meets the hard constraints of the DDNR task.

[0094] The principle of distribution network reconfiguration based on branch switching mechanism is as follows: closing a tie switch i forms a loop, and then opening a sectionalizing switch j within the loop restores the distribution network to a radial configuration. This action can be viewed as a switch pair SP(i, j). Figure 2 Taking an IEEE 14-node distribution network as an example, the system has 14 nodes and 16 branches. Node 1 is a transformer node with 13 normally closed sectionalizing switches and 3 normally open tie switches. Closing the 3 tie switches sequentially can form 3 loops, so feasible switch pairs will be generated in these 3 loops. For example, in the initial state, SP(14,7) (closing switch 14 and opening switch 7) can be selected to complete the topology reconfiguration of the distribution network and realize the transfer of energy between different feeders.

[0095] This invention stores all feasible SP(i, j) at time t in a binary matrix at time t. Mt If SP is feasible, then Mt (i, j) = 1, otherwise 0, therefore at each decision step, a selection can be made from the binary mask matrix M. Mt Reconstruct the swap pairs where (i, j) = 1. Figure 2 In the distribution network, M The dimension is 16×16, and there are not many feasible SPs at each time step, so M can be regarded as a sparse matrix. To match the design of the switch pairs, this invention sets the output of the IPO policy network as a two-dimensional array h. ij (s t ), its dimensions and M Similarly, to output feasible SPs, a security layer based on feasible action masks is added to the output layer of the policy network, which can be represented as:

[0096] (9)

[0097] In the formula, It is a 2D array representing the probability distribution of SP(i, j). This ensures the feasibility of actions taken by reinforcement learning agents; Let represent the state value of switch pair (i, j) at time t, and k and l represent the total number of switches that can be closed and opened. This represents the hidden layer state values ​​of all actions output by the action network at time t. This represents the state value of all switch pairs at time t.

[0098] Traditional branch-swapping methods use heuristics to find the optimal SP (Solution Point). This approach only considers the optimal solution for the current decision step. However, DDNR (Destination-Driven Regression) is a typical sequential decision problem, where the current reconfiguration scheme affects subsequent outcomes. This invention uses the IPO (Initial Programming) algorithm to solve for the optimal SP, which considers the cumulative reports and constraints across the entire control horizon.

[0099] Step S3: Construct a spatiotemporal perception neural network model and apply the spatiotemporal perception neural network to the input layer of the IPO policy network and the value network, so that the agent can better perceive the spatiotemporal change characteristics of the power flow in the distribution network during the DDNR process.

[0100] The spatiotemporal awareness neural network model constructed in this invention consists of a GCN network and a GRU network. The GCN network performs convolution operations on graph-structured data using an invariant convolution kernel, effectively acquiring and updating node feature information from neighboring nodes. Its core idea is to apply graph convolution operations to aggregate the neighborhood information of nodes, obtaining updated node representations. Specifically, by combining the adjacency matrix and feature matrix of the graph and applying a nonlinear activation function, a new node representation is obtained.

[0101] GRU uses gating mechanisms to regulate the transmission and retention of information, thereby learning the current state while preserving important historical information. The GRU structure includes reset gates and update gates. These two gating mechanisms effectively control the influence of past information on the current state, thus adapting to the characteristics of time-series data. The reset gate determines how new inputs are combined with previous hidden states, while the update gate controls the updating and retention process of the hidden states.

[0102] Therefore, the update process of the spatiotemporal awareness neural network can be expressed by the following equation:

[0103] (9)

[0104] (10)

[0105] (11)

[0106] (12)

[0107] (13)

[0108] In the above formula, The feature matrix of PDN, Let PDN be the adjacency matrix. This represents the adjacency matrix plus self-loops. Let be the degree matrix of the adjacency matrix. It is a non-linear activation function. The weight matrix is ​​a learnable matrix. For the new feature representation after GCN, To reset the door, To update the door, and These are the hidden states at the current time and the previous time, respectively. , and These are the weight matrices, , and For bias terms, Let X be a candidate hidden state in a GRU network. t Let t represent the characteristic vector of the distribution network at time t.

[0109] Step S4: Design a specific CMDP model for the DDNR task based on the improved IPO solution framework, including the specific forms of the state space, action space, reward function, constraint function, and state transition function.

[0110] For the DDNR task of PDN, this invention sets up a specific CMDP model as follows:

[0111] State Space: The state space of this invention includes spatial topology information and electrical characteristics representing multiple time periods, let the PDN state at the current time t be... ,in, Pt This refers to the net injection of active power into the node. Qt This refers to the net reactive power injection at the node. Vt The node voltage amplitude, It The amplitude of the branch current. Because of the active power loss of the line, therefore The Xt moving window, composed of past w steps, is used to represent st in order to infer the future trend of each electrical quantity.

[0112] Action Space: In this invention, the action of DDNR is set as a SP, which has been described in step S2. If the closed switch and the open switch are the same, it is assumed that PDN is not performing topology reconfiguration at the current moment and does not incur switching costs.

[0113] Reward Function: The reward function of this invention is the network loss cost and switching action cost of the PDN. The reward function at time t is: ,in For network loss costs, For the cost of switching operation, branch at time t ij Active power loss;

[0114] Constraints: The DDNR process should ensure that node voltages and branch currents do not violate regulations. The node voltage Vi is constrained as follows: The branch current Iij is limited to: Vmin is the lower voltage limit, Vmax is the upper voltage limit, N is the number of distribution network node voltages, Imax is the upper limit of branch current, and ε is the number of distribution network branches.

[0115] State transition probability distribution: State transition probability distribution The DRL agent takes action. at Time state from st Transferred to st The probability of +1, where, This represents the action of the agent at time t. For a set of actions, Indicates the states at times t and t+1. For a set of states, Representing the power flow calculation process, DDNR will trigger power flow changes, and the agent seeks the best strategy from these changes.

[0116] Step S5: Use the improved IPO algorithm constructed in steps S1, S2, and S3 to solve the CMDP model in step S4, design offline training rules to train the agent to learn the optimal DDNR policy, save the trained agent's policy, and apply the policy to the online execution process of the distribution network.

[0117] The improved IPO algorithm proposed in this invention for solving the DDRN problem can be divided into three main phases: exploration, training, and execution. Figure 3 The overall framework of the proposed dynamic network reconfiguration method for distribution networks based on constraint-enhanced security reinforcement learning is presented.

[0118] During the exploration phase, the IPO algorithm is based on the current strategy. It interacts with the PDN environment to generate an experience trajectory, the length of which represents the range of the control horizon.

[0119] During the training phase, the parameters of the actor and critic networks of the IPO agent are adjusted based on the current empirical trajectory. and The purpose of the update is to maximize the cumulative discount reward of the control horizon and ensure that the current policy meets the cumulative constraints. At the same time, the IPO agent can also capture the spatiotemporal dynamic characteristics of the uncertainty of the PDN environment and the DDNR problem from the experience trajectory.

[0120] During the execution phase, the trained policy network is saved, the spatiotemporal features of the environment are extracted through the spatiotemporal graph convolutional network, and the spatiotemporal features are input into the policy network with action masks to generate a secure DDNR policy.

[0121] Reference Figure 4-6 As an embodiment of the present invention, a dynamic network reconfiguration method for distribution networks based on constraint-enhanced security reinforcement learning is provided. To verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.

[0122] use Figure 2 and 4 The effectiveness of the proposed method was verified using the IEEE 14-node and IEEE 33-node test systems. In the IEEE 14-node system, two PVs are located at nodes 2 and 13, and one WP is located at node 7. In the IEEE 33-node system, three PVs are located at nodes 13, 25, and 31, and three WPs are located at nodes 7, 18, and 22. Each node in the test system aggregates a certain amount of electrical load. The PV, WP, and electrical load data in this application were all collected from real-world operating scenarios. To adapt to different test systems, the data was scaled proportionally. One year's worth of data was used to train the DRL agent, and a typical day was selected in each of the two test systems to test the trained agent. This invention used one year's worth of data to train the IPO agent for 1000 rounds. In each round, 10 days of data were randomly selected for interaction, including 240 decision steps. The timestamp of each decision step was 1 hour, corresponding to the data resolution. The upper and lower limits of the node voltage were set to [0.95, 1.05]. The active power loss cost Cp is set at $40 / MWh, and the switching cost is $0.5.

[0123] To verify the effectiveness of the proposed secure reinforcement learning algorithm in solving the DDNR task, the training process of the algorithm is analyzed. Figure 5 The algorithm of the present invention is shown to have average cumulative discount reward and its moving average after 1000 rounds of training on two test systems.

[0124] At the beginning of training, the randomly initialized parameters caused the agents to exhibit low performance, and even power flow non-convergence occurred on the IEEE14 system. This was due to the randomness of the policy network causing invalid actions taken by the DRL agents. As training progressed, the actor and critic network parameters of the algorithm in this invention were continuously updated, and the agents learned the correct policies in this process. After 400 rounds of training, the agents were able to make good actions and obtained high average cumulative discount rewards. Table 1 shows the mean values ​​of average network loss, average switching action cost, average voltage violation rate, and average current violation rate for the last 200 rounds of the two test systems. The results show that the proposed algorithm achieved stable high rewards in the later stages of training. The policy output by the algorithm of this invention can reduce the operating cost of PDN and ensure the satisfaction of operating constraints.

[0125] Table 1

[0126]

[0127] To verify the effectiveness of the proposed algorithm and the advantages of DDNR performance, the trained agent of the present invention was tested on a typical day. At the same time, the brute-force search method (BFS) and the DDNR scheme of the genetic algorithm were implemented to compare the performance with the proposed method. Figure 6 The typical daily operating costs of two test systems under three control schemes and without any control scheme are presented. Figure 6 As can be seen, compared with the scheme without any control, the algorithm of the present invention reduces the total running cost by 33.2% and 41.8% on the two test systems. Compared with the heuristic GA algorithm, the proposed method improves performance by 14.76% and 17%. Compared with the BFS method, the algorithm of the present invention improves performance by 3% and 6.7%. The BFS method has lower running cost in some decision steps, but this method only considers the immediate benefit and does not take into account the benefit of the entire control horizon. Moreover, when the system is large, calculating the power flow for every possible action is very time-consuming. Table 2 shows the actions SP of the proposed algorithm of the present invention in a day.

[0128] Table 2

[0129]

Claims

1. A power distribution network dynamic network reconfiguration method based on constraint reinforcement safety enhanced learning, characterized in that: include, Step S1: Treat the DDNR task of the high-penetration active distribution network of photovoltaic and wind power generation as a CMDP model, and construct a secure reinforcement learning solution framework based on the IPO algorithm. The logarithmic barrier function of IPO enables the agent to satisfy the soft constraints of the DDNR task when learning the policy. Step S2: Design a security layer based on a feasible action mask in conjunction with the branch switching mechanism of the distribution network, and apply the security layer to the output layer of the IPO policy network so that the policy output by the agent always meets the hard constraints of the DDNR task. Step S3: Construct a spatiotemporal perception neural network model and apply the spatiotemporal perception neural network to the input layer of the IPO policy network and the value network, so that the agent can better perceive the spatiotemporal change characteristics of the power flow in the distribution network during the DDNR process; Step S4: Design a specific CMDP model for the DDNR task based on the improved IPO solution framework, including the specific forms of the state space, action space, reward function, constraint function, and state transition function; Step S5: Use the improved IPO algorithm constructed in steps S1, S2, and S3 to solve the CMDP model in step S4, design offline training rules to train the agent to learn the optimal DDNR policy, then save the trained agent's policy and apply the policy to the online execution process of the distribution network. In step S3, the spatiotemporal awareness neural network model includes a GCN network and a GRU network. The GCN network performs convolution operations on graph structure data by defining an invariant convolution kernel. The GRU network regulates the transmission and memory of information through a gating mechanism. The structure of the GRU network includes a reset gate and an update gate. In step S4, the CMDP model includes: State Space: The state space contains spatial topology information and electrical characteristics representing multiple time periods, let the PDN state at the current time t be... ,in, Pt This refers to the net injection of active power into the node. Qt This refers to the net reactive power injection at the node. Vt The node voltage amplitude, It The amplitude of the branch current. Because of the active power loss of the line, therefore The Xt moving window, composed of past w steps, is used to represent st; Action space: The action of DDNR is set as a switch pair SP. If the closed switch and the open switch are the same, it is considered that PDN is not performing topology reconstruction at the current moment and does not incur switching costs. Reward Function: The reward function is the network loss cost and switching action cost of the PDN. The reward function at time t is... ,in For network loss costs, For the cost of switching operation, branch at time t ij Active power loss; Constraint function: The node voltage Vi is constrained as follows: The branch current Iij is limited to: Vmin is the lower voltage limit, Vmax is the upper voltage limit, N is the number of distribution network node voltages, Imax is the upper limit of branch current, and ε is the number of distribution network branches. State transition probability distribution: State transition probability distribution The DRL agent takes action. at Time state from st Transferred to st The probability of +1, where, This represents the action of the agent at time t. For a set of actions, Indicates the states at times t and t+1. For a set of states, This represents the power flow calculation process.

2. The power distribution network dynamic network reconfiguration method based on constraint-enhanced safe reinforcement learning of claim 1, wherein: In step S1, the CMDP includes 6 elements, including a state space. Action space reward function constraint functions State transition probability distribution Discount factor ; The DRL agent, based on the current state st Select Action at By a strategy The goal of the DRL agent is to find a strategy that maximizes the cumulative reward of the discount while satisfying the constraints. The expression is as follows: In the formula, Representative strategy Relative to the cumulative discount cost of the constraint, Representation strategy feasible domain, As a discount factor, Let be the reward function at time t, and d be the constraint. Represents the expectation of a random source. The control horizon of the agent is represented; the goal of CMDP is to select a policy that satisfies the constraints. To maximize cumulative discount returns .

3. The power distribution network dynamic network reconfiguration method based on constraint-enhanced safe reinforcement learning of claim 2, wherein: In the step S1, the policy gradient algorithm uses a policy gradient theory to optimize the objective as follows: In the above formula, For the parameters of the policy network, This represents the current policy function based on... s t Select Action a t The probability, where N is the length of a control horizon, and the state-action value function. Q ( s t , a t Use one parameter as The neural network representation of is calculated as follows: In the above formula, is the state The selected action The reward obtained, is the state value function; The policy gradient algorithm is a near-end policy optimization algorithm, and its objective function is expressed as: in: In the above formula, , Let be the expectation, representing the empirical average over a finite sample; Indicating in strategy Down t The advantage function estimate for the decision step; Indicates the ratio of the old to the new strategies; This is the old strategy before the update; For strategy parameters; This is a truncation function that controls the changes between the old and new strategies within a certain range. Inside, The truncation constant is used to set the range of policy updates.

4. The power distribution network dynamic network reconfiguration method based on constraint-enhanced safe reinforcement learning of claim 3, wherein: In the IPO algorithm of the step S1, for each constraint satisfaction problem, there is an indicator function satisfies the following expression: , That is, when the constraints are satisfied, the problem is solved as an unconstrained policy optimization problem that only considers the reward; when any constraint is violated, the policy needs to be adjusted first to satisfy the constraint. The logarithmic barrier function is defined as follows: Where k is a hyperparameter; At this point, by expanding the objective function using the logarithmic barrier function, the final objective function of the IPO algorithm becomes: 。 5. The power distribution network dynamic network reconfiguration method based on constraint-enhanced safe reinforcement learning of claim 4, wherein: In step S2, the principle of distribution network reconfiguration based on branch exchange mechanism is: close a tie switch i to form a loop, and then open a sectionalizing switch j in the loop to restore the distribution network to a radial shape. This action is regarded as a switch pair SP(i, j). All feasible switch pairs SP(i, j) at time t are stored in a binary matrix Mt If a switch pair SP is feasible, then Mt (i, j) = 1; otherwise 0; At each decision step, a reconstruction operation is performed on the exchange pairs (i, j) for which M(i, j) = 1. Mt (i, j) = 1. The output of the IPO policy network is set as a two-dimensional array h ij (s t ) with the same dimension as M In order to output feasible switch pairs SP, a safety layer based on the actionable action mask is added to the output layer of the policy network, denoted as: In the formula, It is a 2D array representing the probability distribution of SP(i, j). ; Let represent the state value of switch pair (i, j) at time t, and k and l represent the total number of switches that can be closed and opened. This represents the hidden layer state values ​​of all actions output by the action network at time t. This represents the state value of all switch pairs at time t.

6. The power distribution network dynamic network reconfiguration method based on constraint-enhanced safe reinforcement learning of claim 5, wherein: The update process of the spatiotemporal awareness neural network can be represented by the following equation: In the above formula, The feature matrix of PDN, Let PDN be the adjacency matrix. This represents the adjacency matrix plus self-loops. Let be the degree matrix of the adjacency matrix. It is a non-linear activation function. The weight matrix is ​​a learnable matrix. For the new feature representation after GCN, To reset the door, To update the door, and These are the hidden states at the current time and the previous time, respectively. , and These are the weight matrices, , and For bias terms, Let X be a candidate hidden state in a GRU network. t Let t represent the characteristic vector of the distribution network at time t.

7. The power distribution network dynamic network reconfiguration method based on constraint-enhanced safe reinforcement learning of claim 6, wherein: The process is divided into three phases: exploration, training, and execution. In the exploration phase, the IPO algorithm interacts with the current policy and the PDN environment; In the training phase, the parameters of the actor network and the critic network of the IPO agent are updated according to the current experience trajectory, while the IPO agent captures the uncertainty of the PDN environment and the spatiotemporal dynamic characteristics of the DDNR problem from the experience trajectory. and ​ During the execution phase, the trained policy network is saved, the spatiotemporal features of the environment are extracted through the spatiotemporal graph convolutional network, and the spatiotemporal features are input into the policy network with action masks to generate a secure DDNR policy.

8. A computer device, comprising: The computer device includes: one or more processors; and a memory for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to perform the dynamic network reconfiguration method for distribution networks based on constraint-enhanced security reinforcement learning as described in any one of claims 1-7.

9. A computer readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the dynamic network reconfiguration method for distribution networks based on constraint-enhanced security reinforcement learning as described in any one of claims 1-7.