Power distribution network reconfiguration method and system based on topological security constraints and integrated reinforcement learning

By introducing topological security constraints and integrated reinforcement learning into the power system, improving topology coding, and using topology masking mechanisms, the problems of voltage offset and power flow changes in large-scale distributed renewable energy grid integration by traditional algorithms are solved, achieving efficient and economical operation and security assurance of the power system.

CN119891150BActive Publication Date: 2026-01-09HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411713658.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2026-01-09
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Traditional mathematical modeling optimization algorithms struggle to effectively address issues such as voltage offset and power flow changes when dealing with large-scale distributed renewable energy grid integration. Furthermore, reinforcement learning faces challenges in fitting the data and incurring high computational costs in reconstruction optimization tasks.

Method used

By introducing topological safety constraints and ensemble reinforcement learning, the topological encoding method is improved, a topological masking mechanism is used to ensure model security, and a cluster strategy is used to build an action network group for prediction and selection. The action network parameters and policies are optimized by combining a dual-delay-deterministic policy gradient framework for training.

Benefits of technology

It has enabled efficient and economical operation of the power system under large-scale distributed photovoltaic grid connection conditions, improved the solution efficiency and training stability of the model, reduced grid losses and ensured the security of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119891150B_ABST
    Figure CN119891150B_ABST
Patent Text Reader

Abstract

The application discloses a power distribution network reconstruction method and system based on topological safety constraints and integrated reinforcement learning. The method comprises: power grid environment modeling, including system state quantity design, reward function design, power grid constraint embedding and system power flow calculation; constructing a topological safety layer, using a topological mask to correct the topological reconstruction action through a detection algorithm; using an integrated action network, a parameter doping mechanism and a network pruning mechanism to pre-judge different actions; assembling a reinforcement learning training framework to train the initial policy network and obtain network parameters; using the trained policy network to obtain the scheduling strategy of the power distribution network reconstruction according to the state data of the system. The application superimposes a topological safety layer on the basis of reinforcement learning and introduces an action network integration mechanism, solves the problem of lack of safety of the output strategy of the reinforcement learning model, alleviates the volatility of the parameter training process, and can efficiently obtain the reconstruction scheduling strategy of the power distribution network in real time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of power system optimal scheduling, in particular to a distribution network reconstruction method and system based on topological safety constraints and integrated reinforcement learning. BACKGROUND

[0002] With the gradual depletion of fossil energy, the low-carbon transformation of energy structure has attracted widespread attention in the world, prompting large-scale renewable distributed energy to be widely connected to the power grid to meet the growing energy demand. However, the randomness and volatility of new energy generation frequently cause problems such as system voltage deviation and power flow change, seriously threatening the safe and stable operation of the distribution network system.

[0003] To effectively cope with the challenges brought by new energy uncertainty, current scheduling optimization research focuses on integrating energy storage devices, reactive power compensation devices, flexible loads and other diversified source and load adjustable resources, and combining dynamic reconstruction strategies on the network side to realize the coordinated optimization of system operation state by directly controlling the power flow distribution. In the traditional research framework, the mathematical modeling optimization algorithm has a significant decrease in solving efficiency when facing high-dimensional discrete topological variables with uncertain factors, and it is difficult to meet the stringent requirements of real-time scheduling of the power grid. Under this background, the reinforcement learning theory opens up a new path for the optimal scheduling of complex power grids. This method does not rely on accurate system parameters and physical models, and can quickly adapt to environmental changes and directly train a general decision model for complex optimization problems rather than a single strategy, with extremely high solving speed. However, when it comes to reconstruction optimization tasks, the huge discrete topological space and high proportion of invalid actions also make the neural network have obvious fitting difficulties in training. Enumerating the action space not only has high computational cost, but also is difficult to meet the application requirements of practical engineering. Although continuous processing can be tried, it is easy to complicate the problem. Some studies try to filter the effective topology set to reduce the action space, which can alleviate the fitting pressure to some extent, but also indirectly affects the optimization performance of the model. SUMMARY

[0004] The application aims to provide a distribution network reconstruction method and system based on topological safety constraints and integrated reinforcement learning, which improves the coding method of topology, uses a multi-dimensional discrete space for representation, and introduces a topology mask mechanism to ensure the safety and computational efficiency of the model; further, the cluster strategy is used to construct an action network group to realize the pre-judgment and screening of actions, and enhance the stability of the training process.

[0005] Technical scheme: In order to achieve the above application purpose, the distribution network reconstruction method based on topological safety constraints and integrated reinforcement learning provided by the application comprises the following steps:

[0006] Step 1: Construct a power grid environment for interacting with a reinforcement learning model to obtain training data, including system state quantity design, reward function design, power grid constraint embedding and system steady state quantity calculation, the purpose of the reinforcement learning model is to minimize the network loss and the number of line switch actions by adjusting the network structure and energy storage resources of the system; the policy network of the model gives the reconstruction action according to the current state data of the power grid environment, and calculates the reward of the current action according to the reward function, and returns the reward and the next state data of the power grid environment; the state data includes the topology action at the last time of the network, the energy storage capacity state at the last time, and the node power in the last period, and the reconstruction action includes the topology action and the energy storage action, wherein the topology reconstruction action is designed to selectively disconnect a controllable branch in different independent loops;

[0007] Step 2: Construct a topology safety layer, which detects and corrects the topology reconstruction action calculated by the model by using the action mask through the topology detection algorithm, and when it is detected that the branch operated at present is a tree branch, a negative infinite mask element is superimposed on the branch, so that it will not be selected;

[0008] Step 3: Combine multiple action networks into an integrated action group, and pre-judge and screen different reconstruction actions of the integrated action group through a parameter doping mechanism and a network pruning mechanism, the parameter doping mechanism sets a period counter for each action network, counts the optimal action frequency of each action network in the first period, and transfers the parameters of the optimal action network to other action networks for parameter disturbance; the network pruning mechanism records the optimal action frequency of each action network in the second period, and deletes it from the action network group;

[0009] Step 4: Train the reinforcement learning model using a double-delay-deterministic policy gradient framework to obtain network parameters for generating a reconstruction scheduling policy;

[0010] Step 5: Use the trained reinforcement learning model to obtain the scheduling strategy of the distribution network reconstruction according to the state data of the current system.

[0011] The application also provides a distribution network reconstruction system based on topology safety constraints and integrated reinforcement learning, comprising:

[0012] A power grid environment modeling module is configured to construct a power grid environment for obtaining training data by interacting with a reinforcement learning model, including system state quantity design, reward function design, power grid constraint embedding, and system steady-state quantity calculation, the purpose of the reinforcement learning model being to minimize network loss and line switch action times by adjusting the network structure and energy storage resources of the system; a policy network of the model gives a reconstruction action according to current state data of the power grid environment, and calculates the reward of the current action according to the reward function, and returns the reward and the next state data of the power grid environment; the state data includes the topology action at the previous moment, the energy storage capacity state at the previous moment, and the node power in the previous period, and the reconstruction action includes a topology action and an energy storage action, wherein the topology reconstruction action is designed to selectively disconnect a controllable branch in different independent loops in turn;

[0013] A topology safety detection module is configured to construct a topology safety layer, which detects and corrects the topology reconstruction action calculated by the model by using an action mask through a topology detection algorithm, and when it is detected that the branch of the current operation is a tree branch, a negative infinite mask element is superimposed on the branch, so that it will not be selected;

[0014] An action network pre-screening module is configured to combine multiple action networks into an integrated action group, and pre-judge and screen different reconstruction actions of the integrated action group through a parameter doping mechanism and a network pruning mechanism, the parameter doping mechanism sets a period counter for each action network to count the optimal action times of each action network in a first period, and transfers the parameters of the optimal action network to other action networks to perturb the parameters thereof; the network pruning mechanism records the optimal action frequency of each action network in a second period, and deletes it from the action network group;

[0015] A reinforcement learning model training module is configured to train the reinforcement learning model by using a double-delay-deterministic policy gradient framework to obtain network parameters for generating a reconstruction scheduling policy;

[0016] A distribution network reconstruction strategy determination module is configured to use the trained reinforcement learning model to obtain a corresponding distribution network reconstruction scheduling policy according to the state data of the current system.

[0017] The application also provides a computer device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs are executed by the processor to implement the steps of the distribution network reconstruction method based on topology safety constraints and integrated reinforcement learning as described above.

[0018] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the power distribution network reconstruction method based on topology safety constraints and integrated reinforcement learning.

[0019] Advantages: Compared with the prior art, the application has the following significant advantages:

[0020] (1) With the continuous grid connection of large-scale distributed photovoltaics, the power distribution network is facing severe challenges such as voltage fluctuation and network loss increase. Deep reinforcement learning has a great improvement in optimization problem solving efficiency compared with traditional algorithms, but when it comes to collaborative optimization problems involving network reconstruction, the model is usually difficult to converge and the training stability is poor. Therefore, the application provides a power distribution network reconstruction optimization method based on topology safety constraints and integrated reinforcement learning, which embeds a topology safety layer into reinforcement learning, alleviates the stable optimization problem of large topology discrete space through coding design, and realizes strong safety guarantee. At the same time, a special topology connectivity detection algorithm is introduced to meet the topology safety requirements of the power system with high computational efficiency;

[0021] (2) The application provides a training technique of integrated reinforcement learning, which can ensure the rationality of the early action strategy through multi-network prediction, parameter soft update, network reduction, occupy only a small amount of computing power, assist model training, improve the optimization performance of the model, and further guarantee the efficient and economic operation of the power system. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 The figure is a model structure diagram of the power distribution network reconstruction method of the application;

[0023] Figure 2 The figure is a topology connectivity detection diagram of the application;

[0024] Figure 3 The figure is a reward curve diagram of the method of the application and other comparative algorithms;

[0025] Figure 4 The figure is a network loss comparison diagram before and after the reconstruction of the method of the application. DETAILED DESCRIPTION

[0026] The technical solutions of the application will be further described below with reference to the drawings.

[0027] Reference Figure 1 The application provides a power distribution network reconstruction method based on topology safety constraints and integrated reinforcement learning, which includes the following steps:

[0028] Step 1: power grid environment modeling, including system state quantity design, reward function design, power grid constraint embedding and system steady-state calculation;

[0029] Step 2: Construct a topological security layer, correct the reconstruction action of the topology using a topological mask through a detection algorithm;

[0030] Step 3: Introduce an integrated action network, a parameter doping mechanism, and a network pruning mechanism to pre-judge different actions, while ensuring calculation speed and improving the stability of model training;

[0031] Step 4: Assemble a reinforcement learning training framework to train the initial policy network and obtain network parameters;

[0032] Step 5: Use the trained policy network to obtain the scheduling strategy of the power distribution network reconstruction according to the state data of the system.

[0033] The specific implementation process of the power distribution network reconstruction optimization method will be described in detail below with specific examples. The present application takes energy storage and topological coordinated scheduling as the optimization scenario, and selects the IEEE 33-node system as the benchmark case for simulation experiment. In this system, 16 nodes are connected to distributed photovoltaic to simulate the uncertainty of new energy, and 3 energy storage devices and 12 branch switches are configured to ensure the stable operation of the system. The photovoltaic power generation data uses typical summer photovoltaic scene data, and on this basis, proportional noise and bias noise are superimposed to meet the data size requirements of the reinforcement learning training process. The distribution characteristics of these noises are fitted according to the measured photovoltaic data of a certain region in Belgium to ensure that the simulated environment can more realistically reflect the actual operating conditions. Based on the constructed model data, the specific implementation steps of the method are as follows:

[0034] Step (1), build a power distribution network environment for interactive training.

[0035] The power distribution network environment is a virtual simulation environment for interacting with the reinforcement learning model to obtain training data. The input policy network gives the reconstruction action, calculates the reward obtained by the current system state data through the reward function, and returns the reward and the next state data of the power grid environment.

[0036] In general, the environment is a modeling of the task, and the reinforcement learning algorithm (the policy network of the agent) is a tool to solve the task. Unlike traditional algorithms such as using a second-order cone to optimize the power grid flow, then using a stochastic robust algorithm to solve the scheduling optimization problem, the present application first constructs a power distribution network environment, and according to the differences in solving tasks when the power distribution network scheduling optimization task is combined with the actual, it is transformed into a Markov decision environment. Similar reinforcement learning algorithms have better optimization performance after training, such as having a highly condensed state space and a clearly directed reward function.

[0037] The design of the environment includes a reward function (directly determines the orientation of the task), a state (characteristic information of the task decision), and an action (an optimized action space of the intelligent agent). According to an embodiment of the present application, the construction of the power distribution network system environment can be divided into the following steps:

[0038] First, the interactive variables such as the reward function, the state, and the action are designed according to the different tasks and environments. The state data reflects the essential characteristics of the power system, and the reward function determines the optimization direction and task of the strategy model training. The power distribution network reconstruction task set in the present method minimizes the network loss and the number of line switch actions by adjusting the network structure of the system and the load-side resources such as energy storage, and the calculation method is as follows:

[0039]

[0040] wherein S is the system state, A topo,t-1 is the topology action at the previous moment, E t-1 is the energy storage capacity state at the previous moment, P is the node power at the previous period, R is the reward function, k price is the electricity purchase economic coefficient, P ess,j,t is the charge and discharge power of the jth energy storage at the current moment, c pur is the current electricity price, c pur,ess is the average electricity price of the current energy storage, k loss is the network loss coefficient, P loss is the system network loss after the action is performed. k switch is the switch operation coefficient, N switch is the number of switch actions, and A is the strategy network action. A topo and P ess are the topology and energy storage actions, respectively. According to graph theory, the necessary and sufficient condition for realizing the conversion of a ring network to a radial topology is that at least one branch switch is opened in each independent loop. Therefore, the topology reconstruction action of the model is designed to selectively disconnect a controllable branch in different independent loops, L is the set of independent loop numbers, ε switch,l is the set of all controllable branch numbers in the lth independent loop, a l is the discrete action given for the lth independent loop, c l,i is the corresponding action code, and 0 represents connection and 1 represents disconnection.

[0041] Secondly, the physical constraints are incorporated into the power distribution network environment. The power grid environment includes topology constraints, power flow constraints, and energy storage capacity constraints. In addition to reasonably limiting the number of switch actions, the network structure adopted by the system must strictly satisfy the radial network topology, and the system steady state is obtained by solving the power flow equation. The calculation formula of the constraint is as follows:

[0042]

[0043] where G is the total number of actions of all switches in the dispatch cycle, G t,m is the total number of actions of switch m in the dispatch cycle; H m,t is the state of switch m at the t-th moment, 0 indicates on, and 1 indicates off; G max is the upper limit of the total number of actions of all switches in the dispatch cycle, G max,m is the upper limit of the total number of actions of switch m; P i,t , Q i,t are the active power and reactive power injected by the node respectively, G i,j , B i,j are the branch conductance and branch susceptance between nodes ij respectively, V i,t is the voltage of node i at the t-th moment, V min , V max are the upper and lower limits of the node voltage; P ess,min,j , P ess,max,j are the upper and lower limits of the charge and discharge power of the energy storage, η c,j , η d,j are the charge and discharge coefficients of the j-th energy storage, E j,t is the capacity of the j-th energy storage at the t-th moment, E max,j , E min,j are the upper and lower limits of the capacity of the corresponding energy storage, δ SOC,t is the state of charge of the energy storage at the current moment.

[0044] Step (2), constructing a topology safety layer based on safety constraints.

[0045] The present application introduces a topology safety layer for modifying the topology action, which has no effect on the energy storage. This is because the energy storage action can be limited by the limit function (such as the clamp function) according to the input current capacity of the energy storage, which is relatively simple and therefore not the focus of the present application.

[0046] Firstly, the topology mask is introduced to correct the initial topology action (also called action mask) output by the action network. The calculation process of the generation and use of the topology mask is as follows:

[0047]

[0048] wherein, m l is the mask vector of loop l, m l,i is the mask of the i-th controllable branch under the loop, are the original data and the data after action masking (i.e. topology state, [0, 1] indicates on and off) of the i-th controllable branch under the loop respectively; is the original output of the action network, is the action output after masking, which is randomly sampled using the reparameterization trick, gi N is the number of distribution noise to ensure the consistency of the action probability switch,l c is the number of loop branches l,i is the final action code. Here, the reparameterization trick is used to randomly sample the masked action, in order to optimize the action distribution. Assuming that the masked action A has four dimensions, corresponding to four action options, the maximum value is usually selected as the executed action after softmax, but the max() function is gradient-discontinuous, such as 0.4, 0.2, 0.2, 0.2, which will actually always execute the first action, and after the next update, it becomes 0.5, 0.1, 0.2, 0.2, which is still no difference, and is not continuous. The reparameterization trick randomly generates distribution noise (normal) superimposed on the value of the action A, so that the four action distributions are continuous, facilitating training.

[0049] Then, the above specific mask element m l is obtained through a special topological connectivity detection algorithm. The detection algorithm can effectively obtain the constraint relationship between the masks and reduce the calculation amount of detection. The tree node is defined as: a node connected to no more than one non-tree branch; the tree branch is defined as: a branch containing a tree node; in the initial state, all branches are non-tree branches. Referring to Figure 2 , the yellow points are tree nodes, and there are no tree nodes and tree branches at the beginning, and all are non-tree nodes and non-tree branches. After disconnecting a branch, the two ends of the disconnected branch are detected, and the number of non-tree branches connected to a node becomes one, so the node becomes a tree node, and the connection branch of the newly generated tree node becomes a tree branch. The other node of the new tree branch is detected, and it is found that the number of non-tree branches is 2, and the process is ended. Therefore, the above definition is a dynamic and perfect marking process. According to graph theory, it can be understood that the tree branch cannot be disconnected, and once it is disconnected, an island will be formed. The non-tree branch can be disconnected at will, so the code of the tree branch is the mask element. If a branch is a tree branch during the current operation, a negative infinite mask element needs to be superimposed on the branch to make it small enough not to be selected. In other words, the topological monitoring algorithm of the application regards the process of topological generation as the process of separating the tree topology from the original closed topology. Once the tree node and the tree branch are formed, they are included in the final tree topology, and the disconnection of any tree branch will definitely form an island. This means that there is no need to repeatedly judge the connectivity of the tree branch, and under the condition of no island, a closed loop topology will not be formed. Under this mechanism, the tree topology constraint of the grid is further equivalent to the connectivity constraint of the tree branch, and the action mask only needs to ensure that the tree branch is not disconnected by the subsequent loop. By dynamically updating the category of all branches, all mask elements under the corresponding loop can be obtained at one time without repetition. The principle of the topological connectivity detection algorithm is shown in Figure 2 .

[0050] Finally, in order to further improve the calculation speed of the mask action, a hash table is used to store the action mask.

[0051]

[0052] Wherein, N' switch is the total number of switches that can be operated before the current loop, n switch is the total number of effective topologies that can be generated by the branch switch, the matrix B stores all effective states in one-hot encoding format, and the matrix M contains corresponding mask elements.b i,k is the topology code of the current operation, h mask (b i ) is a hash function, and the hash table is retrieved by direct addressing to obtain the mask element mask(b i ).

[0053] Step (3), pre-judging the initial scheduling strategy based on the integrated action network.

[0054] The application avoids the influence of a large number of inferior actions on the training of the model by pre-judging and screening the initial action through integrating multiple action networks. Different initialization parameters are adopted for different action networks, the state is input into the action network group to obtain different scheduling schemes, and the optimal action is selected as the actual scheduling strategy and input into the power grid environment. For each action network, in order to maintain the uniqueness and diversity of the parameters, a period counter is set to count the number of optimal actions of each action network in a period T best , the parameters of the optimal action network are transferred to other action networks in a soft update manner, and the parameters are disturbed. At the same time, since the parameter update only depends on the output of the value network for gradient ascent optimization, the probability of inferior action appearing is significantly reduced as the training proceeds, so the optimal action frequency of each action network is recorded in a longer period T cut , and the optimal action network is deleted from the action network group, and the process is repeated until it degenerates into the original algorithm. The soft update calculation process is represented as:

[0055] φ others = τ best φ ibest +(1-τ best )φ others

[0056] Wherein, φ others is the parameter of other networks excluding the optimal network, φ ibest is the optimal network parameter, and τ best is the corresponding soft update parameter. That is, the model parameter with subscript best is mixed with the model parameter with subscript other at a small proportion τ, which is different from the parameter update using gradient, and is commonly known as soft update in Chinese.

[0057] Step (4), parameter training is performed on the model using a double-delay-determined policy gradient framework, and a reconstructed scheduling policy is output using the trained model.

[0058] The double-delay-determined policy gradient algorithm is a reinforcement learning algorithm under the Actor-Critic framework, which is composed of an Actor network, i.e., a policy network, and a Critic network, i.e., a value network. The Actor network outputs actions according to the state, and the Critic network inputs actions and states to evaluate the actions given by the policy network to obtain the value of the actions. On this basis, the framework uses two Critic networks, takes the smaller one when calculating the target value, and suppresses the overestimation problem of the network. At the same time, a delayed update method is used, and the Actor network is updated after the Critic network is updated multiple times, which ensures the stability of the training of the policy network. The calculation process is as follows:

[0059]

[0060] wherein Q(·) is the value network, γ is the discount coefficient, y is the target value, s' and a' are the state and action at the next time, s and a are the state and action at the current time. θ is the value network parameter, i.e., the parameter to be trained, and N is the batch size of training. φ is the parameter of the policy network, and φ represents the gradient of the policy network, so the parameter update is obtained by the gradient.

[0061] The training process is fine-tuned in the Actor-Critic framework, and different training frequencies of the policy network and the value network are used to achieve double-delay, i.e., fast and slow, to ensure the stability of the training of the policy network.

[0062] The reconstructed scheduling policy is verified using the trained model output, and the state data S is input, and the model gives the reconstructed scheduling policy, i.e., the topology and energy storage action.

[0063] In order to verify the performance of the method, the actual optimization effect of the model is evaluated based on the built optimization scene.

[0064] Figure 3The effect comparison of the method of the application and other methods in the same scene is shown, the solid line represents the reward mean, and the shaded area is the upper and lower limits of the reward fluctuation. It can be observed that the training efficiency of the method of the application is obviously higher than that of other methods in the early stage, because the cluster strategy is used to perform prior screening on the action, so that the action is always kept in a relatively optimal interval. At the same time, the model of the application tends to converge after 10,000 interactions, which is prior to other comparison models. Therefore, although the method of the application consumes more computing power in the early stage, the training speed is also improved. It is noted that the TD3 algorithm and the DDPG algorithm achieve similar effects to the improved algorithm of the application in some periods of some experiments, but the fluctuation of the training reward value is more obvious, and the final reward is less than the reward of the application.

[0065] Figure 4 In order to correspond to the change of the reconstructed system network loss, it can be seen that by changing the system topology at different times, the transmission path of power at different time scenes can be reasonably allocated, and the line loss of the distribution network system is effectively reduced, and the daily total network loss reduction amount reaches 5.7%.

[0066] In summary, the method of the application proposes a reinforcement learning method based on a cluster strategy to solve the reconstruction optimization scheduling strategy of a power system. The effectiveness of the proposed method is verified through example analysis.

[0067] Based on the same technical concept as the method embodiment, the application also provides a distribution network reconstruction system based on topology safety constraints and integrated reinforcement learning, comprising:

[0068] A power grid environment modeling module is used to construct a reinforcement learning model for a distribution network reconstruction task based on a power grid environment, including system state quantity design, reward function design, power grid constraint embedding and system steady-state quantity calculation. The purpose of the reinforcement learning model is to minimize the network loss and the number of actions of the line switch by adjusting the network structure and the energy storage resource of the system. The policy network of the model gives a reconstruction action according to the current state data of the power grid environment, calculates the reward of the action according to the reward function, and returns the reward and the next state data of the power grid environment. The state data includes the topology action at the previous time on the network, the energy storage capacity state at the previous time, and the node power at the previous period. The reconstruction action includes a topology action and an energy storage action, wherein the topology reconstruction action is designed to selectively disconnect a controllable branch in different independent loops.

[0069] A topology safety layer construction module is used to construct a topology safety layer, which is used to detect and correct the topology reconstruction action calculated by the model initially by using a topology mask through a topology detection algorithm.

[0070] The action combination module is configured to combine the plurality of action networks into an integrated action group, which is used to predict different reconstruction actions through a parameter doping mechanism and a network pruning mechanism.

[0071] The network training module is configured to train the initial policy network based on a reinforcement learning training framework to obtain network parameters for generating a reconstruction scheduling policy.

[0072] The network application module is configured to use the trained policy network to obtain a scheduling policy for power distribution network reconstruction according to state data of the system.

[0073] It should be understood that the power distribution network reconstruction system based on topology safety constraints and integrated reinforcement learning in the embodiments of the present application can realize all the technical solutions in the above method embodiments, and the functions of each functional module can be realized according to the methods in the above method embodiments, and the specific implementation process can be referred to the related description in the above embodiments, which will not be described here.

[0074] The present application also provides a computer device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the program is executed by the processor to realize the steps of the power distribution network reconstruction method based on topology safety constraints and integrated reinforcement learning as described above.

[0075] The present application also provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by the processor to realize the steps of the power distribution network reconstruction method based on topology safety constraints and integrated reinforcement learning as described above.

[0076] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, device (system), computer device or computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0077] The present application is described with reference to flowcharts according to the method of the embodiments of the present application. It should be understood that each flow in the flowchart and the combination of the flows in the flowchart can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in the flowchart Figure 1an apparatus that performs the function specified in the flow or flows.

[0078] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow Figure 1 the function specified in the flow or flows.

[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 the function specified in the flow or flows.

Claims

1. A power distribution network reconfiguration method based on topological security constraints and integrated reinforcement learning, characterized in that, Comprising the following steps: Step 1: Constructing a power grid environment for interacting with a reinforcement learning model to obtain training data, including system state quantity design, reward function design, power grid constraint embedding, and system steady-state quantity calculation, the purpose of the reinforcement learning model is to minimize the network loss and the number of line switch actions by adjusting the network structure and energy storage resources of the system; the policy network of the model gives the reconstruction action according to the current state data of the power grid environment, and calculates the reward of this action according to the reward function, and returns the reward and the next state data of the power grid environment; the state data includes the topology action at the last time on the network, the energy storage capacity state at the last time, and the node power in the last period, and the reconstruction action includes topology action and energy storage action, wherein the topology reconstruction action is designed to selectively disconnect a controllable branch in different independent loops; Step 2: Constructing a topology safety layer, which detects and corrects the topology reconstruction action calculated by the model using an action mask through a topology detection algorithm, defines a tree node as a node with no more than one linked non-tree branch; define a tree branch as a branch containing a tree node; all branches are non-tree branches in the initial state; the topology detection algorithm regards the topology generation process as separating a tree topology from the original closed topology, once the tree node and the tree branch are formed, they are included in the final tree topology, and the code of the tree branch is the mask element; when it is detected that the branch to be operated is a tree branch, a negative infinite mask element is added to the branch, so that it will not be selected; Step 3: Combining multiple action networks into an integrated action group, and pre-judging and screening different reconstruction actions of the integrated action group through a parameter doping mechanism and a network pruning mechanism, the parameter doping mechanism sets a period counter for each action network to count the optimal action frequency of each action network in the first period, and transfers the parameters of the optimal action network to other action networks for parameter disturbance; the network pruning mechanism records the optimal action frequency of each action network in the second period, and deletes it from the action network group; Step 4: Training the reinforcement learning model using a double-delay-determined policy gradient framework to obtain network parameters for generating reconstruction scheduling policies; Step 5: Using the trained reinforcement learning model to obtain the scheduling policy of the distribution network reconstruction according to the state data of the current system.

2. The method of claim 1, wherein, The variable design of the power grid environment is as follows: ; in, For system status, This refers to the topological action from the previous time step. This represents the energy storage capacity status at the previous moment. This represents the node power in the previous time period. For the reward function, For the economic coefficient of electricity purchase, For the current moment, the first The charging and discharging power of an energy storage device. At the current electricity price, This represents the current average electricity price for energy storage. This is the network loss coefficient. The system network loss after the action is executed. This is the switching operation coefficient. The number of switch actions. For policy network actions, , These are topology and energy storage actions, respectively. For the set of independent loop numbers, for The set of all controllable branch numbers under the independent loop number. In response to Discrete actions given by independent loops. The corresponding action code is 0 for connection and 1 for disconnection.

3. The method of claim 2, wherein, The power grid environment contains topology constraints, power flow constraints, and energy storage capacity constraints, and the system steady-state is obtained by solving the power flow equation, and the calculation formula of the constraint is: ; in, This represents the total number of times all switches operate within the scheduling cycle. Switching during the scheduling cycle The total number of actions; For the first Switch at any time Status, 0 indicates enabled, 1 indicates disabled; This represents the maximum number of all switching actions within the scheduling period. For switch The maximum number of times; , These represent the active power and reactive power injected into the nodes, respectively. , They are respectively Branch conductance and branch susceptance between nodes for At this moment Node voltage, , These are the upper and lower limits of the node voltage, respectively; , These represent the upper and lower limits of energy storage charging and discharging power, respectively. , The first The charge / discharge coefficient of an energy storage device. For the first At this moment The capacity of the energy storage, , These correspond to the upper and lower limits of energy storage capacity. This represents the state of charge of the stored energy at the current moment.

4. The method of claim 1, wherein, The calculation process of the generation and use of the action mask is represented as: ; wherein, in the formula is the mask vector of the loop l, is the mask of the i-th controllable branch of the loop, , are the original data and the action-masked data of the i-th controllable branch of the loop, respectively; is the original output of the action network, is the masked action output, which is randomly sampled using the reparameterization trick, is the distributed noise used to ensure the consistency of the action probability, is the number of branches of the loop, is the final action encoding.

5. The method of claim 4, wherein, The action mask is stored in a hash table, and the storage and reading process of the hash table is represented as: ; wherein, is the total number of switches operable before the current loop, is the total number of valid topologies that the branch switch can generate, the matrix stores all valid states in one-hot encoding format, the matrix contains the corresponding mask element, is the topology encoding of the current operation, is a hash function, the hash table is retrieved by direct addressing to obtain the mask element .

6. The method of claim 1, wherein, The double-delay-determined policy gradient framework trains the reinforcement learning model, and the calculation process is: ; where, is the value network, is the discount factor, is the target value, , are the state and action at the next time step, respectively, , are the state and action at the current time step, is the value network parameter, i.e., the parameter to be trained, is the number of batches for training, is the parameter of the policy network, denotes the gradient of the policy network.

7. A power distribution network reconfiguration system based on topological security constraints and integrated reinforcement learning, characterized in that, Comprising: A power grid environment modeling module is configured to construct a power grid environment for obtaining training data by interacting with a reinforcement learning model, including system state quantity design, reward function design, power grid constraint embedding, and system steady-state quantity calculation, the purpose of the reinforcement learning model being to minimize network loss and line switch action frequency by adjusting the network structure and energy storage resources of the system; a policy network of the model gives a reconstruction action according to current state data of the power grid environment, and calculates the reward of the action according to a reward function, and returns the reward and the next state data of the power grid environment; the state data includes topology action at the previous time on the network, energy storage capacity state at the previous time, and node power in the previous period, and the reconstruction action includes topology action and energy storage action, wherein the topology reconstruction action is designed to selectively disconnect a controllable branch in different independent loops in turn; A topology safety detection module is configured to construct a topology safety layer, which detects and corrects the topology reconstruction action calculated by the model by using an action mask through a topology detection algorithm, defines a tree node as a node with a linked non-tree branch being less than or equal to 1, and defines a tree branch as a branch containing a tree node; all branches are non-tree branches in the initial state; the topology detection algorithm regards the process of topology generation as the process of separating a tree topology from an original closed topology, and once the tree node and the tree branch are formed, they are included in the final tree topology, and the code of the tree branch is a mask element; when it is detected that the branch of the current operation is a tree branch, a negative infinite mask element is added to the branch, so that the branch will not be selected; An action network pre-screening module is configured to combine multiple action networks into an integrated action group, and pre-judge and screen different reconstruction actions of the integrated action group through a parameter doping mechanism and a network pruning mechanism; the parameter doping mechanism sets a period counter for each action network to count the optimal action frequency of each action network in a first period, and transfers the parameters of the optimal action network to other action networks to perturb the parameters thereof; the network pruning mechanism records the optimal action frequency of each action network in a second period, and deletes the action network from the action network group; A reinforcement learning model training module is configured to train the reinforcement learning model by using a double-delay-deterministic policy gradient framework to obtain network parameters for generating a reconstruction scheduling policy; A distribution network reconstruction strategy determination module is configured to use the trained reinforcement learning model to obtain a corresponding distribution network reconstruction scheduling policy according to the state data of the current system.

8. A computer device, comprising: Comprise: One or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, the programs being executed by the processor to implement the steps of the distribution network reconstruction method based on topology safety constraints and integrated reinforcement learning of any one of claims 1-6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the distribution network reconstruction method based on topology safety constraints and integrated reinforcement learning of any one of claims 1-6.