A control method, system, device and storage medium for a distribution network topology

Through improved pointer network and AC reinforcement learning algorithm, the distribution network topology control model is constructed, which solves the real-time and stability problems of deep reinforcement learning in distribution network topology optimization, and realizes efficient distribution network topology control.

CN114154416BActive Publication Date: 2025-07-08CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111453132.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-07-08
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

The existing deep reinforcement learning technology cannot be effectively applied to distribution network topology optimization, mainly due to the topological changes that lead to environmental instability, the changes in action significance are difficult to evaluate, and traditional methods cannot meet the real-time requirements.

Method used

The improved pointer network and AC reinforcement learning algorithm are used to build a distribution network topology control model. By obtaining static and dynamic information, using AC reinforcement learning algorithm to train the model, generate the switch combination control information to achieve real-time control of the distribution network topology.

Benefits of technology

It realizes efficient and real-time control of distribution network topology, reduces training efficiency and model optimization complexity, and is suitable for neural network self-learning and end-to-end control strategy calculations of multiple types of failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114154416B_ABST
    Figure CN114154416B_ABST
Patent Text Reader

Abstract

The present invention discloses a control method, system, device and storage medium for a distribution network topology, including: acquiring static information and dynamic information of the distribution network topology; inputting the static information and dynamic information of the distribution network topology into a distribution network topology control model trained by using an AC reinforcement learning algorithm to obtain control information of a switch combination in the distribution network topology; and controlling the distribution network topology according to the control information of the switch combination in the distribution network topology, thereby completing the control of the distribution network topology. The method, system, device and storage medium can control the distribution network topology by using deep reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power system automation, and relates to a control method, system, device and storage medium for the topology of a distribution network. Background Art

[0002] In the study of topology optimization problems, it is often necessary to define the states of various control switches as binary variables of 0 and 1. The introduction of integers and the complexity of the system model have caused difficulties in the application of traditional optimization. Therefore, heuristic optimization algorithms such as improved particle swarm optimization and quantum artificial bee colony optimization are often used for solving by means of coding, mapping, etc. These methods are suitable for plan-based control optimization and cannot meet the real-time requirement of the solution. The emergence of artificial intelligence technology provides a new idea for the topology control of a distribution network, and aims to solve the real-time bottleneck of traditional optimization calculation by realizing the end-to-end decision-making from operation characteristics to network control strategies, with the deep reinforcement learning technology as the main research direction.

[0003] However, general deep reinforcement learning requires a strictly stable interaction environment, and the inability to clearly construct an optimization calculation of a Markov decision process makes it lose its advantages. These problems limit the application of this technology in the topology optimization of a distribution network. Firstly, the change of the topology makes the environment unstable, resulting in a change in the meaning of actions. Secondly, the operation mode of the line is a combined effect rather than a long-term benefit, and it is difficult to evaluate its specific value during the formation and change process. Therefore, only single-step decision-making modeling can be used, and thus deep reinforcement learning cannot be used to control the topology of a distribution network. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned disadvantages of the prior art, and provides a control method, system, device and storage medium for the topology of a distribution network, which can use deep reinforcement learning to control the topology of a distribution network.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] On the one hand, the present invention provides a control method for the topology of a distribution network, including:

[0007] Obtain the static information and dynamic information of the topology of the distribution network;

[0008] Input the static information and dynamic information of the topology of the distribution network into the distribution network topology control model trained by using the AC reinforcement learning algorithm, and obtain the control information of the switch combination in the topology of the distribution network;

[0009] Control the topology of the distribution network according to the control information of the switch combination in the topology of the distribution network, and complete the control of the topology of the distribution network.

[0010] A further improvement of the control method for the distribution network topology described in the present invention lies in:

[0011] Before inputting the static information of the distribution network topology into the distribution network topology control model trained by using the AC reinforcement learning algorithm, it further includes:

[0012] Construct a distribution network topology control model by using an improved pointer network;

[0013] Train the distribution network topology control model by using the AC reinforcement learning algorithm to obtain the trained distribution network topology control model.

[0014] The specific process of constructing the distribution network topology control model is as follows:

[0015] Use the improved pointer network to construct a distribution network topology control model based on the current limit values of each line in the distribution network topology, the voltage limit values of each node in the distribution network topology, and a preset objective function.

[0016] The reward function in the process of training the distribution network topology control model by using the improved pointer network and the AC reinforcement learning algorithm is:

[0017]

[0018] Among them, c1 is the target evaluation weight of reliability, c2 is the target evaluation weight of rapidity, G′ is the result of processing the distribution network topology to be controlled through the decision sequence, D is the number of load nodes in the distribution network, γ i is the energized state of the i-th load node in the distribution network, Ω is the decision element sequence, β j is a 0-1 variable representing the change in the switch state. Among them, when the switch state changes, then β j is taken as 1, otherwise, then β j is 0.

[0019] At the t-th moment, the state space in the improved pointer network is:

[0020]

[0021] Among them, s t is the static information of the distribution network topology at the t-th moment, d t is the dynamic information of the distribution network topology at the t-th moment, M is the number of controllable switches in the distribution network topology, s M is the static information of the M-th controllable switch, d t,M The static element position of the dynamic information of the M-th controllable switch.

[0022] At the t-th moment, the dynamic information d t of the distribution network topology is:

[0023]

[0024] Among them, m is the static element position of the dynamic information, and x m is the operation information of the static element of the dynamic information, G is the distribution network topology structure, * is the topology operation, and γ i is the energized state of the i-th load node in the distribution network, and Ω t is the decision element sequence at the t-th moment, D is the number of load nodes in the distribution network, and γ i (G) represents the energized state of the load nodes in topology G.

[0025] At the t-th moment, the calculation logic of the mask in the improved pointer network is as follows:

[0026] According to the mask matrix after the (t - 1)-th moment, all non-zero states are sequentially added to the decision sequence, and it is judged whether the constraint conditions are satisfied. When the constraint conditions are not satisfied, the mask value of the corresponding code is set to 0. Otherwise, the mask value of the state number selected at the t-th moment is set to 0, the mask value of the mutually exclusive state of the state selected at the t-th moment is set to 0, and then the mask value with the mask value of the corresponding code set to 0 is restored.

[0027] In a second aspect of the present invention, the present invention provides a control system for a distribution network topology, including:

[0028] An acquisition module, configured to acquire static information and dynamic information of the distribution network topology;

[0029] A calculation module, configured to input the static information and dynamic information of the distribution network topology into a distribution network topology control model trained by using an AC reinforcement learning algorithm, and obtain control information of the switch combination in the distribution network topology;

[0030] A control module, configured to control the distribution network topology according to the control information of the switch combination in the distribution network topology, and complete the control of the distribution network topology.

[0031] In a third aspect of the present invention, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the control method for the distribution network topology are implemented.

[0032] In a fourth aspect of the present invention, the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the control method for the distribution network topology are implemented.

[0033] The present invention has the following beneficial effects:

[0034] When the control method, system, device, and storage medium of the distribution network topology described in the present invention are specifically operated, the static information of the distribution network topology is input into the distribution network topology control model trained by using the AC reinforcement learning algorithm to obtain the control information of the switch combination in the distribution network topology. Among them, the distribution network topology control model is trained by using the AC reinforcement learning algorithm, thereby reducing the training efficiency and the complexity of model optimization, being applicable to the neural network self-learning for multiple types of faults and the training of the end-to-end control strategy calculation model. Then, according to the control information of the switch combination in the distribution network topology, the distribution network topology is controlled to achieve the purpose of controlling the distribution network topology by using deep reinforcement learning. The operation is convenient, simple, and highly practical. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0036] Figure 1 is the flowchart of the method of the present invention;

[0037] Figure 2 is the system structure diagram of the present invention;

[0038] Figure 3 is the structure diagram of the improved pointer network;

[0039] Figure 4 is the training flowchart of the distribution network topology control model.

[0040] Among them, 1 is the acquisition module, 2 is the calculation module, and 3 is the control module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0042] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0043] The present invention will be further described in detail below with reference to the accompanying drawings:

[0044] Embodiment 1

[0045] Reference Figure 1 , the control method for the distribution network topology described in the present invention includes:

[0046] 1) Construct a distribution network topology control model;

[0047] The specific process of step 1) is as follows:

[0048] For each switch in the distribution network topology, a complete set of feasible solutions is constructed using the number and the corresponding state variable. To ensure the unity of the state expression form, each normal controllable switch is expressed as an arrangement of a set of mutually exclusive elements, that is:

[0049]

[0050] Among them, the current state α of the switch N In the mutually exclusive state of the switch Before, for the faulty switch, its state has no mutually exclusive elements, and two identical elements are used to form a placeholder, that is:

[0051] {(N,0),(N,0)} (2)

[0052] Based on the above modeling method, there are a total of 2|V| elements describing the topological state in the complete set of decision elements. For each element x in the complete set of decision elements, its meaning change is an operation on the original topology G, that is, the switch with the corresponding number N is adjusted to the state α corresponding to the element, and the new topology G' of the distribution network topology is obtained as:

[0053] G' = x n *x n-1 *…*x0*G (3)

[0054] Among them, * represents the topological operation operation.

[0055] Let the line states before and after network topology decision-making be as follows:

[0056]

[0057] Then, with the goal of power supply reliability and rapidity, the constructed objective function is:

[0058]

[0059] Taking the current limits of distribution network line current and node voltage as constraint conditions, that is:

[0060] I jmin ≤I j ≤I jmax (7)

[0061] U imin ≤U i ≤U imax (8)

[0062] 2) Use the AC reinforcement learning algorithm to train the distribution network topology control model to obtain the trained distribution network topology control model;

[0063] The specific process of step 2) is as follows:

[0064] 21) The solution process of the distribution network topology control model can be described as: Based on the decision sequence loop at the current moment, find the element with the maximum output probability of the model, that is:

[0065]

[0066] X t+1 =f(X t ,y t+1 )(9)

[0067] where X t is the input of the decision state space information at time t, y is the decision output information of the model at time t, Y t ={y0,...,y t} is the decision output sequence completed by the model before time t, f is the state transition function for state space update. Since the decision requires T selections to complete, thus t = 1, 2,..., T.

[0068] As can be seen from Equation (9), the predictive variable y is always within the set Φ, and the elements in the set Φ directly constitute the state space X. To study this problem, it is necessary to directly establish a connection between the output and the input of the model. High-dimensional features are constructed through the reward signal and operation process variables feedback by the output strategy, and further a quantization probability model is obtained, that is, the self-mapping process of the input sequence is completed. The pointer network is a type of model for solving the combinatorial selection problem, and its core is to complete feature correspondence and probability calculation through the attention mechanism.

[0069] 22) The improved pointer network is as Figure 3 shown, Figure 1 The state sequence information in is first input into the encoder to implement sequence encoding. The purpose is to convert explicit information into high-dimensional feature vectors. The vector embedding implemented by the encoder is generally completed by convolutional or recurrent neural network structures. Considering the decision state features and data structure input to the encoder, in the present invention, a one-dimensional convolutional network structure is used as the encoder to extract the implicit topological features contained in the state space sequence; then the high-dimensional features output by the encoder are used as part of the input of the decoder, and the original state set is predicted in combination with the attention mechanism. The network topology combination is obtained by solving Equation (9) iteratively. The calculation method of the decoder is the same as that of the RNN network prediction, so the GRU unit is used as the core structure of the module.

[0070] Among them, the dynamic information embedding process in the improved pointer network is as follows:

[0071] The initial version of the pointer network only considers the input situation of the static state space, that is, X t is a fixed value. The improved pointer network constructs the state space by dividing it into static and dynamic parts, that is:

[0072]

[0073] Among them, s t represents static information, d t represents dynamic information, M is the number of controllable switches in the distribution network topology, s M is the static information of the Mth controllable switch, d t,M The static element position of the dynamic information of the Mth controllable switch. Embedding the dynamic information representing the number of changes in the distribution network power supply nodes caused by executing the current element in the model can accurately express the impact of each decision on the reliability of the overall solution. The dynamic information d t is:

[0074]

[0075] Among them, m is the static element position of the dynamic information, and x m is the operation information of the static element.

[0076] The design process of the mask in the improved pointer network is as follows:

[0077] Since the decision of the pointer network is realized by relying on the decision probability distribution calculated by the attention mechanism, reducing the corresponding probability to 0 can avoid the selection of elements. This processing method is called masking. The attention probability after adding the mask is:

[0078] a t = softmax(h t + log(λ t ))(12)

[0079] where λ t represents the mask vector at the current time t, and the value of each bit is either 0 or 1. When the value of a certain bit is 0, the probability of selecting the corresponding element is calculated as 0 and it will not be selected. The basic function of the mask is to control the pointer to select elements in the complete set without repetition, that is, after each prediction, the probability corresponding to the element number is set to zero.

[0080] Taking the model decision at time t as an example, the calculation logic of the mask is as follows:

[0081] 221) According to the mask matrix after time t - 1, sequentially add all non-zero states to the decision sequence and judge whether the constraint conditions are satisfied. When the constraint conditions are satisfied, go to step 22); otherwise, set the mask value of the corresponding number that does not meet the conditions to 0;

[0082] 222) Set the mask of the state number selected at time t to 0;

[0083] 223) Set the mask of the mutually exclusive state of the state selected at time t to 0;

[0084] 224) Restore the mask value set to 0 in step 21);

[0085] 225) When t is equal to the limit number of rounds T, set the mask matrix to all 0; otherwise, save the mask matrix and use it for the decision at step t + 1.

[0086] In addition, in order to represent the initial topological information and the number of operation features, a 0-1 variable β representing the change of the switch state is added to the static state elements. Its value is 1 when the corresponding numbered switch changes, and vice versa, that is, the static state space Ω is a set composed of half of the elements selected from the complete set, which can be expressed as:

[0087] Ω = {(N, β N , α N ) | N = 0...n} (13)

[0088] After adding the β variable, the operation attributes of the elements are not affected, and the element operation is still represented by *. Since the action space is not explicitly defined, its essence is the selection of decision elements, which is expressed by the attention mechanism probability model in the model.

[0089] 23) During the training process, the reward function indirectly expresses the value of the objective function. Since illegal items are masked in the decision-making stage by means of a mask, there is no need to set constraint penalty terms. The reward value mainly judges the satisfaction of the load of each node and supplements it with the statistics of topological operations. The constructed combined evaluation value is:

[0090]

[0091] Among them, c1 represents the target evaluation weight of reliability; c2 represents the target evaluation weight of rapidity. The lower the evaluation value of the reward function, the more the combined strategy scheme meets the optimization requirements. The gradient descent update method is adopted in the training.

[0092] 24) During the training process, the parameters of the neural network are updated by the AC reinforcement learning algorithm. Specifically:

[0093] As Figure 4 shown, as a classic ACTOR-CRITIC architecture algorithm, the ACTOR network itself is a pointer network. The output of the ACTOR network maintains the decision element sequence Ω and the set of logarithmic probability values Π corresponding to the sequence. p , the evaluation function value R of the combined strategy is calculated according to Ω, and the probability of selecting this combination is calculated according to Π p . The ACTOR network is updated by the policy gradient with a baseline, and its purpose is to suppress the variance of network updates. Its objective function is expressed as:

[0094]

[0095] Calculate the policy gradient of the objective function, and then use the baseline correction, that is:

[0096]

[0097] Among them, θ is the network parameter of the pointer network, π is the decision-making strategy, p θ is the probability distribution output by the pointer network according to X, X is the decision state space, μ is the network parameter of the CRITIC network, and b μ is the output result of the CRITIC network. The prediction result of the baseline is calculated by the CRITIC network. According to the static information and dynamic information corresponding to the complete set of the state space, the decoder outputs to judge the decision-making difficulty and the expected evaluation value, and fits and returns the evaluation value of the combined strategy. Its objective function is:

[0098] J μ(X) = E[(R(π|X) - b μ (X)) 2 (17)

[0099] 3) Obtain the static information and dynamic information of the distribution network topology;

[0100] 4) Input the static information and dynamic information of the distribution network topology into the trained distribution network topology control model to obtain the control information of the switch combination in the distribution network topology;

[0101] 5) Control the distribution network topology according to the control information of the switch combination in the distribution network topology to complete the control of the distribution network topology.

[0102] Embodiment 2

[0103] Reference Figure 2 , the control system of the distribution network topology described in the present invention includes:

[0104] An acquisition module 1 for acquiring the static information and dynamic information of the distribution network topology;

[0105] A calculation module 2 for inputting the static information and dynamic information of the distribution network topology into the distribution network topology control model trained by using the AC reinforcement learning algorithm to obtain the control information of the switch combination in the distribution network topology;

[0106] A control module 3 for controlling the distribution network topology according to the control information of the switch combination in the distribution network topology to complete the control of the distribution network topology.

[0107] Embodiment 3

[0108] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the control method of the distribution network topology. Among them, the memory may include a memory, such as a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk memory, etc.; the processor, network interface, and memory are interconnected through an internal bus, and this internal bus can be an Industry Standard Architecture bus, a Peripheral Component Interconnect standard bus, an Extended Industry Standard Architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0109] Embodiment 4

[0110] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the control method for the distribution network topology. Specifically, the computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include read-only memory (ROM), hard disk, flash memory, optical disc, magnetic disk, etc.

[0111] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0112] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0113] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1Steps of the functions specified in one or more boxes.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific implementation manners of the present invention, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A control method for a distribution network topology, characterized in that Including: Obtaining static information and dynamic information of the distribution network topology; Inputting the static information and dynamic information of the distribution network topology into the distribution network topology control model trained by using the AC reinforcement learning algorithm to obtain the control information of the switch combination in the distribution network topology; Controlling the distribution network topology according to the control information of the switch combination in the distribution network topology to complete the control of the distribution network topology; Before inputting the static information and dynamic information of the distribution network topology into the distribution network topology control model trained by using the AC reinforcement learning algorithm, it further includes: Constructing a distribution network topology control model by using an improved pointer network; Training the distribution network topology control model by using the AC reinforcement learning algorithm to obtain the trained distribution network topology control model; The reward function in the process of training the distribution network topology control model by using the AC reinforcement learning algorithm is: Among them, c1 is the target evaluation weight of reliability, c2 is the target evaluation weight of rapidity, G′ is the result of processing the to-be-controlled distribution network topology through the decision sequence, D is the number of load nodes in the distribution network, γ i is the energized state of the i-th load node in the distribution network, Ω is the decision element sequence, β j is a 0-1 variable representing the change in the switch state. Among them, when the switch state changes, then β j is taken as 1, otherwise, β j is 0; During the training process, updating the parameters of the neural network by using the AC reinforcement learning algorithm, specifically: Updating the ACTOR network by using the policy gradient with a baseline, the purpose of which is to suppress the variance of network updates, and its objective function is expressed as: Calculating the policy gradient of the objective function and then using the baseline correction, that is: R(π|X) = R(G′,Ω) where θ is the network parameter of the pointer network, π is the decision-making strategy, p θ is the probability distribution output by the pointer network according to X, X is the decision-making state space, μ is the network parameter of the CRITIC network, b μ is the output result of the CRITIC network. The prediction result of the baseline is calculated by the CRITIC network. According to the static information and dynamic information corresponding to the complete set of state spaces, the decoder outputs the judgment of the decision-making difficulty and the expected evaluation value, and fits and returns the evaluation value of the combined strategy. Its objective function is: J μ (X) = E[(R(π|X) - b μ (X)) 2 .

2. The control method of the distribution network topology according to claim 1, characterized in that The specific process of constructing the distribution network topology control model by using the improved pointer network is: Using the improved pointer network to construct a distribution network topology control model based on the current limit of each line current in the distribution network topology, the voltage limit of each node in the distribution network topology, and the preset objective function.

3. The control method of the distribution network topology according to claim 1, characterized in that At the t-th moment, the state space in the improved pointer network is: Among them, s t is the static information of the distribution network topology at the t-th moment, d t is the dynamic information of the distribution network topology at the t-th moment, M is the number of controllable switches in the distribution network topology, s M is the static information of the M-th controllable switch, d t,M The static element position of the dynamic information of the M-th controllable switch.

4. The control method of the distribution network topology according to claim 3, wherein, At the $t$-th moment, the dynamic information $d$ of the distribution network topology t is as follows: Among them, m is the static element position of the dynamic information, and x m is the operation information of the static element of the dynamic information, G is the distribution network topology structure, * is the topology operation, and γ i is the energized state of the i-th load node in the distribution network, and Ω t is the decision element sequence at the t-th moment, D is the number of load nodes in the distribution network, and γ i (G) represents the energized state of the load nodes in the topology G.

5. The control method for the distribution network topology according to claim 1, wherein At the t-th moment, the calculation logic of the mask in the improved pointer network is: According to the mask matrix after the (t - 1)-th moment, successively add all non-zero states to the decision sequence and judge whether the constraint conditions are satisfied. When the constraint conditions are not satisfied, set the mask value of the corresponding encoding to 0. Otherwise, set the mask value of the state number selected at the t-th moment to 0, set the mask value of the mutually exclusive state of the state selected at the t-th moment to 0, and then restore the mask value whose mask value of the corresponding encoding is set to 0.

6. A control system for a distribution network topology, characterized in that, Including: An acquisition module (1) for obtaining static information and dynamic information of the distribution network topology; A calculation module (2) for inputting the static information and dynamic information of the distribution network topology into the distribution network topology control model trained by using the AC reinforcement learning algorithm to obtain the control information of the switch combination in the distribution network topology; A control module (3) for controlling the distribution network topology according to the control information of the switch combination in the distribution network topology to complete the control of the distribution network topology; Before inputting the static information of the distribution network topology into the distribution network topology control model trained by using the AC reinforcement learning algorithm, it further includes: Constructing a distribution network topology control model by using an improved pointer network; Training the distribution network topology control model by using the AC reinforcement learning algorithm to obtain the trained distribution network topology control model; The reward function in the process of training the distribution network topology control model by using the AC reinforcement learning algorithm is: Among them, c1 is the target evaluation weight of reliability, c2 is the target evaluation weight of rapidity, G′ is the result of processing the to-be-controlled distribution network topology through the decision sequence, D is the number of load nodes in the distribution network, γ i is the energized state of the i-th load node in the distribution network, Ω is the decision element sequence, β j is a 0-1 variable representing the change in the switch state. Among them, when the switch state changes, then β j is taken as 1, otherwise, β j is 0; During the training process, updating the parameters of the neural network by using the AC reinforcement learning algorithm, specifically: The ACTOR network is updated using a policy gradient with a baseline, the purpose of which is to suppress the variance of network updates, and its objective function is expressed as: Calculate the policy gradient of the objective function and then use baseline correction, that is: R(π|X) = R(G′,Ω) where θ are the network parameters of the pointer network, π is the decision-making strategy, p θ is the probability distribution output by the pointer network according to X, X is the decision-making state space, μ are the network parameters of the CRITIC network, b μ is the output result of the CRITIC network. The prediction result of the baseline is calculated by the CRITIC network. According to the static information and dynamic information corresponding to the complete set of state spaces, the decoder outputs the judgment of the decision-making difficulty and the expected evaluation value, and fits and returns the evaluation value of the combined strategy. Its objective function is: J μ (X) = E[(R(π|X) - b μ (X)) 2 .

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the control method for the distribution network topology according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the control method for the distribution network topology according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Power distribution network reconstruction method based on deep reinforcement learning algorithm and source load uncertainty

    CN112488442A