A method, system and related equipment for selecting and pulling single-phase ground fault lines in distribution networks based on deep reinforcement learning

Through the deep reinforcement learning model combined with dynamic graph convolutional neural network, a single-phase grounding fault line selection and pulling decision-making method is constructed in the distribution network, which solves the problem of insufficient fault analysis in the existing technology, realizes rapid identification and efficient decision-making of fault lines, and improves the safety and reliability of the distribution network.

CN120296542BActive Publication Date: 2025-08-26JINCHENG POWER SUPPLY COMPANY OF STATE GRID SHANXI ELECTRIC POWER
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510754653.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-26
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing single-phase grounding fault line selection and pulling method of distribution network fails to effectively combine the line operation conditions, the type of load supplied and the line selection and pulling order table, resulting in inaccurate fault analysis and affecting the reliability of auxiliary decision-making.

Method used

The deep reinforcement learning model is adopted to build the distribution network feeder switch feature matrix and adjacency matrix, combined with the dynamic graph convolution neural network, and generate a fault line selection and pulling strategy to achieve accurate identification and decision-making of fault lines.

Benefits of technology

It improves the accuracy and rapid recovery capabilities of faulty line selection, realizes the efficient integration of expert experience and technical analysis, and provides guarantees for safe, reliable and scientific scheduling and operation of the distribution network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296542B_ABST
    Figure CN120296542B_ABST
Patent Text Reader

Abstract

The present application provides a distribution network single-phase grounding fault line selection decision method, system and related equipment based on deep reinforcement learning, wherein the method includes: constructing a distribution network feeder switch adjacency matrix and a distribution network feeder switch feature matrix, using a dynamic graph convolutional neural network to process the distribution network feeder switch adjacency matrix and the distribution network feeder switch feature matrix to obtain state information feature quantities; generating a fault line selection strategy based on the fault line selection order; based on the state information feature quantities and the fault line selection strategy, constructing a distribution network single-phase grounding fault line selection decision model based on deep reinforcement learning, which is used to implement fault line selection analysis and determine the distribution network single-phase grounding fault line selection; the method provides scientific and efficient auxiliary decision-making, realizes the efficient integration of expert experience and technical analysis, and improves the distribution network single-phase grounding fault line selection efficiency and the distribution network recovery capability; and is applicable to the field of power system research technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of power system research, and specifically to a method, system and related equipment for selecting and making decisions on single-phase grounding fault lines in a distribution network based on deep reinforcement learning. Background Art

[0002] The distribution network is a crucial component of the power system, transporting electricity from power plants to end users, providing power to various electrical devices. As the final link in the power system, it converts electricity transmitted by high-voltage transmission lines into low-voltage power suitable for household, industrial, and commercial use. However, over long-term operation, the distribution network faces various potential fault risks, such as ground faults, busbar faults, and distribution line faults. These faults can cause power outages, equipment damage, and even fires. Therefore, timely and effective analysis and resolution of distribution network faults are crucial to ensuring safe and stable grid operation.

[0003] The selection of fault lines for single-phase grounding faults in distribution networks is a key research topic in their construction and development. From traditional methods based on manual experience and expert knowledge bases to strategies employing multivariate information fusion and deep learning, fault line selection methods are gradually shifting from manual experience to intelligent analysis and selection. The selection of fault lines for single-phase grounding faults in distribution networks not only relies on scientific, rational, and reliable fault line selection strategies but is also closely related to the varying operating conditions of the lines. Currently, the selection of fault lines for single-phase grounding faults in distribution networks primarily utilizes a fault line selection mechanism. This mechanism utilizes a comprehensive analysis of various electrical quantity information on distribution network lines to perform a probabilistic statistical analysis of the fault line. This analysis then provides decision support to grid dispatchers based on the failure probability of each line. When selecting fault lines for single-phase grounding faults in distribution networks, grid dispatchers not only consider the fault analysis results provided by the selection mechanism but also comprehensively consider the line's operating conditions, the type of load being supplied, and the line selection ranking table.

[0004] Common line selection methods include fault electrical quantity information analysis, multi-information fusion line selection, and line selection based on deep learning. Fault electrical quantity information analysis primarily analyzes and selects faulty lines based on steady-state or transient fault information. Multi-information fusion line selection typically uses the post-fault characteristics of multiple types of electrical quantity information to effectively identify the faulty line. It also employs intelligent information fusion methods such as DS evidence theory, genetic algorithms, and fuzzy control theory for comprehensive decision-making and judgment. Line selection mechanisms based on deep learning employ alternative intelligent information fusion methods such as artificial neural networks and deep belief networks to learn and identify the characteristics of multiple types of electrical quantity information associated with single-phase grounding faults in distribution networks, enabling analysis and judgment of faulty lines. These line selection mechanisms for single-phase grounding faults in distribution networks only select faulty lines based on the characteristics of fault electrical quantity information changes. They do not consider the complexities of line operating conditions, load types, and line selection rankings involved in single-phase grounding fault management. Consequently, they fail to truly integrate human experience with technical analysis.

[0005] Deep learning-based line selection analysis for single-phase grounding faults in distribution networks boasts strong representational properties and learning performance. It avoids the difficulties inherent in analyzing electrical quantity changes and line selection due to varying grounding modes in distribution networks, often encountered with intelligent information fusion methods like DS evidence theory and genetic algorithms. This leads to lower reliability and compromises in decision-making. However, the key advantage of deep learning lies in its ability to accurately extract and identify features. However, single-phase grounding fault handling decisions in distribution networks require machine networks to simulate the interaction between human dispatchers and the environment. Therefore, reinforcement learning becomes the primary choice. Combining deep learning and reinforcement learning, a deep reinforcement learning model is employed to model control decisions for line selection for single-phase grounding faults in distribution networks, effectively integrating expert experience and technical analysis. Summary of the Invention

[0006] In order to solve one of the above-mentioned technical defects, the present application provides a single-phase grounding fault line selection decision method, system and related equipment based on deep reinforcement learning.

[0007] According to a first aspect of the present application, a method for selecting a line for a single-phase grounding fault in a distribution network is provided, which is used to select a line where a single-phase grounding fault occurs in the distribution network, comprising:

[0008] Obtain the topological connection relationship of each feeder switch at each moment during the operation of the distribution network single-phase grounding fault, and construct the distribution network feeder switch adjacency matrix G ;

[0009] According to the collected electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network, the distribution network feeder switch characteristic matrix is ​​constructed. ; Among them, the electrical quantity information characteristic data includes: active power mutation quantity, reactive power mutation quantity, current mutation quantity and zero sequence current;

[0010] Adjacency matrix of feeder switches in distribution network is analyzed using dynamic graph convolutional neural network. G and distribution network feeder switch characteristic matrix Processing is performed to obtain characteristic quantities of state information;

[0011] Generate a fault line selection strategy based on the fault line selection order;

[0012] According to the state information characteristics and the fault line selection strategy, a distribution network single-phase grounding fault line selection decision model is constructed based on deep reinforcement learning. It is used to realize the analysis of distribution network single-phase grounding fault line selection and determine the distribution network single-phase grounding fault line selection based on the analysis results.

[0013] Preferably, the distribution network feeder switch characteristic matrix is ​​constructed based on the collected electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network. , specifically including:

[0014] Collect the electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network, and construct the distribution network feeder switch matrix at the corresponding moment; the distribution network feeder switch matrix at time t' for:

[0015] ;

[0016] Where, is the electrical quantity information characteristic data of the nth feeder switch at time t', n is the number of feeder switches in the distribution network, is the active power mutation of the nth feeder switch at time t', is the reactive power mutation of the nth feeder switch at time t', is the current mutation of the nth feeder switch at time t', is the zero-sequence current of the nth feeder switch at time t';

[0017] The distribution network feeder switch matrix at the corresponding time is normalized and calculated to obtain the distribution network feeder switch characteristic matrix; the distribution network feeder switch characteristic matrix at time t' is for:

[0018] ;

[0019] Where, is the normalized electrical quantity information characteristic data of the nth feeder switch at time t', is the normalized active power mutation of the nth feeder switch at time t', is the normalized reactive power mutation of the nth feeder switch at time t', is the normalized current mutation of the nth feeder switch at time t', is the normalized zero-sequence current of the nth feeder switch at time t';

[0020] Constructing the distribution network feeder switch characteristic matrix : , where is the distribution network feeder switch characteristic matrix at time N, .

[0021] Preferably, the dynamic graph convolutional neural network includes an input layer, two hidden layers and an output layer; the two hidden layers are respectively a dynamic graph convolutional network layer and a graph convolutional network layer;

[0022] The dynamic graph convolutional neural network is used to analyze the adjacency matrix of the feeder switch of the distribution network. G and distribution network feeder switch characteristic matrix Processing is performed to obtain state information feature quantities, specifically including:

[0023] Distribution network feeder switch characteristic matrix and the distribution network feeder switch adjacency matrix G Input into the input layer of the dynamic graph convolutional neural network; the number of neurons in the input layer is 4;

[0024] The distribution network feeder switch characteristic matrix of the input layer and the distribution network feeder switch adjacency matrix G In the dynamic graph convolutional network layer of the input hidden layer, the distribution network feeder switch feature matrix is ​​firstly analyzed by the self-attention mechanism. and the distribution network feeder switch adjacency matrix G The dynamic relationship of each feeder switch at each moment is extracted, and then the dynamic graph convolutional network is used to calculate the extracted dynamic relationship to obtain the feeder switch feature information, and the feeder switch feature information is extracted;

[0025] The extracted feeder switch feature information is input into the graph convolutional network layer of the hidden layer, and the feeder switch feature information is calculated using the graph convolutional network to obtain the state information feature, and the state information feature is extracted;

[0026] The extracted state information features are output through the output layer.

[0027] Preferably, the calculation formula of the self-attention mechanism is:

[0028] ;

[0029] Where, M is the weight matrix calculated by the self-attention mechanism, is the normalization function, is the distribution network feeder switch characteristic matrix The transposed matrix of D is the adjacency matrix of the feeder switches in the distribution network B degree matrix of ;

[0030] The calculation formula of the dynamic graph convolutional network in the dynamic graph convolutional network layer is:

[0031] ;

[0032] Where, H is the characteristic information of the feeder switch, is the activation function, is the weight matrix of the dynamic graph convolutional network layer, is the Hadamard product;

[0033] The calculation formula of the graph convolution network in the dynamic graph convolution network layer is:

[0034] ;

[0035] Where: I is the state information characteristic quantity, is the weight matrix of the graph convolutional network layer.

[0036] Preferably, the generation of a faulty line selection strategy based on the faulty line selection order specifically includes:

[0037] Generate a fault line selection table based on the fault line selection order: the fault line selection order is: no-load line (branch) switch, dual power supply line switch, reactive power compensation device, and station transformer; the dual power supply line switch selection order is: general user substation dual power supply line, general public substation dual power supply line, high-risk user substation dual power supply line, and key public substation dual power supply line.

[0038] Obtain the fault phase voltage of a single-phase grounding fault in the distribution network;

[0039] According to the selection order in the fault line selection table, each switch is selected at the same time to determine whether the fault phase voltage is restored after the corresponding switch is selected;

[0040] If it is restored, the single-phase grounding fault line of the distribution network is obtained and the selection is completed;

[0041] Otherwise, the selection and pulling sequence is continued for each switch at the next moment until a single-phase grounding fault line of the distribution network is obtained, and the selection and pulling is terminated.

[0042] Preferably, the construction of a single-phase grounding fault line selection decision model for the distribution network based on deep reinforcement learning according to the state information feature quantity and the fault line selection strategy specifically includes:

[0043] Build a deep reinforcement learning network and experience pool; the deep reinforcement learning network includes a training Q network and a target Q network;

[0044] Initialize the training Q network and the target Q network, and assign values ​​to the parameters in the training Q network and the target Q network;

[0045] According to the extracted state information feature quantity, the initial state space variable is constructed , and the initial state space variables are initialized once per cycle Observation update; wherein, the period is the pull control period, , is the characteristic quantity of state information at time t in each cycle;

[0046] According to the on / off status of each feeder switch, the corresponding action is obtained , and based on each action obtained , according to the fault line selection strategy to form a corresponding action set A ;in, , , is the action at time t in each cycle, is the on / off state of the i-th feeder switch at time t in each cycle, and n is the number of feeder switches in the distribution network;

[0047] Construct reward function r;

[0048] The initial state space variables of each cycle Input the training Q network and iteratively perform the following steps: Strategy, from the corresponding action set A Iterate and select an action , generate the action value corresponding to the initial state , and the selected action Input the fault line selection strategy to execute the selection, obtain the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch, and then calculate the action according to the reward function r The instant reward value ;in, is the neural network weight for training the Q network;

[0049] Based on the SCADA system, the new state information characteristic quantity is obtained according to the updated topological connection relationship of each feeder switch and the electrical quantity information characteristic data, and the state space variables of the next moment of the corresponding cycle are updated and calculated. ;

[0050] Building an empirical dataset , and store it in the experience pool;

[0051] Randomly extract d sets of experience data sets from the experience pool and input them into the training Q network and the target Q network respectively, and calculate the value of the training Q network and the target value of the target Q network; where d <d m , d m is the data capacity of the experience pool;

[0052] According to the loss function, the target value of the target Q network is used to update the neural network weights of the training Q network, so that the value of the training Q network is closer to the target value of the target Q network;

[0053] Determine whether the current number of iterations is equal to the maximum number of iterations;

[0054] If it is equal to, the trained distribution network single-phase grounding fault line selection strategy model is obtained;

[0055] If not, proceed to the next iteration.

[0056] Preferably, the Strategy, from the action set A Select an action , specifically including:

[0057] by The probability of selecting the optimal action corresponding to the maximum Q' value from the action set A is The probability of the action set A Pick a random action from

[0058] The calculation formula of the strategy is:

[0059] ;

[0060] Where, for Policy function, is the random selection probability and ε<1, argmaxQ' (a,s) is the maximum Q' value corresponding to the optimal action a, and |A| is the number of actions in the action set.

[0061] Preferably, the calculation formula of the reward function is:

[0062] ;

[0063] Where, is the fault phase voltage penalty term, is the substation load loss penalty term, is the reactive power imbalance penalty item;

[0064] in, The expression is:

[0065] ;

[0066] Where, is the threshold voltage, is the fault phase voltage when the distribution network has a single-phase ground fault, and d is the penalty term;

[0067] The expression is:

[0068] ;

[0069] Where, For general user substation load loss, For general public substation load loss, Substation load loss for high-risk users, For the loss of load at key utility substations, m 1 ~ m 4 They are general user substation, general public substation, high-risk user substation and key public substation. k 1 ~ k 4 are the load weight parameters of the corresponding substations, and k 4 >k 3 >k 2 >k 1 , l 1 ~ l 4 are the sum of the numbers of corresponding substations or corresponding users respectively;

[0070] The expression is:

[0071] ;

[0072] Where, 、 They are the upper limit and lower limit of reactive power under normal operating conditions of the distribution network system; is the reactive power of the distribution network system at time t, and c is the penalty value.

[0073] The target Q network calculation formula is:

[0074] ;

[0075] Where, is the target value of the target Q network, h is the current iteration number, T is the maximum number of iterations, γ is the discount coefficient, arg max Q '( a h , s h ) is the maximum Q' value corresponding to the optimal action selected in the current iteration;

[0076] The update formula of the neural network weights of the training Q network is:

[0077] ;

[0078] Where, is the loss function, For iteration h The neural network weights of the training Q network after the update; For iteration h The value of training the Q network.

[0079] According to the second aspect of the present application, a distribution network single-phase grounding fault line selection decision system based on deep reinforcement learning is provided, including a module for implementing the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning as described above.

[0080] According to a third aspect of the present application, there is provided a device, comprising:

[0081] Memory;

[0082] processor; and

[0083] computer programs;

[0084] Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning as described above.

[0085] According to the fourth aspect of the present application, a computer-readable storage medium is provided, characterized in that a computer program is stored thereon; the computer program is executed by a processor to implement the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning as described above.

[0086] The single-phase ground fault line selection decision method for distribution network based on deep reinforcement learning provided in this application is constructed by constructing the distribution network feeder switch feature matrix , improve the shortcomings of the traditional matrix, and construct the distribution network feeder switch adjacency matrix G The paper takes the single-phase ground fault line selection and pulling decision model of the distribution network as the object, fully adapts to the establishment of the single-phase ground fault line selection and pulling decision model of the distribution network, and selects the active power mutation, reactive power mutation, current mutation and zero-sequence current of each feeder switch as the feeder switch characteristics, providing reliable and accurate information support for the single-phase ground fault line selection and pulling of the distribution network; This application adopts a dynamic graph convolutional neural network to analyze the adjacency matrix of the feeder switch of the distribution network. G and distribution network feeder switch characteristic matrix Dynamic calculation and feature extraction of the operating topology and electrical quantity information characteristic data in the distribution network are carried out to realize real-time calculation of the operating topology and electrical quantity information characteristic data during the single-phase grounding selection of the distribution network, so as to adapt to the changes in the distribution network topology and flow information caused by the dynamic selection of the single-phase grounding fault line in the distribution network, and improve the adaptability and accuracy of the distribution network single-phase grounding fault line selection decision model; at the same time, the present application adopts a complete and practical fault line selection strategy to play a scientific, reasonable and effective role in the construction of the model, improve the practicality and universality of the distribution network single-phase grounding fault line selection decision model, and improve the accuracy of the method; in addition, the present application performs control modeling of the distribution network single-phase grounding fault line selection based on the deep reinforcement learning model according to the state information characteristic quantity and fault line selection strategy identified by the dynamic graph neural network, which not only realizes the comprehensive analysis and judgment of the line selection of the distribution network single-phase grounding fault, but also realizes the efficient integration of expert experience and technical analysis, improves the rapid recovery capability of the distribution network under the single-phase grounding fault, and provides efficient guarantee for the safe, reliable, scientific and accurate scheduling and operation decision-making of the distribution network.

[0087] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the contents indicated in the written description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0089] Figure 1 A flowchart of a method for selecting and pulling a single-phase grounding fault line in a distribution network based on deep reinforcement learning provided in an embodiment of the present application;

[0090] Figure 2 for Figure 1 A schematic diagram of a process for processing a distribution network feeder switch adjacency matrix and a distribution network feeder switch feature matrix using a dynamic graph convolutional neural network in the embodiment;

[0091] Figure 3 for Figure 1 A schematic diagram of a flow chart for executing a fault line selection strategy in the embodiment;

[0092] Figure 4 for Figure 1 A schematic diagram of a flow chart of obtaining a trained distribution network single-phase grounding fault line selection strategy model in the embodiment;

[0093] Figure 5 A functional structure diagram of a single-phase grounding fault line selection decision system for a distribution network based on deep reinforcement learning provided in an embodiment of the present application;

[0094] Figure 6 for Figure 5 A schematic diagram of the functional structure of the second matrix building module in the embodiment;

[0095] Figure 7 for Figure 5 A functional structure diagram of a state information feature acquisition module in the embodiment;

[0096] Figure 8 for Figure 5 A functional structure diagram of a fault line selection strategy generation module in the embodiment;

[0097] Figure 9 for Figure 5 A schematic diagram of the functional structure of the model building module in the embodiment;

[0098] In the figure: 10 is a topology connection relationship acquisition module, 20 is a first matrix construction module, 30 is a data acquisition module, 40 is a second matrix construction module, 50 is a state information feature acquisition module, 60 is a fault line selection strategy generation module, 70 is a model construction module, 401 is a distribution network feeder switch matrix construction unit, 402 is a first calculation unit, 403 is a distribution network feeder switch feature matrix construction unit, 501 is a first input unit, 502 is a first output unit, 503 is a second input unit, 504 is a self-attention mechanism extraction unit, 505 is a second calculation unit, 506 is a first extraction unit, 507 is a third input unit, 508 is a third calculation unit, 509 is a second extraction unit, 510 is a fourth input unit, 511 is a second output unit, 601 is a fault line selection table generation unit, 602 is a fault phase voltage acquisition unit, 603 is a selection unit, 6 04 is the first judgment unit, 605 is the fault line acquisition unit, 606 is the first execution unit, 701 is the first construction unit, 702 is the initialization and assignment unit, 703 is the second construction unit, 704 is the initial state space variable updating unit, 705 is the action acquisition unit, 706 is the third construction unit, 707 is the fourth construction unit, 708 is the fifth input unit, 709 is the action selection unit, 710 is the data generation unit, 711 is the data updating unit, 712 is the reward value acquisition unit, 713 is the data acquisition unit, 714 is the state updating unit, 715 is the fifth construction unit, 716 is the experience data set storage unit, 717 is the experience data set extraction unit, 718 is the sixth input unit, 719 is the fourth calculation unit, 720 is the neural network weight updating unit, 721 is the second judgment unit, 722 is the model acquisition unit, and 723 is the third execution unit. DETAILED DESCRIPTION

[0099] In order to make the technical solutions and advantages of the embodiments of the present application more clearly understood, the exemplary embodiments of the present application are further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, and are not an exhaustive list of all the embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other unless they conflict.

[0100] In view of some problems existing in the existing technology:

[0101] In a first aspect, embodiments of the present application provide a method for selecting and deciding a single-phase ground fault line in a distribution network based on deep reinforcement learning. This method can be performed by a single-phase ground fault line selection and decision-making system for a distribution network based on deep reinforcement learning, or by components configured within the single-phase ground fault line selection and decision-making system for a distribution network based on deep reinforcement learning, such as a chip or chip system. It can also be implemented by a logic module or software that has some or all of the functions of a single-phase ground fault line selection and decision-making system for a distribution network based on deep reinforcement learning. This application is not limited to this.

[0102] For example, Figure 1 As shown, the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning is used to select the line where the single-phase grounding fault occurs in the distribution network, including:

[0103] Obtain the topological connection relationship of each feeder switch at each moment during the operation of the distribution network single-phase grounding fault, and construct the distribution network feeder switch adjacency matrix G ; Among them, the distribution network feeder switch adjacency matrix G To reflect the distribution network structure, G is an n×n matrix, where n is the number of feeder switches in the distribution network; the adjacency matrix of the feeder switches in the distribution network is G In the equation, if feeder switch i is connected to feeder switch j, the adjacency matrix of feeder switches in the distribution network is G The jth element in the i-th row is 1, otherwise it is 0;

[0104] According to the collected electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network, the distribution network feeder switch characteristic matrix is ​​constructed. ; Among them, the electrical quantity information characteristic data includes: active power mutation quantity, reactive power mutation quantity, current mutation quantity and zero sequence current;

[0105] Adjacency matrix of feeder switches in distribution network is analyzed using dynamic graph convolutional neural network. G and distribution network feeder switch characteristic matrix Processing is performed to obtain characteristic quantities of state information;

[0106] Generate a fault line selection strategy based on the fault line selection order;

[0107] According to the state information characteristics and the fault line selection strategy, a distribution network single-phase grounding fault line selection decision model is constructed based on deep reinforcement learning. It is used to realize the analysis of distribution network single-phase grounding fault line selection and determine the distribution network single-phase grounding fault line selection based on the analysis results.

[0108] Based on the above scheme, the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning provided in this application is constructed by constructing the distribution network feeder switch feature matrix , improve the shortcomings of the traditional matrix, and construct the distribution network feeder adjacency matrix G The paper takes the single-phase ground fault line selection and pulling decision model of the distribution network as the object, fully adapts to the establishment of the single-phase ground fault line selection and pulling decision model of the distribution network, and selects the active power mutation, reactive power mutation, current mutation and zero-sequence current of each feeder switch as the feeder switch characteristics, providing reliable and accurate information support for the single-phase ground fault line selection and pulling of the distribution network; This application adopts a dynamic graph convolutional neural network to analyze the adjacency matrix of the feeder switch of the distribution network. G and distribution network feeder switch characteristic matrix Dynamic calculation and feature extraction of the operating topology and electrical quantity information characteristic data in the distribution network are carried out to realize real-time calculation of the operating topology and electrical quantity information characteristic data during the single-phase grounding selection of the distribution network, so as to adapt to the changes in the distribution network topology and flow information caused by the dynamic selection of the single-phase grounding fault line in the distribution network, and improve the adaptability and accuracy of the distribution network single-phase grounding fault line selection decision model; at the same time, the present application adopts a complete and practical fault line selection strategy to play a scientific, reasonable and effective role in the construction of the model, improve the practicality and universality of the distribution network single-phase grounding fault line selection decision model, and improve the accuracy of the method; in addition, the present application performs control modeling of the distribution network single-phase grounding fault line selection based on the deep reinforcement learning model according to the state information characteristic quantity and fault line selection strategy identified by the dynamic graph neural network, which not only realizes the comprehensive analysis and judgment of the line selection of the distribution network single-phase grounding fault, but also realizes the efficient integration of expert experience and technical analysis, improves the rapid recovery capability of the distribution network under the single-phase grounding fault, and provides efficient guarantee for the safe, reliable, scientific and accurate scheduling and operation decision-making of the distribution network.

[0109] In some possible implementations of the first aspect, the distribution network feeder switch characteristic matrix is ​​constructed based on the collected electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network. , specifically including:

[0110] Collect the electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network, and construct the distribution network feeder switch matrix at the corresponding moment; the distribution network feeder switch matrix at time t' for:

[0111] ;

[0112] Where, is the electrical quantity information characteristic data of the nth feeder switch at time t', n is the number of feeder switches in the distribution network, is the active power mutation of the nth feeder switch at time t', is the reactive power mutation of the nth feeder switch at time t', is the current mutation of the nth feeder switch at time t', is the zero-sequence current of the nth feeder switch at time t';

[0113] According to the formula, the distribution network feeder switch matrix at the corresponding moment is normalized and calculated to obtain the distribution network feeder switch characteristic matrix at the corresponding moment; the distribution network feeder switch characteristic matrix at time t' for:

[0114] ;

[0115] Where, is the normalized electrical quantity information characteristic data of the nth feeder switch at time t', is the normalized active power mutation of the nth feeder switch at time t', is the normalized reactive power mutation of the nth feeder switch at time t', is the normalized current mutation of the nth feeder switch at time t', is the normalized zero-sequence current of the nth feeder switch at time t'.

[0116] Based on the above scheme, the normalized distribution network feeder switch characteristic matrix is All the data in are within the range of [-1, 1], which can clearly reflect the degree of change of the electrical quantity information of the feeder switch and eliminate the adverse effects of singular data.

[0117] Specifically, the calculation formula for the active power mutation is: Where, is the active power mutation of the i-th feeder switch at time t', is the active power of the i-th feeder switch at time t', is the active power of the i-th feeder switch when the distribution network is operating normally without any faults at time t'; this variable has positive and negative properties. A positive value indicates that the feeder switch has power outflow and is the power-side switch; a negative value indicates that the feeder switch has power inflow and is the load-side switch.

[0118] The calculation formula for the reactive power sudden change is: Where, is the reactive power mutation of the ith feeder switch at time t', is the reactive power of the ith feeder switch at time t', The reactive power of the i-th feeder switch corresponding to the normal operation of the distribution network without faults at time t'; this variable has positive and negative properties. A positive value indicates that the feeder switch has power outflow and is the power-side switch; a negative value indicates that the feeder switch has power inflow and is the load-side switch.

[0119] The calculation formula for the current mutation is: Where, is the current mutation of the i-th feeder switch at time t', is the current of the i-th feeder switch at time t', ' is the current of the i-th feeder switch when the distribution network is operating normally without any faults at time t'; this variable has positive and negative properties. A positive value indicates that the current of the feeder switch is outflowing, and it is a power-side switch; a negative value indicates that the current of the feeder switch is inflowing, and it is a load-side switch.

[0120] The calculation formula for zero-sequence current is: , Where, is the zero sequence current; is the system angular frequency, , is the system frequency; is the total capacitance of each component of the distribution network to ground, is the zero-sequence voltage, is the three-phase voltage of A', B', and C' after the distribution network system fails 、 as well as In the distribution network, insulation monitoring devices are usually configured to monitor the zero-sequence voltage. If conditions permit, zero-sequence protection can be configured to obtain the zero-sequence current, thereby realizing the calculation and extraction of zero-sequence current. As an important characteristic quantity, the zero-sequence current is different from the power and current mutation quantities. The zero-sequence current is a vector quantity, including the zero-sequence current amplitude and phase angle, so as to realize the effective identification of the electrical fault characteristics when the distribution network is operating in the grounding mode through the arc suppression coil.

[0121] Optionally, the distribution network feeder switch matrix at the corresponding moment is normalized to obtain the distribution network feeder switch characteristic matrix at the corresponding moment, specifically:

[0122] According to the formula Normalization calculation is performed on each element in the distribution network feeder switch matrix at the corresponding moment to obtain a corresponding normalized element; the element is the active power mutation amount or reactive power mutation amount or current mutation amount or zero-sequence current of each feeder switch;

[0123] Where, is the number of the feeder switch matrix in the distribution network at the corresponding moment. k elements, 、 are the maximum and minimum values ​​of the element respectively. After normalization k elements.

[0124] In some possible implementations of the first aspect, such as Figure 2 As shown, the dynamic graph convolutional neural network includes an input layer, two hidden layers and an output layer; the two hidden layers are respectively a dynamic graph convolutional network layer and a graph convolutional network layer;

[0125] The dynamic graph convolutional neural network is used to analyze the adjacency matrix of the feeder switch of the distribution network. G and distribution network feeder switch characteristic matrix Processing is performed to obtain state information feature quantities, specifically including:

[0126] Distribution network feeder switch characteristic matrix and the distribution network feeder switch adjacency matrix G Input into the input layer of the dynamic graph convolutional neural network; the number of neurons in the input layer is 4;

[0127] The distribution network feeder switch characteristic matrix of the input layer and the distribution network feeder switch adjacency matrix G In the dynamic graph convolutional network layer of the input hidden layer, the distribution network feeder switch feature matrix is ​​firstly analyzed by the self-attention mechanism. and the distribution network feeder switch adjacency matrix G The dynamic relationship of each feeder switch at each moment is extracted, and then the dynamic graph convolution network is used to calculate the extracted dynamic relationship to obtain the feeder switch feature information, and the feeder switch feature information is extracted; wherein, the dynamic relationship is the distribution network feeder switch feature matrix caused by the change of the power grid topology at each moment and the distribution network feeder switch adjacency matrix G the dynamic changes that occur;

[0128] The extracted feeder switch feature information is input into the graph convolutional network layer of the hidden layer, and the feeder switch feature information is calculated using the graph convolutional network to obtain the state information feature, and the state information feature is extracted;

[0129] The extracted state information features are output through the output layer.

[0130] Optionally, the calculation formula of the self-attention mechanism is:

[0131] ;

[0132] Where, Mis the weight matrix calculated by the self-attention mechanism, which is used to characterize the dynamic weight relationship of the feeder switches in the distribution network. is the normalization function, is the distribution network feeder switch characteristic matrix The transposed matrix of D is the adjacency matrix of the feeder switches in the distribution network G The degree matrix of .

[0133] Based on the above scheme, this application uses a self-attention mechanism to dynamically update the calculation based on the position and size of each input variable, thereby obtaining the spatial correlation of the distribution network feeder switch status. In addition, the self-attention mechanism can be used for parallel calculation, which effectively improves computational efficiency.

[0134] Specifically, the degree matrix D is a diagonal matrix, and the elements on its diagonal are the degrees of each vertex in the graph. D The calculation formula is:

[0135] ;

[0136] Where, are diagonal elements, is the degree of the i-th feeder switch, which represents the number of edges connected to the vertex, that is, the sum of all non-zero elements (element 1) in the i-th row, j' are all elements in row i.

[0137] Optionally, the calculation formula of the dynamic graph convolutional network in the dynamic graph convolutional network layer is:

[0138] ;

[0139] Where, H is the characteristic information of the feeder switch, is the activation function, is the weight matrix of the dynamic graph convolutional network layer, is the Hadamard product;

[0140] Optionally, the calculation formula of the graph convolutional network in the graph convolutional network layer is:

[0141] ;

[0142] Where: I is the state information characteristic quantity, is the weight matrix of the graph convolutional network layer.

[0143] Specifically, the activation function is the Sigmoid function, and the calculation formula of the Sigmoid function is:

[0144] .

[0145] In some possible implementations of the first aspect, such as Figure 3 As shown, the fault line selection strategy is generated based on the fault line selection order, specifically including:

[0146] Generate a fault line selection table based on the fault line selection order: the fault line selection order is: no-load line (branch) switch, dual power supply line switch, reactive power compensation device, and station transformer; the dual power supply line switch selection order is: general user substation dual power supply line, general public substation dual power supply line, high-risk user substation dual power supply line, and key public substation dual power supply line.

[0147] Obtain the fault phase voltage at time t of a single-phase grounding fault in the distribution network;

[0148] According to the selection order in the fault line selection table, each switch is selected at time t+1 to determine whether the fault phase voltage has recovered after selecting the corresponding switch. Selecting each switch at the same time according to the selection order avoids random selection at the same time, indiscriminate selection of line types and load losses, and effectively improves the selection accuracy.

[0149] If it is restored, the single-phase grounding fault line of the distribution network is obtained and the selection is completed;

[0150] Otherwise, the selection and pulling is continued for each switch at the next moment according to the selection and pulling sequence until a single-phase grounding fault line of the distribution network is obtained, and the selection and pulling is terminated.

[0151] Optionally, after the dual power supply incoming line switch is pulled, if the single-phase grounding fault does not disappear, the opposite side switch, that is, the dual power supply feeder line node switch is pulled.

[0152] Specifically, after selecting the reactive compensation device and the station transformer, there is no need to determine whether the fault phase voltage has recovered.

[0153] Specifically, the active power, reactive power and current values ​​of the no-load line (branch) switch in normal operation are all 0, that is, , , ;

[0154] The active power, reactive power and current values ​​of the dual power incoming switch under normal operation are the power and current corresponding to the load. Since the grid load is generally inductive, the active power, reactive power and current values ​​are all negative, indicating the direction of the flow, that is, , , ;

[0155] The active power, reactive power and current values ​​of the dual power feeder line node switch under normal operating conditions are the power and current corresponding to the load. At this time, the active power, reactive power and current values ​​are all positive, indicating the outflow direction of the power flow, that is, , , ;

[0156] The active power, reactive power and current values ​​of the reactive compensation device and the station transformer in normal operation are the power and current corresponding to the reactive compensation device and the station transformer. Since the reactive compensation device and the station transformer are capacitive components, the active power of the reactive compensation device and the station transformer is 0, and they inject reactive power and reactive current into the distribution network system. Therefore, the reactive power and current values ​​of the reactive compensation device and the station transformer are negative, indicating the direction of the power flow, that is, , , .

[0157] More specifically, the reactive power compensation device is a capacitor or a high-voltage static VAR generator (SVG).

[0158] The general user substation in this application is the substation for dedicated line users who are not included in the Energy Bureau's high-risk user list; the general public substation is the substation that is not connected to the critical distribution network load (city network users, important users and high-risk users); the high-risk user substation is the substation for dedicated line users included in the Energy Bureau's high-risk user list; the critical public substation is the substation that is connected to the critical distribution network load (city network users, important users and high-risk users).

[0159] In some possible implementations of the first aspect, such as Figure 4 As shown, the distribution network single-phase grounding fault line selection decision model is constructed based on deep reinforcement learning according to the state information feature quantity and the fault line selection strategy, which specifically includes:

[0160] Build a Deep Q Network (DQN) and experience pool. DQN consists of a training Q network and a target Q network, and the two networks have exactly the same structure.

[0161] Initialize the training Q network and the target Q network, and assign values ​​to the parameters in the training Q network and the target Q network;

[0162] According to the extracted state information feature quantity, the initial state space variable is constructed , and the initial state space variables are initialized once per cycle Observation update; wherein, the period is the distribution network selection control period, , is the characteristic quantity of state information at time t in each cycle;

[0163] According to the on / off status of each feeder switch in each cycle, the corresponding action is obtained , and based on each action obtained , according to the fault line selection strategy to form a corresponding action set A , is the action at time t in each cycle, and the breaking state of each feeder switch in each cycle is different; , , is the on / off state of the i-th feeder switch at time t in each cycle, and the on / off state is represented by 0 and 1, where 0 indicates that the switch is not in operation and 1 indicates that it is in operation and disconnected, and n is the number of feeder switches in the distribution network;

[0164] Construct a reward function r. When a single-phase grounding fault occurs in the distribution network, consider whether the selection result eliminates the grounding fault. The electrical quantity that directly reflects a single-phase grounding fault in the distribution network is the phase voltage. When a single-phase grounding fault occurs in the distribution network, the fault phase voltage decreases, and when the voltage recovers, it is within the stable range. When selecting the line for a single-phase grounding fault in the distribution network, consider the load loss caused by selecting the dual-power incoming line, and impose corresponding penalties based on the type of dual-power substation. Also, consider the system reactive imbalance caused by the withdrawal of reactive compensation devices such as capacitors and SVGs, as well as station transformers, due to the grounding selection. Penalize the reactive imbalance caused by selecting capacitors, SVGs, and station transformers.

[0165] The initial state space variables of each cycle Input the training Q network and iteratively perform the following steps: Strategy, from the corresponding action set A Select an action , generate the action value corresponding to the initial state , and the selected action Input the fault line selection strategy to execute the selection, obtain the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch, and then calculate the action according to the reward function r The instant reward value ;in, is the neural network weight for training the Q network;

[0166] Based on the SCADA system, the new state information characteristic quantity is obtained according to the updated topological connection relationship of each feeder switch and the electrical quantity information characteristic data, and the state space variables of the next moment of the corresponding cycle are updated and calculated. ;

[0167] Building an empirical dataset , and store it in the experience pool;

[0168] Randomly extract d sets of experience data sets from the experience pool and input them into the training Q network and the target Q network respectively, and calculate the value of the training Q network and the target value of the target Q network; where d <d m , d m is the data capacity of the experience pool;

[0169] According to the loss function, the target value of the target Q network is used to update the neural network weights of the training Q network, so that the value of the training Q network is closer to the target value of the target Q network;

[0170] Determine whether the current number of iterations is equal to the maximum number of iterations;

[0171] If it is equal to, the trained distribution network single-phase grounding fault line selection strategy model is obtained;

[0172] If not, proceed to the next iteration.

[0173] The SCADA (Supervisory Control And Data Acquisition) system is a data acquisition and monitoring control system that can transmit voltage, current, power and other information of various components such as switches, main transformers, and busbars in the distribution network to the dispatching master station in real time. Dispatching operators can intuitively see the wiring topology and operating information at any moment.

[0174] When a single-phase grounding fault occurs in the distribution network, the grid is typically allowed to operate for two hours. During this period, the grid can be subjected to short-term selective control. The selective control cycle is divided into multiple, 10-second cycles. Each cycle uses the fault line selection strategy provided in this application to select the single-phase grounding fault line in the distribution network. The selection targets are each feeder switch, including branch switches, to determine the single-phase grounding fault line in the distribution network. In addition, the open and closed state of each feeder switch, the topological connection relationship of each feeder switch, and the characteristic data of electrical quantity information are different in each cycle.

[0175] Based on the above scheme, this application breaks the correlation between data through the experience replay mechanism, improves the stability of deep reinforcement learning network training, and improves the reliability and accuracy of the single-phase grounding fault line selection strategy model of the distribution network.

[0176] Optionally, the parameter assignment for the training Q network and the target Q network includes: assigning the number of training Q network layers, the number of target Q network layers, the maximum number of iterations T, the discount coefficient γ and the experience pool data capacity d m .

[0177] Specifically, the DQN includes an input layer, a hidden layer, and an output layer; wherein the input layer is a state space variable , the number of layers is set as the state space variable The number of elements in the dynamic graph neural network is the state information feature after calculation and extraction; the number of hidden layers is set according to the specific training sample data; the output layer is the value of each action in the state , is the action of the i-th feeder switch in the output layer, where , n is the number of feeder switches in the distribution network. In specific implementation, the parameters of the input layer, hidden layer, and output layer of the training Q network and the target Q network are set according to the above parameters.

[0178] Optionally, the basis Strategy, from the action set A Select the corresponding action , specifically including:

[0179] by The probability of selecting the optimal action corresponding to the maximum Q' value from the action set A is The probability of the action set A Pick a random action from

[0180] The calculation formula of the strategy is:

[0181] ;

[0182] Where, for Strategy function, ε is the probability of random selection, and ε is a very small positive number, ε<1, argmax Q' ( a , s ) is the optimal action a The corresponding maximum Q' value, |A| is the number of actions in the action set.

[0183] Optionally, the calculation formula of the reward function is:

[0184] ;

[0185] Where, is the fault phase voltage penalty term, is the substation load loss penalty term, is the reactive imbalance penalty term; when making intelligent decisions, this reward function can minimize the loss cost;

[0186] in, The expression is:

[0187] ;

[0188] Where, is the threshold voltage, is the fault phase voltage when the distribution network has a single-phase grounding fault, and d is the penalty term; when the distribution network has a single-phase grounding fault, the fault phase voltage decreases and the non-fault phase voltage increases. Greater than the threshold voltage When the voltage is normal, no penalty term is added; when the fault phase voltage Below threshold voltage When , the ground fault still does not disappear, and the penalty term d is added;

[0189] The expression is:

[0190] ;

[0191] Where, For general user substation load loss, For general public substation load loss, Substation load loss for high-risk users, For the loss of load at key utility substations, m 1 ~ m 4 They are general user substation, general public substation, high-risk user substation and key public substation. k 1 ~ k 4 are the load weight parameters of the corresponding substations, and k 4 >k 3 >k 2 >k 1 , l 1 ~ l 4 are the sum of the number of corresponding substations or corresponding users respectively. The more important the load loss is, the greater the penalty is.

[0192] The expression is:

[0193] ;

[0194] Where, 、 They are the upper and lower limits of reactive power under normal operating conditions of the distribution network system; is the reactive power of the distribution network system at time t, and c is the penalty value.

[0195] Optionally, the target Q network calculation formula is:

[0196] ;

[0197] Where, is the target value of the target Q network, h is the current iteration number, T is the maximum number of iterations, is the maximum Q' value corresponding to the optimal action selected in the current iteration, is a discount factor to limit the reward value to avoid infinite reward value.

[0198] When it is implemented specifically, h The maximum number of iterations set has been reached T When training the Q network, the immediate reward value is the final reward value; when the maximum number of iterations is not reached T When , the long-term accumulated reward value is the final reward value.

[0199] Based on the above scheme, when constructing the reward function, this application comprehensively considers the order of line selection, load type, power-off load and selection accuracy, thereby improving the accuracy of the comprehensive analysis and judgment of line selection for single-phase grounding faults in the distribution network.

[0200] Optionally, the update formula for the neural network weights of the training Q network is:

[0201] ;

[0202] Where, is the loss function, For iteration h The neural network weights of the training Q network after the update; For iteration h The value of training the Q network.

[0203] In the second aspect, an embodiment of the present application provides a distribution network single-phase grounding fault line selection decision system based on deep reinforcement learning, which includes a module for implementing the aforementioned distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning.

[0204] For example, Figure 5 As shown in FIG, the distribution network single-phase grounding fault line selection decision system based on deep reinforcement learning includes:

[0205] A topology connection relationship acquisition module 10 is used to obtain the topology connection relationship of each feeder switch at each moment during the operation of a single-phase grounding fault in the distribution network;

[0206] The first matrix construction module 20 is used to construct a distribution network feeder switch adjacency matrix according to the topological connection relationship of each feeder switch at each moment during the operation of the distribution network single-phase grounding fault. G ;

[0207] Data acquisition module 30: used to collect electrical quantity information characteristic data of each feeder switch at each moment during the operation of single-phase grounding fault in the distribution network;

[0208] The second matrix construction module 40 is used to construct a distribution network feeder switch characteristic matrix based on the collected electrical quantity information characteristic data of each feeder switch at each moment during the operation of the distribution network single-phase grounding fault ;

[0209] State information feature acquisition module 50: used to use dynamic graph convolutional neural network to analyze the adjacency matrix of distribution network feeder switches G and distribution network feeder switch characteristic matrix Processing is performed to obtain characteristic quantities of state information;

[0210] Fault line selection strategy generation module 60: used to generate a fault line selection strategy based on the fault line selection sequence;

[0211] Model construction module 70: used to construct a distribution network single-phase grounding fault line selection decision model based on deep reinforcement learning according to the state information feature quantity and the fault line selection strategy.

[0212] In some possible implementations of the second aspect, such as Figure 6 As shown, the second matrix building module 40 includes:

[0213] The distribution network feeder switch matrix construction unit 401 at the corresponding moment is used to construct the distribution network feeder switch matrix at the corresponding moment according to the collected electrical quantity information characteristic data of each feeder switch at each moment during the operation of the distribution network single-phase grounding fault;

[0214] The first calculation unit 402 is configured to perform normalization calculation on the distribution network feeder switch matrix at a corresponding moment;

[0215] The distribution network feeder switch characteristic matrix construction unit 403 at the corresponding moment is used to construct the distribution network feeder switch characteristic matrix at the corresponding moment according to the distribution network feeder switch matrix at the corresponding moment after normalization calculation;

[0216] The distribution network feeder switch characteristic matrix construction unit 404 is used to construct the distribution network feeder switch characteristic matrix according to the distribution network feeder switch characteristic matrix at each moment. .

[0217] Alternatively, as Figure 7 As shown, the state information feature quantity acquisition module 50 includes:

[0218] The first input unit 501 is used to convert the distribution network feeder switch characteristic matrix and the distribution network feeder switch adjacency matrix G Input into the input layer of the dynamic graph convolutional neural network;

[0219] First output unit 502: used to output the distribution network feeder switch characteristic matrix of the input layer and the distribution network feeder switch adjacency matrix G ;

[0220] The second input unit 503 is used to input the distribution network feeder switch characteristic matrix of the input layer and the distribution network feeder switch adjacency matrix G Input the hidden layer of the dynamic graph convolutional network layer;

[0221] Self-attention mechanism extraction unit 504: used to extract the distribution network feeder switch feature matrix through the self-attention mechanism and the distribution network feeder switch adjacency matrix G Extract the dynamic relationship of each feeder switch at each moment;

[0222] The second calculation unit 505 is configured to calculate the extracted dynamic relationship using a dynamic graph convolutional network to obtain characteristic information of the feeder switch;

[0223] The first extraction unit 506 is used to extract the characteristic information of the feeder switch;

[0224] The third input unit 507 is used to input the extracted feeder switch feature information into the graph convolutional network layer of the hidden layer;

[0225] The third calculation unit 508 is configured to calculate the feeder switch characteristic information using a graph convolutional network to obtain a state information characteristic;

[0226] The second extraction unit 509 is used to extract the state information feature;

[0227] The fourth input unit 510 is used to input the extracted state information feature into the output layer;

[0228] The second output unit 511 is used to output the state information feature of the output layer.

[0229] Alternatively, as Figure 8 As shown, the fault line selection strategy generation module 60 includes:

[0230] Fault line selection table generating unit 601: used to generate a fault line selection table based on the fault line selection order;

[0231] Fault phase voltage acquisition unit 602: used to acquire the fault phase voltage of a single-phase grounding fault in the distribution network;

[0232] The selection unit 603 is used to select and pull each switch at the same time according to the selection order in the fault line selection table;

[0233] The first judgment unit 604 is used to judge whether the fault phase voltage is restored after the corresponding switch is selected;

[0234] Fault line acquisition unit 605: used to obtain the single-phase grounding fault line of the distribution network when the fault phase voltage recovers, and end the selection;

[0235] The first execution unit 606 is configured to continue the selection and pulling at the next moment according to the selection and pulling sequence when the fault phase voltage has not recovered.

[0236] Alternatively, as Figure 9 As shown, the model building module 70 includes:

[0237] The first construction unit 701 is used to construct a deep reinforcement learning network and an experience pool;

[0238] Initialization and assignment unit 702: used to initialize the training Q network and the target Q network, and assign values ​​to the parameters in the training Q network and the target Q network;

[0239] The second construction unit 703 is used to construct the initial state space variable according to the extracted state information feature quantity ;

[0240] Initial state space variable updating unit 704: used to update the initial state space variable once per cycle Observation update of

[0241] Action acquisition unit 705: used to obtain the corresponding action according to the on / off status of each feeder switch in each cycle ;

[0242] The third construction unit 706 is used to obtain each action , based on the fault line selection strategy, an action set is formed A ;

[0243] The fourth construction unit 707 is used to construct a reward function r;

[0244] The fifth input unit 708 is used to convert the initial state space variables Input into the training Q network;

[0245] Action selection unit 709: used to select Strategy, from the action set A Select an action ;

[0246] Data generation unit 710: used to generate the action value corresponding to the initial state ;

[0247] Data update unit 711: used to update the selected action Input the fault line selection strategy and execute the selection to obtain updated topological connection relationship and electrical quantity information characteristic data of each feeder switch;

[0248] Reward value acquisition unit 712: used to calculate the action according to the reward function r The instant reward value ;

[0249] Data acquisition unit 713: used to acquire new state information characteristic quantities based on the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch based on the SCADA system;

[0250] State updating unit 714: used to update and calculate the state space variables at the next moment based on the new state information feature quantity ;

[0251] Fifth construction unit 715: used to construct an experience data set ;

[0252] Experience data set storage unit 716: used to store the experience data set Deposit into the experience pool;

[0253] Experience data set extraction unit 717: used to randomly extract d groups of experience data sets from the experience pool;

[0254] The sixth input unit 718 is used to input the randomly selected d groups of experience data sets into the training Q network and the target Q network respectively;

[0255] Fourth calculation unit 719: used to calculate the value of the training Q network and the target value of the target Q network;

[0256] Neural network weight updating unit 720: used to update the neural network weights of the training Q network using the target value of the target Q network according to the loss function;

[0257] The second judging unit 721 is used to judge whether the current number of iterations is equal to the maximum number of iterations;

[0258] Model acquisition unit 722: used to obtain a trained distribution network single-phase grounding fault line selection strategy model when the current iteration number is equal to the maximum iteration number;

[0259] The third execution unit 723 is used to continue executing the next iteration when the current number of iterations is not equal to the maximum number of iterations.

[0260] On the third aspect, a device is provided in an embodiment of the present application. The device can be any device that can implement a single-phase grounding fault line selection decision method for a distribution network based on deep reinforcement learning. The device can be various terminal devices, such as: desktop computers, laptops, tablet computers, handheld devices, etc., and can be implemented specifically through software and / or hardware.

[0261] Exemplarily, the device includes:

[0262] Memory;

[0263] processor; and

[0264] computer programs;

[0265] Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning as described above.

[0266] In a fourth aspect, a computer-readable storage medium is provided in an embodiment of the present application. The computer-readable storage medium may be: ROM, RAM, a magnetic disk or an optical disk, etc.

[0267] Exemplarily, a computer program is stored on the computer-readable storage medium; the computer program is executed by a processor to implement the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning as described above.

[0268] In summary, the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning in this application provides scientific and efficient auxiliary decision-making for the handling of single-phase grounding faults in the distribution network, promotes the efficiency of single-phase grounding line selection in the distribution network, and at the same time, improves the recovery capacity of the distribution network.

[0269] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application may be implemented in various computer languages, such as C, VHDL, Verilog, object-oriented programming language Java, and interpreted scripting language JavaScript.

[0270] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0271] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0272] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0273] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. Throughout the description of this application, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0274] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0275] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for selecting a line with a single-phase grounding fault in a distribution network based on deep reinforcement learning, which is used to select a line with a single-phase grounding fault in the distribution network, characterized by: include: Obtain the topological connection relationship of each feeder switch at each moment during the operation of the distribution network single-phase grounding fault, and construct the distribution network feeder switch adjacency matrix G ; According to the collected electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network, the distribution network feeder switch characteristic matrix is ​​constructed. ; Among them, the electrical quantity information characteristic data includes: active power mutation quantity, reactive power mutation quantity, current mutation quantity and zero sequence current; Adjacency matrix of feeder switches in distribution network is analyzed using dynamic graph convolutional neural network. G and distribution network feeder switch characteristic matrix Processing is performed to obtain characteristic quantities of state information; Generate a fault line selection strategy based on the fault line selection order; Based on the state information characteristics and fault line selection strategy, a single-phase ground fault line selection decision model for the distribution network is constructed based on deep reinforcement learning. This model is used to analyze the selection of single-phase ground fault lines in the distribution network and determine the selected line based on the analysis results. Specifically, the model includes: Build a deep reinforcement learning network and experience pool; the deep reinforcement learning network includes a training Q network and a target Q network; Initialize the training Q network and the target Q network, and assign values ​​to the parameters in the training Q network and the target Q network; According to the extracted state information feature quantity, the initial state space variable is constructed s t , and the initial state space variables are initialized once per cycle s t Observation update; wherein, the period is the pull control period, , I t is the characteristic quantity of state information at time t in each cycle; According to the on / off status of each feeder switch in each cycle, the corresponding action is obtained a t , and based on each action obtained a t , according to the fault line selection strategy to form a corresponding action set A ;in, , , a t is the action at time t in each cycle, is the on / off state of the i-th feeder switch at time t in each cycle, and n is the number of feeder switches in the distribution network; Construct a reward function r; the calculation formula of the reward function is: ; Where, r 1 is the fault phase voltage penalty term, r 2 is the substation load loss penalty term, r 3 is the reactive power imbalance penalty item; in, r 1 The expression is: ; Where, is the threshold voltage, is the fault phase voltage when the distribution network has a single-phase ground fault, and d is the penalty term; r 2 The expression is: ; Where, For general user substation load loss, For general public substation load loss, Substation load loss for high-risk users, For the loss of load at key utility substations, m 1 ~ m 4 They are general user substation, general public substation, high-risk user substation and key public substation. k 1 ~ k 4 are the load weight parameters of the corresponding substations, and k 4 >k 3 >k 2 >k 1 , l 1 ~ l 4 are the sum of the numbers of corresponding substations or corresponding users respectively; r 3 The expression is: ; Where, 、 They are the upper and lower limits of reactive power under normal operating conditions of the distribution network system; is the reactive power of the distribution network system at time t, and c is the penalty value; The initial state space variables of each cycle s t Input the training Q network and iteratively perform the following steps: ε- greedy Strategy, from the corresponding action set A Select an action a t , generate the action value Q'( s t , a t ; θ t ) and set the selected action a t Input the fault line selection strategy to execute the selection, obtain the updated topological connection relationship and electrical quantity information characteristic data of each feeder switch, and then calculate the action according to the reward function r a t The instant reward value r t ; where θ t is the neural network weight for training the Q network; Based on the SCADA system, the new state information characteristic quantity is obtained according to the updated topological connection relationship of each feeder switch and the electrical quantity information characteristic data, and the state space variables of the next moment of the corresponding cycle are updated and calculated. s t+1 ; Constructing an empirical dataset ( s t , a t , r t , s t+1 ) and store it in the experience pool; Randomly extract d sets of experience data sets from the experience pool and input them into the training Q network and the target Q network respectively, and calculate the value of the training Q network and the target value of the target Q network; where d <d m , d m is the data capacity of the experience pool; The target Q network calculation formula is: ; Where, y h is the target value of the target Q network, h is the current iteration number, T is the maximum number of iterations, γ is the discount coefficient, arg max Q '( a h , s h ) is the maximum Q' value corresponding to the optimal action selected in the current iteration; According to the loss function, the target value of the target Q network is used to update the neural network weights of the training Q network, so that the value of the training Q network is closer to the target value of the target Q network; The update formula of the neural network weights of the training Q network is: ; Where, is the loss function, θ h For iteration h The neural network weights of the training Q network after the update; Q '( a h , s h ; θ h ) is the iteration h The value of training the Q network times; Determine whether the current number of iterations is equal to the maximum number of iterations; If it is equal to, the trained distribution network single-phase grounding fault line selection strategy model is obtained; If not, proceed to the next iteration.

2. The method for selecting and deciding a single-phase grounding fault line in a distribution network based on deep reinforcement learning according to claim 1 is characterized in that: The distribution network feeder switch characteristic matrix is ​​constructed based on the collected electrical quantity information characteristic data of each feeder switch at each moment during the single-phase grounding fault operation of the distribution network. , specifically including: According to the collected electrical quantity information characteristic data of each feeder switch at each moment during the operation of the distribution network single-phase grounding fault, the distribution network feeder switch matrix at the corresponding moment is constructed; the distribution network feeder switch matrix at time t' X t' for: ; Where, x n,t' is the electrical quantity information characteristic data of the nth feeder switch at time t', n is the number of feeder switches in the distribution network, Δ P n,t' is the active power mutation of the nth feeder switch at time t', Δ Q n,t' is the reactive power mutation of the nth feeder switch at time t', Δ I n,t' is the current mutation of the nth feeder switch at time t', is the zero-sequence current of the nth feeder switch at time t'; The distribution network feeder switch matrix at the corresponding time is normalized and calculated to obtain the distribution network feeder switch characteristic matrix at the corresponding time; the distribution network feeder switch characteristic matrix at time t' is for: ; Where, is the normalized electrical quantity information characteristic data of the nth feeder switch at time t', is the normalized active power mutation of the nth feeder switch at time t', is the normalized reactive power mutation of the nth feeder switch at time t', is the normalized current mutation of the nth feeder switch at time t', is the normalized zero-sequence current of the nth feeder switch at time t'; Constructing the distribution network feeder switch characteristic matrix : , where is the distribution network feeder switch characteristic matrix at time N, .

3. The method for selecting and deciding a single-phase grounding fault line in a distribution network based on deep reinforcement learning according to claim 1 is characterized in that: The dynamic graph convolutional neural network includes an input layer, two hidden layers and an output layer; the two hidden layers are respectively a dynamic graph convolutional network layer and a graph convolutional network layer; The dynamic graph convolutional neural network is used to analyze the adjacency matrix of the feeder switch of the distribution network. G and distribution network feeder switch characteristic matrix Processing is performed to obtain state information feature quantities, specifically including: Distribution network feeder switch characteristic matrix and the distribution network feeder switch adjacency matrix G Input into the input layer of the dynamic graph convolutional neural network; the number of neurons in the input layer is 4; The distribution network feeder switch characteristic matrix of the input layer and the distribution network feeder switch adjacency matrix G In the dynamic graph convolutional network layer of the input hidden layer, the distribution network feeder switch feature matrix is ​​firstly analyzed by the self-attention mechanism. and the distribution network feeder switch adjacency matrix G The dynamic relationship of each feeder switch at each moment is extracted, and then the dynamic graph convolutional network is used to calculate the extracted dynamic relationship to obtain the feeder switch feature information, and the feeder switch feature information is extracted; The extracted feeder switch feature information is input into the graph convolutional network layer of the hidden layer, and the feeder switch feature information is calculated using the graph convolutional network to obtain the state information feature, and the state information feature is extracted; The extracted state information features are output through the output layer.

4. The method for selecting and deciding a single-phase grounding fault line in a distribution network based on deep reinforcement learning according to claim 3 is characterized in that: The calculation formula of the self-attention mechanism is: ; Where, M is the weight matrix calculated by the self-attention mechanism, is the normalization function, is the distribution network feeder switch characteristic matrix The transposed matrix of D is the adjacency matrix of the feeder switches in the distribution network G degree matrix of ; The calculation formula of the dynamic graph convolutional network in the dynamic graph convolutional network layer is: ; Where, H is the characteristic information of the feeder switch, is the activation function, W H is the weight matrix of the dynamic graph convolutional network layer, is the Hadamard product; The calculation formula of the graph convolution network in the graph convolution network layer is: ; Where: I is the state information characteristic quantity, W I is the weight matrix of the graph convolutional network layer.

5. The method for selecting and deciding a single-phase grounding fault line in a distribution network based on deep reinforcement learning according to claim 1 is characterized in that: The generation of a fault line selection strategy based on the fault line selection order specifically includes: Generate a fault line selection table based on the fault line selection order: the fault line selection order is: no-load line (branch) switch, dual power supply line switch, reactive power compensation device, and station transformer; the dual power supply line switch selection order is: general user substation dual power supply line, general public substation dual power supply line, high-risk user substation dual power supply line, and key public substation dual power supply line. Obtain the fault phase voltage of a single-phase grounding fault in the distribution network; According to the selection order in the fault line selection table, each switch is selected at the same time to determine whether the fault phase voltage is restored after the corresponding switch is selected; If it is restored, the single-phase grounding fault line of the distribution network is obtained and the selection is completed; Otherwise, the selection and pulling sequence is continued for each switch at the next moment until a single-phase grounding fault line of the distribution network is obtained, and the selection and pulling is terminated.

6. The method for selecting and deciding a single-phase grounding fault line in a distribution network based on deep reinforcement learning according to claim 1, characterized in that: The basis stated ε-greedy Strategy, from the action set A Select an action a t , specifically including: The optimal action corresponding to the maximum Q' value is selected from the action set A with a probability of 1-ε-ε / |A|, and the optimal action corresponding to the maximum Q' value is selected from the action set A with a probability of ε / |A| A Pick a random action from ε-greedy The calculation formula of the strategy is: ; Where π(a|s) is ε-greedy Strategy function, ε is the probability of random selection and ε<1, arg maxQ' (a,s) is the maximum Q' value corresponding to the optimal action a, and |A| is the number of actions in the action set.

7. A single-phase ground fault line selection decision system for distribution network based on deep reinforcement learning, characterized in that: It includes a module for implementing the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning as described in any one of claims 1-6.

8. A device, characterized in that include: Memory; processor; as well as computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the distribution network single-phase grounding fault line selection decision method based on deep reinforcement learning as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that A computer program is stored thereon; the computer program is executed by a processor to implement a single-phase grounding fault line selection decision method for a distribution network based on deep reinforcement learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Small current grounding fault line selection method and device based on clustering and deep learning

    CN116227538A