A Fault Location Method for Distribution Networks Based on FSGATv2 and Topological Similarity
By adopting a fault positioning method based on FSGATv2 and topological similarity in the distribution network, the fault positioning problem of the distribution network under the high permeability access of distributed power supplies and the changes in the topological structure are solved, and the fault positioning effect with high accuracy and strong robustness is achieved.
Patent Information
- Application Number
- CN202510151049.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-11
AI Technical Summary
In the distribution network, traditional fault positioning methods are difficult to achieve rapid detection and positioning in the operating scenarios of new distribution networks with high permeability access to distributed power supplies, especially under topological changes and small sample conditions, the fault positioning accuracy and robustness are insufficient.
The fault location method of distribution network based on FSGATv2 and topological similarity is adopted, and the distribution network topology is represented through a graph structure. Combined with the zero-sequence current data collected by μPMU, the FSGATv2 model is built, and the topological similarity is used for classification to realize the transfer learning of the model and the optimization of dynamic loss function to adapt to topological structure changes and small sample problems.
It improves the accuracy and robustness of fault positioning in the distribution network, enhances the topological generalization and transfer learning capabilities of the model, and can achieve efficient fault positioning under different topological structures and operating conditions.
Smart Images

Figure CN119619734B_ABST
Abstract
Description
Technical Field
[0001] The present invention provides a distribution network fault location method based on FSGATv2 and topological similarity, belonging to the technical field of distribution network fault location. Background Technique
[0002] Distribution network lines have numerous branches and complex line structures. When a small current grounding system is adopted, there are problems such as weak fault current and unclear characteristics. The information provided by traditional measurement devices is incomplete, the measurement error is large, and data loss and distortion are likely to occur. In addition, due to reasons such as load switching, optimized line loss, and fault reconstruction, the network topological structure changes relatively frequently. The high penetration of distributed generation (DG) such as wind power and photovoltaic power leads to significant differences in the characteristics of active distribution networks from those of traditional distribution networks in terms of power flow distribution, fault current level, and polarity. How to achieve fast fault detection and location in the operation scenario of a new distribution network with high DG penetration is the key to improving the operation reliability of the distribution network. Traditional fault location methods such as impedance method, injection method, and traveling wave method have the advantages of being simple and easy to implement in scenarios where the distribution network structure does not change. However, when the distribution network structure changes, the method of relying on manual extraction of weak characteristic signals depends on empirical values and requires regenerating the network matrix or sample library, which not only greatly increases the calculation amount but also has a high dependence on alarm information, affecting the fault location accuracy.
[0003] In recent years, due to the ability to quickly process complex information and having good fault tolerance, data-driven deep learning methods have been attracting the attention of scholars at home and abroad. Subsequently, with the application of micro phasor measurement units (μPMUs) in distribution networks, it provides a reliable data basis for distribution network fault location technology based on deep learning. Some scholars use a fixed number of interpolations to replace the measurement values of specific branches and apply the gradient boosting tree model to the fault location problem of low-voltage intelligent distribution networks. Some scholars combine variational mode decomposition (VMD) with convolutional neural networks (CNNs) to achieve accurate fault location in networks with different line structures. Some scholars use distributed phasor measurement units (D-PMUs) and Petri nets for distribution network fault diagnosis. Some scholars, based on the phasor distribution characteristics of line voltage and current, use the measurement data of μPMUs to locate and verify single-phase grounding faults in active distribution networks. However, while the above methods using measurement device data and neural networks can accurately locate faults, they cannot take into account the changes in the distribution network topology. The operation modes of active distribution networks are complex and changeable, and the access positions and capacities of distributed power sources are prone to change during the staged development of the power system. How to effectively achieve fault location under few samples, complex topological changes, and different operating conditions is a difficult problem that urgently needs to be solved. Compared with convolutional neural networks, graph neural networks (GNNs) can extend neural network methods to the spatial graph domain and achieve the combination of deep learning and graph data. The most widely used GNNs currently include graph convolutional networks (GCNs) and graph attention networks (GATs), etc. GCNs have already had certain research in the field of power system fault diagnosis. Some scholars consider the spatial fault characteristics between nodes and propose a distribution network fault diagnosis model based on deep GCNs, which has anti-interference and generalization capabilities. Some scholars combine the measurement values of different nodes with the network topology and use GCNs to construct a distribution network fault location framework with high location accuracy. However, the above methods do not consider the generality and practical performance of the model under changes in the network topology.
[0004] Compared with GCN, GAT introduces the attention mechanism in computer vision, making it pay more attention to neighbor nodes to meet the requirements of inductive learning tasks. However, GAT is a restricted static attention mechanism, and each node processes its neighbor nodes with its own representation as the query. In the problem of distribution network fault location tasks where nodes are dynamically selected and the topological structure changes frequently, it has poor robustness to noise and topological generalization ability. Therefore, GATv2 with a dynamic attention mechanism is proposed. GATv2 has a stronger global node weight update ability and topological generalization performance, but the application of GATv2 in the field of power systems is still in its infancy. In addition, the distribution network fault location model based on graph neural network highly depends on a large number of fault data samples. However, the occurrence of faults in the distribution network is a small-sample event, and it is difficult to obtain historical operation data under various fault conditions. Moreover, due to the differences in data distribution, a fault diagnosis model trained with labeled data in one scenario cannot classify unlabeled data obtained from other scenarios. With the development of the new generation of artificial intelligence technology, transfer learning can promote the successful application trained in one scenario to other scenarios for fault diagnosis. On the basis of training with a large number of simulation samples, a machine learning model is trained, and then combined with the actual fault samples of the distribution network, the trained machine learning model is transferred to the actual fault analysis and location tasks. Some scholars considered the actual factors of weak fault current and complex working conditions in the distribution network, proposed a distribution network fault diagnosis model based on sample-based transfer learning, and verified and tested the effectiveness of the model. Some scholars proposed a distribution network fault location framework based on transfer learning, and verified the fault location effect in the target domain using the data collected by the phasor measurement unit (PMU). However, the above methods for the application of transfer learning all assume that the distributions of source domain data and target domain data are the same, without considering the differences between simulation and actual data. Summary of the Invention
[0005] In order to achieve high-precision and strong-robustness location of distribution network faults and solve the small-sample problem of topological structure changes in the process of model application from simulation to practice, the present invention proposes a distribution network fault location method based on FSGATv2 and topological similarity.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: a distribution network fault location method based on FSGATv2 and topological similarity, including the following steps:
[0007] S1: Use a graph structure to represent the actual distribution network topology, collect the zero-sequence current data of the line through μPMU as the target domain data set, and build an FSGATv2 model; wherein the FSGATv2 model is based on the GATv2 network, and an FS module for feature screening is embedded in the GATv2 network;
[0008] S2: Build a simulation model with the same operating conditions as the actual distribution network, simulate various faults and different numbers of μPMUs, collect and extract the statistical features of zero-sequence current, input them into the FS module for feature selection, and generate a labeled source domain data set;
[0009] S3: Classify distribution network topologies based on topological similarity;
[0010] S4: For the new topology distribution network, firstly perform category judgment according to step S3; then select a distribution network with a clear topology structure, complete historical data, and complete fault scenario in the corresponding category as the central distribution network, and use the FSGATv2 model to extract the transferable features of the central distribution network, guide the classifier to output the correct prediction target, and obtain the cross entropy loss function of the classifier on the source domain data set;
[0011] S5: A domain adaptive transfer learning method with optimized dynamic loss function is used to measure the data distribution difference between the source domain and the target domain. The knowledge learned by the FSGATv2 model from the source domain data is adaptively transferred to the target domain data to complete fault location in the target domain.
[0012] The graph structure of the actual distribution network topology in step S1 is represented as: G =( A , X ),in A is the adjacency matrix, which is used to represent the distribution network topology graph N The connection relationship between nodes and edges, X is the node feature matrix. When the topology of the distribution network changes, the adjacency matrix is A Changes to A’ ;
[0013] The node feature matrix X is a sequence composed of 24 types of statistical features of each node, expressed as: , X ∈ R N×24 ,in x i Indicates the node i Class statistical properties, R is a real number, N Indicates the number of nodes. Each node has a label indicating its status. 0 indicates a normal state and 1 indicates a fault state.
[0014] The 24 statistical features include:
[0015] Central tendency statistical characteristics of the time domain, including: mean, median, lower quartile and upper quartile;
[0016] Statistical features of dispersion degree in the time domain, including: minimum value, maximum value, interquartile range, standard deviation, root mean square, and square root amplitude;
[0017] Statistical features of distribution shape in the time domain, including: kurtosis, skewness, shape factor, gap coefficient, crest factor;
[0018] Characteristics of centroid frequency, mean square frequency, root mean square frequency, frequency variance, and frequency standard deviation in the frequency domain;
[0019] Characteristics of power spectrum entropy, singular value entropy, energy entropy, and permutation entropy in the time-frequency domain.
[0020] The FS module assigns high weights to the features related to each node label of the distribution network and low weights to the remaining features, performs weighted combination by multiplying the scalar with the features, and uses the Softmax function to regularize the parameters for feature screening.
[0021] The input of the FS module is the node feature matrix X , and then feature fusion is performed after passing through two network layers, where the two network layers respectively contain a linear transformation layer, a BatchNorm layer, a SumPool layer, a ReLu layer, and a fully connected layer.
[0022] The structure of the FSGATv2 model is as follows: input, FS module, 3 GAL layers, fully connected layer, output, where the input is G =( A , X ), and the output is the fault status of each node. Each GAL layer is composed of a linear transformation layer, a dynamic attention mechanism, and a fully connected layer, where the dynamic attention mechanism is a multi-head attention mechanism, integrating two steps of attention weight calculation and weighted feature aggregation.
[0023] In step S3, the topological similarity is measured by calculating the singular value and Jaccard coefficient of the distribution network topological adjacency matrix, and the Mahalanobis distance between the distribution network samples is used as the clustering criterion to classify the distribution network topology.
[0024] In step S5, the maximum mean discrepancy MMD is used to measure the distribution distance between the source domain and target domain samples, realizing the measurement of the data distribution difference between the source domain and target domain.
[0025] The expression for optimizing the dynamic loss function in step S5 is as follows:
[0026] ;
[0027] In the formula: J is the total loss function, J cis the cross-entropy loss function of the classifier on the source domain dataset, J d is the MMD distance between the source domain and target domain datasets, α is the optimization parameter.
[0028] The expression of the Jaccard coefficient is as follows:
[0029] ;
[0030] In the formula: s is the Jaccard coefficient, F 00 represents A is 0 in A’ and the number of 0s in F 01 represents A is 0 in A’ and the number of 1s in F 10 represents A is 1 in A’ and the number of 0s in F 11 represents A is 1 in A’ and the number of 1s in
[0031] By performing singular value decomposition on the adjacency matrix A ∈ R N×N to obtain the singular value sequence of the adjacency matrix, where where is A the rank of and
[0032] The beneficial effects of the present invention compared with the prior art are as follows:
[0033] 1) The present invention uses 24 types of statistical features of transient zero-sequence current as initial feature quantities, and improves the sensitivity of the model to the fault features of the distribution network with DG through the feature selection module. The GATv2 network adaptively adjusts the attention coefficients between the topological nodes of the distribution network through the dynamic graph attention mechanism, and improves the adaptability to the change of the relationship of the same topological node.
[0034] 2) The present invention classifies the distribution network topology based on topological similarity, judges the actual distribution network topology category, and improves the topological generalization ability of the fault location model to be applicable to distribution networks with completely different topological structures.
[0035] 3) The present invention establishes a general and transferable FSGATv2 distribution network fault location model, migrates the source domain location model to the target domain, uses the t-SNE algorithm to visualize the test results of the calculation examples, provides interpretability analysis for the algorithm, and improves the adaptability of the fault location method based on deep learning under the condition of few unlabeled samples. In addition, on-site experiments are carried out on the distribution network under different operating conditions to test and verify the effectiveness of the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present invention will be further described below with reference to the accompanying drawings:
[0037] Figure 1 It is a schematic structural diagram of a distribution network fault location model based on FSGATv2 and topological similarity;
[0038] Figure 2 It is a schematic structural diagram of a distribution network fault location model based on FSGATv2;
[0039] Figure 3 It is a schematic structural diagram of a feature selection module;
[0040] Figure 4 It is a flow chart of the method of the present invention;
[0041] Figure 5 It is a distribution network topology classification result diagram in an embodiment of the present invention;
[0042] Figure 6 It is a topological structure diagram of the second type of central distribution network system c in an embodiment of the present invention;
[0043] Figure 7 It is a comparison result diagram of fault location of each model in an embodiment of the present invention;
[0044] Figure 8 It is a comparison result diagram of fault location under different permeabilities in an embodiment of the present invention;
[0045] Figure 9 It is a comparison result diagram of the clustering distance and location accuracy between each distribution network and the distribution network system c in an embodiment of the present invention;
[0046] Figure 10 It is a comparison result diagram of the location accuracy of each distribution network under two schemes in an embodiment of the present invention;
[0047] Figure 11 It is a comparison result diagram of model performance in an embodiment of the present invention;
[0048] Figure 12 It is a comparison result diagram of different fault location migration schemes in an embodiment of the present invention;
[0049] Figure 13This is a visualization diagram of the classification results of different solutions in the embodiments of the present invention;
[0050] Figure 14 This is a field diagram of the 10kV distribution network grounding fault experiment in the embodiments of the present invention;
[0051] Figure 15 It is built using Matlab / Simulink Figure 14 The experimental simulation model diagram of the field diagram shown. Detailed implementation manners
[0052] As Figures 1 to 15 shown, the present invention provides a distribution network fault location method based on FSGATv2 and topological similarity, constructs a fault location model. This fault location model is based on the GATv2 network and integrates a feature selection (FS) module and a transfer learning model, and can realize three functions: distribution network topology classification, fault feature extraction, and fault location. The following will detail the existing technologies involved in the model and the improvements made to them.
[0053] 1. Distribution network topology classification
[0054] From the perspective of graph theory, the power network has a natural graph structure. Most distribution networks are radial. Without considering the electrical information of internal components, it can be abstracted as a graph composed of nodes and edges. Among them, the nodes are composed of the electrical nodes of the distribution network, and the edges are composed of overhead lines or cable lines. An accurate topological structure is the basis for fault location. To improve the economy and safety of the power grid system, the distribution network may undergo reconstruction during operation, resulting in changes in the number of nodes and edges. The economic reconstruction does not have faulty lines, and the number of nodes remains unchanged before and after reconstruction. However, the faulty reconstruction may cut off some lines and nodes, and the network topological structure will change under both reconstruction methods.
[0055] Let the given distribution network have its graph structure G =( A , X ), where A is the adjacency matrix, used to represent the connection relationship between N nodes and edges in the distribution network topology graph, A ∈ R N×N , R is a real number, X is the node feature matrix. When the distribution network topological structure changes, the adjacency matrix changes from A to A’ .
[0056] The node feature matrix XIn [the method], a transient zero-sequence current is used to indicate whether there is a fault at the node. In the present invention, a μPMU is used as a measurement device to collect the zero-sequence current of the line, and 24 types of statistical features of the zero-sequence current with rich fault features are extracted in multiple domains, including the central tendency statistical features (mean, median, lower quartile, and upper quartile) in the time domain, the dispersion statistical features (minimum value, maximum value, interquartile range, standard deviation, root mean square, and square root amplitude), and the distribution shape statistical features (kurtosis, skewness, shape factor, gap coefficient, crest factor), the center frequency, mean square frequency, root mean square frequency, frequency variance, and frequency standard deviation in the frequency domain, and the power spectrum entropy, singular value entropy, energy entropy, and permutation entropy features in the time-frequency domain. A sequence composed of 24 types of statistical features of each node is used as the feature input of the model, X ∈ R N×24 , where x i represents the i th type of statistical characteristic of the node, N represents the number of nodes. Each node has a label indicating the state, 0 represents the normal state, and 1 represents the fault state. The output of the model is the state discrimination result of the node, so as to achieve fault location.
[0057] Because the node features will pass through multiple network layers during propagation in the model, when inputting the features, the present invention improves the sensitivity of the GATv2 network to fault features by embedding the FS module. The FS module assigns high weights to the features related to the label of each node in the distribution network and low weights to the remaining features, and performs weighted combination by multiplying the scalar with the features, and uses the Softmax function to regularize the parameters and perform feature screening to reduce the interference of noise features, so as to improve the accuracy of node classification. The structure of the feature selection module is as Figure 3 shown.
[0058] For the distribution network topology graph G =( A , X ), considering the separate mapping relationship between the label and the feature in the classification problem, the set of node features is , x i , i ∈{1, 2, …, d} represents the i th type of statistical feature of the node. In this embodiment, d can take the value of 24, indicating that there are 24 types of statistical features, l k , k ∈{0, 1, …, m} represents the proportion corresponding to each type of statistical feature of each node, that is, the importance of the feature. In this embodimentm Take the value 23. Let be the scalar value multiplied by the X th i component of the node feature matrix. Before the start of model training, initialize , and normalize it with the softmax function to obtain . Multiply the normalized with the feature x i to obtain the weighted combined feature , and its expression is as follows:
[0059] (1);
[0060] (2);
[0061] In the above formula: c is a constant, and in this embodiment, its value is taken as 100, and its function is to prevent from tending to 0. is the Hadamard product of element-wise multiplication. Since the FS module needs gradients when updating parameters, feature soft selection should be used. In this way, during training, the weights of features related to node labels are close to 1, and the weights corresponding to other features are close to 0, and the influence of noise features on the positioning result becomes smaller.
[0062] The specific implementation process of the FS module is as follows: After inputting the node feature matrix X , the output feature is obtained through weighted combination, is the vector concatenation operator. Take as the new feature and input it into the first network layer to obtain the output feature . Then input into the second network layer and concatenate the output results to obtain the node feature X new after passing through the FS module.
[0063] FSGATv2 is a dynamic graph attention GATv2 network (Feature Selection graphattention network, FSGATv2) embedded with the FS module. The distribution network fault location model based on FSGATv2 is as Figure 2 shown. The distribution network topology structure corresponds to the input graph adjacency matrix. Taking the node feature matrix X of the zero-sequence current of the line as the initial feature of each node, features X new. Among them, the two network layers respectively include a linear transformation layer, a BatchNorm layer, a SumPool layer, a ReLu layer, and a fully connected layer. Then X new and the adjacency matrix A are input into 3 graph attention (GAL) layers for node feature aggregation. The GAL layer consists of a linear transformation layer, a dynamic attention mechanism, and a fully connected layer. The linear transformation layer maps X new to a higher-dimensional representation. The dynamic attention mechanism is a multi-head attention mechanism, which combines two steps of attention weight calculation and weighted feature aggregation. Figure 2 The feature aggregation process adopts a multi-head attention operation mode with 4 heads. The number of attention heads in the first two GAL layers of the FSGATv2 model of the present invention is 3, and the number of heads in the last layer is 6, which is used for classification tasks. Through multiple experiments, the learning rate is adjusted to 0.01, the batch size batchsize = 8, and the number of iterations is 50. The fully connected layer completes node classification, and finally outputs whether each node is in a fault state. The algorithm flow of FSGATv2 is shown in Table 1, where I ∈ R d represents a vector of all 1s.
[0064]
[0065] Table 1.
[0066] 2. Transfer learning theory based on topological similarity of distribution network
[0067] The high penetration of distributed power sources into the distribution network makes the network topology change complexly, bringing challenges to the fault location problem. If a general and transferable distribution network fault location model applicable to completely different topological structures can be established, this problem can be solved. When the distribution network fault location algorithm performs transfer learning in different distribution networks, the transfer effect is related to the topological similarity degree of the distribution network. Since the spectral radius of the distribution network is proportional to the number of nodes and only reflects the scale of the distribution network, while the norm reflects the connection situation of the distribution network topology and cannot comprehensively measure the topological similarity. Based on this, the present invention uses singular values and Jaccard coefficients as indicators to measure topological similarity, classifies the distribution network topology structure, and improves the transfer ability of the fault location model.
[0068] Suppose there is an orthogonal matrix , and the adjacency matrix of the distribution network topology graph A ∈ R N×N is subjected to the singular value decomposition shown in Equation (2) to obtain Equation (3). Among them, isA The singular value sequence of is A the rank of
[0069] (2);
[0070] (3).
[0071] The topological adjacency matrix of the distribution network A and A’ The Jaccard coefficient of is shown in Equation (4);
[0072] (4).
[0073] Because A and A’ only have binary data 0 and 1, the simple matching coefficient (SMC) is obtained from Equation (4), as shown in Equation (5).
[0074] (5);
[0075] In the formula: F 00 represents A the number of 0s in and A’ also 0s in, F 01 represents A the number of 0s in and A’ 1s in, F 10 represents A the number of 1s in and A’ 0s in, F 11 represents A the number of 1s in and A’ also 1s in. Generally, the adjacency matrix is sparse, resulting in class imbalance. To solve this problem, ignore the F 00 in the numerator, and the final Jaccard coefficient is shown in Equation (6).
[0076] (6).
[0077] In actual situations, the fault characteristics of distribution networks with different topologies are different, and there are significant differences in the data distributions of the training set and the test set. For a fault location model based on a neural network, the training set and the test set are from the same topology. In the case of no historical fault data or incomplete fault data, the limitations are relatively strong. To address this problem, the present invention proposes a transfer learning model based on topological similarity. This model does not require comprehensive collection of fault data for distribution networks of each topology. It migrates the FSGATv2 fault location model to different distribution networks to quickly and accurately complete the fault location task.
[0078] Transfer learning refers to using the similarity between data, tasks, or models to apply the source domain D s model to the target domain D t in a learning process. For the fault location tasks of distribution networks with different topologies, D s the distribution network with the existing topology is the one whose labels are available, D t and the distribution network with the new topology is the one whose labels are not available. Transfer learning is defined as: Given the source domain dataset and the target domain dataset , x and y are the input data and output data of the samples in the two domains respectively. By using D s the data to learn D t the mapping function on f such that D t has the smallest prediction error (measured by l ), and its definition is shown in Equation (7).
[0079] (7).
[0080] In the formula: f * is the minimum prediction error, and is the expectation.
[0081] Most existing fault location models based on transfer learning are trained using historical data or a large amount of simulation data, and then data under the same learning task is used for prediction or classification in the trained network. This process assumes that the data distributions of the source domain and the target domain are the same. However, in the actual application process, the data distribution of the new structure distribution network in the target domain is different from that of the source domain data that has been trained. It is necessary to use the domain adaptation transfer learning method to make the data distributions of the source domain and the target domain closer to improve the fault location accuracy. For the problem that the source domain and the target domain data distributions are inconsistent in the feature space for the same task, domain adaptation projects the data of the two domains into a new feature space by solving the optimal projection matrix. The present invention selects the Maximum Mean Discrepancy (MMD) to measure the distribution distance between the source domain and the target domain samples, so as to narrow the distribution distance of different data.
[0082] For the source domain dataset and the target domain dataset , the inner product form of the kernel function is used to perform a non-linear mapping on the Hilbert space, as shown in Equation (8).
[0083] (8);
[0084] In the formula: n s and n t are the sample sizes of the source domain and the target domain respectively, is the non-linear mapping function in the reconstructed kernel Hilbert space, which maps each sample to the Hilbert space associated with the kernel H , to obtain Equation (9).
[0085] (9).
[0086] Select the Gaussian kernel function shown in Equation (10) for inter-domain MMD calculation.
[0087] (10).
[0088] In the formula: is the width of the kernel function, which determines the smoothness. The domain adaptation transfer learning method embeds the MMD measurement method as a regularization method into the network training process, and optimizes the supervision criterion by training the network parameters, reducing the difference in the distribution of the source domain and the target domain feature data, thereby improving the transfer effect.
[0089] 3. Distribution network fault location model based on FSGATv2 and topological similarity
[0090] In order to achieve high-precision and strong robustness in locating distribution network faults and solve the problem of small sample sizes when the model changes from simulation to practical application, this paper proposes a distribution network fault location model based on FSGATv2 and topological similarity, such as Figure 1 As shown, the fault location model is used to locate the distribution network fault, which mainly includes the following steps:
[0091] S1: Build the FSGATv2 model; and collect the zero-sequence current data of the actual distribution network as the target domain data set;
[0092] S2: Build a simulation model with the same operating conditions as the actual distribution network, simulate various faults (different fault lines, different fault locations, different fault grounding resistances, different fault initial phase angles, etc.) and different numbers of μPMUs, collect and extract the node feature matrix of zero-sequence current , , input to the FS module for feature selection to generate a labeled source domain dataset.
[0093] S3: Calculate the singular values and Jaccard coefficient of the distribution network topology adjacency matrix to measure the topological similarity, and use the Mahalanobis distance between distribution network samples as the clustering criterion to classify the distribution network topology.
[0094] S4: For the new topology distribution network, firstly perform category judgment according to step S3. Then select a distribution network with a clear topology structure, complete historical data, and complete fault scenario in the corresponding category as the central distribution network, and use the FSGATv2 model to extract the transferable features of the central distribution network, guide the classifier to output the correct prediction target, and obtain the cross entropy loss function of the classifier on the source domain data set. J c ;
[0095] S5: Define the MMD distance between the source domain and target domain datasets J d ,According to the topological structure and data characteristics of the new distribution network in the target domain, the data distribution difference between the source domain and the target domain is measured, as shown in formula (11).
[0096] (11);
[0097] Where: and is a sample of n and The weight of M is the total number of categories in the dataset. Because linear loss is used Learning weights will cause the weights to converge to zero quickly. Therefore, the total loss function needs to be JPerform dynamic optimization, and the optimization result is shown in Equation (12).
[0098] (12);
[0099] In the formula: α is the optimization parameter. The optimized dynamic loss function is smooth and differentiable, which solves the problem that the weight quickly converges to zero. By optimizing the total loss function J transfer the knowledge learned by the model from the source domain data to the target domain data adaptively, and use the test set to complete fault location in the target domain.
[0100] 4. Experimental Verification and Analysis
[0101] 4.1 Example Design
[0102] Since the signal characteristics are weak when a single-phase grounding fault occurs in a distribution network with arc suppression coil grounded at the neutral point, in this embodiment, Matlab / Simulink software is used to build 12 distribution network models with different scales, diverse topologies, and arc suppression coil grounded as shown in Table 2.
[0103]
[0104] Table 2 Topological Structures of Each Distribution Network
[0105] The positive and zero sequence parameters of the overhead line and cable line are shown in Equations (13) and (14).
[0106] (13);
[0107] (14).
[0108] Calculate the first six singular values and Jaccard coefficients of the adjacency matrix of each distribution network, and the resulting bubble matrix diagram of the first six singular values and Jaccard coefficients of the adjacency matrix is as Figure 5 shown.
[0109] Classify the distribution network according to the first six singular values and Jaccard coefficient matrix of the adjacency matrix: a, b are the first category, c, f, i, k, l are the second category, e, g, h, j are the third category, and d is the fourth category. Build as Figure 6The second type of central distribution network system c shown in the figure adopts the neutral point grounding method through the arc suppression coil, the compensation degree is 8%, the arc suppression coil inductance L=0.3885H, there are three feeders, 26 nodes and 23 branches, of which L10, L12, L14-L23 are overhead lines, and the rest are cable lines. Each distributed power source is connected to the feeder through an isolation transformer. Nodes 1, 6, and 17 are connected to photovoltaic power with a capacity of 3MVA, and nodes 10 and 20 are connected to wind power with a capacity of 2MVA. The measurement device μPMU is installed at the feeder outlet and before the terminal nodes 6, 3, 9, 14, 17, 20 and 23 to obtain fault data.
[0110] Each branch is set to have a single-phase grounding fault at 0.8s, and the protection is reliably activated to clear the fault after 0.2s. Data sampling starts 0.1s before the fault occurs and ends 0.2s after the fault is cleared. 300 sampling data points are taken from each branch as samples, and a total of 6900 sets of data from 23 branches constitute the data sample set. In order to solve the problem of sample imbalance, the ratio of fault data to non-fault data is 1:1 after undersampling the data, and then the training set, validation set and test set are divided according to the ratio of 8:1:1, of which the training set has 5520 samples, and the validation set and test set have 690 samples each.
[0111] 4.2 Comparative Analysis of Source Domain Test Results
[0112] This embodiment adopts F 1 The score is used to evaluate the fault location model results, as shown in formula (15).
[0113] (15);
[0114] Where: True Positive T 1 Indicates samples that are actually positive and predicted to be positive; false positive T 2 Indicates samples that are actually negative but predicted to be positive; false negative T 3 Indicates samples that are actually positive but predicted to be negative; F 1 The score takes into account the accuracy of the classification model p And recall r , evaluate the classification effect of the positioning model from multiple dimensions.
[0115] 4.2.1. Analysis of the validity of the source domain model
[0116] The FSGCN (GCN network model embedded with the FS module), FSGAT (GAT network model embedded with the FS module), and FSGATv2 model are respectively used to locate single-phase grounding faults in the distribution network system c, verify the influence of different transition resistances, fault initial phase angles, and fault distances on the model accuracy, as well as the robustness of the model under the influence of noise and data loss. Data loss refers to the situation where some node measurement values are discarded and the node feature values are replaced with 0 for positioning. The positioning accuracies under different conditions are shown in Table 3.
[0117]
[0118] Table 3 Comparison of the results of the source domain fault location models under different conditions.
[0119] The results in Table 3 show that the performance of the FSGATv2 model proposed in the present invention is not affected by the fault initial phase angle and the fault location, and the highest accuracy can reach 99.98%, and it has strong adaptability to the noise of the measurement data.
[0120] 4.2.2, Analysis of the effectiveness of the source domain model
[0121] To verify the effectiveness of the FSGATv2 model of the present invention in the source domain dataset, the fault location accuracies and loss functions of the FSGATv2, GCN, GAT, FSGCN, and FSGAT models are compared, and the results are as Figure 7 shown.
[0122] Figure 7 The results show that the fault location model based on FSGAT converges in 256 rounds, and the accuracy can reach 97.2%, while the FSGATv2 fault location model converges in 124 rounds, and the accuracy can reach 99.98%. This shows that the FSGATv2 model not only speeds up the model convergence speed but also improves the fault location accuracy. Because the level of distributed power generation penetration affects the selection of fault characteristics, thus affecting the fault location accuracy. To verify the fault location capabilities of each model in different DG penetration operation scenarios, a comparative experiment was conducted in the source domain, and the results are as Figure 8 shown.
[0123] Figure 8 The results show that all kinds of fault location models can complete the location task in the initial fault scenario. However, as the DG penetration increases, the GCN and GAT models lack the adaptability to the new scenario data and cannot complete the fault location task. The positioning accuracies of the FSGCN and FSGAT models embedded with the FS module are better than those of the GCN and GAT models, while the FSGATv2 model of the present invention is not affected by the DG penetration, has strong model adaptability and topology generalization ability, and the positioning accuracy still remains at 99.97%.
[0124] 4.3. Comparative Analysis of Target Domain Test Results
[0125] During the process from simulation to actual application, problems such as topological structure changes, grounding method changes, and small sample issues still exist in the field of distribution network fault diagnosis. Collecting fault data for the actual distribution network and relying on a large amount of simulation data for training and diagnosis are time-consuming and laborious. Therefore, considering various fault situations in the target domain, first, a fault scenario is generated as the target domain dataset using the data collection method of the distribution network system c; then, the parameters of the source domain FSGATv2 model are migrated to initially construct the target domain model; finally, the MMD processing of the optimized dynamic loss function is performed on the target domain model using unlabeled target domain data to complete the construction of the target domain model, and thus the fault location of the distribution network is carried out.
[0126] 4.3.1. Analysis of the Migration Effect of Target Domain Fault Location
[0127] To verify the migration effect of the FSGATv2 fault location model on different topological structures, the pre-trained model of the distribution network system c is migrated to distribution networks with different structures. Taking the Mahalanobis distance as the clustering distance, the comparison results of the clustering distance and location accuracy between each distribution network and the distribution network system c are as Figure 9 shown.
[0128] Figure 9 The results show that the FSGATv2 fault location model of the central distribution network system c is migratable in the other 11 distribution networks, and the location accuracy is inversely proportional to the clustering distance of the distribution network system c, that is, the larger the clustering distance, the worse the migration effect. Since the clustering distance depends on the topological similarity degree of the model, for a new distribution network, first perform topological similarity analysis, and better location results will be obtained by migrating between distribution networks with a higher topological similarity degree. To analyze the migration effect of target domain fault location, 4000 groups of fault data of distribution networks a, d, f, and h are collected, and a comparative experiment of direct training and domain adaptive transfer learning is carried out using the FSGATv2 fault location model. The test sets of the two schemes are the same, and the comparison results of the fault location accuracy are as Figure 10 shown.
[0129] Figure 10 The results show that transfer learning has the ability to apply the FSGATv2 fault location model to distribution networks with different structures, and the model trained after migration is faster and more accurate in fault location. To verify the influence of the target domain test data volume on the performance of the FSGATv2 fault location model, for the distribution network topology of 48 nodes and 22 μPMUs, 2500 groups of fault samples are collected, and a comparative experiment of direct training and domain adaptive transfer learning is carried out, and the location effects of different amounts of unlabeled target domain data are verified. The results are as Figure 11 shown.
[0130] Figure 11 The results show that as the amount of data grows, the highest accuracies of the FSGATv2 fault location model are 89.94%, 94.88%, 97.89%, and 99.96% respectively. Moreover, as the amount of data increases, the location accuracy improves and the convergence speed gradually accelerates.
[0131] To verify the effectiveness of the fault location model of the present invention, a comparative experiment was conducted using the domain adaptation transfer method and the fine-tuning transfer method in comparison with the method of the present invention. The fine-tuning transfer method freezes all GAL layers and only trains the parameters of the fully connected layer to ensure the optimal deployment of the new structure distribution network. The comparison results of various schemes are as Figure 12 shown.
[0132] Figure 12 The results show that although the optimal deployment by the fine-tuning method, freezing the hidden layer, and fine-tuning the parameters of the fully connected layer can improve the applicability of the model, the FSGATv2 fault location model of the present invention integrating the domain adaptation transfer learning method has higher generalization ability on the target domain data. Figure 13 The classification results of the model on the target domain are visually displayed through the t-distributed stochastic neighbor embedding algorithm (t-SNE), where the target domain data is blue and the source domain data is red.
[0133] Although the GATv2 fault location model has certain classification and clustering capabilities, its effect is not significant. On this basis, the fine-tuning-based scheme improves the limited generalization ability, while the FSGATv2 fault location model of the present invention integrating the domain adaptation transfer learning method can achieve better classification effects on the target domain. Taking GATv2 as the basic model, the comparison of the effects of different fault location schemes on the target domain is shown in Table 4.
[0134]
[0135] Table 4 Comparison of fault location results of different schemes on the target domain.
[0136] As can be seen from Table 4, the positioning accuracy of Solution 1 on the target domain is still poor after multiple iterations. The reason is that there are significant differences in the network topologies and fault data characteristics between the target domain and the source domain, resulting in low generalization ability of the model on the target domain. In Solution 2, a feature selection module is added to the source domain model, which speeds up the model convergence. However, the positioning accuracy on the target domain has not been significantly improved. Solutions 3 and 4 fine-tune the hyperparameters of the fully connected layer. Compared with Solutions 1 and 2, although the positioning accuracy is improved, the effect is still poor. Solutions 5 and 6 embed the FS module and the domain adaptation transfer learning method, which speeds up the model convergence, improves the generalization of the model, and realizes reliable fault location.
[0137] 4.3.3, Effectiveness Analysis under Different Conditions
[0138] In the actual operation environment of the distribution network, traditional fault location methods are vulnerable to the influence of the neutral grounding method and the asymmetry of line distribution parameters. Therefore, in this embodiment, a target domain dataset is still generated for the distribution network with 48 nodes and 22 μPMUs, considering non-effective grounding methods such as arc suppression coils at the neutral point and the asymmetry of the lines, to verify the effectiveness of the method of the present invention. The fault location accuracies in various situations are shown in Tables 5 and 6.
[0139]
[0140] Table 5 Fault location accuracy with different neutral grounding methods.
[0141]
[0142] Table 6 Fault location results under the condition of asymmetric lines.
[0143] As can be seen from Tables 5 and 6, the method of the present invention can still accurately locate under different neutral grounding methods and the asymmetry of the lines, and is not affected by the unbalanced parameters of the distribution lines, having good generalization ability.
[0144] 4.4, Experimental Verification
[0145] Considering the differences between the simulation test and the actual distribution network, in order to better simulate the fault location scenario of single-phase grounding faults occurring in the actual distribution network, the 10 kV distribution network grounding fault diagnosis experiment as shown in Figure 14 was carried out. A total of 4800 fault samples were collected in the experiment, and the samples included different grid connection methods of DGs and the situation where the grid connection power of DGs at the same point changed, so as to verify the effectiveness of the method of the present invention.
[0146] Use Matlab / Simulink to build Figure 14 the on-site experimental simulation model. This line is composed of a mixture of overhead lines and cable lines, as shown in Figure 15as shown
[0147] The simulation model is used for data collection and model training. The actual data collected from on-site experiments is used to verify the reliability of the method of the present invention. Among them, the photovoltaic power generation adopts a reversible grid-connected system. The fault location results under various fault conditions are shown in Table 7. In the table, the fault type AG is the A-phase ground fault, BG is the B-phase ground fault, and CG is the C-phase ground fault.
[0148]
[0149] Table 7 Fault location results under on-site experimental conditions.
[0150] The results in Table 7 show that the method of the present invention can accurately achieve fault location in the actual distribution network, and the experiment verifies the effectiveness of the proposed algorithm for the fault location problem of distribution networks with few samples and topological changes.
[0151] The present invention aims at a distribution network containing DGs, establishes a new method for distribution network fault location based on topological similarity and FSGATv2 domain adaptation transfer learning, and verifies it by using distribution networks with different structures and fault diagnosis experiments. By means of the FS module, the disadvantages that the input features of the traditional model are single and highly dependent on the input data are eliminated. Under the conditions of different topological structures and non-effective grounding methods, accurate fault location can still be achieved, which has strong robustness and practicability, and provides a new idea for the research on the problems of few samples and topological changes in distribution network fault location.
[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A distribution network fault location method based on FSGATv2 and topology similarity, characterized by: The following steps are involved: S1: Use graph structure to represent the actual distribution network topology, collect line zero-sequence current data as the target domain data set through μPMU, and build FSGATv2 model; FSGATv2 model is based on GATv2 network, and FS module for feature screening is embedded in GATv2 network; S2: Build a simulation model with the same operating conditions as the actual distribution network, simulate various faults and different numbers of μPMUs, collect and extract the statistical features of zero-sequence current, input them into the FS module for feature selection, and generate a labeled source domain data set; S3: Classify distribution network topologies based on topological similarity; S4: For the new topology distribution network, firstly perform category judgment according to step S3; then select a distribution network with a clear topology structure, complete historical data, and complete fault scenario in the corresponding category as the central distribution network, and use the FSGATv2 model to extract the transferable features of the central distribution network, guide the classifier to output the correct prediction target, and obtain the cross entropy loss function of the classifier on the source domain data set; S5: A domain adaptive transfer learning method with optimized dynamic loss function is used to measure the data distribution difference between the source domain and the target domain. The knowledge learned by the FSGATv2 model from the source domain data is adaptively transferred to the target domain data to complete fault location in the target domain.
2. A distribution network fault location method based on FSGATv2 and topology similarity according to claim 1, characterized in that: The graph structure of the actual distribution network topology in step S1 is represented as: G =( A , X ),in A is the adjacency matrix, which is used to represent the distribution network topology graph N The connection relationship between nodes and edges, X is the node feature matrix. When the topology of the distribution network changes, the adjacency matrix is A Changes to A’ ; The node feature matrix X is a sequence composed of 24 types of statistical features of each node, expressed as: , X ∈ R N×24 ,in x i Indicates the node i Class statistical properties, R is a real number, N Indicates the number of nodes. Each node has a label indicating its status. 0 indicates a normal state and 1 indicates a fault state. The 24 statistical features include: Central tendency statistical characteristics of the time domain, including: mean, median, lower quartile and upper quartile; Dispersion statistics in the time domain, including minimum, maximum, interquartile range, standard deviation, root mean square, and square root amplitude; The statistical characteristics of the distribution shape in the time domain include: kurtosis, skewness, shape factor, gap coefficient, and crest factor; The centroid frequency, mean square frequency, root mean square frequency, frequency variance and frequency standard deviation characteristics in the frequency domain; Power spectrum entropy, singular value entropy, energy entropy and permutation entropy characteristics in time-frequency domain.
3. A distribution network fault location method based on FSGATv2 and topology similarity according to claim 2, characterized in that: The FS module assigns high weights to the features related to each node label in the distribution network and low weights to the remaining features. It performs weighted combination by multiplying the scalar with the features, and uses the Softmax function to regularize the parameters for feature screening.
4. A distribution network fault location method based on FSGATv2 and topology similarity according to claim 3, characterized in that: The input of the FS module is the node feature matrix X , and then feature fusion is performed after passing through two network layers, where the two network layers respectively include a linear transformation layer, a BatchNorm layer, a SumPool layer, a ReLu layer, and a fully connected layer.
5. A distribution network fault location method based on FSGATv2 and topology similarity according to claim 4, characterized in that: The structure of the FSGATv2 model is as follows: input, FS module, 3 GAL layers, fully connected layer, output, where the input is G =( A , X ), the output is the fault status of each node. Each GAL layer is composed of a linear transformation layer, a dynamic attention mechanism and a fully connected layer. The dynamic attention mechanism is a multi-head attention mechanism that integrates the two steps of attention weight calculation and weighted feature aggregation.
6. A distribution network fault location method based on FSGATv2 and topology similarity according to claim 2, characterized in that: In step S3, the singular values and Jaccard coefficients of the distribution network topology adjacency matrix are calculated to measure the topological similarity, and the Mahalanobis distance between distribution network samples is used as the clustering criterion to classify the distribution network topology.
7. A distribution network fault location method based on FSGATv2 and topology similarity according to claim 1, characterized in that: In step S5, the maximum mean difference (MMD) is used to measure the distribution distance between the source domain and the target domain samples, thereby measuring the data distribution difference between the source domain and the target domain.
8. A distribution network fault location method based on FSGATv2 and topology similarity according to claim 7, characterized in that: The expression of the optimized dynamic loss function in step S5 is as follows: ; Where: J is the total loss function, J c is the cross entropy loss function of the classifier on the source domain dataset, J d is the MMD distance between the source domain and target domain datasets, α To optimize the parameters.
9. The method for locating distribution network faults based on FSGATv2 and topology similarity according to claim 6, characterized in that: The expression of Jaccard coefficient is as follows: ; Where: s is the Jaccard coefficient, F 00 express A is 0 and A’ The number of 0s in F 01 express A is 0 and A’ The number of 1s in F 10 express A is 1 and A’ The number of 0s in F 11 express A is 1 and A’ The number of 1s in .
10. A distribution network fault location method based on FSGATv2 and topology similarity according to claim 6, characterized in that: By the adjacency matrix A ∈ R N×N conduct The singular value decomposition of the adjacency matrix is obtained by ,in yes A rank, is an orthogonal matrix.