Arc Fault Grounding Identification Method for Distribution Networks Based on Waveform Subsequence Segmentation-Clustering

By improving the LSTM algorithm and feature dimensionality reduction technology, the zero-sequence voltage waveform is segmented and feature extracted, and combined with the three-phase voltage waveform, a fault identification model is built, which solves the problem of low accuracy in arc ground recognition in the existing methods, and achieves higher identification accuracy.

CN115828150BActive Publication Date: 2025-08-05STATE GRID JIANGXI ELECTRIC POWER CO LTD RES INST +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211325243.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2025-08-05
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

The existing arc ground recognition methods of distribution networks have shortcomings in data segmentation and feature extraction, resulting in low recognition accuracy and easy to misjudgment in single waveform feature judgments.

Method used

Using a method based on waveform subsequence segmentation-clustering, the zero-sequence voltage waveform is segmented by improving the LSTM algorithm, combined with the time-frequency feature extraction of the three-phase voltage and zero-sequence voltage waveform, and feature dimensionality reduction is performed to build a fault identification model.

Benefits of technology

It effectively improves the problem of insufficient feature extraction in traditional methods, improves the recognition accuracy of arc grounding faults, reduces the impact of different types of fault data on waveform feature values, and achieves more accurate fault type distinction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115828150B_ABST
    Figure CN115828150B_ABST
Patent Text Reader

Abstract

This invention discloses a method for identifying arc grounding faults in distribution networks based on waveform subsequence segmentation and clustering. The method comprises: waveform subsequence segmentation based on an improved long-short-term memory network; combined time series feature analysis and feature dimensionality reduction of three-phase voltage and zero-sequence voltage waveforms; and arc grounding fault identification based on K-means. An improved LSTM algorithm is used to accurately segment different types of fault data within the same recorded waveform, reducing their impact on waveform eigenvalues. A fault identification model is then established that combines time series feature extraction of three-phase voltage and zero-sequence voltage waveforms. Through experimental data analysis, the boundary conditions between arc grounding faults and other faults are determined, effectively improving the shortcomings of traditional feature extraction based on current analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power distribution automation, and in particular relates to a distribution network arc grounding identification method based on waveform subsequence segmentation-clustering. Background Art

[0002] According to surveys, over 80% of power outage losses are caused by distribution network faults. Therefore, distribution network fault diagnosis has long been a research topic for power supply companies. Arc grounding faults, among other grounding faults, are difficult to detect, pose significant risks, and have complex mechanisms. In recent years, with the development of smart grids in my country, fault feature extraction and fault identification from distribution network fault data have become two key areas of research interest.

[0003] The existing distribution network arc grounding identification method has the following technical defects: 1) It does not consider the data segmentation of developing faults, and the use of overall data analysis will cause confusion in the fault characteristics, resulting in inaccurate feature extraction and reduced identification accuracy; 2) It mainly uses zero-sequence voltage waveform for feature analysis to identify faults, but some fault types have similar characteristics, and single waveform feature judgment is prone to misjudgment. Summary of the Invention

[0004] The present invention provides a distribution network arc grounding identification method based on waveform subsequence segmentation-clustering, which is used to solve the above-mentioned technical problems.

[0005] The present invention provides a distribution network arc grounding identification method based on waveform subsequence segmentation-clustering, comprising: obtaining original zero-sequence voltage waveform data, selecting a zero-sequence voltage sequence with a length of w and calculating its standard deviation, which is recorded as a standard deviation sequence A standard deviation sequence of length u is selected and its short-term memory index is calculated, which is recorded as a short-term memory index sequence h = [h0, h1, ...]; the size of each short-term memory index in the short-term memory index sequence is compared with the fault index threshold δ, and from the first fault moment in the short-term memory index sequence that is greater than the fault index threshold δ, a subsequence of length L in the short-term memory index sequence is selected and stored as a temporary index sequence (h g , h g+L ); Determine the supplementary index sequence of the short-term memory index sequence (h g+L , h g+2L-1 ) whether there is at least one data point with a fault index value |h t | is greater than the fault index threshold δ, and the difference between the fault index value of a fault data point and the fault index value of another fault data point|h t -h t-1 | is greater than the category threshold difference λ, wherein the supplementary indicator sequence (h g+L , hg+2L-1 ) is equal to L; if the supplementary index sequence (h g+L , h g+2L-1 ) has at least one data point fault index value |h t | is greater than the fault index threshold δ, and the difference between the fault index value of a fault data point and the fault index value of another fault data point |h t -h t-1 | is greater than the category threshold difference λ, then the supplementary indicator sequence (h g+L , h g+2L-1 ) before a certain fault data point and the temporary indicator sequence to obtain the final fault indicator sequence, and the supplementary indicator sequence (h g+L , h g+2L-1 ) is used as the starting point of another temporary indicator sequence; the final fault indicator sequence after segmentation is mapped to the zero-sequence voltage sequence to obtain the fault waveform subsequence, where the starting point sp and the end point ep of the final fault indicator sequence (sp, ep) correspond to the zero-sequence voltage sequence as follows: Wherein, start and end represent the starting point and end point of the fault waveform subsequence in the zero-sequence voltage, respectively; the time-frequency features of the fault waveform subsequence are extracted by combining the three-phase voltage waveform and the zero-sequence voltage waveform, and the extracted time-frequency features are subjected to feature dimensionality reduction processing to obtain the fault target features of the fault waveform subsequence; the fault target features are analyzed based on a preset fault identification model to obtain the fault type of the fault waveform subsequence corresponding to the fault target features.

[0006] In some embodiments of the present invention, when determining the supplementary index sequence (h g+L , h g+2L-1 ) whether there is at least one data point with fault value h t | is greater than the fault index threshold δ, and the difference between the fault value of a fault data point and the fault value of another fault data point |h t -h i-1 | is greater than the category threshold difference λ, the method further includes: if the supplementary indicator sequence (h g+L , h g+2L-1 ) has at least one data point in the index value |h t | is greater than the fault index threshold δ, and the difference between the index value of any fault data point and the index value of other fault data points |h t -h t-1 | is not greater than the category threshold difference λ, then the supplementary indicator sequence (h g+L , h g+2L-1) before a certain fault data point and the temporary indicator sequence to obtain the final fault indicator sequence, and the supplementary indicator sequence (h g+L , h g+2L-1 ) is used as the starting point of another temporary indicator series.

[0007] In some embodiments of the present invention, when determining the supplementary index sequence (h g+L , h g+2L-1 ) whether there is at least one data point with fault value h t | is greater than the fault index threshold δ, and the difference between the fault value of a fault data point and the fault value of another fault data point |h t -h t-1 | is greater than the category threshold difference λ, the method further includes: if the supplementary indicator sequence (h g+L , h g+2L-1 ) does not contain at least one fault value of the data point |g t | is greater than the fault index threshold δ, the temporary index sequence is directly defined as the final fault index sequence.

[0008] In some embodiments of the present invention, the fault index value h is calculated according to the improved LSTM algorithm. t The specific calculation process is: perform standard deviation processing on the input time series to obtain the standard deviation series The expression for calculating the standard deviation is: Where Z i and are the ith original zero-sequence voltage signal and the mean of the n consecutively selected original zero-sequence voltage signals respectively; when calculating the short-term memory index, assume that the input at the current moment is The response output of the current input is the short-term memory h t , where x t The value length is u, and the short-term memory h is calculated. t The expression is: Where, f t 、i t are the state results of the forget gate and the input gate respectively, σ is the activation function, W f is the connection weight, x t is the mean sequence input of the zero-sequence voltage signal at the current moment, h t-1 is the short-term memory output at time t-1, tanh is the activation function, h t is the short-term memory output at time t, b f is the bias phase, W c is the input state weight, bc is the bias term, Multiply the corresponding elements of vectors of the same shape. is the unit state input at time t.

[0009] In some embodiments of the present invention, the feature dimensionality reduction processing of the extracted time-frequency features includes: performing feature dimensionality reduction processing on the extracted time-frequency features based on principal component analysis, specifically: inputting m pieces of n-dimensional data, and forming the eigenvalues into an n-row and m-column matrix by column. in, are the m eigenvectors of the nth fault; each row of the matrix φ is zero-meaned, that is, the mean of this row is subtracted; and the covariance matrix E is calculated, where the expression of the covariance matrix E is: Where, is the eigenvector and eigenvectors The covariance of is the eigenvector and eigenvectors covariance; calculate the eigenvalues and eigenvectors of the covariance matrix E, sort the eigenvalues from large to small, select the eigenvectors corresponding to the largest N eigenvalues as row vectors, and form the eigenvector matrix P; transform the eigenvector matrix P into a new space constructed by N eigenvectors, that is, Q = P*φ, where Q is the eigenvector matrix after dimensionality reduction.

[0010] In some embodiments of the present invention, the process of constructing a preset fault identification model includes: determining the value of k, where k means the number of clustered classes; randomly selecting k initial cluster centers, and randomly selecting k centroid vectors {μ1, μ2, ..., μ k}, the coordinate selection formula of the centroid vector is: Where μ [kx] 、μ [ky] 、μ [kz] are the X coordinate, Y coordinate, and Z coordinate of k centroids, mm u 、mm v 、mm w The minimum value in the X coordinate, the minimum value in the Y coordinate, and the minimum value in the Z coordinate, respectively. u 、range v 、range w The difference between the maximum and minimum values of the X coordinate, the difference between the maximum and minimum values of the Y coordinate, and the difference between the maximum and minimum values of the Z coordinate are respectively, and rad() is a random number between (0, 1); assign sample points, update the cluster center, and calculate sample Q ξ (ξ=1,2,...m) The distance from each centroid vector Among them, Q ξ is the ξth sample; assuming that the cluster is divided into (θ1, θ2 , ...θ k ), then Q ξ Record the category θ with the smallest distance k At this time, the center of mass μ is recalculated and updated k , the expression is: Where θ k is the k-th data after clustering, len(θ k ) is θ k The steps of allocating sample points and updating cluster centers are repeated until all sample points are allocated, the categories of all sample points do not change, or the number of iterations reaches the specified maximum value, then clustering is stopped.

[0011] The arc grounding identification method for distribution network based on waveform subsequence segmentation and clustering of the present application has the following beneficial effects:

[0012] 1. The improved LSTM algorithm is used to accurately segment different types of fault data in the same recorded data, reducing the impact of different types of fault data on the waveform characteristic values;

[0013] 2. A fault identification model that combines the time series feature extraction of three-phase voltage and zero-sequence voltage waveforms was established. Through experimental data analysis, the boundary conditions between arc grounding faults and other faults were obtained, effectively improving the problem of insufficient feature extraction based on traditional current analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0015] Figure 1 A flowchart of a method for identifying arc grounding in a distribution network based on waveform subsequence segmentation and clustering is provided in accordance with an embodiment of the present invention;

[0016] Figure 2 A detailed diagram of an improved LSTM unit according to a specific embodiment of the present invention;

[0017] Figure 3 A schematic diagram of waveform subsequence segmentation according to a specific embodiment of the present invention is provided;

[0018] Figure 4A schematic diagram of an arc grounding safety boundary according to a specific embodiment of the present invention;

[0019] Figure 5 A schematic diagram of the relative positions of four types of sample data and security boundaries in a specific embodiment provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] See also Figure 1 , which shows a flow chart of a distribution network arc grounding identification method based on waveform subsequence segmentation-clustering of the present application.

[0022] like Figure 1 As shown in FIG, the distribution network arc grounding identification method based on waveform subsequence segmentation-clustering specifically includes the following steps:

[0023] Step S101: Obtain the original zero-sequence voltage waveform data, select a zero-sequence voltage sequence with a length of w and calculate its standard deviation, which is recorded as the standard deviation sequence And select the standard deviation sequence of length u and calculate its short-term memory index, which is recorded as the short-term memory index sequence h = [h0, h1, ...]

[0024] Step S102: compare the magnitudes of the short-term memory indicators in the short-term memory indicator sequence with the fault indicator threshold δ, and select a subsequence of length L in the short-term memory indicator sequence from the first fault moment in the short-term memory indicator sequence that is greater than the fault indicator threshold δ and store it as a temporary indicator sequence (h g , h g+L ).

[0025] Step S103, determining the supplementary index sequence (h g+L , h g+2L-1 ) whether there is at least one data point with a fault index value |h t | is greater than the fault index threshold δ, and the difference between the fault index value of a fault data point and the fault index value of another fault data point|h i -h t-1 | is greater than the category threshold difference λ, wherein the supplementary indicator sequence (h g+L , hg+2L-1 ) is equal to L.

[0026] In this embodiment, since the ground fault is developmental, it is possible to evolve from one type of fault to another type of fault, such as Figure 2 As shown in Figure 1, if the fault segment data of the oscilloscope is directly analyzed, the fault feature analysis may be confused. Therefore, the waveform subsequence segmentation algorithm based on the improved LSTM is used to segment each fault. First, since the ground fault data has strong periodicity and the long-term impact is small, the long-term memory C is mixed. i and short-term memory t , using only h t Secondly, compared to the traditional LSTM (Long Short-Term Memory) algorithm, the improved LSTM structure no longer requires separate storage for forgotten and remembered information. Instead, a single structure is used. The forget gate's output signal is deducted from the value of 1 to select the state of the memory gate, indicating that only the state of the information to be forgotten is updated. This reduces the complexity of the traditional LSTM model.

[0027] like Figure 3 As shown, the standard deviation of the input time series is processed: The sliding step is S a , Z i and are the ith original zero-sequence voltage signal and the mean of the n consecutively selected original zero-sequence voltage signals. When calculating the short-term memory index, assume that the input at the current moment is The sliding step is S b , x i The value length is u. Then the response output of the current input is the short-term memory h i The calculation formulas for the forget gate and input gate of the gate control unit are as follows:

[0028] f t =σ(W f ·[h t-1, x t ]+b f ),

[0029] i t =1-f t ,

[0030] Where, f t 、i t are the state results of the forget gate and the input gate respectively, σ is the activation function, W f is the connection weight, b f is the bias phase, Multiply the corresponding elements of vectors of the same shape, xt For, h t-1 for.

[0031] In the improved LSTM structure, the input gate determines x t For h t The output gate and the unit state jointly determine the unit output value, and the calculation formula is as follows:

[0032]

[0033]

[0034] Where b c is the bias term, W c is the input state weight is the unit state input at time t.

[0035] Given the fault index threshold δ and the category threshold difference λ, if |h t |>δ, then it is considered that a fault has occurred in the segment. t -h t-1 |>λ, it is considered that different types of faults have occurred.

[0036] Step S104: If the supplementary index sequence (h g+L , h g+2L-1 ) has at least one data point fault index value |h t | is greater than the fault index threshold δ, and the difference between the fault index value of a fault data point and the fault index value of another fault data point |h t -h t-1 | is greater than the category threshold difference λ, then the supplementary indicator sequence (h g+L , h g+2L-1 ) before a certain fault data point and the temporary indicator sequence to obtain the final fault indicator sequence, and the supplementary indicator sequence (h g+L , h g+2L-1 ) is used as the starting point of another temporary indicator series.

[0037] In this embodiment, see Figure 3 , step (1), from the time the fault starts Starting from this, a sequence of length L is selected and temporarily stored as a temporary index sequence (h g , h g+L ).

[0038] Step (2), if the supplementary index sequence (h g+L , h g+2L-1 ) contains fault data, and the fault type is the same as the fault type in the temporary indicator sequence, that is, |ht -h t-1 |≤λ, then all data points before the last fault point (θ+b+1) in the supplementary indicator sequence are merged with the temporary indicator sequence and the temporary indicator sequence is replaced; if there is fault data in the supplementary sequence, but the fault type is different from that in the temporary indicator sequence, that is, |h t -h t-1 |>λ, all data points before the different point are merged with the temporary indicator sequence and the temporary indicator sequence is replaced; if there is no fault point in the supplementary sequence, the temporary indicator sequence is split into the final fault indicator sequence and labeled.

[0039] Step (3): Repeat step (2) until all waveform subsequences are segmented.

[0040] In an application scenario, the fault detection and segmentation accuracy of this method was compared with existing methods such as wavelet transform and modal decomposition for different experimental samples, as shown in Table 1. The results show that this method achieved 100% fault detection accuracy. Variational modal decomposition was unable to detect ferroresonant faults, while wavelet transform was unable to accurately detect general ground faults. Furthermore, when performing sequence segmentation, the wavelet transform was unable to accurately detect the key points of arc-flash ground faults. Consequently, the start of the next arc-flash ground fault was used as the end point of the previous fault, resulting in general ground faults being segmented as arc-flash ground faults.

[0041] Table 1 Comparison of fault segmentation accuracy of different methods

[0042]

[0043] Step S105, the final fault indicator sequence after segmentation is mapped to the zero-sequence voltage sequence to obtain a fault waveform subsequence, wherein the starting point sp and the end point ep of the final fault indicator sequence (sp, ep) correspond to the zero-sequence voltage sequence as follows: Where start and end represent the starting point and end point of the fault waveform subsequence in the zero-sequence voltage, respectively.

[0044] Step S106 , extracting the time-frequency features of the fault waveform subsequence by combining the three-phase voltage waveform and the zero-sequence voltage waveform, and performing feature dimensionality reduction processing on the extracted time-frequency features to obtain the fault target features of the fault waveform subsequence.

[0045] In this embodiment, the segmented fault waveform subsequences are subjected to time-frequency domain feature extraction, primarily including the mean, variance, peak-to-peak value, and kurtosis in the time domain. The mean represents the average level of the data, the variance measures the degree of data dispersion, the peak-to-peak value represents the difference between the highest and lowest signal values within a cycle, and the kurtosis coefficient reflects the distribution characteristics of the vibration signal. Calculating the kurtosis coefficient using the fourth power can reduce the impact of noise and improve the signal-to-noise ratio. In the frequency domain, the harmonic content, center of gravity frequency, frequency standard deviation, and root mean square frequency are also extracted. Harmonic content is obtained by subtracting the fundamental component from the AC quantity. The voltage in the power grid is primarily 50 Hz, but higher frequency signals may appear in certain situations. When the frequency of a harmonic signal is an odd multiple of the fundamental frequency, the harmonic is called an odd harmonic. The center of gravity frequency describes the frequency of the larger signal component in the signal spectrum, reflecting the signal power spectrum. The frequency standard deviation describes the dispersion of the power spectrum energy distribution. The root mean square frequency is the arithmetic square root of the mean square frequency and can be considered the radius of inertia.

[0046] Table 2 shows the time- and frequency-domain characteristic parameters of some fault waveform subsequences. If zero-sequence voltage alone is used for identification, the peak-to-peak and third-harmonic content cannot distinguish ferroresonance faults from arc grounding faults, and the root mean square frequency cannot distinguish arc grounding faults from normal conditions. The fifth and seventh harmonic content cannot distinguish the three types of faults. If three-phase voltage alone is used for identification, the variance cannot distinguish ferroresonance from normal conditions, the center of gravity frequency cannot distinguish arc grounding faults from ferroresonance faults, and the mean cannot distinguish all three types of faults. Therefore, a combined zero-sequence voltage and three-phase voltage approach is used for fault identification. Indicators such as variance and peak-to-peak value can distinguish between several types of faults.

[0047] Table 2 Time domain and frequency domain characteristic parameters of some fault waveform subsequences

[0048]

[0049] Furthermore, principal component analysis uses linear algebra to reduce data dimensionality, converting multiple variables into a small number of uncorrelated composite variables to more comprehensively reflect the entire data set. These composite variables are called principal components, and each principal component is independent of the others, meaning that the information they represent does not overlap. The steps of principal component analysis are as follows:

[0050] Input m pieces of n-dimensional data and organize the eigenvalues into a matrix of n rows and m columns. in, are the m eigenvectors of the nth fault;

[0051] Zero-mean each row of the matrix φ, that is, subtract the mean of this row;

[0052] Calculate the covariance matrix E, where the expression of the covariance matrix E is:

[0053]

[0054] Where, is the eigenvector and eigenvectors The covariance of is the eigenvector and eigenvectors covariance of

[0055] Calculate the eigenvalues and eigenvectors of the covariance matrix E, sort the eigenvalues from large to small, select the eigenvectors corresponding to the largest N eigenvalues as row vectors, and form the eigenvector matrix P:

[0056] The eigenvector matrix P is converted to a new space constructed by N eigenvectors, that is, Q = P*φ, where Q is the eigenvector matrix after dimensionality reduction.

[0057] It should be noted that Table 3 shows the factor loading coefficients for each principal component. The indicators in the table are standardized variance, kurtosis, peak-to-peak value, mean, center of gravity frequency, frequency standard deviation, root mean square frequency, third harmonic, fifth harmonic, and seventh harmonic. As can be seen from the table, principal component 1 has a strong positive correlation with variance; principal component 2 has a strong negative correlation with root mean square frequency and a strong positive correlation with frequency standard deviation and center of gravity frequency; principal component 3 has a strong negative correlation with center of gravity frequency and a strong positive correlation with frequency standard deviation.

[0058] Table 3 Factor load coefficient table

[0059]

[0060] Step S107 : analyzing the fault target feature based on a preset fault identification model to obtain the fault type of the fault waveform subsequence corresponding to the fault target feature.

[0061] In this embodiment, the process of constructing a preset fault identification model includes:

[0062] Determine the value of k, which means the number of aggregated classes.

[0063] Randomly select k initial cluster centers and randomly select k centroid vectors {μ1, μ2, ..., μ k}, the coordinate selection formula of the centroid vector is:

[0064]

[0065] Where μ[kx] 、μ [ky] 、μ [kz] are the X, Y, and Z coordinate values of k centroids, mm u 、mm v 、mm w The minimum value in the X coordinate, the minimum value in the Y coordinate, and the minimum value in the Z coordinate, respectively. u 、range v 、range w are the difference between the maximum and minimum values of the X coordinate, the difference between the maximum and minimum values of the Y coordinate, and the difference between the maximum and minimum values of the Z coordinate, respectively. rand() is a random number between (0, 1);

[0066] Assign sample points, update cluster centers, and calculate sample Q ξ (ξ=1,2,...m) The distance from each centroid vector Among them, Q ξ is the ξth sample;

[0067] Assume that the clusters are divided into (θ1, θ2, ...θ k ), then Q ξ Record the category θ with the smallest distance k At this time, the center of mass μ is recalculated and updated k , the expression is:

[0068]

[0069] Where θ k is the k-th data after clustering, len(θ k ) is θ k The number of samples in ;

[0070] Repeat the steps of assigning sample points and updating cluster centers until all sample points are assigned, the categories of all sample points do not change, or the number of iterations reaches the specified maximum value, then stop clustering.

[0071] Furthermore, after training the training samples, the distance between each sample and each cluster center is shown in Table 4. The distance between the training data and each cluster center is calculated, and the data is divided into the class with the closest distance. The last column is the fault type identification. Class 1 is arc grounding fault, Class 2 is ferromagnetic resonance fault, and Class 3 is normal.

[0072] Table 4 Distance between training samples and cluster centers

[0073]

[0074] It should be noted that if the fault data falls within the safety boundary, the fault type can be accurately identified. The center of the arc grounding fault safety boundary is the cluster center of the data classified as arc grounding fault type after clustering. The equatorial radius and polar radius of the safety boundary are shown as follows:

[0075]

[0076] Where ga is the equatorial radius, gb and gc are the polar radii, max(p / q / r) are the maximum values of the clustered arc grounding fault type data along the coordinate axis X / Y / Z, and min(p / q / r) are the minimum values of the clustered arc grounding fault type data along the coordinate axis X / Y / Z.

[0077] The equatorial radius and polar radius of the safety boundary are calculated. The equatorial radius of the safety boundary is calculated to be 111.0 and 119.5 respectively, and the polar radius is 29.7. Figure 4 The middle spherical area represents the safety boundary for arc grounding faults, and the sample points represent data that has been clustered and classified as arc grounding faults. As can be seen from the figure, the vast majority of arc grounding fault data falls within the safety boundary, indicating that the safety boundary can accurately distinguish arc grounding faults from other faults.

[0078] For arc high-resistance grounding faults, the method proposed in this paper is used to segment, extract features, and reduce the dimension of the fault waveform subsequence to obtain the principal component components of the four types of samples after dimensionality reduction as shown in Table 5. Among them, the high-resistance grounding samples and low-resistance grounding samples are both within the safety boundary, while the normal situation samples and ferromagnetic resonance fault samples are both outside the safety boundary. Figure 5 It can be seen that the method proposed in this paper can also accurately identify arc high-resistance grounding faults.

[0079] Table 5 Principal component components of four types of samples after dimensionality reduction

[0080]

[0081] Table 6 shows the accuracy of arc-fault identification using different methods. The accuracy of the algorithms listed in the table is greater than 90%. The proposed arc-fault identification algorithm achieves an accuracy of 97.12%, which is 3.31% higher than that of the sliding window convolutional neural network (WT-CNN) and 2.38% higher than that of the variational mode decomposition-support vector machine (VMD-SVM), respectively.

[0082] Table 6 Identification performance of different algorithm models

[0083]

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A distribution network arc grounding identification method based on waveform subsequence segmentation and clustering, characterized in that: include: Get the original zero-sequence voltage waveform data, select the zero-sequence voltage sequence with a length of w and calculate its standard deviation, which is recorded as the standard deviation sequence , and select the standard deviation sequence of length u and calculate its short-term memory index, recorded as the short-term memory index sequence ; Compare each short-term memory indicator in the short-term memory indicator sequence with the fault indicator threshold The size of the short-term memory indicator sequence is greater than the fault indicator threshold From the moment of failure, select the short-term memory index sequence with a length of A subsequence of ; Determine the supplementary index sequence of the short-term memory index sequence Is there a fault indicator value for at least one data point in Greater than the fault indicator threshold , and the difference between the fault index value of a certain fault data point and the fault index value of another fault data point Is it greater than the category threshold difference? , wherein the supplementary index sequence The length is equal to , calculate the fault index value according to the improved LSTM algorithm ; If the supplementary index sequence There is at least one fault indicator value in Greater than the fault indicator threshold , and the difference between the fault index value of a certain fault data point and the fault index value of another fault data point Greater than the category threshold difference , then the supplementary index sequence All data points before a certain fault data point are combined with the temporary indicator sequence to obtain the final fault indicator sequence, and the supplementary indicator sequence is A data point after a certain fault data point is used as the starting point of another temporary indicator series; The final fault indicator sequence after segmentation is mapped to the zero-sequence voltage sequence to obtain the fault waveform subsequence, where the final fault indicator sequence is The corresponding relationship between the starting point sp and the end point ep and the zero-sequence voltage sequence is: ,in and They represent the starting point and end point of the fault waveform subsequence in the zero-sequence voltage respectively; Extracting the time-frequency features of the fault waveform subsequence by combining the three-phase voltage waveform and the zero-sequence voltage waveform, and performing feature dimensionality reduction processing on the extracted time-frequency features to obtain the fault target features of the fault waveform subsequence; The fault target feature is analyzed based on a fault identification model to obtain a fault type of a fault waveform subsequence corresponding to the fault target feature.

2. The method for identifying arc grounding in distribution networks based on waveform subsequence segmentation and clustering according to claim 1, characterized in that: In determining the supplementary indicator sequence of the short-term memory indicator sequence Is there a fault value for at least one data point in Greater than the fault indicator threshold , and the difference between the fault value of a fault data point and the fault value of another fault data point Is it greater than the category threshold difference? Afterwards, the method further includes: If the supplementary index sequence There is at least one indicator value in Greater than the fault indicator threshold , and the difference between the index value of any fault data point and the index value of other fault data points Not greater than the category threshold difference , then the supplementary index sequence All data points before a certain fault data point are combined with the temporary indicator sequence to obtain the final fault indicator sequence, and the supplementary indicator sequence is A data point after a certain fault data point is used as the starting point of another temporary indicator series.

3. The method for identifying arc grounding in distribution networks based on waveform subsequence segmentation and clustering according to claim 1, characterized in that: In determining the supplementary indicator sequence of the short-term memory indicator sequence Is there a fault value for at least one data point in Greater than the fault indicator threshold , and the difference between the fault value of a fault data point and the fault value of another fault data point Is it greater than the category threshold difference? Afterwards, the method further includes: If the supplementary index sequence There is no fault value for at least one data point in Greater than the fault indicator threshold , then directly define the temporary indicator sequence as the final fault indicator sequence.

4. The method for identifying arc grounding in distribution networks based on waveform subsequence segmentation and clustering according to claim 1, characterized in that: in, Calculate the fault index value according to the improved LSTM algorithm The specific calculation process is: Perform standard deviation processing on the input time series to obtain the standard deviation series , where the expression for calculating the standard deviation is: , Where, and are the ith original zero-sequence voltage signal and the mean of the n consecutively selected original zero-sequence voltage signals; When calculating the short-term memory index, assume that the current input is , then the response output of the current input is short-term memory ,in, The value length is u, calculating the short-term memory The expression is: , Where, 、 They are the state results of the forget gate and the state results of the input gate, respectively. is the activation function, is the connection weight, is the mean value sequence input of the zero-sequence voltage signal at the current moment, is the short-term memory output at time t-1, is the activation function, is the short-term memory output at time t, is the bias phase, is the input state weight, is the bias term, Multiply the corresponding elements of vectors of the same shape. is the unit state input at time t.

5. The method for identifying arc grounding in distribution network based on waveform subsequence segmentation and clustering according to claim 1, characterized in that: The performing feature dimensionality reduction processing on the extracted time-frequency features includes: The extracted time-frequency features are subjected to feature dimensionality reduction processing based on principal component analysis, specifically: Input m pieces of n-dimensional data and organize the eigenvalues into a matrix of n rows and m columns. ,in, are the m eigenvectors of the nth fault; The matrix Each row of is zero-meaned, that is, the mean of this row is subtracted; Calculate the covariance matrix E, where the expression of the covariance matrix E is: , Where, is the eigenvector and eigenvectors The covariance of is the eigenvector and eigenvectors covariance of Calculate the eigenvalues and eigenvectors of the covariance matrix E, sort the eigenvalues from large to small, select the eigenvectors corresponding to the largest N eigenvalues as row vectors, and form the eigenvector matrix P; Transform the eigenvector matrix P into the new space constructed by N eigenvectors, that is ,in, is the eigenvector matrix after dimensionality reduction.

6. The method for identifying arc grounding in distribution networks based on waveform subsequence segmentation and clustering according to claim 1, characterized in that: in, The process of building a fault identification model includes: Determine the value of k, which means the number of aggregated classes; Randomly select k initial cluster centers from the transformed dataset Randomly select k centroid vectors from , the coordinate selection formula of the centroid vector is: , Where, 、 、 are the X coordinates, Y coordinates, and Z coordinates of the k centroids, 、 、 They are the smallest value in the X coordinate, the smallest value in the Y coordinate, and the smallest value in the Z coordinate, respectively. 、 、 They are the difference between the maximum and minimum values of the X coordinate, the difference between the maximum and minimum values of the Y coordinate, and the difference between the maximum and minimum values of the Z coordinate. is a random number between (0,1); Assign sample points, update cluster centers, and calculate samples Distance from each centroid vector ,in, is the th sample; Assume that the clusters are divided into , then Record the category with the smallest distance At this time, the center of mass is recalculated and updated, and the expression is: , Where, is the k-th category data after clustering, The number of samples in ; Repeat the steps of assigning sample points and updating cluster centers until all sample points are assigned, the categories of all sample points do not change, or the number of iterations reaches the specified maximum value, then stop clustering.