Power distribution network fault positioning method and positioning system

By combining the missing interval length and position distance in the fault location of the distribution network, the interpolation method is selected, and the extreme learning machine ELM is used to repair outliers, the problem of improper handling of missing values and outliers is solved, and more accurate fault location is achieved.

CN120277438AActive Publication Date: 2025-07-08이너 몽골리아 일렉트릭 파워 그룹 컴퍼니 리미티드 이너 몽골리아 일렉트릭 파워 리서치 인스티튜트 브랜치

Patent Information

Application Number
CN202510763950.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

When handling distribution network faults, the missed values and outliers are improperly handled, resulting in inaccurate fault positioning. The traditional interpolation method lacks position dependence, and directly eliminates outliers and ignores their possible fault characteristics. After data processing, it is easy to deviate from the actual trend, resulting in a decrease in positioning accuracy.

Method used

By obtaining the missing intervals of voltage and current sequences, combining the missing interval length and position distance, different interpolation methods are selected for missing value processing, and outliers are repaired using the extreme learning machine ELM bidirectional prediction, and fault location is performed by combining clustering processing and center of mass calculation.

Benefits of technology

It improves the accuracy of missing value filling, explores hidden relationships between data, enhances the accuracy and positioning efficiency of fault judgment, reduces misjudgment, and improves the accuracy and efficiency of fault positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277438A_ABST
    Figure CN120277438A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution network fault positioning method and positioning system, and relates to the technical field of power distribution network fault positioning, and the method comprises the steps: collecting multiple groups of historical data of different faults of a power distribution network, obtaining the missing intervals of voltage and current sequences in each group of historical data, and positioning the missing intervals to the front, middle and back of the whole sequence; different interpolation methods of different missing intervals are obtained by combining the length of the missing interval, the distance from the starting point position of the missing interval to the previous nearest missing interval and the distance from the ending point position of the missing interval to the next nearest missing interval, an abnormal value is marked for a sequence after interpolation is completed, and the sequence is equally divided into S < front > and S < back >; the abnormal values are repaired by utilizing ELM, a standard data set belonging to the fault is obtained, clustering processing is carried out on data in the standard data set, irrelevant clusters are removed, and a specific position is positioned according to the similarity between the fault and each group of faults in the target cluster. The power distribution network fault positioning method effectively improves the power distribution network fault positioning precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network fault location, and specifically to a distribution network fault location method and a location system. Background Art

[0002] During the operation of the distribution network, the occurrence of faults will affect the reliability and stability of power supply. Accurately and quickly locating faults is crucial for reducing power outage losses. Due to the complex operating environment, data such as voltage and current collected by sensors often have missing values and outliers. If not properly processed, it will lead to distortion of fault characteristics. Existing preprocessing methods have obvious defects: when dealing with missing values, traditional interpolation techniques adopt a unified strategy, such as linear interpolation and mean filling, without combining the missing position, interval length, and the correlation characteristics of the front and back data for interpolation, resulting in filling deviation. In the processing of outliers, traditional methods often directly remove abnormal data, such as the interquartile range method, but ignore that it may be a key representation of faults, resulting in the loss of fault characteristic information and the inability to obtain the true fault law. In addition, existing methods do not incorporate the data time series dependence relationship, and the processed data is prone to deviate from the actual trend, further leading to calculation deviation of the fault feature vector, causing clustering centroid shift or cluster division error, and ultimately resulting in false fault judgment or a decrease in location accuracy. Summary of the Invention

[0003] Aiming at the deficiencies of the existing technology, the present invention provides a distribution network fault location method and a location system, which solve the problem of inaccurate fault location caused by improper processing of data missing values and outliers.

[0004] To achieve the above objectives, the present invention is realized through the following technical solutions: A distribution network fault location method, including: S1. Obtain the missing intervals of the voltage and current sequences in each group of historical data, locate them at the front, middle, and back of the entire sequence, and combine the missing interval length L, the distance Dfront from the starting position of the missing interval to the nearest missing interval in front, and the distance Dback from the ending position of the missing interval to the nearest missing interval in the back to obtain different interpolation methods for different missing intervals. Mark the outliers in the interpolated sequence, and divide the sequence into the front group Sfront and the back group Sback of equal length, and use the extreme learning machine ELM bidirectional prediction to repair the outliers respectively; S2. Construct discriminant features based on the processed voltage and current sequences, and perform normalization processing to obtain a standard data set belonging to this type of fault. Cluster the data in the standard data set, find the clusters containing light and heavy faults, eliminate irrelevant clusters according to the minimum time interval and time similarity, calculate the centroid of the remaining clusters, calculate the centroid distance di from the existing fault to the centroids of different fault remaining clusters to locate the target cluster, and then calculate the similarity between this fault and each group of faults in the target cluster to locate the specific position.

[0005] As a further solution of the present invention, before S1, it further includes S0, where S0 is to collect multiple groups of historical data of different faults in the distribution network. The specific historical data includes voltage sequences, current sequences during the fault duration, as well as fault levels, timestamps, and specific locations.

[0006] As a further solution of the present invention, the method for positioning the specific part of the missing interval in the entire sequence is as follows: Calculate the total length M of the sequence, define the start index start and end of the missing interval. If there is only one value in the missing interval, then start is equal to end; If end <= M * A, then position the missing interval at the front of the sequence; If start > M * A and end <= M * (A + B), then position the missing interval in the middle of the sequence; If start > M * (A + B), then position the missing interval at the back of the sequence, where A, B, C ∈ (0, 1) and A + B + C = 1.

[0007] As a further solution of the present invention, when the missing interval is positioned at the front of the sequence, the selected interpolation methods specifically include: If L <= Lth and max{Dfront, Dback} <= Dth, select the nearest neighbor interpolation method; If L <= Lth and max{Dfront, Dback} > Dth, select the linear interpolation method; If L > Lth and max{Dfront, Dback} <= Dth, select the interpolation method used for the missing interval with the closest distance to the current missing interval, that is, the interpolation method used for the missing interval corresponding to min{Dfront, Dback}; If L > Lth and max{Dfront, Dback} > Dth, select the linear extrapolation method.

[0008] As a further solution of the present invention, when the missing interval is positioned in the middle of the sequence, the selected interpolation methods specifically include: If L <= Lth and max{Dfront, Dback} <= Dth, select the nearest neighbor interpolation method; If L <= Lth and max{Dfront, Dback} > Dth, select the weighted moving average method; If L > Lth and max{Dfront, Dback} <= Dth, select the interpolation method used for the missing interval with the closest distance to the current missing interval, that is, the interpolation method used for the missing interval corresponding to min{Dfront, Dback}; If L > Lth and max{Dfront, Dback} > Dth, select Lagrange interpolation or cubic spline interpolation.

[0009] As a further solution of the present invention, when the missing interval is located at the rear of the sequence, the selected interpolation method specifically includes: If L <= Lth and max{Dfront, Dback} <= Dth, the nearest neighbor interpolation method is adopted; If L <= Lth and max{Dfront, Dback} > Dth, the linear interpolation method is adopted; If L > Lth and max{Dfront, Dback} <= Dth, select the interpolation method used for the missing interval with the closest distance to the current missing interval, that is, the interpolation method used for the missing interval corresponding to min{Dfront, Dback}; If L > Lth and max{Dfront, Dback} > Dth, the exponential smoothing method can be adopted.

[0010] As a further solution of the present invention, the specific operation method for repairing outliers by using the extreme learning machine ELM bidirectional prediction is as follows: Set the window size W, calculate the local mean μ and the standard deviation σ. If a certain point xt satisfies |xt - μ| > k * σ, then mark this point as a potential outlier; Divide the sequence into the front group Sfront and the back group Sback with equal length. Detect the outliers within the groups for Sfront and Sback respectively by the interquartile range method, and remove the outliers in Sfront and Sback to obtain the pure training sets Sfront_pure and Sback_pure; Train the ELM model with Sfront_pure and Sback_pure respectively, and use the trained ELM models to predict Sback and Sfront respectively to obtain the predicted value sequences Sback_prediction and Sfront_prediction; For each point yt ∈ Sback in the back group and yt_pred ∈ Sback_prediction, calculate |yt - yt_pred|. If this point exceeds the specified threshold and this point is marked as a potential outlier, then replace yt with yt_pred. Similarly, perform the above operation on Sfront; Merge the repaired Sfront and Sback into a complete sequence, and retain the unmarked normal points, where k is a proportionality coefficient.

[0011] As a further solution of the present invention, clusters where t > Tmax and S’ < S0 and t < Tmin and S’ < S0 in each fault are removed, where t is the minimum time interval between light and heavy faults, S’ is the time similarity, and Tmin, Tmax, and S0 are the lower time threshold, the upper time threshold, and the time similarity threshold respectively.

[0012] As a further solution of the present invention, find the cluster corresponding to min(di) as the target cluster, and the fault corresponding to the target cluster is used as the specific fault type that occurs in the distribution network during this period. Calculate the cosine similarity Si between the fault feature vector and the feature vectors of each historical fault in the target cluster, and obtain the fault location corresponding to min(Si) as the real-time location of the current fault.

[0013] A distribution network fault location system includes: a data acquisition module, a missing value processing module, an outlier processing module, and a fault location module; The data acquisition module collects multiple groups of historical data of different faults in the distribution network and transmits them to the missing value processing module; The missing value processing module obtains the missing intervals of the voltage and current sequences in each group of historical data, locates them at the front, middle, and back of the entire sequence, and combines the length of the missing interval, the distance from the starting position of the missing interval to the nearest missing interval in front, and the distance from the ending position of the missing interval to the nearest missing interval in the back to obtain different interpolation methods for different missing intervals, and transmits the processed sequence to the outlier processing module; The outlier processing module marks the outliers in the interpolated sequence, divides the sequence into the front group S_front and the back group S_back of equal length, and uses the extreme learning machine ELM bidirectional prediction to repair the outliers respectively, and transmits the repaired sequence to the fault location module; The fault location module constructs discriminant features based on the processed voltage and current sequences, obtains a standard data set belonging to this type of fault, performs clustering processing on the data in the standard data set, finds the clusters containing light and heavy faults, eliminates irrelevant clusters according to the minimum time interval and time similarity, calculates the centroid of the remaining clusters, calculates the distance from the existing fault to the centroid of different fault remaining clusters to locate the target cluster, and then calculates the similarity between the fault and each group of faults in the target cluster to locate the specific fault location.

[0014] The present invention provides a distribution network fault location method and a location system, which have the following beneficial effects compared with the prior art: (1) By selecting different interpolation methods for different missing value positions, the present invention can fill the missing values more accurately, reduce data deviation, and improve data quality; (2) When processing outliers, the present invention predicts the sequence forward and backward through the extreme learning machine, and uses the predicted value to replace the outliers, avoiding the problems caused by directly eliminating the outliers, better mining the hidden relationships between data, and improving the accuracy of fault judgment; (3) Through the screening and centroid calculation of the clustering results, and the comparison of the fault feature vectors, the present invention can more accurately locate the fault type and specific location of the distribution network, improving the accuracy and efficiency of fault location. Description of the Drawings

[0015] Figure 1 The steps of the present invention are as follows; Figure 2 This is the principle framework of the system of the present invention. DETAILED DESCRIPTION

[0016] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0017] like Figure 1 , the present application provides a distribution network fault location method, comprising: For each type of fault that occurs in the distribution network, multiple sets of historical data are extensively collected. These data mark the fault level in detail, which can be divided into three levels: light, medium, and severe. They also accurately record the specific location and timestamp of the fault. Each set of data fully covers the voltage [v1, v2, ..., vn] and current [i1, i2, ..., in] after the fault occurs; Taking a city's distribution network as an example, in a fault caused by a tree branch touching the line, the monitoring equipment recorded that the voltage sequence after the fault was [215,210,null,200,195], the current sequence was [80,90,null,110,120], and the timestamp was "2024-10-15 14:30:00". After evaluation, the fault level was mild, and the fault location was near the No. 3 electric pole on a street, where null represents a missing value.

[0018] In the actual monitoring process, due to environmental interference, equipment failure and other factors, missing values ​​often appear in the voltage and current series. One existing method is to directly use various interpolation methods to interpolate the above series, but sometimes there are more than one missing value in the series, and the length of the missing value will also affect the accuracy of interpolation. Another method is to use the existing machine learning model to learn the existing data for interpolation, but it is not easy to mine the time characteristics between the data, and the features learned by the model are incomplete, resulting in large deviations in the subsequent prediction interpolation. Therefore, the present invention divides the sequence into three parts: the front part, the middle part, and the rear part. The specific division rules are as follows: First, calculate the total length M of the sequence. Then, define the start and end indices of the missing interval. If there is only one value in the missing interval, start is equal to end. The front part means that the missing interval is completely within the front A region of the sequence, that is, end <= M * A, close to the beginning of the sequence. The middle part indicates that the missing interval is within the middle B region of the sequence, that is, start > M * A and end <= M * (A + B), in the middle position far from the beginning and end of the sequence. The rear part is that the missing interval is completely within the rear C region of the sequence, that is, start > M * (A + B), close to the end of the sequence. If the missing interval spans multiple regions, the missing interval is divided according to the actual situation, and the front, middle, and rear parts of different divided parts are determined. Among them, A, B, C ∈ (0, 1) and A + B + C = 1; For example, when a fault occurs in a distribution network, the voltage sequence is [215, 210, null, null, null, 233, 244, 255, 213]. The total length M of the sequence is 9, the start index of the missing interval start = 3, end = 5, A = 1 / 3, B = 1 / 3, C = 1 / 3. Then the front part is end <= 3, the middle part is start > 3 and end <= 6, the rear part is start > 6. The missing interval is [3, 5]. Index 3 is in the front part, and indices 4 and 5 are in the middle part. Therefore, the missing interval can be divided, and the missing values in the front and middle parts are processed by different methods to make the interpolation more accurate; After determining the position of the missing interval in the entire sequence, calculate the length L of the missing interval, the front distance D_front from the start position of the missing interval to the nearest missing interval in the front, and the rear distance D_rear from the end position of the missing interval to the nearest missing interval in the rear. Combine these three parameters to determine the difference method applicable to different intervals; The length L of the missing interval reflects the scale of the data volume that needs to be interpolated. When L is small, it means that there is less missing data, and interpolation can rely more on a small amount of adjacent known data. When L is large, it means that there is more missing data, and more macroscopic features such as the overall trend and periodicity of the data need to be considered for interpolation; D_front reflects the richness of the front-side data that can be referred to during interpolation. The larger D_front is, the more front-side data is available for reference, and the better the historical trend of the data can be grasped during interpolation. The smaller D_front is, it indicates that the front-side data is limited, and interpolation may only be based on a few data adjacent to the missing interval; D_rear reflects the amount of rear-side data that can be referred to during interpolation. The larger D_rear is, the richer the rear-side data is, and more information about the subsequent data trend can be provided for interpolation. The smaller D_rear is, the less rear-side data, and the dependence on the rear-side data during interpolation will be reduced; When the index of the missing interval is in the front part of the sequence, if L <= Lth and max{D_front, D_back} <= Dth, the nearest neighbor interpolation method is adopted. The nearest neighbor interpolation method is a simple and direct method that selects the value of the known data point closest to the missing value to fill the missing value. If L <= Lth and max{D_front, D_back} > Dth, it indicates that there is rich data information on the front or back side of the missing interval, and the linear interpolation method can be used. For example, the total length M of a certain voltage sequence is 10, and the missing interval is [2, 3], that is, the second and third data are missing, so L is 2. If the set Lth is 3, D_front is 1, D_back is 2, and the Dth value is 3. At this time, the missing interval is located in the front part of the sequence. At the same time, there is not enough data on the front and back sides of the missing interval for reference, and the missing interval is short. In this case, the nearest neighbor interpolation method is used to fill the second missing value with the value of the first data and the third missing value with the value of the fourth data. This method is simple to calculate and can quickly and effectively fill the missing value in the short missing scenario, ensuring the coherence of the data; If L > Lth and max{D_front, D_back} <= Dth, it indicates that there are many values to be filled in the missing interval, while the amount of data on the front and back sides of the missing interval is very small at this time. Therefore, simply interpolating and inferring through the data on the front and back sides will result in large errors. At this time, select the interpolation method used for the missing interval with the closest distance to the current missing interval, that is, the interpolation method used for the missing interval corresponding to min{D_front, D_back}. If L > Lth and max{D_front, D_back} > Dth, the linear extrapolation method is adopted. The linear extrapolation method predicts the missing value based on the growth or decline trend of the data on the side corresponding to max{D_front, D_back}. For example, the total length M of a certain current sequence is 15, and the missing interval is [3, 6], L is 4, greater than the Lth value of 3, and D_back is 6, greater than the Dth value of 3. If the data on the back side shows a linear growth trend, and the seventh data is known to be 120 and the eighth data is 130, the missing value is predicted through linear calculation. According to the linear growth law, first calculate the growth rate of the data, and then predict the missing value based on this rate and the known data. This method is applicable to the front part missing situation where the missing interval is long but the data trend on the front or back side is obvious, avoiding large errors caused by simple interpolation; When the index of the missing interval is in the middle part of the sequence, if L <= Lth and max{Dfront, Dback} <= Dth, the nearest neighbor interpolation method is adopted. If L <= Lth and max{Dfront, Dback} > Dth, it indicates that there is rich data information on the front or back side of the missing interval, and the weighted moving average method can be used. This method combines the trends and weights of the front and back data to calculate the missing value. By assigning different weights to the data at different time points, it can balance the overall trend and local fluctuations of the data. In the case where the data fluctuates greatly but there is a certain trend, this method can better comprehensively consider various characteristics of the data and obtain a more reasonable estimate of the missing value; If L > Lth and max{Dfront, Dback} <= Dth, select the interpolation method used for the missing interval with the closest distance to the current missing interval, that is, the interpolation method used for the missing interval corresponding to min{Dfront, Dback}. If L > Lth and max{Dfront, Dback} > Dth, Lagrange interpolation or cubic spline interpolation can be used. The Lagrange interpolation method calculates the missing value by constructing a polynomial through the known data points before and after. Its advantage is that it can utilize multiple known data points to fit the data curve by constructing a polynomial, thereby calculating the missing value more accurately. The cubic spline interpolation can ensure the smoothness of the interpolation curve, which is more in line with the actual change of the data. When dealing with data with a certain continuity and smoothness, the cubic spline interpolation can make the interpolated curve smoother, avoid sudden change points, and more accurately reflect the internal law of the data. For example, for a certain sequence with a total length M of 20 and a missing interval of [8, 9], L is 5, which is greater than the Lth value of 3, Dfront is 5, and Dback is 6, both of which are greater than the Dth value of 3. At this time, Lagrange interpolation is used to construct a polynomial through the known points before and after, which can fully consider the change trend of the surrounding data and obtain a more accurate missing value; When the index of the missing interval is in the back part of the sequence, if L <= Lth and max{Dfront, Dback} <= Dth, the nearest neighbor interpolation method is adopted. If L <= Lth and max{Dfront, Dback} > Dth, the linear interpolation method is adopted; If L > Lth and max{Dfront, Dback} <= Dth, select the interpolation method used for the missing interval with the closest distance to the current missing interval, that is, the interpolation method used for the missing interval corresponding to min{Dfront, Dback}. If L > Lth and max{Dfront, Dback} > Dth, the exponential smoothing method can be adopted. This method predicts the missing value based on the historical trend and weight of the front-side or back-side data, and can better adapt to the changes in the data. It assigns a larger weight to the recent data and a smaller weight to the long-term data, and pays more attention to the recent change trend of the data. For example, the total length M of a certain current sequence is 18, the missing interval is [15, 18], L is 4, which is greater than the Lth value of 3, and Dfront is 6, which is greater than the Dth value of 3. In this case, the exponential smoothing method can make full use of the historical trend of the front-side data and predict the missing value through weighted calculation, making the prediction result more in line with the actual change law of the data, especially suitable for the back-side missing scenario where the data has a certain trend and the missing interval is long.

[0019] After reasonably interpolating the voltage and current sequences through the above method, it is also necessary to process the outliers of these sequences. Most traditional outlier processing methods are based on statistical rules, distance metrics, and density, etc., and directly eliminate the outliers that do not conform to the rules. However, in the distribution network data, these outliers are very likely to be caused by faults. If directly eliminated, it will make the subsequent fault judgment more difficult and increase the difficulty of mining the hidden relationships between the data. The present invention adopts a new outlier processing method, and respectively uses the bidirectional prediction of the extreme learning machine ELM to repair the outliers. The specific operation method is as follows: Set the window size W, calculate the local mean μ and the standard deviation σ. If a certain point xt satisfies |xt - μ| > k * σ, then mark this point as a potential outlier; Divide the sequence into the front group Sfront and the back group Sback with equal length, and respectively use the interquartile range method to detect the outliers within the group for Sfront and Sback, and eliminate the outliers in Sfront and Sback to obtain the pure training sets Sfront_pure and Sback_pure; Train the ELM model with Sfront_pure and Sback_pure respectively, and use the trained ELM models to predict Sback and Sfront respectively to obtain the predicted value sequences Sback_prediction and Sfront_prediction; For each point yt ∈ Sback in the back group and yt_pred ∈ Sback_prediction, calculate |yt - yt_pred|. If this point exceeds the specified threshold and this point is marked as a potential outlier, then use yt_pred to replace yt. Similarly, perform the above operation on Sfront; Merge the repaired Sfront and Sback into a complete sequence, and retain the unmarked normal points.

[0020] Construct discriminant features based on the processed voltage and current sequences, including voltage mean, voltage standard deviation, current mean, current standard deviation, and duration. These discriminant features can comprehensively reflect the data characteristics during a fault occurrence, providing a basis for fault judgment. Perform min-max normalization on multiple sets of data to convert data with different ranges and magnitudes to a unified scale, eliminating the influence of data dimensions, and obtaining a standard data set belonging to this type of fault; Each set of data in the standard data set has its corresponding timestamp and specific location. Cluster the data in the standard data set, such as using the K-Means clustering algorithm, to obtain m clusters, and record the corresponding fault levels in each cluster; Find the clusters containing minor and major faults, sort the data in these clusters according to the corresponding timestamps, calculate the minimum time interval t between the minor and major faults, compare t with the specified time threshold T, and calculate the time similarity S' between these two sets of data through dynamic time warping (DTW); Eliminate the clusters when t > Tmax and S' < S0. If the time interval t between the minor fault data and the major fault data is greater than the specified time threshold Tmax, it indicates that the time interval between the occurrence of the minor and major faults is relatively long and their correlation is weak. And the time similarity S is less than the set similarity threshold S0, which means that there are significant differences in the time distributions of these two sets of data. In this case, the minor and major fault data in this cluster may come from different fault events or fault development stages, and they are clustered together possibly due to errors in the clustering algorithm or abnormal fluctuations in the data. Therefore, this cluster needs to be eliminated because it cannot accurately reflect a set of fault data with similar characteristics, and retaining this cluster may interfere with subsequent fault analysis and judgment; Eliminate the clusters when t < Tmin and S' < S0. Although the time interval t between the minor and major fault data is less than the time threshold Tmin, indicating that the fault occurrence times are relatively close, the time similarity S is less than the set threshold S0, indicating that there are obvious differences in their time distribution patterns. This may mean that the minor and major faults occurring in a short period are not different stages of the same fault development process, but independent faults caused by different reasons. In this case, the data characteristics within this cluster are inconsistent and do not conform to the original intention of clustering. If this cluster is retained, it may cause confusion in subsequent analysis. Therefore, it needs to be eliminated to ensure the accuracy and reliability of the clustering results; For example, there are two clusters. One cluster contains minor fault data and major fault data. The minor fault timestamp is 10:00, and the major fault timestamp is 11:30. The specified time threshold Tmax is 1 hour, Tmin is 0.5 hour. The calculated time interval t is 1.5 hours, the time similarity S calculated using the DTW algorithm is 0.6, and the set similarity threshold S0 is 0.8. This cluster needs to be eliminated; After removing all the non-conforming clusters, calculate the centroids of the remaining clusters. When locating a specific fault in the distribution network during a certain period of time, collect the data within this period in real time, calculate the feature vector for this period, and then calculate the Euclidean distance di between it and the centroids of all the clusters in all the faults. The Euclidean distance can measure the similarity between two vectors. Sort the dis to find the minimum di, that is, the fault corresponding to min(di), and thus obtain the specific fault type that occurred in the distribution network during this period; After finding the specific fault type, continue to calculate the similarity Si between the feature vector of this fault and the feature vectors of each fault in the corresponding cluster. For example, use the cosine similarity algorithm to obtain min(Si), and use the fault position marked by fault i as the real-time position of the current fault, so as to realize the accurate location of the distribution network fault; Such as Figure 2 A distribution network fault location system includes: a data acquisition module, a missing value processing module, an outlier processing module, and a fault location module; The data acquisition module collects multiple groups of historical data of different faults in the distribution network and transmits them to the missing value processing module; The missing value processing module obtains the missing intervals of the voltage and current sequences in each group of historical data, locates them at the front, middle, and back of the entire sequence, combines the length of the missing interval, the distance from the starting position of the missing interval to the nearest missing interval in front, and the distance from the ending position of the missing interval to the nearest missing interval in the back, to obtain different interpolation methods for different missing intervals, and transmits the processed sequence to the outlier processing module; The outlier processing module marks the outliers in the interpolated sequence, divides the sequence into the front group S_front and the back group S_back of equal length, and uses the extreme learning machine ELM bidirectional prediction to repair the outliers respectively, and transmits the repaired sequence to the fault location module; The fault location module constructs discriminant features based on the processed voltage and current sequences, obtains the standard data set belonging to this type of fault, performs clustering processing on the data in the standard data set, finds the clusters containing light and heavy faults, removes the irrelevant clusters according to the minimum time interval and time similarity, calculates the centroids of the remaining clusters, calculates the centroid distances from the existing faults to the remaining clusters of different faults to locate the target cluster, and then calculates the similarity between this fault and each group of faults in the target cluster to locate the specific fault position.

[0021] Some of the data in the above formulas are numerically calculated after removing their dimensions, and the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

[0022] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A method for fault location in a distribution network, characterized in that, Including: S1. Obtain the missing intervals of the voltage and current sequences in each group of historical data, locate them at the front, middle, and back of the entire sequence, and combine the length L of the missing interval, the distance D_before from the starting position of the missing interval to the nearest missing interval in front, and the distance D_after from the ending position of the missing interval to the nearest missing interval behind to obtain different interpolation methods for different missing intervals. Mark the outliers in the interpolated sequence, and divide the sequence into a front group S_before and a back group S_after of equal length. Use the bidirectional prediction of the extreme learning machine (ELM) to repair the outliers respectively; S2. Construct discriminant features based on the processed voltage and current sequences, and perform normalization processing to obtain a standard data set belonging to this type of fault. Cluster the data in the standard data set, find the clusters containing light and heavy faults, eliminate irrelevant clusters according to the minimum time interval and time similarity, calculate the centroid of the remaining clusters, calculate the centroid distance di from the existing fault to the remaining clusters of different faults to locate the target cluster, and then calculate the similarity between this fault and each group of faults in the target cluster to locate the specific position.

2. The method for locating faults in a distribution network according to claim 1, wherein Before S1, it also includes S0, where S0 is to collect multiple groups of historical data of different faults in the distribution network. The specific historical data includes the voltage sequence, current sequence during the fault duration, as well as the fault level, timestamp, and specific location.

3. The distribution network fault location method according to claim 1, characterized in that The method for locating the specific position of the missing interval in the entire sequence is as follows: Calculate the total length M of the sequence, define the starting index start and ending index end of the missing interval. If the missing interval contains only one value, then start is equal to end; If end <= M * A, locate the missing interval at the front of the sequence; If start > M * A and end <= M * (A + B), locate the missing interval in the middle of the sequence; If start > M * (A + B), locate the missing interval at the back of the sequence, where A, B, C ∈ (0, 1) and A + B + C = 1.

4. The method for locating faults in a distribution network according to claim 3, wherein When the missing interval is located at the front of the sequence, the selected interpolation methods specifically include: If L <= L_th and max{D_before, D_after} <= D_th, select the nearest neighbor interpolation method; If L <= L_th and max{D_before, D_after} > D_th, select the linear interpolation method; If L > L_th and max{D_before, D_after} <= D_th, select the interpolation method used for the missing interval with the closest distance to the current missing interval, that is, the interpolation method used for the missing interval corresponding to min{D_before, D_after}; If L > L_th and max{D_before, D_after} > D_th, select the linear extrapolation method.

5. The power distribution network fault location method according to claim 3, wherein When the missing interval is located in the middle of the sequence, the selected interpolation methods specifically include: If L <= L_th and max{D_before, D_after} <= D_th, select the nearest neighbor interpolation method; If L <= L_th and max{D_before, D_after} > D_th, select the weighted moving average method; If L > L_th and max{D_before, D_after} <= D_th, select the interpolation method used for the missing interval with the closest distance to the current missing interval, that is, the interpolation method used for the missing interval corresponding to min{D_before, D_after}; If L>Lth and max{Dfront,Dback}>Dth, select Lagrange interpolation or cubic spline interpolation.

6. The method for locating faults in a distribution network according to claim 3, wherein, If the missing interval is located at the end of the sequence, the interpolation method selected includes: If L<=Lth and max{Dbefore, Dafter}<=Dth, use the nearest neighbor interpolation method; If L<=Lth and max{D before, D after}>Dth, use linear interpolation; If L>Lth and max{Dbefore, Dafter}<=Dth, select the interpolation method used for the missing interval that is closest to the current missing interval, that is, the interpolation method used for the missing interval corresponding to min{Dbefore, Dafter}; If L>Lth and max{Dfront, Dback}>Dth, exponential smoothing can be used.

7. The method for locating faults in a distribution network according to claim 1, wherein The specific operation method of repairing outliers using extreme learning machine ELM bidirectional prediction is as follows: Set the window size W, calculate the local mean μ and standard deviation σ, and if a point xt satisfies |xt-μ|>k*σ, mark the point as a potential outlier; The sequences are divided into a front group Spre and a rear group Spost of equal length. The interquartile range method is used to detect abnormal points within the group for Spre and Spost, and the abnormal points in Spre and Spost are removed to obtain the pure training sets Spre_pure and Spost_pure. Use Spre_pure and Spost_pure to train ELM models respectively, and use the trained ELM models to predict Spost and Spre respectively, and obtain the prediction value sequences Spost_prediction and Spre_prediction; For each point yt∈S in the posterior group, yt_pred∈S_pred, calculate |yt-yt_pred|. If the point exceeds the specified threshold and is marked as a potential outlier, yt_pred is used instead of yt. Similarly, the same operation is performed for S. The repaired S before and S after are merged into a complete sequence, and the unmarked normal points are retained, where k is the proportional coefficient.

8. The method for locating a fault in a distribution network according to claim 1, wherein For each fault, t>Tmax and S' <S0和t<Tmin且S’<S0时的簇剔除,其中,t为轻、重故障之间最小的时间间隔,S’为时间相似度,Tmin、Tmax、S0分别为时间阈值下限、时间阈值上限、时间相似度阈值。 9. The method for locating faults in a distribution network according to claim 1, wherein Find the cluster corresponding to min(di) as the target cluster, and the fault corresponding to the target cluster is taken as the specific fault type that occurred in the distribution network during that period of time. Calculate the cosine similarity Si between the fault feature vector and the feature vector of each historical fault in the target cluster, and get the fault location corresponding to min(Si) as the real-time location of the current fault.

10. A distribution network fault location system for performing the distribution network fault location method according to any one of claims 1-9, characterized in that, include: Data acquisition module, missing value processing module, abnormal value processing module, fault location module; A data acquisition module, which collects multiple sets of historical data of different faults in the distribution network and transmits them to the missing value processing module; Missing value processing module, which obtains the missing intervals of the voltage and current sequences in each group of historical data, locates them at the front, middle, and back of the entire sequence, combines the length of the missing interval, the distance from the starting position of the missing interval to the nearest missing interval in the front, and the distance from the ending position of the missing interval to the nearest missing interval in the back, obtains different interpolation methods for different missing intervals, and transmits the processed sequence to the outlier processing module; Outlier processing module, which marks the outliers in the interpolated sequence, divides the sequence into the front group S_front and the back group S_back of equal length, repairs the outliers using the bidirectional prediction of the extreme learning machine ELM respectively, and transmits the repaired sequence to the fault location module; Fault location module, which constructs discriminant features based on the processed voltage and current sequences, obtains the standard data set belonging to this type of fault, performs clustering processing on the data in the standard data set, finds the clusters containing light and heavy faults, eliminates the irrelevant clusters according to the minimum time interval and time similarity, calculates the centroid of the remaining clusters, calculates the distance from the existing fault to the centroid of different fault remaining clusters to locate the target cluster, and then calculates the similarity between the fault and each group of faults in the target cluster to locate the specific fault location.

Citation Information

Patent Citations

  • A method for locate faults in distribution network

    CN109460431A

  • Power distribution network fault positioning method and system based on artificial intelligence

    CN118169510A

  • Power distribution network fault positioning method and system in non-sound communication scene

    CN118584237A

  • Fault diagnosis system and method for photovoltaic power distribution network

    CN119048062A

  • Power transmission line abnormity early warning and fault positioning system

    CN119827916A

Cited By

  • Power distribution network ground fault research and judgment analysis method

    CN120722116A

  • Urban acoustic environment portrait generation method and system capable of adaptively adjusting granularity

    CN121577150A