Intelligent line fault reason identification method based on recorded waveform image
By using an intelligent identification method based on recorded waveform images, and constructing a fault cause identification model using SIFT and SVM, the problems of high resource consumption and reliance on human experience in transmission line fault identification are solved. This achieves fast and accurate fault cause identification and improves the model's generalization ability and transparency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for fault identification in transmission lines suffer from problems such as insufficient reliance on human experience, high resource consumption, poor model generalization ability, and opaque decision-making process, making it difficult to achieve efficient and accurate fault cause identification.
An intelligent identification method based on recorded waveform images is adopted. Features are extracted through scale-invariant feature transform (SIFT) and combined with a support vector machine (SVM) classifier to construct a fault cause identification model. Fault recorded waveform images are used for feature representation and classification.
It enables rapid and accurate identification of transmission line fault causes without relying on electrical parameters, improving the accuracy and timeliness of fault identification, reducing computational resource requirements, and enhancing the model's generalization performance and transparency.
Smart Images

Figure CN121746767A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of power system fault detection, and particularly relates to a line fault cause intelligent identification method based on a recording wave waveform image. BACKGROUND
[0002] The power transmission line is one of the most widely covered power equipment in China, and its stability is crucial to guarantee power supply, enterprise production and resident life. As a core component of the power system, the power transmission line is widely distributed and extends to various regions in China, including complex terrain mountainous areas, harsh climate open fields and the like. These lines are exposed to harsh natural environment for a long time, and are easily disturbed by natural factors such as heavy rain, lightning, snowstorm, birds and beasts. The influence of these natural factors often causes the power transmission line to fail, thereby seriously affecting the normal operation of the power grid, and further affecting social production and people's life.
[0003] Currently, the identification of fault causes mainly relies on on-site inspection and expert experience. When a fault occurs on a transmission line, the power dispatch center usually initiates an emergency response mechanism immediately, arranging staff to inspect the area where the fault is located. Traditional methods are mainly based on the analysis of a small number of fault cases themselves, such as calculating voltage / current amplitude, phase angle, and other single characteristic quantities to determine fault location and cause, analyzing fault recording, etc. For various fault causes and mechanisms on transmission lines, Dong Guangzhe et al. analyzed the waveform changes of current and voltage during the fault process through the typical fault recording graph of high-voltage overhead lines, and summarized the basic correspondence by comparing fault causes and fault types, in order to quickly determine the fault type in future fault finding and analysis [Dong Guangzhe, Qian Meng, Wang Guolong, et al. Typical fault recording analysis of high-voltage overhead lines [J]. Electric World, 2014, 55(10): 1-6.]. Chen Leigang et al. proposed a waveform classification method that can quickly filter out incorrect waveforms from a large number of waveform data, classify and process according to the characteristics of incorrect waveforms, improve the efficiency of waveform detection, and help improve the accuracy of fault point positioning by the distribution master station [Chen Leigang, Qin Minghui, Wang Kaishun. A waveform classification method for transient fault recorder [J]. Electrical Technology, 2020, 21(08): 125-129.]. For external environmental information, Wang Jian et al. proposed a time-varying fault rate calculation method with a monthly time scale to reflect the time-dependent fault law of transmission lines [Wang Jian, Xiong Xiaofu, Li Zhe, et al. Fault time distribution characteristics and simulation of transmission lines related to meteorological environment [J]. Electric Power Automation Equipment, 2016, 36(03): 109-114+123.]. Zhang Wenfeng et al. analyzed the lightning activity law and lightning trip-out situation in Guangdong from 2005 to 2012, and analyzed the lightning protection measures for lightning faults from external factors such as time, terrain, and tower shape [Zhang Wenfeng, Peng Xiangyang, Dou Peng, et al. Lightning activity law in Guangdong and analysis of transmission line lightning trip-out [J]. Guangdong Electric Power, 2014, 27(03): 101-107.]. Comprehensive internal and external characteristics, Huang Xuyong et al. established a bird damage fault risk assessment model for transmission lines, determined the weight coefficients of each influencing factor using comprehensive evaluation method, and realized the early warning of bird damage fault of transmission lines [Huang Xuyong, Shen Zhi, Wang Xin. Risk assessment method of bird damage fault of transmission lines in Yunnan power grid [J]. High Voltage Apparatus, 2020, 56(03): 156-163.].
[0004] Traditional methods have numerous problems and shortcomings when dealing with complex and ever-changing realities. Since transmission lines are mostly located in mountainous areas and other harsh environments, inspection personnel must traverse rugged mountain paths or work in densely vegetated, treacherous terrain. Maintenance personnel may even need to carry heavy tools and equipment on long walks, placing extremely high demands on physical strength and endurance. Furthermore, severe weather conditions, such as torrential rain, strong winds, and extreme heat, significantly increase the difficulty of inspections. In these areas, poor signal strength often hinders communication and coordination among personnel, further complicating fault handling. More importantly, while existing methods based on electrical characteristics can assist in identification to some extent, the electrical characteristics of some fault categories are not obvious or have poor distinguishability, resulting in a very limited range of effectively identifiable fault types.
[0005] Currently, fault identification technologies used in power systems still lack automation and intelligent support, making it impossible to efficiently and accurately identify all types of faults. This not only means that the entire process, from fault occurrence to location and troubleshooting, and then to organizing maintenance departments and relevant professionals for emergency repairs, consumes a significant amount of time, failing to meet the demands for efficient emergency response; more importantly, relying on manual experience and traditional electrical characteristic analysis methods is insufficient to cope with the increasingly complex operating environments and fault modes of power systems. This poses a potential threat to the stable operation of power systems and causes incalculable losses to the socio-economic landscape.
[0006] With the rise of machine learning algorithms, people have begun to design different algorithms to effectively fit recorded waveform data. KM Silva et al. used wavelet transform to extract features of fault signals and combined them with artificial neural networks (ANN) for fault classification [KM Silva, BA Souza and NSD Brito, "Fault detection and classification in transmission lines based on wavelet transform and ANN", in IEEE Transactions on Power Delivery, vol. 21, no. 4, pp. 2058-2063, Oct. 2006, doi: 10.1109 / TPWRD.2006.876659.]. H. Livani and CY Evrenosoglu et al. used discrete wavelet transform (DWT) to extract voltage transient information and used support vector machine (SVM) to identify fault segments and fault half segments. The study conducted tests using transient data from simulated fault types and locations [H. Livani and CY Evrenosoglu, "A Machine Learning and Wavelet-Based Fault Location Method for Hybrid Transmission Lines," in IEEE Transactions on Smart Grid, vol. 5, no. 1, pp. 51-59, Jan. 2014, doi: 10.1109 / TSG.2013.2260421.]. Li et al. analyzed fault mechanisms and waveforms in detail from fault recorders, selecting six influencing factors to characterize six types of power outages.The frequency components of the voltage and current waveforms of the faulted phases were analyzed using Discrete Fourier Transform (DFT), and the combination of these features was used to train and test an SVM architecture to achieve higher classification accuracy [Linan Li, Renfei Che and Hongzhi Zang, "A fault cause identification methodology for transmission lines based on support vector machines", 2016 IEEE PES Asia-Pacific Power and Energy Engineering Conference (APPEEC), Xi'an, 2016, pp. 1430-1434, doi: 10.1109 / APPEEC.2016.7779725.].
[0007] However, these attempts have inevitably been limited by the scarcity of sample resources and the limitations of manual feature engineering. On the one hand, the limited sample resources make it difficult for the model to fully learn the characteristic patterns under various fault conditions, thus affecting the model's generalization ability and identification accuracy. On the other hand, manual feature engineering relies on the experience and knowledge of professionals, which is not only cumbersome and time-consuming, but also often fails to fully cover the complexity of faults, making it difficult to effectively improve the model's accurate identification of fault causes. This, to some extent, restricts the in-depth development of related research and its practical application.
[0008] With the continuous development and maturation of deep learning technology, its advantages in feature extraction have gradually become prominent, leading to the use of training networks to automatically extract features and gradually replacing traditional manual feature engineering that relies on domain knowledge and experience. This has become the mainstream trend in the current machine learning and data processing fields. Feng Yin et al. proposed a power cable fault diagnosis method based on a hybrid model of convolutional neural network and long short-term memory network, which effectively solved the problem of multi-source and heterogeneous original data caused by changes in cable parameters [Feng Yin, Jia Hongtao, Yang Zhenqiang, et al. Research on transmission line fault diagnosis method based on CNN-LSTM [J]. Power Grid and Clean Energy, 2023, 39(11): 59-65.]. Wu Junhong et al., based on the long short-term memory deep network model, directly used the three-phase fault current sampling sequence after the line fault as the input of the model to classify 10 short-circuit fault types of the line [Wu Junhong, Zhang Yin, Li Sha, et al. Research on intelligent line fault diagnosis method based on LSTM algorithm [J]. Large Electric Machine Technology, 2023, (S2): 62-67.]. Zhao Qi et al. proposed a transmission line fault identification method based on Gram angle field and ResNet, which has good fault identification effect [Zhao Qi, Wang Jian, Lin Fengkai, et al. Transmission line fault identification method based on Gram angle field and ResNet [J]. Power System Protection and Control, 2024, 52(10): 95-104.]. However, these deep learning models still reveal a series of problems that need to be solved in practical applications: First, they typically have high demands for computing resources. The large number of parameters and complex calculation processes require powerful hardware support, which makes their deployment on resource-constrained devices challenging and hinders their widespread application in scenarios such as edge computing.
[0009] Secondly, the dataset has poor generalization ability. The model performs well on specific datasets, but often experiences a performance decline when faced with new data. This is mainly because the model relies too much on the distribution of the training data and lacks an understanding of the essential characteristics of the data.
[0010] Finally, deep learning models are often regarded as "black boxes" because their internal decision-making processes are difficult to understand and explain, which brings great difficulties to the debugging, optimization and application of models in the field of power transmission. Summary of the Invention
[0011] The purpose of this invention is to overcome the shortcomings of the prior art and provide an intelligent identification method for line fault causes based on recorded waveform images. This method can directly utilize recorded waveform images of faults to quickly predict the causes of faults without relying on electrical parameters.
[0012] The technical problem solved by this invention is achieved through the following technical solution: A method for intelligent identification of line fault causes based on recorded waveform images, the method comprising the following steps: S1. Optimization of raw waveform recording data: Channel optimization, time optimization, and removal of redundant information are performed on the raw waveform recording data, and a dataset of waveform images of transmission line faults is constructed. S2. The Scale Invariant Feature Transform (SIFT) feature extraction method is used to extract features from the fault recording waveform image to obtain the image feature descriptors. Then, the feature descriptors are clustered to form a visual feature dictionary. Finally, the image pyramid technology is used to achieve multi-scale feature fusion and form a vectorized feature representation of the fault recording waveform image. S3. A support vector machine (SVM) classifier is introduced to classify the extracted features. The dataset constructed in S1 is used to train the fault cause identification network to obtain the line fault cause identification model. S4. Model Evaluation and Performance Optimization: Test the model's performance on a self-built real dataset, and perform data augmentation or data balancing operations on the training set to reduce the probability of overfitting.
[0013] Furthermore, the preferred channel for S1 is as follows: for the faulty line, select the three-phase voltage channel, three-phase current channel, and zero-sequence voltage and zero-sequence current channel corresponding to the faulty line. The above channel data serves as the core element for constructing and analyzing the waveform image. The preferred time for S1 is as follows: select one cycle before and after the fault time, and one cycle before the fault and ten cycles after the fault to completely and clearly capture the waveform changes at the time of the fault. The removal of redundant information in S1 is as follows: remove the fault channel description, the time numerical markers on the horizontal and vertical axes, and other non-waveform feature information.
[0014] Furthermore, during the construction of the S1 fault recording waveform image dataset, each image sample undergoes filtering, normalization, and data augmentation preprocessing.
[0015] Moreover, S2 specifically involves: using the Scale Invariant Feature Transform (SIFT) algorithm to perform key point detection and feature descriptor extraction on the preprocessed fault waveform image; then using the K-means clustering method to divide the extracted feature descriptors into K clusters to form a visual dictionary for the dataset; subsequently, using Gaussian pyramid construction technology to generate image sequences containing different resolution levels to achieve hierarchical fusion of multi-scale features, and finally using the generated feature vector as the input of the model to achieve intelligent identification of fault causes.
[0016] Furthermore, DSIFT feature point detection is performed using a global sliding window to extract feature descriptors that are insensitive to scale, rotation, and illumination changes. Specifically, feature point detection is performed using the dsift function with a fixed step size (step=8), and feature descriptors are extracted within a fixed neighborhood range (size=4), while generating floating-point feature descriptors. K-means clustering is used to divide the feature descriptors of the dataset into K clusters, such that each data point belongs to one and only one cluster, and similar fault feature vectors are grouped into the same class to form a visual dictionary. The Gaussian pyramid construction technique is used to perform multi-scale feature fusion on the extracted floating-point feature descriptors. Specifically, this involves setting pyramid levels and weights for each level, and then weighting and fusing feature descriptors of different scales according to the set weights to generate a pyramid vocabulary feature vector with a unified dimension.
[0017] Furthermore, the SVM classifier of S3 employs a radial basis function (RBF) optimization method, which includes optimizing the selection of kernel function parameters through grid search and introducing a cross-validation mechanism.
[0018] The positive effects that this invention can produce are: This invention can adaptively and automatically extract features from recorded waveform images, without relying on electrical parameters during line faults or requiring additional calculations on the recorded waveform time series data. It can rapidly predict fault causes using fault recorded waveform images, improving the accuracy and timeliness of fault cause identification. Through feature extraction, clustering operations, and spatial pyramid operations, it can more efficiently represent the features of fault waveform images, further improving the accuracy of fault identification. Through cross-validation and monitoring of model performance indicators, it can ensure the generalization performance and stability of the model, providing strong support for fault diagnosis and maintenance in power systems. Attached Figure Description
[0019] Figure 1 This is a flowchart of the present invention; Figure 2 This is an example image of dataset image enhancement for this invention; Figure 3 This is a network structure diagram of the present invention; Figure 4 This is a diagram illustrating the prediction results of the present invention. Detailed Implementation
[0020] The present invention will be further described in detail below through specific embodiments. The following embodiments are merely descriptive and not limiting, and should not be used to limit the scope of protection of the present invention.
[0021] An intelligent method for identifying the causes of line faults based on recorded waveform images is innovative in that the method comprises the following steps: Step 1: Construct the dataset: Collect over 5000 fault waveform images from power grid fault cases. Some example images are shown below. Figure 2 As shown, these images cover a variety of fault causes. To expand the dataset and enhance its diversity, various data augmentation operations were performed on the original images, including but not limited to phase sequence transformation. These operations not only increased the number of samples in the dataset but also improved the model's generalization ability and robustness. To achieve supervised learning and accurately evaluate the classifier's performance, detailed label information was recorded for each image sample. This labeling information was mainly accomplished by marking the image names. The name of each image consists of the true cause of the fault, the case code, and the fault recorder number. This naming convention facilitates data management and retrieval. A complete dataset was formed by converting the raw fault recorder data into image samples. After completing the construction and preprocessing of the dataset, it was divided according to a predetermined ratio, with the dataset divided into a training set and a test set in a 9:1 ratio. The training set was used for model training and parameter optimization, while the test set was used for model performance evaluation and validation.
[0022] Step two involves constructing an intelligent model for identifying the causes of line faults based on recorded waveform images, using Scale Invariant Feature Transform (SIFT) feature extraction combined with a spatial pyramid model. Specifically, firstly, dense SIFT feature extraction is performed on the image to generate a list of feature descriptors. These descriptors are then clustered using K-means clustering to generate a vocabulary representation of the image (e.g., K = 300). Next, a spatial pyramid model is constructed to divide the image into sub-images at different levels (e.g., 1×1, 2×2, 4×4), and a histogram representation is generated for each sub-image. The spatial pyramid model is then used to weight and fuse the histograms from different levels to generate the final pyramid histogram representation, completing all the processes of image feature vectorization. Finally, a Support Vector Machine (SVM) is used to train the efficiently represented multi-scale features to obtain a classification model. Through these steps, an intelligent model for identifying the causes of line faults using recorded waveform images is constructed, which can effectively improve the accuracy and reliability of fault cause identification.
[0023] Table 1 Function parameter configuration of the embodiment The symbols in Table 1 are explained as follows: `dsift` represents the dense SIFT feature extraction algorithm. Its three parameters are: sampling step size (controlling the sampling interval of keypoints); spatial window size of the SIFT descriptor (in pixels); and whether to generate floating-point descriptors (True for floating-point, False for integer). `KMeans` represents the clustering algorithm. Its three parameters are: the specified number of cluster centers; the specified method for initializing cluster centers ('k-means++' is an improved initialization method that can speed up convergence and improve clustering quality); and the seed for the random number generator to ensure the repeatability of the results. The `develop_pyramid_vocabulary` function constructs a vocabulary based on an image pyramid for multi-scale feature fusion. This function divides the image into sub-images of different levels and statistically analyzes the feature descriptors of each sub-image to generate a multi-scale histogram representation. The two parameters `pyramid_num=(1, 2, 4)` represent the number of sub-images in each pyramid layer. Each element of `weights` represents the weight of the corresponding level, used for weighted fusion of features from different levels. `StandardScaler` is used to standardize the data (also known as Z-score standardization). It eliminates the influence of different dimensions on the data by transforming it into a standard normal distribution with a mean of 0 and a standard deviation of 1, making the data more suitable for many machine learning algorithms.
[0024] Step 3: Train the prediction network using the constructed dataset to obtain an optimized fault cause prediction network model. The classifier trained uses the SVC classifier from sklearn, where rbf stands for Radial Basis Function, a commonly used non-linear kernel function that maps data to a high-dimensional space, making originally linearly inseparable data linearly separable. The penalty parameter C controls the model's tolerance for misclassification. When probability is set to True, SVC enables probability estimation, outputting the probability value for each class during prediction. This requires additional computation of the probability model during training.
[0025] The detection model obtained in this embodiment was tested using 168 test set images. The hardware platform used for testing consisted of an i7-14700 CPU with a clock speed of 5.4GHz and an NVIDIA RTX4070 GPU. The model's top-1 accuracy for each category and the average classification accuracy are recorded in Table 2.
[0026] Table 2. Top-1 accuracy of the model on the test set and test results for each category. As shown in Table 2, the intelligent fault cause identification model constructed in this invention can output corresponding predicted probabilities for various fault causes. Specifically, the test results for each category show some differences. For categories with low prediction accuracy, this indicates that using images for fault cause prediction in these categories may be challenging. Furthermore, due to the imbalance between samples from different categories, further model optimization or the collection of more data is needed to improve prediction performance. Overall, the evaluation results on the test set show that the model achieves an average accuracy of 0.875, demonstrating good detection performance.
[0027] To verify the robustness of this method on different datasets, the following robustness experiment was conducted. Specifically, 100 difficult-to-identify waveform data points from different regions, time periods, and fault causes were randomly selected from cases unrelated to the training dataset. After necessary annotation of these data, the trained model was used to identify the fault causes. According to the scoring criteria provided by the power grid, if the model's top two predictions (top 2) contain the actual fault cause, a score of 1 is awarded; if the model's prediction of the fault cause is correct in the first-level category but incorrect in the second-level category, a score of 0.5 is awarded. In this way, the scores for each category and the total score in the two tests were recorded, and the results are shown in Table 3.
[0028] Table 3 shows the model's total score and accuracy for each category on 100 cases. As shown in Table 3, this embodiment still achieves a high score in the prediction of different cases across the country, scoring 60 out of 100, demonstrating good detection performance.
[0029] This invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects as illustrative rather than restrictive, for example: 1) The corresponding fault channel is not limited to the corresponding three-phase voltage and current channel, zero-sequence channel, or switch quantity channel; 2) The operation types are not limited to the eight types mentioned in the examples, such as filtering, adding noise, and sharpening; 3) The network structure proposed in this invention can also be applied to image analysis tasks in other fields, and is not limited to the identification of line fault causes; 4) The selection of various dataset construction parameters and network configuration parameters is not limited to the configurations in the examples.
[0030] Although embodiments and drawings of the present invention have been disclosed for illustrative purposes, those skilled in the art will understand that various substitutions, variations and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the scope of the present invention is not limited to the contents disclosed in the embodiments and drawings.
Claims
1. A method for intelligent identification of line fault causes based on recorded waveform images, characterized in that: The steps of the method are as follows: S1. Optimization of raw waveform recording data: Channel optimization, time optimization, and removal of redundant information are performed on the raw waveform recording data, and a dataset of waveform images of transmission line faults is constructed. S2. The Scale Invariant Feature Transform (SIFT) feature extraction method is used to extract features from the fault recording waveform image to obtain the image feature descriptors. Then, the feature descriptors are clustered to form a visual feature dictionary. Finally, the image pyramid technology is used to achieve multi-scale feature fusion and form a vectorized feature representation of the fault recording waveform image. S3. A support vector machine (SVM) classifier is introduced to classify the extracted features. The dataset constructed in S1 is used to train the fault cause identification network to obtain the line fault cause identification model. S4. Model Evaluation and Performance Optimization: Test the model's performance on a self-built real dataset, and perform data augmentation or data balancing operations on the training set to reduce the probability of overfitting.
2. The intelligent identification method for line fault causes based on recorded waveform images according to claim 1, characterized in that: The channel optimization of S1 specifically involves selecting the three-phase voltage channel, three-phase current channel, and zero-sequence voltage and zero-sequence current channel corresponding to the faulty line. The channel data serves as the core elements for constructing and analyzing the waveform image. The time optimization of S1 specifically involves selecting one cycle before and after the fault, and ten cycles before and after the fault, to completely and clearly capture the waveform changes at the time of the fault. The redundant information removal of S1 specifically involves removing the fault channel description, time numerical markers on the horizontal and vertical axes, and other non-waveform feature information.
3. The intelligent identification method for line fault causes based on recorded waveform images according to claim 1, characterized in that: During the construction of the S1 fault recording waveform image dataset, each image sample undergoes filtering, normalization, and data augmentation preprocessing.
4. The intelligent identification method for line fault causes based on recorded waveform images according to claim 1, characterized in that: Specifically, S2 involves: using the Scale Invariant Feature Transform (SIFT) algorithm to perform key point detection and feature descriptor extraction on the preprocessed fault waveform image; then using the K-means clustering method to divide the extracted feature descriptors into K clusters to form a visual dictionary for the dataset; subsequently, using Gaussian pyramid construction technology to generate image sequences containing different resolution levels to achieve hierarchical fusion of multi-scale features, and finally using the generated feature vector as the input of the model to achieve intelligent identification of fault causes.
5. The intelligent identification method for line fault causes based on recorded waveform images according to claim 4, characterized in that: DSIFT feature point detection is performed using a global sliding window to extract feature descriptors that are insensitive to scale, rotation and illumination changes. Specifically, feature point detection is performed using the dsift function with a fixed step size (step=8), and feature descriptors are extracted within a fixed neighborhood range (size=4), while generating floating-point feature descriptors. K-means clustering is used to divide the feature descriptors of the dataset into K clusters, such that each data point belongs to one and only one cluster, and similar fault feature vectors are grouped into the same class to form a visual dictionary. The Gaussian pyramid construction technique is used to perform multi-scale feature fusion on the extracted floating-point feature descriptors. Specifically, this involves setting pyramid levels and weights for each level, and then weighting and fusing feature descriptors of different scales according to the set weights to generate a pyramid vocabulary feature vector with a unified dimension.
6. The intelligent identification method for line fault causes based on recorded waveform images according to claim 1, characterized in that: The SVM classifier in S3 employs a radial basis function (RBF) optimization method, which includes optimizing the kernel function parameters through grid search and introducing a cross-validation mechanism.