Insulator defect detection method
The integration of S-CBAM attention and GSConv modules with a small target detection layer addresses the limitations of existing insulator defect detection methods, enhancing precision and efficiency in automated power line inspections.
Patent Information
- Application Number
- CN202510500150.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, the insulator defect detection method has insufficient detection accuracy, insufficient extraction of key feature information, and high model complexity and large parameters. The traditional method performs poorly in complex scenarios and is difficult to meet the real-time processing needs.
Using an improved S-CBAM attention mechanism for synchronous channel and spatial attention information extraction, combined with a lightweight GSConv module and a small object detection layer, a lightweight feature fusion network is built to detect insulator defects through a deep learning object detection algorithm.
It realizes high-precision insulator defect detection, reduces false detection and missed detection rates, improves detection efficiency and safety, adapts to complex backgrounds, and is suitable for intelligent upgrades of drone inspection systems.
Smart Images

Figure CN120318212A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning object detection, specifically refers to the insulator defect detection applied to transmission lines, is suitable for automatic maintenance and safety monitoring of transmission lines, and can detect insulator defects in real time and accurately to ensure the safe and stable operation of the power grid. Background Art
[0002] Traditional insulator defect detection methods mainly rely on manual inspection or detection methods based on traditional computer vision. The manual inspection method is limited by the complex spatial distribution of the ultra-large-scale transmission network, and has problems such as low detection accuracy, poor efficiency, high cost, and high safety risks, making it difficult to meet the needs of efficient operation and maintenance of modern power grids. The detection method based on traditional computer vision relies on manually extracted features, and the extraction of these features is usually designed based on specific scenarios, so the generalization ability is weak. In unknown or complex environments, its performance may not be good. At the same time, complex scenarios require the design of complex feature extraction and classification mechanisms, which increases the difficulty of algorithm implementation, and may face efficiency problems for real-time processing application scenarios.
[0003] With the development of intelligent inspection technology, the vision detection technology based on artificial intelligence provides a new direction for the intelligent operation and maintenance of transmission lines. The drone inspection system combined with the deep learning object detection algorithm can efficiently identify insulator defects, significantly improve the inspection accuracy and efficiency, and reduce the operation and maintenance costs and safety risks. Summary of the Invention
[0004] The present invention aims to solve the deficiencies in the existing insulator defect detection methods, such as low detection accuracy for small targets, insufficient extraction of key feature information, and high model complexity and large number of parameters. An insulator defect detection method with an improved attention mechanism and a lightweight feature fusion network is proposed. This method realizes high-precision defect detection and positioning, and reduces the false detection and missed detection rates by integrating the S-CBAM attention mechanism for improved synchronous channel and spatial attention information extraction, constructing a lightweight GSConv module, and building a small target detection layer.
[0005] The present invention provides an insulator defect detection method combining an improved attention mechanism and a lightweight feature fusion network, which has high detection accuracy. The specific technical solution is as follows:
[0006] An insulator defect detection method, characterized by comprising the following steps:
[0007] a) Obtain aerial images of insulators and preprocess the images;
[0008] b) Input the preprocessed image into a deep detection model composed of a backbone network, a neck network, and a detection head. In the backbone network, a synchronous channel-spatial attention mechanism (S-CBAM) is embedded in parallel. The neck network adopts the feature pyramid network (FPN) and path aggregation network (PAN) structures, and replaces the ordinary convolution therein with lightweight GSConv modules. On the basis of the standard three-scale detection branch, the detection head adds a small target detection branch;
[0009] c) Use the stochastic gradient descent (SGD) optimizer to train the model. During the training process, use the cosine annealing learning rate scheduler, and use CIoU loss, Focal Loss, and binary cross-entropy loss respectively;
[0010] d) Export the trained model and convert it into a TensorRT inference engine for accelerated inference on an embedded device;
[0011] e) Perform forward inference on the new aerial image, output the candidate defect boxes and their class confidence levels, and obtain the final detection results through non-maximum suppression (NMS).
[0012] Further, in step a), the preprocessing includes:
[0013] Scale the long side of the image proportionally to 1280 pixels;
[0014] Crop a 1024×1024 pixel area from the center of the scaled image;
[0015] Perform Gaussian smoothing filtering (σ = 0.5) and histogram equalization on the cropped image in sequence.
[0016] Further, in step a), further perform online data augmentation on the image. The augmentation operations are randomly executed according to probability. The augmentation operations include: horizontal or vertical flipping, random cropping (cropping ratio 70% - 100%), random rotation (±30°), brightness / contrast adjustment (0.8 - 1.2 times), hue / saturation transformation, and salt-and-pepper noise addition.
[0017] Further, the S-CBAM attention mechanism includes:
[0018] A channel attention branch, which respectively obtains channel feature descriptions through global average pooling and global max pooling, generates channel weights through a shared multi-layer perceptron (MLP), and multiplies them with the original feature map channel by channel;
[0019] A spatial attention branch, which obtains spatial descriptions through max pooling and average pooling in the channel dimension, generates spatial weights through a 7×7 convolution after splicing, and multiplies them with the original feature map pixel by pixel;
[0020] The channel recalibration result and the spatial recalibration result are concatenated by channel and then fused through a 1×1 convolution to output the final feature map.
[0021] Furthermore, the GSConv lightweight module includes:
[0022] The first branch is a 1×1 ordinary convolution module with the output channel number being half of the target channel number;
[0023] The second branch is a depthwise separable convolution module, including a 3×3 depth convolution and a 1×1 pointwise convolution, with the output channel number being half of the target channel number;
[0024] After concatenating the outputs of the two branches in the channel dimension, a channel shuffle operation is performed to obtain a fused feature map with the same number of channels as the target channel number.
[0025] Furthermore, on the basis of the original three-scale detection branches of P3(80×80), P4(40×40), and P5(20×20), the detection head adds a small target detection branch of P2(160×160). The P2 branch is obtained by concatenating the high-resolution feature output by the second layer of the backbone network and the same-resolution feature upsampled from the deep feature, followed by convolution fusion and prediction of the defect box and category. Small target-specific anchors (8,8), (16,16), and (32,32) are used.
[0026] Furthermore, in step c), the initial learning rate of the SGD optimizer is 0.01, the momentum coefficient is 0.937, and the weight decay coefficient is 5×10 -4 ; The cosine annealing strategy makes the learning rate smoothly decrease from 0.01 to 0.001 within 100 epochs; the batch size is 16; automatic mixed precision (AMP) is enabled during model training to accelerate the operation.
[0027] Furthermore, in step c), the loss functions used include:
[0028] CIoU loss is used for bounding box regression;
[0029] Focal Loss (α = 0.25, γ = 2.0) is used for class classification;
[0030] Binary Cross-Entropy is used for object confidence prediction.
[0031] Furthermore, in step e), the IoU threshold of the non-maximum suppression is 0.5, and the confidence threshold is 0.4, to filter and fuse candidate boxes with too high overlap or low confidence, and output the final defect detection result.
[0032] Furthermore, it also includes oversampling defective samples or synthesizing new defective images using a generative adversarial network (GAN) before or during model training to balance the proportion of positive and negative samples and improve the generalization ability of the model.
[0033] The single-stage YOLO object detection model based on deep learning in the present invention has made various innovations in response to the deficiencies existing in the existing insulator defect detection methods. The specific innovation points and features are as follows:
[0034] (1) Introduction of the S-CBAM attention mechanism:
[0035] CBAM is an attention mechanism module for convolutional neural networks, aiming to improve network performance by focusing on important features and suppressing unimportant features. Its innovation lies in combining two mechanisms: channel attention and spatial attention. It weights the feature map from the channel dimension and the spatial dimension respectively, thereby enhancing the model's ability to capture key information. However, the traditional CBAM module adopts a sequential processing method when extracting channel attention and spatial attention: first, the input feature map is input into the channel attention module to generate channel attention weights; then, the channel attention weights are multiplied element-wise with the original input feature map to obtain the weighted feature map; finally, it is input into the spatial attention module for spatial feature extraction. This sequential processing method may lead to a problem: after channel attention weighting, some spatial feature information may be lost, so that it cannot be fully utilized by the subsequent spatial attention module, limiting the model's ability to model spatial information. To solve this problem, the present invention improves the structure of CBAM and proposes an S-CBAM attention mechanism for synchronously extracting channel and spatial feature information. S-CBAM simultaneously inputs the input feature map into the channel attention module and the spatial attention module respectively, thereby maximizing the retention of the channel information and spatial information of the original input feature map. This synchronous processing method enables channel attention and spatial attention to independently extract key information from the input feature map, avoiding possible information loss in sequential processing.
[0036] (2) Construction of the GSConv lightweight convolution module:
[0037] The GSConv module combines ordinary convolution with depthwise separable convolution, enabling the output of depthwise separable convolution to be as close as possible to that of ordinary convolution. Thus, while reducing the computational cost and the number of parameters of the model, it maintains a high detection accuracy. The present invention makes a lightweight improvement to the feature fusion network and achieves this goal by designing the GSConv module and replacing the ordinary convolution in the Neck structure. The core of this improvement lies in that in the Backbone, during the conversion of the feature map from spatial information to the channel dimension, each spatial compression and channel expansion will result in the loss of semantic information. Dense convolution calculations can maximize the retention of the hidden connections between channels, while sparse convolution will cut off these connections. Therefore, using GSConv in the Backbone will further exacerbate the loss of semantic information. However, in the Neck part, the channel dimension of the feature map has reached the maximum, and the spatial dimension has been compressed to the minimum. At this time, complex conversion operations are no longer required. Therefore, using GSConv in the Neck can reduce the computational cost while retaining sufficient semantic information.
[0038] In the GSConv module structure, the input feature map first passes through a module composed of a two-dimensional convolutional layer, Batch Normalization (BN), and the SiLU activation function to generate a feature map with half the number of channels of the final output channels. Secondly, this feature map is processed through a depthwise separable convolution module, and the output result is concatenated with the input feature map along the channel dimension. Finally, the concatenated feature map is subjected to channel rearrangement through a Shuffle operation to obtain the final output. This design enables the information of ordinary convolution to be fully integrated into the output of depthwise separable convolution, thus retaining rich feature information while reducing the computational complexity.
[0039] (3) Add a small object detection layer:
[0040] Insulator defects in aerial images usually have an extremely low pixel ratio, resulting in insufficient resolution in the target area during the feature extraction process and serious loss of feature information. As a result, it is difficult to be effectively detected and recognized, easily leading to problems of missed detection and false detection. To solve this problem, by deeply analyzing the feature distribution of small targets, it is found that the shallow feature maps of the model contain higher-resolution detailed information, which is crucial for the detection of small targets. Inspired by this, the present invention makes full use of the shallow feature maps of the backbone network and combines the feature pyramid structure for multi-scale feature fusion to enhance the feature representation ability for small targets. Specifically, an independent branch is introduced from the shallow layer of its backbone network to construct an additional detection layer specifically for small target detection. First, the initial layer features are extracted from the second layer of the Backbone to capture the detailed information of small targets, such as edges, textures, etc. However, although the shallow feature maps have high resolution, they lack high-level semantic information, which may affect the discriminative ability of target categories. Therefore, the feature pyramid structure is introduced for multi-scale feature fusion to supplement the high-level semantic information. Since the size of the feature map corresponding to the newly added detection layer is larger than the existing feature maps, it is necessary to perform an upsampling operation on the existing feature maps to achieve the matching of feature map sizes. After the 18th layer in the Neck part, additional convolutional layers and upsampling layers are added to adjust the size and number of channels of the feature maps. Next, at the 21st layer, the upsampled feature maps are fused with the shallow feature maps extracted from the second layer using the Concat operation. This fusion method combines the high-resolution detailed information of the shallow features and the rich semantic information of the deep features, thus generating a more discriminative feature representation. Finally, at the 22nd layer after fusion, the boundary boxes, categories, and confidence predictions of the targets are generated through convolutional layers to form a detection head optimized specifically for small targets. The improved method can significantly improve the detection performance of small targets, especially in complex scenarios with dense distribution of small targets, and the detection effect has been significantly improved.
[0041] (4)Automated precise and efficient detection:
[0042] By introducing the object detection algorithm of deep learning to automatically analyze aerial insulator images, various types of insulator defects, such as cracks, breakages, and dirt accumulations, can be quickly and accurately identified. Compared with traditional manual inspections, the vision detection technology based on artificial intelligence is not only more convenient and fast but also can identify subtle defects that are difficult to detect by the human eye, thus enabling more effective preventive maintenance. At the same time, the risk of manual high-altitude operations is avoided, significantly improving the safety of inspection operations and ensuring the life safety of inspection personnel. The combination of drone inspections and artificial intelligence analysis technology can complete large-scale and high-precision detection tasks in a short time.
[0043] (5)Robustness to adapt to complex backgrounds:
[0044] Traditional detection methods often exhibit high false detection and missed detection rates when faced with an environment with dense distribution and complex and variable backgrounds. The S-CBAM attention mechanism of the present invention, which synchronously extracts channel and spatial attention information, can improve the feature extraction ability, capture key feature information in complex environments, and additionally introduce a small target detection layer to improve the detection accuracy of small targets. On the basis of combining multi-scale feature fusion, the model can still accurately and efficiently detect insulator defects when faced with complex insulator defects, background interference or extreme environments, avoid misdetecting non-insulator defect targets, and provide stable and reliable detection results.
[0045] (6) Scalability for engineering applications:
[0046] The insulator defect detection method of the present invention has strong scalability. It is not only applicable to the detection of insulator defects, expanding the application depth of target detection algorithms in the field of power inspection, but also provides an important theoretical basis and practical reference for the intelligent transformation of related industries, and has broad application prospects and promotion value.
[0047] The innovation and characteristics of the present invention lie in the combination of the single-stage YOLO target detection model of deep learning. On this basis, aiming at the deficiencies of traditional detection methods, improvement strategies are proposed for the backbone feature extraction network, neck feature fusion network, and head detection network. On the premise of maintaining detection real-time performance, the accuracy and robustness of insulator defect detection are greatly improved, providing strong technical support and guarantee for the intelligent upgrade of the UAV inspection system. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is the overall network model structure diagram of the present invention
[0049] Figure 2 is the overall structure diagram of S-CBAM of the present invention
[0050] Figure 3 is the channel attention module diagram of the present invention
[0051] Figure 4 is the spatial attention module diagram of the present invention
[0052] Figure 5 is the GSConv module structure diagram of the present invention DETAILED DESCRIPTION OF THE INVENTION
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] A data set is established, which contains 848 static images, including 600 normal insulator images and 248 images with defects such as cracks, breakages, or fouling. The acquisition environment covers various lighting and climate conditions such as sunny days, cloudy days, rainy days, and low-light at night, ensuring sufficient diversity of the samples.
[0055] Basic image processing
[0056] It includes long-side scaling: Each image is scaled proportionally according to the long-side size so that its long side is exactly 1280 pixels.
[0057] And central cropping: A region of size 1024×1024 is cropped from the central area of the scaled image to unify the input size of all pictures.
[0058] And noise reduction and contrast enhancement: For the cropped image, first apply mild Gaussian smoothing filtering to remove sensor noise, and then perform histogram equalization processing to enhance the contrast of bright and dark details.
[0059] Pixel value standardization
[0060] First, the gray or color channel value of each pixel is converted into a decimal between 0 and 1, and then according to the standard of common pre-trained models, the mean value of each channel is subtracted and corrected, and divided by the corresponding standard deviation value to make the data distribution more in line with the input requirements of the deep network.
[0061] Online data augmentation
[0062] During the training process, in order to prevent the model from overfitting to fixed samples and improve the robustness to defects under different shooting conditions, each image is randomly augmented in various ways with a certain probability, including:
[0063] Horizontal or vertical flipping;
[0064] Random cropping and then restoring to the target size;
[0065] Random angle rotation, up to no more than 30 degrees;
[0066] Randomly adjust the brightness and contrast, varying between 80% and 120%;
[0067] Randomly change the hue and saturation to simulate different light source conditions;
[0068] A small amount of salt-and-pepper noise is added to enhance the resistance to small-scale noise.
[0069] In this way, the model may see slightly different samples in each training iteration, greatly improving the generalization ability of the model.
[0070] Dataset division and annotation
[0071] The processed and enhanced images are expanded to 4716, and divided into a training set (about 3763) and a validation set (about 953) according to an eight-to-two ratio. Use a professional annotation tool to re-check and annotate the rectangular boxes of the defective targets, record the defect category and its position and size in the image for each rectangular box, and finally generate a label file that conforms to the general object detection format, and place the images and label files in the agreed directory structure for subsequent data loading.
[0072] Based on the lightweight YOLOv5s architecture, this invention parallelly introduces an improved attention module S-CBAM after several key layers of its backbone network. This module divides the input features Figure 1 into two parts:
[0073] including the channel attention branch. First, average pooling and max pooling are performed on the feature map in the spatial dimension to obtain two sets of channel descriptions representing different statistical characteristics;
[0074] These two sets of descriptions are respectively processed by the same set of small fully connected layers and then fused to generate a weighted coefficient for the importance of each channel;
[0075] Finally, the original feature map is multiplied by the channel coefficient to highlight the important channels and suppress the irrelevant channels.
[0076] and the spatial attention branch. The maximum value and the average value are respectively taken for the feature map in the channel dimension to obtain two single-channel spatial description maps;
[0077] After splicing these two description maps and performing an operation with a larger convolution kernel, a weight map for the importance of each spatial position is obtained;
[0078] The weight map is multiplied by the original feature map pixel by pixel, making the network more sensitive to key regions.
[0079] Finally, the weighted results of the channel and spatial paths are spliced or added in the channel direction, and then the number of channels is unified through a small convolution to output the final backbone features, which not only retain the semantic information in the channel dimension but also do not lose the spatial structure details.
[0080] Through this parallel rather than traditional sequential attention calculation method, the complete information of the input features is retained to the greatest extent, making the network more accurate in extracting small cracks and weak defects.
[0081] After the backbone network, the classical Feature Pyramid Network (FPN) and Path Aggregation Network (PAN) structures are adopted to complete the interaction of multi-level semantic and detailed information. To reduce the computational cost and model parameters, a self-developed lightweight GSConv module is introduced to replace the original conventional convolution. The GSConv module contains two parallel branches:
[0082] The first branch uses a small-sized ordinary convolutional kernel to compress the input feature map.
[0083] The second branch first extracts spatial information through depthwise separable convolution and then fuses each channel through pointwise convolution;
[0084] Finally, the outputs of the two branches are concatenated in the channel dimension and mixed grouped through a channel rearrangement operation to obtain an expressive ability similar to that of the original conventional convolution, while significantly reducing the number of parameters and the computational amount.
[0085] The standard YOLOv5s outputs feature maps of three different resolutions at the neck for detecting large, medium, and small targets.
[0086] For the insulator cracks and stains to be detected in the present invention, which are often extremely tiny, an additional high-resolution feature channel is introduced at the second level of the backbone network for small target detection:
[0087] High-resolution detailed features are obtained from the shallow layer of the backbone;
[0088] The features of the deepest layer of the network are successively enlarged to the same resolution;
[0089] After the two-way features are concatenated in the channel dimension, they are fused through multiple small convolutions;
[0090] An independent prediction branch is set on the fused features to output the fine-grained bounding box position, class probability, and confidence.
[0091] Through this additional small target branch, the network can better capture tiny defects that only occupy a few pixels, significantly reducing missed detections.
[0092] The optimizer and learning rate scheduler adopt the Stochastic Gradient Descent (SGD) optimizer, and momentum and moderate weight decay are used to improve the convergence performance. The initial learning rate is set to 0.01 and gradually decreases according to the "cosine annealing" principle during training - that is, it first decreases gently to a certain level and then approaches the minimum value of 0.001 in the later stage to avoid the training getting stuck in a local optimum prematurely.
[0093] The loss function design includes using a comprehensive localization loss that takes into account distance, overlap, and aspect ratio for bounding box regression to improve the accuracy of box regression;
[0094] The category classification branch combines the cross - entropy of the classification criterion with the "focal loss" mechanism, applying a higher weight to the hard - to - distinguish negative samples, thereby alleviating the impact of class imbalance on training;
[0095] The confidence branch uses binary cross - entropy to measure the accuracy of predicting the presence or absence of the target.
[0096] In the training set, five - fold cross - validation is used to evaluate the model performance to ensure its robustness to different subsets; meanwhile, an early stopping mechanism is set - when the validation set loss fails to decrease effectively within 20 consecutive training epochs, the training is automatically stopped to prevent overfitting.
[0097] The model finally uses the mean average precision (mAP@0.5), recall rate, and precision rate as the main indicators, where the recall rate aims to be higher than 90%, the precision rate is not less than 85%, and it can reach at least 80 frames per second under single - card half - precision inference conditions to meet the real - time inspection requirements of drones.
[0098] The trained network weights are exported to a common intermediate format, and a semi - precision or integer quantization optimization is performed on the edge device using an inference acceleration framework to reduce the video memory occupancy and speed up the forward inference speed.
[0099] The edge real - time inference process includes
[0100] The drone camera acquires images in real - time and compresses them into a standard format;
[0101] The image decoding, scaling, and normalization are completed on the edge processor;
[0102] The quantized network is used for forward calculation, outputting candidate defect bounding boxes and corresponding confidences;
[0103] Non - maximum suppression is performed on all candidate results to remove redundant boxes with too high overlap and retain the detection results with the highest confidence;
[0104] The final defect bounding boxes and class confidences are drawn on the original image and real - time transmitted back to the ground station via a wireless network.
[0105] When the types of defects to be detected increase, only new categories need to be added to the last prediction branch, and it can be quickly adapted through fine - tuning with a small number of new samples;
[0106] If the computing power or storage of the embedded device is limited, the width or number of layers of the backbone network can be further trimmed, or lighter attention and fusion modules can be selected to obtain lower resource occupancy.
[0107] The following aspects, including the hardware environment, software environment, dataset preparation and preprocessing, model structure design, model training strategy, model deployment and operation, and engineering application considerations, are expected to provide a comprehensive and operable technical basis for the implementation and subsequent popularization and application of the present invention.
[0108] The hardware environment configuration includes a server host, including a GPU card, preferably a high-performance graphics acceleration card with a video memory ≥ 24GB (such as NVIDIA A100 40GB or equivalent models), to meet the training and inference requirements of large-scale deep learning models.
[0109] A CPU processor, at least equipped with a 16-core Intel Xeon or AMD EPYC series, with a main frequency ≥ 2.2GHz, to support multi-threaded data preprocessing and model inference pipelines.
[0110] The system memory is ≥ 128GB DDR4 to ensure the stability of large-scale data loading and multi-process concurrent operations.
[0111] A storage device, configured with a 2TB NVMe SSD, is used to store training datasets, model weight files, and intermediate caches, with a read and write speed ≥ 3000MB / s, to avoid affecting the training efficiency due to I / O bottlenecks.
[0112] And a drone and an edge computing unit, including an aerial photography platform, equipped with a high-resolution camera (pixels ≥ 20 million), with automatic exposure and autofocus functions, and supporting H.265 / H.264 video output.
[0113] An embedded processor, using Jetson AGX Xavier or an equivalent embedded AI processor, with 512 CUDA cores and 64 Tensor cores built-in, supporting TensorRT accelerated inference, and a power consumption ≤ 30W.
[0114] For network communication, a gigabit Ethernet and a 5GHz wireless module are equipped to achieve low-latency and reliable data transmission between the drone and the ground station.
[0115] The software environment setup includes a deep learning framework, including PyTorch, with a version ≥ 1.10, making full use of its dynamic graph mechanism and C++ / CUDA backend optimization;
[0116] And CUDA, with a version of 11.3, compatible with the NVIDIA driver (≥ 450.xx), to achieve efficient GPU computing.
[0117] It also includes an image processing and enhancement library, including OpenCV 4.5 for preprocessing operations such as reading, scaling, color space conversion, and ROI cutting of raw images.
[0118] And Albumentations, which is responsible for image data augmentation, supports the following operations:
[0119] Gaussian noise: Simulate the noise in the transmission environment;
[0120] Random horizontal / vertical flipping;
[0121] Random cropping: The cropping ratio is between 70% and 100%;
[0122] Rotation transformation: ±30°;
[0123] Brightness adjustment: 0.8 to 1.2 times;
[0124] Other auxiliary libraries include NumPy and Pandas for dataset statistics and visualization;
[0125] Used to convert the trained PyTorch model into a high-performance inference engine, specifically TensorRT;
[0126] Matplotlib for drawing the loss curve and accuracy curve during training.
[0127] The original data source of this embodiment uses the "China Power Transmission Line Insulator Dataset" (CPLID), which contains 600 normal insulator images and 248 defective insulator images, for a total of 848 images.
[0128] Through various combinations of augmentation using the Albumentations library, it is amplified to 4,716 images;
[0129] Training / testing division, divided into a training set of 3,763 images and a testing set of 953 images according to an 8:2 ratio;
[0130] Data balancing strategy. For the problem of few defective samples, the defective class can be oversampled in the training set, or synthetic defective images can be generated using GAN to supplement, to prevent model bias caused by class imbalance;
[0131] In this embodiment, all images are uniformly scaled to 640×640 pixels during preprocessing;
[0132] The pixel values are normalized to [0, 1] or standardized based on the mean and variance of ImageNet;
[0133] Random shuffling (shuffle), and multi-threaded data loading (num_workers≥8) is used to improve the training efficiency;
[0134] The backbone network is based on YOLOv5s, and the S-CBAM attention mechanism is embedded in parallel in the deep layers of the backbone.
[0135] Channel attention branch: First, perform average pooling and max pooling on the feature map, generate channel weight vectors through fully connected layers respectively, and then multiply them with the original feature map channel by channel;
[0136] Spatial attention branch: Perform 7×7 convolution on the original feature map to generate a spatial weight map, and multiply it with the original feature map pixel by pixel;
[0137] Parallel fusion: Add or concatenate the two weighted results element-wise, and then input them into the next layer to retain the original channel and spatial information to the greatest extent.
[0138] The Neck network in this embodiment uses the FPN+PAN structure for multi-scale feature fusion; replace the conventional convolution with the self-developed GSConv lightweight module:
[0139] Ordinary convolution: First, use Conv+BN+SiLU activation on the input feature map to generate a feature with C_out / 2 channels;
[0140] Depthwise separable convolution: Process the same input in parallel and retain the spatial receptive field;
[0141] Concat&Shuffle: Concatenate the two outputs in the channel dimension, and then perform Shuffle group rearrangement to generate the final feature map;
[0142] The GSConv module reduces the number of parameters and computational amount by about 40% while maintaining a semantic expression ability similar to that of ordinary convolution.
[0143] In this embodiment, the adopted detection head (Head) adds a fourth detection layer on the basis of the standard three-scale detection head, and is specifically optimized for small targets (such as cracks and stains):
[0144] Feature source: Extract a high-resolution shallow feature map (resolution 160×160) from the second layer of the Backbone;
[0145] Semantic supplement: Upsample the feature map of the 18th deep layer, and then splice it with the shallow feature;
[0146] Additional detection head: Add another group of Conv layers on the spliced feature map to generate bounding box, class, and confidence predictions;
[0147] Multi-scale fusion: The candidate boxes of four scales jointly participate in NMS (IoU threshold 0.5, confidence threshold 0.4) screening, taking into account the accuracy of large, medium, and small targets.
[0148] During the model training process, the optimizer uses SGD, with an initial learning rate of 0.01, momentum of 0.937, and weight decay of 0.0005;
[0149] The learning rate scheduling adopts the cosine annealing strategy (Cosine Annealing), with 100 epochs of training and a minimum learning rate of 0.001;
[0150] Bounding box regression loss: Use CIoU Loss to improve the localization accuracy;
[0151] Classification loss: Combine Focal Loss to alleviate the extremely imbalanced problem between positive and negative samples;
[0152] The training process includes:
[0153] Five-fold cross-validation: Perform 5-fold CV on the training set to evaluate the model's stability and generalization ability;
[0154] If the validation set loss does not decrease significantly for 20 consecutive times, terminate the training early to avoid overfitting;
[0155] Set it to 16 or 32 when the video memory is sufficient. In case of video memory bottlenecks, it can be reduced to 8, and mixed-precision training (FP16) is enabled;
[0156] Use TensorBoard to monitor the training / validation loss, mAP@0.5, Recall, and Precision curves in real time;
[0157] Regularly save the best model weights (based on the highest mAP value of the validation set);
[0158] The log includes GPU occupancy, CPU occupancy, memory occupancy, time consumption per iteration, etc., for subsequent performance tuning.
[0159] Export the trained PyTorch model to the ONNX format, and then optimize it into an efficient inference engine using TensorRT;
[0160] Configure INT8 or FP16 quantization to perform accelerated inference performance testing on edge embedded platforms;
[0161] The drone-side deployment includes:
[0162] Integrate the TensorRT engine into the embedded processor of the drone;
[0163] It can reach ≥120FPS at 640×640 input, meeting the real-time requirements of inspection;
[0164] After the drone camera captures an image, it is automatically scaled to 640×640 and sent to the inference module through the USB3.0 or CSI interface;
[0165] After the inference is completed, the coordinates of the defect bounding box, the class label, and the confidence level are returned. Real-time annotation is performed using a red rectangular box and text, and the results are transmitted back to the ground station via a wireless link together with the GPS coordinates.
[0166] After the ground station receives the detection results, it displays the defect location and type in a graphical interface, and evaluates the severity level according to the confidence level and the size of the defect area.
[0167] An inspection report is automatically generated, including:
[0168] Defect coordinates (tower pole number + longitude and latitude);
[0169] Defect type (crack, breakage, contamination, etc.);
[0170] Severity level (mild, moderate, severe);
[0171] Treatment suggestions (re-inspection, cleaning, repair).
[0172] It should be noted that if there are more defect types in the actual working conditions, the defect categories can be dynamically expanded, and oversampling or GAN can be used to synthesize new samples to continuously improve the generalization ability of the model.
[0173] Under different lighting, seasons, and terrain environments, the dataset should be updated regularly and the model should be retrained or fine-tuned to ensure robustness.
[0174] It should be noted that when the video memory is insufficient, one or more of the following solutions can be adopted:
[0175] Reduce the Batch size (it is recommended not to be less than 4);
[0176] Enable mixed-precision training (automatic mixed precision);
[0177] Crop the model structure to appropriately reduce the channel width or the number of layers of the detection head.
[0178] In addition, in addition to insulator defect detection, this solution is also transferable to the defect detection of other power equipment (such as wires, iron towers, lightning arresters, etc.); it can be combined with inspection platforms such as unmanned vehicles and wall-climbing robots to further expand the application scope; through the cloud-edge collaborative architecture, an organic combination of offline batch training and online real-time inference can be achieved.
[0179] The above describes the present invention and its implementation manners. Such a description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design a structural manner and an embodiment similar to this technical solution without creative work without departing from the gist of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A method for detecting insulator defects, characterized in that, It includes the following steps: a) Obtain aerial insulator images and preprocess the images; b) Input the preprocessed images into a depth detection model composed of a backbone network, a neck network, and a detection head. In the backbone network, a synchronous channel–spatial attention mechanism (S-CBAM) is embedded in parallel. The neck network adopts the Feature Pyramid Network (FPN) and Path Aggregation Network (PAN) structures, and replaces the ordinary convolutions therein with lightweight GSConv modules. On the basis of the standard three-scale detection branches, the detection head adds a small target detection branch; c) Use the Stochastic Gradient Descent (SGD) optimizer to train the model. During the training process, use the cosine annealing learning rate scheduler, and use the CIoU loss, Focal Loss, and binary cross-entropy loss respectively; d) Export the trained model and convert it into a TensorRT inference engine for accelerated inference on an embedded device; e) Perform forward inference on new aerial images, output candidate defect boxes and their class confidence levels, and obtain the final detection results through Non-Maximum Suppression (NMS).
2. The insulator defect detection method according to claim 1, wherein, In the step a), the preprocessing includes: Scale the long side of the image proportionally to 1280 pixels; Crop a 1024×1024 pixel area from the center of the scaled image; Successively perform Gaussian smoothing filtering (σ = 0.5) and histogram equalization on the cropped image.
3. The insulator defect detection method according to claim 1, wherein In the step a), further perform online data augmentation on the image. The augmentation operations are randomly executed according to probabilities. The augmentation operations include: horizontal or vertical flipping, random cropping (cropping ratio 70% - 100%), random rotation (±30°), brightness / contrast adjustment (0.8 - 1.2 times), hue / saturation transformation, and salt-and-pepper noise addition.
4. The insulator defect detection method according to claim 1, wherein The S-CBAM attention mechanism includes: A channel attention branch, which respectively obtains channel feature descriptions through global average pooling and global maximum pooling, generates channel weights through a shared multi-layer perceptron (MLP), and multiplies them with the original feature map channel by channel; A spatial attention branch, which obtains spatial descriptions through maximum pooling and average pooling in the channel dimension, generates spatial weights through a 7×7 convolution after splicing, and multiplies them with the original feature map pixel by pixel; Fuse the channel recalibration result and the spatial recalibration result by concatenating them in the channel dimension and then performing a 1×1 convolution to output the final feature map.
5. The insulator defect detection method according to claim 1, wherein, The GSConv lightweight module includes: The first branch is a 1×1 ordinary convolution module with the output channel number being half of the target channel number; The second branch is a depthwise separable convolution module, including a 3×3 depth convolution and a 1×1 pointwise convolution, with the output channel number being half of the target channel number; After concatenating the outputs of the two branches in the channel dimension, perform a channel shuffle operation to obtain a fused feature map with the same number of channels as the target channel number.
6. The insulator defect detection method according to claim 1, wherein, Based on the original three-scale detection branches of P3(80×80), P4(40×40), and P5(20×20), the detection head adds a small target detection branch of P2(160×160). The P2 branch splices the high-resolution features output by the second layer of the backbone network and the same-resolution features obtained by upsampling from the deep features, then performs convolution fusion and predicts the defect bounding boxes and categories, using dedicated anchors for small targets (8,8), (16,16), and (32,32).
7. The insulator defect detection method according to claim 1, wherein In step c), the initial learning rate of the SGD optimizer is 0.01, the momentum coefficient is 0.937, and the weight decay coefficient is 5×10 -4 ; the cosine annealing strategy smoothly decreases the learning rate from 0.01 to 0.001 within 100 epochs; the batch size is 16; automatic mixed precision (AMP) is enabled during model training to accelerate the operation.
8. The insulator defect detection method according to claim 1, characterized in that In step c), the loss functions used include: CIoU loss is used for bounding box regression; Focal Loss(α = 0.25, γ = 2.0) is used for class classification; Binary Cross-Entropy is used for object confidence prediction.
9. The insulator defect detection method according to claim 1, wherein In step e), the IoU threshold for non-maximum suppression is 0.5, and the confidence threshold is 0.4, to filter and fuse candidate boxes with too high overlap or low confidence, and output the final defect detection result.
10. The insulator defect detection method according to claim 1, wherein, It also includes oversampling the defect samples or synthesizing new defect images using a generative adversarial network (GAN) before or during model training, to balance the ratio of positive and negative samples and improve the generalization ability of the model.
Citation Information
Cited By
Surface defect detection method and system based on deep learning
CN121032953A
Intelligent biopsy area prompting method and system for digestive endoscopy
CN121236371A
Sub-cartridge case surface defect detection method and image training and reasoning integrated platform
CN121304597A
Power scene power transmission line inspection method and system
CN121640353A
Small mechanical part defect visual detection system based on YOLO lightweight
CN121921281A