Lightning arrester infrared thermogram defect detection method based on improved MobileNetV3-SSD
By improving the MobileNetV3-SSD model, the multi-scale feature fusion module and the CBAM attention module are introduced, the accuracy and speed of infrared thermal image defect detection of lightning arrester is solved, and the rapid diagnosis and higher recognition accuracy of infrared thermal image faults of lightning arrester is achieved.
Patent Information
- Application Number
- CN202510136312.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to accurately and quickly detect infrared thermal image defects of lightning arresters, especially in scenarios where the thermal image quality is poor and background interference is high.
Improved MobileNetV3-SSD model, and by introducing a multi-scale feature fusion module and an improved lightweight convolution attention module CBAM, the feature extraction capability of the network is enhanced and the recognition accuracy of infrared thermal image defects of the lightning arrester is improved.
It realizes rapid diagnosis of infrared thermal image failure of lightning arrester, improves the speed and accuracy of image recognition and feature extraction, and provides better guidance for power transmission line maintenance.
Smart Images

Figure CN120070363A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing and power equipment defect detection, and particularly to a method for detecting defects in infrared thermal images of lightning arresters based on improved MobileNetV3-SSD. Background Art
[0002] As an overvoltage protection device in the power system, the function of the lightning arrester is to limit the amplitude of the elevated grid voltage to a certain level, thereby protecting power equipment from overvoltage hazards. To ensure the normal operation of power equipment and avoid risks and losses, it is necessary to regularly inspect the lightning arrester. How to accurately and quickly detect the defects of the lightning arrester is of great significance to the normal operation of the power grid. The infrared thermal imaging method is currently the recommended method for inspecting lightning arresters in the power grid and is suitable for identifying defects that cause local heating. By observing this heating state, the defect state of the lightning arrester can be identified. Therefore, a method for detecting defects in infrared thermal images of lightning arresters is needed. Since a large number of thermal imaging pictures will be generated during the inspection process by drones and the like, batch defect recognition and processing of thermal images are important means to improve the detection efficiency, and the accuracy and efficiency of image processing have become a difficulty and a research hotspot in this technology. Currently, deep learning algorithms are widely used in the defect recognition of lightning arresters, with good generalization ability and the ability to extract features from complex backgrounds, and have been widely applied in the power system.
[0003] There are various artificial intelligence object detection algorithms applied to the image processing of power lightning arresters. Among them, the advantages of the MobileNetV3-SSD model are relatively more suitable, meeting the requirements of applicability, real-time performance, and accuracy for lightning arrester defect detection. However, it is also necessary to improve this model to adapt to the scene requirements where the defect features of infrared thermal images of lightning arresters are not easily distinguishable.
[0004] Therefore, the present invention provides a method for detecting defects in infrared thermal images of lightning arresters based on improved MobileNetV3-SSD, which realizes the detection of weak feature defects such as pollution of lightning arresters in transmission lines, improves the accuracy and efficiency of image processing, and provides new ideas and references for the intelligent inspection of transmission lines.
[0005] Patent CN 118608964A discloses a method for detecting transmission line lightning arresters based on infrared images and rotated target boxes. This method improves the original YOLOv5 feature extraction network, adds an AngleConv module to the feature extraction network, and adopts an improved CBAM attention module with grouped spatial attention and increased input-output direct connection. To a certain extent, it can realize the recognition of transmission line lightning arresters in infrared images and solve the technical problems of poor infrared image quality and many background interferences in horizontal target boxes. This method's model pays more attention to the pre-processing of images.
[0006] Patent CN 115953408A discloses a method for detecting surface defects of lightning arresters based on YOLOv7. This method uses a surface defect dataset of lightning arresters as the training dataset, constructs a defect detection network based on YOLOv7, trains the defect detection network and generates a defect detection model for detecting surface defects of lightning arresters. To a certain extent, it can improve the recognition accuracy of blurred and occluded lightning arrester images and reduce the missed detection rate of target objects at the image edge. This method model pays more attention to the recognition accuracy of defect detection. Summary of the Invention
[0007] The present invention proposes a method for detecting defects in infrared thermal images of lightning arresters based on improved MobileNetV3-SSD (SSD, Single Shot MultiBox Detector). As a lightweight neural network with both speed and accuracy, MobileNetV3-SSD is suitable for deployment on embedded and mobile devices. It can achieve rapid diagnosis of faults in the thermal images of lightning arresters, improve the speed and accuracy of operations such as image recognition and feature extraction, and provide guidance for subsequent transmission line maintenance. To achieve the above invention objectives, the technical solution adopted by the present invention is as follows:
[0008] The present invention provides a method for detecting defects in infrared thermal images of lightning arresters based on improved MobileNetV3-SSD, including the following steps:
[0009] S1. Obtain the initial infrared thermal image of the lightning arrester;
[0010] S2. Supplement the dataset of infrared thermal images of lightning arresters through data augmentation techniques;
[0011] S3. Annotate the infrared thermal images of lightning arresters to form a dataset and divide it into a test set and a training set;
[0012] S4. Improve and model the MobileNetV3-SSD network model;
[0013] S5. Train the improved network model with the dataset and perform performance evaluation to obtain the defect detection results.
[0014] Specifically, in step S1, the infrared thermal image of the lightning arrester is collected by using an infrared thermal imager and combined with various online dataset public resources to form the initial dataset required by the present invention.
[0015] Specifically, in step S2, offline enhancement techniques such as Gaussian blur, randomly adding noise, and color jitter are used to expand the dataset; the Mosaic online enhancement technique is used to generalize and expand the sample data. This data enhancement technique can stitch four different images in the training set together to form a new image. The lightning arrester thermal image dataset is supplemented through these two data enhancement techniques, namely, the combination of online enhancement and offline enhancement.
[0016] Specifically, in step S3, the annotation tool Labelimg is used to annotate the thermal images of defective lightning arresters. After the annotation of the image data is completed, the required dataset is formed and divided into a test set and a training set according to a certain proportion.
[0017] Specifically, in step S4, in the MobileNetV3-SSD backbone network, information at different levels of the network is first fused through a multi-scale feature fusion module. The large-scale feature fusion module and the small-scale feature fusion module are used to establish a connection between the feature layers of convolutional layer 12 and convolutional layer 14.
[0018] By introducing the multi-scale feature fusion module, a feature layer of 19×19×112 is obtained using the large-scale feature fusion module; a feature layer of 10×10×160 is obtained using the small-scale feature fusion module.
[0019] Specifically, in step S4, the improved lightweight convolutional attention module CBAM (Convolutional Block Attention Module) is cited to replace the SE attention module in the original MobileNetV3-SSD network model.
[0020] In the improved CBAM attention module, in the channel attention module, k 1×1 convolutions are used to replace the two fully connected layers originally used for feature mapping to perform channel feature aggregation. Due to the parameter sharing nature of the convolution operation, introducing 1×1 convolution reduces the number of parameters to be calculated to a constant level. For the input feature map F, global average pooling and global max pooling are first performed respectively, and then two 1×1 convolution operations are performed respectively and two different channel attention feature maps are generated through the Sigmoid activation function. The two feature maps are added to form the attention weight, and finally it is multiplied pixel by pixel with the input feature map to obtain the feature map F' adjusted by the channel attention module. The specific process is expressed as:
[0021] where F represents the input feature map; F' represents the feature map adjusted by the channel attention module; δ is the Sigmoid activation function, are k 1×1 convolutions; P avg is global averaging; Pmax is global max pooling.
[0022] Specifically, in step S4, the feature map F' adjusted by the channel attention module is continuously input into the spatial attention module to enhance the feature extraction ability of the network, and the output feature map F'' is obtained.
[0023] The enhancement of the network's feature extraction ability is achieved by using deformable convolution to replace ordinary convolution in the spatial attention module. First, the feature map F' is extracted using a traditional convolution kernel. Second, an adaptively learned horizontal and vertical position offset is added to each convolution kernel unit in the traditional convolution kernel, and the direction of the convolution kernel is adjusted by learning the offset. Finally, the convolution kernel can perform adaptive sampling according to the shape of the lightning arrester, improving the ability to extract infrared defect features of the lightning arrester, and finally outputting the feature map F''.
[0024] Specifically, in step S5, the improved network model is trained and its performance is evaluated to obtain the defect detection result, and a comparative experiment is conducted on the lightning arrester defect detection result with other object detection models.
[0025] The comparative experiment compares the performance of the improved MobileNetV3-SSD of the present invention with other mainstream object detection models in terms of precision, recall, F1 score (the harmonic mean of precision and recall), and the average precision when the intersection over union is 0.5 and the frame rate processing speed representing the model per second.
[0026] Compared with the prior art, the beneficial effects of the present invention are:
[0027] A method for detecting defects in infrared thermal images of lightning arresters based on improved MobileNetV3-SSD provided by the present invention: By collecting a large number of infrared thermal images of lightning arresters, the dataset of infrared thermal images of lightning arresters is expanded through a data augmentation technology that combines online enhancement and offline enhancement, and the image data is annotated to form a dataset and divided into a test set and a training set. Then, the MobileNetV3-SSD network model is used to detect the infrared thermal images of lightning arresters. It is proposed to introduce a multi-scale feature fusion module into the MobileNetV3-SSD backbone network to fuse information at different levels of the network, and then input the feature map after feature fusion into the improved lightweight convolutional attention module CBAM to enhance the feature extraction ability of the network.
[0028] Finally, the extracted features are used for defect detection of infrared thermal images of lightning arresters. The experimental results show that the present invention realizes the rapid diagnosis of faults in infrared thermal images of lightning arresters and can provide a reference for subsequent transmission line maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only the preferred embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0030] Figure 1 It is a schematic diagram of the defect detection process of the lightning arrester thermal image of the present invention.
[0031] Figure 2 It is a schematic diagram of the structure of the large-scale feature fusion module introduced in the present invention. Among them, 2-1: Convolutional layer 14; 2-2: Convolutional layer 12; 2-3: 19×19×112 feature layer.
[0032] Figure 3 It is a schematic diagram of the structure of the small-scale feature fusion module introduced in the present invention. Among them, 3-1: Convolutional layer 14; 3-2: Convolutional layer 12; 3-3: 10×10×160 feature layer.
[0033] Figure 4 It is a structural diagram of the improved lightweight convolutional attention module CBAM introduced in the present invention. Among them, 4-1: Input feature map F; 4-2: Channel attention module; 4-3: Feature map F' adjusted by the channel attention module; 4-4: Spatial attention module; 4-5: Output feature map F”.
[0034] Figure 5 It is a schematic diagram of the structure of the improved MobileNetV3-SSD network model and module of the present invention. Detailed implementation manners
[0035] Combined with Figure 1 , the schematic diagram of the lightning arrester infrared thermal image defect detection method provided by the present invention. The method includes the following steps:
[0036] Step S1, obtain the initial lightning arrester infrared thermal image.
[0037] As an implementable manner, the lightning arrester infrared thermal image is collected by using an infrared thermal imager, and the initial dataset required by the present invention is formed by combining various online dataset public resources.
[0038] Step S2, based on Step S1, expand the lightning arrester infrared thermal image dataset through a data augmentation technology that combines online augmentation and offline augmentation.
[0039] The dataset augmentation is to flip, rotate, adjust brightness, adjust grayscale, randomly crop, and perform mosaics on the images in the initial dataset to increase sample diversity and sample size. The purpose is to improve the generalization ability of the detection network and enhance the robustness of the network model.
[0040] Step S3: Annotate the augmented infrared thermal image atlas of arresters to form a dataset and divide it into a test set and a training set. Both the test set and the training set contain pictures of several typical defects of arresters.
[0041] The specific annotation and classification are as follows: For the atlas of thermal images of arresters with normal and defective arresters, use the annotation tool Labelimg to annotate the thermal imaging.PNG files of defective arresters to obtain.xml files, and then convert them into.txt files adapted to the MobileNetV3-SSD network. After annotating the image data, the required dataset is formed and divided into a test set and a training set according to a ratio, generally 1:10 to 1:20.
[0042] Step S4: Improve and model the MobileNetV3-SSD network model.
[0043] The improvement and modeling are as follows: Select the MobileNetV3-SSD network model as the benchmark model, introduce a multi-scale feature fusion module to fully extract the semantic information of features and reduce the loss of correlation between features; introduce an improved lightweight convolutional attention module CBAM, and input the feature map after feature fusion into the improved CBAM attention module to enhance the model's feature extraction ability for small targets, which helps the model better locate and identify the defective targets of arresters.
[0044] In this example, a multi-scale feature fusion module (a large-scale feature fusion module and a small-scale feature fusion module) is introduced into the backbone feature extraction layer of the MobileNetV3-SSD network model, and the two feature fusion modules establish a connection between the feature layers of convolutional layer 12 and convolutional layer 14.
[0045] See Figure 2 and Figure 3The large-scale feature fusion module first performs an upsampling operation with a factor of 2 on the convolutional layer 14 to make the size of its feature map the same as that of the convolutional layer 12, then performs channel concatenation with the convolutional layer 12 for feature fusion, and uses a 3×3 convolution to eliminate the aliasing effect of the fused features. Then, a 1×1 convolution is used for channel dimensionality reduction operation to obtain a feature layer of 19×19×112; the small-scale feature fusion module performs a max-pooling operation with a stride of 2 and a size of 2×2 on the convolutional layer 12 to reduce the dimensionality of the feature map, obtaining a feature map with a size of 10×10, which is the same size as the convolutional layer 14, and performs channel concatenation with the convolutional layer 14 to obtain a feature layer of 10×10×160.
[0046] In this example, the improved CBAM attention module is used to replace the SE attention module in the original MobileNetV3-SSD network model, which improves the ability to extract detailed features of the arrester, enabling it to better detect subtle defects such as small heating-type salt spots and corona ring discharges at the edges of the upper and lower iron caps of the arrester.
[0047] See Figure 4 In the improved CBAM attention module, in the channel attention module, k 1×1 convolutions are used to replace the two fully connected layers originally used for feature mapping for channel feature aggregation. Due to the parameter sharing property of the convolution operation, introducing 1×1 convolution reduces the number of parameters to be calculated to a constant level. For the input feature map F, global average pooling and global max pooling are first performed respectively, and then two 1×1 convolution operations are performed respectively and passed through the Sigmoid activation function to generate two different channel attention feature maps. The two feature maps are added to form the attention weight, and finally it is multiplied pixel by pixel with the input feature map to obtain the feature map F' adjusted by the channel attention module. The specific process is expressed as:
[0048] where F represents the input feature map; F' represents the feature map adjusted by the channel attention module; δ is the Sigmoid activation function, The activation function performs a non-linear combination of the channel features to enhance the non-linear expression ability of the model; are k 1×1 convolutions; P avg is global average pooling, which calculates the average value of the feature map on each channel's feature map as the output; P max is global max pooling, which is used to find the maximum value on each channel's feature map as the output.
[0049] In this example, the feature map F' adjusted by the channel attention module is continuously input into the spatial attention module to enhance the feature extraction ability of the network, obtaining the output feature map F”.
[0050] SeeFigure 4 and Figure 5 The feature extraction ability of the enhanced network is that in the spatial attention module, deformable convolution is used to replace the ordinary 7×7 convolution. First, the feature map F' is extracted by using a traditional convolution kernel. Secondly, an adaptively learned horizontal and vertical position offset is added to each convolution kernel unit in the traditional convolution kernel, and the direction of the convolution kernel is adjusted by learning the offset. Finally, the convolution kernel can perform adaptive sampling according to the shape of the lightning arrester, improving the ability to extract the infrared defect features of the lightning arrester, and finally outputting the feature map F".
[0051] Step S5: Train and evaluate the performance of the improved network model to obtain the defect detection result;
[0052] Based on the above steps, this example conducts experimental verification. The improved MobileNetV3-SSD of the present invention is compared with other mainstream object detection models in terms of precision P (Precision), recall R (Recall), F1 score (the harmonic mean of precision and recall), and the mean average precision (mAP@0.5) when the intersection over union IoU (Intersection over Union) is 0.5, as well as the frame processing speed representing the number of frames per second of the model.
[0053] The calculation formulas for precision P and recall R are as follows:
[0054] TP (True Positive) represents the correctly predicted positive sample, that is, the sample is detected and the prediction result is correct; FP (False Positive) represents the incorrectly predicted positive sample, that is, the sample is detected but the prediction result is incorrect, also known as false detection; FN (False Negative) represents the incorrectly predicted negative sample, that is, the sample is not detected, and there should be a sample at this position in fact, also known as missed detection.
[0055] The F1 score is the harmonic mean of precision and recall, and the F1 score is used to balance these two evaluation indicators.
[0056] IoU is generally used to compare the overlapping degree of the predicted box and the true box. In some object detection tasks, a threshold is preset for it, and the calculation formula is shown as follows:
[0057] The mean Average Precision (mAP@0.5) is usually used to preset the threshold. The prediction results are sorted in descending order according to the IoU value. After changing the threshold size, repeating this process can obtain the P-R curve, and thus the Average Precision (AP) value can be obtained. The calculation formula is shown as follows:
[0058] In some alternative embodiments, comparative experiments are conducted using other mainstream object detection models, the MobileNetV3-SSD network model, and the improved MobileNetV3-SSD network model provided by the present invention. The experimental results are shown in Table 1 as follows: Table 1 Comparison table of performance results of various object detection models Model P / % R / % F1 Score / % mAP@0.5 / % <![CDATA[Detection speed / frame·s -1 <!-- 5 -->]]> VGG-SSD 71.80 80.68 75.98 78.97 27.39 MobileNetV2-SSD 80.48 87.63 83.90 82.74 34.73 YOLOv3 67.69 88.81 76.83 80.18 41.83 FasterR-CNN 85.27 96.81 90.67 92.26 8.78 MobileNetV3-SSD 81.11 93.44 86.84 84.54 35.56 Improved MobileNetV3-SSD 88.92 95.03 91.87 89.37 29.53
[0059] As can be seen from Table 1, compared with the two-stage detection model Faster R-CNN, the improved MobileNetV3-SSD model has a significant improvement in detection speed and is more suitable for the real-time detection requirements of lightning arresters. Compared with other one-stage object detection models, although YOLOv3 has the fastest detection speed, its F1 score and mAP@0.5 value have decreased by 15.04% and 9.19% respectively compared with the improved MobileNetV3-SSD model. Compared with the traditional SSD model, the improved model has a faster detection speed, higher precision, and recall rate due to the adoption of a lightweight backbone network, the introduction of a multi-scale feature fusion module, and the improvement of the attention model.
[0060] The preferred embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and the devices and structures not described in detail should be understood to be implemented in a common manner in the art. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes, without departing from the scope of the technical solution of the present invention, which does not affect the essence of the present invention. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A lightning arrester infrared thermal image defect detection method based on improved MobileNetV3-SSD, characterized in that: The following steps are involved: S1. Obtaining the initial arrester thermal image; S2, supplement the arrester thermal image dataset through the data enhancement technology combining online enhancement and offline enhancement; S3, annotate the arrester thermal images to form a data set and divide it into a test set and a training set; S4. Introducing a large-scale feature fusion module, a small-scale feature fusion module and a lightweight convolutional attention module CBAM (Convolutional Block Attention Module) into the MobileNetV3-SSD network model; obtaining the improved MobileNetV3-SSD network model; S5. Use the improved network model to perform defect detection and performance evaluation on the image data set to be detected to obtain detection results.
2. According to claim 1, the infrared thermal image defect detection method of lightning arrester based on improved MobileNetV3-SSD is characterized in that: The large-scale feature fusion module, the small-scale feature fusion module and the lightweight convolutional attention module CBAM are introduced into the MobileNetV3-SSD network model; A multi-scale feature fusion module (a large-scale feature fusion module and a small-scale feature fusion module) is introduced into the backbone feature extraction layer (backbone) network of the MobileNetV3-SSD network model; the lightweight convolutional attention module CBAM replaces the conventional convolution Conv in the SE attention module of the MobileNetV3-SSD network model.
3. The method for detecting defects of lightning arrester infrared thermal images based on improved MobileNetV3-SSD according to claim 1 is characterized in that The arrester thermal image is subjected to data preprocessing, including: The arrester thermal imaging image data set is expanded using data enhancement technology, and the arrester thermal imaging image is annotated using an annotation tool. After the annotated image data is completed, the required data set is formed and divided into a test set and a training set at a ratio of 1 / 10 to 1 / 20.
4. The data enhancement technique according to claim 3, characterized in that The dataset is expanded by using offline enhancement techniques such as Gaussian blur, random noise addition and color jitter, and the sample data is generalized and expanded by using Mosaic online enhancement technology. The arrester thermal image dataset is supplemented by the data enhancement technology combining the online enhancement and the offline enhancement.
5. The method for detecting defects of lightning arrester infrared thermal images based on improved MobileNetV3-SSD according to claim 1 is characterized in that: After training, the performance evaluation of the improved MobileNetV3-SSD network model is tested. The model evaluation indicators include precision P (Precision), recall R (Recall), F1 score (the harmonic mean of precision and recall), average precision (mAP@0.5, meanAverage Precision) when the intersection over Union (IoU) is 0.5, and the frame processing speed of the representative model per second as the performance evaluation indicators of the improved MobileNetV3-SSD model.