High-voltage transmission line insulator defect detection method for unmanned aerial vehicle inspection

By improving the model structure on the YOLOv8 model, using multi-scale fusion and shared parameters detection heads, introducing an EMA attention mechanism, using depth separation convolution, and replacing the loss function with Inner-IoU, the problems of low detection accuracy and insufficient real-time performance of the drone for insulator defect detection are solved, and efficient and accurate insulator defect detection are achieved.

CN120088461APending Publication Date: 2025-06-03HEFEI UNIV
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510251359.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing drones are equipped with YOLOv8 models for insulator defect detection with low detection accuracy, high real-time requirements, but low frame rate, and severe interference in complex environments, making it difficult to meet the normal operation needs of the power system.

Method used

Based on YOLOv8, the model structure is improved, multi-scale fusion and shared parameters are used to introduce an EMA attention mechanism, use depth separation convolution, and replace the loss function to Inner-IoU to improve detection accuracy and real-time.

Benefits of technology

It improves the accuracy and efficiency of insulator defect detection, reduces calculation and storage pressure, makes the model more suitable for fast and efficient detection in power systems, and significantly improves detection accuracy and recall.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088461A_ABST
    Figure CN120088461A_ABST
Patent Text Reader

Abstract

The invention provides a high-voltage transmission line insulator defect detection method for unmanned aerial vehicle inspection. Firstly, by designing a multi-scale fusion and parameter sharing detection head, the feature fusion capability is effectively improved; secondly, an EMA attention mechanism is introduced, so that the model is more focused on information of a key area; the use of the depth separable convolution significantly reduces the parameter quantity and calculation complexity of the model. And finally, using Inner-IoU as a frame loss function, and more effectively improving the detection precision of the small target and enhancing the performance of the model in the aspect of prediction frame regression by focusing on the matching degree in the frame. The improved model can be used for accurately detecting the defects of the insulator, the detection precision is improved, and the operand of the model is effectively reduced through the lightweight design, so that the model provided by the invention is more suitable for being deployed to embedded equipment for real-time detection of the defects of the insulator, and the real-time detection of the defects of the insulator is facilitated. A new method is provided for the field of power transmission network insulator defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent inspection of electric power equipment, and in particular to a method for detecting defects of insulators of high-voltage transmission lines by unmanned aerial vehicle inspection. Background Art

[0002] As an important power equipment, the safety performance and stability of insulators have attracted wide attention with the rapid development of the power industry. Due to various reasons such as long-term operation and environmental factors, insulators may have problems such as self-explosion defects. Traditional power inspections mainly rely on manual work, and inspectors need to check equipment one by one, which is very inefficient and cannot meet current inspection needs.

[0003] At present, drone inspections based on deep learning can greatly improve detection efficiency. However, the insulator defect targets photographed by drones are usually small, and limited by the computing and storage capabilities of drones, the high computing requirements and large storage requirements of complex models are difficult to meet, resulting in drones being unable to detect defective insulators in a timely manner, seriously affecting the normal operation of the power system. Research on lightweight detection methods for insulator defects is of great significance to ensuring the safe and stable operation of the power system.

[0004] At present, target detection models can be divided into two categories: two-stage detection models and one-stage detection models. Two-stage detection models such as the R-CNN (Region-based Convolutional Neural Network) series have high detection accuracy, but are slow and difficult to meet the requirements of real-time detection. In contrast, one-stage models such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector) can simultaneously complete target recognition and positioning, have a faster detection speed, and can achieve real-time detection while maintaining high accuracy.

[0005] YOLOv8 is the target detection algorithm in the YOLO series, with many technical improvements and optimizations, designed to provide more efficient and accurate target detection capabilities. YOLOv8 significantly improves detection accuracy while maintaining high speed by adopting a deeper and more complex network structure and improved training techniques. YOLOv8 uses a variety of data enhancement techniques and training strategies during the training process, such as random cutting, color gradient, etc. These methods effectively expand the data set and increase the diversity of data samples. This diversity is crucial to improving the generalization performance of the model. It enables the model to learn more extensive and complex scene features, so that it can make accurate and reliable predictions when encountering new and unseen data.

[0006] However, there are various technical challenges in the practical application of using the YOLOv8 model on drones for insulator defect detection, and these problems need to be specifically optimized. The main difficulties include:

[0007] 1. Low detection accuracy: The defective area of the insulator occupies a very small pixel area in the aerial images captured by the drone, making the defect features not obvious in the images. Moreover, when the drone takes pictures, the changes in the viewing angle and shooting distance will further affect the clarity of the defect features, increasing the difficulty of detection. The ordinary convolution used by YOLOv8 faces multiple limitations, including insufficient local feature extraction ability, weak feature fusion ability, low computational efficiency, and parameter redundancy problems. These drawbacks lead to a decrease in the detection accuracy of the model when dealing with insulator defects and make it difficult to cope with complex backgrounds and real-time requirements.

[0008] 2. High real-time requirement but low frame rate: To improve the detection accuracy, YOLOv8 adopts a deeper and more complex network structure, resulting in an increase in the model size and a decrease in the inference speed. In the scenario of insulator defect detection, the frame rate is a key indicator to measure the real-time detection ability of the model. Due to the increase in model complexity, the frame rate cannot meet the requirements of real-time detection.

[0009] 3. Interference from complex environments: Insulators are usually installed outdoors, and the background environment is complex and changeable. Interferences such as trees, buildings, and weather changes will all affect the detection. The light changes under different times and weather conditions will also affect the image quality and thus the detection effect. Summary of the Invention

[0010] In the research on current insulator defect detection, the present invention uses too many attention mechanism parameters, resulting in an increased computational burden. Most methods often sacrifice speed while pursuing lightweight, and it is difficult to balance detection efficiency and accuracy simultaneously. Based on YOLOv8, the present invention improves the model structure to enhance real-time performance and accuracy, etc., and thus proposes a method for detecting insulator defects in high-voltage transmission lines during drone patrols.

[0011] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0012] A method for detecting insulator defects in high-voltage transmission lines by UAV inspection. This detection method is based on the constructed YOLOv8-SDEI model. First, by designing a detection head with multi-scale fusion and shared parameters, the feature fusion ability is effectively improved. Second, the EMA attention mechanism is introduced to make the model more focused on the information in the key areas. The use of depthwise separable convolution significantly reduces the number of model parameters and computational complexity. Finally, Inner-IoU is used as the bounding box loss function. By focusing on the matching degree inside the box, the detection accuracy of small targets is more effectively improved, and the performance of the model in predicting box regression is enhanced.

[0013] Furthermore, the YOLOv8-SDEI model improves the model's ability to extract and fuse multi-scale features by introducing a Spatial Reconstruction Sub-module (SRU) and a Channel Reconstruction Sub-module (CRU), enabling each convolutional kernel to dynamically adjust according to the features of the input data. This mechanism can better capture the fine-grained spatial and channel information in the image, thereby enhancing the feature representation ability, improving the accuracy of insulator detection, and reducing the spatial redundancy existing in standard convolutions, as well as reducing the computational amount and the number of parameters of the model.

[0014] Furthermore, the overall structure of the Spatial Reconstruction Sub-module (SRU) is as follows:

[0015] First, in the Separate stage, SRU normalizes the input feature X by subtracting the mean μ and dividing by the standard deviation σ, and measures the spatial pixel variance of each batch and channel with the trainable parameter γ ∈ R C for each batch and each channel, there is a trainable parameter to measure the variance; the normalization correlation weight W γ ∈ R C is obtained by formula (1):

[0016]

[0017] Then, the weight values of the weighted feature map are mapped to the range (0, 1) through the sigmoid function and gated by a threshold, which allows the model to focus more on the key information. Finally, the input feature X is multiplied by W γ and W 1 and W 2 to obtain two weighted features: the feature with a large amount of information and the feature with a small amount of information to achieve non-linear transformation of the features;

[0018] In the Reconstruct stage, the feature with a large amount of information and the feature with a small amount of information Sum them up, generate more informative features, and fully fuse the two different weighted information features to enhance the information flow between them; then the cross-reconstructed features and are connected to obtain a spatially refined feature map X w ; The whole process is calculated as:

[0019]

[0020] where, is element-wise multiplication, ⊕ is element-wise summation, and ∪ is concatenation; After the SRU is applied to the intermediate input feature X, not only the features with large amount of information are separated from the features with small amount of information, but also they are reconstructed, which can enhance the representative features and suppress the redundant features in the spatial dimension.

[0021] Furthermore, the overall structure of the channel reconstruction sub-module (CRU) is as follows:

[0022] For the given spatially refined feature X w ∈R c×h×w , first divide the channels of X w into two parts, namely αC channels and (1-α)C channels; then further use 1×1 convolution operation to compress the channels of the feature map and improve the calculation efficiency; then introduce a squeezing ratio r to control the feature channels and reduce the calculation cost of the CRU; finally divide the spatially defined feature X w into the upper part X up and the lower part X low , and process them in different transformation stages. When X up is input into the upper transformation stage, efficient convolution operations (GWC and PWC) are used to replace the standard k×k convolution;

[0023] Perform k×k GWC and 1×1 PWC operations on the same X up , and finally integrate the output and obtain the feature map Y 1 ; When X low is input into the lower transformation stage, PWC is used to generate the feature map Y 2 ; Collect the global spatial information S m ∈R c×1×1 with channel statistics through global average pooling (pooling); 1 Stack S 2 and S 2 ∈R c×1×1 together, and generate the feature importance vector β through SoftMax, β 1 and the lower feature Y 2 are merged in a channel-wise manner to obtain the channel-refined feature Y;

[0024] In the SCConv module, all parameters are concentrated in the conversion stage; the parameters of the standard convolution are Y = M k X can be calculated as:

[0025] P s = k × k × C 1 × C 2 = k 2 C 1 C 2 (3)

[0026] The parameters of the SCConv module can be calculated as:

[0027]

[0028] where k is the kernel size of the convolution; C 1 and C 2 are the numbers of input and output feature channels; α represents the splitting ratio; γ represents the squeezing ratio; g is the group size of the GWC operation.

[0029] Furthermore, the YOLOv8-SDEI model uses depthwise separable convolution (DSC). By independently processing the insulator features of each channel in the depthwise convolution and then integrating all the insulator features through the pointwise convolution, the model can extract fine-grained insulator feature information from each channel, thus more accurately locating the defects of the insulators;

[0030] The ratio of the number of parameters between the depthwise separable convolution and the standard convolution is:

[0031]

[0032] where C K represents the size of the convolution kernel, M represents the number of channels of the input feature map, and N represents the number of channels of the output feature map.

[0033] Furthermore, the YOLOv8-SDEI model introduces the attention mechanism EMA to improve the model's perception ability. EMA constructs local cross-channel interactions in each parallel sub-network without reducing the number of channels; and fuses the output feature maps of the two parallel sub-networks through the cross-space learning method.

[0034] Furthermore, the YOLOv8-SDEI model uses an improved Inner-IoU loss function, specifically:

[0035] The GT box (ground truth box) and the anchor box are respectively denoted as b gtand b, the center point inside the GT box is denoted by (x c , y c ) represents the center point inside the anchor box; the length and width of the GT box are denoted as w gt and h gt , respectively, and the length and width of the anchor are denoted as w and h; denotes the left, right, top, and bottom boundary positions of the actual bounding box; b l , b r , b t , b b denote the left, right, top, and bottom boundary positions of the auxiliary bounding box; inter represents the intersection area between the actual bounding box and the auxiliary bounding box; union represents the union area between the actual bounding box and the auxiliary bounding box; IoU inner denotes the value of Inner-IoU, that is, the intersection area divided by the union area; the value range of Inner-IoU is similar to that of IoU, both within [0, 1]; the following formula describes the calculation process of the Inner-IoU loss:

[0036]

[0037] union = (w gt * h gt ) * (ratio) 2 + (w * h) * (ratio) 2 - inter(11)

[0038]

[0039] where the general value range of the scaling factor ratio is within [0.5, 1.5];

[0040] When ratio is less than 1, the size of the auxiliary bounding box is smaller than that of the actual bounding box, making the effective regression range of high-IoU samples smaller, but the absolute value of the gradient is greater than the gradient obtained by the IoU loss, which can accelerate the convergence of high-IoU samples;

[0041] When ratio is greater than 1, the larger-sized auxiliary bounding box expands the effective regression range and enhances the regression effect of low-IoU samples. This characteristic enables the Inner-IoU loss to more effectively accelerate the convergence of the regression results under different scaling factors, and has a positive impact on both high-IoU and low-IoU samples. Inner-IoU adjusts the scaling factor for various datasets and different object detectors, enhancing the generalization of Inner-IoU and improving the detection performance of the model. The model can capture the key information of insulators faster.

[0042] To address the deficiencies in current insulator defect detection, based on YOLOv8, this invention replaces ordinary convolutions in the network structure with SCConv, introduces depthwise separable convolutions, uses the EMA attention mechanism, and replaces the loss function with inner-IOU to identify and locate insulators in transmission lines, solving the problems of insufficient object detection accuracy and speed when encountering complex backgrounds in current methods. Compared with the prior art, the beneficial effects of this invention are as follows:

[0043] (1) On the baseline model YOLOv8n, this invention replaces standard convolutions with SCConv, enhancing the model's ability to extract and fuse multi-scale features, thereby improving the accuracy of insulator detection. The redesigned detection head reduces the model's computational complexity and parameter count, making the model lightweight. The application of depthwise separable convolutions further reduces the burden of model parameters and computational complexity. The introduction of the EMA attention mechanism enables the model to focus more on key information, while the use of the Inner-IoU loss function enhances the regression effect of the model's prediction boxes.

[0044] (2) The improved YOLOv8 model proposed in this invention effectively reduces the computational and storage pressure while ensuring high accuracy, making it more suitable for rapid and efficient detection of insulators in the power system. Therefore, the lightweight and efficient insulator defect detection method proposed in this invention is of great significance for the safe and stable operation of the power system.

[0045] (3) Experiments were conducted on the insulator defect detection dataset for high-voltage transmission lines. The improved model reduced the computational complexity by 2.2G while achieving an mAP of 90.6%, a 1.6% increase compared to the original model. The model size was only 4.5MB, and the accuracy and recall rates were 83.1% and 94% respectively, a 1.6% and 0.8% increase compared to the original model. The improved model can accurately detect insulator defects. This method not only improves the detection accuracy but also effectively reduces the model's computational workload through lightweight design. Therefore, the model proposed in this invention is more suitable for deployment on embedded devices for real-time insulator defect detection, providing a new method for the field of insulator defect detection in transmission networks. Description of the Drawings

[0046] Figure 1 To improve the YOLOv8-SDEI network structure.

[0047] Figure 2 For the overall SCConv module structure.

[0048] Figure 3 For the SRU module structure.

[0049] Figure 4 For the CRU module structure.

[0050] Figure 5 It is a lightweight shared parameter detection head structure.

[0051] Figure 6 It is a depthwise separable convolution structure.

[0052] Figure 7 It is an EMA attention mechanism structure.

[0053] Figure 8 It is an Inner-IoU structure.

[0054] Figure 9 It is a detection result graph of YOLOv8n and the improved model of the present invention. Detailed implementation manners

[0055] In order to enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0056] Embodiment 1 Improved model structure

[0057] The present invention improves the model structure based on YOLOv8 to improve real-time performance and accuracy, etc., and improve the accuracy and efficiency of insulator target detection. The network structure of the improved YOLOv8-SDEI model is as Figure 1 shown.

[0058] The network structure of YOLOv8-SDEI mainly consists of the following three major parts: Backbone (backbone network), Neck (neck structure), and Head (head). YOLOv8-SDEI uses DWConv to replace the standard convolution in the Backbone to optimize the computational efficiency, enabling the model to still achieve high inference speed on devices with limited computing resources (such as drones, etc.), meeting the requirements of real-time detection. And an EMA attention mechanism is introduced after the SPPF module to enhance the feature focusing ability. By guiding the model to focus on the key areas of insulator defects, the detection accuracy is effectively improved. The EMA mechanism can enhance the model's ability to capture small defect features in complex environments, reduce the interference of complex backgrounds, and enable the model to maintain good detection performance in various environments. Finally, a multi-scale fusion detection head is redesigned in the Head to improve the multi-scale object detection ability, improve the detection accuracy of small targets, and enhance the overall performance of insulator defect detection.

[0059] 1 Multi-scale Fusion Shared Parameter Detection Head

[0060] Since the operations of standard convolution focus on feature extraction in local regions, the ability of the YOLOv8 model to obtain global context information is weakened. As the convolution is stacked, the number of convolution kernels increases and the network depth becomes deeper. The parameters and computational complexity generated by standard convolution will be extremely large, leading to slow model training and inference, and prone to overfitting and other situations. The feature fusion ability between different layers of standard convolution is weak, resulting in a decline in the performance of the YOLOv8 model in processing insulator sample information.

[0061] To solve the above problems, YOLOv8-SDEI uses a new convolution method: Spatial and Channel Reconstruction Convolution (SCConv) (reference: Li J, Wen Y, He L. Scconv: spatial and channel reconstruction convolution for feature redundancy[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023: 6153-6162.). The overall structure is as Figure 2 shown. By introducing a Spatial Reconstruction Sub-module (SRU) and a Channel Reconstruction Sub-module (CRU), the model's ability to extract and fuse multi-scale features is improved, enabling each convolution kernel to dynamically adjust according to the features of the input data. This mechanism can better capture fine-grained spatial and channel information in the image, thereby enhancing the feature representation ability, improving the accuracy of insulator detection, and reducing the spatial redundancy existing in standard convolution, as well as reducing the computational complexity and number of parameters of the model.

[0062] The overall structure of the Spatial Reconstruction Sub-module (SRU) is as Figure 3 shown. First, the SRU normalizes the input feature X by subtracting the mean μ and dividing by the standard deviation σ during the Separate stage, and uses the trainable parameter γ ∈ R in the Group Normalization (GN) layer to C measure the spatial pixel variance for each batch and channel, which means that for each batch and each channel, there is a trainable parameter to measure the variance. The normalization correlation weight W γ ∈ R C is obtained by formula (1):

[0063]

[0064] Then, W is passed through the sigmoid function γThe weight values of the weighted feature map are mapped into the range (0, 1), and then gated by a threshold, which allows the model to focus more on key information. Finally, the input feature X is multiplied by W 1 and W 2 to obtain two weighted features: the feature with a large amount of information and the feature with a small amount of information to achieve non-linear transformation of the features.

[0065] In the reconstruction stage (Reconstruct), the feature with a large amount of information is summed with the feature with a small amount of information to generate a feature with even more information, and the two different weighted information features are fully fused to enhance the information flow between them. Then the cross-reconstructed features and are concatenated to obtain the spatially refined feature map X w . The whole process is calculated as:

[0066]

[0067] where is element-wise multiplication, ⊕ is element-wise summation, and ∪ is concatenation. After the SRU is applied to the intermediate input feature X, not only are the feature with a large amount of information and the feature with a small amount of information separated, but they are also reconstructed, which can enhance the representative features and suppress the redundant features in the spatial dimension.

[0068] The overall structure of the channel reconstruction sub-module (CRU) is as shown in Figure 4 . For the given spatially refined feature X w ∈R c ×h×w , first, the channels of X w are divided into two parts, namely αC channels and (1 - α)C channels. Then, 1×1 convolution operations are further used to compress the channels of the feature map to improve the computational efficiency. Then, a squeezing ratio r is introduced to control the feature channels and reduce the computational cost of the CRU. Finally, the spatially defined feature X w is divided into the upper part X up and the lower part X low , and they are processed in different transformation stages. When X up is input into the upper transformation stage, efficient convolution operations (GWC and PWC) are used to replace the standard k×k convolution to extract high-level representative information and reduce the computational cost.

[0069] Perform k×k GWC and 1×1 PWC operations on the same X up , and finally integrate the outputs to obtain the feature map Y 1 . When X lowIt is input into a lower conversion stage, and PWC is used to generate the feature map Y 2 . Global spatial information S with channel statistics is collected through global average pooling m ∈R c×1×1 . Stack S 1 and S 2 together, and generate the feature importance vector β through SoftMax 1 ,β 2 ∈R c×1×1 . Finally, the upper feature Y 1 and the lower feature Y 2 are merged in a channel-wise manner to obtain the channel-refined feature Y

[0070] In the SCConv module, all parameters are concentrated in the conversion stage. The parameters of the standard convolution Y = M k X can be calculated as:

[0071] P s = k × k × C 1 × C 2 = k 2 C 1 C 2 (3)

[0072] The parameters of the SCConv module can be calculated as:

[0073]

[0074] where k is the kernel size of the convolution; C 1 and C 2 are the numbers of input and output feature channels; α represents the splitting ratio; γ represents the squeezing ratio; g is the group size of the GWC operation. In the experiment, the general parameter set is α = 1 / 2, g = 3, k = 3, C 1 = C 2 = C. When P s / P sc ≈ 5, the number of parameters can be reduced by 5 times. In YOLOv8, the computational cost of Detect accounts for a relatively large proportion. Replace the standard convolution in the head of YOLOv8 with the SCConv convolution, as Figure 5 shown

[0075] YOLOv8-SDEI combines the convolutional layers in the regression and classification branches of Detect. In the initial convolutional layers of the detection head, the regression and classification branches share the same set of convolutional kernels to extract features and only make different predictions at the end. This can reduce unnecessary redundant convolutional operations, make feature extraction more efficient, form a lightweight detection head, and achieve the purpose of parameter sharing. The two branches of the lightweight detection head do not need to learn parameters separately, which can greatly reduce the number of parameters in YOLOv8. Parameter sharing enables the model to focus more on extracting the basic features of insulator defects, improves the generalization ability of the model, and can better adapt to new data. Moreover, sharing parameters does not require a large amount of dataset training and reduces the risk of overfitting in the case of insufficient datasets. By sharing the same set of weights, the feature extraction operations adopted by insulator feature maps of different scales can be kept consistent, enabling the lightweight detection head to achieve the effect of sample balance.

[0076] 2 Depthwise Separable Convolution

[0077] To further lightweight the insulator detection model, YOLOv8-SDEI uses Depthwise Separable Convolution (DSC) to replace the standard convolution in Neck. Depthwise Separable Convolution is a high-speed and effective convolutional neural network structure that divides the traditional convolution operation into Depthwise Convolution and Pointwise Convolution.

[0078] Depthwise Convolution: In this process, each input channel is convolved with the corresponding convolutional kernel. Each convolutional kernel only acts on its corresponding input channel and generates a separate output channel, reducing the number of model parameters and computational amount. Depthwise Convolution performs separate convolution processing on each channel of the input without fusing channels.

[0079] Pointwise Convolution: After Depthwise Convolution, a 1x1 convolution is performed, that is, convolution operations are performed on each pixel. The advantage of this process is to fuse the output of the previous step and enhance the expression ability of the network. Pointwise Convolution can be regarded as a linear combination of all channels at each position.

[0080] Depthwise Separable Convolution independently processes the insulator features of each channel in Depthwise Convolution and then integrates all insulator features through Pointwise Convolution, enabling the model to extract fine-grained insulator feature information from each channel, thereby more accurately locating the defects of insulators.

[0081] The ratio of the number of parameters between Depthwise Separable Convolution and standard convolution is:

[0082]

[0083] Among them, C K represents the size of the convolutional kernel, M represents the number of channels of the input feature map, and N represents the number of channels of the output feature map. The size of the convolutional kernel is generally 3, so the number of parameters of depthwise separable convolution is about 9 times less than that of standard convolution. The insulator detection task requires high real-time performance, especially in real-time inspection devices such as drones. The system needs to quickly process aerial images and detect insulator defects. Depthwise separable convolution can reduce the inference time, enabling the model to process more frames per second, improving the detection speed, and meeting the need for real-time detection of insulator defects.

[0084] 3EMA attention mechanism

[0085] In drone inspection, in the face of complex and changing background environments (such as mountains and villages), the detection accuracy of drones for insulators will be affected. To help drones better focus on the key information of insulators, it is necessary to introduce an attention mechanism to improve the model's perception ability.

[0086] The attention mechanism can automatically adjust the model's attention according to the situation of the insulator, enabling the model to focus on the important information in the insulator, improving the accuracy and efficiency of the model. Compared with traditional neural networks, the attention mechanism can flexibly adjust the weights and overcome the limitations of traditional methods. Its complexity is relatively low and it can be flexibly embedded in the model. In traditional attention mechanisms, the SE attention mechanism mainly focuses on channel attention, so it ignores some information in space to a certain extent, which will affect the overall analysis of insulator information by the model.

[0087] Although the DA attention mechanism uses attention mechanisms in both the spatial and channel dimensions, the computational cost and number of parameters of DA are large, requiring more storage space and computing resources, which is not conducive to lightweighting the model. The CBAM attention mechanism can improve model performance, but it needs to align feature maps of different scales, depends on specific datasets, and has poor generalization ability. The CA attention mechanism embeds precise location information into channels and captures long-range interactions spatially, improving performance. Two 1D global average pooling operations are designed to encode global information along two spatial dimensions and capture long-range interactions spatially along different dimension directions. However, CA ignores the importance of interactions between entire spatial positions, and the limited receptive field of 1x1 convolutions restricts cross-channel interactions and the utilization of context information. YOLOv8-SDEI introduces an efficient multi-scale attention mechanism (EMA). EMA proposes a new cross-space learning method by constructing local cross-channel interactions in each parallel subnetwork without channel dimensionality reduction. By fusing the output feature maps of two parallel subnetworks through the cross-space learning method, EMA can establish short-term and long-term dependencies. EMA considers a general method of reshaping part of the channel dimension into the batch dimension to avoid a certain form of dimensionality reduction through standard convolutions. Compared with the previously mentioned attention mechanisms CBAM, SE, DA, and CA, EMA requires fewer parameters and is more efficient in performance. Therefore, the EMA attention mechanism is introduced in YOLOv8-SDEI. The EMA structure is as Figure 7 shown.

[0088] 4Inner-IoU

[0089] The IoU loss function used in YOLOv8 is particularly sensitive to small localization errors of bounding boxes. If the edge positions of the predicted box and the ground truth box are misaligned, even if the distance between them is extremely small, it will result in a low IoU value. When calculating IoU, bounding boxes of different sizes may lead to significant differences. If the area of the box is relatively large, a small error has a relatively small impact on IoU, but when the area of the box is relatively small, such a small error will significantly reduce the IoU value. For a part of overlapping boxes, the change in the IoU value is not sufficient to accurately describe the specific degree of overlap. Especially when the overlap ratio of the boxes is low, if a small box is completely contained within a large box, it will make the IoU value very low and cannot accurately reflect the situation where the small box is completely predicted and covered.

[0090] The Inner-IoU used in the improvement of YOLOv8-SDEI is as Figure 8As shown, Inner-IoU pays more attention to the inclusion relationship between the predicted box and the ground truth box, is more accurate for tiny localization errors, and performs better when dealing with boxes of different sizes. Even if there are significant differences in box sizes, as long as there is an inclusion relationship, the Inner-IoU value remains accurate, avoiding the situation where IoU ignores small boxes. Therefore, Inner-IoU can effectively reflect the inclusion relationship between boxes. For example, in the case where a small box is completely contained within a large box, Inner-IoU can give a high score, accurately reflecting the good performance of the detection result.

[0091] The GT box (ground truth box) and the anchor box are denoted as b gt and b, and the center point inside the GT box is denoted by . (x c , y c ) represents the center point inside the anchor box. The length and width of the GT box are denoted as w gt and h gt respectively, and the length and width of the anchor are denoted as w and h. denotes the left, right, top, and bottom boundary positions of the actual bounding box; b l , b r , b t , b b denote the left, right, top, and bottom boundary positions of the auxiliary bounding box; inter represents the intersection area between the actual bounding box and the auxiliary bounding box; union represents the union area between the actual bounding box and the auxiliary bounding box; IoU inner denotes the value of Inner-IoU, that is, the intersection area divided by the union area. The value range of Inner-IoU is similar to that of IoU, both within [0, 1]. The following formula describes the calculation process of the Inner-IoU loss:

[0092]

[0093]

[0094] union = (w gt * h gt ) * (ratio) 2 + (w * h) * (ratio) 2 - inter(11)

[0095]

[0096] Among them, the general value range of the scale factor ratio is within [0.5, 1.5].

[0097] When the ratio is less than 1, the size of the auxiliary bounding box is smaller than that of the actual bounding box, which reduces the effective regression range for high IoU samples. However, the absolute value of the gradient is greater than that obtained from the IoU loss, which can accelerate the convergence of high IoU samples.

[0098] When the ratio is greater than 1, the larger-sized auxiliary bounding box expands the effective regression range and enhances the regression effect for low IoU samples. This property enables the Inner-IoU loss to more effectively accelerate the convergence of the regression results under different scaling factors, and it has a positive impact on both high IoU and low IoU samples. Inner-IoU adjusts the scaling factor for various datasets and different object detectors, enhancing its generalization ability and improving the detection performance of the model. The model can capture the key information of the insulator faster.

[0099] Experiment of Example 2:

[0100] 1 Dataset

[0101] The dataset used in this invention is a public dataset, the high-voltage transmission line insulator defect detection dataset (Source: Hong X, Wang F, Ma J. Improved yolov7 model for insulator surface defect detection[C] / / 2022 IEEE 5th Advanced Information Management, Communicates, Electronic and Automation Control Conference (IMCEC). IEEE, 2022, 5: 1667 - 1672.), which contains 800 pictures. Data augmentation can prevent the model from overfitting and improve its adaptability to different scenarios by diversifying the training dataset. In object detection tasks, data augmentation is particularly important because objects may appear at different angles, sizes, and backgrounds. After data augmentation of the dataset, 3000 pictures of the insulator dataset are obtained, and the defects of the insulators in the pictures are labeled. The dataset records the positions and categories of insulator defects in a txt document. In the experiments of this invention, it is randomly divided into a training set and a validation set according to 9:1, and the YOLOv8 series and mainstream object detection algorithms such as RT-DETR-L, YOLOv7n, and YOLOv5n are used to train and validate the data.

[0102] 2 Experimental Environment and Configuration

[0103] This invention was experimented on the Microsoft Windows 10 (64-bit) system. The deep learning framework used in the experiment was PyTorch, the editor was PyCharm, the version number of Python was 3.8, and the version number of CUDA was 11.8. The hardware environment of the system was AMD Ryzen 5 1500X Quad-Core, Processor 3.50GHz, the running memory size was 24GB, the GPU was the NVIDIA GeForce RTX 3060 graphics card, and the video memory was 12GB. In the network parameters of the YOLOv8 model, the input image set for training the model was 640*640, the number of epochs was 200 rounds, the batch size was 64, the model optimizer used was SGD, and the learning rate (lr0) was set to 0.01.

[0104] 3 Evaluation Metrics

[0105] The model evaluation metrics adopted by this invention are: Precision, Recall, mean Average Precision (mAP), number of parameters (Param), computational complexity (GFLOPs), and FPS. The detection time consumption is the time required for the model to process a single image when the batch size is 1, measured in milliseconds (ms). FPS represents the number of image frames that the model can process per second, reflecting the real-time performance of the model. The computational complexity measures the number of floating-point operations required for the model to perform forward propagation, reflecting the computational complexity and running speed of the model. Precision represents the proportion of samples that the model predicts as positive classes and are actually positive classes. Precision measures the accuracy of the model in predicting positive classes. The formula for calculating Precision is:

[0106]

[0107] where α TP represents the number of samples that the model correctly predicts as positive classes, and α FP represents the number of samples that the model wrongly predicts as positive classes.

[0108] Recall represents the proportion of the number of positive samples correctly detected by the model to the total number of actual positive samples, and it is one of the important indicators for measuring the detection performance of the model. α FN is a positive sample predicted as a negative class by the model.

[0109]

[0110] mAP is used to measure the performance of the model on different categories and provides an overall evaluation. Taking the recall rate as the abscissa and the precision rate as the ordinate, a relationship curve representing the precision rate and recall rate of each category is obtained. mAP is calculated as the area under the precision-recall curve for each category, and the average value of the curve areas for all categories is taken. For a dataset with N categories, the calculation method is as follows:

[0111]

[0112] The value of mAP reflects the performance of the model. The larger the mAP of the model, the better the object recognition ability.

[0113] 4 Experimental Result Analysis

[0114] To verify the effectiveness of the YOLOv8-SDEI model, it was tested on the insulator dataset. The experimental results are shown in Table 1. The YOLOv8 series models include YOLOv8n, YOLOv8s, and YOLOv8m. The YOLOv8n model is the lightest version in the YOLOv8 series, with the fewest network layers and the shallowest depth. YOLOv8s enhanced the network width, thus improving the model's expressive ability. YOLOv8m further expanded the network width to achieve higher detection accuracy.

[0115] Table 1 Experimental Results in the YOLOv8 Series

[0116]

[0117] Table 1 shows the experimental results in the YOLOv8 series. In the experiment, the comparison parameters included the mAP value, the number of parameters, GFLOPs, the model size, the training time, and the inference time. The mAP value of YOLOv8m was the highest, reaching 93.7%, but it took the longest time and had the largest number of parameters. The mAP value of YOLOv8-SDEI was slightly lower than that of YOLOv8m, but the number of parameters decreased by 91% and the computational amount decreased by 92%. The overall training time of YOLOv8n and the inference time for a single image were the least, but YOLOv8-SDEI improved the accuracy by 0.16% compared to the baseline model YOLOv8n, and the corresponding number of parameters and computational amount were also smaller. In the improved YOLOv8-SDEI model of the present invention, the multi-scale fusion shared-parameter detection head improved the feature fusion ability. The EMA attention mechanism made the model more focused on key information. The depthwise separable convolution reduced the number of model parameters and the computational amount. Inner-IoU enhanced the regression effect of the model's prediction boxes. The experimental results demonstrated the advantages of the YOLOv8-SDEI model.

[0118] The ablation experiment results are shown in Table 2. "√" indicates that the corresponding improved module is used in the network. The design of the detection head with multi-scale fusion and parameter sharing promotes the effective fusion between features at different levels, enhancing the model's feature extraction and representation capabilities. The EMA attention mechanism guides the model to dynamically focus on the key regions in the image, improving the detection accuracy and efficiency. The DW convolution enables the model to be more easily deployed in resource-constrained environments while maintaining high performance. Inner-IoU improves the localization accuracy of the prediction box and its fit with the true target, and this improvement significantly enhances the model's detection performance in complex scenarios. YOLOv8-SDEI conducts experiments on the Head design of YOLOv8, introducing DW convolution, EMA attention mechanism, and Inner-IoU loss function, and observes the impact of these four factors on the model performance.

[0119] Table 2 Ablation Experiment Results

[0120]

[0121] In the ablation experiment, it can be seen that taking YOLOv8n as the baseline model, after introducing the multi-scale fusion and shared-parameter detection head in YOLOv8-SDEI, the mAP increased by 0.03%, and the computational cost and number of parameters decreased by 11% and 18% respectively. After adding DW, the mAP increased by 0.02%, the computational cost decreased by 13%, and the number of parameters decreased by 15%. After adding the EMA attention mechanism, the number of parameters decreased by 6% and the mAP increased by 1.1%. After replacing the loss function with Inner-IoU, the mAP increased by 0.08%. Finally, compared with the YOLOv8 model, the improved model had an mAP increase of 2.2%, a computational cost decrease of 25%, and a parameter decrease of 34%. This proves the effectiveness of the YOLOv8-SDEI model.

[0122] To deeply analyze the performance of the improved model proposed in this invention in insulator defect detection, a series of comparative experiments were conducted. The YOLOv8-SDEI model of this invention was compared with several other current mainstream object detection algorithms. These comparative algorithm models include: Faster R-CNN, SSD, YOLOv5n, and YOLOv7n. Under the same dataset and training environment, this invention conducted experiments on these models. The experimental results are shown in Table 3.

[0123] Table 3 Experimental Results of Different Algorithms on the Insulator Dataset

[0124]

[0125] According to the information in the table, the YOLOv8-SDEI model has a computational volume of 6.5G. Compared with the SSD model, SSD sacrifices accuracy for detection speed. Although the computational volume of the improved model is higher than that of SSD, the P (precision) and R (recall) have been greatly improved. Compared with YOLOv5n and YOLOv7n, although the number of parameters is slightly inferior to theirs, the YOLOv8-SDEI model has obvious advantages in mAP and FPS. While reducing the computational complexity and the number of parameters for the baseline model YOLOv8, the improved model also achieves the best results in important performance indicators. Considering all indicators comprehensively, the YOLOv8-SDEI model is more suitable for insulator defect detection.

[0126] The insulator detection effects of YOLOv8-SDEI and YOLOv8 are as Figure 9 shown. The overall detection accuracy of the YOLOv8-SDEI model is higher than that of YOLOv8n. YOLOv8n has insufficient detection ability for some small targets, resulting in missed detections, and the accuracy of the labeled detection boxes is relatively low.

Claims

1. A method for detecting defects in insulators of high-voltage transmission lines by drone inspection, characterized in that: This detection method is implemented based on the constructed YOLOv8-SDEI model. First, by designing a detection head with multi-scale fusion and shared parameters, the feature fusion capability is effectively improved; Secondly, the introduction of the EMA attention mechanism makes the model more focused on the information in the key areas; the use of deep separable convolution significantly reduces the number of parameters and computational complexity of the model; finally, using Inner-IoU as the bounding box loss function, by paying attention to the matching degree inside the box, the detection accuracy of small objects is more effectively improved, and the performance of the model in predicting box regression is enhanced.

2. The method for detecting defects in insulators of high-voltage transmission lines according to claim 1, characterized in that: The YOLOv8-SDEI model improves the model's ability to extract and fuse multi-scale features by introducing the spatial reconstruction submodule (SRU) and the channel reconstruction submodule (CRU), so that each convolution kernel can be dynamically adjusted according to the characteristics of the input data. This mechanism can better capture the fine-grained spatial and channel information in the image, thereby enhancing the ability of feature representation and improving the accuracy of detecting insulators. It can also reduce the spatial redundancy existing in standard convolution and reduce the amount of computation and parameters of the model.

3. The method for detecting defects in insulators of high-voltage transmission lines according to claim 2, characterized in that: The overall structure of the spatial reconstruction submodule (SRU) is: First, SRU standardizes the input feature X by subtracting the mean μ and dividing by the standard deviation σ in the separation stage, and the trainable parameters γ∈R in the normalization layer (GN) are C Measure the spatial pixel variance for each batch and channel, which means that for each batch and each channel, there is a trainable parameter to measure the variance; normalize the associated weight W γ ∈R C From formula (1), we can get: Then, W is transformed into γ The weight value of the weighted feature map is mapped to (0, 1), and then gated by the threshold. Gating allows the model to focus more on key information. Finally, the input feature X is multiplied by W1 and W2 to obtain two weighted features: features with large information content and features with low information content Realize nonlinear transformation of features; In the reconstruction phase, features with large amounts of information are Features with low information content Sum to generate features with more information, and fully integrate the two weighted different information features to enhance the information flow between them; then cross-reconstruct the features and Concatenate to obtain the spatially refined feature map X w ; The whole process is calculated as: in, is element-by-element multiplication, ⊕ is element-by-element summation, and ∪ is concatenation; after SRU is applied to the intermediate input feature X, not only are the features with large information content separated from the features with small information content, but they are also reconstructed, which can enhance the representative features and suppress the redundant features in the spatial dimension.

4. The method for detecting defects in insulators of high-voltage transmission lines according to claim 3, characterized in that: The overall structure of the channel reconstruction submodule (CRU) is: For a given spatial refinement feature X w ∈R c×h×w First, X w The channel of CRU is divided into two parts, namely αC channel and (1-α)C channel; then the 1×1 convolution operation is further used to compress the channel of the feature map to improve the calculation efficiency; then a squeezing ratio r is introduced to control the feature channel and reduce the calculation cost of CRU; finally, the feature X defined in the space is w Divided into upper X up and lower X low , and processed at different conversion stages. up It is input into the upper transformation stage, and efficient convolution operations (GWC and PWC) are used to replace the standard k×k convolution; For the same X up Perform k×kGWC and 1×1PWC operations, and finally integrate the outputs to obtain the feature map Y1; when X low It is input to the lower transformation stage, and PWC is used to generate feature map Y2; global spatial information S with channel statistics is collected through global average pooling m ∈R c×1×1 ; Stack S1 and S2 together and generate feature importance vector β1,β2∈R through SoftMax c×1×1 ; Finally, the upper feature Y1 and the lower feature Y2 are merged in a channel manner to obtain the channel refinement feature Y; In the SCConv module, all parameters are concentrated in the conversion stage; the parameters of the standard convolution are Y = M k X can be calculated as: P s =k×k×C1×C2=k 2 C1C2 (3) The parameters of the SCConv module can be calculated as: Among them, k is the kernel size of the convolution; C1 and C2 are the number of input and output feature channels; α represents the segmentation ratio; γ represents the squeezing ratio; and g is the group size of the GWC operation.

5. The method for detecting defects in insulators of high-voltage transmission lines according to claim 4, characterized in that: The YOLOv8-SDEI model uses Depthwise Separable Convolution (DSC), which processes the insulator features of each channel independently in channel-by-channel convolution and then integrates all insulator features through point-by-point convolution, so that the model can extract fine-grained insulator feature information from each channel, thereby locating insulator defects more accurately. The parameter ratio of depth-wise separable convolution to standard convolution is: Among them, C K Represents the size of the convolution kernel, M represents the number of channels of the input feature map, and N represents the number of channels of the output feature map.

6. The method for detecting defects in insulators of high-voltage transmission lines according to claim 5, characterized in that: The YOLOv8-SDEI model introduces the attention mechanism EMA to improve the perception ability of the model. EMA builds local cross-channel interactions in each parallel sub-network without channel dimensionality reduction; and fuses the output feature maps of the two parallel sub-networks through a cross-space learning method.

7. The method for detecting defects in insulators of high-voltage transmission lines according to claim 6, characterized in that: The YOLOv8-SDEI model uses an improved Inner-IoU loss function, specifically: The GT box (ground truth box) and the anchor box (anchor) are denoted as b gt and b, the center point inside the GT box is It means that (x c ,y c ) represents the center point inside the anchor box; the length and width of the GT box are represented by w gt and h gt , the length and width of the anchor are denoted as w and h, respectively; Indicates the left, right, top, and bottom boundary positions of the actual bounding box; b l , b r , b t , b b Indicates the left, right, top, and bottom boundary positions of the auxiliary bounding box; inter indicates the intersection area of ​​the actual bounding box and the auxiliary bounding box; union indicates the union area of ​​the actual bounding box and the auxiliary bounding box; IoU inner Represents the value of Inner-IoU, that is, the intersection area divided by the union area; the value range of Inner-IoU is similar to IoU, both between [0,1]; the following formula describes the calculation process of Inner-IoU loss: union=(w gt *h gt )*(ratio) 2 +(w*h)*(ratio) 2 -inter (11) Among them, the general value range of the proportional factor ratio is [0.5, 1.5]; When the ratio is less than 1, the size of the auxiliary bounding box is smaller than the actual bounding box, which makes the regression effective range of high IoU samples smaller, but the absolute value of the gradient is greater than the gradient obtained by the IoU loss, which can accelerate the convergence of high IoU samples; When the ratio is greater than 1, the larger auxiliary bounding box expands the effective range of regression and enhances the regression effect of low IoU samples. This feature enables the Inner-IoU loss to more effectively accelerate the convergence of regression results under different scaling factors, and has a positive impact on both high IoU and low IoU samples. Inner-IoU adjusts the scaling factor for a variety of data sets and different target detectors, so that the generalization of Inner-IoU is enhanced, the detection performance of the model is improved, and the model can capture the key information of insulators faster.

Citation Information

Cited By

  • Unmanned aerial vehicle aerial photography petroleum leakage intelligent detection method and system fused with MobileNetV4 lightweight network

    CN120388155A

  • An unmanned aerial vehicle aerial oil leakage intelligent detection method and system fusing a MobileNetV4 lightweight network

    CN120388155B

  • Defect detection method of unmanned aerial vehicle for insulator chain

    CN120807402A

  • Partial multi-scale attention lightweight asphalt road defect detection method and system

    CN121053451A

  • Lightweight AI-based distribution line unmanned aerial vehicle edge end real-time visual identification and target detection method and system

    CN121459227A