Transform improvement-based power transmission line multi-target defect detection method
By adopting an improved method based on Transformer in transmission line defect recognition technology, combined with Swin Transformer and CBAM attention mechanism, the YOLO11 model is optimized, and the accuracy and real-time problems of defect recognition in complex environments are solved, and efficient and accurate multi-objective defect detection is achieved.
Patent Information
- Application Number
- CN202510080016.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-19
- Publication Date
- 2025-05-30
AI Technical Summary
The existing transmission line defect identification technology is difficult to accurately and in real time in complex and changeable environments, and it is impossible to accurately distinguish the specific differences between various defects in the line, resulting in mis-checking and missed inspections, affecting the efficiency of line patrol work.
Using a multi-objective defect detection method for transmission lines based on Transformer, the image data of multiple angles and multiple perspectives is collected through the drone camera, combined with the Swin Transformer backbone network and CBAM attention mechanism, the YOLO11 model is optimized to improve detection accuracy and real-time.
It significantly improves detection accuracy and real-timeness, can accurately identify multiple defects in complex backgrounds, reduce false and missed inspections, improve line patrol work efficiency, and adapt to the real-time detection needs of drone patrols.
Smart Images

Figure CN120070853A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and specifically relates to a multi-objective defect detection method for transmission lines improved based on transformers. Background Art
[0002] As a core facility of the power system, the transmission line is responsible for delivering electricity from the power plant to users, and its stable operation directly affects the safety and reliability of power supply. The equipment of the transmission line may have various defects due to the influence of the natural environment, equipment aging or human factors. Common problems include insulator damage, foreign object entanglement (such as balloons and kites), and bird nests on poles and towers. If these defects are not discovered and processed in time, they may have a serious impact on the power system, such as power outages, equipment damage, and even safety accidents. Therefore, it is crucial to identify these defects in a timely and accurate manner to ensure the normal operation of the transmission line.
[0003] Traditional transmission line detection methods mainly rely on manual inspections. Although some problems can be found, the efficiency is low and it is easy to miss some defects. In recent years, automated detection technologies based on computer vision have become an important means to replace manual inspections. In particular, the introduction of object detection algorithms has provided great convenience for the automated detection of transmission line defects. However, due to the complex environment of the transmission line, a large amount of background noise, and the fact that defect targets are usually small and scattered, traditional object detection methods still have difficulty meeting the requirements of real-time performance and high precision.
[0004] To solve the above technical defects, the patent with the application number CN202410921477.7 proposes an insulator defect detection method based on improved yolov5, which includes the following steps: S1: Extract relevant features according to the relevant characteristics of the transmission line data set to obtain sufficient sample data; obtain transmission line defect data, and enrich the data set based on data augmentation of OpenCV. Four forms of image compression, image rotation, image cropping, and image color transformation are performed to augment the overhead line dangerous object data set. The resize function in OpenCV is used to modify the image resolution. When calling the resize function, the interpolation method is determined, and bilinear interpolation is used to complete the adjustment of the image resolution. Bilinear interpolation means performing linear interpolation once in both the x and y directions; S2: Complete the construction of the lightweight model of the convolutional neural network, and at the same time use the obtained transmission line sample data to complete the model training. Based on the yolov5 algorithm, the principle of the attention mechanism is analyzed; S3: Establish a transmission line fault recognition network, and realize the positioning and classification of transmission line defects in complex backgrounds through multiple transmission line training sets and test sets; Although it can realize the recognition of transmission line defects, when detecting line equipment defects in a complex and diverse environment, it is difficult to accurately and real-time identify them, and it cannot accurately distinguish the specific differences of various types of lines.
[0005] Therefore, the applicant has proposed a multi-objective defect detection method for transmission lines improved based on transformers. Summary of the Invention
[0006] The present invention aims to solve the limitations of existing transmission line defect recognition technologies. Specifically, when detecting line equipment in a complex and changeable environment, traditional technologies have technical problems such as difficulty in accurately and real-time identifying equipment defects and inability to precisely distinguish the specific differences of various defects in the line; these technical limitations often lead to false detections and missed detections, resulting in a slowdown in the efficiency of line patrol work; for this reason, the present invention proposes a multi-objective defect detection method for transmission lines improved based on Transformers to improve the reliability of detection and ensure good detection efficiency.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0008] A multi-objective defect detection method for transmission lines improved based on transformers, comprising the following steps:
[0009] Step S1: Use the high-definition camera equipped on the unmanned aerial vehicle to take multi-angle and multi-perspective images of the target area. By adjusting the flight path and camera angle, ensure that the entire research area is covered, and rich image data is collected; these images will form a complete data set, and the images included in the data set should have high resolution, clarity, and sufficient diversity to ensure the accuracy of subsequent processing and analysis. The quality and quantity of the data set are the basis for subsequent analysis and model training;
[0010] Step S2: Use Labelimg to label the target data set, label the collected images, divide the data set according to a certain proportion, and finally obtain a line equipment defect data set, and divide the data set into a training set, a validation set, and a test set;
[0011] Step S3: Obtain the improved YOLO11 model, use this model to identify the obtained transmission line defect images, and use Precision, Recall, mean average precision (mAP), and frames per second (FPS) as evaluation indicators.
[0012] In step S1, use the unmanned aerial vehicle camera to take original images of 4 common and typical transmission line defects in the transmission line area, integrate them with the open-source data set, and screen and sort the images.
[0013] In step S2, distinguish the obtained data set pictures into five types: insulators, insulator self-explosion, insulator flashover, bird nests, and line foreign objects;
[0014] In step S1, the collected transmission line defect images are labeled into five types: insulator, insulator self-explosion, insulator flashover, bird's nest, and line foreign object. Among them, insulator is "insulator", insulator self-explosion is "burst", insulator flashover is "defect", bird's nest is "nest", and line foreign object is "foreign_obj".
[0015] In step S3, the obtained improved YOLO11 model is specifically as follows:
[0016] The initialized input image is input to the Swin Transformer backbone network module. The output of the Swin Transformer backbone network module is connected to the input of the SPPF module. The output of the SPPF module is connected to the input of the C2PSA module. The output of the C2PSA module is connected to the input of the first Upsample module and the input of the fourth C3k2 module. The output of the first Upsample module is connected to the input of the first Concat module. The output of the first Concat module is connected to the input of the first C3k2 module. The output of the first C3k2 module is connected to the input of the first CBAM module. The output of the first CBAM module is connected to the input of the second Upsample module. The output of the second Upsample module is connected to the input of the second Concat module. The output of the second Concat module is connected to the input of the second C3k2 module. The output of the second C3k2 module is connected to the input of the second CBAM module. The output of the second CBAM module is connected to the input of the first 11Detect detection head module and the input of the first Conv module. The output of the first Conv module is connected to the input of the third Concat module. The output of the third Concat module is connected to the input of the third C3k2 module. The output of the third C3k2 module is connected to the input of the third CBAM module. The output of the third CBAM module is connected to the input of the second 11Detect detection head module and the input of the second Conv module. The output of the second Conv module is connected to the input of the fourth Concat module. The output of the fourth Concat module is connected to the input of the fourth C3k2 module. The output of the fourth C3k2 module is connected to the input of the third 11Detect detection head module.
[0017] In step S3, the specific structure of the CBAM module is as follows:
[0018] Initialize the input of the input feature module to connect to the input of the channel attention mechanism module and the input of the first splicing module. The output of the channel attention mechanism module is connected to the input of the first splicing module. The output of the first splicing module is connected to the input of the second splicing module and the input of the spatial attention mechanism. The output of the spatial attention mechanism is connected to the input of the second splicing module. The output of the second splicing module is the refined feature.
[0019] The structure of the channel attention mechanism module is specifically as follows:
[0020] The input feature information is input to the original feature map module. The output of the original feature map module is connected to the input of the max pooling module and the input of the average pooling module. The output of the max pooling module and the output of the average pooling module are connected to the input of the shared fully connected module. The output of the shared fully connected module is connected to the input of the Maxout module and the input of the Avgout module. The output of the Maxout module and the output of the Avgout module are connected to the input of the Sigmoid module.
[0021] The structure of the spatial attention mechanism module is specifically as follows:
[0022] The input feature information is input to the original feature map module. The output of the original feature map module is connected to the input of the max pooling module and the input of the average pooling module. The output of the max pooling module and the output of the average pooling module are connected to the input of the splicing module. The output of the splicing module is connected to the input of the Conv module. The output of the Conv module is connected to the input of the Sigmoid module.
[0023] The Swin Transformer backbone network module is specifically as follows:
[0024] The input feature information is connected to the input of the Patch Partition module. The output of the Patch Partition module is connected to the input of the Linear Embeding module. The output of the Linear Embeding module is connected to the input of the first Swin Transformer Block module. The output of the first Swin Transformer Block module is connected to the input of the second Swin Transformer Block module. The output of the second Swin Transformer Block module is connected to the input of the first Patch Merging module. The output of the first Patch Merging module is connected to the input of the third Swin Transformer Block module. The output of the third Swin Transformer Block module is connected to the input of the fourth Swin Transformer Block module. The output of the fourth Swin Transformer Block module is connected to the input of the second Patch Merging module. The output of the second Patch Merging module is connected to the input of the fifth Swin Transformer Block module. The output of the fifth Swin Transformer Block module is connected to the input of the sixth Swin Transformer Block module. The output of the sixth Swin Transformer Block module is connected to the input of the third Patch Merging module. The output of the third Patch Merging module is connected to the input of the seventh Swin Transformer Block module. The output of the seventh Swin Transformer Block module is connected to the input of the eighth Swin Transformer Block module. The output of the eighth Swin Transformer Block module completes the entire processing process.
[0025] Compared with the prior art, the present invention has the following technical effects:
[0026] The technical effects of the present invention are:
[0027] 1) Significantly improve the detection accuracy and real-time performance: The improved model combines the Swin Transformer backbone network and the CBAM attention mechanism, effectively enhancing the detection ability for small targets and defects in complex backgrounds, and showing higher accuracy and robustness in multi-object detection scenarios;
[0028] 2) Multi-object classification and localization: The present invention can achieve high-precision real-time classification and localization for various defect types such as insulator damage and foreign object entanglement (such as bird nests) on transmission lines, significantly reducing the phenomena of false detection and missed detection;
[0029] 3) Adapt to complex scenarios: By introducing the Inner-EIoU bounding box loss function, the model performs excellently in processing scenarios with large illumination changes and complex backgrounds, especially showing strong adaptability to small defect targets in transmission lines;
[0030] 4) Improve operation efficiency: The present invention optimizes the model structure, adopts a hierarchical Swin Transformer module, reduces the consumption of computing resources, and at the same time maintains the detection effect of high precision and high frame rate;
[0031] 5) High practical application value: The model supports end-to-end detection, is applicable to the drone inspection task of transmission lines, can quickly generate defect detection reports, and provides technical support for the safe operation of the power system;
[0032] 6) Remarkable verification results: Through experimental comparative analysis, the improved model is superior to the traditional YOLO11 model in indicators such as Precision, Recall, and mAP, and has broad application potential and promotion value;
[0033] 7) The present invention proposes a model that improves YOLO11 based on the Swin Transformer backbone network, combines the CBAM attention mechanism to further improve the performance of the model. The improved model can efficiently identify various defects in transmission lines, such as insulator damage, foreign object entanglement, and tower bird nests. Especially in the case of small targets and complex backgrounds, it shows remarkable detection accuracy and real-time performance. By introducing Swin Transformer, the model can better process complex scenarios and long-range dependence information, while the CBAM attention mechanism further enhances the attention to key features, thus effectively improving the detection accuracy and robustness. The model not only has a high detection accuracy rate but also can achieve real-time monitoring, providing strong technical support for the automatic inspection of transmission lines and the safe operation of the power system. Brief Description of the Drawings
[0034] The following further describes the present invention in conjunction with the drawings and embodiments:
[0035] Figure 1 is the schematic diagram of the network structure of the present invention;
[0036] Figure 2 is the schematic diagram of the channel attention mechanism of the present invention;
[0037] Figure 3Schematic diagram of the spatial attention mechanism in the present invention;
[0038] Figure 4 Schematic diagram of the complete process of the CBAM module in the present invention;
[0039] Figure 5 Swin Transformer network structure diagram of the present invention;
[0040] Figure 6 Schematic diagram of the detection results of "insulator" and its "burst" defect in the embodiment of the present invention;
[0041] Figure 7 Schematic diagram of the detection results of "insulator" and its "flashover" defect in the embodiment of the present invention;
[0042] Figure 8 Schematic diagram of the detection results of "bird damage" (Nest) in the embodiment of the present invention;
[0043] Figure 9 Schematic diagram of the detection results of "foreign object" (foreign_obj) in the embodiment of the present invention. Detailed implementation manners
[0044] A multi-objective defect detection method for transmission lines improved based on transformer, comprising the following steps:
[0045] Step S1: Use the high-definition camera equipped on the unmanned aerial vehicle to take multi-angle and multi-view images of the target area. By adjusting the flight path and camera angle, ensure that the entire research area is covered, and rich image data is collected. These images will form a complete data set. The images included in the data set should have high resolution, clarity, and sufficient diversity to ensure the accuracy of subsequent processing and analysis. The quality and quantity of the data set are the basis for subsequent analysis and model training;
[0046] Step S2: Use Labelimg to label the target data set, label the collected images, divide the data set according to a certain ratio, and finally obtain a line equipment defect data set, and divide the data set into a training set, a validation set, and a test set;
[0047] Step S4: Obtain the improved YOLO11 model, use this model to identify the obtained transmission line defect images, and use Precision, Recall, mean average precision (mAP), and frames per second (FPS) as evaluation indicators.
[0048] The improved YOLO11 model includes a Backbone part (the backbone network part), a Feature Fusion Neck part (the neck network), and a Head part (the detection head part). The Backbone part is the part of the model for extracting image features, which includes a group of Swin Transformer backbone network modules, C2SPA modules, and SPPF modules. The Neck part is composed of an Upsample upsampling module, a Concat splicing module, C3k2 modules, and CBAM modules. The Head part is composed of three end-to-end detection heads (11detect modules), including an Inner-EIoU bounding box loss function.
[0049] The structure of this model is as follows: The initialized input image is input to the Swin Transformer backbone network module. The output of the Swin Transformer backbone network module is connected to the input of the SPPF module. The output of the SPPF module is connected to the input of the C2PSA module. The output of the C2PSA module is connected to the input of the first Upsample module and the input of the fourth C3k2 module. The output of the first Upsample module is connected to the input of the first Concat module. The output of the first Concat module is connected to the input of the first C3k2 module. The output of the first C3k2 module is connected to the input of the first CBAM module. The output of the first CBAM module is connected to the input of the second Upsample module. The output of the second Upsample module is connected to the input of the second Concat module. The output of the second Concat module is connected to the input of the second C3k2 module. The output of the second C3k2 module is connected to the input of the second CBAM module. The output of the second CBAM module is connected to the input of the first 11Detect detection head module and the input of the first Conv module. The output of the first Conv module is connected to the input of the third Concat module. The output of the third Concat module is connected to the input of the third C3k2 module. The output of the third C3k2 module is connected to the input of the third CBAM module. The output of the third CBAM module is connected to the input of the second 11Detect detection head module and the input of the second Conv module. The output of the second Conv module is connected to the input of the fourth Concat module. The output of the fourth Concat module is connected to the input of the fourth C3k2 module. The output of the fourth C3k2 module is connected to the input of the third 11Detect detection head module;
[0050] Specifically, step S1 includes:
[0051] S1.1: Use a drone camera to take pictures of 5 types of basic line defects commonly found in transmission lines to obtain an original dataset.
[0052] S1.2: Select and extract open-source images preferentially and stack them in a folder with the original dataset.
[0053] Specifically, step S2 includes:
[0054] Preprocess the dataset to reduce the inaccuracy and contingency of some dataset images. The methods of data preprocessing include rotation, folding, and grayscaling. Finally, 16,995 images are obtained.
[0055] Specifically, step S3 includes:
[0056] S3.1: Perform annotation in the Anaconda environment. Use the labelimg image annotation tool to annotate the collected power transmission line defect classification images. The recognition objects are divided into 5 categories: insulator, insulator flashover, insulator burst, bird's nest, and line foreign object. Among them, the insulator is labeled as "insulator", the insulator flashover is labeled as "defect", the insulator burst is labeled as "burst", the bird's nest is labeled as "nest", and the line foreign object is labeled as "foreign_obj"; obtain the Pascal VOC format data label (XML label), and then convert it into a txt file suitable for training through the data format conversion code written in python. The file includes most of the data, including the x coordinate, y coordinate, width, and height of the center point of the bounding box and other annotation information.
[0057] S3.2: After completing the annotation work, divide the dataset to obtain the training set, validation set, and test set of the power transmission line equipment defect image classification detection dataset, obtaining 14,000 training set images, 2,000 validation set images, and 995 test set images;
[0058] Specifically, step S4 includes:
[0059] S4: Replace the YOLO11s backbone network with the swin transformer backbone network.
[0060] Swin Transformer has successfully applied the Transformer model to the field of computer vision and demonstrated excellent performance in object detection and instance segmentation tasks on the COCO dataset. Swin Transformer uses a hierarchical construction method similar to that in convolutional neural networks, dividing the feature map into multiple scales, and each scale corresponds to a different downsampling rate. This hierarchical construction helps the model better learn the local and global features of the image.
[0061] In addition, Swin Transformer uses Windows Multi-Head Self-Attention (W-MSA) to reduce the computational load. W-MSA divides the feature map into multiple non-overlapping regions and then performs self-attention operations within each region. By restricting the computational load within each region, W-MSA can effectively reduce the overall computational load. Although W-MSA can reduce the computational load, it also isolates the information transfer between different regions. To address this issue, Swin Transformer also proposes Shifted Windows Multi-Head Self-Attention (SW-MSA). SW-MSA shifts the positions of adjacent windows and then performs self-attention operations on these misaligned windows. In this way, information can flow and be transmitted more smoothly between adjacent regions. Such a design not only increases the flexibility of the model when processing images but also improves the overall performance of the model, making it perform more excellently when dealing with complex image tasks.
[0062] Specifically, the operation process (feature extraction process) of the Swin Transformer network includes the following steps:
[0063] Given an image \(x\in\mathbb{R}\) with the number of channels \(H\times W\times3\). First, the image is input into the image patch partition layer for patching operations. Every \(4\times4\) adjacent pixels are divided into one patch, and then flattened in the channel direction. After that, the number of channels of the image is changed through a linear embedding operation, and the image size changes from \(H / 4\times W / 4\times48\) to \(H / 4\times W / 4\times C\). Then, the image will go through four stages in sequence for feature extraction to obtain feature layers of different sizes. In addition to the first stage, in the remaining three stages, the feature layer will first undergo a two-fold downsampling through an image patch merging layer and then perform feature extraction through \(n\) Swin Transformer Block structures.
[0064] The advantages of using Swin Transformer as the backbone network are mainly reflected in the following aspects:
[0065] 1) Efficient local and global feature modeling
[0066] The Swin Transformer divides the image into irregular windows (patches) in a hierarchical manner and models local features through the self-attention mechanism within local windows. At the same time, it introduces cross-window information at different scales to achieve global feature modeling. This strategy enables the Swin Transformer to not only maintain the details of local features but also handle global dependencies, similar to the efficiency of convolutional neural networks (CNNs) in processing local features, while solving the problem of excessive computational complexity of the standard Vision Transformer (ViT) in dealing with long-range dependencies.
[0067] 2) Hierarchical structure
[0068] The Swin Transformer adopts a hierarchical design, starting from smaller patches and gradually increasing the receptive field at each layer. This design is similar to the pyramid structure of traditional convolutional neural networks, enabling it to efficiently perform multi-scale feature learning and capture information in different ranges at different levels. The hierarchical structure also enhances the model's representation ability and generalization ability, especially when dealing with tasks of different resolutions (such as image classification, object detection, semantic segmentation, etc.).
[0069] 3) Advantages of combining with convolutional neural networks
[0070] Although the Swin Transformer is based on the Transformer architecture, it can also be combined with convolutional neural networks (CNNs) to combine the local perception advantages of convolutional neural networks with the global modeling ability of the Transformer, thus obtaining stronger representation ability.
[0071] In summary, the advantages of the Swin Transformer as a backbone network lie in its efficient computational performance, hierarchical structure, good transfer ability, and wide adaptability. Through the local self-attention mechanism and multi-scale feature learning, it not only improves computational efficiency but also enables the model to handle more complex visual tasks.
[0072] The structure of the Swin Transformer backbone network module is as follows: the input feature information is connected to the input of the PatchPartition module, the output of the Patch Partition module is connected to the input of the Linear Embeding module, the output of the Linear Embeding module is connected to the input of the first Swin Transformer Block module, the output of the first Swin Transformer Block module is connected to the input of the second Swin Transformer Block module, the output of the second Swin Transformer Block module is connected to the input of the first Patch Merging module, the output of the first Patch Merging module is connected to the input of the third Swin Transformer Block module, the output of the third Swin Transformer Block module is connected to the input of the fourth Swin Transformer Block module, the output of the fourth Swin Transformer Block module is connected to the input of the second Patch Merging module, the output of the second Patch Merging module is connected to the input of the fifth Swin Transformer Block module, the output of the fifth Swin Transformer Block module is connected to the input of the sixth Swin Transformer Block module, the output of the sixth Swin Transformer Block module is connected to the input of the third Patch Merging module, the output of the third Patch Merging module is connected to the input of the seventh Swin Transformer Block module, the output of the seventh Swin Transformer Block module is connected to the input of the eighth Swin Transformer Block module, and the output of the eighth Swin Transformer Block module completes the entire processing process.
[0073] The structure of the Swin Transformer Block is shown in Figure 5 . It consists of a normalization layer (layer norm, LN), a window based multi-head self-attention (W-MSA), a shifted window based multi-head self-attention (SW-MSA), and a multi-layer perceptron (MLP).
[0074] The locally input feature information is connected to the input of the first LN module. The output of the first LN module is connected to the input of the first W-MSA. The output of the first W-MSA is connected to the output of the original feature information and then to the input of the first concatenation module. The output of the first concatenation module is connected to the input of the second LN module. The output of the second LN module is connected to the input of the first MLP module. The output of the first MLP module and the output of the first concatenation module are connected to the input of the second concatenation module. The output of the second concatenation module is connected to the input of the third LN module and the input of the third concatenation module. The output of the third LN module is connected to the input of the second W-MSA module. The output of the second W-MSA module is connected to the input of the third concatenation module. The output of the third concatenation module is connected to the input of the fourth concatenation module and the input of the fourth LN module. The output of the fourth LN module is connected to the input of the second MLP module. The output of the second MLP module is connected to the input of the fourth concatenation module. The output of the fourth concatenation module is the result after processing the entire process.
[0075] S5: After embedding the CBAM attention mechanism module behind the first C3k2 module, the first C3k2 module, and the first C3k2 module in the neck part of the model, evaluate the model quality. And set up a reference experiment.
[0076] The working principle of CBAM is as follows: The CBAM module enhances the feature expression ability through the channel attention mechanism and the spatial attention mechanism. First, the channel attention mechanism uses global average pooling and max pooling to extract the global features of each channel, generates channel weights through a shared multi-layer perceptron, and weights the input features element-wise by channel. Then, the spatial attention mechanism performs average pooling and max pooling on the channel dimension to obtain a spatial feature map, generates spatial weights through convolution, and weights the feature map element-wise by position. The final output is a feature map that fuses channel and spatial information, thus significantly improving the feature expression ability and performance of the model.
[0077] The structure of the CBAM module is as follows:
[0078] The initialized input feature module is connected to the input of the channel attention mechanism module and the input of the first concatenation module. The output of the channel attention mechanism module is connected to the input of the first concatenation module. The output of the first concatenation module is connected to the input of the second concatenation module and the input of the spatial attention mechanism. The output of the spatial attention mechanism is connected to the input of the second concatenation module. The output of the second concatenation module is the refined feature.
[0079] Among them, the structure of the channel attention mechanism module is as follows:
[0080] The input feature information is input to the original feature map module. The output of the original feature map module is connected to the input of the max pooling module and the input of the average pooling module. The output of the max pooling module and the output of the average pooling module are connected to the input of the shared fully connected module. The output of the shared fully connected module is connected to the input of the Maxout module and the input of the Avgout module. The output of the Maxout module and the output of the Avgout module are connected to the input of the Sigmoid module.
[0081] Among them, the structure of the spatial attention mechanism module is as follows:
[0082] The input feature information is input to the original feature map module. The output of the original feature map module is connected to the input of the max pooling module and the input of the average pooling module. The output of the max pooling module and the output of the average pooling module are connected to the input of the concatenation module. The output of the concatenation module is connected to the input of the Conv module. The output of the Conv module is connected to the input of the Sigmoid module.
[0083] Specifically, CBAM realizes its function in the following ways:
[0084] First, the channel attention mechanism performs global average pooling and max pooling on the input transmission line defect feature map in the spatial dimension to extract global information. Subsequently, a shared multi-layer perceptron is used to generate the importance weights of each channel, and the transmission line defect feature map is adjusted by channel-wise weighting. Then, the spatial attention mechanism performs average pooling and max pooling in the channel dimension on the updated feature map, extracts the spatial information of the transmission line defects, generates the importance weights of each position through convolution after concatenation, and weights the feature map position by position. Through the synergistic effect of these two attention mechanisms, the CBAM module can effectively enhance the attention to the key features of the transmission line defects and improve the expression ability and performance of the network.
[0085] S6: Use the Inner-EIoU loss function as the bounding box loss function of the model proposed in the present invention.
[0086] Use the Inner-EIoU bounding box loss function as the new bounding box loss function of YOLO11s.
[0087] The main idea of Inner-EIoU is as follows: The main idea of Inner-EIoU (Intersection over Union with InnerRegion Emphasis) is that in the object detection task, when evaluating the matching degree between the predicted bounding box and the ground truth bounding box, not only the intersection over union (IoU) is considered, but also the internal distribution of the intersection region is particularly concerned to evaluate the position relationship and shape matching degree of the bounding boxes in a more fine-grained manner. The following is its core idea:
[0088] (1) Emphasize the importance of the internal region:
[0089] Inner-EIoU introduces a new strategy that assigns higher weights to the intersection region between the predicted box and the ground truth box, emphasizing the distribution quality within the intersection region. Compared with ordinary IoU, this method not only focuses on the area of the intersection but also pays attention to the geometric features inside the intersection region, making the evaluation more comprehensive.
[0090] (2) Combine boundary and internal distribution:
[0091] Conventional IoU only considers the ratio of the intersection area to the union area between the predicted box and the ground truth box, ignoring the specific details of the distribution inside the box. While Inner-EIoU incorporates the similarity of pixel distribution or other geometric features (such as center offset, shape difference, etc.) within the intersection region during calculation, making the evaluation criterion more precise.
[0092] (3) Optimize the adjustment of detection boxes:
[0093] During the training process of the object detection model, Inner-EIoU can provide more targeted supervision signals for Bounding Box Regression, encouraging the predicted box to not only better match the shape of the ground truth box but also pay more attention to the internal details of the intersection region, thereby improving the accuracy of box localization.
[0094] (4) More robust performance:
[0095] Compared with traditional metrics such as IoU or GIoU, Inner-EIoU can better handle the situation where the predicted box and the ground truth box highly overlap but have significant offsets or shape differences, avoiding the neglect of errors due to simple area calculation.
[0096] In summary, the main idea of Inner-EIoU is to improve the limitations of the IoU metric by strengthening the influence of the internal features of the intersection region on the box matching degree, thereby providing a more detailed box evaluation criterion in object detection tasks. This method can effectively improve the accuracy requirements of the object detection model for the position and shape of the box.
[0097] The definition of Inner-EIoU is as follows:
[0098] Inner-EIoU is evolved from Inner-IoU. Inner-IoU is an IoU loss function based on auxiliary bounding boxes, introducing a scale factor ratio to control the size of the auxiliary bounding box, and then calculating the loss to accelerate the model convergence. The definition of Inner-IoU is as follows:
[0099]
[0100]
[0101] Union = (wgt * hgt) * ratio 2 + (w * h) * ratio 2 - Inter(11)
[0102]
[0103] Wherein, and are the internal center points of the ground truth box; xc and yc are the internal center points of the predicted box; are the left, right, top, and bottom boundaries of the ground truth box, calculated from the center point, width, height of the ground truth box, and the scale factor ratio; b l b r b t b b are the left, right, top, and bottom boundaries of the predicted box, calculated from the center point, width, height of the predicted box, and the scale factor ratio; Inter represents the intersection part of the ground truth box and the predicted box; Union represents the union part of the ground truth box and the predicted box; Inner-IoU is calculated from the ratio of the two.
[0104] L Inner - EIoU = L EIoU + IoU + IoU Inner (13)
[0105] The advantages of Inner-EIoU are mainly reflected in the following aspects:
[0106] (1) More refined evaluation criteria:
[0107] Inner-EIoU not only considers the intersection over union of the predicted box and the ground truth box, but also strengthens the attention to the internal features of the intersection area, and can more accurately measure the matching degree between boxes, especially when the predicted box and the ground truth box are partially overlapped but have large shape or position deviations.
[0108] (2) Improved localization accuracy:
[0109] By introducing the internal information of the intersection area, Inner-EIoU provides a more detailed supervision signal for bounding box regression, which helps the model to more accurately optimize the position and shape of the predicted box, thus improving the localization performance of the object detection task.
[0110] (3) More robust optimization effect:
[0111] Compared with traditional IoU or its variants (such as GIoU, DIoU, etc.), Inner-EIoU is more robust in dealing with complex scenarios, capable of handling situations where the boxes have a high degree of height overlap but inconsistent shapes or internal distributions, avoiding the limitations of simple area calculations.
[0112] (4) Reduce the problem of gradient vanishing:
[0113] In extreme cases where the boxes hardly overlap or have a high degree of height overlap, the gradient of IoU may be weak, affecting the optimization effect. Inner-EIoU makes the gradient signal richer through the enhancement of internal features, making the optimization process more stable.
[0114] (5) Pay more attention to target shape matching:
[0115] The feature enhancement of the internal region enables Inner-EIoU to better capture the detailed matching between the target shape and the ground truth box, improving the sensitivity of the detection model to the appearance and geometric characteristics of the target.
[0116] (6) Wide range of application scenarios:
[0117] Inner-EIoU is applicable to tasks with higher requirements for box localization, such as fine target detection, evaluation of instance segmentation frameworks, etc., and performs better in scenarios with small targets and complex backgrounds.
[0118] In summary, Inner-EIoU introduces the consideration of the internal information of the intersection region on the basis of traditional IoU, has the advantages of finer granularity, greater robustness and higher accuracy, and has a significant improvement effect on complex target detection tasks. Specifically, step S7 includes:
[0119] S7: Put the power transmission line defect image dataset obtained in S3 into the improved YOLO11 model reconstructed in the present invention for training, the number of training times is 200 times, apply the training set to the improved YOLO11 model for training, finally generate the detection result and obtain the best.pt file. And train the original model, and analyze by comparing comprehensive performance indicators.
[0120] Specifically, step S8 includes:
[0121] S8: To evaluate the model, several key metrics were adopted. The improved YOLO11s model constructed by the present invention was evaluated by using recall (Rec), mean average precision (mAP), and frames per second (FPS). Among them, recall (Rec) reflects the ratio between the number of successfully retrieved relevant documents and the number of omitted relevant documents. Specifically, it is the proportion of the actually relevant documents that are retrieved. Mean average precision (mAP) is the aggregation of the average precision (AP) values for all classes. Frames per second (FPS) refers to the number of image frames that can be processed per second during video processing. The higher the FPS, the smoother the video performance.
[0122] Specifically, step S9 includes:
[0123] Use the trained model to identify and locate the images of line equipment defects collected on-site in the electronics laboratory of the affiliated unit, and be able to accurately classify them.
[0124] Through the ablation experiments shown in Table 1, it can be seen that when adding the Swin Transformer backbone, CBAM module, and Inner EIoU loss function simultaneously, compared with the baseline model and other improved methods, the model of the present invention is much higher than other detection models in terms of accuracy and speed, and can meet the requirements of real-time classification detection.
[0125] The following is the analysis table of the comparison results with the original model
[0126] Table 1 Comparison test result table
[0127]
[0128] The following is the analysis table of the comparison results of the loss functions
[0129] Table 2 Comparison result table of loss functions
[0130]
[0131] From the experimental results in the above table, it can be seen that Inner EIoU has the best performance and the highest accuracy.
[0132] The platform used for the tests of this invention is Pycharm 11.0.15, the model framework is Pytorch 1.13.1, the CPU is 15vCPU Intel(R) Xeon(R) Platinum 8358P CPU @ 2.60GHz, and the GPU is RTX3090, 24GB.
[0133] When setting the hyperparameters, the number of training rounds is set to 200 rounds, the optimizer uses SGD, and the batchsize is set to 16.
[0134] In summary, the YOLO11 model improved based on Swin Transformer in this invention demonstrates significant advantages and extensive application value in the identification of transmission line defects. Compared with traditional models, this model has made breakthroughs in detection accuracy, real-time performance, and small target recognition ability. It can efficiently handle detection tasks with multiple targets and complex backgrounds, and meet the actual needs of transmission line inspections. In practical applications, this model has the ability of real-time classification and precise positioning, effectively reducing the errors of manual detection and improving the efficiency and reliability of defect detection. In addition, by introducing the Swin Transformer backbone network and the CBAM attention mechanism, this model significantly enhances the ability to focus on key features, showing higher accuracy and robustness. This invention not only improves the identification efficiency and detection accuracy of transmission line defects, but also meets the real-time detection requirements of drone inspections, providing an important technical guarantee for the safe operation of the power system, and having significant economic and social significance.
Claims
1. A multi-target defect detection method for power transmission lines based on transformer improvement, characterized in that: The following steps are involved: Step S1: Use the high-definition camera equipped with the drone to take images of the target area from multiple angles and perspectives. By adjusting the flight path and camera angle, ensure that the entire study area is covered and collect rich image data. These images will form a complete data set. The images contained in the data set should have high resolution, clarity, and sufficient diversity to ensure the accuracy of subsequent processing and analysis. The quality and quantity of the data set are the basis for subsequent analysis and model training. Step S2: Use Labelimg to annotate the target data set, annotate the collected images, divide the data set according to a certain ratio, and finally obtain the line equipment defect data set, and divide the data set into a training set, a verification set, and a test set; Step S3: Obtain an improved YOLO11 model, use the model to identify the acquired power transmission line defect image, and use precision Precision, recall rate Recall, mean average precision mAP and transmission frame number per second FPS as evaluation indicators.
2. The method according to claim 1, characterized in that In step S1, the original images of four common and typical transmission line defects are taken in the transmission line area using a drone camera, integrated with open source datasets, and the images are screened and sorted.
3. The method according to claim 1, characterized in that In step S2, the obtained data set images are divided into five types: insulator, insulator self-explosion, insulator flashover, bird's nest, and line foreign matter; In step S1, the collected transmission line defect images are annotated and divided into five types: insulator, insulator burst, insulator flashover, bird's nest, and line foreign objects. Insulator is "insulator", insulator burst is "burst", insulator flashover is "defect", bird's nest is "nest", and line foreign objects are "foreign_obj".
4. The method according to claim 1, characterized in that: In step S3, the obtained improved YOLO11 model is specifically: Initialize the input image and input it to the Swin Transformer backbone network module. The output of the Swin Transformer backbone network module is connected to the input of the SPPF module. The output of the SPPF module is connected to the input of the C2PSA module. The output of the C2PSA module is connected to the input of the first Upsample module and the input of the fourth C3k2 module. The output of the first Upsample module is connected to the input of the first Concat module. The output of the first Concat module is connected to the input of the first C3k2 module. The output of the first C3k2 module is connected to the input of the first CBAM module. The output of the first CBAM module is connected to the input of the second Upsample module. The output of the second Upsample module is connected to the input of the second Concat module. The output of the second Concat module is connected to the input of the second C3k2 module. The output of the block is connected to the input of the second CBAM module, the output of the second CBAM module is connected to the input of the first 11Detect detection head module and the input of the first Conv module, the output of the first Conv module is connected to the input of the third Concat module, the output of the third Concat module is connected to the input of the third C3k2 module, the output of the third C3k2 module is connected to the input of the third CBAM module, the output of the third CBAM module is connected to the input of the second 11Detect detection head module and the input of the second Conv module, the output of the second Conv module is connected to the input of the fourth Concat module, the output of the fourth Concat module is connected to the input of the fourth C3k2 module, and the output of the fourth C3k2 module is connected to the input of the third 11Detect detection head module.
5. The method according to claim 1, characterized in that In step S3, the CBAM module structure is specifically as follows: The initialized input feature module is connected to the input of the channel attention mechanism module and the input of the first splicing module. The output of the channel attention mechanism module is connected to the input of the first splicing module. The output of the first splicing module is connected to the input of the second splicing module and the input of the spatial attention mechanism. The output of the spatial attention mechanism is connected to the input of the second splicing module. The output of the second splicing module is the refined feature.
6. The method according to claim 5, characterized in that The specific structure of the channel attention mechanism module is: The input feature information is input to the original feature map module, the output of the original feature map module is connected to the input of the maximum pooling module and the input of the average pooling module, the output of the maximum pooling module and the output of the average pooling module are connected to the input of the shared fully connected module, the output of the shared fully connected module is connected to the input of the Maxout module and the input of the Avgout module, and the output of the Maxout module and the output of the Avgout module are connected to the input of the Sigmoid module.
7. The method according to claim 5, characterized in that The specific structure of the spatial attention mechanism module is: The input feature information is input to the original feature map module, the output of the original feature map module is connected to the input of the maximum pooling module and the input of the average pooling module, the output of the maximum pooling module and the output of the average pooling module are connected to the input of the splicing module, the output of the splicing module is connected to the input of the Conv module, and the output of the Conv module is connected to the input of the Sigmoid module.
8. The method according to claim 4, characterized in that The Swin Transformer backbone network module is as follows: The input feature information is connected to the input of the Patch Partition module, the output of the Patch Partition module is connected to the input of the Linear Embedding module, the output of the Linear Embedding module is connected to the input of the first Swin TransformerBlock module, the output of the first Swin Transformer Block module is connected to the input of the second Swin TransformerBlock module, the output of the second Swin Transformer Block module is connected to the input of the first Patch Merging module, the output of the first Patch Merging module is connected to the input of the third Swin Transformer Block module, the output of the third Swin Transformer Block module is connected to the input of the fourth Swin Transformer Block module, the output of the fourth Swin Transformer Block module is connected to the input of the second Patch Merging module, the output of the second PatchMerging module is connected to the input of the fifth Swin Transformer Block module, the output of the fifth Swin TransformerBlock module is connected to the input of the sixth Swin Transformer Block module, the output of the sixth Swin TransformerBlock module is connected to the input of the third Patch Merging module, the output of the third Patch Merging module is connected to the input of the seventh Swin Transformer Block module, and the output of the seventh Swin Transformer Block module is connected to the input of the eighth Swin Transformer The entire processing process is completed by inputting the Block module and outputting the eighth Swin Transformer Block module.
Citation Information
Patent Citations
Insulator defect detection method based on improved yolov5
CN118941983A
Cited By
VHR remote sensing image road intersection detection method based on YOLO11
CN121725219A