Improved YOLOv11 power transmission line defect detection method

By introducing Triplet Attention, Shift-Wise Convolution and SEAMHEAD modules into the YOLOv11 model, the problem of low detection accuracy of existing models in complex scenarios is solved, and higher detection accuracy and robustness are achieved, meeting the safety risk control needs of the power industry.

CN120032175APending Publication Date: 2025-05-23SICHUAN POWER TRANSMISSION & TRANSFORMATION CONSTR +2

Patent Information

Application Number
CN202510196749.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing models have poor adaptability to complex scenarios, poor feature expression ability, and low detection accuracy in complex contexts.

Method used

The Triplet Attention module was introduced to optimize the feature extraction capability to improve the model's performance in complex scenarios; the Shift-Wise Convolution (SWC) module was used to spatially offset the feature map in YOLOv11's detection head design to enhance feature correlation at different scales; the SEAMHEAD module was introduced to optimize the feature learning capability of the object detection model through auxiliary supervision strategies and adaptive adjustments.

Benefits of technology

It significantly improves the accuracy and robustness of target detection, improves detection accuracy and training efficiency, and can achieve excellent performance in high-precision requirements and real-time tasks, meeting the needs of safety risk control in the power industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032175A_ABST
    Figure CN120032175A_ABST
Patent Text Reader

Abstract

The invention discloses an improved YOLOv11 power transmission line defect detection method, which introduces an attention mechanism to optimize the feature extraction capability, improves the performance of a model in a complex scene, remarkably enhances the understanding of the model for local and global information, and improves the target detection precision. Spatial offset is carried out on a feature map in a detection head by adopting feature extraction, feature association of different scales is enhanced, the detection precision of a small target is improved, redundant calculation is reduced, the calculation efficiency and performance of a model are further optimized, diversified requirements in a complex environment are met, and the method is suitable for popularization and application. The feature learning ability of a self-integrated attention head optimization target detection model is introduced, and the detection precision and robustness are improved under the scenes of small object detection and complex backgrounds. The method effectively improves the detection precision and training efficiency, can achieve excellent performance in a high-precision requirement and a real-time task, and meets the requirements of safety risk management and control in the power industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric power patrol detection, and in particular to a method for detecting power transmission line defects by improving YOLOv11. Background Art

[0002] As the scale of power transmission lines continues to expand, their operation and maintenance tasks are becoming increasingly arduous. Currently, the status monitoring of power transmission lines is mainly carried out by manual inspection, drone inspection and robot inspection. However, manual inspection has low efficiency and high safety risks. Although drone inspection has improved efficiency, defect analysis still mainly relies on manual processing, which is time-consuming and labor-intensive. Robot inspection is restricted by line structure and high cost, and it is difficult to promote and apply it.

[0003] In this context, intelligent detection technology based on image processing has gradually become a research focus, but existing methods still have many problems, such as insufficient detection robustness in complex scenes, easy omission of small target defects, occlusion and target overlap affecting detection accuracy, and poor real-time performance when processing large-scale high-resolution images. In addition, the reliance of these models on large-scale labeled data makes the training cost high, and the complexity of the model also limits its deployment capabilities in resource-constrained devices. Especially in the complex environment of power transmission lines, interference factors such as strong light, shadows, rain and snow, and occlusion of background vegetation increase the difficulty of detection. Therefore, improving the accuracy, robustness and lightweight performance of defect detection models, especially solving the problem of small target detection in complex scenes, is still a key technical bottleneck that needs to be broken through in the intelligent operation and inspection of power transmission lines.

[0004] In response to the above problems, scholars use drones to patrol transmission lines and collect patrol data, and use deep learning technology to detect targets on the collected data to find out whether there are safety hazards. The patent document with application publication number CN117953395A discloses a method for detecting foreign objects in transmission lines based on a lightweight model. The teacher network is constructed by replacing the backbone network of YOLOv5s with ResNet50. After the foreign object image training set is input, the student network is channel pruned and trained under the guidance of the teacher network, and finally the optimized student network is obtained. Using the trained student model to detect foreign object images in transmission lines can ensure high recognition accuracy while solving the problem of insufficient computing power and resources in embedded devices through a lightweight model structure. The patent document with application publication number CN118298159A discloses a method for detecting defects in transmission line anti-vibration hammers based on improved YOLOv8s. The anti-vibration hammer images are collected through inspection to construct a sample set, which is annotated and preprocessed to form a data set. Improve the YOLOv8s model: replace the downsampling module in Backbone with the CG module, add the CBAM attention module on the Neck side, and replace the loss function with Inner-SIOU. Use the processed data set to train the model, and finally achieve high-precision defect detection of the anti-vibration hammer image, improving the detection accuracy and efficiency of small-size defects.

[0005] Although the above-mentioned public algorithms have achieved good results in image defect detection by improving accuracy and lightweight models, there are still some shortcomings. First, the model may have poor adaptability to complex scenes and may not perform well in multi-target detection. In addition, the computational complexity is still high, which may still affect practical applications. Pruning operations may weaken the ability to express features, resulting in reduced detection accuracy in complex backgrounds. Summary of the invention

[0006] The technical problem to be solved by the present invention is that the existing model has poor adaptability to complex scenes, poor feature expression ability, and low detection accuracy under complex backgrounds. The purpose is to provide a method for detecting power line defects by improving YOLOv11. First, by introducing the Triplet Attention module, the feature extraction ability is optimized, the performance of the model in complex scenes is improved, and the model's understanding of local and global information is significantly enhanced, thereby improving the accuracy of target detection. Secondly, by introducing the Shift-wise Convolution (SWC) module, the feature map is spatially offset in the detection head design of YOLOv11, the feature association of different scales is enhanced, the detection accuracy of small targets is improved, and redundant calculations are reduced, further optimizing the computational efficiency and performance of the model, and adapting to the diverse needs in complex environments. Finally, the SEAMHEAD module is introduced, and the feature learning ability of the target detection model is optimized through auxiliary supervision strategy and adaptive adjustment, forming the YOLOv11-SST algorithm proposed by the present invention, especially in the scene of small object detection and complex background, the algorithm improves the detection accuracy and robustness. It effectively improves detection accuracy and training efficiency, and can achieve excellent performance in high-precision and real-time tasks, meeting the needs of safety risk management in the power industry.

[0007] The present invention is achieved through the following technical solutions:

[0008] The first aspect of the present invention provides a method for detecting power line defects by improving YOLOv11, comprising the following specific steps:

[0009] Obtaining the field data set of the transmission line inspection, preprocessing the field data set, and generating a sample set;

[0010] Build an initial YOLOv11 model, improve the initial YOLOv11 model, and obtain a defect detection model;

[0011] The improvements to the initial YOLOv11 model include:

[0012] The TripletAttention module is introduced into the original YOLOv11 backbone network to model the attention weights of different dimensions of the feature map.

[0013] A feature extraction C3K2-SWC module is added to the neck network of the initial YOLOv11, attention weights are modeled based on different dimensions of the feature map, and feature extraction is performed by combining convolution optimization with multi-scale feature extraction and context information capture;

[0014] The SEAMHEAD module is introduced into the initial YOLOv11 detection head to perform multi-scale feature fusion on the extracted features.

[0015] The sample set is input into the defect detection model for training. The trained defect detection model is optimized through hyperparameter adjustment and loss function to obtain the best model weights. The optimized defect detection model is used for transmission line defect detection.

[0016] Furthermore, the obtaining of a field data set of a transmission line inspection, preprocessing of the field data set, and generating a sample set specifically includes:

[0017] Collect relevant data information of transmission lines, including images of transmission lines, defect annotations and related environmental data, and construct a transmission line defect dataset for experiments;

[0018] Preprocessing the collected transmission line defect dataset, including image normalization and standardization;

[0019] Perform data augmentation operations on the normalized and standardized images to obtain a sample set;

[0020] The sample set is divided into training set, validation set and test set according to the proportion.

[0021] Furthermore, the TripletAttention module of the attention mechanism is introduced into the backbone network of YOLOv11 to model the attention weights of different dimensions of the feature map, specifically including:

[0022] Obtain a feature map, introduce Z-pool(·) to perform maximum pooling and average pooling on the feature map in the spatial dimension, capture the multidimensional dependency relationship in the feature map, and obtain the multidimensional dependency relationship in the feature map. The multidimensional dependency relationship includes three branch information, wherein the dependency relationship includes: channel dependency, spatial dependency, and mixed dependency;

[0023] The three branch information is weightedly fused to obtain the attention weights of different dimensions of the feature map.

[0024] Furthermore, the Z-pool (·) performs maximum pooling and average pooling on the feature map in the spatial dimension, specifically including:

[0025] Z-pool(χ)=[MaxPool 0d (χ),AvgPool 0d (x)]

[0026] Among them, χ represents the input feature map, 0d represents the 0th dimension, and MaxPool 0d (·) represents the maximum pooling operation, AvgPool 0d (·) represents the average pooling operation, and Z-pool(·) represents the output result of the pooling.

[0027] Furthermore, the weighted fusion of the three branch information specifically includes:

[0028]

[0029] Among them, y represents the attention weights modeled in different dimensions of the feature map, It is obtained by rotating the input vector χ 90° counterclockwise along the H axis. It means The tensor obtained after the Z-pool(·) operation, It is obtained by rotating the input vector χ 90° counterclockwise along the W axis. It means The tensor obtained after the Z-pool(·) operation, It is obtained by rotating the input vector χ 90° counterclockwise along the H axis. It means The tensor obtained by the Z-pool(·) operation. For the last branch, the channels of the input tensor χ are reduced to two by Z-pool(·), resulting in a simplified tensor of shape (2×H×W) σ represents the activation function, ψ 1 ,ψ 2 ,ψ 3 represents a 2D convolutional layer, defined by the kernel size k in the three branches of triplet attention.

[0030] Furthermore, the feature extraction C3K2-SWC module is added to the neck network of the initial YOLOv11, and the feature extraction is performed by combining convolution optimization multi-scale feature extraction and context information capture, specifically including:

[0031] Model the attention weights of different dimensions of the feature map by performing a single convolution, a secondary convolution, and a shift operation in sequence;

[0032] The first convolution is used to extract low-dimensional features;

[0033] The secondary convolution is used to enhance the spatial information of the low-dimensional features extracted by the primary convolution;

[0034] The offset operation is used to offset the spatial position of the feature map that enhances the spatial information.

[0035] Furthermore, the one-time convolution specifically comprises the following steps:

[0036] X'=Conv 1×1 (X)

[0037] Among them, X' represents the 1×1 convolution Conv 1×1 The feature map after

[0038] The specific steps of the secondary convolution include:

[0039] X”=Conv 3×3 (X')

[0040] Among them, X" represents the 3×3 convolution Conv 3×3 The feature map after

[0041] The specific steps of the offset operation include:

[0042] X SWC =S(X″)

[0043] X out =α·X SWC +β·X″

[0044] Among them, X SWC represents the feature map after SWC operation, which is used to capture the context information and the dependency between adjacent pixels. out represents the final output feature map, α and β are weight factors used to control the influence of each feature map.

[0045] Furthermore, after the optimized defect detection model is used to detect the defects of the transmission line, the detection results are evaluated. The specific steps of the evaluation include:

[0046]

[0047] Among them, P represents precision, R represents recall, and M represents AP represents the mean average precision, T P Indicates the number of samples that the model predicts to be positive and are actually positive, F P Indicates the number of samples that the model predicts to be positive but are actually negative; F N Indicates the number of samples that the model predicts to be negative but are actually positive; A Pi represents the average precision of the i-th category, and N represents the total number of categories in the dataset.

[0048] A second aspect of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, an improved YOLOv11 method for detecting power transmission line defects is implemented.

[0049] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for detecting power transmission line defects using an improved YOLOv11.

[0050] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0051] Firstly, the Triplet Attention module is introduced. Through three parallel branches, it focuses on channel dependency, spatial dependency and mixed dependency respectively, optimizes the feature extraction capability, and improves the performance of the model in complex scenes. In particular, it significantly enhances the model's understanding of local and global information in the case of small target detection, target overlap and complex background, thereby improving the accuracy of target detection. Secondly, the Shift-Wise Convolution (SWC) module is used to perform spatial shift on the feature map in the detection head design of YOLOv11, enhance the feature association of different scales, improve the detection accuracy of small targets and reduce redundant calculations, further optimize the computational efficiency and performance of the model, and adapt to the diverse needs in complex environments. Finally, the SEAMHEAD module is introduced to optimize the feature learning ability of the target detection model through auxiliary supervision strategy and adaptive adjustment, especially in the scene of small object detection and complex background, improving the detection accuracy and robustness. Overall, the YOLOv11 model using these modules effectively improves the detection accuracy and training efficiency, and can achieve excellent performance in high-precision and real-time tasks, meeting the needs of safety risk management in the power industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative work. In the drawings:

[0053] Figure 1 is a flowchart of a method for detecting defects in a power transmission line in an embodiment of the present invention;

[0054] Figure 2 : is a network structure diagram of YOLOv11 in the power transmission line defect detection method in an embodiment of the present invention;

[0055] Figure 3 is a network structure diagram of Triplet Attention in a power transmission line defect detection method in an embodiment of the present invention;

[0056] Figure 4a The structure of the SWC network in the power transmission line defect detection method in the embodiment of the present invention is Figure 1 ;

[0057] Figure 4b The structure of the SWC network in the power transmission line defect detection method in the embodiment of the present invention is Figure 2 ;

[0058] Figure 4c The structure of the SWC network in the power transmission line defect detection method in the embodiment of the present invention is Figure 3 ;

[0059] Figure 4d FIG4 is a structural diagram of a SWC network in a power transmission line defect detection method in an embodiment of the present invention;

[0060] Figure 5 A network structure diagram of a SEAMHEAD detection head designed in a power transmission line defect detection method in an embodiment of the present invention;

[0061] Figure 6 4 is an overall structural diagram of YOLOv11-SST of the power transmission line defect detection method in an embodiment of the present invention;

[0062] Figure 7 This is a detection effect diagram of the YOLOv11-SST model proposed in an embodiment of the present invention in the detection of transmission line defect targets. DETAILED DESCRIPTION

[0063] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.

[0064] As a possible implementation, Figure 1 As shown, this embodiment provides a method for detecting power line defects by improving YOLOv11, comprising the following steps:

[0065] Step 1: First, collect the image dataset of transmission line equipment and perform preprocessing on these images, including data cleaning, enhancement, and annotation normalization to ensure data quality. Then, divide the dataset into training set, validation set, and test set in a ratio of 7:2:1 to prepare for the subsequent model training. Subsequently, establish a detection model based on the YOLOv11 network structure, and carefully set the training parameters according to the scale and quality of the dataset, thus laying the foundation for accurate defect detection.

[0066] YOLOv11 is optimized on the basis of YOLOv8, improving the Backbone and Neck structures, enhancing the feature extraction capability, and adapting to more complex target detection tasks. Its structure is as follows Figure 2As shown in the figure. In the backbone part, the C2f module of YOLOv8 is replaced with the C3k2 module, and a C2PSA module is added after the SPPF module. The latter consists of two convolutional layers and a multi-head self-attention module to enhance feature extraction. In addition, YOLOv11 also replaces the conventional convolutional layer of the classification branch in the detection head with a depthwise convolution, and adjusts the depth, width, and max_channels ratio parameters of the model.

[0067] The object detection network of YOLOv11 consists of three parts: Backbone, Neck, and Head. Backbone is responsible for extracting low-level features of the input image, such as edges and textures, and constructing high-level semantic information. The main modules include C3K2, SPPF, and the newly added C2PSA modules, that is, the modules connected in sequence: Conv, Conv, C3k2, Conv, C3k2, Conv, C3k2, Conv, C3k2, Conv, C3k2, SPPF, and C2PSA. The C3K2 module consists of convolution, Bottleneck structure, and residual connection to optimize feature extraction and information fusion. The SPPF module uses maximum pooling, splicing, and 1×1 convolution to extract multi-scale information. The C2PSA module integrates channel and spatial attention mechanisms, and by calculating channel and spatial weights, it increases attention to key areas, suppresses background noise interference, and achieves precise positioning.

[0068] When processing images, the backbone network of YOLOv11 first preprocesses the input images, including resizing, normalization, and data enhancement, to ensure that the data adapts to the network and increase the diversity of training data. Next, the image is subjected to preliminary feature extraction by the C3K2 module, which combines small convolution kernels and the cross-stage partial connection (CSP) structure to extract low-level detail features. Subsequently, the residual connection structure effectively fuses features at different levels to further enhance the ability to express features.

[0069] Next, the feature map enters the SPPF module (improved spatial pyramid pooling) to extract multi-scale contextual information through a fixed-size maximum pooling kernel. The pooling result is concatenated with the original feature map and fused through a 1×1 convolution to ensure that the network can capture local information and integrate global information when processing targets of different sizes. The feature map then passes through the C2PSA module (channel and spatial attention), which uses the channel attention mechanism to enhance features with strong semantic information, while using the spatial attention mechanism to focus on the target area and effectively suppress background interference. Finally, these multi-level feature maps will be passed to subsequent parts of the network (such as the neck and detection head) for target classification and positioning, ensuring that detection tasks can be efficiently completed in various complex scenarios.

[0070] Step 2: The Triplet Attention module is introduced at the end of the YOLOv11 backbone network to further improve the feature extraction capability, especially in complex scenarios such as small target detection, target overlap, and complex background. The Triplet Attention module uses three parallel branches to focus on channel dependency, spatial dependency, and mixed dependency, thereby enhancing the network's ability to understand local and global information. Its network structure is as follows: Figure 3 This module can effectively capture the multi-dimensional dependencies in the input feature map, help the network to strengthen the focus on key areas in features at different levels, suppress background interference, and improve target positioning accuracy.

[0071] The Z-pool in Triplet Attention is one of its core modules, which is used to enhance the performance of the channel attention mechanism. The calculation formula of the Z-pool process is:

[0072] Z-pool(χ)=[MaxPool 0d (χ),AvgPool 0d (x)]

[0073] Among them, χ is the input feature, 0d is the 0th dimension, and MaxPool 0d (·) and AvgPool 0d (·) represent the maximum pooling operation and the average pooling operation respectively, and Z-pool(·) is the output result of the pooling.

[0074] Triplet Attention helps improve detection accuracy by weighted fusion of information from three branches, and due to its lightweight design, it hardly increases computational overhead. Its calculation formula is:

[0075]

[0076] Among them, y is the output result of fusing three branches, It is obtained by rotating the input vector χ 90° counterclockwise along the H axis. It means The tensor obtained after the Z-pool(·) operation, It is obtained by rotating the input vector χ 90° counterclockwise along the W axis. It means The tensor obtained after the Z-pool(·) operation, It is obtained by rotating the input vector χ 90° counterclockwise along the H axis. It means The tensor obtained by the Z-pool(·) operation. For the last branch, the channels of the input tensor χ are reduced to two by Z-pool(·), resulting in a simplified tensor of shape (2×H×W) σ is the activation function, and ψ 1 ,ψ 2 ,ψ 3 Represents a standard 2D convolutional layer, defined by the kernel size k in the three branches of triplet attention.

[0077] Step 3: In the YOLOv11 network, the feature extraction process is further optimized by introducing the Shift-Wise Convolution (SWC) module, especially for target detection tasks in complex scenes. The SWC module uses spatial offset operations to enhance the modeling ability of contextual information and can effectively capture the relationship between local features, thereby improving the network's understanding of global and local features. Its structure is as follows: Figure 4a , 4b The core idea of ​​the SWC module is to achieve more efficient feature learning by gradually shifting the position of the convolution kernel in space.

[0078] The design of the SWC module has the advantage of being lightweight. Compared with traditional convolution operations, it reduces computational complexity and the number of parameters. By reducing unnecessary computation, SWC effectively improves the computational efficiency of the network while ensuring feature expression capabilities. Therefore, this paper designs the C3K2-SWC module to replace the last C3K2 module of the neck network, further enhancing the model's multi-scale feature extraction capabilities and improving the accuracy of target detection.

[0079] First is the 1×1 convolution:

[0080] X'=Conv 1×1 (X)

[0081] Among them, X' is the 1×1 convolution Conv 1×1 This operation is usually used to reduce the number of channels and extract low-dimensional features. Next, a 3×3 convolution is applied to process spatial information:

[0082] X”=Conv 3×3 (X')

[0083] Here, X" is the convolution layer after 3×3 convolution. 3×3 The feature map after the image is constructed enhances the expression of spatial information.

[0084] The SWC (Shift-Wise Convolution) operation enhances the local and global information capture capability of the feature map by shifting the spatial position of the feature map. Assume that the spatial shift operation is represented by the S(·) function:

[0085] X SWC =S(X″)

[0086] Among them, X SWC It is the feature map after SWC operation, which can effectively capture the context information and the dependency between adjacent pixels. Finally, feature fusion is performed, and the fusion operation can be achieved through weighted average or addition operation:

[0087] X out =α·X SWC +β·X″

[0088] Among them, X SWC represents the feature map after SWC operation, which is used to capture the context information and the dependency between adjacent pixels. out represents the final output feature map, α and β are weight factors used to control the influence of each feature map.

[0089] Step 4: Introduce the SEAMHEAD module to replace the detection head of the YOLOv11 model and improve the performance of target detection through an adaptive weighting mechanism. The SEAMHEAD (Self-Enhancement Adaptive Module for HEAD) module further improves the performance of the target detection model through adaptive weighting and self-enhancement strategies, especially in complex scenes, occlusions and small object detection. The core principle of this module is to adaptively adjust the weighted distribution of features so that the detection head can focus more on important feature areas in the image, thereby improving detection accuracy when dealing with complex backgrounds and multiple targets. Its network structure is as follows: Figure 5 shown.

[0090] In the SEAMHEAD module, the weight of each feature map is obtained through an adaptive learning process, which can adjust the influence of different features according to the characteristics of the target object and the image content. This weighting strategy enables the model to locate targets more accurately in multi-scale and multi-target scenes without being disturbed by background noise.

[0091] Step 5: Scale the collected transmission line inspection dataset to a preset size and send it to the improved YOLOv11 network model (i.e., YOLOv11-SST model) for training to obtain the optimal model for target detection. The YOLOv11-SST model structure is as follows: Figure 6 shown.

[0092] Step 6: Use the optimal model obtained in step 5 to perform target detection, output the detection results, and analyze and evaluate the results. In step 6, the calculation formulas for each indicator are as follows:

[0093]

[0094] Among them, P represents precision, R represents recall, and M represents AP represents the mean average precision, T P Indicates the number of samples that the model predicts to be positive and are actually positive, F P Indicates the number of samples that the model predicts to be positive but are actually negative; F N Indicates the number of samples that the model predicts to be negative but are actually positive; A Pi M represents the average precision of the i-th category, and N represents the total number of categories in the data set. AP The evaluation is divided into M AP 0.5 and M AP 0.5:0.95 two categories, M AP 0.5 Use a fixed IoU threshold of 0.5 (i.e., detection results with IoU ≥ 0.5 are correct) to calculate the average precision (AP) of each category, and then calculate the average of all categories. Furthermore, for multiple IoU thresholds (from 0.5 to 0.95, with a step size of 0.05, i.e., 0.5, 0.55, 0.6, ..., 0.95), calculate AP separately and take the average of these AP values. M AP 0.5:0.95 This evaluation method has stricter standards. Higher IoU thresholds (such as 0.75, 0.9) require a more precise match between the predicted box and the true box. Usually, these values ​​are lower because high IoU thresholds require higher accuracy, so the performance evaluation of the YOLO model focuses more on M. AP 0.5 for this type of result.

[0095] A defect image detection method for power line inspection provided in this embodiment aims to further promote the intelligence of power inspection through deep learning, improve the safety monitoring capability of power inspection, and ensure the stable operation of the power system. This method is based on the latest YOLOv11 model and makes multiple adjustments to solve the problem of target detection in complex scenes. First, the Triplet Attention module is introduced to improve the feature extraction capability through parallel channel, space and mixed dependency branches, and enhance the processing capability of small targets, target overlap and complex background, so that YOLOv11 can more accurately capture the multi-dimensional information in the input image and significantly improve the detection accuracy. Secondly, the C3K2-SWC module is used to further optimize the convolution structure in the network, improve the computational efficiency and reduce the number of parameters. The C3K2-SWC module reduces redundant calculations through flexible convolution strategies, while retaining key information, and balancing the accuracy and computational efficiency of the model. Finally, the SEAMHEAD module is added to introduce a more efficient feature fusion strategy in the target detection process, and the model's target recognition capability in complex backgrounds is improved through multi-scale context information guidance, ensuring that the detection results are more accurate and robust.

[0096] Through the above improvements, the present invention successfully constructed a high-performance and efficient transmission line defect detection method, namely the YOLOv11-SST model. The improved YOLOv11-SST model showed more outstanding performance in power inspection tasks, significantly improved detection efficiency and accuracy, and could better adapt to the strict requirements of the power industry for safety standardization and defect detection.

[0097] In order to verify the superior performance of the YOLOv11-SST model proposed in this invention, comparative experiments with existing target detection methods were carried out, covering models such as YOLOv3-tiny, YOLOv5s, YOLOv8n and YOLOv11n. The experiments were carried out in a unified experimental environment to ensure that the python and pytorch versions were consistent, and the same training data set was used for training and testing. The experimental results are shown in Table 1. Compared with the latest YOLOv11n model, YOLOv11-SST improved its accuracy by 3.7%, its recall by 1.0%, and its mAP@0.5 by 1.9%. While ensuring the detection accuracy, the number of parameters of YOLOv11-SST was reduced by nearly 4.8%, which significantly optimized the computational efficiency of the model and greatly improved the detection speed (FPS) compared to YOLOv11. The specific detection effects are as follows: Figure 7 As shown, this method has efficient and accurate detection capabilities.

[0098] Table 1 is compared with the original model

[0099]

[0100] In addition, ablation experiments were conducted to analyze the contribution of different modules to the model performance. In the experiment, the key modules in the YOLOv11-SST model were removed and compared with the original model. The experimental results are shown in Table 2. Since the model names of the ablation experiments are long, abbreviations are used instead, such as SE (SEAMHEAD) and TA (Triplet Attention). Through the ablation experiment, it is found that after removing the key modules, YOLOv11-SST still has a good performance in indicators such as precision, recall, and mAP@0.5. This result further proves the synergy of each module in the YOLOv11-SST model and its excellent performance in target detection tasks.

[0101] Table 2 is compared with the original model

[0102]

[0103] As a possible implementation, this embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, an improved YOLOv11 transmission line defect detection method is implemented.

[0104] As a possible implementation, this embodiment provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for detecting power line defects using an improved YOLOv11.

[0105] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for detecting power line defects by improving YOLOv11, characterized in that: The specific steps include: Obtaining the field data set of the transmission line inspection, preprocessing the field data set, and generating a sample set; Build an initial YOLOv11 model, improve the initial YOLOv11 model, and obtain a defect detection model; The improvements to the initial YOLOv11 model include: The TripletAttention module is introduced into the original YOLOv11 backbone network to model the attention weights of different dimensions of the feature map. A feature extraction C3K2-SWC module is added to the neck network of the initial YOLOv11, attention weights are modeled based on different dimensions of the feature map, and feature extraction is performed by combining convolution optimization with multi-scale feature extraction and context information capture; The SEAMHEAD module is introduced into the initial YOLOv11 detection head to perform multi-scale feature fusion on the extracted features. The sample set is input into the defect detection model for training. The trained defect detection model is optimized through hyperparameter adjustment and loss function to obtain the best model weights. The optimized defect detection model is used for transmission line defect detection.

2. The improved YOLOv11 transmission line defect detection method according to claim 1, characterized in that: The obtaining of a field data set of a transmission line inspection, preprocessing of the field data set, and generating a sample set specifically includes: Collect relevant data information of transmission lines, including images of transmission lines, defect annotations and related environmental data, and construct a transmission line defect dataset for experiments; Preprocessing the collected transmission line defect dataset, including image normalization and standardization; Perform data augmentation operations on the normalized and standardized images to obtain a sample set; The sample set is divided into training set, validation set and test set according to the proportion.

3. The improved YOLOv11 transmission line defect detection method according to claim 1, characterized in that: The TripletAttention module is introduced into the initial YOLOv11 backbone network to model the attention weights of different dimensions of the feature map, including: Obtain a feature map, introduce Z-pool(·) to perform maximum pooling and average pooling on the feature map in the spatial dimension, capture the multidimensional dependency relationship in the feature map, and obtain the multidimensional dependency relationship in the feature map. The multidimensional dependency relationship includes three branch information, wherein the dependency relationship includes: channel dependency, spatial dependency, and mixed dependency; The three branch information is weightedly fused to obtain the attention weights of different dimensions of the feature map.

4. The improved YOLOv11 transmission line defect detection method according to claim 3, characterized in that: The Z-pool (·) performs maximum pooling and average pooling on the feature map in the spatial dimension, specifically including: Z-pool(x)=[MaxPool 0d (x),AvgPool 0d (x)] Among them, χ represents the input feature map, 0d represents the 0th dimension, and MaxPool 0d (·) represents the maximum pooling operation, AvgPool 0d (·) represents the average pooling operation, and Z-pool(·) represents the output result of the pooling.

5. The improved YOLOv11 transmission line defect detection method according to claim 4, characterized in that: The weighted fusion of the three branch information specifically includes: Among them, y represents the attention weights modeled in different dimensions of the feature map, It is obtained by rotating the input vector χ 90° counterclockwise along the H axis. It means The tensor obtained by the Z-pool(·) operation, It is obtained by rotating the input vector χ 90° counterclockwise along the W axis. It means The tensor obtained by the Z-pool(·) operation, It is obtained by rotating the input vector χ 90° counterclockwise along the H axis. It means The tensor obtained by the Z-pool(·) operation. For the last branch, the channels of the input tensor χ are reduced to two by Z-pool(·), resulting in a simplified tensor of shape (2×H×W) σ denotes the activation function, ψ1, ψ2, ψ3 denote the 2D convolutional layers defined by the kernel size k in the three branches of triplet attention.

6. The improved YOLOv11 transmission line defect detection method according to claim 1, characterized in that: The feature extraction C3K2-SWC module is added to the neck network of the initial YOLOv11, and the feature extraction is performed by combining convolution optimization multi-scale feature extraction and context information capture, specifically including: Model the attention weights of different dimensions of the feature map by performing a single convolution, a secondary convolution, and a shift operation in sequence; The first convolution is used to extract low-dimensional features; The secondary convolution is used to enhance the spatial information of the low-dimensional features extracted by the primary convolution; The offset operation is used to offset the spatial position of the feature map that enhances the spatial information.

7. The improved YOLOv11 transmission line defect detection method according to claim 6, characterized in that: The specific steps of the first convolution include: X'=Conv 1×1 (X) Among them, X' represents the 1×1 convolution Conv 1×1 The feature map after The specific steps of the secondary convolution include: X”=Conv 3×3 (X') Among them, X" represents the 3×3 convolution Conv 3×3 The feature map after The specific steps of the offset operation include: X SWC =S(X”) X out =α·X SWC +β·X” Among them, X SWC represents the feature map after SWC operation, which is used to capture the context information and the dependency between adjacent pixels. out represents the final output feature map, α and β are weight factors used to control the influence of each feature map.

8. The improved YOLOv11 power transmission line defect detection method according to claim 1, characterized in that: After the optimized defect detection model is used to detect the defects of the transmission line, the detection results are evaluated. The specific steps of the evaluation include: Among them, P represents precision, R represents recall, and M represents AP represents the mean average precision, T P Indicates the number of samples that the model predicts to be positive and are actually positive, F P Indicates the number of samples that the model predicts to be positive but are actually negative; F N Indicates the number of samples that the model predicts to be negative but are actually positive; A Pi represents the average precision of the i-th category, and N represents the total number of categories in the dataset.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the improved YOLOv11 transmission line defect detection method as described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the improved YOLOv11 transmission line defect detection method as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Power transmission line foreign matter image detection method, system and device and storage medium

    CN117953395A

  • Power transmission line stockbridge damper image defect detection method based on YOLOv8s

    CN118298159A

  • PCB defect detection method and device based on improved YOLOv8

    CN119205673A

Cited By

  • Image-based airport runway foreign matter intelligent detection and clearance auxiliary processing system and method

    CN120279511A

  • Photovoltaic panel dust retention detection method and device, electronic equipment and storage medium

    CN120298402A

  • Surface defect small target detection method based on multi-scale feature interaction

    CN120374613A

  • A small target detection method for surface defects based on multi-scale feature interaction

    CN120374613B

  • Improved YOLO11-based water hyacinth target rapid detection method

    CN120635395A