Power transmission line foreign matter detection algorithm based on SCSM-YOLOv8

By improving the backbone and neck network of YOLOv8, SDM, CSPPFS and SSFFE modules are designed, which solves the difficulties in multi-scale object detection and environmental adaptability of YOLOv8 in foreign matter detection in transmission lines, and achieves more efficient foreign matter detection.

CN120339582AActive Publication Date: 2025-07-18GUANGDONG LEINENG POWER GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510433783.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-18
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing YOLOv8 algorithm has difficulties in multi-scale object detection and poor detection effect in unfavorable environments in the detection of foreign objects in transmission lines, and the efficiency of manual inspection is inefficient.

Method used

By designing the SDM module, CSPPFS module and SSFFE module, the backbone and neck network of YOLOv8 are improved, and the switchable expansion convolutional SDconv, CBAM and SimAM attention mechanisms, and the ECA channel attention mechanism are enhanced.

Benefits of technology

It significantly improves the multi-scale object detection performance and the robustness of algorithms in complex environments, improves detection accuracy and speed, and reduces the computing burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339582A_ABST
    Figure CN120339582A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of deep learning target detection and power transmission line foreign matter combination, and particularly relates to a power transmission line foreign matter detection algorithm based on SCSM-YOLOv8, which comprises the following specific steps: creating a user-defined data set for power transmission line foreign matter detection, the training set, the verification set and the test set are divided into three parts through a random sampling method, and the proportions of the training set, the verification set and the test set are 70%, 20% and 10%; in the trunk structure, an SDM module is designed, the designed switchable expansion convolution SDconv is used to replace the original C2f standard convolution, the receptive field is expanded, and the feature extraction capability of the algorithm is improved; in the main structure, a CSPPFS module is designed, a CBAM attention mechanism and a SimAM attention mechanism are introduced into the SPPF, and the target detection performance is improved on the premise that the parameter quantity and the calculation quantity are not increased. According to the method, the detection effect of the algorithm is remarkably improved by using the SDM module, the CSPPFS module and the SSFFE module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning object detection combined with foreign objects on transmission lines, and specifically to a foreign object detection algorithm for transmission lines based on SCSM-YOLOv8. Background Technique

[0002] Transmission lines span complex environments such as mountains, cities, and rural areas, so they are extremely vulnerable to various foreign objects attaching. If these foreign objects are not discovered and cleared in time, it may cause the line to shut down, and then trigger a large-scale power outage accident, resulting in huge economic losses and having a profound impact on society. Detecting foreign objects on transmission lines is an important task in power line inspection. However, manual inspection has a high labor intensity and low efficiency. Currently, intelligent inspection, with its high safety reliability and low cost, has become the mainstream method for transmission line detection. The object detection method of intelligent inspection is mainly to extract the features of the object and perform segmentation and recognition of the object. In recent years, deep learning has developed rapidly, and object detection based on deep learning has been applied in the detection of foreign objects on transmission lines.

[0003] Currently, object detection algorithms based on deep learning can be mainly divided into two types. One is the two-stage object detection algorithm, such as R-CNN, Fast R-CNN, Faster R-CNN, etc., but its slower inference speed does not meet the requirements of the current detection task. The other is the single-stage object detection algorithm. On the basis of maintaining excellent detection accuracy, this type of algorithm can achieve a high-speed inference process. The main representatives include SSD and YOLO. Among them, YOLOv8 performs well in terms of speed and accuracy due to its flexible architecture and free anchor box design. However, in practical applications, YOLOv8 still has some problems, such as the difficulty in detecting multi-scale objects and poor detection effects in adverse environments.

[0004] Therefore, to solve the above problems, this patent proposes a foreign object detection algorithm for transmission lines based on SCSM-YOLOv8. By improving the backbone network and neck network of YOLOv8, the performance of multi-scale object detection is significantly improved, and the robustness of the algorithm in complex environments is enhanced. First, in the backbone structure, the SDM module is designed, and the designed switchable dilated convolution SDconv is used to replace the standard convolution of the original C2f, expanding the receptive field of the algorithm. Different dilation rates of convolution are applied to the same input features, and the results of different convolutions are fused through a switching function, thus expanding the receptive field and capturing more global and local information. Then, in the backbone structure, the CSPPFS module is designed, and the CBAM attention mechanism and SimAM attention mechanism are introduced into the SPPF. Without increasing the number of parameters and computational complexity, the feature expression ability is enhanced, and the object detection performance is further improved. In addition, in the neck network, the SSFFE module is designed, and the SDconv and ECA channel attention mechanisms are introduced into the ASFF to strengthen the feature fusion ability of the neck network in the YOLOv8 algorithm. Summary of the Invention

[0005] To solve the above technical problems, according to one aspect of the present invention, the following technical solutions are provided:

[0006] A foreign object detection algorithm for transmission lines based on SCSM-YOLOv8, which includes the following specific steps:

[0007] S1: Create a custom dataset for foreign object detection of transmission lines, and divide it into three parts by random sampling: training set, validation set and test set, with a ratio of 70%, 20% and 10%;

[0008] S2: In the backbone structure, design the SDM module, use the designed switchable dilated convolution SDconv to replace the standard convolution of the original C2f, expand the receptive field, and improve the feature extraction ability of the algorithm;

[0009] S3: In the backbone structure, design the CSPPFS module, introduce the CBAM attention mechanism and SimAM attention mechanism into the SPPF, and improve the object detection performance without increasing the number of parameters and computational complexity;

[0010] S4: In the neck structure, design the SSFFE module, introduce the SDconv and ECA channel attention mechanisms into the ASFF, and strengthen the feature fusion part of the neck network in the YOLOv8 algorithm;

[0011] S5: Input the dataset into the algorithm for iterative training. After the training is completed, use the optimal algorithm to detect the test set to obtain the final result.

[0012] As a preferred solution of a foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 according to the present invention, wherein: the specific steps of S1 are as follows:

[0013] S11: Create a custom dataset dedicated to foreign object detection in transmission lines, and collect images of transmission lines under variable environmental conditions and diverse perspectives to form an initial dataset;

[0014] S12: In the collected images, exclude images with low quality and unclear blurs, and only retain two-dimensional images that are clearly taken and not affected by compression to ensure the clarity and accuracy of the data;

[0015] S13: After screening, 2157 images containing daily foreign objects are obtained, and their images form a basic dataset. To improve the generalization ability of the algorithm, various data augmentation techniques including rotation, translation, cropping, flipping, and adjustments of brightness, grayscale, and contrast are implemented on the images. Through the data augmentation techniques, the number of images increases to 6382, and a custom dataset for foreign object detection in transmission lines is obtained, which contains different categories of foreign objects;

[0016] S14: Before algorithm training, randomly divide the augmented dataset and allocate it to the training set, validation set, and test set according to the ratio of 7:2:1. After division, the training set contains 4467 images for algorithm training, the validation set contains 1276 images for algorithm performance verification, and the test set contains 639 images for evaluating the final performance of the algorithm.

[0017] As a preferred solution of a foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 according to the present invention, wherein: the switchable dilated convolution SDconv in S2 includes a preposed multi-channel global context module, an SDconv unit, and a postposed multi-channel global context module;

[0018] The steps of the preposed multi-channel global context module are as follows: Receive the input feature map and process it through two pooling layers with different field-of-view ranges; the two pooling layers can capture global context information of different scales; after being processed by the pooling layers, the feature map is further compressed through a 1×1 convolutional layer to reduce the number of channels to half of the original, so as to reduce the computational burden of the algorithm and at the same time extract more refined features; the outputs of the two pooling layers are fused in the channel dimension to integrate context information of different scales and prepare for further processing of the SDconv unit;

[0019] The steps of the SDconv unit are as follows: The input features are processed through two dilated convolutional layers with different dilation rates to capture feature information at different scales; the two dilated convolutional layers respectively use dilation rates d1 and d2 to extract multi-scale features; a switching function is used to dynamically adjust the output weights of the two dilated convolutional layers. The switching function consists of a 5×5 average pooling layer and a 1×1 convolutional layer, which is used to generate a set of weights to adjust the proportion of the contributions of the two convolutional layers; the output of the SDconv unit is the weighted sum of the outputs of the two dilated convolutional layers, which helps the algorithm to perform feature fusion at different scales;

[0020] The steps of the subsequent multi-channel global context module are as follows: After the SDconv unit, the subsequent multi-channel global context module is used again to further refine the features and provide rich context information for the subsequent layers of the network, enhancing the expressive ability of the features.

[0021] As a preferred solution of the transmission line foreign object detection algorithm based on SCSM-YOLOv8 of the present invention, wherein: the steps of the SDM module are as follows: the input features first pass through an SDconv and then enter an SDbottleneck. At the same time, the input features also pass through a traditional convolutional layer; the output of the SDbottleneck is concatenated with the output of the standard convolutional layer in the channel dimension, and then feature integration is performed through another SDconv, and finally used as the output of the SDM module;

[0022] The steps of the SDbottleneck are as follows: the input features are first processed through an SDconv and then passed through another SDconv again. The output of the first SDconv is added to the output of the second SDconv element by element to form a residual connection, which helps to improve the gradient flow and prevent the gradient vanishing problem. The summed feature map is used as the final output of the SDbottleneck.

[0023] As a preferred solution of a foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 of the present invention, wherein: in the CSPPFS module of S3, its input features first pass through the CBAM module, and then through the SimAM module. The SimAM module enhances important features by adaptively adjusting the channel weights of the feature map. The CBAM module enhances feature representation through explicit channel and spatial attention mechanisms. The SimAM module adaptively captures the saliency information in the feature map in a parameter-free manner. The combination of the two enables the model to better extract and utilize features at different levels and dimensions, complementing each other. Specifically, the CBAM module globally enhances features at a higher level, while the SimAM module locally enhances features at a lower level. This multi-level feature enhancement strategy enables the model to better capture information at different scales and improve the performance of the model.

[0024] As a preferred solution of a foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 of the present invention, wherein: the SSFFE module in S4 is composed of feature size adjustment, dynamic feature fusion, and ECA channel attention mechanism weight calibration;

[0025] In the feature size adjustment, the size of the feature map is unified through upsampling or downsampling techniques. During upsampling, a convolutional layer is used to adjust the number of channels, and the resolution is increased through interpolation techniques. For 2-fold downsampling, SDconv with a stride of 2 is used to simultaneously adjust the number of channels and the resolution. For 4-fold downsampling, first, dimensionality reduction is performed through a max-pooling layer with a stride of 2, and then SDconv with a stride of 2 is applied for further adjustment. In SDconv, when the dilation rate is set to 3, the size of the convolutional kernel increases, thereby obtaining a wider receptive field, which helps to more comprehensively capture the image content during the feature fusion stage;

[0026] In the dynamic feature fusion, SSFFE integrates the feature information of the first three levels after size adjustment, and the information is shared among all channels to promote the effective integration of features;

[0027] In the ECA channel attention mechanism weight calibration, the ECA channel attention mechanism first performs global average pooling on each channel of the input feature map, compressing the spatial dimension to 1×1 to generate a channel descriptor containing global information. Then, the ECA channel attention mechanism automatically determines the kernel size of the one-dimensional convolution according to the number of channels and calculates through a non-linear mapping, ensuring that the kernel size is odd to maintain the central symmetry of the convolution operation. The ECA channel attention mechanism uses one-dimensional convolution to process the result of global average pooling to identify the inter-channel relationships. This convolution method avoids the need for dimensionality reduction and increase, reduces the number of parameters, and at the same time preserves the integrity of information. The output of the one-dimensional convolution passes through the Sigmoid activation function to generate attention weights for each channel. The attention weights reflect the importance of each channel. Finally, the attention weights are multiplied element-wise with the corresponding channels of the original feature map to achieve weighted channel features, highlighting key features and suppressing less important information.

[0028] As a preferred solution of a transmission line foreign object detection algorithm based on SCSM-YOLOv8 according to the present invention, wherein: the calculation formula of the ASFF is as follows:

[0029]

[0030] Wherein, represents the feature vector at the (x, y) position of the output feature map of the nth layer, α, β, and γ represent the spatial weights from the mth layer (m ∈ 1, 2, 3) to the nth layer feature map, and A, B, and C respectively represent the feature vectors at the (x, y) position of different feature maps from the mth layer to the nth layer; the ASFF weights and sums the feature vectors with different characteristics according to the corresponding spatial weights, and finally obtains the fused feature F, so as to integrate the effective information of each layer of features into different scales.

[0031] As a preferred solution of a transmission line foreign object detection algorithm based on SCSM-YOLOv8 according to the present invention, wherein: the ECA channel attention mechanism first performs global average pooling on the input feature map, compressing the spatial dimension to 1x1, then captures local cross-channel dependencies through one-dimensional convolution, and then activates the result through the Sigmoid activation function to enhance non-linear features, obtaining the attention weights of each channel. Finally, the obtained attention weight matrix is reshaped into the same shape as the original input feature map and multiplied element-wise with the input feature map to achieve recalibration of channel features;

[0032] The calculation formula of the ECA channel attention mechanism is:

[0033]

[0034] Among them, k is the size of the one-dimensional convolution kernel, b and γ are hyperparameters, set to 1, C is the number of channels of the input feature, and |t| odd represents the odd number closest to t, and its formula ensures that the convolution kernel size is odd.

[0035] Compared with the prior art:

[0036] The present invention aims to achieve more accurate foreign object detection on transmission lines; first, in the backbone structure, an SDM module is designed, and the designed switchable dilated convolution SDconv is used to replace the standard convolution of the original C2f, expanding the receptive field of the algorithm, realizing the application of convolutions with different dilation rates on the same input feature, and fusing different convolution results through a switching function, thereby expanding the receptive field and capturing more global and local information; then, in the backbone structure, a CSPPFS module is designed, and the CBAM attention mechanism and the SimAM attention mechanism are introduced into the SPPF, improving the target detection performance without increasing the number of parameters and the amount of computation; in addition, in the neck network, an SSFFE module is designed, and the SDconv and the ECA channel attention mechanism are introduced into the ASFF to strengthen the feature fusion part of the neck network in the YOLOv8 algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is the flowchart of the method of the present invention;

[0038] Figure 2 is the overall network structure diagram of the present invention;

[0039] Figure 3 is the structure diagram of the SDconv of the present invention;

[0040] Figure 4 is the structure diagram of the SDM of the present invention;

[0041] Figure 5 is the structure diagram of the CSPPFS of the present invention;

[0042] Figure 6 is the structure diagram of the SSFFE of the present invention;

[0043] Figure 7 is the structure diagram of the ECA channel attention mechanism of the present invention;

[0044] Figure 8 is some dataset pictures of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0046] The present invention provides a foreign object detection algorithm for transmission lines based on SCSM-YOLOv8. Please refer to Figures 1-8 , and the specific steps are as follows:

[0047] S1: Create a custom dataset for foreign object detection in transmission lines, and divide it into three parts: training set, validation set, and test set by random sampling, with a ratio of 70%, 20%, and 10%;

[0048] S2: In the backbone structure, design an SDM module, and use the designed switchable dilated convolution SDconv (the SDconv structure is as Figure 3 shown) to replace the standard convolution of the original C2f, expand the receptive field, and improve the feature extraction ability of the algorithm;

[0049] S3: In the backbone structure, design a CSPPFS module, introduce the CBAM attention mechanism and the SimAM attention mechanism into the SPPF, and improve the target detection performance without increasing the number of parameters and the amount of computation;

[0050] S4: In the neck structure, design an SSFFE module, introduce the SDconv and ECA channel attention mechanisms into the ASFF, and strengthen the feature fusion part of the neck network in the YOLOv8 algorithm;

[0051] S5: Input the dataset into the algorithm for iterative training. After the training is completed, use the optimal algorithm to detect the test set to obtain the final result.

[0052] The specific steps of S1 are as follows:

[0053] S11: Create a custom dataset specifically for foreign object detection in transmission lines, and collect images of transmission lines from variable environmental conditions and diverse perspectives to form an initial dataset;

[0054] S12: In the collected images, exclude images with poor quality and unclear clarity, and only retain two-dimensional images that are clearly taken and not affected by compression to ensure the clarity and accuracy of the data;

[0055] S13: After screening, 2157 images containing daily foreign objects are obtained, and their images form a basic dataset. To improve the generalization ability of the algorithm, various data augmentation techniques including rotation, translation, cropping, flipping, and adjustments of brightness, grayscale, and contrast are implemented on the images. Through the data augmentation techniques, the number of images increases to 6382, and a custom dataset for foreign object detection in transmission lines is obtained, which contains different types of foreign objects, such as 1505 balloons, 1765 kites, 1421 bird nests, and 1691 pieces of garbage;

[0056] S14: Before algorithm training, randomly partition the augmented dataset and allocate it to the training set, validation set, and test set according to a ratio of 7:2:1. After partitioning, the training set contains 4,467 images for algorithm training, the validation set contains 1,276 images for validating algorithm performance, and the test set contains 639 images for evaluating the final performance of the algorithm.

[0057] The switchable dilated convolution SDconv in S2 includes a pre - multi - channel global context module, an SDconv unit, and a post - multi - channel global context module;

[0058] The steps of the pre - multi - channel global context module are as follows: Receive the input feature map and process it through two pooling layers with different field - of - view ranges; the two pooling layers can capture global context information at different scales; after being processed by the pooling layers, the feature map is further compressed through a 1×1 convolutional layer to reduce the number of channels to half of the original, so as to reduce the computational burden of the algorithm and extract more refined features at the same time; the outputs of the two pooling layers are fused in the channel dimension to integrate context information at different scales and prepare for further processing of the SDconv unit;

[0059] The steps of the SDconv unit are as follows: The input feature is processed through two dilated convolutional layers with different dilation rates to capture feature information at different scales; the two dilated convolutional layers use dilation rates d1 and d2 respectively to extract multi - scale features; a switching function is used to dynamically adjust the output weights of the two dilated convolutional layers, and the switching function consists of a 5×5 average pooling layer and a 1×1 convolutional layer to generate a set of weights to adjust the proportion of the contributions of the two convolutional layers; the output of the SDconv unit is the weighted sum of the outputs of the two dilated convolutional layers, which helps the algorithm to perform feature fusion at different scales;

[0060] The steps of the post - multi - channel global context module are as follows: After the SDconv unit, the post - multi - channel global context module is used again to further refine the features and provide rich context information for the subsequent layers of the network, enhancing the expressive ability of the features.

[0061] The steps of the SDM module (structure as shown in Figure 4 a) are as follows: The input feature first passes through an SDconv and then enters an SDbottleneck. At the same time, the input feature also passes through a traditional convolutional layer; the output of the SDbottleneck is concatenated with the output of the standard convolutional layer in the channel dimension, and then feature integration is performed through another SDconv, and finally it is used as the output of the SDM module;

[0062] The SDbottleneck (structure as shown inFigure 4 The steps for the case shown in b are as follows: The input features are first processed by an SDconv, and then passed through another SDconv again. The output of the first SDconv is added element-wise to the output of the second SDconv to form a residual connection, which helps to improve the gradient flow and prevent the problem of gradient vanishing. The feature map after the addition is used as the final output of the SDbottleneck.

[0063] Among them, the dilated convolution, as a flexible convolution operation, can effectively increase the receptive field of the convolutional network without changing the size of the feature map; the dilated convolution adjusts the coverage range of the convolutional kernel, enabling the network to capture broader spatial information, which is crucial for understanding the context relationship in the image; the dilated convolution realizes the intermittent expansion of the convolutional kernel by introducing a key parameter - the "dilation rate". This parameter determines the spacing between the elements of the convolutional kernel; in the dilated convolution, the size of the convolutional kernel is adjusted according to the dilation rate; for example, when the dilation rate is d, d - 1 zeros are inserted between each element of the convolutional kernel, thus expanding the originally k×k convolutional kernel to a larger coverage range; compared with the traditional convolution, the dilated convolution can expand the receptive field without reducing the resolution of the feature map, which is very important for maintaining the detailed information of the image; in the dilated convolution, the actual coverage area of the convolutional kernel increases according to the dilation rate, which means that even if the physical size of the convolutional kernel remains unchanged, the range of information it can effectively capture will increase significantly.

[0064] The S3 in the CSPPFS module (with the structure shown in Figure 5 a), its input features first pass through the CBAM module (with the structure shown in Figure 5 b), and then through the SimAM module (with the structure shown in Figure 5 c). The SimAM module enhances the important features by adaptively adjusting the channel weights of the feature map. The CBAM module enhances the feature representation through explicit channel and spatial attention mechanisms, while the SimAM module adaptively captures the saliency information in the feature map in a parameter-free manner. The combination of the two enables the model to better extract and utilize features at different levels and dimensions, complementing each other; specifically, the CBAM module globally enhances the features at a higher level, while the SimAM module locally enhances the features at a lower level. This multi-level feature enhancement strategy enables the model to better capture information at different scales and improve the performance of the model.

[0065] Among them, the CBAM module consists of a spatial attention module (SAM) and a channel attention module (CAM); the feature map will pass through parallel max-pooling and average-pooling layers, reducing the size of the feature map from the original C×H×W ( C is the number of channels, H is the height, W is the width), to C×1×1 . These two pooling operations do not change the size of the feature map, but generate two different feature representations, one is the max-pooling result that emphasizes the most significant features, and the other is the average-pooling result that emphasizes the average features; these two pooling results will then pass through a shared multi-layer perceptron (MLP) module; in this module, the number of channels of the feature map is first compressed to times the original, where R is the reduction rate, and then expanded back to the original number of channels; after passing through the MLP module, two activated feature maps will be obtained; these two feature maps will be merged by element-wise addition; finally, the merged feature map will pass through a Sigmoid activation function to generate channel attention weights; these channel attention weights will then be multiplied element-wise with the original feature map to adjust the weights of each channel, enabling the algorithm to pay more attention to important features;

[0066] The SimAM module first obtains the features of each channel through global average pooling, and then performs interactions between channels through a 1D convolution; this operation method combines the information of the spatial dimension and the channel dimension to generate three-dimensional attention weights; finally, the Sigmoid activation function is used to generate channel weights and multiply them with the original feature map to re-weight the features; this process realizes the adaptive adjustment of features and improves the expression ability of the network;

[0067] The calculation formula for its weight is:

[0068]

[0069] Among them, y represents the input feature, is the dot product operation, λ is a hyperparameter, C is the number of channels, t and x j are the target neuron and the adjacent neuron respectively; the Sigmoid activation function maps the input value to between 0 and 1, which helps to control the output range of the neuron.

[0070] The SSFFE module in S4 (the structure is as Figure 6 shown) consists of feature size adjustment, dynamic feature fusion, and ECA channel attention mechanism weight calibration;

[0071] Regarding the differences in the number of channels and spatial resolution of feature maps in each layer of the YOLOv8 algorithm, in the feature size adjustment, the size of the feature map is unified through upsampling or downsampling techniques; during upsampling, a convolutional layer is used to adjust the number of channels, and the resolution is enhanced through interpolation techniques; for 2x downsampling, SDconv with a stride of 2 is used to simultaneously adjust the number of channels and resolution; for 4x downsampling, first, dimensionality reduction is performed through a max pooling layer with a stride of 2, and then SDconv with a stride of 2 is further applied for adjustment; in SDconv, when the dilation rate is set to 3, the size of the convolutional kernel increases, thereby obtaining a broader receptive field, which helps to capture image content more comprehensively during the feature fusion stage;

[0072] In the dynamic feature fusion, SSFFE integrates the feature information of the first three levels after size adjustment, and its information is shared among all channels to promote the effective integration of features;

[0073] In the ECA channel attention mechanism weight calibration, the ECA channel attention mechanism first performs global average pooling on each channel of the input feature map, compressing the spatial dimension to 1×1 to generate a channel descriptor containing global information. Then, the ECA channel attention mechanism automatically determines the kernel size of the one-dimensional convolution according to the number of channels and calculates through a non-linear mapping, ensuring that the kernel size is odd to maintain the central symmetry of the convolution operation. The ECA channel attention mechanism uses one-dimensional convolution to process the result of global average pooling to identify the inter-channel relationships. This convolution method avoids the need for dimensionality reduction and increase, reduces the number of parameters, and at the same time retains the integrity of information. The output of the one-dimensional convolution passes through the Sigmoid activation function to generate attention weights for each channel. The attention weights reflect the importance of each channel. Finally, the attention weights are multiplied element-wise with the corresponding channels of the original feature map to achieve the weighting of channel features, highlighting key features and suppressing less important information.

[0074] The calculation formula of the ASFF is as follows:

[0075]

[0076] where, represents the feature vector at the (x, y) position of the output feature map of the nth layer, α, β, and γ represent the spatial weights from the mth layer (m ∈ 1, 2, 3) to the nth layer feature map, and A, B, and C respectively represent the feature vectors at the (x, y) position of different feature maps from the mth layer to the nth layer; the ASFF integrates the effective information of each layer of features into different scales by weighted summing the feature vectors with different characteristics according to the corresponding spatial weights, and finally obtains the fused feature F.

[0077] The ECA channel attention mechanism (the structure is as shown in Figure 7 ) realizes the interaction between channels through global average pooling and one-dimensional convolution, avoiding the use of complex fully connected layers, thus greatly reducing the computational complexity; compared with the traditional attention mechanism, it uses a simple and efficient one-dimensional convolution to capture the dependencies between channels, enhances feature aggregation, and effectively improves the computational efficiency; in the detection of foreign objects on transmission lines, small objects may only occupy a few pixels on the feature map. Therefore, adding the ECA channel attention mechanism in front of the small detection head can effectively capture long-range dependencies through local cross-channel interaction, thereby enhancing the detection accuracy of the algorithm for small targets;

[0078] Facing the problem of object scale variation in the detection of foreign objects on transmission lines, the ECA channel attention mechanism helps to better fuse multi-scale features from different levels; the ECA channel attention mechanism first performs global average pooling on the input feature map, compresses the spatial dimension to 1x1, then captures the local cross-channel dependencies through one-dimensional convolution, and then activates the result through the Sigmoid activation function to enhance the non-linear features, obtaining the attention weight for each channel. Finally, the obtained attention weight matrix is reshaped into the same shape as the original input feature map and multiplied element-wise with the input feature map to achieve the recalibration of channel features;

[0079] The calculation formula of the ECA channel attention mechanism is:

[0080]

[0081] where k is the size of the one-dimensional convolution kernel, b and γ are hyperparameters, set to 1, C is the number of channels of the input feature, and |t| odd represents the odd number closest to t, and its formula ensures that the size of the convolution kernel is odd.

[0082] In S5, the specific algorithm training is as follows:

[0083] The experiment was carried out under the Windows 10 operating system, using a device configured with an Intel(R) Core(TM) i9-10900K and equipped with an NVIDIA GeForce RTX 3080, and using PyTorch version 2.3.0 as the main development tool; the specific hyperparameters were set as follows: the batch size was set to 8, the number of training epochs was set to 100, and the Stochastic Gradient Descent (SGD) optimizer was selected for algorithm training, where the initial learning rate was set to 0.01 and the weight decay was set to 0.0005; in addition, all input images were uniformly scaled to a size of 640×640;

[0084] The specific steps to obtain the algorithm detection results in S5 are as follows:

[0085] To comprehensively evaluate the performance of the SCSM-YOLOv8 algorithm, multiple evaluation metrics are selected to ensure the objectivity and accuracy of the evaluation; these metrics include: Precision (P), Recall (R), F1-score, Average Precision (AP), Mean Average Precision (mAP), and Frames Per Second (FPS) as evaluation metrics to detect the network performance. The specific calculation formulas are as follows:

[0086]

[0087] Among them, TP is that the algorithm correctly predicts positive-class samples as positive classes; TN is that the algorithm correctly predicts negative-class samples as negative classes; FP is that the algorithm wrongly predicts negative-class samples as positive classes; FN is that the algorithm wrongly predicts positive-class samples as negative classes;

[0088] AP measures the local performance of the learned algorithm in each category, and mAP measures the overall performance of the learned algorithm in all categories. The specific formulas are as follows:

[0089]

[0090] Among them, n is the number of images in the test set; i is the i-th image in the test set; p i is the accuracy of a certain category of the i-th image; N is the number of target categories; k represents the k-th category; AP k represents the average accuracy of the k-th category on all images in the test set;

[0091] In S5, the improved algorithm is compared with different algorithms, specifically:

[0092] To accurately verify the performance of the SCSM-YOLOv8 algorithm in the task of detecting foreign objects on transmission lines, it is compared with the Faster R-CNN, YOLOv5, YOLOX, YOLOv7, and YOLOv8 algorithms. The above algorithms are trained and tested under the same conditions. The final experimental results are shown in Table 1.

[0093] Table 1 shows the performance of the algorithm in six dimensions: mAP, Recall, Precision, F1, FPS, and the number of parameters. SCSM-YOLOv8 achieved the highest mAP, Recall, Precision, and F1. Compared with the second place, it is 2.56%, 3.47%, 3.14%, and 3% higher respectively. This indicates that SCSM-YOLOv8 has significant performance advantages. This is mainly attributed to the ability of SCSM-YOLOv8 to obtain more feature information and efficiently fuse multi-scale information. Although the number of parameters is relatively large, its performance in terms of frames per second is still competitive, especially when compared with YOLOv7l and YOLOv8s. This shows that the improved SCSM-YOLOv8 algorithm significantly improves the detection accuracy while maintaining a high detection speed, enhances the algorithm's feature extraction ability, and is an algorithm that achieves a good balance between performance and efficiency.

[0094] Table 1 Performance comparison of SCSM-YOLOv8 and other methods on the foreign object dataset of transmission lines

[0095]

[0096] Note: Bold indicates the optimal value; underlined data indicates the sub-optimal value; - indicates no such data.

[0097] Although the present invention has been described above with reference to the embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the embodiments disclosed in the present invention can be combined with each other in any way, and the cases of these combinations are not exhaustively described in this specification only for the sake of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A foreign object detection algorithm for transmission lines based on SCSM-YOLOv8, characterized in that, The specific steps are as follows: S1: Create a custom dataset for detecting foreign objects on transmission lines, and divide it into three parts by random sampling: training set, validation set and test set, with a ratio of 70%, 20% and 10%; S2: In the backbone structure, design the SDM module, and use the designed switchable dilated convolution SDconv to replace the standard convolution of the original C2f to expand the receptive field and improve the feature extraction ability of the algorithm; S3: In the backbone structure, design the CSPPFS module, introduce the CBAM attention mechanism and the SimAM attention mechanism in the SPPF, and improve the object detection performance without increasing the number of parameters and the amount of computation; S4: In the neck structure, design the SSFFE module, introduce the SDconv and ECA channel attention mechanisms in the ASFF, and strengthen the feature fusion part of the neck network in the YOLOv8 algorithm; S5: Input the dataset into the algorithm for iterative training. After the training is completed, use the optimal algorithm to detect the test set to obtain the final result.

2. The foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 according to claim 1, characterized in that, The specific steps of S1 are as follows: S11: Create a custom dataset specifically for detecting foreign objects on transmission lines, and collect images of transmission lines from variable environmental conditions and diverse perspectives to form an initial dataset; S12: In the collected images, exclude images with low quality and unclear clarity, and only retain two-dimensional images that are clearly taken and not affected by compression to ensure the clarity and accuracy of the data; S13: After screening, 2157 images containing daily foreign objects are obtained, and their images form a basic dataset. To improve the generalization ability of the algorithm, various data augmentation techniques including rotation, translation, cropping, flipping, and adjustments of brightness, grayscale and contrast are implemented on the images. Through the data augmentation techniques, the number of images increases to 6382, and a custom dataset for detecting foreign objects on transmission lines is obtained, which contains different types of foreign objects; S14: Before algorithm training, randomly divide the expanded dataset and allocate it to the training set, validation set and test set according to a ratio of 7:2:

1. After division, the training set contains 4467 images for algorithm training, the validation set contains 1276 images for algorithm performance verification, and the test set contains 639 images for evaluating the final performance of the algorithm.

3. A foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 according to claim 1, characterized in that, The switchable dilated convolution SDconv in S2 includes a preposed multi-channel global context module, an SDconv unit, and a postposed multi-channel global context module; The steps of the pre - placed multi - channel global context module are as follows: Receive the input feature map and process it through two pooling layers with different field - of - view ranges; the two pooling layers can capture global context information at different scales; after being processed by the pooling layers, the feature map is further compressed by a 1×1 convolutional layer to reduce the number of channels to half of the original, so as to reduce the computational burden of the algorithm and extract more refined features at the same time; the outputs of the two pooling layers are fused in the channel dimension to integrate context information at different scales and prepare for the further processing of the SDconv unit. The steps of the SDconv unit are as follows: The input feature is processed through two dilated convolutional layers with different dilation rates to capture feature information at different scales; the two dilated convolutional layers use dilation rates d1 and d2 respectively to extract multi - scale features. A switching function is used to dynamically adjust the output weights of the two dilated convolutional layers. The switching function consists of a 5×5 average pooling layer and a 1×1 convolutional layer, which is used to generate a set of weights to adjust the proportion of the contributions of the two convolutional layers; the output of the SDconv unit is the weighted sum of the outputs of the two dilated convolutional layers, which helps the algorithm to perform feature fusion at different scales. The steps of the post - placed multi - channel global context module are as follows: After the SDconv unit, the post - placed multi - channel global context module is used again to further refine the features and provide rich context information for the subsequent layers of the network, enhancing the expressive ability of the features.

4. The foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 according to claim 3, wherein, The steps of the SDM module are as follows: The input feature first passes through an SDconv and then enters an SDbottleneck. At the same time, the input feature also passes through a traditional convolutional layer; the output of the SDbottleneck is concatenated with the output of the standard convolutional layer in the channel dimension, and then feature integration is performed through another SDconv, and finally it is used as the output of the SDM module. The steps of the SDbottleneck are as follows: The input feature is first processed through an SDconv and then passes through another SDconv. The output of the first SDconv is added element - by - element to the output of the second SDconv to form a residual connection, which helps to improve the gradient flow and prevent the problem of gradient disappearance. The feature map after the addition is used as the final output of the SDbottleneck.

5. A foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 according to claim 1, characterized in that, The above-mentioned S3 is in the CSPPFS module. Its input features first pass through the CBAM module and then through the SimAM module. The SimAM module enhances important features by adaptively adjusting the channel weights of the feature map. The CBAM module enhances feature representation through explicit channel and spatial attention mechanisms. The SimAM module adaptively captures the saliency information in the feature map in a parameter-free manner. The combination of the two enables the model to better extract and utilize features at different levels and dimensions, complementing each other. Specifically, the CBAM module globally enhances features at a higher level, while the SimAM module locally enhances features at a lower level. This multi-level feature enhancement strategy enables the model to better capture information at different scales and improve the performance of the model.

6. The foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 according to claim 1, wherein In the above-mentioned S4, the SSFFE module consists of feature size adjustment, dynamic feature fusion, and ECA channel attention mechanism weight calibration. In the feature size adjustment, the size of the feature map is unified through upsampling or downsampling techniques. During upsampling, a convolutional layer is used to adjust the number of channels, and the resolution is increased through interpolation techniques. For 2-fold downsampling, an SDconv with a stride of 2 is used to simultaneously adjust the number of channels and the resolution. For 4-fold downsampling, first, a max pooling layer with a stride of 2 is used for dimensionality reduction, and then an SDconv with a stride of 2 is applied for further adjustment. In the SDconv, when the dilation rate is set to 3, the size of the convolutional kernel increases, thereby obtaining a wider receptive field, which helps to more comprehensively capture the image content during the feature fusion stage. In the dynamic feature fusion, the SSFFE integrates the feature information of the first three levels after size adjustment, and the information is shared among all channels to promote the effective integration of features. In the ECA channel attention mechanism weight calibration, the ECA channel attention mechanism first performs global average pooling on each channel of the input feature map, compresses the spatial dimension to 1×1, and generates a channel descriptor containing global information. Then, the ECA channel attention mechanism automatically determines the kernel size of the one-dimensional convolution according to the number of channels and calculates it through a non-linear mapping, ensuring that the kernel size is odd to maintain the central symmetry of the convolution operation. The ECA channel attention mechanism uses one-dimensional convolution to process the result of global average pooling to identify the inter-channel relationships. This convolution method avoids the need for dimensionality reduction and increase, reduces the number of parameters, and at the same time retains the integrity of the information. The output of the one-dimensional convolution passes through the Sigmoid activation function to generate attention weights for each channel. The attention weights reflect the importance of each channel. Finally, the attention weights are multiplied element-wise with the corresponding channels of the original feature map to achieve weighted channel features, highlighting key features and suppressing less important information.

7. An algorithm for detecting foreign objects on transmission lines based on SCSM-YOLOv8 according to claim 6, characterized in that, The calculation formula of the above-mentioned ASFF is as follows: Among them, represents the feature vector of the output feature map of the nth layer at the (x, y) position. α, β, and γ represent the spatial weights from the feature map of the mth layer (m ∈ 1, 2, 3) to the nth layer feature map. A, B, and C respectively represent the feature vectors of different feature maps at the (x, y) position from the mth layer to the nth layer. The ASFF obtains the fused feature F by weighted summing the feature vectors with different characteristics according to the corresponding spatial weights, thereby integrating the effective information of each layer of features into different scales.

8. A foreign object detection algorithm for transmission lines based on SCSM-YOLOv8 according to claim 6, characterized in that, The ECA channel attention mechanism first performs global average pooling on the input feature map to compress the spatial dimension to 1x1, then captures local cross-channel dependencies through one-dimensional convolution, and then activates the result through the Sigmoid activation function to enhance non-linear features, obtaining the attention weight for each channel. Finally, the obtained attention weight matrix is reshaped into the same shape as the original input feature map, and element-wise multiplication is performed with the input feature map to achieve recalibration of channel features; The calculation formula of the ECA channel attention mechanism is as follows: where k is the size of the one-dimensional convolutional kernel, b and γ are hyperparameters set to 1, C is the number of channels of the input feature, and |t| odd represents the odd number closest to t, and its formula ensures that the convolutional kernel size is odd.

Citation Information

Patent Citations

  • X-ray dangerous object detection method

    CN116758376A

  • PCBA surface defect detection method and system based on image recognition

    CN119273682A

  • ATP-YOLOv8-based power transmission line channel hidden danger remote sensing image target detection algorithm

    CN119723060A