Power transmission line small target detection method based on improved YOLOv5s
By improving the YOLOv5s network, sub-pixel convolution and multi-head self-attention module are adopted, combined with shift window attention and CIoU_Loss loss function, the problem of insufficient feature expression and missed detection in small object detection of transmission lines is solved, and higher accuracy and robust detection are achieved.
Patent Information
- Application Number
- CN202510394563.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-22
AI Technical Summary
The existing small-objective detection algorithms have problems such as insufficient feature expression capabilities, susceptibility to noise, high false detection and missed detection rates in power transmission line inspections, especially in complex backgrounds and long-distance situations, which are difficult to effectively model the feature information of small-objectives.
The improved YOLOv5s network is adopted to downsample through sub-pixel convolution inverse operation, combining multi-head self-attention and shift window attention, enhance the feature expression and modeling capabilities of small targets, and introduce a bidirectional feature fusion strategy and CIoU_Loss loss function to optimize the target box positioning.
It improves the accuracy and robustness of small-object detection, reduces false detection and missed detection, enhances the model's detection capabilities in complex backgrounds, and is suitable for UAV inspection and power equipment monitoring.
Smart Images

Figure CN120355893A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a small target detection method for transmission lines based on improved YOLOv5s. Background Art
[0002] The inspection and maintenance of transmission lines are crucial for the stable operation of the power system. Traditional inspection methods for transmission lines mainly rely on manual inspection or ground monitoring equipment, which have problems such as low detection efficiency, high labor costs, and poor environmental adaptability. In recent years, unmanned aerial vehicle (UAV) inspection technology has gradually become an important means of intelligent inspection of transmission lines, enabling high-altitude shooting and covering a large area. However, during UAV inspection, due to factors such as shooting angle, lighting conditions, and target size, the detection of small targets (such as insulators, wire clips, fittings, etc.) on transmission lines still faces great challenges.
[0003] Existing small target detection algorithms, such as Faster R-CNN, YOLO (You Only Look Once) series, and SSD (Single Shot MultiBox Detector), etc., although they have achieved good results in general target detection tasks, still have certain deficiencies in the scenario of small target detection for transmission lines. First, since small targets account for a relatively small proportion in the image and have insufficient feature expression ability, it is difficult for the detection model to effectively learn the feature information of small targets. Second, when existing detection algorithms process long-distance or complex background interference, they are easily affected by noise, resulting in high false detection and missed detection rates. In addition, although existing lightweight target detection networks such as YOLOv5s have advantages in computing efficiency and detection speed, there is still much room for improvement in small target detection ability.
[0004] In the prior art, Chinese Patent CN116843649B discloses a method for intelligent defect detection of transmission lines based on an improved YOLOv5 network, including the following steps: improving the YOLOv5 network to construct a target detection network, the target detection network includes a backbone network, a path fusion network, and an output module, using a transmission line database to construct a data set to train the target detection network, and saving the trained target detection network; inputting the image of the transmission line to be detected into the trained target detection network, and outputting the corresponding target detection result. However, this patent lacks special detection optimization for small targets, resulting in weak feature expression ability of small targets, easy occurrence of missed detection or false detection, inability to effectively model the global dependence between long-distance pixels, especially in complex backgrounds and target occlusion situations, it is easy to cause loss or confusion of target features, and traditional target box regression mainly relies on IoU loss or GIoU loss, which may lead to a decrease in positioning accuracy in small target detection tasks. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a small target detection method for transmission lines based on improved YOLOv5s.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] On the one hand, the present invention provides a small target detection method for transmission lines based on improved YOLOv5s, including the following steps:
[0008] Construct a small target detection model, train the small target detection model, and detect small targets on the transmission line through the trained small target detection model; wherein the small target detection model includes an input end, a backbone network, a connection layer, and a prediction layer connected in sequence. The input end uses the Mosaic method for data augmentation, randomly scales, crops, and arranges the input image, and splices it into a new input image; the backbone network includes a sampling module, multiple groups of convolutional modules, residual modules, convolutional modules, and pooling modules connected in sequence; the connection layer includes multiple attention modules, convolutional modules, multiple groups of upsampling modules, and fusion modules, multiple groups of residual modules, and convolutional attention modules, upsampling modules, and four output channels connected in sequence; the prediction layer includes multiple output modules, and the output modules are connected to the output channels at the end of the connection layer.
[0009] Further, the four output channels of the connection layer include a first output channel, a second output channel, a third output channel, and a fourth output channel. The output channel includes a fusion module, a shift attention module, and a convolutional attention module connected in sequence. The output ends of the multiple attention modules and convolutional modules in the connection layer are connected to the fusion module of the fourth output channel. The output ends of the first group of residual modules and convolutional attention modules in the connection layer are connected to the fusion module of the third output channel. The output ends of the second group of residual modules and convolutional attention modules in the connection layer are connected to the fusion module of the second output channel. The output end of the last upsampling module in the connection layer is connected to the fusion module of the first output channel. The first output channel, the second output channel, the third output channel, and the fourth output channel are connected in sequence.
[0010] Further, the sampling module of the backbone network downsamples the input image through the inverse operation of sub-pixel convolution, specifically including:
[0011] The input image is divided into four parts. After each part of the data is downsampled by 2 times, it is concatenated in the channel dimension and a convolution operation is performed, and the number of output channels is one-fourth of the number of input channels.
[0012] Further, the output ends of the first group of convolutional modules and residual modules of the backbone network are connected to the fusion module of the first output channel of the connection layer, and the output end of the pooling module of the backbone network is connected to the input end of the multi-attention module of the connection layer.
[0013] Further, detecting small targets on the transmission line by the trained small target detection model specifically includes:
[0014] After an input image is input at the input end, it is processed by a sampling module, convolutional modules, and residual modules to obtain a first feature map and a first calculation graph. After the first feature map passes through two groups of convolutional modules and residual modules of the backbone network, it is then input into a convolutional module and a pooling module and outputs a second feature map. The second feature map is input into the multi-attention module and convolutional module of the connection layer;
[0015] The first calculation graph is input into the fusion module of the first output channel of the connection layer;
[0016] The second feature map is processed by the multi-attention module and convolutional module and outputs a second calculation graph. The second calculation graph is respectively input into the first group of upsampling modules and fusion module of the connection layer, the first group of residual modules, and the convolutional attention module to output a third calculation graph. The third calculation graph outputs a fourth calculation graph through the second group of upsampling modules and fusion module, the second group of residual modules, and the convolutional attention module. The fourth calculation graph is processed by the last upsampling module of the connection layer to output a fifth calculation graph;
[0017] After the fifth calculation graph is fused with the first calculation graph through the fusion module of the first output channel of the connection layer, it is input into the shift attention module of the first output channel. After being processed by the shift attention module, it outputs a first channel graph. The first channel graph is input into the first output module of the prediction layer, and the first output module of the prediction layer outputs a first prediction result;
[0018] After the first channel graph is processed by the convolutional attention module of the first output channel, it is fused with the fourth calculation graph through the fusion module of the second output channel, and then output a second channel graph after being processed by the shift attention module of the second output channel. The second channel graph is input into the second output module of the prediction layer, and the second output module of the prediction layer outputs a second prediction result;
[0019] After the second channel graph is processed by the convolutional attention module of the second output channel, it is fused with the third calculation graph through the fusion module of the third output channel, and then output a third channel graph after being processed by the shift attention module of the third output channel. The third channel graph is input into the third output module of the prediction layer, and the third output module of the prediction layer outputs a third prediction result;
[0020] After the third-channel graph is processed by the convolutional attention module of the third output channel, it is fused with the second computational graph through the fusion module of the fourth output channel, and then processed by the shift attention module of the fourth output channel to output the fourth-channel graph. The fourth-channel graph is input into the fourth output module of the prediction layer, and the fourth prediction result is output by the fourth output module of the prediction layer.
[0021] Further, the multi-attention module includes a multi-head attention module and a multi-layer perceptron module, and each sub-layer is connected by a residual block.
[0022] Further, the processing process of the multi-attention module includes the following steps:
[0023] The input image is divided into N image patches, and the resolution of each image patch is P×P, where N = HW÷P^2, H and W are the height and width of the input image respectively, and C is the number of channels of the input image;
[0024] Each image patch is reorganized into a vector of P^2×C dimensions, and each vector is mapped into a fixed-length vector of dimension D through linear projection to generate an embedding layer. The initialization formula of the embedding layer is:
[0025]
[0026] Among them, Z0 represents the initialization vector of the embedding layer, x class represents the image patch matrix of different lengths, represents the corresponding image patch, N represents the number of image patches, E represents the dimension of different-length image patches, E pos represents the position of different-length image patches, R represents real numbers, P represents the image patch resolution, C represents the number of channels, and D represents the dimension of the fixed-length vector;
[0027] According to the density threshold ρ t of the image patches, the splicing quantity P of the image patches is dynamically adjusted 2 , and the calculation formula is:
[0028]
[0029] Among them, P 2 represents the splicing quantity of the image patches, ρ t represents the density threshold of the image patches, which is used to determine the number of spliced image patches;
[0030] The image patch vectors after adjusting the splicing quantity are spliced to obtain a spliced vector matrix. In the spliced vector matrix, the position component representing the image position is added to the slice component to obtain a picture slice component with position information. The following operations are sequentially performed on the picture slice component with position information:
[0031] For the layer, first perform LayerNorm normalization on the output result of the previous layer , then input the normalized result into the multi-head self-attention module MSA for processing, and add the output result of MSA to residually to obtain the multi-head self-attention residual output of the layer Perform LayerNorm normalization on , then input it into the multi-layer perceptron module MLP, and add the output result of MLP to residually to obtain the multi-layer perceptron residual output of the layer Perform LayerNorm normalization on the multi-layer perceptron residual output z of the last layer L to obtain the final output result y, and the formula is:
[0032]
[0033]
[0034] y = LN(z L )
[0035] where is the multi-layer perceptron residual output of the layer, is the multi-head self-attention residual output of the layer, LN is the LayerNorm normalization operation, z L is the multi-layer perceptron residual output of the L-th layer, L is the total number of layers of the multi-attention module, MSA represents the processing of the multi-head self-attention module, and MLP represents the processing of the multi-layer perceptron module.
[0036] Furthermore, the shifted attention module includes a multi-head self-attention module W-MSA based on moving windows and a multi-layer perceptron module MLP. Each sub-layer is regularized by a normalization layer LayerNorm and uses the GELU non-linear activation function.
[0037] Furthermore, the processing process of the shifted attention module includes the following steps:
[0038] Divide the input image into a set of non-overlapping image patches, each with a size of 4×4, a feature dimension of 4×4×3 = 48, and the number of image patches is 160×160;
[0039] Apply linear embedding to the original feature layer to map the feature dimension of the divided image patches to an arbitrary dimension C;
[0040] Adjacent image patches within a 2×2 range are merged through image patch stitching to generate hierarchical representations of feature layers at different scales.
[0041] For the features processed as above, self-attention calculation is performed on them using the multi-head self-attention module W-MSA based on a moving window, and the multi-layer perceptron module MLP processes the output of the multi-head self-attention module to obtain the final output features of this layer.
[0042] Furthermore, the loss function of the small object detection model is:
[0043]
[0044] Among them, CIoU Loss is the loss function of the small object detection model, IoU represents the intersection over union, which is used to measure the overlapping degree of the predicted box and the ground truth box; Box pre represents the predicted box; Box gt represents the ground truth object box; Intersection(Box pre , Box gt ) represents the intersection area of the predicted box and the ground truth box; Union(Box pre , Box gt ) represents the union area of the predicted box and the ground truth box; represents the square of the diagonal distance of the smallest enclosing rectangle box, represents the square of the Euclidean distance between the center points of the predicted box and the object box, ν is a parameter measuring the aspect ratio consistency, W gt and h gt respectively represent the width and height of the ground truth object box, W p and h p respectively represent the width and height of the ground truth object box.
[0045] Compared with the prior art, the present invention has the following advantages:
[0046] (1) In the backbone network of the present invention, the inverse operation of sub-pixel convolution is adopted for downsampling. After dividing the input image into multiple sub-regions, convolution and channel stitching are performed, thereby reducing the loss of feature information and improving the feature expression ability for small objects. This technical means enhances the discrimination ability of the model for small objects, avoids the information loss problem caused by traditional downsampling, and thus improves the detection accuracy of small objects.
[0047] (2) By combining multi-head self-attention (MSA) and shifted window attention (W-MSA), the present invention improves the model's ability to model distant targets and enhances the learning ability of small target features in complex backgrounds. Compared with the prior art that only uses the CBAM module for local attention calculation, the multi-layer attention mechanism of the present invention can perform global information modeling on feature maps of different scales, enabling the model to more accurately focus on small target features and improve detection robustness.
[0048] (3) The present invention adopts a bidirectional feature fusion strategy and an adaptive weighted fusion module in the connection layer. Through multiple upsamplings and residual connections, the expression ability of features of different scales is enhanced, enabling high-level semantic information to be effectively transmitted to the low level and improving the perception ability of small targets. At the same time, the weighting mechanism of the fusion module can dynamically adjust the contribution ratio of features of different layers, making feature fusion more balanced and improving the overall detection effect.
[0049] (4) The present invention adopts a dynamic image block division strategy, dynamically adjusts the block size according to the target density, and combines the dynamically adjusted image block resolution, enabling the detection model to adapt to small targets of different sizes and improving the generalization ability of detection. This strategy effectively solves the problem of insufficient adaptability to target sizes in the prior art, enabling the model to more accurately locate small targets of different scales and improve detection accuracy.
[0050] (5) The present invention introduces the CIoU_Loss loss function in the target box regression process. While optimizing the target box positioning, it comprehensively considers IoU, the distance between the target center points, and the aspect ratio consistency, enabling the model to more accurately fit the true boundary of small targets during the optimization process and improving the precision and stability of small target detection. Compared with traditional IoU or GIoU losses, the optimization strategy of the present invention enables the model to more stably handle small target detection tasks and reduce the occurrence of missed detections and false detections. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is the network structure diagram of the small target detection model according to an embodiment of the present application;
[0052] Figure 2 is the schematic diagram of the multi-attention module structure according to an embodiment of the present application;
[0053] Figure 3 is the schematic diagram of the convolutional block attention module CBAM structure according to an embodiment of the present application;
[0054] Figure 4 is the comparison of the FPN, PANet, and BiFPN structures according to an embodiment of the present application;
[0055] Figure 5It is the overall structure diagram of the shift attention module according to an embodiment of the present application;
[0056] Figure 6 It is the structure diagram of the shift attention module according to an embodiment of the present application;
[0057] Figure 7 It is the schematic diagram of the principle of image segmentation based on SW-MSA according to an embodiment of the present application;
[0058] Figure 8 It is the schematic diagram of the confusion matrix of the object detection model with the best detection performance according to an embodiment of the present application;
[0059] Figure 9 It is the P-R curve graph of the object detection model with the best detection performance according to an embodiment of the present application;
[0060] Figure 10 It is the schematic diagram of the test results of the object detection model with the best performance according to an embodiment of the present application. Specific implementation manners
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] Embodiment 1:
[0063] This embodiment provides a method for detecting small targets on transmission lines based on improved YOLOv5s, including the following steps:
[0064] Construct a small target detection model, train the small target detection model, and detect small targets on the transmission line through the trained small target detection model;
[0065] As Figure 1 shown, the small target detection model includes an input end, a backbone network, a connection layer, and a prediction layer connected in sequence. The input end uses the Mosaic method for data augmentation, randomly scales, crops, and arranges the input pictures, and splices them into a new input picture; the backbone network includes a sampling module, multiple groups of convolutional modules, a residual module, a convolutional module, and a pooling module connected in sequence; the connection layer includes a multi-attention module, a convolutional module, multiple groups of upsampling modules, a fusion module, multiple groups of residual modules, a convolutional attention module, an upsampling module, and four output channels connected in sequence; the prediction layer includes multiple output modules, and the output modules are connected to the output channels at the end of the connection layer.
[0066] In the embodiment, the architecture improvement of the small target detection model has obvious improvement in the detection performance, especially in the small target detection performance, and the real-time performance can also meet the requirements of real-time detection and path planning. In the task of UAV inspection of transmission lines, it not only significantly improves the performance of the target detection model, but also can be extended to other applications involving small target detection or long-distance target detection. The detection ability and perception robustness of small targets on transmission lines have been better improved, and it has a positive impact on the detection of all-size targets on transmission lines, so as to provide better planning and decision-making strategies for intelligent unmanned inspection of transmission lines and defect identification of transmission lines.
[0067] The small target detection model uses the Mosaic method for data augmentation at the input end, randomly scales, crops and arranges the pictures, and stitches them into a new input picture. Through the Mosaic method, medium and large targets are randomly scaled into small targets, the data distribution of small targets is balanced, the robustness of the network is stronger, and the training effect for small target detection is also better. From the construction of the 500KV transmission line inspection image dataset used in this application, it can be found that there are a large number of small target instances, and most of these small targets are concentrated in two categories: wire clamps and fittings, and the proportion of small targets in the two categories of glass insulators and composite insulators is relatively small.
[0068] Further, the four output channels of the connection layer include a first output channel, a second output channel, a third output channel, and a fourth output channel. The output channel includes a fusion module, a shift attention module, and a convolutional attention module connected in sequence. The multi-attention module and the convolutional module output ends of the connection layer are connected to the fusion module of the fourth output channel. The first group of residual modules and the convolutional attention module output ends of the connection layer are connected to the fusion module of the third output channel. The second group of residual modules and the convolutional attention module output ends of the connection layer are connected to the fusion module of the second output channel. The output end of the last upsampling module of the connection layer is connected to the fusion module of the first output channel. The first output channel, the second output channel, the third output channel, and the fourth output channel are connected in sequence.
[0069] Further, the sampling module of the backbone network downsamples the input image through the inverse operation of sub-pixel convolution, specifically including:
[0070] The input image is divided into four parts. After each part of the data is downsampled by 2 times, it is stitched in the channel dimension and a convolution operation is performed. The number of output channels is one-fourth of the number of input channels.
[0071] Furthermore, the pooling module uses multiple sliding windows to perform pooling on images of different input sizes to obtain pooled features of the same size. The pooling module significantly improves the scale invariance of the images and effectively reduces the overfitting situation. The pooling module enables the detection network to converge more easily and has less impact on the network complexity. In the implementation process of the pooling module in the small object detection model image detection network, the whole picture is first convolved, and then the target window is pooled, and the obtained result is used as the input of the fully connected layer. The pooling method using the pooling module can more effectively increase the receptive field of the backbone features than simply using the max pooling method, and can separate important context features to the greatest extent. In the pooling module, the max pooling method with three scales of k = {1*1, 5*5, 9*9, 13*13} is used, and after obtaining three feature maps of different scales, a tensor concatenation operation is performed.
[0072] Furthermore, the output ends of the first group of convolutional modules and residual modules of the backbone network are connected to the fusion module of the first output channel of the connection layer, and the output end of the pooling module of the backbone network is connected to the input end of the multi-attention module of the connection layer.
[0073] Furthermore, detecting small objects on the transmission line by using the trained small object detection model specifically includes:
[0074] After an image is input at the input end, it is processed by the sampling module, convolutional module and residual module to obtain a first feature map and a first calculation graph. After the first feature map passes through two groups of convolutional modules and residual modules of the backbone network, it is then input into a convolutional module and a pooling module and then outputs a second feature map. The second feature map is input into the multi-attention module and convolutional module of the connection layer;
[0075] The first calculation graph is input into the fusion module of the first output channel of the connection layer;
[0076] The second feature map is processed by the multi-attention module and convolutional module and then outputs a second calculation graph. The second calculation graph is respectively input into the first group of upsampling modules and fusion module of the connection layer, the first group of residual modules and convolutional attention module to output a third calculation graph. The third calculation graph passes through the second group of upsampling modules and fusion module, the second group of residual modules and convolutional attention module to output a fourth calculation graph. The fourth calculation graph passes through the last upsampling module of the connection layer to process and output a fifth calculation graph;
[0077] After the fifth calculation graph is fused with the first calculation graph through the fusion module of the first output channel of the connection layer, it is input into the shift attention module of the first output channel. After being processed by the shift attention module, it outputs a first channel graph. The first channel graph is input into the first output module of the prediction layer, and the first prediction result is output by the first output module of the prediction layer;
[0078] After the first-channel graph is processed by the convolutional attention module of the first output channel, it is fused with the fourth computational graph through the fusion module of the second output channel, and then processed by the shift attention module of the second output channel to output the second-channel graph. The second-channel graph is input to the second output module of the prediction layer, and the second prediction result is output by the second output module of the prediction layer;
[0079] After the second-channel graph is processed by the convolutional attention module of the second output channel, it is fused with the third computational graph through the fusion module of the third output channel, and then processed by the shift attention module of the third output channel to output the third-channel graph. The third-channel graph is input to the third output module of the prediction layer, and the third prediction result is output by the third output module of the prediction layer;
[0080] After the third-channel graph is processed by the convolutional attention module of the third output channel, it is fused with the second computational graph through the fusion module of the fourth output channel, and then processed by the shift attention module of the fourth output channel to output the fourth-channel graph. The fourth-channel graph is input to the fourth output module of the prediction layer, and the fourth prediction result is output by the fourth output module of the prediction layer.
[0081] Furthermore, the multi-attention module includes a multi-head attention module and a multi-layer perceptron module, and each sub-layer is connected by a residual block.
[0082] Furthermore, the processing process of the multi-attention module includes the following steps:
[0083] The input image is divided into N image patches, and the resolution of each image patch is P×P, where N = HW÷P^2, H and W are the height and width of the input image respectively, and C is the number of channels of the input image;
[0084] Each image patch is reorganized into a vector of P^2×C dimensions, and each vector is mapped into a fixed-length vector of dimension D through linear projection to generate an embedding layer. The initialization formula of the embedding layer is:
[0085]
[0086] where, Z0 represents the initialization vector of the embedding layer, x class represents the matrix of image patches with different lengths, represents the corresponding image patch, N represents the number of image patches, E represents the dimension of image patches with different lengths, E pos represents the position of image patches with different lengths, R represents real numbers, P represents the image patch resolution, C represents the number of channels, and D represents the dimension of the fixed-length vector;
[0087] According to the density threshold ρ of the image blocks t dynamically adjust the number P of the spliced image patches 2, the calculation formula is:
[0088]
[0089] Among them, P 2 represents the number of spliced image blocks, and ρ t represents the density threshold of image segmentation, which is used to determine the number of spliced image blocks;
[0090] Splice the image block vectors after adjusting the number of splices to obtain a spliced vector matrix. In the spliced vector matrix, add the position component representing the image position to the slice component to obtain a picture slice component with position information. Perform the following operations on the picture slice component with position information in sequence:
[0091] For the layer, first perform a LayerNorm normalization operation on the output result of the previous layer , then input the normalized result into the multi-head self-attention module MSA for processing, and add the output result of the MSA to by residual to obtain the multi-head self-attention residual output of the layer Perform a LayerNorm normalization operation on , then input it into the multi-layer perceptron module MLP, and add the output result of the MLP to by residual to obtain the multi-layer perceptron residual output of the layer Perform a LayerNorm normalization operation on the multi-layer perceptron residual output z of the last layer L to obtain the final output result y. The formula is:
[0092]
[0093]
[0094] y = LN(z L )
[0095] Among them, is the multi-layer perceptron residual output of the layer, is the multi-head self-attention residual output of the layer, LN is the LayerNorm normalization operation, and z L is the multi-layer perceptron residual output of the L layer. L is the total number of layers of the multi-attention module. MSA represents the processing of the multi-head self-attention module, and MLP represents the processing of the multi-layer perceptron module.
[0096] Furthermore, the structure of the multi-attention module is as Figure 2As shown in the figure. The multi-attention module pays more attention to capturing global information while also being able to obtain rich context information. The multi-attention module improves the ability to capture information from different feature layers. It can also enhance the feature extraction ability by using the self-attention mechanism. In this application, the multi-attention module is built at the prediction layer and the end of the backbone network. The resolution of the feature map at the end of the backbone network is relatively low. Applying the multi-attention module on the low-resolution feature map can reduce the computational cost, improve the computational efficiency, and compress the model size.
[0097] As Figure 2 shown, the multi-attention module mainly includes a multi-head attention module (Muliti-Head Attention, MSA) and a multi-layer perceptron module (MLP). Residual blocks are used to connect between each sub-layer. The normalization layer (LayerNorm) and dropout layer in it help to accelerate the network convergence and effectively prevent the network from overfitting. The multi-head attention module can not only help the current node focus on the current pixel information but also obtain the context semantic information of adjacent regions.
[0098] The convolutional block attention module refers to an operation that, for a feature map, sequentially performs attention mapping inference along two independent dimensions, the channel dimension and the spatial dimension, and multiplies the two attention mapping spectra with the backbone feature map respectively to achieve adaptive feature refinement. After integrating the convolutional block attention module into the small target detection model on different classification and detection datasets, the performance of the model has been greatly improved, which proves the effectiveness of this module. On the power transmission line inspection images taken by drones, using the convolutional block attention module can extract the attention area, help the small target detection model distinguish the chaotic background element information, and make the network more focused on the small targets to be detected.
[0099] The convolutional block attention module is composed of a cascaded channel attention module and a spatial attention module. Its structural schematic diagram is as Figure 3 shown. The channel attention map is generated from the color channel relationship of the input feature. Since the input feature dimension is large and the calculation is complex, it is necessary to compress the spatial dimension of the input feature through pooling operations. Through the average pooling operation, the target distribution range is understood; through the max pooling operation, the target features are collected. The average pooling operation and the max pooling operation are used to aggregate the spatial information and context representation in the feature map, and are represented by and respectively representing the average pooling feature and the max pooling feature. After being processed by a multi-layer perceptron (MLP) and a hidden layer of size R C / r×1×1 , a one-dimensional channel attention map M c ∈R C×1×1 is obtained, and its calculation method is as follows:
[0100]
[0101] Among them, σ represents the sigmoid activation function, AvgPool(F) represents the average pooling operation, maxPool(F) represents the max pooling operation, and r represents the scaling factor. W0 ∈ R C / r×C and W1 ∈ R C×C / r respectively represent the weights of the multi-layer perceptron (MLP), and the weights are connected by ReLU as the activation function.
[0102] The spatial attention feature map is inferred from the spatial relationship of different features in different channels of the channel attention map. Different from the channel attention map, the spatial attention map pays more attention to the position information of the image, and the two complement each other. During calculation, average pooling and max pooling operations are performed along the channel axis to generate two two-dimensional mappings and They are connected and convolved through a standard convolutional layer to obtain the two-dimensional spatial attention map M s ∈ R 1 ×G×W . Performing pooling operations on the channel axis can effectively highlight the information area. Its calculation method is as follows:
[0103]
[0104] Among them, σ represents the sigmoid activation function, and f 7×7 represents the convolution operation with a convolution kernel of 7*7.
[0105] In the convolutional attention module, an anti-occlusion unit is added. When inputting an image, first remove the occlusion, then calculate the channel attention, and finally calculate the spatial attention. The channel attention and spatial attention respectively focus on the content and position of the target. Specifically, taking a feature map F ∈ R C×H×W as the input, the convolutional block attention module CBAM sequentially infers the one-dimensional channel attention map M c ∈ R C×1×1 and the two-dimensional spatial attention map M s ∈ R 1×H×W . The entire process of the convolutional block attention module CBAM can be summarized by the following formula:
[0106] L out = Conv(MaxPool(F)) + CAttn(DConv(F))
[0107]
[0108] Among them: Conv is deformable convolution to enhance the geometric adaptability of occluded targets, and CAttn is channel attention. High-frequency edge features are strengthened through local maximum pooling, and deformable convolution is used to handle occlusion scenarios such as wire winding. F ′ The feature map after channel attention weighting, and F″ is the feature map after double weighting of channel attention and spatial attention. w i is the dynamic weight coefficient, denotes element-wise multiplication, and Attn(·) is the hybrid attention function (including channel + spatial attention). Cross-scale feature interaction is achieved through adaptive weight allocation. The CBAM module that combines multi-scale channel attention captures the details of small targets in shallow features.
[0109] The dynamic weight coefficient is determined by the following formula:
[0110]
[0111] where ReLU(w j ) means calculating the previous weight w j using the activation function, ∈ represents the scale-sensitive factor, where A target represents the area of the small target, and A image represents the image area. Thus, the detail capture of small targets can be improved.
[0112] One of the main difficulties in small target detection is how to effectively represent and process multi-scale feature fusion. Early object detectors usually directly use the feature pyramid network (FPN) extracted from the backbone network for prediction. This top-down method combines multi-scale features, but has a small weight for shallow features during feature fusion, ignoring the rich location information of shallow features. Based on the improvement of the feature pyramid network (FPN), the path aggregation network (PANet) adds an additional bottom-up feature aggregation network on this basis, constructing a new network structure for cross-scale feature fusion. However, due to different input features having different resolutions, during the process of upsampling, downsampling, and tensor splicing, the weights of the output fused features are inconsistent, affecting the feature fusion of the learning features of small targets.
[0113] Therefore, in this application, the fusion module introduces learnable weights to distinguish the importance of different input features, so as to strengthen the influence of the learning features of small targets on the feature fusion network.
[0114] The biggest feature of the fusion module is to achieve efficient bidirectional cross-scale connection and weighted feature fusion. The structure diagram of the fusion module (BiFPN) is different from that of the feature pyramid network and the path aggregation network as Figure 4 shown.
[0115] The node connection method of the fusion module is different from that of the path aggregation network. The cross-scale connection optimization methods adopted are mainly as follows:
[0116] (1) Remove the only input node in the path aggregation network. Since the contribution of the node lacking feature fusion to the feature network transfer calculation is very limited, the intermediate nodes of P3 and P6 can be removed. This operation is equivalent to forming a small-scale simplified bidirectional network.
[0117] (2) Add skip connections from the input node to the output node at the same scale. The skip connections in the same feature layer fuse more features at different levels with a limited increase in computational cost.
[0118] (3) Different from the path aggregation network which has only one top-down feature path and one bottom-up feature path, the fusion module regards each bidirectional path as a feature network layer. Repeating this feature network layer multiple times can achieve higher-dimensional feature fusion.
[0119] When fusing features with different resolutions, the common method is to first adjust all features to the same resolution and then add the features. However, since different input features contribute unequally to the output feature at different resolutions, for example, in this application, the input feature weight of small targets should be strengthened to make the output feature more sensitive to small target detection. Therefore, each input needs to be weighted to enable the detection network to understand the importance of each input. The fusion module integrates the bidirectional cross-connection and fast normalization methods for feature fusion. The formula for fast normalization fusion is as follows:
[0120]
[0121] In the formula, O is the fused output feature, w i is the weight of the input feature of the i-th layer, w j is the weight of the input feature of the j-th layer, I i is the input feature of the i-th layer. w i Ensures w i ≥0 through the ReLU activation function. ε = 0.0001 is a small additional value to maintain the stability of the O value. After normalization, the weight remains in the range of 0 to 1. Thus, the calculation formula for a single layer of the fusion module is:
[0122]
[0123] In the formula, represents the intermediate feature of the i-th layer in the top-down path, represents the input feature of the i-th layer. Represents the output feature of the i-th layer from bottom to top. In the feature fusion stage of this application, depthwise separable convolution is used for operation, and batch normalization and activation functions are added after each convolution to improve the computational efficiency.
[0124] At the end of the connection layer, the above-mentioned multi-attention module is used. However, in the object detection task, the computational complexity of the self-attention of the multi-attention module is the square of the image size. Therefore, it is necessary to construct a hierarchical feature map with a computational complexity linearly related to the image size to perform dense image patch prediction. This makes directly using the multi-attention module for high-resolution images bring extremely large computational volume and high resource occupancy.
[0125] To solve the above problems, this application further modifies the multi-attention module into a shifted attention module. The shifted attention module is a hierarchical multi-attention module calculated using shifted windows instead of traditional moving windows. It performs self-attention calculation on non-overlapping local feature layers and realizes neighborhood feature aggregation through cross-layer connections. The shifted attention module constructs a hierarchical feature map by merging adjacent small-sized image patches as the depth deepens. Since the number of image patches in each feature layer is fixed, the computational complexity shows a linear relationship with the image size.
[0126] In this application, the shifted attention module introduces a hierarchical construction method in the convolutional neural network and the idea of image regions to perform self-attention calculation on non-overlapping image windows. The ViT structure calculates self-attention globally from beginning to end, while the shifted attention module continuously enlarges the window and calculates self-attention in units of the window, which is equivalent to introducing local aggregated information. Different from the convolution process in the convolutional neural network CNN, CNN performs convolution operations on each window, and the obtained value represents the feature of that window, while the shifted attention module performs self-attention calculation on each window, and what is obtained is the updated window. Through the operation of image patch merging, self-attention calculation is performed on the merged window.
[0127] The structural diagram of the shifted attention module is as Figure 5 shown. First, the picture is divided into a set of non-overlapping image patches through image patch partitioning. The size of each image patch is 4×4, the feature dimension of each image patch is 4×4×3 = 48, and the number of image patches is 160×160. In the original feature layer, the linear embedding method is applied to change the feature dimension of the divided image patches to any dimension (denoted as C). In the shifted attention module, first, adjacent image patches within a 2×2 range are merged through image patch splicing to obtain hierarchical representations of feature layers at different scales.
[0128] The shifted attention module consists of a multi-head attention model (MSA) based on a moving window, and is sequentially connected to a two-layer multi-layer perceptron (MLP) with a GELU non-linear activation function. LayerNorm is applied for regularization before both the multi-head attention model (MSA) and the multi-layer perceptron (MLP).
[0129] The multi-attention module architecture performs self-attention calculations on the entire image, consuming relatively more computing resources. In contrast, the shifted attention module focuses on self-attention calculations for non-overlapping local windows. If the image is divided into non-overlapping M×M image patches, the computational complexities of the global multi-head self-attention module MSA and the moving-window based multi-head self-attention module W-SMA are as follows:
[0130] Ω(MSA) = 4hwC 2 + 2(hw) 2 C
[0131] Ω(W-MSA) = 4hwC 2 + 2M 2 hwC
[0132] As can be seen from the above equations, the computational complexity of the multi-attention module is proportional to the square of the number of image patches hw, while the computational complexity of the moving-window based multi-head self-attention module has a linear relationship with the number of image patches hw.
[0133] As Figure 6 shown, by using the shifted window partitioning method, the calculation formula for the shifted attention module of consecutive windows is as follows:
[0134]
[0135] where and z l represent the output features of the multi-head self-attention module and the multi-layer perceptron module for image patch l, respectively.
[0136] When calculating self-attention, the similarity for each multi-head is calculated by computing the relative position bias The formula is as follows:
[0137]
[0138] In the formula, Q, K, are the query, key, and value matrices in sequence, d is the dimension of the key, M 2 is the number of image patches in the window. Since the relative position of each center is in the range of [-M + 1, M - 1], the bias matrix is B is The value of
[0139] In window-based self-attention calculation, the image is divided into 9 windows (as shown in the left figure below), and the middle feature region A is the region for information interaction. By moving the window and moving the yellow block to the lower right corner and then dividing it into 4 windows, the feature region A is independently divided, as shown in the right figure below. While further reducing the complexity of self-attention calculation, it also increases the receptive field. Figure 7 In window-based self-attention calculation, the image is divided into 9 windows (as shown in the left figure below), and the middle feature region A is the region for information interaction. By moving the window, moving the yellow block to the lower right corner and then dividing it into 4 windows, the feature region A is independently divided, as shown in the right figure below. While further reducing the complexity of self-attention calculation, it also increases the receptive field. Figure 7 In window-based self-attention calculation, the image is divided into 9 windows (as shown in the left figure below), and the middle feature region A is the region for information interaction. By moving the window, moving the yellow block to the lower right corner and then dividing it into 4 windows, the feature region A is independently divided, as shown in the right figure below. While further reducing the complexity of self-attention calculation, it also increases the receptive field.
[0140] Furthermore, the shifted attention module includes a multi-head self-attention module W-MSA based on moving windows and a multi-layer perceptron module MLP. Regularization is performed between each sub-layer through a normalization layer LayerNorm, and the GELU non-linear activation function is used.
[0141] Furthermore, the processing process of the shifted attention module includes the following steps:
[0142] The input image is divided into a set of non-overlapping image patches, each with a size of 4×4, a feature dimension of 4×4×3 = 48, and the number of image patches is 160×160;
[0143] Apply linear embedding to the original feature layer to map the feature dimension of the divided image patches to an arbitrary dimension C;
[0144] Adjacent image patches within a 2×2 range are merged through image patch stitching to generate hierarchical representations of feature layers at different scales;
[0145] For the features processed above, the multi-head self-attention module W-MSA based on moving windows performs self-attention calculation on them, and the multi-layer perceptron module MLP processes the output of the multi-head self-attention module to obtain the final output features of this layer.
[0146] Furthermore, the loss function of the small object detection model is:
[0147]
[0148] where CIoU Loss is the loss function of the small object detection model, IoU represents the intersection over union, which is used to measure the overlap degree between the predicted box and the ground truth box; Box pre represents the predicted box; Box gt represents the ground truth object box; Intersection(Box pre ,Box gt) represents the intersection area of the predicted bounding box and the ground truth bounding box; Union(Box pre , Box gt ) represents the union area of the predicted bounding box and the ground truth bounding box; represents the square of the diagonal distance of the minimum external rectangle, represents the square of the Euclidean distance between the center points of the predicted bounding box and the target bounding box, v is a parameter to measure the aspect ratio consistency, W gt and h gt respectively represent the width and height of the ground truth target bounding box, W p and h p respectively represent the width and height of the ground truth target bounding box.
[0149] Example 2:
[0150] Parts not mentioned in this example are the same as those in Example 1.
[0151] The small target detection model in this application can effectively alleviate the negative impact brought by the interpolation factor during image scaling and scale change. It is obtained by performing multi-attention module transformation on the low-level high-resolution feature map, and the feature map at this level is more sensitive to small target detection. The performance of small target detection is improved.
[0152] Experimental analysis: This application uses 500KV transmission line inspection photos as the dataset to train the small target detection network for transmission line inspection. The original dataset has a total of 4935 images, and 3938 images of different sizes after small target super-resolution processing are additionally expanded. The training set and test set are divided according to a ratio of 4:1.
[0153] This application designs five models with different module combinations, and the definitions of different types of models are as shown in Table 1 below:
[0154] Table 1 Definitions of the target detection models in this application
[0155]
[0156] In the table, P2 represents the additional prediction head and anchor box settings for small target detection. Model 5 is the optimal model in this application.
[0157] This application uses 4 NVIDIA GeForce RTX 3090 graphics cards (with a video memory of 24G) to train the model. The system is Ubuntu 20.04 LTS 64-bit system, and the processor is Xeon(R) Gold 6242R CPU @ 3.10GHz × 20, the memory is 128G, the Python version is 3.7.0, and the PyTorch version is 1.7.1.
[0158] In this application, the object detection performance of YOLOv5s on the same dataset is used as a benchmark. A modified YOLOv5s network is adopted for comparison, which adds a small object prediction head, a prediction head module based on a multi-attention module, an attention mechanism module based on a convolutional attention module, and a fusion module. At the same time, a prediction head based on a shifted attention module is added as a supplementary experiment, focusing on the detection of objects of all sizes in transmission line inspection images. According to the annotation of the dataset, it is divided into four categories: "clamp, glass insulator, composite insulator, and hardware".
[0159] Taking the basic network model YOLOv5s and the advanced model YOLOv5m as the comparison references for the experiment, the mAP and mAP@.5:.95 for objects of all sizes in this application are shown in Tables 1 and 2 below:
[0160] Table 1 mAP@.5 for object detection of all sizes in transmission lines
[0161]
[0162]
[0163] Table 2 mAP@.5:.95 for object detection of all sizes in transmission lines
[0164]
[0165] All structural improvements in this application are made by modifying the YOLOv5s network. It can be seen from the above table that the mAP@.5 performance of the improved model 5, namely the small object detection network, on the 500KV transmission line inspection dataset has comprehensively led the benchmark YOLOv5s network, and even outperformed the deeper and wider YOLOv5m network. The mAP@.5 for object detection of all sizes in transmission lines reaches 91.4%, 4.1% higher than the YOLOv5s network and 1.6% higher than the YOLOv5m network. Limited by the network depth and the number of convolutional kernels, its performance in mAP@.5:.95 is slightly inferior to the YOLOv5m network. The network layers, number of parameters, computational amount, and inference time of all models in this application are shown in Table 3 below:
[0166] Table 3 Comparison of network layers, number of parameters, computational amount, and inference time of all models
[0167]
[0168] As can be seen from the table above, even the model 5 with the most changes has a similar amount of computation as YOLOv5m, but the introduction of new modules has affected the model's reasoning speed. However, compared with the accuracy improvement gain of target detection, the model reasoning speed has a lower weight, and the reasoning speed of 12.8ms plus 0.1ms preprocessing and 0.9ms non-maximum suppression NMS processing time can still achieve a prediction speed of 71FPS, which is higher than the real-time requirement of 60FPS for target detection.
[0169] By comparing Model 2, Model 3, and Model 4, it can be seen that the number of network layers in Model 4 with the addition of the multi-attention module is reduced from 330 to 299, and the number of parameters is also reduced by 0.02M. This is because the multi-attention module and the convolutional block attention module share some network layers and parameter transfers. Therefore, the introduction of the multi-attention module can actually reduce the size of the network.
[0170] Comparing Model 4 and Model 5, we can see that the introduction of the shifted attention module has the greatest impact on the model's computational complexity, parameter quantity, and number of network layers. This is due to the shifted window calculation and self-attention calculation. However, Model 5 using the shifted attention module is only 0.5% higher than Model 4 using the multi-attention module in terms of mAP@.5 for all-size power line target detection, but the computational complexity is 50% higher. From the perspective of all-size target detection alone, this improvement is a bit costly.
[0171] Model 5 with the best detection performance is selected for data set verification. When the IoU threshold is 0.4 and the confidence threshold is 0.25, the confusion matrix is created as follows: Figure 8 As shown in the figure, since a smaller anchor box is set for small target detection, the confidence threshold of the predicted box is low and the underlying feature map is resampled, some Background FP instances are introduced into the network, but the overall impact on the detection network is small. The detection confusion matrix shows that the model is robust and stable.
[0172] The precision recall curve PR curve of the model is as follows Figure 9 As shown in the figure, when the IoU threshold is 0.5, the average precision of all types reaches 0.914. The model's detection ability for wire clamps and composite insulators is higher than that for glass insulators and hardware. This is because glass insulators are easily affected by environmental factors such as lighting, while hardware has more small target instances.
[0173] Taking YOLOv5s as the benchmark and referring to the definition of relative size of small target detection, targets with an image area of less than 0.12% of the area can be called relatively small targets. The performance comparison of small target detection on power transmission lines using such small targets as reference indicators is shown in the following table:
[0174] Table 4 Comparative test of small target detection performance on transmission lines
[0175]
[0176] It can be seen that the enhanced improvement algorithm for small target detection based on YOLOv5s has made better improvements in the task of small target detection on transmission lines. The specialized improvement for small targets has increased the mAP value of small target detection by 5.5%, and also improved the performance of target detection for all sizes of transmission lines to a certain extent.
[0177] The test results of Model 5, i.e., the small target detection model, are as Figure 10 shown. It can be Figure 10 seen that for the small fitting parts targets (such as the left side of Figure 003464) and the small glass insulator targets (such as the lower right corner of Figure 003254 and above Figure 003443) whose areas in the test set are significantly less than 1% of the image area, this model has good detection capabilities.
[0178] Combined with the experimental results, the defect elimination experiment is analyzed as follows:
[0179] (1) Adding an additional small target prediction head for the dataset increases the number of network layers of the original YOLOv5s from 213 to 270, and the computational cost increases from 15.8 GFLOPS to 24 GFLOPS. Although the network depth and computational cost increase, the mAP@.5 for small target detection increases by 1.4%, and at the same time, the mAP@.5 for target detection of all sizes of transmission lines increases by 1.8%. This improvement is effective.
[0180] (2) Introducing the multi-attention module and the shifted attention module, it can be seen that in terms of mAP@.5 for small target detection, the latter increases by 0.6% compared to the former. Compared with the YOLOv5s algorithm without adding the multi-attention module, the mAP@.5 for small target detection increases by 5.5%, and the effect is very significant. At the same time, using the multi-attention module can not only increase the feature mapping of small targets but also play a role in target detection of all sizes. However, the computational cost of the shifted attention module is larger than that of the multi-attention module, which seems to be overkill in the network structure of YOLOv5s. In deeper network structures such as YOLOv5l or YOLOv5x, the shifted attention module will significantly improve the detection accuracy of small targets.
[0181] (3) By introducing the Convolutional Block Attention Module, it is not difficult to see from the comparison between Model 1 and Model 3 that there is not much difference in the number of parameters and computational volume of the models, and there is no prominent impact on the mAP@.5 of object detection for all sizes, but it has a greater impact on the mAP@.5 value of small object detection, reaching a difference of 1.8%. Therefore, using the Convolutional Block Attention Module can extract the attention area, help YOLOv5 distinguish the chaotic background element information, and make the network more focused on the small objects to be detected.
[0182] In this application, the architecture improvement of the small object detection model has obvious improvement in the detection performance, especially in the small object detection performance, and the real-time performance can also meet the requirements of real-time detection and path planning. In the task of UAV inspection of transmission lines, this application not only significantly improves the performance of the object detection model, but also can be extended to other applications involving small object detection or long-distance object detection. It has a better improvement in the detection ability and perception robustness of small objects on transmission lines, and has a positive impact on the detection of objects of all sizes on transmission lines, so as to provide better planning and decision-making strategies for intelligent unmanned inspection of transmission lines and defect identification of transmission lines.
[0183] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0184] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A small target detection method for transmission lines based on improved YOLOv5s, characterized in that, The steps include: Construct a small target detection model, train the small target detection model, and detect small targets on the transmission line through the trained small target detection model. The small target detection model includes an input end, a backbone network, a connection layer, and a prediction layer connected in sequence. The input end uses the Mosaic method for data augmentation, randomly scales, crops, and arranges the input image, and splices it into a new input image. The backbone network includes a sampling module, multiple groups of convolution modules, residual modules, convolution modules, and pooling modules connected in sequence. The connection layer includes multiple attention modules, convolution modules, multiple groups of upsampling modules, a fusion module, multiple groups of residual modules, a convolutional attention module, an upsampling module, and four output channels. The prediction layer includes multiple output modules, and the output modules are connected to the output channels at the end of the connection layer.
2. The small target detection method for transmission lines based on improved YOLOv5s according to claim 1, wherein, The four output channels of the connection layer include a first output channel, a second output channel, a third output channel, and a fourth output channel. Each output channel includes a fusion module, a shift attention module, and a convolutional attention module connected in sequence. The output ends of the multiple attention modules and convolution modules in the connection layer are connected to the fusion module of the fourth output channel. The output ends of the first group of residual modules and convolutional attention modules in the connection layer are connected to the fusion module of the third output channel. The output ends of the second group of residual modules and convolutional attention modules in the connection layer are connected to the fusion module of the second output channel. The output end of the last upsampling module in the connection layer is connected to the fusion module of the first output channel. The first output channel, the second output channel, the third output channel, and the fourth output channel are connected in sequence.
3. A small target detection method for transmission lines based on improved YOLOv5s according to claim 1, characterized in that, The sampling module of the backbone network downsamples the input image through the inverse operation of sub-pixel convolution, specifically including: The input image is divided into four parts. After each part of the data is downsampled by a factor of 2, they are concatenated in the channel dimension and a convolution operation is performed, and the number of output channels is one-fourth of the number of input channels.
4. A small target detection method for transmission lines based on improved YOLOv5s according to claim 1, characterized in that, The output ends of the first group of convolution modules and residual modules in the backbone network are connected to the fusion module of the first output channel in the connection layer. The output end of the pooling module in the backbone network is connected to the input end of the multiple attention modules in the connection layer.
5. A small target detection method for transmission lines based on improved YOLOv5s according to claim 1, characterized in that The detection of small targets on the transmission line through the trained small target detection model specifically includes: After the input image is input through the input end, it is processed by the sampling module, convolution module, and residual module to obtain a first feature map and a first calculation graph. The first feature map passes through two groups of convolution modules and residual modules in the backbone network, and then is input into a convolution module and a pooling module and outputs a second feature map. The second feature map is input into the multiple attention modules and convolution modules in the connection layer. The first calculation graph is input into the fusion module of the first output channel in the connection layer. The second feature map is processed by a multi-attention module and a convolutional module to output a second computational graph. The second computational graph is respectively input into the first upsampling module and the fusion module of the connection layer, the first group of residual modules and the convolutional attention module to output a third computational graph. The third computational graph is output by the second group of upsampling modules and the fusion module, the second group of residual modules and the convolutional attention module to output a fourth computational graph. The fourth computational graph is processed by the last upsampling module of the connection layer to output a fifth computational graph; After the fifth computational graph is fused with the first computational graph through the fusion module of the first output channel of the connection layer, it is input into the shifted attention module of the first output channel. After being processed by the shifted attention module, a first channel map is output. The first channel map is input into the first output module of the prediction layer, and the first prediction result is output by the first output module of the prediction layer; After the first channel map is processed by the convolutional attention module of the first output channel, it is fused with the fourth computational graph through the fusion module of the second output channel, and then output a second channel map after being processed by the shifted attention module of the second output channel. The second channel map is input into the second output module of the prediction layer, and the second prediction result is output by the second output module of the prediction layer; After the second channel map is processed by the convolutional attention module of the second output channel, it is fused with the third computational graph through the fusion module of the third output channel, and then output a third channel map after being processed by the shifted attention module of the third output channel. The third channel map is input into the third output module of the prediction layer, and the third prediction result is output by the third output module of the prediction layer; After the third channel map is processed by the convolutional attention module of the third output channel, it is fused with the second computational graph through the fusion module of the fourth output channel, and then output a fourth channel map after being processed by the shifted attention module of the fourth output channel. The fourth channel map is input into the fourth output module of the prediction layer, and the fourth prediction result is output by the fourth output module of the prediction layer.
6. A small target detection method for transmission lines based on improved YOLOv5s according to claim 1, characterized in that The multi-attention module includes a multi-head attention module and a multi-layer perceptron module, and each sub-layer is connected by a residual block.
7. A small target detection method for transmission lines based on improved YOLOv5s according to claim 1, characterized in that, The processing process of the multi-attention module includes the following steps: The input image is divided into N image patches, and the resolution of each image patch is P×P, where N = HW÷P^2, H and W are the height and width of the input image respectively, and C is the number of channels of the input image; Each image patch is reorganized into a vector of P^2×C dimensions, and each vector is mapped into a fixed-length vector of dimension D through linear projection to generate an embedding layer. The initialization formula of the embedding layer is: Among them, Z0 represents the initialization vector of the embedding layer, and x class represents the matrix of image patches with different lengths, represents the corresponding image patch, N represents the number of image patches, E represents the dimension of image patches with different lengths, and E pos represents the positions of image patches with different lengths, R represents real numbers, P represents the image patch resolution, C represents the number of channels, and D represents the dimension of the fixed-length vector; According to the density threshold ρ of image segmentation t Dynamically adjust the number P of spliced image blocks 2 , and the calculation formula is as follows: Among them, P 2 represents the number of spliced image blocks, and ρ t represents the density threshold of image segmentation, which is used to determine the number of spliced image blocks; The image patch vectors after adjusting the splicing quantity are spliced to obtain a spliced vector matrix. In the spliced vector matrix, the position component representing the image position is added to the slice component to obtain a picture slice component with position information. The following operations are sequentially performed on the picture slice component with position information: For the layer, first perform a LayerNorm normalization operation on the output result of the previous layer , then input the normalized result into the multi-head self-attention module MSA for processing, add the output result of the MSA to residually, and obtain the multi-head self-attention residual output of the layer Perform a LayerNorm normalization operation on , then input it into the multi-layer perceptron module MLP, add the output result of the MLP to residually, and obtain the multi-layer perceptron residual output of the layer Perform a LayerNorm normalization operation on the multi-layer perceptron residual output z of the last layer L to obtain the final output result y, and the formula is: y = LN(z L ) Among them, is the residual output of the multi-layer perceptron of the th layer, is the residual output of the multi-head self-attention of the th layer, LN is the LayerNorm normalization operation, z L is the residual output of the multi-layer perceptron of the Lth layer, L is the total number of layers of the multi-attention module, MSA represents the processing of the multi-head self-attention module, and MLP represents the processing of the multi-layer perceptron module.
8. A small target detection method for transmission lines based on improved YOLOv5s according to claim 1, characterized in that The shifted attention module includes a multi-head self-attention module based on a moving window W-MSA and a multi-layer perceptron module MLP. Each sub-layer is regularized through a normalization layer LayerNorm and uses a GELU non-linear activation function.
9. A small target detection method for transmission lines based on improved YOLOv5s according to claim 1, characterized in that, The processing procedure of the shift attention module includes the following steps: Divide the input image into a set of non-overlapping image patches, each with a size of 4×4, a feature dimension of 4×4×3 = 48, and the number of image patches is 160×160; Apply linear embedding on the original feature layer to map the feature dimension of the divided image patches to any dimension C; Merge adjacent image patches within a 2×2 range through image patch splicing to generate hierarchical representations of feature layers at different scales; For the features processed above, perform self-attention calculation on them using the window-based multi-head self-attention module W-MSA, and the multi-layer perceptron module MLP processes the output of the multi-head self-attention module to obtain the final output features of this layer.
10. A small target detection method for transmission lines based on improved YOLOv5s according to claim 1, characterized in that, The loss function of the small object detection model is: Among them, CIoU Loss is the loss function of the small target detection model. IoU represents the intersection over union, which is used to measure the overlapping degree between the predicted bounding box and the ground truth bounding box; Box pre represents the predicted bounding box; Box gt represents the ground truth target bounding box; Intersection(Box pre , Box gt ) represents the intersection area between the predicted bounding box and the ground truth bounding box; Union(Box pre , Box gt ) represents the union area between the predicted bounding box and the ground truth bounding box; represents the square of the diagonal distance of the minimum enclosing rectangle, represents the square of the Euclidean distance between the center points of the predicted bounding box and the target bounding box, v is a parameter to measure the consistency of the aspect ratio, W gt and h gt respectively represent the width and height of the ground truth target bounding box, W p and h p respectively represent the width and height of the ground truth target bounding box.
Citation Information
Patent Citations
Intelligent defect detection method for transmission lines based on improved YOLOv5 network
CN116843649B
Cited By
Distribution line unmanned aerial vehicle inspection target identification method fused with multi-feature extraction module
CN120997715A
Urban low-altitude weak unmanned aerial vehicle identification method based on self-attention mechanism
CN121582552A
Unmanned aerial vehicle self-adaptive route planning and dynamic obstacle avoidance method and system for distribution line inspection
CN121657730A
Unmanned aerial vehicle adaptive route planning and dynamic obstacle avoidance method and system for power distribution line inspection
CN121657730B