Improved unmanned aerial vehicle small target detection method based on YOLOv8n
By improving the YOLOv8n network structure, introducing deformable convolution and TA attention mechanisms, and using PIoU loss function, the recognition problem of drone small object detection in complex background is solved, and efficient and accurate small object detection is achieved, suitable for real-time deployment.
Patent Information
- Application Number
- CN202510170160.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-13
AI Technical Summary
The existing drone small-objective detection algorithms do not perform well in complex contexts, making it difficult to accurately identify small targets, and the algorithm is complex and has a large amount of calculations, making it difficult to achieve real-time processing and deployment.
Based on YOLOv8n's improved drone small object detection method, the DSFTA module is designed by optimizing the network structure, introducing deformable convolution and TA attention mechanism, and introducing PIoU loss function to improve detection accuracy and reduce calculation amount.
It improves the accuracy and efficiency of small-target detection of drones, reduces the amount of model parameters and calculations, and is suitable for real-time deployment, especially in complex backgrounds and dense small-target scenarios.
Smart Images

Figure CN119992391A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) small target detection, and in particular to an improved UAV small target detection method based on YOLOv8n. Background Art
[0002] With the rapid development of drone technology, drone target detection, as one of the important application fields, is also constantly improving. Due to its flexibility and wide field of view, drones have brought convenience to humans in many fields, such as urban traffic planning, power transmission line insulator defect detection, and forest fire detection. Therefore, improving the performance of drone target detection algorithms can further expand the application value and benefits of drones in various practical scenarios.
[0003] In the field of small target detection in drones, there are significant differences between traditional target detection algorithms and deep learning-based algorithms. Traditional methods usually rely on hand-designed feature extraction (such as SIFT, HOG, etc.) and classifiers (such as support vector machines SVM), which have certain recognition capabilities for specific types of objects, but their generalization and accuracy are often limited when faced with complex natural scenes. In addition, traditional algorithms are particularly difficult to detect small targets because they are easily affected by factors such as background clutter, lighting changes, and posture diversity, and their computational efficiency is low, making it difficult to meet real-time requirements.
[0004] In contrast, deep learning algorithms, especially convolutional neural networks (CNNs), automatically learn feature representations through large amounts of data, showing greater adaptability and higher detection accuracy. Deep learning models can capture more abstract and high-level features, which is particularly important for the detection of small targets in complex backgrounds. At the same time, with the advancement of hardware performance and the development of optimization technologies such as GPU acceleration and lightweight network design, deep learning algorithms can achieve fast detection while ensuring high accuracy, making them suitable for actual deployment.
[0005] Due to the particularity of the aerial photography perspective of drones, existing deep learning algorithms perform poorly when faced with this task. On the one hand, the scale of targets in aerial images varies greatly, especially the pixel values of small targets in the images are low, making it difficult for the algorithm to extract the features of small targets; on the other hand, the interference of complex backgrounds affects the detection effect, such as the intensity of light and a large number of objects irrelevant to the detection. These factors make it difficult for drones to accurately identify targets during target detection.
[0006] Although deep learning has brought significant improvements in the detection of small targets in drones, it has also introduced new challenges. For example, in order to improve detection performance, scholars have tried to improve existing models by adding parallel network structures, increasing upsampling scales, introducing attention mechanisms, replacing loss functions, etc., but these measures often lead to increased algorithm complexity and increased computational complexity, which increases the requirements for hardware performance and is not conducive to real-time processing and deployment. In addition, in the process of improving detection accuracy, it is difficult to balance efficiency and accuracy, and cross-layer information fusion needs to be strengthened. Therefore, while pursuing high accuracy, how to reduce algorithm complexity, reduce computational complexity and maintain efficient cross-layer information transmission is a key issue that needs to be addressed in current research. Summary of the invention
[0007] In view of the shortcomings of the prior art, the present invention provides an improved UAV small target detection method based on YOLOv8n. For different deployment platforms and application scenarios, YOLOv8n provides models of different scales such as n / s / m / l / x. Considering that UAV small target detection needs to comprehensively consider the accuracy and the difficulty of hardware deployment, YOLOv8n is selected as the benchmark model for improvement. The specific scheme is as follows:
[0008] A method for detecting small targets of unmanned aerial vehicles based on improved YOLOv8n includes the following steps:
[0009] Step 1: Perform image preprocessing on the drone small target dataset and randomly divide it into a validation set and a training set; the drone small target dataset contains static photos taken by the drone in different scenes;
[0010] Specifically, the images are uniformly resized to a set size, the training set is expanded through data augmentation techniques, and finally the images are converted into tensor format and the bounding box coordinates are adjusted;
[0011] Step 2: Optimize the YOLOv8n network structure; the details are as follows:
[0012] The YOLOv8n backbone network downsamples through convolution and pooling, gradually reducing the size of the feature map, while continuously learning the feature information of the target; then the feature maps with input dimensions of 20×20×1024, 40×40×512 and 80×80×256 in the YOLOv8n backbone network are sent to the Neck part for feature fusion, wherein the YOLOv8n network structure is optimized to reduce the downsampling multiples of the backbone network, and the scales of the feature maps used for feature fusion of the Neck part in the backbone network are adjusted to 40×40×512, 80×80×256 and 160×160×128.
[0013] Step 3: Introduce deformable convolution into the YOLOv8n backbone network;
[0014] Improve the C2f module in the YOLOv8n network; Specifically, the deformable convolution DCNV2 is used to improve the C2f module to form the CD2 module, which is dynamically adjusted according to the characteristics of the detection target; Specifically, the deformable convolution learns a set of offsets and modulatable items Mask through an additional convolution layer. The offsets correspond to the horizontal and vertical offsets respectively. The modulatable items give each sampling point a new weight, and then the pixel value of the new sampling point is obtained through bilinear interpolation, and it is multiplied and summed with the convolution kernel to finally get the output result;
[0015] The operational expression of the deformable convolution DCNV2 is as follows:
[0016]
[0017] Where F() represents the output feature map, x i,j represents the pixel value of the input i-th row and j-th column, x i+l,j+r It represents the pixel value after position shift, l represents the displacement in the horizontal direction, r represents the displacement in the vertical direction, and w k is the weight corresponding to the kth channel, M is the modulatable term, and K represents the number of convolution kernels at the sampling position;
[0018] Step 4: Introduce the parameter-free TA attention mechanism to improve the bidirectional weighted pyramid BiFPN; the details are as follows:
[0019] The TA attention mechanism adopts a three-branch design to capture the interaction results between H dimension and C dimension, C dimension and W dimension, and W dimension and H dimension respectively; each branch processes the input tensor, and the input tensor is rotated, Z-pooled, convolved, batch normalized, and Sigmoid activated to obtain the interaction results between different dimensions. Finally, the processing results of the three branches are summed and averaged to form a comprehensive attention map; the specific calculation formula is as follows:
[0020] Z-pool(x)=[Maxpool 0d , Avgpool 0d ]
[0021]
[0022] Where Maxpool 0d 、Avgpool 0d Represent the maximum pooling layer and the average pooling layer respectively, R1 and R2 represent the rotation operations of the first branch and the second branch respectively, They represent the convolutional layers of the three branches respectively, σ() represents the sigmoid activation function, is the tensor obtained by the first branch through Z-pool, is the tensor obtained by the second branch through Z-pool, is the tensor obtained by the third branch through Z-pool, and X is the input tensor;
[0023] The DSFTA module is designed by combining the TA attention mechanism with the bidirectional weighted pyramid BiFPN. The DSFTA module uses the CBS module to downsample the shallow feature map, where the CBS module includes Conv convolution, BatchNorm batch normalization, and Silu activation function. The downsampled shallow feature map is then merged with the deep feature map in the channel dimension. Subsequently, the CBS module is used to integrate the merged feature map, and the number of output channels is set to 256. Finally, the TA attention mechanism is introduced to enhance the feature information of small targets in the feature map.
[0024] Step 5: Introduce PIoU loss function;
[0025] The specific formula is as follows:
[0026]
[0027] L PIoU =1-PIoU
[0028] q=e -P
[0029]
[0030] L PIoUv2 =u(λq)×L PIou
[0031] Where, d w1 d w2 d h1 d h2 is the absolute value of the distance between the corresponding edge of the predicted box and the real box, w gt 、h gt represents the height and width of the target box; P is the penalty factor; IoU is the intersection over union ratio, L PIoU is the loss function of PIoUv1, q is the hyperparameter for measuring the quality of the prediction box, u(x) is the attention function, and λ is the hyperparameter for controlling the attention mechanism of the function.
[0032] Step 6: Use the training set to train the improved model, adjust the parameters through the optimization algorithm to minimize the loss function, and regularly use the validation set to evaluate the performance to ensure the generalization ability of the model. After the training is completed, save the weight file with the best performance on the validation set to form the target detection model, integrate the optimal weight file into the reasoning environment of the drone, and the drone collects data in real time and applies the target detection model to perform target detection;
[0033] The beneficial effects of adopting the above technical solution are:
[0034] The present invention provides a method for detecting small targets of unmanned aerial vehicles (UAVs) based on an improved YOLOv8n. Aiming at the difficulties in the existing small target detection tasks of unmanned aerial vehicles, the present invention improves YOLOv8n and designs an improved YOLOv8n network for detecting small targets of unmanned aerial vehicles. The specific structure is as follows: Figure 1 As shown in the figure. First, the model structure is optimized to reduce the feature loss caused by downsampling, greatly reduce the number of model parameters, and improve the accuracy; second, the CD2 module is designed, combined with deformable convolution, to enhance the backbone network's ability to extract features for small targets; third, the TA-BiFPN structure is designed, which retains more shallow feature information that is beneficial to small targets, and introduces the TA attention mechanism to make the model focus on target information, effectively improving the efficiency of feature fusion; finally, the PIoUv2 loss function is cited to improve the performance of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is the overall network structure diagram in the embodiment of the present invention:
[0036] Figure 2 is a deformable convolution structure diagram in an embodiment of the present invention;
[0037] Figure 3 TA attention mechanism diagram in an embodiment of the present invention;
[0038] Figure 4 is a diagram of an improved Neck network structure in an embodiment of the present invention;
[0039] Figure 5 It is a comparison of thermal images in the embodiments of the present invention;
[0040] Figure (a) is the original image, Figure (b) is before TA attention is introduced, and Figure (c) is after TA attention is introduced;
[0041] Figure 6 Graphs showing detection effects in different scenarios in an embodiment of the present invention;
[0042] in Figure 6 (a) is a comparison chart of different algorithms for scenes photographed at night. Figure 6 (b) is a comparison chart of algorithms for complex background and dense small target detection scenarios. DETAILED DESCRIPTION
[0043] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0044] A small target detection method for drones based on improved YOLOv8n, such as Figure 1 As shown, the following steps are included:
[0045] Step 1: Perform image preprocessing on the drone small target dataset and randomly divide it into a validation set and a training set; the drone small target dataset contains static photos taken by the drone in different scenes;
[0046] Specifically, the images are uniformly resized to a set size, the training set is expanded through data augmentation techniques, and finally the images are converted to tensor format and the bounding box coordinates are adjusted to ensure accuracy. The preprocessing steps help improve the generalization ability and detection performance of the model.
[0047] In this example, the existing publicly available Visdrone2019 dataset is used. The dataset contains a total of 8629 static photos, of which 6471 photos are used as training sets and 548 photos are used as validation sets. As shown in the figure, the dataset contains 10 categories, namely pedestrians, people, bicycles, cars, vans, trucks, tricycles, awning-tricycles, buses, and motorcycles.
[0048] Step 2: Optimize the YOLOv8n network structure as follows:
[0049] The YOLOv8n backbone network downsamples through convolution and pooling, gradually reducing the size of the feature map, while continuously learning the feature information of the target; then the feature maps with input dimensions of 20×20×1024, 40×40×512 and 80×80×256 in the YOLOv8n backbone network are sent to the Neck part for feature fusion and finally used for prediction. However, small targets have less feature information, and too large a sampling scale may cause the loss of feature information related to small targets, thereby affecting the detection effect. In response to the above problems, the YOLOv8n network structure is optimized to reduce the downsampling multiple of the backbone network, and the feature map scale used for feature fusion in the Neck part of the backbone network is adjusted to 40×40×512, 80×80×256 and 160×160×128. After optimizing the sampling scale, the loss of feature information of small targets is reduced, and the number of parameters of the model is reduced.
[0050] Step 3: Introduce deformable convolution into the YOLOv8n backbone network;
[0051] Improve the C2f module in the YOLOv8n network; conventional convolution uses a convolution kernel with fixed sampling points to extract features, which is suitable for detecting objects with regular structures. However, since the targets detected by drones are usually irregularly shaped objects and the targets occlude each other, it is difficult for conventional convolution layers to extract effective features. Different from conventional convolution, the deformable convolution DCNV2 is used to improve the C2f module to form the CD2 module, which introduces an additional transformation step so that the position of the convolution kernel is no longer fixed and can be dynamically adjusted according to the characteristics of the detected target; the structural diagram of the deformable convolution is shown in the figure. Figure 2 As shown in the figure, the deformable convolution learns a set of offsets and modulatable items through additional convolution layers. The offsets correspond to the horizontal and vertical offsets, respectively, so that the convolution can introduce this set of offsets at a fixed sampling position. The modulatable items give each sampling point a new weight, so that the convolution can achieve a dynamic adjustment process, and the sampling is more flexible and more suitable for target detection tasks from the perspective of drones. Then the pixel value of the new sampling point is obtained through bilinear interpolation, and it is multiplied and summed with the convolution kernel to finally get the output result; the operation expression of the deformable convolution DCNV2 is as follows:
[0052]
[0053] Where F() represents the output feature map, x i,j represents the pixel value of the input i-th row and j-th column, x i+l,j+r It represents the pixel value after position shift, l represents the displacement in the horizontal direction, r represents the displacement in the vertical direction, and w kis the weight corresponding to the kth channel, M is the modulatable term, and K represents the convolution kernel of K sampling positions;
[0054] The C2f module is a key component in YOLOv8, mainly used for feature extraction and fusion. The DCNv2 module is introduced to improve the original C2f module, and the CD2 module is designed. On the basis of the original C2f module, the flexible sampling of deformable convolution is added to enhance the feature extraction capability of the backbone network and improve the detection accuracy.
[0055] Step 4: Introduce the parameter-free TA attention mechanism to improve the bidirectional weighted pyramid BiFPN, as follows:
[0056] Bidirectional Feature Pyramid Network (BiFPN) is an advanced feature fusion architecture designed to improve multi-scale feature representation in target detection tasks. It not only enhances the interaction between shallow and deep features, but also ensures efficient fusion of features of different scales by introducing bidirectional information flow between bottom-up and top-down paths. BiFPN can effectively capture the details of objects from small to large, reduce the impact of background noise, and its lightweight design makes it easy to integrate into various detection models, significantly improving detection accuracy and speed, especially when dealing with small target detection.
[0057] The TA attention mechanism is as follows Figure 3 As shown in the figure, the structure adopts a three-branch design to capture the interaction results between H dimension and C dimension, C dimension and W dimension, and W dimension and H dimension respectively; each branch processes the input tensor, and the input tensor is rotated, Z-pooled, convolved, batch normalized, and Sigmoid activated to obtain the interaction results between different dimensions. Finally, the processing results of the three branches are summed and averaged to form a comprehensive attention map; the specific calculation formula is as follows:
[0058] Z-pool(x)=[Maxpool 0d , Avgpool 0d ]
[0059]
[0060] Where Maxpool 0d 、Avgpool 0d Represent the maximum pooling layer and the average pooling layer respectively, R1 and R2 represent the rotation operations of the first branch and the second branch respectively, They represent the convolutional layers of the three branches respectively, σ() represents the sigmoid activation function, is the tensor obtained by the first branch through Z-pool, is the tensor obtained by the second branch through Z-pool, is the tensor obtained by the third branch through Z-pool, and X is the input tensor.
[0061] The improved Neck network structure is shown in the figure below: Figure 4 As shown in the figure, the TA attention mechanism is combined with the bidirectional weighted pyramid BiFPN, and the DSFTA module is designed; the DSFTA module uses the CBS module to downsample the shallow feature map, where the CBS module contains Conv convolution, BatchNorm batch normalization and Silu activation function, which retains important edge and texture information while preparing for subsequent feature fusion, and then merges the downsampled shallow feature map with the deep feature map in the channel dimension to achieve effective communication of cross-level features. Subsequently, the CBS module is used to integrate the merged feature map, and the number of output channels is set to 256 to reduce redundant information and reduce model complexity. Finally, the TA attention mechanism (TripleAttention) is introduced to enhance the feature information of small targets in the feature map, so that the model can focus more on important features, significantly improving detection accuracy and robustness.
[0062] The shallow feature map in this embodiment refers to the feature map generated by the convolution layer close to the input end of the neural network, which mainly contains low-level feature information, such as edges, textures, colors, etc. The deep feature map refers to the feature map generated by the convolution layer close to the output end of the neural network, which contains high-level semantic information, such as categories, structures, etc.
[0063] Step 5: Introduce PIoU loss function;
[0064] In the detection of small targets in drones, bounding box regression (BBR) is a key step to ensure high-precision detection. Although the CIoU loss function adopted by YOLOv8 performs well in many aspects, it still has some limitations when dealing with complex backgrounds and multi-scale small targets, such as anchor box expansion problems and slow convergence speed.
[0065] In comparison, PIoU introduces an adaptive penalty factor for target size, which solves the problem of imbalance of targets of different scales in the regression process. This penalty factor is dynamically adjusted according to the actual size of the target, avoiding unnecessary expansion of the anchor frame due to excessive penalty, which is especially important for small targets photographed by drones because these targets often have complex backgrounds and changeable postures. In addition, PIoU designs a gradient adjustment function based on the quality of the anchor frame, which evaluates the quality of each anchor frame and adjusts the direction and strength of the gradient update, so that the anchor frame moves closer to the true bounding box more quickly and stably, reducing the influence of interference factors in complex backgrounds. After further development into PIoUv2, a non-monotonic attention layer was introduced, which is specifically used to enhance the attention to medium-quality anchor frames, helping the model to better focus on those targets that have potential but have not yet been fully matched, and promoting these anchor frames to converge to the optimal state faster, which is particularly suitable for the precise positioning of small targets in drone aerial images. The specific formula is as follows:
[0066]
[0067] L PIoU =1-PIoU
[0068] q=e -P
[0069]
[0070] L PIoUv2 =u(λq)×L PIou
[0071] Where, d w1 d w2 d h1 d h2 is the absolute value of the distance between the corresponding edge of the predicted box and the real box, w gt 、h gt represents the height and width of the target box; P is the penalty factor; IoU is the intersection over union ratio, L PIoU is the loss function of PIoUv1, q is the hyperparameter for measuring the quality of the prediction box, u(x) is the attention function, and λ is the hyperparameter for controlling the attention mechanism of the function.
[0072] Step 6: Use the training set to train the improved model, adjust the parameters through the optimization algorithm to minimize the loss function, and regularly evaluate the performance with the validation set to ensure the generalization ability of the model. After the training is completed, save the weight file that performs best on the validation set, integrate the optimal weight file into the reasoning environment of the drone, and the drone collects data in real time and applies the target detection model for target detection; achieving efficient target recognition and positioning in tasks such as environmental monitoring and search and rescue.
[0073] In order to verify the effectiveness of the improved algorithm, ablation experiments were carried out on the various modules proposed in this paper. In order to ensure the consistency of the experiments, on the Visdrone2019 dataset, YOLOv8n was used as the benchmark model, and the improved modules were added one by one for experiments. Each experiment used the same experimental parameters. The experimental results are shown in the following table:
[0074]
[0075]
[0076] In the table, Baseline, A, B, C, and D represent YOLOv8n, CD2, TA-BiFPN, optimized network structure, and PioUv2, respectively. As can be seen from the table, after adding CD2, the deformable convolution strengthens the backbone network's ability to extract irregular small target features. In terms of accuracy indicators, the precision, recall, mAP 50, and mAP 50-95 increased by 1.6%, 2.2%, 1.9%, and 1.3%, respectively. The model size increased by 0.5M, and the amount of calculation decreased by 0.2G. After adding TA-BiFPN, compared with the original Neck network, the algorithm's complexity indicators were reduced, and the amount of calculation and model size decreased by 0.5G and 1.2M, respectively. At the same time, the precision decreased by 0.8%, but the recall, mAP 50, and mAP After optimizing the sampling scale, the model structure is simplified, and the model size is reduced from 6.2M to 2.2M, which is a significant decrease. The algorithm increases the amount of calculation, from 8.2G to 10.6G, the precision rate increases by 1.6%, the recall rate increases by 1.2%, and the mAP 50 and mAP 50-95 increase by 1.2% and 1% respectively. After the introduction of PioUv2, the performance of the algorithm is improved, and the precision rate, recall rate, mAP 50 and mAP 50-95 increase by 0.1%, 0.9%, 0.4% and 0.1% respectively. Ultimately, the improved algorithm has improved accuracy indicators, with precision, recall, mAP 50 and mAP 50-95 increased by 3.9%, 3.6%, 3.8% and 2.3% respectively. Although the algorithm's computational complexity increased by 2.6G, the model size decreased by 60%, meeting the requirements for small target detection deployment in drones.
[0077] like Figure 5 As shown in the figure, the recognition effect before and after adding the TA attention mechanism is compared through heat map visualization. Figure 5 (a) is the original image. From the comparison, we can clearly see that the heat map after the introduction of the attention mechanism is more concentrated in the target area, with higher heat, which convexly shows the degree of attention paid by the model to the target. Figure 5As shown in (c), in contrast, heat maps without attention mechanisms may have heat dispersion and background interference, resulting in inaccurate or insufficiently prominent target positioning, such as Figure 5 (b) As shown in the figure, the comparison of heat maps shows that the introduction of the attention mechanism improves the effect of the heat map, which improves the model's attention to the target and the positioning accuracy.
[0078] In order to intuitively feel the effectiveness of the improved algorithm in this embodiment, pictures in different scenes are introduced for comparison, such as Figure 6 The figure shows the comparison of detection effects before and after the improvement. Figure 6 (a) is a comparison chart of algorithms in night aerial photography scenes. Figure 6 (b) is a comparison chart of algorithms in complex background and dense small target scenes. The left side is the detection effect of YOLOv8n, and the right side is the improved detection effect.
[0079] from Figure 6 The first row of comparison pictures in (a) shows that YOLOv8n has poor target recognition in the case of insufficient light at night, and many vehicles are missed. The improved algorithm greatly reduces the number of missed detections. Figure 6 (b) It can be seen that in the scene of dense small target detection, the targets occlude each other and cannot be well identified before the improvement. For example, when a person and a car stand together, YOLOv8n only detects the car, while the improved algorithm can distinguish between the two and the number of missed targets is also reduced.
[0080] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.
Claims
1. A UAV small target detection method based on YOLOv8n improvement, characterized in that: The following steps are involved: Step 1: Perform image preprocessing on the drone small target dataset and randomly divide it into a validation set and a training set; the drone small target dataset contains static photos taken by the drone in different scenes; Step 2: Optimize the YOLOv8n network structure; Step 3: Introduce deformable convolution into the YOLOv8n backbone network; Step 4: Introduce the parameter-free TA attention mechanism to improve the bidirectional weighted pyramid BiFPN; Step 5: Introduce PIoU loss function; Step 6: Use the training set to train the improved model, adjust the parameters through the optimization algorithm to minimize the loss function, and regularly evaluate the performance with the validation set to ensure the generalization ability of the model. After the training is completed, save the weight file with the best performance on the validation set to form a target detection model, integrate the optimal weight file into the reasoning environment of the drone, and the drone collects data in real time and applies the target detection model for target detection.
2. According to the improved UAV small target detection method based on YOLOv8n according to claim 1, it is characterized in that: The preprocessing described in step 1 is specifically as follows: uniformly resizing the images to a set size, expanding the training set through data augmentation technology, and finally converting the images into a tensor format and adjusting the bounding box coordinates.
3. The improved UAV small target detection method based on YOLOv8n according to claim 1 is characterized in that: The step 2 is specifically as follows: the YOLOv8n backbone network performs downsampling through convolution and pooling to gradually reduce the size of the feature map, while continuously learning the feature information of the target; then, the feature maps with input dimensions of 20×20×1024, 40×40×512 and 80×80×256 in the YOLOv8n backbone network are sent to the Neck part for feature fusion, wherein the YOLOv8n network structure is optimized to reduce the downsampling multiple of the backbone network, and the scale of the feature map used for feature fusion of the Neck part in the backbone network is adjusted to 40×40×512, 80×80×256 and 160×160×128.
4. The improved UAV small target detection method based on YOLOv8n according to claim 1 is characterized in that: The step 3 is specifically as follows: improving the C2f module in the YOLOv8n network; specifically using deformable convolution DCNV2 to improve the C2f module to form a CD2 module, and dynamically adjusting according to the characteristics of the detection target; specifically, the deformable convolution learns a set of offsets Offsets and modulatable items Mask through an additional convolution layer, the offsets correspond to horizontal and vertical offsets respectively, and the modulatable items give each sampling point a new weight, and then obtain the pixel value of the new sampling point through bilinear interpolation, and multiply and sum it with the convolution kernel to finally obtain the output result.
5. The improved UAV small target detection method based on YOLOv8n according to claim 4 is characterized in that: The operational expression of the deformable convolution DCNV2 is as follows: Where F() represents the output feature map, x i,j Represents the pixel value of the input row i and column j, x i+l,j+r It represents the pixel value after position shift, l represents the displacement in the horizontal direction, r represents the displacement in the vertical direction, and w k is the weight corresponding to the kth channel, M is the modulatable term, and K represents the number of convolution kernels at the sampling position.
6. The improved UAV small target detection method based on YOLOv8n according to claim 1 is characterized in that: The specific steps of step 4 are as follows: the TA attention mechanism adopts a three-branch design to capture the interaction results between H dimension and C dimension, C dimension and W dimension, and W dimension and H dimension respectively; each branch processes the input tensor, and the input tensor is rotated, Z-pooled, convolved, batch normalized, and Sigmoid activated to obtain the interaction results between different dimensions. Finally, the processing results of the three branches are summed and averaged to form a comprehensive attention map; the specific calculation formula is as follows: Z-pool(x)=[Maxpool 0d ,Avgpool 0d ] Where Maxpool 0d 、Avgpool 0d Represent the maximum pooling layer and the average pooling layer respectively, R1 and R2 represent the rotation operations of the first branch and the second branch respectively, They represent the convolutional layers of the three branches respectively, σ() represents the sigmoid activation function, is the tensor obtained by the first branch through Z-pool, is the tensor obtained by the second branch through Z-pool, is the tensor obtained by the third branch through Z-pool, and X is the input tensor; The TA attention mechanism and the bidirectional weighted pyramid BiFPN are combined to design a DSFTA module; the DSFTA module uses the CBS module to downsample the shallow feature map, where the CBS module includes Conv convolution, BatchNorm batch normalization and Silu activation function, and then merges the downsampled shallow feature map with the deep feature map in the channel dimension. Subsequently, the CBS module is used to integrate the merged feature map, and the number of output channels is set to 256. Finally, the TA attention mechanism is introduced to enhance the feature information of small targets in the feature map.
7. The improved UAV small target detection method based on YOLOv8n according to claim 1 is characterized in that: The PIoU loss function is as follows: L PIoU =1-PIoU q=e -P L PIoUv2 =u(λq)×L PIou Where, d w1 d w2 d h1 d h2 is the absolute value of the distance between the corresponding edge of the predicted box and the real box, w gt 、h gt represents the height and width of the target box; P is the penalty factor; IoU is the intersection over union ratio, L PIoU is the loss function of PIoUv1, q is the hyperparameter for measuring the quality of the prediction box, u(x) is the attention function, and λ is the hyperparameter for controlling the attention mechanism of the function.
Citation Information
Cited By
Cell therapy product magnetic bead residue automatic detection system based on artificial intelligence
CN120672710A