A focused attention and fine fusion coding unmanned aerial vehicle small target detection method
Patent Information
- Application Number
- CN202410789495.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-06-19
AI Technical Summary
小目标在复杂的背景中容易被遮挡或混淆,如城市景观中的无人机可能与建筑物、树木或其他结构物混合在一起,从而难以分辨
[0026]将 C2f_DCNv2 和 所述FCA 聚焦协调注意力结合应用于颈部网络的第二至第五个 C2f 模块可获得最佳结果,通过对高度和宽度信息的多路径分析,实现了对空间维度的精细控制,从而提高了识别小尺度目标的能力。
Smart Images

Figure CN118675069B_ABST
Abstract
Description
Technical Field
[0001] This application relates to a method for small target detection in unmanned aerial vehicles (UAVs) that focuses on coordinated attention and fine fusion coding. Background Technology
[0002] Unmanned aerial vehicles (UAVs) are widely used in both civilian and military fields due to their flexibility, maneuverability, and cost-effectiveness. In civilian applications, UAVs are widely used for video surveillance, cargo transportation, and pesticide spraying, demonstrating unique advantages, especially when accessing remote or inaccessible areas such as mountains and densely populated urban areas. Militarily, UAVs are primarily used for reconnaissance, surveillance, and target location; their miniaturized design allows them to perform missions undetected by the enemy. While the development of UAV technology has brought tremendous convenience and benefits, the challenges of small target detection technology for UAVs are equally significant. The miniaturization of UAVs themselves and their use in complex environments necessitate efficient and accurate target detection capabilities. Traditional target detection techniques, such as feature-based image processing and deep learning methods, while effective to some extent, exhibit limitations when dealing with small targets captured by UAV aerial photography. Small targets are easily obscured or confused in complex backgrounds; for example, UAVs in urban landscapes may blend into buildings, trees, or other structures, making them difficult to distinguish. Background interference is another major challenge. In natural environments, weather conditions such as rain, fog, and clouds can affect image clarity; while in urban environments, multiple light sources, billboards, and moving vehicles can all become sources of interference. In conclusion, although drone technology has demonstrated unique advantages in many fields, the limitations of small target detection technology remain a problem that urgently needs to be addressed in current technological development. Summary of the Invention
[0003] In view of this, this application provides a method for detecting small targets on UAVs using focused coordinated attention and fine-grained fusion coding. By utilizing an Inner IoU loss function based on adjustable auxiliary bounding boxes, the regression process is effectively accelerated, improving the model's perception performance for small targets. A refined fusion coding (RFE) module is designed to perform detailed feature fusion and processing, which not only helps in the detection of small targets but also effectively controls the computational cost and number of parameters. This design aims to maximize detection performance while maintaining low computational complexity. A focused coordinated attention (FCA) module is designed, combining C2f_DCNv2 and focused coordinated attention (FCA) and applying it to the second to fifth C2f modules of the neck network. Through multi-path analysis of height and width information, fine-grained control of spatial dimensions is achieved, thereby improving the ability to identify small-scale targets.
[0004] The method for small target detection in UAVs that focuses on coordinated attention and fine-grained fusion coding in this application specifically includes the following steps:
[0005] (1) Obtain the publicly available dataset VisDrone2019 for drone aerial photography, and divide the dataset VisDrone2019 into a training set, a validation set, and a test set;
[0006] (2) Construct a small object detection model and train it using the partitioned VisDrone2019 dataset to obtain a trained small object detection model; the small object detection model includes: calculating the loss using the InnerIoU loss function based on adjustable auxiliary bounding boxes; designing a refined fusion coding RFE module and introducing it into the P2 small object detection layer; designing a focused coordinated attention FCA to analyze the height and width information of multiple paths, and combining the focused coordinated attention FCA with C2f_DCNv2 and applying it to the second to fifth C2f modules of the neck network;
[0007] (3) Input the real-time collected drone aerial photography data into the trained small target detection model to obtain the target detection results.
[0008] The VisDrone2019 public dataset for drone aerial photography as described in claim 1 includes:
[0009] The image sizes of the VisDrone2019 public drone aerial photography dataset were uniformly adjusted to 640×640 pixels, and the mosaic data augmentation technique was turned off in the last 10 rounds of training.
[0010] According to claim 1, the method for detecting small targets in unmanned aerial vehicles using focused coordinated attention and fine-grained fusion coding is characterized in that the small target detection model comprises:
[0011] An Inner IoU loss function based on adjustable auxiliary bounding boxes is adopted. The generation of auxiliary bounding boxes is controlled by a scaling factor, thereby calculating the loss and accelerating convergence. This effectively speeds up the regression process and improves the model's perception performance for small targets.
[0012] A refined fusion coding RFE module was designed, employing a scaled sequence feature fusion (SSFF) module to fuse the outputs of layers P3, P4, and P5 extracted from the Backbone network. Through the SSFF module, this module can efficiently merge feature maps of different sizes captured by layers P3, P4, and P5, capturing feature diversity from fine-grained to coarse-grained.
[0013] The feature maps of P3, P4 and P5 are first normalized to a uniform size, then upsampled and stacked, and finally integrated through 3D convolution. This module also optimizes the feature fusion process of the path aggregation network PANet structure.
[0014] A triple feature encoder (TFE) module is used to segment large, medium and small features, add large-size feature maps, and enlarge features to improve detailed feature information.
[0015] Before feature encoding, the number of each feature channel is adjusted to match the characteristics of the main scale;
[0016] For larger feature maps, after processing with a convolutional module that adjusts the number of channels to 1, a hybrid downsampling strategy (combining max pooling and average pooling) is employed, which helps to preserve both high-resolution details and the effectiveness and diversity of the cell image.
[0017] For smaller feature maps, the number of channels is adjusted through the convolution module, and then the nearest neighbor interpolation method is applied for upsampling to maintain detailed local features in the low-resolution image and prevent the loss of information of small target features.
[0018] Finally, the large, medium, and small feature maps of the same size after adjustment are convolved once, followed by channel concatenation; this improves the model's ability to represent targets using fine-grained processing mechanisms and ensures efficient encoding of multi-scale features.
[0019]
[0020] in This represents the output feature map of the TFE module. , and These represent feature maps for large, medium, and small sizes, respectively.
[0021] Depend on , and obtained by piecing together resolution and Same, but the number of channels is Three times;
[0022] A refined fusion coding RFE module is designed to perform detailed feature fusion and processing by introducing SSFF and TFE modules into the P2 small target detection layer, thereby improving target detection performance. By effectively combining feature maps at different levels, this multi-level feature fusion strategy enables the model to capture a wide range of features from acceptable to coarse, thus significantly improving target detection capabilities at all scales. The model performs feature fusion at layers P2 to P5 to ensure a balanced response to targets at all scales. In layer P2, no TFE module is introduced to upsample and fuse more refined features, which not only helps in the detection of small targets but also effectively controls the computational cost and number of parameters. This design aims to maximize detection performance while maintaining low computational complexity.
[0023] A Focused Coordinating Attention (FCA) technique was designed, which uses adaptive average pooling to reduce the input feature map into two paths along height H and width W. This operation effectively compresses the global information in each spatial dimension.
[0024] Subsequently, the obtained feature maps are reconnected in two dimensions, and convolutional operations are used to merge this information to encode the high-dimensional and wide-dimensional features, enabling the model to allocate different levels of attention to information at different spatial locations.
[0025] Next, the feature map is further refined using a 1 × 1 convolutional kernel and converted into an attention weight map using a sigmoid activation function. Then, it is split into two weight maps corresponding to the height and width directions. Through this element-wise multiplication operation, the Focus Coordination Attention (FCA) gives the network higher expressiveness on different paths and enables it to focus on the features with the most information.
[0026] Combining C2f_DCNv2 with the FCA focused attention and applying it to the second to fifth C2f modules of the neck network yields optimal results. Through multipath analysis of height and width information, fine control over spatial dimensions is achieved, thereby improving the ability to identify small-scale targets. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below.
[0028] Figure 1 This is a network structure diagram of the UAV small target detection method that focuses on coordinated attention and fine fusion coding in this application;
[0029] Figure 2These are schematic diagrams of two structures for constructing finely fused coded RFE in this application;
[0030] Figure 3 This application focuses on the Coordinating Attention (FCA) structure diagram. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0032] This invention belongs to the field of small target detection in unmanned aerial vehicles (UAVs), specifically relating to a method for small target detection in UAVs that focuses on coordinated attention and fine fusion coding.
[0033] The UAV small target detection method proposed in this application, which focuses on coordinated attention and fine fusion coding, is an improved detection method based on YOLOv8s. It uses the Inner IoU loss function of adjustable auxiliary bounding boxes to calculate the loss and accelerate convergence. A refined fusion coding RFE module is designed and introduced into the P2 detection layer to optimize the network structure and achieve higher performance with fewer parameters. Combining C2f_DCNv2 and the FCA focused coordinated attention and applying it to the second to fifth C2f modules of the neck network can obtain the best results. Through multi-path analysis of height and width information, fine control of spatial dimensions is achieved, thereby improving the ability to identify small-scale targets.
[0034] The small target detection method in this application includes:
[0035] (1) Obtain the publicly available dataset VisDrone2019 for drone aerial photography, and divide the dataset VisDrone2019 into a training set, a validation set, and a test set.
[0036] Specifically, the image size in the training dataset was uniformly adjusted to 640×640 pixels, and the mosaic data augmentation technique was turned off in the last 10 rounds of training.
[0037] (2) Construct a small target detection model and train it using the partitioned VisDrone2019 dataset to obtain a trained small target detection model. The small target detection model includes: the network structure diagram of the specific focus coordination attention and fine fusion coding UAV small target detection method is shown in Figure 1. The loss is calculated using the Inner IoU loss function based on adjustable auxiliary bounding boxes. A refined fusion coding RFE module is designed and introduced into the P2 small target detection layer. A focus coordination attention FCA is designed to analyze the height and width information of multiple paths and is combined with C2f_DCNv2 and applied to the second to fifth C2f modules of the neck network.
[0038] Specifically, this embodiment uses the Inner IoU loss function, which controls the generation of auxiliary bounding boxes through a scaling factor, thereby calculating the loss and accelerating convergence. This effectively speeds up the regression process and improves the model's perception performance for small targets. The calculation method of the Inner IoU loss function is as follows:
[0039]
[0040]
[0041]
[0042] in, , , , These are the left, right, top, and bottom coordinates of the anchor point, respectively. , , , These are the corresponding coordinates of the auxiliary bounding box; This is a scaling factor that controls the size of the auxiliary bounding box. The ground reality GT box and anchor points are represented as follows: and The center of the ground view frame and the center inside the ground view frame are respectively used as and It means, and and The center of the anchor point and the center of the internal anchor point are represented respectively; the width and height of the ground view GT frame are represented by... and express.
[0043] A refined fusion-encoded RFE module was designed and introduced into the P2 detection layer to optimize the network structure and achieve higher performance with fewer parameters. To further improve the accuracy of the small target detection model, a refined fusion-encoded RFE module was designed and introduced into the P2 small target detection layer. The SSFF and TFE modules are respectively described below. Figure 2As shown in (a) and (b), detailed feature fusion and processing are performed to improve target detection performance. By effectively combining feature maps at different levels, this multi-level feature fusion strategy enables the model to capture a variety of features from acceptable to coarse, thereby significantly improving the target detection capability at all scales. The model performs feature fusion at layers P2 to P5 to ensure a balanced response to targets at all scales. At layer P2, no TFE module is introduced to upsample and fuse more refined features, which not only helps in the detection of small targets but also effectively controls the amount of computation and the number of parameters. This design aims to maximize detection performance while maintaining low computational complexity.
[0044] The structure diagram of the fine fusion coding network is as follows: Figure 1 As shown,
[0045] A Focused Coordinating Attention (FCA) structure was designed, and the structure diagram of the Focused Coordinating Attention (FCA) is shown below. Figure 3As shown, the Coordinated Attention (FCA) model uses adaptive average pooling techniques X Avg Pool and Y Avg Pool. X Avg Pool performs average pooling in the horizontal direction, while Y Avg Pool performs it in the vertical direction, reducing the input feature map to two paths along height H and width W. GAP performs global average pooling on the input, effectively compressing global information in each spatial dimension. The permute operation reconnects the obtained feature maps in both dimensions, and convolutional (Conv) operations are used to merge this information, encoding high-dimensional and wide-dimensional features. This allows the model to assign different levels of attention to information at different spatial locations. Next, a 1 × 1 convolutional (Cov) is used to further refine the feature map, and a sigmoid activation function is used to convert it into an attention weight map. This weight map is then split into two weight maps corresponding to the height and width directions. The split operation on the left divides the output of the convolutional (Conv) into two parts, X and Y. The split operation on the right divides the output of the sigmoid into two parts, X Weight and Y Weight. The right-hand multiplication Mul: Multiply X and X Weight to obtain the weighted X. Multiply Y and Y Weight to obtain the weighted Y. The middle multiplication Mul: Multiply the right-hand convolution and the weighted feature map. The left-hand multiplication Mul: Multiply the outputs of X and Y. The final multiplication Mul: Multiply the results of the left and middle multiplications to obtain the final output. Through this element-wise multiplication operation Mul, the Focused Coordinating Attention (FCA) gives the network higher expressiveness on different paths and enables it to focus on the features with the most information. Combining C2f_DCNv2 and the FCA attention mechanism and applying them to the second to fifth C2f modules of the neck network yields the best results. Through multipath analysis of height and width information, fine control of spatial dimensions is achieved, thereby improving the ability to identify small-scale targets.
[0046] (3) Input the real-time collected drone aerial photography data into the trained small target detection model to obtain the target detection results.
[0047] The target detection method of this application maintains high detection performance while having low computational burden and storage requirements. Moreover, compared with other algorithms in the prior art, the detection method of this application embodiment can more accurately identify small and distant targets, with fewer missed detections and false detections.
[0048] To further verify the effectiveness of the algorithm in this embodiment, the detection performance of the proposed method was tested on the VisDrone2019 dataset. The proposed model was compared with other classic models, and the comparative experimental results are shown in Tables 1 and 2. The experimental results in Tables 1 and 2 show that the proposed method outperforms other models in terms of detection capability. In this application, the applicant comprehensively evaluated the performance of the proposed small object detection model with several existing YOLO series and classic models through comparative experiments. The models compared in the experiments included different versions of the YOLO series and classic models such as Faster R-CNN, SSD, RetinaNet, FSAF, ATSS, and Fcos. The performance of each model was evaluated by detection precision, recall, mean precision (mAP0.5 vs. mAP0.5:0.95), and model size (parameters). Our model outperforms other comparable models across all key metrics. Specifically, our model achieves 55.3% precision, 42.5% recall, 45.1% mAP0.5, and 27% mAP0.5:0.95, with a parameter count of 9.1 MB. In contrast, while YOLOv4 boasts a higher recall of 48.5%, its performance in precision, mAP0.5, and mAP0.5:0.95 is less impressive. Furthermore, our model has a significantly lower parameter count than all models except the YOLOv5s series. This demonstrates that our small target detection model maintains high detection performance while requiring less computation and storage. The results also highlight the potential for practical applications, particularly in scenarios with limited computational resources. The algorithm in this embodiment can more accurately identify small and distant targets with fewer false negatives and false positives.
[0049] Table 1 shows the detection results of the classic YOLO series models and our proposed method.
[0050] YOLOv3 53.8 43.2 41.7 23.1 103.7 YOLOv4 36.2 48.5 42.2 25.6 64.4 YOLOv5s 46.5 34.7 34.5 19.1 7.2 FE-YOLOv5 ﹣ ﹣ 37.0 21.0 9.14 YOLOv5s_MSES ﹣ ﹣ 41.9 23.7 5.6 YOLOv7 51.3 42.3 40.1 21.7 64 YOLOv8s 50.8 38.1 39.3 23.5 11.1 Algorithm in this embodiment 55.3 42.5 45.1 27 9.1
[0051] Table 2 shows the detection results of the classic model and our proposed method.
[0052] Faster-R-CNN 21.7 36.9 22.6 11.3 27.8 37.6 Cascade R-CNN 24.1 39.3 26.3 13.6 28.1 40.2 SSD 15.1 35.6 24.5 15.2 30.2 38.7 RetinaNet 15.2 26.2 15.3 6.1 25.3 34.4 FSAF 20.7 36.2 20.6 10.4 29.5 38.3 ATSS 20.8 34.2 22.8 11.5 31.5 36.6 Fcos 17.9 30.4 18.3 9.2 27.6 35.4 Algorithm in this embodiment 27.0 45.1 26.5 16.9 37.2 47.3
Claims
1. A method for small target detection in unmanned aerial vehicles (UAVs) that focuses on coordinated attention and fine-grained fusion coding, characterized in that, Includes the following steps: (1) Obtain the publicly available dataset VisDrone2019 for drone aerial photography, and divide the dataset VisDrone2019 into a training set, a validation set, and a test set; (2) Construct a small object detection model and train it using the divided VisDrone2019 dataset to obtain a trained small object detection model; The small target detection model includes: using an Inner IoU loss function based on adjustable auxiliary bounding boxes, controlling the generation of auxiliary bounding boxes through a scaling factor, and thus calculating the loss; The Scale Sequence Feature Fusion (SSFF) module is used to fuse the outputs of layers P3, P4, and P5 extracted from the Backbone network. The feature maps of P3, P4, and P5 are first normalized to a uniform size, then upsampled and stacked, and finally integrated through 3D convolution. The Triple Feature Encoder (TFE) module is used to segment large, medium, and small features and add large-size feature maps. Before feature encoding, the number of channels for each feature is adjusted to match the characteristics of the main scale. For large feature maps, a convolutional module with the number of channels adjusted to 1 is used, followed by a hybrid downsampling strategy that combines max pooling and average pooling. For small feature maps, the number of channels is adjusted using a convolutional module, and then upsampling is performed using nearest neighbor interpolation. Finally, the large, medium, and small feature maps of the same adjusted size are convolved once, and then the channels are concatenated. in This represents the output feature map of the TFE module. , and These represent feature maps for large, medium, and small sizes, respectively. Depend on , and obtained by piecing together resolution and Same, but the number of channels is Three times; A refined fusion coding RFE module was designed by introducing SSFF and TFE modules into the P2 small target detection layer. By effectively combining feature maps at different levels, the model performs feature fusion in layers P2 to P5. In layer P2, the TFE module was not introduced to upsample and fuse more refined features. A Focused Coordinating Attention (FCA) mechanism is designed. FCA uses adaptive average pooling to reduce the input feature map to two paths along height H and width W. Subsequently, the obtained feature maps are reconnected in both dimensions, and convolutional operations are used to merge this information. Next, a 1 × 1 convolutional kernel is used to further refine the feature map, and a sigmoid activation function is used to convert it into an attention weight map, which is then split into two weight maps corresponding to the height and width directions. The C2f_DCNv2 and the FCA attention mechanism are combined and applied to the second through fifth C2f modules of the neck network. (3) Input the real-time collected drone aerial photography data into the trained small target detection model to obtain the target detection results.
2. The method for small target detection in UAVs based on focused coordinated attention and fine fusion coding according to claim 1, characterized in that, It also includes image preprocessing for the VisDrone2019 public drone aerial photography dataset: The image sizes of the VisDrone2019 public drone aerial photography dataset were uniformly adjusted to 640×640 pixels, and the mosaic data augmentation technique was turned off in the last 10 rounds of training.
Citation Information
Patent Citations
Multi-scale attention-fused traffic helmet small target detection system and method
CN116665156A
Unmanned aerial vehicle aerial photography small target detection method based on multi-level feature fusion
CN118072146A