Rapid small target-oriented unmanned aerial vehicle image detection method and related equipment
By using the method of fine feature extraction module and multiple feature aggregation network in the image detection of drone, the problem of accuracy and robustness of small object detection in aerial images of drone is solved, and more efficient and accurate object detection effects are achieved.
Patent Information
- Application Number
- CN202510192345.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-13
AI Technical Summary
The existing object detection algorithms have problems of reduced accuracy and insufficient robustness when processing aerial images of drones, especially in small object detection, and it is difficult to effectively deal with complex backgrounds and lighting changes.
A fast drone image detection method for small-objectives is proposed, using fine feature extraction module and multiple feature aggregation network, and the efficiency and accuracy of feature extraction and fusion are improved through multi-branch structure and receptive field convolution weighting optimization.
It significantly improves the detection performance of the model for small targets, reduces the risks of false detection and missed detection, enhances the processing ability of complex scenarios, and achieves more efficient feature fusion and detection efficiency.
Smart Images

Figure CN120147900A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to a fast UAV image detection method for small targets and related devices. Background Art
[0002] Traditional manual target detection methods lack generality and generalization ability and are difficult to accurately detect targets. With the booming development of artificial intelligence (AI) in many fields, object detection algorithms based on deep learning have been widely applied. Compared with traditional methods, these algorithms have powerful feature extraction capabilities and have become the mainstream of object detection research and application. Currently, deep learning-based methods are mainly divided into two categories: two-stage and one-stage detection algorithms.
[0003] The two-stage detection algorithm divides the detection process into two main stages. The first stage is to generate candidate regions in the input image, and the second stage is to perform detailed classification and precise localization on these generated candidate regions. Classic two-stage algorithms mainly include the R-CNN series, such as Faster R-CNN, Mask R-CNN, and Cascade R-CNN. The two-stage algorithm ensures excellent detection performance, but due to the two independent stages, it requires more inference time and computing resources. Different from the two-stage detection algorithm, the one-stage detection algorithm does not need to generate candidate regions, but directly predicts the input feature map through a dense grid or anchor boxes. Classic one-stage algorithms mainly include the YOLO series, SSD, and RetinaNet. Although the one-stage detection algorithm is slightly less accurate, it has a faster detection speed. Therefore, the one-stage algorithm has been widely used in real-time detection tasks.
[0004] Existing general object detection algorithms (such as R-CNN, YOLO, and SSD) have achieved remarkable results in various vision tasks. However, due to the uniqueness of UAV aerial images, these algorithms have obvious limitations when dealing with UAV aerial images. First, these algorithms usually rely on specific input image quality and feature distribution, but due to the special perspective, flight altitude, and complex background information of UAV aerial images, the algorithms face greater challenges in feature extraction and information recognition. In addition, such images often contain extremely small targets, which occupy fewer pixels in the image and are easily interfered by noise or background. When existing algorithms detect such tiny targets, their accuracy drops significantly, and they may even be completely missed. At the same time, due to the changing lighting conditions and constantly changing shooting angles in UAV images, the robustness of existing algorithms has also been tested, and false detection and missed detection problems may occur in different scenarios.
[0005] In view of the characteristics of UAV aerial images, a large number of related studies have made progress. For example, Chalavadi et al. proposed a novel multi-scale object detection network (mSO-DANet), which learns the context information of different objects by using hierarchical dilated convolution to improve the detection effect. However, the strategy of uniformly processing different regions in this method wastes a large amount of computing resources in unimportant regions, limiting the further improvement of detection performance. Zhang et al. proposed a CFA structure for parallel fusion of feature maps of different scales to achieve high-quality feature fusion results, and used the LASPP module to expand the receptive field and maintain sensitivity to different receptive fields. However, these improved methods still have limitations in small object detection, especially in the extraction of feature detail information, resulting in weak detection performance for small-scale objects. Summary of the Invention
[0006] In view of the problems such as a large proportion of small targets and complex background interference in UAV images, the present invention proposes a fast UAV image detection method and related devices for small targets, effectively overcoming the limitations of existing algorithms in processing UAV images.
[0007] In a first aspect, the present invention provides a fast UAV image detection method for small targets, including:
[0008] Obtaining a UAV image to be measured;
[0009] Inputting the UAV image to be measured into a preset target detection model to output a detection result; the detection result includes the target type and the corresponding confidence; the preset target detection model includes a backbone network, a neck network, and a detection head; the backbone network is used to extract multi-level features of the UAV image to be measured; the neck network is used to perform weighted fusion on the multi-level features output by the backbone network along the top-down and bottom-up fusion paths respectively to generate multi-scale fusion features; the detection head is used to perform target detection on each scale of fusion features.
[0010] Wherein, a fine feature extraction module is used in the backbone network to extract multi-level features of the UAV image to be measured; the processing process of the fine feature extraction module for the input feature map includes: using multiple parallel branches to process the input feature map respectively; wherein, different dilated rates of convolution are used in different branches; splicing the outputs of the multiple parallel branches to generate a first feature map; using group convolution to process the first feature map to obtain receptive field spatial features, using average pooling and group convolution to process the first feature map successively to generate receptive field weights; convolving the product of the receptive field weights and the receptive field spatial features to obtain the output of the fine feature extraction module.
[0011] In this embodiment, the refined feature extraction module designed by the present invention is adopted in the backbone network to efficiently mine multi-scale target information. This module adopts a multi-branch structure and enhances the network's expression ability for different scales and context semantics by dynamically adjusting the feature capture range. The captured feature information is processed by the receptive field convolution. The receptive field convolution extracts the spatial features of its receptive field and performs weighted calculation, highlighting the foreground information and suppressing the background interference, significantly improving the integrity and extraction efficiency of the target features, breaking through the limitations of the traditional fixed receptive field convolution, and providing richer and more accurate feature support for subsequent detection tasks.
[0012] Further, in the neck network, a multi-level feature aggregation network is used to perform weighted fusion on the multi-level features output by the backbone network along the top-down and bottom-up fusion paths respectively to generate multi-scale fusion features; the multi-level feature aggregation network includes multiple multi-level feature fusion modules, and the multi-level feature fusion modules are used to assign weights to all input feature maps and perform weighted fusion;
[0013] The bottom-up fusion path includes a first multi-level feature fusion module, a second multi-level feature fusion module and two convolutional layers; among them, the output of the sub-high-level feature map output by the backbone network after passing through one of the convolutional layers, the sub-high-level feature map and the highest-level feature map serve as the inputs of the first multi-level feature fusion module; the output of the low-level feature map output by the backbone network after passing through the other convolutional layer, the low-level feature map and the output of the first multi-level feature fusion module serve as the inputs of the second multi-level feature fusion module; the low-level feature map and the sub-high-level feature map are adjacent layer feature maps;
[0014] The top-down fusion path includes a third multi-level feature fusion module, a fourth multi-level feature fusion module and a fifth multi-level feature fusion module; among them, the output of the low-level feature map output by the backbone network after passing through the other convolutional layer and the output of the second multi-level feature fusion module serve as the inputs of the third multi-level feature fusion module; the output of the sub-high-level feature map output by the backbone network after passing through one of the convolutional layers, the output of the first multi-level feature fusion module and the output of the third multi-level feature fusion module serve as the inputs of the fourth multi-level feature fusion module; the highest-level feature map output by the backbone network and the output of the fourth multi-level feature fusion module serve as the inputs of the fifth multi-level feature fusion module; the outputs of the third multi-level feature fusion module, the fourth multi-level feature fusion module and the fifth multi-level feature fusion module are the multi-scale fusion features.
[0015] This embodiment introduces a multi - feature aggregation network into the neck network. In this network, by adding new aggregation paths and adjusting the input size, operations such as feature recombination of the input information are performed to achieve multi - level integration of features. The multi - feature fusion module in this aggregation network weighs the contributions and importance of the input feature maps, and uses these weights to perform weighted fusion on the input feature maps, thereby obtaining the final fused feature map to achieve more effective feature fusion and obtain better feature representations.
[0016] Further, the multi - feature fusion module is used to assign weights to all input feature maps and perform weighted fusion, specifically including: converting each input feature map into a corresponding vector through a tiling operation; processing each vector using a fully - connected layer to generate a scalar for each input feature map; normalizing each scalar using the softmax function to generate the weight for each input feature map; and performing weighted fusion on all input feature maps based on the weights of each input feature map to obtain the output of the multi - feature fusion module.
[0017] Further, the backbone network includes a convolutional layer, three feature extraction layers, and a spatial pyramid pooling module connected in sequence; the feature extraction layer includes a fine - grained feature extraction module and a C2f module; the outputs of the first two feature extraction layers and the output of the spatial pyramid pooling module are the multi - level features output by the backbone network.
[0018] Further, the spatial pyramid pooling module adopts the SPPF module.
[0019] Further, the object detection model uses the Yolov8 network as the basic framework, adds a small - scale detection head P2, and removes the large - scale detection head P5.
[0020] Further, the number of C2f modules in the last two feature extraction layers is reduced from 6 to 3.
[0021] In a second aspect, the present invention provides a fast unmanned aerial vehicle (UAV) image detection device for small targets, including:
[0022] An acquisition unit for acquiring the UAV image to be measured;
[0023] A detection unit for inputting the UAV image to be measured into a preset object detection model and outputting a detection result; the detection result includes the target type and the corresponding confidence level; the preset object detection model includes a backbone network, a neck network, and a detection head; the backbone network is used to extract multi - level features of the UAV image to be measured; the neck network is used to perform weighted fusion on the multi - level features output by the backbone network along the top - down and bottom - up fusion paths respectively to generate multi - scale fusion features; the detection head is used to perform object detection on each scale of the fusion features.
[0024] Among them, a fine feature extraction module is used in the backbone network to extract multi-level features of the UAV image to be measured; the processing process of the fine feature extraction module for the input feature map includes: using multiple parallel branches to process the input feature map respectively; among them, convolutions with different dilation rates are adopted in different branches; the outputs of the multiple parallel branches are spliced to generate a first feature map; group convolution is used to process the first feature map to obtain receptive field spatial features, and average pooling is used to process the first feature map to generate receptive field weights; the product of the receptive field weights and the receptive field spatial features is convolved to obtain the output of the fine feature extraction module.
[0025] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in the first aspect is implemented.
[0026] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0027] The beneficial effects of the present invention are as follows:
[0028] (1) Improve detection performance: The present invention realizes the efficient extraction of detailed features through the fine feature information extraction module, enhances the attention to the target area, significantly improves the model's ability to capture important foreground information, and thus significantly reduces the risks of false detection and missed detection. In addition, the proposed multi-feature aggregation network fully integrates multi-scale features, improves the perception ability of the target, and integrates rich feature information, thereby effectively improving the detection accuracy.
[0029] (2) Improve detection efficiency: Since it is difficult to identify targets with a large proportion of small pixels in UAV images, the present invention proposes to add a small target detection head, and at the same time remove the last two feature extraction layers of the backbone network and the large target detection head, and reduce the number of times the C2f module in the backbone network is used. These improvements significantly improve the detection accuracy while greatly reducing the number of parameters, effectively avoiding feature degradation, improving the detection efficiency, and making the model more lightweight while maintaining high accuracy.
[0030] (3) It can handle complex scenarios: Since it is difficult to identify targets with a large proportion of small pixels in UAV images, the present invention proposes a design of adding a small target detection head. At the same time, the last two feature extraction layers of the backbone network and the large target detection head are removed, and the number of uses of the C2f module in the backbone network is reduced. These optimization measures significantly improve the detection accuracy while greatly reducing the number of parameters, effectively avoiding feature degradation, and improving the detection efficiency, making the model more lightweight while maintaining high accuracy.
[0031] (4) Achieve efficient fusion: A multiple feature fusion module is designed in the neck network for concatenating input feature maps and efficiently fusing them. This module performs tiling and normalization processing on the input feature maps, and assigns weights according to the contributions of each feature, thereby achieving efficient feature fusion. Description of the Drawings
[0032] Figure 1 It is a schematic flow chart of a fast UAV image detection method for small targets provided by an embodiment of the present invention;
[0033] Figure 2 It is the fine feature extraction module DFE provided by an embodiment of the present invention;
[0034] Figure 3 It is the multiple feature aggregation network MFIN provided by an embodiment of the present invention;
[0035] Figure 4 It is the receptive field convolution RFConv provided by an embodiment of the present invention;
[0036] Figure 5 It is the network architecture of a target detection model provided by an embodiment of the present invention;
[0037] Figure 6 It is the change curve of mAP0.5 and mAP0.5:0.95 of each model provided by an embodiment of the present invention with the increase of the number of training rounds;
[0038] Figure 7 It is a schematic structural diagram of a fast UAV image detection device for small targets provided by an embodiment of the present invention;
[0039] Figure 8 It is a structural block diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0040] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] The present invention combines a Detailed Feature Extraction Module (DFE), a Multi level feature integration network (MFIN), and a Multi-Level Feature Integration Network (MFF), and significantly improves the network's perception ability for small targets by improving the detection head design. While improving the detection head, the present invention reduces the number of times the C2f module is used in the backbone network and removes the last two feature extraction layers of the backbone network, greatly improving the detection performance while keeping the model lightweight, and effectively overcoming the limitations of traditional object detection methods in specific scenarios.
[0042] In one embodiment, as Figure 1 shown, the embodiment of the present invention provides a fast UAV image detection method for small targets, including the following steps:
[0043] S101: Obtain the UAV image to be measured;
[0044] S102: Input the UAV image to be measured into a preset object detection model and output a detection result; the detection result includes the target type and the corresponding confidence level; the preset object detection model includes a backbone network, a neck network, and a detection head; the backbone network is used to extract multi-level features of the UAV image to be measured; the neck network is used to perform weighted fusion on the multi-level features output by the backbone network along the top-down and bottom-up fusion paths respectively to generate multi-scale fusion features; the detection head is used to perform object detection on each scale of fusion features;
[0045] Among them, a multi-level feature extraction module is used in the backbone network to extract features of the UAV image to be measured; the processing process of the multi-level feature extraction module for the input feature map includes: using multiple parallel branches to process the input feature map respectively; among them, convolutions with different dilation rates are adopted in different branches; the outputs of the multiple parallel branches are concatenated to generate a first feature map; group convolution is used to process the first feature map to obtain receptive field spatial features, and average pooling is used to process the first feature map to generate receptive field weights; the product of the receptive field weights and the receptive field spatial features is convolved to obtain the output of the multi-level feature extraction module.
[0046] Specifically, the object detection model can adopt the YOLO series network (including YOLOv8) as the benchmark framework, and then perform the above improvements on the backbone network of the selected YOLO network.
[0047] As an implementable manner, as Figure 2 shown in a multi-level feature extraction module. First, the feature map is input into 3 parallel branches to obtain different receptive fields, while establishing a rich context relationship and improving the feature extraction ability for small targets, and then the information obtained from each branch is concatenated; then, these features are input into the Receptive Field Convolution (RFConv) module. This process can be expressed by formulas (1) and (2) as follows:
[0048] F 1 = Concat(DConv 1 (F in ), DConv 3 (F in ), DConv 5 (F in )) (1)
[0049] F 2 = RFConv(F 1 ) (2)
[0050] Among them, F in represents the input feature map, F 1 represents the feature obtained by concatenating the features extracted under multiple receptive fields. F 2 represents the output result of the DFE module.
[0051] Furthermore, the internal processing process of RFConv includes: first, the input feature map F 1Extract receptive field spatial features through group convolution; subsequently, use average pooling (AvgPool2d) and group convolution (GConv) to calculate the weights of each receptive field feature. Finally, multiply these weights by the corresponding receptive field features to significantly highlight foreground information. This process can be expressed by formulas (3), (4), and (5) as follows:
[0052] λ = GConv(AvgPool2d(F 1 ))(3)
[0053] F' = RELU(BN(GConv(F 1 )))(4)
[0054] F 2 = Conv(λ × F')(5)
[0055] Wherein, BN represents normalization, and RELU is an activation function. F' represents the captured receptive field spatial features, and λ represents the calculated weights.
[0056] In the embodiment of the present invention, the DFE module captures information from different receptive fields of the feature map through a multi-branch structure and applies convolutions with different dilation rates in each branch to construct rich context relationships. Subsequently, the features captured by the multi-branch are concatenated and input into the receptive field convolution (RFConv). RFConv highlights foreground features through weight calculation and suppresses background interference, significantly improving the feature extraction ability in object detection. RFConv enhances foreground features through a weight calculation mechanism while effectively suppressing the interference of complex backgrounds. In addition, since RFConv operates on the features within each receptive field separately, the weight values of each receptive field are different from each other. Therefore, this mechanism not only solves the problem of parameter sharing of traditional convolutional kernels but also further improves the feature expression ability.
[0057] In one embodiment, the backbone network includes a convolutional layer, three feature extraction layers, and a spatial pyramid pooling module connected in sequence; the feature extraction layer includes a fine feature extraction module and a C2f module; the outputs of the first two feature extraction layers and the output of the spatial pyramid pooling module are the multi-level features output by the backbone network.
[0058] Due to the significant scale variation problem in UAV images, in order to further improve the object detection performance, in one embodiment, as Figure 3As shown in the figure, an embodiment of the present invention designs a new neck network, namely, the Multi-Feature Aggregation Network (MFIN). This network uses multiple multi-feature fusion modules, which assign weights to all input feature maps and perform weighted fusion to fully exploit and reuse the feature maps containing detailed information, effectively alleviating the problem of information loss in the fusion path of feature maps. In the neck network, the multi-feature aggregation network is used to perform weighted fusion on the multi-level features output by the backbone network along the top-down and bottom-up fusion paths respectively, generating multi-scale fusion features. Among them, bottom-up means transmitting information from high-level to low-level; similarly, top-down means transmitting information from low-level to high-level. The specific process is as follows:
[0059] The bottom-up fusion path includes a first multi-feature fusion module, a second multi-feature fusion module, and two convolutional layers; among them, the output of the sub-high-level feature map output by the backbone network after passing through one of the convolutional layers, the sub-high-level feature map, and the high-level feature map are used as the inputs of the first multi-feature fusion module; the output of the low-level feature map output by the backbone network after passing through the other convolutional layer, the low-level feature map, and the output of the first multi-feature fusion module are used as the inputs of the second multi-feature fusion module; the low-level feature map and the sub-high-level feature map are adjacent layer feature maps;
[0060] Specifically, for example, taking the multi-level features output by the backbone network as feature maps C2, C3, and C4 respectively, for the first multi-feature fusion module MFF3', its three inputs are: (1) The output of the new fusion path: The feature map C3 is recombined through a 1×1 convolution to extract key information and input it into the fusion module MFF3'. (2) The directly input feature map: The feature map C3 output by the backbone network is directly input into the fusion module MFF3' to retain low-level features and avoid information loss caused by convolutional operations. (3) The feature map from a higher level: The feature map C4 is upsampled to adjust the spatial size to match the feature map C3 and then input into the fusion module MFF3'.
[0061] The top-down fusion path includes the third multi-feature fusion module, the fourth multi-feature fusion module, and the fifth multi-feature fusion module. Among them, the output of the low-level feature map output by the backbone network after passing through another convolutional layer and the output of the second multi-feature fusion module serve as the input of the third multi-feature fusion module. The output of the sub-high-level feature map output by the backbone network after passing through one of the convolutional layers, the output of the first multi-feature fusion module, and the output of the third multi-feature fusion module serve as the input of the fourth multi-feature fusion module. The output of the top-level feature map output by the backbone network and the output of the fourth multi-feature fusion module serve as the input of the fifth multi-feature fusion module. The outputs of the third multi-feature fusion module, the fourth multi-feature fusion module, and the fifth multi-feature fusion module are the multi-scale fusion features.
[0062] Specifically, still taking the multi-level features output by the backbone network as the feature maps C2, C3, and C4 as an example, for the fourth multi-feature fusion module MFF3”, its three inputs are: (1) The feature map from the low level: the output of the third multi-feature fusion module is downsampled to adjust the size to match the feature map C3, and then input into the fusion module MFF3”; (2) The output of the new fusion path: the feature map C3 undergoes 1×1 convolution to reorganize the features, and the key information extracted (i.e., Conv(C3)) is input into the fusion module MFF3”; (3) The output of the fusion module MFF3'. By fusing these three inputs, the module realizes the effective combination of multi-level features and provides richer and more diverse feature expressions for the subsequent network layers.
[0063] In the embodiments of the present invention, for two adjacent intermediate-level feature maps, MFIN respectively adds a new fusion path, and this path reorganizes the features of the input feature map through 1×1 convolution to extract key information. The output of the new fusion path is input into the multi-feature fusion module (MFF) in the bottom-up path, and is also input into the multi-feature fusion module (MFF) in the top-down path, for providing more key semantic information and detailed features.
[0064] In the entire feature fusion network, the multi-feature fusion module (MFF module) is responsible for concatenating the feature maps input from different branches, and realizes weighted fusion by evaluating the importance of each input feature and assigning weights. In one embodiment, the embodiments of the present invention provide a multi-feature fusion module.
[0065] Taking Figure 3 the shown fusion module MFF3' as an example, the process of calculating weights and performing weighted fusion of the multi-feature fusion module provided by the embodiments of the present invention is as shown in 4. For the 3 input feature maps, they are respectively denoted as F 1 、F 2 、F 3, it is converted into a vector through a tiling operation, denoted as f 1 , f 2 , f 3 respectively; subsequently, these vectors pass through a fully connected layer respectively to generate the scalar f 1 ' of each feature map, f 2 '; then, these feature map scalars are normalized through the softmax function to obtain the weights ω 3 of each input feature map, ω 1 , ω 2 , ω 3 respectively. Finally, the final fused feature map is obtained through weighted fusion, and the specific calculation process is shown in Formula (6) and Formula (7).
[0066] ω 1 , ω 2 , ω 3 = softmax(f 1 ', f 2 ', f 3 ')(6)
[0067] P 3 = ω 1 × F 1 + ω 2 × F 2 + ω 3 × F 3 (7)
[0068] Among them, P 3 is the final output result of the fusion module MFF3'.
[0069] In drone aerial images, the targets are usually small in size and densely distributed, which poses great challenges to the detection algorithm. On the one hand, the high proportion of small targets leads to the easy compression and loss of their features in the deep convolutional network; on the other hand, drone detection requires fast response to meet the needs of real-time applications, and it is difficult for traditional methods to balance speed and accuracy. To address the above problems, based on the above embodiments, the embodiments of the present invention further improve the detection head. Taking YOLOv8 as an example, in the embodiments of the present invention, first, a smaller-scale detection head P2 is added to specifically handle the feature extraction and detection tasks of small targets. By introducing a finer-grained feature scale, the detection ability of the network for small targets is effectively enhanced. At the same time, to reduce the complexity of the network, the last two feature extraction layers in the backbone network and the detection head P5 with the largest detection scale are removed. This modification improves the detection performance for small targets, reduces the number of parameters, and thus optimizes the computational efficiency.
[0070] In addition, in the backbone network of the baseline model YOLOv8, the C2f module is reused. The present invention believes that this approach may lead to feature degradation. Therefore, the present invention reduces the number of C2f modules in the fourth and sixth layers of the backbone network from 6 to 3.
[0071] In the embodiments of the present invention, on the one hand, reducing the design of the detection head and the feature extraction layer significantly reduces the number of parameters of the network, making the model more lightweight and suitable for deployment on resource-constrained devices. On the other hand, the newly added small-scale detection head improves the feature capture ability for small targets, significantly enhancing the detection speed while maintaining high detection accuracy. This balance between speed and accuracy is particularly suitable for the real-time detection requirements in the UAV aerial photography scenario.
[0072] A fast and efficient UAV image detection method for small targets provided by the present invention designs an efficient detailed feature extraction module (DFE), which captures feature information under different receptive fields through a multi-branch structure to construct a context relationship with rich semantics. The extracted features are further weighted and optimized through receptive field convolution (RFConv) to highlight foreground features and suppress background interference, significantly enhancing the pertinence and distinctiveness of feature expression. To fully fuse multi-level features and reduce information loss during the feature fusion process, a multi-feature aggregation network (MFIN) is proposed. This network introduces an additional aggregation path and combines a multi-feature fusion module (MFF) to perform weighted fusion on the feature maps, balancing the contributions of feature maps at different scales and optimizing the overall feature expression. In addition, to enhance the model's detection ability for small targets, the present invention optimizes the scale design of the detection head, removes the last two layers of feature extraction layers in the backbone network, and reduces the number of times the C2f module is used to avoid feature degradation and improve detection efficiency. These improvements significantly reduce the number of parameters while taking into account detection accuracy and operating efficiency.
[0073] To verify the effectiveness of the detection method of the present invention, the present invention uses Yolov8 as the baseline model and improves Yolov8 with all the above improvement points of the detection method of the present invention to form the network architecture of the object detection model as shown in Figure 5 The network architecture of the object detection model shown in the figure includes key modules such as a detailed feature extraction module, a multi-feature aggregation network, and a multi-feature fusion module, aiming to effectively improve the accuracy and efficiency of small target detection. In this object detection model, the multi-level features extracted by the DFE are input into the multi-feature aggregation network (MFIN) for fusion. By adding additional branch connections, the MFIN can effectively retain the detailed information in the feature map and ensure the full combination of multi-level features. In addition, a multi-feature fusion module (MFF) is further introduced to weigh the importance of feature maps at different levels and assign adaptive weights to strengthen the contribution of key features. Finally, the fused feature maps are respectively sent to the corresponding detection heads for target prediction to generate the final detection results.
[0074] The experimental setup of the present invention is as follows: All experiments were conducted on the Ubuntu 20.04.6 LTS operating system to verify the performance of the proposed object detection model. The hardware configuration includes two NVIDIA GeForce RTX 3090 GPUs, each with 24 GB of memory. The experiments were implemented using the PyTorch deep learning framework, with Python version 3.9.19, PyTorch version 2.1.0, and CUDA version 12.4. In the training phase, the input image size was uniformly set to 640×640 pixels. The Stochastic Gradient Descent (SGD) optimizer with momentum was used, with an initial learning rate of 0.01, a momentum parameter of 0.937, and a weight decay coefficient of 0.0005. The batch size was set to 8. The present invention uses the publicly available drone dataset VisDrone2019-DET, and the total number of training iterations is 300. In the experimental result table, all data are presented as percentages except for the parameters expressed in megabytes. The experimental results are shown in Table 1. In addition, in this experiment, the curves of mAP0.5 and mAP0.5:0.95 with the increase in the number of training rounds were also plotted. As Figure 6 shown.
[0075] Table 1 Detection results of each model on the VisDrone 2019-DET dataset
[0076]
[0077] A comprehensive comparison was made between the present invention and the baseline model Yolov8 series. As can be seen from Table 1, the present invention achieved the best results in the detection accuracy of most of the 10 categories in the VisDrone2019-DET dataset, and the mAP0.5 detection accuracy of the present invention also reached the highest level. It is worth mentioning that compared with the baseline model Yolov8s, the present invention reduced the number of parameters by 30%, thus achieving faster drone target detection.
[0078] From Figure 6 Figures (a) and (b) below, it can be seen that although the Yolov8 series has a relatively fast convergence speed, its continuous learning ability is limited. This analysis shows that the present invention surpasses almost the entire Yolov8 series in the drone target detection task, while significantly reducing the number of parameters, demonstrating good practical application potential.
[0079] Based on the same inventive concept, as Figure 7 shown, an embodiment of the present invention also provides a fast drone image detection device for small targets, including: an acquisition unit and a detection unit.
[0080] Among them, the acquisition unit is used to acquire the drone image to be detected; the detection unit is used to input the drone image to be detected into a preset target detection model and output a detection result; the detection result includes a target type and a corresponding confidence level; the preset target detection model includes a backbone network, a neck network, and a detection head; the backbone network is used to extract multi-level features of the drone image to be detected; the neck network is used to perform weighted fusion on the multi-level features output by the backbone network along top-down and bottom-up fusion paths respectively to generate multi-scale fusion features; the detection head is used to perform target detection on each scale of fusion features; among them, a fine feature extraction module is used in the backbone network to extract multi-level features of the drone image to be detected; the processing process of the fine feature extraction module for the input feature map includes: using multiple parallel branches to process the input feature map respectively; among them, convolutions with different dilation rates are adopted in different branches; the outputs of the multiple parallel branches are spliced to generate a first feature map; the first feature map is processed using group convolution to obtain receptive field spatial features, and the first feature map is processed using average pooling to generate receptive field weights; the product of the receptive field weights and the receptive field spatial features is convolved to obtain the output of the fine feature extraction module.
[0081] It should be noted that the fast drone image detection device for small targets provided in the embodiments of the present invention is to implement the above method, and its functions can be specifically referred to the above method embodiments, which will not be elaborated here.
[0082] Figure 8 An example of the physical structure diagram of an electronic device is shown in Figure 8As shown in the figure, the electronic device may include: a processor 801, a communications interface 802, a memory 803, and a communication bus 804. Among them, the processor 801, the communications interface 802, and the memory 803 complete communication with each other through the communication bus 804. The processor 801 may call the logical instructions in the memory 803 to execute the UAV image detection method, which includes: obtaining the UAV image to be detected; inputting the UAV image to be detected into a preset target detection model and outputting a detection result; the detection result includes the target type and the corresponding confidence level; the preset target detection model includes a backbone network, a neck network, and a detection head; the backbone network is used to extract multi-level features of the UAV image to be detected; the neck network is used to perform weighted fusion on the multi-level features output by the backbone network along the top-down and bottom-up fusion paths respectively to generate multi-scale fusion features; the detection head is used to perform target detection on each scale of fusion features; among them, a fine feature extraction module is used in the backbone network to extract multi-level features of the UAV image to be detected; the processing process of the fine feature extraction module for the input feature map includes: using multiple parallel branches to process the input feature map respectively; among them, different dilation rates of convolution are used in different branches; splicing the outputs of the multiple parallel branches to generate a first feature map; using group convolution to process the first feature map to obtain receptive field spatial features, and using average pooling and group convolution to process the first feature map successively to generate receptive field weights; convolving the product of the receptive field weights and the receptive field spatial features to obtain the output of the fine feature extraction module.
[0083] In addition, when the logical instructions in the above-mentioned memory 803 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0084] An embodiment of the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the drone image detection method provided by each of the above method embodiments.
[0085] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the drone image detection method provided by each of the above method embodiments is implemented.
[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A fast small-target drone image detection method, characterized in that: include: Obtain the image of the drone to be tested; Input the image of the drone to be tested into a preset target detection model and output the detection result; The detection result includes the target type and the corresponding confidence level; the preset target detection model includes a backbone network, a neck network and a detection head; the backbone network is used to extract multi-level features of the image of the drone to be tested; the neck network is used to perform weighted fusion of the multi-level features output by the backbone network along the top-down and bottom-up fusion paths respectively to generate multi-scale fusion features; the detection head is used to perform target detection on each scale fusion feature; Among them, a fine feature extraction module is used in the backbone network to extract multi-level features of the image of the drone to be tested; the processing process of the input feature map by the fine feature extraction module includes: using multiple parallel branches to process the input feature map respectively; wherein convolutions with different expansion rates are used in different branches; the outputs of multiple parallel branches are spliced to generate a first feature map; the first feature map is processed by group convolution to obtain receptive field spatial features, and the first feature map is processed successively by average pooling and group convolution to generate receptive field weights; the product of the receptive field weights and the receptive field spatial features is convolved to obtain the output of the fine feature extraction module.
2. A fast small target oriented drone image detection method according to claim 1, characterized in that: In the neck network, a multi-feature aggregation network is used to perform weighted fusion on the multi-level features output by the backbone network along top-down and bottom-up fusion paths to generate multi-scale fusion features; the multi-feature aggregation network includes a plurality of multi-feature fusion modules, and the multi-feature fusion modules are used to assign weights to all input feature maps and perform weighted fusion; The bottom-up fusion path includes a first multiple feature fusion module, a second multiple feature fusion module and two convolutional layers; wherein the output of the second high-level feature map output by the backbone network after passing through one of the convolutional layers, the second high-level feature map and the highest-level feature map are used as inputs of the first multiple feature fusion module; the output of the low-level feature map output by the backbone network after passing through another convolutional layer, the low-level feature map and the output of the first multiple feature fusion module are used as inputs of the second multiple feature fusion module; the low-level feature map and the second high-level feature map are feature maps of adjacent layers; The top-down fusion path includes a third multiple feature fusion module, a fourth multiple feature fusion module and a fifth multiple feature fusion module; wherein, the output of the low-level feature map output by the backbone network after passing through another convolutional layer and the output of the second multiple feature fusion module serve as the input of the third multiple feature fusion module; the output of the second high-level feature map output by the backbone network after passing through one of the convolutional layers, the output of the first multiple feature fusion module and the output of the third multiple feature fusion module serve as the input of the fourth multiple feature fusion module; the highest-level feature map output by the backbone network and the output of the fourth multiple feature fusion module serve as the input of the fifth multiple feature fusion module; the outputs of the third multiple feature fusion module, the fourth multiple feature fusion module and the fifth multiple feature fusion module are multi-scale fusion features.
3. A fast small target oriented drone image detection method according to claim 2, characterized in that: The multiple feature fusion module is used to assign weights to all input feature maps and perform weighted fusion, specifically including: converting each input feature map into a corresponding vector through a tiling operation; processing each vector using a fully connected layer to generate a scalar for each input feature map; normalizing each scalar using a softmax function to generate a weight for each input feature map; and performing weighted fusion of all input feature maps based on the weight of each input feature map to obtain the output of the multiple feature fusion module.
4. A fast small target oriented drone image detection method according to any one of claims 1 to 3, characterized in that: The backbone network includes a convolutional layer, three feature extraction layers and a spatial pyramid pooling module connected in sequence; the feature extraction layer includes a fine feature extraction module and a C2f module; the outputs of the first two feature extraction layers and the output of the spatial pyramid pooling module are the multi-level features output by the backbone network.
5. A fast small-target drone image detection method according to claim 4, characterized in that: The spatial pyramid pooling module adopts the SPPF module.
6. A fast small target oriented drone image detection method according to claim 4, characterized in that: The target detection model adopts the Yolov8 network as the basic framework, adds a small-scale detection head P2, and removes the large-scale detection head P5.
7. A fast small-target drone image detection method according to claim 6, characterized in that: The number of C2f modules in the last two feature extraction layers is reduced from 6 to 3.
8. A fast small-target drone image detection device, characterized in that: include: An acquisition unit, used for acquiring images of the drone to be tested; A detection unit, used to input the image of the drone to be tested into a preset target detection model and output a detection result; The detection result includes the target type and the corresponding confidence level; the preset target detection model includes a backbone network, a neck network and a detection head; the backbone network is used to extract multi-level features of the image of the drone to be tested; the neck network is used to perform weighted fusion of the multi-level features output by the backbone network along the top-down and bottom-up fusion paths respectively to generate multi-scale fusion features; the detection head is used to perform target detection on each scale fusion feature; Among them, a fine feature extraction module is used in the backbone network to extract multi-level features of the drone image to be tested; the processing process of the input feature map by the fine feature extraction module includes: using multiple parallel branches to process the input feature map respectively; wherein convolutions with different expansion rates are used in different branches; the outputs of multiple parallel branches are spliced to generate a first feature map; the first feature map is processed by group convolution to obtain receptive field spatial features, and the first feature map is processed by average pooling to generate receptive field weights; the product of the receptive field weights and the receptive field spatial features is convolved to obtain the output of the fine feature extraction module.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.