Unmanned aerial vehicle small target detection method based on multi-scale feature enhancement network

By adopting multi-scale feature enhancement network in drone image object detection, combining lightweight backbone network and attention mechanism, the problem of small object detection is solved, and high-precision and fast detection effects are achieved.

CN120164128APending Publication Date: 2025-06-17ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510172876.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

It is difficult to detect small objects in the drone image, and it is difficult to effectively detect low resolution and small objects in the prior art. The Transformer-based detection model has problems such as insufficient feature refinement and high computing requirements.

Method used

Using a multi-scale feature enhancement network method, we deeply explore features through lightweight backbone networks and attention mechanisms, combine multi-scale feature fusion structure to splice shallow and deep feature maps, and improve model accuracy by modifying the loss function.

Benefits of technology

It significantly enhances the discrimination of local features, removes redundant information, improves the accuracy and speed of small object detection, and is suitable for real-time detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164128A_ABST
    Figure CN120164128A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle small target detection method based on a multi-scale feature enhancement network, and belongs to the technical field of unmanned aerial vehicle image target detection. Comprising the steps of obtaining a to-be-detected unmanned aerial vehicle image, inputting the unmanned aerial vehicle image into a pre-constructed multi-scale feature enhancement network, outputting detection, and completing unmanned aerial vehicle small target detection; the multi-scale feature enhancement network firstly uses a lightweight backbone network and an attention mechanism to deeply explore features. Secondly, a multi-scale feature fusion structure is adopted, and shallow-layer and deep-layer feature maps are spliced; and finally, the precision of the model is further improved by modifying the loss function. According to the method, the distinguishability of local features is remarkably enhanced, and redundant information is effectively removed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of UAV image target detection, and more specifically, to a method for detecting small targets of UAVs based on a multi-scale feature enhancement network. Background Art

[0002] The task of UAV (Unmanned Aerial Vehicle) image target detection aims to identify and locate specific objects in images captured by UAVs through computer vision technology. With the wide application of UAVs in fields such as security monitoring, emergency search and rescue, and cable inspection, it has become increasingly important to improve the accuracy of UAV image target detection.

[0003] Different from traditional images, UAV images usually contain a large number of small and low-resolution objects. Due to the low resolution of these small objects and the lack of sufficient context information, their detection becomes extremely challenging. In addition, small objects often coexist with large objects in the same image, which makes the features of larger objects dominant in the feature learning process, resulting in small objects not being effectively detected. At the same time, in the task of UAV image target detection, the features of low-pixel small objects are often ignored due to the lack of details.

[0004] In recent years, the rapid development of deep learning technology has promoted significant progress in the field of target detection. Object detection methods based on deep learning are generally divided into two categories: two-stage methods and one-stage methods. Two-stage algorithms, such as Faster R-CNN and Cascade RCNN, have slow inference speeds and large computational overheads due to the existence of independent region proposal and classification steps, making them unsuitable for real-time detection. In contrast, one-stage detectors, such as SSD and YOLO, achieve faster speeds through end-to-end designs, but usually rely on shallower features for detection, thus affecting the detection accuracy.

[0005] Currently, in order to meet the requirements of real-time performance and high detection accuracy, a new real-time detection Transformer model (RT-DETR) has been proposed. This model efficiently processes multi-scale features through internal scale interaction and multi-scale feature fusion, and is superior to advanced YOLO detectors of the same scale in terms of speed and accuracy. However, end-to-end detection models based on Transformer still face some challenges in UAV image target detection, such as insufficient feature refinement, global feature covering local features, and high computational requirements.

[0006] Therefore, how to provide a method for detecting small targets of UAVs based on a multi-scale feature enhancement network is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a method for detecting small targets of unmanned aerial vehicles (UAVs) based on a multi-scale feature enhancement network. First, a lightweight backbone network and an attention mechanism are used to deeply explore features. Second, a multi-scale feature fusion structure is adopted to splice shallow and deep feature maps. Finally, the accuracy of the model is further improved by modifying the loss function. The present invention not only significantly enhances the distinguishability of local features but also effectively removes redundant information.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] A method for detecting small targets of unmanned aerial vehicles (UAVs) based on a multi-scale feature enhancement network, comprising:

[0010] Obtaining a UAV image to be detected, inputting the UAV image into a pre-constructed multi-scale feature enhancement network, and outputting a detection to complete the detection of small targets of UAVs;

[0011] Wherein, the multi-scale feature enhancement network includes: a backbone network, a multi-scale feature fusion encoder, and a decoder;

[0012] The backbone network extracts key information from the input UAV image and generates multi-scale feature maps from the last four stages;

[0013] The multi-scale feature maps are fused through the multi-scale feature fusion encoder;

[0014] The Inner-SIoU-aware query selection mechanism selects a fixed number of image features therefrom as the starting query of the decoder. Through the auxiliary head, the decoder refines the query to generate a bounding box with confidence; and completes the detection of small targets of UAVs.

[0015] Further, the backbone network adopts ResNet-18 as the basic backbone network, and optimizes the downsampling module in the basic backbone network by introducing an improved FasterBlock. An SEA module is added to the improved FasterBlock to replace the second 3x3 convolutional layer of the FasterBlock in the original basic network, and a new FS-BasicBlock is constructed as the feature extraction module.

[0016] Further, the SEA module includes three steps: squeezing operation, excitation operation, and scaling operation; wherein, the squeezing operation compresses the spatial dimension of the feature map; the excitation operation learns the importance of each channel; and the scaling operation adjusts the original feature map according to the learned weights.

[0017] Furthermore, the improved FasterBlock introduces partial convolution (PConv), replaces the standard convolutional layer with PConv, and performs conventional convolution operations only on a part of the input channels for spatial feature extraction, leaving other channels unchanged.

[0018] Furthermore, the expression of the partial convolution (PConv) is:

[0019] F PConv = h × w × k 2 × C P 2 ;

[0020] In the formula, k represents the convolutional kernel size, and h, w, and C p represent the length, width of the input feature map, and the number of channels of the corresponding PConv, respectively.

[0021] Furthermore, the P2 feature layer in the backbone network is introduced into the multi-scale feature fusion encoder using spatial-to-depth convolution.

[0022] Furthermore, the multi-scale feature fusion encoder further includes an enhanced feature expression module, which is composed of a large-scale branch, a local branch, and a global branch. The output feature maps of the large-scale branch, local branch, and global branch are integrated through additive fusion and output feature maps after passing through a convolutional layer.

[0023] Furthermore, the multi-scale feature enhancement network is trained based on a preset loss function, and the expression of the loss function is:

[0024] inter = (min(b r gt , b r ) - max(b l gt , b l )) * (min(b b gt , b b ) - max(b t gt , b t ))

[0025] union = (w gt * h gt ) * (ratio) 2 + (w * h) * (ratio) 2 - inter

[0026]

[0027] In the formula, b gta and b respectively represent the internal true bounding box and the internally predicted anchor box, w gt and h gt respectively represent the width and height of the internal true bounding box, while w and h represent the width and height of the internally predicted anchor box, and ratio represents the size ratio of the true bounding box and the internally predicted anchor box.

[0028] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method for detecting small targets of unmanned aerial vehicles based on a multi-scale feature enhancement network. First, a lightweight backbone network and an attention mechanism are used to deeply explore features. Second, a multi-scale feature fusion structure is adopted to splice shallow and deep feature maps. Finally, the accuracy of the model is further improved by modifying the loss function. The present invention not only significantly enhances the distinguishability of local features but also effectively removes redundant information. The specific beneficial effects are as follows:

[0029] (1) The proposed faster squeeze-and-excitation (FSE) module is integrated into the backbone network and assigns weights to features of different channels, making the model more lightweight, so that the model can focus more on key features.

[0030] (2) The designed feature enhancement (FE) module can more effectively capture and utilize the global and local features of the image by adjusting attention and selecting frequencies in different domains.

[0031] (3) The Inner-SIoU loss function is introduced, which dynamically adjusts the size of the auxiliary bounding box to improve the detection performance of small objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0033] Figure 1 It is a schematic structural diagram of the multi-scale feature enhancement network (eRT-DETR model) provided by the embodiment of the present invention;

[0034] Figure 2 It is a schematic diagram of a partial convolution structure provided by the embodiment of the present invention;

[0035] Figure 3 It is a schematic structural diagram of the FS-BasicBlock provided by the embodiment of the present invention;

[0036] Figure 4Schematic diagram of the space-to-depth convolution structure provided by the embodiments of the present invention;

[0037] Figure 5 Schematic diagram of the enhanced feature expression module structure provided by the embodiments of the present invention;

[0038] Figure 6(a) is a schematic diagram of the confusion matrix of the RT-DETR network provided by the embodiments of the present invention on VisDrone2019-Val;

[0039] Figure 6(b) is a schematic diagram of the confusion matrix of eRT-DETR provided by the embodiments of the present invention on VisDrone2019-Val;

[0040] Figure 7 Schematic diagram of the visualization results on the VisDrone2019 dataset provided by the embodiments of the present invention. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0042] The embodiments of the present invention disclose a method for detecting small UAV targets based on a multi-scale feature enhancement network, including:

[0043] Obtain the UAV image to be detected, input the UAV image into a pre-constructed multi-scale feature enhancement network, output detections, and complete the detection of small UAV targets;

[0044] Among them, the multi-scale feature enhancement network includes: a backbone network, a multi-scale feature fusion encoder, and a decoder;

[0045] The backbone network extracts key information from the input UAV image and generates multi-scale feature maps from the last four stages;

[0046] The multi-scale feature maps are fused by the multi-scale feature fusion encoder;

[0047] The Inner-SIoU-aware query selection mechanism selects a fixed number of image features therefrom as the starting queries for the decoder. Through the auxiliary head, the decoder refines the queries to generate bounding boxes with confidence; complete the detection of small UAV targets.

[0048] In a specific embodiment, the multi-scale feature enhancement network proposed by the present invention is an improved RT-DETR model, named eRT-DETR. To enhance the recognition ability of small objects in UAV images, in this embodiment, a faster squeeze-and-excitation (FSE) module is introduced into the backbone network of eRT-DETR, aiming to improve the feature extraction performance and reduce the hardware cost. This embodiment optimizes the original multi-scale feature fusion encoder and adds a feature enhancement (FE) module to better capture and utilize global and local image features and promote the fusion of meaningful feature representations. In addition, the Inner-SIoU loss function dynamically adjusts the size of the auxiliary bounding box to improve the detection performance of small objects.

[0049] Specifically, as Figure 1 shown, the multi-scale feature enhancement network (eRT-DETR) consists of three main parts: a backbone network, a multi-scale feature fusion encoder, and a decoder. First, the backbone network extracts key information from the input UAV image and generates multi-scale feature maps from the last four stages (P2, P3, P4, P5). Then, these feature maps are fused through the multi-scale feature fusion encoder. Next, the Inner-SIoU-aware query selection mechanism selects a fixed number of image features from them as the starting queries for the decoder. Through the auxiliary head, the decoder gradually refines these queries to generate bounding boxes with confidence.

[0050] In a specific embodiment, to reduce the computational redundancy brought by complex models in simple tasks and improve the detection speed, in this embodiment, a relatively lightweight ResNet-18 is used as the basic backbone network. To further reduce the computational cost and improve the running speed, an improved method is proposed by introducing a lightweight module, FasterBlock, to optimize the BasicBlock in the backbone network. FasterBlock is the core component of the FS-BasicBlock downsampling module. This module replaces the conventional convolution with partial convolution (PConv, as Figure 2 ) and only performs conventional convolution operations (Conv) on some input channels for spatial feature extraction, while the remaining channels remain unchanged. This method has a significant effect on reducing the computational complexity and memory access requirements. For PConv to achieve memory access optimization for continuous channels, the module regards the first or last continuous channel as the representative of the feature map for calculation. In typical settings, the number of channels of the input and output feature maps usually remains the same. The computational complexity of conventional convolution is shown in Equation (1), while the complexity of partial convolution is shown in Equation (2). When the number of channels C of partial convolution PWhen it is set to 1 / 4 of the number of input channels C, the computational complexity can be reduced to 1 / 16 of the original. Compared with traditional convolution, this optimization not only significantly reduces the number of model parameters by reducing redundant calculations, but also effectively reduces the memory access requirements, thereby improving the efficiency of the model in capturing spatial features.

[0051] F Conv = h × w × k 2 × C 2 (1)

[0052] F PConv = h × w × k 2 × C P 2 (2)

[0053] In the formula, k represents the size of the convolution kernel, and h, w, C, and C p represent the length, width, the corresponding ordinary convolution, and the number of channels of the corresponding PConv of the input feature map, respectively.

[0054] Specifically, although PConv reduces the model parameters, the sequential feature extraction method of the convolutional layer (Conv) still makes the inference accuracy unsatisfactory. In addition, in order to improve the performance and efficiency of the model in the UAV image detection task, in this embodiment, on the basis of FasterBlock, a lightweight module SEAttention (Squeeze-and-Excitation Attention, SEA) is introduced to enhance the network's ability to capture the features of small-sized objects, and FasterSqueeze-and-ExcitationBlock (FS-Block) is proposed. The structure is shown in Figure 3 as shown in (b). Then, FS-Block is used to replace the second 3×3 convolution in the original BasicBlock to obtain FS-BasicBlock, and the detailed structure is shown in Figure 3 as shown in (a).

[0055] The Squeeze-and-Excitation Attention module consists of three steps: squeeze, excitation, and scale. As Figure 3 shown in (c). First, the squeeze operation compresses the spatial dimension of the input feature map to 1×1 through global average pooling. Then, the excitation operation uses two fully connected layers (FC layers) to learn the weights of each channel. The first FC layer realizes dimensionality reduction, and the second FC layer generates a weight vector matching the number of channels to judge the importance of each channel. Finally, the scale operation weights the original feature map using these weights to highlight the key features.

[0056] This structure aims to reduce the number of model parameters while achieving more efficient feature extraction and more accurate object detection by combining the lightweight characteristics of FasterBlock and the feature enhancement effect of SEA.

[0057] In a specific embodiment, to solve the problem of blurred target information in the P3, P4, and P5 feature layers, this embodiment uses space-to-depth convolution (SPDConv) to introduce the P2 feature layer from the backbone into the multi-scale fusion decoder, enhancing the model's ability to learn small-size features. In addition, the multi-scale fusion decoder combines with the enhanced feature expression module, which can effectively improve the capture and utilization rate of global and local image features.

[0058] In a specific embodiment, traditional convolution performs well in image recognition, but when detecting small-size objects, due to the limited receptive field, details are lost, making it difficult to handle low-resolution and small-target tasks. To address this problem, this embodiment integrates space-to-depth (SPD) convolution (SPDConv), as Figure 4 shown. SPDConv effectively replaces traditional strided convolution and pooling operations by first performing space-to-depth conversion and then convolution, enhancing the model's ability to capture details of small objects in complex backgrounds.

[0059] Specifically, the SPD (Space-to-Depth) module aims to achieve spatial downsampling by rearranging the spatial dimensions of the input feature map while retaining the information in the channels. Specifically, the input feature map (with size C1×S×S) is divided into r×r small blocks (where r is the downsampling factor), and these small blocks are rearranged into the channel dimension to generate a new feature map with size S / r×S / r×(C×r2). For example, when r = 2, the size of the new feature map is S / 2×S / 2×4C.

[0060] In a specific embodiment, the Feature Enhancement (FE) module consists of three key branches: the large-scale branch, the local branch, and the global branch. The output feature maps of these branches are integrated through additive fusion and further refined through a convolutional layer to generate high-quality feature maps, as Figure 5 shown in (a) of

[0061] Specifically, in the large branch, by introducing depthwise separable convolution (DSConv) with a kernel size of 31×31, the receptive field of the model is significantly expanded. The local branch aims to capture local information in the image, and this branch integrates local signals through 1×1 depth convolution (DConv) to construct a concise but effective local processing mechanism.

[0062] Specifically, considering that the input degraded images in the inference stage are usually larger than those in training, the 31×31 convolutional kernel is difficult to cover the entire image. To address this issue, the global branch introduces dual-domain processing to achieve global modeling. As shown in (b)-(c) of Figure 5 , this branch includes a Frequency Channel Attention (FCA) module and a Frequency-based Spatial Attention Module (FSAM). The FCAM applies frequency channel attention (FCA) to the input global feature X Global , Equation (3);

[0063]

[0064] Through Fourier processing, the global feature is effectively refined according to the spectral convolution theorem. and are the Fast Fourier Transform and its inverse operation respectively; X FCA is the output of FCA; is a 1×1 convolutional layer, GAP represents global average pooling, represents element-wise multiplication. Then, apply the frequency-based attention module to X Global to refine the spectrum at the fine-grained level to obtain the result X FSA of the FSAM module, Equation (4);

[0065]

[0066] After the processing of FCAM and FSAM and adding with other branches, it optimizes the extraction and utilization of image features, strengthens the feature representation of small targets and improves the detection accuracy, and significantly enhances the model's ability to recognize small targets in complex backgrounds.

[0067] In a specific embodiment, when the original GIoU model is used to evaluate the overlap between the predicted bounding box and the actual bounding box, it faces several challenges. These challenges include insensitivity to small object detection, neglect of the differences in the shape and position of the bounding boxes, and the inability to provide effective information in the absence of overlap. To address these issues, the present invention proposes an Inner-IoU loss, which optimizes the loss calculation by controlling the size of the auxiliary bounding box using a scaling factor. This method accelerates the convergence of the model, improves the accuracy of small object detection, and enhances the robustness of the model to low-quality samples. The Inner-IoU loss is defined as follows:

[0068] inter = (min(b r gt , b r ) - max(b l gt , b l )) * (min(b bgt , b b ) - max(b t gt , b t )) (5)

[0069] union = (w gt * h gt ) * (ratio) 2 + (w * h) * (ratio) 2 - inter (6)

[0070]

[0071] where b gt and b respectively represent the internal true bounding box and the internally predicted anchor box. w gt and h gt respectively represent the width and height of the internal true bounding box, while w and h represent the width and height of the internally predicted anchor box, and ratio represents the size ratio between the true bounding box and the internally predicted anchor box.

[0072] This embodiment further introduces the combination of SIoU

[38] and Inner - IoU, proposes the Inner - SIoU loss function, simultaneously calculates the IoU loss using the auxiliary bounding box, and considers the angular difference between the anchor box and the true bounding box. This method not only improves the accuracy of object detection but also enhances the model's ability to adapt to changes in object size and position in complex environments. The Inner - SIoU loss is defined as follows:

[0073]

[0074] Loss inner-SIoU = L SIoU + IoU - IoU inner (9)

[0075] where Δ represents the shape loss, which is used to measure the aspect ratio difference between the predicted bounding box and the true bounding box, and Ω represents the angular loss, which is used to quantify the angular difference between the predicted bounding box and the true bounding box.

[0076] In a specific embodiment, experiments and analyses are carried out based on the above - disclosed method for detecting small UAV targets based on a multi - scale feature enhancement network.

[0077] Specifically, the network experimental environment is based on Ubuntu 20.04, Python 3.8.19, and Pytorch 2.4.1. The experiment uses a computer with an Intel Xeon Platinum 8375C processor, 73 GB of memory, an RTX 3090 graphics card, and 24 GB of video memory. The training batch size is set to 4, the number of training epochs is set to 300, and the learning rate is set to 1×10-4. An adaptive image size of 640×640 is selected for the experiment. Other experimental settings are consistent with RT-DETR.

[0078] The VisDrone2019 dataset is a drone aerial image object detection dataset constructed by the AISKYEYE team at Tianjin University. This dataset is designed specifically for understanding visual data captured by drones and covers complex environments and challenges such as different scenarios and different lighting conditions. The dataset contains 7,019 images, of which the training set VisDrone2019-train contains 6,471 images and the validation set VisDrone2019-val contains 548 images. The dataset has 10 predefined object categories: pedestrian, person, bicycle, car, van, truck, tricycle, awning tricycle, bus, and motorcycle.

[0079] Specifically, the HIT-UAV dataset was created by a research team at Harbin Institute of Technology (HIT) and focuses on object detection in drone infrared thermal imaging images. The dataset contains 2,898 infrared thermal imaging images from various real-world scenarios, including schools, parking lots, roads, and playgrounds, covering both daytime and nighttime lighting conditions during image acquisition. The dataset contains five categories of objects, mainly consisting of dense and small objects, including people, vehicles, and bicycles.

[0080] In a specific embodiment, by maintaining consistent experimental conditions, the performance changes in image detection before and after model optimization are compared to measure the efficiency of the algorithm. During the evaluation process, the following criteria are used: Precision, Recall, mean Average Precision (mAP), GigaFLOPS (GFLOPs), and Frames Per Second (FPS).

[0081]

[0082] Specifically, precision refers to the proportion of actual positive cases among the cases predicted as positive by the network model, that is, the ratio of the number of correctly identified positive samples to the number of all samples predicted as positive. Recall reflects the proportion of all actual positive samples that are correctly predicted as positive by the model. In this evaluation system, TP represents the number of true positives, that is, the number of positive samples correctly determined by the model; FP represents the number of false positives, that is, the number of negative samples misjudged as positive; FN represents the number of false negatives, that is, the number of positive samples not detected by the model.

[0083]

[0084] The average precision (AP) of a single class reflects the average of the highest precision that can be achieved at different recall levels for that class. The overall mean average precision (mAP) is the average of the AP values of all classes, providing a comprehensive evaluation metric for the model's performance. In this calculation process, M represents the total number of classes, and the denominator of mAP is the cumulative result of the AP values of these classes. In addition, GFLOPs is used to measure the computational complexity of the model.

[0085] In a specific embodiment, it also includes conducting ablation experiments on the loss function.

[0086] Specifically, to evaluate the effectiveness of the Inner-SIoU loss function, it was compared with other mainstream loss functions, including GIoU, DIoU, CIoU, SIoU, and EIoU. The comparison results are shown in Table 1. The results show that when the ratio is set to 1.20, the Inner-SIoU loss function achieves the best results. Compared with the original model using the GIoU loss, the precision is increased by 1.7%, and mAP50 is increased by 0.8%. In addition, compared with the baseline model using the SIoU loss, the precision is increased by 1.5%, and mAP50 is increased by 1.1%. These results indicate that selecting the appropriate ratio and adopting the Inner-SIoU loss function can significantly enhance the consistency of bounding box regression, thus improving the detection accuracy.

[0087] Table 1. Performance comparison table of the improved loss function model on the VisDrone2019-Val dataset

[0088]

[0089]

[0090] Specifically, module ablation experiments: To evaluate the impact of the proposed improved module on the model performance, eight ablation experiments were conducted. The RT-DETR network was used in the experiments, and the improved module was gradually introduced. The specific changes included replacing the feature extraction basic module with the lightweight FS-BasicBlock module, using the FE module to optimize the feature fusion network, and updating the loss function to Inner-SIoU. The results in Table 2 show that Experiments 1, 2, and 3 verified that these three methods significantly improved the performance. Among them, the impact of the FE module was the most obvious, with the mAP50 increasing by 1.9% and the mAP50:95 increasing by 1.5%. With the introduction of more modules, the performance continued to improve. Compared with the baseline model, the proposed model increased the mAP50 by 1.9% and the mAP50:95 by 1.5%, while reducing the number of parameters by 12.6%. These results verified the effectiveness of the proposed method in improving the small object detection performance.

[0091] Table 2. Ablation experiment results of the algorithm on the VisDrone2019-Val dataset ("√" symbol indicates the improvement of the corresponding module structure)

[0092]

[0093]

[0094] In addition, the confusion matrices of RT-DETR and eRT-DETR were generated, as shown in Figures 6(a)-(b). By comparing these matrices, it can be found that eRT-DETR has significantly improved the detection performance in each category, especially in the correct classification rate of bicycles. These findings further prove the improvement of the improved model in category recognition ability.

[0095] In a specific embodiment, the comparative experiment on the Visdrone2019 dataset is as follows:

[0096] To evaluate the superior performance of the eRT-DETR algorithm in drone image object detection, this embodiment comprehensively compares the method of the present invention with several advanced existing technologies, including Faster RCNN, Cascade RCNN, Tood, RetinaNet, and the YOLO series. Using the VisDrone2019 dataset, this embodiment calculates the mAP50 values for 10 target categories, as shown in Table 3. The experimental results show that the method of the present invention achieves the highest average detection accuracy among all detection algorithms, surpassing all other methods. Especially in the car category, eRT-DETR reaches an accuracy rate of 85.2%, showing the best result among all tested algorithms. In addition, eRT-DETR performs excellently in detecting challenging small objects (such as bicycles). Although its accuracy in the tricycle category does not reach the highest, eRT-DETR demonstrates excellent overall performance in the drone image object detection task.

[0097] Table 3 Performance comparison table of different models on the VisDrone2019 dataset

[0098]

[0099] In a specific embodiment, the comparative experiment on the Visdrone2019 dataset is as follows:

[0100] The present invention comprehensively evaluates eRT-DETR and other mainstream detection algorithms using the HIT-UAV dataset, and the results are shown in Table 4. The experimental results show that eRT-DETR demonstrates significant advantages in drone infrared thermal imaging detection. Although its performance is slightly lower than that of YOLOv5 in terms of the mAP50 metric, eRT-DETR is superior to the comparison models in terms of both precision and recall, showing higher true positive recognition accuracy and fewer missed detection targets. These improvements further prove the excellent performance of eRT-DETR in object detection.

[0101] Table 4 Performance comparison table of different models on the HIT-UAV dataset

[0102]

[0103] Specifically, to better demonstrate the effectiveness of the method of the present invention in actual scenarios, this embodiment conducts a comparative test using samples with different lighting conditions and environmental factors in the VisDrone2019 dataset, and the results are as Figure 7As shown, where (a) is the true label, and (b) and (c) are the results predicted by RT-DETR and eRT-DETR respectively. Under low-light conditions, both RT-DETR and eRT-DETR exhibit high recognition accuracy. However, under normal light and weak light, the accuracy of RT-DETR in small target detection decreases, and the incidence of false positives and missed detections is relatively high. In contrast, eRT-DETR performs well in these situations. In scenarios involving occlusion, both RT-DETR and eRT-DETR experience detection failures or false detections, but eRT-DETR gives a higher confidence score when detecting the same object. In addition, eRT-DETR also performs well in detecting unlabeled targets. Overall, eRT-DETR demonstrates strong robustness in detecting small targets in various complex aerial environments.

[0104] Specifically, to address the challenges in drone image target detection, especially the problem of small target recognition, the present invention proposes a multi-scale feature enhancement network (eRT-DETR model). First, a novel fast squeeze-and-excitation (FSE) module is designed. By combining the fast module with the squeeze-and-excitation attention mechanism, the backbone feature extraction network is optimized. This innovation effectively captures small target features while reducing the model complexity and computational cost. Second, to address the problem that small targets are easily interfered by complex background information, the present invention establishes a multi-layer feature fusion structure, splicing shallow and deep feature maps, and the designed feature enhancement module further improves the effective capture and utilization of global and local image features. In addition, the Inner-SIoU loss function is introduced to replace the traditional GIoU loss, and the auxiliary bounding box is optimized by adjusting the scaling factor, thereby improving the detection accuracy. Finally, the present invention is evaluated on multiple datasets, and compared with existing models, it achieves better detection accuracy. In the drone image target detection task, the method of the present invention outperforms the RT-DETR baseline network. Future research will focus on optimizing the computational efficiency of the model and further improving the detection performance to better meet the resource constraints and real-time requirements in practical applications.

[0105] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for relevant parts.

[0106] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting small targets of unmanned aerial vehicles based on a multi-scale feature enhancement network, characterized in that: include: Obtain the drone image to be detected, input the drone image into a pre-built multi-scale feature enhancement network, output the detection, and complete the drone small target detection; Wherein, the multi-scale feature enhancement network includes: a backbone network, a multi-scale feature fusion encoder and a decoder; The backbone network extracts key information from the input drone image and generates multi-scale feature maps from the last four stages; The multi-scale feature map is fused by the multi-scale feature fusion encoder; The Inner-SIoU-aware query selection mechanism selects a fixed number of image features as the starting query for the decoder. Through the auxiliary head, the decoder refines the query and generates a bounding box with confidence, completing the detection of small drone targets.

2. According to claim 1, a method for detecting small targets of unmanned aerial vehicles based on a multi-scale feature enhancement network is characterized in that: The backbone network includes ResNet-18 as the basic backbone network, and optimizes the downsampling module in the basic backbone network by introducing an improved FasterBlock, adding a SEA module to the improved FasterBlock to replace the second 3x3 convolutional layer of FasterBlock in the original basic network, and constructing a new FS-BasicBlock as a feature extraction module.

3. The method for detecting small drone targets based on a multi-scale feature enhancement network according to claim 2 is characterized in that: The SEA module includes three steps: a squeeze operation, an excitation operation, and a scaling operation; wherein the squeeze operation compresses the spatial dimension of the feature map; the excitation operation learns the importance of each channel; and the scaling operation adjusts the original feature map according to the learned weights.

4. The method for detecting small drone targets based on a multi-scale feature enhancement network according to claim 3 is characterized in that: The improved FasterBlock introduces partial convolution PConv, replaces the standard convolution layer with the partial convolution PConv, performs conventional convolution operations on only a part of the input channels to extract spatial features, and keeps other channels unchanged.

5. The method for detecting small targets of unmanned aerial vehicles based on a multi-scale feature enhancement network according to claim 4 is characterized in that: The expression of the partial convolution PConv is: F PConv =h×w×k 2 ×C P 2 ; In the formula, k represents the convolution kernel size, h, w and C p They respectively represent the length, width and the number of channels of the corresponding PConv of the input feature map.

6. The method for detecting small drone targets based on a multi-scale feature enhancement network according to claim 1, characterized in that: The P2 feature layer in the backbone network is introduced into the multi-scale feature fusion encoder using space-to-depth convolution.

7. The method for detecting small drone targets based on a multi-scale feature enhancement network according to claim 1, characterized in that: The multi-scale feature fusion encoder also includes an enhanced feature expression module, which consists of a large-scale branch, a local branch and a global branch. The output feature maps of the large-scale branch, the local branch and the global branch are integrated by additive fusion, and the feature maps are output after passing through a convolutional layer.

8. The method for detecting small drone targets based on a multi-scale feature enhancement network according to claim 1, characterized in that: It also includes training the multi-scale feature enhancement network based on a preset loss function, where the expression of the loss function is: inter=(min(b r gt ,b r )-max(b l gt ,b l ))*(min(b b gt ,b b )-max(b t gt ,b t )) union=(w gt *h gt )*(ratio) 2 +(w*h)*(ratio) 2 -inter Where b gt and b represent the internal ground-truth bounding box and the internal predicted anchor box, respectively, and w gt and h gt They represent the width and height of the internal true bounding box, respectively, while w and h represent the width and height of the internal predicted anchor box, and ratio represents the size ratio of the true bounding box and the internal predicted anchor box.

Citation Information

Cited By

  • Low-illumination unmanned aerial vehicle target detection method based on double-domain contrast learning

    CN121095823A

  • Sparse gating double-domain target detection method for unmanned aerial vehicle image

    CN121527405A