Multi-scale context enhancement small target detection method based on improved RT-DETR

By improving the CSP-GFCG and ConvGLU modules of the RT-DETR framework, the problem of insufficient feature extraction in UAV small target detection is solved, the detection accuracy and computational efficiency are improved, and it is suitable for UAV aerial photography scenarios.

CN120599503APending Publication Date: 2025-09-05SOUTHWEAT UNIV OF SCI & TECH
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510765419.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing drone target detection technology lacks the ability to capture small target features and has weak multi-scale adaptability, resulting in low detection accuracy and high computing resource consumption, especially poor performance in complex aerial photography scenarios.

Method used

An improved RT-DETR framework is adopted, combined with the CSP-GFCG feature extraction method and the ConvGLU module. Feature extraction is enhanced through global filtering and convolutional gated linear units to achieve multi-scale context perception, improve the small target detection accuracy and reduce the computational complexity.

Benefits of technology

It significantly improves the ability to detect edges and texture details of small targets, reduces the number of model parameters and computing resource consumption, and enhances the detection efficiency and accuracy of UAV platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005441419940000011
    Figure HDA0005441419940000011
  • Figure HDA0005441419940000012
    Figure HDA0005441419940000012
  • Figure HDA0005441419940000021
    Figure HDA0005441419940000021
Patent Text Reader

Abstract

The invention discloses a multi-scale context enhancement small target detection method based on an improved RT-DETR (Reverse Transcription DET Rate). The method comprises the following steps: preprocessing unmanned aerial vehicle image data, and then constructing an improved RT-DETR model; the core is that a CSP-GFCG feature extraction module is used for modulating a feature map through frequency domain transform (DFT / IDFT) and a learnable global filter by using GFNet to realize global context modeling; and then the processed features are input into a ConvGLU module, and local features are enhanced in combination with depth separable convolution and a gating linear unit. GFNet and ConvGLU cooperate with each other, and challenge is effectively reserved for scale change and details in small target detection. The method aims at optimizing a feature extraction mechanism, reducing redundancy and improving the detection performance of a small target under a complex background. Meanwhile, the calculation efficiency is improved, the resource consumption is reduced, and the problem of missing detection of small targets is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and in particular is a multi-scale context-enhanced small target detection method based on improved RT-DETR. Background Art

[0002] Accurately detecting small-scale objects remains a challenging frontier in computer vision research. Due to the limited area covered by target pixels, information degradation is common during cross-scale feature representation. This is particularly true when the distance of the target significantly alters its morphology, often resulting in blurry or distorted images of distant objects.

[0003] At the same time, the semantic similarity between targets and backgrounds in complex scenes can cause feature confusion, significantly reducing detection reliability. With the widespread application of target detection technology in security inspections, autonomous driving systems, medical image analysis, and drone situational awareness, existing detection solutions based on the RTDETR framework still face technical bottlenecks such as insufficient ability to capture small target features and weak multi-scale adaptability. There is an urgent need to improve detection accuracy and scenario robustness through algorithm optimization.

[0004] As the intelligence level of drone platforms continues to improve, the target detection systems they carry place higher demands on feature expression capabilities. Although drones' aerial perspectives offer the advantage of wide-area perception, existing feature extraction architectures have significant limitations: for example, cross-level feature mismatch: the dynamic height changes of drones cause the target scale to span more than 10 times, making it difficult for traditional convolutional networks to efficiently fuse shallow detail features (such as edge textures of small targets) with deep semantic features; insufficient capture of local details: in the case of extremely limited target coverage areas, the fixed receptive field design of traditional convolutional kernels makes it difficult to effectively capture the edge and texture features of small targets, resulting in the continuous degradation of key detail information during feature transmission; occlusion feature confusion: the feature coupling effect caused by occlusion in dense target areas makes it difficult for existing attention mechanisms to achieve target instance-level feature decoupling. The above problems lead to the degradation of feature representation capabilities of existing models in aerial photography scenarios, seriously restricting the improvement of detection accuracy.

[0005] While current mainstream single-stage detection frameworks (typically the YOLO series) offer outstanding inference efficiency, they suffer from significant accuracy degradation in aerial drone photography scenarios. With the introduction of the Transformer architecture, detection paradigms such as DETR-derived models (such as Deformable DETR and RT-DETR) have gradually broken through the performance boundaries of traditional convolutional networks. Their end-to-end nature successfully eliminates the non-maximum suppression (NMS) post-processing step, demonstrating advantages in detection integrity and real-time performance.

[0006] It is worth noting that although RT-DETR has application potential in the field of drone detection, its feature extraction mechanism has inherent defects: the resolution of the feature map generated by the backbone network is too low, resulting in serious loss of high-frequency details of small targets; at the same time, the fixed receptive field design is difficult to adapt to the drastic multi-scale changes in aerial images, and the model complexity caused by the high parameter density significantly restricts its lightweight deployment on embedded platforms. Summary of the Invention

[0007] The purpose of the present invention is to provide a multi-scale context-enhanced small target detection method based on improved RT-DETR, which aims to improve the computational efficiency and recognition accuracy of the model by optimizing the feature extraction mechanism, reducing redundant structures, and effectively improving the detection performance of small targets in complex backgrounds through global filtering and convolutional gated linear unit modules. Especially in drone aerial photography scenes, it significantly improves the ability to capture the edge and texture details of small targets, improves detection accuracy, reduces computing resource consumption, and has strong practical application promotion value.

[0008] The present invention is implemented as follows: a multi-scale context-enhanced small target detection method based on improved RT-DETR comprises the following steps:

[0009] Step 1: Obtain and preprocess UAV aerial image data to support multi-scale small target detection tasks and build a model based on the improved RT-DETR;

[0010] Step 2: In the CSP-GFCG feature extraction method, GFNet implements global context modeling through frequency domain transformation, including two-dimensional discrete Fourier transform (DFT) to convert the input feature map to the frequency domain; element-by-element product modulation of the spectrum through a learnable global filter; and two-dimensional inverse Fourier transform (IDFT) to convert the modulated spectrum back to the spatial domain;

[0011] Step 3: By passing the features processed by GFNet into the ConvGLU module, ConvGLU enhances local features through depthwise separable convolution and gated linear units. The two work together to address the object scale change and detail preservation problems in UAV small target detection;

[0012] Step 4: Build a UAV small target detection model and use it to train and test the UAV aerial image dataset, observe the detection results, and verify the model performance;

[0013] The beneficial effects of the present invention are as follows: (1) The present invention significantly improves the edge and texture detail modeling capability of small targets through a multi-scale context-aware mechanism and the fusion of global filtering and convolutional gated linear unit modules. It also improves detection accuracy while preserving spatial resolution, and outperforms traditional feature extraction networks in complex aerial photography scenes. (2) The feature extraction method involved in the present invention takes into account both global and local feature modeling, effectively reduces the number of model parameters and computational complexity, improves detection efficiency, and enhances applicability in resource-constrained platforms (such as drone embedded devices). (3) The feature extraction method involved in the present invention has strong versatility and modular design. In addition to being able to be integrated with the RT-DETR framework, it can also be applied to mainstream target detection models such as YOLO and Faster R-CNN. It has strong algorithm compatibility and wide potential for practical application and promotion. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a process of a feature extraction method based on multi-scale context enhancement improved by RT-DETR provided by an embodiment of the present invention;

[0015] Figure 2 This is an example of a network backbone structure diagram provided by an embodiment of the present invention;

[0016] Figure 3 This is an example of a CSP-GFCG feature extraction module structure diagram provided by an embodiment of the present invention;

[0017] Figure 4 is a training result diagram of the training data set provided by an embodiment of the present invention;

[0018] Figure 5 This is a visualization result diagram of a data set provided by an embodiment of the present invention; DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0020] In this example, the above method was applied to the VisDrone2019 dataset. This dataset covers 10 categories, including pedestrians, vehicles, bicycles, and tricycles. Its images feature large variations in object scale and complex backgrounds. This dataset effectively demonstrates the effectiveness of the proposed method in improving detection accuracy and reducing computational resource consumption in small object detection tasks.

[0021] Figure 1The flow chart of the present invention shows a feature extraction method based on multi-scale context enhancement improved by RT-DETR. As shown in the figure, the present invention is implemented as follows, comprising the following steps:

[0022] S1: Acquire and preprocess UAV aerial image data to support multi-scale small target detection tasks and build a CSP-GFCG-based model;

[0023] S2: In the CSP-GFCG feature extraction method, GFNet implements global context modeling through frequency domain transformation, including two-dimensional discrete Fourier transform to convert the input feature map to the frequency domain; element-wise product modulation of the spectrum through a learnable global filter, and two-dimensional inverse Fourier transform (IDFT) to convert the modulated spectrum back to the spatial domain;

[0024] S3: By passing the features processed by GFNet into the ConvGLU module, ConvGLU enhances local features through depthwise separable convolution and gated linear units. The two work together to address the object scale change and detail preservation issues in UAV small target detection;

[0025] S4: Input the drone aerial image to be detected into the trained optimal detection model, which outputs the target's category and location information. Use the test set to verify the model's performance and ensure the accuracy and reliability of the detection results.

[0026] Furthermore, the GFNet in step S2 is specifically as follows: GFNet is a computationally efficient and conceptually elegant architecture designed to model long-range spatial dependencies in the frequency domain, providing a compelling alternative to traditional self-attention mechanisms in visual tasks.

[0027] By leveraging the inherent properties of the Fourier transform, GFNet achieves global context modeling with log-linear complexity, making it particularly suitable for applications requiring high-resolution feature processing, such as small object detection in drone imagery.

[0028] The core of GFNet is to replace the quadratic complexity of the self-attention layer with a lightweight but powerful frequency domain operation. The frequency domain operation consists of three main steps:

[0029] Step 1: First, two-dimensional discrete Fourier transform Input feature map Convert to the frequency domain to get a complex tensor This step converts the features from the spatial domain to the frequency domain to facilitate subsequent global information interaction. Represents the 2D DFT operator, H and W represent the width and height of the feature map respectively, and D represents the number of channels;

[0030] Step 2: Then, the frequency domain representation X is modulated in the frequency domain, and a learnable global filter K is used to achieve the Hadamard product (⊙), i.e., the element-by-element product. The global filter K is responsible for adjusting the spectrum. Its parameters control the response of different frequency components and are the key to model learning.

[0031] Step 3: Finally, the modulated spectrum is transformed by two-dimensional inverse Fourier transform (2D IDFT) Convert back to the spatial domain to complete the feature update This approach not only reduces computational overhead but also ensures that the model can effectively capture both long-term and short-term spatial interactions. By adaptively modulating frequency domain features, GFNet can flexibly capture interactions between different spatial scales, effectively addressing the significant scale variations of objects in drone images.

[0032] In addition, the low computational complexity of GFNet ensures that the module can remain efficient when processing high-resolution feature maps, which is crucial for preserving detailed information of small objects.

[0033] Furthermore, the ConvGLU module in step S3 is specifically:

[0034] ConvGLU is a channel mixer that combines 3x3 depthwise separable convolution and gated linear unit (GLU).

[0035] This design enables each token to obtain channel attention based on its neighboring image features, thereby enhancing the local modeling ability and the robustness of the model.

[0036] The uniqueness of ConvGLU is that each token has a unique gating signal based on the fine-grained features of its nearest neighbors, which solves the problem of too coarse information caused by global average pooling in the SE mechanism.

[0037] Furthermore, the CSP-GFCG module is specifically:

[0038] The CSP-GFCG module organically combines the efficient global context modeling capability of GFNet, the efficient information transfer of the CSP architecture, and the local feature enhancement capability of ConvGLU.

[0039] This combination enables the model to effectively handle object scale changes, retain detailed information of small targets, and achieve more efficient calculations in drone small target detection scenarios.

[0040] Furthermore, in step S4, specifically:

[0041] The feature extraction method described in steps S2 and S3 is integrated with the AIFI module in RT-DETR to enhance spatial information. The rest of the method uses the native module of RT-DETR to maintain the integrity of the framework.

[0042] Furthermore, step S4 constructs a drone aerial image dataset, and uses a new target detection model to train and test the dataset, observe the target detection results, and monitor the performance indicators of the model. After training is completed, the specific performance of indicators such as P, R, mAP50, and mAP50-90 can be further analyzed, and the recognition effect of the target image can be evaluated.

[0043] Figure 4 The PR graph generated during model training is shown to demonstrate the model's recognition performance for different target categories. Figure 5 The model's detection results for targets in drone aerial images are presented.

[0044] Existing technologies for small object detection face numerous limitations. Combining traditional feature extraction methods with deep learning methods can lead to structural redundancy, resulting in insufficient feature representation, reduced detection and recognition rates, and increased computational resource consumption. Furthermore, small object detection itself faces numerous challenges. For one thing, capturing the global context of small objects is difficult, especially when the contrast between the target and the background is low, leading to frequent missed detections.

[0045] On the other hand, the lack of flexibility in multi-scale feature modeling further increases the difficulty of detection. The feature extraction method designed in this paper can effectively utilize the edge information of small targets while ensuring the preservation of spatial information. Applying CSP-GFCG to a UAV small target detection dataset and combining it with the optimized RT-DETR to construct a new model can significantly improve computational efficiency and resource consumption while reducing the number of model parameters, thereby improving the accuracy and reliability of UAV small target detection in complex backgrounds.

[0046] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without the need for creative work are still within the scope of protection of the present invention.

Claims

1. A method for multi-scale context-enhanced small target detection based on improved RT-DETR, characterized in that: The method for multi-scale context-enhanced small target detection based on improved RT-DETR comprises the following steps: Step 1: Obtain a preprocessed drone aerial image dataset, which contains target images of various categories, with a unified image resolution of 640×640 pixels. The dataset is divided into a training set (70%), a validation set (15%), and a test set (15%) to ensure the diversity and representativeness of the dataset and support the detection task of small targets at multiple scales. Step 2: Build a drone-viewpoint target detection model based on the improved RT-DETR. The backbone network of the model adopts the improved CSP-GFCG module, which implements global context modeling through a unique frequency domain transformation mechanism. It also combines the adaptive global filter network (GF-Net) and the convolutional gated linear unit (ConvGLU) module to extract multi-scale features of the input image, thereby enhancing feature expression capabilities and effectively reducing computing resource consumption. Step 3: Use the training set to train the drone-viewpoint target detection model based on the improved RT-DETR, adjust the model hyperparameters through the validation set, and use the optimization algorithm AdamW for training to obtain the optimal target detection performance; Step 4: Input the drone aerial images to be detected into the optimal detection model, output the target category and location information, and verify the performance of the model through the test set to ensure the accuracy and reliability of the detection results.

2. The method according to claim 1, wherein In step 2, a GFNet network based on a frequency response adjustment mechanism is introduced, Fourier frequency domain information is used for cross-region context perception, and asymmetric feature enhancement is achieved through an adaptive filter kernel, thereby improving the spatial distribution recognition capability of the target.

3. The method according to claim 1, wherein In step 2, the ConvGLU module is composed of a depth-wise separable convolutional layer and a gated linear unit, which is used to adaptively learn local features. Specifically, it includes the following sub-steps: Step 1: Perform linear transformation on the input features to generate two branches representing the target features and the control signal; Step 2: In the control signal branch, a 3×3 depthwise separable convolution operation is used to adaptively model the features of the local area and generate a dynamic adjustment gating signal through an activation function; Step 3: Combine the output of the control signal branch with the target feature branch by point-by-point multiplication to complete the selective weighting of the feature channel; Step 4: The weighted features are fused with the original features through jump connections to enhance the interaction between local and global features. The ConvGLU module combines adaptive local feature extraction with global feature enhancement mechanism to effectively suppress irrelevant background interference and highlight target features, thereby enhancing the robustness of the model in complex backgrounds.

4. The method according to claim 1, wherein In step 2, the CSP-GFCG module processes the input feature map through the residual structure, specifically including: Step 1: The input feature map is divided into two parts and sent to the depthwise separable convolution and ConvGLU processing modules for feature extraction respectively; Step 2: Integrate the output features of the two processing paths across scales and generate channel-level weighting coefficients using a global pooling operation. Step 3: Through the adaptive routing mechanism, the optimal combination of multi-scale features is selected to enhance the edge preservation effect of small target areas.

5. The method according to claim 1, wherein The feature extraction strategy introduces the AIFI unit in the RT-DETR structure, and enhances the preservation of spatial structure and edge details through the feature transfer mechanism within the same scale range, thereby optimizing the recognition performance of the model on small-sized targets.

6. The method according to claim 1, wherein The step 6 includes: Step 1: Build a drone image dataset containing different categories of targets; Step 2: Set hyperparameters for the fused RT-DETR model and perform multiple rounds of training and validation to optimize model performance. The hyperparameters include an initial learning rate of 0.0002, a batch size of 4, an image size of 640, 300 training rounds, a data augmentation method such as mosaic enhancement of 1.0, and an early stopping strategy of 30. Step 3: Observe the detection effect of the model on the test set and further adjust the model to ensure high-precision detection.

Citation Information

Cited By

  • Intelligent identification and classification method and system for traditional Chinese medicine decoction pieces based on deep learning

    CN120997827A

  • Intelligent identification and classification method and system for traditional Chinese medicine decoction pieces based on deep learning

    CN120997827B

  • Automatic driving small target detection enhancement method and device based on RT-DETR framework

    CN121170758A

  • Unmanned aerial vehicle small target aerial photography detection method and detection system based on improved RT-DETR, and computer equipment

    CN121191025A

  • An unmanned aerial vehicle aerial small target detection method, detection system and computer equipment based on improved RT-DETR

    CN121191025B