Lightweight small target detection method based on improved RT-DETR
By improving the lightweight module, multi-branch structure and feature fusion network of RT-DETR, the problems of high computing complexity and large model size in small object detection are solved, and efficient lightweight and accurate detection are achieved, suitable for devices with resource limitations.
Patent Information
- Application Number
- CN202510489838.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
AI Technical Summary
The existing small object detection methods have challenges in terms of high computational complexity and large model size, making it difficult to deploy lightweight on resource-constrained devices, and ignore the details of shallow networks, resulting in limited detection performance.
The improved lightweight module (SRFM), multi-branched neural network block (DMSFM) and PAN-based feature fusion network (AEFN) are used to enhance feature representation and position information, and the model is optimized through residual connections, partial convolution operations and channel attention mechanisms, reducing parameter amount and computational complexity.
It realizes efficient and lightweight small object detection, improves detection accuracy and generalization capabilities of model, reduces computing costs and resource consumption, and is suitable for real-time application on edge devices.
Smart Images

Figure CN120339714A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of object detection in computer vision, and specifically relates to a lightweight small object detection method based on improved RT-DETR. Background Art
[0002] The rapid development of computer vision technology has provided broad space for the research and application of object detection technology. As one of the core tasks in the field of computer vision, object detection plays a crucial role in practical applications. Through in-depth analysis of images or videos, object detection can accurately identify and locate specific objects in the image, providing important support for key fields such as autonomous driving, video surveillance, and intelligent security. Small object detection refers to the process of specifically detecting and locating small-sized and low-pixel-density objects in the object detection task. In actual scenarios, small objects are often interfered by various factors such as noise and deformation. At the same time, complex background information and occluders may also pose challenges to the detection of small objects. Therefore, the detection of small objects has become a challenging task.
[0003] In the field of small object detection, researchers have proposed various methods to solve these problems. Among them, Deformable DETR, as an improved version of DETR, introduces a deformable attention mechanism to better handle the non-rigid deformation of objects and improve the detection effect of small objects. Sparse R-CNN combines a sparse attention mechanism and Transformer, and improves the detection performance of small objects through a carefully designed sparse attention mechanism while reducing the computational complexity. Dynamic-DETR is an improved version of dynamic feature prediction, which can adaptively capture object features at different scales, thereby improving the detection performance of small objects. However, there are still some challenges in practical applications. For example, some methods adopt complex network structures and attention mechanisms, resulting in a high computational complexity of the model, which is not conducive to deployment and real-time applications on resource-constrained devices. In addition, some methods have a large number of model parameters, making the model volume large and not conducive to lightweight deployment on edge devices.
[0004] This problem prompts thinking about how to extend DETR to scenarios involving real-time detection of small objects and achieve model lightweighting. To solve this problem, the model structure of RT-DETR was re-conceived to improve the performance of small object detection and achieve model lightweighting. Specifically, in RT-DETR, it was found that a hybrid encoding of Transformer and CNN was used, but only the last layer of features output from the Backbone was input into the hybrid encoding for real-time object detection. This results in the input to the Transformer being semantic information from a deep network, while ignoring the detailed information crucial for small object detection from the shallow network, such as edges, textures, and color gradients. In addition, in RT-DETR, simple concatenation operations are performed on feature maps of different levels, which may lead to insufficient information fusion because feature maps of different sizes may capture semantic information at different levels and scales, and simple summation operations are difficult to effectively handle the complex relationships between multi-scale information, thereby potentially affecting the detection performance of small objects. Additionally, in terms of the backbone network, RT-DETR-r18 adopted a resnet network structure, and as the network depth increases, the computational complexity will increase accordingly, resulting in a need for more computational resources and training time. Summary of the Invention
[0005] To solve these problems, the object of the present invention is to provide a lightweight small object detection method based on improved RT-DETR. First, a lightweight module based on a residual structure (SRFM) is proposed to improve the training effect and expression ability of the model, reduce the number of parameters, and achieve model lightweighting. Second, a neural network block with a multi-branch structure (DMSFM) is proposed to fuse feature representations at multiple levels and angles and enhance the expression ability of the network. Third, an attention module (CHM) and a feature fusion network based on the PAN structure (AEFN) are proposed to enhance the location and semantic information of small objects. Finally, the network is proposed.
[0006] The object of the present invention is achieved through the following technical solutions:
[0007] Step 1: RT-DETR is used as the object detection framework, and the BasicBlock module in resnet18 is replaced with an improved lightweight module (SRFM) in the extraction backbone network. This module utilizes the residual connection mechanism to effectively transmit gradients and optimize information. In addition, the input features are localized through partial convolution operations, thereby further reducing the number of parameters of the model;
[0008] Step 2: Replace the RepBlock in the Fusion module of the baseline model with the proposed neural network block (DMSFM) with a multi-branch structure to fuse feature representations from multiple levels and angles; use the fused main branch for inference in the inference stage to reduce the model complexity and computational cost.
[0009] Step 3: Adopt the proposed feature fusion network (AEFN) based on the PAN structure. Through the proposed attention module, important features are enhanced and unimportant features are suppressed to enhance the location and semantic information of small targets.
[0010] Step 4: Use the URPC2020, NWPU VHR-10, and RSOD public dataset object detection datasets to train the improved RT-DETR object detection network model, configure the training parameters, and set the number of training epochs.
[0011] Step 5: Input the small target detection dataset into the trained improved RT-DETR object detection network model. The model outputs the detection results of various types of small targets after feature extraction and multi-scale feature fusion, realizing small target detection based on the improved RT-DETR.
[0012] 2. The lightweight small target detection method based on the improved RT-DETR according to claim 1, wherein in step 1, in order to overcome the shortcomings of network degradation and high computational complexity of the ResNet structure, a lightweight module (SRFM) that can better transmit gradients and optimize feature information is proposed. This module uses the BasicBlock as the basic residual block and embeds the FasterNet Block module inside each residual block to achieve faster feature processing. The FasterNet Block module includes an MLP structure, partial convolution operations (PConv), and a channel adjustment function for feature fusion and channel adaptation. This module adds the input features to the channel-adapted features through the residual connection mechanism and combines the DropPath operation, thereby achieving more efficient information transmission and feature learning.
[0013] 3. A lightweight small target detection method based on improved RT-DETR according to claim 1, characterized in that the neural network block (DMSFM) with a multi-branch structure proposed in step 2 can fuse feature representations at multiple levels and angles by introducing convolution operations and pooling operations in different branches, so as to achieve a comprehensive description of the input data. Finally, a non-linear transformation is performed through a non-linear activation function to improve the expression ability of the model, capture complex data relationships, solve classification problems, and avoid the problem of gradient disappearance, thereby obtaining the final output. In the inference stage, the fused main branch is adopted. Through six structure reparameterization transformation methods, equivalent convolution kernels and bias parameters are obtained, and a new convolution layer is created accordingly. This new convolution layer represents the merged result of the original branches and becomes the main branch in the inference stage, effectively reducing the complexity and computational cost of the model.
[0014] 4. A lightweight small target detection method based on improved RT-DETR according to claim 1, characterized in that the feature fusion structure AEFN proposed in step 3, on the basis of inheriting the bottom-up path of PAN, innovatively embeds a specially designed channel attention module (CHM) in its bottom path, aiming to enhance the attention of features. For the feature map input to the channel attention module, the module first performs global max pooling and global average pooling operations on each channel, calculates the maximum eigenvalue and average eigenvalue of each channel, and inputs these two values into a shared 2D convolution layer. This convolution layer first compresses the number of channels to 1 / r (i.e., the reduction rate) of the original, then expands it to the original number of channels, and obtains two activated results through the Hardswish activation function, namely the global maximum feature vector and the average feature vector. These two vectors are used to learn the attention weights of each channel, enabling the network to adaptively determine which channels are more important for the current task. Subsequently, the global maximum feature vector and the average feature vector are interacted to obtain the final attention weight vector. To ensure that the attention weights are between 0 and 1, a Sigmoid activation function is used for processing. Finally, the obtained attention weights are multiplied by each channel of the original feature map to obtain the attention-weighted channel feature map. This process strengthens the channel features beneficial to the task while suppressing the influence of irrelevant channels.
[0015] 5. A lightweight small target detection method based on improved RT-DETR according to claim 1, characterized in that the specific method of step 5 is as follows:
[0016] Step 5.1, Input the preprocessed and enhanced data into the backbone network. Extract the main features of the image through the lightweight module (SRFM). The residual structure in this module can retain the key information of the input features, while PConv (Partial Convolution) performs spatial feature extraction by applying conventional convolution only on some input channels, reducing redundancy and the memory access frequency. As the network depth increases, higher-level semantic features can be extracted and combined with low-level semantic features.
[0017] Step 5.2, Select the output features of the last stage of the backbone network as the input of the encoder. By introducing the attention-based intra-scale feature interaction (AIFI) module, the self-attention mechanism can process high-level features rich in semantic information. This helps the model better understand the correlations between different features in the image, thereby improving the accuracy of object detection. At the same time, considering that low-level features lack clear semantic concepts and may cause confusion when interacting with high-level features, it is not necessary to perform low-level feature interaction within the same scale.
[0018] Step 5.3, After completing the processing of the intra-scale feature interaction module, pass its output to the CNN-based cross-scale feature fusion module. First, fuse the feature representations from different levels and angles through upsampling operations and neural network blocks with multi-branch structures. Then, use downsampling operations and attention modules to enhance and transform the features. Finally, fuse the intermediate layer feature information in the FPN path and the enhanced low-level feature information in the feature fusion network (AEFN) based on the PAN structure.
[0019] Step 5.4, After the features are processed by the encoder, use the IoU-aware query mechanism to select a fixed number of image features from the feature sequence output by the encoder as the initial object queries for the decoder. Through the constraints during the training process, the model can assign higher classification scores to features with higher IoU scores and lower classification scores to features with lower IoU scores. Therefore, the prediction boxes corresponding to the top K encoder features selected by the model show excellent performance in both classification scores and IoU scores.
[0020] Step 5.5, The decoder with an auxiliary prediction head iteratively optimizes the object queries to generate accurate confidence boxes and corresponding confidence scores. Description of the Drawings
[0021] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0022] Figure 1 is the network architecture diagram improved based on RT-DETR;
[0023] Figure 2 is the structural diagram of the lightweight module (SRFM);
[0024] Figure 3 is the structural diagram of the neural network block with multi-branch structure (DMSFM);
[0025] Figure 4 is the structural diagram of the attention module (CHM) and the feature fusion network (AEFN) based on the PAN structure;
[0026] Figure 5 is the flowchart of a lightweight small target detection method based on the improved RT-DETR of the present invention; Detailed implementation manners
[0027] The present invention will be further described in detail below with reference to the accompanying drawings of the specification.
[0028] The main improvements of a lightweight small target detection method based on the improved RT-DETR are as follows:
[0029] Step 1: Combine the appendix Figure 1 as shown. The steps are as follows:
[0030] Step 1.1, preprocess the initial dataset, including transformations such as random rotation, flipping, and translation, to increase the diversity and generalization of the data. Subsequently, use the multi-scale retinal color restoration algorithm to optimize the brightness, contrast, and color balance of the image, thereby achieving a more natural and effective image enhancement effect.
[0031] Step 1.2, input the preprocessed and enhanced data into the backbone network. Extract the main features of the image through the lightweight module (SRFM) that can better transmit gradients and optimize feature information. The residual structure in this module can retain the key information of the input features, while PConv (Partial Convolution) extracts spatial features by applying conventional convolution only on some input channels, reducing redundancy and the memory access frequency. As the network depth increases, higher-level semantic features can be extracted and combined with low-level semantic features.
[0032] Step 1.3, select the output features of the last stage of the backbone network as the input of the encoder. By introducing the attention-based intra-scale feature interaction (AIFI) module, the self-attention mechanism can process high-level features rich in semantic information. This helps the model to better understand the associations between different features in the image, thereby improving the accuracy of target detection. At the same time, considering that low-level features lack clear semantic concepts and may be confused when interacting with high-level features, it is not necessary to perform low-level feature interaction within the same scale.
[0033] Step 1.4, after completing the processing of the intra-scale feature interaction module, its output is passed to the CNN-based cross-scale feature fusion module. First, the feature representations from different levels and perspectives are fused through an upsampling operation and a neural network block with a multi-branch structure. Then, the features are enhanced and transformed using a downsampling operation and an attention module. Finally, the intermediate layer feature information in the FPN path and the enhanced low-level feature information in the feature fusion network (AEFN) based on the PAN structure are fused.
[0034] Step 1.5, after the features are processed by the encoder, the IoU-aware query mechanism is used to select a fixed number of image features from the feature sequence output by the encoder as the initial object queries for the decoder. Through the constraints in the training process, the model assigns higher classification scores to the features with higher IoU scores and lower classification scores to the features with lower IoU scores. Therefore, the predicted bounding boxes corresponding to the top K encoder features selected by the model exhibit excellent performance in both classification scores and IoU scores.
[0035] Step 1.6, the decoder with an auxiliary prediction head iteratively optimizes the object queries to generate accurate confidence bounding boxes and corresponding confidence scores.
[0036] Step Two: Combine the appendix Figure 2 , replace the BasicBlock module in the backbone resnet18 with a lightweight module (SRFM). This module uses BasicBlock as the basic residual block and embeds the FasterNet Block module inside each residual block to achieve faster feature processing. The FasterNetBlock module contains an MLP structure, partial convolution operations (PConv), and a channel adjustment function for feature fusion and channel adaptation. The lightweight module adds the input features and the channel-adapted features through a residual connection mechanism and combines the DropPath operation, thus achieving more efficient information transmission and feature learning. This module can be expressed as:
[0037] Output = mlp_layer(PConv(Conv 3×3 (S)))+Conv 3×3 (S)+S
[0038] S represents the input feature information, and the mlp_layer contains two 1×1 convolutional layers. Given that overusing normalization layers and activation layers throughout the network may suppress the diversity of features and thus affect performance, a BatchNorm layer and a ReLu activation function layer are introduced only after the first convolutional layer. In addition, this module uses PConv, which performs conventional convolution only on some input channels to extract spatial features, while the other channels remain unchanged. To ensure the coherence or regularity of memory access, the first or the last consecutive several channels are selected as the representatives of the entire feature map for calculation. Therefore, the computational complexity (FLOPs) of PConv is only
[0039]
[0040] Typical partial ratio So the FLOPs of PConv are only 1 / 16 of those of conventional Conv. In addition, the memory access amount of PConv is small, that is
[0041]
[0042] For r = 1 / 4, this is only 1 / 4 of that of conventional Conv. Therefore, the adoption of PConv in this module greatly reduces the computational complexity and the number of parameters of the model, achieving the lightweight of the model.
[0043] The output of the PConv layer is processed by two 1×1 convolutional layers (mlp_layer), which can further extract and transform features, enabling the model to focus more on the key feature information for classification and regression tasks, while filtering out unnecessary redundant information and noise. On this basis, the lightweight module introduces a 3×3 convolution and a residual structure, which not only more effectively retains the information of the input features, but also significantly improves the accuracy and stability of the model. The module successfully achieves the lightweight of the model by adopting the FasterNetBlock module, DropPath operation, and combining channel adaptation and residual structure. These designs not only optimize the model performance, but also reduce the computational complexity and shrink the model size, thus reducing the computational cost and resource consumption while maintaining excellent performance, meeting the actual needs of lightweight model design.
[0044] Role of components in the lightweight module structure (SRFM): The design of the local receptive field of PConv makes the network pay more attention to the correlation between local features, which helps to capture information and structure finely. The output of the PConv layer will be processed by two 1×1 convolutional layers, and the functions of these two convolutional layers can be summarized into two aspects: feature mapping and feature transformation. The first 1×1 convolutional layer is responsible for mapping the output feature map of the PConv layer to a new feature space, so that the model can better learn the features required for classification and regression tasks; while the second 1×1 convolutional layer transforms the mapped features to obtain more accurate classification and regression results. In addition, the introduced 3×3 convolution and residual structure can further extract and process the information of the input features, strengthen the model's learning ability of features, make the model show a certain robustness when facing the noise and interference in the training data, and thus improve the stability of the model.
[0045] Step 3: Combine the attached Figure 3 , replace the RepBlock in the Fusion module of the baseline model with the neural network block (DMSFM) with a multi-branch structure, and send the initial features to the multi-branch structure. The functions of each branch are as follows:
[0046] dbb_origin: This branch maintains the feature representation of the original input and plays a role in transmitting information.
[0047] dbb_1x1: Use a 1×1 convolution kernel to reduce the dimension of the input features and achieve channel reduction in order to extract key features more precisely.
[0048] dbb_avg: Downsample the input through average pooling operation, reduce the spatial dimension, and then expand the receptive field to achieve the capture of global features.
[0049] dbb_1x1_kxk: Combine 1×1 and k×k convolution kernels to achieve combined extraction of features. Among them, the 1×1 convolution is responsible for adjusting the number of channels and reducing the computational burden; while the k×k convolution focuses on feature extraction within the local area. This combination enables the network to capture the features of the input data more comprehensively, thereby enhancing the expression ability of the network.
[0050] In this module, there are two identical dbb_avg and dbb_1x1_kxk branches. The advantage of this design is that even though their structures are the same, due to the randomness of the learning process, they may extract different feature representations. This design strategy ensures that when one branch extracts unstable features, the other identical branch can provide supplementation for it, thereby optimizing the overall performance of the model.
[0051] Accumulating the results of multiple different branches can fuse feature representations at multiple levels and angles, thereby enhancing the expressive power of the network. Finally, a non-linear transformation is performed through a non-linear activation function to improve the expressive power of the model, capture complex data relationships, solve classification problems, and avoid the vanishing gradient problem, and then obtain the final output.
[0052] In the inference stage, the fused main branch is adopted. Through six structural reparameterization transformation methods, equivalent convolution kernels and bias parameters are obtained, and a new convolutional layer dbb_reparam is created accordingly. This new convolutional layer represents the merged result of the original branches and becomes the main branch in the inference stage, effectively reducing the complexity and computational cost of the model.
[0053] When the model is in the deployment state, it can be described as:
[0054] output = nonlinear(dbb_reparam)(inputs)
[0055] When the model is in the non-deployment state, it can be described as:
[0056] output = nonlinear{(dbb_origin)(inputs)+(dbb_1x1)(inputs)+(dbb_avg)(inputs)+(dbb_avg2)(inputs)+(dbb_1x1_kxk)(inputs)+(dbb_1x1_kxk2)(inputs)}
[0057] Among them, (dbb_reparam) represents the convolutional operation in the deployment state, nonlinear represents the activation function operation, and (dbb_origin), (dbb_1x1), (dbb_avg), (dbb_avg2), (dbb_1x1_kxk), (dbb_1x1_kxk2) respectively represent the operations of each branch in the non-deployment state.
[0058] This module significantly enhances the representational ability and generalization performance of the network by virtue of its multi-branch operation and feature integration mechanism. This enables the network to exhibit excellent performance when dealing with complex tasks and demonstrates a high degree of flexibility and performance advantages.
[0059] Step Four: Combine the appendix Figure 4, based on inheriting the bottom-up path of PAN, the structure (AEFN) innovatively embeds a specially designed channel attention module (CHM) in its bottom path to enhance the attention of features. In the channel attention module, the bottom features are enhanced and transformed. Subsequently, through a concatenation operation, the intermediate layer feature information in the FPN path is fused with the enhanced bottom feature information in the feature fusion network (AEFN) based on the PAN structure, thereby ensuring that the fused features contain rich semantic and location information.
[0060] For the feature map input to the attention module, the module first performs global max-pooling and global average-pooling operations on each channel, calculates the maximum eigenvalue and average eigenvalue of each channel, and inputs these two values into a shared 2D convolutional layer. This convolutional layer first compresses the number of channels to 1 / r (i.e., the reduction rate) of the original, then expands it back to the original number of channels, and obtains two activated results through the Hardswish activation function, namely the global maximum feature vector and the average feature vector. These two vectors are used to learn the attention weights of each channel, enabling the network to adaptively determine which channels are more important for the current task. Subsequently, the global maximum feature vector and the average feature vector are interacted to obtain the final attention weight vector. To ensure that the attention weights are between 0 and 1, the Sigmoid activation function is used for processing. Finally, the obtained attention weights are multiplied by each channel of the original feature map to obtain the attention-weighted channel feature map. This process strengthens the channel features beneficial to the task while suppressing the influence of irrelevant channels.
[0061] Channel attention (CHM) formula:
[0062] Conv(x) = Conv2d(Hordsuish(Conv2d(x, W1)), W2)
[0063] M ch (F org ) = σ(Conv(Avgpool(F org )) + Conv(Maxpool(F org ))
[0064] M ch (F org ) = σ(W2(W1(F cang )) + W2(W1(F cmax )))
[0065] In the shared 2D convolutional layer, the traditional ReLU activation function is abandoned, and the Hardswish activation function is adopted instead. Hardswish is an unbounded, non-monotonic, and smooth function, which helps to improve the accuracy of the model and alleviate the problems of gradient explosion and disappearance that may occur during training. By using the Hardswish activation function, the accuracy loss caused by the approximation of the sigmoid function in different implementation methods can be avoided. Selecting the Hardswish activation to normalize the feature map can further enrich the difference and expressiveness of the features, thus helping the attention mechanism to more accurately focus on important regions and suppress non-critical information.
[0066] The feature fusion network based on the PAN structure not only strengthens the underlying features but also retains the position information, significantly improving the detection performance of small objects. This module effectively solves the problem of detecting small-sized objects under complex background interference and significantly enhances the overall performance of object detection.
[0067] Step 5: Build an improved RT-DETR object detection model architecture for training and validation. The steps are as follows:
[0068] Step 5.1: The example model of the present invention uses a series of parameter settings for experimental evaluation. First, a CPU equipped with an Intel Core i5-8264U processor is used, and an NVIDIA GeForce RTX 3090Ti graphics card is installed for accelerated computing, with CUDA version 11.7. The code environment is built based on Python 3.8.13 and PyTorch 1.12.1. The image size is set to 640×640 pixels, AdamW is selected as the optimizer, the batch size is set to 8, and the momentum is 0.9. 300 iteration training cycles are carried out, and the weight decay is set to 0.0001. The initial learning rate is 0.0001, and the warm-up momentum is 0.8.
[0069] Step 5.2: Evaluate the trained model, and use complexity GFLOPs, mean average precision mAP, recall, number of parameters Params, and precision P to evaluate the performance of the network. Specifically, AP, respectively describe the overall accuracy performance of the detector at different IOU thresholds 0.5:0.05:0.95, 0.5, 0.75. AP s ,AP m ,AP l is used to evaluate the detection effect of the detector on objects of different scales (small objects: area < 322, medium objects: 322 < area < 962, large objects: area > 962).
[0070] Table 1
[0071] Ablation experiments were conducted on the URPC2020 dataset. The effectiveness of the three methods was verified in five evaluation metrics: Recall, P, AP, Params, and GFLOPs.
[0072]
[0073] Ablation experiments were conducted on the URPC2020 dataset. In Experiment A, the unadjusted baseline was used as a reference, and all other experiments were compared and analyzed based on this baseline. In Experiment B, the neural network module with a multi-branch structure was used to replace the RepBlock module of the baseline, resulting in a 1.1% increase in recall rate, a 1.8% increase in the P metric, a 0.9% increase, and an 0.8% increase in AP. In Experiment C, the BasicBlock in the backbone network resnet18 of the baseline was replaced with a lightweight module. Compared with the baseline, the recall rate increased by 0.7%, the P metric increased by 0.4%, a 0.8% increase, and an 0.8% increase in AP. The number of parameters Params decreased by 3.09M, and the complexity GFLOPs decreased by 7.5. In Experiment D, the feature fusion network module based on the PAN structure was used to replace the original PAN structure of the baseline, resulting in a 1% increase in recall rate, a 0.5% increase in the P metric, a 0.8% increase, and an 0.6% increase in AP. In Experiment E, the neural network module with a multi-branch structure was used to replace the RepBlock module, and the lightweight module was used to replace the BasicBlock module in the backbone resnet18, resulting in a 1.7% increase in recall rate, a 2% increase in the P metric, a 1% increase, and an 1.1% increase in AP. In Experiment F, the neural network module with a multi-branch structure was used to replace the RepBlock module, the lightweight module was used to replace the BasicBlock module in the backbone resnet18, and the feature fusion network module based on the PAN structure was used to replace the PAN module, resulting in a 2.3% increase in recall rate, a 2.4% increase in the P metric, a 1.1% increase, and an 1.3% increase in AP. The number of parameters Params decreased by 2.96M, and the complexity GFLOPs decreased by 7.3.
[0074] Table 2
[0075] Ablation experiments were conducted on the NWPUVHR-10 dataset. The effectiveness of the three methods was verified in five evaluation metrics: Recall, P, AP, Params, and GFLOPs.
[0076]
[0077] Ablation experiments were conducted on the NWPU VHR-10 dataset. Experiment A used the unadjusted baseline as a reference, and all other experiments were compared and analyzed based on this baseline. In Experiment B, the neural network module with a multi-branch structure was used to replace the RepBlock module of the baseline, resulting in a 0.8% increase in recall rate, a 0.6% increase in the P metric, a 0.6% increase, and a 1.1% increase in AP. In Experiment C, the BasicBlock in the backbone network resnet18 of the baseline was replaced with a lightweight module. Compared with the baseline, the recall rate increased by 0.6%, a 0.7% increase, a 1% increase in AP, the number of parameters Params decreased by 3.09M, and the complexity GFLOPs decreased by 7.5. In Experiment D, the feature fusion network module based on the PAN structure was used to replace the original PAN structure of the baseline, resulting in a 1.2% increase in recall rate, a 0.4% increase in the P metric, a 0.7% increase, and a 0.3% increase in AP. In Experiment E, the feature fusion network module based on the PAN structure was used to replace the original PAN structure of the baseline, and the BasicBlock module in the backbone resnet18 was replaced with a lightweight module, resulting in a 1.6% increase in recall rate, a 0.6% increase in the P metric, a 0.7% increase, and a 1.3% increase in AP. In Experiment F, the neural network module with a multi-branch structure was used to replace the RepBlock module of the baseline, and the feature fusion network module based on the PAN structure was used to replace the PAN module, resulting in a 1.8% increase in recall rate, a 1.1% increase in the P metric, a 1.1% increase, and a 0.4% increase in AP. In Experiment G, the neural network module with a multi-branch structure was used to replace the RepBlock module, the BasicBlock module in the backbone resnet18 was replaced with a lightweight module, and the feature fusion network module based on the PAN structure was used to replace the PAN module, resulting in a 2.9% increase in recall rate, a 2.7% increase in the P metric, a 1.4% increase, a 1.4% increase in AP, the number of parameters Params decreased by 2.96M, and the complexity GFLOPs decreased by 7.3.
[0078] The performance of the proposed method and RT-DETR-r18 on URPC2020 and NWPU VHR-10 was compared in Tables 1 and 2. The method improved based on RT-DETR ensured It increased by more than 1%, more than 1% on the NWPUVHR-10 dataset. The experiments, under more stringent and fairer metrics, demonstrated the good generalization ability of the improved method based on RT-DETR on the URPC2020 and NWPU VHR-10 datasets. In addition, through experimental verification, the method showed significant performance improvements in multiple aspects such as Recall, P, AP, Params, and GFLOPs, further confirming the effectiveness of the method.
[0079] Table 3
[0080] Comparison of different models on the RSOD dataset for categories and mAP
[0081]
[0082]
[0083] Table 4 Comparison of different models on the RSOD dataset for mAP, FPS, and number of parameters Params
[0084]
[0085] Table 3 shows the comparison of the results of this study with other studies in the field that used the RSOD dataset to evaluate the effectiveness of the proposed model. For easy identification, the most successful results in each category in the table are marked in bold black, and the second-best results are presented in bold green. Through this comparison, the relative advantages and status of this study in the field can be clearly seen. For the aircraft class, the proposed model reached 97.6%, with the highest detection accuracy, a 3.9% increase compared to the latest detector RT-DETR, and a 6.6% increase for the overpass class. The mAP of the proposed model also reached the highest, a 5.1% and 2.4% increase compared to Mobile Vit and TRCNet, which are also based on the transformer model, and a 0.9% increase compared to the latest detector RT-DETR. Table 4 compares with other methods in the field in terms of mAP, FPS, and number of parameters. From the results, it can be seen that the proposed model still has the highest accuracy, reaching 95.3%, 1.65% higher than the second-place MobileNetV2. The lightweight implementation reduced the number of parameters of the proposed model to 16.93M. At the same time, the FPS of the model still remained at a high level, reaching 65.6. This performance stands out among the research methods in the same field and is better than most competitors.
[0086] Table 5 Comparison of different models on the NWPUVHR-10 dataset for categories and mAP
[0087]
[0088] Comparison of Different Models on mAP, FPS, and Parameter Quantity Params on the NWPU VHR-10 Dataset
[0089]
[0090] Table 5 shows the comparison of the results of this study with those of other studies in this field that used the NWPU VHR-10 dataset to evaluate the effectiveness of the proposed model. For easy identification, the most successful results in each category in the table are marked in bold black, and the second-best results are presented in bold green. It can be seen that the proposed method shows excellent detection accuracy in 5 out of 10 categories. Notably, the Baseball Diamond category has the highest detection accuracy of 99.1%, with a 16.97% performance improvement compared to COPD. For the Tennis Court and Bridge categories, the proposed model also shows significant performance improvements. Specifically, compared with RT-DETR, the detection accuracy of the Tennis Court category has increased by 3.1%; compared with FMSSD, it has increased by 4.4%; compared with YOLOv8, it has increased by as much as 18.28%. For the Bridge category, compared with RT-DETR, the performance of the proposed model has increased by 2.2%; compared with FMSSD, it has increased by 7%; compared with RICAO, it has increased by 18.57%. In addition, the method improved based on RT-DETR reached the highest level of 91.4% in terms of mAP. By comparing with other methods in the same field in Table 6, it can be clearly seen that the proposed model shows obvious competitive advantages in terms of mAP, FPS, and parameter quantity.
[0091] Comparison Experiments on the URPC2020 Dataset in Table 7
[0092]
[0093] Comparison of the URPC2020 Dataset: The comparison results of the URPC2020 dataset reveal that the method improved based on RT-DETR outperforms the state-of-the-art method RT-DETR in terms of performance. As shown in Table 7, the method improved based on RT-DETR reached an average precision AP of 50.2, which is better than 48.9 of RT-DETR; in In terms of metrics, the 84.8 of the method improved based on RT-DETR also exceeds the 83.4 of RT-DETR, showing obvious advantages. Compared with other models in the DETR series, the method improved based on RT-DETR also performs excellently. In the evaluation of different target scales, the method improved based on RT-DETR is 10.1% higher than Deformable DETR on small-scale targets, 4.3% ahead on medium-scale targets, and 1.8% higher on large-scale targets. In addition, the method improved based on RT-DETR is 1.2% higher than the state-of-the-art method in APl and APm respectively, proving its comprehensive advantages in object detection at different scales.
[0094] It is worth mentioning that while maintaining high object localization and classification capabilities, the method improved based on RT-DETR has the fewest parameters. This feature enables the method improved based on RT-DETR to achieve a good balance between model complexity and performance, further highlighting its superiority in practical applications.
[0095] Step Six: Combine the appendix Figure 5 As shown below. Implement the above steps as follows:
[0096] Step 1: RT-DETR is used as the object detection framework, and the backbone network extraction uses an improved lightweight module (SRFM) to replace the BasicBlock module in resnet18. This module uses the residual connection mechanism to effectively transmit gradients and optimization information. In addition, the input features are localized through partial convolution operations, thereby further reducing the number of parameters of the model;
[0097] Step 2: Use the proposed neural network block with a multi-branch structure (DMSFM) to replace the RepBlock in the Fusion module of the baseline model to fuse feature representations at multiple levels and angles; in the inference stage, use the fused main branch for inference to reduce the model complexity and computational cost;
[0098] Step 3: Use the proposed feature fusion network based on the PAN structure (AEFN). Through the proposed attention module, important features are enhanced and unimportant features are suppressed to enhance the position and semantic information of small targets;
[0099] Step 4: Use the URPC2020, NWPU VHR-10, and RSOD public dataset object detection datasets to train the improved RT-DETR object detection network model, configure the training parameters, and set the number of training epochs;
[0100] Step 5: Input the small target detection dataset into the improved RT-DETR target detection network model that has been trained. After feature extraction and multi-scale feature fusion by the model, the detection results of small targets of various categories are output, realizing small target detection based on the improved RT-DETR.
[0101] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A lightweight small target detection method based on improved RT-DETR, characterized in that, It includes the following steps: Step 1: RT-DETR is a target detection framework. In the backbone network extraction, the improved lightweight module (SRFM) is used to replace the BasicBlock module in ResNet18. This module utilizes the residual connection mechanism to effectively transmit gradients and optimization information. In addition, local processing of the input features is performed through partial convolution operations, thereby further reducing the number of model parameters; Step 2: The proposed neural network block with a multi-branch structure (DMSFM) is used to replace the RepBlock in the Fusion module of the baseline model to fuse feature representations at multiple levels and angles; In the inference stage, the fused main branch is used for inference to reduce the model complexity and computational cost; Step 3: The proposed feature fusion network based on the PAN structure (AEFN) is adopted. Through the proposed attention module, important features are enhanced and unimportant features are suppressed to enhance the location and semantic information of small targets; Step 4: The improved RT-DETR target detection network model is trained using the URPC2020, NWPU VHR-10, and RSOD public dataset target detection datasets, configuring the training parameters and setting the number of training epochs; Step 5: The small target detection dataset is input into the trained improved RT-DETR target detection network model. After feature extraction and multi-scale feature fusion, the model outputs the detection results of various small targets, realizing small target detection based on the improved RT-DETR.
2. The lightweight small target detection method based on improved RT-DETR according to claim 1, wherein In Step 1, to overcome the shortcomings of network degradation and high computational complexity of the ResNet structure, a lightweight module (SRFM) that can better transmit gradients and optimize feature information is proposed. This module uses BasicBlock as the basic residual block, and the FasterNet Block module is embedded inside each residual block to achieve faster feature processing. The FasterNet Block module contains an MLP structure, partial convolution operations (PConv), and a channel adjustment function for feature fusion and channel adaptation. This module adds the input feature to the channel-adapted feature through the residual connection mechanism and combines the DropPath operation, thereby achieving more efficient information transmission and feature learning.
3. The lightweight small target detection method based on improved RT-DETR according to claim 1, characterized in that The neural network block (DMSFM) with a multi-branch structure proposed in step 2 can fuse feature representations at multiple levels and angles by introducing convolutional operations and pooling operations in different branches, so as to achieve a comprehensive description of the input data. Finally, a non-linear transformation is performed through a non-linear activation function to improve the expressive power of the model, capture complex data relationships, solve classification problems, and avoid the problem of gradient disappearance, thereby obtaining the final output. In the inference stage, the fused main branch is adopted. Through six structural reparameterization transformation methods, equivalent convolutional kernels and bias parameters are obtained, and a new convolutional layer is created based on this. This new convolutional layer represents the merged result of the original branches and becomes the main branch in the inference stage, effectively reducing the complexity and computational cost of the model. The six structural reparameterization transformation methods are Transform I: a conv for conv-BN, Transform II: a conv for branch addition, Transform III: a conv for sequential convolutions, Transform IV: a conv for depth concatenation, Transform V: a conv for average pooling, Transform VI: a conv for multi-scale convolutions.
4. A lightweight small target detection method based on improved RT-DETR according to claim 1, characterized in that The feature fusion structure AEFN proposed in step 3, while inheriting the bottom-up path of PAN, innovatively embeds a specially designed channel attention module (CHM) in its bottom path to enhance feature attention. For the feature map input to the channel attention module, the module first performs global max pooling and global average pooling operations on each channel, calculates the maximum eigenvalue and average eigenvalue of each channel, and inputs these two values into a shared 2D convolutional layer. This convolutional layer first compresses the number of channels to 1 / r (i.e., the reduction rate) of the original, and then expands it back to the original number of channels, and obtains two activated results through the Hardswish activation function, namely the global maximum feature vector and the average feature vector. These two vectors are used to learn the attention weights of each channel, enabling the network to adaptively determine which channels are more important for the current task. Subsequently, the global maximum feature vector and the average feature vector are interacted to obtain the final attention weight vector. To ensure that the attention weights are between 0 and 1, a Sigmoid activation function is used for processing. Finally, the obtained attention weights are multiplied by each channel of the original feature map to obtain the attention-weighted channel feature map. This process strengthens the channel features beneficial to the task while suppressing the influence of irrelevant channels.
5. A lightweight small target detection method based on improved RT-DETR according to claim 1, characterized in that The specific method of step 5 is as follows: Step 5.1, Input the preprocessed and enhanced data into the backbone network. Extract the main features of the image through the lightweight module (SRFM). The residual structure in this module can retain the key information of the input features, while PConv (Partial Convolution) performs spatial feature extraction by applying conventional convolution only on some of the input channels, reducing redundancy and the memory access frequency. As the network depth increases, higher-level semantic features can be extracted and combined with low-level semantic features. Step 5.2, Select the output features of the last stage of the backbone network as the input of the encoder. By introducing the attention-based intra-scale feature interaction (AIFI) module, the self-attention mechanism can process the high-level features rich in semantic information. This approach helps the model to better understand the correlations between different features in the image, thereby improving the accuracy of object detection. At the same time, considering that low-level features lack clear semantic concepts and may be confused when interacting with high-level features, it is not necessary to perform low-level feature interaction within the same scale. Step 5.3, After completing the processing of the intra-scale feature interaction module, pass its output to the CNN-based cross-scale feature fusion module. First, fuse the feature representations from different levels and perspectives through the upsampling operation and the neural network block with a multi-branch structure. Then, use the downsampling operation and the attention module to enhance and transform the features. Finally, fuse the intermediate layer feature information in the FPN path and the enhanced low-level feature information in the feature fusion network (AEFN) based on the PAN structure. Step 5.4, After the features are processed by the encoder, use the IoU-aware query mechanism to select a fixed number of image features from the feature sequence output by the encoder as the initial object queries for the decoder. Through the constraints during the training process, the model can assign higher classification scores to the features with higher IoU scores and lower classification scores to the features with lower IoU scores. Therefore, the prediction boxes corresponding to the top K encoder features selected by the model show excellent performance in both classification scores and IoU scores. Step 5.5, The decoder with an auxiliary prediction head iteratively optimizes the object queries to generate accurate confidence boxes and corresponding confidence scores.