Small target detection method for unmanned aerial vehicle image based on improved DETR

By improving the DETR method, using a dual-domain hybrid encoder and enhanced query selection mechanism, combined with knowledge distillation strategy, the problem of insufficient detection accuracy of small targets in drone images is solved, and efficient and accurate small target detection is achieved.

CN120032277AActive Publication Date: 2025-05-23FUDAN UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510102214.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art has problems of decreasing accuracy and increasing missed detection rates when detecting small targets in drone images, especially in small target dense scenarios.

Method used

Using an improved DETR method, the spatial domain and frequency domain characteristics are integrated through the dual-domain hybrid encoder module, the query selection mechanism is enhanced to optimize query resource allocation, and the model complexity is reduced through knowledge distillation strategies.

Benefits of technology

It significantly improves the accuracy and efficiency of small target detection, reduces missed and missed detection, and ensures efficient operation on resource-limited drone platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032277A_ABST
    Figure CN120032277A_ABST
Patent Text Reader

Abstract

The invention discloses a small target detection method for an unmanned aerial vehicle image based on an improved DETR. The invention is constructed based on an end-to-end attention target detection model DETR, and provides three key technical improvements: a double-domain hybrid encoder module, an enhanced query selection mechanism and a knowledge distillation strategy. Wherein the double-domain hybrid encoder module is used for integrating spatial domain and frequency domain features, only operating high-level features through a self-attention mechanism, and combining low-level features by using the double-domain fusion module; secondly, an enhanced query selection mechanism is introduced, anchor frames are generated on the multi-scale feature map, and the high-score anchor frames are dynamically screened in combination with the expanded intersection-union ratio to serve as query; in addition, through a knowledge distillation strategy, knowledge of the teacher model is migrated to the lightweight student model, and accurate detection and efficient reasoning of a small target are realized. The method has obvious advantages in the aspects of small target detection precision and calculation amount, and can be used for various unmanned aerial vehicle image processing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a small target detection method for unmanned aerial vehicle images based on improved DETR. Background Art

[0002] With the continuous development of drone technology, drones are being used more and more widely in surveying and mapping, inspection, security, agriculture and other fields. However, in the high-altitude images obtained by drones, small targets are usually large in number, small in size, and have variable shapes, which poses a great challenge to the target detection algorithm. Existing target detection methods, especially algorithms for general scenarios, often have problems such as decreased detection accuracy and increased missed detection rate when facing small targets in drone images.

[0003] DETR (Detection Transformer) is an object detection method based on the attention mechanism that has emerged in recent years. By treating object detection as a set prediction problem, it simplifies the dependence of many traditional detectors on prior anchor boxes. However, DETR still has the following shortcomings when detecting small objects and multi-object dense scenes:

[0004] 1. Spatial resolution loss: Direct processing in the feature pyramid or deeper network layers often results in insufficient information of small objects in high-dimensional semantic feature maps;

[0005] 2. Single query mechanism: A fixed number of queries is difficult to dynamically adapt to the diversity of target distribution in the image. Especially for areas with dense small targets, query resource allocation is prone to waste or insufficient allocation.

[0006] 3. Model complexity: On drone platforms or other scenarios with limited computing resources, deploying larger models will result in higher inference latency and memory overhead. Summary of the invention

[0007] The purpose of the present invention is to overcome the problems of insufficient accuracy and efficiency in the prior art for detecting small targets in drone images, and to provide a method for detecting small targets in drone images based on improved DETR. The method is implemented by the following technical solutions:

[0008] S100: receiving a drone image to be subjected to target detection;

[0009] S200: Use the backbone network to extract features from the input drone image and output the extracted features to the dual-domain hybrid encoder module;

[0010] S300: Dual-domain mixed encoding steps:

[0011] S310: Self-attention operation: In the dual-domain hybrid encoder module, the highest level features (s 5 Layer) performs self-attention operation;

[0012] S320: Dual-domain fusion operation: The dual-domain fusion module is used to convert s 2 and 3 The features of the layers are fused and the fused features are output to the enhanced query selection mechanism;

[0013] S400: Enhanced query selection operation:

[0014] S410: Anchor box generation: generating anchor boxes on multi-scale feature maps in an enhanced query selection mechanism;

[0015] S420: high-scoring anchor box selection: using the classification head and the regression head, dynamically selecting high-scoring anchor boxes as queries through the extended intersection-over-union ratio;

[0016] S500: an auxiliary prediction head is provided in the decoder to iteratively optimize the selected query to predict the category and anchor box of the target;

[0017] S600: Knowledge distillation step:

[0018] S610: Build a teacher model based on RT-DETR-R50 and a student model derived from RT-DETR-R18. The student model uses EfficientFormerV2 to replace the traditional ResNet-18 backbone network.

[0019] S620: Calculate the distillation loss for classification and positioning through binary cross entropy loss, L1 loss and intersection-over-union loss;

[0020] S630: Use linear decay strategy to adjust the weight of distillation loss;

[0021] S700: Output the detection result including the target category and positioning information.

[0022] Preferably, the so-called input includes an input drone image to be subjected to target detection.

[0023] Preferably, the backbone network, including the backbone network of the teacher model, is ResNet50, and the backbone network of the student model is EfficientFormerV2.

[0024] Preferably, the specific operation of the dual-domain fusion module in the dual-domain hybrid encoder module is as follows: for a given feature map This module first performs a convolution operation and then performs channel segmentation:

[0025] [X1 ,X 2 ]=Split(Conv(X) , [C 1 ,C 2 ], dim = 0);

[0026] In the above formula, Conv(·) represents a convolutional layer, which adjusts the channel dimension while keeping the spatial dimension unchanged, and C 1 +C 2 =C, C 1 Set to one-fourth of the total number of channels; the first branch X 1 After another convolution, it is passed through a Gaussian error linear unit to introduce nonlinearity:

[0027] X conv =GELU(Conv(X 1 ));

[0028] In order to capture complex feature dependencies, the convolutional feature X conv The following processing is done in the frequency domain:

[0029] X out =α 1 |IFFT(FFT(Conv(|X conv |)))·Conv(|X conv |)|+Conv(ReLU(X 1 +Conv(X conv )+β 1 ·|X conv |));

[0030] In the above formula, α 1 and β 1 are learnable parameters that balance the contributions from frequency domain enhancement features and spatial residual connections; the outputs of the two branches are concatenated and passed through the final convolutional layer to integrate the refined features:

[0031] X final =Conv([X out ,X 2 ]).

[0032] Preferably, the enhanced query selection mechanism includes first generating anchor boxes on the multi-scale feature map using a fixed grid, and then optimizing them through position transformation and selection in logarithmic space. During training, the regression head dynamically adjusts the anchor boxes to match the position and scale of the target. In order to improve the performance of small object detection, the classification head dynamically selects the top k anchor boxes with the highest scores as queries in combination with the extended IoU indicator.

[0033] Preferably, the expanded IoU in the enhanced query selection mechanism is defined as: for the predicted anchor box B' of the model p and the true anchor box B' gt , the expanded IoU is the scaled predicted anchor box B' p and the scaled ground truth anchor box B' gt The intersection-over-union ratio between the two anchor boxes is proportional to the factor α. 2 >1 while keeping their centers unchanged. The scaled anchor boxes are:

[0034] B' p =(x p ,y p ,w p ×α 2 ,h p ×α 2 );

[0035] B′ gt =(x gt ,y gt ,w gt ×α 2 ,h gt ×α 2 );

[0036] Where (x p ,y p ) and (x gt ,y gt ) is the center coordinate, w p ,h p ,w gt ,h gt Represent the width and height of the predicted box and the real box respectively, and the expanded IoU is defined as:

[0037]

[0038] Preferably, the knowledge distillation strategy includes constructing a teacher model based on RT-DETR-R50 and a student model derived from RT-DETR-R18, wherein the derived model replaces the traditional ResNet-18 backbone network with EfficientFormerV2; calculating the distillation loss for classification and positioning through binary cross entropy loss, L1 loss and intersection-over-union loss; and using a linear attenuation strategy to adjust the weight of the distillation loss.

[0039] Preferably, the decoder includes an auxiliary prediction head that iteratively optimizes queries to predict target categories and anchor boxes.

[0040] Preferably, the binary cross entropy loss calculation formula in the knowledge distillation strategy is:

[0041]

[0042] In the above formula, s c and t c They represent the classification scores of the final decoder layer output of the student model and the teacher model respectively; the L1 loss calculation formula is:

[0043]

[0044] Among them, s b and t b Represent the anchor box coordinates predicted by the student model and the teacher model, respectively, o represents the target confidence score of the teacher model; the intersection-over-union loss calculation formula is:

[0045]

[0046] Among them, the calculation formula of the extended SIoU is:

[0047] Expanded_SIoU=SIoU-IoU+Expanded_IoU;

[0048] The calculation formula for distillation loss is:

[0049]

[0050] In the above formula, α, β and γ are constant coefficients that balance the contribution of each loss component.

[0051] Preferably, in the model inference stage, a multi-threaded or asynchronous data acquisition strategy is adopted, combined with software and hardware collaborative optimization, to effectively reduce detection delays, significantly improve the real-time performance of the system, and meet the needs of UAV application scenarios for rapid detection.

[0052] Preferably, by adopting compression techniques such as model pruning and quantization, the model is optimized to reduce the number of model parameters and the amount of calculation, thereby overcoming the limitations of the UAV system in terms of flight time and computing resources, so that the method can run efficiently on UAV equipment with limited resources.

[0053] Preferably, the linear decay strategy includes assigning a higher weight to the distillation loss at the beginning of training. As the training progresses, the weight of the distillation loss gradually decreases.

[0054] Preferably, the so-called output includes outputting detection results, including target category and positioning information.

[0055] Beneficial Effects

[0056] Compared with the prior art, the present invention provides a small target detection method for drone images based on improved DETR, which has the following beneficial effects:

[0057] Based on the improved DETR small target detection method for drone images, the dual-domain hybrid encoder module plays a key role in improving the detection accuracy. It innovatively integrates the spatial domain and frequency domain features, and performs fine processing on the different levels of features output by the backbone network. 5 layer), the self-attention operation enables it to focus on the key details in the high-level semantic information and strengthen the deep understanding of the small target features, while the dual-domain fusion module 2 and 3 The fusion of layer features cleverly supplements the low-level spatial information. The synergy of high- and low-level information is excellent in enhancing high-resolution feature representation. For example, in complex urban mapping drone images, this module can accurately capture the features of small targets such as tiny traffic signs on the street or small ancillary facilities on buildings, avoiding the drawback of traditional methods that cause the loss of small target information due to direct processing in high-dimensional semantic feature maps. Through rigorous testing on the VisDrone-2019-DET and UAVVaste datasets, the accuracy (AP) and precision (AP50) of small target detection have been significantly improved under different computational complexity models, which strongly proves its outstanding performance in improving accuracy.

[0058] In terms of query resource allocation optimization, the enhanced query selection mechanism has brought about a breakthrough. Traditional methods are limited by a fixed number of queries and are difficult to cope with the diversity of target distribution in images, especially in areas with dense small targets, which often fall into the dilemma of query resource allocation. This mechanism of the present invention first uses a fixed grid to generate anchor frames on a multi-scale feature map, and then optimizes them through position transformation and selection in logarithmic space. During the training process, the regression head dynamically adjusts the anchor frame to accurately match the position and scale of the target, and the classification head combines the extended IoU indicator for screening. In the scenario of agricultural plant protection drones monitoring crop pests and diseases, when faced with densely distributed small pest and disease targets in large farmlands, the mechanism can flexibly allocate query resources according to the actual situation of the target, accurately select high-scoring anchor frames as queries, greatly improve the detection efficiency, effectively reduce the occurrence of missed detections and false detections, and ensure the reliability and stability of the detection results.

[0059] The knowledge distillation strategy has achieved remarkable results in model performance optimization and adaptation to drone platforms. By carefully constructing a teacher model based on RT-DETR-R50 and a RT-DETR-R18-derived student model that replaces the traditional backbone network with EfficientFormerV2, the binary cross entropy loss is used to accurately measure the deviation between the student model and the teacher model in classification prediction. The L1 loss is used to weight the anchor box coordinate difference based on the target confidence score of the teacher model. The intersection-over-union loss is introduced to help the student model align the small target anchor box. The linear attenuation strategy is then used to reasonably adjust the distillation loss weight, thus achieving efficient transfer of knowledge from the teacher model to the student model. In the application of drone inspection infrastructure, the knowledge distillation strategy, on the one hand, greatly reduces the model complexity while maintaining high-precision detection capabilities, effectively reducing the model's inference delay and memory overhead; on the other hand, by adopting multi-threaded or asynchronous data acquisition strategies in the model inference stage and combining software and hardware collaborative optimization, as well as using compression technologies such as model pruning and quantization, the knowledge distillation strategy further reduces detection delays, reduces the number of model parameters and the amount of calculation, and successfully overcomes the limitations of drone systems in terms of flight time and computing resources, ensuring that the method can run efficiently and stably on resource-limited drone equipment, greatly expanding its feasibility and scope of application in practical applications, and providing strong technical support for drone intelligent detection in security monitoring, infrastructure inspection, agricultural plant protection, environmental monitoring and other fields, with immeasurable economic and social value. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 This is a model architecture diagram of an improved DETR proposed by the present invention;

[0061] Figure 2 This is the architecture diagram of the dual-domain fusion module;

[0062] Figure 3 A data schematic diagram of the verification results of the method provided by the present invention on the VisDrone-2019-DET dataset;

[0063] Figure 4 A data schematic diagram of the verification results of the method provided by the present invention on the UAVVaste data set;

[0064] Figure 5 A schematic diagram of qualitative comparison of the method provided by the present invention in terms of detection results and attention heat maps. DETAILED DESCRIPTION

[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] The object of the present invention is to provide a small target detection method for drone images based on improved DETR. Figure 1-Figure 5 The present invention is further described in detail. The specific implementation method is as follows:

[0067] See also Figure 1 As can be seen from the overall architecture of the model, a small target detection method for drone images based on improved DETR includes at least the following steps:

[0068] S100: receiving a drone image to be subjected to target detection;

[0069] S200: Use the backbone network to extract features from the input drone image, and output the extracted features to the dual-domain hybrid encoder module. The teacher model uses ResNet50, and the student model uses EfficientFormerV2. The extracted features are output to the dual-domain hybrid encoder module;

[0070] S300: Dual-domain mixed encoding steps:

[0071] S310: Self-attention operation: In the dual-domain hybrid encoder module, the highest level features (s 5 layer) to perform self-attention operation;

[0072] S320: Dual-domain fusion operation: The dual-domain fusion module is used to convert s 2 and 3 The features of the layers are fused and the fused features are output to the enhanced query selection mechanism;

[0073] See also Figure 2 , in the dual-domain fusion module, for a given feature map The model

[0074] The block first performs a convolution operation and then performs channel splitting:

[0075] [X 1 ,X 2 ]=Split(Conv(X),[C 1 ,C 2 ], dim = 0);

[0076] In the above formula, Conv(·) represents a convolutional layer, which adjusts the channel dimension while keeping the spatial dimension unchanged. The above segmentation operation divides the convolutional features into two different branches: X 1 With C 1 Channels, X 2 With C 2 channels, of which C 1 +C 2 =C. Specifically, C 1 Set to one-fourth the total number of channels.

[0077] First Branch X 1 After another convolution, it is passed through a Gaussian error linear unit to introduce nonlinearity:

[0078] X conv =GELU(Conv(X 1 ));

[0079] In order to capture complex feature dependencies, the convolutional feature X conv The following processing is done in the frequency domain:

[0080] X out =α 1 |IFFT(FFT(Conv(|X conv |)))·Conv(|X conv |)|+Conv(ReLU(X 1 +Conv(X conv )+β 1 ·|X conv |));

[0081] In the above formula, α 1 and β 1 are learnable parameters that balance the contributions from frequency domain enhancement features and spatial residual connections. The outputs of the two branches are concatenated together and passed through the final convolutional layer to integrate the refined features:

[0082] X final =Conv([X out ,X 2 ]);

[0083] The generated features are then fed into the fusion module, which outputs the features to the enhanced query selection mechanism.

[0084] S400: Enhanced query selection operation:

[0085] S410: Anchor box generation: generating anchor boxes on multi-scale feature maps in an enhanced query selection mechanism;

[0086] S420: high-scoring anchor box selection: using the classification head and the regression head, dynamically selecting high-scoring anchor boxes as queries through the extended intersection-over-union ratio;

[0087] For the multi-scale feature maps output by the dual-domain hybrid encoder module, anchor boxes are generated on the multi-scale feature maps using a fixed grid. Subsequently, they are optimized through position transformation and selection in logarithmic space. During training, the regression head dynamically adjusts the anchor boxes to match the position and scale of the object. In order to improve the performance of small object detection, the classification head combines the extended IoU metric to dynamically select the top k anchor boxes with the highest scores as queries.

[0088] For the model's predicted anchor box B p and the real anchor box B gt , the expanded IoU is defined as the scaled predicted anchor box B' p and the scaled ground truth anchor box B' gt The intersection-over-union ratio between them. Both anchor boxes are scaled by factor α 2 >1 while keeping their centers unchanged. The scaled anchor boxes are:

[0089] B' p =(x p ,y p ,w p ×α 2 ,h p ×α 2 );

[0090] B' gt =(x gt ,y gt ,w gt ×α 2 ,h gt ×α 2 );

[0091] Among them, (x p ,y p ) and (x gt ,y gt ) is the center coordinate, w p ,h p ,w gt ,h gt Represent the width and height of the predicted box and the true box respectively.

[0092] The expanded IoU is defined as:

[0093]

[0094] Similarly, the extended SIoU is defined as:

[0095] Expanded_SIoU=SIoU-IoU+Expanded_IoU.

[0096] S500: Equip the decoder with an auxiliary prediction head to iteratively optimize the selected query to predict the category and anchor box of the target; for the query generated by the above-mentioned enhanced query selection mechanism, the decoder iteratively optimizes the query to predict the target category and anchor box by equipping the auxiliary prediction head.

[0097] S600: Knowledge distillation step:

[0098] S610: Build a teacher model based on RT-DETR-R50 and a student model derived from RT-DETR-R18. The student model uses EfficientFormerV2 to replace the traditional ResNet-18 backbone network. Both the teacher model and the student model are embedded with the above-mentioned dual-domain hybrid encoder module and enhanced query selection mechanism.

[0099] S620: Calculate the distillation loss for classification and positioning through binary cross entropy loss, L1 loss and intersection-over-union loss;

[0100] S630: Use linear decay strategy to adjust the weight of distillation loss;

[0101] The binary cross entropy loss is used to calculate the deviation between the classification prediction of the student model and the teacher model:

[0102] In the above formula, s c and t c Represent the classification scores of the final decoder layer output of the student model and the teacher model, respectively.

[0103] The L1 loss is used to calculate the difference in the anchor box coordinates and is weighted by the object confidence score of the teacher model:

[0104] Among them, s b and t b Represent the anchor box coordinates predicted by the student model and the teacher model, respectively, o represents the target confidence score of the teacher model.

[0105] The intersection-over-union loss is used to help the student model align the anchor boxes of small objects, thereby improving its performance in small object detection. Its calculation formula is:

[0106] The calculation formula for distillation loss is:

[0107] In the above formula, α, β and γ are constant coefficients that balance the contribution of each loss component.

[0108] The linear decay strategy consists in assigning a higher weight to the distillation loss at the beginning of training, and gradually reducing the weight of the distillation loss as training progresses.

[0109] S700: Output the detection result including the target category and positioning information.

[0110] See also Figure 3 The method provided by this solution is verified on the VisDrone-2019-DET dataset. The models are divided into three groups according to the computational complexity: low computational complexity (Low Computation) models with a computational complexity of less than 50GFLOPs, medium computational complexity (Medium Computation) models with a computational complexity of 50-100GFLOPs, and high computational complexity (High Computation) models with a computational complexity of more than 100GFLOPs. The verification results show that on the VisDrone-2019-DET dataset, the method provided by this solution (SO-DETR) achieved the highest AP and AP in the low, medium, and high computational complexity categories. 50 Fraction.

[0111] See also Figure 4 , the method provided by this solution is verified on the UAVVaste dataset. The verification results show that on the UAVVaste dataset, the method (SO-DETR) provided by this solution can effectively enhance the detection ability of small targets while maintaining competitive computational efficiency. In addition, the verification results on the above two datasets show that the method provided by this solution has shown better performance and robustness on different datasets.

[0112] See also Figure 5 , which shows a heat map of the small target detection results in the VisDrone-2019-DET dataset. As shown in the figure, compared with RT-DETR-EV2, the method provided by this solution (SO-DETR-EV2) has significantly reduced errors in small target detection. The method provided by this solution shows a higher accuracy in small target detection, highlighting its excellent performance in complex scenes.

[0113] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0114] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

[0115] Technical advantages and application prospects

[0116] The small target detection method for UAV images based on improved DETR proposed in the present invention effectively integrates spatial domain and frequency domain features through a dual-domain hybrid encoder, significantly improving the detection accuracy of small targets. At the same time, an enhanced query selection mechanism is used to optimize query resource allocation to ensure efficient detection in small target dense scenarios. Combined with the knowledge distillation strategy, the model is lightweight while maintaining high accuracy, and is suitable for real-time deployment on UAV platforms with limited computing resources. This method has high adaptability and versatility, and is widely used in security monitoring, infrastructure inspection, agricultural plant protection, environmental monitoring, logistics distribution, disaster emergency response, military reconnaissance and other fields. It can significantly improve the intelligent detection capabilities of UAVs in complex environments, and has broad application prospects and significant economic and social value.

[0117] Implementation Effect

[0118] Through experimental verification on the VisDrone-2019-DET and UAVVaste datasets, the small target detection method for drone images based on improved DETR proposed in the present invention shows significant superiority. In addition, the dual-domain hybrid encoder effectively integrates the spatial domain and frequency domain features, enhancing the model's ability to recognize small targets in complex backgrounds; the enhanced query selection mechanism optimizes the query resource allocation to ensure efficient detection in small target dense scenes; the knowledge distillation strategy enables the model to be lightweight while maintaining high precision, making it perform well in a variety of application scenarios such as security monitoring, infrastructure inspection, and agricultural plant protection. In summary, the method of the present invention not only significantly improves the performance of small target detection, but also optimizes computational efficiency and resource utilization, meets the requirements of drone image processing for efficiency and accuracy, and shows broad application prospects and significant economic and social benefits.

[0119] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise one" do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0120] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A small target detection method for UAV images based on improved DETR, characterized by: The following steps are involved: S100: receiving a drone image to be subjected to target detection; S200: Use the backbone network to extract features from the input drone image and output the extracted features to the dual-domain hybrid encoder module; S300: Dual-domain mixed encoding steps: S310: Self-attention operation: In the dual-domain hybrid encoder module, self-attention operation is performed on the highest layer features (s5 layer) from the backbone network; S320: dual-domain fusion operation: the features of the s2 and s3 layers are fused through the dual-domain fusion module, and the fused features are output to the enhanced query selection mechanism; S400: Enhanced query selection operation: S410: Anchor box generation: generating anchor boxes on multi-scale feature maps in an enhanced query selection mechanism; S420: high-scoring anchor box selection: using the classification head and the regression head, dynamically selecting high-scoring anchor boxes as queries through the extended intersection-over-union ratio; S500: an auxiliary prediction head is provided in the decoder to iteratively optimize the selected query to predict the category and anchor box of the target; S600: Knowledge distillation step: S610: Build a teacher model based on RT-DETR-R50 and a student model derived from RT-DETR-R18. The student model uses EfficientFormerV2 to replace the traditional ResNet-18 backbone network. S620: Calculate the distillation loss for classification and positioning through binary cross entropy loss, L1 loss and intersection-over-union loss; S630: Use linear decay strategy to adjust the weight of distillation loss; S700: Output the detection result including the target category and positioning information.

2. A small target detection method for drone images based on improved DETR according to claim 1, characterized in that: The specific operation of the dual-domain fusion module in the dual-domain hybrid encoder module is as follows: This module first performs a convolution operation and then performs channel segmentation: [X1,X2]=Split(Conv(X),[C1,C2],dim=0); In the above formula, Conv(·) represents a convolutional layer, which keeps the spatial dimension unchanged while adjusting the channel dimension, and C1+C2=C, C1 is set to one-fourth of the total number of channels; the first branch X1 is convolved again and then passes through a Gaussian error linear unit to introduce nonlinearity: X conv =GELU(Conv(X1)); In order to capture complex feature dependencies, the convolutional feature X conv The following processing is done in the frequency domain: X out =α1·|IFFT(FFT(Conv(|X conv |)))·Conv(|X conv |)|+Conv(ReLU(X1+Conv(X conv )+β1·|X conv |)); In the above formula, α1 and β1 are learnable parameters, which are used to balance the contribution from frequency domain enhancement features and spatial residual connections respectively; the outputs of the two branches are concatenated together and passed through the final convolutional layer to integrate the refined features: X final =Conv([X out ,X2])。 3. A small target detection method for drone images based on improved DETR according to claim 1, characterized in that: The extended IoU in the enhanced query selection mechanism is defined as: for the model's predicted anchor box B' p and the true anchor box B' gt , the expanded IoU is the scaled predicted anchor box B' p and the scaled ground truth anchor box B' gt The intersection-over-union ratio between them, both anchor boxes are scaled by a factor α2>1 while keeping their centers unchanged. The scaled anchor boxes are: B′ p =(x p ,y p ,w p ×α2,h p ×α2); B′ gt =(x gt ,y gt ,w gt ×α2,h gt ×α2); Where (x p ,y p ) and (x gt ,y gt ) is the center coordinate, w p ,h p ,w gt ,h gt Represent the width and height of the predicted box and the real box respectively, and the expanded IoU is defined as:

4. The small target detection method for drone images based on improved DETR according to claim 1 is characterized in that: The binary cross entropy loss calculation formula in the knowledge distillation strategy is: In the above formula, s c and t c They represent the classification scores of the final decoder layer output of the student model and the teacher model respectively; the L1 loss calculation formula is: Among them, s b and t b Represent the anchor box coordinates predicted by the student model and the teacher model, respectively, o represents the target confidence score of the teacher model; the intersection-over-union loss calculation formula is: Among them, the calculation formula of the extended SIoU is: Expanded_SIoU=SIoU-IoU+Expanded_IoU; The calculation formula for distillation loss is: In the above formula, α, β and γ are constant coefficients that balance the contribution of each loss component.

5. The small target detection method for drone images based on improved DETR according to claim 1 is characterized in that: During the model inference stage, a multi-threaded or asynchronous data collection strategy is adopted, combined with software and hardware collaborative optimization to effectively reduce detection delays, significantly improve the real-time performance of the system, and meet the needs of drone application scenarios for rapid detection.

6. The small target detection method for drone images based on improved DETR according to claim 1, characterized in that: By adopting compression technologies such as model pruning and quantization, the model is optimized, the number of model parameters and the amount of calculation are reduced, and the limitations of the UAV system in terms of flight time and computing resources are overcome, so that this method can run efficiently on UAV equipment with limited resources.

7. The small target detection method for drone images based on improved DETR according to claim 1, characterized in that: This method can be applied to a variety of drone usage scenarios, including security monitoring, inspections, and agricultural plant protection, and can accurately detect different types of small targets, such as vehicles, personnel, animals, and plants, demonstrating the wide applicability and good detection performance of the method in practical applications.

8. A small target detection system for drone images based on improved DETR, characterized by: include: S100: a detection input module, used to input an image to be detected; S200: backbone network, used for preliminary feature extraction; S300: dual-domain hybrid encoder for feature extraction and fusion; S400: an enhanced query selection module for querying resource allocation; S500: decoder, used to predict target categories and anchor boxes; S600: knowledge distillation module, used for model optimization; S700: a detection output module, used to output the detection result.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, the small target detection method for drone images based on improved DETR as described in claims 1 to 4 is implemented, and a carrier for software storage and execution is provided for the implementation of the method.

Citation Information

Patent Citations

  • Unmanned aerial vehicle small sample weak target increment detection and identification method and system

    CN115761549A

  • Cross-scene multi-domain fusion small sample remote sensing target robust identification method

    CN118918476A

  • An optimized approach for disease detection in tomato leaf using deep learning and internet of things.

    IN202011028567A

  • Deep-learning-based method for small target detection in unmanned aerial vehicle scenario

    WO2024108857A1

  • Multi-task joint sensing network model and detection method for traffic road surface information

    WO2024138993A1