Unmanned aerial vehicle aerial image small target detection method based on RT-DETR

By introducing adaptive frequency cavity convolution, context and spatial feature calibration networks and wavelet pooling technology on the RT-DETR algorithm, the problem of difficulty in detecting small objects in aerial images of drones is solved, and higher detection accuracy and robustness are achieved.

CN120032273AActive Publication Date: 2025-05-23UNIV OF ELECTRONICS SCI & TECH OF CHINA

Patent Information

Application Number
CN202510075654.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-23
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

It is difficult to detect small and medium-sized objects in aerial images of drones, and is affected by problems such as shooting height, angle, weather conditions, small target size, and difficult target occlusion and background.

Method used

Using an improved algorithm FCW-RTDETR based on RT-DETR, the network is extracted and fused with image features through adaptive frequency cavity convolution, context and spatial feature calibration network, and wavelet pooling technology, IoU-aware query selection and decoder iterative optimization are performed to generate bounding box and confidence scores.

Benefits of technology

It improves the detection accuracy and robustness of small targets in aerial images of drones, and can effectively detect small targets such as pedestrians and cars in different traffic environments and density scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032273A_ABST
    Figure CN120032273A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle aerial image small target detection method based on RT-DETR, an end-to-end target detection algorithm RT-DETR is improved, an FCW-RTDETR algorithm is provided, and the method comprises the following steps: improving a backbone network by using adaptive frequency cavity convolution, and extracting unmanned aerial vehicle aerial image features; performing intra-scale feature interaction AIFI on a P5 layer in the backbone network, and inputting obtained features into a cross-scale feature fusion network CCFM; the method comprises the following steps: respectively improving a CCFM structure and an up-and-down sampling method by using a context and spatial feature calibration network CSFCN and wavelet pooling, and fusing features of different scales in the CCFM to obtain an image feature sequence; selecting a feature from a feature sequence output by an encoder as an initial target query of a decoder by using IoU-perceived query selection; and the decoder iteratively optimizes the target query through the auxiliary prediction header, generates a bounding box and a confidence score, and completes small target detection. According to the scheme, the small target detection efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and to target detection, deep learning and other technologies, and specifically to a method for detecting small targets in unmanned aerial vehicle aerial images based on RT-DETR. Background Art

[0002] The traffic scenes captured by drones are complex, and drone aerial images are usually affected by many factors such as shooting height, angle, and weather conditions. At the same time, the detected targets have problems such as small size, mutual occlusion between targets, and difficulty in distinguishing from the surrounding background.

[0003] Small target detection has considerable practical value in the fields of drone traffic inspection, drone emergency rescue, traffic flow monitoring and road safety monitoring. Therefore, it is necessary to design a solution for target detection of different targets (pedestrians, cars, bicycles, etc.) in different traffic environments (urban and rural areas) and different densities (sparse and crowded scenes) based on the images taken by drones. Summary of the invention

[0004] In view of the above technical problems, the present invention provides a small target detection method for UAV aerial images based on RT-DETR.

[0005] The present invention is implemented by adopting the following technical scheme: a method for detecting small targets in unmanned aerial images based on RT-DETR, which improves the end-to-end target detection algorithm RT-DETR and proposes the FCW-RTDETR algorithm, including the following steps: Step S1: Use adaptive frequency dilated convolution to improve the backbone network and extract the features of drone aerial images; Step S2: Perform intra-scale feature interaction AIFI on the P5 layer in the backbone network, and input the obtained features into the cross-scale feature fusion network CCFM; Step S3: Use the context and spatial feature calibration network CSFCN to improve the CCFM structure, and use wavelet pooling to improve the up and down sampling methods in CCFM, fuse features of different scales in CCFM and obtain an image feature sequence; Step S4: Using IoU-aware query selection, a fixed number of features are selected from the feature sequence output by the encoder as the initial target query for the decoder; Step S5: The decoder iteratively optimizes the target query through the auxiliary prediction head, generates bounding boxes and confidence scores, and completes small target detection.

[0006] Specifically, the improvement of the adaptive frequency dilated convolution in step S1 includes adaptive dilated rate, adaptive kernel and frequency selection, and the adaptive dilated rate adjusts the dilated rate spatially, specifically including the following steps: Step A1: Using Discrete Fourier Transform The feature map Transformed to the frequency domain, it is expressed as: ; in, An array of complex numbers representing the DFT output; and Represents the feature maps The height and width of and Represents the feature maps Coordinates of; feature map The normalized frequencies in height and width are given by |u| and |v|. After moving the low frequency part to the center, u is taken from the set Take the value from Take the value in Step A2: Introduce an adaptive hole rate strategy to assign different hole rates to each pixel, expressed as: ; Make predictions through a convolutional layer with parameter θ; Step A3: A modulation mechanism is adopted and a ReLU layer is introduced to maximize the receptive field while minimizing the frequency information lost for each pixel. A local feature with a window size of s and a center of p is defined as , its receptive field and There is a positive correlation; Step A4: Calculate the high frequency power Measure the missing frequency information and optimize directly ,exist The lower position increases the void rate and increases the receptive field. A higher position suppresses the hole rate to reduce the loss of frequency information, which is expressed as: ; in, and denote the pixels with the highest and lowest high frequency power, respectively.

[0007] Specifically, the adaptive kernel operates the convolution kernel weights, specifically including the following steps: Step B1: Before introducing dynamic weighting to adjust the frequency response, the convolution kernel parameters are decomposed into low-frequency and high-frequency components. For a static convolution kernel, its weight This can be broken down as follows: ; in, Represents the average weight of each core , as a low-pass Mean filter, using The parameters defined are used for 1 × 1 convolution. represents the residual part; Step B2: After decomposition, the adaptive kernel dynamically adjusts the high-frequency and low-frequency components, expressed as: ; in, , is the dynamic weight of each channel, predicted by a global pooling + convolution layer; dynamically adjusted according to the input context ratio, the network focuses on a specific frequency band.

[0008] Specifically, the frequency selectively balances the frequency power of the input feature to expand the receptive field, and specifically comprises the following steps: Step C1: Frequency Selection The features are decomposed into different frequency bands by applying different masks in the Fourier domain, expressed as: ; in, represents inverse fast Fourier transform; It is a binary mask designed to extract the corresponding frequency, and its expression is: ; in, , is composed of B+1 predefined frequency thresholds obtained; Step C2: Perform spatial dynamic reweighting on the frequency components of different frequency bands, expressed as: ; in, is the frequency equalization feature after frequency selection learning, Indicates the selection mapping of the bth frequency band.

[0009] Specifically, the context and spatial feature calibration network CSFCN in step S3 adopts an asymmetric encoder-decoder architecture, introduces a context feature calibration module to construct a private context for each pixel to enhance its distinguishability, and uses a spatial feature calibration module to perform spatial feature calibration; the context feature calibration module specifically includes: Given features And Processing to capture highly abstract multi-scale context ; Calculate pixel context similarity ,in and denote the total number of pixels and the total number of contexts, respectively, and ; use As a guide, we aggregate context for each pixel and further adjust the response value of each semantic context to generate fine-grained context and achieve context recalibration. Context calibration is defined as: ; in, denote input, output, rescaling factor and context respectively, The value range is , Represents a pairwise function used to calculate the affinity between features.

[0010] Specifically, the spatial feature calibration module specifically includes: Feature resampling is used to reconstruct features, and the spatial coordinates of each position on the feature map are defined as , the learned 2D offset map is ; Calibration function It is expressed as: ; from The sample features at the output ; Adaptively fusing and calibrating semantic features via gating strategy and fine-grained features , bridging the calibration semantic features and fine-grained features The representation gap between: ; in, and represents a gate mask; The calibration and fusion processes are integrated into a single spatial feature calibration module.

[0011] Specifically, the single spatial feature calibration module calibration and fusion process specifically includes: Given features and , unify the channels to the same number C through two convolutional layers; Bilinear interpolation is used Upsample, and then and After splicing, the two images are input into a convolutional block to predict two sets of offset maps. and , aligning the features of the two levels, , Control the flow of feature information at two levels; The calibrated cross-level features are summed element by element to get the output. The spatial feature calibration module is expressed as: ; in, represents the bilinear upsampling function, and Represents a convolutional layer with BN and ReLU; right , Using 1+tanh activation, becomes an identity map and , the spatial feature calibration module is expressed as: .

[0012] Specifically, the improved wavelet pooling downsampling specifically includes: In the wavelet domain, the second-order decomposition and pooling features are performed according to the fast wavelet transform, which is expressed as: ; ; in is an approximate function, is the detail function, , They are called approximation coefficient and detail coefficient respectively; and are the scale vector and wavelet vector of time reversal, n represents the samples in the vector, and j represents the resolution; The image is processed by applying transformations to rows and columns respectively, and the detail subbands LH, HL and HH of each decomposition layer and the approximate subband LL of the highest decomposition layer are obtained; after performing the second-order decomposition, the image features are reconstructed using the second-order wavelet subbands, and the image features are pooled twice using the inverse FWT: .

[0013] Specifically, the improved wavelet pooling upsampling includes: reversing the downsampling process, performing a first-order wavelet decomposition on the upsampled feature map; upsampling the decomposed detail coefficient subband by a factor of 2 to form a new first-layer decomposition, the initial decomposition becomes a second-layer decomposition, the new second-order wavelet decomposition reconstructs the image features, and upsampling is performed using IDWT.

[0014] The beneficial effect of the present invention is that the present invention uses frequency-adaptive dilated convolution to improve the modules in the backbone network of the benchmark model, and introduces a context and spatial feature calibration network in the encoder part of the network to achieve context feature calibration and spatial feature calibration. Finally, since the Nyquist sampling theorem is usually ignored when downsampling using ordinary convolution operations, resulting in aliasing and performance degradation, the present method uses a wavelet pooling method to effectively suppress aliasing. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.

[0016] Figure 1 This is a structural diagram of RT-DETR in an embodiment of the present invention; Figure 2 RT-DETR flow chart in one embodiment of the present invention; Figure 3 This is a diagram of the FCW-RTDETR algorithm model architecture in one embodiment of the present invention; Figure 4 A schematic diagram of frequency adaptive dilated convolution in an embodiment of the present invention; Figure 5 A schematic diagram of a context space and spatial feature calibration network in an embodiment of the present invention; Figure 6 A schematic diagram of a context feature calibration module in an embodiment of the present invention; Figure 7 A schematic diagram of a spatial feature calibration module in an embodiment of the present invention; Figure 8 A schematic diagram of wavelet pooling downsampling in an embodiment of the present invention; Fig. 9 Schematic diagram of wavelet pooling upsampling in an embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0018] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0019] The following is combined with Figures 1 to 9 , some embodiments of the present invention are described in detail. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0020] The present invention proposes a method for detecting small targets in UAV aerial images based on RT-DETR. In a preferred embodiment, Figure 2 As shown, the specific steps include: Step S1: Use adaptive frequency dilated convolution to improve the backbone network and extract the features of drone aerial images; Step S2: Perform intra-scale feature interaction AIFI on the P5 layer in the backbone network, and input the obtained features into the cross-scale feature fusion network CCFM; Step S3: Use the context and spatial feature calibration network CSFCN to improve the CCFM structure, and use wavelet pooling to improve the up and down sampling methods in CCFM, fuse features of different scales in CCFM and obtain an image feature sequence; Step S4: Using IoU-aware query selection, a fixed number of features are selected from the feature sequence output by the encoder as the initial target query for the decoder; Step S5: The decoder iteratively optimizes the target query through the auxiliary prediction head, generates bounding boxes and confidence scores, and completes small target detection.

[0021] In this embodiment, the algorithm architecture of this solution is as follows Figure 1 As shown, the content is as follows: 1. Add adaptive frequency-dilated convolution to the backbone network ResNet18; 2. Add context and spatial feature calibration network to cross-scale feature fusion CCFM; 3. Use wavelet pooling to replace the ordinary pooling operation in the network.

[0022] First of all, as the most popular target detection method at present, the YOLO method has achieved a good balance between computing speed and accuracy. However, since the YOLO algorithm is a single-stage target detection method, its non-maximum suppression post-processing process will not only affect the computing speed, but also introduce hyperparameters, which will bring uncertainty to the model performance. Therefore, this method improves and innovates a Transformer-based end-to-end target detection algorithm RT-DETR (Real-Time Detection Transformer) and proposes the FCW-RTDETR algorithm. Its architecture is as follows: Figure 3 . Secondly, traditional convolution has problems such as loss of spatial hierarchical information and inability to reconstruct small object information, while dilated convolution increases the receptive field without reducing the resolution, reducing the information loss of up and down sampling. Therefore, this method uses adaptive frequency dilated convolution to improve the modules in the backbone network of the benchmark model. Next, in the efficient hybrid encoder part of the network, although the multi-level feature fusion method can effectively improve the performance of small target detection, it cannot specifically deal with the problems of pixel context mismatch and feature misalignment. Therefore, the present invention introduces a context and spatial feature calibration network (Context and Spatial Feature Calibration) in the encoder part of the network to achieve context feature calibration and spatial feature calibration. Finally, since the Nyquist sampling theorem is usually ignored when downsampling using ordinary convolution operations, resulting in aliasing and performance degradation, this method uses wavelet pooling to effectively suppress aliasing.

[0023] 1. Frequency Adaptive Dilated Convolution In this embodiment, frequency adaptive dilated convolution (FADC) is introduced to improve the BasicBlock in Backbone. Figure 4 As shown in Figure 2, FADC includes three key strategies, namely adaptive dilation rate (AdaDR), adaptive kernel (AdaKern) and frequency selection (FreqSelect), which aim to improve each stage of dilated convolution, such as Figure 4 As shown in Figure 2, AdaDR adjusts the dilation rate spatially, AdaKern operates on the convolution kernel weights, and FreqSelect directly balances the frequency power of the input features to expand the receptive field.

[0024] 1. Adaptive Dilation Rate The currently widely used hole convolution can be expressed as follows:

[0025] in, Represents the output feature map The pixel value of the position, represents the kernel size, represents the kernel weight parameter, Represents the input feature map with The pixel value at the corresponding position, the offset is .variable Represents the first positions (-1, -1), (-1, 0), (-1, +1), ..., (+1, +1). This can be achieved by increasing the void ratio To expand the receptive field.

[0026] Previous studies have found that increasing the hole rate will lead to a decrease in the ability to capture frequency information. Specifically, according to the scaling characteristics of the Fourier transform, increasing the hole rate from 1 to When , the magnification of the convolution kernel is Therefore, the response frequency of the convolution kernel is reduced to , causing the frequency response to shift from high frequencies to low frequencies. Convolution effectively This means that it works at a sampling rate of However, this also means that it cannot capture frequencies above the Nyquist frequency, which is half the sampling rate ( ). Therefore, the dilated convolution may not be able to adequately represent high-frequency information beyond this limit.

[0027] Specifically, this method first uses discrete Fourier transform (DFT) The feature map Transformed to the frequency domain, it can be expressed as:

[0028] in Array of complex numbers representing the DFT output. and Represents the feature maps The height and width of and Represents the feature maps The coordinates of . Feature map The normalized frequencies in height and width are given by |u| and |v|. After moving the low frequency part to the center, u is removed from the set Take the value from Therefore, the high frequency set greater than the Nyquist frequency It cannot be captured accurately, which limits its bandwidth.

[0029] Based on the above analysis, the choice of dilation rate can be regarded as a trade-off between large receptive field and effective bandwidth. Considering that the input feature map varies in space, the optimal dilation rate for each pixel can be different. Therefore, this method introduces an adaptive dilation rate (AdaDR) strategy to achieve a better balance. It assigns different dilation rates to each pixel:

[0030] It can be predicted by a convolutional layer with parameter θ. This method introduces a ReLU layer to ensure the non-negativity of the hole, and also adopts a modulation mechanism. It aims to maximize the receptive field while minimizing the frequency information lost for each pixel. For a local feature centered at p and with a window size of s, this method is called Its receptive field and A set The frequency in the image cannot be accurately captured. Therefore, the high frequency power can be calculated to measure the missing frequency information. Therefore, The optimization can be written as:

[0031] However, due to the frequency set Due to the discrete nature of and the non-differentiable nature of HP, direct optimization may be unrealistic. Therefore, this method chooses to directly optimize , that is, by The lower position increases the void rate to increase the receptive field. The higher position suppresses the hole rate to reduce the loss of frequency information. This method expresses it as:

[0032] here, and denote pixels with the highest / lowest (e.g., 25 %) high frequency power, respectively.

[0033] 2. Adaptive Kernel AdaDR achieves a delicate balance between effective bandwidth and receptive field by assigning a dilation rate to each pixel individually, optimizing both factors together. The effective bandwidth is closely related to the weight of the convolution kernel and plays a key role. Traditional convolution kernels learn features across different frequency bands, which are crucial for understanding complex visual patterns, but they become static once trained. To further enhance the effective bandwidth, this method decomposes the convolution kernel parameters into low-frequency and high-frequency components before introducing dynamic weighting to adjust the frequency response. This process adds only a small amount of additional parameters and computational overhead. For a static convolution kernel, its weight W can be decomposed as follows:

[0034] here, represents the core-by-core average It acts as a low-pass The mean filter is then used A 1 × 1 convolution is performed with the defined parameters. A higher mean value is more likely to attenuate high frequency components. Represents the residual part, captures local differences, and extracts high-frequency components. After decomposition, AdaKern dynamically adjusts the high-frequency and low-frequency components and can be formally expressed as:

[0035] in , is the dynamic weight of each channel, predicted by a simple and lightweight global pooling + convolution layer. Dynamically adjusted according to the input context The ratio of , enables the network to focus on specific frequency bands, adapting to the complexity of visual patterns in features. This dynamic frequency adaptation method enhances the network's ability to capture both low-frequency context and high-frequency local details. This in turn increases the effective bandwidth, thereby improving performance in segmentation tasks that require different feature extraction at different frequencies.

[0036] 3. Frequency Selection Traditional convolution usually acts as a high-pass filter. Therefore, the resulting features tend to show a higher proportion of high-frequency components. This tendency leads to a smaller overall dilation rate to maintain a high effective bandwidth, but unfortunately compromises the size of the receptive field. FreqSelect aims to enhance the receptive field by balancing the high-frequency and low-frequency components in the feature representation.

[0037] Specifically, FreqSelect first decomposes the features into different frequency bands by applying different masks in the Fourier domain:

[0038] in stands for Inverse Fast Fourier Transform. is a binary mask designed to extract the corresponding frequencies:

[0039] here, , is composed of B+1 predefined frequency thresholds Then, FreqSelect performs spatial dynamic reweighting on the frequency components of different frequency bands. It can be expressed as:

[0040] in is the frequency equalization characteristic after FreqSelect learning. represents the selection mapping of the bth frequency band. Specifically, the frequency is decomposed into 4 frequency bands in an octave manner, namely .

[0041] 2. Context Space and Spatial Feature Calibration Network The overall architecture of the context space and spatial feature calibration network used in the present invention is as follows: Figure 5 As shown, an asymmetric encoder-decoder architecture is adopted. The present invention uses a lightweight backbone network to extract multi-level features. After that, a context feature calibration module (Context Feature Calibration Module, CFC) is introduced to build a private context for each pixel to enhance its distinguishability. Furthermore, the method uses a spatial feature calibration module (Spatial Feature Calibration Module, SFC) for spatial feature calibration to produce strong semantic features with precise boundaries.

[0042] 1. Context Feature Calibration Module Contextual information can provide rich scene category priors to correct unexpected misclassifications. However, previous methods aggregate context within a predefined region, ignoring that not all context contributes equally to the classification of a given pixel, which inevitably leads to the problem of context mismatch. In addition, since large objects contain more pixels, the captured context is heavily biased towards large objects and leads to over-smoothing of small objects. Based on the above insights, our method adopts a Context Feature Calibration module to crop and refine the semantic context for each pixel.

[0043] Typically, large objects dominate images, and global context tends to be similar to these large objects rather than spatial details. Our method matches context to each pixel to capture the context that is most instructive for its classification. That is, our method aggregates context from regions with closer semantic distances rather than regions with closer spatial distances. Specifically, given a feature , first of all Processing to capture high-level, multi-scale context Then, the pixel context similarity is calculated , which is the spatial attention map, where and Represent the total number of pixels and the total number of contexts respectively. This method uses As a guide, contexts are aggregated for each pixel to achieve context feature calibration. Finally, our method further adjusts the response value of each semantic context to generate fine-grained contexts, thereby achieving context recalibration. Mathematically, CFC can be defined as:

[0044] In the formula denote input, output, rescaling factor and context respectively, The value range is , Represents a pairwise function used to calculate the affinity between features.

[0045] Figure 6 The CFC module is shown in . In order to pursue efficiency, this method uses the Cascaded Pyramid Pooling (CPP) module to reuse the pooling results of the previous layers and reduce unnecessary redundant calculations. Figure 6 As shown, given a feature , this method first uses a 1×1 convolutional layer to generate a dimensionality reduction feature , where C′ is much smaller than C (by default, C′ = 32, C = 256). Then, the CPP block is used to obtain multi-scale context In particular, when the pooling layer outputs height (Generally, the output width is equal to the height), Default settings . Since the pooling layer performs pooling on a homogeneous spatial grid (e.g., 3×3), this may lead to information redundancy in the pooled features when the aspect ratio of the input is not 1. To this end, our method keeps the aspect ratio of the pooled output size equal to the input. For example, when training on square crops, our method's minimum output size is 1×1, while for a 1024×2048 input, our method's minimum output size is 1×2. Finally, our method feeds Z into two convolutional layers (with BN and ReLU) to produce two forms of contextual representations, namely, and .

[0046] 2. Spatial Feature Calibration Module To compensate for the loss of spatial details caused by downsampling, previous methods use cross-level feature fusion to enhance high-level semantic features with low-level details. and high-resolution features , firstly use standard bilinear interpolation to Upsample, and then and after upsampling The output is obtained by adding or concatenating them. However, due to the spatial misalignment and the huge representation gap, directly fusing them still cannot achieve satisfactory performance. In addition, since the feature maps encode various semantic information, performing unified feature alignment along the channel dimension will also hurt the performance. To alleviate these problems, this method groups the channel dimension into multiple sub-features to perform calibration operations separately, and seamlessly integrates the gating mechanism to adaptively fuse cross-layer features.

[0047] Regarding feature calibration, this method proposes to reconstruct features by using feature resampling. Specifically, assuming that the spatial coordinates of each position on the feature map are , the learned 2D offset map is The calibration function of this method is It can be expressed as: from The sample features at the output Since p represents an arbitrary position, the above formula enumerates all integral positions and uses a bilinear interpolation kernel to obtain the characteristics of the sampling position. Intuitively, Figure 5 As shown in the SFC sampling process, this method only needs to sample the most favorable feature (blue point) at the current position to replace the feature at the current position (green point) for calibration. In addition, in order to perform more precise calibration, this method divides the feature F into G groups along the channel dimension, and then aligns the features in each group separately, as shown in Figure 7 shown.

[0048] because 'usually contains rich spatial details, while 'contains more semantic information, and simply calibrating and fusing them cannot achieve satisfactory results. This is because feature calibration alone cannot handle the huge representation differences between features. Therefore, this method further proposes to adaptively fuse and calibrate semantic features through a gating strategy and fine-grained features , to bridge the representation gap between them:

[0049] In the formula and Represents a gate mask.

[0050] To improve efficiency, this method integrates the calibration and fusion processes into a single spatial feature calibration module. Figure 7 As shown, given the feature and , this method first unifies their channels to the same number C (128 by default) through two convolutional layers. Secondly, this method uses bilinear interpolation to Then, the upsampled and Then, it is input into a convolutional block to predict two sets of offset maps. and , used to align two levels of features, two gate masks , , which is used to control the flow of feature information at two levels. Finally, the calibrated cross-level features are summed element by element to obtain the output. In summary, the SFC module can be formally written as:

[0051] in represents the bilinear upsampling function, and Represents a convolutional layer with BN and ReLU.

[0052] It is worth noting that our method adopts the idea of ​​residual to alleviate the negative impact of large offset and mask prediction errors in the initial stage, which allows our method to insert SFC into the network without compromising its original performance (if the convolutional block is constructed as a zero map). That is, our method initializes the weights of the last convolutional layer of the convolutional block to zero to gradually learn more accurate offsets and masks. In addition, our method chooses to use 1+tanh activation for the gate mask. Therefore, at the beginning becomes an identity map and . Now, SFC can be expressed as: .

[0053] 3. Wavelet Pooling Compared with general object detection, the effect of aliasing in tiny object detection is more significant. This is because tiny objects themselves have more high-frequency features, and their pixel values ​​change rapidly in a very small range, making them more sensitive to aliasing effects. Unfortunately, the convolutional neural network that the current backbone network for tiny target detection relies on often ignores the Nyquist sampling theorem when performing downsampling operations, which directly leads to serious aliasing problems, thus affecting the accuracy and reliability of detection. In order to suppress the aliasing effect, this method replaces the up- and down-sampling methods in the backbone network and the neck network with wavelet pooling methods.

[0054] 1. Downsampling The adopted wavelet pooling method pools features by performing a 2nd-order decomposition in the wavelet domain according to the Fast Wavelet Transform (FWT), which is a more efficient implementation of the 2D Discrete Wavelet Transform (DWT):

[0055]

[0056] in is the approximate function, ψ is the detail function, , They are called the approximation coefficient and detail coefficient respectively. and are the time-reversed scale vector and wavelet vector, n represents the samples in the vector, and j represents the resolution. When the FWT is used to process the image, this method applies it twice (once on the rows and then again on the columns). Through this combination, the method obtains the detail subbands (LH, HL, HH) of each decomposition layer and the approximate subband (LL) of the highest decomposition layer. After performing the second-order decomposition, the method reconstructs the image features, but only uses the second-order wavelet subbands. Based on the inverse DWT (IDWT), this method uses the inverse FWT (IFWT) to pool the image features by 2 times:

[0057] Figure 8 A schematic diagram of the wavelet pooling downsampling algorithm is given.

[0058] 2. Upsampling The introduced wavelet pooling algorithm performs upsampling by reversing the downsampling process. First, the upsampled feature map is subjected to a 1st-order wavelet decomposition. The decomposed detail coefficient subbands are upsampled by a factor of 2 to form a new 1st-level decomposition. The initial decomposition becomes a 2nd-level decomposition. Finally, this new 2nd-order wavelet decomposition reconstructs the image features for further upsampling using IDWT. Fig. 9 The upsampling algorithm of wavelet pooling is introduced.

[0059] In a specific embodiment data, the dataset used is the VisDrone2019 dataset. This method is compared with a representative target detection algorithm on this dataset to verify the effectiveness of the model proposed by this method.

[0060] The dataset was collected by the AISKYEYE team of the Machine Learning and Data Mining Laboratory of Tianjin University. The benchmark dataset contains 288 video clips, consisting of 261,908 frames of video and 10,209 static images. Among the image data used in this method, 6471 are used for training, 548 are used for verification, and 3190 are used for testing. These data come from various types of drone cameras and cover a wide range, including different cities (from 14 different cities in China, thousands of kilometers apart), different environments (urban and rural), different objects (pedestrians, vehicles, bicycles, etc.) and different densities (sparse and crowded scenes). It should be noted that the dataset was collected using different models of drone platforms in different scenes, weather and lighting conditions.

[0061] The biggest difference between the VisDrone2019 dataset and other datasets is that the object scale is relatively small. Table 1 summarizes the target scale results of the dataset. According to the standard of the MS COCO dataset, targets with a resolution less than 32×32 are small targets. As can be seen from the table, 59.6% of the targets in the VisDrone2019 dataset are small targets, which is undoubtedly the biggest difficulty and challenge of this dataset.

[0062] Table 1. VisDrone2019 target scale statistics

[0063] Visual analysis: In complex environments, small targets are affected by various factors, such as different environments, uneven lighting, different heights of drone aerial photography, and occlusion caused by overlapping objects. This method evaluates the performance of the FCW-RTDETR model in various scenarios. The test results show that this scheme can accurately identify small targets from drone aerial images and has good robustness.

[0064] The comparison of the effects of aerial photography small target detection under different conditions shows that the algorithm proposed in this method shows good adaptability in various environments, and can effectively detect small targets in scenes such as parking lots and parks, and the detection results have high accuracy and stability. The comparison of detection effects under different light intensities shows that the change of light has a certain impact on the detection results, but overall, the algorithm shows strong robustness to light changes. The analysis of the impact of aerial photography altitude on the detection effect shows that the algorithm can still detect the target well at different altitudes. The detection results under dense occlusion show that the algorithm can detect most of the occluded targets, but a small number of targets are still not correctly identified, indicating that there is still room for further optimization when dealing with severe occlusion problems.

[0065] In one embodiment, the FCW-RTDETR and RT-DETR-R18 models and the YOLOv8m model are compared for small target detection using drone aerial images under various environmental conditions; FCW-RTDETR outperforms the baseline model RT-DETR-R18 and the current mainstream target detection model YOLOv8m in small target detection performance. FCW-RTDETR fully detects pedestrians with very small resolution in the original image, while RT-DETR-R18 has some missed detection of pedestrians, and YOLOv8m almost does not detect pedestrians, indicating that FCW-RTDETR has a high sensitivity to small targets. RT-DETR-R18 misdetects the car at the top of the original image as a van, and the trash can at the bottom left as people, and like YOLOv8m, it misses the motor at the top and right. While FCW-RTDETR has a low missed detection rate, it also outperforms other models in detection accuracy. Overall, FCW-RTDETR shows excellent detection performance for small targets in various complex aerial images, making it suitable for high-altitude UAV detection tasks. For the aforementioned embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily required by the present application.

[0066] The above embodiments describe the basic principles and main features of the present invention and the advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the changes and modifications made by those skilled in the art shall be within the scope of protection of the appended claims of the present invention without departing from the spirit and scope of the present invention.

Claims

1. A method for detecting small targets in UAV aerial images based on RT-DETR, characterized in that: The end-to-end target detection algorithm RT-DETR is improved and the FCW-RTDETR algorithm is proposed, which includes the following steps: Step S1: Use adaptive frequency dilated convolution to improve the backbone network and extract the features of drone aerial images; Step S2: Perform intra-scale feature interaction AIFI on the P5 layer in the backbone network, and input the obtained features into the cross-scale feature fusion network CCFM; Step S3: Use the context and spatial feature calibration network CSFCN to improve the CCFM structure, and use wavelet pooling to improve the up and down sampling methods in CCFM, fuse features of different scales in CCFM and obtain an image feature sequence; Step S4: Using IoU-aware query selection, a fixed number of features are selected from the feature sequence output by the encoder as the initial target query for the decoder; Step S5: The decoder iteratively optimizes the target query through the auxiliary prediction head, generates bounding boxes and confidence scores, and completes small target detection.

2. A method for detecting small targets in unmanned aerial images based on RT-DETR as claimed in claim 1, characterized in that: The improvement of the adaptive frequency dilated convolution in step S1 includes adaptive dilated rate, adaptive kernel and frequency selection. The adaptive dilated rate adjusts the dilated rate in space, specifically including the following steps: Step A1: Using Discrete Fourier Transform The feature map Transformed to the frequency domain, it is expressed as: ; in, An array of complex numbers representing the DFT output; and Represents the feature maps The height and width of and Represents the feature maps Coordinates of; feature map The normalized frequencies in height and width are given by |u| and |v|. After moving the low frequency part to the center, u is taken from the set Take the value from Take the value in Step A2: Introduce an adaptive hole rate strategy to assign different hole rates to each pixel, expressed as: ; Make predictions through a convolutional layer with parameter θ; Step A3: A modulation mechanism is adopted and a ReLU layer is introduced to maximize the receptive field while minimizing the frequency information lost for each pixel. A local feature with a window size of s and a center of p is defined as , its receptive field and There is a positive correlation; Step A4: Calculate the high frequency power Measure the missing frequency information and optimize directly ,exist The lower position increases the void rate and increases the receptive field. A higher position suppresses the hole rate to reduce the loss of frequency information, which is expressed as: ; in, and denote the pixels with the highest and lowest high frequency power, respectively.

3. A method for detecting small targets in unmanned aerial images based on RT-DETR as claimed in claim 2, characterized in that: The adaptive kernel operates the convolution kernel weights, specifically including the following steps: Step B1: Before introducing dynamic weighting to adjust the frequency response, the convolution kernel parameters are decomposed into low-frequency and high-frequency components. For a static convolution kernel, its weight This can be broken down as follows: ; in, Represents the average weight of each core , as a low-pass Mean filter, using The parameters defined are used for 1 × 1 convolution. represents the residual part; Step B2: After decomposition, the adaptive kernel dynamically adjusts the high-frequency and low-frequency components, expressed as: ; in, , is the dynamic weight of each channel, predicted by a global pooling + convolution layer; dynamically adjusted according to the input context ratio, the network focuses on a specific frequency band.

4. The method for detecting small targets in unmanned aerial images based on RT-DETR as claimed in claim 2, characterized in that: The frequency selective balanced input feature frequency power expansion receptive field specifically comprises the following steps: Step C1: Frequency Selection The features are decomposed into different frequency bands by applying different masks in the Fourier domain, expressed as: ; in, represents inverse fast Fourier transform; It is a binary mask designed to extract the corresponding frequency, and its expression is: ; in, , is composed of B+1 predefined frequency thresholds obtained; Step C2: Perform spatial dynamic reweighting on the frequency components of different frequency bands, expressed as: ; in, is the frequency equalization feature after frequency selection learning, Indicates the selection mapping of the bth frequency band.

5. The method for detecting small targets in unmanned aerial images based on RT-DETR as claimed in claim 1, characterized in that: The step S3 context and spatial feature calibration network CSFCN adopts an asymmetric encoder-decoder architecture, introduces a context feature calibration module to construct a private context for each pixel to enhance its distinguishability, and uses a spatial feature calibration module to perform spatial feature calibration; The context feature calibration module specifically includes: Given features And Processing to capture highly abstract multi-scale context ; Calculate pixel context similarity ,in and denote the total number of pixels and the total number of contexts, respectively, and ; use As a guide, we aggregate context for each pixel and further adjust the response value of each semantic context to generate fine-grained context and achieve context recalibration. Context calibration is defined as: ; in, denote input, output, rescaling factor and context respectively, The value range is , Represents a pairwise function used to calculate the affinity between features.

6. A method for detecting small targets in unmanned aerial images based on RT-DETR as claimed in claim 5, characterized in that: The spatial feature calibration module specifically includes: Feature resampling is used to reconstruct features, and the spatial coordinates of each position on the feature map are defined as , the learned 2D offset map is ; Calibration function It is expressed as: ; from The sample features at the output ; Adaptively fusing and calibrating semantic features via gating strategy and fine-grained features , bridging the calibration semantic features and fine-grained features The representation gap between: ; in, and represents a gate mask; The calibration and fusion processes are integrated into a single spatial feature calibration module.

7. The method for detecting small targets in unmanned aerial images based on RT-DETR as claimed in claim 6, characterized in that: The single spatial feature calibration module calibration and fusion process specifically includes: Given features and , unify the channels to the same number C through two convolutional layers; Bilinear interpolation is used Upsample, and then and After splicing, the two images are input into a convolutional block to predict two sets of offset maps. and , aligning the features of the two levels, , Control the flow of feature information at two levels; The calibrated cross-level features are summed element by element to get the output. The spatial feature calibration module is expressed as: ; in, represents the bilinear upsampling function, and Represents a convolutional layer with BN and ReLU; right , Using 1+tanh activation, becomes an identity map and , the spatial feature calibration module is expressed as: 。 8. The method for detecting small targets in unmanned aerial images based on RT-DETR as claimed in claim 1, characterized in that: The improved wavelet pooling downsampling specifically includes: In the wavelet domain, the second-order decomposition and pooling features are performed according to the fast wavelet transform, which is expressed as: ; ; in is an approximate function, is the detail function, , They are called approximation coefficient and detail coefficient respectively; and are the scale vector and wavelet vector of time reversal, n represents the samples in the vector, and j represents the resolution; The image is processed by applying transformations to rows and columns respectively, and the detail subbands LH, HL and HH of each decomposition layer and the approximate subband LL of the highest decomposition layer are obtained; after performing the second-order decomposition, the image features are reconstructed using the second-order wavelet subbands, and the image features are pooled twice using the inverse FWT: 。 9. The method for detecting small targets in unmanned aerial images based on RT-DETR as claimed in claim 8, characterized in that: The improved wavelet pooling upsampling specifically includes: reversing the downsampling process, performing a first-order wavelet decomposition on the upsampled feature map; upsampling the decomposed detail coefficient subband by a factor of 2 to form a new first-layer decomposition, the initial decomposition becomes a second-layer decomposition, the new second-order wavelet decomposition reconstructs image features, and upsampling is performed using IDWT.

Citation Information

Patent Citations

  • Unmanned aerial vehicle aerial photography small target detection method based on improved YOLOv8s algorithm and electronic equipment

    CN118230194A

  • Unmanned aerial vehicle aerial photography small target detection method based on improved RT-DETR network

    CN118521929A

Cited By

  • VHPC-DETR-based violent target detection method

    CN120411736A

  • Unmanned aerial vehicle image small target detection method fusing space-frequency features

    CN120894537A

  • Vehicle defect detection method based on fusion frequency adaptive expansion convolution

    CN121074005A

  • Unmanned aerial vehicle aerial target detection method based on improved RT-DETR

    CN121095814A

  • Image feature enhancement model target detection method based on RTDETR network

    CN121582559A