Small target unmanned aerial vehicle detection method and system for decoupling spatial-temporal feature fusion
Through the spatiotemporal feature fusion method optimized by differential resolution processing and self-attention mechanism, the problems of limited feature extraction capability and poor robustness in drone detection are solved, and high-precision real-time detection of small targets at long distances is achieved, which is suitable for complex environments.
Patent Information
- Application Number
- CN202510685318.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-09
AI Technical Summary
Existing drone detection technology has problems in long-distance small target scenarios, such as limited feature extraction capabilities, poor robustness, and difficulty in balancing real-time and accuracy. In particular, the false detection rate is high in complex environments, the lightweight model is not accurate enough, and the high-precision model has high computational complexity, making it difficult to meet real-time requirements.
Differentiated resolution processing is used to decompose the video stream into high-definition key frames and low-resolution frame sequences, which are input into two-dimensional and three-dimensional convolution branches respectively. Combined with two-dimensional discrete Haar wavelet transform and 3D backbone network, through the spatiotemporal feature decoupling and fusion module, the self-attention mechanism is used to optimize feature fusion, achieving efficient spatiotemporal feature extraction and classification positioning.
It significantly improves the accuracy and robustness of small target drone detection, reduces the false detection rate, takes into account both real-time and high precision, adapts to complex dynamic scenes, optimizes the feature fusion strategy, and improves the adaptability and detection performance of the model.
Smart Images

Figure CN120612518A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of small target UAV detection. It is a small target UAV detection method and system that decouples spatiotemporal feature fusion and can be used to solve the problem of real-time high-precision detection of long-range, small-sized UAV targets. Background Art
[0002] With the rapid development of drone technology, it has been widely used in fields such as environmental monitoring, agricultural plant protection, and emergency rescue. Accordingly, real-time, high-precision drone detection methods are needed to detect and identify drone targets for drone management. However, drone target detection faces many challenges: First, in long-range monitoring scenarios, the imaging area of drone targets typically accounts for less than 0.1% of the image area, resulting in a severe lack of feature information, making accurate identification difficult. This places higher demands on the feature extraction capabilities of detection methods. Second, dynamic backgrounds, lighting changes, and interference from similar objects in complex environments significantly increase the probability of false and missed detections of drone targets, necessitating improvements in the robustness of detection methods. Third, in applications with high real-time requirements, such as security monitoring, detection processing speeds of at least 30 frames per second are typically required. Therefore, detection methods that balance real-time performance and accuracy are crucial for practical applications.
[0003] Among existing drone target detection methods, image object detection (IOD) based on still images mainly includes two-stage methods, single-stage methods, and anchor-free methods. Two-stage methods offer high detection accuracy but slow processing speed. Single-stage methods support high-resolution image input and have fast detection speed, but limited detection accuracy. Anchor-free methods eliminate the reliance on traditional anchors, achieving both high detection efficiency and accuracy. However, some models are complex. For example, high-precision models such as YOLOv7x have up to 70M parameters, resulting in generally low inference speed. Lowering the model complexity, such as the lightweight MobileNet-YOLO model, can achieve real-time detection at over 30 FPS, but its mAP@0.5 on a drone small target test set is only 46%, indicating insufficient accuracy. Because drone targets are typically characterized by high speed, small size, and motion blur, these IOD methods suffer from difficulties in feature extraction and poor robustness. Furthermore, their lack of utilization of temporal information limits their detection accuracy.
[0004] Video object detection (VOD) methods analyze video streams to identify and locate moving objects, addressing the shortcomings of static image object detection methods in their lack of utilization of temporal information. Some VOD methods employ a dual-branch network architecture, with a two-dimensional branch focusing on spatial feature extraction and a three-dimensional branch on temporal feature extraction. However, processing high-resolution frame sequences with the three-dimensional convolutional branch suffers from high resource overhead and low computational efficiency. Reducing the input video resolution so that both branches share the same low-resolution frame sequence limits the spatial feature extraction capability of the two-dimensional convolutional branch. Furthermore, spatiotemporal feature fusion (SFF), a key approach in VOD design, often employs methods such as channel concatenation, convolution transformation, and element-wise summation. While these methods can achieve certain improvements in accuracy and speed, they suffer from high computational complexity and memory overhead, making it difficult to achieve a balance between real-time performance and accuracy. For example, channel splicing will lead to high feature dimension, thereby increasing computational complexity and memory consumption; the fixed receptive field of traditional convolution transformation makes it difficult to capture a wide range of spatial context information; although element addition is simple and intuitive, when there are significant differences in the data distribution of different features, the feature fusion efficiency will be low and the complementarity will be poor, resulting in information loss or redundancy.
[0005] In summary, existing drone detection technologies face the following challenges when applied to long-range, small-target scenarios: First, image-based detection (IOD) methods rely on single-frame data for feature extraction and lack the ability to model the temporal correlation between multiple frames, limiting their feature extraction capabilities. Second, existing detection methods easily confuse drone targets with background interference objects such as clouds and birds, and their false detection rates significantly increase under conditions such as dynamic blur and partial occlusion. They are not adaptable to complex environments, and their robustness needs to be improved. Third, existing algorithms struggle to balance real-time performance, accuracy, and complexity. While lightweight models can achieve real-time detection, their small-target detection accuracy is typically low, making them impractical. High-precision models, while capable of achieving high detection performance, often suffer from low processing rates due to their excessive complexity, making them unable to meet dynamic, real-time requirements. These issues hinder the development of real-time small-target drone detection technology, and innovative solutions are urgently needed. Summary of the Invention
[0006] In response to the problems existing in the existing technology, the present invention provides a small target UAV detection method with decoupled spatiotemporal feature fusion, aiming to achieve real-time and high-precision detection of small target UAVs in complex scenes by enhancing the feature extraction capability of the model.
[0007] The present invention is a small-target drone detection method that decouples spatiotemporal feature fusion and is applicable to a variety of scenarios that require high-precision drone target detection. In the field of security monitoring, the present invention can be used for illegal intrusion detection in areas such as airports, borders, and important facilities. By monitoring the video stream in real time, it can accurately identify and track long-range, small-sized drone targets, effectively prevent the intrusion of illegal drones, and ensure the safety of key areas. In terms of traffic management, the present invention can monitor and manage drone activities in the airspace, identify and track drones in real time, and maintain airspace safety. The small-target drone detection method that decouples spatiotemporal feature fusion includes: first, decomposing the video stream into high-definition key frames and low-resolution frame sequences through differential resolution processing, and inputting them into two-dimensional convolution branches and three-dimensional convolution branches respectively. The two-dimensional convolution branch uses a 2D backbone network combined with a two-dimensional discrete Haar wavelet transform to achieve high- and low-frequency decomposition and decoupling of spatial features. The three-dimensional convolution branch uses a 3D backbone network to extract multi-scale temporal features and fuses them using temporal compression (TC), transposed convolution (TConv), and element-wise addition fusion (EA). The features output by both branches undergo parallel point-by-point convolution for dimensionality reduction. Scaled dot-product self-attention (SDPA) is then used to calculate channel correlation weights. Ultimately, the decoupled detection head completes the classification and localization of drone targets in the video, outputting the drone target classification results.
[0008] Furthermore, the specific steps of the small target drone detection method of decoupling spatiotemporal feature fusion are as follows:
[0009] Step 1: Perform differential resolution processing on the input video stream to obtain high-resolution key frames and low-resolution frame sequences;
[0010] Step 2: In the two-dimensional convolution branch, a 2D backbone network is used to extract multi-scale spatial features of high-resolution keyframes, and a feature decoupling module based on wavelet transform is used to decompose the high- and low-frequency information of the spatial features to obtain spatially decoupled classification features and spatially decoupled regression features suitable for classification and regression.
[0011] Step 3: In the 3D convolution branch, the 3D backbone network is used to extract the spatiotemporal features of the low-resolution frame sequence, and the deconvolution operation is used to upsample the spatiotemporal features, construct a feature pyramid structure, and output multi-scale spatiotemporal features;
[0012] Step 4: Perform spatiotemporal feature fusion based on the channel encoder on the spatially decoupled classification features and spatially decoupled regression features output by the two-dimensional convolution branch, as well as the multi-scale spatiotemporal features output by the three-dimensional convolution branch. The channel encoder uses a parallel point-by-point convolution structure to perform parallel dimensionality reduction and fusion processing on the features output by the two-dimensional convolution branch and the three-dimensional convolution branch. At the same time, a scaled dot product self-attention mechanism is used to calculate the correlation weights of different channel feature vectors.
[0013] In step 5, the fused features are input into the decoupling detection head to classify and locate the drone targets in the video, and the detection results are output, including the target category, confidence level, and bounding box coordinates.
[0014] The design idea of this invention is to comprehensively utilize the spatial and temporal information of the video by adopting differentiated resolution processing, parallel convolution branches and a feature fusion module based on an improved channel self-attention mechanism, enhance the feature extraction capability of the model, and achieve high-precision real-time detection of small target drones.
[0015] Compared with the prior art, the present invention has the following advantages:
[0016] 1. Efficient extraction of spatiotemporal features: Traditional drone target detection methods primarily rely on the spatial information of a single frame, ignoring the temporal information inherent in video frame sequences. This paper utilizes differentiated resolution preprocessing, feeding high-resolution key frames and low-resolution frame sequences into the 2D and 3D convolution branches, respectively, to fully utilize temporal information. This not only significantly improves detection accuracy but also enhances the system's adaptability to complex dynamic scenes.
[0017] 2. Robustness in Complex Environments: Video frame sequence detection methods that rely solely on spatial features, while leveraging temporal motion information, suffer from high false positive rates in complex environments with dynamic backgrounds (such as swaying leaves or moving clouds) or changing lighting, and are unable to identify hovering drones. This invention, through feature decoupling and multi-scaling techniques, effectively utilizes temporal information to filter out background interference, avoiding the degradation of detection performance caused by complex backgrounds or dynamic scenes, significantly reducing false positives and missed detections, and thus achieving superior detection performance.
[0018] 3. Adaptive Feature Fusion Optimization: Existing methods often use fixed strategies for feature fusion, which can lead to high computational complexity, difficulty ensuring real-time performance, and poor adaptability to different scenarios. This paper introduces a scaled dot-product self-attention mechanism to achieve dimensionality reduction and fusion of input features within the channel encoder. This adaptively adjusts the feature fusion strategy based on the complexity of the detection scenario, enhancing the model's spatiotemporal modeling capabilities and improving detection performance.
[0019] 4. Balancing real-time performance and accuracy: Existing high-precision models struggle to meet real-time requirements due to their high computational complexity. While lightweight models can achieve real-time detection, their accuracy is insufficient. This invention optimizes feature extraction and fusion to ensure high detection accuracy while supporting real-time reasoning, effectively resolving the difficulty of existing methods in balancing efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of a small target UAV detection method by decoupling spatiotemporal feature fusion provided by the present invention;
[0021] Figure 2 This is a structural diagram of a small target UAV detection system model with decoupled spatiotemporal feature fusion provided by the present invention;
[0022] Figure 3 This is a structural diagram of the two-dimensional convolution branch module in the present invention;
[0023] Figure 4 This is the structure diagram of the MSFI (temporal compression TC+deconvolution TConv+element addition fusion EA) model in the three-dimensional convolution branch module of the present invention;
[0024] Figure 5 This is the PR curve diagram of the present invention and other mainstream VOD and IOD models for drone category detection targets;
[0025] Figure 6 It is a PR curve diagram of the present invention and other mainstream VOD and IOD models for bird detection targets. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0027] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] Example 1
[0029] In recent years, drones have been widely used in numerous fields due to their convenience. However, their misuse also poses a threat to critical infrastructure and flight safety. Therefore, effective drone detection and identification is essential. With the rapid development of computer vision technology, deep learning-based object detection methods have demonstrated significant advantages in drone detection. To fully utilize spatiotemporal information, some video object detection methods employ a dual-branch structure for feature extraction. The two-dimensional convolution branch, built on a 2D convolutional neural network (2D CNN), focuses on extracting spatial features by processing a single frame to capture static information such as an object's shape and texture. The three-dimensional convolution branch, on the other hand, focuses on extracting temporal features by analyzing the differences between consecutive frames to capture the object's motion. The 3D convolutional neural network (3D CNN) is a commonly used method for spatiotemporal feature extraction. It utilizes a three-dimensional convolution kernel to perform convolution operations simultaneously in the spatial and temporal dimensions, capable of extracting spatiotemporal features from video frame sequences. It has been widely used in video analysis and behavior recognition. However, the three-dimensional convolutional branches built on 3D CNNs suffer from high resource overhead and low computational efficiency when processing high-definition video frame sequences. Although some methods have adopted lightweight branch structures, such as fast and slow branches, to improve efficiency, these approaches share the same low-resolution input between the two branches, limiting the spatial feature extraction capabilities of the two-dimensional convolutional branches based on 2D CNNs and thus the overall performance of the model. The features extracted by the two branches are typically fused in subsequent stages to accurately locate and classify objects in the video.
[0030] In summary, this paper designs a decoupled spatio-temporal feature fusion detection network (DS2F-DN) for high-precision real-time detection of small-target drones. The network uses two-dimensional convolution branches and three-dimensional convolution branches to process high-resolution key frames and low-resolution frame sequences in parallel, and comprehensively utilizes spatial and temporal information. Among them, in order to adapt to the semantic differences between classification and regression tasks, we adopt a feature decoupling module based on wavelet transform decomposition. The decoupled features contain the semantic information required for classification and regression tasks, which can improve the recognition ability of small targets. In addition, for the multi-scale spatio-temporal features extracted by the two-dimensional convolution branch and the three-dimensional convolution branch, a feature fusion module based on the improved channel self-attention mechanism is adopted, which can realize the efficient fusion of spatio-temporal features and cross-dimensional information interaction, and enhance the spatio-temporal modeling ability of the model. See Figure 1 , the present invention includes the following steps:
[0031] (1) Perform differential resolution processing on the input video stream to obtain high-resolution key frames and low-resolution frame sequences;
[0032] In step (1), see Figure 2 , including the following steps to achieve differential resolution processing:
[0033] (1a) Read the input video stream and capture t consecutive video frames to form a frame sequence V = {I1, I2, ... I t}, and perform data enhancement preprocessing on the frame sequence V, where the input video stream includes multiple video frames.
[0034] (1b) Adjust the resolution of the frame sequence V to obtain a low-resolution frame sequence V′, and extract a specific single frame from a fixed position of V and adjust the resolution of the specific single frame to obtain a high-resolution key frame I key , the key frame I key The resolution of the low-resolution frame sequence V′ is twice the resolution of the low-resolution frame sequence V′.
[0035] (1c) Take V′ as the input of the 3D convolution branch and I key As the input of the 2D convolution branch.
[0036] (2) In the two-dimensional convolution branch, a 2D backbone network is used to extract multi-scale spatial features of high-resolution key frames, and a feature decoupling module based on wavelet transform is used to decompose the high- and low-frequency information of spatial features to obtain spatially decoupled classification features and spatially decoupled regression features for classification and regression;
[0037] (3) In the 3D convolution branch, the 3D backbone network is used to extract the spatiotemporal features of the low-resolution frame sequence, and the deconvolution operation is used to upsample the spatiotemporal features, construct a feature pyramid structure, and output multi-scale spatiotemporal features;
[0038] (4) The spatial decoupled classification features and spatial decoupled regression features output by the two-dimensional convolution branch, as well as the multi-scale spatiotemporal features output by the three-dimensional convolution branch, are fused using a channel encoder. The channel encoder uses a parallel point-by-point convolution structure to perform parallel dimensionality reduction and fusion processing on the features output by the two-dimensional convolution branch and the three-dimensional convolution branch. At the same time, a scaled dot product self-attention mechanism is used to calculate the correlation weights of different channel feature vectors:
[0039] (4a) The spatial features and multi-scale spatiotemporal compression features output by the 2D backbone network are first processed by parallel point-wise convolution (PWConv), and then spliced in the channel dimension to form a new feature tensor.
[0040] (4b) Perform Same Convolution (SConv) on the concatenated feature tensor, and then vectorize the output of the Same Convolution in the spatial dimension to obtain a matrix containing the feature vectors of all channels.
[0041] (4c) Using scaled dot product attention score function Calculate the attention score between the feature vectors to reflect the similarity between the two channels, while effectively avoiding the numerical overflow problem caused by the dot product result being too large, and preventing the subsequent Softmax gradient from disappearing or exploding.
[0042] (4d) Perform Softmax processing on the attention score to obtain the attention weight matrix and the output of the scaled dot product channel self-attention module based on the Gram matrix
[0043] (5) The fused features are input into the decoupled detection head, which classifies and locates the drone target in the video and outputs the detection results. The decoupled detection head is a structural design that separates the target classification and position regression tasks. Traditional detection heads usually share some network layers for these two tasks, while the decoupled detection head processes the classification and regression tasks separately through two independent sub-network layers.
[0044] Example 2
[0045] A small target drone detection method with decoupled spatiotemporal feature fusion is the same as in Example 1. In step (2), a 2D backbone network is used to extract multi-scale spatial features of high-resolution key frames. Figure 3 , including the following steps:
[0046] (2a) The spatial feature map of the 2D backbone network output in the 2D convolution branch is obtained by using 2D discrete Haar wavelet transform Perform fine decomposition of spatial dimensions, where i∈{2,3,4} indicates small, medium, and large scales, c and α indicate the number of channels and the downsampling multiple, and h and w indicate the height and width of the channel, respectively. Specifically, a depthwise convolution module (DWConv) with a stride of 2 is used to process the , the convolution kernel of DWConv is composed of a low-pass filter and a set of high-pass filters and composition, Each channel of is processed by the above four filters to generate 4 new outputs. The output groups of the low-pass filter and the high-pass filter are grouped according to the filter category. Specifically, The channel that performs deep convolution operation with a low-pass filter (convolution kernel) obtains a low-frequency feature map, and the channel that performs deep convolution operation with a high-pass filter obtains a high-frequency feature map. Finally, a low-frequency feature map can be obtained. and 3 high-frequency feature maps and Represent the approximate, horizontal edge, vertical edge and diagonal features of the image respectively;
[0047] (2b) Low-frequency feature map X LL After being processed by n convolution-batch normalization-activation function (CBL) modules of the classification branch, the spatial decoupled classification feature map is obtained. Each CBL module consists of a convolution module, a batch normalization module, and an activation module;
[0048] (2c) The three high-frequency feature maps X LH , X HL and X HH Stacking along the channel dimension, we get Stack dim=c (·) indicates that the feature maps are stacked in the channel dimension, and then the linear projection operator LP(·) is used to stack the feature maps. dim=c (X LH ,X HL ,X HH ) Project from 3c dimension to c dimension to obtain the high-frequency feature map after dimensionality reduction
[0049] (2d) X HF After being processed by n CBL modules of the regression branch, a spatial decoupled regression feature map is generated
[0050] Example 3
[0051] A small target drone detection method with decoupled spatiotemporal feature fusion is the same as in Example 1. In step (3), a 3D backbone network is used to extract the spatiotemporal features of a low-resolution frame sequence. Figure 4 ,The MSFI model in the 3D backbone network includes the following steps:
[0052] (3a) Use the 3D backbone network to extract features from V′ and output multi-scale features F stage 2. F stage 3 and F stage 4. F stage 2. F stage 3 and F stage4. Use TC for temporal compression to achieve dimensionality reduction. The implementation of TC includes two steps: first, average pooling is performed on the spatiotemporal feature map in the time dimension (only when the time dimension is greater than 1) to compress its time dimension to 1; second, the time dimension is eliminated by the tensor reshape operation to further reduce the feature map dimension while keeping the dimensions of other dimensions unchanged. TC converts a 5-dimensional feature map with a shape of d1×d2×d3×d4×d5 into a 4-dimensional feature map with a shape of d1×d2×d4×d5, where d1, d2, d3, d4, and d5 represent the values of batch, channel, time, height, and width dimensions, respectively. For example, when t≤16, F stage The time dimension of 4 is After removing its time dimension, the spatiotemporal compression feature can be obtained When t>16, F stage The time dimension of 4 At this time, the time dimension will be The shape is The sub-feature graph of F is summed up and averaged according to the elements at the corresponding positions, so as to reduce the time dimension of the feature graph to 1. stage 4 After removing the time dimension, we can get F s ' tage 2 and F s ' tage 3 The calculation process of is similar. The feature obtained by TC processing is recorded as F s ' tage 2 、F s ' tage 3 and F s ' tage 4 ;
[0053] (3b) Use the deconvolution module TConv to s ' tage 4 Upsample to obtain the feature map TConv r-2 (F s ' tage 4 ) so that zeros are inserted into the input feature map and the convolution kernel is applied;
[0054] (3c) The feature map TConv obtained by upsampling TConv r-2 (F s ' tage 4 ) with the original feature map F of the same scale s ' tage 3 and F s ' tage 2 Perform element addition operation to obtain the medium and small scale space-time compression characteristics: and It can be expressed as:
[0055]
[0056] Example 4
[0057] A small target drone detection method with decoupled spatiotemporal feature fusion is the same as that in Example 1.
[0058] 1. Simulation conditions:
[0059] The operating system is Windows 10, configured with an Anaconda environment running Python 3.8. The deep learning framework used is Pytorch 1.13.0, paired with CUDA 11.7 for parallel computing support. The hardware environment is a host computer equipped with an Intel i7-14700KF CPU and an Nvidia RTX 3090 graphics card. Due to Pytorch and CUDA version limitations, the training and validation of the MEGA, DFF, FGFA, and RDN models was performed on an Ubuntu host equipped with an RTX 2080ti.
[0060] This example uses the Drone-Detection-Dataset drone video detection dataset for model training and testing. This dataset contains four target categories: airplanes, birds, drones, and helicopters, totaling 89,151 annotated frames. We split the dataset into a training set and a test set in a 7:3 ratio. The training set contains 197 videos and a total of 61,622 annotated frames, while the test set contains 88 videos and a total of 27,529 annotated frames.
[0061] 2. Simulation content:
[0062] The performance of the proposed decoupled spatiotemporal feature fusion detection network (DS2F-DN) is compared with that of the video detection models MEGA, FGFA, and DFF, as well as the image-based Faster R-CNN and YOLOv8 models (Faster R-CNN-50 and Faster R-CNN-101 represent variants using ResNet-50 and ResNet-101 as feature extractors, respectively). The experimental results are shown in Table 1.
[0063] Table 1
[0064]
[0065] We use mAP as the detection accuracy indicator, use single-frame inference latency (FIL) to evaluate the detection speed, and use floating-point operations (FLOPs) and model parameters (Params) as evaluation indicators of model complexity. In the experiment, the input image size of IOD is uniformly set to 600×600, and the input image size of YOWOv2 and DS2F-DN is set to 224×224 (this is a choice after comprehensive consideration of the resource overhead of the three-dimensional convolution branch). For the input image of VOD, we obtained it by scaling the original image proportionally and fixing its shortest side to 600 pixels. The training rounds of the YOLOv8 model are 150 rounds, the training rounds of YOWOv2 and DS2F-DN are 10 rounds, and the number of iterations of the other models (regardless of rounds) is 1.2×10 5 Second-rate.
[0066] As can be seen from Table 1, MEGA can effectively aggregate the local details and global semantic information of the target, and the drone detection accuracy reaches 78.31%, which is the best performance among all the compared models. However, its FIL is as high as 109.2ms, which does not meet the real-time detection requirements, and MEGA has a high number of parameters and computational complexity. In comparison, YOWOv2 has the lowest model complexity in VOD due to the use of lightweight 2D and 3D backbone networks, and its computational complexity is lower than other models, but it does not meet the real-time monitoring requirements. The DS2F-DN proposed in the present invention has the best mAP performance and excellent detection performance for small target birds and drone categories. In addition, the single-frame inference delay (FIL) of DS2F-DN is 23.4ms, which supports real-time detection.
[0067] Figure 5 and Figure 6 Based on the true positives (TP) and false positives (FP) in the detection results, we plotted the performance ratio (PR) curves of different models for each detection category. We found that the PR curve of DS2F-DN was in the upper middle range of all curves, indicating that its detection accuracy was superior to that of most of the compared models. In particular, DS2F-DN performed best in bird detection, demonstrating its superior ability to distinguish birds, which are easily confused.
[0068] In summary, DS2F-DN demonstrates excellent performance in terms of mAP, small target detection capability, robustness, model complexity, and real-time performance. Its overall performance outperforms existing methods and is suitable for small target drone detection tasks in complex scenarios.
[0069] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Components not specified in this embodiment may be implemented using existing technologies.
Claims
1. A small target drone detection method based on decoupled spatiotemporal feature fusion, characterized by: The following steps are involved: (1) Perform differential resolution processing on the input video stream to obtain high-resolution key frames and low-resolution frame sequences; (2) In the two-dimensional convolution branch, a 2D backbone network is used to extract multi-scale spatial features of high-resolution key frames, and a feature decoupling module based on wavelet transform is used to decompose the high- and low-frequency information of spatial features to obtain spatially decoupled classification features and spatially decoupled regression features for classification and regression; (3) In the 3D convolution branch, the 3D backbone network is used to extract the spatiotemporal features of the low-resolution frame sequence, and the deconvolution operation is used to upsample the spatiotemporal features, construct a feature pyramid structure, and output multi-scale spatiotemporal features; (4) performing spatiotemporal feature fusion based on a channel encoder on the spatially decoupled classification features and spatially decoupled regression features output by the two-dimensional convolution branch, as well as the multi-scale spatiotemporal features output by the three-dimensional convolution branch. The channel encoder uses a parallel point-by-point convolution structure to perform parallel dimensionality reduction and fusion processing on the features output by the two-dimensional convolution branch and the three-dimensional convolution branch, and uses a scaled dot product self-attention mechanism to calculate the correlation weights of different channel feature vectors; (5) The fused features are input into the decoupling detection head to classify and locate the drone targets in the video and output the detection results.
2. The small target drone detection method based on decoupled spatiotemporal feature fusion according to claim 1 is characterized in that: Step (1) includes the following steps: (1a) Read the input video stream and capture t consecutive video frames to form a frame sequence V = {I1, I2, ... I t }, and performing data enhancement preprocessing on the frame sequence V, wherein the input video stream includes multiple video frames; (1b) Adjust the resolution of the frame sequence V to obtain a low-resolution frame sequence V′, and extract a specific single frame from a fixed position of V and adjust the resolution of the specific single frame to obtain a high-resolution key frame I key , the key frame I key The resolution is twice that of the low-resolution frame sequence V′; (1c) Take V′ as the input of the 3D convolution branch and I key As the input of the 2D convolution branch.
3. The small target drone detection method based on decoupled spatiotemporal feature fusion according to claim 1 is characterized in that: Step (2) includes the following steps: (2a) I key Input 2D backbone network to get spatial feature map The two-dimensional discrete Haar wavelet transform is used to Perform a fine decomposition of the spatial dimension, where the superscript i∈{2,3,4} indicates the small, medium and large scales respectively, c and α indicate the number of channels and the downsampling multiple respectively, h and w indicate the height and width of the channel respectively; The decomposed output is grouped to obtain a low-frequency feature map and 3 high-frequency feature maps and Represent the approximate, horizontal edge, vertical edge and diagonal features of the image respectively; (2b) Low-frequency feature map X LL After being processed by n convolution-batch normalization-activation function CBL modules of the classification branch, the spatial decoupled classification feature map is obtained. Each CBL module consists of a convolution module, a batch normalization module, and an activation module; (2c) The three high-frequency feature maps X LH , X HL and X HH Stacking along the channel dimension, we get Stack dim=c (·) indicates that the feature maps are stacked in the channel dimension, and then the linear projection operator LP(·) is used to stack the feature maps. dim=c (X LH ,X HL ,X HH ) Project from 3c dimension to c dimension to obtain the high-frequency feature map after dimensionality reduction c is a positive integer; (2d) X HF After being processed by n CBL modules of the regression branch, a spatial decoupled regression feature map is generated 4. The small target drone detection method based on decoupled spatiotemporal feature fusion according to claim 1 is characterized in that: Step (3) includes the following steps: (3a) Use the 3D backbone network to extract features from V′ and output a multi-scale feature map F stage 2 、F stage 3 and F stage 4 ; for F stage 2 、F stage 3 and F stage 4 The time series compression module TC is used to perform time series compression to achieve dimensionality reduction. The feature map obtained by TC processing is recorded as F′ stage 2 , F′ stage 3 and F′ stage 4 ; (3b) Use the deconvolution module TConv to stage 4 Upsample to obtain the feature map TConv r-2 (F′ stage 4 ) so that zeros are inserted into the input feature map and the convolution kernel is applied; (3c) The feature map TConv obtained by upsampling TConv r-2 (F′ stage 4 ) with the original feature map F′ of the same scale stage 3 and F′ stage 2 Perform element-wise addition operation.
5. The small target UAV detection method based on decoupled spatiotemporal feature fusion according to claim 3 is characterized in that: In step (2a), the two-dimensional discrete Haar wavelet transform step specifically includes: Use the deep convolution module DWConv with a step size of 2 for processing The convolution kernel of DWConv is composed of a low-pass filter and a set of high-pass filters and composition, Each channel of is processed by the above four filters to generate four output feature maps.
6. The small target drone detection method based on decoupled spatiotemporal feature fusion according to claim 4 is characterized in that: In step (3a), the steps of the timing compression module TC specifically include: TC converts a 5-dimensional feature map of shape d1×d2×d3×d4×d5 into a 4-dimensional feature map of shape d1×d2×d4×d5, where d1, d2, d3, d4, and d5 represent the values of batch, channel, time, height, and width dimensions, respectively.
7. A system using the small target drone detection method of decoupled spatiotemporal feature fusion according to any one of claims 1 to 6, comprising: (1) Video preprocessing module, used to perform differential resolution processing on the input video stream to obtain high-resolution key frames and low-resolution frame sequences; (2) A two-dimensional convolutional branch module is used to extract multi-scale spatial features of high-resolution key frames using a 2D backbone network, and use a feature decoupling module based on wavelet transform to decompose the high- and low-frequency information of spatial features to obtain spatially decoupled classification features and spatially decoupled regression features for classification and regression; (3) A three-dimensional convolutional branch module, which is used to extract the spatiotemporal features of low-resolution frame sequences using a 3D backbone network, upsample the spatiotemporal features using a deconvolution operation, construct a feature pyramid structure, and output multi-scale spatiotemporal features; (4) a spatiotemporal feature fusion module, which is used to perform spatiotemporal feature fusion based on a channel encoder on the spatially decoupled classification features and spatially decoupled regression features output by the two-dimensional convolution branch, as well as the multi-scale spatiotemporal features output by the three-dimensional convolution branch. The channel encoder uses a parallel point-by-point convolution structure to perform parallel dimensionality reduction and fusion processing on the features output by the two-dimensional convolution branch and the three-dimensional convolution branch, and uses a scaled dot product self-attention mechanism to calculate the correlation weights of different channel feature vectors; (5) Detection output module, which is used to input the fused features into the decoupling detection head, classify and locate the drone targets in the video, and output the detection results.