Unmanned aerial vehicle tiny target detection method based on SCCA-YOLO
By optimizing the YOLOv1 backbone network and using the SCCA-YOLO model, the problem of insufficient computing efficiency of the drone micro-object detection algorithm in complex road environments is solved, and efficient and accurate small-object detection is achieved.
Patent Information
- Application Number
- CN202510784417.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-25
AI Technical Summary
The existing drone micro-objective detection algorithms are insufficient in complex road environments, which are difficult to meet real-time requirements, and the characteristics of small targets are easily lost, resulting in serious missed detection and missed detection.
Using the SCCA-YOLO model, the YOLOv1 backbone network is optimized through parallel heterogeneous convolution-multi-scale context aggregation downsampling module MCAD, space and channel collaborative attention module SCCA, multi-scale feature fusion module MSFI and spatial channel aggregation module SCC, and the YOLOv1 backbone network is optimized to enhance the retention and detection capabilities of small target features.
It improves the accuracy and adaptability of the drone's micro-object detection and complex scenarios, reduces the computational complexity, reduces deployment costs, and enhances the feature retention and detection performance of small targets.
Smart Images

Figure CN120375243A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and specifically relates to a method for detecting tiny targets of drones based on SCCA-YOLO. Background Technique
[0002] With the wide application of drones in intelligent traffic monitoring, the demand for real-time detection of vehicles and pedestrians in complex road environments from high-altitude perspectives is becoming increasingly urgent. Due to the relatively high shooting altitude of drone aerial images, the detected targets are generally small, so they are characterized by low pixel occupancy, dense distribution, and being easily affected by the environment.
[0003] In the prior art, although mainstream detection algorithms such as Faster R-CNN and Mask R-CNN have advantages in positioning accuracy and can detect and present small targets in a relatively clear state, their multi-stage processing processes lead to insufficient computational efficiency and are difficult to meet the real-time requirements of embedded devices; another stage of mainstream detection algorithms such as SSD and YOLO series usually adopt a deep downsampling strategy to balance speed and accuracy, resulting in serious loss of detailed features of tiny targets in the deep network. Especially in road scenes with frequent target occlusion and complex backgrounds, the phenomena of missed detection and false detection are particularly prominent; Therefore, there is an urgent need for a small target detection method that can balance cost, speed, and accuracy to provide a new solution for the detection and deployment of drones. Summary of the Invention
[0004] Purpose of the Invention: To solve the problems mentioned in the background technique, the present invention discloses a method for detecting tiny targets of drones based on SCCA-YOLO. The SCCA-YOLO model is obtained by improving yolov11. While ensuring the improvement of detection performance, the increase in the number of model parameters and computational complexity is controlled within an acceptable range, thereby controlling costs and reducing the deployment threshold.
[0005] Technical Solution:
[0006] The present invention discloses a method for detecting tiny targets of drones based on SCCA-YOLO, and the method includes the following steps:
[0007] S1: Obtain the aerial image of the drone to be detected;
[0008] S2: Preprocess the aerial image to be detected so that the preprocessed aerial image conforms to the usage format of the target detection model, and divide it into a training set, a validation set, and a test set;
[0009] S3: Construct the SCCA-YOLO target detection model:
[0010] S3.1: Replace the convolution module of the backbone network with the MCAD module that uses parallel heterogeneous convolution - multi - scale context aggregation for downsampling
[0011] S3.2: Introduce the spatial and channel collaborative attention module SCCA that combines the intelligent dimension adaptation mechanism PSAS at the terminal layer of the backbone network as the last layer of the backbone network;
[0012] S3.3: Replace the Bottleneck in the C3k2 module of the backbone network with Bottleneck - SCCA to construct the C3k2 - SCCA module;
[0013] S3.4: For the corresponding neck network, the 14th and 18th layers of the SCCA - YOLO model adopt the multi - scale feature fusion module MSFI operation, and the 26th layer deploys the spatial channel aggregation module SCC;
[0014] S4: Input the training set to perform object detection training on the SCCA - YOLO model;
[0015] S5: Use the trained network model to detect the test set images to verify the model performance.
[0016] Furthermore, the MCAD module described in S3.1 includes an initial downsampling unit of the ConvBNPReLU layer, a dual - path feature extraction unit, and a multi - modal attention enhancement unit MMA;
[0017] When the image feature map is input into the MCAD module, it undergoes spatial dimension compression and channel dimension expansion through the initial downsampling unit, followed by batch normalization and an activation function, and the output is the feature map U;
[0018] The dual - path structure extraction unit is used to perform parallel modeling on the downsampled feature map U to obtain the fused feature map Z; the fused feature map Z sequentially undergoes batch normalization, PReLU activation, and 1×1 point convolution operations for channel adjustment to obtain the compressed feature map V; the output V is input into the multi - modal channel attention unit MMA.
[0019] Furthermore, the Spatial and Channel Collaborative Attention Module (SCCA) described in S3.2 consists of three parts connected in series: the intelligent dimension adaptation mechanism (PSAS), the spatial attention priority calculation module (SAU), and the channel self-attention module (CAU). When the input features are processed by the SCCA module, the intelligent dimension adaptation mechanism (PSAS) performs dynamic channel dimension adjustment to achieve the optimal matching with the multi-head attention mechanism structure. Then, it enters the spatial attention priority calculation module (SAU), where multi-scale depth convolution groups are used to extract spatial features and fuse cross-axis attention weights to generate spatially enhanced features. The features are then passed to the channel self-attention module (CAU), where after downsampling, the inter-channel dependencies are modeled through the multi-head self-attention mechanism, and then weighted by the dynamic gating system and upsampled to restore the spatial dimensions. The output of the SCCA module is obtained by adding the channel attention result and the initial input after dimension adjustment through residual connection, forming a feature representation with dual spatial-channel enhancement.
[0020] Furthermore, the intelligent dimension adaptation mechanism (PSAS) realizes the optimal matching between the number of attention heads and the feature dimensions through a three-stage processing flow of legality verification - dynamic adjustment - dimension reorganization:
[0021] Legality verification: Receive the dimension parameter C of the input features in , and verify the legality of the preset number of heads h pre . If the input feature dimension is divisible by the preset number of heads, that is, C in mod h pre = 0, then this preset number of heads is used as the dynamically optimized number of heads h act ; otherwise, PSAS will automatically search for the largest legal number of heads that satisfies the divisibility condition C pre mod h in = 0 from the candidate number of heads set not greater than h pre .
[0022] Dynamic adjustment: After obtaining the legal number of heads h act , the system further calculates the single-head feature dimension d and reorganizes the multi-head attention parameter matrix to ensure that the total output dimension C out is strictly aligned with the original input dimension C in , achieving lossless information mapping.
[0023] Dimension reorganization: After the adjustment is completed, PSAS outputs a new set of attention configuration parameters, including the legal number of heads, single-head dimension, and the reconstructed parameter tensor layout, directly driving the multi-head attention mechanisms in the spatial attention priority calculation module (SAU) and the channel attention module (CAU).
[0024] Furthermore, both SAU and CAU rely on PSAS to achieve the initialization of the multi-head structure and the alignment of input and output dimensions, ensuring seamless docking of the entire attention flow from spatial guidance to channel modeling:
[0025] SAU performs multi-scale spatial attention operations based on the feature dimension C output by PSAS. Specifically, the number of groups of depthwise convolutions in SAU is strictly set to C out to ensure the consistency of subsequent concatenation dimensions, guarantee the legal execution of multi-scale depthwise convolution groups, avoid the failure of the concatenation operation due to channel dimension mismatch, and enhance the model's response to spatial location features; out
[0026] CAU uses h act and d provided by PSAS to construct the multi-head attention mechanism. The total number of projection channels of the QKV tensors generated by CAU is 3C out to ensure that C out mod h act = 0 to maintain the divisibility of the QKV tensors in the multi-head dimension, that is, the output channels can be divided evenly by the number of heads, ensuring that each attention head in the multi-head attention can obtain consistent and independent information partitioning.
[0027] Furthermore, in SAU of the spatial attention priority calculation module, a collaborative mechanism that fuses multi-scale local perception and cross-axis global modeling is proposed: parallel depthwise separable convolutions are applied to the input feature map, and the output of each branch is processed by the GELU activation function and batch normalization to obtain the corresponding local perception feature representation A k ;
[0028] Max pooling operations are performed along the height axis and width axis to obtain two one-dimensional structural attention maps Subsequently, the pooled results are interpolated to the original size of the input feature map respectively, and after concatenation in the channel dimension, they are sent to a 1×1 convolution for fusion, and finally a cross-axis spatial attention weight map A cross is generated through the Sigmoid activation function;
[0029] The above-mentioned outputs A k of each branch and the cross-axis spatial attention map A cross are weighted and fused item by item to form the final spatial attention weight matrix A final .
[0030] Furthermore, in the channel self-attention module CAU, an average pooling operation is applied to the spatially enhanced feature map X ′ , and the spatial size is reduced from H×W to H ′ ×W ′ to obtain the compressed feature Perform group normalization on the downsampled feature y; use depthwise separable 1×1 convolutions to generate query Q, key K, and value V tensors respectively.
[0031] Calculate scaled dot-product attention to obtain weighted features of shape Restore to the spatial dimension through a reshaping operation. Obtain the channel descriptor through average pooling, feed it into a multi-layer perceptron for non-linear modeling, and then generate the channel attention weight through the Sigmoid activation function. Multiply with the spatially enhanced feature map X of the spatial attention SAU ′ to obtain the final output feature, realizing channel-level recalibration.
[0032] Furthermore, the Bottleneck-SCCA described in S3.3 embeds the SCCA module on the basis of the traditional residual bottleneck structure, performs spatial-channel joint attention modeling on the intermediate features, including using 1×1 convolution for channel compression, connecting to the 3×3 convolution of the backbone network for local feature modeling, introducing the SCCA module to parallelly model spatial attention and channel attention, enhancing the intermediate features, restoring the channel dimension through 1×1 convolution, and performing residual connection.
[0033] Furthermore, the structure of the multi-scale feature fusion module MSFI described in S3.4 is as follows:
[0034] Before feature encoding, adjust the feature maps of different scales to the same channel dimension through 1×1 convolution to make them consistent with the main-scale features; adopt a bimodal pooling strategy for the high-resolution feature map, compress the largest feature map F l to the medium-scale feature F through adaptive max pooling and average pooling m in terms of spatial size, and perform element-wise summation fusion; for the low-resolution feature map F s adopt the nearest neighbor value μ nearest to upsample to the resolution of F m ; finally, perform channel dimension concatenation fusion, and concatenate the three aligned scale features in the channel dimension.
[0035] The structure of the spatial channel aggregation module SCC is as follows:
[0036] First, perform 1×1 convolutions on the features of different levels of the backbone network respectively to make each level unified to the same dimension C in the channel dimension, and align the lower-resolution p3 and p4 to the spatial size of p2 through nearest neighbor interpolation; after alignment, expand the dimension in the scale dimension through the unsqueeze operation, and stack the three-way features to form a three-dimensional tensor F; perform three-dimensional convolution fusion, perform 1×1×1 three-dimensional convolution on the stacked tensor to extract the semantic correlation between scales, and realize cross-scale and spatial local linear transformation; for the output F ′After three-dimensional batch normalization and LeakyReLU activation, through the maximum pooling operation along the scale dimension, the information on the scale dimension is condensed into the most significant response F ” ; Finally, the scale dimension is removed through dimension compression, restored to a three-dimensional feature tensor, and the fused feature tensor is output.
[0037] Beneficial effects:
[0038] 1. The present invention designs a multi-scale context aggregation downsampling module (MCAD module) to replace the traditional convolutional layer. Through the MMA multi-modal attention enhancement mechanism, channel attention (CA), spatial attention (SA), and position attention (PA) dynamically adjust the feature contributions through learnable weights, realizing attention-guided feature selection, enhancing the feature learning of small-scale targets, reducing information loss in the downsampling process, improving the feature retention of dense tiny targets, and thus improving the detection accuracy of small targets.
[0039] 2. The present invention designs a spatial and channel collaborative attention module (SCCA), and combines the intelligent dimension adaptation mechanism (PSAS) to achieve the optimal division configuration of the number of attention heads and the number of input channels, ensuring the information integrity and resource utilization efficiency in the attention calculation process. Compared with the traditional attention mechanism based on a fixed grouping strategy, the PSAS system can avoid information loss caused by truncation and computational redundancy and noise introduction caused by dimension padding while ensuring the structural rationality. Through the combination of the three, the SCCA module effectively improves the network's attention ability to key regions and significant semantics through multi-scale context awareness and spatial-channel collaborative modeling, while maintaining a low computational complexity and good structural generality, taking into account both performance and cost.
[0040] 3. The present invention deeply integrates the SCCA module with the backbone network C3K2 module to propose the C3K2_SCCA module, realizing the effective combination of cross-stage feature processing ability and feature enhancement ability, making full use of multi-semantic space information to guide spatial and channel feature extraction, rich category information and context information, enabling the model to learn higher-quality feature representations from multiple perspectives, and enhancing the model's detection ability for complex scenes;
[0041] 4. In the present invention, a multi-scale feature fusion module (MSFI) and a spatial-channel aggregation module (SCC) are designed to replace the traditional upsampling as the new Neck structure, enabling the model to focus on small targets related to high-value channels and spatial positions and cross-scale feature map fusion, enhancing the representation of detailed information, and further reducing the interference of the small target detection ability in complex scenes. Description of the drawings
[0042] Figure 1 It is a schematic diagram of the overall network framework of SCCA-YOLO in the present invention;
[0043] Figure 2 is the network structure diagram of the multi-scale context aggregation downsampling MCAD module in the present invention;
[0044] Figure 3 is the network structure diagram of the spatial channel collaborative attention module SCCA in the present invention;
[0045] Figure 4 is the network structure diagram of the C3k2-SCCA module in the present invention;
[0046] Figure 5 is the network structure diagram of the multi-scale feature fusion module MSFI and the spatial channel aggregation module SCC of the present invention;
[0047] Figure 6 is a schematic diagram for comparing the vehicle and pedestrian detection results from the perspective of the UAV in the embodiment of the present invention with the detection results of YOLOv11;
[0048] Figure 7 is a schematic diagram for comparing the average precision map50 of SCCA-YOLO with the YOLOv11 model of the present invention. Specific implementation manners
[0049] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0050] As Figure 1 shown, the present invention discloses a UAV micro-target detection method based on SCCA-YOLO, and the method steps are as follows:
[0051] S1: The UAV carries a high-resolution camera and collects images of different positions, environments, object types, and scene densities along a preset route and altitude to obtain the aerial images to be detected by the UAV.
[0052] S2: The collected aerial images are labeled and format-converted to meet the image input requirements of the YOLO model, and the data set is divided into a training set, a test set, and a validation set according to 7:2:1.
[0053] S3: Construct an SCCA-YOLO target detection model:
[0054] S3.1: The backbone network uses the MCAD module with parallel heterogeneous convolution - multi-scale context aggregation downsampling to replace the convolution module of the backbone network;
[0055] As Figure 2As shown, the MCAD module includes an initial downsampling unit (ConvBNPReLU layer), a dual-path feature extraction unit, and a multi-modal attention enhancement unit (MMA);
[0056] When the image feature map is input into the MCAD module, first, the feature map is compressed in the spatial dimension and expanded in the channel dimension through the initial downsampling unit. The specific processing includes a common convolutional layer with a convolutional kernel size of 3×3 and a stride of 2, followed by batch normalization (Batch Normalization, BN) and the Parametric ReLU (PReLU) activation function. The output at this time is U;
[0057]
[0058] Where X is the input feature map, BN is the batch normalization layer, and PReLU is the non-linear activation function.
[0059] To enrich the feature representation, the module uses a dual-path structure to parallelly model the downsampled feature map U. One path uses a standard depthwise separable convolution to extract local fine-grained features F loc , and the other path uses a depthwise separable convolution with a dilation rate of 2 to obtain context information F under a larger receptive field sur ;
[0060]
[0061] Where DConv 3×3 is the depthwise separable convolution, with a kernel size of 3×3, a stride s = 1, d is the dilation rate of the convolutional kernel, F loc is local feature extraction, and F sur is context feature extraction
[0062] The output feature maps of the two branches are fused element-wise with weights to obtain the fused feature map Z, which is used to jointly represent local and context semantics;
[0063]
[0064] The fused feature Z is successively passed through batch normalization, PReLU activation, and 1×1 point convolution operations for channel adjustment to obtain the compressed feature map V;
[0065]
[0066] The final output V will be input into the subsequent multi-modal channel attention mechanism (MMA) to further enhance the network's ability to model the importance of features in different dimensions;
[0067] Among them, the MMA module includes Channel Attention (CA), Spatial Attention (SA), and Position Attention (PA). The three calculate their respective attentions A c , A s and A p , and through the learnable fusion weight w = [ω c , ω s , ω p = [1, 1, 1]), after normalization by the Softmax function, they are weighted and combined, and finally the enhanced feature map Y is output:
[0068]
[0069] where X is the input feature map, and A i (X) represents the attention map obtained after the i-th attention mechanism acts on X, represents element-wise multiplication, and α i is the normalized fusion coefficient of each attention branch, with the initial value of w = [1, 1, 1] and can be adaptively optimized through network training.
[0070] The Channel Attention (CA) layer aims to model the dependencies between channels to improve the network's perception of the importance differences between different channels, and realizes efficient compression and restoration through bottleneck transformation;
[0071] First, perform global average pooling operation on the input feature map to compress the two-dimensional spatial information of each channel and extract the channel-level global statistical features The calculation formula is:
[0072]
[0073] where X c (i, j) represents the pixel value of the c-th channel in the input feature map at the position (i, j), and S represents the statistical description results of all channels
[0074] Secondly, perform bottleneck mapping. Through two layers of 1×1 convolutions to form a bottleneck structure (channel reduction rate r = 16), and cooperate with the ReLU activation function to establish a non-linear mapping between channels; finally, perform channel recalibration to obtain the channel attention A c
[0075]
[0076] where is the learnable weight parameter, and the Sigmoid function generates a channel weight vector in [0 - 1];
[0077] The CA layer uses a pair of bottleneck convolutions (1×1) to complete the learning from global description to channel weights. Among them, W1 is responsible for compressing the C-dimensional global statistics to dimensions; W2 is responsible for restoring the activated structure back to C dimensions, obtaining channel weights in the interval (0,1) through the Sigmoid transformation, and finally weighting the original feature map by channel to achieve adaptive channel recalibration.
[0078] Adopting the bottleneck structure of "dimensionality reduction - activation - dimensionality increase" can effectively capture the dependencies between channels while ensuring the lightweight of the model, and improve the network's ability to express key information.
[0079] The SA spatial attention layer aims to model the importance of the spatial positions of the feature map, enhance the response of key regions in an adaptive manner, and suppress redundant background interference at the same time to improve the discriminative ability of feature expression;
[0080] For the input feature map X, the SA layer generates the spatial attention weight map A through convolution operation s :
[0081]
[0082] where Conv 7×7 is a convolution with a large receptive field, and the Sigmoid function generates a channel weight vector in [0-1]
[0083] The SA layer generates a spatial position weight map through a layer of convolution with a large receptive field and Sigmoid activation, and multiplies it element-wise with the input features to achieve adaptive enhancement or suppression of different spatial regions, thus effectively focusing on the key target regions.
[0084] The PA position attention layer aims to perform refined modeling and adaptive recalibration on the cross-channel features of each spatial position in the feature map. Through channel dimensionality reduction, the number of channels is compressed to 1 / 8 of the original using 1×1 convolution (dimensionality reduction parameter d = 8), and grouped dilated convolution (group = C / d) is adopted. 3×3 grouped convolution is used to capture the local spatial model and establish pixel-level dependencies. Finally, Sigmoid normalization is performed to generate an attention map A with the same shape as the input, and the expression is: p , the expression is:
[0085]
[0086] where DWConv 3×3 represents depthwise separable grouped convolution, and the Sigmoid activation function is used to generate a normalized attention map.
[0087] This design can reduce the loss of spatial information while enhancing the model's understanding of multi-scale contexts, expanding the receptive field, capturing more extensive context information, and better improving the feature discriminability in pedestrian and vehicle occlusion areas.
[0088] S3.2: Introduce the Spatial and Channel Cooperative Attention Module (SCCA) combined with the intelligent dimension adaptation mechanism (PSAS) at the terminal layer of the backbone network as the last layer of the backbone network;
[0089] The SCCA module mainly consists of three parts: the intelligent dimension adaptation mechanism (PSAS), the spatial attention priority calculation module (SAU), and the channel self-attention module (CAU);
[0090] When the input feature is processed by the SCCA module, first, the intelligent dimension adaptation mechanism (PSAS) performs dynamic channel dimension adjustment to achieve adaptive alignment with the multi-head attention mechanism structure; then it enters the spatial attention priority calculation module (SAU) part, where multi-scale depth convolutional groups are used to extract spatial features and fuse cross-axis attention weights to generate spatially enhanced features; next, the feature is passed to the channel self-attention module (CAU) part, where after downsampling, the inter-channel dependence relationship is modeled through the multi-head self-attention mechanism, and then weighted by the dynamic gating system and upsampled to restore the spatial size;
[0091] The above three sub-modules are connected in a serial connection manner to process the input feature in sequence. Finally, the module output is obtained by adding the channel attention result and the initial input after dimension adjustment through a residual connection, thus forming a feature representation with dual spatial-channel enhancement. The overall structure is as Figure 3 shown.
[0092] In the intelligent dimension adaptation mechanism (PSAS) part, this mechanism first performs a legality verification, receives the dimension parameter C in of the input feature, and verifies the legality of the preset number of heads h pre . If the input feature dimension is divisible by the preset number of heads, that is, C in mod h pre = 0, then the preset number of heads is used as the dynamically optimized number of heads h act . Otherwise, PSAS will automatically search for the largest legal number of heads that satisfies the divisibility condition C pre mod h in = 0 from the set of candidate numbers of heads not greater than h pre :
[0093]
[0094] where h act is the dynamically optimized number of heads, and C in is the input feature dimension
[0095] Next, dynamic adjustment is performed. After obtaining the legal number of heads h act the system further calculates the single-head feature dimension d and reorganizes the multi-head attention parameter matrix to ensure that the total output dimension C out is strictly aligned with the original input dimension C in to achieve lossless information mapping;
[0096]
[0097] where: d is the single-head feature dimension, C in is the input feature map dimension, and C out is the effective output dimension after system adjustment
[0098] Finally, dimension reorganization is carried out. After the adjustment is completed, PSAS outputs a new group of attention configuration parameters, including the legal number of heads, single-head dimension, and the reconstructed parameter tensor layout, directly driving the multi-head attention mechanism in the subsequent spatial attention priority calculation module (SAU) and channel attention module (CAU). Both SAU and CAU rely on this mechanism to achieve multi-head structure initialization and input-output dimension alignment, thus ensuring seamless docking of the entire attention flow from spatial guidance to channel modeling and avoiding information fragmentation and channel resource waste;
[0099] The PSAS module undertakes the key tasks of dynamically calculating the legal number of heads h act and the single-head dimension d;
[0100] The spatial attention priority calculation module (SAU) performs multi-scale spatial attention operations based on the feature dimension C output by PSAS out Specifically, the number of groups of depth convolutions in SAU is strictly set to C out to ensure the consistency of subsequent splicing dimensions, thereby ensuring the legal execution of the multi-scale depth convolution group and avoiding the failure of the splicing operation due to channel dimension mismatch, and enhancing the model's response to spatial position features;
[0101] In the channel self-attention module (CAU), the h act and d provided by the PSAS module are used to construct the multi-head attention mechanism. The total number of projection channels of the QKV tensors generated by CAU is 3C out To ensure the divisibility of the subsequent QKV tensors in the multi-head dimension, the system strictly ensures that C out mod h act = 0, that is, the output channel number is divisible by the number of heads, effectively ensuring that each attention head in the multi-head attention can obtain consistent and independent information partitioning, thus avoiding information fragmentation problems caused by dimension mismatch or incomplete segmentation;
[0102] In the Spatial Attention Priority Computation Module (SAU), to balance spatial information enhancement and computational efficiency, a collaborative mechanism that integrates multi-scale local perception and cross-axis global modeling is proposed. First, the input feature map is processed using parallel depthwise separable convolutions (Conv1D, with kernel sizes of 3, 5, and 7 respectively) to cover different receptive fields, thereby capturing multi-scale spatial context information and alleviating the sensitivity of the traditional single-kernel structure to scale changes. The output of each branch is processed through the GELU activation function and batch normalization (Batch Normalization, BN) to obtain the corresponding local perception feature representation A k ;
[0103] A k = BN(GELU(Conv1D k (X)))
[0104] where Conv1D k is a depthwise separable convolution with a kernel size of k×k, k = {3, 5, 7} is a preset multi-scale kernel set, GELU is the activation function, and BN is the batch normalization layer.
[0105] To further enhance the network's ability to model the row and column structures of images, max-pooling operations are performed along the height axis and width axis respectively to obtain two one-dimensional structural attention maps Subsequently, the pooled results are interpolated to the original size of the input feature map respectively, and after concatenation in the channel dimension, they are fed into a 1×1 convolution for fusion, and finally, a cross-axis spatial attention weight map A is generated through the Sigmoid activation function cross ;
[0106] A h = Interp(MaxPool h (X))
[0107] A w = Interp(MaxPool w (X))
[0108] A cross = sigmoid(Conv 1×1 (A h ||A w ))
[0109] where MaxPool h is the max-pooling along the height axis, MaxPool w is the max-pooling along the width axis, Sigmoid generates a [0-1] spatial attention weight, || is channel concatenation, and Interp represents the bilinear interpolation function
[0110] Subsequently, the outputs A of the above branchesk With the cross-axis spatial attention map A cross Perform element-by-element weighted fusion to form the final spatial attention weight matrix A final ;
[0111]
[0112] Where K = {3, 5, 7} is the set of convolution kernels.
[0113] To obtain the final spatially enhanced feature map, use the Sigmoid function to control the attention intensity and perform element-by-element multiplication with the original input feature map X to obtain the feature output X ′ , enabling the network to selectively emphasize important spatial row and column structures, achieving feature emphasis in significant regions and suppression in redundant regions.
[0114]
[0115] Where represents element-by-element multiplication, and A final is the spatial attention weight matrix, and Sigmoid generates [0 - 1] spatial attention weights.
[0116] In the channel self-attention module (CAU) part, aiming to further model the long-range dependence relationship between channels after completing spatial modulation (i.e., SAU spatial attention enhancement), to effectively control the computational complexity and improve the channel modeling efficiency, first apply average pooling operation to the spatially enhanced feature map X ′ , reducing the spatial size from H×W to H ′ ×W ′ , obtaining the compressed feature
[0117] After downsampling, perform group normalization on the downsampled feature y to enhance the stability of the feature distribution. Then use depthwise separable 1×1 convolutions to generate query Q, key K, and value V tensors respectively, for modeling the dependencies between different positions in the feature map, enhancing the model's ability to capture global context information:
[0118] Q, K, V = Conv q (y), Conv k (y), Conv v (y)
[0119] The initial shapes of the above tensors Where B is the batch size, C is the number of channels, and H ′ , w ′ are the spatial dimensions after downsampling. To implement the multi-head attention mechanism, reshape each tensor into a three-dimensional structure Q, K, where h is the number of attention heads, is the channel dimension of each attention head, and N = H ′ ×W ′ is the spatial flattening length of each channel.
[0120] Next, calculate the Scaled Dot-Product Attention to explore the intra-channel and cross-channel dependency relationships:
[0121]
[0122] where Q is the query tensor, K is the key tensor, and V is the value tensor, is the scaling factor.
[0123] Get the weighted feature Attention with shape and then restore it to the spatial dimension through a reshaping operation On this basis, extract its global channel statistical information, obtain the channel descriptor through average pooling (AvgPooling), and feed it into a Multi-Layer Perceptron (MLP) for non-linear modeling and then generate the channel attention weight through the Sigmoid activation function Multiply it with the feature output X of the spatial attention SAU part ′ to obtain the final output feature and achieve channel-level recalibration.
[0124] M c = Sigmoid(MLP(Avg(Attention)))
[0125] Subsequently, upsample the channel attention weight M C back to the size of the downsampled feature map by bilinear interpolation (Interp) and apply the per-channel weighting operation (Hadamard Product) to obtain the channel-modulated feature, upsample it to the original size and apply gating to get M ′ c .
[0126]
[0127] Finally, perform channel-level element-wise multiplication on the channel attention output M ′ c and the previous SAU spatial enhancement output X ′ to obtain the final co-attention enhanced feature representation:
[0128]
[0129] Among them, X ′ is the output feature of the spatial attention part, M ′ c is the channel attention weight, is the channel-level element-wise multiplication, MLP is the multi-layer perceptron, and Interp is the bilinear interpolation.
[0130] The SCCA module improves the feature expression of images through spatial-channel collaborative processing. The former captures spatial priors with multi-scale row-column attention to provide a more focused input for the follow-up, and the latter captures long-range dependencies between channels with scaled dot-product self-attention and converts the results into global channel attention weights to recalibrate the features again. It provides both accurate and efficient feature enhancement for tasks sensitive to fine-grained features such as small object detection in the downstream, and at the same time, to improve the adaptability of the module in multi-tasks and multi-network structures, a PSAS is introduced to provide a dynamically adaptable attention configuration mechanism, enabling the module to dynamically adjust the fusion strategy of spatial and channel attention according to the context environment and task requirements, and realizing flexible scheduling and structural generalization at the module level.
[0131] S3.3: Replace the Bottleneck in the C3k2 module in the backbone network with the Bottleneck-SCCA module to propose the C3k2-SCCA module; this module simultaneously has the cross-stage feature processing ability of C3k2 and the feature enhancement ability of SCCA, improving the adaptability of the model to complex traffic scenarios. The specific structure is as Figure 4 shown;
[0132] The Bottleneck-SCCA structure embeds the SCCA module on the basis of the traditional residual bottleneck structure to perform spatial-channel joint attention modeling on the intermediate features and enhance the response ability of key target regions. The key operations include using 1×1 convolution for channel compression, connecting to the backbone 3×3 convolution for local feature modeling, introducing the SCCA module to parallelly model spatial attention and channel attention, enhancing the intermediate features, restoring the channel dimension through 1×1 convolution and performing residual connection;
[0133] This structure not only retains the lightweight nature and residual propagation advantages of the traditional Bottleneck, but also introduces a multi-dimensional attention mechanism to strengthen the discriminability of feature representation.
[0134] S3.4: Corresponding to the neck network, the 14th and 18th layers of the SCCA-YOLO model adopt the multi-scale feature fusion module MSFI operation, and the 26th layer deploys the spatial-channel aggregation module SCC;
[0135] Specifically, the MSFI module is mainly responsible for the scale alignment and context integration tasks of feature maps at different levels. It first appears in the 14th layer of the SCCA-YOLO network and is used to fuse high-semantic features (12th layer), middle-level features (6th layer), and fine-grained low-level features (13th layer) from the backbone network. After multi-scale scaling and alignment operations, these feature maps are concatenated and integrated to form fused features at a unified scale. The second MSFI operation is located in the 18th layer, and the inputs include high-resolution detail features (17th layer), middle-level features (4th layer), and upstream downsampled features (16th layer) from the backbone network. MSFI sequentially performs scale normalization, bimodal pooling, and parameter-free concatenation fusion operations on large, medium, and small-scale feature maps to achieve scale unification and context modeling, enhancing the discriminative ability of the fused features;
[0136] On this basis, the SCC module is deployed in the 26th layer of the network to further model the correlation and complementarity between multi-scale features. SCC receives feature maps from the 4th layer (P3), 6th layer (P4), and 8th layer (P5) of the backbone network, generates global features through channel compression and cross-scale attention weighting, and fuses them with the P3 detection head features (19th layer) through an element-wise addition operation in the 27th layer to form small object detection features with enhanced context;
[0137] The two jointly optimize the detection performance of multi-scale objects, especially improving the detection accuracy of small objects. The specific structure is as Figure 5 shown:
[0138] The structure of the MSFI module is as follows:
[0139] Before feature encoding, different-scale feature maps are adjusted to a unified channel dimension through 1×1 convolution to make them consistent with the main-scale features.
[0140] For high-resolution feature maps, a bimodal pooling strategy is adopted, and the largest feature map F l is compressed to the spatial size of the medium-scale feature F m through adaptive max pooling and average pooling, and element-wise summation fusion is performed. The high-resolution feature map F l performs the operation as:
[0141]
[0142] where ρ max represents adaptive max pooling, ρ avg represents adaptive average pooling, is element-wise summation.
[0143] This design combines max pooling ρ max and average pooling ρ avg, introduce context information while maintaining the significant response, effectively suppressing the loss of high-frequency information during the feature downsampling process.
[0144] For the low-resolution feature map F s Use the nearest neighbor value μ nearest Upsample it to the resolution of F m :
[0145] F′ s = μ nearest (F s , H m , W m )
[0146] where μ nearest represents nearest neighbor interpolation upsampling
[0147] Use the nearest neighbor interpolation method to maintain geometric structure consistency and avoid the artifact effect caused by linear interpolation.
[0148] Finally, perform channel dimension concatenation fusion, and concatenate the aligned three-scale features in the channel dimension:
[0149]
[0150] where C total = C l + C m + C s , Φ concat represents the channel dimension concatenation operation
[0151] The MSFI module realizes the collaborative optimization of cross-scale semantic information and the retention of multi-granularity context through a hierarchical feature alignment framework. During the feature space transformation process, a bimodal pooling strategy is adopted to fuse the local significance extraction ability of max pooling and the global context awareness characteristics of average pooling, and the high-frequency feature attenuation is effectively suppressed through operator superposition. For low-resolution feature reconstruction, a topology-preserving interpolation mechanism is designed to maintain the integrity of the underlying geometric structure while avoiding spatial distortion.
[0152] The cross-level parameter-free fusion achieved through channel dimension concatenation constructs a direct gradient path while retaining the characteristics of the original feature distribution, significantly improving the backpropagation efficiency. The unified scale feature space output by the module realizes the hierarchical calibration of high-level semantic abstraction and low-level detail representation. Especially for complex visual tasks, it exhibits excellent scale invariance. This integration mechanism constructs a discriminative and robust composite representation base by coordinating the semantic conflicts between multi-scale features, providing cross-level collaborative optimization feature support for downstream tasks.
[0153] To further integrate the feature information from different semantic levels, a spatio-channel consolidator (SCC) module is proposed. Based on the three-dimensional modeling of stacked features and the attention convergence strategy, this module realizes the integrated aggregation of cross-scale features.
[0154] The structure of the spatio-channel consolidator (SCC) module is as follows:
[0155] First, 1×1 convolutions are performed on the features of different levels of the backbone network (p2, p3, p4 in the figure) respectively, so that each level is unified to the same dimension C in the channel dimension, and the lower-resolution p3 and p4 are aligned to the spatial size of p2 through nearest-neighbor interpolation;
[0156]
[0157] After alignment, the dimension is expanded in the scale dimension through the unsqueeze operation, and the three-way features (P2, P3, P4) are stacked to form a three-dimensional tensor F;
[0158]
[0159] Among them, unsqueeze is the newly added scale dimension.
[0160] Subsequently, three-dimensional convolution fusion is performed. A 1×1×1 three-dimensional convolution is performed on the stacked tensor to extract the semantic correlation between scales and realize the local linear transformation across scales and in space;
[0161]
[0162] Among them, Conv3D is the three-dimensional convolution kernel.
[0163] Then, for the output F ′ After three-dimensional batch normalization and LeakyReLU activation, through the max-pooling operation along the scale dimension, the information in the scale dimension is condensed into the most significant response F”;
[0164]
[0165] Among them, BN is three-dimensional batch normalization, LEakyReLU is the activation function with a leakage coefficient, and Maxpool3D is max-pooling along the scale dimension.
[0166] Finally, the scale dimension is removed through dimension compression, and it is restored to a three-dimensional feature tensor, and the fused feature tensor is output;
[0167]
[0168] Among them, Squeeze is to remove the scale dimension.
[0169] With such a design, SCC can take into account the unique semantics of each pyramid layer during the cross-scale interaction. For high-level features, it focuses on the overall structure and semantic consistency of the target, and dynamically allocates network capacity through adaptive weight learning; for low-level features, it emphasizes fine-grained textures and edge details to enhance the response of small target regions. After the features of each layer are subjected to normalized projection and learnable transformation in a unified semantic space, they are weighted and fused based on the association strength, which not only retains the unique semantic information of each level but also avoids feature conflicts caused by simple splicing or smooth weighting. At the same time, with the help of the local receptive field of 3D convolution and the saliency selection of pooling, the dynamic changes and detailed differences of the target are captured in parallel in the time and space dimensions, effectively improving the recognition ability of small targets in occlusion and similar backgrounds and achieving efficient and orderly fusion of multi-level features.
[0170] S4: Input the training set to train the SCCA-YOLO model for object detection;
[0171] S5: Use the trained network model to detect the test set images to verify the model performance. The vehicle and pedestrian detection results from the perspective of the UAV in this embodiment are compared with the detection results of YOLOv11 as Figure 6 shown.
[0172] The embodiments of the present invention are carried out under unified hardware and software configurations. The graphics processor is selected as NVIDIA GeForce RTX 4090, the software environment is deployed on the Ubuntu 22.04 operating system, the programming environment is based on the Python 3.12 interpreter and the PyTorch 2.3.0 deep learning framework, and the parallel computing architecture adopts the CUDA 12.1 acceleration library. The specific parameters for model training are as follows: The number of epochs is set to 300 times, the base learning rate adopts a dynamic adjustment strategy (initial value 0.01, final value 0.20), and the weight decay coefficient (weight decay) is configured as 0.0005. The optimization algorithm selects the momentum stochastic gradient descent method (Stochastic Gradient Descent, SGD), and the momentum factor is set to 0.937. The input data dimension is uniformly adjusted to a resolution of 640×640 pixels after standardization preprocessing, and the batch size is dynamically adjusted according to the video memory capacity. All experiments are carried out with the same hyperparameters, without loading any pre-trained weight coefficients, and the model is trained on the VisDrone2019 dataset. During the training process, SCCA-YOLO is compared with the original YOLOv11n model as Figure 7As shown, the experimental results are analyzed as shown in Table 1. The mAP50 and mAP50-95 of SCCA-YOLO reach 35.5% and 20.8% respectively, which are 2.4 and 1.5 percentage points higher than those of the baseline model YOLOv11n, and the accuracy rate is increased by 2.5 percentage points to 47%.
[0173] Table 1
[0174] Name Accuracy Rate / % mAP50 / % mAP50 - 95 / % SCCA - YOLO 47.0 35.5 20.8 YOLOv11 44.5 33.1 19.3
[0175] The above description of the embodiments enables those skilled in the art to implement or use the present invention. Various modifications to the embodiments will be obvious to those skilled in the art. The general principles of the present invention can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention should not be limited to the embodiments shown herein, but should cover the widest scope that conforms to the principles and novel features disclosed in the present invention.
Claims
1. A method for detecting tiny targets of drones based on SCCA-YOLO, characterized in that, The method includes the following steps: S1: Obtain the aerial images of the drones to be detected; S2: Preprocess the aerial images to be detected so that the preprocessed aerial images conform to the usage format of the object detection model, and divide them into a training set, a validation set, and a test set; S3: Construct the SCCA-YOLO object detection model: S3.1: Replace the convolutional module of the backbone network with the MCAD module that uses parallel heterogeneous convolution - multi-scale context aggregation downsampling in the backbone network S3.2: Introduce the spatial and channel collaborative attention module SCCA that combines the intelligent dimension adaptation mechanism PSAS at the terminal layer of the backbone network as the last layer of the backbone network; S3.3: Replace the Bottleneck in the C3k2 module of the backbone network with Bottleneck-SCCA to construct the C3k2-SCCA module; S3.4: Corresponding to the neck network, the 14th and 18th layers of the SCCA-YOLO model adopt the multi-scale feature fusion module MSFI operation, and the 26th layer deploys the spatial channel aggregation module SCC; S4: Input the training set to perform object detection training on the SCCA-YOLO model; S5: Use the trained network model to detect the test set images to verify the model performance.
2. The method for detecting tiny targets of drones based on SCCA-YOLO according to claim 1, wherein, The MCAD module described in S3.1 includes an initial downsampling unit of the ConvBNPReLU layer, a dual-path feature extraction unit, and a multi-modal attention enhancement unit MMA; When the image feature map is input into the MCAD module, it is subjected to spatial dimension compression and channel dimension expansion through the initial downsampling unit, followed by batch normalization and an activation function, and the output is the feature map U; Adopt the dual-path structure extraction unit to perform parallel modeling on the downsampled feature map U to obtain the fused feature map Z; the fused feature map Z is sequentially subjected to batch normalization, PReLU activation, and 1×1 point convolution operations for channel adjustment to obtain the compressed feature map V; The output V is input into the multi-modal channel attention unit MMA.
3. The method for detecting tiny targets of an unmanned aerial vehicle based on SCCA-YOLO according to claim 1, wherein, The spatial and channel collaborative attention module SCCA described in S3.2 is serially connected by three parts: the intelligent dimension adaptation mechanism PSAS, the spatial attention priority calculation module SAU, and the channel self-attention module CAU; when the input feature is processed by the SCCA module, the intelligent dimension adaptation mechanism PSAS performs dynamic channel dimension adjustment to achieve the optimal match with the structure of the multi-head attention mechanism; enter the spatial attention priority calculation module SAU, extract spatial features through a multi-scale depth convolution group and fuse cross-axis attention weights to generate spatially enhanced features; The feature is transmitted to the channel self-attention module CAU, where the inter-channel dependence relationship is modeled through the multi-head self-attention mechanism after downsampling, and then weighted by the dynamic gating system and restored to the spatial size through upsampling; the output of the SCCA module is formed by adding the channel attention result and the initial input after dimension adjustment through a residual connection to form a feature expression with dual spatial-channel enhancement.
4. The method for detecting tiny targets of drones based on SCCA-YOLO according to claim 3, characterized in that, The intelligent dimension adaptation mechanism PSAS realizes the optimal match between the number of attention heads and the feature dimension through a three-stage processing flow of legality verification - dynamic adjustment - dimension reorganization: Legality verification: Receive the dimension parameter C of the input feature in , and verify the legality of the preset number of heads h pre . If the input feature dimension is divisible by the preset number of heads, i.e., C in mod h pre = 0, then adopt this preset number of heads as the dynamically optimized number of heads h act ; otherwise, PSAS will automatically search for the largest legal number of heads that satisfies the divisibility condition C pre mod h in = 0 from the set of candidate numbers of heads not greater than h pre ; Dynamic adjustment: After obtaining the legal number of heads h act The system further calculates the single-head feature dimension d and reorganizes the multi-head attention parameter matrix to ensure that the total output dimension C out is strictly aligned with the original input dimension C in to achieve lossless information mapping; Dimension Reorganization: After the adjustment is completed, PSAS outputs a new set of attention configuration parameters, including the legal number of heads, the dimension of a single head, and the reconstructed parameter tensor layout, directly driving the multi-head attention mechanism in the Spatial Attention Priority Calculation Module (SAU) and the Channel Attention Module (CAU).
5. The method for detecting tiny targets of drones based on SCCA-YOLO according to claim 4, wherein Both SAU and CAU rely on PSAS to achieve the initialization of the multi-head structure and the alignment of input and output dimensions, ensuring seamless docking of the entire attention flow from spatial guidance to channel modeling: The feature dimension C output by the SAU with PSAS out is used as the basis to perform multi-scale spatial attention operations. Specifically, the depth convolution group number of the SAU is strictly set to C out , ensuring the consistency of the subsequent splicing dimensions, guaranteeing the legal execution of the multi-scale depth convolution group, avoiding the failure of the splicing operation due to channel dimension mismatch, and realizing the enhancement of the model's response to spatial position features; The CAU uses h provided by PSAS act and d to construct the multi-head attention mechanism. The total number of projection channels of the QKV tensor generated by the CAU is 3C out , ensuring that C out mod h act = 0 to maintain the divisibility of the QKV tensor in the multi-head dimension, that is, the number of output channels is divisible by the number of heads, ensuring that each attention head in the multi-head attention can obtain a consistent and independent information partition.
6. The method for detecting tiny targets of an unmanned aerial vehicle based on SCCA-YOLO according to claim 5, wherein, In the Spatial Attention Unit (SAU) of the spatial attention priority calculation module, a collaborative mechanism that integrates multi-scale local perception and cross-axis global modeling is proposed: parallel depthwise separable convolutions are applied to the input feature map, and the output of each branch is processed by the GELU activation function and batch normalization to obtain the corresponding local perception feature representation A k ; Perform max pooling operations along the height axis and the width axis to obtain two one-dimensional structural attention maps With Subsequently, the pooled results are respectively interpolated to the original size of the input feature map, and after concatenation in the channel dimension, they are fed into a 1×1 convolution for fusion. Finally, a cross-axis spatial attention weight map A is generated through the Sigmoid activation function cross ; Output each of the above branches A k and the cross-axis spatial attention map A cross are weighted and fused item by item to form the final spatial attention weight matrix A final .
7. The method for detecting tiny targets of drones based on SCCA-YOLO according to claim 5, characterized in that In the channel self-attention module CAU, an average pooling operation is applied to the spatially enhanced feature map X′, and the spatial dimensions are reduced from H×W to H′×W′ to obtain the compressed feature The downsampled feature y is subjected to group normalization; depthwise separable 1×1 convolutions are used to generate query Q, key K, and value V tensors respectively, Calculate the scaled dot - product attention to obtain weighted features of shape and restore them to the spatial dimension through a reshaping operation Obtain the channel descriptor through average pooling, send it into a multi - layer perceptron for non - linear modeling, and then generate channel attention weights through the Sigmoid activation function Multiply it with the spatially enhanced feature map X′ of the spatial attention SAU to obtain the final output feature, realizing channel - level recalibration.
8. The method for detecting tiny targets of an unmanned aerial vehicle based on SCCA-YOLO according to claim 1, wherein The BottleNeck-SCCA described in S3.3 embeds the SCCA module in the traditional residual bottleneck structure to perform spatial-channel joint attention modeling on intermediate features, including using 1×1 convolution for channel compression, accessing the 3×3 convolution of the backbone network for local feature modeling, introducing the SCCA module to parallelly model spatial attention and channel attention to enhance intermediate features, and restoring the channel dimension through 1×1 convolution and performing residual connection.
9. The method for detecting tiny targets of drones based on SCCA-YOLO according to claim 1, wherein, The structure of the Multi-Scale Feature Fusion Module (MSFI) described in S3.4 is as follows: Before feature encoding, the feature maps of different scales are adjusted to the same channel dimension through 1×1 convolution to make them consistent with the main-scale features; for the high-resolution feature maps, a dual-modal pooling strategy is adopted to compress the largest feature map F l to the medium-scale feature F through adaptive max pooling and average pooling m in terms of spatial size, and perform element-wise summation fusion; for the low-resolution feature map F s use the nearest neighbor value μ nearest to upsample to the resolution of F m ; finally, perform channel dimension concatenation fusion, and concatenate the three-scale features after alignment in the channel dimension; The structure of the Spatial Channel Aggregation Module (SCC) is as follows: First, perform 1×1 convolution on the features of different levels of the backbone network respectively to unify each level to the same dimension C in the channel dimension, and align the lower-resolution p3 and p4 to the spatial size of p2 through nearest neighbor interpolation; After the alignment is completed, expand the dimension in the scale dimension through the unsqueeze operation, and stack the three-way features to form a three-dimensional tensor F; perform three-dimensional convolution fusion, perform 1×1×1 three-dimensional convolution on the stacked tensor to extract the semantic correlation between scales, and achieve cross-scale and spatial local linear transformation; after the output F′ is batch-normalized in three dimensions and activated by LeakyReLU, condense the information in the scale dimension into the most significant response F” through the max pooling operation along the scale dimension; finally, remove the scale dimension through dimension compression and restore it to a three-dimensional feature tensor to output the fused feature tensor.
Citation Information
Cited By
Lightweight target detection method
CN120807955A
Multi-core feature representation learning system and method for enhancing remote sensing image based on content retrieval
CN120853029A
Small target detection method for unmanned aerial vehicle data
CN120953859A
Unmanned aerial vehicle target detection method and system based on YOLOv12
CN121330561A
High-resolution unmanned aerial vehicle multi-scale remote sensing image target detection method and device
CN121640199A