Unmanned aerial vehicle image target detection algorithm based on cross-layer single-head attention mechanism and context information enhancement
By introducing cross-layer single-head attention, multi-scale context extraction and adaptive local context enhancement modules into the YOLOv8 algorithm, the problems of complex occlusion and high density distribution of targets in the drone images are solved, which significantly improves detection accuracy and reduces false detection and missed detection rates.
Patent Information
- Application Number
- CN202510103623.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The YOLOv8 algorithm is difficult to effectively solve the problems of complex occlusion and high density distribution of targets in drone images, resulting in low detection accuracy.
Based on the YOLOv8 algorithm, cross-layer single head attention (CLSHA) module, multi-Scale Context Extraction (MSCE) module and adaptive local context enhancement (ALCE) module are introduced to enhance the modeling ability of global context information, extract multi-scale context information, and accurately perceive information about the surrounding areas of the target.
Through the combination of these modules, the detection accuracy of densely occluded targets in drone images is significantly improved, the error detection rate and missed detection rate are reduced, and good adaptability and performance are shown in experiments on different drone data sets.
Smart Images

Figure CN120014237A_ABST
Abstract
Description
Technical Field
[0001] This invention provides a drone aerial target detection algorithm based on the YOLOv8 algorithm. In view of the limitations of YOLOv8 in the field of drone target detection, the present invention adopts a series of optimization strategies to improve the detection performance. This technology can be integrated into the drone system to accurately identify the predetermined target. Background Art
[0002] Drones have been widely used in military and civilian fields due to their small size, flexible maneuverability and economical cost. With the advancement of deep learning and hardware technology, computer vision tasks such as image classification, object detection and semantic segmentation have achieved significant performance improvements. This makes the combination of drones and object detection a frontier area of research.
[0003] In the scenes shot by drones, target detection faces different challenges from conventional scenes, especially the complex occlusion problems between the target and the background, between targets, and the high density distribution of targets. In order to improve the detection accuracy, it is necessary to design and improve the existing general target detection algorithms specifically for the unique scenes of drones to adapt to the characteristics of drone images.
[0004] In order to solve the problem of overlapping occlusion between targets and occlusion between targets and background in drone images, many studies have modeled the global context and used context information to distinguish similar targets and supplement the missing target information. For example, the paper "TPH-YOLOv5:Improved YOLOv5 based on transformer prediction head for object detection on drone-captured scenarios[C].Proceedings ofthe IEEE / CVFInternational Conference on Computer Vision,2021:2778-2788" introduced the Transformer prediction head, the paper "YOLOv7 based on multi-scale for object detection on UAV aerialphotography[J].Drones,2023,7(3):188" adopted the Swin Transformer prediction head, and the paper "Densetiny object detection:A scene context guided approach and a unified benchmark[J].IEEE Transactions on Geoscience and Remote Sensing,2024" designed a scene context fusion module. These methods use multi-head attention to model long-range spatial dependencies, thereby effectively obtaining global context information and improving the detection network's ability to identify overlapping and occluded targets. However, although the multi-head attention mechanism performs well in capturing global information, its computational complexity is large, which significantly increases the computational complexity of the model and thus affects the detection efficiency.
[0005] Therefore, based on the YOLOv8 algorithm, the present invention improves the target density and occlusion problems, and proposes a drone image target detection algorithm based on a cross-layer single-head attention mechanism and context information enhancement. The ablation experiment proves the effectiveness of the three modules proposed in the present invention; the false detection analysis experiment proves the effect of the present invention in improving the false detection of densely occluded targets; the comparative experiments on three drone image data sets verify the advancement, adaptability and effectiveness of the present invention in densely occluded target detection. Summary of the invention
[0006] Aiming at the shortcomings of the original YOLOv8 algorithm in the detection of complex occlusion and high-density distribution of targets in drone images, the present invention optimizes and designs the YOLOv8 algorithm based on the detection targets of drone aerial images and the characteristics of the deployment equipment. Firstly, in view of the fact that the YOLOv8 algorithm cannot effectively solve the problem of dense occlusion of aerial images in drone scenes in general scenarios, a Cross Layer Single Head Attention (CLSHA) module is proposed to enhance the ability of long-distance dependency modeling and extracting global context information, help identify the association between dense targets, avoid misdetecting multiple targets as a single target, and assist in inferring the position of the occluded target; secondly, a Multi-Scale Context Extraction (MSCE) module is designed to improve the ability to extract multi-scale context information and enhance the model's ability to distinguish dense similar targets; finally, the present invention designs an Adaptive Local Context Enhancement (ALCE) module to achieve accurate perception of the area around the target and supplement the key information missing from the occluded target.
[0007] The technical solution adopted by the present invention to solve the technical problem comprises the following steps:
[0008] Step 1: Construct a drone image target detection algorithm based on cross-layer single-head attention mechanism and context information enhancement;
[0009] Step 1-1: Add the CLSHA cross-layer single-head attention module before the small object detection head of the YOLOv8 object detection algorithm; this module combines the high-level feature map with the low-level feature map as the input of the attention mechanism, and only uses part of the channels for a single attention calculation;
[0010] Step 1-2: Add the MSCE multi-scale context extraction module after the YOLOv8 backbone network; first, pass the input feature map through a 1×1 convolution layer; then, divide the feature map into two parts by channel: one part captures multi-scale context information through a set of parallel depth-separable convolutions of different scales and fuses them, and finally adjusts the fused feature map through a 1×1 convolution; the other part is directly channel-joined with the adjusted fused feature map to obtain the final output of the entire module;
[0011] Step 1-3: Streamline the YOLOv8 algorithm detection head: remove the large target detection head in the original algorithm, and add the ALCE adaptive local context enhancement module in front of the YOLOv8 medium target detection head; this module extracts the local context information around the target through a regular convolution and a dilated convolution with a dilation rate of 3; next, the spatial adaptive weights corresponding to each context are calculated and weighted to achieve adaptive context adjustment for different inputs.
[0012] Step 2: Train the drone image target detection algorithm based on the cross-layer single-head attention mechanism and context information enhancement;
[0013] Set the training parameters: input image size, batch size, impulse size, learning rate and maximum number of iterations; use VisDrone2019, UAVDT and aerial photography small target datasets to train the network; after training, the final drone image target detection model based on cross-layer single-head attention mechanism and contextual information enhancement is obtained.
[0014] In order to obtain stable detection performance and training speed of the model, the present invention needs to optimize and adjust the relevant parameters according to the characteristics of each data set to find the optimal weight;
[0015] Preferably, the present invention sets the input image resolution of the VisDrone2019 and UAVDT datasets to 960×960, and sets the input image resolution of the aerial small target dataset to 640×640, which can achieve a balance between efficiency and accuracy;
[0016] Considering the GPU memory capacity of the training model, the Batch Size of the present invention is set to 2;
[0017] The impulse of the present invention is set to 0.937;
[0018] The initial learning rate of the present invention is set to 0.01, and the final learning rate is set to 0.001;
[0019] According to the learning difficulty and size of the data set used, the present invention preferably sets the training epochs of the VisDrone2019 data set to 200, the training epochs of the UAVDT data set to 15, and the training epochs of the aerial photography small target data set to 150.
[0020] Step 3: Load the trained detection model and perform model evaluation;
[0021] The weight best.pt is obtained through training, and the performance of the model is evaluated using the VisDrone2019, UAVDT datasets, and aerial photography small target datasets.
[0022] Step 4: Input the drone aerial image to be tested into the final drone image target detection model based on cross-layer single-head attention mechanism and context information enhancement, and output the target detection result.
[0023] The beneficial effects of the present invention are as follows:
[0024] 1. In order to enhance the global context information, model long-distance dependencies, and enhance the fusion of cross-layer information, the present invention proposes a cross-layer single-head attention module to improve the detection accuracy of densely occluded small targets with lower computational redundancy.
[0025] 2. In order to improve the ability to extract multi-scale context information, the present invention proposes a multi-scale context extraction module, which improves the multi-scale context perception ability of the model with a small increase in computational complexity.
[0026] 3. In order to accurately obtain effective context, the present invention proposes an adaptive local context enhancement module to improve the model's context perception of a small area around the target and assist in the detection of densely occluded targets.
[0027] 4. Extensive experiments were conducted on the VisDrone2019 dataset to fully verify the effectiveness of the proposed algorithm in improving the accuracy of densely occluded target detection. At the same time, experiments on two other drone datasets further proved that the proposed algorithm has good adaptability and can maintain high performance in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 : The overall network structure diagram of the present invention
[0029] Figure 2 :Structure diagram of the cross-layer single-head attention module in the algorithm of this invention
[0030] Figure 3 :Structure diagram of the multi-scale context extraction module in the algorithm of the present invention
[0031] Figure 4 :Structure diagram of the adaptive local context enhancement module in the algorithm of the present invention
[0032] Figure 5 : A comparison of partial channel single-head attention, multi-head attention, and full channel single-head attention used in the algorithm of the present invention. Figure 5 (a) is multi-head attention, Figure 5 (b) is full-channel single-head attention, Figure 5 (c) Partial channel single-head attention used in the present invention.
[0033] Figure 6 :Overall steps flow chart
[0034] Figure 7 : Comparison chart of the results of the baseline algorithm and the algorithm of the present invention on the VisDrone2019 test set. The left half is the baseline algorithm, and the right half is the algorithm of the present invention. DETAILED DESCRIPTION
[0035] The present invention will now be further described with reference to the embodiments and the accompanying drawings:
[0036] The specific implementation is as follows:
[0037] Step 1: Construct a drone image target detection algorithm based on cross-layer single-head attention mechanism and context information enhancement;
[0038] Step 1-1: Add the CLSHA cross-layer single-head attention module before the small object detection head of the YOLOv8 object detection algorithm; this module combines the high-level feature map with the low-level feature map as the input of the attention mechanism, and only uses part of the channels for a single attention calculation;
[0039] Step 1-1-1: The overall structure of the cross-layer single-head attention module can be expressed as:
[0040] Y=CLSHA(X1,X2) (1)
[0041] Among them, X1 comes from the underlying feature map, and X2 comes from the feature map of the current layer. The two are processed as input through the cross-layer single-head attention module, and the final output is Y;
[0042] Step 1-1-2: The low-level feature map X1 is divided into two parts in the channel dimension: X att1 and X res1 , where only X att1 Is used for subsequent attention calculations:
[0043] X att1 ,X res1 =Split(X1,[C a ,CC a ]) (2)
[0044] Among them, X att1 is the feature map used for subsequent attention calculation, X res1 Remove X from X1 att1 The remaining feature map, C a Yes X att1 is the number of channels of X1, C is the number of channels of X1;
[0045] Step 1-1-3: The current layer feature map X2 adjusts the number of channels through the convolution operation to obtain X att2 , this feature map is used to participate in subsequent attention calculations:
[0046] X att2 =Conv(X2) (3)
[0047] Step 1-1-4: Place X att1 Perform linear transformation to obtain KeyK and Value V, and convert X att2 Perform linear transformation to obtain Query Q. K , W V , W Q There are three trainable parameter matrices:
[0048] K=X att1 W K
[0049] V=X att1 W V
[0050] Q=X att2 W Q (4)
[0051] Step 1-1-5: Perform attention calculation. First, scale the dot product result to get the attention score. qk is the dimension of the Key vector, is the scaling factor; then, the calculated attention score is normalized by the Softmax function to obtain the attention weight between each Query and all Keys, and the calculated attention weight is weighted summed with the Value V to obtain the attention output X att :
[0052]
[0053] Step 1-1-6: Output the obtained attention to X att The feature map X segmented from the low-level feature map res1 After concatenation in the channel dimension, the weight matrix W O Perform projection mapping; then, perform residual connection between the projection output and the current layer feature map X2 to obtain the intermediate output Y1:
[0054] Y1=X2+Concat(X att ,X res1 )W O (6)
[0055] Step 1-1-7: Perform a residual connection between the intermediate output Y1 and the output processed by the feed-forward neural network (FFN) to obtain the final output Y:
[0056] Y=FFN(Y1)+Y1 (7)
[0057] Step 1-2: Add the MSCE multi-scale context extraction module after the YOLOv8 backbone network;
[0058] First, the channel of the input feature map is divided into two parts: one part extracts local information through a 1×1 convolution layer, and then uses a set of parallel depth-separable convolutions of different scales to capture multi-scale context information. Subsequently, the context information of different scales is fused and the fused feature map is adjusted through a 1×1 convolution. The other part is directly spliced with the adjusted fused feature map through channels, and the final output of the entire module is obtained after channel splicing;
[0059] Its structure can be expressed as:
[0060]
[0061] Among them, X represents the input of the module, C represents the number of channels of the input feature map, L represents the output of the input feature map after passing through the depthwise separable convolution layer with a convolution kernel size of 3×3, Z represents the output after a series of parallel depthwise separable convolutions and fusion, and Y represents the final output of the entire module;
[0062] Step 1-3: Streamline the YOLOv8 algorithm detection head: remove the large target detection head in the original algorithm, and add the ALCE adaptive local context enhancement module before the medium target detection head of YOLOv8;
[0063] First, the feature maps that have undergone different convolutions are processed by 1×1 convolution, and then the channels of these feature maps are spliced to prepare for the subsequent spatial weight calculation.
[0064] F=Concat[Conv(X1,C,C,1),Conv(X2,C,C,1)] (9)
[0065] Among them, X1 and X2 are feature maps obtained by traditional convolution and dilated convolution respectively. The first C in Conv represents the number of input channels of the feature map, and the second C represents the number of output channels of the feature map. 1 means that the convolution kernel size is 1×1, and F is the feature map after channel splicing. Secondly, the spliced feature map is subjected to a 1×1 convolution operation to reduce the number of channels to the number of input feature maps (i.e., 2). The two channels of the obtained feature map correspond to two adaptive spatial weights respectively.
[0066] T=Conv(F,2×C,2,1) (10)
[0067] Among them, T represents the feature map with a channel number of 2 obtained after the 1×1 convolution operation.
[0068] Again, the obtained spatial weights are normalized through a Softmax operation, which normalizes all channel values at each spatial position and converts them into a probability distribution so that the sum of the activation values of all channels at each spatial position is 1.
[0069] W = softmax(T) (11)
[0070] Among them, W represents the spatial adaptive weight with the same size as the input feature map X and the number of channels as 2.
[0071] Finally, the feature maps with different receptive fields are multiplied by their corresponding spatial adaptive weights pixel by pixel, and then summed up to obtain the final feature map Y with an adaptive receptive field.
[0072]
[0073] Step 2: Train the drone image target detection algorithm based on the cross-layer single-head attention mechanism and context information enhancement, and adjust the training parameters to obtain the optimal weights;
[0074] Set the training parameters: input image size, batch size, impulse size, learning rate, and maximum number of iterations; use VisDrone2019, UAVDT datasets, and aerial photography small target datasets to train the network; after training, the final improved model is obtained;
[0075] In order to obtain stable detection performance and training speed of the model, the present invention needs to optimize and adjust the relevant parameters according to the characteristics of each data set of drone aerial photography to find the optimal weight;
[0076] (1) Set the input image size
[0077] The resolution of drone aerial images will affect the computational workload and memory consumption of the GPU hardware platform used in the training process. On the one hand, if the input resolution of the image is too small, some important information may be lost during training, resulting in a decrease in model performance. On the other hand, if the input resolution of the aerial image is too large during training, the computational workload and memory consumption may increase, resulting in slower model training and prediction speeds, or even exceeding hardware limitations. Preferably, according to the characteristics of drone aerial images and the hardware platform, the present invention sets the input image resolution of VisDrone2019 and UAVDT datasets to 960×960, and sets the input image resolution of the aerial small target dataset to 640×640, so as to achieve a balance between efficiency and accuracy.
[0078] (2) Set the training batch size
[0079] Batch size refers to the number of data samples simultaneously input into the model during a training process; the size of the batch size will affect the training effect and speed of the model, and it should not be too large or too small; if the batch size is too large, the gradient direction of each update is too smooth, and it may exceed the GPU memory capacity, resulting in overfitting of the training or decreased generalization ability; if the batch size is too small, the gradient direction of each update is unstable, and the parallel computing capability of the GPU cannot be fully utilized, resulting in non-convergence of training or very slow convergence speed; preferably, considering the GPU memory capacity of the training model, the batch size of the present invention is set to 2;
[0080] (3) Set the impulse size
[0081] Momentum is a technique for optimizing the gradient descent method. It can accelerate the convergence of the model, avoid falling into local optimality or saddle points, and improve the generalization ability of the model. The magnitude of the momentum determines the degree of influence of the last update direction on the current update direction. It is generally between 0 and 1, and the commonly used value is 0.9. Preferably, the momentum of the present invention is set to 0.937.
[0082] (4) Setting the learning rate
[0083] The learning rate is a hyperparameter that adjusts the parameter update step size at each iteration of the optimization algorithm. The size of the learning rate will affect the convergence speed and effect of the model. If the learning rate is too large, the loss function may oscillate or diverge. If the learning rate is too small, it may lead to slow convergence or fall into a local optimum. Preferably, the initial learning rate of the present invention is set to 0.01, and the final learning rate is set to 0.001.
[0084] (5) Set training rounds
[0085] A training epoch refers to a single training iteration of all batches in forward and backward propagation; an epoch means that each sample in the training dataset has the opportunity to update the internal model parameters; the size and quality of the training dataset determine how much information the model needs to learn, as well as the difficulty and effect of learning; preferably, the present invention sets the training epochs of the VisDrone2019 dataset to 200, the training epochs of the UAVDT dataset to 15, and the training epochs of the aerial photography small target dataset to 150 according to the learning difficulty and size of the dataset used.
[0086] Step 3: Load the trained detection model and perform model evaluation;
[0087] The weight best.pt is obtained through training, and the performance of the model is evaluated using the VisDrone2019, UAVDT datasets, and aerial small target datasets.
[0088] Step 4: Input the drone aerial image to be tested into the final drone image target detection model based on cross-layer single-head attention mechanism and context information enhancement, and output the target detection result. Specific embodiment:
[0090] 1. Experimental conditions
[0091] In order to test the performance of the drone image target detection algorithm model based on the cross-layer single-head attention mechanism and context information enhancement, the present invention uses the VisDrone2019 dataset for ablation experiments and comparative experiments, and uses the UAVDT dataset and the aerial photography small target dataset for comparative experiments. Both VisDrone2019 and UAVDT are publicly used drone image datasets with high frequency, which are suitable for studying target detection tasks based on drone images.
[0092] Experimental test environment: The system is Ubuntu 20.04, the CPU is Intel (R) Core (TM) i9-10940X CPU @ 3.30GHz, the memory is 48GB, the GPU processor is NVIDIA GeForce GTX 2080Ti, and the video memory is 11GB.
[0093] 2. Experimental content
[0094] The present invention first designs and optimizes the YOLOv8 algorithm from three different perspectives, namely, designing a cross-layer single-head attention module, an adaptive local context enhancement module, and a multi-scale context extraction module. The designed algorithm structure is as follows: Figure 1 shown.
[0095] In order to verify the performance of the algorithm proposed in this invention in terms of algorithm detection accuracy and computational efficiency, ablation experiments and comparative experiments of different algorithms were carried out on the VisDrone2019 dataset, and comparative experiments were carried out on the UAVDT dataset and the aerial photography small target dataset.
[0096] 3. Classification evaluation indicators
[0097] The detection quality evaluation index used in the present invention is the mean average precision (mAP). The basic concepts involved in the calculation of mAP are introduced below.
[0098] (1) Intersection over Union (IoU): Intersection over Union ratio. In object detection, the detection box and the true value box are both rectangles. The intersection of the two divided by their union is the IoU value. The IoU value represents the degree of overlap between the detection box and the true value box, and is used to measure the accuracy of positioning. When the overlap rate is greater than the set threshold, the positioning is considered accurate, otherwise it is considered incorrect. When the two boxes completely overlap, the IoU reaches the maximum value of 1.
[0099] (2) True Positive (TP): The number of samples that are correctly identified as positive samples, indicating those samples that are successfully detected.
[0100] (3) True Negatives (TN): The number of negative samples that are correctly identified as negative samples, indicating that the background is not misidentified as a target.
[0101] (4) False Positives (FP): The number of negative samples that are mistakenly identified as positive samples, indicating that the background is identified as the target, i.e., false detection.
[0102] (5) False Negatives (FN): The number of positive samples that are mistakenly identified as negative samples, indicating that the target is mistakenly classified as the background and not detected, that is, missed detection.
[0103] (6) Precision: Accuracy, also known as precision. It indicates the proportion of samples that are actually positive examples among those samples that are positive examples in the detection output results. The calculation formula is:
[0104]
[0105] (7) Recall: Recall rate, also known as recall rate. It indicates the proportion of positive examples that are actually identified to all positive examples. The calculation formula is:
[0106]
[0107] In VOC2010 and later, the average precision (AP) is the area enclosed by the PR (Precision-Recall) curve and the coordinate axis, which is used to measure the detection quality of a class of targets. The calculation formula is as follows:
[0108]
[0109] When detecting multiple categories of targets, averaging the AP of all categories can obtain the mean average precision (mAP) used to measure the quality of multi-category target detection. The calculation formula is as follows:
[0110]
[0111] Where: M represents the number of categories, i∈(1,M).
[0112] Another important evaluation index of the target detection algorithm is speed. Only when the speed meets the real-time requirements can it be applied to industry, which is of great significance for the real-time detection of drones. The commonly used measurement index is the number of frames processed per second (Frame Per Second, FPS), that is, the number of images that can be processed per second, which is defined as follows:
[0113]
[0114] Where: Tot represents the total time spent on detecting an image or video; FC refers to the number of image frames processed. Usually, the FPS obtained varies greatly depending on the hardware configuration used.
[0115] 4. Simulation test
[0116] In order to test and verify the detection performance of the algorithm proposed in this invention for drone images, experiments were carried out on the VisDrone2019 dataset, the UAVDT dataset, and the aerial photography small target dataset, and the actual detection results were visualized and analyzed.
[0117] (1) Ablation experiments on the VisDrone dataset
[0118] ① Analysis of detection results of cross-layer single-head attention module
[0119] The experimental results of recall rate and mAP of adding cross-layer single-head attention module to the baseline model are shown in the data of "+CLSHA" in Table 1. As can be seen from Table 1, after adding cross-layer single-head attention module, mAP and mAP 50 The mAP is improved by 1.0% and 1.3% respectively. S 、mAP M and mAP LThe results show that after adding the cross-layer single-head attention module, the recall rate of the model increased by 1.9%. In terms of detection efficiency, the experimental results of the detection efficiency of adding the cross-layer single-head attention module are shown in the data of "+CLSHA" in Table 2. The experimental results show that after adding this module, the number of model parameters increased from 11.2M to 11.3M, an increase of only 0.1M. In terms of computational complexity, it increased from 28.5GFLOPs to 29.6GFLOPs, an increase of 1.1GFLOPs. The number of detection frames per second decreased from 106 to 70. In order to analyze the role of partial channel attention calculation and cross-layer input in the CLSHA module, relevant experiments were carried out on the VisDrone2019 validation set, and the results are shown in Tables 3 and 4. Among them, "+A" means that the single-head partial channel attention calculation of the CLSHA module is changed to a single-head full channel attention calculation and then added to the baseline model, and "+B" means that the cross-layer feature map input is changed to a single-layer feature map input and then added to the baseline model.
[0120] Table 1 Recall rate and mAP ablation test results of each model module on the VisDrone2019 validation set
[0121]
[0122] Table 2 Ablation test results of detection efficiency of each module of the model on the VisDrone2019 validation set
[0123]
[0124] Table 3 Recall and mAP ablation test results of the CLSHA module and its variants on the VisDrone2019 validation set
[0125]
[0126] Table 4 Detection efficiency ablation experimental results of the CLSHA module and its variants on the VisDrone2019 validation set
[0127]
[0128] It can be seen from the data in Table 3 that the CLSHA module proposed in this invention outperforms the two variant modules in terms of recall rate and various mAP indicators. Compared with the single-head full-channel attention module, the recall rate of the CLSHA module is increased by 0.5%, and the mAP and mAP are 50 The mAP is improved by 0.8% and 1.3% respectively. S 、mAP M and mAP LCompared with the single-layer feature map input module, the recall rate of the CLSHA module is improved by 1.7%, mAP and mAP 50 Improved mAP by 1.5% and 2.2% respectively S 、mAP M and mAP L An increase of 1.2%, 1.7% and 2.3% respectively.
[0129] It can be seen from the data in Table 4 that in terms of detection efficiency, the CLSHA module proposed in the present invention is comparable to the single-layer feature map input variant module, but is better than the full-channel attention variant module. Compared with the full-channel attention variant module, the overall computational complexity of the CLSHA module is reduced by 0.2GFLOPs, indicating that the use of partial channel single-head attention calculation effectively reduces computational redundancy and improves operating efficiency.
[0130] ②Analysis of ablation experimental results of multi-scale context extraction module
[0131] The experimental results of recall rate and detection accuracy of adding the multi-scale context extraction module to the baseline model are shown in the data of "+MSCE" in Table 1. As can be seen from the table, after adding the multi-scale context extraction module, mAP and mAP 50 The mAP is improved by 0.9% and 1.5% respectively. S 、mAP M and mAP L They are increased by 1.2%, 0.6% and 1.5% respectively. After adding the MSCE module, the recall rate of the model is increased by 2.2%, which indicates that the missed detection rate of the model is reduced after adding the MSCE module.
[0132] The experimental results of detection efficiency are shown in the data of "+MSCE" in Table 2. After adding the MSCE module, the model parameters increased from 11.2M to 11.5M, an increase of 0.3M. In terms of computational complexity, it increased from 28.5GFLOPs to 28.8GFLOPs, an increase of 0.3GFLOPs. The number of detection frames per second decreased from 106 to 100. The number of parameters increased by only 0.3M and the computational complexity increased by only 0.3GFLOPs. Overall, the MSCE module has little impact on the detection efficiency of the entire model and has good detection performance.
[0133] ③Analysis of ablation experimental results of adaptive local context enhancement module
[0134] The experimental results of recall rate and detection accuracy of adding adaptive local context enhancement module to the baseline model are shown in the data of "+ALCE" in Table 1. The experimental results show that after adding ALCE module, mAP and mAP 50The mAP is improved by 0.9% and 1.0% respectively. S 、mAP M and mAP L They are improved by 0.9%, 0.8% and 2.4% respectively. In addition, after adding the ALCE module, the recall rate of the model is improved by 1.9%, which indicates that the missed detection rate of the model is reduced.
[0135] The experimental results of the detection efficiency of adding the ALCE module to the baseline model are shown in the data of "+ALCE" in Table 2. The experimental results show that this improvement reduces the number of model parameters from the original 11.2M to 8.9M, a reduction of 2.3M. In terms of computational complexity, it increases from the original 28.5GFLOPs to 30.0GFLOPs, an increase of 1.5GFLOPs. The number of detection frames per second decreases from the original 106 to 103. In addition, the improvement has little impact on computational complexity and detection speed, and has good detection efficiency.
[0136] (2) Comparative experiments were conducted on VisDrone2019, UAVDT dataset, and aerial photography small target dataset
[0137] The algorithm proposed in the present invention is compared with different algorithms on the VisDrone2019 dataset, the UAVDT dataset and the aerial photography small target dataset, including the YOLO series of target detection algorithms and other target detection algorithms for drone images: FCOS, YOLOv5, Swin Transformer, YOLOv7, YOLOv8, GCGE-YOLO, C3TB-YOLOv5, Ni's model, YOLOv11 and DTSSNet, etc.
[0138] Table 5 Detection results of the proposed model and other models on the VisDrone2019 validation set
[0139]
[0140] From the experimental results in Table 5, it can be seen that in terms of the average accuracy of detection, the model proposed by the present invention exceeds the compared classic target detection model and the advanced drone image target detection model in the past two years. In addition, the present model also performs quite well in terms of detection efficiency, and the model has small parameters and computational complexity. Compared with the drone image target detection model DTSSNet in 2024, the model of the present invention is significantly higher in detection accuracy than the model, while its parameters are slightly lower than DTSSNet, and the computational complexity is greatly reduced, showing a good balance between accuracy and efficiency.
[0141] The results obtained on the UAVDT dataset are shown in Table 6.
[0142] Table 6 Detection results of the model of the present invention and other models on UAVDT
[0143]
[0144] The comparison results of the UAVDT dataset are shown in Table 6. The comparison models in the table cover multiple classic target detection algorithms and advanced algorithms in recent years. Since the UAVDT official only provides undivided datasets, it is necessary to divide the training set and test set by yourself. In order to ensure the fairness of the experiment, the test results of the UAVDT dataset in this experiment are all derived from local experiments, and no experimental data from other literature is directly cited.
[0145] The data in Table 6 show that the proposed model is 6.4% higher than the baseline model YOLOv8-s in terms of mAP. 50 In terms of detection accuracy, the model proposed by the present invention outperforms all the comparison models and demonstrates its good detection performance. 50 They are 12.0% higher than FCOS, 7.9% higher than YOLOv5-s, 6.4% higher than YOLOv7-tiny, and 7.6% higher than YOLOv11-s. These results fully demonstrate the advantages of the model proposed in the target detection task. In addition, compared with the baseline model, this model successfully achieved a significant improvement in detection accuracy while reducing the number of parameters by 1.8M and increasing the amount of calculation by only 2.9GFLOPs, showing a good balance between accuracy and efficiency.
[0146] (3) Visual analysis of actual test results
[0147] In order to more intuitively compare the performance gap between the proposed model and the baseline model, the prediction results of the two models on the VisDrone2019 test set are visualized, as shown in the figure. Figure 7As shown, the left half is the detection result of the baseline model, and the right half is the detection result of the model of the present invention. At the same time, in order to make the visual effect better, the target category and confidence score in the detection result are removed and the area that can highlight the performance gap between the two models is enlarged and displayed. In the first scene, the present model successfully detected multiple densely arranged targets blocked by leaves, while the baseline model failed to detect them. In the second scene, the baseline model failed to detect the car that was partially blocked by leaves and monitoring poles, while the present model successfully detected it. In addition, the baseline model mistakenly detected the background part as two targets, while the present model did not have similar false detections. In the third scene, although the baseline model detected the car blocked by leaves, there was a case of repeated detection, while the present model not only detected the target detected by the baseline model, but also avoided repeated detection, and successfully detected a relatively hidden car. The above visualization results show that after targeted improvements, this model has improved the detection ability of targets occluded by background objects and densely arranged in UAV images, reduced the missed detection rate and false detection rate of densely occluded targets, and has a great advantage in improving the detection ability of densely occluded targets.
[0148] In summary, the algorithm of the present invention can improve the target detection accuracy in UAV scenarios and reduce the missed detection rate and false detection rate of densely occluded targets, and has good application potential.
Claims
1. A drone image target detection algorithm based on cross-layer single-head attention mechanism and context information enhancement, characterized in that: The following steps are involved: Step 1: Construct a drone image target detection algorithm based on cross-layer single-head attention mechanism and context information enhancement; Step 1-1: Add a CrossLayer Single Head Attention (CLSHA) module before the small object detection head of the YOLOv8 object detection algorithm; this module combines the high-level feature map with the low-level feature map as the input of the attention mechanism, and only uses part of the channels for a single attention calculation; Step 1-2: Add a Multi-Scale Context Extraction (MSCE) module after the YOLOv8 backbone network; first pass the input feature map through a 1×1 convolutional layer; Subsequently, the feature map is divided into two parts by channel: one part captures multi-scale contextual information through a set of parallel separable convolutions of different scales and depths, and then is fused, and then the fused feature map is adjusted through a 1×1 convolution; the other part is directly channel-joined with the adjusted fused feature map to obtain the final output of the entire module; Step 1-3: Streamline the YOLOv8 algorithm detection head: remove the large target detection head in the original algorithm, and add an Adaptive Local Context Enhancement (ALCE) module in front of the medium target detection head of YOLOv8; this module extracts local context information around the target through a regular convolution and a dilated convolution with a dilation rate of 3; next, the spatial adaptive weights corresponding to each context are calculated and weighted to achieve adaptive context adjustment for different inputs. Step 2: Train the drone image target detection algorithm based on the cross-layer single-head attention mechanism and context information enhancement; Set training parameters: input image size, batch size, impulse size, learning rate, and maximum number of iterations; The network is trained using VisDrone2019, UAVDT, and aerial photography small target datasets; After training, the final drone image target detection model based on cross-layer single-head attention mechanism and context information enhancement is obtained. Step 3: Load the trained detection model and perform model evaluation. Step 4: Input the drone aerial image to be tested into the final drone image target detection model based on cross-layer single-head attention mechanism and context information enhancement, and output the target detection result.
2. According to claim 1, a drone image target detection algorithm based on cross-layer single-head attention mechanism and context information enhancement is characterized in that: The cross-layer single-head attention module in step 1-1 includes the following steps: Step 1-1-1: The overall structure of the cross-layer single-head attention module can be expressed as: Y=CLSHA(X1,X2) (1) Among them, X1 represents the low-level feature map input, and X2 represents the current-level feature map input. The two are processed jointly through the cross-layer single-head attention module, and the final output is Y; Step 1-1-2: The low-level feature map X1 is divided into two parts in the channel dimension: X att1 and X res1 , where only X att1 Is used for subsequent attention calculations: X att1 ,X res1 =Split(X1,[C a ,C-C a ]) (2) Among them, X att1 is the feature map used for subsequent attention calculation, X res1 Remove X from X1 att1 The remaining feature map, C a Yes X att1 is the number of channels of X1, C is the number of channels of X1; (1) Step 1-1-3: The current layer feature map X2 adjusts the number of channels through the convolution operation to obtain X att2 , this feature map is used for subsequent attention calculations: X att2 =Conv(X2) (3) Step 1-1-4: Place X att1 Perform linear transformation to obtain KeyK and Value V; att2 Perform a linear transformation and get Query Q: K=X att1 W K V=X att1 W V Q=X att2 W Q (4) Among them, W K , W V , W Q are three trainable parameter matrices; Step 1-1-5: Perform attention calculation; first, scale the dot product result to get the attention score; second, normalize the calculated attention score through the Softmax function to get the attention weight between each Query and all Keys; finally, perform weighted summation of the calculated attention weight and Value V to get the output X of the attention mechanism. att : X att =Attention(Q,K,V) Among them, d qk is the dimension of the Key vector, is the scaling factor; Step 1-1-6: Output the obtained attention mechanism X att The feature map X segmented from the low-level feature map res1 , after splicing in the channel dimension, and then through the weight matrix W O Perform projection mapping; then, perform residual connection between the projection output and the current layer feature map X2 to obtain the intermediate output Y1: Y1=X2+Concat(X att ,X res1 )W O (6) Step 1-1-7: Perform a residual connection between the intermediate output Y1 and the output processed by the feed-forward neural network (FFN) to obtain the final output Y: Y=FFN(Y1)+Y1 (7) 3. The drone image target detection algorithm based on cross-layer single-head attention mechanism and context information enhancement according to claim 1 is characterized in that: The training parameters are set as follows: the input size of the VisDrone2019 and UAVDT datasets is 960×960, and that of the aerial small target dataset is 640×640, the batch size is 2, the impulse size is 0.937, the initial learning rate is 0.01, the final learning rate is 0.001, the maximum number of iterations of the VisDrone2019 dataset is 200, the maximum number of iterations of the UAVDT dataset is 15, and the maximum number of iterations of the aerial small target dataset is 150.