Unmanned aerial vehicle small target detection model establishment and detection method based on YOLO11s improvement

By improving the YOLO11s model, deleting the downsampling layer and introducing a variable simple parameterless attention mechanism and a space-to-depth residual convolution module, the problem of accuracy loss in the detection of small targets of drones is solved, and the detection capability and adaptability of the model in complex backgrounds is improved.

CN120298937AActive Publication Date: 2025-07-11QUANZHOU INST OF EQUIP MFG +1

Patent Information

Application Number
CN202510789238.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

The existing lightweight convolutional neural network model has accuracy loss in the detection of small targets by UAVs, especially in complex backgrounds and low-resolution image conditions, and lacks targeted small target feature retention design.

Method used

The YOLO11s model is improved, by deleting the last downsampling fusion layer, replacing the C3k2 module in the backbone network as a variable simple parameterless attention mechanism module, and using the space-to-depth residual convolution module and the content-aware feature recombination upsampling module to enhance the retention and detection capabilities of small-target features.

Benefits of technology

It significantly improves the accuracy and adaptability of small-object detection of drones, reduces model complexity, while maintaining real-time and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298937A_ABST
    Figure CN120298937A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision, in particular to an improved unmanned aerial vehicle small target detection model establishment and detection method based on YOLO11s, a YOLO11s-UAV model is established, the YOLO11s-UAV model is an improved YOLO11s model, the last down-sampling fusion layer in a backbone network of the YOLO11s model is deleted, the down-sampling fusion layer comprises a Conv module and a C3k2 module, and the Conv module is connected with the C3k2 module. Replacing the residual C3k2 module in the backbone network of the YOLO11s model and the C3k2 module in the neck network of the YOLO11s model with a variable simple parameter-free attention mechanism module, and replacing the residual Conv module in the backbone network of the YOLO11s model with a space-to-depth residual convolution module; replacing an Upsample module in a neck network of the YOLO11s model with a content awareness feature recombination up-sampling module; the last downsampling fusion layer in the backbone network of the YOLO11s model is deleted, so that the whole network retains more fine feature information of small targets, and meanwhile, the complexity of the model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a detection method for establishing a small target detection model of an unmanned aerial vehicle (UAV) improved based on YOLO11s. Background Art

[0002] As a mobile embedded platform, the computing unit of the UAV has strict power consumption and computing power limitations. Existing lightweight solutions, such as dynamic convolution and depthwise separable convolution, although they can reduce the number of model parameters or the amount of computation by more than 20%, generally show a significant decline in various accuracy indicators, forming a vicious cycle of lightweight and accuracy loss.

[0003] Existing models generally lack the design for retaining small target features specifically. Firstly, the downsampling process of traditional convolutional neural networks will continuously lose the fine-grained features of small targets. Secondly, mainstream object detection architectures often do not deploy a dedicated small target attention mechanism, resulting in the model ignoring the extraction of the saliency features of small targets in the shallow network.

[0004] Single-stage models, such as the YOLO series, with their concise structure and fast inference speed, can meet the real-time requirements of various application scenarios while maintaining high accuracy. However, in the UAV scenario, due to problems such as complex backgrounds, small target pixel ratios, and low-resolution images, the detection ability of these general single-stage models for small targets is still insufficient. Summary of the Invention

[0005] The purpose of the present invention is to provide a detection method for establishing a small target detection model of an unmanned aerial vehicle (UAV) improved based on YOLO11s, which can reduce the loss of feature information.

[0006] To achieve the above purpose, the present invention adopts the following technical solution: A method for establishing a small target detection model of an unmanned aerial vehicle (UAV) improved based on YOLO11s, which establishes a YOLO11s-UAV model. The YOLO11s-UAV model is an improved YOLO11s model, and the YOLO11s-UAV model includes a backbone network for feature extraction of the input image, a neck network for feature extraction and fusion of the feature map, and a head structure for detecting and classifying the fused feature map output by the neck network; Delete the last downsampling fusion layer in the backbone network of the YOLO11s model. The downsampling fusion layer includes a Conv module and a C3k2 module. Replace the remaining C3k2 modules in the backbone network of the YOLO11s model and the C3k2 modules in the neck network of the YOLO11s model with variable simple parameterless attention mechanism modules, and replace the remaining Conv modules in the original backbone network with spatial-to-depth residual convolution modules; Replace the Upsample module in the neck network of the YOLO11s model with a content-aware feature recombination upsampling module; A first convolution module, a first splicing module, and a variable simple parameterless attention mechanism module are sequentially arranged along the data transmission direction between the C2PSA module of the backbone network of the YOLO11s model and the first content-aware feature recombination upsampling module of the neck network.

[0007] Preferably, the YOLO11s-UAV model further includes a second convolution module, a third convolution module, and a fourth convolution module; The second convolution module is used to perform a convolution operation on the fused features output by the first variable simple parameterless attention mechanism module of the backbone network and then deliver them to the splicing module after the first content-aware feature recombination upsampling module of the neck network; The third convolution module is used to perform a convolution operation on the fused features output by the second variable simple parameterless attention mechanism module of the backbone network and then deliver them to the first splicing module of the neck network; The fourth convolution module is used to perform a convolution operation on the fused features output by the first variable simple parameterless attention mechanism module after the content-aware recombination and interactive feature pyramid network of the neck network and then deliver them to the last splicing module of the neck network.

[0008] Preferably, the spatial-to-depth residual convolution module includes a spatial-depth conversion module that divides the input feature map into multiple pixel blocks and rearranges them into the depth dimension, a non-strided convolution module that performs an operation to convert the number of channels of the feature map of the pixel blocks, and an atrous residual module that extracts multi-scale context information from the feature map output by the non-strided convolution module, and performs a combined batch normalization process on the feature map output by the non-strided convolution module and the feature map output by the atrous residual module, and applies an activation function to output a feature map.

[0009] Preferably, the variable simple parameterless attention mechanism module sequentially performs a channel dimension transformation operation and a feature segmentation operation on the input feature map to obtain segmentation features. The segmentation features are path-selected according to the value of the boolean parameter c3k. When the value of the boolean parameter c3k is true, the segmentation features are processed through a branch path composed of multiple C3kSimAM modules. When the value of the boolean parameter c3k is false, the segmentation features are processed through multiple Bottleneck modules, and the segmentation features are subjected to residual connection and fusion with the feature maps output by the corresponding branch paths, and then the fused feature maps are subjected to feature fusion and dimension reduction through a convolutional layer and then output.

[0010] Preferably, replace the Bottleneck module in the C3k module with the BottleneckSimAM module to obtain the C3kSimAM module.

[0011] Preferably, the content-aware feature recombination upsampling module includes a kernel prediction module and a content-aware recombination module. The kernel prediction module sequentially performs channel compression, content encoding, and generates a reassembled kernel on the input feature map, and applies an activation function to the kernel to output a predicted kernel. The content-aware recombination module uses a weighted sum operator to recombine the predicted kernel and outputs a recombined kernel.

[0012] A method for detecting small targets of an unmanned aerial vehicle includes the following steps executed in sequence: S1: Obtain the pictures taken by the unmanned aerial vehicle, perform normalization target box coordinate annotation and category label annotation on the taken pictures, and divide the annotated pictures into a training set, a validation set, and a test set; S2: Input the training set into the YOLO11s-UAV model established by the method for establishing a small target detection model of an unmanned aerial vehicle improved based on YOLO11s as described in any one of the above for training; S3: Input the test set into the trained YOLO11s-UAV model for detection and output the detection result.

[0013] By adopting the foregoing design scheme, the beneficial effects of the present invention are: In view of the task characteristics of small target detection, delete the last downsampling fusion layer in the original backbone network, so that the entire network retains more fine feature information of small targets, and at the same time reduces the complexity of the model; Replace the C3k2 module in the network with a variable simple parameter-free attention mechanism module, which can not only further reduce the model complexity, but also significantly expand the perception range of the shallow network for small target features by using real three-dimensional attention weights; Replace the remaining Conv modules in the original backbone network with spatial-to-depth residual convolution modules. This design eliminates the information loss caused by the operation of traditional stride convolution reducing the spatial dimension of the feature map, and at the same time greatly reduces the difficulty of the network capturing multi-scale context information, providing a new technical solution for the effective retention of small target features. Description of the Drawings

[0014] Figure 1 It is the network architecture diagram of the YOLO11s-UAV model of the present invention; Figure 2 It is the variable simple parameter-free attention mechanism module of the present invention; Figure 3 It is the processing flow chart of the spatial-to-depth residual convolution module of the present invention for pictures; Figure 4 This is the processing flow chart of the content-aware feature recombination upsampling module of the present invention for pictures; Figure 5 This is the heat map of the shallow features of the comparison network before and after adopting the variable simple parameterless attention mechanism module of the present invention; Figure 6 This is the detection effect diagram under the night and exposure scenes of the present invention; Figure 7 This is the scatter plot comparing the spatial-to-depth residual convolution module of the present invention with other downsampling modules. Specific implementation manners

[0015] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0016] The terms "first", "second", "third", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0017] A method for establishing an improved UAV small target detection model based on YOLO11s, establishing a YOLO11s-UAV model as Figure 1 shown. The YOLO11s-UAV model is an improved YOLO11s model. The YOLO11s-UAV model includes a backbone network for feature extraction of the input image, a neck network for feature extraction and fusion of the feature map, and a head structure for detecting and classifying the fused feature map output by the neck network.

[0018] Delete the last downsampling fusion layer in the backbone network of the YOLO11s model. This downsampling fusion layer includes a Conv module and a C3k2 module. On this basis, according to the task characteristics of small target detection, unnecessary downsampling fusion layers in the backbone network are removed, so that the entire network retains more fine feature information of small targets and further reduces the model complexity.

[0019] Replace both the remaining C3k2 modules in the backbone network of the YOLO11s model and the C3k2 modules in the neck network of the YOLO11s model with a variable simple parameter-free attention mechanism module (FlexSimAM). This variable simple parameter-free attention mechanism module combines the C3k2 module with a simple parameter-free attention module (SimAM), that is, replaces the Bottleneck module in the C3k module with a BottleneckSimAM module to obtain a C3kSimAM module. The specific structure is as Figure 2 shown.

[0020] The simple parameter-free attention mechanism module is a lightweight attention mechanism. Its core lies in using the local self-similarity calculation of the feature map itself to efficiently generate real three-dimensional attention weights without introducing additional training parameters. This feature makes it particularly suitable for the small target detection task of drones with limited airborne resources. The core innovation of the variable simple parameter-free attention mechanism module lies in its unique dynamic module switching mechanism, which flexibly switches between the bottleneck module (Bottleneck) and the C3kSimAM module in the model through the c3k boolean value.

[0021] The variable simple parameter-free attention mechanism module sequentially performs a channel dimension transformation operation on the input feature map through a 1×1 Conv module and a feature segmentation operation through a feature segmentation module (Split) to obtain segmented features. These segmented features make a path selection according to the value of the boolean parameter c3k. When the value of the boolean parameter c3k is true, the segmented features are processed through a branch path composed of multiple C3kSimAM modules. When the value of the boolean parameter c3k is false, the segmented features are processed through multiple Bottleneck modules, and the segmented features are connected and fused with the feature maps output by the corresponding branch paths. Then, the fused feature maps are passed through a 1×1 Conv module to complete feature fusion and dimension reduction before output.

[0022] The variable simple parameter-free attention mechanism module has the following three significant advantages: First, according to the image features of the dataset to be detected, the number configuration of the C3kSimAM modules can be manually adjusted. This quantity-driven adaptive mechanism greatly enhances the adaptability and generalization ability of the YOLO11s-UAV model of this application in different scenarios; Second, as Figure 5 shown, aiming at the difficulties of small target detection in drone aerial images, this mechanism enhances the sensitivity of the shallow network to the feature extraction of small targets by generating real three-dimensional attention weights, significantly improving the recognition rate of small targets; Third, the lightweight design of this mechanism enables the YOLO11s-UAV model to effectively control the model complexity while ensuring the detection accuracy.

[0023] Table 1 shows multiple groups of tests on the method proposed in this application using the variable simple parameterless attention mechanism module (FlexSimAM) with the VisDrone-DET2019 dataset, for demonstrating the influence of the c3k boolean parameter in the variable simple parameterless attention mechanism module at each layer.

[0024] Table 1 Influence of the c3k boolean parameter in the variable simple parameterless attention mechanism module (FlexSimAM) proposed in this application at each layer (VisDrone-DET2019 dataset)

[0025] As can be seen from the above table, for the VisDrone-DET2019 dataset, three groups of comparative experiments were conducted on the setting of the c3k parameter boolean value (true: √; false: ×) of the FlexSimAM module in different layers of the method in this application. The baseline architecture means retaining all the original C3k2 modules in the corresponding layer of the method in this application. The experimental results show that Experiment 3 achieved the optimal detection accuracy. Therefore, for detecting UAV datasets in different scenarios, the model can achieve better detection performance by dynamically adjusting the c3k parameter boolean value in FlexSimAM, and at the same time, the model can be made somewhat lightweight.

[0026] Replace the remaining Conv modules in the backbone network of the YOLO11s model with the space-to-depth residual convolution module (S2DResConv). The space-to-depth residual convolution module of this application is designed based on the space-depth conversion module (SPD) and combined with the dilated residual module (DWR), as Figure 3 shown. The space-to-depth residual convolution module includes a space-depth conversion module that divides the input feature map into multiple pixel blocks and rearranges them to the depth dimension, a non-strided convolution module that performs an operation to convert the number of channels of the feature map of the pixel blocks, and a dilated residual module that extracts multi-scale context information from the feature map output by the non-strided convolution module, and combines and batch-normalizes the feature map output by the non-strided convolution module and the feature map output by the dilated residual module, and applies the Gaussian error linear unit (GELU) activation function to output the feature map; among them, the normalization process makes the activation values smoother, ensures stable output and is conducive to gradient flow, and the smooth non-linear transformation provided by the GELU activation function not only avoids the neuron death problem related to the ReLU activation function but also enhances the expression ability of the entire model.

[0027] In this embodiment, it is assumed that the feature map M processed by the spatial depth conversion module belongs to (S, S, C1), and is divided into several pixel blocks with a stride of 2, and these feature blocks are combined into four sub-blocks along the spatial dimension , , , , as shown in formula (1): (1); The size of each sub-block is (S / 2, S / 2, C1). After the spatial-to-depth conversion, the number of channels increases to 4C1, and the length and width are halved. Finally, a feature map with a size of (S / 2, S / 2, 4C1) is obtained.

[0028] The non-stride convolution module applies an operation to convert the number of channels of the feature map output by the spatial depth conversion module, adjusting the channel dimension of the feature map to the specified value C2. This non-stride convolution module does not perform any operation to reduce the spatial dimension of the feature map, avoiding the loss of target feature information.

[0029] The dilated residual module uses a two-step method to extract multi-scale context information: one is regional residualization (RR) to generate residual features from the input features, and the other is semantic residualization (SR) to apply multi-rate dilated depth convolution to perform morphological filtering on the features of different-sized regions. Finally, the features are merged through pointwise convolution to form a residual, which is then added to the input feature map to create a more comprehensive feature representation.

[0030] As Figure 7 shown, the VisDrone-DET2019 dataset is used to test the method of this application using the spatial-to-depth residual convolution module, and comparative experiments are conducted using other conventional general convolution modules. This scatter plot intuitively shows that the S2DResConv downsampling module proposed by the method of this application demonstrates significant advantages in the small target detection task of drones for the VisDrone-DET2019 dataset: it is comprehensively superior to other comparative downsampling modules in terms of detection accuracy indicators, and at the same time only introduces a relatively small number of model parameters, which reflects its excellent balance between detection accuracy and model efficiency.

[0031] Replace the Upsample module in the neck network of the YOLO11s model with a content-aware feature recombination upsampling module (CARAFE); this content-aware feature recombination upsampling module demonstrates significant advantages in small target detection tasks through large receptive field context aggregation and dynamic kernel generation mechanisms.

[0032] The content-aware feature reorganization upsampling module includes a kernel prediction module and a content-aware reorganization module. The kernel prediction module sequentially performs channel compression, content encoding, and generation of a reassembled kernel on the input feature map, and applies an activation function to the kernel to output a predicted kernel. The content-aware reorganization module uses a weighted sum operator to reorganize the predicted kernel and outputs the reorganized kernel.

[0033] In this embodiment, as Figure 4 shown, the working process of the content-aware feature reorganization upsampling module is as follows: Given an input feature map X ∈ (C × H × W) and an upsampling factor σ, by implementing the kernel prediction module and the content-aware reorganization module, the generation process of the output feature map ∈ (C × σH × σW) is realized, and the specific implementation includes: The kernel prediction module consists of three sub-modules, namely a channel compressor, a content encoder, and a kernel normalizer. They respectively perform channel compression, content encoding, and generation of a reassembled kernel on the input feature map, and apply the softmax function to this kernel. In the k×k neighborhood centered on the feature , this kernel prediction module predicts its position-related kernel for each target position : ; Among them, represents the kernel prediction module, represents the kernel size of the content encoder, The content-aware reorganization module reassembles the predicted kernel in the local area to obtain the output feature map : ; Among them, represents the operator of the content-aware reorganization module, represents the size of the reorganized kernel, The content-aware reorganization module uses a simple weighted sum operator to perform reorganization, which can pay more attention to the information provided by the relevant points in the local area, and the semantics of the reorganized feature map may be stronger than that of the original feature map.

[0034] A first convolution module, a first splicing module, and a variable simple parameter-free attention mechanism module are sequentially arranged in the data transmission direction between the C2PSA module of the backbone network of the YOLO11s model and the first content-aware feature reorganization upsampling module of the neck network of the YOLO11s model. In this embodiment, the first convolution module is a 1×1 Conv module.

[0035] The YOLO11s-UAV model further includes a second convolutional module, a third convolutional module, and a fourth convolutional module; and the second convolutional module, the third convolutional module, and the fourth convolutional module are all 3×3 Conv modules; The second convolutional module is used to perform a convolutional operation on the fused features output by the first variable simple parameter-free attention mechanism module of the backbone network and then send them to the splicing module after the first content-aware feature recombination upsampling module of the neck network; The third convolutional module is used to perform a convolutional operation on the fused features output by the second variable simple parameter-free attention mechanism module of the backbone network and then send them to the first splicing module of the neck network; The fourth convolutional module is used to perform a convolutional operation on the fused features output by the first variable simple parameter-free attention mechanism module after the content-aware recombination and interactive feature pyramid network of the neck network and then send them to the last splicing module of the neck network; The fused features output by the first variable simple parameter-free attention mechanism module of the backbone network are sent to the splicing module after the second content-aware feature recombination upsampling module (CARAFE) of the neck network, and the fused features output by the second variable simple parameter-free attention mechanism module of the backbone network are sent to the splicing module after the first content-aware feature recombination upsampling module (CARAFE) of the neck network; The fused features output by the variable simple parameter-free attention mechanism module between the C2PSA module and the first content-aware feature recombination upsampling module of the neck network of the YOLO11s model are sent to the last splicing module of the neck network, and the fused features output by the variable simple parameter-free attention mechanism module before the second content-aware feature recombination upsampling module (CARAFE) of the neck network are sent to the penultimate splicing module of the neck network.

[0036] Introducing the first convolutional module, the first splicing module, and the variable simple parameter-free attention mechanism module, as well as the settings of the second convolutional module, the third convolutional module, and the fourth convolutional module between the backbone network and the neck network is equivalent to introducing a skip connection layer in the YOLO11s-UAV model and adaptively aligning the channel dimension differences of the multi-scale feature maps, significantly improving the fusion efficiency of feature information at different scales in the same depth network layer.

[0037] Table 2 shows multiple groups of tests on the method proposed in this application using different upsampling modules with the VisDrone-DET2019 dataset to compare the content-aware feature recombination upsampling module (CARAFE) with other upsampling modules. Among them, Nearest and Bilinear represent different upsampling methods in the Upsample module of the original YOLO11s model.

[0038] Table 2 Content-Aware Reassembly and Upsampling Module (CARAFE) introduced by the method of this application Comparison experiment with other upsampling modules (VisDrone-DET2019 dataset)

[0039] As can be seen from the above table, six groups of replacement tests were carried out on the upsampling module in the method of this application. The results show that, compared with other upsampling modules, the Content-Aware Reassembly and Upsampling Module has achieved better detection performance improvement on the premise of introducing a small number of model parameters and computational costs.

[0040] This embodiment also provides a detection method for performing detection in the YOLO11s-UAV model established by the above method.

[0041] A method for detecting small targets of unmanned aerial vehicles, comprising the following steps executed in sequence: S1: Obtain the pictures taken by the unmanned aerial vehicle, perform normalized annotation of the target box coordinates and category labels on the taken pictures, and divide the annotated pictures into a training set, a validation set, and a test set; S2: Input the training set into the YOLO11s-UAV model established by the method for establishing a small target detection model of unmanned aerial vehicles improved based on YOLO11s described in any one of the above for training; S3: Input the test set into the trained YOLO11s-UAV model for detection, and output the detection results, as Figure 6 shown.

[0042] Table 3 shows the ablation experiment on the three methods proposed in this application using the VisDrone-DET2019 dataset, where the newly introduced skip connection layer and the replaced Content-Aware Reassembly and Upsampling Module (CARAFE) together represent the improvement of the feature pyramid network structure of the original YOLO11s model into the Content-Aware Reassembly and Interaction Feature Pyramid Network Structure (CARIFPN).

[0043] Table 3 Ablation experiment of the method of this application (VisDrone-DET2019 dataset)

[0044] As can be seen from the above table, the ablation experiment of the method of this application verified the effectiveness of the improvement superposition by gradually integrating three improvements: the Content-Aware Reassembly and Interaction Feature Pyramid Network Structure (CARIFPN), the Variable Simple Parameter-Free Attention Mechanism Module (FlexSimAM), and the Spatial-to-Depth Residual Convolution Module (S2DResConv) on the baseline model YOLO11s.

[0045] Table 4 Comparative experiment of mAP@0.5 index between the method of this application and the baseline method (VisDrone-DET2019 dataset)

[0046] As can be seen from the above table, compared with the baseline model YOLO11s, the method of this application has achieved significant improvements on both the validation set and the test set of the VisDrone-DET2019 dataset. Among them, the key index of mAP@0.5 has increased by 7.8% on the validation set and 5.9% on the test set. This result fully demonstrates that this method is not only effective for the validation set, but also shows excellent detection performance for unknown test set data.

[0047] Table 5 Generalization experiment on other UAV aerial photography datasets

[0048] As can be seen from the above table, compared with the baseline model YOLO11s, the method of this application has shown significant advantages on the two typical UAV shooting datasets of TinyPerson and UAVDT-DET. In the validation set part, mAP@0.5 of all categories has increased by 7.2% and 3.0% respectively. This result shows that this method can achieve obvious improvements in the detection accuracy of small targets in different UAV scenarios, showing excellent dataset generalization performance. The specific embodiments described above have further detailed the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for establishing a small target detection model for drones improved based on YOLO11s, characterized in that: Establish a YOLO11s-UAV model, which is an improved YOLO11s model. The YOLO11s-UAV model includes a backbone network for extracting features from an input image, a neck network for extracting and fusing features from a feature map, and a head structure for detecting and classifying a fused feature map output by the neck network; Delete the last downsampling fusion layer in the backbone network of the YOLO11s model, which includes the Conv module and the C3k2 module. Replace the remaining C3k2 modules in the backbone network of the YOLO11s model and the C3k2 modules in the neck network of the YOLO11s model with variable simple parameter-free attention mechanism modules. Replace the remaining Conv modules in the original backbone network with spatial-to-depth residual convolution modules. Replace the Upsample module in the neck network of the YOLO11s model with a content-aware feature reorganization upsampling module; The first convolution module, the first splicing module and the variable simple parameter-free attention mechanism module are arranged in sequence between the C2PSA module of the backbone network of the YOLO11s model and the first content-aware feature reorganization upsampling module of the neck network along the data transmission direction.

2. The method for establishing a small target detection model of an unmanned aerial vehicle improved based on YOLO11s according to claim 1, characterized in that: The YOLO11s-UAV model also includes a second convolution module, a third convolution module, and a fourth convolution module; The second convolution module is used to perform a convolution operation on the fusion features output by the first variable simple parameter-free attention mechanism module of the backbone network and transmit the convolution operation to the splicing module after the first content-aware feature reorganization upsampling module of the neck network; The third convolution module is used to perform a convolution operation on the fusion features output by the second variable simple parameter-free attention mechanism module of the backbone network and then transmit the convolution operation to the first splicing module of the neck network; The fourth convolution module is used to perform a convolution operation on the fused features output by the first variable simple parameter-free attention mechanism module after the content-aware reorganization of the neck network and the interactive feature pyramid network, and then transmit them to the last splicing module of the neck network.

3. The method for establishing a small target detection model of an unmanned aerial vehicle improved based on YOLO11s according to claim 2, wherein: The space-to-depth residual convolution module includes a space-to-depth conversion module that divides the input feature map into multiple pixel blocks and rearranges them to the depth dimension, a non-stride convolution module that performs feature map channel number conversion operations on the pixel blocks, and a dilated residual module that extracts multi-scale context information from the feature map output by the non-stride convolution module, as well as merging and batch normalizing the feature map output by the non-stride convolution module and the feature map output by the dilated residual module, and applying an activation function to output the feature map.

4. The method for establishing a small target detection model of an unmanned aerial vehicle improved based on YOLO11s according to claim 3, characterized in that: The variable simple parameter-free attention mechanism module performs a channel dimension transformation operation and a feature segmentation operation on the input feature map in sequence to obtain segmented features. The segmented features perform path selection according to the value of the boolean parameter c3k. When the value of the boolean parameter c3k is true, the segmented features are processed through a branch path composed of multiple C3kSimAM modules. When the value of the boolean parameter c3k is false, the segmented features are processed through multiple Bottleneck modules, and the segmented features are subjected to residual connection and fusion with the feature maps output by the corresponding branch paths. Then, the fused feature maps are subjected to feature fusion and dimension reduction through a convolutional layer and then output.

5. The method for establishing a small target detection model of an unmanned aerial vehicle improved based on YOLO11s according to claim 4, characterized in that: Replace the Bottleneck module in the C3k module with a BottleneckSimAM module to obtain the C3kSimAM module.

6. The method for establishing a small target detection model of an unmanned aerial vehicle improved based on YOLO11s according to claim 5, characterized in that: The content-aware feature recombination upsampling module includes a kernel prediction module and a content-aware recombination module. The kernel prediction module sequentially performs channel compression, content encoding, and generation of a reassembled kernel on the input feature map, and applies an activation function to the kernel to output a predicted kernel. The content-aware recombination module uses a weighted sum operator to recombine the predicted kernel and outputs a recombined kernel.

7. A method for detecting small targets of an unmanned aerial vehicle, characterized in that: Including the following steps executed in sequence: S1: Obtain the pictures taken by the drone, perform normalized target box coordinate annotation and category label annotation on the taken pictures, and divide the annotated pictures into a training set, a validation set, and a test set; S2: Input the training set into the YOLO11s-UAV model established by the method for establishing an improved small target detection model of a drone based on YOLO11s described in any one of claims 1-6 above for training; S3: Input the test set into the trained YOLO11s-UAV model for detection and output the detection results.

Citation Information

Patent Citations

  • Power grid transmission line smoke detection method and system, electronic equipment and medium

    CN117036361A

  • Light unmanned aerial vehicle aerial image small target detection method and system based on TFA-YOLO11

    CN119810409A

  • Road crack detection method, medium and product

    US20250174019A1

Cited By

  • Rice and crab target detection device and method for rice field complex scene pictures and training method of target detection network

    CN121459392A

  • Improved unmanned aerial vehicle small target detection method based on YOLO26s

    CN121904640A

  • Unmanned aerial vehicle small target detection method based on improved YOLO26s

    CN121904640B