Lightweight Aerial Image Object Detection Method Based on Improved YOLOv11
By improving the YOLOv11 model, using the DBRDown module and the MSFLB module, a lightweight aerial image detection model is built, which solves the problem of target detection accuracy and efficiency on resource-constrained devices, and achieves more efficient and accurate detection effects.
Patent Information
- Application Number
- CN202510302342.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-14
AI Technical Summary
When a single-stage aerial image detection model runs on resource-constrained edge devices, it is difficult to achieve efficient and accurate object detection under the limitations of computing power and storage capacity.
Based on the improved YOLOv11 lightweight aerial image object detection method, by replacing the Conv module in the backbone network with the DBRDown module and adding the MSFLB module to the neck network, a lightweight aerial image detection model is constructed.
The detection accuracy of aerial image detection models on edge devices is improved, the calculation complexity and parameter quantity is reduced, the detection accuracy of small targets is significantly improved, and the regression quality of bounding boxes is improved by optimizing the loss function.
Smart Images

Figure CN119832219B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV aerial image data processing, and specifically to a lightweight aerial image target detection method based on improved YOLOv11. Background Art
[0002] With the rapid development of computer technology and UAV technology, the number of aerial images has been continuously increasing, and the application scope has also been expanding day by day. Nowadays, aerial images have been widely used in many fields such as ecological protection and monitoring, traffic management, disaster prevention and control, emergency rescue, and smart cities. In the target detection task, aerial image detection technology plays a crucial role, aiming to accurately identify, locate, and classify various interested targets, such as vehicles, pedestrians, buildings, etc., from complex aerial images.
[0003] With the vigorous development of deep learning technology, target detection has gradually shifted from traditional methods that rely on manually designed features to an automated mode based on deep neural networks. These deep learning-based algorithms not only significantly improve the detection accuracy and processing efficiency but also show stronger generalization ability. Currently, target detection algorithms are mainly divided into two categories: two-stage detection algorithms and one-stage detection algorithms. Two-stage algorithms (such as Mask R-CNN) have attracted much attention due to their excellent detection accuracy, but their complex network structure and high computational cost limit their application in real-time scenarios. In contrast, one-stage algorithms (such as YOLOv8, RetinaNet, and NanoDet) have been widely used in real-time scenarios and resource-constrained embedded systems due to their simple and efficient design, which significantly improves the detection speed while maintaining high detection performance.
[0004] However, one-stage aerial image detection models usually need to run on resource-constrained edge devices, which poses severe challenges to computing power and storage capacity. To achieve efficient and accurate target detection in these restricted environments, it is urgent to develop more lightweight detection algorithms to meet the performance requirements in practical applications. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a lightweight aerial image target detection method based on improved YOLOv11, so that the constructed aerial image detection model can effectively improve the accuracy of aerial image target detection under the condition of limited resources of edge devices.
[0006] The technical solution of the present invention is as follows:
[0007] A lightweight aerial image target detection method based on improved YOLOv11 specifically includes the following steps:
[0008] (1) Collect the aerial image dataset and convert the format of the aerial image dataset into the YOLO format;
[0009] (2) Build an aerial image detection model. The aerial image detection model uses the YOLOv11 model as the backbone network, including a backbone network, a neck network, and a head network. In the backbone network of the aerial image detection model, replace the Conv module in the backbone network of the YOLOv11 model with the DBRDown module; the neck network of the aerial image detection model is to add the MSFLB module to the neck network of the YOLOv11 model;
[0010] The DBRDown module is a lightweight downsampling convolutional module, which reduces the computational amount and the number of parameters, and improves the detection accuracy at the same time;
[0011] The MSFLB module is a multi-scale feature learning module, which realizes the extraction and fusion of multi-scale features through a multi-branch structure and dilated convolutions with different dilation rates;
[0012] (3) Iteratively train the built aerial image detection model to obtain a trained aerial image detection model, and use the trained aerial image detection model to detect the target to be detected in the aerial image to obtain the detection result.
[0013] The backbone network of the aerial image detection model includes one Conv module, four DBRDown modules, four C3K2 modules, one SPPF module, and one C2PSA module. The processing process of the backbone network is shown in the following formula (1):
[0014] (1);
[0015] In formula (1), is the input of the backbone network, that is, the collected aerial image, , , are three feature maps with different scales extracted by the backbone network.
[0016] The processing process of the DBRDown module specifically includes the following steps:
[0017] S11: The input feature map first extracts local features through the depthwise separable convolution DWConv module, and then the feature map output by the DWConv module is processed by the GeLU activation function and batch normalization, and the output feature map is then connected with the input feature map in a residual connection;
[0018] After the feature map after residual connection is processed by the Split function, the features are split into two branches and processed through two paths respectively: one path extracts fine-grained local features through a 3×3 convolution, and the other path realizes downsampling and global feature fusion through a max pooling layer and a 1×1 convolution in sequence. Finally, the outputs of the two paths are concatenated by Concat to integrate multi-scale information and output the concatenated feature map.
[0019] The parameter calculation formula of the DBRDown module is as shown in the following formula (2):
[0020] (2);
[0021] In formula (2), is the convolution kernel size, represents the number of channels of the input feature map, represents the number of channels of the output feature map.
[0022] The neck network of the aerial image detection model includes two Upsample modules, two MSFLB modules, four C3K2 modules, two Conv modules and four Concat modules. The processing process of the neck network is as shown in the following formula (3):
[0023] (3);
[0024] In formula (3), is the intermediate feature map of the neck network, , and are three feature maps output by the neck network, , , are three feature maps with different scales extracted by the backbone network;
[0025] The three feature maps output by the neck network , and are input into the head network for detection to output the final detection result.
[0026] The processing process of the MSFLB module specifically includes the following steps:
[0027] S21. First, the input feature map is adjusted in channels through a 1×1 convolution. After the output feature map is processed by the Split function, the features are split into three branches. Each branch sequentially performs 3×3 convolutions to extract local spatial features. Then, after activation by the ReLU activation function, dilated convolutions with different dilation rates are used to expand the receptive field and capture context information at different scales. The outputs of the three branches are then connected to the corresponding branch inputs through residual connections;
[0028] S22. The outputs after the residual connections of the three branches are concatenated in the channel dimension. Finally, the output after concatenation is further fused in channel information through a 1×1 convolution to reduce the dimension of the output feature map, obtaining the feature map after multi-scale fusion.
[0029] The dilation rates of the dilated convolutions in the three branches are 1, 2, and 3 respectively, and the original convolution kernel sizes of the dilated convolutions in the three branches are all 3×3.
[0030] The aerial image detection model is iteratively trained, and the loss function for iterative training is the FPIoU loss function. The calculation formula is shown in the following formula (4):
[0031] (4);
[0032] In formula (4), represents the IoU loss function, and represent the width and height of the ground truth box respectively, represents the distance between the corners, and the calculation formula of is shown in the following formula (5):
[0033] (5);
[0034] In formula (5), , represent the horizontal and vertical coordinates of a certain corner of the predicted box respectively, , represent the horizontal and vertical coordinates of a certain corner of the ground truth box respectively.
[0035] Advantages of the present invention:
[0036] (1). In the present invention, the Conv module in the backbone network of the YOLOv11 model is replaced with the DBRDown module, adopting a lightweight downsampling structure, which reduces the computational complexity of the model while maintaining the stability of detection accuracy, so as to better retain the target information features.
[0037] (2) In the present invention, the MSFLB module is added to the neck network of the YOLOv11 model. Aiming at the problem of insufficient extraction of small target feature information, the ability of the aerial image detection model to extract and fuse features of small targets is strengthened, and the detection accuracy of the aerial image detection model for small targets is improved.
[0038] (3) The loss function for iterative training of the aerial image detection model in the present invention is the FPIoU loss function. By combining the IoU loss function and the normalized vertex distance penalty, it not only measures the overlap degree between the predicted box and the ground truth box, but also precisely optimizes the position and shape of the box. The logarithmic function smooths the distance penalty, reduces extreme errors, and improves the robustness to targets of different sizes through normalization, effectively making up for the deficiency of the IoU loss function in position sensitivity and significantly improving the regression quality of the bounding box. Description of the Drawings
[0039] Figure 1 is the flowchart of the present invention.
[0040] Figure 2 is the network framework diagram of the aerial image detection model of the present invention.
[0041] Figure 3 is the network framework diagram of the DBRDown module of the present invention.
[0042] Figure 4 is the network framework diagram of the MSFLB module of the present invention. Detailed Embodiments
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] See Figure 1 , the lightweight aerial image target detection method based on the improved YOLOv11 specifically includes the following steps:
[0045] (1) Collect the aerial image data set and convert the format of the aerial image data set into the YOLO format;
[0046] (2) Construct an aerial image detection model (see Figure 2 ), the aerial image detection model takes the YOLOv11 model as the benchmark network and includes a backbone network, a neck network, and a head network;
[0047] The backbone network of the aerial image detection model is used to extract features at multiple scales from the input aerial image, including a Conv module, four DBRDown modules, four C3K2 modules, an SPPF module, and a C2PSA module. The processing process of the backbone network is shown in the following formula (1):
[0048] (1);
[0049] In formula (1), is the input of the backbone network, that is, the collected aerial image, , , are three feature maps with different scales extracted by the backbone network;
[0050] The neck network of the aerial image detection model is responsible for multi-scale feature fusion. By fusing the feature maps from different stages of the backbone network, the feature representation ability is enhanced, including two Upsample modules, two MSFLB modules, four C3K2 modules, two Conv modules, and four Concat modules. The processing process of the neck network is shown in the following formula (3):
[0051] (3);
[0052] In formula (3), is the intermediate feature map of the neck network, , and are three feature maps output by the neck network, , , are three feature maps with different scales extracted by the backbone network;
[0053] The three feature maps output by the neck network , and are input into the three detection heads Head of the head network for detection, and the final detection results are output, that is, bounding boxes are generated and the objects within the bounding boxes are classified;
[0054] See Figure 3 , the DBRDown module is a lightweight downsampling convolution module, which reduces the amount of computation and the number of parameters, and at the same time improves the detection accuracy. The processing process of the DBRDown module specifically includes the following steps:
[0055] S11. First, the input feature map extracts local features through a depthwise separable convolution DWConv 3×3 module. Then, the feature map output by the DWConv 3×3 module is processed by the GeLU activation function and batch normalization BatchNorm to enhance the non-linear expression ability and feature stability. The output feature map is then connected with the input feature map through a residual connection to alleviate the problem of gradient disappearance and improve the optimization efficiency;
[0056] S12. After the feature map after the residual connection is split by the Split function, the features are split into two branches and processed through two paths respectively: one path extracts fine-grained local features through a 3×3 convolution (Conv 3×3), and the other path realizes downsampling and global feature fusion through a max pooling layer (MaxPool2d) and a 1×1 convolution (Conv 1×1) in sequence. Finally, the outputs of the two paths are concatenated through a Concat process to integrate multi-scale information and output the concatenated feature map;
[0057] The parameter calculation formula for a common downsampling convolution Conv module is shown in the following formula (6):
[0058] (6);
[0059] The parameter calculation formula for the DBRDown module is shown in the following formula (2):
[0060] (2);
[0061] In formulas (2) and (6), is the convolution kernel size, represents the number of channels of the input feature map, represents the number of channels of the output feature map;
[0062] When , , ; then the ratio of the number of parameters of the Conv module to the DBRDown module is: , from which it can be seen that the number of parameters of the DBRDown module is greatly reduced;
[0063] See Figure 4 , the MSFLB module is a multi-scale feature learning module, which realizes the extraction and fusion of multi-scale features through a multi-branch structure and dilated convolutions with different dilation rates. The processing process of the MSFLB module specifically includes the following steps:
[0064] S21. First, the input feature map is adjusted in channels through a 1×1 convolution (Conv 1×1) to enhance the feature expression ability. By increasing the number of channels, the representation space of the feature map is expanded, enabling the network to capture richer contour information. Then, after the output feature map is processed by the Split function, the features are split into three branches. Each branch sequentially performs a 3×3 convolution (Conv 3×3) to extract local spatial features, followed by activation with the ReLU activation function to have better non-linear expression ability. Then, dilated convolutions (DConv 3×3) with different dilation rates are used to expand the receptive field and capture context information at different scales. The dilation rates d of the dilated convolutions in the three branches are 1, 2, and 3 respectively. The outputs of the three branches are then connected to the corresponding branch inputs through residual connections, effectively retaining detailed features and enhancing the network's feature learning ability;
[0065] S22. The outputs after the residual connections of the three branches are concatenated (Concat) in the channel dimension, enabling the fusion of multi-level information from different dimensions to form rich semantic expressions. Finally, the concatenated output is further fused with channel information through a 1×1 convolution (Conv 1×1) to reduce the dimension of the output feature map and obtain the multi-scale fused feature map;
[0066] (3). The constructed aerial image detection model is iteratively trained, and the hyperparameters for model training are set. That is, the SGD optimizer is used to optimize the aerial image detection model. The initial learning rate is 0.01, the final learning rate is 0.0001, the momentum is 0.937, the batch normalization size is 16, and the number of iterations is 200 times. The training set is used to iteratively train the aerial image detection model, and only the weight information of the optimal model and the model of the last round are saved during the training process;
[0067] The loss function for iterative training is the FPIoU loss function, and the calculation formula is shown in the following formula (4):
[0068] (4);
[0069] In formula (4), represents the IoU loss function, and represent the width and height of the ground truth box respectively, represents the distance between the corners, and the calculation formula of is shown in the following formula (5):
[0070] (5);
[0071] In formula (5), 、 represent the horizontal and vertical coordinates of a certain corner of the predicted box respectively, 、 respectively represent the horizontal and vertical coordinates of a certain corner point of the ground truth box.
[0072] Then, the trained aerial image detection model is used to detect the target to be detected in the aerial image, and the detection result is obtained.
[0073] Performance analysis:
[0074] A performance detection experiment is conducted on the aerial image detection model of the present invention and the existing YOLO11n model (YOLOv11 training model). The experimental results are shown in the following table. It can be seen from Table 1 below that the aerial image detection model of the present invention has improved by 3.6%, 1.2%, 2.1% and 4.3% respectively compared with the YOLO11n model in terms of Precision, Recall, mAP@0.5 and mAP@0.5:0.95. At the same time, the Params (number of parameters) index of the aerial image detection model has decreased by 16.2%. These experimental results show that the aerial image detection model of the present invention has achieved significant performance improvement while taking into account the model scale and detection accuracy.
[0075] Table 1
[0076]
[0077] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A lightweight aerial image target detection method based on improved YOLOv11, characterized by: The specific steps include: (1) Collect aerial image data sets and convert the format of the aerial image data sets into YOLO format; (2) Construct an aerial image detection model. The aerial image detection model uses the YOLOv11 model as the benchmark network, including a backbone network, a neck network, and a head network. In the backbone network of the aerial image detection model, the Conv module in the backbone network of the YOLOv11 model is replaced with a DBRDown module. The neck network of the aerial image detection model is the neck network of the YOLOv11 model with the MSFLB module added. The DBRDown module is a lightweight downsampling convolution module that reduces the amount of calculation and parameters while improving detection accuracy; The processing process of the DBRDown module specifically includes the following steps: S11, the input feature map first extracts local features through the depthwise separable convolution DWConv module, and then the feature map output by the DWConv module is processed by the GeLU activation function and batch normalization, and the output feature map is then residually connected with the input feature map; S12. After the feature map after residual connection is segmented by the Split function, the features are split into two branches and processed by two paths respectively: one path extracts fine-grained local features through 3×3 convolution, and the other path sequentially realizes downsampling and global feature fusion through the maximum pooling layer and 1×1 convolution. Finally, the outputs of the two paths are concatenated and Concat processed to integrate multi-scale information and output the concatenated feature map; The MSFLB module is a multi-scale feature learning module, which realizes the extraction and fusion of multi-scale features through a multi-branch structure and dilated convolutions with different dilation rates; (3) Iteratively train the constructed aerial image detection model to obtain a trained aerial image detection model, and use the trained aerial image detection model to detect the target to be detected in the aerial image to obtain the detection result.
2. The lightweight aerial image target detection method based on improved YOLOv11 according to claim 1, characterized in that: The backbone network of the aerial image detection model includes a Conv module, four DBRDown modules, four C3K2 modules, a SPPF module and a C2PSA module. The processing process of the backbone network is shown in the following formula (1): (1); In formula (1), As the input of the backbone network, that is, the collected aerial images, , , Feature maps of three different scales extracted by the backbone network.
3. The lightweight aerial image target detection method based on improved YOLOv11 according to claim 1, characterized in that: The parameter calculation formula of the DBRDown module is as follows (2): (2); In formula (2), is the convolution kernel size, Represents the number of channels of the input feature map, Represents the number of channels of the output feature map.
4. The lightweight aerial image target detection method based on improved YOLOv11 according to claim 1, characterized in that: The neck network of the aerial image detection model includes two Upsample modules, two MSFLB modules, four C3K2 modules, two Conv modules and four Concat modules. The processing process of the neck network is shown in the following formula (3): (3); In formula (3), is the intermediate feature map of the neck network, , and Three feature maps output by the neck network, , , Three feature maps of different scales extracted by the backbone network; Three feature maps output by the neck network , and Input into the head network for detection and output the final detection result.
5. The lightweight aerial image target detection method based on improved YOLOv11 according to claim 4 is characterized in that: The processing process of the MSFLB module specifically includes the following steps: S21, the input feature map is first adjusted by 1×1 convolution for channel adjustment, and the output feature map is segmented by Split function, and the features are split into three branches. Each branch is sequentially subjected to 3×3 convolution to extract local spatial features, and then activated by ReLU activation function, and then the receptive field is expanded by using dilated convolution with different dilation rates to capture contextual information of different scales. The outputs of the three branches are then residually connected with the corresponding branch inputs; S22: The outputs of the three branch residual connections are concatenated in the channel dimension. Finally, the concatenated output is further convolved with a 1×1 convolution to further fuse the channel information, reduce the dimension of the output feature map, and obtain a multi-scale fused feature map.
6. The lightweight aerial image target detection method based on improved YOLOv11 according to claim 5 is characterized in that: The expansion rates of the dilated convolutions in the three branches are 1, 2, and 3, respectively, and the original convolution kernel sizes of the dilated convolutions in the three branches are all 3×3.
7. The lightweight aerial image target detection method based on improved YOLOv11 according to claim 1, characterized in that: The aerial image detection model is iteratively trained, and the loss function of the iterative training is the FPIoU loss function, and the calculation formula is shown in the following formula (4): (4); In formula (4), represents the IoU loss function, and Represent the width and height of the real frame respectively, represents the distance between corner points, The calculation formula is shown in the following formula (5): (5); In formula (5), , Respectively represent the horizontal and vertical coordinates of a corner point of the prediction box, , They represent the horizontal and vertical coordinates of a corner point of the real box respectively.
Citation Information
Patent Citations
Tracheal intubation positioning method and device based on deep learning and storage medium
CN112907539A