Unmanned aerial vehicle aerial photography mangrove forest target detection method based on improved YOLOv11

By improving the YOLOv11 algorithm, GhostConv, EMA attention mechanism and MSAA module were introduced, which solved the problem of insufficient detection performance in scenarios such as dense and overlapping small and medium-sized mangroves, and achieved more efficient and accurate object detection.

CN120182871APending Publication Date: 2025-06-20GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510344552.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In mangrove ecological monitoring, traditional object detection algorithms are difficult to achieve high accuracy and efficient detection in scenarios where small targets are dense and overlapping.

Method used

The UAV aerial mangrove object detection method based on improved YOLOv11 was adopted, and GhostConv replaced the conventional convolution, EMA attention mechanism was introduced to replace the ordinary attention mechanism in C2PSA, and MSAA module was added to the Neck layer to enhance feature fusion and detection performance.

Benefits of technology

It improves the accuracy and efficiency of target detection, can more effectively deal with small targets and overlapping situations, and meets the detection needs of mangrove ecological monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182871A_ABST
    Figure CN120182871A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle aerial photograph mangrove forest target detection method based on an improved YOLOv11 algorithm, and belongs to the technical field of unmanned aerial vehicle aerial photograph target detection. In the invention, aiming at the characteristics of small, dense and overlapped to-be-detected targets in mangrove forest ecological monitoring, a GhostConv module is utilized to replace the original common convolution, the calculation amount and model parameters are reduced, an EMA attention mechanism is introduced to replace a common attention mechanism in a C2PSA module, the attention distribution is smoother and more stable, an MSAA module is embedded in a Neck layer, and the detection accuracy is improved. The feature information of different scales is effectively integrated, the multi-scale feature fusion is enhanced, and the detection performance of the small target under the complex background is effectively improved. Compared with the prior art, the detection accuracy of dense small targets in the mangrove forest is improved while efficient real-time detection is guaranteed, and the method is suitable for unmanned aerial vehicle aerial photography mangrove forest ecological monitoring scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) aerial photography target detection, and particularly to a method for detecting mangrove targets in UAV aerial photography based on improved YOLOv11. Technical Background

[0002] With the rapid development of UAV technology, especially in the field of ecological monitoring, UAVs can conduct aerial photography of large areas of forest areas with high-definition cameras, collect target image data in real time, and have become an important tool for efficiently obtaining large-scale image information, providing strong support for ecological environmental protection and precision forestry. However, in the application scenario of mangrove ecological monitoring, mangrove targets in the high-resolution images collected by UAVs are often small, dense, and have complex backgrounds. Coupled with the limited computing resources of the UAV platform, traditional target detection algorithms are difficult to meet the actual requirements in terms of accuracy, and there are still many challenges in detecting small targets in UAV aerial photography of mangroves.

[0003] Deep learning has significant advantages in small target detection. Among them, the YOLO series is well-known for the significant balance between speed and accuracy. YOLOv11 is the latest version of this series, and still maintains high detection accuracy while using fewer parameters. However, if YOLOv11 is directly used for detecting dense small targets, problems such as the adaptation of the default configuration, small and overlapping targets in the application scenario of mangrove ecological monitoring, and the computing requirements for processing high-resolution aerial photography images will lead to insufficient detection correctness and accuracy, and insufficient feature fusion. Therefore, it is urgent to design a method for detecting UAV aerial photography targets based on improved YOLOv11 for small target dense scenarios. Summary of the Invention

[0004] The present invention proposes a method for detecting mangrove targets in UAV aerial photography based on improved YOLOv11, aiming to solve the problem that the detection correctness and accuracy are insufficient in the case of dense distribution of small targets and it is difficult to be applied to the mangrove ecological monitoring in the real scenario.

[0005] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for detecting mangrove targets in UAV aerial photography based on improved YOLOv11, including the following steps:

[0006] S1: Conduct data collection, use a UAV to take a video of mangroves in the mangrove nature reserve at low altitude, and perform frame extraction on the video to construct a dataset based on the mangrove scenario as the Mangrov dataset. And use the LabelImg annotation tool to annotate the dataset, annotate mangrove targets in different growth states, and save the label file in YOLO format. The dataset is divided into a training set, a validation set, and a test set, with a ratio of 7:2:1;

[0007] S2: Improve the YOLOv11 algorithm. The improvement based on the improved YOLOv11 algorithm is obtained from the improvement of the network structure of YOLOV11. The overall structure of the YOLOv11 algorithm includes an input and preprocessing end, a Backbone end, a Neck end, and a detection head. To address the challenges of small, dense, and overlapping targets in the UAV mangrove ecological monitoring task, GhostConv is introduced into the YOLOv11 network structure to replace the conventional convolution, and the EMA attention mechanism is introduced to replace the ordinary attention mechanism in C2PSA, forming C2PSA_EMA. At the same time, MSAA is introduced after the Concat module in the Neck part.

[0008] S3: Use the pre-trained YOLOv11 model as the base model, load the pre-trained weights, set the initial learning rate and weight decay, and adopt staged decay. Set the training batch size and the number of training epochs according to the GPU memory size, and train the model with the Mangrov dataset.

[0009] S4: Infer the trained improved YOLOv11 model on the validation set, identify and detect images, output the detection results with detection boxes, observe whether there are problems such as missed detections and false detections in the detection results, and analyze metrics such as precision, recall, and the Loss curve on the data to evaluate the object detection performance.

[0010] Furthermore, in step S1, when collecting data, it is necessary to select typical and representative mangrove areas to ensure that the images cover samples in different growth states (such as healthy, damaged, pest-infected, etc.). At the same time, it is necessary to pay attention to the weather and lighting conditions to avoid too dark and overexposed situations. In addition, it is necessary to set an appropriate flight altitude and flight path and use a camera with a high enough resolution to ensure that the images have sufficient resolution and can cover a large area.

[0011] Furthermore, in step S2, GhostConv is introduced into the YOLOv11 network structure to replace the conventional convolution. First, a small number of intrinsic feature maps are generated from the input data X through a 1×1 standard convolution, and then each intrinsic feature map is linearly transformed through a 3×3 depthwise separable convolution DWConv to generate the output features.

[0012] Furthermore, in step S2, the ordinary attention mechanism in C2PSA is replaced with EMA to form C2PSA_EMA, enhancing the multi-scale features of the backbone. EMA adopts three parallel paths to extract the attention weight descriptors of the grouped feature maps. Two of the branches are 1×1 branches, which decompose the original input tensor into two parallel 1D feature encoding vectors. These two parallel 1D feature encoding vectors share a 1×1 convolution for dimensionality reduction, and the output is decomposed into two parallel 1D feature encoding vectors. Then, they respectively pass through the non-linear Sigmoid function to fit the 2D bivariate normal distribution on the linear convolution. Finally, the original intermediate feature map is aggregated into the output using the attention map weights learned by the two parallel paths; the other branch is a 3×3 branch. The 3×3 branch first captures local cross-channel interactions through a 3×3 convolution, and then uses 2D global average pooling to encode global spatial information in the 1×1 branch. The output is directly converted into the corresponding dimensional shape, and then the spatial attention map is derived. Finally, the spatial attention weight values generated by the two groups are aggregated respectively, and through the Sigmoid function, the pixel-level pairwise relationship is captured and the global context of all pixels is highlighted.

[0013] Furthermore, in step S2, an MSAA module is introduced between the Concat module and the C3K2 module at the Neck end, and the features of different scales are fused through spatial and channel attention mechanisms respectively, enhancing the feature expression. On the spatial aggregation path, the multi-scale spatial information is fused by summing the convolutions with different kernel sizes and spatial feature aggregation; on the channel aggregation path, the dimensionality is reduced through global average pooling, and then convolution and ReLU activation are performed to generate the channel attention map, which is combined with the attention map obtained from the spatial refinement path, enhancing the feature expression in both the spatial and channel dimensions.

[0014] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:

[0015] Aiming at the problems of small, dense and overlapping targets in the application scenario of mangrove ecological monitoring, the present invention proposes a method for detecting mangrove ecological targets by drone aerial photography based on improved YOLOv11. Specifically, this method uses GhostConv to replace the conventional convolution in the original network structure of YOLOv11, reducing the computational load and improving the inference speed; adding the EMA attention mechanism to replace the ordinary attention mechanism in C2PSA to stabilize the attention distribution and enhance the multi-scale features of the Backbone; and adding an MSAA module at the Neck layer to effectively integrate the feature information of different scales and strengthen the multi-scale feature fusion. This method improves the deficiencies in the detection performance of ordinary YOLOv11 when facing small, dense and overlapping targets in mangrove ecological monitoring. Description of the Drawings

[0016] Figure 1 Flowchart of the method for detecting mangrove targets in UAV aerial photography by improving YOLOv11 of the present invention;

[0017] Figure 2 Network structure diagram of the improved YOLOv11 model of the present invention;

[0018] Figure 3 Structure diagram of the Ghostsconv convolution introduced by the present invention;

[0019] Figure 4 Structure diagram of the EMA attention mechanism introduced by the present invention;

[0020] Figure 5 Structure diagram of the MSAA introduced by the present invention. Detailed implementation manners

[0021] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The following examples are only used to illustrate the technical solutions of the present invention more clearly, and cannot be used to limit the protection scope of the present invention

[0022] The present invention discloses a method for detecting mangrove targets in UAV aerial photography based on improved YOLOv11. Referring to Figure 1 , the steps for obtaining a model for detecting mangrove targets in UAV aerial photography based on improved YOLOv11 are as follows:

[0023] S1: Perform data collection. Use a UAV to take low-altitude videos of mangroves in the mangrove nature reserve, and perform frame extraction on the videos to construct a dataset based on the mangrove scene (Mangrov dataset). Use the LabelImg annotation tool to annotate the dataset, annotate mangrove targets in different growth states, and save the label file in YOLO format. The dataset is divided into a training set, a validation set and a test set, with a ratio of 7:2:1;

[0024] S2: Improve the YOLOv11 algorithm. The overall structure of the YOLOv11 algorithm includes an input and preprocessing end, a Backbone end, a Neck end, and a detection head; as Figure 2 shown, in response to the challenges of small, dense, and overlapping detection targets in the mangrove ecological detection task, replace the CBS module in the Backbone network of the YOLOv11 model with a GhostConv module, and replace the 17th Conv module in the Neck with a GhostConv module. Refer to Figure 3, GhostConv divides the operations of a regular convolutional layer into two parts: First, it performs batch normalization (BN) on it and introduces the SiLU activation function. The first part generates a small number of intrinsic feature maps Y′ from the input data X through a 1×1 standard convolution, Y′ = X * f′, where f′ is the convolutional kernel for generating intrinsic features; subsequently, for each intrinsic feature map, a linear transformation is performed through a 3×3 depthwise separable convolution (DWConv), BN processing is performed on it, and the SiLU activation function is introduced to adapt to the stacking of the two parts of the feature maps. The core idea of GhostConv is to utilize the redundant information existing in the input feature maps to generate the same number of output features as traditional convolution with less computation, while greatly reducing the number of parameters and the amount of computation;

[0025] The EMA attention mechanism is introduced to replace the ordinary attention mechanism in C2PSA, forming C2PSA_EMA, so that the spatial semantic features are well distributed within each feature group, enhancing the multi-scale features of the backbone. Refer to Figure 4 , EMA uses three parallel paths to extract the attention weight descriptors of the grouped feature maps. Two of the paths are 1×1 branches (i.e., 1×1 branches), and the third path is a 3×3 branch: In the 1×1 branch, first, the original input tensor is decomposed into two parallel 1D feature encoding vectors. One of the parallel routes comes from 1D global average pooling in the horizontal dimension direction. Let the original input tensor X ∈ R C×H×W denotes the intermediate feature map, where C represents the number of input channels, and H and W represent the spatial dimensions of the input features respectively. Then, the 1D global average pooling that encodes the global information in the horizontal dimension direction of C at height H is expressed as Similarly, the other path comes from 1D global average pooling in the vertical dimension direction. Then, the pooling output in C with width W is Subsequently, these two parallel 1D feature encoding vectors share a 1×1 convolution for dimensionality reduction. The output is decomposed into two parallel 1D feature encoding vectors, and then they respectively pass through the nonlinear Sigmoid function to fit the 2D bivariate normal distribution on the linear convolution. Finally, the original intermediate feature map is aggregated into the output using the attention map weights learned from the two parallel paths; meanwhile, the other 3×3 branch captures local cross-channel interactions through a 3×3 convolution to expand the feature space. Subsequently, 2D global average pooling is used to encode the global spatial information in the 1×1 branch. The output of the 3×3 branch is directly converted into the corresponding dimensional shape, and then the spatial attention map is derived. Finally, the spatial attention weight values generated by the two groups are aggregated respectively, and through the Sigmoid function, the pixel-level pairwise relationships are captured and the global context of all pixels is highlighted. Therefore, without reducing the channel dimension, EMA captures features on a larger scale through a 3×3 convolution and aggregates the output feature maps of multiple parallel sub-networks through a cross-space learning method.

[0026] An MSAA module is introduced between the Concat module and the C3K2 module in the Neck. Refer to Figure 5 , and there are two parallel paths, spatial and channel, in the MSAA module for feature aggregation. On the spatial refinement path, the input combined features are subjected to feature aggregation, and the channel C1 is restored to C2 through a 1×1 convolution, where Subsequently, multi-scale fusion is performed, which respectively passes through convolutions with kernel sizes of 3×3, 5×5, and 7×7, and then the spatial features are aggregated using mean and max pooling, and then element-wise multiplication is performed with the feature map activated by Sigmoid using a 7×7 convolution kernel. At the same time, in the channel aggregation path, global average pooling is used to reduce the dimension to C1×1×1, then 1×1 convolution and ReLU activation are performed, and then a channel attention map is generated and expanded to match the input size, and combined with the attention map of spatial refinement. Introducing the MSAA module enhances the feature expression ability in both spatial and channel dimensions and improves the feature resolution of the model.

[0027] S3: Use the pre-trained YOLOv11 model as the base model, load the pre-trained weights, set the initial learning rate between 0.001 and 0.01, adopt staged decay, set the weight decay between 0.0001 and 0.0005, set the batch size and the number of training epochs according to the GPU memory size, and use the Mangrov dataset for model training to train and evaluate the model;

[0028] S4: Infer the trained improved YOLOv11 model on the validation set, identify and detect images, output the detection results with detection boxes, observe whether there are problems such as missed detections and false detections in the detection results, and analyze metrics such as precision, recall, and the Loss curve on the data to evaluate the object detection performance. Among them, precision reflects the proportion of samples that are truly positive among the samples predicted as positive, measuring how accurate the objects predicted by the model are, and its calculation formula is where TP represents true positives (the number of correctly predicted objects), and FP represents false positives (the number of incorrectly predicted objects); recall reflects the proportion of objects that are successfully detected by the model among all truly existing objects, measuring the degree of missed detections of the model, and its calculation formula is where FN represents false negatives (the number of objects that actually exist but are not detected); and the Loss curve reflects the change trend of the loss function value with the number of iterations during the model training process, used to evaluate the convergence of the training process and the model performance. When the Loss curve shows a gradually decreasing trend, it indicates that the model is constantly being optimized, the error is decreasing, and the feature extraction and prediction effects are continuously improving.

[0029] The above description is for the specific implementation of the present invention and not a limitation thereof. Those skilled in the relevant technical field can also make their respective equivalent technical solutions without departing from the scope of the present invention. Therefore, all equivalent technical solutions should be included in the scope of the patent protection of the invention.

Claims

1. A method for detecting mangrove targets in drone aerial photography based on improved YOLOv11, characterized by: The method comprises the following steps: S1: Data collection: Use drones to shoot mangrove videos in mangrove nature reserves at low altitudes, extract frames from the videos, and construct a dataset based on mangrove scenes, namely the Mangrov dataset. Use the LabelImg annotation tool to annotate the dataset, annotate mangrove targets in different growth states, and save the label file in YOLO format. The dataset is divided into training set, validation set, and test set in a ratio of 7:2:

1. S2: Improve the YOLOv11 algorithm. The improvement based on the improved YOLOv11 algorithm is obtained based on the improvement of the network structure of YOLOV11. The overall structure of the YOLOv11 algorithm includes an input and preprocessing end, a Backbone end, a Neck end, and a detection head. In view of the challenges of small, dense, and overlapping targets in the drone aerial photography of mangrove ecological monitoring tasks, GhostConv is introduced into the YOLOv11 network structure to replace conventional convolution, and the EMA attention mechanism is introduced to replace the ordinary attention mechanism in C2PSA to form C2PSA_EMA. At the same time, MSAA is introduced after the Concat module of the Neck part. S3: Use the pre-trained YOLOv11 model as the base model, load the pre-trained weights, set the initial learning rate and weight decay, and use phased decay. Set the training batch and number of training rounds according to the GPU memory size, and use the Mangrov dataset for model training. S4: Perform inference on the validation set using the trained improved YOLOv11 model to identify the detection image and output the detection result including the detection box. Observe whether there are any missed detection or false detection problems in the detection result. Analyze the precision rate, recall rate and loss curve on the data to evaluate the target detection performance.

2. The method for detecting mangrove targets by drone aerial photography based on improved YOLOv11 as claimed in claim 1, characterized in that: In step S1, when collecting data, it is necessary to select a typical and representative mangrove area to ensure that the image covers samples in three different growth states: healthy, damaged, and pest-infested. At the same time, attention should be paid to weather and lighting conditions to avoid being too dark or overexposed. In addition, a suitable flight altitude and route should be set, and a camera with sufficiently high resolution should be used to ensure that the image has sufficient resolution and can cover a large area.

3. The method for detecting mangrove targets by drone aerial photography based on improved YOLOv11 as claimed in claim 1, characterized in that: In the step S2, GhostConv is introduced into the YOLOv11 network structure to replace the conventional convolution. A small number of intrinsic feature maps are first generated from the input data X through a 1×1 standard convolution, and then each intrinsic feature map is linearly transformed through a 3×3 depth-separable convolution DWConv to generate output features.

4. The method for detecting mangrove targets by drone aerial photography based on improved YOLOv11 as claimed in claim 1, characterized in that: In the step S2, EMA is used to replace the ordinary attention mechanism in C2PSA to form C2PSA_EMA, and the multi-scale features of the backbone are enhanced. EMA uses three parallel paths to extract the attention weight descriptors of the grouped feature maps, two of which are 1×1 branches, which decompose the original input tensor into two parallel 1D feature encoding vectors. The two parallel 1D feature encoding vectors share a 1×1 convolution for dimensionality reduction, and the output is decomposed into two parallel 1D feature encoding vectors, which are then respectively subjected to a nonlinear Sigmoid function to fit the 2D binormal distribution on the linear convolution. Finally, the original intermediate feature map is aggregated as the output using the attention map weights learned by the two parallel paths; the other branch is a 3×3 branch, which first captures local cross-channel interactions through 3×3 convolutions, and then uses 2D global average pooling to encode global spatial information in the 1×1 branch. The output is directly converted to the corresponding dimensional shape, and then the spatial attention map is derived. Finally, the spatial attention weight values ​​generated by the two groups are aggregated, and the pixel-level pairwise relationship is captured through the Sigmoid function and the global context of all pixels is highlighted.

5. The method for detecting mangrove targets by drone aerial photography based on improved YOLOv11 as claimed in claim 1, characterized in that: In the step S2, an MSAA module is introduced between the Concat module and the C3K2 module in the Neck part, and features of different scales are fused through spatial and channel attention mechanisms to enhance feature expression; On the spatial aggregation path, multi-scale spatial information is fused by summing convolutions with different kernel sizes and aggregating spatial features. On the channel aggregation path, dimensionality reduction is performed through global average pooling, and convolution and ReLU activation are performed to generate a channel attention map, which is combined with the attention map obtained by the spatial refinement path to enhance the feature expression in both spatial and channel dimensions.

Citation Information

Cited By

  • Surface defect small target detection method based on multi-scale feature interaction

    CN120374613A

  • Training method of micronucleus image deep learning detection model

    CN120766059A