Target detection method in unmanned aerial vehicle aerial photography scene based on improved YOLOv8

The improved YOLOv8 model with LHGNet backbone and fusion modules addresses the challenge of detecting small targets in UAVs by improving feature extraction and fusion, resulting in enhanced accuracy and efficiency.

CN120318582APending Publication Date: 2025-07-15JIANGSU XINYANG NEW MATERIALS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510450992.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

It is difficult for drones to detect tiny target objects in aerial photography scenarios. Due to the influence of different perspectives and complex backgrounds, it is difficult for the prior art to achieve efficient and accurate target detection.

Method used

Using the improved YOLOv8 model, by building the LHGNet backbone network and optimizing the Neck part, combining deep separable convolution and channel shuffle technology, feature extraction capabilities are enhanced, and LGS bottleneck and LGSCSP fusion modules are introduced to optimize the feature fusion process and improve the model's detection accuracy and efficiency of the detection of small targets.

Benefits of technology

It significantly improves the accuracy and efficiency of target detection in drone aerial photography scenarios, can efficiently identify tiny objects, and maintain the lightweight characteristics of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318582A_ABST
    Figure CN120318582A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method in an unmanned aerial vehicle aerial photography scene based on improved YOLOv8 in the technical field of target detection. The method comprises the following steps: S1, data acquisition and preprocessing; s2, labeling and constructing a data set; s3, training and optimizing the model; for the improved YOLOv8 model, training and optimizing the model by using a training data set with a label, and learning features of a target in an image; s4, performing model evaluation; the performance of the model is verified on an independent test set, evaluation is carried out through accuracy and recall rate indexes, and fine adjustment is carried out; s5, actual detection and result acquisition; inputting a to-be-detected aerial photography data set into the trained and optimized detection model, carrying out real-time analysis on the image, detecting small objects in the image, and determining the positions of the small objects; the output result covers the category and position of the object, necessary information is provided for subsequent processing, and the problem that edge equipment carried by the unmanned aerial vehicle (UAV) is difficult to detect the tiny object from multiple angles is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and particularly relates to a target detection method. Background Art

[0002] With the progress of unmanned aerial vehicle (UAV) and target detection technologies, UAV target detection has become an important part of many fields such as agriculture, surveillance, disaster response, infrastructure inspection, etc. The use of UAVs can efficiently and flexibly collect data in vast or inaccessible areas, making it a valuable tool for tasks benefiting from an aerial perspective. Therefore, researchers and practitioners in the field of computer vision seek to develop lightweight, robust, and accurate UAV target detection algorithms to enhance the capabilities of these aerial systems in different applications.

[0003] However, due to various factors such as different perspectives of UAVs and large imaging distances, small target objects in images are represented by only a few pixels. In addition, complex backgrounds and dynamic random noise can have a negative impact on target detection accuracy. Therefore, considering the existence of these interference factors, detecting tiny targets remains a formidable challenge. Summary of the Invention

[0004] Aiming at the deficiencies in the prior art, the present invention provides a target detection method in an aerial photography scene of a UAV based on improved YOLOv8, which solves the problem that edge devices carried by UAVs face difficulties in detecting tiny objects from multiple angles.

[0005] The object of the present invention is achieved as follows: A target detection method in an aerial photography scene of a UAV based on improved YOLOv8, comprising the following steps: S1. Data acquisition and preprocessing: Using a UAV to collect aerial photography picture data under different conditions, and preprocessing the collected pictures, including cropping, scaling, and color correction, to adapt to the model input requirements; S2. Dataset annotation and construction: Classifying the processed data according to whether it is annotated, constructing a training dataset containing the positions and categories of target objects, and a dataset for model validation and testing; S3. Model training and optimization: For the improved YOLOv8 model, using the labeled training dataset to train and optimize the model, and learning the features of the targets in the images; S4. Model evaluation: Verifying the performance of the model on an independent test set, evaluating through accuracy and recall rate metrics, and making fine adjustments to the model based on the evaluation results; S5. Actual Detection and Result Acquisition: Input the aerial photography dataset to be measured into the trained and optimized detection model. The model performs real-time analysis on the input aerial images, detects small objects in the images, and determines their positions. The output results will cover the categories and positions of the objects, providing necessary information for subsequent processing.

[0006] Furthermore, the specific design of the improved YOLOv8 model in step 3) is as follows: S31. Construct the LHGNet backbone network. Enhance the feature extraction ability of the backbone network by introducing the residual technique, and on the basis of the existing HGNetv2 backbone network, incorporate depthwise separable convolutions into the backbone network. S32. Optimize the Neck part; Expand the PAN-FPN module, add a feature fusion layer for small object detection after the first LHG block of the backbone network. Subsequently, delete the feature fusion layer in the Neck responsible for detecting large objects, and retain three feature maps for the detection head to output. S33. Head part; The Head part is responsible for the final detection of the target based on the feature maps passed from the previous network. It infers the coordinates of the bounding box, the confidence of the target, and the classification probability of the object through a series of convolutional operations, and then filters out overlapping prediction boxes through non-maximum suppression to output the final detection results.

[0007] In the improved YOLOv8 model, the data processing flow is as follows: First, the input image is subjected to feature extraction through the constructed LHGNet backbone network, which uses the residual technique and depthwise separable convolutions to enhance the feature extraction ability. Then, the optimized Neck part specifically processes small objects by expanding the PAN-FPN module and adding a feature fusion layer, while deleting the feature fusion layer for detecting large objects to retain the key feature maps for the detection head. Finally, the Head part receives these feature maps, performs convolutional operations to infer the bounding box coordinates, target confidence, and classification probability, and filters out overlapping prediction boxes through non-maximum suppression to output the final detection results. This process ensures the accuracy and efficiency of the model in target detection in the UAV aerial photography scenario.

[0008] Furthermore, in step S31, the LHGNet backbone network includes the following modules: 1) HGStem module; The initial image is processed through the HGStem module, which is mainly used for upsampling and downsampling the features of the input image. The HGStem module receives the input image and processes it through consecutive convolutional and pooling steps, finally generating a set of feature maps to provide useful feature representations for subsequent network layers. 2) LHG module; All 3x3 standard convolution units in the original HGNetv2 architecture are replaced by two groups of depthwise separable convolution units. The LHG module consists of two depthwise separable convolution units arranged in sequence, which perform a cascading operation, and finally are processed by two 1x1 standard convolution units; 3) ConvModule convolution module; Responsible for extracting features of the input image; 4) SPPF module; The SPPF module is used to process the feature map output by the Backbone. Through a series of convolution and pooling operations, the input feature map is transformed into a multi-scale feature representation.

[0009] In step S31, the data is first preliminarily processed by the HGStem module to generate a basic feature map. Subsequently, these feature maps sequentially pass through multiple LHG modules to perform efficient feature extraction using depthwise separable convolution units. Then, the ConvModule further refines the features. Finally, the SPPF module integrates multi-scale features to provide rich feature information for subsequent object detection. This coherent processing flow ensures that the LHGNet backbone network can effectively extract and integrate key features from the input image.

[0010] Furthermore, the design of the Neck part in step S32 is optimized as follows: 1) Upsample module; In the Neck stage of network training, the feature maps of each layer are upsampled and fused; 2) Concat module; The Concat module merges feature maps from different levels in the channel dimension, integrates information of different scales and semantic levels, enhances the model's ability to understand objects. By concatenating feature maps of different levels, the Concat module can provide more comprehensive information to subsequent network layers; 3) LGSCSP fusion module; Designed by a one-time aggregation method, while simplifying the calculation and network structure, it ensures the accuracy of the model. The LGSCSP fusion module improves the model's performance in identifying multi-size objects by optimizing the feature fusion process. The LGS bottleneck module is integrated into the LGSCSP fusion module; 4) ConvModule convolution module; Responsible for extracting feature data from the input image, and performing multiple convolution operations on the input image through convolution operations to capture the features in the image.

[0011] In the Neck stage, the data processing flow first upsamples the feature map through the Upsample module to enhance the spatial resolution of the features. Then, the Concat module merges these upsampled feature maps in the channel dimension to integrate information from different levels. Next, the LGSCSP fusion module further optimizes feature fusion through its one-shot aggregation method, improving the model's ability to recognize multi-scale targets. Meanwhile, the integrated LGS bottleneck module helps simplify the calculation. Finally, the ConvModule extracts and refines features through multiple convolutional operations, providing rich feature information for object detection in the Head part. This process ensures that the model can efficiently process and understand the object information in the image.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes to significantly improve the detection accuracy and the operating efficiency of the model by deeply mining and fully utilizing the local details and channel features in the image.

[0013] To enhance the ability to recognize tiny objects and improve the efficiency of the model, the present invention introduces LHGNet as the backbone network, which is an advanced feature extraction architecture. LHGNet uses LHG blocks and combines depthwise separable convolution and channel shuffle techniques to construct a deep feature extraction framework that can comprehensively capture and integrate local details and channel features at different levels, thus achieving more accurate feature extraction.

[0014] To further improve the efficiency of the model, the present invention introduces the "LGS bottleneck" and "LGSCSP" fusion modules. The "LGS bottleneck" aims to maximize the retention of channel information by adding additional convolutional layers and using connection operations to replace traditional addition operations, in order to extract more refined features at multiple scales.

[0015] These improvements greatly enhance the detection accuracy of the model for tiny targets, making it perform excellently in practical applications. Through these innovative structural designs and optimization strategies, the proposed detection method achieves efficient and accurate object detection while maintaining light weight.

[0016] In summary, this solution improves the design of the YOLOv8 model, especially the introduction of the LHGNet backbone network, the LGS bottleneck, and the LGSCSP fusion module. These innovation points together constitute the core of the present invention, and they cooperate with each other to improve the accuracy and efficiency of object detection in the UAV aerial photography scenario. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.

[0018] Figure 1 The step flow of a target detection method in an aerial photography scene of an unmanned aerial vehicle based on improved YOLOv8 of the present invention.

[0019] Figure 2 The overall network structure of the Yolov8-UAV model based on the YOLOv8 network of the present invention.

[0020] Figure 3 The schematic diagram of the depthwise separable convolution introduced in the present invention.

[0021] Figure 4 The schematic diagram of the HGStem module in the present invention.

[0022] Figure 5 The schematic diagram of the LHG module in the present invention.

[0023] Figure 6 The schematic diagram of the LGSCSP fusion module in the present invention. Detailed implementation manners

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0025] As Figure 1 shown, a target detection method in an aerial photography scene of an unmanned aerial vehicle based on improved YOLOv8 includes the following steps: S1. Data collection and preprocessing. Use an unmanned aerial vehicle to collect aerial photography picture data under different conditions, and preprocess the collected pictures, including cropping, scaling, color correction, etc., to adapt to the model input requirements.

[0026] S2. Dataset annotation and construction. Classify the processed data according to whether it is annotated, and construct a training dataset containing the positions and categories of target objects, as well as a dataset for model verification and testing.

[0027] S3. Model Training and Optimization. For the improved YOLOv8 model (Yolov8-UAV), use the training dataset with labels to train and optimize the model and learn the features of the targets in the images. By combining depthwise separable convolution and channel shuffle modules, construct the LHGNet backbone network, aiming to mine and utilize the local detail information and channel features in the up and down spaces, improve the detection accuracy of small targets and achieve model lightweighting. At the same time, introduce the LGS bottleneck and LGSCSP fusion modules to further optimize and lightweight the model. The overall network structure of the Yolov8-UAV model is as Figure 2 shown as follows: S31. Construct the LHGNet backbone network (Backbone). By introducing residual technology to enhance the feature extraction ability of the backbone network, and on the basis of the existing HGNetv2 backbone network, merge depthwise separable convolution into the backbone network to address the challenge of a significant increase in parameters as the network scale grows. As Figure 3 shown, depthwise separable convolution consists of two parts: depth convolution and point convolution, both of which play crucial roles in the overall architecture. Due to the addition of pointwise convolution, the output channels can be freely adjusted, and at the same time, it allows channel fusion of the feature maps output by depth convolution. Therefore, compared with standard convolution, it shows superior generalization ability and reduces the computational complexity.

[0028] In step S31, it mainly includes the following modules: 1) HGStem module. The initial image is processed through the HGStem module, as Figure 4 shown, which is mainly used for upsampling and downsampling the features of the input image. The HGStem module receives the input image and processes it through consecutive convolution and pooling steps, and finally generates a set of feature maps to provide useful feature representations for subsequent network layers.

[0029] 2) LHG module. To simplify the calculation and construct a deeper network, thus exploring and utilizing the features between the spatial domain and the channel domain more deeply, all 3x3 standard convolution units in the original HGNetv2 architecture are replaced with two groups of depthwise separable convolution units, and this depthwise separable convolution helps to enhance the fusion effect between channels in the feature map. As Figure 5As shown, the LHG module consists of two sequentially arranged depthwise separable convolution units that perform a cascading operation and are finally processed by two 1x1 standard convolution units. These two 1x1 standard convolution modules play a crucial role in the LHG module. They not only help to conveniently adjust the dimension of the output feature map but also play a key role in extracting local information, especially related to detecting tiny and sparse targets. This helps to improve the detection accuracy of capturing tiny objects. To enhance channel information fusion, a channel shuffle module is integrated between these two 1x1 convolution modules. In addition, by combining the MPS strategy, the LHG block obtains richer gradient flow information. This enables the acquisition of feature information at different depths and receptive fields, thereby further enhancing its ability to extract features from tiny and sparse targets. In this way, the LHG block can extract more detailed features from each feature map under the guidance of local and channel information through two 1x1 standard convolutions and channel shuffle.

[0030] 3) ConvModule convolution module. The convolution module is the basic unit that constitutes the network and is responsible for extracting the features of the input image. The nonlinear expression ability and feature recognition accuracy of the model are enhanced through components such as the Conv2d layer, BatchNorm2d layer, and activation function.

[0031] 4) Spatial Pyramid Pooling Fast (SPPF) module. The SPPF module is used to process the feature map output by the Backbone. Through a series of convolution and pooling operations, the input feature map is transformed into a multi-scale feature representation, enabling the model to more effectively handle multi-scale object detection tasks, thereby improving the overall detection performance.

[0032] S32. Optimize the Neck part. The Neck structure is an important part of this model and plays a crucial role in feature extraction and fusion. The Neck part follows the concept of PAN-FPN in YOLOv8, integrating the feature maps output from multiple stages of the feature extraction backbone network into the PAN-FPN network structure to achieve robust multi-scale feature fusion and improve object detection performance. Considering that the feature maps generated by the feature extraction backbone network are relatively small and may not be easily able to detect tiny and dense objects, the PAN-FPN module is extended, and a feature fusion layer dedicated to tiny object detection is added after the first LHG block of the backbone network. Subsequently, the feature fusion layer responsible for detecting large objects in the Neck is removed, but three feature maps are still retained for the detection head to output in order to identify tiny, small, and medium-sized objects. This is beneficial for feature extraction and fusion across multiple stages, thereby improving the model's feature representation ability. At the same time, this method integrates GSConv and introduces a more efficient and lightweight LGSbottleneck and LGSCSP fusion module to make the model more lightweight. The specific design of the Neck part is as follows: 1) Upsample module. In the Neck stage of network training, by upsampling and fusing the feature maps of each layer, the model's ability to recognize small objects is enhanced.

[0033] 2) Concat module. The main function of the Concat module is to merge the feature maps from different levels in the channel dimension. This can integrate information of different scales and semantic levels, enhance the model's understanding ability of objects. By splicing the feature maps of different levels, the Concat module can provide more comprehensive information to the subsequent network layers, contributing to improving the accuracy and robustness of object detection.

[0034] 3) LGSCSP fusion module. It is designed through a one-time aggregation method, ensuring the accuracy of the model while simplifying the calculation and network structure. The LGSCSP fusion module improves the model's performance in recognizing multi-size objects by optimizing the feature fusion process. As Figure 6 shown, the LGS bottleneck module is integrated into the LGSCSP fusion module as a key component inside it. Inside the LGSbottleneck, the initial feature map is respectively subjected to feature extraction through standard convolution and GSConv, and then the obtained feature maps are concatenated. This not only retains the image features and channel information as much as possible but also further reduces the computational complexity. In this solution, the LGS bottleneck is responsible for further compressing and extracting features during the feature fusion process, while the LGSCSP fusion module is responsible for overall feature fusion and cross-scale information integration.

[0035] 4) ConvModule convolution module. It is mainly responsible for extracting feature data from the input image, and performing multiple convolution operations on the input image through convolution operations to capture the features in the image.

[0036] S33. Head part. The Head part is the output layer of the model, which is responsible for performing the final detection of the target based on the feature map passed from the previous network. It infers the coordinates of the bounding box, the confidence of the target, and the classification probability of the object through a series of convolution operations, and then filters the overlapping prediction boxes through non-maximum suppression (NMS) to output the final detection result.

[0037] S4. Model evaluation. Verify the performance of the model on an independent test set, evaluate it through indicators such as accuracy and recall, and make fine adjustments to the model based on the evaluation results to ensure its accuracy and robustness in different aerial photography scenarios.

[0038] S5. Actual detection and result acquisition. Input the aerial photography dataset to be tested into the trained and optimized detection model, and perform real-time analysis on the input aerial photography image through the model to detect small objects in the image and determine their positions. The output results will cover the categories and positions (bounding boxes) of the objects, providing necessary information for subsequent processing.

[0039] The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A target detection method in an aerial drone photography scene based on improved YOLOv8, characterized in that, It includes the following steps: S1. Data collection and preprocessing: Use a drone to collect aerial image data under different conditions, and preprocess the collected images, including cropping, scaling, and color correction, to adapt to the model input requirements; S2. Dataset annotation and construction; Classify the processed data according to whether it is annotated, and construct a training dataset containing the positions and categories of target objects, as well as a dataset for model validation and testing; S3. Model training and optimization: For the improved YOLOv8 model, use the labeled training dataset to train and optimize the model, and learn the features of the targets in the images; S4. Model evaluation: Validate the performance of the model on an independent test set, evaluate it through accuracy and recall metrics, and make minor adjustments to the model based on the evaluation results; S5. Actual detection and result acquisition: Input the aerial image dataset to be tested into the trained and optimized detection model, and use the model to perform real-time analysis on the input aerial images, detect small objects in the images, and determine their positions; the output results will cover the categories and positions of the objects, providing necessary information for subsequent processing.

2. The object detection method in the UAV aerial photography scene based on the improved YOLOv8 according to claim 1, characterized in that, The specific design of the improved YOLOv8 model in step 3) is as follows: S31. Construct the LHGNet backbone network, enhance the feature extraction ability of the backbone network by introducing residual technology, and on the basis of the existing HGNetv2 backbone network, incorporate depthwise separable convolutions into the backbone network; S32. Optimize the Neck part: Expand the PAN-FPN module, add a feature fusion layer for small target detection after the first LHG block of the backbone network, and then delete the feature fusion layer in the Neck responsible for detecting large objects, retaining three feature maps for the detection head to output; S33. Head part: The Head part is responsible for the final detection of targets based on the feature maps passed from the previous network, infers the coordinates of the bounding boxes, the confidence of the targets, and the classification probabilities of the objects through a series of convolutional operations, and then filters out overlapping prediction boxes through non-maximum suppression to output the final detection results.

3. The object detection method in the UAV aerial photography scene based on the improved YOLOv8 according to claim 2, wherein, In step S31, the LHGNet backbone network includes the following modules: 1) HGStem module: The initial image is processed by the HGStem module, which is mainly used for upsampling and downsampling the features of the input image; the HGStem module receives the input image and processes it through consecutive convolutional and pooling steps, finally generating a set of feature maps to provide useful feature representations for subsequent network layers; 2) LHG module: All 3x3 standard convolutional units in the original HGNetv2 architecture are replaced with two sets of depthwise separable convolutional units. The LHG module consists of two depthwise separable convolutional units arranged in sequence, which perform a cascading operation, and finally are processed by two 1x1 standard convolutional units; 3) ConvModule convolutional module: Responsible for extracting the features of the input image; 4) SPPF module; The SPPF module is used to process the feature map output by the Backbone. Through a series of convolution and pooling operations, the input feature map is transformed into a multi-scale feature representation.

4. A target detection method in an aerial drone photography scene based on improved YOLOv8 according to claim 2, characterized in that, In step S32, the design of the Neck part is optimized as follows: 1) Upsample module; In the Neck stage of network training, the feature maps of each layer are upsampled and fused; 2) Concat module; The Concat module merges the feature maps from different levels in the channel dimension, integrates information at different scales and semantic levels, enhances the model's understanding ability of the target. By splicing the feature maps of different levels, the Concat module can provide more comprehensive information to the subsequent network layers; 3) LGSCSP fusion module; Designed by a one-time aggregation method, while simplifying the calculation and network structure, it ensures the accuracy of the model; The LGSCSP fusion module improves the performance of the model in identifying multi-size targets by optimizing the feature fusion process; The LGS bottleneck module is integrated into the LGSCSP fusion module; 4) ConvModule convolutional module; Responsible for extracting feature data from the input image, and performing multiple convolution operations on the input image through convolution operations to capture the features in the image.