YOLOv8-based optical remote sensing image small target detection method
Through the improved method based on YOLOv8, the model of small object detection of optical remote sensing images is optimized, which solves the problem of limited detection accuracy of the existing technology in complex backgrounds, and achieves higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510258012.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art has poor detection effect on small targets in optical remote sensing images, especially in complex backgrounds, and it is difficult to effectively capture the multi-scale features and complex details of small targets, resulting in limited detection accuracy.
Based on YOLOv8's improved optical remote sensing image small object detection method, the model's perception ability to small objects and adaptability in complex remote sensing scenarios includes the introduction of D-C2f modules and deformable convolution modules with feature fusion operations in the Backbone part, the introduction of a hybrid dual attention mechanism module in the Neck part, and the rotation bounding box and a new WDA IoU loss function are adopted in the Head part.
It significantly improves the accuracy and robustness of small object detection in optical remote sensing images, reduces missed and missed detection, enhances the ability to identify and position small objects, and is suitable for complex environments and real-time application scenarios.
Smart Images

Figure CN120088463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a small target detection method for optical remote sensing images improved based on YOLOv8, belonging to the field of target detection in optical remote sensing images. Background Technique
[0002] Optical remote sensing images can obtain stable data without being interfered by the ground environment, and with their wide single-shot imaging coverage, they can quickly capture a large amount of target information. In addition, the unique imaging angle of remote sensing images enables them to clearly show the relative position relationships between objects, such as the spatial layout between vehicles, roads, and buildings. This characteristic enables remote sensing images to accurately classify and locate targets in target detection, and they are widely used in military reconnaissance, maritime rescue, environmental monitoring, urban planning and other fields, having significant practical application value.
[0003] Remote sensing images have significant characteristics such as large format, high resolution, and high contrast. Especially for optical remote sensing imaging, its application scenarios are the most extensive and the target categories are the richest. However, target detection in remote sensing images itself is of high difficulty, especially facing many challenges in small target detection. For example, small targets are not only numerous and densely distributed, their texture features are often not obvious, the scene is complex and changeable, and there is a lot of background interference at the same time. These factors make the target detection task more difficult. In addition, there are often significant performance differences between existing methods when detecting tiny targets and regular-sized targets, mainly because the feature expression of small targets is relatively weak, and the complex remote sensing scene further exacerbates this problem, posing higher requirements on the existing technology.
[0004] The existing target detection technologies still have significant limitations in detecting small targets such as cars in remote sensing images. These methods usually rely on single-scale feature representations and are difficult to effectively capture the multi-scale features and their complex details of small targets, especially performing poorly in complex backgrounds. Since the single-scale attention mechanism cannot fully distinguish small targets from their surrounding environments, the detection accuracy is significantly restricted. Therefore, conducting in-depth research on small target detection in optical remote sensing images has important academic value and practical application significance.
[0005] This research combines the characteristics of small targets in remote sensing images and proposes an improved detection algorithm based on YOLOv8, aiming to improve the detection accuracy and robustness and achieve more accurate small target recognition and positioning. This method not only optimizes the model's perception ability for small targets but also enhances its adaptability in complex remote sensing scenes, having important value for promoting the development of small target detection technology in remote sensing images. Summary of the Invention
[0006] The present invention provides a detection method for small targets in optical remote sensing images improved based on YOLOv8, which has more accurate detection results when dealing with the characteristics of complex backgrounds, small target sizes, dense distributions, and arbitrary orientations in optical remote sensing images.
[0007] In view of the deficiencies of the prior art, this application proposes an improved detection method for small targets in optical remote sensing images, which is optimized and implemented based on YOLOv8. The method includes the following steps: First, obtain the optical remote sensing image dataset and preprocess it, including image cropping, removing blurred and duplicate images, to improve data quality. Subsequently, classify and label the preprocessed dataset, and the labeled categories cover common remote sensing targets, such as ships, airplanes, cars, trucks, and other target categories. Finally, divide the dataset into a training set, a validation set, and a test set to ensure the scientificity and effectiveness of model training and evaluation.
[0008] After completing the dataset preparation, build and train the YOLOv8 model. Adjust the hyperparameters (such as learning rate, batch size, number of training epochs, etc.) according to the task requirements, and use the training set to iteratively optimize the model weights to minimize the prediction error. During the training process, evaluate the model performance in real time through the validation set to prevent overfitting. Finally, use the test set to evaluate the detection accuracy, recall rate, and mAP of the model. The optimized YOLOv8 model has significant advantages in the detection of small targets in optical remote sensing images. It improves the detection accuracy, reduces missed detections and false detections, enhances the recognition ability of targets such as ships and airplanes, and provides reliable data support for subsequent analysis. At the same time, the model has a fast detection speed, is suitable for real-time application scenarios, and has good generalization ability, can adapt to different resolutions and complex environments, and improves the overall level and application value of small target detection in optical remote sensing images.
[0009] In the Backbone part, replace the C2f modules of the last two layers (P3 and P4 layers) with D-C2f modules introducing feature fusion operations, and replace the conventional convolution in the detection layer with a deformable convolution module. At the same time, perform weighted fusion on the output feature maps of each layer of the backbone network, and input the generated new feature maps into the SPPF structure to enhance the feature expression ability. In the Neck part, introduce a hybrid dual attention mechanism module (MDA) integrating deformable convolution after each upsampling to enhance the recognition and localization accuracy of the model for small targets. In the Head part, use the OBB (Oriented Bounding Box) bounding box to replace the original horizontal bounding box to reduce the anchor box area and improve the adaptability of the target box. In addition, introduce a new WDA IoU loss function to optimize the deviation between the candidate box and the ground truth box, and further improve the regression accuracy.
[0010] The D-C2f module splices the feature maps output by two adjacent Bottleneck modules and applies the SiLU activation function for further fusion. This method is simple to operate, has a relatively small computational cost, can smooth the feature maps, and can more effectively fuse feature information at different scales. By replacing the last two C2f modules in the backbone network with the improved D-C2f module and performing weighted accumulation on the output feature maps of each layer, a feature map that fuses high-level and low-level information is finally generated and input into the SPPF structure. This method effectively alleviates the problems of false detection and missed detection caused by insufficient low-level feature information and improves the accuracy of small target detection.
[0011] The attention mechanism module not only considers the optimization of light and shadow changes and target multi-scale problems, but also fully considers the dependence relationship between spatial attention and channel attention in optical remote sensing images with complex backgrounds and small-scale targets. Based on the structure of DA-Net, in the improved module, the feature map output by the adder of the spatial attention module is input to the multiplication calculation position of the channel attention module. At the same time, the feature map output by the adder of the channel attention module is also input to the corresponding position of the spatial attention module to achieve two-way interaction. Finally, the outputs of the two modules are weighted and fused to enhance the feature expression ability. In addition, in the Neck part, the MDA module is introduced after each upsample structure, which helps to alleviate the problem of feature aliasing, reduce the interference weight of background information, and strengthen the attention to the target area.
[0012] The deformable convolution module can automatically adjust the receptive field size according to needs, improving the adaptability of the network to irregular targets. In the prediction box regression stage, this module can more accurately optimize the regression parameters, enhance the attention to small targets, and thus improve the overall detection performance and robustness of the model.
[0013] The WDA IoU loss function models the rotated bounding box using the differentiable Wasserstein distance. Compared with the IoU loss, the main advantage of the Wasserstein distance is that it can accurately measure the distribution similarity of two bounding boxes, even if they do not overlap or contain each other.
[0014] As can be seen from the above, the advantages of the present invention are as follows: The present invention provides a small target detection method for optical remote sensing images improved based on YOLOv8. This method performs data augmentation processing on remote sensing image data to obtain multiple small target images with annotations. Subsequently, the enhanced image dataset is input into the small target detection model for optical remote sensing images, and this model uses an algorithm improved based on YOLOv8 for training and parameter optimization. Finally, the predicted image is input into the trained small target detection model for optical remote sensing images to detect the predicted small target information and obtain the detected image with annotations. The improved model can improve the model detection accuracy, thus solving the problem that traditional algorithms cannot accurately detect in the face of characteristics such as complex backgrounds, small and densely distributed target sizes, and variable target orientations in optical remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 It is a schematic flow chart of a small target detection method for optical remote sensing images improved based on YOLOv8 provided by the present invention;
[0017] Figure 2 It is a schematic diagram of a network structure provided by the present invention;
[0018] Figure 3 It is a schematic diagram of the structure of a D-C2f module provided by the present invention;
[0019] Figure 4 It is a schematic diagram of the structure of a convolutional attention mechanism MDA module provided by the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In the following description, specific details such as specific system structures and technologies are proposed only to help better understand the embodiments of the present application, and are not limitations on the scope of implementation. However, those skilled in the art should understand that other embodiments of the present application can also be implemented without relying on these specific details. In some cases, to avoid unnecessary details from affecting the description, the present application omits the detailed descriptions of well-known systems, devices, circuits, and methods.
[0021] It should be understood that the term "comprising" used in this specification and the appended claims is intended to indicate the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof. Therefore, the term "comprising" should be understood as having an open nature, allowing the addition of other relevant features or components without departing from the scope of the present application.
[0022] Next, the technical solutions will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the present application, not all embodiments. Based on the embodiments in the present application, other embodiments obtained by those skilled in the art without creative efforts all fall within the scope of protection of the present application.
[0023] Embodiment
[0024] Dataset: The DOTA dataset is used as the dataset for small target detection in remote sensing images in the present invention. The dataset has a total of 2,806 images of 4000×4000, containing a total of 188,282 targets.
[0025] Data augmentation is performed on the images in the dataset. First, operations such as adding Gaussian noise and color transformation are performed on the image data machine, which helps to improve the robustness of the model; through operations such as sliding cropping, geometric rotation, and scaling transformation, the number of training samples can be increased.
[0026] The Labellmg image annotation tool is used to annotate the targets of the images, annotating the category and location information of the targets. Finally, the dataset is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1 for model training, tuning, and evaluation.
[0027] In this embodiment, we adopt YOLOv8 as the baseline model and optimize and improve its structure to enhance the detection ability of small targets in remote sensing images.
[0028] Backbone part: Replace the C2f modules in the last two layers (P3 and P4 layers) with D-C2f modules that introduce feature fusion operations. Fusing the output feature maps of two adjacent bottleneck modules has simple operations, small computational complexity, and smoother results. Replace the ordinary convolution in the detection layer with a deformable convolution module, using the same prediction offsets on different channels of the same feature map to enhance the model's ability to learn the invariance of complex targets. At the same time, weighted fusion is performed on the output feature maps of each layer of the backbone network to generate new feature maps and input them into SPPF to strengthen the multi-scale feature expression ability.
[0029] Neck part: Introduce a Mixed Dual Attention mechanism module (MDA) that incorporates deformable convolutions to enhance the recognition and localization accuracy of small targets. This operation is performed every time there is an upsampling.
[0030] The MDA attention mechanism module consists of two parallel parts: spatial attention and channel attention, to fully exploit and utilize key information and improve the feature representation ability. The specific process is as follows:
[0031] Feature segmentation: After the input feature map is processed by the Conv module, the channels are divided into two parts and fed into the spatial attention mechanism and the channel attention mechanism respectively.
[0032] Attention calculation: The spatial attention mechanism focuses on the spatial distribution information of the feature map to enhance the model's ability to locate the target area. The channel attention mechanism can adaptively adjust the importance of different channels to improve the discrimination ability of feature representation.
[0033] Feature fusion: The output feature maps of the two modules are fused through element-wise multiplication to strengthen the interaction of multi-level information.
[0034] Final integration: The fused features are further subjected to the Sum fusion operation to fully aggregate the information of spatial and channel attention, so as to enhance the overall perception ability of the model.
[0035] The Head part will use Oriented Bounding Boxes (OBBs) instead of horizontal bounding boxes to reduce the anchor box area and enhance the adaptability of the detection boxes. Additionally, a new WDA IoU loss function is introduced to optimize the matching deviation between the candidate boxes and the ground truth boxes to improve the regression accuracy and detection accuracy.
[0036] The WDA-IoU loss function combines a two-dimensional Gaussian distribution and a differentiable Wasserstein distance to model the oriented bounding boxes. Compared with the traditional IoU loss, this method can accurately measure the distribution similarity between bounding boxes even when there is no overlap or inclusion relationship between them, thereby improving the accuracy and robustness of object localization.
[0037] By combining model learning with precision measurement, it is possible to solve the problems of inconsistent measurement and loss function in detection and the boundary problem of rotated rectangular boxes, and improve the model performance and detection accuracy. The calculation formula is as follows:
[0038] The hardware of the experimental environment for this experiment is as follows: CPU processor Intel Core i9-12900H; memory is 64GB DDR4 (3200MHZ); GPU graphics card GeForce RTX 3090; storage hard disk is 1024GB PCIe 4.0 NVMe solid state drive; the software of the experimental environment is: operating system Ubuntu 20.04 development environment: programming software is Python 3.8, and the model framework is pytorch 2.0.1.
[0039] In summary, this application proposes a small target detection method for optical remote sensing images improved based on YOLOv8. Aiming at the problems of misdetection, missed detection, and excessive regression loss of targets caused by factors such as complex backgrounds, small target sizes, dense distributions, and variable directions in optical remote sensing images, we designed a new improved model. This model provides an innovative detection method, which can significantly improve the detection accuracy and efficiency of optical remote sensing images, has high practical application value, and is especially suitable for real-time target detection tasks.
[0040] The above embodiments only show a specific implementation manner of the present invention. Although it has been described in detail, it does not limit the patent protection scope of the present invention. It should be particularly noted that for those of ordinary skill in the art, without departing from the basic concept of the present invention, various deformations and improvements can be made to it, and all these deformations and improvements should be regarded as the content within the protection scope of the present invention.
Claims
1. A small target detection method for optical remote sensing images based on improved YOLOv8, characterized in that: include: S1: A dataset of small targets in remote sensing images is manually annotated and divided into training and test sets in proportion. The images are then preprocessed and data enhanced to create a high-quality optical remote sensing image dataset. S2: Construct a small target detection network for optical remote sensing images, replace the P3 and P4 layer C2f modules of the YOLOv8 backbone network with the D-C2f module, add the output feature maps of each layer and input them into SPPF, add the MDA attention mechanism module after upsampling at the neck end, and introduce WDA IoU as a new loss function; S3: Pre-train the constructed optical remote sensing image small target detection network and optimize the model parameters; S4: Send the processed test set to the target network model for testing, and then output the detection results of small targets in the optical remote sensing image.
2. The preprocessing requirements for the training set images as described in claim 1 include data augmentation operations, such as rotation, scaling, color transformation, etc., to enhance the generalization ability of the model and improve the detection effect of small targets.
3. The optical remote sensing image small target detection method according to claim 1, characterized in that: The D-C2f module includes: performing pairwise concatenation operations on the feature maps output by adjacent bottleneck modules and applying SiLu activation functions. In the P3 and P4 layers, deformable convolution is introduced as the detection layer to enhance the model's detection performance and adaptability to small targets.
4. The optical remote sensing image small target detection method according to claim 1, characterized in that: After adding the output feature maps of each layer of the backbone network, the obtained feature maps integrating high- and low-level information are input into the SPPF.
5. The optical remote sensing image small target detection method according to claim 1, characterized in that: The feature extraction network includes two upsampling operations, and an attention mechanism module MDA is added after each upsampling at the neck end to enhance the model's ability to recognize and locate small targets.
6. The optical remote sensing image small target detection method according to claim 5, characterized in that: The MDA attention mechanism module includes Conv module, spatial attention, channel attention and Sum fusion.
7. The optical remote sensing image small target detection method according to claim 5, characterized in that: In the MDA module, after the feature map passes through the Conv module, the channel is divided into two parts, which are respectively sent to the spatial attention mechanism and the channel attention mechanism. The output feature maps of the two parts are fused by element-by-element multiplication, and finally summed (Sum fusion) is performed to obtain the final feature map output.
8. The optical remote sensing image small target detection method according to claim 1, characterized in that: The loss function module WDAIoU replaces the CIoU loss function in the original YOLOv8 network and is used for regression loss calculation of the rotation detection frame.
9. The optical remote sensing image small target detection method according to claim 1, characterized in that: The training set after data enhancement is input into the network model for training, and the NAG optimizer is used for multiple rounds of iterative optimization to finally obtain a converged network model.