High-resolution remote sensing image target detection method based on improved YOLOv8s
By improving the feature extraction and fusion methods of the YOLOv8s model, the problems of accuracy and computational cost in small target detection in remote sensing images were solved, and efficient and accurate small target detection was achieved.
Patent Information
- Application Number
- CN202510100422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-01-22
AI Technical Summary
In high-resolution remote sensing images, the detection accuracy of small targets is insufficient. Traditional methods are difficult to adapt to diverse image data, feature extraction and feature fusion are insufficient, and the computational load is large. The YOLOv8s model is prone to missing or misjudging small targets and consumes too much computational resources.
We employ the C2f-PD module based on PConv and DSConv for feature extraction, introduce the En-Fusion module based on spatial attention for feature fusion, and design a lightweight detection head module that replaces some convolution operations with PConv to reduce computational complexity.
It significantly improves the detection accuracy of small targets, reduces the computational load, and enhances the model's detection performance in remote sensing images, making it suitable for edge terminal devices with limited computing resources.
Smart Images

Figure CN120032266B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of remote sensing images and target detection, and relates to a high-resolution remote sensing image target detection method based on an improved YOLOv8s. BACKGROUND
[0002] Target detection, as a core research hotspot in the field of computer vision, is of paramount importance. It not only plays a key role in various cutting-edge application fields such as autonomous driving, object recognition, video surveillance, and robot navigation, but also cannot be separated from the continuous innovation and optimization of target detection technology for the further development of these fields. The core task of target detection is to accurately identify and locate specific objects in images or video frames, which may include pedestrians, vehicles, animals, buildings, and various entities, while requiring algorithms to effectively distinguish these objects from complex background information.
[0003] Although the detection technology of large and medium-sized targets has made great progress, small target detection remains a highly challenging problem. In particular, in the special field of high-resolution remote sensing images, the difficulty of small target detection is further magnified. The targets in remote sensing images often exhibit large scale span, diverse morphology, and dense distribution, and the imaging method from a high altitude makes the occlusion relationship between targets more complex. These factors greatly increase the difficulty of detection, making it difficult for traditional large target detection technology to be directly applied to small target detection tasks.
[0004] Given the importance of small target detection in the field of remote sensing images, especially its wide application prospects in environmental monitoring, urban planning, disaster assessment, resource investigation, and other aspects, improving the detection performance of small targets in remote sensing images undoubtedly has extremely important research significance. This not only contributes to the further development of computer vision technology, but also provides more accurate and reliable data support for related fields, thereby facilitating the innovation and application of related technologies. Therefore, how to design a more efficient and accurate small target detection algorithm based on the characteristics of remote sensing images has become one of the key problems to be solved in the current computer vision field.
[0005] Traditional object detection algorithms usually manually design features and then use classifiers for object recognition. The core idea is to design a series of features that can describe the characteristics of the target based on human understanding and prior knowledge of the target, such as edges, textures, shapes, colors, etc. These manually designed features can capture some local information of the target and perform object recognition and positioning through classifiers. However, this method has obvious shortcomings in dealing with complex scenes. Because the traditional feature extraction method relies on expert knowledge, it is difficult to adapt to diverse image data, and the manually designed features are not suitable for targets with large scale and shape changes or occlusion. The end-to-end learning paradigm of deep learning enables the model to automatically learn deep features in images, significantly improving detection accuracy. In recent years, deep learning has made significant progress in remote sensing image target detection, and numerous improved algorithms have emerged. Deep learning target detection algorithms can be roughly divided into two categories: one is two-stage detection algorithms, and the other is single-stage detection algorithms. Two-stage algorithms, represented by Faster-RCNN, Mask-RCNN, and Cascade-RCNN, are known for their high accuracy but slow detection speed. To balance speed and accuracy, single-stage algorithms have emerged. YOLO and SSD are typical representatives of this category. They can complete target detection in one forward propagation by directly predicting the image, saving the step of generating candidate boxes, and have faster inference speed. This is particularly important for processing high-resolution remote sensing images, as remote sensing images are usually large in data volume and require real-time performance. YOLO algorithm is widely used in practical production activities, but its detection accuracy for small targets in complex remote sensing image data is limited, so further improvement is needed.
[0006] To solve the problem of insufficient detection accuracy of small targets, scholars have conducted a lot of research. For example, Kim et al. designed an adaptive scale perception network to solve the problem of target detection model size variation; Akyon et al. used a segmentation detection fusion method to improve the detection accuracy of small targets, but there is a difficult balance between accuracy and speed; Qu et al. combined feature fusion and dilated convolution, which enhanced the extraction of shallow features and improved the detection performance of small targets; Zhu et al. added a CBAM (Convolu-tional Block Atten-tion Module) module to increase the number of channels and used a small target detection head based on Transformer to improve detection accuracy, but the computational complexity and algorithm complexity increased significantly.
[0007] The above algorithms are based on general frameworks, so they have good generalization ability for target detection on regular datasets. However, for small target detection, due to the small size of the target, feature loss may occur during feature extraction.
[0008] Target detection in remote sensing images, such as cars, ships, etc., has important application value in environmental monitoring, traffic management, disaster assessment, and national security, etc. With the rapid development of remote sensing technology, high-resolution remote sensing images provide rich data sources for target detection. However, targets in remote sensing images, such as cars and ships, usually have small size, dense distribution, and complex background, which brings great challenges to target detection. Therefore, how to improve the detection accuracy of YOLOv8 algorithm while reducing the computational complexity. Because the model, in feature extraction, feature fusion, etc. There are some problems.
[0009] ①Problems in feature extraction
[0010] YOLOv8 has the following challenges for remote sensing image target detection: first, remote sensing images contain a large number of vehicle, ship and other targets in the form of strips, which are densely distributed in the dataset images, small in area, and vary in length; second, the receptive field of YOLOv8s model is relatively large, which causes the features of small targets to be submerged in background information when detecting small targets, resulting in missed detection or misjudgment as background. The C2f module in YOLOv8s is used for feature extraction and performs well in many target detection tasks, but its convolution kernel size is fixed, which may be limited by the fixed pattern in processing specific scenarios. Moreover, the receptive field is limited, making it difficult to fully express the complex features of the target and model global or long-range dependencies, affecting detection performance. This results in the C2f module not being perfectly suitable for various scenarios.
[0011] ②Problems in feature fusion
[0012] Feature fusion is a key part of target detection. In convolutional neural networks, feature extraction gradually transitions from low-level local features to high-level semantic features as the network depth increases. With the continuous deepening of the down-sampling operation, the receptive field of the network gradually expands, the spatial resolution of the feature map gradually decreases, and the spatial geometry and other details of the feature map are gradually lost, while the semantic information expression ability of the feature map is higher. At the same time, in the problem of remote sensing image detection, the low-level feature map has edge and texture features, which can provide more accurate target position information. The edge information of cars, ships and other targets is more conducive to network detection and more suitable for small-scale target detection, but it lacks sufficient semantic content. Therefore, it is necessary to capture the special relevance between different layers of features through feature fusion and select more important feature layers to fully utilize low-level local features and high-level semantic features to improve detection effectiveness.
[0013] However, the traditional concat splicing operation is used in the original model, which is only a simple superposition of the input in the channel direction. Obviously, this fusion method is not sufficient for the fusion between high-level and low-level semantic features, and cannot meet the above requirements, which will directly affect the detection effect of the model. YOLOv8s needs other feature fusion methods for further processing. Especially in remote sensing images, there are many small targets, and they are easily blocked. Simple channel superposition only splices the feature maps in the spatial dimension, which does not remove feature redundancy, does not emphasize important features, and cannot fully capture the complex relationships and interactions between different layer features. This will make it difficult to extract the features of small targets with fewer pixels, making small target detection more difficult.
[0014] ③Large amount of calculation
[0015] Unlike the coupling structure of the YOLOv5 model detection head, the YOLOv8s detection head cancels the previous target branch and uses a decoupled classification branch and a regression branch, so that the target position and category information can be extracted separately and then learned through different network branches. Although this design of decoupled head can more effectively improve the detection accuracy, it also brings obvious challenges, resulting in a substantial increase in parameter quantity and calculation amount. Experiments show that the parameter quantity of the YOLOv8s detection head is 2.15M, and the calculation amount is 8.32FLOPs, accounting for 19.2% and 28.9% of the entire network, respectively. Obviously, there is a lot of room for improvement, and the calculation amount needs to be reduced. SUMMARY
[0016] Therefore, the purpose of the present application is to provide a high-resolution remote sensing image target detection method based on improved YOLOv8s. In view of the problems of small target size, dense distribution and complex background region of remote sensing targets, the YOLOv8s model is optimized and improved.
[0017] To achieve the above purpose, the present application provides the following technical solutions:
[0018] A high-resolution remote sensing image target detection method based on improved YOLOv8s, which performs the following steps based on the improved YOLOv8s model:
[0019] S1, obtaining high-resolution satellite optical remote sensing map data, and preprocessing the data;
[0020] S2, inputting the preprocessed data into the backbone network of the improved YOLOv8s model for feature extraction, wherein a plurality of C2f-PD modules based on PConv and DSConv are used for multi-scale feature extraction;
[0021] S3, input the multi-scale features into the neck network of the improved YOLOv8s model for feature fusion, wherein an enhanced fusion module with double inputs and single output based on spatial attention is used to strengthen the fusion degree between high-level and low-level features;
[0022] S4, input the fused features into the head network of the improved YOLOv8s model for target detection, wherein a lightweight detection head based on PConv is used for target detection.
[0023] Further, in step S1, the preprocessing step includes manual screening, target labeling, format conversion, data set enhancement, and data set division; wherein, first, download high-resolution satellite optical remote sensing map data and manually screen out target-containing regions in various backgrounds, crop and save the target regions, then label, save the labeling information as a YOLO format file, and perform Mosaic data enhancement on the image; build a network for training to obtain a weight file; finally, divide the image into a training set, a test set, and a validation set according to a preset ratio.
[0024] Further, in step S2, the C2f-PD module is used for feature extraction in the last three convolutions of the backbone network, wherein the C2f-PD module includes an input Conv block, a Split block, a plurality of PD-Bottleneck blocks, a Concat block, and an output Conv block, each PD-Bottleneck block includes a partial convolution PConv and a dynamic snake convolution DSConv, the PConv convolution preliminarily extracts spatial features, and then the DSConv convolution enhances the feature extraction capability; a plurality of PD-Bottleneck blocks are connected in series to extract multi-scale feature information.
[0025] Further, in each PD-Bottleneck block, the PConv convolution utilizes the redundancy of feature mapping between different feature maps to perform convolution operation on only part of the channels of the input feature map to extract local spatial features; and the calculation complexity is reduced by taking the first or last continuous channel as a representative of the entire feature map, and the calculation process is as follows:
[0026] First, mask processing is performed on the input features:
[0027] x m =X⊙M
[0028] Where X is the input feature, M is a binary mask, and represents element-wise multiplication.
[0029] Then, partial convolution operation is performed on the mask-processed features:
[0030]
[0031] where p is the current position, W is the convolution kernel weight, k is the convolution kernel size, M(p) is the mask value at position p, i represents the index of the element within the convolution kernel, M(p) i represents the mask value at the i-th position within the neighborhood of position p, (·) T represents the transpose operation;
[0032] Mask update rule:
[0033]
[0034] where the mask update is usually performed after the partial convolution operation, and the updated mask will be used for the next partial convolution operation, gradually expanding the effective area.
[0035] Further, in each PD-Bottleneck block, the DSConv convolution dynamically adjusts the size and direction of the convolution kernel according to the shape and geometric characteristics in the image to enhance the feature extraction capability, wherein the DSConv convolution continuously adjusts the shape and position of the convolution kernel by introducing a learnable deformation offset, adopts an iterative learning method, so that the convolution kernel can adaptively adjust the shape and size, thereby capturing targets of different scales and shapes, and the specific process is as follows:
[0036] First, a standard two-dimensional convolution (N x N) with a coordinate P, denoted as the center coordinate P i = (x i , y i ), where i ∈ [0, N], and the calculation formula of a 3 x 3 convolution kernel coordinate is:
[0037] P = {(x-1, y-1), (x-1, y), …, (x+1, y+1)}
[0038] In the formula, P is the coordinate of the standard two-dimensional convolution; x and y represent the horizontal and vertical offsets of the convolution kernel relative to the center position, respectively;
[0039] The dynamic snake convolution adjusts the convolution kernel in the x-axis and y-axis directions, wherein the change in the x-axis direction is:
[0040]
[0041] In the formula, c represents the horizontal distance from the center grid, Δy represents the offset in the vertical direction, and subscript i represents the index of the convolution kernel on the input feature map.
[0042] The change in the y-axis direction is:
[0043]
[0044] wherein subscript j represents the index of the convolution kernel on the input feature map, and Δx represents the offset in the horizontal direction;
[0045] The floating-point coordinates in the offset results in the x direction and the y direction are mapped to the image coordinates by using a bilinear interpolation method, which is represented as:
[0046] P = ∑ P′ B(P', P)·P'
[0047] wherein P represents the decimal position in the offset results in the x direction and the y direction, and P' enumerates all integer space positions; B is a bilinear interpolation kernel, which can be decomposed into two one-dimensional kernels, i.e.:
[0048] B(P, P') = b(P x ,P x ′)b(P y ,P y ′)
[0049] wherein b(·) represents a one-dimensional interpolation kernel function, P x ,P x ′ respectively represent the decimal part and the integer part of P in the horizontal direction, and P y ,P y ′ respectively represent the decimal part and the integer part of P in the vertical direction;
[0050] With the changes of the x axis and the y axis, the receptive field of the dynamic snake convolution kernel is adjusted, so as to adaptively capture the geometric features of the target.
[0051] Further, in step S3, in the neck network of the YOLOv8s model, an enhanced fusion module is used to replace the Concat operation before the small target, and the processing process is as follows:
[0052] First, check whether the channel numbers c1, c2 of the input feature maps x1 and x2 are consistent, if consistent, directly splice, that is:
[0053] x = Concat(x1, x2)
[0054] wherein x is the spliced feature map;
[0055] If not consistent, adjust the channel number of x1 by 1×1 convolution to match x2, that is, ensure the compatibility of the two feature maps in the dimension, and then splice in the channel, that is:
[0056] x = Concat(Conv 1×1 (x1, c2), x2), if (c1≠c2)
[0057] After splicing, the feature map x is processed by a spatial grouping enhancement module:
[0058] x SGE = SGE(x)
[0059] where SGE(·) denotes a spatial group enhancement module;
[0060] Then, the enhanced feature map is split into two parts respectively corresponding to the number of channels of the original input feature map:
[0061] Out1 = Split(x SGE , [c2, c2])[0]
[0062] Out2 = Split(x SGE , [c2, c2])[1]
[0063] where Split(·) denotes a split function in the channel dimension;
[0064] Then, the split feature map is multiplied with the original feature map to obtain a further enhanced feature map, and the multiplication operation in this process introduces a nonlinear factor; finally, the first original feature map is added to the second enhanced feature map, and the second original feature map is added to the first enhanced feature map, and then they are spliced to obtain the integrated enhanced feature:
[0065] EnhanceFusion(x1, x2) = Concat(x1 + x2 ⊙ Out1, x2 + x1 ⊙ Out1)
[0066] where ⊙ denotes element-wise multiplication, and Concat denotes splicing in the channel dimension.
[0067] Further, in the SGE module of the feature fusion module, the SGE module groups the features into multiple sub-features along the channel direction, and processes each feature group in parallel, so as to capture different semantic features, and the steps include:
[0068] Firstly, the feature map with a dimension of CxHxW is divided into G groups along its channel, and each group is represented as a vector in the spatial dimension:
[0069] X = {x1, x2, …, x i , x m}, x i ∈ R G , m = HxW
[0070] where m denotes the total number of spatial positions;
[0071] Then, the spatial average function is used to strengthen the semantic modeling of the key area by using global statistical features as supplementary information:
[0072]
[0073] In the formula, F g (·) is a spatial average function;
[0074] Then, the importance weight coefficient of each feature is calculated, and the similarity between the global semantic feature g and the local feature x is quantified by a dot product operation, that is: i
[0075] c i = g·x i
[0076] The coefficient c is spatially normalized, and two learnable parameters are introduced for each coefficient, and the normalized value is linearly changed by adjusting the two parameters to realize the identity transformation, that is:
[0077] a i = αx i + β
[0078] In the formula, alpha is a scaling factor, and beta is an offset;
[0079] Finally, the output of the Sigmoid activation is multiplied by the original feature map to obtain an enhanced feature vector Finally, the result feature group is formed:
[0080]
[0081] In the formula, the enhanced feature vector is expressed as
[0082] Further, in step S4, the designed lightweight detection head module includes a Pconv convolution, a Conv standard convolution and a Conv2d convolution, wherein the PConv convolution is first used to perform convolution operation on only part of the channels of the input feature map, the PConv convolution removes the redundant information in the feature map on the basis of preserving the feature information flow, then the second Conv standard convolution is used for complete feature information extraction, and the last Conv2d convolution is used for channel adjustment to complete target detection.
[0083] The beneficial effects of the present application are:
[0084] The improved YOLOv8s model of the application significantly improves the feature extraction capability by introducing a C2f-PD module based on PConv (partial convolution) and DSConv (dynamic snake convolution), especially for the challenges of small target detection in remote sensing images. PConv effectively reduces redundant calculations and improves feature extraction efficiency by only performing convolution operations on part of the input feature map channels, while retaining sufficient feature information flow. DSConv enhances the model's feature extraction capability for small elongated structure targets by dynamically adjusting the shape and direction of the convolution kernel. This combination enables the C2f-PD module to more fully express the complex features of targets when processing complex targets in remote sensing images, solving the problem of limited receptive field of the original C2f module, which is difficult to fully model global or long-range dependencies. Experimental results show that this improvement increases mAP50 by 2.2% and mAP50:95 by 2.0%, significantly improving the detection accuracy of small targets.
[0085] In terms of feature fusion, the application designs an enhanced fusion module (En-Fusion) based on spatial attention with double input and single output. This module realizes better fusion of high and low level feature maps through flexible feature adjustment capability. The En-Fusion module first checks whether the channel numbers of the input feature maps are consistent. If not, it adjusts the channel number through 1x1 convolution to ensure dimension compatibility. Then, the spatial grouping enhancement module (SGE) is used to process the spliced feature map to enhance the performance of the feature map in spatial information, especially the attention to key regions in the image. Finally, element-wise multiplication and addition operations are used to realize the interactive fusion of different levels of features. This fusion method not only suppresses noise but also improves feature expression capability, solving the problem of insufficient fusion of traditional concat splicing operations. Experiments show that the introduction of the En-Fusion module further improves the detection performance of the model, especially when dealing with small targets and complex backgrounds in remote sensing images.
[0086] The application designs a lighter-detect module, which significantly reduces the computational complexity by introducing PConv to replace the first standard convolution in the detection head. PConv removes redundant information in the feature map by only performing convolution operations on part of the input feature map channels, while retaining enough feature information flow for subsequent convolution extraction. This design reduces the amount of calculation while ensuring the stability of detection accuracy. Experimental results show that the introduction of the lighter-detect module reduces the total computational amount of the model by 20.6%, while the detection accuracy hardly decreases. This improvement makes it possible to deploy the improved YOLOv8s model on edge terminal devices with limited computing resources, greatly improving the practicality and application range of the model. At the same time, compared with other mainstream target detection algorithms, the improved YOLOv8s performs well in both accuracy and computational amount, especially in processing high-resolution remote sensing images.
[0087] Other advantages, objects, and features of the application will be set forth in part in the following specification taken in conjunction with the accompanying drawings, and in part will become apparent to those skilled in the art from a consideration of the following specification and drawings. The objects and other advantages of the application will be realized and attained by means of the instrumentalities and combinations pointed out in the following specification. BRIEF DESCRIPTION OF DRAWINGS
[0088] In order to make the objects, technical solutions and advantages of the application clearer, the preferred detailed description of the application will be combined with the drawings to describe the application, in which:
[0089] Figure 1 The overall structure diagram of the improved YOLOv8s of the improved YOLOv8s-based high-resolution remote sensing image target detection method of the application;
[0090] Figure 2 The structure diagram of the C2f-PD module in the improved YOLOv8s model of the application;
[0091] Figure 3 The structure diagram of the PConv convolution in the C2f-PD module of the application;
[0092] Figure 4 The structure diagram of the enhanced fusion module in the improved YOLOv8s model of the application;
[0093] Figure 5 The structure diagram of the lighter-detect module in the improved YOLOv8s model of the application;
[0094] Figure 6 The overall flowchart of verifying the effect of the improved YOLOv8s model of the application in the embodiment;
[0095] Figure 7 Figure 7 is a schematic diagram of a heat map comparison result of the improved YOLOv8s model of the target one under the embodiment and the existing YOLOv8s model, wherein 7(a) is an input image, 7(b) is a heat map of the original YOLOv8s model, and 7(c) is a heat map of the improved model;
[0096] Figure 8 Figure 8 is a schematic diagram of a detection result comparison of the improved YOLOv8s model of the target two under the embodiment and the existing YOLOv8s model, wherein 8(a) is an input image, 8(b) is a heat map of the original YOLOv8s model, and 8(c) is a heat map of the improved model. DETAILED DESCRIPTION
[0097] The advantages and effects of the present application can be easily understood by those skilled in the art from the disclosure of the present specification. The present application can also be implemented or applied by different specific embodiments, and various modifications or changes can be made to the details in the specification based on different views and applications without departing from the spirit of the present application. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0098] The drawings are only used for illustrative explanation, and the representation is only a schematic diagram, not a physical diagram, and should not be understood as a limitation of the present application. In order to better illustrate the embodiments of the present application, some components in the drawings may be omitted, enlarged or reduced, and do not represent the actual size of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions in the drawings may be omitted.
[0099] The same or similar reference numerals in the drawings of the embodiments of the present application correspond to the same or similar components; in the description of the present application, it should be understood that the orientation or position relationship indicated by the terms "upper", "lower", "left", "right", "front", "back" and the like is based on the orientation or position relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore the terms describing the position relationship in the drawings are only used for illustrative explanation, and should not be understood as a limitation of the present application, for those skilled in the art, the specific meaning of the above terms can be understood according to the specific situation.
[0100] Please refer to Figures 1-8 , a high-resolution remote sensing image target detection method based on improved YOLOv8s.
[0101] Embodiment
[0102] The embodiment gives a detailed implementation process of a high-resolution remote sensing image target detection method based on an improved YOLOv8s, as shown in the network architecture diagram of the improved YOLOv8s model, based on the improved YOLOv8s model, the method comprises the following steps:
[0103] S1, obtaining high-resolution satellite optical remote sensing map data, and preprocessing the data;
[0104] S2, inputting the preprocessed data into the backbone network of the improved YOLOv8s model for feature extraction, wherein a plurality of C2f-PD modules based on PConv and DSConv are used for multi-scale feature extraction;
[0105] S3, inputting the multi-scale features into the neck network of the improved YOLOv8s model for feature fusion, wherein a dual-input single-output enhanced fusion module based on spatial attention is used to strengthen the fusion degree between high and low layer features;
[0106] S4, inputting the fused features into the head network of the improved YOLOv8s model for target detection, wherein a lightweight detection head based on PConv is used for target detection.
[0107] In step S1 of the embodiment, the preprocessing steps for high-resolution satellite optical remote sensing map data include manual screening, target labeling, format conversion, data set enhancement, and data set division.
[0108] Specifically, first, download the high-resolution satellite optical remote sensing map data and manually screen out the areas containing ship targets in various backgrounds, crop and save these areas, then label them, save the labeling information as a YOLO format file, and perform Mosaic data enhancement on the images; build a network for training to obtain a weight file; finally, select 80% of the images as the training set, 10% as the test set, and 10% as the validation set.
[0109] In step S2 of the embodiment, the C2f module in YOLOv8s is used for feature extraction and performs well in many target detection tasks, but its convolution kernel size is fixed, which may be limited by the fixed mode in processing specific scenarios, and the receptive field is limited, which is difficult to fully express the complex features of the target and difficult to model global or long-range dependencies, affecting the detection performance, which leads to the C2f module cannot be perfectly applied to various scenarios. In order to solve the above problems of the C2f module in the backbone network of YOLOv8s, the present application constructs a C2f-PD module based on PConv and DSConv.
[0110] The PConv (Partial Conv) and DSConv (Dynamic Snake Conv) are introduced to replace the first and second convolutions in the Bottlelock in the C2f module, respectively, to construct a C2f-PD module, and then the C2f-PD module is placed in the last three convolutions in the YOLOv8s backbone network to improve the backbone network.
[0111] As Figure 2 The structure diagram of the C2f-PD module is shown. The PConv partial convolution can effectively reduce redundant calculation and memory access, optimize the use efficiency of resources, and at the same time reduce the calculation cost while retaining the original channel information to the greatest extent, improve the inference speed, and show good effect in extracting spatial features. The DSConv dynamic snake convolution enhances the perception ability through a dynamic shape convolution kernel, controls the movement of the convolution kernel through an iterative strategy, and has stronger feature extraction capability. The first convolution in the Bottleneck in the C2f is replaced by PConv, and the second convolution in the Bottleneck is replaced by DSConv to form a PD-Bottleneck, and then the PD-Bottleneck is used to replace all Bottlenecks in the C2f, thereby designing a better feature extraction module C2f-PD. The PD-Bottleneck first speeds up the calculation process through PConv, and at the same time preliminarily extracts spatial features, and then perceives key areas through the dynamic convolution kernel of the dynamic snake convolution, more flexibly processes local information in the feature map, and more fully extracts the spatial features of the target. The two work together, and after the different scale features of the series PD-Bottleneck are fused, the multi-level information is used to improve the expression ability of the feature map, so that the C2f-PD module learns more rich and complex features, and solves the problem that the C2f module is difficult to fully extract target features in the scene of remote sensing images.
[0112] More specifically, for the PConv convolution, since different feature maps have high similarity, PConv utilizes the redundancy of these feature mappings. PConv only convolves a part of the channels of the input feature map, thereby effectively extracting local spatial features. For continuous or regular memory access, PConv reduces the computational complexity and improves the feature extraction efficiency by taking the first or last continuous channel as a representative of the entire feature map. Compared with the traditional convolution operation, PConv is more suitable for processing the characteristics of remote sensing images, such as Figure 3 The structure diagram of the PConv convolution is shown, and the calculation process is as follows:
[0113] First, the input features are masked:
[0114] xm =X⊙M
[0115] Where X is the input feature, M is the binary mask, and ⊙ represents element-wise multiplication;
[0116] Then, perform partial convolution operations on the masked features:
[0117]
[0118] Where p is the current position, W is the kernel weight, k is the kernel size, M(p) is the mask value at position p, and i represents the index of the element within the kernel. i Let (·) represent the mask value at position i within the neighborhood of position p. T Indicates the transpose operation;
[0119] Mask update rules:
[0120]
[0121] Mask updates are typically performed after partial convolution operations. The updated mask is then used in the next partial convolution operation to gradually expand the effective region.
[0122] For DSConv, because the receptive field of traditional convolutional neural networks is fixed, its feature extraction capability is insufficient when dealing with targets of varying scales and shapes, especially small targets with thin, elongated structures, thus limiting its generalization performance in target detection tasks. DSConv, on the other hand, can dynamically adjust the size and orientation of the convolutional kernel based on the shape and geometric characteristics of the image, thereby better handling targets in complex scenes and improving the feature extraction effect for difficult-to-detect targets. Replacing the original second ordinary convolution with a DSConv convolution in the Bottleneck of the C2f module enables it to more adaptively handle strip-shaped structures and complex backgrounds, which enhances the model's detection capability, performing particularly well when dealing with small targets.
[0123] Dynamic snake convolution introduces learnable deformation offsets to continuously adjust the shape and position of the convolution kernel. By adopting an iterative learning approach, the convolution kernel can adaptively adjust its shape and size, thereby flexibly capturing targets of different scales and shapes and effectively improving the model's feature extraction capability for complex targets.
[0124] First, a standard two-dimensional convolution (N×N) with coordinates P is given, denoted by center coordinate P. i =(x i ,y i ), where i∈[0,N]. The formula for calculating the coordinates of a 3×3 convolution kernel is:
[0125] P = {(x - 1, y - 1), (x - 1, y),..., (x + 1, y + 1)}
[0126] where P is the coordinate of the standard two-dimensional convolution; x, y represent the horizontal and vertical offset of the convolution kernel relative to the center position.
[0127] The dynamic snake convolution adjusts the convolution kernel in both the x-axis and y-axis directions. For a 9x9 convolution kernel, the specific position of each grid in p is represented as: P i±c = (x i±c ,y i±c ) where c = 0, 1, 2, 3, 4, represents the horizontal distance from the center grid. The selection of each grid position P i±c in the convolution kernel P is a cumulative process. After a 3x3 convolution, the offset map is obtained, and then the offsets are limited to [-1, 1] by the tanh function to obtain the final initial offset value δ. Starting from the center position K, the position away from the center grid depends on the position of the previous grid: P i+1 is increased by an offset Δ = {δ | δ ∈ [-1, 1]} relative to P i . Therefore, the offset needs to be accumulated to ensure that the convolution kernel conforms to the linear morphological structure. The change in the x-axis direction is:
[0128]
[0129] where Δy represents the offset in the vertical direction, and subscript i represents the index of the convolution kernel on the input feature map.
[0130] The change in the y-axis direction is:
[0131]
[0132] where subscript j represents the index of the convolution kernel on the input feature map, and Δx represents the offset in the horizontal direction.
[0133] Since image coordinates are discrete integers, and offsets Δ are usually floating-point numbers, a method is needed to map floating-point coordinates to image coordinates. Bilinear interpolation is a commonly used method that calculates the pixel value corresponding to the floating-point coordinate by performing linear interpolation on four adjacent pixel points. It can be represented as:
[0134] P = ∑ P′ B(P', P) · P'
[0135] where P represents the decimal position in the offset result in the x and y directions, P' lists all integer space positions, and B is the bilinear interpolation kernel, which can be decomposed into two one-dimensional kernels, i.e.:
[0136] B(P,P′)=b(P x ,P x ′)b(P y ,P y ′)
[0137] Where b(·) represents the one-dimensional interpolation kernel function, P x ,P′ x Let P represent the decimal and integer parts of P in the horizontal direction, respectively. y ,P′ y These represent the decimal and integer parts of P in the vertical direction, respectively;
[0138] As the x-axis and y-axis change, the receptive field of the dynamic serpentine convolution kernel can be flexibly adjusted, covering a maximum area of 9×9, thereby adaptively capturing the geometric features of the target, especially suitable for the recognition of strip-shaped structures. Based on these characteristics of dynamic serpentine convolution, this invention makes innovative improvements to the C2f module of the YOLOv8s model. Replacing the second ordinary convolution in the backbone module of the C2f module, while extracting feature map information from the PConv identity mapping, the shape of the convolution kernel can be flexibly adjusted to adapt to different target shapes and sizes. This is crucial for improving the model's ability to detect small targets and targets in complex backgrounds. Combining the characteristics of PConv and DSConv, this invention designs the C2f-PD module, which not only inherits the efficiency of PConv and the flexibility of DSConv, but also better captures complex structural features in images, thereby improving model accuracy.
[0139] In step S3 of this embodiment, an enhanced fusion module En-Fusion is used in the neck network, such as... Figure 4 The illustrated enhancement fusion module aims to fully utilize low-level feature information to achieve better fusion of high- and low-level feature maps, suppress noise, and thus improve feature representation capabilities, thereby enhancing model performance. This module possesses flexible feature adjustment capabilities, adapting to the processing needs of data at different scales while reducing the risk of overfitting and enhancing the model's expressive and generalization abilities in various visual tasks. When using the En-Fusion module, directly replacing all concat operations would lead to performance degradation and increased computational complexity. Given that the dataset in this embodiment contains a large number of small targets, only the concat operation before the small targets is replaced, thus achieving a balance between inference speed and accuracy.
[0140] The En-Fusion module first checks whether the channel numbers c1, c2 of the input feature maps x1 and x2 are consistent. If they are consistent, the feature maps are directly spliced, that is, x = Concat(x1, x2), where x is the spliced feature map; if they are not consistent, the channel number of x1 is adjusted through a 1x1 convolution to match x2, that is, to ensure the compatibility of the two feature maps in the dimension, and then the splicing in the channel is performed, that is, x = Concat(Conv 1×1 (x1, c2), x2), if (c1≠c2).
[0141] After splicing, the feature map x is processed by the spatial grouping enhancement module, which helps the deep network to learn the remote sensing image more carefully, strengthens the performance of the feature map in the spatial information, especially the attention to the key regions in the image, thereby improving the expression of the features. Then, the enhanced feature map is split into two parts, and the two parts correspond to the channel number of the initial input feature map respectively. Then, the split feature map is multiplied with the initial feature map to obtain a further enhanced feature map. The multiplication operation in this process introduces a nonlinear factor, so that the model can capture more complex interactions between features. Finally, in order to make the feature maps at different levels fully integrated, the first original feature map is added to the second enhanced feature map, and the second original feature map is added to the first enhanced feature map, so as to realize the interactive fusion of features at different levels, and then they are spliced to complete the effective integration and enhancement of the features. The formula is as follows:
[0142] x = Concat(x1, x2);
[0143] x SGE = SGE(x);
[0144] Out1 = Split(x SGE , [c2, c2])[0];
[0145] Out2 = Split(x SGE , [c2, c2])[1];
[0146] EnhanceFusion(x1, x2) = Concat(x1 + x2⊙Out1, x2 + x1⊙Out1)
[0147] where⊙ represents element-wise multiplication, Concat represents splicing in the channel dimension, and Split represents splitting in the channel dimension.
[0148] The SGE module is lightweight and helps deep networks enhance fine-grained image learning. The SGE module groups features into multiple sub-features along the channel direction and processes each feature group in parallel, thereby capturing different semantic features. The SGE module effectively reduces noise interference and enhances feature utilization efficiency, thus highlighting key semantic regions. The SGE module's workflow includes the following steps:
[0149] First, the feature map of dimension C×H×W is divided into G groups along its channels. Each group has a vector representation in the spatial dimension, i.e., X={x1,x2,…,x i ,…,x m},x i ∈R G m = H × W, thus enabling the learning of different semantic features, such as the outline and texture of a car. Subsequently, the spatial averaging function F... g (·), using global statistical features as supplementary information to enhance semantic modeling of key regions, as shown in the formula.
[0150]
[0151] Subsequently, the importance weight coefficient of each feature is calculated, and the global semantic feature g and the local feature x are quantified by dot product operation. i The degree of similarity between them, i.e., c i =g·x i To reduce the differences in coefficients among different samples, the coefficient c is spatially normalized to obtain c0. i Two learnable parameters are introduced for each coefficient: a scaling factor α and an offset β. By adjusting these two parameters, the normalized value is linearly varied, thereby achieving an identity transformation, i.e., α i =αx i +β.
[0152] Finally, the output of the sigmoid activation is multiplied by the original feature map to obtain the enhanced feature vector. Represented as The final result feature group is shown in the formula.
[0153]
[0154] In step S4 of the embodiment, a lightweight detection head is designed to reduce the computational complexity. Specifically, in the target detection network, there is redundant information in the multi-channel feature map, especially in the first convolution in the detection head during feature extraction. The full-channel 3x3 convolution does not fully utilize this redundancy and needs a second convolution for further extraction. PConv plays a role in improving the C2f module, and similarly, it can also be used to optimize the detection head module. Because PConv significantly reduces the computational complexity by only convolving a part of the input feature map channels compared to traditional convolution, while also reducing the number of memory accesses. More importantly, although PConv removes the redundant information of the feature map, PConv also retains enough feature information flow and does not remove effective features, which are fully extracted by the subsequent convolution through the identity mapping method. In the detection head module, there are three standard convolutions, so the first standard convolution is replaced by PConv to reduce the computational load, and the second standard convolution is used to fully extract feature information, which significantly reduces the floating point computation while ensuring that the accuracy does not change significantly. The improved detection head module is shown in FIG. 8, which includes a Pconv convolution, a Conv standard convolution, and a Conv2d convolution. First, the PConv convolution only convolves a part of the input feature map channels, and the PConv convolution removes the redundant information in the feature map while retaining the feature information flow. Then, the second Conv standard convolution is used to extract complete feature information, and the last Conv2d convolution is used for channel dimension reduction processing to complete the target detection. Figure 5
[0155] The embodiment performs experimental evaluation on the above-mentioned scheme. According to the evaluation standard of the YOLO series model, the following several widely used evaluation indicators are selected. The average precision value (mAP) is an index that comprehensively evaluates the precision and recall rate, which is calculated by the intersection over union (IoU) of the predicted box and the labeled box. mAP50 represents the mAP value of the true box and the predicted box IoU when the confidence is 0.5. mAP50:95 represents the average precision of the model in the IoU threshold range from 0.50 to 0.95. In addition, the network computation of the model is also a key indicator of the experiment, and GFLOPs (billion floating point operations) is used to measure the size of the network computation. At the same time, we also use the parameter quantity (Params) to measure the memory resource consumption of the model. The specific formula is as follows:
[0156]
[0157] where AP is the precision of a single class, and N is the actual number of classes.
[0158] Table 1 shows the improved YOLOv8s model of the application. The smiling experiment evaluation comparison results are as follows:
[0159] Table 1
[0160]
[0161] In Table 1, A represents the C2f-PD feature extraction module; B represents the En-Fusion feature fusion module; and C represents the Lighter-Detect lightweight detection head module. Compared with the baseline model, the mAP50 is improved by 2.2%, the mAP50:95 is improved by 2.0%, and the calculation amount is reduced by 20.6%. The experimental results show that the improved measures significantly improve the recognition accuracy of the model to the target, and jointly promote the improvement of the network performance.
[0162] In order to comprehensively evaluate the detection effect of the improved YOLOv8s algorithm of the application on remote sensing images, under the condition of ensuring the consistency of the experimental environment and various hyperparameters, the improved algorithm and the current mainstream target detection algorithm are compared on the SIMD dataset, including YOLOv5s, YOLOv7s, YOLOX-l, YOLOv10s, YOLOv11s, Faster-RCNN, RetinaNet, RVSA, SCRDet, S 2 Anet and YOLOv8s, and the experimental results are shown in Table 2.
[0163] Table 2
[0164]
[0165] As shown in Table 2, compared with YOLOv5s, YOLOv7s, YOLOX-l, YOLOv10s, YOLOv11s algorithms which are also single-stage, the mAP50 is improved by 2.5%, 11%, 9.8%, 2.1%, 1.7% respectively, and the mAP50:95 is higher by 2.1%, 10.4%, 9.8%, 1.8%, 1.2% respectively. Among them, the YOLOv11 algorithm is the latest algorithm of the same company as YOLOv8s. Obviously, the improved YOLO of the application has a significant advantage in mAP50 and mAP50:95, greatly improves the detection effect, and the calculation amount is only 22.8 GFLOPs, which is only 7% different from the smallest YOLOv11s algorithm, can be deployed on devices with limited computing performance, and has strong practicality. Compared with the two-stage algorithms Faster-RCNN, RetinaNet, RVSA, SCRDet, and S 2 Anet, the precision is greatly improved, and the calculation amount is obviously reduced, which has a considerable advantage.
[0166] The improved model YOLOv8s has an mAP50 of 2.2% higher than the baseline model YOLOv8s, an mAP50:95 of 2.0% higher, and a computational load of 20.6% lower, which better balances the requirements of computational load and accuracy compared to YOLOv8s, providing greater possibilities for model deployment on edge terminal devices.
[0167] As Figure 6 The validation flowchart of the improved model in the present embodiment is shown. First, high-resolution satellite optical remote sensing map data is downloaded and manually screened to obtain areas containing ship targets in various backgrounds. These areas are cropped and saved, then labeled, and the labeling information is saved as a YOLO format file. 80% of the images are selected as the training set, 10% as the test set, and 10% as the validation set. Then the images are subjected to Mosaic data augmentation. The network is built for training, and the weight file is obtained. The trained model file data is saved to test the test set. The model is validated, and ablation experiments are performed to generate heat maps and model detection effect comparison graphs. Finally, comparative experiments are performed with other models to evaluate precision and computational load.
[0168] Finally, the improved YOLOv8s model and the existing YOLOv8s model are used to detect target one and target two. The heat map comparison of target one is shown in Figure 7 , and the heat map comparison of target two is shown in Figure 8 , where 7(a), 8(a) are input images, 7(b), 8(b) are the heat maps of the original YOLOv8s model, and 7(c), 8(c) are the heat maps of the improved model. As can be seen from Figure 7 and Figure 8 , the improved heat map shows higher activation intensity in the target area, indicating that the model's attention to target features has significantly improved.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions, and they should be included in the scope of the claims of the present application.
Claims
1. An improved YOLOv8s-based high-resolution remote sensing image target detection method, characterized in that: The method performs the following steps based on the improved YOLOv8s model: S1, obtaining high-resolution satellite optical remote sensing map data, and preprocessing the data; S2, inputting the preprocessed data into the backbone network of the improved YOLOv8s model for feature extraction, wherein a plurality of C2f-PD modules based on PConv and DSConv are used for multi-scale feature extraction; S3, inputting the multi-scale features into the neck network of the improved YOLOv8s model for feature fusion, wherein a double-input single-output enhanced fusion module based on spatial attention is used to strengthen the fusion degree between high and low layer features; In step S3, in the neck network of the YOLOv8s model, the enhanced fusion module is used to replace the Concat operation before the small target, and the processing process is as follows: First check the input feature maps and the number of channels , are consistent, if consistent, directly spliced, that is: wherein is the concatenated feature map; If not, adjust by 1x1 convolution Number of channels to match That is, ensure the compatibility of the two feature maps in the dimension, and then splice in the channel, that is: stitched feature maps processed by a spatial grouping enhancement module: wherein, represents a spatial grouping enhancement module; Then, the enhanced feature map is split into two parts corresponding to the channel number of the initial input feature map: In the formula, denotes a partition function in the channel dimension; Then, the split feature map is multiplied by the initial feature map to obtain a further enhanced feature map, and the multiplication operation in this process introduces a nonlinear factor; finally, the first original feature map and the second enhanced feature map are added, and the second original feature map and the first enhanced feature map are added, and then they are spliced to obtain the integrated enhanced feature: wherein denotes element-wise multiplication, and Concat denotes concatenation in the channel dimension; S4, inputting the fused features into the head network of the improved YOLOv8s model for target detection, wherein a lightweight detection head based on PConv is used for target detection.
2. The high-resolution remote sensing image target detection method based on improved YOLOv8s according to claim 1, characterized in that: In step S1, the preprocessing steps include: manual screening, target labeling, format conversion, data set enhancement and data set division; wherein, first, download high-resolution satellite optical remote sensing map data and manually screen out target-containing regions in various backgrounds, crop and save the target regions, then label, save the labeling information as a YOLO format file, and perform Mosaic data enhancement on the image; build a network for training to obtain a weight file; finally, divide the image into a training set, a test set and a validation set according to a predetermined proportion.
3. The high-resolution remote sensing image target detection method based on improved YOLOv8s according to claim 1, characterized in that: In step S2, the C2f-PD module is used for feature extraction in the last three convolutions of the backbone network, wherein the C2f-PD module includes an input Conv block, a Split block, a plurality of PD-Bottleneck blocks, a Concat block and an output Conv block, each PD-Bottleneck block includes a partial convolution PConv and a dynamic snake convolution DSConv, the PConv convolution preliminarily extracts spatial features, and then the DSConv convolution enhances the feature extraction capability; a plurality of PD-Bottleneck blocks are connected in series to extract multi-scale feature information.
4. The high-resolution remote sensing image target detection method based on improved YOLOv8s according to claim 3, characterized in that: In each PD-Bottleneck block, the PConv convolution utilizes the redundancy of feature mapping between different feature maps to perform convolution operation on only part of the channels of the input feature map to extract local spatial features; and the first or last continuous channel is used as a representative of the entire feature map to reduce the computational complexity, and the calculation process is as follows: First, the input features are masked: wherein is an input feature, is a binary mask, denotes element-wise multiplication; Then, the masked features are subjected to partial convolution operation: wherein is the current position, is a convolution kernel weight, is a convolution kernel size, is a mask value at position , denotes an index of an element within the convolution kernel, denotes the mask value of the th position within the neighborhood of position , denotes a transposition operation; Mask update rule: The mask is updated after the partial convolution operation, and the updated mask is used for the next partial convolution operation, gradually expanding the effective area.
5. The high-resolution remote sensing image target detection method based on improved YOLOv8s according to claim 4, characterized in that: In each PD-Bottleneck block, the DSConv convolution dynamically adjusts the size and direction of the convolution kernel according to the shape and geometric characteristics in the image to enhance the feature extraction capability. The DSConv convolution continuously adjusts the shape and position of the convolution kernel by introducing a learnable deformation offset, and uses an iterative learning method to make the convolution kernel adaptively adjust the shape and size, so as to capture targets of different scales and shapes. The specific process is as follows: First, a coordinate is Standard 2D Convolution Record the center coordinates ,in ,one The formula for calculating the convolution kernel coordinates is: In the formula: is the coordinate of the standard two-dimensional convolution; respectively represent the horizontal and vertical offsets of the convolution kernel relative to the center position. The dynamic snake convolution adjusts the convolution kernel in the x-axis and y-axis directions. In the x-axis direction, the change is: wherein, represents the horizontal distance from the center grid, represents the offset in the vertical direction, subscript represents the index of the convolution kernel on the input feature map; In the y-axis direction, the change is: wherein the subscript represents an index of the convolution kernel on the input feature map, represents an offset in the horizontal direction; The floating-point coordinates in the offset results of the x-axis and y-axis directions are mapped to the image coordinates using the bilinear interpolation method, which is represented as: wherein, denotes the decimal position in the offset result in the x- and y-direction, enumerating all integer spatial positions; is a bilinear interpolation kernel, which can be decomposed into two one-dimensional kernels, i.e.: wherein, denotes a one-dimensional interpolation kernel function, denotes a fractional part and an integer part of P in the horizontal direction, respectively, denotes a fractional part and an integer part of P in the vertical direction, respectively; With the changes in the x-axis and y-axis, the receptive field of the dynamic snake convolution kernel is adjusted to adaptively capture the geometric features of the target.
6. The high-resolution remote sensing image target detection method based on improved YOLOv8s according to claim 1, characterized in that: In the SGE module of the feature fusion module, the SGE module groups the features and parallel processes each feature group along the channel direction to capture different semantic features. The steps include: First, the feature map of dimension is divided into groups along its channels, each group represented as a vector in the spatial dimensions: In the formula, denotes the total number of spatial positions; Then, the spatial average function is used to utilize the global statistical features as supplementary information to enhance the semantic modeling of the key areas: In the formula, for the spatial average function; Subsequently, the importance weight coefficient of each feature is calculated, and the global semantic features are quantified by dot product operation. and local features The degree of similarity between them, that is: coefficient Spatial normalization is performed and two learnable parameters are introduced for each coefficient. The normalized value is linearly changed by adjusting the two parameters to achieve the identity transformation, i.e.: In the formula, is a scaling factor, is an offset; Finally, the output of the sigmoid activation is multiplied with the original feature map to obtain the enhanced feature vector The final result feature group is formed: where the enhanced eigenvectors are represented as .
7. The high-resolution remote sensing image target detection method based on improved YOLOv8s according to claim 1, characterized in that: In step S4, the designed lightweight detection head module includes a Pconv convolution, a Conv standard convolution and a Conv2d convolution. First, the PConv convolution is used to perform convolution operation on only part of the channels of the input feature map. The PConv convolution removes redundant information in the feature map while preserving the feature information stream. Then, the second Conv standard convolution is used to extract complete feature information, and the last Conv2d convolution is used for channel dimension reduction processing to complete target detection.
Citation Information
Patent Citations
Marine target detection method based on high-efficiency channel priority attention
CN118941984A
Traffic remote sensing target detection method
CN119169268A