Remote sensing target detection method based on partial convolution and multi-scale feature fusion
By improving the backbone network and feature fusion process of the YOLOv8 model, partial convolution and multi-scale feature fusion methods are used to solve the problems of high computational complexity and correlation in optical remote sensing image processing in the prior art, and more efficient and accurate remote sensing object detection is achieved.
Patent Information
- Application Number
- CN202411863162.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing YOLOv8 model has high computing complexity and memory requirements when processing optical remote sensing images, making it difficult to apply on embedded systems with limited hardware resources, and feature fusion cannot take advantage of the correlation between all feature maps.
The remote sensing object detection method based on partial convolution and multi-scale feature fusion is adopted to improve the backbone network, feature fusion, detection head and loss function of the YOLOv8 model. By introducing the FEC module and an efficient multi-scale attention mechanism, the feature extraction and fusion process is optimized.
The convergence speed and generalization ability of the model are improved, the calculation requirements of feature extraction are reduced, the detection accuracy and feature expression ability of optical remote sensing images are enhanced, and the application problem on embedded systems with limited hardware resources is solved.
Smart Images

Figure CN120014473A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing target detection method based on partial convolution and multi-scale feature fusion. Background Art
[0002] With the continuous popularization and application of artificial intelligence technology, target detection technology plays a role in various fields. As an efficient and accurate target detection model, YOLOv8 is widely used in various fields. The YOLOv8 model includes a backbone network, a feature pyramid network, a detection head, and a loss function; the backbone network is used to extract image features and reduce computational complexity; the feature pyramid network fuses features at different levels to improve the model's ability to detect targets of different scales; the detection head performs multi-scale predictions to improve detection flexibility.
[0003] However, the existing YOLOv8 model has the following defects in the field of optical remote sensing images (ORSI), where objects of a certain category usually show large scale changes and the background is complex and diverse: feature extraction through convolutional neural networks has high requirements for memory and computing, which is impractical for embedded systems with limited hardware resources in actual application scenarios; feature fusion through feature pyramid structures relies on simple operations such as summation or concatenation to merge pyramid features, and cannot utilize the correlation between all feature maps. For example, when detecting ships on the water, the extracted redundant information is usually related to the water background, which may contribute little to the detection task.
[0004] Therefore, a remote sensing target detection method based on partial convolution and multi-scale feature fusion is needed. Summary of the invention
[0005] In view of this, the present invention provides a remote sensing target detection method based on partial convolution and multi-scale feature fusion, adopts the YOLOv8 model framework, improves the backbone network, feature fusion, detection head and loss function of the model, thereby improving the convergence speed and generalization ability of the model.
[0006] To this end, the present invention provides the following technical solutions:
[0007] A remote sensing target detection method based on partial convolution and multi-scale feature fusion, comprising:
[0008] Receiving remote sensing images;
[0009] Extract remote sensing image features through the backbone network to generate initial multi-level features;
[0010] The initial multi-level features are fused through the neck network to obtain a multi-scale feature map;
[0011] Detect objects of different sizes based on multi-scale feature maps;
[0012] The backbone network includes a c2f module and a FEC module.
[0013] Furthermore, the neck network comprises:
[0014] A first scale sequence feature fusion module, a second scale sequence feature fusion module, and a triple feature encoder module;
[0015] The triple feature encoder module and the scale sequence feature fusion module transmit feature data through the PANet framework to obtain a multi-scale feature map.
[0016] Furthermore, the first scale sequence feature fusion module extracts multi-scale feature data based on the initial multi-level features.
[0017] Furthermore, the multi-scale feature map includes:
[0018] The triple feature encoder module generates comprehensive feature data as a P4 feature map and a P5 feature map based on the initial multi-level features;
[0019] Fusing the multi-scale feature data with the P4 feature map and the P5 feature map to obtain a P3 feature map;
[0020] The second-scale sequence feature fusion module fuses the P3 feature map and the comprehensive feature data to obtain the P2 feature map.
[0021] Furthermore, the shallow features of the remote sensing image are extracted by the c2f module; and the deep features of the remote sensing image are extracted by the FEC module.
[0022] Further, the FEC module includes a Faster-EMA unit;
[0023] The Faster-EMA includes: partial convolution and efficient multi-scale attention mechanism.
[0024] Furthermore, the object detection loss is calculated using the auxiliary bounding box:
[0025] L Inner-SIoU =L SIoU +IoU-IoU inner
[0026] Among them, IoU is used to measure the overlap between the predicted box and the real box, which is defined as the ratio of the area of the intersection of the two boxes to the area of the union;
[0027] IoUinner is the IoU calculated by the auxiliary bounding box. A scaling factor is introduced to control the size of the auxiliary bounding box, which is used to generate auxiliary bounding boxes of different scales to calculate the loss, thereby accelerating the bounding box regression process.
[0028] Lsiou considers the influence of the angle between the anchor box and the true box, and introduces the angle loss into the bounding box regression loss function;
[0029] Linner-siou applies IoUinner to the loss function obtained by Lsiou.
[0030] Advantages and positive effects of the present invention:
[0031] The present invention adopts the YOLOv8 model framework with the advantages of high precision and high speed; and improves the feature extraction and feature fusion of the YOLOv8 model to improve the detection accuracy and generalization ability of optical remote sensing images.
[0032] In the backbone network, the FEC module is used to replace the c2f module for deep feature extraction, and partial convolution is used to perform normal convolution on only a subset of the input stream to obtain spatially extracted features; a deep neural network with a balance of speed, accuracy, and deployability is constructed; the problem that the existing technology has high requirements for memory and computing and cannot be applied to embedded systems with limited hardware resources is solved; and an efficient multi-scale attention mechanism (EMA) is introduced. On the one hand, the two encoded features are connected along the height direction of the image and share the same 1×1 convolution through a process similar to CA. On the other hand, the 3×3 branch expands the feature field by using 3×3 convolution to capture local cross-channel connections; redundant image information is eliminated to improve the contribution of feature extraction to the detection task.
[0033] The small target detection head, scale sequence feature fusion module and triple feature encoder module are introduced into the neck network; and the comprehensive feature data generated by the triple feature encoder module and the multi-scale feature data extracted by the scale sequence feature fusion module are deeply fused by using the structure and information transmission mechanism of the PANet framework. The multi-scale feature data extracted by the scale sequence feature fusion module and the comprehensive feature data of the triple feature encoder module complement each other to form a more powerful feature representation. The fused feature data is input into the small target detection head for detection, which integrates the advantages of different modules in scale processing and feature extraction, and further enhances the ability to express target features, so as to fully understand and identify targets from multiple angles and multiple scales; and solves the problem that feature fusion in the prior art cannot utilize the correlation between all feature maps.
[0034] Through the improvement of the backbone network and the neck network, the extraction of deep features is optimized and the computational requirements of feature extraction are reduced. At the same time, a more comprehensive feature fusion is performed to improve the accuracy of small target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0036] Figure 1 This is a flow chart of improving the YOLOv8 model in an embodiment of the present invention;
[0037] Figure 2 This is a diagram of the improved YOLOv8 model architecture in an embodiment of the present invention;
[0038] Figure 3 This is a diagram of the FEC module architecture in an embodiment of the present invention;
[0039] Figure 4 This is a structural diagram of a Faster-EMA unit in an embodiment of the present invention;
[0040] Figure 5 This is a schematic diagram of the detection effect of the improved YOLOv8 model in an embodiment of the present invention;
[0041] Figure 6 Flow chart of the method in the embodiment of the present invention. DETAILED DESCRIPTION
[0042] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0044] The invention provides a remote sensing target detection method based on partial convolution and multi-scale feature fusion, which is conceived as follows: adopting the YOLOv8 model framework; using the FEC (faster and more efficient convolution) module to replace the two C2f modules under the backbone network to perform deep feature extraction; the FEC module uses Faster-EMA to replace the BottleNeck in C2f to reduce redundant image information; Faster-EMA is composed of Faster block and EMA; adding a small target detection head and a scale sequence feature fusion module (SSFF) and a triple feature encoding module (TFE) in the neck network feature fusion process; and then using the correlation between multi-scale feature maps to perform multi-scale fusion. And using the auxiliary bounding box to calculate the loss to speed up the bounding box regression process, thereby improving the convergence speed and generalization ability.
[0045] like Figure 1 As shown, this embodiment uses the improved YOLOv8 model to perform remote sensing target detection, and the specific process includes the following steps:
[0046] S1: Obtain a remote sensing image dataset with noise, and divide it into training set, validation set and test set in the ratio of 7:2:1.
[0047] S2: Extract the initial multi-scale features of the input image through the backbone network; the backbone network includes: c2f module, FEC module and SPPF module;
[0048] Optical remote sensing images contain rich and diverse information, among which deep features contain more abstract and semantically rich information about the target object, which is of irreplaceable importance for accurately identifying the category of the target and accurately locating the target position. However, due to the complexity of the image background and the variability of the target scale, the difficulty of extracting deep features has increased significantly. The Bottleneck unit in the existing C2f module cannot fully mine more abstract and semantically rich information when processing deep features, resulting in some important deep features being lost or masked by background information during the extraction process. This affects the accurate recognition of the target and reduces the detection capability of multi-scale targets in complex scenes.
[0049] Therefore, the existing C2f module is easy to extract shallow feature information, but the deep feature extraction information is insufficient; the FEC module is introduced to optimize the extraction process of deep features in the backbone network; that is, the two C2f modules used to extract deep features in the backbone network are replaced by the FEC module; the shallow features are extracted by the C2f module, and the deep features are replaced by the FEC module, so as to achieve collaborative optimization in the extraction of features at different levels. The shallow features extracted by the C2f module can quickly provide basic image information for improving the YOLOv8 model and quickly locate the approximate area of the target, while the deep features, under the optimization of the FEC module, provide more accurate and semantically rich target feature representation, thereby achieving accurate recognition and positioning of the target.
[0050] S21, FEC module is obtained by replacing the BottleNeck unit of C2f module with Faster-EMA unit; the structure diagram of FEC module is as follows Figure 3 shown.
[0051] Combination Figure 4 The structure of the Faster-EMA unit in this embodiment is further described. The Faster-EMA unit in this embodiment includes: a partial convolution PConv3×3 layer, two Conv1×1 layers and an EMA layer; a normalization layer and an activation layer are included between the two Conv1×1 layers.
[0052] Specifically, in the feature extraction process, the input features are first processed preliminarily through partial convolution operations in Faster-EMA, which effectively reduces the amount of calculation and retains key information. Subsequently, the features are integrated and mapped through 1×1 convolution. On this basis, the efficient multi-scale attention mechanism encodes and processes the feature map in the height direction, while paying attention to the target features at different scales. In this process, the convolution operation of the 3×3 branch further expands the feature field by capturing local cross-channel connections, thereby enhancing the improved YOLOv8's perception of target features.
[0053] The efficient multi-scale attention mechanism also includes cross-space learning technology, which simulates the long-range interdependence between features by linearly transforming and softmaxing the 2D global average pooling output. This fusion of global information combined with the fine processing of local features enables the model to accurately locate the target object in a wider context and effectively suppress the influence of background noise. Ultimately, the attention map generated by the efficient multi-scale attention mechanism can accurately guide the model to focus on the key areas of the target object, thereby improving the quality and effectiveness of feature extraction.
[0054] S22, splicing the initial multi-level features through the SPPF module of the last layer of the backbone network for subsequent extraction of more spatial information.
[0055] S3: The neck network fuses the initial multi-level features extracted by the backbone network to generate multi-scale feature maps for object detection of different sizes.
[0056] The neck network includes a scale sequence feature fusion module and a triple feature encoder module; the path aggregation network (PANet) is used to integrate and transmit feature data to complete the initial multi-level feature fusion and generate a multi-scale feature map.
[0057] S31, the scale sequence feature fusion module combines the high-dimensional data of the deep feature map with the specific details of the shallow feature map. In the SSFF module, feature maps P3, P4 and P5 are extracted for processing. These feature maps represent information of different scales, spaces and forms. In order to be able to effectively fuse these features, feature maps P3, P4 and P5 are normalized to the same dimension. This operation ensures that in subsequent processing, feature maps of different scales can be operated on the same basis.
[0058] The triple feature encoder module of S32 and TFE modules divides the large, medium and small feature maps, adds a large-size feature map, uses feature amplification to enhance specific feature data, and generates comprehensive feature data; utilizes the structure and information transmission mechanism of the PANet framework to efficiently propagate the comprehensive information of features of different scales in the TFE module to each feature branch; enables the comprehensive feature data extracted by the TFE module to be quickly and accurately transmitted to the entire network, and interact and fuse with other modules in the network.
[0059] S33. Integrate and transfer feature data through the PANet framework. Under the PANet framework, the comprehensive feature data generated by the TFE module is deeply integrated with the multi-scale feature data extracted by the SSFF module.
[0060] The multi-scale feature data extracted by the SSFF module and the comprehensive feature data of the TFE module complement each other to form a more powerful feature representation. This fusion not only integrates the advantages of different modules in scale processing and feature extraction, but also further enhances the model's ability to express target features, enabling the model to fully understand and identify targets from multiple angles and scales.
[0061] S34, the result of fusing the multi-scale feature data extracted by the SSFF module in S33 with the comprehensive feature data of the TFE module is input into the P3 fork.
[0062] S35, the fused data in S34 and the multi-scale feature data extracted by the SSFF module are combined again into the P2 branch to form a detection head specifically for small target detection. The small target detection head integrates the feature data of the SSFF module and the TFE module, and has the ability to fuse multi-scale features and enhance small target features. It can accurately focus on small targets in complex optical remote sensing images and effectively identify densely overlapping small targets, providing a solid foundation for achieving high-precision optical remote sensing target detection.
[0063] S4: In this embodiment, auxiliary bounding boxes are used to calculate the loss and speed up the bounding box regression process, and Inner-IoU is used for calculation. Compared with the IoU loss, when the ratio is less than 1 and the size of the auxiliary bounding box is smaller than the actual bounding box, the effective regression range is smaller than the range of the IoU loss. However, the absolute value of the gradient is larger, which accelerates the convergence of high IoU samples. On the contrary, when the ratio exceeds 1, the larger scale auxiliary bounding box expands the effective regression range and enhances the regression of low IoU samples. The Inner-IoU loss is applied to the current IoU-based bounding box regression equation, referred to as Inner-SIoU, which is defined as follows:
[0064] L Inner-SIoU =L SIoU +IoU-IoU inner
[0065] Among them, IoU (Intersection over Union) is used to measure the overlap between the predicted box and the ground truth box (GT box), which is defined as the ratio of the area of the intersection of the two boxes to the area of the union;
[0066] IoUinner (Inner Intersection over Union): It is the intersection over union ratio calculated by auxiliary bounding boxes. A scaling factor is introduced to control the size of the auxiliary bounding boxes, which is used to generate auxiliary bounding boxes of different scales to calculate the loss, thereby accelerating the bounding box regression process. Its calculation method is based on the coordinates and size of the auxiliary bounding boxes;
[0067] Lsiou (SIoU Loss): It is a bounding box regression loss function. Based on previous studies, it considers the influence of the angle between the anchor box and the true box, and introduces the angle loss into the bounding box regression loss function.
[0068] Linner-siou (Inner SIoU Loss): is the loss function obtained by applying IoUinner to Lsiou.
[0069] S5: Use the calculation results of the loss function to perform backpropagation and gradient update on the improved YOLOv8 model, so that the improved YOLOv8 model can better learn the characteristics and rules of target detection and ultimately obtain more accurate detection results.
[0070] like Figure 2 As shown in FIG. 1 , it is an architecture diagram based on the improved YOLOv8 model in this embodiment.
[0071] In this embodiment, the backbone network uses the FEC module to replace the C2f module in the YOLOv8 model backbone network to extract deep features; the FEC module includes: Faster-EMA unit. The neck network introduces a small target detection head and a scale-order feature fusion module combination and a triple encoder; integrating the advantages of different modules in scale processing and feature extraction, further enhancing the improved YOLOv8 model's ability to express target features, so that the improved YOLOv8 model can fully understand and identify targets from multiple angles and multiple scales.
[0072] In this embodiment, the detection effect of the improved YOLOv8 model on the DOTA dataset is as follows Figure 5 shown.
[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A remote sensing target detection method based on partial convolution and multi-scale feature fusion, characterized in that: include: Receiving remote sensing images; Extract remote sensing image features through the backbone network to generate initial multi-level features; The initial multi-level features are fused through the neck network to obtain a multi-scale feature map; Detect objects of different sizes based on multi-scale feature maps; The backbone network includes a c2f module and a FEC module.
2. According to claim 1, a remote sensing target detection method based on partial convolution and multi-scale feature fusion is characterized in that: The neck network comprises: A first scale sequence feature fusion module, a second scale sequence feature fusion module, and a triple feature encoder module; The triple feature encoder module and the scale sequence feature fusion module transmit feature data through the PANet framework to obtain a multi-scale feature map.
3. According to claim 1, a remote sensing target detection method based on partial convolution and multi-scale feature fusion is characterized in that: The first scale sequence feature fusion module extracts multi-scale feature data based on the initial multi-level features.
4. A remote sensing target detection method based on partial convolution and multi-scale feature fusion according to claim 2, characterized in that: The multi-scale feature map includes: The triple feature encoder module generates comprehensive feature data as a P4 feature map and a P5 feature map based on the initial multi-level features; Fusing the multi-scale feature data with the P4 feature map and the P5 feature map to obtain a P3 feature map; The second-scale sequence feature fusion module fuses the P3 feature map and the comprehensive feature data to obtain the P2 feature map.
5. A remote sensing target detection method based on partial convolution and multi-scale feature fusion according to claim 1, characterized in that: The c2f module is used to extract shallow-level features of remote sensing images; and the FEC module is used to extract deep-level features of remote sensing images.
6. A remote sensing target detection method based on partial convolution and multi-scale feature fusion according to claim 5, characterized in that: The FEC module includes a Faster-EMA unit; The Faster-EMA includes: partial convolution and efficient multi-scale attention mechanism.
7. The remote sensing target detection method based on partial convolution and multi-scale feature fusion according to claim 1, characterized in that: The object detection loss is calculated using the auxiliary bounding box: L Inner-SIoU =L SIoU +IoU-IoU inner Among them, IoU is used to measure the overlap between the predicted box and the real box, which is defined as the ratio of the area of the intersection of the two boxes to the area of the union; IoUinner is the IoU calculated by the auxiliary bounding box. A scaling factor is introduced to control the size of the auxiliary bounding box, which is used to generate auxiliary bounding boxes of different scales to calculate the loss, thereby accelerating the bounding box regression process. Lsiou considers the influence of the angle between the anchor box and the true box, and introduces the angle loss into the bounding box regression loss function; Linner-siou applies IoUinner to the loss function obtained by Lsiou.