A Detection Method for Small and Occluded Targets Suitable for Remote Sensing Images

By optimizing the backbone and neck network of the YOLOv10 model, using the C2f-CPCC module and the DAPD module, combined with the Repulsion-IoU loss function, the problem of difficult detection and occlusion of small objects in remote sensing images is solved, and efficient and accurate object detection effect is achieved.

CN119850890BActive Publication Date: 2025-06-20ANHUI UNIV OF SCI & TECH GUOZHEN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510318445.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-20
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

Small objects in remote sensing images are difficult to detect, there are occlusion problems, and high computing requirements are difficult to meet the needs of real-time processing.

Method used

By optimizing the backbone and neck network of the YOLOv10 model, using the C2f-CPCC module and the DAPD module, combined with the Repulsion-IoU loss function, the model's detection performance for small targets and occlusion targets is improved.

Benefits of technology

The model's performance on remote sensing image small targets and occlusion target detection is significantly improved, achieving efficient and accurate object detection effects, especially in handling complex and occlusion scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850890B_ABST
    Figure CN119850890B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting small targets and occluded targets applicable to remote sensing images. First, a high-resolution remote sensing image dataset is collected and converted into the YOLO format. Then, a remote sensing image detection model is constructed for target detection. The remote sensing image detection model uses the YOLOv10 model as the benchmark network, and the C2f-CPCC module is used to replace the C2f module in the backbone network and the neck network, combining the detailed position information of the low-order feature map and the rich semantic information of the high-order feature map to improve the detection ability. The DAPD module is used to replace the convolutional module, dynamically process according to the regional and context information of the input features, extract the key features in the remote sensing image, and retain the overall regional information, thereby enhancing the model's ability to capture global features. The Repulsion-IoU loss function is introduced to solve the problem of mutual occlusion between small targets during the detection process and improve the model's recognition ability for occluded targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image data processing, and specifically to a method for detecting small targets and occluded targets suitable for remote sensing images. Background Art

[0002] With the rapid development of remote sensing technology, the application of remote sensing images has become increasingly widespread in multiple fields such as geographic information, environmental monitoring, urban planning, and agricultural monitoring. Remote sensing images are usually obtained through satellites, drones, or other remote sensing devices. These images can provide a large amount of valuable information about the earth's surface and are of extremely high value. In the fields of large-scale analysis of earth surface changes, disaster monitoring, land use planning, urban expansion, and agricultural development, remote sensing images are undoubtedly an important basis for analysis and decision-making. However, the complexity of remote sensing images also poses significant challenges to target detection, especially problems such as large differences in target sizes, complex backgrounds, and poor image quality.

[0003] The application of target detection technology in remote sensing images is mainly used for extracting, locating, and classifying ground object features such as buildings, roads, farmlands, water bodies, forests, etc. The accurate extraction of these ground object features not only helps environmental monitoring, urban planning, and land use management, but also plays an important role in precision agriculture, climate change monitoring, and ecological protection. In recent years, the rapid progress of deep learning technology has brought significant breakthroughs to remote sensing image target detection. Traditional target detection methods rely on manual feature extraction, with low efficiency and difficulty in adapting to complex changes in remote sensing images. Deep learning, especially convolutional neural networks (CNNs) and region extraction algorithms, can automatically learn and extract effective features in images, greatly improving the accuracy, speed, and robustness of target detection.

[0004] Remote sensing image target detection algorithms can generally be divided into two categories: two-stage detection algorithms and one-stage detection algorithms. Two-stage detection algorithms such as Faster R-CNN usually provide higher detection accuracy, but due to their high computational complexity and slow processing speed, they are not very suitable for real-time monitoring tasks. One-stage detection algorithms, such as YOLO, SSD, RetinaNet, etc., have higher processing speeds and lower computational overheads, so they are more suitable for real-time detection and the processing of large-scale remote sensing images. With the gradual popularization of computing resources and the continuous optimization of algorithms, one-stage algorithms have become a common choice in remote sensing image target detection.

[0005] However, object detection in remote sensing images still faces some challenges, mainly reflected in the following aspects: (1) Difficulty in detecting small objects: Remote sensing images are usually taken by satellites or drones, with relatively low image resolution. Due to the scale differences of surface objects, especially the detection of some small ground object targets such as small buildings, farmlands, and water bodies, it often becomes difficult due to the low image clarity. The low resolution causes small objects to be blurred or distorted in the image, and traditional object detection methods are difficult to accurately identify and locate small objects from these low-quality images. (2) Occlusion problem: In remote sensing images, target objects are often occluded. Especially in urban environments, targets such as buildings, roads, and traffic facilities are often partially occluded or overlapped by other objects. In addition, targets in natural environments such as forests, mountains, or agricultural areas may also be affected by trees, land, other ground objects, or weather conditions, resulting in partial loss or invisibility of the targets, which poses a great challenge to object detection algorithms. (3) High computational requirements: Remote sensing images usually cover a large area and contain a large amount of ground object information. To ensure detection accuracy, object detection models must process and analyze a large amount of image data, which places high requirements on computational resources. Especially in real-time detection applications, the image processing speed and the computational efficiency of the model become an important bottleneck. Currently, many high-precision object detection models have a large amount of calculations and consume a large amount of computational resources and storage space when deployed and executed, making it difficult to be efficiently deployed on resource-constrained devices. Especially in environments with limited computational resources such as drones or edge devices, it is difficult to meet the requirements of real-time processing. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method for detecting small objects and occluded objects applicable to remote sensing images. By optimizing the backbone network and neck network of the YOLOv10 model, efficient feature extraction and fusion are achieved, thereby significantly improving the performance of the model in detecting small objects and occluded objects in remote sensing images, and achieving an efficient and accurate object detection effect.

[0007] The technical solution of the present invention is as follows:

[0008] A method for detecting small objects and occluded objects applicable to remote sensing images specifically includes the following steps:

[0009] (1) Collect a high-resolution remote sensing image dataset and convert the format to the YOLO format;

[0010] (2) Construct a remote sensing image detection model. The remote sensing image detection model uses the YOLOv10 model as the backbone network, including a backbone network, a neck network, and a head network. In the backbone network and the neck network, the C2f-CPCC module is used to replace the C2f module. The C2f-CPCC module replaces the Bottleneck block in the C2f module with the CPCC module. The CPCC module performs max pooling, multi-scale convolution, and fusion operations on the input feature map. Through multi-scale convolution and feature fusion, the model's perception ability for targets of different scales is improved. In the backbone network and the neck network, the DAPD module is used to replace the Conv and SCDown modules. The DAPD module performs grouped convolution, depthwise separable convolution, and average pooling operations on the input feature map, thereby retaining global information and avoiding over-reliance on local features.

[0011] (3) Iteratively train the constructed remote sensing image detection model. The loss function for training is the Repulsion-IoU loss function, and its calculation formula is as follows (3):

[0012] (3);

[0013] In formula (3), represents the Repulsion-IoU loss function, is the attraction term of the ground truth box to the target detection box, is the repulsion term of the non-target ground truth box to the target detection box, is the IoU loss function, and are hyperparameters;

[0014] (4) Use the trained remote sensing image detection model to detect the target to be detected in the remote sensing image and obtain the detection result.

[0015] The backbone network includes two Conv modules, two C2f modules, two C2f-CPCC modules, three DAPD modules, one SPPF module, and one PSA module. The processing process of the backbone network is shown in the following formula (1):

[0016] (1);

[0017] In formula (1), is the input of the backbone network, that is, the collected remote sensing image, , , are three feature maps of different scales extracted by the backbone network;

[0018] The neck network includes two Upsample modules, one C2f module, three C2f-CPCC modules, one DAPD module, one Conv module and four Concat modules. The processing process of the neck network is shown in the following formula (2):

[0019] (2);

[0020] In formula (2), is the intermediate feature map of the neck network, , and are the three feature maps output by the neck network;

[0021] The three feature maps output by the neck network , and are input into the head network for detection to output the final detection result.

[0022] The processing process of the C2f-CPCC module specifically includes the following steps:

[0023] S11. The input feature map is transformed through the first convolutional layer to generate an intermediate feature map. Then, the generated intermediate feature map is split into two parts by the Split function: one part is passed to multiple CPCC modules for further processing. Finally, the multiple feature maps obtained after being processed by multiple CPCC modules are concatenated with the other part after splitting to form a fused feature map. The fused feature map is then passed through the second convolutional layer to output a feature map;

[0024] S12. The processing process of each CPCC module is as follows: the input feature map is divided along the channel dimension by the Split function and then input into two branches respectively. The number of channels of the feature map input into each branch is half of the input feature map. The first branch first performs a max pooling operation, and then uses a 1×1 convolution to fuse the channel features to output a feature map. The second branch uses convolutional kernels of different sizes for multi-scale feature extraction to obtain multiple multi-scale feature maps. After the multiple multi-scale feature maps are concatenated and activated by the ReLU activation function, a feature map is output. Finally, the feature map output by the first branch and the feature map output by the second branch are concatenated and then further fused by a 1×1 convolution to output a feature map.

[0025] The processing process of the DAPD module specifically includes the following steps:

[0026] S21. The input feature map is divided along the channel dimension and then input into two branches respectively. The number of channels of the feature map input into each branch is half of the input feature map. After the first branch performs grouped convolution first, the output feature map is then divided into two branches. After these two branches perform depthwise separable convolution and average pooling operations respectively, both pass through the BN layer for batch normalization processing, and then two feature maps are output;

[0027] S22. After the second branch performs depthwise separable convolution, a feature map is output;

[0028] S23. The two feature maps output by the first branch and the feature map output by the second branch are concatenated, and after 1×1 convolution is performed, different features are fused to output a feature map.

[0029] The attractive term of the target ground truth box to the target detection box The calculation process is shown in the following formula (4):

[0030] (4);

[0031] In formula (4), is the predicted box, is the ground truth box, is the smooth L1 loss function.

[0032] The repulsive term of the non-target ground truth box to the target detection box The calculation process is shown in the following formula (5):

[0033] (5);

[0034] In formula (5), is the weight of the occluded background sample, which is used to measure the degree of repulsion of the non-target area to the target detection box, is the ground truth box, is the predicted box.

[0035] The IoU loss function mentioned above The calculation process is shown in the following formula (6):

[0036] (6);

[0037] In formula (6), is the ground truth box, is the predicted box.

[0038] Advantages of the present invention:

[0039] (1). In the present invention, some C2f modules in the backbone network and neck network of the existing YOLOv10 model are replaced with C2f-CPCC modules. In the backbone network, the C2f-CPCC module is mainly used to enhance the feature extraction ability. Through multi-scale convolution and feature fusion, the model's perception ability for targets of different scales is improved. In the neck network, the role of the C2f-CPCC module is to optimize feature fusion. By fusing and enhancing the multi-scale features output by the backbone network, the accuracy of object detection is further improved, especially in the performance when dealing with complex and occluded scenarios. That is, while extracting multi-scale information, the C2f-CPCC module effectively reduces redundant calculations, thereby improving the detection performance.

[0040] (2). In the present invention, some Conv and SCDown modules in the backbone network and neck network of the existing YOLOv10 model are replaced with DAPD modules. The deep average pooling downsampling DAPD module has obvious advantages compared with the traditional convolution Conv and SCDown modules. First, the DAPD module can better retain global information. By calculating the average value within the region, the loss of local features is reduced, and it performs more stably especially when dealing with targets with large scale differences. Second, the DAPD module has a strong inhibitory effect on noise, can smooth outliers, and avoid over-reliance on local features. Third, the DAPD module is relatively simple to calculate. Especially when dealing with large-size inputs, it can significantly reduce the computational amount. In contrast, the SCDown module may have a high computational cost when resources are limited. Finally, DAPD has better adaptability and generalization ability, can effectively process targets of different scales, and is not easy to lose important information, while the traditional convolution Conv and SCDown modules may perform poorly in extracting features for small-scale targets or in complex backgrounds.

[0041] (3). The loss function trained in the present invention is the Repulsion-IoU loss function. Based on the traditional IoU loss function, a repulsion mechanism between targets is introduced, effectively solving the detection problems of small target occlusion and dense target regions in remote sensing images. While calculating the overlap of target boxes, this loss function adds a repulsion term to handle the mutual interference between occluded targets, avoiding false detections or missed detections caused by occlusion, and enhancing the model's discrimination ability between dense targets and small targets. Especially when there is occlusion or overlap between targets, it can better optimize the model's detection ability. By measuring the overlap degree between the predicted box and the ground truth box (IoU loss function) to optimize the model, the model performs well in the precise positioning between large targets and target boxes. Description of the Drawings

[0042] Figure 1 is the flowchart of the present invention.

[0043] Figure 2 It is the network framework diagram of the remote sensing image detection model of the present invention.

[0044] Figure 3 It is the network framework diagram of the C2f-CPCC module of the present invention.

[0045] Figure 4 It is the network framework diagram of the CPCC module part in the C2f-CPCC module of the present invention.

[0046] Figure 5 It is the network framework diagram of the DAPD module of the present invention. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] See Figure 1 , a method for detecting small targets and occluded targets applicable to remote sensing images, specifically including the following steps:

[0049] (1), collect a high-resolution remote sensing image dataset and convert the format to the YOLO format;

[0050] (2), see Figure 2 , construct a remote sensing image detection model. The remote sensing image detection model takes the YOLOv10 model as the benchmark network and includes a backbone network, a neck network, and a head network;

[0051] The backbone network includes two Conv (convolution) modules, two C2f modules, two C2f-CPCC modules, three DAPD modules, one SPPF module, and one PSA module. The processing process of the backbone network is shown in the following formula (1):

[0052] (1);

[0053] In formula (1), is the input of the backbone network, that is, the collected remote sensing image, , , are three feature maps with different scales extracted by the backbone network;

[0054] The neck network includes two Upsample modules, one C2f module, three C2f-CPCC modules, one DAPD module, one Conv (convolution) module, and four Concat (concatenation) modules. The processing process of the neck network is shown in the following formula (2):

[0055] (2);

[0056] In formula (2), is the intermediate feature map of the neck network, , and are the three feature maps output by the neck network;

[0057] The three feature maps output by the neck network , and are input into the head network for detection to output the final detection result;

[0058] See Figure 3 and Figure 4 , the processing process of the C2f-CPCC module specifically includes the following steps:

[0059] S11. The input feature map is transformed through the first convolutional layer to generate an intermediate feature map, and then the generated intermediate feature map is split into two parts by the Split function: one part is passed to multiple CPCC modules for further processing, and finally the multiple feature maps obtained after processing by multiple CPCC modules are concatenated with the other part after splitting to form a fused feature map. The fused feature map is then passed through the second convolutional layer to output a feature map;

[0060] S12. The processing process of each CPCC module is as follows: the input feature map is divided along the channel dimension by the Split function and then input into two branches respectively. The number of channels of the feature map input into each branch is half of the input feature map. The first branch first performs a max pooling (Max-Pool) operation, and then uses a 1×1 convolution to fuse the channel features to output a feature map. The second branch uses three different-sized convolutional kernels (1×1 convolution, 3×3 convolution, 5×5 convolution) for multi-scale feature extraction to obtain three multi-scale feature maps. After the three multi-scale feature maps are concatenated and activated by the ReLU activation function, a feature map is output. Finally, the feature map output by the first branch and the feature map output by the second branch are concatenated and then further fused by a 1×1 convolution to output a feature map;

[0061] See Figure 5 , the DAPD module is a deep average pooling downsampling module, and its processing process specifically includes the following steps:

[0062] S21. After the input feature map is divided along the channel dimension by the Split function, it is respectively input into two branches. The number of channels of the feature map input into each branch is half of that of the input feature map. After the first branch first performs grouped convolution (Group-Conv), the output feature map is then divided into two branches. After these two branches respectively perform depthwise separable convolution (DWConv) and average pooling (Average-Pool) operations, they both go through the BN layer for batch normalization processing, and then two feature maps are output;

[0063] S22. After the second branch performs depthwise separable convolution (DWConv), a feature map is output;

[0064] S23. Concatenate the two feature maps output by the first branch and the feature map output by the second branch, and then perform 1×1 convolution to fuse different features and output a feature map;

[0065] (3). Iteratively train the constructed remote sensing image detection model. The loss function for training is the Repulsion-IoU loss function, and its calculation formula is as follows (3):

[0066] (3);

[0067] In formula (3), represents the Repulsion-IoU loss function, is the attraction term of the ground truth box to the object detection box, is the repulsion term of the ground truth box of non-object to the object detection box, is the IoU loss function, and are hyperparameters;

[0068] Among them, the calculation process of the attraction term of the ground truth box to the object detection box is shown in the following formula (4):

[0069] (4);

[0070] In formula (4), is the predicted box, is the ground truth box, is the smooth L1 loss function;

[0071] The calculation process of the repulsion term of the ground truth box of non-object to the object detection box is shown in the following formula (5):

[0072] (5);

[0073] In formula (5), is the weight of the occluded background sample, which is used to measure the degree of rejection of the non-target area for the target detection box and can be determined in various ways, including: ① IoU weighting, calculated based on the IoU between the target box and the occluded area, the higher the IoU, the greater the weight; ② Area ratio weighting, the larger the area ratio of the occluded area, the higher the weight; ③ Confidence weighting, the lower the target confidence, indicating more serious occlusion, and the weight is correspondingly increased; ④ Attention mechanism, using channel or spatial attention modules to dynamically adjust the weight. The above weight determination methods can be used alone or in combination to more accurately measure the impact of occlusion on target detection; is the ground truth box; is the predicted box;

[0074] IoU loss function The calculation process is shown in the following formula (6):

[0075] (6);

[0076] In formula (6), is the ground truth box, is the predicted box;

[0077] (4) Use the trained remote sensing image detection model to detect the target to be detected in the remote sensing image and obtain the detection result.

[0078] Performance analysis:

[0079] Perform performance detection experiments on the remote sensing image detection model (Ours) of the present invention and other existing detection models (Faster RCNN, DETR, Tood, DDQ, ATSS, DINO, YOLOv10). The experimental results are shown in Table 1 below. It can be seen from Table 1 that the detection model (Ours) of the present invention leads other existing detection models with an accuracy of 0.817, indicating that it performs best in correctly identifying positive samples and has the lowest Params (number of parameters) index, greatly improving the detection performance.

[0080] Table 1

[0081]

[0082] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for detecting small targets and occluded targets in remote sensing images, characterized in that: The specific steps include: (1) Collect high-resolution remote sensing image datasets and convert them into YOLO format; (2) Construct a remote sensing image detection model. The remote sensing image detection model uses the YOLOv10 model as the benchmark network, including a backbone network, a neck network, and a head network. The C2f-CPCC module is used to replace the C2f module in the backbone network and the neck network. The C2f-CPCC module replaces the Bottleneck block in the C2f module with the CPCC module. The CPCC module performs maximum pooling, multi-scale convolution, and fusion operations on the input feature map. Through multi-scale convolution and feature fusion, the model's perception ability of targets of different scales is improved. The DAPD module is used to replace the Conv and SCDown modules in the backbone network and the neck network. The DAPD module performs grouped convolution, depthwise separable convolution, and average pooling operations on the input feature map. The processing process of the C2f-CPCC module specifically includes the following steps: S11, the input feature map is transformed through the first convolution layer to generate an intermediate feature map, and then the generated intermediate feature map is split into two parts by the Split function: one part is passed to multiple CPCC modules for further processing, and finally the multiple feature maps obtained by the multiple CPCC modules are spliced ​​with the other part after the split to form a fused feature map, and the fused feature map is then passed through the second convolution layer to output the feature map; S12. The processing process of each CPCC module is as follows: the input feature map is divided along the channel dimension by the Split function and then input into two branches respectively. The number of channels of the feature map input into each branch is half of the input feature map. The first branch first performs a maximum pooling operation, and then uses a 1×1 convolution to fuse the channel features and outputs a feature map. The second branch uses convolution kernels of different sizes to extract multi-scale features to obtain multiple multi-scale feature maps. After the multiple multi-scale feature maps are spliced ​​and activated by the ReLU activation function, the feature map is output. Finally, the feature map output by the first branch and the feature map output by the second branch are spliced, and the features are further fused by a 1×1 convolution to output the feature map. (3) The constructed remote sensing image detection model is iteratively trained. The training loss function is the Repulsion-IoU loss function, and its calculation formula is as follows (3): (3); In formula (3), Represents the Repulsion-IoU loss function, is the attraction item of the target real box to the target detection box, It is the exclusion item of the non-target real box to the target detection box. is the IoU loss function, and is a hyperparameter; (4) Use the trained remote sensing image detection model to detect the target to be detected in the remote sensing image and obtain the detection result.

2. The method for detecting small targets and occluded targets in remote sensing images according to claim 1, characterized in that: The backbone network includes two Conv modules, two C2f modules, two C2f-CPCC modules, three DAPD modules, one SPPF module and one PSA module. The processing process of the backbone network is shown in the following formula (1): (1); In formula (1), As the input of the backbone network, that is, the collected remote sensing image, , , Three feature maps of different scales extracted by the backbone network; The neck network includes two Upsample modules, one C2f module, three C2f-CPCC modules, one DAPD module, one Conv module and four Concat modules. The processing process of the neck network is shown in the following formula (2): (2); In formula (2), is the intermediate feature map of the neck network, , and Three feature maps output by the neck network; Three feature maps output by the neck network , and Input into the head network for detection and output the final detection result.

3. The method for detecting small targets and occluded targets in remote sensing images according to claim 2, characterized in that: The processing process of the DAPD module specifically includes the following steps: S21, the input feature map is divided along the channel dimension and input into two branches respectively. The number of feature map channels input into each branch is half of the input feature map. The first branch first performs group convolution, and then the output feature map is divided into two branches. After the two branches perform depthwise separable convolution and average pooling operations respectively, they are batch normalized by the BN layer and then output two feature maps. S22, the second branch performs depth-wise separable convolution and outputs a feature map; S23: Concatenate the two feature maps output by the first branch and the feature map output by the second branch, perform 1×1 convolution on the concatenated maps, fuse different features, and output a feature map.

4. The method for detecting small targets and occluded targets in remote sensing images according to claim 1, characterized in that: The attraction item of the target real box to the target detection box The calculation process of is shown in the following formula (4): (4); In formula (4), is the prediction box, It is a real frame. is the smooth L1 loss function.

5. The small target and occluded target detection method applicable to remote sensing images according to claim 1, characterized in that: The non-target real frame excludes the target detection frame The calculation process of is shown in the following formula (5): (5); In formula (5), is the weight of the occluded background sample, which is used to measure the degree of rejection of the target detection frame by the non-target area. It is a real frame. is the prediction box.

6. The small target and occluded target detection method applicable to remote sensing images according to claim 1, characterized in that: The IoU loss function The calculation process of is shown in the following formula (6): (6); In formula (6), It is a real frame. is the prediction box.

Citation Information

Patent Citations

  • Ultraviolet image target detection method and device, equipment and storage medium

    CN118485901A

  • Road defect detection method based on improved YOLOv10 network

    CN119273637A