Remote Sensing Image Small Target Detection Method Based on Noise Adaptation and Context Awareness

Through noise adaptation and context-aware methods, the candidate bounding box is expanded and position correction is performed, which solves the offset sensitivity problem in high-score remote sensing small object detection, and improves detection accuracy and recall.

CN116883872BActive Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310784167.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-07-11
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

The existing high-score remote sensing small object detection algorithm fails to effectively solve the offset sensitivity problem of small objects, resulting in increased difficulty in detecting small objects and frequent occurrence of false detection and missed detection.

Method used

Using noise adaptive and context-aware methods, multi-scale features are acquired by enlarging candidate bounding boxes, using densely connected expanded convolutional networks, and position corrections are performed, and the two-stage detection framework is improved to alleviate offset sensitivity problems.

Benefits of technology

It significantly improves the detection accuracy and recall of small targets, reduces errors, corrects missed detection, and optimizes the position of the target box.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116883872B_ABST
    Figure CN116883872B_ABST
Patent Text Reader

Abstract

The present invention provides a method for detecting small targets in remote sensing images based on noise adaption and context awareness. It mainly includes processes such as convolutional feature extraction, context-aware multi-scale feature acquisition, noise-adaptive candidate bounding box acquisition, candidate region feature extraction with position correction, object recognition, and position regression. The present invention adds noise adaption, context awareness, and position correction processing on the basis of the existing two-stage detection framework based on deep learning, which can effectively alleviate the problem of offset sensitivity in small target detection and significantly improve the detection accuracy of small targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information processing, and particularly relates to a small target detection method for remote sensing images based on noise adaptation and context awareness. Background Art

[0002] In recent years, with the development of aerospace remote sensing platforms and deep learning technologies, high-resolution remote sensing image processing algorithms have made great progress. As one of the most challenging problems, small target detection in high-resolution remote sensing images has many application scenarios, such as marine search and rescue, national defense security, urban planning, autonomous driving, etc. In practical applications, due to the wide ground range covered by remote sensing images, the signal responses of small targets are often relatively weak, resulting in a significant increase in the difficulty of detecting small targets. There are often phenomena of false detection and missed detection of small targets. Therefore, it is extremely challenging to identify and locate small targets.

[0003] Existing high-resolution remote sensing small target detection algorithms are based on the general object detection paradigm of deep learning and carefully design modules to improve the performance of small target detection. These methods can be mainly divided into the following four categories: data augmentation methods, multi-scale feature fusion methods, super-resolution methods, and context modeling methods.

[0004] The data augmentation-based methods expand the scale of small targets in the dataset through different data augmentation strategies, and weaken the influence caused by the relatively small proportion of current small targets in the dataset. Typical data augmentation strategies such as the method of oversampling small targets by copy-pasting proposed by Kisantal et al. in the literature "M. Kisantal, Z. Wojna, J. Murawski, J. Naruniec, and K. Cho, Augmentation for Small Object Detection. arXiv preprint, arXiv:1902.07296, 2019." However, the performance gain obtained by such data augmentation methods depends on the dataset and needs to be designed and optimized according to the target characteristics.

[0005] The method based on multi-scale feature fusion includes the Feature Pyramid Network proposed by T. Lin et al. in the literature "T. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie. Feature Pyramid Networks for Object Detection. In Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2017, pp. 2117–2125." This method constructs a top-down path on hierarchical features, enabling high-resolution low-level representations to have rich semantics at the same time, and greatly improving the localization accuracy of small objects.

[0006] The method based on super-resolution mainly enriches the information of small objects by super-resolving the input image or features. There is the "J. Li, X. Liang, Y. Wei, T. Xu, J. Feng, and S. Yan, Perceptual Generative Adversarial Networks for Small Object Detection. In Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2017, pp. 1951–1959." proposed by J. Li et al., which obtains high-quality representations of small objects based on generative adversarial networks. However, such methods need to balance heavy computations and overall performance.

[0007] The method based on context modeling was proposed by Y. Gong et al. in the literature "Y. Gong, Z. Xiao, X. Tan, H. Sui, C. Xu, H. Duan, and D. Li. Context-Aware Convolutional Neural Network for Object Detection in VHR Remote Sensing Imagery. IEEE Transactions on Geoscience and Remote Sensing, vol. 58, no. 1, pp. 34–44, 2019." It assists object detection by generating a context region of interest (Context-RoI) containing surrounding potential context information for each region of interest (RoI).

[0008] None of the above methods consider the impact caused by the extremely small size of small targets, that is, the problem of small target offset sensitivity. The problem of small target offset sensitivity mainly causes two interferences to the method: 1) In the candidate region generation stage, the anchor boxes prepared for small targets are more likely to be filtered, resulting in a small number of small targets in the positive samples, which is not conducive to the detection ability of the method for small targets; 2) In the positioning stage, it is more difficult to obtain the accurate detection boxes of small targets. Compared with medium-sized targets or large-scale targets, a slight deviation of the predicted boxes of small targets will lead to a significant decrease in the intersection over union (IoU) between the predicted results and the labels, resulting in a large error in the detection results. Summary of the Invention

[0009] In order to overcome the deficiencies of the prior art, the present invention provides a method for detecting small targets in remote sensing images based on noise adaption and context awareness. It mainly includes processes such as convolutional feature extraction, context-aware multi-scale feature acquisition, noise-adaptive candidate bounding box acquisition, candidate region feature extraction with position correction, object recognition, and position regression. The present invention adds noise adaption, context awareness, and position correction processing on the basis of the existing two-stage detection framework based on deep learning, which can effectively alleviate the problem of offset sensitivity in small target detection and significantly improve the detection accuracy of small targets.

[0010] A method for detecting small targets in remote sensing images based on noise adaption and context awareness, characterized by the following steps:

[0011] Step 1, extract the convolutional feature map: Use a convolutional neural network to extract features from the input image and output convolutional feature maps of different scales;

[0012] Step 2, obtain the feature pyramid: For the convolutional feature maps of different resolutions obtained in Step 1, starting from the feature map with the lowest resolution, upsample it to the same size as the feature map with the next higher resolution, and then fuse the two to obtain one layer of pyramid features. Repeat this process until the highest resolution feature map to complete the construction of the pyramid features;

[0013] Step 3, obtain context-aware multi-scale features: Input the feature layer with the middle resolution of the pyramid features obtained in Step 2 into the context awareness module to output context-aware multi-scale features; The context awareness module is a network composed of four groups of dilated convolutions with different dilation rates that are densely connected. Among them, the dilated convolution is obtained by injecting holes into the standard convolution kernel, and the dense connection means that the outputs of the dilated convolutions of the current group and all previous groups are added together and used as the input of the next group of dilated convolutions;

[0014] Step 4, Obtain candidate bounding boxes adapted to noise: Use a regional candidate network to extract target candidate regions from the input image to obtain the bounding boxes of the target candidate regions. Let the label of the bounding box of the target candidate region be (x c , y c , w, h), where (x c , y c ) represents the coordinates of the center point of the bounding box, w represents the width of the bounding box, and h represents the height of the bounding box; use a noise adaptation method to expand the bounding box of the target candidate region by k times to obtain the expanded bounding box of the target candidate region. The label of the expanded candidate region bounding box is (x c , y c , kw, kh); where the value range of k is 1.5 to 4;

[0015] Step 5, Obtain a candidate region feature map with position correction: First, intercept the corresponding region in the context-aware multi-scale features output in Step 3 according to the candidate region bounding box obtained in Step 4 to obtain the convolutional features of the candidate region. Then, scale the candidate region features to a fixed size of 7×7 through a pooling operation to obtain the candidate region feature map. Finally, use a position correction module to perform spatial position correction;

[0016] The specific process of using the position correction module to perform spatial position correction is as follows: First, use 1×1 convolution to enhance the candidate region feature map in the x direction and the y direction respectively to obtain two attention mapping graphs M x and M y , M x represents the attention mapping graph in the x direction, and M y represents the attention mapping graph in the y direction; then, multiply the candidate region feature map and the attention mapping graph pixel by pixel to obtain two three-dimensional feature maps, and use 1×1 convolution to compress the three-dimensional feature maps in the channel dimension respectively to obtain two two-dimensional feature maps; then use global maximum pooling operations to aggregate these two two-dimensional feature maps respectively to obtain feature vectors in the x direction and the y direction, and use 1×3 convolution and 3×1 convolution to perform attention operations on the feature vectors respectively to obtain attention vectors A x and A y in the x direction and the y direction; finally, multiply the two attention vectors by the candidate region feature map respectively and then add them. The added result is the candidate region feature after position correction;

[0017] Step 6, Object recognition: Use a convolutional layer plus a fully connected layer to learn the position-corrected candidate region features obtained in Step 5 to obtain a global feature vector, and then use a fully connected layer to perform type prediction on the global feature vector to obtain the class probability vector of the object;

[0018] Step 7, Location Regression: Use a convolutional layer followed by a fully connected layer to learn the candidate region features with location correction obtained in Step 5, obtaining a global feature vector. Then, use a fully connected layer to perform bounding box prediction on the global feature vector, obtaining the regression values of the predicted region bounding box in the x and y directions and the regression values of the width w and height h, thus completing object detection.

[0019] The beneficial effects of the present invention are as follows: (1) Due to the introduction of noise adaptation processing, that is, expanding the range of the bounding box, it can prevent the loss of small target anchor boxes during the extraction of candidate region bounding boxes. At the same time, combined with location correction, it further refines the positions of small targets, and can better solve the problem of small target offset sensitivity and reduce its adverse effects; (2) Due to the adoption of a compact and efficient context-aware module composed of dilated convolutions with dense connections, on the basis of increasing the receptive field of the model, it can capture the relationships between targets from different observation perspectives, make full use of the spatial context information of the feature parameter layer, and improve the detection accuracy of small targets; (3) Since the present invention is a method based on deep learning and is improved on the basis of the existing two-stage detection framework, it can effectively alleviate the problem of offset sensitivity commonly existing in small target detection and can significantly improve the recall rate and detection accuracy of small objects. Brief Description of the Drawings

[0020] Figure 1 is a flowchart of a method for detecting small targets in remote sensing images based on noise adaption and context awareness of the present invention;

[0021] Figure 2 is an experimental result image of different methods on the HRRSD dataset;

[0022] Figure 3 is an experimental result image of different methods on the ITCVD dataset;

[0023] In the figure, (a) is the detection result image of the Faster R-CNN method in Scenario 1, (b) is the detection result image of the method of the present invention in Scenario 1, (c) is the detection result image of the Faster R-CNN method in Scenario 2, and (d) is the detection result image of the method of the present invention in Scenario 2. Detailed Embodiments

[0024] The present invention will be further described below in conjunction with the drawings and embodiments. The present invention includes but is not limited to the following embodiments.

[0025] As Figure 1 shown, the present invention provides a method for detecting small targets in remote sensing images based on noise adaption and context awareness, and its specific implementation process is as follows:

[0026] Step 1, extract convolutional feature maps: Use a convolutional neural network to extract features from the input image and output convolutional feature maps of different scales. Specifically, the convolutional neural network can adopt a ResNet50 or ResNet101 network.

[0027] Step 2, obtain a feature pyramid: For the convolutional feature maps of different resolutions obtained in Step 1, upsample the low-resolution feature maps to the same size as the high-resolution feature maps and perform information fusion to gradually complete the construction of the feature pyramid.

[0028] Step 3, obtain context-aware multi-scale features: Input the intermediate-resolution feature layer (with both rich spatial and semantic information) in the pyramid features obtained in Step 2 into the context-aware module to output context-aware multi-scale features; the context-aware module is a network composed of four groups of dilated convolutions with different dilation rates. Among them, the dilated convolution is obtained by injecting holes into the standard convolution kernel, which not only ensures that the number of parameters does not increase but also allows the receptive field to expand exponentially. In addition, dense connection means that the outputs of the dilated convolutions of the current group and all previous groups are added together and used as the input of the next group of dilated convolutions, so as to obtain the context features of the target under multiple observation perspectives on the basis of a larger receptive field.

[0029] Step 4, obtain noise-adaptive candidate bounding boxes: Use a region proposal network to extract target candidate regions from the input image to obtain the bounding boxes of the target candidate regions. Let the label of the bounding box of the target candidate region be (x c ,y c ,w,h), where (x c ,y c ) represents the coordinates of the center point of the bounding box, w represents the width of the bounding box, and h represents the height of the bounding box; use a noise-adaptive method to expand the bounding box of the target candidate region by k times to obtain the expanded bounding box of the target candidate region. The label of the expanded candidate region bounding box is (x c ,y c ,kw,kh); where the value range of k is 1.5 to 4.

[0030] Step 5, obtain the candidate region feature map with position correction: First, intercept the corresponding region in the context-aware multi-scale features output in Step 3 according to the candidate region bounding box obtained in Step 4 to obtain the convolutional features of the candidate region. Then, scale the candidate region features to a fixed size of 7×7 through a pooling operation to obtain the candidate region feature map. Finally, use the position correction module to perform spatial position correction along the x-axis and y-axis to filter out the pure noise in the background around small targets and refine the positions of small targets.

[0031] The specific process of using the position correction module for spatial position correction is as follows: First, 1×1 convolutions are used to enhance the candidate region feature map in the x and y directions respectively, obtaining attention mapping maps M x and M y , M x represents the attention mapping map in the x direction, and M y represents the attention mapping map in the y direction; then, the candidate region feature map and the attention mapping maps are multiplied pixel by pixel to obtain two three-dimensional feature maps, and 1×1 convolutions are respectively used to compress the three-dimensional feature maps in the channel dimension to obtain two two-dimensional feature maps; then, global max pooling operations are respectively used to aggregate these two feature maps to obtain feature vectors in the x and y directions, and 1×3 convolution and 3×1 convolution are respectively used to perform attention operations on the feature vectors to obtain attention vectors A x and A y , A x represents the attention vector in the x direction, and A y represents the attention vector in the y direction; finally, the obtained attention vectors are respectively multiplied by the candidate region feature map, and the results are added to obtain the candidate region feature map after position correction.

[0032] Step 6, object recognition: Use a convolutional layer plus a fully connected layer to learn the candidate region features after position correction obtained in step 5 to obtain a global feature vector, and then use a fully connected layer to perform type prediction on the global feature vector to obtain the class probability vector of the object.

[0033] Step 7, position regression: Use a convolutional layer plus a fully connected layer to learn the candidate region features after position correction obtained in step 5 to obtain a global feature vector, and then use a fully connected layer to perform bounding box prediction on the global feature vector to obtain the regression values of the predicted region bounding box in the x and y directions and the regression values of the width w and height h, completing object detection.

[0034] To verify the effectiveness of the method of the present invention, two GeForce RTX 2080Ti GPUs were used on the Ubuntu operating system, and the Mmdetection framework was used for simulation experiments in the PyTorch 1.7 environment. Two public remote sensing datasets, namely HRRSD and ITCVD, were used in the experiments. Among them, the HRRSD dataset was released by the Chinese Academy of Sciences in 2017 and contains 13 types of remote sensing ground objects for studying the object detection of high-resolution remote sensing images. This dataset includes 21,761 color images downloaded from Google Maps with a resolution of 0.15 m to 1.2 m, and 4,961 images downloaded from Baidu Maps with a resolution of 0.6 m to 1.2 m. The images in the ITCVD dataset were taken by an aircraft at an altitude of about 330 m above Enschede, the Netherlands, with nadir views and oblique views. The tilt angle of the oblique view is 45 degrees, and the ground sampling distance (GSD) of the nadir image is 10 cm. This dataset contains a total of 173 images and 29,088 vehicle targets. In the experiment, each image was cropped into 16 parts for convenient training, and all images were divided into a training set, a validation set, and a test set according to the ratio of 25%, 25%, and 50%.

[0035] The evaluation metrics of the method of the present invention include the average detection precision AP (Average Precision) and the average recall rate AR (Average Recall). Among them, the metrics related to the adopted IoU threshold are AP_50:95, AP_50, AP_75, AR_50:95. The numbers after the underscore respectively represent: the average value at intervals of 0.05 when the IoU threshold is 0.5 to 0.95, the IoU threshold is 0.5, and the IoU threshold is 0.75; the metrics related to the target area size are AP_S, AP_M, AP_L, AR_S, AR_M, AR_L, which respectively represent the performance of the detector on small targets S (area less than 32 2 ), medium targets M (area greater than 32 2 and less than 96 2 ), and large targets L (area greater than 96 2 ).

[0036] First, on the HRRSD dataset, the indicators of the method of the present invention are compared with those of methods such as Faster R-CNN, ATSS, FoveaBox, RetinaNet, Libra R-CNN, Guided Anchoring, and Double-Head R-CNN, as shown in Table 1. It can be seen that the method of the present invention has good performance in detection accuracy, and there are significant improvements in indicators such as AP_50:95, AP_50, AP_S, AP_M, and AP_L. Among them, the recall rate and accuracy improvement of small targets are particularly obvious.

[0037] Table 1

[0038]

[0039]

[0040] Table 2

[0041]

[0042] Table 2 presents the experimental results of comparing the evaluation indicators of the method of the present invention with those of methods such as Faster R-CNN, ATSS, TOOD, RetinaNet, FoveaBox, PANet, and Double-Head R-CNN on the ITCVD dataset. It can be seen that the method of the present invention has great competitive advantages. Except for the AP_L and AR_L indicators that need to be improved, it exceeds other comparison methods in all other indicators.

[0043] Figure 2 Some experimental result images on the HRRSD dataset are given. Among them, the first column is the original image, the second column is the detection result image of Faster R-CNN, and the third column is the detection result image of the present invention. It can be seen that for small targets, the present invention can correct the phenomenon of missed detection, and significantly correct the positions of some target boxes. Multiple overlapping target boxes are optimized into the only correct target box. Figure 3 Some experimental result images on the ITCVD dataset are given. Among them, (a) and (c) are the detection result images of Faster R-CNN in different scenarios, and (b) and (d) are the detection result images of the method of the present invention in different scenarios. It can be seen that the method of the present invention can detect the missed small targets. In summary, through experiments, it is verified that the method of the present invention can effectively alleviate the problem of missed detection of small targets and obtain more accurate position regression results.

Claims

1. A small target detection method for remote sensing images based on noise adaptation and context awareness, characterized in that The steps are as follows: Step 1, extract convolution feature maps: Use a convolutional neural network to extract features from the input image and output convolution feature maps of different scales; Step 2, obtain a feature pyramid: For the convolution feature maps of different resolutions obtained in Step 1, starting from the feature map with the lowest resolution, upsample it to the same size as the feature map with the next higher resolution, and then fuse the two to obtain one layer of pyramid features. Do this until the highest resolution feature map to complete the construction of the pyramid features; Step 3, obtain context-aware multi-scale features: Input the feature layer with the middle resolution of the pyramid features obtained in Step 2 into the context-aware module to output context-aware multi-scale features; The context-aware module is a network composed of four groups of dilated convolutions with different dilation rates. Among them, the dilated convolution is obtained by injecting holes into the standard convolution kernel, and the dense connection means that the outputs of the dilated convolutions in the current group and all previous groups are added together and used as the input for the next group of dilated convolutions; Step 4, obtain noise-adaptive candidate bounding boxes: Use a region proposal network to extract target candidate regions from the input image to obtain the bounding boxes of the target candidate regions. Let the label of the bounding box of the target candidate region be (x c , y c , w, h), where (x c , y c ) represents the coordinates of the center point of the bounding box, w represents the width of the bounding box, and h represents the height of the bounding box; use a noise-adaptive method to expand the bounding box of the target candidate region by k times to obtain the expanded bounding box of the target candidate region. The label of the expanded candidate region bounding box is (x c , y c , kw, kh); where the value range of k is 1.5 to 4; Step 5, obtain the candidate region feature map with position correction: First, intercept the corresponding region in the context-aware multi-scale features output in Step 3 according to the candidate region bounding box obtained in Step 4 to obtain the convolution features of the candidate region. Then, scale the candidate region features to a fixed size of 7×7 through a pooling operation to obtain the candidate region feature map. Finally, use the position correction module for spatial position correction; The specific process of using the position correction module for spatial position correction is as follows: First, 1×1 convolution is used to enhance the candidate region feature map in the x and y directions respectively, obtaining attention mapping maps M x and M y . M x represents the attention mapping map in the x direction, and M y represents the attention mapping map in the y direction. Then, the candidate region feature map is multiplied pixel by pixel with the attention mapping map to obtain two three-dimensional feature maps, and 1×1 convolution is respectively used to compress the three-dimensional feature maps in the channel dimension to obtain two two-dimensional feature maps. Then, global max pooling operations are respectively used to aggregate these two two-dimensional feature maps to obtain feature vectors in the x and y directions, and 1×3 convolution and 3×1 convolution are respectively used to perform attention operations on the feature vectors to obtain attention vectors A x and A y . Finally, the two attention vectors are respectively multiplied with the candidate region feature map and then added together, and the added result is the candidate region feature after position correction. Step 6, object recognition: Use a convolutional layer plus a fully connected layer to learn the candidate region features with position correction obtained in Step 5 to obtain a global feature vector, and then use a fully connected layer to perform type prediction on the global feature vector to obtain the class probability vector of the object; Step 7, position regression: Use a convolutional layer plus a fully connected layer to learn the candidate region features with position correction obtained in Step 5 to obtain a global feature vector, and then use a fully connected layer to perform bounding box prediction on the global feature vector to obtain the regression values of the predicted region bounding box in the x and y directions and the regression values of the width w and height h to complete object detection.

Citation Information

Patent Citations

  • Small target detection method and device, electronic equipment and storage medium

    CN110782430A

  • Remote sensing image target detection method based on global context perception

    CN114519819A