A multi-level feature fusion target detection method for small targets in remote sensing images
By adopting an object detection method that integrates multi-level features in remote sensing images, using convolutional neural networks and multi-level features, the problem of weak object detection accuracy in remote sensing images is solved, and higher detection accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202110690152.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-11
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-06-11
AI Technical Summary
The prior art is difficult to accurately detect weak targets in remote sensing images, with low detection accuracy and slow rate.
A fusion multi-level feature object detection method for weak remote sensing images is proposed. Based on the multi-level features of weak remote sensing objects, the detection accuracy of weak remote sensing images is improved through the end-to-end target detection network. Specific steps include the formation of image feature pyramids, cross-level feature fusion, spatial feature aggregation and the use of non-maximum suppression algorithms.
It effectively improves the detection rate and accuracy of weak target detection in remote sensing images, and overcomes the problems of low detection accuracy and slow detection rate in weak target detection methods.
Smart Images

Figure CN113723172B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and machine learning, and in particular relates to a fusion multi-level feature target detection method for small targets in remote sensing images. Background Art
[0002] With the development of remote sensing image target detection, small target detection has gradually become a research focus of remote sensing earth observation. Small targets in remote sensing images have weak features and small size. The weak features are mainly reflected in the unclear outline of the small targets, the unprominent texture features, and the high similarity with the adjacent background features. The small size is mainly reflected in the small number of pixels occupied by the small targets in the remote sensing images, such as dozens of pixels. Therefore, only a small number of effective features of small targets can be extracted in the remote sensing images, making it difficult to detect small targets in remote sensing images.
[0003] In recent years, deep learning methods have shown great advantages in the field of target detection. Deep learning methods can learn features from massive data sets in advance and fully extract shallow surface features and deep semantic features for the target to be detected. When there are sufficient training samples, the trained network model has strong generalization ability and can cope with target detection in various complex environments.
[0004] For targets with clear features and large size, conventional target detection methods are mainly used, such as Faster R-CNN, SSD, YOLO series methods, etc., which can achieve high target detection accuracy; however, in order to accurately detect weak and small targets, conventional target detection methods can no longer meet current requirements. It is necessary to propose key technologies that are conducive to improving the detection accuracy of weak and small targets based on the characteristics of weak and small targets. Summary of the invention
[0005] In order to overcome the problem of low detection accuracy of conventional target detection methods when detecting small and weak targets in remote sensing images, the present invention proposes a fusion multi-level feature target detection method for small and weak targets in remote sensing images.
[0006] The present invention proposes a fusion multi-level feature target detection method for small targets in remote sensing images. The method is based on a convolutional neural network and multi-level features of small targets in remote sensing images. Through an end-to-end target detection network, the detection accuracy of small targets in remote sensing images is improved. The specific process of the detection method is as follows:
[0007] Step 1: Input the remote sensing image with small targets into the convolutional neural network, and the backbone feature extraction network samples the remote sensing image to form an image feature pyramid;
[0008] Step 2: performing double upsampling and downsampling operations on some levels in the image feature pyramid, and fusing upper and lower level features of adjacent feature extraction layers to extract fused features;
[0009] Step 3: Aggregate features through the spatial feature aggregation module to predict weak and small targets in the dual-branch feature map; finally, use the non-maximum suppression algorithm NMS to obtain the weak and small target detection results.
[0010] Furthermore, the step 1 is specifically as follows:
[0011] The remote sensing image with small targets is input into the convolutional neural network, and the backbone feature extraction network downsamples the remote sensing image four times in a row, with the downsampling multiple being twice, thereby extracting the shallow to deep features of the image, and the corresponding level number is p i (i=0,1,2,3,4), forming an image feature pyramid.
[0012] Furthermore, the step 2 is specifically as follows:
[0013] The p2, p3, and p4 layers in the image feature pyramid are upsampled by the nearest neighbor interpolation method, downsampled by convolution, and the upper and lower features of adjacent feature extraction layers are fused to extract more effective features with target texture information and semantic information; at the same time, the p2, p3, and p4 layers of the feature pyramid all have input and output nodes for feature fusion, and an additional feature fusion path is added between the input and output nodes of the middle layer to fuse more channel features; finally, a dual-branch feature map after cross-level multi-channel fusion is output;
[0014] The feature fusion is expressed by the following formula:
[0015]
[0016]
[0017] In the formula, ↑ 2× Indicates that the feature map is upsampled twice by the nearest neighbor interpolation method;↓ 2× Indicates that the feature map is downsampled twice through convolution; and It is the feature map output after cross-level channel feature fusion; The third term p3 in the formula is an added cross-level channel feature fusion path.
[0018] Furthermore, the step 3 is specifically as follows:
[0019] After the cross-level channel features are fused, a dual-branch feature map is output to obtain a stronger weak target feature; a spatial feature aggregation module is introduced, which is composed of two upper and lower parallel position attention mechanisms CA, corresponding to the upper and lower branches of the dual-branch feature map formed after the cross-level channel features are fused; for the position attention mechanism CA on each branch, features are aggregated along two spatial directions respectively to further retain the precise position information of the target, and finally a weak target dual-branch feature map with prominent features is obtained. Weak target prediction is performed on the weak target dual-branch feature map with prominent features, and finally the weak target detection result is obtained by using non-maximum suppression NMS.
[0020] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:
[0021] The present invention can be applied to small target detection in remote sensing images, overcoming the problems of low detection accuracy and slow rate in small target detection by conventional learning target detection methods, and effectively improving the recall and precision of small target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the overall process of the target detection method of the present invention;
[0023] Figure 2 This is an example of a remote sensing image with a small target;
[0024] Figure 3 It is a schematic diagram of the backbone feature extraction network and cross-level channel feature fusion in the target detection method of the present invention;
[0025] Figure 4 It is a schematic diagram of a position attention space feature aggregation and a dual-branch feature weak small target prediction module in the target detection method of the present invention;
[0026] Figure 5 This is an example diagram of the results of detecting weak and small targets in remote sensing images in the target detection method of the present invention. DETAILED DESCRIPTION
[0027] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings:
[0028] The present invention is a fusion multi-level feature target detection method for small and weak targets in remote sensing images. The small and weak target detection method is based on convolutional neural networks and multi-level features of small and weak targets in remote sensing images. The convolutional neural network performs target image feature extraction, cross-level channel feature fusion, position attention mechanism feature aggregation and dual-branch feature map prediction layer by layer to form an end-to-end target detection network, thereby improving the detection accuracy of small and weak targets in remote sensing images. The process is as follows Figure 1The present invention can accurately detect small and weak targets in remote sensing images, and effectively improve the recall and precision of small and weak target detection in remote sensing images.
[0029] The specific process of the fusion multi-level feature target detection method for small and weak targets in remote sensing images is as follows:
[0030] Step 1: Input the remote sensing image with small targets into the convolutional neural network, and the backbone feature extraction network downsamples the remote sensing image four times in a row, with the downsampling multiple being twice, thereby extracting the shallow to deep features of the image, and the corresponding level number is p i (i=0,1,2,3,4), forming an image feature pyramid; Step 2: up-sample the p2, p3, and p4 layers in the image feature pyramid by the nearest neighbor interpolation method, down-sample by convolution, and fuse the upper and lower features of adjacent feature extraction layers to extract more effective features with target texture information and semantic information; At the same time, the p2, p3, and p4 layers of the feature pyramid all have input and output nodes for feature fusion, and an additional feature fusion path is added between the input and output nodes of the middle level to fuse more channel features; Finally, a dual-branch feature map after cross-level multi-channel fusion is output;
[0031] The feature fusion is expressed by the following formula:
[0032]
[0033]
[0034] In the formula, ↑ 2× Indicates that the feature map is upsampled twice by the nearest neighbor interpolation method;↓ 2× Indicates that the feature map is downsampled twice through convolution; and It is the feature map output after cross-level channel feature fusion; The third term p3 in the formula is an added cross-level channel feature fusion path;
[0035] Step 3: After the cross-level channel features are fused, a dual-branch feature map is output to obtain a stronger weak target feature; a spatial feature aggregation module is introduced, which is composed of two upper and lower parallel position attention mechanisms CA, which correspond to the upper and lower branches of the dual-branch feature map formed after the cross-level channel features are fused; for the position attention mechanism CA on each branch, features are aggregated along two spatial directions respectively to further retain the precise position information of the target, and finally a weak target dual-branch feature map with prominent features is obtained. Weak target prediction is performed on the weak target dual-branch feature map with prominent features, and finally the weak target detection result is obtained by using non-maximum suppression NMS.
[0036] The specific example of small target detection in remote sensing images is as follows:
[0037] Figure 2 The present invention is an example of a remote sensing image with small and weak targets. As can be seen from the figure, the small and weak targets in the remote sensing image have the characteristics of weak features and small size. The weak features are mainly reflected in the unclear outline of the small and weak targets, the unprominent texture features, and the high similarity with the adjacent background features. The small size is mainly reflected in the small number of pixels occupied by the small and weak targets in the remote sensing image. When conventional target detection methods are used to detect small and weak targets, multiple downsampling operations will cause the loss of target feature information, resulting in low detection accuracy. The present invention proposes key technologies for detecting small and weak targets, which effectively improve the recall and precision of small and weak target detection.
[0038] Application step 1: Input the remote sensing image with weak targets into the convolutional neural network, and the backbone feature extraction network downsamples the remote sensing image four times in a row, with the downsampling multiple being twice, thereby extracting the shallow to deep features of the image, and the corresponding level number is p i (i=0,1,2,3,4), forming an image feature pyramid;
[0039] Application step 2: Upsample the p2, p3, and p4 layers in the image feature pyramid by the nearest neighbor interpolation method, downsample by convolution, and fuse the upper and lower features of adjacent feature extraction layers to extract more effective features with target texture information and semantic information; at the same time, the p2, p3, and p4 layers of the feature pyramid all have input and output nodes for feature fusion, and an additional feature fusion path is added between the input and output nodes of the intermediate layer to fuse more channel features; finally, as Figure 3 As shown, the dual-branch feature map after cross-level multi-channel fusion is output;
[0040] The feature fusion is expressed by the following formula:
[0041]
[0042] In the formula, ↑ 2× Indicates that the feature map is upsampled twice by the nearest neighbor interpolation method;↓ 2× Indicates that the feature map is downsampled twice through convolution; and It is the feature map output after cross-level channel feature fusion; The third term p3 in the formula is an added cross-level channel feature fusion path;
[0043] Application step 3: After cross-level channel feature fusion, output a dual-branch feature map to obtain stronger weak target features; introduce a spatial feature aggregation module, such as Figure 4 As shown, the spatial feature aggregation module is composed of two upper and lower parallel position attention mechanisms CA, which correspond to the upper and lower branches of the dual-branch feature map formed after the cross-level channel feature fusion; for the position attention mechanism CA on each branch, the features are aggregated along two spatial directions respectively, and the precise position information of the target is further retained, and finally a dual-branch feature map of weak targets with prominent features is obtained. Weak target prediction is performed on the dual-branch feature map of weak targets with prominent features, and finally the non-maximum suppression NMS is used to obtain the weak target detection result. The specific detection result example is as follows Figure 5 The specific implementation methods described above further describe the purpose, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for detecting small targets in remote sensing images by integrating multi-level features, characterized in that: The method is based on convolutional neural networks and multi-level features of small and weak targets in remote sensing images. It extracts target features layer by layer through convolutional neural networks, fuses target features across layers, aggregates target features through position attention mechanisms, and predicts targets through dual-branch feature maps, thus forming an end-to-end target detection network and achieving accurate detection of small and weak targets. The specific process of the detection method is: Step 1: Input the remote sensing image with small targets into the convolutional neural network, and the backbone feature extraction network samples the remote sensing image to form an image feature pyramid; the backbone feature extraction network continuously downsamples the remote sensing image four times, and the corresponding level number is p i (i=0,1,2,3,4); Step 2: performing double upsampling and downsampling operations on some levels in the image feature pyramid, and fusing upper and lower level features of adjacent feature extraction layers to extract fused features; Performing two-fold upsampling and two-fold downsampling operations on the image feature pyramid p2, p3, and p4 layers, and fusing upper and lower level features of adjacent feature extraction layers; adding a cross-level channel feature fusion path, performing cross-level channel feature fusion, and further extracting fusion features; The cross-level channel feature fusion is specifically as follows: The p2, p3, and p4 layers in the image feature pyramid are upsampled by the nearest neighbor interpolation method, downsampled by convolution, and the upper and lower features of adjacent feature extraction layers are fused to extract more effective features with target texture information and semantic information; at the same time, the p2, p3, and p4 layers of the feature pyramid all have input and output nodes for feature fusion, and an additional feature fusion path is added between the input and output nodes of the middle layer to fuse more channel features; finally, a dual-branch feature map after cross-level multi-channel fusion is output; The feature fusion is expressed by the following formula: In the formula, ↑ 2× Indicates that the feature map is upsampled twice by the nearest neighbor interpolation method;↓ 2× Indicates that the feature map is downsampled twice through convolution; and It is the feature map output after cross-level channel feature fusion; The third term p3 in the formula is an added cross-level channel feature fusion path; Step 3: Aggregate features through the spatial feature aggregation module to predict small targets in the dual-branch feature map; finally, use the Non-Maximum Suppression NMS algorithm to obtain small target detection results.
2. The method for detecting small targets in remote sensing images by integrating multi-level features according to claim 1, characterized in that: The step 1 is specifically as follows: The remote sensing image with small targets is input into the convolutional neural network, and the backbone feature extraction network downsamples the remote sensing image four times in a row, with the downsampling multiple being twice, thereby extracting the shallow to deep features of the image, and the corresponding level number is p i (i=0,1,2,3,4), forming an image feature pyramid.
3. The method for detecting small targets in remote sensing images by integrating multi-level features according to claim 2, characterized in that: In step 3, the features are aggregated by the spatial feature aggregation module as follows: After cross-level channel feature fusion, a dual-branch feature map is output to obtain stronger weak target features; a spatial feature aggregation module is introduced, which consists of two upper and lower parallel position attention mechanisms CoordinateAttention CA, which correspond to the upper and lower branches of the dual-branch feature map formed after cross-level channel feature fusion; For each branch, the position attention mechanism CA aggregates features along two spatial directions respectively to further retain the precise position information of the target, and finally obtains a dual-branch feature map of weak targets with prominent features.
Citation Information
Patent Citations
Deep learning small target detection method and device based on cascade fusion and attention mechanism
CN112801158A