Remote Sensing Image Target Detection Method Based on Foreground Attention Network

Through multi-scale fusion feature extraction and a foreground attention network combined with a cosine distance generation network, a stable prototype is generated for remote sensing image object detection, which solves the problem of low detection accuracy caused by target scale changes and complex background in remote sensing images, and achieves higher detection accuracy.

CN116543315BActive Publication Date: 2025-07-29HARBIN ENG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310596184.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2025-07-29
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

The existing small sample image object detection method that combines the idea of prototype networks in remote sensing images has large changes in the target scale and complex background, and the generated prototype is unstable, resulting in low accuracy of detection results.

Method used

A multi-scale fusion feature extraction network, a foreground attention network and a prototype generation network based on cosine distance are used to generate more stable prototypes for object detection through multi-scale feature extraction, feature enhancement and prototype generation.

Benefits of technology

The accuracy of target detection of small sample remote sensing images is improved, adapts to target scale changes, reduces background impact, and generates more stable prototypes for subsequent detection, improving detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543315B_ABST
    Figure CN116543315B_ABST
Patent Text Reader

Abstract

A remote sensing image target detection method based on a foreground attention network, specifically involving a small-sample remote sensing image target detection method based on multi-scale feature fusion, a foreground attention network, and a prototype generation network. To solve the problem that the small-sample remote sensing image target detection results of existing small-sample image target detection methods combined with the prototype network idea have low accuracy. It uses a multi-scale fusion feature extraction network to extract multi-scale features of each image, and then uses a foreground attention network to obtain enhanced features. For the same target, a prototype generation network based on cosine distance assigns different weights to different enhanced features of the current target, and performs weighted averaging to obtain the prototype of each type of target. The multi-scale features of the image to be queried are obtained, and the proposed bounding boxes of the targets in each scale feature are obtained using RPN. The target prototype with the highest similarity to the proposed bounding box target is used as the target of the image to be queried, and the target category and location are obtained. It belongs to the field of remote sensing image target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for target detection of remote sensing images, in particular to a method for small-sample remote sensing image target detection based on multi-scale feature fusion, foreground attention network and prototype generation network, belonging to the field of remote sensing image target detection. Background Technique

[0002] In the field of remote sensing image interpretation, target detection, as a hot task with extensive applications, has developed rapidly in recent years. However, in many remote sensing scenarios, there are problems such as a scarcity of remote sensing images of specific targets or insufficient annotation of remote sensing images, which makes the effect of directly using deep learning models for target detection unsatisfactory. Therefore, some small-sample image target detection methods (such as Meta R-CNN, FSWD) combine the idea of prototype network in small-sample classification, and generate prototypes as prior knowledge to assist the model in judging the position and category of the detection target. Although this method has achieved some results in small-sample target detection of natural images, due to the large scale change of targets and more complex background in remote sensing images compared with natural images, the prototypes generated by this method are not stable enough, resulting in poor effect when using the prototypes as prior knowledge for target detection subsequently, and leading to low accuracy of small-sample remote sensing image target detection results. Summary of the Invention

[0003] In order to solve the problem that when the existing small-sample image target detection method combined with the idea of prototype network is applied to remote sensing images, due to the large scale change of remote sensing image targets and more complex background, the generated prototypes are not stable, and the effect is poor when using the prototypes as prior knowledge for target detection subsequently, resulting in low accuracy of small-sample remote sensing image target detection results, the present invention further proposes a method for target detection of remote sensing images based on a foreground attention network.

[0004] The technical solution adopted by the present invention is as follows:

[0005] It includes the following steps:

[0006] S1. Obtain the images of each target in the DIOR dataset. The type and position of the target have been annotated in each image. For each type of target, define K annotated samples of the current target as the support set, and the remaining annotated samples as the query set. The images in the support set are called support images, and the images in the query set are called query images;

[0007] S2. Construct a multi-scale fusion feature extraction network. The multi-scale fusion feature extraction network sequentially includes a backbone network, a global context module, an FPN network and a feature fusion module. Input the support set into the multi-scale fusion feature extraction network to obtain the multi-scale features of each support image;

[0008] S3. Construct a foreground attention network, input the multi-scale features of each support image into the foreground attention network to obtain corresponding enhanced features;

[0009] S4. Construct a prototype generation network based on cosine distance, input the enhanced features of all support images into the prototype generation network based on cosine distance. For the same target, different weights are assigned to different enhanced features of the current target through cosine distance, and all weights are weighted and averaged to obtain the prototype of the current target, and thus the prototypes of each class of targets can be obtained;

[0010] S5. Use the prototypes of all targets to detect the targets corresponding to each query image in the query set, and obtain the target category and location corresponding to each query image;

[0011] S6. Based on the target category and target location corresponding to each query image, repeat S2 - S5 to update the prototypes of each class of targets until the upper limit of the number of iterations is satisfied, and obtain the prototypes of each class of targets;

[0012] S7. Use S2 to extract features from the image to be queried to obtain the multi-scale features of the image to be queried, input the multi-scale features into the RPN network to obtain the proposal boxes of the targets in each scale feature;

[0013] Calculate the similarity between the targets corresponding to the proposal boxes and the prototypes of each class of targets obtained in S6, and respectively correct the positions of the proposal boxes according to the prototypes of each class of targets. Take the prototype of the target with the highest similarity as the target in the image to be queried, and obtain the target category and location in the image to be queried according to the target and the corrected proposal box.

[0014] Furthermore, the specific process of S2 is as follows:

[0015] S21. Input the support set into the backbone network of the multi-scale fusion feature extraction network. Set the backbone network to include 5 output layers, and the output feature scales of different output layers are different. Select the last three output layers as the final output layers to obtain the features output by each support image through each final output layer, that is, obtain the features of different scales of each support image;

[0016] Record the output layer in the backbone network where the feature of each scale of each support image belongs as C i , i = 1, 2, 3, 4, 5;

[0017] S22. Input the output feature of each support image at the fifth output layer C5 of the backbone network into the global context module, respectively obtain feature maps representing context information of different receptive fields through dilated convolutions with different dilation rates, splice all the obtained feature maps, and then generate a feature map with 256 channels through 1x1 convolution;

[0018] S23. Treat the features of different scales as feature maps at different times. Input the output features of the output layer C3, the output features of the output layer C4, and the feature map of the output layer C5 obtained in S22 corresponding to each support image into the FPN network. Use 256 1×1 convolutional kernels to make the number of channels of all features consistent, and obtain the new feature P corresponding to each feature. i , i = 3, 4, 5. That is, for each support image, three new feature maps are obtained, namely P3, P4, and P5.

[0019] Fix the sizes of the three new feature maps P3, P4, and P5 of each support image to the size of P3 through upsampling processing, and obtain all the feature maps with modified sizes of each support image.

[0020] S24. Input all the feature maps with modified sizes of each support image into the feature fusion module. Expand the scale dimension t of each feature map, splice all the expanded feature maps in the scale dimension to obtain the spliced feature map. After the spliced feature map passes through 3D convolution with a size of 3×3×3, batch normalization, and ReLU function activation processing in sequence, perform average pooling in the time dimension to obtain a multi-scale fusion feature map F. fusion ;

[0021] S25. According to S21 - S24, obtain a multi-scale fusion feature map for each support image. The size of the multi-scale fusion feature map is the same as the size of P3.

[0022] Pool each multi-scale fusion feature map, and adjust the size of the pooled multi-scale fusion feature map to be the same as the sizes of P4 and P5 respectively. Then, for each support image, obtain the pooled multi-scale fusion feature maps with sizes of P3, P4, and P5 respectively.

[0023] For each support image, splice the pooled multi-scale fusion feature map with a size of P3 and the feature map with a size of P3 obtained in S23 in channels, splice the pooled multi-scale fusion feature map with a size of P4 and the feature map with a size of P4 obtained in S23 in channels, splice the pooled multi-scale fusion feature map with a size of P5 and the feature map with a size of P5 obtained in S23 in channels to obtain three spliced feature maps. Use 1×1 convolution to obtain the new scale feature F of each spliced feature map. i scale Then, features of three scales are obtained.

[0024] Furthermore, the dilated convolution of the global context module in S22 is 3×3 convolution, and the dilation rates are 1, 3, and 5 respectively.

[0025] Further, in S23, the sizes of the three new feature maps P3, P4, and P5 of each support image are fixed to the size of P3 through upsampling processing, and all the feature maps of each support image after modifying the size are obtained. The specific process is as follows:

[0026]

[0027] Among them, P i '(x, y) represents the feature map obtained through upsampling processing, (x, y) represents the coordinates in the upsampled feature map, (m, n) represents the coordinates in the feature map before upsampling, M and N are both upsampling factors, representing the magnification multiples of the length and width respectively, and w() represents the weight function of bilinear interpolation.

[0028] Further, the specific process of S3 is as follows:

[0029] Input the multi-scale features of each support image obtained in S2 into the foreground attention network;

[0030] Calculate the average feature F sa of all the images in the support set, perform global average pooling on the average feature F sa to obtain a global average pooling feature of C×1×1. Take the global average pooling feature as the linear matrix ω of the foreground attention network. After multiplying the linear matrix ω with each support image, successively pass through a 3×3 convolution and a Sigmoid activation function to generate the corresponding spatial attention map. Use the spatial attention map to perform foreground enhancement on the multi-scale features of the corresponding support image obtained in S2 to obtain the corresponding enhanced features, that is, obtain the enhanced features of all support images.

[0031] Further, the specific process of S4 is as follows:

[0032] Input the enhanced features of all support images obtained in S3 into the prototype generation network based on cosine distance, calculate the distances between every two of the K enhanced features of the same class of objects to obtain a K×K distance metric matrix, sum the distance metric matrix by columns to obtain a 1×K distance vector, pass the distance vector through the Softmax operation to obtain a 1×K weight vector, assign weight coefficients to the K enhanced features according to the weight vector, and finally perform weighted average to obtain the prototype of the corresponding object.

[0033] Further, the distance metric matrix D is:

[0034] D = [dist(F` s,i , F` s,j )] K×K

[0035] Among them, dist(F` s,i,F` s,j ) is the feature F` s,i and F` s,j The cosine distance between them, i, j = 1, 2,..., K.

[0036] Furthermore, the distance vector d is:

[0037] d = [d i 1×K

[0038] where d i is the sum of all elements in the i-th column of the distance metric matrix D, i = 1, 2,..., K.

[0039] Furthermore, the weight vector w is:

[0040] w = [w i 1×K

[0041] where w i is the weight of each enhanced feature, i = 1, 2,..., K, and e is the natural constant.

[0042] Furthermore, the prototype is:

[0043]

[0044] Beneficial effects:

[0045] ​​The present invention uses the DIOR dataset as the training set, obtains the images of each target in the DIOR dataset, and the target category and target location have been marked in each image. For each category of target, K marked samples of the current target are defined as the support set, and the remaining marked samples are used as the query set; a multi-scale fusion feature extraction network is constructed, the support set is input, and the multi-scale features of each support image are output. A foreground attention network is constructed, the multi-scale features of each support image are input, and the corresponding enhanced features are output. A prototype generation network based on cosine distance is constructed, the enhanced features of all support images are input, and the prototype of each category of target is output. The prototypes of all targets are used to detect the targets corresponding to each query image in the query set, and the target category and location corresponding to each query image are obtained, and the target prototypes are tested. The multi-scale fusion feature extraction network is used to extract features from the query set, and the multi-scale features of each query image in the query set are obtained. The multi-scale features are input into the RPN network for target detection, and the proposal boxes of the targets in each scale feature are obtained. The similarity between the targets corresponding to the proposal boxes and the prototypes of each category of target is calculated, and the positions of the proposal boxes are corrected according to the prototypes of each category of target. The target prototype with the highest similarity is used as the target in the image to be queried, and the category and location of the target in the image to be queried are obtained according to the target and the corrected proposal box.

[0046] The present invention performs calculations from three aspects: feature extraction, feature enhancement, and prototype generation, so as to better extract information related to the target from remote sensing images. First, the enhanced multi-scale features of the support image or query image are obtained through the multi-scale fusion feature extraction network to adapt to the scale change of the target. Secondly, through the foreground attention network, the target information of the support image is enhanced, and the influence of the background is eliminated. Finally, through the prototype generation network based on cosine distance, more stable prototypes are generated for subsequent detection of query images, which improves the effect when the target prototype is used as prior knowledge for target detection, and improves the accuracy of the target detection results of small-sample remote sensing images. Description of the Drawings

[0047] Figure 1 is the flowchart of this application;

[0048] Figure 2 is the experimental result diagram of the 1-shot bar chart of the DIOR dataset in the embodiment;

[0049] Figure 3 is the experimental result diagram of the 3-shot bar chart of the DIOR dataset in the embodiment;

[0050] Figure 4 is the experimental result diagram of the 5-shot bar chart of the DIOR dataset in the embodiment;

[0051] Figure 5It is the experimental result diagram of the bar chart of the 10-shot DIOR dataset in the embodiment; Specific implementation manner

[0052] Specific implementation manner 1: In combination with Figure 1 To illustrate this implementation manner, the method for remote sensing image target detection based on the foreground attention network described in this implementation manner includes the following steps:

[0053] S1. Obtain the images of each target in the DIOR dataset. The target type and target position have been marked in each image. For each target, define K labeled samples of the current target as the support set, and the remaining labeled samples as the query set. The images in the support set are called support images, and the images in the query set are called query images.

[0054] The DIOR dataset is one of the largest and most diverse remote sensing image target detection datasets currently publicly available, containing 20 common targets, namely airplane, airport, baseball field, basketball court, bridge, chimney, dam, expressway service area, expressway toll station, harbor, golf course, groundtrack field, overpass, ship, stadium (storage tank), tennis court, train station, vehicle, windmill. The dataset was manually annotated by Google Earth experts, mainly by labeling horizontal bounding boxes for the targets. Since the data in the DIOR database was collected from multiple remote sensing satellites, multiple time points, and multiple locations, the DIOR dataset has four significant characteristics:

[0055] (1) Large scale, a total of 23,463 remote sensing images are included, with large differences in target sizes, pixel size of 800×800, and spatial resolution ranging from 0.5m to 30m;

[0056] (2) The range of target sizes varies greatly. Due to the influence of sensor spatial resolution and category size changes, the target sizes in the DIOR dataset vary significantly;

[0057] (3) The images vary greatly, and the dataset contains remote sensing images from more than 80 countries, with significant differences in image quality, weather, seasons, imaging conditions, etc.;

[0058] (4) The similarity between classes is high, and the differences within classes are large. The dataset includes many object categories with high similarity and large differences.

[0059] Table 1 Statistical information of the DIOR dataset

[0060]

[0061] The support set contains a small number of labeled samples for training the model or providing prior knowledge, and the query set contains unlabeled samples for testing the model or evaluating the generalization ability. In the training stage, randomly sample N (N ≤ 20) target category (such as ships, windmills, airplanes, etc.) images from the DIOR dataset. For each target category, there are K labeled samples as the support set, and the remaining samples of each target category are used as the query set (this organization method is called N-way K-shot). In the testing stage, each time randomly sample N target category images from the test set. For each target category, there are K labeled samples as the support set, and the remaining samples of each category are used as the query set.

[0062] Data preprocessing: Since the N-way K-shot method needs to be used for training, the DIOR dataset needs to be reorganized to generate the corresponding few-shot dataset. First, adjust the images in the DIOR dataset to the Pascal VOC format for unified processing. Read the trainval.txt file to obtain the ID list of all image files in the DIOR dataset. For each image file ID, read the corresponding xml annotation file containing the real bounding box information of each target in the image. Create a category annotation list and add the annotation file path to it according to the category labels of all targets in the image. For each category label, create a dictionary, where the keys correspond to the number of images to be sampled (i.e., the so-called shots), and the values correspond to the list of image file paths. According to the file list of each category, randomly select different numbers of files to represent the training set data under different shots and store them in the category label dictionary. For each shot value, the number of sampled images is different. To ensure that the categories between different shots do not overlap, check whether the selected data already contains samples of the current category each time data is selected. If it already contains, do not select it repeatedly. Finally, write the training data numbers corresponding to different shot values for each category label into a separate text for subsequent reading.

[0063] S2. Construct a multi-scale fusion feature extraction network, which successively includes a backbone network, a global context module, an FPN network, and a feature fusion module. Input the support set into the multi-scale fusion feature extraction network to obtain the multi-scale features of each support image. The specific process is as follows:

[0064] S21. Input the support set into the ResNet101 backbone network of the multi-scale fusion feature extraction network. The ResNet101 backbone network generally divides the entire network into 5 output layers, each output layer corresponding to several residual blocks, and each residual block contains 3 convolutional layers and skip connections. The output feature scales of different output layers are different. Select the last three output layers as the final output layers to obtain the features output by each final output layer for each support image, that is, obtain the features of different scales of each support image.

[0065] Denote the output layer in the backbone network where the feature of each scale of each support image belongs as C i , i = 1, 2, 3, 4, 5.

[0066] S22. Since the pixels of the target are small in the detection of small targets, in order to improve the resolution ability of the target through context information, input the output features of each support image at the fifth output layer C5 of the backbone network into the global context module, and obtain feature maps representing context information of different receptive fields through dilated convolutions with different dilation rates. Since the receptive field at C5 itself is already 483×483, there is no need to set too many dilation rates like in Atrous Spatial Pyramid Pooling. Therefore, in the present invention, the dilated convolution is set as a 3×3 convolution with dilation rates of 1, 3, and 5 respectively. Each dilated convolution corresponds to a feature map, and three feature maps are obtained. Concatenate all the obtained feature maps and then generate a feature map with 256 channels through a 1x1 convolution.

[0067] F context = Conv 1×1 (Concat(DilatedConv 3×3 (C5, rate = 1, 3, 5)))

[0068] S23. Consider the features of different scales as feature maps at different times. Input the output features of the output layer C3, the output features of the output layer C4, and the feature map of the output layer C5 obtained in S22 corresponding to each support image into the FPN network, and use 256 1×1 convolutional kernels to make the number of channels of all features the same, obtaining the new feature P corresponding to each feature i , i = 3, 4, 5, that is, for each support image, three new feature maps are obtained, namely P3, P4, and P5.

[0069] P i = Conv 1×1 (C i )

[0070] A small sample refers to a sample with a small amount of data. A small amount of data will cause the traditional deep learning-based object detection model to have poor detection accuracy for objects with scale changes, especially the detection performance for small objects is not good. A small object refers to an object with a small pixel ratio, corresponding to a small scale in the scale change problem of remote sensing images. Therefore, in order to ensure the performance of small objects, the sizes of the three new feature maps P3, P4, and P5 of each support image are fixed to the size of P3 through upsampling processing. The size of P3 is Obtain all the feature maps after modifying the size of each support image:

[0071]

[0072] Among them, P i '(x, y) represents the feature map obtained through upsampling processing. (x, y) represents the coordinates in the upsampled feature map, (m, n) represents the coordinates in the feature map before upsampling. M and N are both upsampling factors, representing the magnification of length scaling and width scaling respectively. w() represents the weight function of bilinear interpolation, and the weight value is calculated according to the relative position. Since the resolutions of the feature maps P3, P4, and P4 gradually decrease and the number of channels gradually increases, the spatial information (for localization) of the feature maps gradually decreases, and the semantic information (for classification) gradually increases. Because there is no information about small objects when the resolution is too low, and the resolution of P3 is relatively large, it can retain more spatial information of small objects.

[0073] S24. Input all the feature maps after modifying the size of each support image into the feature fusion module. Expand the scale dimension t of each feature map, and splice all the expanded feature maps in the scale dimension to obtain the spliced feature map F 3d , and after sequentially passing the spliced feature map through 3D convolution with a size of 3×3×3, batch normalization, and ReLU function activation processing, perform average pooling in the time dimension to obtain a multi-scale fusion feature map F with a size of . fusion .

[0074] F 3d = Concat(Expand(P i ', dim = t), i = 3, 4, 5)

[0075] F fusion = AvgPool(ReLU(BatchNorm(Conv 3×3×3 (F 3d ))))

[0076] S25. According to S21 - S24, for each support image, a multi - scale fusion feature map can be obtained. The size of the multi - scale fusion feature map is the same as that of P3. Pool each multi - scale fusion feature map. Since the size of the pooled multi - scale fusion feature map is the same as that of P3, the size of the pooled multi - scale fusion feature map is respectively adjusted to be the same as that of P4 and P5. Then, for each support image, pooled multi - scale fusion feature maps with sizes of P3, P4, and P5 are obtained. Respectively splice each adjusted pooled multi - scale fusion feature map with the corresponding feature maps P`3, P`4, and P`5 of the corresponding support image according to the channels. That is, for each support image, splice the pooled multi - scale fusion feature map with size P3 with the corresponding feature map with size P3` according to the channels, splice the pooled multi - scale fusion feature map with size P4 with the corresponding feature map with size P4` according to the channels, and splice the pooled multi - scale fusion feature map with size P5 with the corresponding feature map with size P5` according to the channels, obtaining three spliced feature maps. Use 1×1 convolution to obtain the new - scale feature F of each spliced feature map i scale :

[0077] F i scale = Conv 1×1 (Concat(AvgPool(F fusion ), C i ), i = 3, 4, 5

[0078] Then, for each support image, three scale features with scale robustness are obtained That is, the multi - scale features of each support image are obtained.

[0079] When using multi - scale features for object detection traditionally, feature fusion at different scales is usually carried out in FPN. However, directly adding features at different scales improves the multi - scale performance but ignores the problem of conflicting information between different - scale features. Therefore, to solve the problem of scale mismatch between the query set and the support set, the present invention proposes a method for fusing different - scale features through 3D convolution of the feature fusion module. With the help of 3D convolution and dilated convolution of the global context module, features at each scale have a certain scale invariance and global information. 3D convolution is usually used in video tasks, and its convolution kernel has one more scale dimension than 2D convolution.

[0080] S3. Construct a foreground attention network for feature enhancement. Input the multi - scale features of each support image into the foreground attention network to obtain the corresponding enhanced features.

[0081] The foreground attention network improves the spatial attention module. The basic structure remains unchanged (underlined part), only the linear matrix has changed. With the overall information of the support set, spatial attention is generated by per-channel convolution to enhance the foreground. The previous method was to learn a linear matrix ω of dimension C×1 and multiply it with the support set features of C×H×W through matrix multiplication. Then, a matrix of 1×H×W is obtained through Softmax as the spatial attention. However, this method does not directly utilize the information of the support set images and does not perform well on the support images of new classes. Therefore, in the foreground attention network of the present invention, the average feature F of all images in the support set is calculated sa , for the average feature F sa Perform global average pooling to obtain a global average pooling feature of C×1×1, and use the global average pooling feature as the linear matrix ω of the foreground attention network. Since most of the pixels in the vector intercepted by the target box are the pixels corresponding to the detected target, the main proportion of the feature after average pooling should be the feature of the target, so the activation value in the foreground area will be higher. After multiplying the linear matrix ω with each support image F s , through a 3×3 convolution and a Sigmoid activation function, the corresponding spatial attention map is generated, and the spatial attention map is used to perform foreground enhancement operations on the multi-scale features of the corresponding support images obtained by S2 to obtain the corresponding enhanced features:

[0082] F s ' = F s ·Sigmoid(Conv 1×1 (F s ⊙GlobalAvgPool(F sa )))

[0083] Feature enhancement is performed for each scale feature. Then, each support image has three scale features, and through the enhanced features, the corresponding three enhanced features are obtained.

[0084] S4. Construct a prototype generation network based on cosine distance. Input the enhanced features of all support images obtained in S3 into the prototype generation network based on cosine distance. For the same target, different weights are assigned to different enhanced features of the current target through cosine distance, and the weighted average of all weights is calculated to obtain the prototype of the current target, and the prototype of each target can be obtained.

[0085] Since the differences among multiple support images of the same target may lead to significant differences in the feature maps of the support images of the same target. If we want to make good use of all the information in the support set, the most straightforward method is not to use a prototype to represent all the information in the support set, but to process each image separately. Although this method can improve the performance of the network, it causes the computational complexity to increase exponentially with the increase in the number and variety of small samples. Therefore, to balance performance and cost, in order to better retain the similarities among all the support set images and downplay the differences, the present invention generates a corresponding prototype from multiple images of the same target (airplane or ship or airport) for subsequent small sample detection to adapt to the performance changes brought about by different support images in the case of small samples. Specifically, all the enhanced features F corresponding to the support images obtained by the foreground attention network in S3 s ` are input into the prototype generation network based on cosine distance, and the K enhanced features F of the same class of targets are calculated s ` to obtain a K×K distance metric matrix D (the diagonal elements are 0):

[0086] D = [dist(F` s,i , F` s,j )] K×K

[0087] where dist(F` s,i , F` s,j ) is the cosine distance between the features F` s,i and F` s,j , and i, j = 1, 2,..., K.

[0088] The distance metric matrix D is summed by column to obtain a 1×K distance vector d, and the distance vector d is passed through the Softmax operation to obtain a 1×K weight vector w:

[0089] d = [d i 1×K

[0090] i = 1, 2,..., K

[0091] w = [w i 1×K

[0092]

[0093] where d i is the sum of all elements in the i-th column of the distance metric matrix D, and w i is the weight of each enhanced feature, and e is the natural constant.

[0094] Assign weight coefficients to the K enhanced feature maps according to the weight vector, so that the features clustered together have greater weights and the relatively discrete features have smaller weights, and finally perform weighted averaging to obtain the prototype corresponding to the target:

[0095]

[0096] S5. Detect the targets corresponding to each query image in the query set using the prototypes of all targets, and obtain the target categories and positions corresponding to each query image.

[0097] S6. Based on the target categories and target positions corresponding to each query image, repeat steps S2 - S5 to update the prototypes of each type of target until the upper limit of the iteration times is met, and obtain the prototypes of each type of target.

[0098] S7. Use S2 to extract the features of the image to be queried, obtain the multi-scale features of the image to be queried, input the multi-scale features into the RPN network, obtain the proposed bounding boxes of the targets in each scale feature, calculate the similarity between the targets corresponding to the proposed bounding boxes and the prototypes of each type of target obtained in S6, and correct the positions of the proposed bounding boxes according to the prototypes of each type of target respectively. Take the target prototype with the highest similarity as the target in the image to be queried, and obtain the target category and position in the image to be queried based on the target and the corrected proposed bounding box.

[0099] The above feature extraction and target detection are consistent with the classical algorithm Faster RCNN. Only after detecting the proposed bounding boxes, directly judging the target category is changed to comparing with the prototypes of different target categories to determine the target category and location in the image. The present invention obtains better and more accurate prototypes through features, and uses the prototypes for the detection of the multi-scale feature maps of the subsequent query set, which can improve the detection performance.

[0100] Embodiment

[0101] To verify the effect of the present invention, it is compared with several classical few-shot object detection models FSRW, FSODM, TFA, and DeFRCN, where FSRW and FSODM use the method of meta-learning, and TFA and DeFRCN use the method of fine-tuning. The specific performance test is to compare the AP values and mAP values of these methods on five categories including airplane, baseball field, tennis court, railway station, and windmill when the number of samples (shot) is 1, 3, 5, and 10.

[0102] Figure 2The detection performance comparison of the present invention and each of the above models under the 1-shot condition is shown. It can be seen that under the 1-shot condition, according to the mAP, the method of the present invention reaches 10.33, which is the best, 0.58 higher than the second place. The method of meta-learning has a better detection effect on the tennis court than transfer learning. This may be because the tennis court is not large and the texture is not very clear. It is difficult to learn the features of this category through fine-tuning under the 1-shot condition, while the meta-learning method can better adapt to these changes. For types such as airplanes, train stations, and windmills, the effects of various methods are not very ideal. This may be because the objects themselves are small or the shapes are too extreme, so it is difficult to learn in the case of a single example.

[0103] Figure 3 The detection performance comparison of the present invention and each of the above models under the 3-shot condition is shown. Compared with 1-shot, with the increase in the number of support set images, the performance of the models has been greatly improved. Even if only two instances are added, the mAP values of all models have reached twice the original. This shows that under the small sample condition, the performance improvement brought by a good network design may not be as good as adding more data. The method of the present invention improves the mAP index by 1.3 compared with the second place DeFRCN. Especially in the tennis court and airplane categories, the performance has been improved by 13.85 and 1.26 respectively.

[0104] Figure 4 The detection performance comparison of the present invention and each of the above models under the 5-shot condition is shown. It can be seen that the method of transfer learning has obvious performance improvement in all categories. Among them, DeFRCN is 4.08 and 8.28 higher than the method of the present invention in the baseball field, train station, and windmill, and the overall mAP is also 0.12 higher than the method of the present invention. This may be because DeFRCN finely adjusts the learning rate, so it can better learn the features corresponding to the samples. And because the transfer learning method does not require preprocessing of the support set images, the original images can better maintain the semantic and positional relationships between the target and the context, so the model can learn how to locate the corresponding category from the pictures. However, because the present invention uses multi-scale attention in the region proposal network and uses a rotation strategy during matching, it is 1.29 and 18.88 higher than DeFRCN in the tennis court and airplane respectively.

[0105] Figure 5Shows the comparison of the detection performance of each model under the 10-shot condition. The performance of the present invention is still the highest among these methods, and the mAP value is 1.38 higher than that of the second place. Due to the increase in samples, compared with the three meta-learning methods, the performance of the transfer learning method has improved significantly on the tennis court, so the performance gap in this category has been greatly reduced. For other categories, the first two meta-learning methods are inferior to the transfer learning method, and the method of the present invention has increased the AP value by 3.20 and 9.87 respectively compared with the second place on the airplane and the tennis court.

[0106] It can be seen from the experimental results that the performance of the few-shot object detection model without improvement for remote sensing images is uneven on remote sensing images, and the performance needs to be improved compared with natural images. And our model has been significantly improved. This is sufficient to prove the effectiveness of our method.

[0107]

[0108]

Claims

1. A remote sensing image target detection method based on a foreground attention network, characterized in that: It includes the following steps: S1. Obtain the images of each target in the DIOR dataset. The target category and location are labeled in each image. For each category of target, define K labeled samples of the current target as the support set, and the remaining labeled samples as the query set. The images in the support set are called support images, and the images in the query set are called query images; S2. Construct a multi-scale fusion feature extraction network. The multi-scale fusion feature extraction network sequentially includes a backbone network, a global context module, an FPN network, and a feature fusion module. Input the support set into the multi-scale fusion feature extraction network to obtain the multi-scale features of each support image; S3. Construct a foreground attention network. Input the multi-scale features of each support image into the foreground attention network to obtain the corresponding enhanced features; S4. Construct a prototype generation network based on cosine distance. Input the enhanced features of all support images into the prototype generation network based on cosine distance. For the same target, assign different weights to the different enhanced features of the current target through cosine distance, and perform weighted average on all weights to obtain the prototype of the current target, that is, the prototype of each category of target can be obtained; S5. Use the prototypes of all targets to detect the targets corresponding to each query image in the query set, and obtain the target category and location corresponding to each query image; S6. Based on the target category and target location corresponding to each query image, repeat S2 - S5 to update the prototype of each category of target until the upper limit of the number of iterations is met, and obtain the prototype of each category of target; S7. Use S2 to extract features from the image to be queried, obtain the multi-scale features of the image to be queried, and input the multi-scale features into the RPN network to obtain the proposed bounding boxes of the targets in each scale feature; Calculate the similarity between the targets corresponding to the proposed bounding boxes and the prototypes of each category of target obtained in S6, and correct the positions of the proposed bounding boxes according to the prototypes of each category of target. Take the prototype of the target with the highest similarity as the target in the image to be queried, and obtain the category and location of the target in the image to be queried according to the target and the corrected proposed bounding box.

2. The method for remote sensing image target detection based on the foreground attention network according to claim 1, characterized in that: The specific process of S2 is as follows: S21. Input the support set into the backbone network of the multi-scale fusion feature extraction network. Set the backbone network to include 5 output layers. The output feature scales of different output layers are different. Select the last three output layers as the final output layers, and obtain the features output by each support image through each final output layer, that is, obtain the features of different scales of each support image; Denote the output layer in the backbone network to which the features of each scale of each support image belong as C i , where i = 1, 2, 3, 4, 5; S22. Input the output features of each support image at the fifth output layer C5 of the backbone network into the global context module. Obtain the feature maps representing context information of different receptive fields through dilated convolutions with different dilation rates, splice all the obtained feature maps, and then generate a feature map with 256 channels through 1x1 convolution; S23. Treat the features of different scales as feature maps at different times. Input the output features of the output layer C3, the output features of the output layer C4, and the feature map of the output layer C5 obtained in S22 corresponding to each support image into the FPN network. Use 256 1×1 convolutional kernels to make the number of channels of all features consistent, and obtain the new feature P corresponding to each feature. i , where i = 3, 4, 5. That is, for each support image, three new feature maps are obtained, namely P3, P4, and P5. Through upsampling processing, fix the sizes of the three new feature maps P3, P4, and P5 of each support image to the size of P3, and obtain all the feature maps with modified sizes of each support image; S24. Input all the feature maps after resizing each support image into the feature fusion module. Expand the dilation scale dimension t of each feature map, and concatenate all the expanded feature maps in the scale dimension to obtain the concatenated feature map. After the concatenated feature map passes through 3D convolution with a size of 3×3×3, batch normalization, and ReLU function activation processing in sequence, perform average pooling in the time dimension to obtain a multi-scale fusion feature map F fusion ; S25. According to S21 - S24, obtain a multi-scale fusion feature map for each support image. The size of the multi-scale fusion feature map is the same as the size of P3; Pool each multi-scale fusion feature map, and adjust the sizes of the pooled multi-scale fusion feature maps to be the same as those of P4 and P5 respectively. Then, for each support image, obtain the pooled multi-scale fusion feature maps with sizes of P3, P4, and P5 respectively. For each supported image, the pooled multi-scale fusion feature map with size P3 is concatenated with the feature map of size P3 obtained by S23 processing along the channels, the pooled multi-scale fusion feature map with size P4 is concatenated with the feature map of size P4 obtained by S23 processing along the channels, and the pooled multi-scale fusion feature map with size P5 is concatenated with the feature map of size P5 obtained by S23 processing along the channels, resulting in three concatenated feature maps. The new-scale feature F of each concatenated feature map is obtained using 1×1 convolution i scale , thus obtaining features F3 at three scales scale , F4 scale , F5 scale .

3. The remote sensing image target detection method based on the foreground attention network according to claim 2, characterized in that: The dilated convolution of the global context module in S22 is a 3×3 convolution with dilation rates of 1, 3, and 5 respectively.

4. The remote sensing image target detection method based on the foreground attention network according to claim 3, characterized in that: In S23, through upsampling processing, fix the sizes of the three new feature maps P3, P4, and P5 of each support image to the size of P3, and obtain all the feature maps of each support image after modifying the sizes. The specific process is as follows: Among them, P i '(x, y) represents the feature map obtained after upsampling processing, (x, y) represents the coordinates in the upsampled feature map, (m, n) represents the coordinates in the feature map before upsampling, M and N are both upsampling factors, representing the magnification of length scaling and width scaling respectively, and w() represents the weight function of bilinear interpolation.

5. The remote sensing image target detection method based on the foreground attention network according to claim 4, characterized in that: The specific process of S3 is as follows: Input the multi-scale features of each support image obtained in S2 into the foreground attention network. Calculate the average feature F of all images in the support set sa , for the average feature F sa Perform global average pooling to obtain a global average pooling feature of C×1×1. Take the global average pooling feature as the linear matrix ω of the foreground attention network. After multiplying the linear matrix ω with each support image, successively pass through a 3×3 convolution and a Sigmoid activation function to generate the corresponding spatial attention map. Use the spatial attention map to perform foreground enhancement on the multi-scale features of the corresponding support image obtained by S2 to obtain the corresponding enhanced features, that is, obtain the enhanced features of all support images.

6. The method for remote sensing image target detection based on a foreground attention network according to claim 5, characterized in that: The specific process of S4 is as follows: Input the enhanced features of all support images obtained in S3 into the prototype generation network based on cosine distance, calculate the distances between every two of the K enhanced features of the same class of targets to obtain a K×K distance metric matrix, sum the distance metric matrix by column to obtain a 1×K distance vector, pass the distance vector through the Softmax operation to obtain a 1×K weight vector, assign weight coefficients to the K enhanced features according to the weight vector, and finally perform weighted average to obtain the prototype of the corresponding target.

7. The method for remote sensing image target detection based on a foreground attention network according to claim 6, wherein: The distance metric matrix D is: D = [dist(F s,i `, F s,j `)] K×K where dist(F s `, F ,i `) is the cosine distance between feature F s ` ,j and F s `, i, j = 1, 2, ..., K. ,i and F s `, ,j and i, j = 1, 2, ..., K.

8. The method for remote sensing image target detection based on the foreground attention network according to claim 7, wherein: The distance vector d is: d = [d i 1×K ​ where d i is the sum of all elements in the i-th column of the distance metric matrix D, 9. The method for remote sensing image object detection based on a foreground attention network according to claim 8, wherein: The weight vector w is: w = [w i 1×K ​ where, w i is the weight of each enhancement feature, e is the natural constant.

10. The remote sensing image target detection method based on a foreground attention network according to claim 9, characterized in that: The prototype is:

Citation Information

Patent Citations

  • Hyperspectral image semi-supervised classification method based on small sample learning

    CN113408605A

  • Method for recognizing distribution network equipment based on raspberry pi multi-scale feature fusion

    US11631238B1