Few-sample remote sensing image target identification method and device

By introducing channel and spatial attention modules and multi-head attention mechanisms into the ResNet-50 network, combined with the Vision Transformer encoder, the problem of underfitting in the recognition of remote sensing images with few samples was solved, and the recognition effect with high accuracy was achieved.

CN120198709APending Publication Date: 2025-06-24PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510173912.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing remote sensing image target recognition method is prone to underfitting when training data is limited, and it is difficult to capture rich information of the target image from a small number of samples, resulting in unsatisfactory recognition performance.

Method used

A small sample remote sensing image target recognition network based on ResNet-50 is adopted, combining channel and spatial attention modules, multi-head attention mechanism and Vision Transformer encoder are used to extract multi-scale convolutional features and fuse top-level semantic features to generate rich target recognition features.

Benefits of technology

The recognition effect of remote sensing images with few samples is improved. The generated features are both detailed and semantic of the underlying features, which significantly improves the classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198709A_ABST
    Figure CN120198709A_ABST
Patent Text Reader

Abstract

The invention provides a few-sample remote sensing image target identification method and device, and the method comprises the steps: obtaining a few-sample remote sensing image, randomly selecting the remote sensing image as training data, and carrying out the preprocessing; determining enhanced features corresponding to the convolution features of the four Block layers by using the configured ResNet-50 network in combination with a channel and a space attention module; respectively calculating multi-head attention guidance information by utilizing the enhancement characteristics of the first three Block layers, and determining the characteristics of the embedded layer by utilizing the enhancement characteristics of the third Block layer; the multi-head attention features are calculated through the embedded layer features and the multi-head attention guiding information of the third Block layer, the second Block layer and the first Block layer in sequence; fusing the multi-head attention features with the enhanced features of the fourth Block layer, and determining target recognition features of the few-sample remote sensing image; and inputting the target recognition features into a recognition unit, and determining an optimal recognition model of the few-sample remote sensing image target. The method can improve the recognition effect of few-sample remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to a few-shot remote sensing image target recognition method and device. Background Art

[0002] Few-shot remote sensing image target recognition is one of the important requirements and means in the field of earth observation. By implementing sufficient and detailed feature learning on limited remote sensing images, instant monitoring and dynamic feature analysis of various targets on the earth's surface can be realized, so as to provide detailed and diversified data support for earth observation research and achieve "seeing the whole leopard from one spot" based on remote sensing image observation. Currently, with the continuous evolution of artificial intelligence technology, the demand in the field of remote sensing image target recognition is increasing day by day. Especially in achieving high-accuracy target recognition using a small number of remote sensing images, due to limited data resources available for model training, network learning is often insufficient, and underfitting often occurs during the model training process. As a result, it is difficult to capture rich information of target images from a relatively small number of samples to obtain ideal image target recognition performance.

[0003] Current few-shot remote sensing image target recognition methods can be divided into methods based on shallow features and methods based on deep features. The shallow feature method is based on features such as color, texture, and shape of remote sensing images, and discriminates the category of image targets by comparing parameters of a classifier composed of, for example, gray-level co-occurrence matrix, histogram of oriented gradients, and Gabor filters. This type of method not only consumes a large amount of manpower but also has a large error in recognizing complex remote sensing images. The deep feature method is to extract features of few-shot remote sensing images using machine learning methods, and perform iterative training through a classifier composed of a multi-layer neural network, support vector machine, or K-nearest neighbor algorithm, etc., to obtain the best classification model for remote sensing image target recognition.

[0004] CN116109629A discloses an image defect classification model. A training set and a validation set are constructed using product defect images, and the learning of important features of product defect images by the model is improved by adding an attention mechanism in the ResNet-50 network, achieving a relatively high fine-grained classification performance.

[0005] CN115953592A discloses a terahertz security inspection image recognition method based on a variational autoencoder. The terahertz security inspection image is input into the variational autoencoder for data reconstruction, and then an attention module is connected after DenseNet-201 to guide the network to focus on useful information. At the same time, metric learning is used to increase the inter-class distance and reduce the intra-class distance to alleviate the problem of high feature similarity between images, thereby improving the model recognition performance.

[0006] CN115984578A discloses a skin image recognition method based on DenseNet and Transformer. First, DenseNet is used to extract local features of skin images, then Transformer is utilized to learn the local features to obtain global features of skin images, and finally, the recognition is carried out by using the fusion information of the two kinds of features.

[0007] The above methods are basically image target recognition models obtained when the data volumes of training data and validation data are quite equal and the image contents are extremely similar. Therefore, the effect of its model in actual image recognition remains to be tested. In addition, the existing methods use relatively simple ways to extract few-shot image features by neural networks, and the utilization rate of deep features for few-shot data is limited. Either only the top-layer features of the network are used, or the attention mechanism is added, etc., without fully exerting the advantages of convolutional networks and Transformer in extracting image target features. Summary of the Invention

[0008] The technical solution adopted by the present invention is to design a few-shot remote sensing image target recognition network based on ResNet-50 to solve the above problems and improve the recognition effect of few-shot remote sensing images. In view of this, the present invention provides a few-shot remote sensing image target recognition method and device.

[0009] The technical solution of the present invention proposes a few-shot remote sensing image target recognition method, including: Step S1, obtain few-shot remote sensing images, randomly select a relatively small number of images as training data, and perform preprocessing; Step S2, use the configured ResNet-50 network, combined with channel and spatial attention modules, to determine the enhanced features corresponding to the convolutional features of the first Block layer, the second Block layer, the third Block layer, and the fourth Block layer; Step S3, use the enhanced features of the first 3 Block layers to calculate multi-head attention guidance information respectively, and use the enhanced features of the third Block layer to determine the embedded layer features; Step S4, sequentially use the embedded layer features, the multi-head attention guidance information of the third Block layer, the second Block layer, and the first Block layer to calculate multi-head attention features; Step S5, fuse the multi-head attention features with the enhanced features of the fourth Block layer to determine the target recognition features of few-shot remote sensing images; Step S6, input the target recognition features into the recognition unit to determine the best recognition model for few-shot remote sensing image targets.

[0010] In one embodiment, the step S1 includes: Step S101, randomly crop the training data, where the crop covers at least 80% of the training data picture; Step S102, randomly rotate the training data, where the rotation range is between -45 degrees and 45 degrees; Step S103, horizontally flip the current training data with a probability of 0.5; Step S104, crop from the center of the current training data to both sides; Step S105, convert the current training data into tensor data of a preset shape; Step S106, normalize the tensor data channel by channel, change the mean to 0 and the standard deviation to 1 to obtain the preprocessed training data.

[0011] In one embodiment, the step S2 includes: Step S201, build a ResNet-50 network and import the pre-trained parameter "resnet50-11ad3fa6.pth" obtained on the ImageNet dataset; Step S202, input the preprocessed training data; Step S203, use the ResNet-50 network to extract the convolutional features of the 1st Block layer, the 2nd Block layer, the 3rd Block layer, and the 4th Block layer; Step S204, use the channel and spatial attention modules to calculate the enhanced features of the 1st Block layer, the 2nd Block layer, the 3rd Block layer, and the 4th Block layer corresponding to the 4 convolutional features.

[0012] In one embodiment, the step S3 includes: Step S301, calculate the embedding layer features of the 3rd Block layer; Step S302, calculate the guiding information of the 3rd Block layer; Step S303, calculate the guiding information of the 2nd Block layer; Step S304, calculate the guiding information of the 1st Block layer.

[0013] In one embodiment, the step S4 includes: Step S401, calculate the multi-head attention features of "Transformer Encoder1"; Step S402, calculate the multi-head attention features of "Transformer Encoder2"; Step S403, calculate the multi-head attention features of "Transformer Encoder3".

[0014] In one embodiment, step S5 includes: Step S501: Fuse the multi-head attention features of "Transformer Encoder3" with the enhanced features of the 4th Block layer to obtain the target recognition features of the few-shot remote sensing image.

[0015] In one embodiment, step S6 includes: Step S601: Use the global average pooling function to calculate the mean of all pixel values of each channel of the target recognition features to obtain the target pooling features; Step S602: Flatten the target pooling features to obtain one-dimensional vector features; Step S603: Input the one-dimensional vector features into a fully connected layer to output fully connected features.

[0016] Step S604: Input the features of the fully connected layer into the Softmax function to obtain the prediction probability; Step S605: Use the cross-entropy loss function to determine the distance loss between the true probability and the prediction probability; Step S606: Determine the model corresponding to the minimum value of the distance loss as the best target recognition model for the few-shot remote sensing image.

[0017] On the other hand, the present invention also provides a few-shot remote sensing image target recognition device, including: A few-shot remote sensing image preprocessing unit, configured to acquire remote sensing image training data and preprocess it; A multi-scale feature enhancement unit, configured to use the configured ResNest-50 network to determine the convolutional features corresponding to the current training data; An embedded layer feature calculation unit, configured to perform "image patch segmentation" and "trainable parameter embedding" operations using the enhanced features of the 3rd Block layer to determine the embedded layer features; A multi-head attention guidance information calculation unit, configured to determine the guidance information using maximum pooling, average pooling, channel dimensionality reduction, image patch embedding feature calculation, and embedded class token operations; A multi-head attention feature calculation unit, configured to perform multi-head self-attention feature calculation and multi-head attention guidance feature calculation using "Transformer Encoder" to determine the multi-head attention features; A target recognition feature calculation unit, configured to fuse the multi-head attention features with the enhanced features of the 4th Block layer to obtain the classification features of the few-shot remote sensing image target; The few-shot remote sensing image recognition unit is configured to determine the optimal model for few-shot remote sensing image recognition based on the classification features through a fully connected network, a Softmax function, and a cross-entropy loss function.

[0018] Adopting the above technical solution, the present invention has at least the following advantages: 1) This application uses the multi-scale convolutional features extracted by the first 3 Block layers of the ResNet-50 network to calculate the attention guidance information, and uses the Vision Transformer to learn the convolutional features extracted by the ResNet-50 network. Under the guidance of attention information at different layers, continuous deep learning is carried out through the Vision Transformer encoder. In addition, the top-level semantic features extracted by the ResNet-50 network are fused. Finally, the obtained few-shot remote sensing image features are rich in content, having both the details of the underlying features and the semantics of the high-level features. Therefore, the method of this application has excellent classification accuracy.

[0019] 2) This application selects the ResNet-50 convolutional network and combines the encoder part of the Vision Transformer without a more complex neural network. And the feature output layer of ResNet-50 is used as the input of the Vision Transformer after passing through the Embedding layer. In the encoder, the attention guidance information of different convolutional features is simply added without a particularly large number of parameters. Therefore, the method of this application is simpler and more effective. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 is the basic flowchart of the few-shot remote sensing image recognition method according to an embodiment of the present invention; Figure 2 is the schematic diagram of the network architecture of the few-shot remote sensing image recognition method according to an embodiment of the present invention; Figure 3 is the schematic diagram of the calculation process of the multi-head attention guidance information according to an embodiment of the present invention; Figure 4 is the schematic diagram of the calculation process of the multi-head attention guidance features according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the present invention will be described in detail as follows in combination with the accompanying drawings and preferred embodiments.

[0022] In the accompanying drawings, although exemplary embodiments of the present invention are shown, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art. The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0023] The first embodiment of the present invention proposes a few-shot remote sensing image target recognition method and device, as Figure 1 shown, including the following steps: Step S1, obtain a few-shot remote sensing image, randomly select a small number of images as training data, and preprocess them; Step S2, input the data to be predicted into the constructed ResNet-50 network to extract the convolutional features of the 1st Block layer, the 2nd Block layer, the 3rd Block layer, and the 4th Block layer using the ResNet-50 network; for each layer of the extracted convolutional features, use the channel and spatial attention modules to obtain the enhanced features corresponding to each layer of convolutional features; Step S3, use the enhanced features of the 1st Block layer, the 2nd Block layer, and the 3rd Block layer to calculate the multi-head attention guiding information respectively; perform image block segmentation based on the enhanced features of the 3rd Block layer, and embed the class token and position information to obtain the embedded layer features; Step S4, input the multi-head attention mechanism guiding information of the 3 Block layers into "Transformer Encoder1", "Transformer Encoder2", and "Transformer Encoder3" respectively, and connect them in sequence. The input of "Transformer Encoder1" also includes the embedded layer features; Step S5, fuse the output of "Transformer Encoder3" and the enhanced features of the 4th Block layer in the channel dimension to obtain the target recognition features of the few-shot remote sensing image; Step S6, input the target recognition features of the few-shot remote sensing image into the recognition unit to determine the best recognition model for the few-shot remote sensing image target.

[0024] Referring to Figure 2 , the method provided in this embodiment will be described in detail step by step below.

[0025] Step S1: Obtain few-shot remote sensing images, randomly select a small number of images as training data, and perform preprocessing. Data preprocessing is to reduce noise and increase the amount of training data, making the model more generalizable. Use the `transforms` function in PyTorch to perform the following processing on the few-shot remote sensing image training set: Step S101: Randomly crop the training data to 256×256 pixels, ensuring that at least 80% of the picture is covered during the cropping process; Step S102: Randomly rotate the training data within the range of -45 degrees to 45 degrees; Step S103: Horizontally flip the training data with a probability of 0.5; Step S104: Crop from the center of the training data along both sides, and the size of the cropped picture is 224×224 pixels; Step S105: Convert the training data into tensor format; Step S106: Normalize the tensor data channel by channel, changing the mean to 0 and the standard deviation to 1; Step S2: Use the configured ResNet-50 network to determine the convolutional features of the first Block layer, the second Block layer, the third Block layer, and the fourth Block layer corresponding to the current training data. Among them, the ResNet-50 network includes pre-trained parameters imported on the ImageNet dataset.

[0026] In this embodiment of Step S201, a ResNet-50 network is built, the pre-trained parameter "resnet50-11ad3fa6.pth" obtained on the ImageNet dataset is imported, and then the few-shot remote sensing image training data is input. The ResNet-50 network performs convolutional operations on the training data, and sequentially extracts the output convolutional features of 4 Block layers 、 、 、 。

[0027] Step S202: Use the channel and spatial attention module to perform feature enhancement on the convolutional layer features of four different scales 、 、 、 respectively, to obtain the enhanced features 、 、 、 。

[0028] Step S3: Calculate the multi-head attention guiding information respectively using the enhanced features of the first three Block layers; perform image patch segmentation on the enhanced feature of the third layer, and embed class tokens and positional information to obtain the embedded layer features. Step S301: Calculate the embedded layer features of the third Block layer. As Figure 2 shown, an exemplary process is as follows: First, perform "image patch segmentation" and "embedding of trainable parameters" operations on the enhanced features of the ResNet-50 network. Perform "image patch segmentation" and "embedding of trainable parameters" operations on the enhanced features of the ResNet-50 network.

[0029] (1) "Image patch segmentation" Segment the enhanced features into 196 image patches, each with a size of 1×1 pixel.

[0030] (2) "Embedding of trainable parameters" Map each image patch to a one-dimensional vector, process the last two dimensions using the Flatten function in PyTorch; then embed class tokens and positional information to finally obtain the embedded layer features .

[0031] Step S302: Calculate the guiding information of the first Block layer, as Figure 3 shown; Perform max pooling and average pooling operations respectively on the features obtained after feature enhancement to obtain local features and global features , and then perform channel dimensionality reduction, calculation of image patch embedding features, and embedding of class tokens to finally obtain the guiding information. Perform max pooling and average pooling operations respectively on the features obtained after feature enhancement to obtain local features

[0032] and global features and then perform channel dimensionality reduction, calculation of image patch embedding features, and embedding of class tokens to finally obtain the guiding information. and local features using 48 convolutional kernels with a size of 4×4 and a stride of 4 for channel dimensionality reduction to obtain the reduced-dimensional features of the global features and the reduced-dimensional features of the local features .

[0033] (2)Image patch embedding features Stretch the reduced-dimensional features of the local features into one-dimensional vectors respectively using the Flatten function, and use the transforse function to transform the first and second dimensions to finally obtain the local image patch embedding features after max pooling .

[0034] For the reduced-dimensional features of the global features The Flatten function is used to stretch it into a one-dimensional vector, and the transforse function is used to transform the first and second dimensions, and finally the global image block embedding feature after average pooling is obtained. .

[0035] (3) Embedding category token Initialize a trainable parameter token and use the Contact function to embed it with the local image block. Splicing to obtain the local image block embedding features of the embedded category token ; Initialize a trainable parameter token and use the Contact function to embed it with the global image block feature Splicing to obtain the global image block embedding features of the embedded category token .

[0036] (4) Obtaining guidance information The different semantic information that may be generated by different image block positions is added to the image block embedding tensor to make up for the lack of position information and obtain convolution features. Corresponding guidance information and .

[0037] Step S303, calculating the guidance information of the second Block layer; Enhanced Features Perform maximum pooling and average pooling operations respectively to obtain local features and global features , and then perform channel dimension reduction, image block embedding feature calculation, embedding category token, and finally obtain guidance information.

[0038] (1) Channel Dimensionality Reduction For global features and local features Use 48 convolution kernels with a size of 1×1 and a step size of 2 to perform channel dimensionality reduction operations to obtain the dimensionality reduction features of the global features. And the dimension reduction features of local features .

[0039] (2) Image block embedding features Dimensionality reduction of local features The Flatten function is used to stretch them into one-dimensional vectors, and the transforse function is used to transform the first and second dimensions, and finally the local image block embedding features after maximum pooling are obtained. .

[0040] Dimensionality reduction of global features The Flatten function is used to stretch it into a one-dimensional vector, and the transforse function is used to transform the first and second dimensions, and finally the global image block embedding feature after average pooling is obtained. .

[0041] (3) Embedding category token Initialize a trainable parameter token and use the Contact function to embed it with the local image block. Splicing to obtain the local image block embedding features of the embedded category token ; Initialize a trainable parameter token and use the Contact function to embed it with the global image block feature Splicing to obtain the global image block embedding features of the embedded category token .

[0042] (4) Obtaining guidance information The different semantic information that may be generated by different image block positions is added to the image block embedding tensor to make up for the lack of position information and obtain convolution features. Corresponding guidance information and .

[0043] Step S304, calculating the guidance information of the third Block layer; Enhanced Features Perform maximum pooling and average pooling operations respectively to obtain local features and global features , and then perform channel dimension reduction, image block embedding feature calculation, embedding category token, and finally obtain guidance information.

[0044] (1) Channel Dimensionality Reduction For global features and local features Use 48 convolution kernels with a size of 1×1 and a step size of 1 to perform channel dimensionality reduction to obtain the dimensionality reduction features of the global features. And the dimension reduction features of local features .

[0045] (2) Image block embedding features Dimensionality reduction of local features The Flatten function in PyTorch is used to stretch it into a one-dimensional vector, and the transforse function is used to transform the first and second dimensions, and finally the local image block embedding features after maximum pooling are obtained. .

[0046] Similarly, the global image patch embedding feature of the feature is obtained. .

[0047] (3) Embedding class token Initialize a trainable parameter token, and use the Contact function to concatenate it with the local image patch embedding feature to obtain the local image patch embedding feature with the embedded class token ; Similarly, initialize a trainable parameter token, and use the Contact function to concatenate it with the global image patch embedding feature to obtain the global image patch embedding feature with the embedded class token .

[0048] (4) Obtain guiding information Add different semantic information that may be generated at different image patch positions to the image patch embedding tensor to make up for the lack of position information, and obtain the corresponding guiding information of the convolutional feature and .

[0049] In step S4, the guiding information of the multi-head attention mechanism of the 3 Block layers is respectively input into "Transformer Encoder1", "Transformer Encoder2", and "Transformer Encoder3", and they are connected in sequence. The input of "Transformer Encoder1" also includes the embedding layer feature, as Figure 4 shown; In step S401, calculate the multi-head attention feature of "Transformer Encoder1"; Input the attention guiding information and of the third layer feature together with the embedding layer feature into "Transformer Encoder1" for multi-head attention calculation, which is divided into multi-head self-attention feature calculation and multi-head attention guiding feature calculation: (1) Multi-head self-attention feature calculation Normalize the input feature , and then perform multi-head self-attention calculation to obtain the multi-head self-attention information .

[0050] Combine the multi-head self-attention information and the input feature Add element by element and then perform normalization to obtain the multi-head self-attention features 。

[0051] (2) Multi-head attention-guided feature calculation Perform multi-head self-attention calculation on the multi-head self-attention features along with the guidance information and to obtain the multi-head attention-guided information 。

[0052] Add the multi-head attention-guided information and the multi-head self-attention features element by element, and then perform normalization to obtain the multi-head attention-guided features 。

[0053] Step S402: Calculate the multi-head attention features of "Transformer Encoder2"; Input the guidance information of the second-layer features 、 together with the calculation result of the previous step into "Transformer Encoder2" for multi-head attention calculation, which is divided into multi-head self-attention feature calculation and multi-head attention-guided feature calculation: (1) Multi-head self-attention feature calculation Normalize the input features and then perform multi-head self-attention calculation to obtain the multi-head self-attention information 。

[0054] Add the multi-head self-attention information and the input features element by element, and then perform normalization to obtain the multi-head self-attention features 。

[0055] (2) Multi-head attention-guided feature calculation Perform multi-head self-attention calculation on the multi-head self-attention features along with the calculated guidance information and to obtain the multi-head attention-guided information 。

[0056] Add the multi-head attention-guided information and the multi-head self-attention features element by element, and then perform normalization to obtain the multi-head attention-guided features 。

[0057] Step S403: Calculate the multi-head attention features of "Transformer Encoder3". The guiding information of the first-layer features , and the calculation result of the previous step are jointly input into "Transformer Encoder3" for multi-head attention calculation, which is divided into multi-head self-attention feature calculation and multi-head attention guiding feature calculation: (1) Multi-head self-attention feature calculation Normalize the input features , and then perform multi-head self-attention calculation to obtain multi-head self-attention information .

[0058] Add the multi-head self-attention information and the input features element-wise, and then perform normalization to obtain multi-head self-attention features .

[0059] (2) Multi-head attention guiding feature calculation Perform multi-head self-attention calculation on the multi-head self-attention features , the calculated guiding information and together to obtain multi-head attention guiding information .

[0060] Add the multi-head attention guiding information and the multi-head self-attention features element-wise, and then perform normalization to obtain multi-head attention guiding features .

[0061] Step S5: Fuse the output of "Transformer Encoder3" with the enhanced features of the 4th Block layer to obtain target recognition features; Step S501: Perform adaptive average pooling on the enhanced features to obtain the highest-layer adaptive average pooling features , and then use the flatten function in PyTorch to "flatten" to obtain convolutional layer features .

[0062] Step S502: Use the multi-head attention guiding features to extract class tokens, and denote the extracted class tokens as .

[0063] Step S503: Combine with the features Fusion is performed using the Concat function to obtain fused features .

[0064] Step S6: Input the target recognition feature into the recognition unit to determine the target recognition result of the few-shot remote sensing image target.

[0065] After obtaining the fused features successively determine the distance between the true probability and the predicted probability through a fully connected network, a Softmax function, and a cross-entropy loss function. The predicted probability corresponding to the smaller loss is the optimal recognition model for the few-shot remote sensing image target.

[0066] Compared with the prior art, this embodiment has at least the following advantages: 1) This application calculates attention guidance information using multi-scale convolutional features extracted from the first 3 Block layers of the ResNet-50 network, and uses Vision Transformer to learn the convolutional features extracted by the ResNet-50 network. Under the guidance of attention information at different layers, continuous deep learning is performed through the Vision Transformer encoder. In addition, the top-level semantic features extracted by the ResNet-50 network are fused. Finally, the obtained few-shot remote sensing image features are rich in content, having both the detail of the underlying features and the semantics of the high-level features. Therefore, the method of this application has excellent classification accuracy.

[0067] 2) This application selects the ResNet-50 convolutional network and combines the encoder part of Vision Transformer without a more complex neural network. The output features of the 3rd Block layer of ResNet-50 are processed as the input of Vision Transformer, and attention guidance information of different convolutional features is simply added in the encoder without a particularly large number of parameters. Therefore, the method of this application is simpler and more effective.

[0068] The second embodiment of the present invention corresponds to the first embodiment. This embodiment introduces a few-shot remote sensing image target recognition device, which includes the following components: A few-shot remote sensing image preprocessing unit, configured to obtain remote sensing image training data and perform preprocessing; A multi-scale feature enhancement unit, configured to use the configured ResNest-50 network to determine the convolutional features corresponding to the current training data; An embedded layer feature calculation unit, configured to perform "image patch splitting" and "trainable parameter embedding" operations using the enhanced features of the configured 3rd Block layer to determine the embedded layer features; The multi-head attention guidance information calculation unit is configured to determine the guidance information by using operations such as maximum pooling, average pooling, channel dimensionality reduction, image patch embedding feature calculation, and embedding category tokens; The multi-head attention feature calculation unit is configured to perform multi-head self-attention feature calculation and multi-head attention guidance feature calculation by using "Transformer Encoder" to determine the multi-head attention features; The target recognition feature calculation unit is configured to fuse the multi-head attention features with the enhanced features of the 4th Block layer to obtain the classification features of the few-shot remote sensing image targets; The few-shot remote sensing image recognition unit is configured to determine the optimal model for few-shot remote sensing image recognition based on the classification features through a fully connected network, Softmax function, and cross-entropy loss function.

[0069] It should be noted that in the embodiments of the present application, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element.

[0070] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0071] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0072] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims. These all belong to the protection scope of the present application.

Claims

1. A method for object recognition in remote sensing images with few samples, characterized in that: include: Step S1, obtaining a few sample remote sensing images, randomly selecting remote sensing images as training data, and preprocessing them; Step S2, using the configured ResNet-50 network, combined with the channel and spatial attention modules, determine the enhanced features corresponding to the convolutional features of the first Block layer, the second Block layer, the third Block layer, and the fourth Block layer; Step S3, using the enhanced features of the first three Block layers, respectively calculate the multi-head attention guidance information, and use the enhanced features of the third Block layer to determine the embedding layer features; Step S4, calculating the multi-head attention features by using the embedding layer features, the multi-head attention guidance information of the third Block layer, the second Block layer, and the first Block layer in sequence; Step S5, fusing the multi-head attention features with the enhanced features of the fourth Block layer to determine the target recognition features of the few-sample remote sensing image; Step S6: input the target recognition features into a recognition unit to determine the best recognition model for the target in the few-sample remote sensing image.

2. The method for object recognition in a remote sensing image using a small number of samples as claimed in claim 1, characterized in that: The step S1 comprises: Step S101, randomly cropping the training data, wherein the random cropping covers at least 80% of the training data image; Step S102, randomly rotating the training data, wherein the range of random rotation is between -45 degrees and 45 degrees; Step S103, horizontally flipping the current training data with a probability of 0.5; Step S104, cutting from the current training data center to both sides; Step S105, converting the current training data into tensor data of a preset shape; Step S106, normalizing the tensor data channel by channel, changing the mean to 0 and the standard deviation to 1, so as to obtain preprocessed training data.

3. The method for object recognition in a remote sensing image using a small number of samples as claimed in claim 2, characterized in that: The step S2 comprises: Step S201, build a ResNet-50 network and import the pre-trained parameters "resnet50-11ad3fa6.pth" obtained on the ImageNet dataset; Step S202, inputting preprocessed training data; Step S203, using the ResNet-50 network to extract convolutional features of the first Block layer, the second Block layer, the third Block layer, and the fourth Block layer; Step S204, using the channel and spatial attention modules to calculate the enhanced features of the first Block layer, the second Block layer, the third Block layer and the fourth Block layer corresponding to the four convolutional features.

4. The method for object recognition in a remote sensing image with a small number of samples as claimed in claim 3, characterized in that: The step S3 comprises: Step S301, calculating the embedding layer features of the third Block layer; Step S302, calculating the guidance information of the third Block layer; Step S303, calculating the guidance information of the second Block layer; Step S304, calculating the guidance information of the first Block layer.

5. The method for object recognition in a remote sensing image with a small number of samples as claimed in claim 4, characterized in that: The step S4 comprises: Step S401, calculating the multi-head attention features of "Transformer Encoder1"; Step S402, calculating the multi-head attention features of "Transformer Encoder2"; Step S403, calculate the multi-head attention features of "Transformer Encoder3".

6. A method for object recognition in a remote sensing image with a small number of samples as claimed in claim 5, characterized in that: The step S5 comprises: Step S501, the multi-head attention features of "Transformer Encoder3" are fused with the enhanced features of the 4th Block layer to obtain the target recognition features of the few-sample remote sensing image.

7. The method for object recognition in a remote sensing image with a small number of samples as claimed in claim 6, characterized in that: The step S6 comprises: Step S601, using a global average pooling function to average all pixel values ​​of each channel of the target recognition feature to obtain a target pooling feature; Step S602, flattening the target pooling feature to obtain a one-dimensional vector feature; Step S603: input the one-dimensional vector feature into a fully connected layer to output a fully connected feature. Step S604, inputting the features of the fully connected layer into the Softmax function to obtain the predicted probability; Step S605, using the cross entropy loss function to determine the distance loss between the true probability and the predicted probability; Step S606: determine the model corresponding to the minimum value of the distance loss as the optimal target recognition model for the few-sample remote sensing image.

8. A device for object recognition in remote sensing images with a small number of samples, characterized in that: include: A few-sample remote sensing image preprocessing unit is configured to obtain remote sensing image training data and preprocess it; A multi-scale feature enhancement unit is configured to determine a convolutional feature corresponding to the current training data using a configured ResNest-50 network; The embedding layer feature calculation unit is configured to use the configured third Block layer enhancement feature to perform "image block segmentation" and "trainable parameter embedding" operations to determine the embedding layer feature; A multi-head attention guidance information calculation unit is configured to determine guidance information using maximum pooling and average pooling, channel dimension reduction, image patch embedding feature calculation, and embedding category token operation; A multi-head attention feature calculation unit, configured to use the "Transformer Encoder" to perform multi-head self-attention feature calculation and multi-head attention guidance feature calculation to determine the multi-head attention feature; A target recognition feature calculation unit is configured to fuse the multi-head attention feature with the fourth Block layer enhancement feature to obtain a classification feature of a few-sample remote sensing image target; The few-sample remote sensing image recognition unit is configured to determine the best model for few-sample remote sensing image recognition based on the classification features through a fully connected network, a Softmax function and a cross entropy loss function.

Citation Information

Patent Citations

  • Skin image feature extraction method based on series fusion of DenseNet and Transformer

    CN115984578A

  • Defect classification method based on fine-grained recognition and attention mechanism

    CN116109629A