A Small-Sample Classification Method for Remote Sensing Images Based on Multi-View Feature Fusion

Through the multi-view feature fusion method, the rotation insensitive characteristics and self-supervised tasks of remote sensing images are used to improve the classification accuracy and generalization ability of the remote sensing scene recognition model, and solve the problem of insufficient generalization in small-sample remote sensing image recognition.

CN116543192BActive Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310070568.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-07
Publication Date
2025-07-18
Estimated Expiration
2043-02-07

AI Technical Summary

Technical Problem

The existing remote sensing scene recognition model lacks generalization capabilities in small samples, and does not fully utilize the rotation-insensitive characteristics of remote sensing images and the sharing of information from multiple perspectives.

Method used

The multi-view feature fusion method is adopted to process images through rotation enhancement, and combined with a full-connection layer rotation angle classifier, basic semantic classifier and multi-view feature fusion semantic classifier, self-supervised auxiliary tasks and class distribution consistency loss function are designed to improve the model's migratorable feature extraction and classification accuracy.

Benefits of technology

Effectively suppress semantic irrelevant content in remote sensing images, improve the classification accuracy and generalization capabilities of the model, simplify the model training process, and realize end-to-end efficient classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543192B_ABST
    Figure CN116543192B_ABST
Patent Text Reader

Abstract

The present invention provides a small-sample classification method for remote sensing images based on multi-view feature fusion. First, the input training set images are processed by rotation enhancement; then, features are extracted from all images under multiple views; next, the extracted features are input into a classification model network for training, and the model includes three parallel branches: a fully connected layer rotation angle classifier, a basic semantic classifier, and a multi-view feature fusion semantic classifier, and corresponding loss functions are designed respectively; finally, the trained network is used to classify and predict the remote sensing image data set to be processed. The present invention can solve the problem of insufficient generalization existing in the model during the training process of remote sensing scene recognition when the number of labeled samples is scarce, and has the beneficial effects of promoting the model to learn transferable knowledge, suppressing semantically irrelevant content in remote sensing images, and strengthening the correlation information of nearest neighbor prototype matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and image processing, and particularly relates to a small-sample classification method for remote sensing images based on multi-view feature fusion. Background Art

[0002] Remote sensing scene recognition has high application value in various fields, such as natural disaster prediction, land use monitoring, and autonomous perception of on-site security. In recent years, deep models driven by labeled data have greatly improved the performance of remote sensing scene classification algorithms with their powerful learning ability. However, with the development of various high-resolution sensors, the size of remote sensing scene images is increasing day by day and the types are diverse, which makes the annotation task of remote sensing scene images extremely difficult. Inspired by the rapid knowledge transfer ability of humans, few-shot classification aims to identify target samples using extremely small amounts of labeled data, which effectively alleviates the problems of data annotation and sample collection difficulties. Related research on few-shot learning can be roughly divided into the following three categories: few-shot learning based on data augmentation, few-shot learning based on meta-learning, and few-shot learning based on metric learning. Among them, algorithms that combine metric learning and meta-learning have achieved remarkable results.

[0003] The literature "H. Ji, Z. Gao, Y. Zhang, Y. Wan, C. Li, and T. Mei. Few-Shot Scene Classification of Optical Remote Sensing Images Leveraging Calibrated Pretext Tasks. IEEE Transactions on Geoscience and Remote Sensing, 60, 1-13, 2022." proposed a few-shot classification method for remote sensing scenes based on self-supervised auxiliary tasks. This method first introduced a rotation angle prediction task to improve the transferable feature extraction ability of the model; then used contrastive learning as an auxiliary task to aggregate features of the same category and make heterogeneous features far away from each other, improving the feature expression ability of the model; finally, to further alleviate the overfitting problem of the model and improve its generalization ability, this method used a regularization method based on AMP to correct the parameters of the model.

[0004] The literature "Q. Zeng and J. Geng. Task-Specific Contrastive Learning for Few-Shot Remote Sensing Image Scene Classification. ISPRS Journal of Photogrammetry and Remote Sensing, 191, 143-154, 2022." proposed a task-related improved contrastive learning method to enhance the performance of the model for few-shot remote sensing scene classification. First, the paper designed a feature enhancement module of "self-attention + mutual-attention" to filter out the background noise in remote sensing images and assist the model in capturing the potential association between test samples and class centers. In addition, the paper optimized traditional contrastive learning, expanded the screening range of positive and negative sample pairs by introducing semantic labels, and proposed a task-related contrastive learning method.

[0005] However, these methods all have limitations. They only use rotation prediction as an auxiliary task to enhance the feature extraction ability of the model and do not fully utilize the rotation-insensitive characteristics of remote sensing images. On the one hand, since remote sensing images do not have clear azimuth and pose information, the semantic prediction probability distribution of remote sensing images under different rotation perspectives should be consistent. On the other hand, since the content in different views of each remote sensing image contributes almost equally to the semantic label, the potential shared information between them can provide important value for the matching of test samples and class centers. Summary of the Invention

[0006] To overcome the deficiencies of the prior art, the present invention provides a few-shot remote sensing image classification method based on multi-view feature fusion. First, perform rotation enhancement processing on the input training set images. Then, extract features from all images under multiple views. Next, input the extracted features into the classification model network for training. The model includes three parallel branches: a fully connected layer rotation angle classifier, a basic semantic classifier, and a multi-view feature fusion semantic classifier, and corresponding loss functions are designed respectively. Finally, use the trained network to classify and predict the remote sensing image data set to be processed. The present invention can solve the problem of insufficient generalization of the model in the training process of remote sensing scene recognition when the number of labeled samples is scarce, and has the beneficial effects of promoting the model to learn transferable knowledge, suppressing semantically irrelevant content in remote sensing images, and strengthening the matching association information of the nearest neighbor prototype.

[0007] A few-shot remote sensing image classification method based on multi-view feature fusion, characterized by the following steps:

[0008] Step 1: Input the training image dataset and perform rotation augmentation on all images in the dataset. The rotation augmentation means rotating each image by 0°, 90°, 180°, and 270° respectively to obtain images with corresponding perspectives. For the i-th image in the dataset, denote the multi-perspective image set obtained after its rotation augmentation as where correspond to the images under four perspectives respectively, i = 1, 2, …,, represents the training image dataset, and represents the total number of images contained in the dataset;

[0009] Step 2: Use the ResNet-12 feature extraction network to extract features from all the multi-perspective images obtained in Step 1, and obtain the features corresponding to the images under each perspective. All features are one-dimensional vectors with a length of d = 640;

[0010] Step 3: Input the features of all the multi-perspective images into the classification model, and perform overall optimization training of the model in an end-to-end form to obtain a trained model. Among them, the classification model includes three parallel branches: a fully connected layer rotation angle classifier, a basic semantic classifier, and a multi-perspective feature fusion semantic classifier;

[0011] The fully connected layer rotation angle classifier adopts a single-layer fully connected + relu activation function structure, with an input dimension of 640 and an output dimension of 4. After passing through the fully connected layer rotation angle classifier, the features are mapped to the angle category space, and its corresponding rotation angle prediction loss function is as follows:

[0012]

[0013] where represents the rotation angle prediction loss, θ represents the network parameters of the feature extractor, represents the parameters of the fully connected layer rotation angle classifier, is the cross-entropy loss function calculated according to the following formula:

[0014]

[0015] where R = 4, represents the four perspectives of rotation, r represents the r-th perspective of rotation, f θ () represents the feature extraction operation, represents the fully connected layer rotation angle classification operation, [·] r represents taking the r-th element in the vector;

[0016] The basic semantic classifier adopts the nearest neighbor prototype representation principle, selects the category center closest to the features of the test image as the semantic category of the test image, and obtains the semantic probability distribution of the test image under each perspective through it. Its corresponding category distribution consistency loss function is as follows:

[0017]

[0018] Among them, represents the class distribution consistency loss, and P r represents the class probability distribution of all query set images after r×90° rotation augmentation. The query set Q refers to the set of all test images. N represents the number of classes in each mini-batch during training. represents the class probability distribution of the i-th test image in the query set after r×90° rotation augmentation, where i = 1, 2, …, |Q|, and |Q| represents the number of images contained in the set Q; P represents the average class probability distribution under all perspectives, calculated according to Calculated, D KL (·||·) represents calculating the KL divergence of vectors; The c-th element in The calculation expression of is as follows:

[0019]

[0020] Among them, τ is a scaling factor, and its value range is from 128 to 512. represents the class center of class c obtained by averaging the features of all support set images under r×90° rotation augmentation. The support set S represents the training images with known labels in each mini-batch during training. represents the Euclidean distance between the feature of the i-th test image and the class center c under r×90° rotation augmentation, and the value range of c is from 1 to N;

[0021] The calculation formula of the KL divergence is as follows:

[0022]

[0023] The multi-view feature fusion semantic classifier described above adopts a transformer structure. Through it, the fused test image features and class center features are obtained, and then based on the principle of nearest neighbor prototype representation, the class probability distribution of each test image is output. The specific process is as follows: Concatenate R feature vectors to obtain a multi-view feature map F i , for the support set images, denote the obtained multi-view feature map as F i S , for the query set images, denote the obtained multi-view feature map as F i Q , where i represents the image serial number in the set; then, for each test image in the query set, concatenate its corresponding multi-view feature map and all class centers row by row to obtain the corresponding augmented multi-view feature map

[0024]

[0025] Among them, is the multi-view category center feature map of category c, which is obtained by averaging the support set image features under all views;

[0026] Then, feature fusion is performed through the transformer structure, and the specific expression is as follows:

[0027]

[0028]

[0029] Among them, (Q, K, V) is the received triple input of the transformer structure, and W Q , W K and W V are three fully connected layers, is the fused feature;

[0030] For , it is evenly split by row to obtain two features, denoted as and Taking and , they are expanded in the second and third dimensions and their Euclidean distance D i is calculated according to the following formula, which is the distance between the fused test image feature map and the fused category center feature map:

[0031]

[0032] Among them, d(·) represents the Euclidean distance function, and row j represents taking the j-th row of the matrix;

[0033] Adopting the nearest neighbor prototype representation principle, the nearest category center is selected as the predicted category of this test image;

[0034] The loss function corresponding to the multi-view feature fusion semantic classifier is as follows:

[0035]

[0036] Among them, represents the main classification loss of multi-view feature fusion, y i represents the true semantic label of the i-th test image in the query set, and [D i c represents the Euclidean distance between the feature of the i-th test image in the query set and the category center of category c;

[0037] The total loss function of the classification model is as follows:​

[0038]

[0039] wherein, represents the total loss of the classification model network, and β is the weight hyperparameter of the rotation angle prediction loss term, and its value range is from 1 to 5, and γ is the weight hyperparameter of the class distribution consistency loss term, and its value range is from 10 to 50;

[0040] Step 4: Input the remote sensing image dataset to be processed into the classification model trained in Step 3, and the output of the multi-view feature fusion semantic classifier in the classification model is the final class prediction result of each image.

[0041] The beneficial effects of the present invention are as follows: Due to the adoption of the fully connected layer rotation angle classifier, the rotation-insensitive characteristics of remote sensing images are fully utilized, and a class distribution consistency loss function is designed, which can effectively suppress the features irrelevant to semantics in remote sensing images; Since both the rotation angle classification task and the class distribution consistency task belong to self-supervised auxiliary tasks, the model's transferable feature extraction ability is better improved; Due to the design of a new multi-view attention capture module and its embedding into the supervised few-shot classifier to form a multi-view feature fusion semantic classifier, the shared information in multi-view features and the strong correlation information between the query set samples and the class center in the nearest neighbor matching are extracted simultaneously, which can effectively eliminate redundant information and capture the strong correlation information between samples and the class center, improving the classification accuracy of the model; The classification model proposed by the present invention is a multi-task deep neural network, which can achieve end-to-end training without a redundant pre-training process, and the entire model framework is more concise and efficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a flowchart of the method for few-shot classification of remote sensing images based on multi-view feature fusion according to the present invention;

[0043] Figure 2 is a classification confusion matrix diagram on the NWPU-RESISC45 dataset using the method of the present invention;

[0044] wherein, (a) is a schematic diagram of the results of the 5-way 1-shot task, and (b) is a schematic diagram of the results of the 5-way 5-shot task;

[0045] Figure 3 is a classification confusion matrix diagram on the WHU-RS19 dataset using the method of the present invention;

[0046] wherein, (a) is a schematic diagram of the results of the 5-way 1-shot task, and (b) is a schematic diagram of the results of the 5-way 5-shot task. Detailed implementation manners

[0047] The present invention will be further described below in conjunction with the accompanying drawings and embodiments, and the present invention includes but is not limited to the following embodiments.

[0048] The present invention provides a small-sample classification method for remote sensing images based on multi-view feature fusion. The core is to construct a multi-task deep neural network classification model. By designing self-supervised auxiliary tasks, nearest neighbor prototype representations, supervised classifiers for multi-view feature fusion, etc., and adopting a multi-task training method, a concise and efficient classification model applicable to small-sample remote sensing images is obtained. As Figure 1 shown, the specific implementation process of the present invention is as follows:

[0049] 1. Generate multi-view remote sensing images

[0050] Input the training image dataset, and perform rotation augmentation processing on all images in the dataset. The rotation augmentation processing refers to rotating each image by (r - 1)×90°, where r = 1, 2, 3, 4, that is, rotating by 0°, 90°, 180°, and 270° respectively to obtain corresponding perspective images. For the i-th image in the dataset, denote the multi-view image set obtained after its rotation augmentation as where, correspond to the images under four perspectives respectively, i = 1, 2, …, |ε|, ε represents the training image dataset, and |ε| represents the total number of images contained in the dataset.

[0051] 2. Extract original features

[0052] Use the ResNet-12 feature extraction network to extract features from all the multi-view images obtained after the processing in step 1. After multi-level deep feature extraction, that is, after convolution and pooling operations of different layers, the feature output of the image under each perspective is a one-dimensional vector with a length of 640.

[0053] The ResNet-12 network is described in the literature "K. He, X. Zhang, S. Ren, et al. Deep Residual Learning for Image Recognition. In Proceeding of IEEE Conference on Computer Vision and Pattern Recognition, 770 - 778, 2016."

[0054] 3. Construct a classification model and model training

[0055] The features of all multi-view images are input into a classification model, and the overall model is optimized and trained in an end-to-end manner to obtain a trained model. Among them, the classification model includes three parallel branches: a fully connected layer rotation angle classifier, a basic semantic classifier, and a multi-view feature fusion semantic classifier.

[0056] (1) Fully connected layer rotation angle classifier and rotation angle prediction loss

[0057] Since using different rotation angles as task-irrelevant labels for supervised learning is beneficial to improving the model's transferable feature extraction ability, the present invention introduces a self-supervised task of rotation angle prediction.

[0058] Among them, the fully connected layer rotation angle classifier has a structure of "single-layer fully connected + relu activation function", with an input dimension of 640 and an output dimension of 4. After mapping the features to the angle category space through it, the corresponding rotation angle prediction loss function is as follows:

[0059]

[0060] Among them, represents the rotation angle prediction loss, θ represents the network parameters of the feature extractor, represents the parameters of the fully connected layer rotation angle classifier, is the cross-entropy loss function calculated according to the following formula:

[0061]

[0062] Among them, R = 4, representing four perspectives of rotation, r represents the r-th perspective of rotation, f θ (·) represents the feature extraction operation, represents the fully connected layer rotation angle classification operation, [·] r represents taking the r-th element in the vector.

[0063] The above rotation angle prediction task introduces different rotation perspectives of the original image as supervision signals, and designs a fully connected layer rotation angle classifier to map the original features output by the feature extractor to the angle category space.

[0064] (2) Basic semantic classifier and class distribution consistency loss

[0065] The present invention designs another loss function in the self-supervised auxiliary task to optimize the model so that the semantic prediction probability distributions of different perspectives of the same original picture are consistent, that is, the basic semantic classifier and the class distribution consistency loss.

[0066] The basic semantic classifier adopts the principle of nearest neighbor prototype representation, that is, the class center closest to the features of the test image is selected as the semantic class of the test image, and the semantic probability distribution of the test image under each perspective is obtained through it.

[0067] To keep the class probability distribution of the test images in the query set consistent under different rotation perspectives, the present invention minimizes the Kullback–Leibler (KL) divergence between the probability distribution of each perspective and the mean of all perspectives, and the corresponding class distribution consistency loss function is designed as follows:

[0068]

[0069] Among them, represents the class distribution consistency loss, and P r represents the class probability distribution of all query set images after being enhanced by r×90° rotation. The query set Q refers to the set of all test images. N represents the number of classes in each mini-batch. represents P r is a probability matrix of |Q| rows and N columns. represents the class probability distribution obtained after the i-th test image in the query set is enhanced by r×90° rotation, where i = 1, 2, …, |Q|, and |Q| represents the number of images contained in the set Q; P represents the average class probability distribution under all perspectives, calculated according to and D KL (·||·) represents calculating the KL divergence of vectors; The c-th element in is calculated as follows:

[0070]

[0071] Among them, τ is a scaling factor, and its value range is from 128 to 512. represents the class center of class c obtained by averaging the features of all support set images under r×90° rotation. The support set S represents the training images with known labels in each mini-batch. represents the Euclidean distance between the features of the i-th test image and the class center c under r×90° rotation, and the value range of c is from 1 to N.

[0072] The calculation formula of the KL divergence is as follows:

[0073]

[0074] (3) Multi-view feature fusion module and main classification loss

[0075] To fuse multi-view features, the present invention embeds a new multi-view attention module into a supervised classifier based on nearest neighbor prototype representation, namely a multi-view feature fusion semantic classifier. This classifier can extract the shared information of each test image under different rotation perspectives, effectively suppressing semantically irrelevant content in remote sensing images. In addition, this classifier can capture the strong correlation information between the test image and the category center in the nearest neighbor prototype matching, thereby improving the accuracy of the nearest neighbor prototype matching;

[0076] The multi-view feature fusion semantic classifier described above is of a transformer structure. After passing through it, the fused test image features and category center features are obtained, and then based on the nearest neighbor prototype representation principle, the category probability distribution of each test image is output. Specifically, for the multi-view set corresponding to the i-th test image in the query set After passing through the feature extractor, R one-dimensional feature vectors with a length of d = 640 can be obtained. By concatenating the R feature vectors, a multi-view feature map is obtained Similarly, a multi-view support set feature map and a multi-view query set feature map can be represented as F i S and F i Q , for each test image in the query set, its corresponding multi-view feature map and all category centers are concatenated by rows to obtain the corresponding augmented multi-view feature map:

[0077]

[0078] Among them, The multi-view category center feature map of category c is calculated by averaging the support set image features under all perspectives, and it can be expressed as K is the number of images in the support set. The transformer structure in this classifier receives the triple input (F, F, F) as (Query, Key, and Value) respectively. For the i-th test image in the query set, the corresponding augmented multi-view feature map can be calculated by Equation (17), and the process of transformer feature fusion can be expressed as follows:

[0079]

[0080]

[0081] Among them, W Q , W K and W V represent three fully connected layers respectively. By evenly splitting by rows, a fused test image feature map can be obtained and a fused class center feature map

[0082] Next, and are expanded along the second and third dimensions to where M = R×d. Furthermore, the distance vector between the fused test image feature map and the fused class center feature map can be expressed as:

[0083]

[0084] where, d(·) represents the Euclidean distance function, and row j represents taking the j-th row of the matrix;

[0085] Finally, using the nearest neighbor prototype representation principle, the class center with the closest distance is selected as the predicted class of this test image.

[0086] The loss function corresponding to the multi-view feature fusion semantic classifier is as follows:

[0087]

[0088] where, represents the main classification loss of multi-view feature fusion, y i represents the true semantic label of the i-th test image in the query set, [D i c represents the Euclidean distance between the feature of the i-th test image in the query set and the class center of class c.

[0089] (4) Total loss of the model

[0090] For the calculation of the total loss of the classification model, the present invention adopts a joint loss of weighted summation of the rotation angle prediction loss function, the class distribution consistency loss function, and the main classification loss function of multi-view feature fusion, that is:

[0091]

[0092] where, represents the total loss of the classification model network, β is the weight hyperparameter of the rotation angle prediction loss term, and its value range is from 1 to 5, γ is the weight hyperparameter of the class distribution consistency loss term, and its value range is from 10 to 50.

[0093] 4. Remote sensing image classification and prediction

[0094] ​Input the remote sensing image dataset to be processed into the classification model trained in step 3, where the final class prediction result of each image is obtained by the multi-view feature fusion semantic classifier in the classification model.

[0095] The effects of the present invention can be further illustrated by the following simulation experiments:

[0096] 1. Simulation conditions

[0097] The specific configuration of the experimental equipment in this embodiment is: central processing unit i7 - 6800K@3.40GHz, 64GB of memory, graphics processing unit GeForce GTX 1080Ti, operating system Ubuntu2016, and use the PyTorch deep learning framework for simulation.

[0098] There are two datasets used in the simulation: NWPU - RESISC45 dataset and WHU - RS19 dataset. The NWPU - RESISC45 dataset was proposed by Cheng et al. in the literature "G.Cheng, J.Han, and X.Lu. Remote Sensing Image Scene Classification: Benchmark and State of the Art. In Proceeding of the IEEE, 105(10), 1865 - 1883, 2017.", including 31,500 high - definition remote sensing images of 256×256, covering 45 categories of scenes, with 700 images in each scene. The WHU - RS19 dataset was proposed by Sheng et al. in the literature "G.Sheng, W.Yang, T.Xu, et al. High - Resolution Satellite Scene Classification Using a Sparse Coding Based Multiple Feature Combination. International Journal of Remote Sensing, 33(8), 2395 - 2412, 2012.", including 1,005 high - definition remote sensing images of 600×600, covering 19 categories of scenes, with at least 50 images in each scene.

[0099] In the experiment, the 45 categories of the NWPU - RESISC45 dataset are divided into 3 groups of 25 / 10 / 10 as the training set, validation set, and test set respectively; the 19 categories of the WHU - RS19 dataset are divided into 3 groups of 9 / 5 / 5 as the training set, validation set, and test set respectively.

[0100] 2. Simulation content

[0101] First, use the data in the training set to train the end-to-end multi-task deep learning framework proposed by the present invention, and obtain and store the trained model; then, use the data in the test set to test the trained model. Divide all the images in the test set into several mini-batches, and then divide each mini-batch into a support set (with known categories) and a query set (with unknown categories). In order to evaluate the performance of the model under different numbers of labeled images, design 5-way 1-shot tasks (the support set contains 5 categories, and there is 1 sample in each category) and 5-way 5-shot tasks (the support set contains 5 categories, and there are 5 samples in each category) respectively, and calculate the average classification accuracy of the 95% confidence interval as the final evaluation performance.

[0102] To prove the effectiveness of the method of the present invention, a total of 5 existing algorithms, namely the ProtoNet algorithm, the SPNet algorithm, the DANet algorithm, the SGMNet algorithm, and the TSC algorithm, were selected as comparison algorithms respectively.The ProtoNet algorithm was proposed in the literature "J. Snell, K. Swersky, and R. Zemel. Prototypical Networks for Few-Shot Learning. Advances in Neural Information Processing Systems, 2017."; the SPNet algorithm was proposed in the literature "G. Cheng, L. Cai, C. Lang, X. Yao, et al. SPNet: Siamese-Prototype Network for Few-Shot Remote Sensing Image Scene Classification. IEEE Transactions on Geoscience and Remote Sensing, 60, 1–11, 2022."; the DANet algorithm was proposed in the literature "M. Gong, J. Li, Y. Zhang, et al. Two-Path Aggregation Attention Network With Quad-Patch Data Augmentation for Few-Shot Scene Classification. IEEE Transactions on Geoscience and Remote Sensing, 60, 1–16, 2022."; the SGMNet algorithm was proposed in the literature "B. Zhang, S. Feng, X. Li, et al. SGMNet: SceneGraph Matching Network for Few-Shot Remote Sensing Scene Classification. IEEE Transactions on Geoscience and Remote Sensing, 60, 1–15, 2022."; the TSC algorithm was proposed in the literature "Q. Zeng and J. Geng. Task-Specific Contrastive Learning for Few-Shot Remote Sensing Image Scene Classification. ISPRS Journal of Photogrammetry and Remote Sensing 191, 143-154, 2022."

[0103] The calculation results of the average classification accuracy of different algorithms are shown in Table 1. It can be seen that on the two datasets, the average classification accuracy of the present invention in the two task settings of 5-way 1-shot and 5-way 5-shot is higher than that of other algorithms.

[0104] Table 1

[0105]

[0106] Figure 2 and Figure 3 respectively give the classification confusion matrices of the method of the present invention on the two datasets. Among them, Figure 2 (a) and Figure 2 (b) are the classification confusion matrices of the 5-way 1-shot task and the 5-way 5-shot task on the NWPU-RESISC45 dataset respectively. Figure 3 (a) and Figure 3 (b) are the classification confusion matrices of the 5-way 1-shot task and the 5-way 5-shot task on the WHU-RS19 dataset respectively. In the figure, both the horizontal and vertical coordinates are the class labels in the test set, including several classes such as Airport. The element in the i-th row and j-th column represents the probability that the model predicts an image belonging to class i as class j. From Figure 2 and Figure 3 it can be seen that the probability of correct classification for each class is relatively stable and close to the overall average classification accuracy of the dataset, indicating that the method of the present invention can obtain a high classification accuracy.

[0107] The present invention uses a few-shot classification framework based on multi-view feature fusion under nearest neighbor prototype representation to fully exploit the rotation-insensitive characteristics of remote sensing images. First, a fully convolutional network is used to extract rich deep features in the remote sensing images, and then the transferable feature extraction ability and generalization ability of the model are improved through two self-supervised auxiliary tasks and a multi-view feature fusion main classifier.

Claims

1. A small-sample classification method for remote sensing images based on multi-view feature fusion, characterized in that The steps are as follows: Step 1: Input the training image dataset and perform rotation augmentation on all images in the dataset. The rotation augmentation means rotating each image by 0°, 90°, 180°, and 270° respectively to obtain images with corresponding perspectives. For the i-th image in the dataset, denote the multi-perspective image set obtained after its rotation augmentation as where correspond to the images under four perspectives respectively, i = 1, 2, …, |ε|, ε represents the training image dataset, and |ε| represents the total number of images in the dataset; Step 2: Use the ResNet-12 feature extraction network to extract features from all the multi-view images obtained after the processing in Step 1, and obtain the features corresponding to the images in each view. All the features are one-dimensional vectors with a length of d = 640; Step 3: Input the features of all the multi-view images into the classification model, and adopt an end-to-end form to optimize and train the whole model to obtain a trained model. Among them, the classification model includes three parallel branches: a fully connected layer rotation angle classifier, a basic semantic classifier, and a multi-view feature fusion semantic classifier; The fully connected layer rotation angle classifier adopts a single-layer fully connected + relu activation function structure, with an input dimension of 640 and an output dimension of 4. After passing through the fully connected layer rotation angle classifier, the features are mapped to the angle category space, and its corresponding rotation angle prediction loss function is as follows: Among them, represents the rotation angle prediction loss, θ represents the network parameters of the feature extractor, represents the parameters of the fully connected layer rotation angle classifier, is the cross-entropy loss function calculated by the following formula: Among them, R = 4 represents four perspectives of rotation, r represents the r-th perspective of rotation, and f θ (·) represents the feature extraction operation, represents the fully connected layer rotation angle classification operation, [·] r represents taking the r-th element in the vector; The basic semantic classifier adopts the nearest neighbor prototype representation principle, selects the category center closest to the feature of the test image as the semantic category of the test image, and obtains the semantic probability distribution of the test image in each view through it. Its corresponding category distribution consistency loss function is as follows: Among them, represents the class distribution consistency loss, and P r represents the class probability distribution of all query set images after being augmented by r×90° rotation. The query set Q refers to the set of all test images. N represents the number of classes in each mini-batch during training. represents the class probability distribution of the i-th test image in the query set after being augmented by r×90° rotation, where i = 1, 2, …, |Q|, and |Q| represents the number of images contained in the set Q; P represents the average class probability distribution under all perspectives, calculated according to and D KL (·||·) represents calculating the KL divergence of vectors. the c-th element in has the following calculation expression: where τ is a scaling factor with a value range from 128 to 512, represents the class center of class c obtained by averaging all support set image features under r×90° rotation augmentation. The support set S represents the training images with known labels in each mini-batch during training, represents the Euclidean distance between the feature of the i-th test image and the class center c under r×90° rotation augmentation, where c ranges from 1 to N; The calculation formula of the KL divergence is as follows: The multi-view feature fusion semantic classifier described above adopts a Transformer structure. After obtaining the fused test image features and class center features through it, and then based on the principle of nearest neighbor prototype representation, the class probability distribution of each test image is output. The specific process is as follows: Concatenate R feature vectors to obtain a multi-view feature map F i , for the support set images, denote the obtained multi-view feature map as F i S , for the query set images, denote the obtained multi-view feature map as F i Q , where i represents the image serial number in the set; then, for each test image in the query set, concatenate its corresponding multi-view feature map and all class centers row by row to obtain the corresponding augmented multi-view feature map Among them, is the multi-view class center feature map of class c, which is obtained by averaging the support set image features under all views; Then, feature fusion is performed through the transformer structure, and the specific expression is as follows: Among them, (Q, K, V) is the received triple input of the transformer structure, and W Q , W K and W V are three fully connected layers, is the fused feature; Pairwise Split equally by row to obtain two features, denoted as and Expand and in the 2nd and 3rd dimensions and calculate their Euclidean distance D according to the following formula i , which is the distance between the fused test image feature map and the fused class center feature map: where d(·) represents the Euclidean distance function, and row j denotes taking the j-th row of the matrix; Adopt the nearest neighbor prototype representation principle, and select the closest category center as the predicted category of the test image; The loss function corresponding to the multi-view feature fusion semantic classifier is as follows: Among them, represents the multi-view feature fusion main classification loss, and y i represents the true semantic label of the i-th test image in the query set, [D i c represents the Euclidean distance between the feature of the i-th test image in the query set and the class center of class c;​ The total loss function of the classification model is as follows: Among them, represents the total loss of the classification model network, β is the weight hyperparameter of the rotation angle prediction loss term, with a value range of 1 to 5, and γ is the weight hyperparameter of the class distribution consistency loss term, with a value range of 10 to 50; Step 4: Input the remote sensing image dataset to be processed into the classification model trained in Step 3. Among them, the output of the multi-view feature fusion semantic classifier in the classification model is the final category prediction result of each image.

Citation Information

Patent Citations

  • Micro-video popularity prediction method based on attributive classification and multi-angle feature fusion

    CN107609570A

  • Image mask filter method, device, system, and storage medium

    WO2021042549A1