Semi-Supervised Remote Sensing Image Retrieval Method under the Condition of Limited Annotated Samples

By designing a semi-supervised multitasking dual-branch network, combined with supervised deep metric learning and self-supervised comparison learning, the problems of scarce labeled data and unused data information in the existing technology are solved, and high-performance remote sensing image retrieval under limited labeled samples are realized.

CN115937671BActive Publication Date: 2025-07-01XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211440678.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-07-01
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

The existing semi-supervised remote sensing image retrieval methods have deteriorated performance when labeled data is scarce, and have failed to fully utilize the structural and semantic information in unlabeled data.

Method used

A semi-supervised multitasking dual-branch network is designed. Through joint training of supervised branches and self-supervised branches, supervised deep metric learning loss function and improved self-supervised comparison learning loss function, fully mining semantic information in labeled data.

Benefits of technology

Under the condition of limited labeled sample, the performance of remote sensing image retrieval is improved, the distinction and generalization of feature representation are enhanced, and the overfitting problem caused by scarcity of labeled data is effectively solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937671B_ABST
    Figure CN115937671B_ABST
Patent Text Reader

Abstract

The present invention discloses a semi-supervised remote sensing image retrieval method under the condition of limited labeled samples, mainly solving the problems that it is difficult to fully exploit the semantic information hidden in unlabeled samples in the prior art and the poor generalization performance under the condition of limited labeled samples. The solution is as follows: constructing a semi-supervised multi-task dual-branch network composed of a supervised branch and a self-supervised branch; dividing a high-resolution remote sensing data set into labeled images and unlabeled images; extracting the feature vectors and class probability vectors of the labeled images, and extracting the feature vectors of the enhanced views of the unlabeled images; training the semi-supervised multi-task dual-branch network with the extracted feature vectors and class probability vectors; inputting the query image and the retrieval database to the trained semi-supervised multi-task dual-branch network to obtain the retrieved remote sensing images. The present invention enhances the generalization performance of the semi-supervised multi-task dual-branch network, improves the retrieval performance of semi-supervised remote sensing images, and can be used for the management of large-scale remote sensing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a semi-supervised remote sensing image retrieval method, which can be used for the management of large-scale remote sensing data. Background Art

[0002] With the rapid development of remote sensing technology, the number of remote sensing images is increasing at a rate of tens of thousands per day. Systematically managing these large amounts of data is crucial for different remote sensing applications. Remote sensing image retrieval, as an effective tool to solve this problem, has always attracted the attention of researchers.

[0003] Existing remote sensing image retrieval methods can be roughly divided into two categories, namely, handcrafted-based methods and deep learning-based methods. Handcrafted-based methods are designed to represent feature descriptors of visual attributes in remote sensing images. Commonly used handcrafted descriptors include global features and local features. In addition, local aggregation descriptors and bag-of-words of visual words, which are combined according to specific rules, are also commonly used. The implementation of handcrafted descriptors is very simple, but they cannot fully represent the complex content in remote sensing images, thus limiting their performance in remote sensing image retrieval. Currently, with the development of deep learning, especially convolutional neural networks, deep learning-based feature learning methods dominate the remote sensing field. With the help of the powerful non-linear fitting ability of convolutional neural networks, more and more convolutional neural network-based feature learning methods have been developed for remote sensing image retrieval.

[0004] Although the features obtained by convolutional neural network-based methods have greatly improved the performance of remote sensing image retrieval, training deep neural networks requires a large number of labeled samples. This is unacceptable and even impractical for remote sensing images because of their complex content. Manually annotating remote sensing images requires professional knowledge and high time costs. A natural solution to the above problems is to develop unique deep models in the unsupervised learning paradigm. However, due to the lack of prior knowledge, the performance of the features obtained by deep unsupervised models cannot meet expectations. Therefore, it is ideal to be able to use a training set with a small amount of labeled data to train the deep model. For this purpose, semi-supervised learning has entered the field of remote sensing image retrieval, aiming to train the model using both labeled and unlabeled data. Many semi-supervised feature learning networks have been proposed for remote sensing image retrieval and other remote sensing tasks. Although their performance is satisfactory, there is still room for improvement. First, for most semi-supervised models, the proportion of labeled data is small. Nevertheless, when the number of labeled data is extremely scarce, such as 5, 8, or 10 per class, their performance will be greatly weakened. Second, in the traditional semi-supervised learning framework, labeled samples are used to learn semantics, while unlabeled samples are used to obtain the internal structure information of the data. Therefore, the rich semantic knowledge hidden in the unlabeled samples is wasted, which limits the discriminability of the final feature representation.

[0005] Zhang et al. proposed a semi-supervised learning method based on center loss in the IEEE Journal of Applied Earth Observations and Remote Sensing (JSTAR), vol. 13, pp. 1362–1373, 2020. This method consists of a supervised branch and an unsupervised branch. The supervised branch uses a small number of labeled samples to generate the original class centers, and takes the original class centers as the initial centers for clustering in the unsupervised branch. Then, the unsupervised clustering algorithm generates corrected class centers based on the initial class centers and the features of the unlabeled data, and feeds the corrected class centers back to the supervised branch to update the original class centers. Finally, the updated class centers are used to calculate the center loss to guide the network training. However, the unsupervised clustering algorithm of this method will greatly reduce the unlabeled data participating in model training, resulting in a large waste of the structural information and hidden semantic information in the unlabeled data, thus limiting the performance of the semi-supervised model.

[0006] Kang et al. published a high-rank regularized semi-supervised deep metric learning method for remote sensing images in IEEE Remote Sensing vol.12, no.16, p.2603, 2020. This method proposes a normalized softmax loss with margins to learn a metric space with high intra-class compactness and inter-class diversity by using the supervision of a small amount of labeled data, while using high-rank regularization on unlabeled data to maintain the discrimination and diversity of unlabeled data. However, since high-rank regularization only maintains the discrimination of unlabeled data by increasing the rank of the class probability matrix of unlabeled data during the learning process, it does not fully explore the semantic information hidden in the unlabeled data, thus limiting the performance of the semi-supervised model.

[0007] Sohn et al. proposed a semi-supervised learning method based on consistency and confidence at the conference on progress in neural information processing systems, Neurips. vol. 33, pp. 596–608, 2020. The loss function of this method consists of two cross entropy functions. The cross entropy function of the supervised part uses a small amount of labeled data to learn the semantic information of the image. The unsupervised part first performs two types of data enhancement, weak data enhancement and strong data enhancement, on the unlabeled data. Then, the class probability vector generated by the weak data enhancement view is used to generate the corresponding pseudo label of the strong data enhancement view according to the threshold. Then, the cross entropy loss function of the strong data enhancement view is calculated based on the pseudo label, and the two cross entropy functions are used to jointly train the semi-supervised model. Since this method does not directly close the distance between images of different classes or distance the distance between images of different classes in the feature space learning process, it will lead to large differences in the feature space between some images of the same class and large similarities in the feature space between features of different classes, thus limiting the performance of the semi-supervised model. Summary of the invention

[0008] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a semi-supervised remote sensing image retrieval method under the condition of limited annotated samples, so as to make full use of unlabeled data to capture structural and semantic information, achieve good generalization under the condition of limited annotated samples, and improve the performance of the semi-supervised model.

[0009] To achieve the above object, the implementation steps of the present invention include the following:

[0010] (1) Construct a semi-supervised multi-task dual-branch network consisting of a supervised branch and a self-supervised branch. Both branches include a backbone network, a multi-scale attention module, and two fully connected layers. The weights of the backbone network and the multi-scale attention module of the two branches are shared.

[0011] (2) Input the images of the labeled samples and their corresponding labels into the supervised branch. Extract the feature vectors of each image through the first fully-connected layer of this branch, and generate the class probability vectors of each image through the second fully-connected layer;

[0012] (3) Perform two different random data augmentations on each image of the unlabeled samples to obtain two different augmented views, and then input all the generated augmented views into the self-supervised branch. Extract the feature vectors of each augmented view through the non-linear projection head composed of two fully-connected layers of the self-supervised branch;

[0013] (4) Train the semi-supervised multi-task dual-branch network:

[0014] 4a) Calculate the supervised deep metric learning loss function \(L_{s}\) using the feature vectors extracted by the first fully-connected layer of the supervised branch, calculate the cross-entropy loss function \(L_{c}\) using the class probability vectors generated by the second fully-connected layer, and add these two loss functions to form the total loss function \(L_{sup}\) of the supervised branch; SDML and use the class probability vectors generated by the second fully-connected layer to calculate the cross-entropy loss function \(L_{c}\), CE and add these two loss functions to form the total loss function \(L_{sup}\) of the supervised branch; sup ;

[0015] 4b) Calculate the improved self-supervised contrastive learning loss function \(L_{ssl}\) using the feature vectors extracted by the self-supervised branch, and use this function as the loss function of the self-supervised branch; ICSL and use this function as the loss function of the self-supervised branch;

[0016] 4c) Set the hyperparameter \(\lambda\), and obtain the loss function \(L\) of the semi-supervised multi-task dual-branch network by weighted summing the loss functions of the supervised branch and the self-supervised branch through the parameter \(\lambda\);

[0017] 4d) Use the stochastic gradient descent algorithm to iteratively solve the loss function \(L\) of the semi-supervised multi-task dual-branch network, and update the network parameters by backpropagation at the same time until the loss function converges to obtain the trained semi-supervised multi-task dual-branch network;

[0018] (5) Use the trained semi-supervised multi-task dual-branch network for remote sensing image retrieval:

[0019] 5a) Use the entire high-resolution remote sensing dataset as the database to be retrieved, and use the unlabeled samples in the high-resolution remote sensing dataset as the test set;

[0020] 5b) Input all the images in the test set and the database to be retrieved into the trained semi-supervised multi-task dual-branch network, and extract the feature vectors of each image in the test set and the database to be retrieved through the first fully-connected layer of the supervised branch;

[0021] 5c) Calculate the Euclidean distances between the feature vectors of each image in the test set and the feature vectors of all images in the database to be retrieved. Take the images corresponding to the top 20 feature vectors with the smallest Euclidean distance for each image as the retrieval results for each image.

[0022] Compared with the prior art, the present invention has the following advantages:

[0023] First, through the designed semi-supervised multi-task dual-branch network, the present invention fully exploits the multi-scale information of remote sensing images and highlights the information of key regions from complex image content, solves the problems of multi-scale and complex image content of remote sensing images, enhances the representation ability of remote sensing image features, and improves the performance of the semi-supervised model.

[0024] Second, through the supervised deep metric learning loss function set in the present invention, the problems of intra-class diversity and inter-class similarity of remote sensing images are solved, the distinguishability of remote sensing image features is enhanced, and the performance of the semi-supervised multi-task dual-branch network is further improved.

[0025] Third, through the improved self-supervised contrast learning loss function set in the present invention, the hidden semantic information in the unlabeled data is fully exploited, the overfitting problem caused by limited labeled samples is effectively solved, and the generalization of the semi-supervised multi-task dual-branch network is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is the implementation process framework diagram of the present invention;

[0027] Figure 2 is the semi-supervised multi-task dual-branch network model diagram of the present invention;

[0028] Figure 3 is the multi-scale attention module framework diagram of the present invention;

[0029] Figure 4 is the simulation result diagram of remote sensing image retrieval of the present invention and the prior art on different data sets. DETAILED DESCRIPTION OF THE INVENTION

[0030] The following further details the examples and effects of the present invention with reference to the accompanying drawings.

[0031] Refer to Figure 1 , under the condition of limited labeled samples in this example, the semi-supervised remote sensing image retrieval method sequentially includes constructing a semi-supervised multi-task dual-branch network, dataset division, extracting the feature vectors and class probability vectors of labeled images, extracting the feature vectors of enhanced views of unlabeled images, training the semi-supervised multi-task dual-branch network, and performing remote sensing image retrieval. The specific implementation is as follows:

[0032] Step 1, construct a semi-supervised multi-task dual-branch network.

[0033] Reference Figure 2 , the semi-supervised multi-task dual-branch network constructed in this step consists of a supervised branch and a self-supervised branch. Both branches include a backbone network, a multi-scale attention module, and two fully connected layers, and the weights of the backbone network and the multi-scale attention module of the two branches are shared. Its construction is as follows:

[0034] 1.1) Establish a backbone network composed of a cascaded convolution block, a max pooling layer, and four residual convolution blocks in sequence, where:

[0035] The convolution block consists of a cascaded 7×7 standard convolution layer, a batch normalization layer, and a ReLU activation layer;

[0036] The max pooling layer is a max pooling layer with a kernel size of 3×3, a stride of 2, and a padding of 1;

[0037] Each of the four residual convolution blocks is composed of two cascaded residual units;

[0038] 1.2) Establish a multi-scale attention module composed of a cascaded multi-scale sub-module and attention sub-module:

[0039] Reference Figure 3 , the multi-scale sub-module is composed of a cascade of four standard convolution layers with kernel sizes of 1×1, 3×3, 5×5, and 7×7 in parallel and a 1×1 standard convolution layer; the attention sub-module is composed of a cascade of a 1×1 standard convolution layer, a batch normalization layer, a Sigmid activation function layer, and a global average pooling layer;

[0040] 1.3) Establish two sequentially connected fully connected layers for the supervised branch and the self-supervised branch, where:

[0041] The two fully connected layers of the supervised branch are composed of a fully connected layer with a dimension of 128 and a fully connected layer with a dimension equal to the number of dataset categories in cascade;

[0042] The two fully connected layers of the self-supervised branch are composed of a fully connected layer with a dimension of 256 and a fully connected layer with a dimension of 128 in cascade.

[0043] Step 2, dataset division.

[0044] Randomly select 5, 8, or 10 images from the images of each category in the dataset as the labeled images of the training set, and use the remaining images in the dataset as the unlabeled images of the training set.

[0045] Step 3, extract the feature vectors and class probability vectors of the labeled images.

[0046] 3.1) Iteratively sample the labeled images in the training set, that is, in each sampling process, first randomly select N categories from the C categories of the labeled images in the training set, and then select M images from each of the N categories; where C is the total number of categories of the labeled images in the training set, and N < C;

[0047] 3.2) Input the images sampled in 3.1) into the supervised branch of the semi-supervised multi-task dual-branch network, extract the feature vector of each image through the first fully connected layer of the supervised branch, and generate the class probability vector of each image through the second fully connected layer.

[0048] Step four, extract the feature vectors of the augmented views of the unlabeled images.

[0049] 4.1) Sample the unlabeled images:

[0050] Set B u as the batch size for sampling the unlabeled images. In each iteration, randomly select B u images from the unlabeled images in the training set;

[0051] 4.2) Perform data augmentation on the unlabeled images to obtain augmented views:

[0052] Perform data augmentation on the unlabeled images in sequence with two operations consisting of random horizontal flipping, random color jittering, random Gaussian blurring, and random grayscaling, so that each unlabeled image obtains two different augmented views;

[0053] 4.3) Input the augmented views obtained in 4.2) into the self-supervised branch of the semi-supervised multi-task dual-branch network, and extract the feature vector of each augmented view through the second fully connected layer of the self-supervised branch.

[0054] Step five, train the semi-supervised multi-task dual-branch network.

[0055] 5.1) Calculate the supervised deep metric learning loss function L SDML :

[0056] 5.1.1) According to the feature vectors extracted by the first fully connected layer of the supervised branch and the corresponding class labels, calculate the class center of each class:

[0057]

[0058] where p i represents the class center of the i-th class, f ij represents the j-th feature vector of the i-th class, and M represents that each category has M images;

[0059] 5.1.2) Calculate the deep metric learning loss function L center :

[0060]

[0061] Among them, N represents the number of randomly sampled categories, N c represents the set composed of N randomly sampled categories, N c -i represents the set composed of the other N - 1 categories in the set N c except the i-th category, D(x, y) represents the Euclidean distance between two feature vectors, Max(x, 0) represents the operation of finding the larger value between x and 0, m1 is a margin hyperparameter, and m1 > 0;

[0062] 5.1.3) Perform hard sampling based on the feature vectors and corresponding class labels extracted from the first fully connected layer of the supervised branch. That is, for the i-th sample, select a sample with the closest distance as the negative sample from the samples of the N - 1 categories other than the category to which the i-th sample belongs, and select a sample with the farthest distance as the positive sample from the samples of the same category;

[0063] 5.1.4) Calculate the deep metric learning loss function L sample :

[0064]

[0065] where m2 is a margin hyperparameter and m2 > 0, i C represents the category of the i-th sample, N C -i C represents the set composed of the other N - 1 categories in the set N c except category i C outside, f i represents the i-th vector, f k represents the k-th vector;

[0066] 5.1.5) Calculate the supervised deep metric learning loss function L according to the results of 5.1.2) and 5.1.4) SDML :

[0067] L SDML = L center + L sample ;

[0068] 5.2) Calculate the cross-entropy loss function L CE :

[0069] Calculate the cross-entropy loss function L according to the class probability vector and corresponding label of the labeled image extracted from the semi-supervised multi-task dual-branch network CE :

[0070]

[0071] where is the class probability vector of the \(i\)-th image; \(y\) i is the label of the \(i\)-th image, in the form of a one-hot vector;

[0072] 5.3) Calculate the total loss function \(L\) of the supervised branch according to the results of 5.1) and 5.2) sup :

[0073] \(L\) sup =\(L\) SDML +\(L\) CE ;

[0074] 5.4) Calculate the improved self-supervised contrastive learning loss function \(L\) ICSL :

[0075] 5.4.1) Let the set of feature vectors of the views of the unlabeled images extracted by the self-supervised branch be: where \(B\) u is the number of unlabeled images, \(z\) j and represent the feature vectors of two different augmented views of the \(j\)-th image after random data augmentation, i.e., the first feature vector is \(z\) j , and the second feature vector is

[0076] 5.4.2) Take the second feature vector as the positive sample of the first feature vector \(z\) j ;

[0077] 5.4.3) Form a non-positive sample set \(T\) from the feature vectors in the feature vector set \(Z\) other than \(z\) j and . Select the top \(k\) feature vectors \(T\) j closest to the first feature vector \(z\) k from this set \(T\):

[0078] \(T\) k =\(\min\) k \{d(z j ,z m )|z m \in T\}

[0079] where \(\min\) k is a function for finding the top \(k\) minimum values in a sequence, \(z\) m is a feature vector in the non-positive sample set \(T\), and \(d(z j ,z m ) represents the cosine distance between \(z j and \(z m , and \(k\) is a hyperparameter;

[0080] 5.4.4) Use the feature vectors in the unpositive sample set \(T\) except \(T\) k as the first feature vector \(z\) j for the negative sample set \(T - T\) k , and calculate the improved self-supervised contrastive learning loss function \(L\) based on this negative sample set and the positive samples constructed in 5.4.2) ICSL :

[0081]

[0082] where \(\tau\) is the temperature parameter, and \(z\) k is the feature vector in \(T - T\) k ;

[0083] 5.5) Calculate the loss function \(L\) of the semi-supervised multi-task dual-branch network according to the results obtained in 5.3) and 5.4):

[0084] \(L = L\) sup +\(\lambda L\) ICSL

[0085] where \(\lambda\) is the hyperparameter used to weight and sum the two parts of the loss function.

[0086] 5.6) Iteratively solve the loss function \(L\) of the semi-supervised multi-task dual-branch network:

[0087] 5.6.1) Set \(\omega\) as the parameter of the semi-supervised multi-task dual-branch network, and calculate the derivative of the loss function \(L\) with respect to \(\omega\)

[0088] 5.6.2) Set \(\alpha\) as the learning rate, and update the parameter \(\omega\) according to the derivative to:

[0089]

[0090] where \(\omega'\) is the parameter after the network update;

[0091] 5.6.3) Repeat 5.6.1) and 5.6.2) until the loss function \(L\) converges to obtain the trained semi-supervised multi-task dual-branch network.

[0092] Step Six, perform remote sensing image retrieval.

[0093] 6.1) Use the entire high-resolution remote sensing data set as the database to be retrieved, and use the unlabeled images in the high-resolution remote sensing data set as the test set;

[0094] 6.2) Input all the images in the test set and the database to be retrieved into the trained semi-supervised multi-task dual-branch network, and extract the feature vectors of each image in the test set and the database to be retrieved through the first fully-connected layer of the supervised branch;

[0095] 6.3) Calculate the Euclidean distances between the feature vectors of each image in the test set and the feature vectors of all images in the database to be retrieved:

[0096] 6.3.1) Set n as the dimension of the feature vectors output by the semi-supervised multi-task dual-branch network. Let q i be the feature vector of the i-th image in the test set, and p j be the feature vector of the j-th image in the database to be retrieved, where:

[0097] q i = (q i1 , q i2 ,......, q in ), p j = (p j1 , p j2 ,......, p jn )

[0098] 6.3.2) Calculate the Euclidean distance d i between q j and p ij :

[0099]

[0100] 6.3.3) Calculate the Euclidean distances between the feature vectors of each image in the test set and the feature vectors of all images in the database to be retrieved according to the formula for calculating the Euclidean distance between two feature vectors in 6.3.2);

[0101] 6.4) Take the images corresponding to the top 20 feature vectors with the smallest Euclidean distance for each image as the retrieval results of each image, and complete the semi-supervised remote sensing image retrieval under the condition of limited labeled samples.

[0102] The effects of the present invention can be further illustrated by the following simulations:

[0103] I. Simulation conditions

[0104] The simulation is completed using the python3.7.9 + pytorch1.7.1 framework on a workstation with a GPU with 12G of memory on TITAN Xp and UbuntuOS 16.04;

[0105] The three datasets used in the simulation are the UCM dataset, the AID dataset, and the NWPU dataset. Among them,

[0106] The UCM dataset was released by the University of California, Merced, and contains a total of 2,100 images in 21 scene categories. Each category contains 100 images with a size of 256×256 pixels, and the spatial resolution of each pixel is 0.3m;

[0107] The AID dataset was released by Wuhan University and contains a total of 10,000 images in 30 scene categories, with the number of images in each category ranging from 220 to 420. The size of each image is 600×600 pixels, and the spatial resolution range of each pixel is from 0.5m to 8m;

[0108] The NWPU dataset was released by Northwestern Polytechnical University and contains a total of 31,500 images in 45 scene categories. Each category consists of 700 images with a size of 256×256 pixels, and the spatial resolution of each pixel ranges from 0.2m to 30m.

[0109] II. Simulation Content

[0110] Simulation 1, under the above simulation conditions, the proposed method of the present invention and five existing methods, SSCL, HR-S 2 DML, FixMatch, MixMatch, and ReMixMatch were used to conduct remote sensing image retrieval simulations on three datasets. The visualization results are shown in Figure 4, where:

[0111] Figure 4(a) is the precision-recall curve of the present invention and the existing five methods on the UCM dataset;

[0112] Figure 4(b) is the precision-recall curve of the present invention and the existing five methods on the AID dataset;

[0113] Figure 4(c) is the precision-recall curve of the present invention and the existing five methods on the NWPU dataset;

[0114] As can be seen from Figure 4, the precision-recall curve of the present invention covers a larger area than the existing five methods on all three datasets, indicating that the retrieval performance of the present invention is better, fully demonstrating the superiority of the method proposed in the present invention.

[0115] To quantitatively illustrate the performance of the proposed method, the average retrieval precision, a commonly used numerical performance index for remote sensing image retrieval, was selected to measure the differences between the above existing methods and the present invention. The results are shown in the following table, where:

[0116] Table 1 shows the numerical results of the present invention and the existing five methods on the UCM dataset;

[0117] Table 2 shows the numerical results of the present invention and the existing five methods on the AID dataset;

[0118] Table 3 shows the numerical results of the present invention and five existing methods on the NWPU dataset.

[0119] Table 1 UCM Dataset Results

[0120]

[0121] Table 2 AID Dataset Results

[0122]

[0123] Table 3 NWPU Dataset Results

[0124]

[0125] As can be seen from Table 1 to Table 3, the average retrieval accuracy of the present invention is higher, further demonstrating the superiority of the method proposed by the present invention.

[0126] The sources of the five existing technologies are as follows:

[0127] SSCL is a method for semi-supervised remote sensing image feature learning published by Zhang et al. in IEEE, that is: J. Zhang, M. Zhang, B. Pan, and Z. Shi, “Semisupervised center loss for remote sensing image scene classification,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 13, pp. 1362–1373, 2020;

[0128] HR-S 2 DML is a method for semi-supervised deep metric learning for remote sensing images published by Kang et al. in IEEE, that is: J. Kang, R. Fernandez-Beltran, Z. Ye, X. Tong, P. Ghamisi, and A. Plaza, “High-rankness regularized semi-supervised deep metric learning for remote sensing imagery,” Remote Sensing, vol. 12, no. 16, p. 2603, 2020;

[0129] FixMatch is a method for semi-supervised feature learning of images published by Sohn et al. at the NeurIPS (Advances in Neural Information Processing Systems) conference, namely: K. Sohn, D. Berthelot, N. Carlini, Z. Zhang, H. Zhang, C. A. Raffel, E. D. Cubuk, A. Kurakin, and C.-L. Li, "Fixmatch: Simplifying semi-supervised learning with consistency and confidence," Advances in Neural Information Processing Systems, vol. 33, pp. 596–608, 2020;

[0130] MixMatch is a method for semi-supervised feature learning of images published by Berthelot et al. at the NeurIPS (Advances in Neural Information Processing Systems) conference, namely: D. Berthelot, N. Carlini, I. Goodfellow, N. Papernot, A. Oliver, and C. A. Raffel, "Mixmatch: A holistic approach to semi-supervised learning," Advances in Neural Information Processing Systems, vol. 32, 2019;

[0131] ReMixMatch is a method for semi-supervised feature learning of images published by Berthelot et al. at the ICLR (International Conference on Learning Representations) conference, namely: D. Berthelot, N. Carlini, E. D. Cubuk, A. Kurakin, K. Sohn, H. Zhang, and C. Raffel, "Remixmatch: Semi-supervised learning with distribution matching and augmentation anchoring," in International Conference on Learning Representations, 2019.

Claims

1. A semi-supervised remote sensing image retrieval method under the condition of limited labeled samples, characterized in that It includes the following steps: (1) Construct a semi-supervised multi-task dual-branch network composed of a supervised branch and a self-supervised branch. Both branches include a backbone network, a multi-scale attention module, and two fully connected layers, and the weights of the backbone network and the multi-scale attention module of the two branches are shared; (2) Input the images of the labeled samples and their corresponding labels into the supervised branch. Extract the feature vectors of each image through the first fully connected layer of this branch, and generate the class probability vectors of each image through the second fully connected layer; (3) Perform two different random data augmentations on each image of the unlabeled samples to obtain two different augmented views, and then input all the generated augmented views into the self-supervised branch. Extract the feature vectors of each augmented view through the non-linear projection head composed of the two fully connected layers of the self-supervised branch; (4) Train the semi-supervised multi-task dual-branch network: 4a) Calculate the supervised deep metric learning loss function \(L\) using the feature vectors extracted by the first fully-connected layer of the supervised branch SDML , and calculate the cross-entropy loss function \(L\) using the class probability vectors generated by the second fully-connected layer CE , and add these two loss functions to form the total loss function \(L\) of the supervised branch sup ; 4b) Calculate the improved self-supervised contrastive learning loss function L using the feature vectors extracted by the self-supervised branch ICSL , and use this function as the loss function of the self-supervised branch; 4c) Set the hyperparameter λ, and obtain the loss function L of the semi-supervised multi-task dual-branch network by weighted summation of the loss functions of the supervised branch and the self-supervised branch through the parameter λ; 4d) Use the stochastic gradient descent algorithm to iteratively solve the loss function L of the semi-supervised multi-task dual-branch network, and at the same time backpropagate to update the network parameters until the loss function converges to obtain the trained semi-supervised multi-task dual-branch network; (5) Use the trained semi-supervised multi-task dual-branch network for remote sensing image retrieval: 5a) Use the entire high-resolution remote sensing dataset as the database to be retrieved, and use the unlabeled samples in the high-resolution remote sensing dataset as the test set; 5b) Input all the images in the test set and the database to be retrieved into the trained semi-supervised multi-task dual-branch network, and extract the feature vectors of each image in the test set and the database to be retrieved through the first fully connected layer of the supervised branch; 5c) Calculate the Euclidean distances between the feature vectors of each image in the test set and all the feature vectors of the images in the database to be retrieved, and use the images corresponding to the top 20 feature vectors with the smallest Euclidean distance for each image as the retrieval results of each image.

2. The method according to claim 1, characterized in that, The backbone network in step (1) includes a convolutional block, a max pooling layer, and four residual convolutional blocks connected in sequence: This convolutional block is used to directly downsample the input image to retain as much original image information as possible; This max pooling layer is used to further reduce the dimension of the information extracted by the convolutional layer to reduce the computational amount, and strengthen the invariance of the image features to enhance the robustness of the image in terms of offset, rotation, etc.; These four residual convolutional blocks are used to further extract the features of the output features of the previous layer and perform downsampling to obtain the output features of the backbone network.

3. The method according to claim 1, characterized in that The multi-scale attention module in step (1) includes a multi-scale sub-module and an attention sub-module connected in sequence: This multi-scale sub-module is used to extract the features of multiple scales of the image and fuse them to obtain the multi-scale fusion features of the image; This attention sub-module is used to extract the attention weights at each position in the multi-scale fusion features, and multiply the attention weights element-wise with the multi-scale fusion features to obtain the output features of the multi-scale attention module.

4. The method according to claim 1, characterized in that, The two fully connected layers included in both the supervised branch and the self-supervised branch in step (1) have different dimensions and are connected in sequence: The first fully connected layer of the supervised branch is used to extract features of the labeled image, and the second fully connected layer is used to extract the class probability vector of the labeled image; The two fully connected layers of the self-supervised branch are used to extract features of the unlabeled image.

5. The method according to claim 1, characterized in that, Calculate the supervised deep metric learning loss function \(L\) in step 4a) SDML , which is implemented as follows: 4a1) Calculate the class center of each class according to the feature vector extracted by the first fully connected layer of the supervised branch and the corresponding class label: where p i represents the class center of the i-th class, and f ij represents the j-th feature vector of the i-th class, and M represents that each class has M images; 4a2) Calculate the depth metric learning loss function L based on class centers center :[[]]END]] where N represents the number of randomly sampled categories, N c represents the set composed of N randomly sampled categories, N c -i represents the set N c excluding the i-th category, the set composed of the other N - 1 categories, D(x, y) represents the Euclidean distance between two feature vectors, Max(x, 0) is used to find the larger value between x and 0, and m1 is a margin hyperparameter and m1 > 0; 4a3) Perform hard sampling according to the feature vector extracted by the first fully connected layer of the supervised branch and the corresponding class label, that is, for the i-th sample, select a sample with the closest distance as the negative sample from the samples of N - 1 classes other than the class to which the i-th sample belongs, and at the same time select a sample with the farthest distance as the positive sample from the samples of the same class; 4a4) Calculate the deep metric learning loss function L based on hard sampling sample : where m2 is a margin hyperparameter and m2 > 0, i C represents the class of the i-th sample, N C -i C set N c except for class i C in the set consisting of the other N - 1 classes, f i represents the i-th vector, f k represents the k-th vector, 4a5) Calculate the supervised deep metric learning loss function L based on the results of 4a2) and 4a3) SDML :[[]]END]] L SDML = L center + L sample 。 6. The method according to claim 1, characterized in that, Calculate the cross-entropy loss function L in step 4a) CE , and the formula is as follows: where is the class probability vector of the i-th image, and y i is the ground truth scene label in the form of a one-hot vector.

7. The method according to claim 1, characterized in that The calculation of the improved self-supervised contrastive learning loss function L in step 4b) ICSL is implemented as follows: 4b1) Set the set represents the feature vector of the view of the unlabeled image extracted by the self-supervised branch, where B u is the number of unlabeled images, z j and represent the feature vectors of two different augmented views of the j-th image after random data augmentation, that is, the first feature vector is z j , and the second feature vector is 4b2) Use the second eigenvector as the positive sample of the first eigenvector z j ; 4b3) Set T to represent the set of (2B j - 2) eigenvectors generated from views of different images, and select the top k eigenvectors from set T that are closest to the first eigenvector z u : j distance T k = min k {d(z j , z m ) | z m ∈ T} where T k represents a set consisting of the top k eigenvectors in set T that are closest to the first eigenvector z j , min k is a function for finding the top k minimum values in a sequence, z m is an eigenvector in set T, d(z j , z m ) represents the cosine distance between z j and z m , and k is a hyperparameter; 4b4) Use the eigenvectors in T other than T k as the first set of eigenvectors z j of the negative eigenvectors, and calculate the improved self-supervised contrastive learning loss function L based on this set of negative samples and the positive samples constructed in 4b2 ICSL : where τ is the temperature parameter, T - T k denotes the set composed of the eigenvectors in T except for T k and z k is the eigenvector in T - T k .

8. The method according to claim 1, characterized in that The loss function of the semi-supervised multi-task double-branch network in step 4c) is implemented as follows: 4c1) Calculate the loss function L of the semi-supervised multi-task double-branch network according to the results obtained in 4a7) and 4b4): L = L sup + λL ICSL where λ is a hyperparameter used to weight and sum the two parts of the loss function.

9. The method according to claim 1, wherein The iterative solution of the loss function L of the semi-supervised multi-task double-branch network in step 4d) is implemented as follows: 4d1) Set ω as the parameter of the semi-supervised multi-task dual-branch network, and calculate the derivative of the loss function L with respect to ω 4d2) Set α as the learning rate and update the parameter ω according to the derivative as follows: ω′ is the parameter of the network after update; 4d3) Repeat 4d1) and 4d2) until the loss function L converges.

10. The method according to claim 1, characterized in that Calculate the Euclidean distance between the feature vector of each image in the test set and the feature vectors of all images in the database to be retrieved in step 5c) as follows: 5c1) Set \(n\) as the dimension of the feature vector output by the semi-supervised multi-task dual-branch network, and let \(q\) i be the feature vector of the \(i\)-th image in the test set, and \(p\) j be the feature vector of the \(j\)-th image in the database to be retrieved, where: q i =(q i1 ,q i2 ,......,q in ), p j =(p j1 , p j2 ,......, p jn ) 5c2) Calculate q i and p j to obtain the Euclidean distance d ij : 5c3) Calculate the Euclidean distance between the feature vector of each image in the test set and the feature vectors of all images in the database to be retrieved according to the formula for calculating the Euclidean distance between two feature vectors in 5c2).

Citation Information

Patent Citations

  • Remote sensing image road segmentation method based on convolutional neural network weak supervised learning

    CN112070779A

  • Hyperspectral image classification method based on double-branch spectrum multi-scale attention network

    CN113486851A