A cross-domain multimodal remote sensing image classification method based on contrastive learning

Through comparative learning and cross-domain training strategies, the problem of domain offset in remote sensing image classification is solved, unsupervised classification is achieved in the target domain, classification accuracy is improved, and network convergence is accelerated, which is suitable for the field of remote sensing image classification.

CN116912595BActive Publication Date: 2025-08-12XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310959584.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-01
Publication Date
2025-08-12
Estimated Expiration
2043-08-01

AI Technical Summary

Technical Problem

The prior art has problems in remote sensing image classification that the target domain classification performance is degraded due to domain offset and non-aligned distribution, especially in target domains where labels are difficult to obtain. Existing domain adaptation methods such as based on maximum average differences, adversarial training and reconstruction methods have complexity or sensitivity problems.

Method used

A cross-domain multimodal remote sensing image classification method based on contrast learning is adopted, and the source domain data is normalized and data preprocessed through the pre-training stage. Combined with the two-step training strategy in the cross-domain comparison learning stage, contrast learning is used to reduce the differences between domains, and features are extracted using convolutional layers and comparative losses are calculated to achieve unsupervised classification of the target domain.

Benefits of technology

It effectively narrows the domain differences between the source domain and the target domain, improves the classification accuracy of remote sensing images in the target domain, realizes unsupervised target domain classification, reduces hardware requirements and speeds up the network convergence speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116912595B_ABST
    Figure CN116912595B_ABST
Patent Text Reader

Abstract

A cross-domain multimodal remote sensing image classification method based on contrastive learning includes: a pre-training phase, in which source domain data is pre-processed; source domain data features are extracted to obtain source domain fusion features; the source domain fusion features are input into a classifier to obtain classification results; a cross-domain contrastive learning phase, in which source and target domain data are re-input and pre-processed; the source and target domain networks are initialized using the optimal parameters of the source domain network obtained in the pre-training phase, and feature queues for each category in the target domain are initialized; source and target domain data features are extracted to obtain source domain and target domain fusion features; classification results and high-dimensional features are obtained, and domain adaptive alignment is performed using contrastive learning; back-propagation is used to update the source domain network, momentum is used to update the target domain network, and the parameters of the optimal target domain network and its classification results are saved. The present invention can achieve unsupervised classification in the target domain, reduce the inter-domain differences between the source and target domains, and significantly improve the classification accuracy of remote sensing images in the target domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image classification, and in particular relates to a cross-domain multimodal remote sensing image classification method based on contrastive learning. Background Art

[0002] In the field of remote sensing, multi-source remote sensing data based on hyperspectral imagery has been systematically applied to land use / land cover classification, target detection, and environmental change monitoring. Specifically, hyperspectral imagery provides detailed spectral information, while remote sensing data from other sources can provide complementary information, such as LiDAR data, which provides meaningful elevation and spatial information. Multimodal data, with its unique advantages over single-modality data, improves classification accuracy.

[0003] In recent years, deep learning has achieved tremendous success in many image processing applications. Deep networks, such as convolutional layers, possess a strong ability to extract high-level features for pattern recognition. Inspired by this, numerous works have applied deep neural networks to remote sensing image classification, demonstrating superior performance when large amounts of labeled data are available. These methods are known as supervised learning methods. To integrate spatial and spectral information for more accurate classification, two-dimensional and three-dimensional convolutional layers have also been applied to extract deep features from remote sensing images.

[0004] However, labeling large-scale training data to obtain supervised data is laborious and expensive. A potential solution is to migrate the model trained on the labeled source domain to the desired unlabeled target domain. However, direct model migration often leads to a decline in the classification performance of the target domain due to phenomena such as domain offset or non-aligned distribution.

[0005] Domain adaptation is an effective measure to solve the above problems. There are two main types of domain adaptation methods: traditional domain adaptation methods and domain adaptation methods based on deep learning.

[0006] Traditional domain adaptation methods mainly include: feature-based adaptation, which adjusts source domain samples and target domain samples to the same feature space using a mapping f, so that the samples of the two can be aligned in this feature space; instance-based adaptation, considering that there are always some samples in the source domain that are very similar to target domain samples, multiplies the loss of all source domain samples by a weight during training. The more similar the sample is to the target domain, the larger the weight; model parameter-based adaptation, by finding a new model parameter θ', through parameter migration, the model can work better in the target domain.

[0007] Domain adaptation methods based on deep learning: The maximum mean difference-based method reduces the target domain generalization error by reducing the difference between the two domains. Common methods include transfer component analysis, which maps the source domain and the target domain into a reproducing kernel Hilbert space and uses the maximum mean difference to measure the difference in the data distribution of the two mapped domains. The maximum mean difference is used to construct a regularization term during feature learning to constrain the learned representation, so that the features on the two domains are as similar as possible, thereby reducing the distribution deviation; the adversarial method generates features through the generator, and then lets the discriminator determine whether it is a feature of the source domain or the target domain. If it cannot be determined, it means that the source domain and the target domain are consistent in this feature space; the reconstruction-based method, such as DRCN, encodes the source domain and target domain samples through the encoder, and then uses a classifier to classify the source domain features and a decoder to decode the target domain features, so that the target domain samples can be restored as much as possible. In this way, the feature space where the generated features are located is similar in the source domain and target domain samples.

[0008] Among deep learning-based domain adaptation methods, designing an inter-domain distance representation based on the maximum mean difference method is very complex; when the feature extraction network after adversarial training is too strong, the adversarial method will lead to a mismatch in the distribution of different categories in the feature space between the source domain and the target domain, making the classifier poor in classifying target domain samples; the reconstruction-based method relies on a powerful feature extractor, and this method is sensitive to noise and outliers, which may cause the model to deviate when processing target domain data. Summary of the Invention

[0009] In order to overcome the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a cross-domain multimodal remote sensing image classification method based on contrastive learning, which can realize unsupervised remote sensing image classification in the target domain, effectively solving the problem of difficulty in obtaining remote sensing image labels; the present invention performs mean-variance normalization on both source domain and target domain data, effectively reducing the inter-domain differences; the present invention proposes a two-step training strategy of source domain pre-training followed by source domain-target domain network contrastive learning cross-domain training, effectively accelerating the network convergence speed; the present invention proposes integrating contrastive learning into cross-domain training, effectively reducing the inter-domain differences between the source domain and the target domain, and effectively improving the classification accuracy of remote sensing images in the target domain.

[0010] In order to achieve the above object, the technical solution adopted by the present invention is:

[0011] A cross-domain multimodal remote sensing image classification method based on contrastive learning includes the following steps:

[0012] S101: In the pre-training stage, the source domain hyperspectral and source domain lidar image data to be classified are first input, and the source domain data are pre-processed;

[0013] S102: Use the convolution layer to extract the features of the pre-processed source domain data, flatten the features and then concatenate them to obtain the source domain fusion features;

[0014] S103: Input the source domain fusion features into the classifier to obtain the classification results, repeat S102 to S103, train the network multiple times, and save the optimal parameters of the source domain network;

[0015] S104: In the cross-domain comparative learning stage, the source domain hyperspectral and source domain lidar image data are first re-input, and the target domain hyperspectral and target domain lidar image data are input, and the data are preprocessed; the optimal parameters of the source domain network are used as the initial parameters of the source domain network and the target domain network in this stage, and the feature queues corresponding to each category in the target domain are initialized;

[0016] S105: Use the convolution layer to extract the features of the source domain and target domain data preprocessed in S104, flatten and concatenate the source domain and target domain features respectively, and obtain the source domain fusion features and the target domain fusion features;

[0017] S106: Input the source domain fusion features and the target domain fusion features into the classifier and mapper to obtain the source domain and target domain classification results and high-dimensional features, update the feature queue of S104 according to the classification results of the target domain, and perform comparative learning at the same time;

[0018] S107: Back propagation updates the source domain network, momentum updates the target domain network, repeats S104 to S107, and saves the parameters of the optimal target domain network and its classification results. As a further technical solution of the present invention, the step S101 is specifically as follows:

[0019] In the pre-training stage, the input source domain hyperspectral and source domain lidar image data are first normalized to obtain the source domain hyperspectral image H S and the source domain lidar image L S , respectively, are obtained through the following two formulas:

[0020]

[0021] where H' S is the unstandardized source domain hyperspectral image, is the average value of the source domain hyperspectral image, is the standard deviation of the source domain hyperspectral image;

[0022]

[0023] L' S is the unnormalized source domain lidar image, is the average value of the source domain lidar image, is the standard deviation of the source domain lidar image;

[0024] The source domain hyperspectral image H S and the source domain lidar image L S Perform edge filling, taking each pixel before filling as the center, and construct one-to-one corresponding source domain hyperspectral image blocks and source domain lidar image blocks;

[0025] Next, 200 pairs of labeled source domain hyperspectral image patches and source domain lidar image patches are randomly selected from each category according to the category of the center pixel point as the training set, and the rest are used as the test set.

[0026] As a further technical solution of the present invention, step S102 is specifically as follows:

[0027] In the source domain network, the hyperspectral image processing branch and the lidar image processing branch each apply two convolutional layers for feature extraction;

[0028] Assuming source domain hyperspectral image patches The size is C×11×11, where C is the number of channels of the hyperspectral image block. Two convolution layers are constructed with a convolution kernel size of 64×3×3, a step size of 2, and a padding of 1, and a convolution kernel size of 32×3×3, a step size of 2, and a padding of 1. The output of each layer is input into the activation function ReLU, and finally the source domain hyperspectral image features are obtained. Its dimensions are 32×3×3;

[0029] Assuming source domain lidar image patches The size is 1×11×11, and two convolution layers are constructed with a convolution kernel size of 16×3×3, a step size of 2, and a padding of 1, and a convolution kernel size of 32×3×3, a step size of 2, and a padding of 1. The output of each layer is input into the activation function ReLU, and finally the source domain lidar image features are obtained. Its dimensions are 32×3×3;

[0030] The convolution of the source domain hyperspectral image features and source domain lidar image features Under the condition that the channel dimension remains unchanged, the features are flattened to obtain the source domain hyperspectral image features with a size of 32×9. and source domain lidar image features of size 32×9

[0031] The obtained source domain hyperspectral image features and source domain lidar image features Perform splicing and fusion in the channel dimension to obtain the source domain fusion features Its size is 64×9;

[0032] As a further technical solution of the present invention, step S103 is specifically as follows:

[0033] A network consisting of a linear layer, a batch normalization layer, and a ReLu layer is selected as the classifier. The last layer of the classifier is a linear layer, and the number of its output channels is the number of ground object categories. Assume that the output is Y S , the true label is Y S , the cross entropy loss is as follows:

[0034]

[0035] Where M is the number of categories; y ic Is a sign function (0 or 1), if the true category of sample i is equal to c, it takes 1, otherwise it takes 0; P S Y S The probability vector of the predicted sample obtained by the Softmax function; is the predicted probability that the observed sample i belongs to category c;

[0036] The network learning is supervised by the cross-entropy loss function, and the source domain network parameters are updated using back propagation and stochastic gradient descent methods, and the source domain network parameters with the best performance on the source domain test set are saved.

[0037] As a further technical solution of the present invention, the specific steps of step S104 are as follows:

[0038] In the cross-domain comparative learning stage, the preprocessing of the source domain data is consistent with the pre-training stage described above. The input target domain hyperspectral image and target domain lidar image are normalized by mean variance to obtain the target domain hyperspectral image H T and the target domain lidar image L T , obtained by the following two formulas:

[0039]

[0040] where H' T is the target domain hyperspectral image that has not been normalized, is the average value of the target domain hyperspectral image, is the standard deviation of the target domain hyperspectral image;

[0041]

[0042] L' T is the target domain lidar image that has not been normalized, is the average value of the target domain lidar image, is the standard deviation of the target domain lidar image;

[0043] Then, the target domain hyperspectral image H T and the target domain lidar image L T Perform edge filling, centering on each pixel before filling, to construct a one-to-one correspondence between the target domain hyperspectral image patch and the target domain lidar image patch. Since the target domain data does not have true labels, there is no need to divide it into training and test sets; all samples can be directly used for training.

[0044] Loading the optimal parameters of the source domain network saved in S103 as the initial parameters of the source domain network and the target domain network in subsequent training;

[0045] The present invention sets a queue with a capacity of K for each category of target domain data to store corresponding features and provide data for the subsequent calculation of contrastive learning loss values. The initialized target domain network is used to test all samples in the target domain, and the index value of the maximum value of the probability vector obtained by the classifier output after the Softmax function is used as a pseudo label. If the confidence of the pseudo label is greater than the set threshold, the corresponding feature output by the mapper will be included in the queue of the corresponding category according to the pseudo label. After all tests are completed, the initial feature queue of each category is obtained.

[0046] As a further technical solution of the present invention, step S105 is specifically as follows:

[0047] In the cross-domain contrast learning stage, the structures of the source domain and target domain networks are exactly the same, so the convolution kernels used to extract features are the same, and their parameters are consistent with the source domain convolution kernel parameters in the pre-training stage, so the extracted target domain hyperspectral image features can be obtained. and target domain lidar image features The dimensions are 32×3×3 and 32×3×3;

[0048] The target domain flattening and splicing operation is consistent with the source domain, and both are spliced in the channel dimension, and the source domain fusion features at this stage can be obtained. Fusion features with the target domain The size is 64×9.

[0049] As a further technical solution of the present invention, step S106 is specifically as follows:

[0050] In this stage, the target domain network uses the same network composed of linear layer, batch normalization layer and ReLu layer as the source domain network as the classifier. The last layer of the classifier is a linear layer, and the number of its output channels is the number of ground object categories. However, since the target domain data has no real labels, there is no need to calculate the target domain cross entropy loss. Only the source domain cross entropy loss needs to be calculated. The calculation method of the source domain cross entropy loss in this stage is consistent with that in the pre-training stage.

[0051] In this stage, both the source and target domain networks use a network consisting of a linear layer, a batch normalization layer, and a ReLu layer as a mapper. The mapper maps features to a high-dimensional space. The feature queue stores the corresponding high-dimensional features output by the mapper based on the classification results of the target domain network and dequeues some earlier features.

[0052] The present invention proposes that during contrastive learning cross-domain training, based on the true label of the current sample in the source domain, a feature queue of the corresponding category in the target domain is selected, and all features in the queue are averaged to obtain a high-dimensional feature. The corresponding high-dimensional features output by the source domain sample through the mapper are regarded as positive samples of the high-dimensional feature, and the high-dimensional features in the feature queues corresponding to other categories are regarded as negative samples. The contrastive loss is calculated using the InFoNCE loss function, as shown in the following formula:

[0053]

[0054] Where q is the high-dimensional feature after the mean is obtained for the corresponding queue, k + is the high-dimensional feature of the source domain sample, N is the number of all samples (including positive and negative samples), and τ is the temperature hyperparameter;

[0055] As a further technical solution of the present invention, step S107 is specifically as follows:

[0056] At this stage, the source domain cross entropy loss function L CE and contrast loss L CL Guide the source domain network learning and update the source domain network parameters using back propagation and stochastic gradient descent methods;

[0057] The target domain network gradient will not be back-propagated, and the target domain network parameters are updated through the following momentum update method:

[0058] θ=m·θ+(1-m)·ξ

[0059] The target domain network parameter is θ, the source domain network parameter is ξ, and m is a hyperparameter;

[0060] Train multiple times until the network converges, and save the parameters and classification results of the target domain network that has the best test effect on the target domain samples.

[0061] Beneficial effects of the present invention:

[0062] 1. The present invention uses image blocks for training, which reduces hardware requirements while ensuring the integrity of spatial information.

[0063] 2. The present invention uses mean-variance normalization to process the source domain and target domain data, so that the source domain and target domain data generally approximately obey the 0-1 normal distribution, thereby reducing the difference between domains.

[0064] 3. The present invention proposes a two-step training strategy, namely, pre-training the source domain first, and then loading the optimal parameters of the source domain network obtained from the source domain pre-training into the source domain and target domain networks as their initial parameters in the second stage of cross-domain comparative learning training, which effectively speeds up the network training speed.

[0065] 4. The present invention proposes to integrate contrastive learning into cross-domain training, so that the encoded representation can capture the information shared between the same type of two domains, while effectively reducing the inter-domain differences between the source domain and the target domain, and effectively improving the classification accuracy of remote sensing images in the target domain.

[0066] 5. The present invention can realize unsupervised remote sensing image classification in the target domain, effectively solving the problem of difficulty in obtaining remote sensing image labels. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is a flow chart of a cross-domain multimodal remote sensing image classification method based on contrastive learning provided by an embodiment of the present invention.

[0068] Figure 2 Schematic diagram of the source domain network structure in the pre-training phase provided by an embodiment of the present invention.

[0069] Figure 3 Schematic diagram of the cross-domain contrastive learning training network structure provided by an embodiment of the present invention.

[0070] Figure 4 is the target domain remote sensing image classification result diagram and the real label diagram of the method proposed in the present invention provided by the embodiment of the present invention, wherein Figure 4 (a) is the real label map, Figure 4 (b) is the target domain remote sensing image classification result diagram obtained by the method proposed in this invention. DETAILED DESCRIPTION

[0071] The present invention will be described in further detail below with reference to the accompanying drawings.

[0072] like Figure 1 As shown, the cross-domain multimodal remote sensing image classification method based on contrastive learning provided by the present invention includes the following steps

[0073] S101: In the pre-training stage, the source domain hyperspectral and source domain lidar image data to be classified are first input, and the source domain data are pre-processed;

[0074] S102: Use the convolution layer to extract the features of the pre-processed source domain data, flatten the features and then concatenate them to obtain the source domain fusion features;

[0075] S103: Input the source domain fusion features into the classifier to obtain the classification results, repeat S102 to S103, train the network multiple times, and save the optimal parameters of the source domain network;

[0076] S104: In the cross-domain comparative learning phase, the source domain hyperspectral and source domain lidar image data are first re-input, and the target domain hyperspectral and target domain lidar image data are input to preprocess the data. The optimal parameters of the source domain network obtained in the pre-training phase in S103 are used as the initial parameters of the source domain network and the target domain network in this phase, and the feature queues corresponding to each category in the target domain are initialized.

[0077] S105: Use the convolution layer to extract the features of the source domain and target domain data preprocessed in S104, flatten and concatenate the source domain and target domain features respectively, and obtain the source domain fusion features and the target domain fusion features.

[0078] S106: Input the source domain fusion features and the target domain fusion features into the classifier and mapper to obtain classification results and high-dimensional features, update the feature queue according to the classification results of the target domain, and perform comparative learning at the same time;

[0079] S107: Back propagation updates the source domain network, momentum updates the target domain network, repeats S104 to S107, and saves the parameters of the optimal target domain network and its classification results.

[0080] like Figure 1 As shown, the cross-domain multimodal remote sensing image classification method based on contrastive learning provided by the present invention has the following implementation process:

[0081] (1) Figure 2 The figure shows the source domain network model in the pre-training phase. In the pre-training phase, the source domain hyperspectral and source domain lidar image data to be classified are first input and the source domain data is pre-processed.

[0082] In order to make the features that may have large distribution differences have the same weight influence on the model, the input source domain hyperspectral image and source domain lidar image data are normalized by mean variance so that the features meet the normal distribution with mean 0 and standard deviation 1. S and the source domain lidar image L S , respectively, are obtained by the following two formulas:

[0083]

[0084] where H' S is the unstandardized source domain hyperspectral image, is the average value of the source domain hyperspectral image, is the standard deviation of the source domain hyperspectral image;

[0085]

[0086] L' Sis the unnormalized source domain lidar image, is the average value of the source domain lidar image, is the standard deviation of the source domain lidar image;

[0087] Hyperspectral images usually contain a lot of information. Using the entire image directly for training requires high hardware requirements. However, using a single pixel for training ignores the spatial correlation between pixels. S and the source domain lidar image L S Edge filling is performed, and each pixel before filling is used as the center to construct a one-to-one corresponding source domain hyperspectral image block and source domain lidar image block. The image block size is C×11×11, where C is the number of channels of the source domain hyperspectral image block or source domain lidar image block. This reduces the hardware requirements while ensuring that the information is basically complete.

[0088] Next, 200 pairs of labeled source domain hyperspectral image patches and source domain lidar image patches are randomly selected from each category according to the category of the center pixel point as the training set, and the rest are used as the test set.

[0089] (2) Use the convolutional layer to extract the features of the preprocessed source domain data, flatten the features and splice them to obtain the source domain fusion features.

[0090] In the source domain network, the hyperspectral image processing branch and the lidar image processing branch apply two convolutional layers to extract spatial, spectral information and spatial, elevation information;

[0091] Source domain hyperspectral image patches The size is C×11×11, and two convolution layers are constructed with a convolution kernel size of 64×3×3, a stride of 2, and a padding of 1, and a convolution kernel size of 32×3×3, a stride of 2, and a padding of 1. The activation function ReLU is input after the output of each layer to improve the nonlinear relationship between the layers of the neural network and enhance the expression ability of the network. Finally, the source domain hyperspectral image features are obtained. Its dimensions are 32×3×3;

[0092] Source domain lidar image patches The size is 1×11×11, and two convolution layers are constructed with a convolution kernel size of 16×3×3, a stride of 2, and a padding of 1, and a convolution kernel size of 32×3×3, a stride of 2, and a padding of 1. The output of each layer is input into the activation function ReLU, and finally the source domain lidar image features are obtained. Its dimensions are 32×3×3;

[0093] The convolution of the source domain hyperspectral image features and source domain lidar image features Under the condition that the channel dimension remains unchanged, the features are flattened to obtain the source domain hyperspectral image features with a size of 32×9 and source domain lidar image features of size 32×9

[0094] Hyperspectral images usually contain rich spectral information, but their spatial information is relatively scarce. LiDAR images contain rich spatial-elevation information. However, using only hyperspectral features or LiDAR features for classification tasks is not ideal. Therefore, the source domain hyperspectral image features are obtained. and source domain lidar image features Perform splicing and fusion in the channel dimension to obtain the source domain fusion features Its size is 64×9;

[0095] (3) Input the source domain fusion features into the classifier to obtain the classification results, repeat S102 to S103, train the network multiple times, and save the optimal parameters of the source domain network.

[0096] A network consisting of a linear layer, a batch normalization layer, and a ReLu layer is selected as the classifier. The last layer of the classifier is a linear layer, and the number of its output channels is the number of object categories. The source domain fusion features are input into the classifier to obtain the classification result. Assume that the output is Y S , the true label is Y S , the cross entropy loss is as follows:

[0097]

[0098] Where M is the number of categories; y ic Is a sign function (0 or 1), if the true category of sample i is equal to c, it takes 1, otherwise it takes 0; P S Y S The probability vector of the predicted sample obtained by the Softmax function, is the predicted probability that the observed sample i belongs to category c;

[0099] The source domain network learning is supervised by the cross-entropy loss function, and the source domain network parameters are updated using back propagation and stochastic gradient descent methods, and the source domain network parameters with the best performance on the source domain test set are saved.

[0100] (4) Figure 3The figure shows the cross-domain comparative learning phase. In this phase, the source domain hyperspectral and source domain lidar image data are first re-input, and the target domain hyperspectral and target domain lidar image data are input and preprocessed. The optimal parameters of the source domain network obtained in the pre-training phase in S103 are used as the initial parameters of the source domain network and the target domain network in this phase, and the feature queues corresponding to each category in the target domain are initialized.

[0101] (4a) In the cross-domain contrastive learning stage, the preprocessing of the source domain data is consistent with the pre-training stage described above. The input target domain hyperspectral image and target domain lidar image are normalized by mean variance to obtain the target domain hyperspectral image H T and the target domain lidar image L T , obtained by the following two formulas:

[0102]

[0103] where H' T is the target domain hyperspectral image that has not been normalized, is the average value of the target domain hyperspectral image, is the standard deviation of the target domain hyperspectral image;

[0104]

[0105] L' T is the target domain lidar image that has not been normalized, is the average value of the target domain lidar image, is the standard deviation of the target domain lidar image;

[0106] After the present invention performs mean-variance normalization on both source domain and target domain data, the source domain and target domain data generally approximately obey a 0-1 normal distribution, thereby reducing the difference between domains.

[0107] Then, the target domain hyperspectral image H T and the target domain lidar image L T Perform edge filling, centering on each pixel before filling, to construct a one-to-one correspondence between the target domain hyperspectral image patch and the target domain lidar image patch. Since the target domain data does not have true labels, there is no need to divide it into training and test sets; all samples can be directly used for training.

[0108] (4b) Loading the optimal parameters of the source domain network in the pre-training phase as the initial parameters of the source domain network and the target domain network in subsequent training;

[0109] (4c) The present invention sets a queue with a capacity of K for each category of target domain data to store the corresponding features and provide data for the subsequent calculation of the comparative learning loss value; uses the initialized target domain network to test all samples in the target domain, and uses the index value of the maximum value of the probability vector obtained by the classifier output after the Softmax function as a pseudo-label. If the maximum value of the probability vector, that is, the confidence of the pseudo-label is greater than the set threshold, the corresponding feature output by the mapper will be included in the queue of the corresponding category according to the pseudo-label. After completing all tests, the initial feature queue of each category is obtained. During the subsequent training process, the feature queue will be dynamically updated, the latest features will be included, and some features that were included earlier will be dequeued to ensure that the difference between the features that were added early and those that were added later in the queue is not large.

[0110] (5) Use the convolutional layer to extract the features of the source domain and target domain data, flatten and concatenate the source domain and target domain features respectively, and obtain the source domain fusion features and target domain fusion features.

[0111] The structures of the source domain and target domain networks in the cross-domain contrast learning phase are exactly the same, so the convolution kernels used to extract features are the same, and their parameters are consistent with the source domain convolution kernel parameters in the pre-training phase, so the extracted target domain hyperspectral image features can be obtained. and target domain lidar image features The dimensions are 32×3×3 and 32×3×3;

[0112] The flattening and splicing operations of the target domain hyperspectral image features and the target domain lidar image features are consistent with those of the source domain. Both are spliced in the channel dimension, and the source domain fusion features at this stage can be obtained. Fusion features with the target domain The size is 64×9;

[0113] (6) The source domain fusion features and the target domain fusion features are input into the classifier and mapper to obtain the classification results and high-dimensional features. The feature queue is updated according to the classification results of the target domain, and comparative learning is performed at the same time.

[0114] In this stage, the target domain network uses the same network composed of linear layer, batch normalization layer and ReLu layer as the source domain network as the classifier. The last layer of the classifier is a linear layer, and the number of its output channels is the number of ground object categories. However, since the target domain data has no real label, there is no need to calculate the target domain cross entropy loss. Only the source domain cross entropy loss needs to be calculated. The calculation method of the source domain cross entropy loss is consistent with that in the pre-training stage.

[0115] In contrastive learning, each sample is typically mapped to a projection space, where positive samples are brought closer together and negative samples are pushed further apart, forcing the representation model to ignore surface factors and learn the inherently consistent structural information of the samples. The present invention employs a linear layer, a batch normalization layer, and a ReLu layer as mappers in the source and target domain networks. The mappers map features to a high-dimensional space. The target domain's feature queue stores the corresponding high-dimensional features based on the target domain network's classification results and dequeues some earlier features.

[0116] The present invention proposes that during contrastive learning cross-domain training, based on the true label of the current sample in the source domain, a feature queue of the corresponding category in the target domain is selected, and all features in the queue are averaged to obtain a high-dimensional feature anchor. The corresponding high-dimensional features output by the source domain sample through the mapper are regarded as positive samples of the high-dimensional features, and the high-dimensional features in the feature queues corresponding to other categories are regarded as negative samples. The contrast loss is calculated using the InFoNCE loss function, as shown in the following formula:

[0117]

[0118] Where q is the high-dimensional feature after the mean is obtained for the corresponding queue, k + is the high-dimensional feature of the source domain sample output by the mapper, N is the number of all samples (including positive and negative samples), and τ is the temperature hyperparameter;

[0119] Through this contrast loss function, the present invention shortens the distance between the anchor and the positive sample and expands the distance between the anchor and the negative sample in the high-dimensional mapping space, so that the encoded representation can capture the information shared between the same type in the two domains.

[0120] (7) Use backpropagation to update the source domain network, momentum to update the target domain network, repeatedly train the network, and save the optimal target domain network parameters and its classification results.

[0121] At this stage, the source domain cross entropy loss function L CE and contrast loss L CL Guide the source domain network learning and update the source domain network parameters using back propagation and stochastic gradient descent methods;

[0122] The target domain network gradient will not be back-propagated, and the target domain network parameters are updated through the following momentum update method:

[0123] θ=m·θ+(1-m)·ξ

[0124] The target domain network parameter is θ, the source domain network parameter is ξ, and m is a hyperparameter;

[0125] Momentum updates can ensure the consistency of the source and target domain networks, allowing the target domain network to evolve more smoothly. This prevents drastic changes in the target domain network parameters from causing large differences in features in the queue and destroying the consistency of representation.

[0126] Repeated training until the network converges, save the parameters of the target domain network with the best test effect on the target domain sample and its classification results. The following is a detailed description of the technical effects of the present invention in conjunction with simulation experiments:

[0127] (1) Simulation experiment conditions:

[0128] The hardware platforms for the simulation experiments of the present invention are: NVIDIA GeForce 3090 and Intel(R) Core(TM) i9-10900X CPU@3.70GHz.

[0129] The software platforms for the simulation experiment of the present invention are: operating system Ubuntu 18.06, Python 3.7 and Pytorch 1.12.

[0130] The datasets used in the simulation experiments of this invention include the Houston 2013 LiDAR and Houston 2018 LiDAR image data, both of which are hyperspectral images and their corresponding LiDAR images. Houston 2013 and Houston 2018 are hyperspectral images of the University of Houston campus and its surrounding scenes acquired at different times by different sensors. The Houston 2013 image data consists of 349 × 1905 pixels, contains 144 spectral bands with a wavelength range of 380–1050 nm, and has an image spatial resolution of 2.5 m. The Houston 2018 image data has the same wavelength range as the Houston 2013 image data but only contains 48 spectral bands with a spatial resolution of 1 m. Both scenes contain seven common ground objects. Forty-eight spectral bands (wavelength range 0.38–1.05 μm) were extracted from the Houston 2013 image data corresponding to the Houston 2018 image data scene, and the overlapping region was selected, with a size of 209 × 955 pixels. We select the Houston 2013 LiDAR image data as the source domain data and the Houston 2018 LiDAR image data as the target domain data. The categories and number of samples are listed in Table 1.

[0131] Table 1 Number of source and target domain data samples

[0132]

[0133] (2) Experimental content and results analysis

[0134] To validate the effectiveness of the proposed method, we selected three widely used cross-domain remote sensing image classification methods: DeepCoral, DAAN, and DSAN. We performed cross-domain remote sensing image classification on the input Houston 2013 LiDAR and Houston 2018 LiDAR image data, respectively, and obtained the final target domain remote sensing image classification results.

[0135] The existing technology used in this invention to compare cross-domain remote sensing image classification methods refers to:

[0136] The prior art DeepCoral cross-domain remote sensing image classification method refers to the cross-domain remote sensing image classification method proposed by Sun et al. in the document “Deep CORAL: Correlation Alignment for Deep Domain Adaptation. In: Hua, G., Jégou, H. (eds) Computer Vision–ECCV 2016 Workshops. ECCV 2016. Lecture Notes in Computer Science (), vol 9915. Springer, Cham.”

[0137] The existing technology DAAN cross-domain remote sensing image classification method refers to the cross-domain remote sensing image classification method proposed by Yu et al. in the document "Transfer Learning with Dynamic Adversarial Adaptation Network, 2019 IEEE International Conference on Data Mining (ICDM), Beijing, China, 2019, pp. 778-786."

[0138] The existing technology DSAN cross-domain remote sensing image classification method refers to the cross-domain remote sensing image classification method proposed by Zhu et al. in the document "Deep Subdomain Adaptation Network for Image Classification," in IEEE Transactions on Neural Networks and Learning Systems, vol. 32, no. 4, pp. 1713-1722, April 2021."

[0139] Two evaluation metrics (overall accuracy, OA, and chi-square coefficient, Kappa) were used to objectively evaluate the target domain image classification results obtained by the four methods. Overall accuracy, OA, represents the proportion of correctly classified samples to the total number of samples. A value closer to 1 indicates higher detection accuracy. The chi-square coefficient, Kappa, measures the consistency between the obtained results and the reference image. A value closer to 1 indicates better method performance. The values of the various evaluation metrics are plotted in Table 2.

[0140] Table 2 Quantitative analysis of target domain remote sensing image classification results obtained by cross-domain remote sensing image classification of Houston 2013-LiDAR and Houston 2018-LiDAR image data using the present invention and the prior art

[0141] DeepCoral DAAN DSAN Proposed OA 52.30 53.27 57.75 67.76 Kappa 33.94 35.81 39.30 53.46

[0142] From Table 2, it can be seen that the total accuracy OA of the present invention reaches 67.76%, and the Kappa value reaches 53.46, which are 10.01% and 14.16 higher than the best comparison method (DSAN) among the currently listed comparison methods, respectively. Both are significantly higher than the existing technology methods, proving that the present invention can effectively improve the classification accuracy of remote sensing images in the target domain. Figure 4 (b) is the target domain remote sensing image classification result obtained by the method proposed in this invention, Figure 4 (a) is its true label map.

[0143] The above simulation experiments show that the cross-domain multimodal remote sensing image classification method based on contrastive learning provided by the present invention realizes unsupervised remote sensing image classification in the target domain, integrates contrastive learning into cross-domain training, effectively reduces the inter-domain differences between the source domain and the target domain, and effectively improves the classification accuracy of remote sensing images in the target domain.

[0144] The above content is only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.

Claims

1. A cross-domain multimodal remote sensing image classification method based on contrastive learning, characterized in that: The following steps are involved: S101: In the pre-training stage, the source domain hyperspectral and source domain lidar image data to be classified are first input, and the source domain data are pre-processed; S102: Use the convolution layer to extract the features of the pre-processed source domain data, flatten the features and then concatenate them to obtain the source domain fusion features; S103: Input the source domain fusion features into the classifier to obtain the classification results, repeat S102 to S103, train the network multiple times, and save the optimal parameters of the source domain network; S104: In the cross-domain comparative learning stage, the source domain hyperspectral and source domain lidar image data are first re-input, and the target domain hyperspectral and target domain lidar image data are input, and the data are preprocessed; the optimal parameters of the source domain network are used as the initial parameters of the source domain network and the target domain network in this stage, and the feature queues corresponding to each category in the target domain are initialized; S105: Use the convolution layer to extract the features of the source domain and target domain data preprocessed in S104, flatten and concatenate the source domain and target domain features respectively, and obtain the source domain fusion features and the target domain fusion features; S106: Input the source domain fusion features and the target domain fusion features into the classifier and mapper to obtain the source domain and target domain classification results and high-dimensional features, update the feature queue of S104 according to the classification results of the target domain, and perform comparative learning at the same time; S107: Back propagation updates the source domain network, momentum updates the target domain network, repeats S104 to S107, and saves the parameters of the optimal target domain network and its classification results.

2. The cross-domain multimodal remote sensing image classification method based on contrastive learning according to claim 1 is characterized in that: The step S101 is specifically as follows: In the pre-training stage, the input source domain hyperspectral and source domain lidar image data are first normalized to obtain the source domain hyperspectral image H S and the source domain lidar image L S , respectively, are obtained through the following two formulas: where H' S is the source domain hyperspectral image that has not been normalized. is the average value of the source domain hyperspectral image, is the standard deviation of the source domain hyperspectral image; L' S is the unnormalized source domain lidar image, is the average value of the source domain lidar image, is the standard deviation of the source domain lidar image; The source domain hyperspectral image H S and the source domain lidar image L S Perform edge filling, taking each pixel before filling as the center, and construct one-to-one corresponding source domain hyperspectral image blocks and source domain lidar image blocks; Next, 200 pairs of labeled source domain hyperspectral image patches and source domain lidar image patches are randomly selected from each category according to the category of the center pixel point as the training set, and the rest are used as the test set.

3. The cross-domain multimodal remote sensing image classification method based on contrastive learning according to claim 1, characterized in that: The step S102 is specifically as follows: In the source domain network, the hyperspectral image processing branch and the lidar image processing branch each apply two convolutional layers for feature extraction; Assuming source domain hyperspectral image patches The size is C×11×11, where C is the number of channels of the hyperspectral image block. Two convolution layers are constructed with a convolution kernel size of 64×3×3, a step size of 2, and a padding of 1, and a convolution kernel size of 32×3×3, a step size of 2, and a padding of 1. The output of each layer is input into the activation function ReLU, and finally the source domain hyperspectral image features are obtained. Its dimensions are 32×3×3; Assume that the source domain lidar image patch The size is 1×11×11, and two convolution layers are constructed with a convolution kernel size of 16×3×3, a step size of 2, and a padding of 1, and a convolution kernel size of 32×3×3, a step size of 2, and a padding of 1. The output of each layer is input into the activation function ReLU, and finally the source domain lidar image features are obtained. Its dimensions are 32×3×3; The convolution of the source domain hyperspectral image features and source domain lidar image features Under the condition that the channel dimension remains unchanged, the features are flattened to obtain the source domain hyperspectral image features with a size of 32×9. and source domain lidar image features of size 32×9 The obtained source domain hyperspectral image features and source domain lidar image features Perform splicing and fusion in the channel dimension to obtain the source domain fusion features Its size is 64×9.

4. The cross-domain multimodal remote sensing image classification method based on contrastive learning according to claim 1, characterized in that: The step S103 is specifically as follows: A network consisting of a linear layer, a batch normalization layer, and a ReLu layer is selected as the classifier. The last layer of the classifier is a linear layer, and the number of its output channels is the number of ground object categories. Assume that the output is Y S , the true label is Y S , the cross entropy loss is as follows: Where M is the number of categories; y ic Is a sign function (0 or 1), if the true category of sample i is equal to c, it takes 1, otherwise it takes 0; P S Y S The probability vector of the predicted sample obtained by the Softmax function; is the predicted probability that the observed sample i belongs to category c; The network learning is supervised by the cross-entropy loss function, and the source domain network parameters are updated using back propagation and stochastic gradient descent methods, and the source domain network parameters with the best performance on the source domain test set are saved.

5. The cross-domain multimodal remote sensing image classification method based on contrastive learning according to claim 1, characterized in that: The specific steps of step S104 are as follows: In the cross-domain comparative learning stage, the preprocessing of the source domain data is consistent with the pre-training stage above. The input target domain hyperspectral image and target domain lidar image are normalized by mean variance to obtain the target domain hyperspectral image H T and the target domain lidar image L T , obtained by the following two formulas: where H' T is the target domain hyperspectral image that has not been normalized, is the average value of the target domain hyperspectral image, is the standard deviation of the target domain hyperspectral image; L' T is the target domain lidar image that has not been normalized, is the average value of the target domain lidar image, is the standard deviation of the target domain lidar image; Then, the target domain hyperspectral image H T and the target domain lidar image L T Perform edge filling, centering on each pixel before filling, to construct one-to-one correspondence between the target domain hyperspectral image block and the target domain lidar image block. The target domain data has no real labels, so there is no need to divide it into training and test sets. All samples can be directly used for training. The optimal parameters of the source domain network saved in S103 are loaded as the initial parameters of the source domain network and the target domain network in subsequent training.

6. The cross-domain multimodal remote sensing image classification method based on contrastive learning according to claim 5, characterized in that: A queue with a capacity of K is set for each category of the target domain data to store the corresponding features and provide data for the subsequent calculation of the contrastive learning loss value; the initialized target domain network is used to test all samples in the target domain, and the index value of the maximum value of the probability vector obtained by the classifier output after the Softmax function is used as the pseudo label. If the confidence of the pseudo label is greater than the set threshold, the corresponding feature output by the mapper will be included in the queue of the corresponding category according to the pseudo label. After completing all tests, the initial feature queue of each category is obtained.

7. The cross-domain multimodal remote sensing image classification method based on contrastive learning according to claim 1, characterized in that: The step S105 is specifically as follows: In the cross-domain contrast learning stage, the structures of the source domain and target domain networks are exactly the same, so the convolution kernels used to extract features are the same, and their parameters are consistent with the source domain convolution kernel parameters in the pre-training stage, so the extracted target domain hyperspectral image features can be obtained. and target domain lidar image features The dimensions are 32×3×3 and 32×3×3; The target domain flattening and splicing operation is consistent with the source domain, and both are spliced in the channel dimension, and the source domain fusion features at this stage can be obtained. Fusion features with the target domain The size is 64×9.

8. The cross-domain multimodal remote sensing image classification method based on contrastive learning according to claim 1, characterized in that: The step S106 is specifically as follows: In this stage, the target domain network uses the same network composed of linear layer, batch normalization layer and ReLu layer as the source domain network as the classifier. The last layer of the classifier is a linear layer, and the number of its output channels is the number of ground object categories. However, since the target domain data has no real labels, there is no need to calculate the target domain cross entropy loss. Only the source domain cross entropy loss needs to be calculated. The calculation method of the source domain cross entropy loss in this stage is consistent with that in the pre-training stage. At this stage, both the source and target domain networks use a network consisting of a linear layer, a batch normalization layer, and a ReLu layer as a mapper. The mapper maps features to a high-dimensional space. The feature queue stores the corresponding high-dimensional features output by the mapper based on the classification results of the target domain network and dequeues some earlier features.

9. The cross-domain multimodal remote sensing image classification method based on contrastive learning according to claim 8, characterized in that: According to the true label of the current sample in the source domain, the feature queue of the corresponding category in the target domain is selected, and the average value of all features in the queue is obtained to obtain a high-dimensional feature. The corresponding high-dimensional feature output by the source domain sample through the mapper is regarded as the positive sample of the high-dimensional feature, and the high-dimensional features in the feature queues corresponding to other categories are regarded as its negative samples. The contrast loss is calculated using the InFoNCE loss function, as shown in the following formula: Where q is the high-dimensional feature after the mean is obtained for the corresponding queue, k + is the high-dimensional feature of the source domain samples, N is the number of all samples (including positive and negative samples), and τ is the temperature hyperparameter.

10. The cross-domain multimodal remote sensing image classification method based on contrastive learning according to claim 1, characterized in that: The step S107 is specifically as follows: At this stage, the source domain cross entropy loss function L CE and contrast loss L CL Guide the source domain network learning and update the source domain network parameters using back propagation and stochastic gradient descent methods; The target domain network gradient will not be back-propagated, and the target domain network parameters are updated through the following momentum update method: θ=m·θ+(1-m)·ξ The target domain network parameter is θ, the source domain network parameter is ξ, and m is a hyperparameter; Train multiple times until the network converges, and save the parameters and classification results of the target domain network that has the best test effect on the target domain samples.

Citation Information

Patent Citations

  • Multi-sensor remote sensing image fusion classification method of hierarchical dense fusion network

    CN113255727A

  • Cross-domain remote sensing scene classification and retrieval method based on self-supervised contrast learning

    CN115471739A