Domain adaptation hyperspectral image classification method and system based on mask consistency
Through the mask consistency regularization method, strong and weak enhancement strategies are used to screen high-confidence pseudo labels and combined with consistency loss to solve the problems of difficulty in obtaining labeled samples and spectral drift in hyperspectral image classification, and improve the classification performance of the model in the target domain.
Patent Information
- Application Number
- CN202510630032.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-10-03
AI Technical Summary
Existing hyperspectral image classification methods fail to fully utilize target domain data when faced with the difficulty of obtaining labeled samples and spectral drift problems. In addition, existing pseudo-labeling methods have problems such as uncertainty in pseudo-labels of difficult samples and insufficient model generalization ability.
The mask consistency regularization method is adopted to screen high-confidence pseudo labels through strong and weak enhancement strategies, and the target domain data is used for training. The confidence sample consistency and class-related consistency loss are combined to improve the model's discriminative feature extraction ability for the target domain.
It effectively improves the accuracy and generalization ability of hyperspectral image classification, especially in the case of spectral drift, and improves the adaptability and classification performance of the model to the target domain data.
Smart Images

Figure CN120747576A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image classification, and in particular relates to a domain-adaptive hyperspectral image classification method and system based on mask consistency. Background Art
[0002] Hyperspectral images, as a typical example of spatial-spectral multidimensional data, can capture diagnostic spectral information about ground objects, reflecting their inherent characteristics and serving as a crucial basis for their identification and discrimination. Due to their high spectral resolution, excellent material recognition capabilities, relatively high spatial resolution, and potential for detecting hidden information, hyperspectral images hold broad application prospects in numerous fields, such as geological and mineral exploration, land cover classification, agricultural monitoring, and military reconnaissance. However, in real-world scenarios, acquiring labeled samples is often time-consuming and costly, resulting in collected data often having only a few labels or even no labels at all. This poses a significant challenge to the accurate classification of hyperspectral images. Furthermore, due to varying acquisition conditions, such as time, location, and sensor, hyperspectral images often exhibit significant spectral drift. Even if a model is fully trained on a well-labeled hyperspectral image dataset, direct application to another dataset with spectral drift often yields unsatisfactory classification performance. Therefore, mitigating the impact of spectral drift and improving model performance remain key challenges that need to be addressed.
[0003] In recent years, domain adaptation techniques have garnered increasing attention in the field of computer vision. Research has shown that they can improve the performance of hyperspectral image classification to a certain extent. The core goal of domain adaptation is to leverage knowledge from a richly labeled source domain to learn a model applicable to a target domain with scarce or no labels. This extends the model's generalization capabilities from the source domain to the target domain while mitigating the challenges of spectral drift. Traditional domain adaptation methods can be categorized into three main categories: instance-based, feature-based, and classifier-based. Instance-based methods adjust the marginal distribution of samples in the source or target domain to reduce inter-domain differences and achieve distribution alignment. Feature-based methods map source and target domain data into a common feature space, aligning the feature distributions of the two domains. Classifier-based methods adapt classifiers trained in the source domain to the target domain by taking into account the characteristics of unlabeled samples in the target domain. There are also other deep domain adaptation methods. Domain adaptation methods based on deep learning can achieve features with both stronger generalization and better transferability. While these methods have achieved some success in image classification tasks, they often fail to fully utilize target domain data, resulting in insufficient discriminative performance.
[0004] To address the limitation of underutilized target domain data, some authors have adopted pseudo-labeling techniques. Pseudo-labeling is a widely used technique in domain adaptation. It involves assigning pseudo-labels to unlabeled data in the target domain using a model originally trained on the source domain. It leverages unlabeled data to improve the model's performance in the target domain, allowing the model to generalize to the target domain. While pseudo-labeled examples from the target domain can provide additional supervisory information to enhance the model's adaptive capabilities, existing domain adaptation methods based on pseudo-labeling still face two challenges. The first challenge lies in how to utilize the obtained pseudo-labels for domain distribution alignment. Most existing domain adaptation methods uniformly apply all obtained pseudo-labels to domain distribution alignment, ignoring the accuracy of their pseudo-labels. Specifically, the target domain often contains examples that differ significantly from the source domain data and are difficult for the model to correctly classify, namely hard examples. The reason these hard examples are difficult to correctly classify is that they contain discriminative features that are difficult to extract, resulting in uncertain pseudo-labels. Pseudo-labels for examples that are highly similar to the source domain data and easily classified by the model (simple examples) are generally more reliable. Therefore, treating all pseudo-labels equally for classification may lead to a significant performance drop. The second challenge is that since the pseudo-labels corresponding to difficult samples have high uncertainty, it is difficult to determine the correctness of the pseudo-labels for difficult samples. This problem ultimately hinders the generalization ability of the model. Summary of the Invention
[0005] To address the aforementioned technical issues, a mask consistency regularization method is proposed, which allows the target domain data to be fully utilized during training. During training, a batch of target domain data is first slightly perturbed (flipped) to obtain the predicted probabilities for this batch of samples. Then, a threshold τ is set to select high-confidence samples, which serve as simple samples and are used to derive highly accurate pseudo-labels. At the same time, these high-confidence samples are strongly perturbed, i.e., data augmentation with flipping and masking, allowing the model to learn discriminative features that are difficult to extract in the target domain. Furthermore, this method robustly improves the class separability of easily confused classes in the target domain through the strongly augmented samples.
[0006] The present invention provides a domain-adaptive hyperspectral image classification method based on mask consistency, comprising the following steps:
[0007] Step 1: Use the image encoder to extract the global features of the original target domain samples. Based on this global feature, use a dual classifier to classify the image samples and align the distribution of similar objects in the source domain and target domain.
[0008] Step 2: The original target domain samples are strongly and weakly enhanced by using strong and weak enhancement strategies to obtain processed target domain samples; the discriminative features in the target domain are extracted by using the processed target domain samples to calculate the confidence samples and class-related consistency loss values, while distinguishing the confusion categories;
[0009] Step 3: Image classification is achieved by aligning the distribution of similar categories in the source domain and the target domain and extracting discriminative features in the target domain to distinguish these confused categories at the same time, thereby enhancing the performance of domain-adapted hyperspectral image classification.
[0010] Furthermore, the distribution of the same type of data in the source domain and the target domain is aligned in step 1, which specifically includes the following steps:
[0011] (1-1) The image is fed into the image encoder, and the attention-based spatial features and attention-based spectral features of the hyperspectral image are extracted through two branches respectively;
[0012] (1-2) fusing the attention-based spatial features and the attention-based spectral features of the extracted hyperspectral image to obtain fused image features;
[0013] (1-3) The fused image features are fed into the dual classifier to obtain the predicted probability;
[0014] (1-4) Calculate the source classification loss as follows: The source classification loss is the cross entropy loss. The calculation formula of the source classification loss is as follows:
[0015]
[0016] Among them, N s is the number of source domain samples, is the label, G(·) is the output of the image encoder, is the image feature obtained by the image encoder, F(·) is the output of the two classifiers, is the predicted probability; L ce (·) represents the cross entropy loss;
[0017] (1-5) The steps to update the classifier using the adversarial loss function are as follows: The outputs of the target domain samples obtained by the two classifiers are measured using the absolute value indicator to align the distribution of the same type in the source domain and the target domain. The adversarial loss function is as follows:
[0018]
[0019] where N t is the number of target domain samples, is the predicted probability of the i-th sample obtained by the two classifiers.
[0020] Furthermore, in step 2, the target domain samples obtained by strong and weak enhancement are calculated using the following formula:
[0021] X t_strong =A(X t ),X t_weak =a(X t ) (3)
[0022] Among them, A(·), a(·) refers to the strong and weak enhancement strategy, X t is a sample of the target domain.
[0023] Furthermore, step 2 extracts the discriminative features in the target domain, which specifically includes the following steps:
[0024] (2-1) Calculate the confidence sample consistency loss. The calculation formula is as follows:
[0025]
[0026] Among them, N t is the number of target domain samples, It is the probability prediction of the strong and weak enhancement data obtained by the image encoder and the first classifier. It is the probability prediction obtained by weakly enhanced data The resulting pseudo-labels are calculated; τ is a key scalar hyperparameter used as a confidence threshold to filter and retain high-quality pseudo-labels;
[0027] (2-2) Calculate the class-related consistency loss. First, perform probability adjustment. The calculation formula is as follows:
[0028]
[0029] Where C represents the number of categories, p ij represents the probability that the i-th sample belongs to the j-th class, and T is a hyperparameter representing the temperature;
[0030] Secondly, calculate the class correlation, the calculation formula is as follows:
[0031]
[0032] where z ·j′ is obtained from formula (5);
[0033] The entropy function is used to measure the level of uncertainty associated with a sample, and is calculated as follows:
[0034]
[0035] Where C represents the number of categories;
[0036] SoftMax is applied; the weight expression after conversion is:
[0037]
[0038] Where B is the batch size;
[0039] Using this weighting mechanism, class confusion can be rewritten as (6):
[0040]
[0041] Then the category normalization method used in the random walk algorithm was adopted. The specific expression is as follows:
[0042]
[0043] Where C represents the number of categories. Similarly, probability calibration is also used as (5) as z for the strongly enhanced target domain samples. ij ; and the uncertainty reweighting strategy in (8) is adopted to have a higher credibility for the sample; each element in the strongly enhanced class confusion matrix is:
[0044]
[0045] Finally, the class-related consistency loss is calculated, and the calculation formula is as follows:
[0046]
[0047] Among them, the domain alignment module extracts the features of hyperspectral images through an image encoder and classifies them using a dual classifier; the mask consistency regularization module includes a confidence sample consistency module and a class-related consistency module, which are used to mine discriminative features that are difficult to extract in the target domain and alleviate class confusion problems.
[0048] Furthermore, in step 3, the total loss consisting of four loss functions is minimized to enhance the performance of domain-adapted hyperspectral image classification. The calculation formula is as follows:
[0049]
[0050] in represents the source classification loss, Denotes the adversarial loss, L csc represents the confidence sample consistency loss, L ccc represents the class-dependent consistency loss.
[0051] According to another aspect of the present invention, a domain-adaptive hyperspectral image classification system based on mask consistency is provided, which is characterized by comprising:
[0052] The domain alignment module extracts global features of the original target domain samples through the image encoder, classifies the image samples based on the global features using a dual classifier, and aligns the distribution of similar objects in the source domain and the target domain;
[0053] The mask consistency regularization module: performs strong and weak enhancement on the original target domain samples through strong and weak enhancement strategies to obtain processed target domain samples; calculates the confidence sample consistency loss value and the class-related consistency loss value using the processed target domain samples to extract discriminative features in the target domain and distinguish confusion categories at the same time;
[0054] Performance enhancement module for hyperspectral image classification: It achieves image classification by aligning the distribution of similar classes in the source and target domains and extracting discriminative features in the target domain to distinguish these confused classes at the same time, thereby enhancing the performance of domain-adapted hyperspectral image classification.
[0055] Advantages of the present invention: The problems that the present invention needs to solve are: it is difficult to obtain labeled samples of hyperspectral images, there is spectral drift between the hyperspectral images of the source scene and the target scene, and the existing technology still fails to fully utilize the information of the unlabeled hyperspectral images in the target domain. The present invention provides a domain-adaptive hyperspectral image classification method based on mask consistency, which can make full use of the target domain data to participate in training. During the training process, we first slightly perturb (flip) the target domain data in a batch to obtain the prediction probability of this batch of samples, and then use the set threshold τ to select those samples with high confidence, let these samples be simple samples, and use simple samples to obtain pseudo labels with high accuracy. At the same time, for these high-confidence samples, we use strong perturbation, that is, flipping and masking data enhancement, to let the model learn discriminative features that are difficult to extract in the target domain. At the same time, this method robustly improves the category separability of easily confused classes in the target domain through strongly enhanced samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A framework diagram of a domain-adaptive hyperspectral image classification method and system method based on mask consistency provided by the embodiments disclosed in the present invention;
[0057] Figure 2 A diagram of the image encoder framework provided by the disclosed embodiment of the present invention;
[0058] Figure 3 A diagram showing the classification results of the Houston dataset provided in the embodiment disclosed in the present invention;
[0059] Figure 4 A diagram showing the classification results of the YanCheng dataset provided in the embodiments disclosed in the present invention;
[0060] Figure 5This is a diagram of the classification results of the HyRANK dataset provided in the embodiment disclosed in the present invention. DETAILED DESCRIPTION
[0061] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0062] like Figure 1 As shown, the present invention discloses a domain adaptation hyperspectral image classification method and system based on mask consistency. This solves the problem that existing unsupervised domain adaptation methods fail to fully utilize target domain data. The present invention adopts the following technical solutions:
[0063] Step 1: Build a domain alignment network, extract global features through the image encoder, use a dual classifier to classify image samples, align the distribution of similar objects in the source domain and the target domain, and generate pseudo labels for the target domain.
[0064] Step 2: Construct a mask consistency regularization module, and use the confidence sample consistency module and class-related consistency module mining model to mine the discriminative features that are difficult to extract in the target domain and prevent the misclassification of the target domain data and the enhanced target domain data during the domain alignment process while distinguishing these confusing categories.
[0065] In order to fully extract image features, the image encoder of this embodiment adopts a dual-branch network to extract the attention-based spatial features and attention-based spectral features of the hyperspectral image respectively, and fuses the two.
[0066] To fully utilize the target domain data, a pixel mask augmentation operation is performed on the target domain data. Confidence sample consistency and class confusion matrix consistency resulting from confusion classes in the target domain are also tested. A cross-entropy loss function is used to force the model to generalize to discriminative features that are difficult to extract in the target domain. This allows the model to identify these difficult-to-extract but highly discriminative features during subsequent training, mitigating the model's tendency to mispredict target domain data during testing.
[0067] In step 1, the image is fed into the image encoder, and the attention-based spatial features and attention-based spectral features of the hyperspectral image are extracted through two branches respectively, and the two are fused. After the image features are fed into the dual classifier, the source classification loss is calculated. The loss is the cross entropy loss, and the source classification loss is shown in formula (1):
[0068]
[0069] in is the label, G(·) is the output of the image encoder, is the image feature obtained by the image encoder, F(·) is the output of the two classifiers, is the predicted probability. L ce (·) represents the cross entropy loss.
[0070] Furthermore, adversarial loss is used to update the classifier. The output of the target domain samples obtained by the two classifiers is measured using the absolute value indicator to measure the difference. The adversarial loss is shown in formula (2):
[0071]
[0072] The step 2 is specifically as follows: In the mask consistency regularization module, it is divided into three parts to explain
[0073] (1) Strong and weak data augmentation strategies: Since the center pixel holds the classification information of the entire sample, it is hoped that when the center pixel is masked, the model can be forced to learn the relationship between the adjacent pixels and the center pixel in the sample, so as to extract those discriminative features that are difficult to extract in the target domain. First, strong and weak augmentation are performed on the unlabeled target domain data. Weak augmentation is a random flipping strategy for the sample. Strong augmentation first randomly flips the sample and then randomly masks each pixel. In this process, the center pixel may also be randomly masked. Samples obtained through strong and weak augmentation.
[0074] X t_strong =A(X t ),X t_weak =a(X t ) (3)
[0075] Where A(·), a(·) represent strong and weak enhancement strategies. The target domain samples and the samples after strong and weak enhancement are fed into the image encoder and classifier to obtain probability predictions.
[0076]
[0077] (2) Confidence sample consistency module: Use the probability prediction obtained from the weakly enhanced data obtained by the first classifier to calculate the label Then the predicted probability of the strongly enhanced data obtained by the first classifier is Cross entropy loss is enforced to enable the model to exploit discriminative features that are difficult to extract in the target domain. The confidence sample consistency loss is defined as:
[0078]
[0079] Among them, τ is a key scalar hyperparameter that serves as a confidence threshold to screen and retain high-quality pseudo-labels. By introducing τ, samples with higher prediction confidence can be selected, thereby ensuring the reliability of pseudo-labels. This mechanism is particularly important in unsupervised domain adaptation tasks because, in the absence of true labels, it uses consistency regularization to mine discriminative features that are difficult to extract in the target domain, thereby enhancing the adaptability of the model. However, the value of τ has a significant impact on model performance: if the threshold is set too high, it may lead to an insufficient number of valid samples; if it is too low, unreliable noisy pseudo-labels may be introduced. Therefore, in practical applications, τ needs to be carefully tuned through experiments to achieve the best balance between sample utilization and pseudo-label quality.
[0080] The core concept of this module is to improve the model's adaptability to the target domain data distribution by utilizing a contrastive learning strategy of strong and weak augmentation. Weak augmentation generates stable baseline predictions through slight perturbations, providing a reliable foundation for pseudo-labeling; while strong augmentation forces the model to learn more robust feature representations by simulating more complex noise scenarios. On this basis, the confidence sample consistency loss further strengthens the training process, gradually improving the model's classification ability for unlabeled data through pseudo-label screening and consistency constraints. This approach not only fully utilizes the complementary properties of strong and weak augmentation, but also effectively reduces the impact of pseudo-label noise through a confidence screening mechanism, thereby achieving higher performance and stability in unsupervised domain adaptation.
[0081] (3) Class-related consistency module: predicting probability for strongly enhanced data And the predicted probability of the original data of the target domain A class-dependent consistency loss is designed. This loss aims to enhance the class separability of the target domain data and its augmented version, while effectively alleviating class confusion. By imposing a consistency constraint on the predictions of both, the model can better capture the discriminative features between classes, thereby improving classification performance.
[0082] First, probability adjustment: In previous studies, deep neural networks often produce overconfident predictions, which may hinder their ability to effectively resolve class confusion. Therefore, temperature adjustment in (MCC) is adopted to mitigate the adverse effects of overconfident predictions:
[0083]
[0084] Among them, p ij represents the probability that the i-th sample belongs to the j-th class, and T is a hyperparameter representing the temperature.
[0085] Then we get the class correlation: ij The relationship between the i-th instance and the j-th class is revealed, and the class correlation between two classes j and j′ is defined as:
[0086]
[0087] It explains the possibility that the classifier classifies B samples into both class j and class j′.
[0088] Next, we need to derive the uncertainty weight: Note that examples are not equally important for quantifying class confusion. When the prediction is closer to a uniform distribution and lacks a clear peak (the probability of some classes is significantly greater), the classifier is considered to be ignorant of the example. In contrast, when the prediction shows multiple peaks, it indicates the ambiguity of the classifier between several confusing classes. Obviously, examples that make the classifier ambiguous across classes are more suitable for describing class confusion. Use the entropy function to measure the level of uncertainty associated with a sample. The expression is as follows:
[0089]
[0090] On the other hand, the goal of is to emphasize the distribution with higher accuracy, so SoftMax is applied. The weight expression after conversion is:
[0091]
[0092] Using this weighting mechanism, class confusion can be rewritten as (5):
[0093]
[0094] In addition, due to the large number of classes, there will be a serious class imbalance in each batch. To solve this problem, a class normalization technique widely used in random walk algorithms is adopted, which is expressed as follows:
[0095]
[0096] Then the class confusion matrix generated from the original data is enforced to be consistent with the enhanced class confusion matrix. Specifically, in each batch, the classifier output of the strongly enhanced target image is Similarly, we also use the probability calibration as (4) as a strong enhancement of z ij In order to mitigate the impact of unreliable predictions, the uncertainty reweighting strategy in (8) is adopted to have a higher confidence in the samples. Each element in the strongly enhanced class confusion matrix is:
[0097]
[0098] Finally, the class-related consistency loss is calculated:
[0099]
[0100] In this process, the class-related consistency module can prevent the misclassification of target domain data and enhanced target domain data during the domain alignment process while distinguishing these confusing categories.
[0101] To summarize, during training, the total loss consisting of four loss functions is minimized to achieve higher classification accuracy:
[0102] L total =L cls +L dis +L csc +L ccc (13)
[0103] Example 1
[0104] Image encoder and dual classifier
[0105] like Figure 2 As shown in Figure 1, the attention-based image encoder is divided into two parts: spectral feature extraction and spatial feature extraction. Taking the Houston dataset as an example, the input data block size is 7×7×48, where 7×7 is the spatial neighborhood size and 48 is the number of bands.
[0106] The spectral feature extraction branch consists of three Conv3d-BN-ReLU units, a concatenated Conv3d, and an SE attention module. The outputs of the first Conv3d-BN-ReLU and Conv3d are connected using a residual connection. After passing through the third Conv3d-BN-ReLU, the resulting data is input into the spectral SE attention module. The learned attention weights are multiplied by the spectral features to generate attention-based spectral features. This considers the importance of each spectral band, enabling the model to focus more on spectral features.
[0107] The spatial feature extraction branch focuses on extracting spatial information from the image. It consists of two Conv3d-BN rectified linear units (ReLUs), a series of Conv3ds, and a SE attention module. Similarly, after the data passes through two Conv3d-BN-ReLUs and a residual operation, it is input into the spatial SE attention module. This module enables the model to focus on and utilize spatial features.
[0108] Finally, the spectral and spatial attention features obtained by the spectral feature extraction branch and the spatial feature extraction branch are fused to obtain the final spatial-spectral features as image features. This processing enables the model to fully understand all aspects of the sample, improving classification accuracy and model generalization capabilities.
[0109] The dual classifier consists of a multi-layer FC-BN-ReLU-Dropout structure, where FC, BN, and ReLU represent the fully connected, batch normalized, and activation layers, respectively. The two classifiers generate probability vectors from the image features obtained by the image encoder and then apply the SoftMax function to obtain logistic vectors, representing the predicted probabilities of the corresponding classifiers for the image category. This structure not only enhances the model's nonlinear expressiveness but also prevents overfitting and enhances its generalization capabilities.
[0110] Finally, the entire network is optimized by alternately updating the image encoder and the dual classifier.
[0111] Example 2
[0112] Confidence Sample Consistency Module
[0113] The confidence sample consistency module uses the structure of the image encoder and the first classifier. The cross entropy loss is enforced in the confidence sample consistency module to explore the model's ability to mine discriminative features that are difficult to extract in the target domain.
[0114] Example 3
[0115] Class-related consistency modules
[0116] The class-related consistency module uses the same image encoder and dual classifier as the domain alignment stage. It imposes consistency constraints on the prediction results of the original target domain samples and the strongly enhanced target domain samples based on the prediction probabilities. This allows the model to better capture the discriminative features between categories, thereby improving classification performance.
[0117] Example 4
[0118] This embodiment presents a domain-adaptive hyperspectral image classification method and system based on mask consistency, consisting of two main components: a domain alignment component and a mask consistency regularization component. In the domain alignment component, an image encoder obtains image features, and a dual classifier generates predictions. Strong and weak augmentation schemes are then used to mine discriminative features that are difficult to extract in the target domain from the augmented target domain data. This enhances the model's adaptability, prevents misclassification of the target domain data and the augmented target domain data during the domain alignment process, and distinguishes between these confusing categories.
[0119] Specifically, the confidence sample consistency module in the mask consistency regularization is only performed in stage A of domain adaptation. In stage A, the image encoder and dual classifier are updated using labeled source domain samples, enabling the image encoder to fully learn the discriminative features of the source domain. This allows the dual classifier to be trained to accurately classify source domain samples while minimizing the confidence sample consistency loss. Simultaneously, a class-dependent consistency loss is applied to the prediction probabilities of the strongly augmented data obtained by the two classifiers and the prediction probabilities of the original data. This loss prevents misclassification of target and augmented target data during domain alignment and distinguishes between these confused classes. In stage B, the image encoder trained in stage A is updated by increasing the difference between the dual classifier outputs, while also minimizing the class-dependent consistency loss. In stage C, the image encoder is optimized using the fixed dual classifier, minimizing the difference between the dual classifier outputs on target samples, so that the predictions of the target samples from the two classifiers are as consistent as possible, while also minimizing the class-dependent consistency loss.
[0120] In addition, this paper selects three different hyperspectral datasets (Houston, YanCheng, HyRANK) and compares them with seven related work algorithms (Domain Adaptation Network, DAN; Domain-Adversarial Neural Networks, DANN; Class-wise Distribution Adaptation, CDA; Deep metric learning-based feature embedding Domain Adaptation, ED-DMM-UDA; Two-branch Attention Adversarial DA, TAADA; Mind the Gap: Multilevel Unsupervised Domain Adaptation, MLUDA; Masked Self-Distillation Domain Adaptation, MSDA). The results are as follows: Figure 3-5 As shown, the effectiveness of the present invention is demonstrated.
[0121] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A domain-adaptive hyperspectral image classification method based on mask consistency, characterized in that: The following steps are included: Step 1: Use the image encoder to extract the global features of the original target domain samples. Based on this global feature, use a dual classifier to classify the image samples and align the distribution of similar objects in the source domain and target domain. Step 2: The original target domain samples are strongly and weakly enhanced by using strong and weak enhancement strategies to obtain processed target domain samples; the discriminative features in the target domain are extracted by using the processed target domain samples to calculate the confidence samples and class-related consistency loss values, while distinguishing the confusion categories; Step 3: Image classification is achieved by aligning the distribution of similar categories in the source domain and the target domain and extracting discriminative features in the target domain to distinguish these confused categories at the same time, thereby enhancing the performance of domain-adapted hyperspectral image classification.
2. The domain-adaptive hyperspectral image classification method based on mask consistency according to claim 1, characterized in that: In step 1, the distribution of similar data in the source domain and the target domain is aligned, which specifically includes the following steps: (1-1) The image is fed into the image encoder, and the attention-based spatial features and attention-based spectral features of the hyperspectral image are extracted through two branches respectively; (1-2) fusing the attention-based spatial features and the attention-based spectral features of the extracted hyperspectral image to obtain fused image features; (1-3) The fused image features are fed into the dual classifier to obtain the predicted probability; (1-4) Calculate the source classification loss as follows: The source classification loss is the cross entropy loss. The calculation formula of the source classification loss is as follows: Among them, N s is the number of source domain samples, is the label, G(·) is the output of the image encoder, is the image feature obtained by the image encoder, F(·) is the output of the two classifiers, is the predicted probability; L ce (·) represents the cross entropy loss; (1-5) The steps to update the classifier using the adversarial loss function are as follows: The outputs of the target domain samples obtained by the two classifiers are measured using the absolute value indicator to align the distribution of the same type in the source domain and the target domain. The adversarial loss function is as follows: where N t is the number of target domain samples, is the predicted probability of the i-th sample obtained by the two classifiers.
3. The domain-adaptive hyperspectral image classification method based on mask consistency according to claim 1, characterized in that: In step 2, the target domain samples obtained by strong and weak enhancement are calculated using the following formula: X t_strong =A(X t ),X t_weak =a(X t ) (3) Among them, A(·), a(·) refers to the strong and weak enhancement strategy, X t is a sample of the target domain.
4. The domain-adaptive hyperspectral image classification method based on mask consistency according to claim 1, characterized in that: Step 2 extracts discriminative features in the target domain, which specifically includes the following steps: (2-1) Calculate the confidence sample consistency loss. The calculation formula is as follows: Among them, N t is the number of target domain samples, It is the probability prediction of the strong and weak enhancement data obtained by the image encoder and the first classifier. It is the probability prediction obtained by weakly enhanced data The resulting pseudo-labels are calculated; τ is a key scalar hyperparameter used as a confidence threshold to filter and retain high-quality pseudo-labels; (2-2) Calculate the class-related consistency loss. First, perform probability adjustment. The calculation formula is as follows: Where C represents the number of categories, p ij represents the probability that the i-th sample belongs to the j-th class, and T is a hyperparameter representing the temperature; Secondly, calculate the class correlation, the calculation formula is as follows: where z ·j′ is obtained from formula (5); The entropy function is used to measure the level of uncertainty associated with a sample, and is calculated as follows: Where C represents the number of categories; SoftMax is applied; the weight expression after conversion is: Where B is the batch size; Using this weighting mechanism, class confusion can be rewritten as (6): Then the category normalization method used in the random walk algorithm was adopted. The specific expression is as follows: Where C represents the number of categories. Similarly, probability calibration is also used as (5) as z for the strongly enhanced target domain samples. ij ; and the uncertainty reweighting strategy in (8) is adopted to have a higher credibility for the sample; each element in the strongly enhanced class confusion matrix is: Finally, the class-related consistency loss is calculated, and the calculation formula is as follows: Among them, the domain alignment module extracts the features of hyperspectral images through an image encoder and classifies them using a dual classifier; the mask consistency regularization module includes a confidence sample consistency module and a class-related consistency module, which are used to mine discriminative features that are difficult to extract in the target domain and alleviate class confusion problems.
5. The domain-adaptive hyperspectral image classification method based on mask consistency according to claim 1, characterized in that: In step 3, the total loss composed of four loss functions is minimized to enhance the performance of domain-adapted hyperspectral image classification. The calculation formula is as follows: in represents the source classification loss, Denotes the adversarial loss, L csc represents the confidence sample consistency loss, L ccc represents the class-dependent consistency loss.
6. A domain-adaptive hyperspectral image classification system based on mask consistency, characterized in that: include: the domain alignment module; The global features of the original target domain samples are extracted through the image encoder. Based on this global feature, the image samples are classified using a dual classifier to align the distribution of similar samples in the source domain and the target domain. The mask consistency regularization module: performs strong and weak enhancement on the original target domain samples through strong and weak enhancement strategies to obtain processed target domain samples; calculates the confidence sample consistency loss value and the class-related consistency loss value using the processed target domain samples to extract discriminative features in the target domain and distinguish confusion categories at the same time; Performance enhancement module for hyperspectral image classification: It achieves image classification by aligning the distribution of similar classes in the source and target domains and extracting discriminative features in the target domain to distinguish these confused classes at the same time, thereby enhancing the performance of domain-adapted hyperspectral image classification.