Domain generalization hyperspectral image classification method based on uncertainty perception and language guidance
The dismuted samples were extracted through uncertainty perception and language guidance, and an intermediate subdomain that took into account both diversity and authenticity, solving the diversity and authenticity balance of cross-scene hyperspectral image classification in the existing methods, achieving a robust generalization effect under target-free domain data.
Patent Information
- Application Number
- CN202510289990.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-11
AI Technical Summary
The existing domain adaptation method requires the training of target domain data, which is difficult to adapt to the new target domain. The existing domain generalization method cannot effectively balance diversity and authenticity when generating intermediate subdomains, resulting in insufficient classification performance of cross-scene hyperspectral images.
The undeterministic samples are extracted through the uncertainty perception module, combined with language guidance to generate intermediate subdomains, use the diversity of the undeterministic samples and text description labels to ensure the authenticity of the subdomains, and use the contribution decision-making balance module and the multi-loss function optimization classifier.
It realizes robust generalization in the absence of target domain data, and improves the accuracy and robustness of cross-scene hyperspectral image classification.
Smart Images

Figure CN120298752A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to a domain generalization hyperspectral image classification method based on uncertainty perception and language guidance. Background Art
[0002] In recent years, domain adaptation techniques have been applied to cross-scene hyperspectral image classification tasks in the prior art. This method aims to reduce the distribution difference between the source domain and the target domain to promote cross-domain knowledge transfer. Although the hyperspectral image classification method based on domain adaptation has achieved impressive generalization performance, these models require target domain samples to participate in training during the learning process, which has great limitations: these models are customized for specific target domains, and when facing other new target domains, the models need to be retrained. However, with the increasing development of remote sensing technology and the rapid growth of hyperspectral image data volume, it is difficult to implement customizing models for each new target domain in practical application scenarios. For example, in practical application scenarios where target domain data such as geological disaster relief, fire extinguishing, or military reconnaissance cannot be effectively obtained, the domain adaptation method will no longer be applicable.
[0003] To address the above challenges, domain generalization methods can be considered to complete the cross-scene hyperspectral image classification task. Domain generalization methods aim to learn domain-invariant knowledge under the condition that target domain data is not visible during training and achieve robust generalization to unseen target domains. Currently, mainstream domain generalization methods simulate and generate intermediate sub-domains through data augmentation methods to expand the diversity of the source domain distribution, fully cover the feature distribution of the target domain, and expect to improve the generalization performance of the model in unseen target domains. However, the above domain generalization methods have two problems. Problem 1: These methods usually use data augmentation methods for all data without selection to generate intermediate sub-domains with diversity. However, there are always some distorted samples in the source domain due to some influencing factors. These non-selective data augmentation methods will cause the intermediate sub-domains obtained based on the distorted samples to cross the category feature distribution boundary and no longer be effective diversity information. In addition, since the environmental factors that cause the inherent differences between cross-scene hyperspectral images are similar to the reasons for generating distorted samples in the source domain, most of the distorted samples are distributed in the intersection part of the source domain and the target domain, which makes it possible to use the distorted samples to fully simulate the target domain data. Problem 2: These domain generalization methods usually customize prior constraint conditions for the domain to ensure the authenticity of the generated intermediate sub-domains. Some researchers try to use morphological constraints or manifold constraints to ensure the authenticity of the intermediate sub-domains. Morphological constraints rely on the assumption that the spatial morphology of the ground objects in the source domain and the target domain is highly consistent. However, in the actual application scenario of cross-scene hyperspectral image classification, the spatial morphology differences are likely to be too large, and it is difficult to meet this assumption at this time. The effectiveness of manifold constraints highly depends on the artificial prior of the data manifold, and these predefined artificial priors are difficult to achieve optimality in complex real-world scenarios. Summary of the Invention
[0004] Object of the Invention: Aiming at the problems existing in the above-mentioned background technology, the present invention provides a domain generalization hyperspectral image classification method based on uncertainty perception and language guidance, which takes into account the diversity and authenticity of sub-domains through distorted sample mining and language prior guidance.
[0005] Summary of the Invention: To achieve the above object, the technical solution adopted by the present invention is: A domain generalization hyperspectral image classification method based on uncertainty perception and language guidance, comprising the following steps:
[0006] Step 1, splitting the hyperspectral image source domain data into category sub-images according to the label category mask;
[0007] Extracting a random sequence of bands from the category sub-images and inputting them into the uncertainty perception module to obtain the Gaussian distribution prediction results of each category sub-image where K represents the total number of category sub-images;
[0008] After obtaining the Gaussian distribution prediction result U, extract the class distortion sample mask, that is, obtain the corresponding single-class distortion sample mask, and take the union of all class distortion sample masks to obtain the global distortion sample mask. The position corresponding to the source domain dataset is the global distortion data group, and then extract the global distortion sample data group from the source domain dataset;
[0009] Step 2: Randomly sample distortion samples from the distortion sample data group, perform batch masking on the distortion samples to obtain masked data; input the masked data into the spectral feature extraction block and the spatial feature extraction block respectively to obtain the spectral feature and the spatial feature of the distortion samples, linearly combine the spectral feature and the spatial feature, and perform residual connection with the masked data. Finally, obtain the spatial-spectral feature of the distortion samples after pooling;
[0010] Step 3: Text-image fusion learning
[0011] First, randomly sample data from the source domain data of the hyperspectral image, input the sampled data into the image encoder to obtain image features, and at the same time input the class text description label of the sampled data into the text encoder to obtain text features. Finally, obtain the text-image fusion features through visual-language alignment driven by supervised contrastive loss.
[0012] Step 4: After randomizing the spatial-spectral feature of the distortion samples and the text-image fusion feature respectively, input them into the Contribution Decision Balance Module (CDBM) for decision balance to generate an intermediate sub-domain. Finally, input the intermediate sub-domain into the classifier for classification.
[0013] Furthermore, the uncertainty perception module includes a multi-level feature perception (MFA) encoder, two feature aggregation triangle (FPT) decoders, and two segmentation heads, where one feature aggregation triangle (FPT) decoder and one segmentation head form a group; the feature aggregation triangle (FPT) decoders in the two groups have the same structure, and the segmentation heads in the two groups have the same structure but different initialization parameters; input the extracted random sequence of bands into the multi-level feature perception (MFA) encoder to obtain a multi-level feature sequence; input the multi-level feature sequence into the two feature aggregation triangle (FPT) decoders and segmentation heads respectively to obtain the mean of the learnable Gaussian distribution and variance Then, based on Gaussian distribution modeling, obtain the uncertainty prediction result U of each pixel point in the k-th class sub-map (k) .
[0014] Furthermore, after inputting the multi-level feature sequence into the two feature aggregation triangle (FPT) decoders respectively to obtain aggregation features, use the segmentation heads to reduce the dimension of the aggregation features to one channel respectively, and activate them with the Sigmoid activation function, then the extracted random sequence of bands is represented as the mean of the learnable Gaussian distribution Sum of variances Furthermore, based on the uncertainty prediction result U of each pixel point in the k-th category subgraph obtained by Gaussian distribution modeling (k) ;
[0015]
[0016] where, ò is obtained from the standard Gaussian distribution, and k represents the current k-th category subgraph;
[0017] Furthermore, the balanced mean squared error loss BMSE function and the weighted binary cross-entropy loss WBCE function are used to optimize the uncertainty prediction result, that is, to train the uncertainty perception module; the uncertainty prediction result refers to the Gaussian distribution prediction result U.
[0018] Furthermore, the balanced mean squared error loss BMSE is expressed as:
[0019]
[0020] where, the weight embedding M i is expressed as:
[0021] M i = αY i +(1 - α)(1 - Y i )
[0022] The calculation method of the weight α is as follows:
[0023]
[0024] where, represents the variance of the i-th pixel point, represents the predicted variance of the i-th pixel point, Y i represents the label of the i-th pixel point, represents the number of pixel points, and represent the negative sample and the positive sample in the k-th category subgraph respectively.
[0025] Furthermore, the weighted binary cross-entropy loss WBCE is expressed as:
[0026]
[0027] where, β t = t / T represents the asymptotic factor, t represents the current training times, T represents the total number of training times, and T = B / 3. where, B represents the number of spectral bands.
[0028] Balanced binary cross-entropy loss L BCE is expressed as:
[0029]
[0030] Among them, and respectively represent the prediction result of the \(i\)-th pixel point in the \(k\)-th category subgraph and the category label.
[0031] Furthermore, the spatial-spectral feature of the distorted sample in step 2 is represented as:
[0032] DS fus = DS i + w spa DS spa + w spe DS spe
[0033] Among them, w spa and w spe respectively represent the fusion weights of the spatial feature and the spectral feature, DS i represents the embedded data, DS spa represents the spectral feature, DS spe represents the spatial feature, DS fus represents the spatial-spectral feature of the distorted sample.
[0034] Furthermore, in step 3, the visual-language alignment driven by the supervised contrastive loss is used to obtain the text-image fusion feature, and the loss is the text-image fusion loss L CLIP :
[0035] L CLIP = λ2((1 - α)L coarse + αL fine )
[0036] Among them, λ2 represents the hyperparameter for balancing the alignment loss, and α represents the learnable parameter for controlling the sharing of coarse-grained and fine-grained text features;
[0037] The visual-semantic alignment loss L vsa represents L coarse or L fine
[0038] The visual-semantic alignment loss L vsa is represented as:
[0039]
[0040] Among them, P I (i) and N I (i) respectively represent the positive sample set and the negative sample set in the image feature I, P T (i) and N T(i) respectively represent the positive sample set and the negative sample set in the text feature T. The image features and text features of the same category are represented as positive samples, and the image features and text features of different categories are represented as negative samples. |P(i)| represents the number of positive samples, and |P(i)| = |P I (i)| = |P T (i)|, the superscript T represents the transpose operation, and τ represents the temperature hyperparameter.
[0041] When the visual-semantic alignment loss L vsa is L coarse , the above formula represents the coarse-grained visual-semantic alignment loss L coarse , and the text feature T in the formula represents the coarse-grained text feature;
[0042] When the visual-semantic alignment loss L vsa is L fine , the above formula represents the fine-grained visual-semantic alignment loss L fine , and the text feature T in the formula represents the fine-grained text feature;
[0043] Furthermore, in step 4, randomizing the distorted sample spatial-spectral feature and the text-image fusion feature respectively means performing style randomization on the distorted sample spatial-spectral feature DS fus to obtain the feature X E1 , and performing content randomization on the text-image fusion feature to obtain the feature X E2 .
[0044] Furthermore, in step 4, through the contribution metric mapping method, weights are dynamically assigned to the diversity and authenticity information, better balancing the diversity and authenticity information of the generated subdomains. The specific steps are as follows:
[0045] Construct a training set, which includes pixel points and their corresponding classification labels; divide the training set into N groups, with the number of samples in each group being N1. In the specific implementation, the number of samples in each group is N1 = 256, that is, 256 pixels are divided into a group, and traversing all groups is one round of training. The specific steps for training the classifier are as follows:
[0046] Step 4.1, for the feature X E2 in each group of the training set, in the first round of training, it is randomly linearly combined with X E1 to obtain the intermediate subdomain
[0047]
[0048] where, w0 represents the weight of the first linear combination in training;
[0049] Step 4.2, after inputting the intermediate sub-domain into the classifier, the obtained prediction result is while the feature X E2 after passing through the classifier, the obtained prediction result is expressed as then use the function v to map X E1 to its contribution:
[0050]
[0051] where y i represents the label of the i-th group, a contribution of 1 indicates a positive contribution, a contribution of -1 indicates a negative contribution, and i represents the i-th group of X E2 .
[0052] After having the contribution mapping, sort the categories by contribution, and let Π K represent the set of all permutations of the total number of categories K, that is, |Π K | = K!.
[0053] For the feature X E1 sorted by contribution category as π ∈ Π K , use to represent the set of all the first n groups of iterations of the feature in π, then for the marginal contribution of the n-th linear combination X E1 it is expressed as:
[0054]
[0055] where the value range of n is 2 - N, represents the feature X E1 in the n-th group of the training set;
[0056] Step 4.3: After combining with all N groups of features X E2 to obtain the corresponding marginal contribution of the feature X E1 , normalize it and take the mean to project it into the probability space to obtain the lazy weight w to complete one training;
[0057] Then, the intermediate sub-domain X ID obtained in the next training is:
[0058] X ID = wX E1 +(1 - w)X E2
[0059] Subsequently, perform the following contribution balance strategy: In the early stage of training, in order to enable the classification model to quickly learn authenticity information, use the vision-language fusion features to guide the generation of the intermediate sub-domain, improve the cross-scene recognition ability, and use the lazy weight w to balance X E1 , that is, at this time X ID= wX E1 +(1 - w)X E2 。
[0060] In the middle stage of training, in order to enable the classification model to learn diverse knowledge during the cross - domain process, explore more unseen domains, and thus improve the classification ability in difficult cross - domain scenarios, the lazy weight w is used to balance X at a decreasing interval E2 , forcing the model to learn more complex sample data, that is, at this time X ID = (1 - w)X E1 + wX E2 。
[0061] In the late stage of training, in order to integrate authenticity information and diversity information and further improve the classification ability of the model, the lazy weight w is alternately assigned at a fixed interval.
[0062] To relieve the learning pressure of the classification model, the intermediate sub - domain X ID is randomly linearly combined with the source - domain data X to obtain the mixed sub - domain X MIX :
[0063] X MIX = αX ID +(1 - α)X
[0064] where α is a random number randomly sampled from a uniform distribution from 0 to 1.
[0065] Finally, a classification operation is performed. The cross - entropy loss is used to measure the class consistency between X ID and X MIX :
[0066]
[0067] where N1 represents the number of samples in each training set, and Cls represents the classification operation.
[0068] Then, the cross - entropy loss between X E2 and X is increased:
[0069]
[0070] The cross - entropy loss L CE is overall expressed as:
[0071] L CE = L X + L E2 + L ID + L MIX
[0072] In addition, in order to capture the cross-domain consistency information of categories between the enhanced data at each stage and learn the category authenticity information, the KL divergence loss is used to measure the deviation degree of the category semantic information.
[0073] The KL divergence loss is expressed as:
[0074]
[0075] where p(y i |x i ) represents the predicted probability distribution obtained by processing the sample x i in X through the classifier.
[0076] The total KL divergence loss L KL is expressed as:
[0077]
[0078] In summary, the total classification loss L cls is expressed as:
[0079] L cls = λ3L KL + L CE
[0080] where λ3 represents the hyperparameter that balances the KL divergence loss and the cross-entropy loss.
[0081] Beneficial effects: Aiming at the problem that the existing methods use data augmentation methods to generate diverse intermediate subdomains for all data without selection, the present invention focuses on distorted samples and utilizes the diversity contained therein to generate intermediate subdomains with effective diversity to achieve full coverage of the feature distribution of the target domain. Aiming at the problem that the existing methods customize prior constraints for the domain to ensure the authenticity of the generated intermediate subdomains, the present invention customizes text description labels between a single category and different categories, and uses the text description labels that are more in line with human cognitive consistency as reliable guiding priors across scenarios with the help of a large-scale vision-language pre-training model, thereby ensuring the authenticity of each category in the intermediate subdomain. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 is the principle block diagram of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0083] The present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0084] The hyperspectral image classification method based on uncertainty perception and language guidance provided by the present invention has the following specific principle Figure 1 as shown. First, a global distortion sample group is extracted through uncertainty perception. Then, the spatial-spectral features of the global distortion samples are extracted using uncertainty learning encoding. At the same time, the text description label is embedded into the image features to obtain the visual-language fusion features, that is, Figure 1 the text-image fusion features in
[0085] Step 1: Uncertainty perception;
[0086] In uncertainty perception, first, the source domain data is split into class subgraphs according to the label category mask. Then, a random sequence of bands is extracted from the class subgraphs and input into a multi-level feature-aware (MFA) encoder to obtain a multi-level feature sequence. Next, these multi-level feature sequences are input into two feature aggregation triangles (FPT) decoders and a segmentation head with the same structure but different weights to obtain Gaussian distribution prediction results. Finally, a distortion sample mask is obtained from the prediction results, and then a global distortion sample data group is extracted from the source domain dataset.
[0087] Further, Step 1 also includes performing effective uncertainty perception on the source domain data and precisely quantifying it, and then selecting distortion samples from the source domain data. The specific steps include:
[0088] First, randomly sample 3 bands from the class subgraphs to form band random data and input it into the multi-level feature-aware encoder (MFA Encoder) to obtain a multi-level feature sequence where represents 1 / 32 of the original image, and represents 1 / 2 of the original image.
[0089] Next, the feature aggregation triangle FPT takes the multi-level feature sequence as input, and through progressive upsampling between different scales and short-range and long-range dense skip connections between the same-scale levels, full interaction between data is achieved. Convolution operations of different scales are performed on the concatenated data to integrate information from different-scale receptive fields and achieve progressive aggregation of multi-scale features.
[0090] After that, after two FPT Decoders with the same structure and shared weights obtain the aggregated features respectively, the segmentation head is used to reduce their dimensions to one channel respectively, and the Sigmoid activation function is used to activate them. Thus, the band random data is represented as the mean and variance The uncertainty prediction result U of each point obtained by Gaussian distribution modeling (k) :
[0091]
[0092] where, ò is obtained from the standard Gaussian distribution, and k represents the k-th category subgraph currently represents the mean of the learnable Gaussian distribution represents the standard deviation of the learnable Gaussian distribution
[0093] Next, the Gaussian distribution information is optimized by the balanced mean square error loss (BMSE) function and the weighted binary cross-entropy loss (WBCE) function. For the balanced mean square error loss BMSE, since the variance in the Gaussian distribution can be regarded as an index to measure the uncertainty of each pixel point, the variance can be further used to guide the training of the uncertainty perception module. Using the label set variance σ 2 of the category subgraph as the supervision of the prediction variance can be used to represent the degree of interference under different bands. However, for a single-category pixel that accounts for no more than 20% of the total number of image pixel points, its distribution is very sparse. Therefore, an adaptive weighting scheme is adopted to balance the loss, making the proposed uncertainty perception module more focused on category pixels. Since the number of positive samples is much smaller than the number of negative samples, the balanced weight assigns a higher weight to the uncertainty perception module for category pixels, and finally it is expressed as:
[0094]
[0095] where, the weight embedding M i is expressed as:
[0096] M i = αY i +(1 - α)(1 - Y i )
[0097] The calculation method of the weight α is as follows:
[0098]
[0099] where, represents the variance of the i-th pixel point represents the predicted variance of the i-th pixel point, Y i represents the label of the i-th pixel point represents the number of pixel points and Respectively represent the negative and positive samples in the k-th category sub-graph. The negative samples are represented as non-current category labels, and the positive samples are the current category labels.
[0100] For the weighted binary cross-entropy loss WBCE, separating distorted samples is regarded as a binary classification task, so the binary cross-entropy loss is used to optimize this task. Since the number of positive and negative samples in the data is unbalanced, weight embedding is added for balancing, and the balanced binary cross-entropy loss is expressed as:
[0101]
[0102] Among them, and respectively represent the prediction result and the i-th pixel point in the class label Y (k) , and k represents the k-th category sub-graph.
[0103] In the uncertainty map represented by the prediction result, points with higher uncertainty tend to be closer to the edge. Therefore, these pixels should be given priority during training to strengthen the attention of the uncertainty perception module to these edge-distorted samples. Therefore, in order to make samples with higher uncertainty in the prediction result have higher weights and prevent samples with higher uncertainty from confusing the model in the early stage of training, a progressive uncertainty-driven weighting strategy is proposed, where samples with higher uncertainty prediction results are gradually given higher weights, and finally expressed as:
[0104]
[0105] Among them, β t =t / T represents the progressive factor, t represents the current training iteration, T represents the total number of training iterations, and T = B / 3. Among them, B represents the number of spectral bands.
[0106] Step 2, learning of distorted samples;
[0107] In the learning of distorted samples, first, distorted samples are randomly sampled from the global distorted sample data group and batch-embedded to obtain embedded data. Then, the embedded data are respectively input into the spectral feature extraction block and the spatial feature extraction block to obtain the spectral features and spatial features of the distorted samples. Next, these two types of features are linearly combined and connected through a residual connection with the batch-embedded data. Finally, after pooling, the spatial-spectral features of the distorted samples are obtained.
[0108] First, the distorted sample data sampled from the global distorted data group is first processed by Embedding to obtain the pixel embedding DS i .
[0109] Then, DSi They are respectively input into the spectral feature extraction block and the spatial feature extraction block to process and obtain the spatial features DS of the distorted samples spa and the spectral features DS of the distorted samples spe .
[0110] Finally, DS spa and DS spe are weighted and fused, and after residual connection with DS i and pooling, the spatio-spectral features DS of the distorted samples are obtained fus . The fusion process can be expressed as:
[0111] DS fus = DS i + w spa DS spa + w spe DS spe
[0112] where w spa and w spe respectively represent the fusion weights of the spatial features and the spectral features, which are obtained by random initialization and updated through backpropagation
[0113] Step 3, text-image fusion learning
[0114] In text-image fusion learning, first, a batch of data is randomly sampled from the source domain data. Then, these data are input into the image encoder to obtain image features, and at the same time, the class text description labels related to these data are input into the text encoder to obtain text features. Finally, the text-image fusion features are obtained by driving with the supervised contrast loss
[0115] Specifically, first, the global HSI batch data X sampled from the source domain data is input into the image encoder, and the output is the image feature I
[0116] Then, the text description labels are tokenized through the vocabulary and input into the text encoder and projected into the semantic space to obtain the text feature T
[0117] Subsequently, in order to minimize the text-image fusion features of the same category and maximize the text-image fusion features of different categories in the semantic space, by jointly optimizing the coarse-grained visual-semantic alignment loss L coarse and the fine-grained visual-semantic alignment loss L fine the text-image fusion loss L CLIP is obtained:
[0118] L CLIP = λ2((1 - α)L coarse + αL fine )
[0119] Among them, λ2 represents the hyperparameter of the balanced alignment loss, α represents the learnable parameter used to control the sharing of coarse-grained and fine-grained text features, and the coarse-grained visual-semantic alignment loss L coarse can be expressed as:
[0120]
[0121] Among them, P I (i) and N I (i) respectively represent the positive sample set and the negative sample set in the image feature I, P T (i) and N T (i) respectively represent the positive sample set and the negative sample set in the coarse-grained text feature T. The image features and text features of the same category are represented as positive samples, and the image features and text features of different categories are represented as negative samples. |P(i)| represents the number of positive samples, and |P(i)| = |P I (i)| = |P T (i)|, T represents the transpose operation, and τ represents the temperature hyperparameter.
[0122] The fine-grained visual-semantic alignment loss L fine is similar to the setting of L coarse , only changing the coarse-grained text feature to the fine-grained text feature.
[0123] Finally, according to α, the image feature I and the text feature T are fused to obtain the visual-language fusion feature containing cross-domain real information
[0124] Step 4, weighted classification.
[0125] In weighted classification, first, the style of the distorted sample's spatial-spectral feature is randomized to obtain X E1 , and the content of the text-image fusion feature is randomized to obtain X E2 . Then, X E1 and X E2 are input into the Contribution Decision Balance Module (CDBM) for decision balance to generate intermediate sub-domains. Finally, these intermediate sub-domains are input into the classifier for classification.
[0126] Specifically, first, perform AdaIN-based style randomization on DS fus to obtain X E1 , and perform AdaIN-based content randomization on to obtain X E2 .
[0127] Then, perform the following contribution calculation steps:
[0128] Step 4.1: For each group of XE2 At the first time of training, it is randomly linearly combined with X E1 to obtain an intermediate sub-domain
[0129]
[0130] Step 4.2: The prediction result obtained by the classifier is while X E2 The prediction result obtained by the classifier is Then X can be mapped to its contribution by the function v E1 as follows:
[0131]
[0132] Among them, a contribution of 1 indicates a positive contribution, a contribution of -1 indicates a negative contribution, and i represents the i-th group of X E2 .
[0133] After the contribution mapping, sort by category contribution and let Π K represent the set of all permutations of the total number of categories K, that is, |Π K | = K!; For the contribution category sorting of the feature X E1 is π ∈ Π K , use to represent the set of all the first n group iterations of it in π. Then, for the n-th linear combination X E1 the marginal contribution can be expressed as:
[0134]
[0135] Step 4.3: After combining with all N groups of X E2 to obtain the corresponding marginal contribution of X E1 , normalize it and take the mean to project it into the probability space to obtain the lazy weight w. Then, the intermediate sub-domain X ID obtained in the next training is:
[0136] X ID = wX E1 + (1 - w)X E2
[0137] Subsequently, execute the following contribution balance strategy: In the early stage of training, in order to enable the model to quickly learn authenticity information, use visual-language fusion features to guide the generation of intermediate sub-domains and improve the cross-scene recognition ability. Use the lazy weight w to balance X E1 , that is, at this time X ID = wX E1 + (1 - w)X E2. In the middle stage of training, in order to enable the model to learn diverse knowledge during the cross-domain process, explore more unseen domains, and thus improve the cross-domain difficult scenario classification ability, the lazy weight w is used to balance X at a decreasing interval E2 , forcing the model to learn more complex sample data, that is, at this time X ID =(1 - w)X E1 +wX E2 . In the later stage of training, in order to integrate authenticity information and diversity information and further improve the classification ability of the model, the lazy weight w is alternately assigned at fixed intervals.
[0138] In this specific embodiment, the early stage of training is the first 30% of the rounds, the middle stage is the 30% - 70% of the rounds, and the later stage is the remaining rounds”;
[0139] To relieve the learning pressure of the model, the intermediate sub-domain and the source domain data X are randomly linearly combined to obtain the mixed sub-domain X MIX :
[0140] X MIX =αX ID +(1 - α)X
[0141] where α is a random number randomly sampled from a uniform distribution from 0 to 1.
[0142] Finally, a classification operation is performed. The cross-entropy loss is used to measure the class consistency between X ID and X MIX :
[0143]
[0144] where N1 represents the number of samples in each group, and Cls represents the classification operation.
[0145] Then, the cross-entropy loss between X E2 and X is increased:
[0146]
[0147] The cross-entropy loss L CE can be generally expressed as:
[0148] L CE =L X +L E2 +L ID +L MIX
[0149] In addition, in order to capture the cross-domain class consistency information between the augmented data at each stage to learn the class authenticity information, the KL divergence loss is used to measure the deviation degree of the class semantic information. The KL divergence loss can be expressed as:
[0150]
[0151] where p(y i | x i ) represents the predicted probability distribution obtained by processing the sample x in X i through the classifier.
[0152] The total KL divergence loss L KL can be expressed as:
[0153]
[0154] In summary, the total classification loss L cls can be expressed as:
[0155] L cls = λ3L KL + L CE
[0156] where λ3 represents the hyperparameter that balances the KL divergence loss and the cross-entropy loss.
[0157] Finally, a trained classifier is obtained, and the target domain hyperspectral data is input into the trained classifier for classification.
Claims
1. A hyperspectral image classification method based on uncertainty perception and language guidance, characterized in that Step 1: Split the source domain data of the hyperspectral image into category subgraphs according to the label category mask; Extract the band random sequence from the class subgraphs and input it into the uncertainty perception module to obtain the Gaussian distribution prediction results of each class subgraph where K represents the total number of class subgraphs; After obtaining the Gaussian distribution prediction result U, extract the category distortion sample mask, that is, obtain the corresponding single-category distortion sample mask, and take the union of all category distortion sample masks to obtain the global distortion sample mask. The position corresponding to the source domain data set is the global distortion data group, and then extract the global distortion sample data group from the source domain data set; Step 2: Randomly sample distortion samples from the distortion sample data group, perform batch masking on the distortion samples to obtain masked data; input the masked data into the spectral feature extraction block and the spatial feature extraction block respectively to obtain the spectral features and spatial features of the distortion samples, linearly combine the spectral features and the spatial features, and connect them with the masked data through residual connection; finally, obtain the spatial-spectral features of the distortion samples after pooling; Step 3: First, randomly sample data from the source domain data of the hyperspectral image, input the sampled data into the image encoder to obtain image features, and at the same time input the category text description label of the sampled data into the text encoder to obtain text features. Finally, obtain the text-image fusion features through visual-language alignment driven by supervised contrastive loss; Step 4: After randomizing the spatial-spectral features of the distortion samples and the text-image fusion features respectively, input them into the Contribution Decision Balance Module (CDBM) for decision balance to generate an intermediate sub-domain; finally, input the intermediate sub-domain into the classifier for classification.
2. The hyperspectral image classification method based on uncertainty perception and language guidance for domain generalization according to claim 1, characterized in that The uncertainty perception module includes a multi-level feature perception (MFA) encoder, two feature aggregation triangle (FPT) decoders, and two segmentation heads, where one feature aggregation triangle (FPT) decoder and one segmentation head form a group; The structures of the feature aggregation triangle (FPT) decoders in the two groups are the same, and the structures of the segmentation heads in the two groups are the same, but the initialization parameters are different; Input the extracted band random sequence into the multi-level feature-aware MFA encoder to obtain a multi-level feature sequence; input the multi-level feature sequence into two groups of feature aggregation triangular FPT decoders and a segmentation head respectively to obtain the mean of the learnable Gaussian distribution and variance Then, based on the Gaussian distribution modeling, obtain the uncertainty prediction result U of each pixel point in the k-th category subgraph (k) .
3. The method for hyperspectral image classification based on uncertainty perception and language guidance for domain generalization according to claim 2, wherein After the multi-level feature sequences are respectively input into two feature aggregation triangular FPT decoders to obtain aggregated features, the aggregated features are respectively reduced to one channel using a segmentation head and activated using a Sigmoid activation function, then the extracted band random sequence is represented as the mean of a learnable Gaussian distribution and variance Furthermore, the uncertainty prediction result U of each pixel point in the k-th class subgraph obtained by Gaussian distribution modeling (k) ; Among them, ò is obtained from the standard Gaussian distribution, and k represents the kth category subgraph at present.
4. The method for hyperspectral image classification based on uncertainty perception and language guidance for domain generalization according to claim 3, characterized in that, The balanced mean square error loss (BMSE) function and the weighted binary cross entropy loss (WBCE) function are used to optimize the uncertainty prediction result, that is, to train the uncertainty perception module.
5. The method for hyperspectral image classification based on uncertainty perception and language guidance for domain generalization according to claim 4, wherein The balanced mean square error loss (BMSE) is expressed as: Among them, the weight embedding M i is expressed as: M i = αY i + (1 - α)(1 - Y i ) The calculation method of the weight α is as follows: Among them, represents the variance of the i-th pixel point, represents the predicted variance of the i-th pixel point, Y i represents the label of the i-th pixel point, represents the number of pixel points, and respectively represent the negative sample and the positive sample in the k-th category sub-graph.
6. The method for hyperspectral image classification based on uncertainty perception and language guidance for domain generalization according to claim 4, wherein The weighted binary cross entropy loss (WBCE) is expressed as: Among them, β t = t / T represents the progressive factor, where t represents the current number of training times, T represents the total number of training times, and T = B / 3. Here, B represents the number of spectral bands; Balanced binary cross-entropy loss L BCE It is expressed as: Among them, and Y i (k) respectively represent the prediction result of the i-th pixel point in the k-th category sub-graph and the category label.
7. The method for hyperspectral image classification based on uncertainty perception and language guidance for domain generalization according to claim 1, characterized in that, The spatial-spectral features of the distortion samples in Step 2 are expressed as: DS fus = DS i + w spa DS spa + w spe DS spe Among them, w spa and w spe represent the fusion weights of the spatial feature and the spectral feature respectively, DS i represents the embedded data, DS spa represents the spectral feature, DS spe represents the spatial feature, DS fus represents the spatial-spectral feature of the distorted sample.
8. The method for hyperspectral image classification based on uncertainty perception and language guidance for domain generalization according to claim 1, wherein In step 3, the text-image fusion feature is obtained through vision-language alignment driven by the supervised contrastive loss, and the loss is the text-image fusion loss L CLIP : L CLIP = λ2((1 - α)L coarse + αL fine ) Among them, λ2 represents the hyperparameter of the balanced alignment loss, and α represents the learnable parameter used to control the sharing of coarse-grained and fine-grained text features; Adopt the visual semantic alignment loss L vsa Denote L coarse Or L fine Visual semantic alignment loss L vsa It is expressed as: Among them, P I (i) and N I (i) respectively represent the positive sample set and the negative sample set in the image feature I, P T (i) and N T (i) respectively represent the positive sample set and the negative sample set in the text feature T. The image features and text features of the same category are represented as positive samples, and the image features and text features of different categories are represented as negative samples. |P(i)| represents the number of positive samples, and |P(i)| = |P I (i)| = |P T (i)|, the superscript T represents the transpose operation, and τ represents the temperature hyperparameter; When the visual-semantic alignment loss L vsa = L coarse The above equation represents the coarse-grained visual-semantic alignment loss L coarse , and the text feature T in the equation represents the coarse-grained text feature; When the visual-semantic alignment loss L vsa = L fine at this time, the above formula represents the fine-grained visual-semantic alignment loss L fine , and the text feature T in the formula represents the fine-grained text feature.
9. The hyperspectral image classification method based on uncertainty perception and language guidance for domain generalization according to claim 1, characterized in that In step 4, randomizing the spatial-spectral features of the distorted samples and the fused features of the text images respectively means performing style-based randomization on the spatial-spectral features DS of the distorted samples fus to obtain feature X E1 , and performing content randomization on the fused features of the text images to obtain feature X E2 .
10. The hyperspectral image classification method based on uncertainty perception and language guidance for domain generalization according to claim 9, wherein, For the training of the classifier, the specific steps are as follows: Construct a training set, which includes pixel points and their corresponding classification labels; divide the training set into N groups, and the number of samples in each group is N1. In the specific implementation, the number of samples in each group is N1 = 256, that is, 256 pixels are divided into a group, and traversing all groups is one round of training; Step 4.1, for the feature X in each training set E2 , during the first round of training, it is randomly linearly combined with X E1 to obtain an intermediate sub-domain Among them, w0 represents the weight of the first linear combination in training; Step 4.2, input the intermediate sub-domain into the classifier, and the obtained prediction result is while the feature X E2 the prediction result obtained through the classifier is expressed as then use the function v to map X E1 to its contribution: Among them, y i represents the label of the i-th group. A contribution of 1 indicates a positive contribution, and a contribution of -1 indicates a negative contribution. i represents the i-th group of X E2 ; After having the contribution mapping, sort the categories by contribution and let Π K represent the set of all permutations of the total number of categories K, that is, |Π K | = K!; For feature X E1 The contribution categories are sorted as π ∈ Π K , using to represent the feature in the set of all the first n group iterations in π, then the marginal contribution of the nth linear combination X E1 is expressed as: wherein, the value range of n is 2 - N, denotes the feature X in the n-th training set E1 ; Step 4.3: After combining with all N groups of feature X E2 to obtain the corresponding marginal contribution of feature X E1 normalize it and project the mean value onto the probability space to obtain the lazy weight w to complete one training; Then the intermediate sub-domain X obtained during the next training ID : X ID = wX E1 + (1 - w)X E2 Subsequently, execute the following contribution balance strategy: In the early stage of training, X ID = wX E1 + (1 - w)X E2 ; In the middle of training, X ID = (1 - w)X E1 + wX E2 ; In the later stage of training, alternately assign the lazy weights w and 1 - w to X E1 and X E2 at fixed intervals; Randomly linearly combine the intermediate subdomain X ID with the source domain data X to obtain the mixed subdomain X MIX : X MIX = αX ID + (1 - α)X Among them, α is a random number randomly sampled from a uniform distribution from 0 to 1; Finally, perform a classification operation; use cross-entropy loss to measure the class consistency of X ID and X MIX : Among them, N1 represents the number of samples in each training set, and Cls represents the classification operation; Next, X is added E2 and the cross-entropy loss of X: Cross-entropy loss L CE It is generally expressed as: L CE = L X + L E2 + L ID + L MIX The KL divergence loss is used to measure the deviation degree of the category semantic information; The KL divergence loss is expressed as: Among them, p(y i |x i ) represents the predicted probability distribution obtained by processing the sample x i in X through the classifier; Total KL divergence loss L KL Is expressed as: In summary, the total classification loss L cls is expressed as: L cls = λ3L KL + L CE Among them, λ3 represents the hyperparameter that balances the KL divergence loss and the cross-entropy loss.