Text perception-based cross-domain small sample learning hyperspectral image classification method and system
By utilizing text-aware methods in cross-domain hyperspectral image classification, combined with a dual-branch transformer model and an adaptive strategy, the problem of existing methods failing to fully utilize text information is solved, and better model generalization performance and unseen category recognition capabilities are achieved.
Patent Information
- Application Number
- CN202411669356.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-11-21
AI Technical Summary
Existing cross-domain hyperspectral image classification methods fail to fully utilize text information and ignore language modality knowledge, resulting in insufficient model generalization performance, especially difficulties in identifying categories that have not been seen in the target domain.
A text-aware cross-domain small-sample learning method was designed. Through domain-agnostic prior semantic information description and a dual-branch transformer model, domain-invariant visual and textual spatial-spectral features were extracted. A text-aware spatial-spectral domain adaptation strategy was adopted to achieve visual and language alignment and reduce domain shift.
The model's ability to recognize unseen categories in the target domain is significantly improved, and the generalization performance of the model is enhanced. Experimental results show that it outperforms existing methods on multiple hyperspectral image datasets.
Smart Images

Figure CN119649213B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and particularly relates to a hyperspectral image classification method and system based on text perception cross-domain few-shot learning. BACKGROUND
[0002] Hyperspectral image (HSI) has unique advantages due to its ability to provide rich spectral information of ground objects, making it possible to identify ground objects with high precision. HSI has been widely used in many fields, such as environmental monitoring, resource management, military reconnaissance and urban planning. HSI classification is a key technology in HSI processing, which aims to assign all pixels in HSI to a specific class. However, the complexity and risk of the acquisition environment bring the small sample problem to HSI classification.
[0003] In recent years, it has been found that cross-domain information can help solve the problem of limited labeled samples, so the research and development of cross-domain HSI classification has become a current hotspot. In cross-domain HSI classification, the source domain contains sufficient labeled samples, while the labeled samples of the target domain are less. In addition, the classes in the source domain are called seen classes, and the classes that do not exist in the source domain in the target domain are called unseen classes. The core idea of cross-domain HSI classification is to learn the transfer knowledge of the source domain and train a model that can effectively generalize to the small sample labeled data of the target domain.
[0004] Few-shot learning (FSL) is a typical cross-domain learning method, which obtains a universal model by constructing a cross-domain task. Few-shot learning (FSL) is an effective solution to cross-domain hyperspectral image (HSI) classification. The principle behind this is that such a model trained from source domain knowledge can easily generalize to the target domain, requiring only a few labeled samples. FSL methods are mainly divided into two categories: meta-learning and metric learning. In the aspect of HSI processing, the current research mainly adopts the method of metric learning to realize cross-domain HSI classification. For example, deep few-shot learning (DFSL) and deep cross-domain few-shot learning (DCFSL) are both classic applications of FSL in cross-domain HSI classification, which learn the metric space in the source domain, and then use the metric space to classify the HSI pixels of the target domain.
[0005] A core goal of few-shot learning is to train a classifier on a source domain that can generalize to new classes based on a few labeled samples in the target domain. Since text information is considered independent of the data domain, many text-based few-shot learning methods have been proposed in the prior art to improve the generalization performance of the model. Most existing text-based few-shot learning methods first form a set of positive and negative sample pairs. Then, the goal of these methods is to align the features of the positive sample pairs while distinctly separating the negative sample pairs in the projection space. Under this guiding principle, researchers have actively developed various text-based few-shot learning methods to improve the efficiency of cross-scene hyperspectral image (HSI) classification. Although many text-based few-shot learning methods have been proposed in the prior art for remote sensing image scene classification or cross-scene HSI classification, the research of text-based few-shot learning methods in cross-domain HSI classification is still in its early stages. Current methods have not considered utilizing semantic information to assist in the visual feature learning of unseen classes in the target domain.
[0006] While FSL methods have made significant progress in cross-domain HSI classification, most FSL methods only focus on learning domain-invariant representations at the visual level, ignoring the importance of text information. Text information can provide rich prior knowledge of ground objects, which helps to achieve object recognition and classification. Moreover, this information is considered independent of the data domain. With the rapid development of deep learning and natural language processing techniques, multi-modal learning models, especially image-text models, have shown excellent performance in various image processing tasks, including few-shot transfer tasks. For example, the contrastive language-image pre-training model CLIP performs well in few-shot and zero-shot tasks. Recently, the language-aware domain generalization network (LDGnet) has also been successfully applied to HSI classification. However, the application of multi-modal learning-based models in cross-domain HSI classification is still in its early stages, and existing methods do not consider unseen classes in the target domain, failing to utilize their language modalities to learn domain-invariant visual representations.
[0007] Recently, multi-modal learning models, especially image-text models, have been proven to have many advantages in remote sensing image processing. More importantly, text information contains various prior knowledge about ground objects, which is considered independent of the data domain. However, existing FSL methods for hyperspectral images ignore the knowledge of the language modality, failing to utilize it to improve the generalization performance of the model.
[0008] Hyperspectral image classification is a key technology in the field of hyperspectral image processing, with the goal of classifying all pixels in a hyperspectral image into specific classes. Cross-domain information can effectively address the problem of limited labeled samples. Although cross-domain HSI classification algorithms have made some progress, existing methods still lack exploration of prior semantic information and have not fully utilized text knowledge to enhance the domain invariance of the model. SUMMARY
[0009] In order to solve the technical problems existing in the prior art, the present application proposes a text perception based cross-domain small sample learning hyperspectral image classification method and system, a domain independent prior semantic information description method is designed to effectively represent the intra-class relationship and cross-class relationship; the image-text model used can extract domain invariant spatial-spectral features to effectively extract cross-domain visual and text spatial-spectral features, and these features are used to realize the alignment of vision and language; a text perception spatial-spectral domain adaptive strategy is used to generalize the model to the unseen classes of the target domain, reduce the domain shift, and further improve the generalization performance of the model.
[0010] In the embodiment of the present application, a text perception based cross-domain small sample learning hyperspectral image classification method specifically comprises the following steps:
[0011] S1, a domain independent prior semantic information description method is designed, the secondary classification of the class names of the source domain and the target domain is performed, the secondary classification semantic information of the class names is embedded into the text template of each hyperspectral image pixel, and the text template is combined with the intra-class and inter-class relationship to adapt to different data distributions;
[0012] S2, a hyperspectral image classification model based on a double-branch transformer is constructed to realize the alignment of vision and language, obtain domain invariant visual representation, and extract cross-domain visual spatial-spectral features and text spatial-spectral features;
[0013] S3, a text perception spatial-spectral domain adaptive strategy is designed to improve the generalization ability of the hyperspectral image classification model based on the double-branch transformer;
[0014] S4, the total loss of the training phase is calculated, the hyperspectral image classification model based on the double-branch transformer and the conditional domain discriminator are jointly trained, and the total loss L of the hyperspectral image classification model is calculated as the sum of the feature classification contrast loss of the hyperspectral image classification model and the loss of the text perception spatial-spectral domain adaptive strategy; total defined as the sum of the feature classification contrast loss of the hyperspectral image classification model and the loss of the text perception spatial-spectral domain adaptive strategy;
[0015] S5, the hyperspectral image classification model based on the double-branch transformer after training is applied to classify the cross-domain hyperspectral image.
[0016] The embodiment of the present application also provides a text perception based cross-domain small sample learning hyperspectral image classification system, which is realized based on the aforementioned hyperspectral image classification method; comprising the following modules:
[0017] The information description module is used for describing the domain-agnostic prior semantic information, and the secondary classification semantic information of the category name is embedded into the text template of each hyperspectral image pixel by performing secondary classification on the category names of the source domain and the target domain, so that the text template combines the intra-class and inter-class relationships to adapt to different data distributions.
[0018] The hyperspectral image classification model based on the dual-branch transformer is used to realize the alignment of vision and language, obtain the domain-invariant visual representation, and extract the cross-domain visual-spectral features and text-spectral features.
[0019] The text-aware spectral domain adaptation strategy module is used to improve the generalization capability of the hyperspectral image classification model based on the dual-branch transformer.
[0020] The total loss calculation module is used for joint training of the hyperspectral image classification model based on the dual-branch transformer and the conditional domain discriminator, and the total loss L of the hyperspectral image classification model is calculated. total The total loss L is defined as the sum of the feature classification contrast loss of the hyperspectral image classification model and the loss of the text-aware spectral domain adaptation strategy.
[0021] The classification module applies the trained hyperspectral image classification model based on the dual-branch transformer to classify the cross-domain hyperspectral images.
[0022] From the above technical solutions, the present application proposes a novel text-aware few-shot learning (TAFSL) method for cross-domain hyperspectral image classification. Compared with the prior art, the present application achieves the following technical effects:
[0023] 1. The present application is the first to use an image-text model to process unseen classes in cross-domain HSI classification, fully utilizing the advantages of the language mode to assist visual representation learning; a field-independent prior semantic information description method is designed to effectively represent intra-class relationships and cross-class relationships.
[0024] 2. The present application uses a dual-stream transformer-based model as an image-text model, which can extract domain-invariant spectral features, i.e., spectral features for capturing vision and text across domains, and use these features to align vision and language to obtain domain-invariant visual representations.
[0025] 3. The present application proposes a text-aware spectral domain adaptation strategy to generalize the model to unseen classes in the target domain, which reduces the domain shift through adversarial learning. This strategy enables the model to generalize to unseen classes in the target domain, further improving the generalization performance of the model.
[0026] 4. The experimental results of cross-domain HSI classification performed on four hyperspectral image datasets show that the TAFSL proposed in the present application is more superior than the most advanced existing methods. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1 A flowchart of the text-aware cross-domain few-shot learning hyperspectral image classification method in the embodiments of the present application. DETAILED DESCRIPTION
[0028] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the accompanying drawings are exemplary and are only used to explain the present application and cannot be understood as a limitation of the present application.
[0029] Embodiments
[0030] In the text-aware cross-domain few-shot learning hyperspectral image classification method provided in the present embodiment, first, the domain prior semantic information of each class is given to effectively describe the intra-class and inter-class relationship; then, a hyperspectral image classification model based on a double-branch transformer is used to extract cross-domain visual and text spectral features, a visual language alignment technology is used to enhance the domain invariance of visual representation, and a domain-invariant visual representation is obtained; a text-aware spectral domain adaptive strategy is designed to further improve the generalization ability of the hyperspectral image classification model based on the double-branch transformer.
[0031] As shown in Figure 1 The hyperspectral image classification method of the present embodiment specifically includes the following steps:
[0032] S1, a domain-agnostic prior semantic information description method is designed, the class names of the source domain and the target domain are classified twice, the secondary classification semantic information of the class names is embedded into each hyperspectral image pixel text template, and the text template can adapt to different data distributions by combining intra-class and inter-class relationships.
[0033] In previous image-text models for hyperspectral images (HSI), the class name is mainly taken as the semantic information keyword of each HSI pixel. Specifically, these methods first construct a coarse-grained text template expressed as "hyperspectral image of <class name>", and then provide some fine-grained text information containing adjacency relations and other prior knowledge, such as "compressed grassland is located beside the road". However, this fine-grained prior semantic information description method has challenges in adapting to the distribution of unseen classes in different target domains. In addition, only using coarse-grained prior semantic information cannot express the relationship between classes, nor can it extract the domain invariance of text information from the source domain and the target domain.
[0034] To this end, the embodiment designs a domain-agnostic prior semantic information description method, which enables the text template to combine intra-class and inter-class (i.e., cross-class) relationships by performing secondary classification on the class names of the source domain and the target domain.
[0035] Specifically, this step mainly uses the text understanding and generation capabilities of a large language model (LLM) to enhance the association between seen classes and unseen classes. First, all class names of the source domain and the target domain are input into the large language model LLM at one time; the large language model LLM classifies the input class names into three categories based on prior geographical knowledge: "natural class", "man-made class", and "agricultural class" features. The man-made class features are also called artificial class features. Then, the secondary classification semantic information of the class names is embedded into the text template of each hyperspectral image HSI pixel, enabling the text template to effectively utilize intra-class and inter-class relationships. Taking "grassland" as an example, its corresponding text template can be expressed as "hyperspectral image of [natural class] grassland".
[0036] In the subsequent part, the text representation of each pixel (whether from the source domain or the target domain) is defined as X test , which is the basis for text spectral feature extraction and visual-linguistic domain discriminators.
[0037] S2, a hyperspectral image classification model based on a double-branch transformer is constructed to realize the alignment of vision and language, obtain domain-invariant visual representations, and extract cross-domain visual spectral features and text spectral features.
[0038] In the field of computer vision, traditional deep learning methods mainly use CNN modules to extract visual features. CNN modules can provide local receptive fields and have the advantage of parameter sharing. In recent years, transformers have shown extraordinary capabilities in text processing due to their ability to capture long-range dependencies and contextual relationships in data. The latest developments in applying transformers to image processing include the development of visual transformers, which can provide better performance than CNN-based models on multiple benchmark datasets. These advances highlight the potential of transformers to provide a more comprehensive and context-aware approach to understanding visual data, revolutionizing image processing. However, existing image-text models only use transformers for text feature extraction, ignoring the advantages of transformers in understanding global image context.
[0039] Therefore, in order to effectively capture the visual and textual spectral features of hyperspectral images, the embodiment constructs a novel hyperspectral image classification model based on dual-branch transformers, namely a dual-flow transformer architecture model, which can be used to align vision and language and obtain domain-invariant visual representations. The hyperspectral image classification model based on dual-branch transformers mainly includes a visual branch transformer encoder and a text branch transformer encoder.
[0040] In terms of the visual branch transformer encoder, first, a three-dimensional local block is segmented from the hyperspectral image to represent each pixel of the hyperspectral image; then, each three-dimensional local block is further divided into a center sequence X cen and a neighborhood sequence X nbr . Specifically, let X 3D be a three-dimensional local block of a hyperspectral image, then the segmentation function of the three-dimensional local block X 3D can be represented as follows:
[0041]
[0042] where ω represents the segmentation rate. Let the size of the three-dimensional local block X 3D be x1×y1×λ1, then the size of the center sequence X cen and the neighborhood sequence X nbr are x2×λ1 and x3×λ1, respectively, where x2+x3=x1×y1.
[0043] Then, the obtained center sequence X cen and neighborhood sequence X nbr as input of the visual branch transformer encoder.
[0044] The visual branch transformer encoder mainly includes a multi-head attention module and a feed-forward network (FFN). As for the multi-head attention module, it includes three inputs of query (i.e., Q), key (i.e., K) and value (i.e., V), as shown below:
[0045]
[0046] wherein d k represents a scale factor, the size of which is based on the dimension size of the query. As can be seen from formula (2), the attention mechanism of the multi-head attention module is realized by calculating the dot product between the query and the key, and the result of the dot product is further used as the weight of the value through the softmax function and the scale factor. In this embodiment, the center sequence X cen is selected as the query, and the neighborhood sequence X nbr is selected as the key and the value; by calculating the difference distribution of the local neighborhood, the neighborhood correlation can be fully utilized to obtain the three inputs of the multi-head attention module, which are specifically as follows:
[0047]
[0048] wherein, represents the weight matrix for calculating the query (Q) of the visual feature, represents the weight matrix for calculating the key (K) of the visual feature, and represents the weight matrix for calculating the value (V) of the visual feature.
[0049] Therefore, the single-head attention score with neighborhood correlation can be calculated as follows:
[0050] head i =Attention(Q′ i , K′ i , V′ i ) (6)
[0051] The multi-head attention is formed by concatenating the single-head attention, and the calculation of the multi-head attention score is as shown below:
[0052] MultiHead(Q′, K′, V′)=Concat(head1, head2, …, head h )w 0 (7)
[0053] Then, the multi-attention score is concatenated with the center sequence X cenThe local visual-spectral feature is input into a feed-forward network FFN to calculate the local visual-spectral feature (i.e., the local visual-spectral feature of the three-dimensional local block of the hyperspectral image). The feed-forward network FFN includes two fully connected layers, the activation function of the first fully connected layer is ReLU, and the second fully connected layer is a linear transformation layer. In this embodiment, the local visual-spectral feature is calculated by the following feature extraction function:
[0054] F(X 3D )=FFN(X nbr +MultiHead img )
[0055] =max{0,(((X nbr +MultiHead img )w1+b1)w2+b2)) (8)
[0056] where w1, b1 are the weights and bias of the first fully connected layer, w2, b2 are the weights and bias of the second fully connected layer; X nbr is the neighborhood sequence, and MultiHead img is the multi-head attention score.
[0057] Finally, the local visual-spectral feature in the feature metric is one-dimensional, so the following flattening transformation is needed to obtain the visual-spectral feature vector of the three-dimensional local block of the hyperspectral image:
[0058] f image (X 3D )=faltten(F image ) (9)
[0059] where F image is the local visual-spectral feature of the three-dimensional local block of the hyperspectral image calculated by formula (8), and f image (X 3D ) is the visual-spectral feature vector of the three-dimensional local block of the hyperspectral image after flattening transformation.
[0060] As for the text branch transformer encoder, some components are similar to those of the visual branch transformer encoder, but there are still some differences, which are described as follows:
[0061] First, the text is tokenized, and then encoded into a vector using methods such as Byte Pair Encoding (BPE) to convert it into a vector for the text branch transformer encoder model, which is defined as follows:
[0062] X emb =Embedding(X text ) (10)
[0063] where X emb is the encoded text vector. Then, position encodings are added to the encoded text vector:
[0064] X′ emb = X emb + position(X emb ) (11)
[0065] For the text attention mechanism, their Q′ text , K′ text and V′ text are linearly mapped from the text vector X′ emb with position encodings added, and the calculation method is similar to that of equations (3), (4), (5).
[0066]
[0067] where denotes the weight matrix of the query (Q) for computing the text feature, denotes the weight matrix of the key (K) for computing the text feature, denotes the weight matrix of the value (V) for computing the text feature.
[0068] In addition, their multi-head attention calculation method is the same as that of equations (6), (7), and can be represented as follows:
[0069] MultiHead text = MultiHead(Q′ text , K′ text , V′ text ) (15)
[0070] Unlike the visual branch transformer encoder, the text branch transformer encoder is mainly composed of multiple residual blocks (ResBlocks). The calculation method of a residual block is as follows:
[0071] resblock(X text ) = Ln(FFN(Ln(Q′ text + MuliHead text ))) (16)
[0072] where Ln() denotes layer normalization. Then, the text spectral feature vector f text (X text ) is obtained from the stacked residual blocks:
[0073] f text (X text) = Stack(resblock1, resblock2,..., resblock l ) (17)
[0074] In addition, the hyperspectral image classification model of the dual-branch transformer of the embodiment can be divided into two stages of updating: FSL and visual-linguistic alignment.
[0075] In the FSL stage, the data of each batch will be divided into a support set and a query set:
[0076]
[0077] where x i represents a support sample, y i represents the label of the support sample, S s represents a support set of K samples labeled with class C; x j represents a query sample, y j represents the label of the query sample, Q s represents a query set of K samples labeled with class C.
[0078] In the visual-linguistic alignment stage, the data of each batch will be displayed as the sum of the support set and the query set:
[0079]
[0080] where Img s represents all hyperspectral images of c classes in the current batch, each class containing K+N samples, x u represents an image sample of the u-th class, y u represents the label corresponding to the image sample of the u-th class.
[0081] In the training stage, the embodiment will calculate the distance metric between the text set and the image set, and the distance metric between the support set and the query set, respectively, to obtain the class probability distribution. According to the calculation result of the distance metric, the parameters of each network will be updated through the softmax function. The class probability distribution based on the distance metric between the support set and the query set can be defined as follows:
[0082]
[0083] Similarly, the class probability distribution based on the distance metric between the text set and the image set can be represented as follows:
[0084]
[0085] where d(·) represents the Euclidean distance metric function, C K represents the kth a category, text k is the k th th category feature encoding, denotes the visual branch transformer encoder with parameter λ, denotes the text branch transformer encoder with parameter β, x j is a query sample from the sample query set Q s or a hyperspectral image Img s , y j is the label of the query sample x j .
[0086] The category probability distribution based on the distance metric between the support set and the query set, the category probability distribution based on the distance metric between the text set and the image set, are used to further calculate the loss.
[0087] According to the negative logarithmic probability of the query sample with the true label, the support-query sample classification loss (also known as FSL loss) of each set can be represented as:
[0088]
[0089] Similarly, the image-text contrast loss calculated by the visual-linguistic alignment can be calculated as follows:
[0090]
[0091] Finally, the feature classification contrast loss L feat of the hyperspectral image classification model based on the dual-branch transformer is the sum of the support-query sample classification loss and the image-text contrast loss:
[0092]
[0093] In summary, in the process of visual and linguistic registration and alignment, the text feature is considered as a prototype, and the cross-entropy loss generated thereby is combined with the FSL loss, representing the multiple alignment of the empty spectral features in the semantic space.
[0094] S3, design a text-aware empty spectral domain adaptation strategy to improve the generalization ability of the hyperspectral image classification model based on the dual-branch transformer.
[0095] This embodiment defines the visual domain and the language domain, the source domain and the target domain. The designed text-aware empty spectral domain adaptation strategy includes the adversarial strategy between the source domain and the target domain, and the adversarial strategy between the visual domain and the language domain.
[0096] In order to enhance the perception of prior semantic information, the existing data distribution is adjusted through the spatial-spectral domain adaptation strategy, which helps to reduce the domain offset between the source domain and the target domain, as well as the domain offset between the visual domain and the language domain. The adversarial strategy between the source domain and the target domain is also used to improve the generalization ability of the target domain, while the adversarial strategy between the visual domain and the language domain adjusts the distribution of visual spatial-spectral features and the distribution of text spatial-spectral features, and uses the domain independence of text to help the visual branch transformer encoder improve its generalization ability.
[0097] S31, through the conditional domain discriminator D on the source distribution and target distribution s-t , reducing the offset of source and target domain feature extraction in the dual-branch transformer-based hyperspectral image classification model to set up an adversarial strategy between the source and target domains.
[0098] This embodiment proposes a method for source distribution P s (x) and the target distribution P t Conditional domain discriminator D of (x) s-t , which is used for adversarial training with a dual-branch transformer-based hyperspectral image classification model (also known as a dual-branch transformer feature extraction model) to reduce the offset of the dual-branch transformer feature extraction model in the source domain and target domain features.
[0099] In this step, the data of the source domain and the target domain are first taken as input, and the empty spectrum feature vector from the source domain dataset is obtained through the dual-branch transformer feature extraction model. And the prediction results of the classification labels of the input source domain dataset At the same time, the empty spectrum feature vector from the target domain dataset is also obtained And the prediction results of the classification labels of the input target domain dataset Among them, the prediction results of the classification labels By the empty spectrum characteristic vector The input classifier G (a learnable linear mapping layer) is obtained. Then, the prediction results of the empty spectrum feature vector and the classification label are used as the conditional domain discriminator D s-t input.
[0100] The prediction result g of the classification label contains discriminative information, which can condition the adversarial adaptability of the feature representation. The empty spectrum feature vector f (specifically including the empty spectrum feature vector from the source domain dataset) Null spectral feature vector from target domain dataset ) and the prediction results of the classification labels g (Specifically including the prediction results of the classification labels of the source domain dataset Prediction results of the classification labels of the target domain dataset ) can be simultaneously used in the conditional domain discriminator D s-t On modeling.
[0101] In this step, the conditional domain discriminator D s-t The input is the data of the source domain and the target domain, so it is also called the source domain target domain discriminator. The domain adversarial loss of the source domain target domain discriminator is calculated as follows:
[0102]
[0103] in represents a random variable According to the source domain data distribution P s (x) Expected value calculation during sampling, represents a random variable According to the target domain data distribution P t (x) Expected calculation during sampling; Receive the empty spectrum feature vector of the source domain data And the prediction results of the classification labels As input, the computation is done on samples from the source domain probability; Receive the empty spectrum feature vector of the target domain data and classification results As input, a sample from the target domain During the training process, the conditional domain discriminator D s-t Minimize the domain adversarial loss function, that is, the training goal of the conditional domain discriminator is to judge the correct result as much as possible (correctly judge whether the input data comes from the source domain or the target domain); and the dual-branch transformer feature extraction model and classifier G maximize it, that is, let the conditional domain discriminator D s-t By using this adversarial approach, we can train a visual branch transformer encoder that extracts approximate representations in different scenarios.
[0104] When the empty spectrum feature vector f and the predicted classification label g are simultaneously modeled as the domain decision information model h(f, g), a multilinear mapping is used for modeling. However, a multilinear mapping requires the outer product of f and g, which can lead to dimensionality explosion. Therefore, the following adjustment strategy is used to avoid dimensionality explosion:
[0105]
[0106] in, represents the element-wise product of f and g, d f and d gdimension size of f and g respectively, d represents the settable dimension size, which is set to 1024 in this example, and represents dot product, and represent two random matrices, which are sampled only once and fixed in the training phase, and d << d f xd g If the dimension of the linear mapping is greater than the maximum number of units of the deep network 1024, use The domain-invariant encoding feature loss from the source domain to the target domain is denoted as:
[0107]
[0108] wherein represents the domain judgment information model of the source domain, represents the domain judgment information model of the target domain, represents the domain judgment information model of the source domain passing through the dimension explosion prevention model, represents the domain judgment information model of the target domain passing through the dimension explosion prevention model, which is derived from formula (29).
[0109] So far, the setting of the adversarial strategy between the source domain and the target domain is completed.
[0110] S32, through the conditional domain discriminator D img-text about the visual domain data and the text domain data, reduce the deviation of the hyperspectral image classification model based on the double-branch transformer in extracting visual features and text features, to set the adversarial strategy between the visual domain and the language domain.
[0111] The embodiment proposes a conditional domain discriminator D img-text about the visual domain data and the text domain data, to better utilize the language information to improve the generalization ability of the transformer branch encoder.
[0112] In this step, first, the visual domain data and the text domain data are input, and the empty spectral feature vector f img from the visual domain data and the prediction result g img of the classification label of the input visual domain data set are obtained through the double-branch transformer feature extraction model; at the same time, the empty spectral feature vector f text from the text domain data and the prediction result g text of the classification label of the input text domain data are also obtained; wherein the prediction result g img of the classification label is obtained by inputting the empty spectral feature vector f text into the classifier G (which is a layer of learnable linear mapping layer). Then, the empty spectral feature vector and the prediction result of the classification label are input into the conditional domain discriminator D img-text .
[0113] In this step, the input of the conditional domain discriminator is data from the visual domain and text domain, so it is also called the visual language domain discriminator. The loss of the visual language domain discriminator is calculated as follows:
[0114]
[0115] in Represents the random variable Img i According to the image data distribution P img (x) Mean calculation during sampling, Represents the random variable text j According to the text data distribution P text (x) Mean calculation during sampling; D img-text (f img , g img ) Receive the spatial spectral feature vector f of the visual domain data img And the prediction result g of the classification label img , calculates the probability from the visual domain; while 1-D img-text (f text , g text ) Receive the empty spectrum feature vector f of the text domain data text And the prediction result g of the classification label text , computing the probability from the language domain.
[0116] Similarly, the spatial spectral feature vector f from the visual domain data img And the prediction result g of the classification label img Modeled to be dimensionality-explosion-resistant The empty spectrum feature vector f from the text domain data text And the prediction result g of the classification label text Modeled to be dimensionality-explosion-resistant The loss of the adversarial strategy from the visual domain to the language domain is expressed as follows:
[0117]
[0118] This completes the adversarial strategy between the visual and language domains. This adversarial strategy allows us to train a dual-branch transformer encoder that extracts domain-invariant feature vector representations across different modalities.
[0119] S33. Couple the adversarial strategy between the source domain and the target domain and the adversarial strategy between the visual domain and the language domain to obtain a text-aware spatial-spectral domain adaptation strategy.
[0120] Loss L for text-aware spatial-spectral domain adaptation adapt is the conditional domain discriminator Ds-t domain-adversarial loss of the domain pair img-text a weighted sum of the adversarial losses of the domains, specifically expressed as follows:
[0121] L adapt = L s-t + λL img-text (33)
[0122] The conditional domain discriminator D s-t , the conditional domain discriminator D img-text have the same structure, except that the input data is different. In the network setting of the corresponding conditional domain discriminator in the above strategy, after stacking multiple perceptron layers, a multi-layer linear mapping is applied. The first four operation layers each consist of a fully connected layer, a ReLU activation layer and a dropout layer; the last layer is a fully connected layer containing a softmax activation function to determine the input domain dependence.
[0123] S4, calculate the total loss of the training phase.
[0124] In TAFSL, the dual-branch transformer-based hyperspectral image classification model and the conditional domain discriminator are jointly trained. Combined with the above loss function, the total loss L total of the dual-branch transformer-based hyperspectral image classification model is defined as:
[0125] L totl = L feat + L adapt (34)
[0126] Wherein, L feat is the feature classification contrast loss of the dual-branch transformer-based hyperspectral image classification model calculated in step S2.
[0127] S5, apply the trained dual-branch transformer-based hyperspectral image classification model to classify cross-domain hyperspectral images.
[0128] This embodiment first switches the hyperspectral image classification model to evaluation mode, which will disable behaviors specific to the training phase to ensure the stability and consistency of performance measurement.
[0129] The three-dimensional local block of the hyperspectral image is input into the visual branch transformer encoder to obtain the spatial-spectral feature vector. The minimum and maximum values of the spatial-spectral feature vector are calculated, and then linear normalization operation is performed using these extreme values to map the spatial-spectral feature vector to the range of [0, 1].
[0130] Then, a KNN (K-Nearest Neighbors) classifier is initialized and trained using the normalized spectral feature vectors and the corresponding training labels. The classifier provides class predictions for the input test data through the nearest neighbor strategy. For the test data, the test features are first generated by the visual branch transformer encoder, and normalized based on the extreme value of the KNN classifier training stage; the class of the test data is predicted by the KNN classifier, and the prediction result is compared with the real label to generate a reward value sequence.
[0131] Finally, the reward values of the correct samples are accumulated, and the average accuracy is calculated based on the total number of test samples.
[0132] Based on the same inventive concept, the embodiment also provides a hyperspectral image classification system based on text-aware cross-domain small sample learning, which is realized based on the aforementioned hyperspectral image classification method; and specifically includes the following modules:
[0133] An information description module is configured to describe the domain-agnostic prior semantic information, embed the secondary classification semantic information of the class names into the text template of each hyperspectral image pixel by performing secondary classification on the class names of the source domain and the target domain, and make the text template adapt to different data distributions by combining the intra-class and inter-class relationships;
[0134] A hyperspectral image classification model based on a dual-branch transformer is configured to realize the alignment of vision and language, obtain domain-invariant visual representations, and extract cross-domain visual and text spectral features.
[0135] A text-aware spectral domain adaptation strategy module is configured to improve the generalization ability of the hyperspectral image classification model based on the dual-branch transformer.
[0136] A total loss calculation module is configured to jointly train the hyperspectral image classification model based on the dual-branch transformer and the conditional domain discriminator, calculate the total loss L total defined as the sum of the feature classification contrast loss of the hyperspectral image classification model and the loss of the text-aware spectral domain adaptation strategy.
[0137] A classification module is configured to apply the trained hyperspectral image classification model based on the dual-branch transformer to classify the cross-domain hyperspectral images.
[0138] Each of the above modules is used to realize each step of the aforementioned hyperspectral image classification method, and the detailed implementation process is described with reference to steps S1-S5.
[0139] Experiments were conducted on a computer equipped with a 2.1 GHz Intel Xeon Silver 4130 processor, 32 GB DDR4 memory, and a NVIDIA GeForce RTX 4090 graphics processing unit (GPU). The training and testing experiments used the open-source PyTorch framework.
[0140] To verify the effectiveness of the method of the embodiment, the TAFSL method is compared with four supervised methods (support vector machine (SVM), spectral-spatial residual network (SSRN), etc.), and six FSL methods (DFSL+SVM, DCFSL, CMFSL, Gia-FSL, MRLF and HyMuT). As for the FSL method, the cross-domain duration strategy is adopted in the training stage. Specifically, 200 labeled samples are selected from the source domain for each class to obtain transferable knowledge, and a limited number of target domain labeled samples are used for training. The remaining samples of the target domain dataset are left for testing.
[0141] In the experiment, the spatial size of the three-dimensional local tensor of the astrolabe is defined as 9x9, and the spectral size e of S0 is set to 128. The open-minded initialization method is used for parameter initialization. In addition, when extracting text features, the text encoder is initialized as a model with 33 million parameters, three layers, and a width of 512, with eight attention heads. It uses a transformer architecture similar to CLIP, encodes text in lowercase bytes (BPE), and has a vocabulary of 49,152. To maintain computational efficiency, the maximum sequence length is limited to 76. The overall accuracy, average accuracy (mean accuracy), and kappa coefficient (kappa coefficient) are used to evaluate the classification performance of all methods in the embodiment.
[0142] Hyperparameter experiment: The parameter sensitivity analysis is performed in the embodiment to evaluate the sensitivity of TAFSL to the change of the regularization parameter λ of the visual language domain adaptive discriminator. The basic regularization parameter λ is regarded as an adjustable hyperparameter, which is selected from the set {0.9, 0.95, 1.0, 1.05, 1.1}. After adjusting the weight λ, the gradient of the loss function is used to estimate the model weight hyperparameter during the descent process. Table 1 lists the classification results corresponding to different basic learning rates for the four datasets. For the four datasets, the weight λ of 1.0 is the most ideal.
[0143] Table 1 Overall accuracy of parameters λ for classification of four datasets under 5-SHOT
[0144]
[0145] Ablation experiments: The domain discriminator is a key component of TAFSL, and an important strategy to enhance the model's domain generalization ability is to use a combination of strong text and visual language domain discriminators. Pruning analysis is used to evaluate the important contribution of TAFSL components by removing each component from the entire framework. There are three variants of ablation analysis:
[0146] (1) "TAFSL (remove text)": remove the visual language alignment and visual language domain adaptation discriminators from FAFSL;
[0147] (2) "TAFSL (remove image-text modality discriminator)": remove the visual language domain adaptation discriminator from FAFSL;
[0148] (3) "TAFSL (remove prior knowledge text)": remove the prior text information used to enhance the domain discriminant ability of the domain discriminator from FAFSL.
[0149] Table 2
[0150]
[0151] As can be seen from Table 2, the TAFSL method of the present embodiment is significantly better than the existing variants, which marks a significant progress. Specifically, compared with the TAFSL (remove text) baseline, the integration of the visual language alignment architecture improves the TAFSL (remove image-text modality discriminator) by about 3%. This highlights the positive impact of language patterns on enhancing visual representation learning. In addition, compared with TAFSL (remove image-text modality discriminator), TAFSL improves by more than 0.6%, which shows that the visual language domain discrimination strategy effectively utilizes the domain agnostic of text, thereby enhancing the cross-domain generalization ability of the visual branch transformer model. In addition, compared with TAFSL (remove prior knowledge text), TAFSL improves by nearly 0.5%, which shows that the land cover type classification effectively establishes the relationship within and between classes, thereby improving the cross-domain generalization ability of the dual-branch transformer model.
[0152] The TAFSL proposed in the embodiment combines visual and language modalities together, and is performed under the condition of small amount of learning; a double-branch transformer model is used to simultaneously extract visual and language features; text enhanced by twice classification of category names is regarded as shared information irrelevant to the domain, aiming to reduce the domain shift between the source domain and the target domain. Specifically, the text features obtained by the text encoder are regarded as prototypes, and the distance between the query and the text is calculated. By minimizing the distance measure, the gap between the visual features and the language features will be narrowed. In addition, the embodiment also designs the adaptation of the source-target domain and the visual-language domain, and obtains the cross-domain invariant visual feature representation by using adversarial learning. Finally, the comprehensive experiments on four data sets verify the effectiveness of the proposed TAFSL in the aspect of domain generalization.
[0153] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and purposes of the present application, and the scope of the present application is defined by the claims and their equivalents.
Claims
1. A text perception-based cross-domain small sample learning hyperspectral image classification method, characterized in that, The method comprises the following steps: S1, designing a domain-agnostic prior semantic information description method, embedding the secondary classification semantic information of the class names into the text template of each hyperspectral image pixel by twice classifying the class names of the source domain and the target domain, so that the text template combines the intra-class and inter-class relationships to adapt to different data distributions; S2, constructing a hyperspectral image classification model based on a double-branch transformer, for realizing the alignment of vision and language, obtaining domain-invariant visual representation, and extracting cross-domain visual and text spectral features; S3, designing a text-aware spectral domain adaptation strategy to improve the generalization ability of the hyperspectral image classification model based on the double-branch transformer; S4, calculate the total loss of the training stage, jointly train the hyperspectral image classification model based on the double-branch transformer and the conditional domain discriminator, and calculate the total loss L of the hyperspectral image classification model total defined as the sum of the feature classification contrast loss of the hyperspectral image classification model and the loss of the text-aware empty spectral domain adaptive strategy; S5, applying the trained hyperspectral image classification model based on the double-branch transformer to classify the cross-domain hyperspectral images; Step S1 utilizes the text understanding and generation capabilities of the large language model LLM to enhance the association between seen classes and unseen classes; The text-aware spectral domain adaptation strategy designed in step S3 includes an adversarial strategy between the source domain and the target domain, and an adversarial strategy between the visual domain and the language domain; Step S3 comprises: S31, discriminating by a conditional domain discriminator D about the source distribution and the target distribution s-t , reduce the offset of the hyperspectral image classification model based on the double-branch transformer for source domain and target domain feature extraction, to set the confrontation strategy between the source domain and the target domain; The data of the source domain and the target domain are taken as inputs, and through a double-branch transformer feature extraction model, an empty spectrum feature vector from a source domain dataset is obtained and a prediction result of a classification label of the input source domain dataset Meanwhile, an empty spectrum feature vector from a target domain dataset is also obtained and a prediction result of a classification label of the input target domain dataset Wherein, the prediction result of the classification label is obtained by inputting the empty spectrum feature vector into a classifier G; then, the empty spectrum feature vector and the prediction result of the classification label are taken as inputs of a conditional domain discriminator D s-t S32, determining the condition domain of the visual domain data and the text domain data by the condition domain discriminator D img-text , reduce the deviation of the hyperspectral image classification model based on the double-branch transformer in extracting visual features and text features, to set the confrontation strategy between the visual domain and the language domain; S33, coupling the adversarial strategy between the source domain and the target domain and the adversarial strategy between the visual domain and the language domain to obtain a text-aware spectral domain adaptation strategy; taking a weighted sum of the domain adversarial loss of the conditional domain discriminator D s-t and the adversarial loss of the conditional domain discriminator D img-text as the loss of the text-aware spectral domain adaptation.
2. The method of claim 1, wherein, Step S1 comprises: All class names of the source domain and the target domain are input into the large language model LLM at one time; The large language model LLM divides the input class names into "natural classes", "man-made classes", and "agricultural classes" based on prior geographical knowledge; The secondary classification semantic information of the class names is embedded into the text template of each hyperspectral image pixel, so that the text template effectively utilizes the intra-class and inter-class relationships.
3. The method of claim 1, wherein, The hyperspectral image classification model based on the double-branch transformer constructed in step S2 comprises a visual branch transformer encoder and a text branch transformer encoder; The visual branch transformer encoder segments three-dimensional local patches from the hyperspectral image to represent each hyperspectral image pixel; each three-dimensional local patch is further divided into a center sequence X cen and a neighborhood sequence X nbr as input to the visual branch transformer encoder; The text branch transformer encoder tokenizes the text and encodes the text to convert it into a vector; position encoding is added to the encoded text vector.
4. The method of claim 3, wherein, The visual branch transformer encoder comprises a multi-head attention module and a feed-forward network FFN; select center sequence X cen for query, neighborhood sequence X nbr for key and value; three inputs of multi-head attention module are obtained by calculating the difference distribution of local neighborhood, using neighborhood correlation The multi-attention scores are combined with the center sequence X cen An input feedforward network FFN is applied to compute local visual-spectral features of a three-dimensional local patch of the hyperspectral image.
5. The method of hyperspectral image classification according to claim 4, characterized in that, The feed-forward network FFN comprises two fully connected layers, the activation function of the first fully connected layer is ReLU, and the second fully connected layer is a linear transformation layer; The local visual spectral feature is calculated by the following feature extraction function: F(X 3D ) = FFN(X nbr + MultiHead img ) = max{0, (((X nbr + MultiHead img )w1+b1)w2+b2)} wherein w1, b1 are the weights and bias of the first fully connected layer, w2, b2 are the weights and bias of the second fully connected layer; X nbr is the neighborhood sequence, MultiHead img is the multi-head attention score.
6. The method of claim 4, wherein, The local visual-spectral feature is flattened and transformed to obtain a visual-spectral feature vector f of a three-dimensional local block of the hyperspectral image image (X 3D ): f image (X 3D )=faltten(F image ) where F image is the local visual spectral feature.
7. The method of claim 1, wherein, The hyperspectral image classification model based on the double-branch transformer is updated in two stages of FSL and visual language alignment; In the FSL stage, the data of each batch will be divided into a support set and a query set; where x i represents support samples, y i represents labels of support samples, S s represents a set of K samples supporting class C; x j represents query samples, y j represents labels of query samples, Q s represents a set of N samples querying class C; In the visual language alignment stage, the data of each batch will be displayed as the sum of the support set and the query set: wherein, Img s represents all hyperspectral images of C classes in the current batch, each class contains K+N samples, x u represents the image sample of the u-th class, y u represents the label corresponding to the image sample of the u-th class.
8. The method of claim 1, wherein, Step S32 further comprises: The visual domain data and the text domain data are input, and the hyperspectral feature vector f from the visual domain data is obtained through a double-branch transformer feature extraction model img and the prediction result g of the classification label of the input visual domain data set img ; meanwhile, the hyperspectral feature vector f from the text domain data is also obtained text and the prediction result g of the classification label of the input text domain data text ; wherein the prediction result g of the classification label img is obtained by inputting the hyperspectral feature vector f text into the classifier G; then, the hyperspectral feature vector and the prediction result of the classification label are input into the conditional domain discriminator D img-text .
9. A text-aware based cross-domain few-shot learning hyperspectral image classification system, characterized in that, The hyperspectral image classification method based on any one of claims 1-8 is realized; It comprises the following modules: The information description module is used for describing the domain-agnostic prior semantic information, and the secondary classification semantic information of the class names of the source domain and the target domain is embedded into the text template of each hyperspectral image pixel by performing secondary classification on the class names, so that the text template combines the intra-class and inter-class relationships to adapt to different data distributions. The hyperspectral image classification model based on the dual-branch transformer is used for realizing alignment of vision and language, obtaining domain-invariant visual representation, and extracting cross-domain visual and text spectral features. The text-aware spectral domain adaptive strategy module is used for improving the generalization ability of the hyperspectral image classification model based on the dual-branch transformer. A total loss calculation module, based on joint training of a hyperspectral image classification model based on a double-branch transformer and a conditional domain discriminator, calculates a total loss L total defined as the sum of a feature classification contrast loss of the hyperspectral image classification model and a loss of a text-aware empty-spectrum domain adaptive strategy; The classification module applies the hyperspectral image classification model based on the dual-branch transformer after training to classify the cross-domain hyperspectral images.
Citation Information
Patent Citations
Hyperspectral image classification method combining spatial spectral domain self-adaption and ensemble learning
CN115393719A
Transform-based cross-domain double-branch adversarial domain adaptive image classification method
CN116740434A