Cross-scene hyperspectral image classification method based on multi-modal domain generalization

By constructing an optimized semantic space and combining image and text features, the problem of poor model adaptability in cross-scene hyperspectral image classification is solved, and higher classification accuracy and model stability are achieved.

CN119992324AActive Publication Date: 2025-05-13HARBIN NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510062655.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

In cross-scene hyperspectral image classification, it is difficult for the prior art to effectively utilize multimodal correlation information between images and text, resulting in a degradation in the performance of the model when it is generalized to other scenarios.

Method used

By building an optimized semantic space, combining image encoder and text encoder, the features of images and text are extracted, and deep interactive fusion is carried out through a multimodal fusion network, supervised contrast losses are optimized to reduce cross-domain differences.

Benefits of technology

The accuracy of cross-scene hyperspectral image classification and the adaptability and stability of the model are significantly improved, and a richer and more diverse image feature extraction and cross-domain invariant representation learning are achieved through multimodal fusion network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992324A_ABST
    Figure CN119992324A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-scene hyperspectral image classification method based on multi-modal domain generalization. The method comprises the following steps: step 1, data preparation; step 2, image feature extraction; step 3, constructing an optimized semantic space; step 4, calculating loss and optimizing the model; 5, testing the model; according to the method, the classification algorithm is optimized by using the multi-modal fusion network, and the network enables the two feature extraction modules to cooperatively run in parallel, so that richer and multivariate image features can be mined; an optimized semantic space is constructed by processing prior text knowledge of each category, so that image features can be accurately aligned according to the categories, and the classification accuracy is remarkably improved; according to the method, the semantic space is optimized, deep interactive fusion of two modes of images and texts is realized, learning of cross-domain invariant representation is promoted, and through targeted optimization of supervised contrast loss, features are prompted to be aligned according to classes in the semantic space, so that cross-domain differences are further reduced, and adaptability and stability of the model are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a cross-scene hyperspectral image classification method based on multimodal domain generalization. Background Art

[0002] In the field of hyperspectral image classification, deep learning technology, especially convolutional neural networks, has made significant progress and can extract and apply implicit information in hyperspectral images in environmental monitoring, medical imaging, agriculture, geological exploration and other fields. However, in cross-scene hyperspectral image classification, since training samples and test samples often come from different scenes and have different spectral and spatial characteristics, domain shift or distribution mismatch occurs, making it difficult to generalize the model trained in a single scene to other scenes, which seriously affects its performance.

[0003] In the existing technology, transfer learning is generally used to solve the problem of poor adaptability of deep learning in cross-scene hyperspectral image classification, and domain adaptation in transfer learning is the most commonly used method. For example, feature adaptation aligns source and target domain samples to a unified feature space by mapping to ensure alignment between different domains. Deep adaptation networks use the maximum mean difference to alleviate domain differences, while deep adversarial neural networks have pioneered adversarial technology to solve domain adaptation problems. It focuses on allowing the training network to extract similar features from the source and target domains. However, the above methods all focus on feature alignment and do not fully explore and utilize other potential related information. Some transfer learning technologies use domain generalization methods. Domain generalization currently involves data enhancement, representation learning, and the formulation of learning strategies. Among them, representation learning is the most important method. Representation learning mainly considers how to train the model to reduce the representation differences of one or more source domains, but it does not consider the role of language features in promoting visual representation learning. Moreover, for hyperspectral images, since they themselves lack textual description information that can reflect feature categories, how to introduce text and achieve efficient image and text fusion becomes a difficult problem that needs to be explored urgently. If this problem is not solved, it will hinder the accuracy and comprehensiveness of hyperspectral image classification tasks in cross-scenario applications. Summary of the invention

[0004] The purpose of the present invention is to provide a cross-scene hyperspectral image classification method based on multimodal domain generalization to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solution: a cross-scene hyperspectral image classification method based on multimodal domain generalization, comprising the following steps: step one, data preparation; step two, image feature extraction; step three, construction of optimized semantic space; step four, calculation of loss and optimization of model; step five, model testing;

[0006] In the above step 1, the hyperspectral image in the data set is cut into image blocks and sent to the image encoder to construct text prior knowledge for the data set to form a knowledge base;

[0007] In the above step 2, the following steps are specifically included:

[0008] 2.1 Extracting spatial and spectral features: In the image encoder, the SSEM based on three-dimensional convolution extracts the spatial and spectral features of the hyperspectral image block through the three-dimensional spectral information;

[0009] 2.2 Extracting multi-scale features: In the image encoder, the MSEM based on the two-dimensional residual pyramid structure is used to extract the multi-scale features of the hyperspectral image blocks through the Conv2d-BN2d module, three pyramid residual modules and average pooling layer;

[0010] 2.3 Feature concatenation: The features extracted by SSEM and MSEM are concatenated and sent to the classification head and the optimized semantic space;

[0011] In the above step 3, the following steps are specifically included:

[0012] 3.1 Text processing: The optimized semantic space construction module obtains text prior knowledge from the knowledge base according to the label, and then encodes the text prior knowledge through the text encoder;

[0013] 3.2 Extracting text features: The text encoder uses the Transformer architecture to extract text features from the encoded text prior knowledge;

[0014] 3.3 Constructing optimized semantic space and projecting image features: The optimized semantic space construction module constructs the optimized semantic space using the extracted text features, and projects the image features extracted by the image encoder into the optimized semantic space;

[0015] In the above step 4, the following steps are specifically included:

[0016] 4.1 Calculate classification loss: Input the image features extracted by the image encoder into the classification head, which outputs the classification probability and uses the true value to calculate the cross entropy loss of supervised learning and obtain the classification loss of the network;

[0017] 4.2 Calculate contrast loss: Calculate contrast loss based on the similarity between image features and semantic space features;

[0018] 4.3 Calculating the overall loss: the overall loss during model training Defined as:

[0019]

[0020] in, is the classification loss, is the contrast loss, λ is a hyperparameter;

[0021] 4.4 Model parameter optimization: The gradient of the model parameters is calculated based on the overall loss. Based on the gradient of the model parameters, the Adam optimizer is used to optimize the image encoder and text encoder.

[0022] In the above step five, SSEM and MSEM are used to extract image features and classify them, and the performance of the model optimized in step four in the target domain is evaluated based on OA and KC.

[0023] Preferably, in step 1, the image block size is h×w×b, wherein h and w are both the patch sizes of the image block, and b represents the number of spectral channels.

[0024] Preferably, in step 1, for the knowledge base Each text prior knowledge t i Both with tags One-to-one correspondence, text prior knowledge describes the spatial relationship and image features of the category corresponding to the label. Given a label The following mapping relationship can be used Find the corresponding text t i :

[0025]

[0026] Preferably, in step 2.1, SSEM uses a serial three-dimensional convolution residual block, a maximum pooling layer and a three-dimensional convolution to deeply extract information, the three-dimensional convolution residual block consists of two Conv3d-BN3d-ReLU modules and one Conv3d module, and the three-dimensional convolution formula is:

[0027]

[0028] Among them, h, v, and c represent the length, width, and number of channels of the three-dimensional convolution respectively. represents the value of the neuron at the jth feature sample (x, y, z) in the i-th layer network, y represents the weight of the (i-1)th feature map in the (h, v, c) convolution kernel, and b ij Indicates network bias.

[0029] Preferably, in step 2.2, the three pyramid residual modules P i In (i=1, 2, 3), P1 contains three residual modules with the same structure. Each residual module contains three BN2d-Conv2d blocks and one BN2d-ReLU block. The kernel size of the second convolution layer is 7, and the kernel size of the rest is 3. The first residual module in P2 and P3 The kernel size of the second convolutional layer is 8, the stride is 2, and the rest of the structure is the same as P1. In order to systematically increase the depth of the feature maps of each layer, MSEM adopts the following rules:

[0030]

[0031] Where A is the initial number of channels, is the depth of the jth unit in the i-th pyramid residual module, N (net) is the total number of residual units in the entire network. At this time, the depth of each layer of the feature map depends on A and α.

[0032] Preferably, in step 3.1, the encoding formula is as follows:

[0033] T=Embed[BPE(T)]+PE

[0034] Among them, BPE is lowercase byte pair encoding, the vocabulary size is 49152, Embed is the token embedding operation, and PE is the position encoding.

[0035] Preferably, in step 3.2, the Transformer architecture uses a multi-head self-attention mechanism to extract features and tokenize the text T∈R S×B×D is divided into h parts, each of which is set to Among them, B represents the batch size, S represents the maximum sentence length, D represents the embedding dimension size, and h represents the number of heads. After linear transformation, we get Q, K and V. The specific calculation process is as follows:

[0036]

[0037]

[0038] Where W q , W k , W v All are learnable matrices;

[0039] Then calculate the attention weight matrix for each head and merge the attention matrices of all heads:

[0040]

[0041] Attention(T)∈R S×B×D =Concat(A1, A2, …, A h )

[0042] Among them, Softmax is used to perform Softmax operation on each data sequence, and Concat merges the attention matrix of each data head into one attention matrix;

[0043] The equation of the entire feature extraction process is simplified as follows:

[0044] T∈R S×B×D =LayerNorm1(Attention(T)+T)

[0045] T∈R S×B×D =LayerNorm2(FFN(T)+T)

[0046] Among them, LayerNorm1 and LayerNorm2 are layer normalization operations, FFN is a feed-forward neural network, the above feature extraction process involves only one layer, the number of repetitions is the same as the number of layers, and the Transformer architecture has 3 layers and 8 heads, with a width of 512.

[0047] Preferably, in step 3.3, the specific formula is as follows:

[0048] Z oss =T·w oss +b oss

[0049] Z ssem =Linear[f ssem (X)]

[0050] Z msem =Limear[f msem (X)]

[0051] Among them, w oss is the learnable matrix, b oss is the bias term, and the linear layer is responsible for mapping image features to the semantic space.

[0052] Preferably, in step 4.1, the cross entropy loss The formula is:

[0053]

[0054] Classification loss of the network The formula is:

[0055]

[0056] Among them, y i is x i One-hot encoding of category information, P i and C(x i ) is the predicted probability, and N is the number of samples.

[0057] Preferably, in step 4.2, the similarity measurement formula is:

[0058]

[0059] In order to highlight the key role of similarity, Sim is squared and the optimized supervised contrast loss is:

[0060]

[0061]

[0062] Among them, P oss (i) Corresponding image features, N oss (i) do not correspond to image features, namely positive features and negative features, and represents one of the positive and negative samples, |P(i)|=|P oss (i)|, α is the contribution weight.

[0063] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention utilizes a multimodal fusion network to optimize the classification algorithm. The network operates two feature extraction modules in a coordinated and parallel manner, and can mine richer and more diverse image features; by processing the prior text knowledge of each category to construct an optimized semantic space, image features can be accurately aligned by category, thereby significantly improving the accuracy of classification; the optimized semantic space realizes the deep interactive fusion of the two modalities of image and text, promotes the learning of cross-domain invariant representations, and through targeted optimization of supervised contrast loss, promotes the alignment of features by category in the semantic space, further reduces cross-domain differences, and enhances the adaptability and stability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is the overall training framework diagram of the image classification method of the present invention;

[0065] Figure 2 is a flow chart of the image encoder SSEM;

[0066] Figure 3 This is a schematic diagram of the pyramid bottleneck residual structure;

[0067] Figure 4 is a flowchart of the image encoder MSEM;

[0068] Figure 5 This is a pseudo-color image of the Houston dataset;

[0069] Figure 6 This is a pseudo-color image of the Pavia dataset;

[0070] Figure 7 This is a pseudo-color image of the Indiana dataset;

[0071] Figure 8 OA line graph of the proposed method using different base learning rates on three datasets;

[0072] Fig. 9 OA line graphs of the proposed method using different regularization parameters on three datasets;

[0073] Fig.10 OA line graphs of the proposed method using different contribution weights on three datasets;

[0074] Fig.11 OA line graphs of the proposed method using different tribute patch sizes on three datasets;

[0075] Fig.12 The true value graph and model prediction results of Houston2018, including (a) GT; (b) SDEnet; (c) S2ECnet; (d) LLURnet; (a) GT; (b) GroupDRO; (c) ANDMask; (d) VREx; (e) DIFEX; (f) SDEnet; (g) S2ECnet; (h) LLURnet; (i) the method of the present invention;

[0076] Fig.13 The truth graph and model prediction results of PaviaCenter, including (a) GT; (b) SDEnet; (c) S2ECnet; (d) LLURnet; (a) GT; (b) GroupDRO; (c) ANDMask; (d) VREx; (e) DIFEX; (f) SDEnet; (g) S2ECnet; (h) LLURnet; (i) the method of the present invention;

[0077] Fig.14 The true value map and model prediction results of IndianaTD, including (a) GT; (b) SDEnet; (c) S2ECnet; (d) LLURnet; (a) GT; (b) GroupDRO; (c) ANDMask; (d) VREx; (e) DIFEX; (f) SDEnet; (g) S2ECnet; (h) LLURnet; (i) the method of the present invention;

[0078] Fig.15 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0079] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0080] Please see attached Figure 1 -Attached Fig.15 , a technical solution provided by the present invention:

[0081] Example:

[0082] The cross-scene hyperspectral image classification method based on multimodal domain generalization includes the following steps: step one, data preparation; step two, image feature extraction; step three, construction of optimized semantic space; step four, calculation of loss and optimization of model; step five, model testing;

[0083] In the above step 1, the hyperspectral image in the data set is cut into image blocks and sent to the image encoder to build text prior knowledge for the data set to form a knowledge base; the image block size is h×m×b, where h and w are the patch sizes of the image block, b represents the number of spectral channels, and for the knowledge base Each text prior knowledge t i Both with tags One-to-one correspondence, text prior knowledge describes the spatial relationship and image features of the category corresponding to the label. Given a label The following mapping relationship can be used Find the corresponding text t i :

[0084]

[0085] In the above step 2, the following steps are specifically included:

[0086] 2.1 Extracting spatial and spectral features: In the image encoder, the SSEM based on three-dimensional convolution extracts the spatial and spectral features of the hyperspectral image block through the three-dimensional spectral information. SSEM uses serial three-dimensional convolution residual blocks, maximum pooling layers and three-dimensional convolutions to deeply extract information. The three-dimensional convolution residual block consists of two Conv3d-BN3d-ReLU modules and one Conv3d module. The outputs of the first and last Conv3d-BN3d-ReLU modules are used for residual merging. The three-dimensional convolution formula is:

[0087]

[0088] Among them, h, v, and c represent the length, width, and number of channels of the three-dimensional convolution respectively. represents the value of the neuron at the jth feature sample (x, y, z) in the i-th layer network, y represents the weight of the (i-1)th feature map in the (h, v, c) convolution kernel, and b ij Indicates network bias;

[0089] 2.2 Extracting multi-scale features: In the image encoder, the MSEM based on the two-dimensional residual pyramid structure extracts the multi-scale features of the hyperspectral image block through the Conv2d-BN2d module, three pyramid residual modules and average pooling layer. i In (i=1, 2, 3), P1 contains three residual modules with the same structure. Each residual module contains three BN2d-Conv2d blocks and one BN2d-ReLU block. The kernel size of the second convolution layer is 7, and the kernel size of the rest is 3. The first residual module in P2 and P3 The kernel size of the second convolutional layer is 8, the stride is 2, and the rest of the structure is the same as P1. In order to systematically increase the depth of the feature maps of each layer, MSEM adopts the following rules:

[0090]

[0091] Where A is the initial number of channels, is the depth of the jth unit in the i-th pyramid residual module, N (net) is the total number of residual units in the entire network. At this time, the depth of each layer of the feature map depends on A and α;

[0092] 2.3 Feature concatenation: The features extracted by SSEM and MSEM are concatenated and sent to the classification head and the optimized semantic space;

[0093] In the above step 3, the following steps are specifically included:

[0094] 3.1 Text processing: The optimized semantic space construction module obtains text prior knowledge from the knowledge base according to the label, and then encodes the text prior knowledge through the text encoder. The encoding formula is as follows:

[0095] T=Embed[BPE(T)]+PE

[0096] Among them, BPE is lowercase byte pair encoding, the vocabulary size is 49152, Embed is the token embedding operation, and PE is the position encoding;

[0097] 3.2 Extracting text features: The text encoder uses the Transformer architecture to extract text features from the encoded text prior knowledge. The Transformer architecture uses a multi-head self-attention mechanism for feature extraction and tokenizes the text T∈RS×B×D is divided into h parts, each of which is set to Among them, B represents the batch size, S represents the maximum sentence length, D represents the embedding dimension size, and h represents the number of heads. After linear transformation, we get Q, K and V. The specific calculation process is as follows:

[0098]

[0099] Where W q , W k , W v All are learnable matrices;

[0100] Then calculate the attention weight matrix for each head and merge the attention matrices of all heads:

[0101]

[0102] Attention(T)∈R S×B×D =Concat(A1, A2, …, A h )

[0103] Among them, Softmax is used to perform Softmax operation on each data sequence, and Concat merges the attention matrix of each data head into one attention matrix;

[0104] The equation of the entire feature extraction process is simplified as follows:

[0105] T∈R S×B×D =LayerNorm1(Attention(T)+T)

[0106] T∈R S×B×D =LayerNorm2(FFN(T)+T)

[0107] Among them, LayerNorm1 and LayerNorm2 are layer normalization operations, FFN is a feed-forward neural network, the above feature extraction process involves only one layer, and the number of repetitions is the same as the number of layers. The Transformer architecture has 3 layers and 8 heads, with a width of 512;

[0108] 3.3 Constructing optimized semantic space and projecting image features: The optimized semantic space construction module uses the extracted text features to construct the optimized semantic space and projects the image features extracted by the image encoder into the optimized semantic space. The specific formula is as follows:

[0109] Z oss =T·w oss +b oss

[0110] Zssem =Linear[f ssem (X)]

[0111] Z msem =Linear[f msem (X)]

[0112] Among them, w oss is the learnable matrix, b oss is the bias term, and the linear layer is responsible for mapping image features to semantic space;

[0113] In the above step 4, the following steps are specifically included:

[0114] 4.1 Calculate classification loss: Input the image features extracted by the image encoder into the classification head, which outputs the classification probability and uses the true value to calculate the cross entropy loss of supervised learning, and obtain the classification loss and cross entropy loss of the network. The formula is:

[0115]

[0116] Classification loss of the network The formula is:

[0117]

[0118] Among them, y i is x i One-hot encoding of category information, P i and C(x i ) is the predicted probability, N represents the number of samples;

[0119] 4.2 Calculate contrast loss: Calculate contrast loss based on the similarity between image features and semantic space features. The similarity measurement formula is:

[0120]

[0121] In order to highlight the key role of similarity, Sim is squared and the optimized supervised contrast loss is:

[0122]

[0123] Among them, P oss (i) Corresponding image features, N oss (i) do not correspond to image features, namely positive features and negative features, and represents one of the positive and negative samples, |P(i)|=|P oss (i)|, the contribution weight α is used to adjust the influence of the image features of the two branches in the optimized semantic space;

[0124] 4.3 Calculating the overall loss: the overall loss during model training Defined as:

[0125]

[0126] in, is the classification loss, is the contrast loss, λ is a hyperparameter, and the role of λ is to balance the impact of optimizing the semantic space on the overall model;

[0127] 4.4 Model parameter optimization: The gradient of the model parameters is calculated based on the overall loss. Based on the gradient of the model parameters, the Adam optimizer is used to optimize the image encoder and text encoder.

[0128] In the above step five, SSEM and MSEM are used to extract image features and classify them, and the performance of the model optimized in step four in the target domain is evaluated based on OA and KC.

[0129] Experimental Example 1:

[0130] To verify the effectiveness of the method proposed in the embodiment, the following experiments were conducted: three publicly available datasets were used: Houston, Pavia, and Indiana. The Houston dataset is shown in Table 1, the Pavia dataset is shown in Table 2, and the Indiana dataset is shown in Table 3. All experiments were run in the same environment, using python 3.6 and pytorch 11.3 frameworks. The default weight decay of L2 regularization for all modules was set to 1e-4. The image encoder and text encoder were optimized by Adam, and the CPU was AMD EPYC7642. 48-core processor, GPU is RTX3090 with 24GB memory, knowledge bases are created for three data sets, each category corresponds to a prior knowledge, Houston's knowledge base is shown in Table 4, Pavia's knowledge base is shown in Table 5, and Indiana's knowledge base is shown in Table 6; In order to evaluate the sensitivity of the network to the three target domains, parameter sensitivity analysis is performed; the four adjustable hyperparameters are the base learning rate η, the regularization parameter λ, the contribution weight α, and the patch size; η is selected from 1e-1, 1e-2, 1e-3, 1e-4, and 1e-5; λ is selected from 1e+2, 1e+1, 1e+0, 1e-1, and 1e-2; α is selected from 0.1, 0.3, 0.5, 0.7, and 0.9; the patch size is selected from 9, 11, 13, 15, and 17; Figure 8 The classification results of different base learning rates on three datasets are shown; Fig. 9 The classification results for different regularization parameters on three datasets are shown; Fig.10The classification results of different weight contributions on three datasets are shown; Fig.11 The classification results of different patch sizes on three datasets are shown; on these three datasets, a basic learning rate of 1e-2, a contribution weight of 0.5, and a patch size of 13 are used; for the Houston and Pavia datasets, the regularization parameter is set to 1e-0, which can effectively prevent overfitting and improve the stability of the model; for the Indiana dataset, setting the regularization parameter to 1e+1 will help solve the problem of large feature differences in this dataset, thereby improving the classification performance.

[0131] Experimental Example 2:

[0132] In order to verify the effectiveness of the proposed SSEM, MSEM and optimized semantic space, an ablation experiment was designed; in the experiment, the first two modules are SSEM and MSEM, and the last two modules are SS and OSS; SS means the use of original supervised contrastive learning, while OSS introduces optimized supervised contrastive learning; by comparing the experimental results of SS and OSS, the effectiveness of the proposed contrastive learning loss optimization method is proved; Table 7 shows that the experimental results show that the proposed method of the present invention has better performance than these variants; specifically, SSEM and MSEM are designed as two branches of a parallel network, and training them together can give full play to their respective advantages and achieve complementary advantages; the optimization of semantic space and supervised contrastive loss further improves the generalization ability of the model.

[0133] Comparative Example:

[0134] To demonstrate the superiority of the method proposed in the embodiment in the cross-domain classification task, a comparative experiment was conducted with other methods, specifically: SDEnet, S2ECnet and LLURnet were used as comparison methods; since the existing cross-scene hyperspectral classification methods are relatively limited, four RGB image classification methods, namely GroupDRO, ANDMask, VREx and DIFEX, were selected to further verify the effectiveness of the method; all data samples and labels in the source domain were used as training instances, 80% of which were used for training and 20% for verification; at the same time, all data in the target domain were used for testing; in order to minimize the impact of random sampling, each algorithm took the average of ten runs to determine the results; in order to measure the accuracy and consistency between the true label and the predicted value, the overall accuracy (OA) and Kappa coefficient (KC) of the target domain were used as evaluation indicators of the algorithm; the results are shown in Tables 8, 9 and 10; in H On the ouston dataset, LLURNet has the best performance among all the compared methods; the OA performance of the proposed method is 4.22 higher than that of LLURNet, and the KC performance is 7.16 higher than that of LLURNet; on the Pavia dataset, LLURNet has the best performance among all the compared methods; the OA performance of the proposed method is 0.86 higher than that of LLURNet, and the KC performance is 0.96 higher than that of LLURNet; on the Indiana dataset, DIFEX has the best performance among all the compared methods; the OA performance of the proposed method is 2.08 higher than that of DIFEX, and the KC performance is 3.76 higher than that of DIFEX; in the Indiana dataset, the standard deviation of some categories is greater than the correct rate, indicating that the experimental results fluctuate greatly, which reflects the challenge of classifying these feature categories; the proposed method effectively alleviates this problem; in general, the classification effect of the proposed method is better than that of the compared methods; the classification diagrams of the three datasets are shown in Figure 2. Fig.12 , Fig.13 and Fig.14 As shown in Figure 2, each pixel with an unspecified category is used as the background, while pixels with a category are predicted for comparison. Each image contains a true value image and the prediction results of all models. Fig.14 It can be clearly seen that the method of the present invention has a better prediction effect on the second and fifth categories.

[0135] Table 1 Houston dataset

[0136]

[0137] Table 2 Pavia dataset

[0138]

[0139]

[0140] Table 3 Indiana dataset

[0141]

[0142] Table 4 Knowledge base of the Houston dataset

[0143]

[0144] Table 5 Knowledge base of Pavia dataset

[0145]

[0146]

[0147] Table 6 Knowledge base of Indiana dataset

[0148]

[0149] Table 7 Ablation experiment results

[0150]

[0151] Table 8 Comparison of accuracy, OA (%) and KC (κ) of different methods for target scene Houston2018

[0152]

[0153] Table 9 Comparison of different methods on the target scene Pavia Center class accuracy, OA (%) and KC (κ)

[0154]

[0155]

[0156] Table 10 Comparison of accuracy, OA (%) and KC (κ) of different methods for target scene IndianaTD class

[0157]

[0158] Based on the above, the advantage of the present invention is that the method proposed in the present invention is a multimodal fusion network, which reduces the influence of data distribution differences by introducing language modalities, thereby assisting in learning cross-domain invariant representations of hyperspectral images. The method of the present invention includes a parallel architecture of a spatial-spectral feature extraction module (SSEM) and a multi-scale feature extraction module (MSEM) to capture richer image information in the feature extraction stage. In addition to being used for classification, the extracted features are also mapped to an optimized semantic space (OSS). The optimized semantic space brings similar features closer and separates different features, thereby realizing a deep interactive fusion of the two modalities of image and text, and promoting the learning of cross-domain invariant representations. Compared with existing methods, the method of the present invention has better performance in cross-scene hyperspectral image classification tasks.

[0159] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

Claims

1. A cross-scene hyperspectral image classification method based on multimodal domain generalization includes the following steps: Step 1: data preparation; Step 2: image feature extraction; Step 3: construct an optimized semantic space; Step 4: calculate the loss and optimize the model; Step 5: test the model; it is characterized by: In the above step 1, the hyperspectral image in the data set is cut into image blocks and sent to the image encoder to construct text prior knowledge for the data set to form a knowledge base; In the above step 2, the following steps are specifically included: 2.1 Extracting spatial and spectral features: In the image encoder, the SSEM based on three-dimensional convolution extracts the spatial and spectral features of the hyperspectral image block through the three-dimensional spectral information; 2.2 Extracting multi-scale features: In the image encoder, the MSEM based on the two-dimensional residual pyramid structure is used to extract the multi-scale features of the hyperspectral image blocks through the Conv2d-BN2d module, three pyramid residual modules and average pooling layer; 2.3 Feature concatenation: The features extracted by SSEM and MSEM are concatenated and sent to the classification head and the optimized semantic space; In the above step 3, the following steps are specifically included: 3.1 Text processing: The optimized semantic space construction module obtains text prior knowledge from the knowledge base according to the label, and then encodes the text prior knowledge through the text encoder; 3.2 Extracting text features: The text encoder uses the Transformer architecture to extract text features from the encoded text prior knowledge; 3.3 Constructing optimized semantic space and projecting image features: The optimized semantic space construction module constructs the optimized semantic space using the extracted text features, and projects the image features extracted by the image encoder into the optimized semantic space; In the above step 4, the following steps are specifically included: 4.1 Calculate classification loss: Input the image features extracted by the image encoder into the classification head, which outputs the classification probability and uses the true value to calculate the cross entropy loss of supervised learning and obtain the classification loss of the network; 4.2 Calculate contrast loss: Calculate contrast loss based on the similarity between image features and semantic space features; 4.3 Calculating the overall loss: the overall loss during model training Defined as: in, is the classification loss, is the contrast loss, λ is a hyperparameter; 4.4 Model parameter optimization: The gradient of the model parameters is calculated based on the overall loss. Based on the gradient of the model parameters, the Adam optimizer is used to optimize the image encoder and text encoder. In the above step five, SSEM and MSEM are used to extract image features and classify them, and the performance of the model optimized in step four in the target domain is evaluated based on OA and KC.

2. The cross-scene hyperspectral image classification method based on multimodal domain generalization according to claim 1 is characterized by: In the step 1, the image block size is h×w×b, where h and w are both the patch sizes of the image block, and b represents the number of spectral channels.

3. The cross-scene hyperspectral image classification method based on multimodal domain generalization according to claim 1 is characterized in that: In the step 1, for the knowledge base Each text prior knowledge t i Both with tags One-to-one correspondence, text prior knowledge describes the spatial relationship and image features of the category corresponding to the label. Given a label The following mapping relationship can be used Find the corresponding text t i :

4. The cross-scene hyperspectral image classification method based on multimodal domain generalization according to claim 1 is characterized in that: In step 2.1, SSEM uses serial three-dimensional convolution residual blocks, maximum pooling layers and three-dimensional convolutions to deeply extract information. The three-dimensional convolution residual block consists of two Conv3d-BN3d-ReLU modules and one Conv3d module. The three-dimensional convolution formula is: Among them, h, v, and c represent the length, width, and number of channels of the three-dimensional convolution respectively. represents the value of the neuron at the jth feature sample (x, y, z) in the i-th layer network, y represents the weight of the (i-1)th feature map in the (h, v, c) convolution kernel, and b ij Indicates network bias.

5. The cross-scene hyperspectral image classification method based on multimodal domain generalization according to claim 1 is characterized in that: In step 2.2, the three pyramid residual modules P i In (i=1, 2, 3), P1 contains three residual modules with the same structure. Each residual module contains three BN2d-Conv2d blocks and one BN2d-ReLU block. The kernel size of the second convolution layer is 7, and the kernel size of the rest is 3. The first residual module in P2 and P3 The kernel size of the second convolutional layer is 8, the stride is 2, and the rest of the structure is the same as P1. In order to systematically increase the depth of the feature maps of each layer, MSEM adopts the following rules: Where A is the initial number of channels, is the depth of the jth unit in the i-th pyramid residual module, N (net) is the total number of residual units in the entire network. At this time, the depth of each layer of the feature map depends on A and α.

6. The cross-scene hyperspectral image classification method based on multimodal domain generalization according to claim 1 is characterized by: In step 3.1, the encoding formula is as follows: T=Embed[BPE(T)]+PE Among them, BPE is lowercase byte pair encoding, the vocabulary size is 49152, Embed is the token embedding operation, and PE is the position encoding.

7. The cross-scene hyperspectral image classification method based on multimodal domain generalization according to claim 1 is characterized by: In step 3.2, the Transformer architecture uses a multi-head self-attention mechanism to extract features and tokenize the text T∈R S×B×D is divided into h parts, each of which is set to α = 1, 2, ..., h, where B represents the batch size, S represents the maximum sentence length, D represents the embedding dimension size, and h represents the number of heads. After linear transformation, we get Q, K and V. The specific calculation process is as follows: Where W q , W k , W v All are learnable matrices; Then calculate the attention weight matrix for each head and merge the attention matrices of all heads: Attention(T)∈R S×B×D =Concat(A1,A2,…,A h ) Among them, Softmax is used to perform Softmax operation on each data sequence, and Concat merges the attention matrix of each data head into one attention matrix; The equation of the entire feature extraction process is simplified as follows: T∈R S×B×D =LayerNorm1(Attention(T)+T) T∈R S×B×D =LayerNorm2(FFN(T)+T) Among them, LayerNorm1 and LayerNorm2 are layer normalization operations, FFN is a feed-forward neural network, the above feature extraction process involves only one layer, the number of repetitions is the same as the number of layers, and the Transformer architecture has 3 layers and 8 heads, with a width of 512.

8. The cross-scene hyperspectral image classification method based on multimodal domain generalization according to claim 1 is characterized by: In step 3.3, the specific formula is as follows: Z oss =T·w oss +b oss Z ssem =Linear[f ssem (X)] Z msem =Linear[f msem (X)] Among them, w oss is the learnable matrix, b oss is the bias term, and the linear layer is responsible for mapping image features to the semantic space.

9. The cross-scene hyperspectral image classification method based on multimodal domain generalization according to claim 1, characterized in that: In step 4.1, the cross entropy loss The formula is: Classification loss of the network The formula is: Among them, y i is x i One-hot encoding of category information, P i and C(x i ) is the predicted probability, and N is the number of samples.

10. The cross-scene hyperspectral image classification method based on multimodal domain generalization according to claim 1, characterized in that: In step 4.2, the similarity measurement formula is: By squaring Sim, the optimized supervised contrast loss is: Among them, P oss (i) Corresponding image features, N oss (i) do not correspond to image features, namely positive features and negative features, and represents one of the positive and negative samples, |P(i)|=|P oss (i)|, α is the contribution weight.

Citation Information

Patent Citations

  • Hyperspectral image classification method based on double-branch spectrum multi-scale attention network

    CN113486851A

  • Apparatus and method for image-guided interventions with hyperspectral imaging

    US20200237229A1