Hyperspectral image small sample classification method based on joint domain adaptive weight self-learning
The joint domain adaptation weight self-learning method addresses the challenges of high-spectral image classification by enhancing common features and suppressing domain differences, improving classification accuracy and adaptability with limited labeled samples.
Patent Information
- Application Number
- CN202510500208.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-15
AI Technical Summary
In cross-domain learning, the existing hyperspectral image classification methods are limited in model performance and adaptability due to the complex distribution differences between source and target domains, especially in the case of scarce labeled samples.
Using a method based on joint domain adaptive weight self-learning, features are extracted through a spectral-space attention dual-branch dense network, combined with a domain discriminator and a domain projector to enhance common features, suppress differential features, and adjust the loss weight through an adaptive learner to construct a small sample classification model of hyperspectral images.
It improves the accuracy and robustness of the classification of small samples of hyperspectral images, and performs better especially when there is less prior information, improving the generalization ability of the model in cross-domain learning.
Smart Images

Figure CN120318589A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hyperspectral image few-shot classification, and particularly relates to a hyperspectral image few-shot classification method based on joint domain adaptation weight self-learning. Background Technique
[0002] Hyperspectral imaging is an important remote sensing technology that captures rich spectral information through the electromagnetic spectrum to provide high-resolution surface images. Hyperspectral classification aims to classify each pixel according to its spectral characteristics to accurately depict the land cover distribution.
[0003] Due to its advantages in automatic feature extraction and high-dimensional data processing, deep learning has achieved remarkable success in hyperspectral image classification tasks. Specifically, the convolutional neural network significantly reduces the number of parameters through local connection and weight sharing, making it widely used in hyperspectral image classification. However, since the deep learning hyperspectral image classification method requires a large number of labeled samples to train the model to reach the optimal performance, and the annotation of hyperspectral images usually requires a large amount of labor cost, and the annotation process is very expensive and time-consuming, the scarcity of labeled samples is still the main obstacle to achieving accurate classification.
[0004] When the annotation of target domain data is scarce, cross-domain learning can reduce the need for labeled data in the target domain by transferring the label information and knowledge of the source domain. In this way, the model can still achieve good classification results in the case of scarce labels. However, there are often significant differences between the source domain and target domain distributions, such as the mismatch of the number of spectral bands. When the source domain features cannot be transferred to the target domain, it will lead to a decline in the performance of the model.
[0005] Existing cross-domain learning models usually adopt adversarial domain adaptation techniques to balance the distribution differences between the source domain and the target domain. These methods usually only focus on suppressing the differential features between domains, and enhancing the common features is often a by-product. However, the actual distribution differences between the source domain and the target domain are often diverse and complex. Merely suppressing the differential features between domains may not be sufficient to handle these distribution differences, resulting in poor transfer effects, and the performance and adaptability of the model will also be significantly limited. Summary of the Invention
[0006] The purpose of the present invention is to propose a hyperspectral image few-shot classification method based on joint domain adaptation weight self-learning, which enhances the common features while suppressing the differential features between domains, and sets an adaptive learner to balance the weights between different domain adaptation losses, so as to improve the model performance and ensure the rationality and quality of the hyperspectral image few-shot classification results.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A small-sample classification method for hyperspectral images based on joint domain adaptation weight self-learning, comprising the following steps:
[0009] Step 1. Build a small-sample classification model for hyperspectral images based on joint domain adaptation weight self-learning, which includes the following structure:
[0010] An embedded feature extraction module, which includes a mapping layer and a feature extraction network; the mapping layer is used to ensure that the data dimensions of the source domain and target domain hyperspectral images input into the feature extraction network are consistent; the feature extraction network extracts the spectral-spatial discriminant features of the source domain and target domain through a spectral-spatial attention double-branch dense network;
[0011] A small-sample learning module, which is used to perform small-sample learning on the source domain and target domain according to the distances between the support set and query set samples in the embedded space, and obtain the class distribution of the query set samples;
[0012] A joint domain adaptation module, which includes a domain discriminator and a domain projector; the domain projector is used to enhance the common features of the source domain and target domain, and the domain discriminator is used to suppress the differential features of the source domain and target domain;
[0013] A weight self-learning module, which is used to adjust the weight between the domain discrimination loss and the domain projection loss through an adaptive learner;
[0014] A multi-label classification module, which is used to classify the unlabeled samples in the target domain according to the spectral-spatial discriminant features output by the embedded feature extraction module;
[0015] Step 2. Combine the small-sample learning loss and domain adaptation loss of the source domain and target domain to construct a total loss function, and train the small-sample classification model for hyperspectral images based on joint domain adaptation weight self-learning;
[0016] Step 3. Use the trained small-sample classification model for hyperspectral images based on joint domain adaptation weight self-learning to classify the input target domain hyperspectral images.
[0017] In addition, on the basis of the above small-sample classification method for hyperspectral images based on joint domain adaptation weight self-learning, the present invention also proposes a computer device, which includes a memory and one or more processors.
[0018] The memory stores executable code, and when the processor executes the executable code, it is used to implement the steps of the above-mentioned small-sample classification method for hyperspectral images based on joint domain adaptation weight self-learning.
[0019] The present invention has the following advantages:
[0020] As described above, the present invention relates to a small-sample classification method for hyperspectral images based on joint domain adaptation weight self-learning. Different from most current small-sample classification methods for hyperspectral images based on adversarial domain adaptation techniques, the method of the present invention sets up a spectral-spatial attention dual-branch dense network to achieve efficient extraction of spectral-spatial discriminative features; and through a joint domain discriminator and a domain projector, better aligns the features between different domains in two different directions; the present invention also designs an adaptive learner, which has a significant improvement effect on the problem of poor out-of-domain label quality caused by inconsistent feature spaces in different hyperspectral images, making the model perform more excellently in cross-domain learning. The method of the present invention improves the performance of the small-sample classification model for hyperspectral images, improves the rationality and quality of the small-sample classification results of hyperspectral images, especially in the case of less prior information, the method of the present invention has more advantages, and the classification performance of the small-sample classification model for hyperspectral images proposed by the present invention is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of the small-sample classification method for hyperspectral images based on joint domain adaptation weight self-learning in an embodiment of the present invention.
[0022] Figure 2 It is a partial model architecture diagram of the small-sample classification model for hyperspectral images based on joint domain adaptation weight self-learning in an embodiment of the present invention.
[0023] Figure 3 It is a schematic diagram of classifying a target domain hyperspectral image by using the trained small-sample classification model for hyperspectral images based on joint domain adaptation weight self-learning in an embodiment of the present invention.
[0024] Figure 4 It is a schematic diagram of the structure of the embedded feature extraction module in an embodiment of the present invention.
[0025] Figure 5 It is a schematic diagram of the structure of the small-sample learning module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0027] Embodiment 1
[0028] In this embodiment, a small-sample classification method for hyperspectral images based on joint domain adaptation weight self-learning is disclosed. First, a mapping layer is constructed to align the dimensions of the source domain and the target domain. The output of the mapping layer serves as the input to the feature extraction network, and a spectral-spatial double-branch dense network is used to avoid gradient vanishing. At the same time, a joint channel and spatial attention mechanism is used to efficiently extract the spectral-spatial fusion features of hyperspectral images. Then, a domain projector and a domain discriminator are introduced to solve the domain adaptation problem in different directions simultaneously to handle the significant distribution differences existing in different domains of hyperspectral images. Subsequently, an adaptive learner is designed to adjust the weights to achieve an effective unification of joint domain adaptation, thereby efficiently achieving domain distribution alignment. Finally, metric learning is used to perform small-sample classification of hyperspectral images in the target domain. The present invention uses joint domain adaptation technology to align and adjust features in different directions to better handle the distribution differences between domains, thereby improving the generalization ability and robustness of the model when facing unknown domains.
[0029] As Figure 1 shown, the small-sample classification method for hyperspectral images based on joint domain adaptation weight self-learning includes the following steps:
[0030] Step 1. Build a small-sample classification model for hyperspectral images based on joint domain adaptation weight self-learning, which includes an embedded feature extraction module, a small-sample learning module, a joint domain adaptation module, a weight self-learning module, and a multi-label classification module.
[0031] The embedded feature extraction module includes a mapping layer and a feature extraction network; the mapping layer is used to ensure that the data dimensions of the source domain and the target domain input into the feature extraction network are the same; the feature extraction network is used to extract the spectral-spatial discriminant features of the source domain and the target domain through a spectral-spatial attention double-branch dense network.
[0032] The mapping layer of the present invention is implemented based on a two-dimensional convolutional neural network (2D CNN) and is used to make the number of bands of hyperspectral images in the source domain and the target domain equal. Specifically, hyperspectral images in the source domain and the target domain are respectively selected. After being processed by the mapping layer, the number of bands of hyperspectral images in the source domain and the target domain is the same to ensure that the data dimensions of the subsequent source domain and the target domain input into the feature extraction network are the same.
[0033] The processing flow of the signal in the mapping layer is as follows:
[0034] Assume that the hyperspectral image input into the mapping layer is cube I, where w×w is the size of the cube, and L represents the number of bands. Then, the feature output of the mapping layer is:
[0035] I′ = I × T (1)
[0036] where T represents the mapping transformation matrix, l is the number of channels after the mapping transformation; I′ is the hyperspectral image after the mapping transformation. In this embodiment, w = 9 and l = 100. There are L×l learnable parameters in the mapping transformation.
[0037] Subsequently, I′ is used as the input of the batch normalization layer, and the batch normalization layer is utilized to make all the reflectance values of the hyperspectral images in the source domain and the target domain fall between 0 and 1.
[0038] Specifically, as Figure 2 shown, in this embodiment, the mapping layer includes a source domain mapping layer, a target domain mapping layer, and a batch normalization layer. A hyperspectral image with 128 bands in a source domain dataset such as the Chikusei dataset is selected and input into the source domain mapping layer, and a hyperspectral image with 200 bands in a target domain dataset such as the Indian Pines dataset is selected and input into the target domain mapping layer. After the mapping transformation, hyperspectral images of the source domain and the target domain with the same number of bands are output respectively, and the batch normalization layer is used to process the reflectance values of the output hyperspectral images of the source domain and the target domain to make them fall between 0 and 1. Finally, the output of the mapping layer is obtained.
[0039] The feature extraction network is used to extract the spectral-spatial discriminative features of the source domain and the target domain through a spectral-spatial attention double-branch dense network. The processing flow of the signal in the feature extraction network is as follows: The hyperspectral image output by the mapping layer is used to extract spectral features and spatial features by the spectral-spatial double-branch dense network, and the spectral weight and the spatial weight are obtained through the attention mechanism. Subsequently, the spectral features and the spatial features are weighted and fused to obtain the spectral-spatial discriminative features of the source domain and the target domain.
[0040] As Figure 2 shown, the spectral-spatial attention double-branch dense network is used as the feature extraction network to extract spectral-spatial discriminative features, namely the spectral-spatial features of the source domain and the spectral-spatial features of the target domain. The spectral-spatial attention double-branch dense network includes a spectral-spatial double-branch dense network and a feature fusion module, where the spectral-spatial double-branch dense network includes a spectral branch and a spatial branch.
[0041] The output of the mapping layer is used as the input of the feature extraction network. The spectral branch can effectively distinguish the spectral features of different ground objects, while the spatial branch is responsible for capturing the global and local spatial features of the image. The input data extracts spectral features and spatial features through the spectral and spatial branches respectively, generates a spectral feature map and a spatial feature map, and calculates the spectral weight and the spatial weight through the spectral attention mechanism module and the spatial attention mechanism module respectively. Then, the spectral-spatial features of the source domain and the target domain are obtained through weighted fusion by the feature fusion module, that is, the final feature vector is generated by fusion, providing richer ground object discriminative information for the model, thereby improving the classification accuracy.
[0042] In this embodiment, the detailed structure of the feature extraction network is shown in Table 1. Among them, the spectral branch includes five dense spectral blocks, three feature splicing modules, and a spectral attention mechanism module. The five dense spectral blocks are respectively defined as the first dense spectral block, the second dense spectral block, the third dense spectral block, the fourth dense spectral block, and the fifth dense spectral block. Each dense spectral block includes a 3D convolutional layer, a 3D batch normalization layer, and a mish layer. The three feature splicing modules of the spectral branch are defined as the first feature splicing module, the second feature splicing module, and the third feature splicing module. The spectral attention mechanism module includes a spectral attention layer, a 3D batch normalization layer, a mish layer, and a global average pooling layer.
[0043] The spatial branch includes four dense spatial blocks, three feature splicing modules, and a spatial attention mechanism module. The four dense spatial blocks are respectively defined as the first dense spatial block, the second dense spatial block, the third dense spatial block, and the fourth dense spatial block. Each dense spatial block includes a 3D convolutional layer, a 3D batch normalization layer, and a mish layer. The three feature splicing modules of the spatial branch are respectively defined as the fourth feature splicing module, the fifth feature splicing module, and the sixth feature splicing module. The spatial attention mechanism module includes a spatial attention layer, a 3D batch normalization layer, a mish layer, and a global average pooling layer.
[0044] Table 1 Detailed structure of the feature extraction network
[0045]
[0046] The processing flow of the signal in the feature extraction network is as follows:
[0047] Take the output of the mapping layer as the input of the feature extraction network.
[0048] The hyperspectral images of the source domain and the target domain input to the spectral branch are successively subjected to feature extraction through the first dense spectral block and the second dense spectral block, and then the outputs of the first dense spectral block and the second dense spectral block are feature-spliced through the first feature splicing module, and the output of the first feature splicing module is used as the input of the third dense spectral block. The second feature splicing module feature-splices the outputs of the first dense spectral block, the first feature splicing module, and the third dense spectrum, and the output of the second feature splicing module is used as the input of the fourth dense spectral block. The third feature splicing module feature-splices the outputs of the first dense spectral block, the first feature splicing module, the second feature splicing module, and the fourth dense spectrum, and the output of the third feature splicing module is used as the input of the fourth dense spectral block. The output of the fourth dense spectral block is the spectral feature map of the source domain and the target domain.
[0049] The high-spatial images of the source domain and the target domain in the input space branch are successively passed through the first dense spatial block and the second dense spatial block for feature extraction, and then the outputs of the first dense spatial block and the second dense space are feature-stitched through the fourth feature stitching module, and the output of the fourth feature stitching module is used as the input of the third dense spatial block. The fifth feature stitching module stitches the outputs of the first dense spatial block, the fourth feature stitching module, and the third dense space, and the output of the fifth feature stitching module is used as the input of the fourth dense spatial block. The sixth feature stitching module stitches the outputs of the first dense spatial block, the fourth feature stitching module, the fifth feature stitching module, and the fourth dense space, and the output of the sixth feature stitching module is the spatial feature map of the source domain and the target domain.
[0050] The spectral feature maps of the source domain and the target domain output by the spectral branch are input into the spectral attention layer to calculate the spectral weights. The spatial feature maps of the source domain and the target domain output by the spatial branch are input into the spatial attention layer to calculate the spatial weights. In the spectral attention layer, the input features generate query Q, key K, and value V through 1×1 convolution, and then self-attention calculation is performed on the channel dimension after flattening the spatial dimension. The channel similarity matrix is obtained through the dot product of Q and K, and after Softmax normalization, it is used to weight V, thereby obtaining the spectral features that emphasize important bands; the spatial attention layer also uses the self-attention mechanism. After aggregating the input features along the channel dimension, Q, K, and V at spatial positions are generated simultaneously, and then the correlation between positions is calculated after flattening the spatial dimension to obtain the spatial attention weights, which are used to weight and enhance the feature map in the spatial dimension. The two attention modules work together to effectively improve the model's ability to focus on key information. The spectral feature map and the spatial feature map are weighted and fused to generate the final feature vector.
[0051] The few-shot learning module is used to perform few-shot learning of the source domain and the target domain based on the spectral-spatial discriminative features and according to the distances between the support set and query set samples in the embedding space.
[0052] Few-shot learning is performed in the source domain and the target domain in a metric learning manner to learn the transferable knowledge in the source domain and generate an embedding space with high sample separation in the target domain. In this embedding space, similar samples are closer and dissimilar samples are farther apart, so as to achieve the final division of samples by optimizing the sample distances.
[0053] The few-shot learning module aims to effectively train the model through a small number of labeled samples. In this module, few-shot learning is simultaneously performed in the source domain and the target domain through metric learning, which can not only learn the transferable knowledge in the source domain but also generate an embedding space with high sample separation in the target domain.
[0054] The processing flow of the signal in the few-shot learning module is as follows:
[0055] Since the few-shot learning execution processes in the source domain and the target domain are similar, the few-shot learning in the source domain is taken as an example here.
[0056] First, the support set and the query set are partitioned. K s samples are selected from the labeled samples of each class in the source domain to form the source domain support set where C is the total number of classes, i ∈ [1, C×K s , x i is the i-th labeled sample in the source domain support set, and y i is the class label corresponding to x i Subsequently, N s samples are selected from the remaining samples of each class in the source domain to form the source domain query set where j ∈ [1, C×N s , x j is the j-th labeled sample in the source domain query set, and y j is the class label corresponding to x j Then, the samples of S s and Q s are input into the embedding feature extraction module, dimensionality reduction is performed using the mapping layer, and discriminative features are obtained through the feature extraction network. Finally, distance metrics are calculated based on the Euclidean distance, few-shot learning in the source domain and the target domain is performed, transferable knowledge in the source domain is learned, and an embedding space with a high sample separation degree is generated in the target domain.
[0057] In the source domain, the class distribution of the query set samples x s in the source domain query set Q j is:
[0058]
[0059] where P(y j = u|x j ∈ Q s ) represents the class distribution of the query set samples x s in the source domain query set Q j , d represents the Euclidean distance, v c is the embedding prototype vector of the k-th class in the source domain support set, is the embedding feature extraction module with parameter , v u represents the prototype vector of the true class corresponding to x j , c ∈ [1, C].
[0060] According to the negative log probability of the feature vector in the source domain query set and its true class k, a source set The small-sample learning loss of all query set samples is expressed as:
[0061]
[0062] where k represents the true class, x and y represent the sample and its corresponding label, represents the feature vector in the embedding space belonging to the true class k.
[0063] The execution process of few-shot learning in the target domain is similar to that in the source domain. It should be noted that the number of labeled samples in the target domain is extremely small, and data augmentation such as random flipping and adding Gaussian noise needs to be performed on the labeled samples in the target domain before dividing the support set and query set. From the target domain support set S t and the target domain query set Q t The target set composed of the few-shot learning loss is expressed as:
[0064]
[0065] where S t and Q t are the target domain support set and the target domain query set respectively. In this way, the model can more effectively learn the correlation relationship between samples to accurately distinguish samples of different classes.
[0066] A joint domain adaptation module, which includes a domain discriminator and a domain projector; the domain projector is used to enhance the common features of the source domain and the target domain, and the domain discriminator is used to suppress the different features of the source domain and the target domain.
[0067] Use the domain projector to enhance the common features of the source domain and the target domain, and use the domain discriminator to suppress the different features of the source domain and the target domain to balance the domain differences between different hyperspectral images.
[0068] Construct the domain discriminator and the domain projector to adjust the domain shift in different directions, so as to guide the model to complete domain adaptation in different directions; on the one hand, the domain projector is used to learn to map the source domain and target domain data into a shared feature space to maximize the common features of the source domain and the target domain. On the other hand, the domain discriminator prompts the model to learn a feature representation that is insensitive to the domain differences, and blurs the differences between the source domain and the target domain. By combining the domain projector and the domain discriminator to enhance the common features and suppress the different features to solve the domain distribution differences, the model can better align the features between domains and perform more excellently in the cross-domain adaptation process. During the domain adaptation process, the domain projection loss L dp and the domain discrimination loss L ddTo train the embedding feature extraction module, domain projector, and domain discriminator. The specific implementation is as follows.
[0069] For the domain projection loss L dp :
[0070] The domain projector aims to project the source domain and target domain data onto a unified new nominal domain, that is, map the source domain and target domain data into a shared feature space to achieve domain alignment and minimize the inter-domain difference. This nominal domain can be regarded as a one-class classification OCC task problem.
[0071] The processing flow of the signal in the domain projector is as follows:
[0072] Take the output of the embedding feature extraction module as the input of the domain projector. In this embodiment, the input of the domain projector is a feature vector with a shape of 1×120. The input feature vector of the domain projector is processed through a fully connected layer, ReLU layer, Dropout layer, and Sigmoid layer in sequence to transform the source domain and target domain features, and finally map to a shared feature space to output a feature vector with a shape of 1×1, realizing the projection of the source domain and target domain data to the new nominal domain.
[0073] The domain projector can predict the probability p o (·) of each sample belonging to the new nominal domain. Accordingly, based on the domain projection loss L dp of the one-class classification task is defined as:
[0074]
[0075] where y o represents the domain label, and its value is set to 1, that is, the domain label of the new nominal domain mapped by the source domain and target domain is 1. n s represents the total number of samples in a source set, that is, the sum of the number of samples in the source domain support set and the source domain query set, represents the i s -th sample in the source set, i s ∈[1, n s . n t then represents the total number of samples in a target set, that is, the sum of the number of samples in the target domain support set and the target domain query set, represents the i t -th sample in the target set, i t ∈[1, n t . x j represents the j-th sample mapped into the new nominal domain, j ∈[1, n s + n t .
[0076] The domain projector loss L dpEssentially, it is the cross-entropy loss in the OCC task. By minimizing L dp the purpose of projecting the source domain and target domain features into a unified new nominal domain can be achieved, thereby reducing the inter-domain difference.
[0077] For the domain discrimination loss L dd :
[0078] Different from the domain projector, the domain discriminator further blurs the difference between the source domain and the target domain by suppressing the differential features of the two, thereby reducing the inter-domain difference. The present invention adopts a conditional domain discriminator D with respect to the source distribution P s (x) and the target distribution P t (x) to reduce the inter-domain difference. After feature extraction, a linear classifier is used to obtain the classification discriminant information of the features.
[0079] The processing flow of the signal in the domain discriminator is as follows:
[0080] To prevent dimensional explosion, the present invention selects the result of the random bilinear mapping M(·) as the input of the domain discriminator, where g and are respectively the domain variances of the feature extraction network and the linear classifier in the embedded feature extraction module. After the random bilinear mapping, a domain discriminator composed of five fully connected layers is used to predict the domain label of the input sample. Specifically, the feature vector with the input shape of 1×120 passes through four fully connected layers in sequence. Each fully connected layer contains a ReLU layer and a Dropout layer to enhance the non-linear expression ability and prevent overfitting. The dimension of the output feature vector remains 1×1024. Subsequently, the feature vector is reduced to 1×2 through a fully connected layer and normalized through a Softmax layer, and finally the probability distributions of different domains are obtained, that is, the probability p d that the domain label of the input sample is predicted to be the source domain is obtained, which is used to judge whether the sample comes from the source domain or the target domain.
[0081] The domain discrimination loss L dd is defined as:
[0082]
[0083] where, represents any source set, represents any target set, p d and 1-p d respectively represent the probabilities that the domain discriminator predicts the input sample to be the source domain and the target domain, and represent the results of the random bilinear mapping. By minimizing L dd the purpose of making the source domain and the target domain indistinguishable can be achieved, thereby further reducing the inter-domain difference on the basis of the domain projector.
[0084] The weight self - learning module is used to adjust the weight between the domain discrimination loss and the domain projection loss through an adaptive learner. The adaptive learner is used to adjust the weight between the domain discrimination loss and the domain projection loss, achieve an effective unification of joint domain adaptation, and guide model optimization.
[0085] The adaptive learner adopts an adaptive weighting strategy, enabling the model to autonomously learn the weights α and β of the domain adaptation losses L dp and L dd gradually to reach the optimum. First, define and
[0086]
[0087]
[0088] where η represents the learning rate, and the initial values of α and β are both set to 1 / 2. After normalization, the values of α and β are:
[0089]
[0090] The multi - label classification module is used to classify the unlabeled samples in the target domain according to the spectral - spatial discriminant features output by the embedding feature extraction module. In this embodiment, a nearest - neighbor (NN) classifier is used, which classifies the hyperspectral image based on the spectral - spatial discriminant features generated by the embedding feature extraction module.
[0091] When training the model, the nearest - neighbor classifier is not used. Instead, the few - shot learning module classifies based on the Euclidean distance calculated between each prototype. Compared with directly using the nearest - neighbor classifier to obtain labels, the few - shot learning module finally obtains soft labels, that is, probability distributions, which can be used to calculate the few - shot learning loss to optimize the model. After completing the training of the model, the labeled samples are used to fit the nearest - neighbor classifier to classify the unlabeled samples. Therefore, in this embodiment, the multi - label classification module finally only contains one nearest - neighbor classifier and is only used when performing few - shot hyperspectral image classification after completing the training of the model.
[0092] Step 2. Construct a total loss function by combining the few - shot learning loss and the domain adaptation loss of the source domain and the target domain, and train the hyperspectral image few - shot classification model based on joint domain adaptation weight self - learning.
[0093] Define the total loss L total of the model as:
[0094]
[0095] where Denote the total source domain loss, and the total target domain loss. Specifically, define the total source domain loss as:
[0096]
[0097] where is the few-shot learning loss of the source domain, and is the domain adaptation loss of the source domain.
[0098]
[0099] Similarly, combining the few-shot learning loss of the target domain and the domain adaptation loss the total target domain loss is defined as:
[0100]
[0101] where is the few-shot learning loss of the target domain, and is the domain adaptation loss of the target domain.
[0102]
[0103] When the model finishes optimizing according to L total a support set is constructed using a small number of labeled samples in the target domain, and the spectral-spatial discriminative features generated by the feature extraction network are used for the nearest neighbor (NN) classifier, and then it is applied to the final classification of the unlabeled samples in the target domain.
[0104] Step 3. Classify the input hyperspectral image using the trained few-shot hyperspectral image classification model based on joint domain adaptation weight self-learning.
[0105] It should be noted that all samples in the training set used in the training phase are labeled samples. The training set includes a support set and a query set. Only a part of the labeled samples are used to form the support set as the model input to train the model, so as to predict the query set composed of the remaining labeled samples. When performing few-shot hyperspectral image classification after completing the training of the model, all labeled samples in the training set are used as the model input to predict the query set composed of the remaining unlabeled samples.
[0106] After completing the overall training of the model in Step 2, enter the classification phase of the unlabeled samples of the target domain hyperspectral image. At this time, the weights of the few-shot hyperspectral image classification model based on joint domain adaptation weight self-learning will be fixed, and the knowledge learned during training is used to process the input data, that is, the target domain hyperspectral image. The specific process is as follows:
[0107] First, a small number of labeled samples in the target domain are used as support set samples, and a large number of unlabeled samples to be classified in the target domain are used as query set samples. The support set samples and query set samples are respectively input into the trained embedding feature extraction module to extract their corresponding spectral-spatial discriminant feature representations. Then, for each category, the features of the support set samples of this category are aggregated in the embedding space, and the central feature of this category is represented by calculating its average vector, i.e., the category prototype. Subsequently, for each query set sample to be classified, the model measures the similarity between its embedding representation and the category prototypes based on the Euclidean distance, and classifies this query set sample into the category with the closest distance. Generally speaking, the classification decision is completely based on the relative position relationship between the query set samples and the category prototypes in the embedding space.
[0108] In addition, to verify the effectiveness of the method proposed in the present invention, the following specific experiments are also given.
[0109] 1. Select sample data.
[0110] The Indian pines hyperspectral dataset was taken by the airborne visible infrared imaging spectrometer AVIRIS in the Indian pine forest in the northwest of Indiana, USA. After removing 20 absorption bands (104 - 105, 150 - 163, 220), it contains 200 bands, and the spectral range covers 0.4μm to 2.5μm. The hyperspectral image contains 145×145 pixels and 16 types of target ground objects, including no-till corn, cultivated soybeans, and woods, etc. The dataset contains a total of 21025 pixels, among which the number of target pixels is 10249 and the number of background pixels is 10776.
[0111] 2. Experimental settings.
[0112] In this embodiment, 6 small sample classification methods are selected and compared with the method of the present invention on the Indian pines hyperspectral dataset. Four quantitative evaluation indicators are used for evaluation, including the classification accuracy CA, overall accuracy OA, average accuracy AA, and Kappa coefficient for each specific land cover category, to verify the effectiveness of the invention method. The comparison methods include the nearest neighbor classifier NN, the hybrid 3D - 2D convolution + nearest neighbor classifier HybridSN+NN, the deep cross - domain small sample learning DCFSL, the heterogeneous small sample learning HFSL, the dual - metric multi - stage relationship network DM - MRN, and the hyperspectral small sample method FSCF - SSL based on self - supervised learning.
[0113] To comprehensively evaluate the effectiveness of the proposed hyperspectral image few-shot classification method JDAWSL based on joint domain adaptation weight self-learning in the present invention and to conduct a thorough comparison with other few-shot classification methods, supervised samples are randomly selected in the target domain in the present invention, and all experiments are carried out 10 times to eliminate the influence of random sampling. The higher the mean value of OA in the 10 experiments, the better the classification performance. The present invention uses 9×9 image patches as input, the learning rate is fixed at 0.001, the training is iterated 5000 times, and the Adam optimizer is used to guide the parameter update and optimization. The initial values of α and β are both set to 0.5. The present invention is implemented using a framework based on Python 3.11 and PyTorch 2.2. All experiments are run on an Intel I7-13700 CPU and an Nvidia RTX4090 GPU.
[0114] 3. Experimental results.
[0115] Table 2 Classification results of the Indian Pines dataset (mean ± standard deviation)
[0116]
[0117] To more intuitively compare the performance of each few-shot classification method, Table 2 shows the classification results of different target domain datasets with only 5 labeled samples per class on the Indian Pines dataset. Among them, the few-shot hyperspectral image classification method JDAWSL based on joint domain adaptation weight self-learning in the present invention achieves the best results, reaching 76.67%, 85.48% and 73.67% respectively, which may be due to the highest classification accuracy of JDAWSL in classes 2, 5, 6, 7 and 10. Followed by FSCF-SSL, HFSL and the DCFSL method using traditional domain adaptation DM-MRN performs poorly. The above classification results can verify that the use of the few-shot hyperspectral image classification method based on joint domain adaptation weight self-learning in the present invention has a positive impact.
[0118] Example 2
[0119] This Example 2 describes a computer device, which includes a memory and one or more processors.
[0120] Executable code is stored in the memory, and when the processor executes the executable code, it is used to implement the steps of the few-shot hyperspectral image classification method based on joint domain adaptation weight self-learning in the above Example 1.
[0121] The computer device in this example is any device or apparatus with data processing capabilities, which will not be elaborated here.
[0122] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to the enumerated embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.
Claims
1. Hyperspectral image small-sample classification method based on joint domain adaptation weight self-learning, characterized in that It includes the following steps: Step 1. Build a small-sample classification model for hyperspectral images based on joint domain adaptation weight self-learning, which includes the following structure: An embedded feature extraction module, which includes a mapping layer and a feature extraction network; the mapping layer is used to ensure that the data dimensions of the source domain and target domain hyperspectral images input into the feature extraction network are consistent; the feature extraction network extracts spectral-spatial discriminant features of the source domain and target domain through a spectral-spatial attention double-branch dense network; A small-sample learning module, which is used to perform small-sample learning of the source domain and target domain according to the distances between the support set and query set samples in the embedded space, and obtain the class distribution of the query set samples; A joint domain adaptation module, which includes a domain discriminator and a domain projector; the domain projector is used to enhance the common features of the source domain and target domain, and the domain discriminator is used to suppress the differential features of the source domain and target domain; A weight self-learning module, which is used to adjust the weight between the domain discrimination loss and the domain projection loss through an adaptive learner; A multi-label classification module, which is used to classify the unlabeled samples in the target domain according to the spectral-spatial discriminant features output by the embedded feature extraction module; Step 2. Construct a total loss function by combining the small-sample learning loss and domain adaptation loss of the source domain and target domain, and train the small-sample classification model for hyperspectral images based on joint domain adaptation weight self-learning; Step 3. Use the trained small-sample classification model for hyperspectral images based on joint domain adaptation weight self-learning to classify the input target domain hyperspectral images.
2. The small-sample classification method for hyperspectral images based on self-learning of joint domain adaptation weights according to claim 1, wherein In the above Step 1, The mapping layer is implemented based on a two-dimensional convolutional neural network and is used to make the number of bands of the hyperspectral images in the source domain and target domain equal; The processing flow of the signal in the mapping layer is as follows: Assume that the hyperspectral image of the input mapping layer is cube I, where w×w is the size of the cube, and L represents the number of bands. Then the characteristic output of the mapping layer is: I′ = I × T; where T represents the mapping transformation matrix, l is the number of channels after the mapping transformation; I′ is the hyperspectral image after the mapping transformation, There are L×l learnable parameters in the mapping transformation; Subsequently, I′ is used as the input of the batch normalization layer, and the batch normalization layer is used to make all reflectance values of the source domain and target domain hyperspectral images after the mapping transformation between 0 and 1.
3. The small-sample classification method for hyperspectral images based on self-learning of joint domain adaptation weights according to claim 1, characterized in that In the above Step 1, the processing flow of the signal in the feature extraction network is as follows: The feature extraction network extracts spectral features and spatial features from the hyperspectral images output by the mapping layer respectively, and obtains spectral weights and spatial weights through an attention mechanism, and then performs weighted fusion on the spectral features and spatial features to obtain spectral-spatial discriminant features of the source domain and target domain.
4. The small-sample classification method for hyperspectral images based on self-learning of joint domain adaptation weights according to claim 1, wherein In the above Step 1, the processing flow of the signal in the small-sample learning module is as follows: Select K samples from various labeled samples in the source domain s to form a source domain support set, and then select N samples from the remaining various labeled samples in the source domain s to form a source domain query set; select K samples from various labeled samples in the target domain t to form a target domain support set, and then select N samples from the remaining various labeled samples in the target domain t to form a target domain query set; Feature aggregation is performed on the labeled samples of each class in the source domain support set and target domain support set in the embedded space learned by the embedded feature extraction module, and the average vector, that is, the class prototype, is calculated. The class prototype is used to represent the central feature of each class in the support set samples; Then, similarity measurement based on the Euclidean distance is performed between the embedded representations of the source domain query set samples and target domain query set samples and each class prototype, and thus the class distribution of the query set samples is obtained.
5. The small-sample classification method for hyperspectral images based on self-learning of joint domain adaptation weights according to claim 1, characterized in that, In the above Step 1, The processing flow of the signal in the domain projector is as follows: The output of the embedded feature extraction module is used as the input of the domain projector. The input feature vector of the domain projector is processed through a fully connected layer, a ReLU layer, a Dropout layer and a Sigmoid layer to project the source domain and target domain data into a new nominal domain; The processing flow of the signal in the domain discriminator is as follows: Result of selecting random bilinear mapping As the input of the domain discriminator, where g and Are respectively the domain variances of the feature extraction network and the linear classifier in the embedded feature extraction module; The feature vectors in the input domain discriminator are processed through a fully connected layer, a ReLU layer, and a Dropout layer to enhance the non-linear expression ability and prevent overfitting; then, the feature vectors are dimensionally reduced through a fully connected layer and normalized through a Softmax layer, and finally, the probability distributions of the samples in the source domain and the target domain are obtained.
6. The small-sample classification method for hyperspectral images based on self-learning of joint domain adaptation weights according to claim 1, characterized in that In the said step 1, the processing flow of the signal in the weight self-learning module is as follows: Definition and where α and β represent the weights of the domain discrimination loss and the domain projection loss, L dp represents the domain projection loss, L dd represents the domain discrimination loss, and η represents the learning rate; After normalization, the values of α and β are obtained as:
7. The small-sample classification method for hyperspectral images based on self-learning of joint domain adaptation weights according to claim 1, characterized in that, In the said step 1, The multi-label classification module adopts a nearest neighbor classifier.
8. The small-sample classification method for hyperspectral images based on joint domain adaptation weight self-learning according to claim 4, characterized in that In the said step 2, Source set Supported by the source domain support set S s and the source domain query set Q s constitute the source set The small-sample learning loss of the is expressed as: where x and y represent the sample and its corresponding label, and k represents the true class, denotes the feature vector in the embedding space belonging to the true class k; Target set Supported by the target domain support set S t And the target domain query set Q t Composed, the target set The small sample learning loss of Is expressed as: Domain projection loss L dp is defined as: Among them, is the source domain projection loss, is the target domain projection loss, y o represents the domain label, p o (·) represents the probability that a sample belongs to the new nominal domain; represents the i-th s sample in the source set, i s ∈ [1, n s , n s represents the number of samples in the source set; represents the i-th t sample in the target set, i t ∈ [1, n t , n t represents the number of samples in the target set; x j represents the j-th sample mapped into the new nominal domain, j ∈ [1, n s + n t ; Domain discrimination loss L dd is defined as: Among them, is the source domain discriminative loss, is the target domain discriminative loss, M(·) represents a random bilinear mapping, represents any source set, represents any target set, p d and 1 - p d respectively represent the probabilities that the domain discriminator predicts the input sample as the source domain and the target domain; Define the total loss function L total as follows: Among them, represents the total source domain loss, represents the total target domain loss; the total source domain loss is defined as: Among them, the source domain adaptation loss is defined as: Combined with the target domain few-shot learning loss and the domain adaptation loss The total loss of the target domain is defined as: Among them, the target domain adaptation loss is defined as:
9. The small-sample classification method for hyperspectral images based on self-learning of joint domain adaptation weights according to claim 1, characterized in that The said step 3 is specifically: All the labeled samples and the unlabeled samples to be classified in the target domain are respectively input into the trained embedded feature extraction module to extract spectral-spatial discriminant features, and an embedding space is learned through the embedded feature extraction module; Feature aggregation is performed on the labeled samples of each category in the target domain in the embedding space, and the prototypes of each category are calculated respectively; For each unlabeled sample to be classified, the multi-label classification module measures the similarity by calculating the Euclidean distance between the embedding representation of the unlabeled sample and the prototypes of each category, and then, according to the principle of the minimum distance, assigns the unlabeled sample to the category with the closest distance to it.
10. A computer device, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the said processor executes the said executable code, the steps of the hyperspectral image small sample classification method based on joint domain adaptation weight self-learning as described in any one of claims 1 to 9 are implemented.
Citation Information
Cited By
Hyperspectral image classification method and system based on spectrum-spatial diffusion generation
CN121121306A