A Cross-Domain Few-Shot Hyperspectral Image Classification Method Based on Semantic Consistency-Oriented Prototype Networks

CN122676211APending Publication Date: 2026-09-01KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610493840.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-15
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

但现有方法多采用静态原型构建策略,对采样偏差敏感,难以适配任务特定语义;且原型通常在固定或孤立的特征空间中学习,对域偏移与语义差异的鲁棒性不足

Benefits of technology

[0044]1、本发明提出了一种基于语义一致性导向原型网络的跨域少样本高光谱图像分类方法,通过构建原型驱动的跨域精炼机制与特征校准机制,有效融合源域与目标域信息,实现类别判别性与跨域语义一致性的统一提升。相比传统方法,本发明能够在极少标注样本条件下显著增强模型的跨域泛化能力与分类精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676211A_ABST
    Figure CN122676211A_ABST
Patent Text Reader

Abstract

This invention discloses a cross-domain few-shot hyperspectral image classification method based on a semantic consistency-guided prototype network, comprising: determining the dataset: selecting several publicly available hyperspectral image datasets; hyperspectral data preprocessing; network construction: constructing a semantic consistency-guided cross-domain few-shot hyperspectral image classification prototype network; network training: inputting all samples from the source domain and training samples from the target domain into the constructed network for training; sample classification: after completing a fixed number of training iterations, inputting test samples from the target domain into the semantic consistency-guided cross-domain few-shot hyperspectral image classification prototype network to obtain the classification result. This invention can fully exploit the spectral-spatial intrinsic features of hyperspectral images, effectively mitigate the impact of cross-domain distribution differences and sample scarcity, achieve high-precision identification of target domain land cover categories under few-shot conditions, and is suitable for hyperspectral image land cover classification tasks in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a cross-domain few-shot hyperspectral image classification method based on semantic consistency-guided prototype networks, and particularly to a cross-domain few-shot hyperspectral image classification method related to remote sensing technology, belonging to the field of image processing technology. Background Technology

[0002] Hyperspectral images provide detailed representations of ground features within a continuous high-dimensional spectral space, offering rich and discriminative information for remote sensing tasks such as land cover identification, environmental monitoring, agricultural assessment, and disaster analysis. Compared to natural images, hyperspectral images possess higher spectral resolution, enabling more accurate differentiation of land cover categories with highly similar spectral characteristics. However, the inherent high dimensionality and strong spectral-spatial correlation of hyperspectral images necessitate that classification models rely on large-scale, high-quality labeled samples for stable operation. In practical applications, the high cost of manual annotation, the requirement for specialized domain knowledge, and the complex and variable imaging conditions result in an extreme scarcity of labeled samples across different scenarios, severely limiting the large-scale deployment of hyperspectral image classification models.

[0003] In recent years, deep learning technology has driven significant progress in hyperspectral image classification. Convolutional neural networks, attention mechanisms, and Transformers have been widely used for end-to-end joint modeling of spectral features and spatial context information, effectively improving feature discriminativeness and classification performance. However, these models generally have a large number of parameters and are highly dependent on labeled data. In scenarios with extremely limited labeled samples, they are prone to overfitting, resulting in a significant decrease in generalization ability.

[0004] To address the problem of scarce labeled samples, classification of low-sample hyperspectral images has become a research hotspot. Existing methods improve feature robustness and discriminativeness under limited supervision through self-supervised contrastive learning, self-pooling Transformers, multi-view calibration prototype learning, and adaptive subspace modeling. However, these methods typically assume that training and testing data follow the same distribution. When significant domain shifts occur due to differences in sensors, scenes, and imaging conditions, the model's applicability drops sharply, making it difficult to meet practical remote sensing needs.

[0005] To address this, cross-domain few-sample hyperspectral image classification has been proposed, aiming to transfer transferable knowledge from a well-labeled source domain to achieve accurate classification in a target domain with scarce labeling. Existing methods alleviate the cross-domain distribution mismatch problem to some extent through strategies such as domain bridging, bias reduction, decoupled knowledge distillation, dual-prototype learning, global-local graph attention, and multimodal prototype correction. However, in real-world scenarios, the distribution of target domain data is difficult to estimate accurately, and over-reliance on insufficiently calibrated source domain representations leads to instability in category-level semantic representations, ultimately limiting the model's discriminative ability and cross-domain adaptive performance. Furthermore, prototype-based representation learning is widely used in hyperspectral image classification to capture category-level semantics and reduce the impact of noise and sample imbalance. However, existing methods often employ static prototype construction strategies, which are sensitive to sampling bias and difficult to adapt to task-specific semantics; moreover, prototypes are usually learned in fixed or isolated feature spaces, lacking robustness to domain shifts and semantic differences.

[0006] Against this background, in order to achieve fast and accurate hyperspectral pixel classification under domain offset and few-sample constraints, this invention proposes a simple and effective cross-domain few-sample hyperspectral image classification method. Summary of the Invention

[0007] This invention proposes a cross-domain few-sample hyperspectral image classification method based on a semantic consistency-guided prototype network. It can simultaneously achieve class discriminativeness, cross-domain semantic consistency, and domain adaptation capability, effectively alleviate the distribution offset and semantic differences between the source and target domains, stabilize the target domain feature structure under few-sample annotation conditions, and improve cross-domain generalization performance, thus achieving high-precision and robust cross-domain few-sample hyperspectral image classification.

[0008] The technical solution of this invention is: a cross-domain few-shot hyperspectral image classification method based on semantic consistency-guided prototype networks, the specific steps of which are as follows:

[0009] Step 1: Prepare several general and publicly available hyperspectral image datasets for network training, with one dataset as the source domain and the others as the target domains.

[0010] Step 2: Preprocess the hyperspectral image, then extract hyperspectral image patches centered on each pixel, and divide the target domain hyperspectral image patches into a non-overlapping hyperspectral training sample set and a hyperspectral test sample set;

[0011] Step 3: Construct a prototype network for cross-domain few-sample hyperspectral image classification based on semantic consistency. The entire network consists of a feature extractor, a prototype-driven cross-domain refinement module, a target domain prototype calibration branch, and a multi-angle data augmentation module. The prototype-driven cross-domain refinement module consists of a bias-aware prototype modeling submodule and a target-guided prototype attention submodule. The target domain prototype calibration branch consists of two progressive feature calibration modules connected in series.

[0012] Step 4: Train the prototype network for cross-domain few-sample hyperspectral image classification based on semantic consistency using the source domain sample set and the target domain training sample set.

[0013] Step 5: Input the target domain test sample set into the trained prototype network for cross-domain few-shot hyperspectral image classification based on semantic consistency guidance to obtain the category label of each pixel in the test sample, thus completing the hyperspectral image classification.

[0014] As a further aspect of the present invention, the specific steps of Step 2 are as follows:

[0015] Step 2.1: Hyperspectral images have three-dimensional characteristics, and their data are represented as S∈R H×W×C Fill the edges of the original hyperspectral image with pixels of size n and a pixel value of 0; extract hyperspectral image blocks from the filled image.

[0016] Step 2.2: Classify the hyperspectral image patch into the category set to which the image patch belongs based on the category of its center pixel;

[0017] Step 2.3: All source domain samples are used for model training; a fixed number of samples from each class of the target domain dataset are selected as training samples, and the remaining samples are test samples.

[0018] As a further aspect of the present invention, the specific steps of Step 3 are as follows:

[0019] Step 3.1: Construct a feature extractor consisting of a 3D convolutional layer, a normalization layer, a ReLU activation function layer, and an average pooling layer connected in series.

[0020] Step 3.2: Construct a prototype-driven cross-domain refining module consisting of a deviation-aware prototype modeling submodule and a goal-guided prototype attention submodule;

[0021] Step 3.3: Construct a target domain prototype calibration branch consisting of two progressive feature calibration modules connected in series;

[0022] Step 3.4: Construct a multi-angle data augmentation module consisting of neighboring pixel masking, spectral resampling, and mixed noise enhancement operations.

[0023] As a further aspect of the present invention, the specific steps of Step 3.2 are as follows:

[0024] Step 3.2.1: Construct a deviation-sensing prototype modeling submodule consisting of parallel spectral head branches and spatial head branches;

[0025] Step 3.2.2: Construct a goal-guided prototype attention submodule consisting of a self-attention mechanism.

[0026] As a further embodiment of the present invention, in Step 3.1, the feature extractor includes a three-dimensional convolutional block and an average pooling layer, and the structure of the convolutional block is fixed as: three-dimensional convolutional layer → normalization layer → activation function layer;

[0027] The kernel size of the 3D convolutional layer is set to 3×3×3, and the activation function of each activation layer is set to the ReLU activation function.

[0028] As a further embodiment of the present invention, in Step 3.2, the prototype-driven cross-domain refinement module is composed of a deviation-aware prototype modeling submodule and a target-guided prototype attention submodule, and its fixed structure is: deviation-aware prototype modeling submodule → target-guided prototype attention submodule;

[0029] The deviation-aware prototype modeling submodule consists of parallel spectral head branches and spatial head branches. The structure of both branches is fixed as: linear layer → activation function layer → linear layer → activation function layer.

[0030] The target-guided prototype attention submodule is composed of a self-attention mechanism. In addition, the output feature map of the self-attention mechanism is element-wise added to the output of the deviation-aware prototype modeling submodule in a skip connection manner.

[0031] As a further embodiment of the present invention, in Step 3.3, the target domain prototype calibration branch is composed of two progressive feature calibration modules connected in series, and its fixed structure is: progressive feature calibration module → progressive feature calibration module.

[0032] The progressive feature calibration module consists of a Softmax layer and a bidirectional soft matching calibration layer, and its structure is fixed as follows: Softmax layer → bidirectional soft matching calibration layer.

[0033] As a further aspect of the present invention, in Step 4:

[0034] The network parameters are updated using stochastic gradient descent. The total loss is calculated through weighted joint optimization with multiple losses. The total loss and each sub-loss are expressed by the following formula:

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041] in Indicates the total loss. This represents the balancing weight coefficient for each loss term; This represents the few-shot classification loss based on prototype metric learning. Indicates the query sample. Indicates querying sample tags, Represents sample embedding features. Indicates the first Class prototype, Indicates Euclidean distance; This represents the source domain prototype consistency regularization loss. This represents the optimized source domain prototype. Represents the static prototype of the source domain; This represents the target domain feature alignment loss. , They represent the first After the next iteration, the target domain supports set and query set features. Indicates the total number of iterations; This indicates the learning loss in supervised comparison. Indicates sample features, Indicates similar enhanced viewpoint features, Indicates the temperature coefficient; This represents the cross-domain prototype contrast loss. This represents the prototype of the target domain.

[0042] The proposed cross-domain few-shot hyperspectral image classification prototype network based on semantic consistency consists of a source domain data branch and a target domain data branch, which operate in parallel. In the source domain data branch, hyperspectral image patches from the source domain are first input into a feature extractor for feature extraction. Then, static and dynamic prototypes are constructed based on the extracted features. The static prototype is obtained from global samples in the source domain, while the dynamic prototype is calculated from the support set samples of the current task. Subsequently, the static and dynamic prototypes are input into the prototype-driven cross-domain refinement module. In this module, a deviation-aware prototype modeling submodule, composed of a spectral head and a spatial head, corrects prototype deviations. Further, the corrected source-domain prototype and target-domain prototype are input into the target-guided prototype attention submodule. A self-attention mechanism enables cross-domain semantic information interaction and enhancement, resulting in a source-domain refined prototype with target-domain adaptability, ultimately used for classification of source-domain samples. In the target-domain data branch, the target-domain hyperspectral image is first input into the multi-angle data augmentation module. Multi-view augmented samples are generated through operations such as neighborhood pixel masking, spectral band resampling, and noise perturbation. Each augmented sample is then input into a feature extractor to obtain feature representations. Subsequently, the target-domain support set features and query set features are input into the progressive feature calibration module. A bidirectional soft-matching mechanism based on attention iterates through multiple rounds to progressively align and optimize the target-domain feature distribution. Finally, the calibrated features are used to construct a target-domain category prototype, and the target-domain query samples are classified based on a prototype metric learning method.

[0043] The beneficial effects of this invention are:

[0044] 1. This invention proposes a cross-domain few-sample hyperspectral image classification method based on a semantic consistency-driven prototype network. By constructing a prototype-driven cross-domain refinement mechanism and feature calibration mechanism, it effectively fuses source and target domain information, achieving a unified improvement in class discriminativeness and cross-domain semantic consistency. Compared with traditional methods, this invention can significantly enhance the model's cross-domain generalization ability and classification accuracy under conditions of very few labeled samples.

[0045] 2. This invention designs a prototype-driven cross-domain refinement module, which dynamically corrects and semantically enhances the source domain prototype through the synergistic effect of bias-aware prototype modeling and target-guided prototype attention mechanism. On the one hand, by jointly modeling static and dynamic prototypes, it effectively alleviates sampling bias and task bias problems under few-sample conditions; on the other hand, by introducing the target domain prototype as a semantic anchor, cross-domain semantic alignment is achieved through a self-attention mechanism, enabling the source domain prototype to have stronger target domain adaptability, thereby significantly improving cross-domain classification performance.

[0046] 3. This invention proposes a progressive feature calibration mechanism, which achieves gradual alignment and optimization of the target domain feature distribution through bidirectional soft matching and iterative updates between the support set and the query set. This mechanism can effectively improve intra-class compactness and inter-class separability, alleviate the problem of target domain feature instability, and thus enhance the model's discriminative ability and robustness under conditions of few samples.

[0047] 4. This invention constructs a multi-angle data augmentation module that expands the target domain samples from three dimensions: spatial structure, spectral characteristics, and noise perturbation. It can enrich the data distribution without additional annotation, reduce the risk of overfitting, and improve the model's adaptability to complex imaging conditions and domain shifts. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of the overall process of this method;

[0049] Figure 2 This is a diagram of the prototype-driven cross-domain refining module in this method;

[0050] Figure 3 This is a diagram illustrating the progressive feature calibration module in this method;

[0051] Figure 4 This is a diagram illustrating the multi-angle data augmentation module in this method;

[0052] Figure 5 The figures show the classification results of this method and two other state-of-the-art methods on the Pavia University dataset.

[0053] Figure 6 The figures show the classification results of this method and two other advanced methods on the WHU-Hi-LongKou dataset. Detailed Implementation

[0054] The embodiments and effects of the present invention will be further described below with reference to the accompanying drawings.

[0055] Example 1, referring to Figure 1 A cross-domain few-shot hyperspectral image classification method based on semantic consistency-guided prototype networks, the specific steps of which are as follows:

[0056] Step 1: Prepare several general and publicly available hyperspectral image datasets for network training, with one dataset as the source domain and the others as the target domains.

[0057] Step 2: Preprocess the hyperspectral image, then extract hyperspectral image patches centered on each pixel, and divide the target domain hyperspectral image patches into a non-overlapping hyperspectral training sample set and a hyperspectral test sample set;

[0058] Step 3: Construct a prototype network for cross-domain few-sample hyperspectral image classification based on semantic consistency. The entire network consists of a feature extractor, a prototype-driven cross-domain refinement module, a target domain prototype calibration branch, and a multi-angle data augmentation module. The prototype-driven cross-domain refinement module consists of a bias-aware prototype modeling submodule and a target-guided prototype attention submodule. The target domain prototype calibration branch consists of two progressive feature calibration modules connected in series.

[0059] Step 4: Train the prototype network for cross-domain few-sample hyperspectral image classification based on semantic consistency using the source domain sample set and the target domain training sample set.

[0060] Step 5: Input the target domain test sample set into the trained prototype network for cross-domain few-shot hyperspectral image classification based on semantic consistency guidance to obtain the category label of each pixel in the test sample, thus completing the hyperspectral image classification.

[0061] Furthermore, the specific steps of Step 2 are as follows:

[0062] Step 2.1: Hyperspectral images have three-dimensional characteristics, and their data can be represented as S∈R H×W×C Where H and W represent the length and width in the hyperspectral image space, and C represents the number of hyperspectral image channels; the edges of the original hyperspectral image are filled with pixels of size n and a pixel value of 0; hyperspectral image patches are extracted from the filled image, and their data can be represented as X∈R P×P×C , where P represents the size of the extracted hyperspectral image patch, that is, a hyperspectral image patch with a spatial size of (2n+1)×(2n+1) and a channel number of C is selected with each original pixel as the center, where in this example n=3 but is not limited to n=3.

[0063] Step 2.2: Classify the hyperspectral image patch into the category set to which the image patch belongs based on the category of its center pixel;

[0064] Step 2.3: All source domain samples are used for model training; for each class of the target domain dataset, 5 samples are randomly selected as training samples, and the remaining samples are test samples.

[0065] Furthermore, the specific steps of Step 3 are as follows:

[0066] Step 3.1: Build the feature extractor;

[0067] This feature extractor consists of a 3D convolutional layer, a normalization layer, a ReLU activation function layer, and an average pooling layer connected in series, wherein:

[0068] The feature extractor includes a three-dimensional convolutional block and an average pooling layer. The structure of the convolutional block is fixed as follows: three-dimensional convolutional layer → normalization layer → activation function layer.

[0069] The kernel size of the 3D convolutional layer is set to 3×3×3, and the activation function of each activation layer is set to the ReLU activation function.

[0070] Step 3.2: Build a prototype-driven cross-domain refining module;

[0071] The prototype-driven cross-domain refinement module consists of a deviation-aware prototype modeling submodule and a goal-guided prototype attention submodule, with a fixed structure of: deviation-aware prototype modeling submodule → goal-guided prototype attention submodule.

[0072] Furthermore, the specific steps of Step 3.2 are as follows:

[0073] Step 3.2.1: Build the deviation perception prototype modeling submodule;

[0074] Reference Figure 2 The deviation-aware prototype modeling submodule consists of parallel spectral head branches and spatial head branches. The structure of both branches is fixed as: linear layer → activation function layer → linear layer → activation function layer; the inputs to both branches include: source domain static prototype. Source Domain Dynamic Prototype and prototype deviation The prototype deviation is expressed as:

[0075]

[0076] in It is the source domain identifier. Indicates the first One category;

[0077] The source domain prototype representation obtained by spectral head branching and spatial head branching is as follows:

[0078]

[0079]

[0080]

[0081] in, , , , Both represent linear layer weight matrices. This indicates a splicing operation. Represents the ReLU activation function. and These are learnable weight parameters.

[0082] Step 3.2.2: Build the goal-guided prototype attention submodule;

[0083] Reference Figure 2 The target-guided prototype attention submodule is composed of a self-attention mechanism, and its input Key and Value are obtained through the source domain prototype. and target domain prototype The process is represented as follows:

[0084]

[0085] in It is the target domain identifier;

[0086] The query, on the other hand, uses the source domain prototype. and target domain prototype The sum of the pooling results can be represented as:

[0087]

[0088] in Indicates pooling operation;

[0089] The target-guided prototype attention submodule outputs a jump connection to the source domain prototype. The modified prototype can be obtained by adding elements together. , can be represented as:

[0090]

[0091] in This represents the dimension of the feature vector.

[0092] Step 3.3: Construct the target domain prototype calibration branch;

[0093] The target domain prototype calibration branch consists of two progressive feature calibration modules connected in series, with a fixed structure of: progressive feature calibration module → progressive feature calibration module.

[0094] Furthermore, the specific steps of Step 3.3 are as follows:

[0095] Step 3.3.1: Build a progressive feature calibration module;

[0096] Reference Figure 3 The progressive feature calibration module consists of a Softmax layer and a bidirectional soft matching calibration layer, with a fixed structure: Softmax layer → bidirectional soft matching calibration layer. This module first obtains the target domain support set through matrix multiplication and the Softmax layer. and query set Similarity matrix , means as follows:

[0097]

[0098] in, It is the target domain identifier. It supports set identifiers. It is the query set identifier. The dimension of the feature vector. express function;

[0099] Subsequently, bidirectional soft matching calibration operations are performed, primarily using the support set and query set, respectively. The results are then used as the new target domain support set and query set. This process can be represented as follows:

[0100]

[0101]

[0102] in, Indicates the current iteration number. and The first Features of the support set and query set after the next iteration This is the similarity matrix corresponding to the number of iterations. and This represents the learnable weight parameters. This indicates a two-way soft-match calibration operation; in this embodiment... Take 2.

[0103] Step 3.4: Build a multi-angle data augmentation module;

[0104] This multi-angle data augmentation module consists of parallel neighbor pixel masking, spectral band resampling, and hybrid noise enhancement operations, and is used to expand the target domain data from three dimensions: spatial structure, spectral characteristics, and noise perturbation.

[0105] Furthermore, the specific steps of Step 3.4 are as follows:

[0106] Step 3.4.1, Neighborhood pixel masking operation;

[0107] Reference Figure 4 The neighborhood pixel masking operation is used to simulate spatial structure perturbations. By randomly occluding the local neighborhood of the input hyperspectral image, the robustness of the model to spatial information loss is enhanced. Let the target domain input sample be:

[0108]

[0109] in, Indicates the number of spectral segments. and These represent the height and width of the space, respectively.

[0110] Construct a random mask matrix The mask value of 0 indicates that the pixel at that location is occluded, and 1 indicates that it is preserved. The enhanced sample after masking the neighboring pixels is represented as follows:

[0111]

[0112] in, This indicates element-wise multiplication.

[0113] Step 3.4.2: Band resampling operation;

[0114] Reference Figure 4 The spectral resampling operation is used to simulate spectral response changes under different sensor or imaging conditions. By randomly sampling or interpolating the spectral dimensions, it enhances the model's adaptability to spectral shifts. Let the input sample be... The band resampling operation is represented as:

[0115]

[0116] in, This represents a mapping function from the original band index to the resampled band index.

[0117] Step 3.4.3: Hybrid noise enhancement operation;

[0118] Reference Figure 4 The hybrid noise enhancement operation includes two steps: random pruning and noise perturbation. First, a random pruning operation is performed on the input samples:

[0119]

[0120] in, This represents a random pruning function used to prune samples of size 1 from the original sample. Local area;

[0121] Subsequently, Gaussian noise was superimposed on the cropped sample:

[0122]

[0123] in, This indicates that the mean is 0 and the variance is 0. Gaussian distributed random noise.

[0124] The proposed cross-domain few-shot hyperspectral image classification prototype network based on semantic consistency consists of a source domain data branch and a target domain data branch, which operate in parallel. In the source domain data branch, hyperspectral image patches from the source domain are first input into a feature extractor for feature extraction. Then, static and dynamic prototypes are constructed based on the extracted features. The static prototype is obtained from global samples in the source domain, while the dynamic prototype is calculated from the support set samples of the current task. Subsequently, the static and dynamic prototypes are input into the prototype-driven cross-domain refinement module. In this module, a deviation-aware prototype modeling submodule, composed of a spectral head and a spatial head, corrects prototype deviations. Further, the corrected source-domain prototype and target-domain prototype are input into the target-guided prototype attention submodule. A self-attention mechanism enables cross-domain semantic information interaction and enhancement, resulting in a source-domain refined prototype with target-domain adaptability, ultimately used for classification of source-domain samples. In the target-domain data branch, the target-domain hyperspectral image is first input into the multi-angle data augmentation module. Multi-view augmented samples are generated through operations such as neighborhood pixel masking, spectral band resampling, and noise perturbation. Each augmented sample is then input into a feature extractor to obtain feature representations. Subsequently, the target-domain support set features and query set features are input into the progressive feature calibration module. A bidirectional soft-matching mechanism based on attention iterates through multiple rounds to progressively align and optimize the target-domain feature distribution. Finally, the calibrated features are used to construct a target-domain category prototype, and the target-domain query samples are classified based on a prototype metric learning method.

[0125] Furthermore, in Step 4:

[0126] The network parameters are updated using stochastic gradient descent. The total loss is calculated through weighted joint optimization with multiple losses. The total loss and each sub-loss are expressed by the following formula:

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133] in Indicates the total loss. This represents the balancing weight coefficient for each loss term; This represents the few-shot classification loss based on prototype metric learning. Indicates the query sample. Indicates querying sample tags, Represents sample embedding features. Indicates the first Class prototype, Indicates Euclidean distance; This represents the source domain prototype consistency regularization loss. This represents the optimized source domain prototype. Represents the static prototype of the source domain; This represents the target domain feature alignment loss. , They represent the first After the next iteration, the target domain supports set and query set features. Indicates the total number of iterations; This indicates the learning loss in supervised comparison. Indicates sample features, Indicates similar enhanced viewpoint features, Indicates the temperature coefficient; This represents the cross-domain prototype contrast loss. This represents the prototype of the target domain.

[0134] Furthermore, in this embodiment, the learning rate of the network is set to 0.001, the batch size is 64, and the training is performed for 5000 iterations to obtain the final network's paradoxical parameters and weight file.

[0135] The hardware platform for the simulation experiment in this embodiment is: a 12th Gen Intel(R) Core(TM) i9-12900KF CPU and an NVIDIA GeForce RTX 3090 GPU with 24GB of memory; the operating system is Ubuntu 20.04.5LTS, and the virtual environment configured includes: Python 3.9, PyTorch 1.13.1, CUDA 11.6, etc.

[0136] Furthermore, the specific steps of Step 5 are as follows:

[0137] The target domain test sample set is fed into the hyperspectral classification network that has been trained above to calculate three general evaluation metrics: overall classification accuracy (OA), average accuracy (AA), and Kappa coefficient (K). The larger these three metrics are, the better the classification performance.

[0138] To evaluate the effectiveness of this method, two existing state-of-the-art methods, DPL-MSF and CDFS-CASCL, were used to classify ground objects in two public hyperspectral datasets, Pavia University and WHU-Hi-LongKou.

[0139] The DPL-MSF method mentioned above refers to the hyperspectral classification method proposed by Y.Li et al. in "Dual-Prototype Learning With Multisemantic Fusion for Cross-Domain Few-Shot Hyperspectral Image Classification", abbreviated as DPL-MSF;

[0140] The CDFS-CASCL method refers to the hyperspectral classification method proposed by Z.Li et al. in "Cross-domain few-shot hyperspectral image classification with cross-modal alignment and supervised contrastive learning", abbreviated as CDFS-CASCL.

[0141] Table 1 Comparison of classification results of the three networks on two datasets.

[0142] Experiments on two mainstream datasets demonstrate that the proposed method outperforms state-of-the-art cross-domain few-shot hyperspectral image classification methods, and can more accurately predict the pixel sample category of the target domain hyperspectral image.

[0143] Figure 5 (a) is a diagram showing the classification results of the DPL-MSF method on the Pavia University dataset;

[0144] Figure 5 (b) is a diagram showing the classification results of the CDFS-CASCL method on the Pavia University dataset;

[0145] Figure 5 (c) The classification results of this method on the Pavia University dataset are shown in the figure;

[0146] Figure 6 (a) is a diagram showing the classification results of the DPL-MSF method on the WHU-Hi-LongKou dataset;

[0147] Figure 6 (b) is a diagram showing the classification results of the CDFS-CASCL method on the WHU-Hi-LongKou dataset;

[0148] Figure 6 (c) is a diagram showing the classification results of this method on the WHU-Hi-LongKou dataset;

[0149] As can be clearly seen from the figure, this method has the fewest misclassified pixels and achieves good classification results in noisy classes and boundary regions, demonstrating strong classification capabilities compared to other methods.

[0150] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A cross-domain few-shot hyperspectral image classification method based on a semantic consistency oriented prototype network, characterized in that: Includes the following steps: Step 1: Prepare several general and publicly available hyperspectral image datasets for network training, with one dataset as the source domain and the others as the target domains. Step 2: Preprocess the hyperspectral image, then extract hyperspectral image patches with each pixel as the center point, and divide the target domain hyperspectral image patches into non-overlapping target domain training sample sets and target domain test sample sets. Step 3: Construct a prototype network for cross-domain few-sample hyperspectral image classification based on semantic consistency. The overall framework includes a feature extractor, a prototype-driven cross-domain refinement module, a target domain prototype calibration branch, and a multi-angle data augmentation module. The prototype-driven cross-domain refinement module consists of a bias-aware prototype modeling submodule and a target-guided prototype attention submodule. The target domain prototype calibration branch consists of two progressive feature calibration modules connected in series. Step 4: Train the prototype network for cross-domain few-sample hyperspectral image classification based on semantic consistency using the source domain sample set and the target domain training sample set. Step 5: Input the target domain test sample set into the trained prototype network for cross-domain few-shot hyperspectral image classification based on semantic consistency guidance to obtain the category label of each pixel in the test sample, thus completing the hyperspectral image classification.

2. The cross-domain few-shot hyperspectral image classification method based on the semantic consistency guided prototype network according to claim 1, characterized in that: The specific steps of Step 2 are as follows: Step2.1, Hyperspectral image has three-dimensional characteristics, and its data is represented as S ∈ R H×W×C Step2.2, Fill the pixels with size n and pixel value 0 around the edges of the original hyperspectral image; extract the hyperspectral image block from the filled image; Step 2.2: Classify the hyperspectral image patch into the category set to which the image patch belongs based on the category of its center pixel; Step 2.3: All source domain samples are used for model training; a fixed number of samples from each class of the target domain dataset are selected as training samples, and the remaining samples are test samples.

3. The cross-domain few-shot hyperspectral image classification method based on semantic consistency-guided prototyping networks according to claim 1, characterized in that: The specific steps of Step 3 are as follows: Step 3.1: Construct a feature extractor consisting of a 3D convolutional layer, a normalization layer, a ReLU activation function layer, and an average pooling layer connected in series. Step 3.2: Construct a prototype-driven cross-domain refining module consisting of a deviation-aware prototype modeling submodule and a goal-guided prototype attention submodule; Step 3.3: Construct a target domain prototype calibration branch consisting of two progressive feature calibration modules connected in series; Step 3.4: Construct a multi-angle data augmentation module consisting of neighboring pixel masking, spectral resampling, and mixed noise enhancement operations.

4. The cross-domain few-shot hyperspectral image classification method based on semantic consistency-guided prototyping networks according to claim 3, characterized in that: The specific steps of Step 3.2 are as follows: Step 3.2.1: Construct a deviation-sensing prototype modeling submodule consisting of parallel spectral head branches and spatial head branches; Step 3.2.2: Construct a goal-guided prototype attention submodule consisting of a self-attention mechanism.

5. The cross-domain few-shot hyperspectral image classification method based on semantic consistency-guided prototyping networks according to claim 3, characterized in that: In Step 3.1, the feature extractor includes a three-dimensional convolutional block and an average pooling layer. The structure of the convolutional block is fixed as follows: three-dimensional convolutional layer → normalization layer → activation function layer. The kernel size of the 3D convolutional layer is set to 3×3×3, and the activation function of each activation layer is set to the ReLU activation function.

6. The cross-domain few-shot hyperspectral image classification method based on semantic consistency-guided prototyping networks according to claim 3, characterized in that: In Step 3.2, the prototype-driven cross-domain refinement module consists of a deviation-aware prototype modeling submodule and a target-guided prototype attention submodule, with a fixed structure of: deviation-aware prototype modeling submodule → target-guided prototype attention submodule; The deviation-aware prototype modeling submodule consists of parallel spectral head branches and spatial head branches. The structure of both branches is fixed as follows: linear layer → activation function layer → linear layer → activation function layer. The target-guided prototype attention submodule is composed of a self-attention mechanism. Furthermore, the output feature map of the self-attention mechanism is element-wise added to the output of the deviation-aware prototype modeling submodule in a skip connection manner.

7. The cross-domain few-shot hyperspectral image classification method based on semantic consistency-guided prototyping networks according to claim 3, characterized in that: In Step 3.3, the target domain prototype calibration branch is composed of two progressive feature calibration modules connected in series, and its fixed structure is: progressive feature calibration module → progressive feature calibration module. The progressive feature calibration module consists of a Softmax layer and a bidirectional soft matching calibration layer, and its structure is fixed as follows: Softmax layer → bidirectional soft matching calibration layer.

8. The cross-domain few-shot hyperspectral image classification method based on semantic consistency-guided prototyping networks according to claim 1, characterized in that: In Step 4: The network parameters are updated using stochastic gradient descent. The total loss is calculated through weighted joint optimization with multiple losses. The total loss and each sub-loss are expressed by the following formula: ; ; ; ; ; ; in, Indicates the total loss. This represents the balancing weight coefficient for each loss term; This represents the few-shot classification loss based on prototype metric learning. Indicates the query sample. Indicates querying sample tags, Represents sample embedding features. Indicates the first Class prototype, Indicates Euclidean distance; This represents the source domain prototype consistency regularization loss. This represents the optimized source domain prototype. Represents the static prototype of the source domain; This represents the target domain feature alignment loss. , They represent the first After the next iteration, the target domain supports set and query set features. Indicates the total number of iterations; This indicates the learning loss in supervised comparison. Indicates sample features, Indicates similar enhanced viewpoint features, Indicates the temperature coefficient; This represents the cross-domain prototype contrast loss. This represents the prototype of the target domain.