Cross-scene hyperspectral image classification method based on global-local clustering and comparative learning

Through the global-local clustering and contrast learning method, combined with morphological Transformer and adversarial domain adaptive training, the problems of low accuracy and poor adaptability in cross-scene hyperspectral image classification are solved, and more efficient feature extraction and classification performance improvement are achieved.

CN120339837APending Publication Date: 2025-07-18HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510420808.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-06
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing hyperspectral image classification methods have problems with low classification accuracy and poor model adaptability in cross-scene applications, especially when the categories of the source and target domains are inconsistent. The existing methods have failed to make full use of the dynamic correlation between spectral and spatial information and have failed to effectively deal with unknown categories.

Method used

Using a method based on global-local clustering and contrast learning, a morphological Transformer network is constructed, combining spectral and spatial morphological convolutions, adversarial domain adaptive training is performed, and a positive and negative sample pair is constructed to optimize the discriminant and robustness of features using global-local clustering algorithm and contrast learning strategy.

Benefits of technology

It significantly improves the accuracy and robustness of cross-scene hyperspectral image classification, can effectively handle landform classification in different scenarios, and improves the model's adaptability in the target domain and the ability to identify unknown categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339837A_ABST
    Figure CN120339837A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-scene hyperspectral image classification method based on global-local clustering and comparative learning, and belongs to the technical field of remote sensing image classification. The problem that an existing cross-scene classification method is poor in classification performance when the categories of the source domain and the target domain are inconsistent is solved. The method comprises the following steps: extracting local morphological features and global dependency features through a morphological Transform network in combination with spectrum-spatial morphological convolution and a self-attention mechanism; a global clustering strategy and a local KNN clustering strategy are adopted, and cross-domain negative migration is reduced; contrast learning is introduced to construct cross-domain positive and negative sample pairs, the unknown category separation capability is enhanced, and the cross-scene classification robustness is remarkably improved. The method can be applied to cross-domain hyperspectral image classification scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image classification, and particularly relates to a cross-scene hyperspectral image classification method based on global-local clustering and contrast learning. Background Art

[0002] With the continuous development and optimization of remote sensing technology, hyperspectral image (HSI), as an important milestone in the optical remote sensing information chain, has shown great application potential in the fields of precision agriculture, land cover analysis, ocean hydrology detection, geological exploration, etc. by virtue of its "spectrum-map integration" feature. HSI forms the triple information fusion advantage of spatial distribution characteristics, radiation intensity characteristics, and material fingerprint spectral characteristics by integrating the reflections of visible light, near-infrared, infrared, and other bands, and can provide fine reflection information of ground objects in multiple spectral bands, which is of great significance for distinguishing ground object types with subtle spectral differences.

[0003] However, the existing hyperspectral image classification (HSIC) methods still face many challenges and limitations in cross-scene applications. On the one hand, traditional spectral feature-based classification methods, such as spectral angle mapping and parallelepiped classification, mainly rely on spectral features and ignore the correlation of ground objects at the spatial structure and semantic levels. This leads to a significant reduction in classification accuracy when dealing with ground objects with similar spectral features but different spatial distributions. For example, in urban areas, the rooftops of buildings and road surfaces may have similar spectral reflection characteristics, but their spatial layouts and functional uses are very different, and it is difficult to effectively distinguish them only relying on spectral features.

[0004] On the other hand, in recent years, hyperspectral image classification methods based on deep learning have been widely used. This method can automatically learn and extract complex features in hyperspectral data, thereby improving classification performance. However, in cross-scene classification tasks, existing methods generally face the domain adaptation problem, resulting in limited model generalization ability. Specifically, existing methods often ignore the dynamic correlation between spectral and spatial information during feature extraction and fail to fully utilize the interaction between global and local features. Existing methods also have the following limitations in cross-scene classification: the ground object categories in the source domain (SD) and the target domain (TD) may partially overlap or be completely different. Existing methods usually assume that the source domain and the target domain share exactly the same category space, which is difficult to meet in practical applications. In the target domain, due to the difference in category distribution, pseudo-label assignment may introduce incorrect information, leading to a decline in model performance. Existing adversarial training methods do not fully consider the attributes and distribution characteristics of "unknown classes", resulting in difficult effective classification of private classes in the target domain. Existing methods fail to fully combine morphological features and global dependence features during feature extraction, resulting in insufficient mining of complex spectral-spatial information.

[0005] In summary, existing hyperspectral image ground object classification methods still have deficiencies in cross-scene classification tasks. There is an urgent need for a new classification method to overcome the above problems, improve the accuracy and efficiency of hyperspectral image ground object classification, and better meet the actual application requirements. Summary of the Invention

[0006] The purpose of the present invention is to solve the problems of low classification accuracy and poor model adaptability of existing cross-scene hyperspectral image classification methods in scenarios with inconsistent categories, and to propose a cross-scene hyperspectral image classification method based on global-local clustering and contrast learning.

[0007] The technical solutions adopted by the present invention to solve the above technical problems are as follows:

[0008] A cross-scene hyperspectral image classification method based on global-local clustering and contrast learning, the method specifically includes the following steps:

[0009] Step 1, obtain hyperspectral image data of the source domain and the target domain, and preprocess the hyperspectral images of the source domain and the target domain;

[0010] Step 2, construct a morphological Transformer (morphFormer) feature extraction network, and use spectral and spatial morphological convolution operations to improve the interaction between the structure and shape information of the HSI;

[0011] Step 3, perform adversarial domain adaptation training to align the feature distributions of the source domain and the target domain;

[0012] Step 4: Use the global-local clustering algorithm to assign pseudo-labels to the target domain data;

[0013] Step 5: Construct positive and negative sample pairs through a contrastive learning strategy to optimize the discriminability and robustness of features;

[0014] Step 6: Input the target domain data into the trained model and output the classification result.

[0015] The working process of the morphological Transformer (morphFormer) is as follows:

[0016] Step 1: Use sequential layers of Conv3D and Conv2D to extract robust and discriminative features from the HSI. The original data is arranged in sub-cubes, and the sub-cubes are reshaped using padding to ensure that the spatial dimensions of the output image are the same as those of the input image;

[0017] Step 2: A Conv2D block composed of two parallel working Conv2D layers follows the Conv3D layer. One layer performs group convolution operations, and the other layer performs point convolution operations to extract multi-scale information using two different-sized kernels;

[0018] Step 3: Batch normalization (BN) and ReLU activation layers are used after the Conv3D layer and the Conv2D block. By introducing non-linearity, ReLU helps to smooth the backpropagation of the loss;

[0019] Step 4: Flatten the HSI sub-cubes to obtain patches;

[0020] MorphFormer can effectively represent the inherent shape, structure, and size of objects in the image, and includes two operations: dilation and erosion.

[0021] Furthermore, the dilation operation can expand the range of bright regions in the image and enhance the features of the target ground objects by comparing and operating the structural element with the pixels in the image.

[0022] Furthermore, the erosion operation shrinks the bright regions and suppresses interference information such as noise. Through these morphological convolution operations, morphFormer can effectively extract the local morphological features of the data and deeply mine the information in the structure and shape of the image, which is crucial for accurately identifying the ground object categories in hyperspectral images.

[0023] The working process of the adversarial domain adaptation training is as follows:

[0024] Step 1): Train a domain classifier using the sample features of the source domain and the target domain, and generate embedded features by the feature extraction network;

[0025] Step 2): Output the probability that the sample belongs to the source domain through the domain classifier, and calculate the domain classification loss;

[0026] Step 3): Train the feature extraction network, and at the same time use the labeled samples of the source domain and the target domain, as well as the unlabeled samples of the target domain to output and calculate the cross-entropy loss;

[0027] Step 4): Repeat Step 1), Step 2) and Step 3) until the model converges;

[0028] The working process of the global-local clustering algorithm is as follows:

[0029] Step a: Initialize the target domain clustering, and map the target domain data to the embedding space through the feature extraction network.

[0030] Step b: Construct positive and negative subsets. For each category, select the sample with the highest confidence from the source domain to construct the positive subset, and classify the remaining samples as the negative subset, indicating that they do not belong to the category;

[0031] Step c: Calculate the cosine similarity between the target domain samples and the positive and negative prototypes for pseudo-label assignment. If the similarity of the sample to the positive prototype is significantly higher than that of the negative prototype, assign a pseudo-label;

[0032] Step d: Unsupervised adaptive training, calculate the cross-entropy loss using the pseudo-labels to optimize the model's classification ability, and combine the kernel triplet module loss and the local KNN clustering strategy to force similar samples to be close and dissimilar samples to be far away, enhancing the feature separability;

[0033] Step e: Repeat Step b to Step d, gradually correct the pseudo-labels and optimize the model until the shared classes are aligned and the unknown classes are separated in the feature space.

[0034] The working process of the contrastive learning is as follows:

[0035] Step (1): At the dataset level, generate positive sample pairs through clustering, estimate the expected number of samples for each category, and select the non-expected samples most similar to the target sample as negative samples;

[0036] Step (2): Calculate the similarity of the positive and negative pairs using the cosine similarity function;

[0037] Step (3): Optimize the model through backpropagation to enhance the ability to capture the internal structure of the data, and loop through Step (1) to Step (2) to gradually improve the model's performance in distinguishing "known" and "unknown" classes. Description of the Drawings

[0038] Figure 1 is the flowchart of the cross-scene hyperspectral image classification method based on global-local clustering and contrastive learning;

[0039] Figure 2 It is a framework diagram of a hyperspectral image classification model based on global-local clustering and contrast learning;

[0040] Figure 3 It is a diagram of the morphFormer model;

[0041] Figure 4 It is a schematic diagram of dilation and erosion operations;

[0042] Figure 5 It is a schematic diagram of adversarial domain adaptation;

[0043] Figure 6 It is a schematic diagram of the kernel triplet hard loss;

[0044] Figure 7 It is a ground truth map of the PaviaU dataset;

[0045] Figure 8 It is a ground truth map of the PaviaC dataset;

[0046] Figure 9 It is a classification result map of the PaviaU→PaviaC dataset;

[0047] Figure 10 It is a classification result map of the PaviaC→PaviaU dataset. Detailed implementation manners

[0048] The present application will be further described in detail below in conjunction with the accompanying drawings through specific implementation manners. Obviously, the described implementation manners are only a part of the implementation manners of the present invention, rather than all the implementation manners. All other implementation manners obtained by those of ordinary skill in the art based on the implementation manners in the present invention without creative efforts fall within the scope of protection of the present invention.

[0049] Detailed implementation manner 1. In combination with Figure 1 and Figure 2 This implementation manner will be described. A cross-scene hyperspectral image classification method based on global-local clustering and contrast learning described in this implementation manner specifically includes the following steps:

[0050] Detailed implementation manner 1: A cross-scene hyperspectral image classification method based on global-local clustering and contrast learning described in this implementation manner specifically includes the following steps:

[0051] Step 1, obtain hyperspectral image data of the source domain and the target domain, and preprocess the hyperspectral images of the source domain and the target domain;

[0052] Step 2: Construct a morphological Transformer (morphFormer) feature extraction network and use spectral and spatial morphological convolution operations to improve the interaction between the structure and shape information of HSI;

[0053] Step 3: Perform adversarial domain adaptation training to align the feature distributions of the source domain and the target domain;

[0054] Step 4: Use the global-local clustering algorithm to assign pseudo-labels to the target domain data;

[0055] Step 5: Construct positive and negative sample pairs through a contrastive learning strategy to optimize the discriminability and robustness of the features;

[0056] Step 6: Input the target domain data into the trained model and output the classification results.

[0057] The present invention designs a hyperspectral general set image classification model combining morphFormer, adversarial training, and the global-local clustering algorithm. The convolutional neural network is used for local feature extraction, and the global feature representation ability of morphFormer is combined for feature extraction. At the same time, the feature extractor and the classifier are trained separately by means of adversarial training to train the data feature extraction ability of the model. The global-local clustering algorithm improves the model domain adaptation ability. It has a high recognition ability for "unknown classes" on the general set, and the model performance is very superior in the cross-scene classification scenario, improving the classification performance.

[0058] Specific Embodiment 2: Combination Figure 3 Describe this embodiment. The difference between this embodiment and Specific Embodiment 1 is that the morphological Transformer (morphFormer), wherein:

[0059] HSI has many spectral bands, usually represented by the number of channels. Therefore, CNN can be used to control the depth of the output feature map. As Figure 3 shown, the model proposed in this chapter uses CNN to extract high-level abstract features for morphFormer. This model uses sequential layers of Conv3D and Conv2D to extract robust and discriminative features from HSI. The original data is arranged in sub-cubes, and the sub-cubes are reshaped using padding to ensure that the spatial dimensions of the output image are the same as those of the input image. The subsequent Conv2D blocks follow the Conv3D layer and consist of two Conv2D layers working in parallel. One layer is used for group convolution, and the other layer is used for point convolution. One layer performs group convolution operations, and the other layer performs point convolution operations, using two kernels of different sizes to extract multi-scale information. The outputs obtained from these two convolutions are combined element-wise and returned as the output:

[0060]

[0061] Where X in represents the input HSI data, K1 = 3, and K2 = 1 represent different convolutional kernels. Batch Normalization (BN) and ReLU activation layers are used after the Conv3D layer and the Conv2D block. By introducing non-linearity, ReLU helps smooth the backpropagation of the loss. After that, the HSI sub-cube is flattened to obtain a patch:

[0062] X flat = T(Flatten(X out ))

[0063] Where T(·) is the transpose function. Considering the great correlation between patches at adjacent positions, the sum of the elements of the patch at this position and the patch at the previous position is used as the patch at this position, and position embedding is added, which can be expressed as:

[0064]

[0065] For the dilation operation, by combining the patch tokens of the HSI input with the SE, the pixels with the maximum value in the local neighborhood are selected to generate an extended image, that is, the boundary is broadened. For the erosion operation, the output of the convolution with the SE selects the pixels with the minimum value in the local neighborhood. This operation reduces the shape of the HSI background objects, can eliminate tiny details, expand holes, and make them distinguishable from each other in different texture regions, as Figure 4 shown.

[0066] Other steps and parameters are the same as those in the first specific implementation manner.

[0067] In cross-scene hyperspectral images, the distribution of ground objects and background characteristics under different scenes vary greatly. It is difficult to comprehensively understand the overall structure of the image and the relationship between different regions only by convolution operations. morphFormer can not only finely extract local morphological features through morphological operations, but also use the self-attention mechanism of the transformer to obtain global dependence features, so as to extract data features more comprehensively and deeply.

[0068] Specific implementation manner three: Combine Figure 5 to illustrate this implementation manner. The difference between this implementation manner and one of the first and second specific implementation manners is the adversarial domain adaptation training, where:

[0069] During the adversarial training process, the feature extractor and the discriminator engage in an adversarial game, making it difficult for the domain discriminator to accurately determine whether a sample comes from the source domain or the target domain, thus achieving a "confusion" effect; while the discriminator is continuously optimized to improve its ability to distinguish between SD and TD data. Through this adversarial training, the feature extractor gradually learns a feature representation that can eliminate domain differences and realizes the effective transfer of source domain knowledge to the target domain. The adversarial training has two processes, namely, alternately training the domain classifier and the feature extraction network, which corrects the problem that the feature extraction network does not update parameters, and at the same time the domain classifier is learnable.

[0070] The input of the domain classifier is D S and D T the embedded features output by the feature extraction network of the sample. The domain classifier outputs a value between 0 and 1, indicating the probability that the sample comes from D S or D T . Samples with an output value close to 1 are more likely to come from D S . The loss L D is defined as follows:

[0071]

[0072] L D makes the output of the domain classifier for the sample as close to 1 as possible and the output for the D T sample as close to 0 as possible.

[0073] Other steps and parameters are the same as those in the first or second specific implementation manner.

[0074] This strategy realizes the feature aggregation of similar samples in TD and SD by optimizing the feature space distribution, and at the same time expands the feature differences between different category samples. This dual effect enhances the intra-class compactness and inter-class separability of the features, thereby improving the classification performance of the model.

[0075] Specific implementation manner four: Combine Figure 6 to illustrate this implementation manner. The difference between this implementation manner and one of the first to third specific implementation manners is that the global-local clustering algorithm, where:

[0076] First, identify the category with the highest K of δ(f T (x t (x t )) scores from D c to construct the positive subset P c , and at the same time designate the rest as the negative subset N t (x t )) represents the maximum probability that the target instance x t belongs to the c-th category. Assume K = N c / Ct 。

[0077] Then, the positive prototype P is obtained through K-means c (which is the c-th class) and the negative prototype (which is not the c-th class).

[0078]

[0079] After that, the following formula can be used to determine whether the data sample x t belongs to the c-th class. Among them, represents that the data sample x t is pseudo-labeled to the c-th class. Finally, the above process is iterated to obtain the pseudo-labels of all target domain classes c ∈ C. Using the obtained pseudo-labels, unsupervised model adaptation is performed in combination with the cross-entropy loss. Among them, represents the pseudo-label of the data sample .

[0080]

[0081] After that, the source domain and target domain features are mapped to a common feature space. First, the source domain features are fixed, and the pseudo-label cluster centers of the target domain are approximated to the sample centers of each class in the source domain, and only the class center with the largest similarity is used as the target. For the shared classes of the two, that is, the "known" classes in the target domain are close to each other; for the "unknown" classes, they are far from each other. Considering all possible classes in the training batch, and mapping the samples to a higher-dimensional Gaussian kernel space to obtain better classification, combined with the kernel triplet hard loss, as Figure 6 shown. In the original feature space, the distance between the most difficult positive pairs (connected by the black dashed line) is much greater than the distance between the most difficult negative pairs (connected by the red solid line). The proposed kernel triplet hard loss can map the samples to a high-dimensional feature space, making it easier to classify the samples.

[0082] To alleviate the situation of incorrect pseudo-label assignment, the present invention further introduces a local KNN clustering strategy. Specifically, during the model adaptation process, a memory bank G t ={g t (x t ), σ(f t (x t ))} is maintained, which contains the target features and the corresponding prediction scores. Then, local KNN consensus clustering is achieved through the formula and. Thereby improving the accuracy and robustness of pseudo-label assignment.

[0083] Other steps and parameters are the same as those in any one of the specific embodiments 1 to 3.

[0084] Specific Embodiment 5: The difference between this embodiment and any one of the first to fourth specific embodiments is that in the contrast learning, where:

[0085] At the dataset level, a local consensus clustering method is adopted to construct positive sample pairs and build a memory bank based on this, rather than relying on traditional instance-level data augmentation strategies. This method effectively avoids the defects of class-level data expansion and at the same time completely preserves the intrinsic semantic structure of the dataset, showing significant advantages when dealing with "unknown" data. For the construction of negative sample pairs, a simple and efficient hard negative sample mining strategy is designed, abandoning the practice of simply treating all other samples in a mini-batch as negative samples. Specifically, first estimate the expected number of samples for each category in the mini-batch, and then select the non-expected sample that is most similar to the target sample as the most representative negative sample.

[0086]

[0087] and are the positive and negative pairs of data respectively. S(·) represents the cosine similarity function. For simplicity, the sizes of the positive pairs, negative pairs, and local clustering neighbors are balanced, that is

[0088] Other steps and parameters are the same as any one of the first to fourth specific embodiments.

[0089] Experimental Section

[0090] All experiments of the present invention were carried out in a unified computing environment. The specific hardware configuration includes: Intel Core i7-6850K central processing unit, Nvidia GeForce GTX 1080Ti, and 64GB of memory; the software configuration uses Python 3.7 and PyTorch 1.5.0 to implement the encoding of the proposed morphFormer. In order to make full use of the spatial information of hyperspectral images, patch-level input is adopted in this paper, and the number of obtained HSI patch labels is taken as four. During training and testing, batch sizes of 64 and 500 are used respectively. Patches of size 11×11×B are extracted from the HSI and used as the input of the model. SGD and Adam optimizers are also used to train the model, with a weight decay of 5e-3 and a learning rate of 5e-4. In addition, these methods use a step scheduler with a gamma of 0.9 and a step size of 50, and are trained within 500 epochs. The average value of each experiment is calculated based on three repetitions.

[0091] Dataset Introduction: The Pavia dataset was collected by the University of Pavia, Italy in 2003. This dataset contains two subsets: Pavia University (PaviaU) and Pavia Center (PaviaC). The PaviaU dataset was collected by the ROSIS-03 sensor, contains 115 spectral bands, with a wavelength range of 430–860 nm, a spatial resolution of 1.3 m, and an image size of 610×340 pixels, as Figure 7 shown. The PaviaC dataset was also collected by the ROSIS-03 sensor, covers the central area of Pavia, Italy, contains 102 spectral bands, with a wavelength range of 430–860 nm, a spatial resolution of 1.3 m, and an image size of 1096×715 pixels, as Figure 8 shown.

[0092] The subjective classification results of the cross-scene hyperspectral image classification model based on global-local clustering and contrast learning used in this invention on the PaviaU→PaviaC dataset and PaviaC→PaviaU are as Figure 9 and Figure 10 shown. The objective classification results are shown in Table 1 and Table 2 respectively:

[0093] Table 1 Classification Results of PaviaU→PaviaC under the Method of this Invention

[0094]

[0095] Table 2 Classification Results of PaviaC→PaviaC under the Method of this Invention

[0096]

[0097] The above examples of this invention are only to illustrate in detail the calculation model and calculation process of this invention, rather than to limit the implementation manner of this invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or variations derived from the technical solutions of this invention still fall within the protection scope of this invention.

Claims

1. A cross-scene hyperspectral image classification method based on global-local clustering and contrast learning, characterized in that It includes the following steps: Step 1: Obtain the hyperspectral image data of the source domain and the target domain, and preprocess the hyperspectral images of the source domain and the target domain; Step 2: Construct a morphological Transformer (morphFormer) feature extraction network, and use spectral and spatial morphological convolution operations to improve the interaction between the structural and shape information of hyperspectral images; Step 3: Perform adversarial domain adaptation training to align the feature distributions of the source domain and the target domain; Step 4: Use the global-local clustering algorithm to assign pseudo-labels to the target domain data; Step 5: Construct positive and negative sample pairs through a contrastive learning strategy to optimize the discriminability and robustness of features; Step 6: Input the target domain data into the trained model and output the classification result.

2. The cross-scene hyperspectral image classification method based on global-local clustering and contrast learning according to claim 1, wherein The construction of the morphological Transformer (morphFormer) feature extraction network includes the following steps: Step 1: Use sequential layers of Conv3D and Conv2D to extract robust and discriminative features from the HSI. The original data is arranged in sub-cubes, and the sub-cubes are reshaped using padding to ensure that the spatial dimensions of the output image are the same as those of the input image; Step 2: A Conv2D block composed of two parallel working Conv2D layers follows the Conv3D layer. One layer performs group convolution operations, and the other layer performs point convolution operations to extract multi-scale information using two different-sized kernels; Step 3: Batch normalization and ReLU activation layers are used after the Conv3D layer and the Conv2D block. By introducing non-linearity, ReLU helps to smooth the backpropagation of the loss; Step 4: Flatten the hyperspectral image sub-cubes to obtain patches.

3. A cross-scene hyperspectral image classification method based on global-local clustering and contrast learning according to claim 1, characterized in that, The adversarial domain adaptation training includes the following steps: Step 1): Train a domain classifier using the sample features of the source domain and the target domain, and generate embedded features by the feature extraction network; Step 2): Calculate the domain classification loss by outputting the probability that the sample belongs to the source domain by the domain classifier; Step 3): Train the feature extraction network, and at the same time use the labeled samples of the source domain and the target domain, and the unlabeled samples of the target domain to output and calculate the cross-entropy loss; Step 4): Repeat steps 1), 2) and 3) until the model converges; 4. A cross-scene hyperspectral image classification method based on global-local clustering and contrastive learning according to claim 1, characterized in that, The global-local clustering algorithm includes the following steps: Step a: Initialize the target domain clustering, and map the target domain data to the embedding space through the feature extraction network. Step b: Construct positive and negative subsets. For each category, select the sample with the highest confidence from the source domain to construct the positive subset, and classify the remaining samples as the negative subset, indicating that they do not belong to the category; Step c: Calculate the cosine similarity between the target domain samples and the positive and negative prototypes for pseudo-label assignment. If the similarity of the sample to the positive prototype is significantly higher than that of the negative prototype, assign a pseudo-label; Step d: Unsupervised adaptive training, calculate the cross-entropy loss using the pseudo-labels to optimize the classification ability of the model, and combine the kernel triplet module loss and the local KNN clustering strategy to force similar samples to be close and dissimilar samples to be far away, enhancing the feature separability; 5. A cross-scene hyperspectral image classification method based on global-local clustering and contrast learning according to claim 1, characterized in that, The contrastive learning algorithm includes the following steps: Step (1): At the dataset level, generate positive sample pairs through clustering, estimate the expected number of samples for each category, and screen the non-expected samples that are most similar to the target samples as negative samples; Step (2): Calculate the similarity of positive and negative pairs using the cosine similarity function; Step (3): Optimize the model through backpropagation to enhance the ability to capture the internal structure of the data. Loop through steps (1) to (2) to gradually improve the model's performance in distinguishing between "known" and "unknown" categories.

Citation Information

Cited By

  • Cross-domain image classification method and device

    CN122067030A

  • A cross-domain image classification method and device

    CN122067030B