Training sample selection method for hyperspectral remote sensing image classification
Feature extraction through GSCVIT model and combined with dynamic distribution balance loss optimization training, the high annotation cost and category imbalance of hyperspectral remote sensing images are solved, and the classification accuracy and stability of the model are improved, especially the recognition ability of a few categories.
Patent Information
- Application Number
- CN202510754096.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The data labeling of hyperspectral remote sensing images is high and the categories are unbalanced. The existing sample selection strategies are difficult to balance representativeness, discrimination and category balance, resulting in low recognition accuracy of the model in a few categories, affecting the decision accuracy of practical applications.
The pre-trained GSCVIT model is used to extract features, combine multi-headed attention weights to calculate spatial attention entropy, filter core set samples through K-Center greedy algorithm, and optimize training with dynamic distribution balance loss DDB Loss, dynamically adjust the category weights to improve classification performance.
It significantly improves the overall accuracy and stability of hyperspectral remote sensing image classification, especially the ability to identify a few and weak targets, effectively alleviating the problem of category imbalance.
Smart Images

Figure CN120388253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification, and in particular to a method for selecting training samples for hyperspectral remote sensing image classification. Background Art
[0002] Hyperspectral remote sensing images (HRSIs) have hundreds of continuous narrow bands, with high spectral resolution, which can finely depict the spectral characteristics of ground objects. They are widely used in fields such as land classification, vegetation monitoring, and environmental assessment, and have important practical value and economic potential. However, the data annotation of HRSIs still highly depends on manual work, resulting in a series of challenges such as low sample utilization efficiency and weak model generalization ability. In practical applications, the sample annotation cost of hyperspectral images is extremely high, relying on professional personnel for ground field measurements and manual classification, making the number of available labeled samples extremely limited. In this case, how to select "which samples" for annotation and training becomes crucial - different selection strategies will directly affect the learning effect and final performance of the model.
[0003] In addition, there are often serious class imbalance problems in hyperspectral remote sensing scenes, that is, the number of samples in the majority class is much larger than that in the minority class. This not only causes the model to be biased during training, performing well on the dominant class but poorly on rare classes, but may also cause serious consequences in practical applications. For example, in ecological monitoring and disaster identification, the minority class often represents key targets (such as wetlands, fire areas, etc.), and misjudgment of them will directly affect the decision-making accuracy and response efficiency.
[0004] In current sampling techniques, although a variety of sample selection strategies have been proposed, including random sampling, stratified sampling, clustering sampling, active learning, and core set sampling, these methods still have many limitations when facing hyperspectral images with high annotation costs, uneven class distributions, and complex feature expressions. For example, although random sampling is simple, it cannot identify heterogeneous regions with discriminative value; clustering or core set methods focus on representativeness but ignore the model's perception of the discriminative difficulty of samples; and existing imbalance learning methods such as Focal Loss or class balance loss only perform static weighting at the loss level and are difficult to dynamically adjust the training strategy according to the actual performance of the model.
[0005] Therefore, there is an urgent need for a new sample selection and training optimization mechanism that takes into account sample representativeness, discriminability, and class balance to improve the recognition accuracy and generalization ability of the model for various ground objects in hyperspectral images under the condition of limited labeled samples. This need widely exists in many practical engineering fields such as forestry resource surveys, farmland management, environmental monitoring, geological exploration, and urban security, and has extremely high application and promotion value. Summary of the Invention
[0006] The object of the present invention is to provide a method for selecting training samples for hyperspectral remote sensing image classification, aiming to effectively select representative and discriminative training samples in the actual application scenario with limited annotation cost, not only improving the overall classification accuracy, but also significantly enhancing the recognition ability of minority and weak class targets, and effectively alleviating the performance bottleneck caused by class imbalance.
[0007] To achieve the above object, the present invention provides a method for selecting training samples for hyperspectral remote sensing image classification, including the following steps:
[0008] S1. Given an unlabeled data set where represents a hyperspectral image sample with a spatial dimension of H×W and a spectral band number of B;
[0009] S2. Use the pre-trained GSCVIT model to extract features from the data set to obtain classification features F and multi-head attention weights;
[0010] S3. Calculate the spatial attention entropy SSAE based on the attention weights in step S2, and splice the classification features and the spatial attention entropy to generate enhanced features;
[0011] S4. Adopt the K-Center greedy algorithm to screen out the core set sample subset D from the original data set based on the enhanced features c ;
[0012] S5. Use a group sampling loader to load training data, and adopt the dynamic distribution balance loss DDB Loss to train the model to optimize the classification performance in the case of class imbalance.
[0013] Preferably, in step S2, the input feature map size of the pre-trained GSCVIT model is C×8×8, where C is the spectral band number, the input image patch is normalized to the [0,1] interval, and data augmentation is performed by horizontal flipping, vertical flipping and random rotation.
[0014] Preferably, for each batch of data x b ∈D, the classification features and multi-head attention weights are obtained through backpropagation of the GSCVIT model, where L is the number of layers, H is the number of heads, and P is the number of positions;
[0015] The attention weights of the l-th layer and the h-th head are normalized by softmax:
[0016]
[0017] Calculate the attention entropy:
[0018]
[0019] Preferably, in step S4, the K-Center greedy sampling step includes:
[0020] Initialize the core set S and select initial samples according to the sample distribution stratification;
[0021] Iteratively select samples: For the unselected sample x j ∈D / S, calculate the minimum Euclidean distance from its enhanced features to the core set
[0022] Combined with the attention entropy SSAE j , calculate the score: S j =λ·d j +(1 - λ)·SSAE j , where λ∈[0,1] is the balance parameter;
[0023] Select the sample with the highest score and add it to the core set until the size of the core set reaches the preset value N.
[0024] Preferably, in step S4, use the group sampling loader to load the training data, which specifically includes the following steps:
[0025] Randomly shuffle the order of samples within each category;
[0026] Group the shuffled samples by the group size n, and the remaining samples form a separate group;
[0027] Randomly shuffle the grouped samples of all categories again to generate training batches.
[0028] Preferably, in step S4, the construction steps of the dynamic distribution balance loss DDB Loss include:
[0029] Define the recall rate and the precision where, TP i , FN i and FP i are the number of true positive, false negative, and false positive samples of the i-th category respectively;
[0030] Initialize the category weight matrix C is the number of categories;
[0031] Dynamically update the weights based on the recall rate and precision of the validation set after each round of training:
[0032] Recall rate weight:
[0033] Precision weight:
[0034] Recall i is the recall rate for the i-th class, SF is the scaling factor, Precision i is the precision for the i-th class;
[0035] Normalize the dynamic weights and introduce the negative class sensitivity to construct the loss function, specifically:
[0036] Normalize after fusing the recall rate and precision weights:
[0037] The dynamic distribution balance loss is calculated as follows:
[0038]
[0039] Among them, is the predicted score of the i-th class of the k-th sample, is the true label of the i-th class of the k-th sample, and the true label adopts one-hot encoding;
[0040] C is the total number of classes, w i is the dynamic weight matrix, and λ is the hyperparameter used to adjust the negative class loss.
[0041] Therefore, the present invention adopts a training sample selection method for hyperspectral remote sensing image classification with the above structure, having the following beneficial effects:
[0042] The dynamic collaborative balance sampling method proposed by the present invention can effectively identify and preferentially select key samples in hyperspectral remote sensing data in practical application scenarios with limited annotation costs, while dynamically optimizing the class distribution, significantly improving the classification accuracy and stability of the model under unbalanced data.
[0043] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Brief Description of the Drawings
[0044] Figure 1 is a schematic flow chart of a training sample selection method for hyperspectral remote sensing image classification according to the present invention. Detailed Embodiments
[0045] The technical solutions of the present invention will be further described below through the drawings and embodiments.
[0046] Unless otherwise defined, the technical terms or scientific terms used in this invention shall have the ordinary meanings as understood by those of ordinary skill in the field to which this invention pertains. The terms "first", "second" and similar terms used in this invention do not denote any order, quantity or importance, but are only used to distinguish different components. Terms such as "comprising" or "including" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. Terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Terms such as "upper", "lower", "left", "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0047] Embodiment
[0048] As Figure 1 shown, the present invention provides a method for selecting training samples for hyperspectral remote sensing image classification, including the following steps:
[0049] S1. Given an unlabeled data set where represents a hyperspectral image sample with a spatial dimension of H×W and a spectral band number of B;
[0050] S2. Use a pre-trained GSCVIT model to extract features from the data set to obtain classification features F and multi-head attention weights;
[0051] S3. Calculate the spatial attention entropy SSAE based on the attention weights in step S2, and splice the classification features with the spatial attention entropy to generate enhanced features;
[0052] S4. Adopt the K-Center greedy algorithm to screen out the core set sample subset D from the original data set based on the enhanced features c ;
[0053] S5. Use a group sampler loader to load training data, and adopt the dynamic distribution balance loss DDB Loss to train the model to optimize the classification performance in the case of class imbalance.
[0054] Among them, the dynamic distribution balance loss (DDB Loss) adopts a dual-index dynamic weight adjustment strategy based on recall and precision to optimize the classification performance of the model under a fixed training sample size. The specific process is as follows:
[0055] Index definition:
[0056] Recall measures the ability of the model to identify positive class samples, representing the proportion of correctly identified positive class samples among all positive class samples. The calculation formula is as follows:
[0057]
[0058] Among them, TP (True Positives) is the number of samples that the model correctly predicts as positive class, and FN (False Negatives) is the number of samples that the model incorrectly predicts as negative class but are actually positive class.
[0059] Precision measures the proportion of actual positive class samples among all samples predicted as positive class by the model. Its mathematical expression is:
[0060]
[0061] Among them, FP is the number of samples that the model incorrectly predicts as positive class but are actually negative class.
[0062] Dynamic class weight calculation:
[0063] Initialize the weight matrix: Construct a weight array according to the class distribution of the initial training data (C is the number of classes);
[0064] Dynamic update rule: After each round of training, update the weights based on the performance of the current model on the validation set:
[0065] Recall weight:
[0066]
[0067] Precision weight:
[0068]
[0069] Ri is the recall rate of the i-th class, SF (Scaling Factor) is the scaling factor, and Pi is the precision of the i-th class.
[0070] Weight normalization: Normalize after fusing the recall and precision weights:
[0071]
[0072] Loss function construction:
[0073] Dynamic Distribution Balanced Loss (DDB Loss) enables the model to better handle the balance problem between classes by introducing dynamic class weights and negative class sensitivity in loss calculation.
[0074] Dynamic class weights: By loading the dynamic class weight array w i , the model can adjust the learning weights during training and increase the learning proportion of weak classes.
[0075] Negative class sensitivity adjustment: The negative class loss is adjusted by hyperparameters, which are used to control the impact of negative samples on the overall loss and prevent over-punishment, thereby reducing the bias of the model towards negative classes.
[0076] The calculation of the dynamic distribution balance loss is as follows:
[0077]
[0078] where is the predicted score of the i-th class of the k-th sample, is the true label of the i-th class of the k-th sample, and the true label uses one-hot encoding;
[0079] C is the total number of classes, w i is the dynamic weight matrix, and λ is the hyperparameter used to adjust the negative class loss.
[0080] Core set sampling with attention entropy enhancement
[0081] (1) Feature and attention entropy extraction:
[0082] For each batch of data x b ∈ D:
[0083] Backpropagation to obtain classification features:
[0084]
[0085] Extract multi-head attention weights where where H is the number of heads, P is the number of positions, and L is the number of layers. The normalized weight of the h-th head at the p-th position in the l-th layer:
[0086]
[0087] Calculate the attention entropy:
[0088]
[0089] (2) Enhanced feature construction:
[0090] Concatenate the classification features and the attention entropy to generate enhanced features to simultaneously capture the semantic information and uncertainty of the samples:
[0091]
[0092] (3) K-Center greedy sampling
[0093] Based on enhanced features, an improved K-Center greedy algorithm is used to select the core set The specific steps are as follows:
[0094] Initialization: Samples x are selected by stratification according to the sample distribution init Add
[0095] Iterative selection (for a = 2 to A):
[0096] For each unselected sample Calculate its minimum Euclidean distance to the core set:
[0097]
[0098] Combining the distance term and the attention entropy term, adjust the weight through the balance parameter λ ∈ [0, 1]:
[0099] s j = λ·d j +(1 - λ)·SSAE j ;
[0100] iii. Select the N samples with the highest scores a and add them to the core set:
[0101]
[0102] Termination: When output the core set
[0103] Group sampling loader
[0104] Load samples through the group sampling loader, and the model can access a richer combination of samples in each iteration, thereby improving training stability and reducing the risk of overfitting. The detailed steps are as follows:
[0105] Shuffle by class: Randomly shuffle the order of samples within each class c i to reduce the model's dependence on the fixed order and improve the generalization ability.
[0106] Group processing: Group the shuffled samples according to the group size. If the number of samples is insufficient or not divisible, the remaining samples are grouped separately without resampling to ensure balanced distribution between groups.
[0107] Shuffle between groups: After aggregating the small groups of all classes, shuffle them randomly again to avoid the fixed order of mini-batches during training and enhance training diversity.
[0108] This invention was carried out on a personal computer equipped with a 12th generation Intel i9-12900H processor (2.50 GHz), 16 GB of memory, and a Windows 10 operating system. The experimental environment was built based on Python 3.8 and the PyTorch framework. The GSC-ViT model was used as the baseline model for training and testing, and the input feature map size was initialized to C×8×8, where C represents the number of spectral bands. The input image patches were normalized to the range [0, 1] and extended through data augmentation methods, including horizontal flipping, vertical flipping, and random rotations of 90°, 180°, and 270°. The model was trained using the AdamW optimizer with 200 training epochs, a learning rate of 0.001, a weight decay coefficient of 0.05, and a batch size of 128. To ensure the reliability of the experimental results, each experiment was repeated 10 times, and the average value was recorded as the final result.
[0109] Experimental verification:
[0110] To comprehensively evaluate the effectiveness of the proposed method, this invention conducted systematic experiments on four typical hyperspectral remote sensing datasets (Indian Pines, Salinas, Pavia University, and WHU-Hi-LongKou). On each dataset, this invention set unified limited labeled sample conditions and used 512 (Indian Pines), 270 (Salinas), 44 (Pavia University), and 102 (WHU-Hi-LongKou) samples for training, respectively. On this basis, this invention compared the proposed dynamic class boundary sampling method (DCBS) with common sampling strategies such as random sampling, stratified sampling, entropy sampling, MC Dropout sampling, and Coreset sampling, and systematically analyzed the classification performance differences of each method in the class imbalance scenario. In addition, to verify the independent contributions of each module in the proposed method, this invention constructed ablation experiments based on the GSCViT model, gradually introduced the grouped sampling loader, entropy-guided coreset sampling, and dynamic distribution balance loss (DDB Loss), and analyzed the performance improvement effects of different modules on the model. The experiments used the overall accuracy (OA), average accuracy (AA), Kappa coefficient, and F1 score as the main evaluation indicators, and focused on analyzing the performance of each method under weak category and extremely small sample conditions.
[0111] This invention uses the Indian Pines dataset, Salinas dataset, Pavia University dataset, and WHU-Hi-LongKou dataset to demonstrate as follows:
[0112] 1. Indian Pines dataset (512 training samples):
[0113] In this dataset, random sampling leads to significant class imbalance. In particular, the accuracies of classes 1, 7, and 9 are much lower than those of other classes, dragging down the overall average accuracy (AA). The DCBS method significantly improves the recognition effects of these disadvantaged classes by dynamically balancing the weights between classes, with improvements of 27.54%, 27.43%, and 56.1% respectively. The overall AA and OA are improved by 9.89% and 3.38% respectively, performing best among all sampling methods and fully demonstrating its adaptability to small samples and difficult-to-classify classes.
[0114] 2. Salinas dataset (270 training samples):
[0115] In the Salinas dataset with a large number of classes and complex distributions, random sampling shows instability in multiple classes, and entropy sampling even regresses due to its poor ability to guide between classes. In contrast, the DCBS method can effectively improve the classification performance of classes 3, 8, and 11, with a significant increase in average accuracy, and leads comprehensively in the four indicators of OA, AA, Kappa, and F1, showing its strong robustness to the problems of sample scarcity and inter-class confusion.
[0116] 3. PaviaUniversity dataset (44 training samples):
[0117] In the PaviaUniversity data dominated by urban features, the overall performance of each method is similar, but DCBS still demonstrates the ability to accurately model difficult-to-classify classes. In particular, for classes 6, 7, 8, etc., which perform poorly in other methods, there is a significant improvement. The AA is increased by 1.86%, and the OA reaches 98.63%, achieving better performance at both the micro and macro levels, verifying its stable generalization ability.
[0118] 4. WHU-Hi-LongKou dataset (102 training samples):
[0119] This dataset presents challenging features such as subtle differences in ground objects and blurred boundaries. Random sampling is difficult to obtain representative samples, resulting in an AA of only 89.19%. In this context, DCBS effectively guides the sample selection strategy, significantly improving the accuracies of classes 4 and 8 by 16.45% and 10.73% respectively, with the AA increased by 8.55% and the OA reaching 98.75%, demonstrating the ability to construct a balanced training set in complex semantic scenarios.
[0120] The proposed dynamic collaborative balance sampling method in the present invention shows significant advantages in all four remote sensing datasets. On the Indian Pines dataset, compared with random sampling, the classification accuracies of classes 1, 7, and 9 are significantly improved, and the overall AA and OA are increased by 9.89% and 3.38% respectively; in the Salinas dataset, the classification difficulties of classes 3, 8, and 11 are effectively alleviated, and AA is increased to a maximum of 95.93%; in the Pavia University dataset, even with extremely few samples, more stable recognition of core classes is still achieved. Compared with the Coreset method, OA and AA are increased by 1.37% and 1.43% respectively; on the WHU-Hi-LongKou dataset, the proposed method achieves the best performance in terms of OA, AA, Kappa, and F1 metrics. Especially in the case of obvious differences in the distribution among classes, excellent classification consistency is still maintained. These effects fully verify the effectiveness and robustness of the proposed method in the scenarios of limited samples and class imbalance.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for selecting training samples for hyperspectral remote sensing image classification, characterized in that: It includes the following steps: S1. Given an unlabeled dataset where represents a hyperspectral image sample with a spatial dimension of H×W and a spectral band number of B; S2. Use the pre-trained GSCVIT model to extract features from the dataset to obtain classification features F and multi-head attention weights; S3. Calculate the spatial attention entropy SSAE based on the attention weights in step S2, and splice the classification features with the spatial attention entropy to generate enhanced features; S4. Use the K-Center greedy algorithm to screen out the core set sample subset D from the original dataset based on the enhanced features c ; S5. Use a group sampling loader to load the training data, and adopt the dynamic distribution balance loss DDB Loss to train the model to optimize the classification performance in the case of class imbalance.
2. The training sample selection method for hyperspectral remote sensing image classification according to claim 1, characterized in that: In step S2, the input feature map size of the pre-trained GSCVIT model is C×8×8, where C is the number of spectral bands, the input image patches are normalized to the [0,1] interval, and data augmentation is performed by horizontal flipping, vertical flipping, and random rotation.
3. A method for selecting training samples for hyperspectral remote sensing image classification according to claim 1, characterized in that: For each batch of data x b ∈ D, the classification features are obtained through backpropagation of the GSCVIT model and the multi-head attention weights where L is the number of layers, H is the number of heads, and P is the number of positions; Perform softmax normalization on the attention weights of the h-th head in the l-th layer: Calculate the attention entropy:
4. A method for selecting training samples for hyperspectral remote sensing image classification according to claim 1, characterized in that: In step S4, the K-Center greedy sampling step includes: Initialize the core set S, and select initial samples hierarchically according to the sample distribution; Iteratively select samples: For the unselected sample x j ∈ D / S, calculate the minimum Euclidean distance from its enhanced features to the core set Combined Attention Entropy SSAE j , calculate the score: S j = λ·d j +(1 - λ)·SSAE j , where λ ∈ [0,1] is the balance parameter; Select the sample with the highest score and add it to the core set until the size of the core set reaches the preset value N.
5. A method for selecting training samples for hyperspectral remote sensing image classification according to claim 1, characterized in that: In step S4, using a group sampling loader to load the training data specifically includes the following steps: Randomly shuffle the order of samples within each category; Group the shuffled samples into groups of size n, and the remaining samples form a separate group; Randomly shuffle the grouped samples of all categories a second time to generate training batches.
6. A method for selecting training samples for hyperspectral remote sensing image classification according to claim 1, characterized in that: In step S4, the construction steps of the dynamic distribution balance loss DDB Loss include: Define recall and precision where TP i 、FN i and FP i are the numbers of true positives, false negatives, and false positives of the i-th class, respectively; Initialize the class weight matrix C is the number of classes; Dynamically update the weights based on the recall rate and precision of the validation set after each round of training: Recall rate weight: Precision weight: Recall i is the recall rate for the i-th class, SF is the scaling factor, Precision i is the precision for the i-th class; Normalize the dynamic weights and introduce the negative class sensitivity to construct the loss function, specifically: Normalize after fusing the recall rate and precision weight: The dynamic distribution balance loss is calculated as follows: wherein, is the predicted score of the i-th category of the k-th sample, is the true label of the i-th category of the k-th sample, and the true label is encoded using one-hot encoding; C is the total number of categories, and w i is the dynamic weight matrix, and λ is the hyperparameter used to adjust the loss of negative classes.
Citation Information
Patent Citations
Spatial spectrum attention hyperspectral image classification method based on Octave convolution
CN110516596A
Hyperspectral remote sensing image classification method
CN113705526A
Hyperspectral image classification method and device combining random shielding and BYOL structure
CN115115878A
Hyperspectral remote sensing image classification method based on attention joint network
CN115564996A
Multi-agent wave band selection method based on clustering and various experience pools
CN119445384A