An image classification method and system
The pseudo-label selection through feature extraction network PCET-Net and dynamic threshold strategy solves the problems of insufficient training samples and inaccurate pseudo-labels in hyperspectral image classification, and improves classification accuracy and robustness.
Patent Information
- Application Number
- CN202510517559.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing hyperspectral image classification models are difficult to obtain excellent classification results when training samples are insufficient, and the accuracy and consistency of pseudo-labels affect model performance. The fixed receptive field of CNN limits the ability to capture a wider contextual information, and the acquisition cost of labeled samples is high.
The feature extraction network PCET-Net is used, combined with superpixel segmentation and label propagation technology, and the output confidence of the module is evaluated, and the SSMF module is used to fuse spectral and spatial information, and the dynamic threshold strategy is used to select pseudo-labels to build a self-trained semi-supervised deep learning model.
It improves the accuracy and robustness of hyperspectral image classification, alleviates the problems of data distribution imbalance and category differences, and selects reliable pseudo-label samples, which enhances the feature extraction ability of the model.
Smart Images

Figure CN120047754B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and particularly to an image classification method and system. Background Art
[0002] Hyperspectral images (HSIs) record continuous and detailed high-resolution spectral bands of dozens or even hundreds of targets. They can not only observe information such as the shape and size of substances like optical images, but also utilize the characteristics that the physical components of substances and the internal interactions within the components have specific responses to specific wavelengths of electromagnetic wave radiation to distinguish the physical and chemical compositions and even structural information inside the substances. Hyperspectral images contain rich spatial-spectral features, which are beneficial to the recognition and distribution evaluation of objects in the target area, providing information basis for geological resource exploration, agricultural management, environmental monitoring, medical assistance, and military security defense, etc. Due to the extremely rich information contained in hyperspectral images, how to effectively extract meaningful spectral and spatial features has become the key to improving the classification accuracy and application effect of HSIs.
[0003] With the rapid development of deep learning technology and the improvement of computing power, many deep learning network models have been successfully applied to hyperspectral image classification, which can automatically learn deep features from data and avoid the process of manually designing feature extraction, such as Stacked Autoencoder (SAE), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Generative Adversarial Network (GAN), and so on. Among them, CNN performs excellently in hyperspectral image classification by virtue of its characteristics such as local connection, parameter sharing, and hierarchical structure. However, the fixed receptive field size of CNN limits its ability to capture more extensive context information and cannot effectively capture medium- and long-term dependencies. In addition, due to the high-dimensional spectral information of hyperspectral images, which can be regarded as continuous features, there are information losses and insufficient feature expressions when using CNN to process high-dimensional spectral data. In recent years, the Transformer model based on the attention mechanism has provided a more effective mechanism for processing long-range dependencies and global context, but most current research focuses on improving the Transformer model from the perspective of feature extraction. In current research, the input samples are usually the neighborhoods centered on each labeled pixel. When the neighborhood contains category information inconsistent with the label of the central pixel, it is likely to affect the classification results of the Transformer for hyperspectral images.
[0004] In addition, in practical applications, it is difficult to obtain labeled samples of hyperspectral images. The cost of manual annotation is very high and time-consuming and laborious. A typical problem in applying deep learning networks is that a large number of labeled samples are required to train a high-precision classification model. Therefore, how to ensure that the model can still obtain excellent classification results under the condition of insufficient training samples is also an issue worthy of consideration. A commonly used solution is to use limited training labeled samples to enhance the training of the model through semi-supervised learning, thereby improving classification accuracy. However, this method depends on the quality of pseudo-label samples, making its performance limited by the accuracy and consistency of pseudo-labels. Summary of the Invention
[0005] The object of the present invention is to provide an image classification method and system for improving classification accuracy.
[0006] To achieve the above object, the present invention adopts the following technical solution:
[0007] An image classification method, comprising the following steps executed in sequence:
[0008] S1: Obtain labeled samples in the hyperspectral image, construct a labeled data set, the labeled data set includes a training set and a test set, perform superpixel segmentation on the hyperspectral image, propagate the labels of the labeled samples in the training set to the unlabeled samples, and select unlabeled samples with balanced data distribution from them as the unlabeled data set;
[0009] S2: Construct a feature extraction network PCET-Net, the feature extraction network PCET-Net includes an evaluation module for evaluating the quality of image patches and outputting the confidence of image patches, an SSMF module for fusing spectral information and spatial information, and a classification module. Input the training set into the feature extraction network PCET-Net for training, output the classification category probability, perform weak augmentation processing on the samples of the unlabeled data set to obtain, perform strong augmentation processing on the samples of the unlabeled data set to obtain, respectively input and into the feature extraction network PCET-Net for training, compare the output category probability with the dynamic threshold adjusted by the dynamic threshold strategy. If the output category probability is greater than the dynamic threshold, the category obtains the corresponding pseudo-label, otherwise, the category does not obtain the pseudo-label; perform weak augmentation processing to obtain and perform strong augmentation processing on the samples of the unlabeled data set to obtain processing to obtain and input and into the feature extraction network PCET-Net for training respectively, and compare the output category probability with the dynamic threshold adjusted by the dynamic threshold strategy. If the output category probability is greater than the dynamic threshold, the category obtains the corresponding pseudo-label, otherwise, the category does not obtain the pseudo-label;
[0010] S3: Calculate the loss functions of the training set and the unlabeled dataset respectively according to the classification category probabilities output in step S2, and use the total loss function to reversely adjust the feature extraction network PCET-Net. The calculation method of the total loss function is as follows:
[0011] S3-1: Calculate the loss function of the training set using the following formula :
[0012] ;
[0013] where is the confidence level output by the evaluation module, represents the predicted value of the target category, represents the predicted value for category ; represents the exponential function;
[0014] S3-2: Calculate the loss function of the unlabeled dataset using the following formula :
[0015] ;
[0016] where represents the pseudo-label of the unlabeled sample, represents the corresponding dynamic threshold, is the weak augmentation of the th unlabeled sample, is the strong augmentation of the th unlabeled sample, is the total number of unlabeled samples, s.t. is the abbreviation of the English phrase "subject to", indicating "subject to" or "satisfying the following conditions", and max represents the maximization operation;
[0017] S3-3: The total loss function is calculated using the following formula:
[0018] ;
[0019] S4: Input the test set into the adjusted feature extraction network PCET-Net for image classification and output the classification result.
[0020] Preferably, the label propagation in step S1 specifically includes the following steps executed in sequence:
[0021] S1-1: Use simple linear iterative clustering to segment the hyperspectral image into a set of superpixels. The number of superpixel segments is obtained by the following formula:
[0022] ;
[0023] Among them, represents the impact factor, and its value affects the number of pixels in each superpixel. and respectively represent the length and width of the hyperspectral image;
[0024] S1-2: According to whether the superpixel contains labeled samples, these superpixels are divided into two cases. One case is that the segmentation region consists of superpixels including both labeled and unlabeled samples. Calculate the number of labeled samples, and propagate the label with the most corresponding labels to other unlabeled samples in the same superpixel. The other case is that the segmentation region only contains superpixels of unlabeled samples. The nearest neighbor algorithm is used to assign labels to the unlabeled samples in this superpixel. The distance between each superpixel and the labeled samples of each category is calculated through their average spectra. The distance between each superpixel and the labeled samples of each category is calculated by the following formula:
[0025] ;
[0026] Among them, represents the average spectrum of the th superpixel that only contains unlabeled samples, represents the norm of the average spectrum, represents the average spectrum of the labeled samples of the th category; the category label corresponding to the maximum distance is assigned to all pixels in the th superpixel that only contains unlabeled samples;
[0027] S1-3: According to the calculation in S1-2, all unlabeled superpixels obtain propagated labels through label propagation, and unlabeled samples are uniformly selected from each category as the unlabeled dataset.
[0028] Preferably, the evaluation module described in step S2 includes two branches. One branch fuses the spectral information and spatial information of the image block through the convolutional block attention module (CBAM) and outputs a fused feature map. The other branch sequentially performs feature extraction on the image block through the first pointwise convolution (PW), depthwise convolution (DW), and the second pointwise convolution (PW) to obtain a preliminary feature map. The preliminary feature map and the fused feature map are added pointwise and then average pooled. The feature map after average pooling is flattened and batch-normalized in sequence, and some neurons in the feature map are randomly discarded. The features are mapped using a fully connected layer mapping, and the mapping result is processed by a sigmoid function to obtain the confidence. , the specific operation process is as follows:
[0029] Assume the input image patch , denotes the image patch size, denotes the number of spectral dimensions. This convolutional block attention module CBAM includes a spectral attention module and a spatial attention module. The spectral attention module uses average pooling AvgPool and max pooling MaxPool operations to aggregate spatial information, thereby generating two context information feature maps with different spaces and . Input and into a shared multi-layer perceptron MLP, and fuse these two features through element-wise summation to generate a spectral attention map :
[0030] ;
[0031] Among them, denotes the sigmoid function, , denote average pooling and max pooling operations respectively;
[0032] The spectral attention map and the input image patch are weighted to obtain a feature map :
[0033] ;
[0034] This spatial attention module applies average pooling and max pooling operations along the spectral dimension to aggregate spectral information, generating two feature maps and , and and are concatenated in the spectral dimension to generate an effective feature description, and a convolutional layer is used to process the concatenated feature description to generate a spatial attention map :
[0035] ;
[0036] In the formula, denotes the sigmoid function, denotes 7 ×7 convolutional kernel, denotes the concatenation operation of features;
[0037] The spatial attention map and the feature map Obtain the fused feature map output by the convolutional attention module CBAM through weighting :
[0038] ;
[0039] The operations of the other branch are represented by the following formula:
[0040] ;
[0041] ;
[0042] ;
[0043] ;
[0044] In the formula, and respectively represent the first and second pointwise convolutional PW operations, represents the depth convolution operation, represents the result of the input passing through a series of and operations, represents the fused feature map output by the CBAM module, represents and after performing pointwise addition and average pooling result, represents the result of performing batch normalization and dropout on the flattened , represents the confidence level, represents the fully connected layer mapping, represents the sigmoid activation function, represents randomly discarding some neurons, represents batch normalization, represents flattening.
[0045] Preferably, the SSMF module includes four branches. The first branch uses convolution kernels to extract the spectral features of the image patch to obtain the spectral feature map. The second branch uses convolution kernels to extract the spatial information of the image patch to obtain the spatial feature map. The third branch performs average pooling operation, convolution operation and upsampling operation on the image patch in sequence to obtain the global description feature map. The fourth branch uses the global attention mechanism GAM to obtain the image patch Global context information is obtained to acquire the global feature map. The spectral feature map, spatial feature map, global descriptive feature map, and global feature map are feature-connected, and after performing a convolution operation on the concatenated feature map, a feature map is output , and the operation of the SSMF module is represented by the following formula:
[0046] ;
[0047] In the formula, represents the connection operation of features, , respectively represent and convolution operations, represents that the features sequentially pass through average pooling, convolution operation, and upsampling operation, represents the feature map output by the global attention mechanism.
[0048] Preferably, the global attention mechanism GAM stores the three-dimensional information of the image patch in a three-dimensional arrangement, enhances the spectral-spatial correlation using a multi-layer perceptron MLP with a two-layer encoder-decoder structure, and outputs the spectral attention feature using an inverse arrangement and a sigmoid activation function . The spectral attention feature is weighted with the image patch to obtain the feature map :
[0049] ;
[0050] Among them, represents element-wise multiplication;
[0051] The feature map is fused using two-layer convolution operations to output the spatial attention feature . The spatial attention feature is weighted with the feature map to obtain the feature map :
[0052] ;
[0053] Among them, represents element-wise multiplication.
[0054] Preferably, the classification module includes a sequentially connected convolutional tokenization, linear projection layer, EA 2 T module, and classification multi-layer perceptron module;
[0055] The convolutional tokenization and the linear projection layer perform operations on the feature map For mapping, the neighborhood around the central pixel is divided into different spatial tokens: , , and these spatial tokens are concatenated with the classification token , and the position information is embedded into the above spatial tokens to generate a token sequence suitable for processing by the EA 2 T module, which is represented by the following formula:
[0056] ;
[0057] The EA 2 T module first performs layer normalization on the token sequence to generate , where , and the input feature is transformed into and using two matrices and , where , , , is the dimension of the feature map, and the query matrix is multiplied by the learnable parameter vector to learn the query attention weights, resulting in the global attention query vector which is represented by the following formula:
[0058] ;
[0059] The global attention query vector is multiplied by the query matrix and pooled to generate the global query vector which is represented by the following formula:
[0060] ;
[0061] The global query vector is multiplied element-wise by the matrix to construct the global context, introducing the interaction between the linear layer processing to learn the hidden representation of the tokens, resulting in the matrix 2 output by the EA whose output is shown by the following formula:
[0062] ;
[0063] where represents the normalized query matrix, and represents the linear mapping;
[0064] Input the matrix into the classification multi-layer perceptron for classification mapping, and output the classification category probability.
[0065] Preferably, the dynamic threshold strategy includes the following calculation steps:
[0066] Evaluate the classification difficulty of the th class by calculating the number of samples with classification probability greater than the category adaptive threshold :
[0067] ;
[0068] where is the indicator function, which takes the value of 1 when the condition in the square brackets is satisfied, and 0 otherwise, is the total number of pseudo-labeled samples, is a preset threshold, and s.t. is the abbreviation of the English phrase "subject to", indicating "constrained by" or "satisfying the following conditions", means that the category predicted by the model is equal to ;
[0069] Normalize to the range using the following formula:
[0070] ;
[0071] where represents the normalized , represents maximizing the classification difficulty of the c-th class;
[0072] Construct weights using the mapping to scale the initial threshold , and the dynamic threshold is obtained using the following formula:
[0073] ;
[0074] For , the mapping is a non-linear convex function and is mapped using the following formula:
[0075] ;
[0076] where is used to determine the value of when.
[0077] An image classification system includes a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, the image classification method described in any one of the above is implemented.
[0078] By adopting the foregoing design scheme, the beneficial effects of the present invention are as follows:
[0079] 1. This application uses self-training semi-supervised deep learning. The feature extraction network PCET-Net is adopted to select reliable unlabeled samples to enhance the performance of the model, thereby improving the classification accuracy.
[0080] 2. This application adopts superpixel-guided label propagation and dynamic threshold strategies to alleviate the problems of unbalanced data distribution and large class differences, thereby better selecting reliable pseudo-label samples.
[0081] 3. This application constructs a feature extraction network PCET-Net for feature extraction and classification, introduces an evaluation module to evaluate the quality of the input image patches and output confidence as a strategy for training weights, reduces the interference of heterogeneous pixels at the class boundary, and proposes an SSMF module to further enhance the feature expression ability of the feature extraction network PCET-Net, improving the classification accuracy and robustness. Description of the Drawings
[0082] Figure 1 It is a block diagram of the overall network model architecture of the present invention;
[0083] Figure 2 It is a comparison diagram of the original image of the present invention and SLIC superpixel segmentation;
[0084] Figure 3 It is an example diagram of the label propagation process of the present invention;
[0085] Figure 4 It is a schematic diagram of the model structure of the feature extraction network PCET-Net of the present invention;
[0086] Figure 5 It is a schematic diagram of the structure of the convolutional attention module CBAM of the present invention;
[0087] Figure 6 It is a schematic diagram of the structure of the SSMF module of the present invention;
[0088] Figure 7 It is a schematic diagram of the structure of the global attention mechanism GAM of the present invention;
[0089] Figure 8 It is the EA 2 T module and a comparison diagram of different attention mechanisms. Detailed Embodiments
[0090] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0091] The terms "first", "second", "third", etc. in the specification, claims, and above-mentioned drawings of the present invention are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.
[0092] An image classification method, as Figure 1 shown, includes the following steps executed in sequence:
[0093] S1: Obtain labeled samples in a hyperspectral image, construct a labeled data set, where the labeled data set includes a training set and a test set, perform superpixel segmentation on the hyperspectral image, propagate the labels of the labeled samples in the training set to unlabeled samples, and select unlabeled samples with balanced data distribution therefrom as an unlabeled data set;
[0094] As Figure 3 shown, the label propagation in step S1 specifically includes the following steps executed in sequence:
[0095] S1-1: Use Simple Linear Iterative Clustering (SLIC) to segment the hyperspectral image (HSI) into a set of superpixels, which are composed of non-overlapping regions of different shapes and sizes. As Figure 2 shown, the number of superpixel segments is obtained by the following formula:
[0096] ;
[0097] where represents an influence factor, and its value affects the number of pixels in each superpixel, usually set to 50, and represent the length and width of the hyperspectral image respectively;
[0098] S1-2: According to whether the superpixels contain labeled samples, where the labeled samples in this step refer to the labeled samples in the training set, these superpixels are divided into two cases. One case is that the segmentation region consists of superpixels including both labeled samples and unlabeled samples, such as Figure 3 in and . The pixels within the superpixels are highly similar and may belong to the same category. Therefore, calculate the number of superpixels with known labeled samples, and propagate the label with the most corresponding labels to other unlabeled samples in the same superpixel.
[0099] The other case is that the segmentation region only contains superpixels of unlabeled samples, such as Figure 3 in , and . The nearest neighbor algorithm is used to assign labels to the unlabeled samples of this superpixel. The distance between each superpixel and the labeled samples of each category is calculated through their average spectra. The distance between each superpixel and the labeled samples of each category is calculated by the following formula:
[0100] ;
[0101] where, represents the average spectrum of the th superpixel containing only unlabeled samples, represents the modulus of the average spectrum, represents the average spectrum of the labeled samples of the th category; the category label corresponding to the maximum distance is assigned to all pixels in the th superpixel containing only unlabeled samples; represents the cosine function of the included angle; the larger the distance , the higher the similarity between and , and vice versa. Therefore, the label c corresponding to the maximum distance is assigned to the superpixel , and this label is propagated to all pixels in the superpixel region.
[0102] S1-3: According to the calculations in S1-2, all unlabeled superpixels obtain propagated labels through label propagation, and unlabeled samples are uniformly selected from each category as the unlabeled dataset.
[0103] Since random selection may lead to unbalanced data distribution, it is crucial to reasonably select unlabeled samples for semi-supervised learning. To mitigate the negative impact of the unbalanced distribution, label propagation is adopted. Guided by the hyperspectral superpixel segmentation technology based on spectral-spatial similarity, the limited known labeled samples are propagated to the entire hyperspectral image to achieve balanced sampling of samples of each category.
[0104] S2: Construct the feature extraction network PCET-Net, as Figure 4 shown. The feature extraction network PCET-Net includes an evaluation module for evaluating the quality of image patches and outputting the confidence of image patches, an SSMF module for fusing spectral information and spatial information, and a classification module.
[0105] Input the training set into the feature extraction network PCET-Net for training, output the classification category probability, and perform weak augmentation processing on the samples of the unlabeled dataset to obtain , and perform strong augmentation processing on the samples of the unlabeled dataset to obtain , where weak augmentation is to randomly flip the samples vertically and horizontally, and strong augmentation is to add random noise on the basis of weak augmentation. Input and into the feature extraction network PCET-Net for training respectively, compare the output category probability with the dynamic threshold adjusted by the dynamic threshold strategy. If the output category probability is greater than the dynamic threshold, the category will obtain the corresponding pseudo-label; otherwise, the category will not obtain the pseudo-label.
[0106] In this embodiment, the evaluation module in step S2 includes two branches. One branch fuses spectral information and spatial information of the image patch through the convolutional block attention module CBAM and outputs a fused feature map. The other branch extracts features from the image patch through the first pointwise convolution PW, depthwise convolution DW, and the second pointwise convolution PW in sequence to obtain a preliminary feature map. Add the preliminary feature map and the fused feature map pointwise and perform average pooling. Perform flattening and batch normalization on the feature map after average pooling in sequence, and randomly discard some neurons in the feature map. Map the features using the fully connected layer mapping, and the mapping result is processed by the sigmoid function to obtain the confidence , and the specific operation process is as follows:
[0107] Assume the input image patch , represents the size of the image patch, represents the number of spectral dimensions, such as Figure 5As shown, the convolutional attention module CBAM includes a spectral attention module and a spatial attention module. The spectral attention module utilizes the inter-band relationships between features to generate a spectral attention map, and aggregates spatial information through average pooling (AvgPool) and max pooling (MaxPool) operations to generate two context information feature maps in different spaces. and , input and into a shared multi-layer perceptron (MLP) which has a hidden layer with a hidden activation size of , representing the compression ratio of the MLP, and fuse these two features through element-wise summation to generate a spectral attention map :
[0108] ;
[0109] wherein, represents the sigmoid function, , represent the average pooling and max pooling operations respectively.
[0110] The spectral attention map and the input image patch are weighted to obtain a feature map :
[0111] ;
[0112] The spatial attention module aggregates spectral information by applying average pooling and max pooling operations along the spectral dimension to generate two feature maps and , and and are concatenated in the spectral dimension to generate an effective feature description, and the concatenated feature description is processed using a convolutional layer to generate a spatial attention map :
[0113] ;
[0114] In the formula, represents the sigmoid function, represents a 7 ×7 convolutional kernel, represents the concatenation operation of features;
[0115] The spatial attention map and the feature map are weighted to obtain the fused feature map output by the convolutional attention module CBAM:
[0116] ;
[0117] By generating corresponding attention maps in the spectral and spatial dimensions, CBAM can dynamically adjust the weights of each part in the feature map. Subsequently, these attention maps are fused with the input feature map, enabling the model to pay more attention to the important local information in the input feature map.
[0118] The operations of the other branch are expressed by the following formula:
[0119] ;
[0120] ;
[0121] ;
[0122] ;
[0123] In the formula, and represent the first and second pointwise convolution PW operations respectively, represents the depth convolution operation, represents the result of the input passing through a series of and operations, represents the fused feature map output by the CBAM module, represents the result of performing pointwise addition on and and then performing average pooling on the result, represents performing batch normalization and dropout processing on the flattened , represents the confidence, represents the fully connected layer mapping, represents the sigmoid activation function, represents randomly discarding some neurons, represents batch normalization, represents flattening.
[0124] This evaluation module is used to evaluate the input image patch and output its confidence Confidence. A higher confidence indicates a greater positive contribution of the image patch to the training, and this confidence is used in the calculation of the loss function so that the model can dynamically adjust the learning focus according to the quality of the input image patch.
[0125] Depthwise convolution DW performs independent convolution on each spectral dimension using corresponding convolution kernels, greatly reducing the number of parameters compared to standard convolution; pointwise convolution PW then uses convolution kernels to merge and remap the feature information of different channels, thereby realizing the fusion of feature information in different spectral dimensions.
[0126] In this embodiment, in the feature extraction stage, in order to make full use of the rich spectral and spatial information contained in the hyperspectral image and capture complementary and representative features from the image patches, an SSMF module is designed to fuse the spectral and spatial information, which can enhance the model's ability to distinguish different classes, so that the model can better capture the comprehensive features of substances.
[0127] As Figure 6 shown, the SSMF module includes four branches, which perform different processes on the input image patch to generate 4 independent feature maps, aiming to extract feature information from different perspectives. The first branch uses convolution kernels to extract the spectral features of the image patch to obtain a spectral feature map. The second branch uses convolution kernels to extract the spatial information of the image patch to obtain a spatial feature map. The third branch performs average pooling operation, convolution operation and upsampling operation on the image patch in sequence to obtain a global description feature map. The fourth branch uses the global attention mechanism GAM to obtain the global context information of the image patch to obtain a global feature map. The spectral feature map, spatial feature map, global description feature map and global feature map are feature-connected, and after performing a convolution operation on the concatenated feature map, the feature map is output. The operation of the SSMF module is represented by the following formula:
[0128] ;
[0129] In the formula, represents the operation of feature connection, , respectively represent the convolution operations of and , represents that the feature passes through average pooling, convolution operation and upsampling operation in sequence, represents the feature map output by the global attention mechanism.
[0130] The SSMF module integrates spectral and spatial information, and at the same time considers local details and global background, so as to be able to extract multi-scale feature representations of hyperspectral images and form a high-resolution feature set.
[0131] The global attention mechanism GAM comprehensively considers the cross-dimensional information of features and enhances the interaction between global dimensions, such as Figure 7 shown. The global attention mechanism GAM stores the three-dimensional information of image patches in a three-dimensional arrangement and uses a multi-layer perceptron MLP with a two-layer encoder-decoder structure to enhance the spectral-spatial correlation, and its hidden activation size is , which represents its compression ratio, and outputs spectral attention features using an inverse arrangement and a sigmoid activation function . The spectral attention features are weighted with the image patches to obtain a feature map :
[0132] ;
[0133] wherein, represents element-wise multiplication;
[0134] The feature map is fused using two-layer convolutional operations, and spatial attention features are output. The spatial attention features are weighted with the feature map to obtain a feature map :
[0135] ;
[0136] wherein, represents element-wise multiplication.
[0137] In this embodiment, the classification module includes a convolutional tokenization, a linear projection layer, an EA 2 T module, and a classification multi-layer perceptron module connected in sequence;
[0138] The convolutional tokenization and the linear projection layer map the feature map . The neighborhood around the central pixel is divided into different spatial tokens: , , and these spatial tokens are concatenated with the classification token which is used to perform the classification task. In order to provide location information and embed the location information into the above spatial tokens, a token sequence 2 suitable for processing by the EA T module is generated and is represented by the following formula:
[0139] ;
[0140] such as Figure 8As shown, EA 2 The T module first performs layer normalization on the token sequence to generate , where , taking as the input feature of the EA 2 T module, using two matrices and to transform the input feature into and , where , , is the dimension of the feature map, the query matrix multiplies the learnable parameter vector to learn the query attention weights, obtaining the global attention query vector represented by the following formula:
[0141] ;
[0142] The global attention query vector multiplies the query matrix and is pooled to generate the global query vector represented by the following formula:
[0143] ;
[0144] The global query vector and the matrix are element-wise multiplied to construct the global context, introducing linear layer processing between the interactions to learn the hidden representation of the tokens, obtaining the matrix 2 output by the EA output is shown by the following formula:
[0145] ;
[0146] Among them, represents the normalized query matrix, represents the linear mapping;
[0147] The matrix is input into the classification multi-layer perceptron for classification mapping, and the classification category probability is output.
[0148] The dynamic threshold strategy includes the following calculation steps:
[0149] By calculating the number of samples with classification probability greater than the class adaptive threshold , to evaluate the classification difficulty of the th class :
[0150] ;
[0151] wherein, represents the indicator function, which takes the value of 1 when the condition in the square brackets is satisfied, and 0 otherwise. is the total number of unlabeled samples. is the preset threshold, and s.t. is the abbreviation of the English phrase "subject to", indicating "constrained by" or "satisfying the following conditions". means that the class predicted by the model is equal to ;
[0152] The following formula is used to normalize to the range:
[0153] ;
[0154] wherein, represents the normalized , represents maximizing the classification difficulty of the c-th class; it can be seen that the class with fewer samples corresponds to a lower . To ensure that each class can obtain sufficient learnable labeled samples, for the class with fewer samples, its sample quantity should be increased; for the class with more samples, information redundancy should be reduced. This method can balance the sample quantity of each class and improve the discrimination and robustness of the model.
[0155] Weights are constructed using the mapping to scale the initial threshold , and the dynamic threshold is obtained using the following formula:
[0156] ;
[0157] It can be seen that the threshold is positively correlated with the weight . Since the "easy" class with a large number of samples has a high threshold, while the "difficult" class with a small number of samples has a lower threshold, when , the mapping should be a monotonically increasing function. In addition, for the "difficult" class, the weight rapidly decreases as decreases, in order to allocate more pseudo-labels for learning. For the "easy" class, the weight approaches the maximum value of 1, meaning that the weight increases slowly as increases. Therefore, for , the mapping should be a non-linear convex function.
[0158] Based on the above considerations, for , the mapping is a non-linear convex function, and the following formula is used for the mapping:
[0159] ;
[0160] Among them, is used to determine the value when , and the value of its parameter can be selected through ablation experiments.
[0161] S3: Calculate the loss functions of the training set and the unlabeled dataset respectively according to the classification category probabilities output in step S2, and use the total loss function to adjust the feature extraction network PCET-Net backward. The calculation method of the total loss function is as follows:
[0162] S3-1: Calculate the loss function of the training set using the following formula :
[0163] ;
[0164] Among them, is the confidence level output by the evaluation module, represents the predicted value of the target category, represents the predicted value for the category , represents the exponential function;
[0165] S3-2: Calculate the loss function of the unlabeled dataset using the following formula :
[0166] ;
[0167] Among them, represents the pseudo-label of the unlabeled sample, represents the corresponding dynamic threshold, is the weak augmentation of the th unlabeled sample, is the strong augmentation of the th unlabeled sample, is the total number of unlabeled samples, s.t. is the abbreviation of the English phrase "subject to", which means "subject to" or "satisfies the following conditions", and max represents the maximization operation;
[0168] S3-3: The overall loss function consists of two parts: the supervised cross-entropy loss on the labeled sample training set and the unsupervised cross-entropy loss on unlabeled samples , the total loss function is calculated using the following formula:
[0169] ;
[0170] S4: Input the test set into the adjusted feature extraction network PCET-Net for image classification and output the classification result.
[0171] In this embodiment, a system for implementing the above method is also provided.
[0172] An image classification system includes a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, the image classification method described in any one of the above is implemented.
[0173] In summary, this application only requires a small number of labeled samples, alleviates the data distribution imbalance problem through superpixel-guided label propagation, constructs a high-precision and high-robustness feature extraction network PCET-Net based on consistency regularization, uses dynamic thresholds to select highly reliable unlabeled samples to enhance model training in a supervised manner, and improves the classification accuracy.
[0174] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. An image classification method, characterized in that: It includes the following steps executed sequentially: S1: Obtain the labeled samples in the hyperspectral image, construct a labeled data set, the labeled data set includes a training set and a test set, perform superpixel segmentation on the hyperspectral image, propagate the labels of the labeled samples in the training set to the unlabeled samples, and select the unlabeled samples with balanced data distribution from them as the unlabeled data set; S2: Construct a feature extraction network PCET-Net, which includes an evaluation module for evaluating the quality of image patches and outputting the confidence of image patches, an SSMF module for fusing spectral information and spatial information, and a classification module. Input the training set into the feature extraction network PCET-Net for training, output the classification category probability, and for the samples of the unlabeled dataset perform weak augmentation processing to obtain , and for the samples of the unlabeled dataset perform strong augmentation processing to obtain , and input and into the feature extraction network PCET-Net for training respectively. Compare the output category probability with the dynamic threshold adjusted by the dynamic threshold strategy. If the output category probability is greater than the dynamic threshold, the category will obtain the corresponding pseudo label; otherwise, the category will not obtain the pseudo label. S3: Calculate the loss functions of the training set and the unlabeled data set respectively according to the classification category probabilities output in step S2, and use the total loss function to reversely adjust the feature extraction network PCET-Net. The calculation method of the total loss function is as follows: S3-1: Calculate the loss function of the training set using the following formula : ; Among them, is the confidence level output by the evaluation module, represents the predicted value of the target category, represents for the category the predicted value, represents the exponential function; S3-2: Calculate the loss function of the unlabeled dataset using the following formula : ; Among them, represents the pseudo-label of the unlabeled sample, represents the corresponding dynamic threshold, is the weak augmentation prediction probability of the is the strong augmentation prediction probability of the is the total number of unlabeled samples, s.t. is the abbreviation of the English phrase "subject to", indicating "constrained by" or "satisfying the following conditions", and max represents the maximization operation; S3-3: Total loss function It is calculated using the following formula: ; S4: Input the test set into the adjusted feature extraction network PCET-Net for image classification and output the classification result.
2. The image classification method according to claim 1, wherein: The label propagation in step S1 specifically includes the following steps executed sequentially: S1-1: Segment the hyperspectral image into a set of superpixels using simple linear iterative clustering, and the number of superpixel segments is obtained by the following formula: ; Among them, represents the influence factor, and its value affects the number of pixels of each superpixel, and represent the length and width of the hyperspectral image respectively; S1-2: According to whether the superpixels contain labeled samples, these superpixels are divided into two cases. In one case, the segmentation region consists of superpixels including both labeled and unlabeled samples. Calculate the number of labeled samples, and propagate the label with the most corresponding labels to other unlabeled samples in the same superpixel. In the other case, the segmentation region only contains superpixels of unlabeled samples. The nearest neighbor algorithm is used to assign labels to the unlabeled samples of this superpixel. The distance between each superpixel and the labeled samples of each category is calculated through their average spectra. The distance between each superpixel and the labeled samples of each category is calculated by the following formula: ; Among them, represents the average spectrum of the th superpixel containing only unlabeled samples, represents the modulus of the average spectrum, represents the average spectrum of the labeled samples of the th class; the class label corresponding to the maximum distance is assigned to all pixels in the th superpixel containing only unlabeled samples; S1-3: According to the calculation in S1-2, all unlabeled superpixels obtain propagated labels through label propagation, and unlabeled samples are uniformly selected from each category as the unlabeled data set.
3. The image classification method according to claim 1, wherein: The evaluation module described in step S2 includes two branches. One branch fuses the spectral information and spatial information of the image patch through the Convolutional Block Attention Module (CBAM) and outputs a fused feature map. The other branch sequentially extracts features from the image patch through the first pointwise convolution (PW), depthwise convolution (DW), and the second pointwise convolution (PW) to obtain a preliminary feature map. The preliminary feature map and the fused feature map are added pointwise and then average pooled. The feature map after average pooling is flattened and batch-normalized in sequence, and some neurons in the feature map are randomly discarded. The features are mapped using a fully connected layer mapping, and the mapping result is processed by a sigmoid function to obtain a confidence level. , and the specific operation process is as follows: Assume the input image patch , denotes the image patch size, denotes the number of spectral dimensions. The convolutional block attention module CBAM includes a spectral attention module and a spatial attention module. The spectral attention module uses average pooling AvgPool and max pooling MaxPool operations to aggregate spatial information, thereby generating two context information feature maps with different spaces and . Feed and into a shared multi-layer perceptron MLP, and fuse these two features by element-wise summation to generate a spectral attention map : ; Among them, represents the sigmoid function, , represent average pooling and max pooling operations respectively; Spectral attention map With the input image patch The feature map obtained by weighting : ; Among them, the spatial attention module aggregates spectral information by applying average pooling and max pooling operations along the spectral dimension to generate two feature maps and , and and are concatenated in the spectral dimension to generate an effective feature description, and a convolutional layer is used to process the concatenated feature description to generate a spatial attention map : ; In the formula, represents the sigmoid function, represents 7 7 convolutional kernels, represents the concatenation operation of features; Spatial attention map and the feature map are weighted to obtain the fused feature map output by the convolutional attention module CBAM : ; Among them, the operation of the other branch is represented by the following formula: ; ; ; ; In the formula, and respectively represent the first and second pointwise convolution PW operations, represents the depth convolution operation, represents the result of the input passing through a series of and operations, represents the fused feature map output by the CBAM module, represents the result of and being pointwise added and then average pooled , represents batch normalization and dropout processing on the flattened , represents the confidence, represents the fully connected layer mapping, represents the sigmoid activation function, represents randomly discarding some neurons, represents batch normalization, represents flattening.
4. The image classification method according to claim 3, wherein: The SSMF module includes four branches. The first branch uses a convolutional kernel to extract the spectral features of the image patch and obtain a spectral feature map. The second branch uses a convolutional kernel to extract the spatial information of the image patch and obtain a spatial feature map. The third branch performs average pooling operation, convolutional operation, and upsampling operation on the image patch in sequence to obtain a global descriptive feature map. The fourth branch uses the global attention mechanism GAM to obtain the global context information of the image patch and obtain a global feature map. The spectral feature map, spatial feature map, global descriptive feature map, and global feature map are feature-connected, and after performing a convolutional operation on the spliced feature map, a feature map is output. The operation of the SSMF module is represented by the following formula: ; In the formula, represents the connection operation of features, , respectively represent and convolution operations of, represents that the features sequentially pass through average pooling, convolution operation and upsampling operation, represents the feature map output by the global attention mechanism.
5. The image classification method according to claim 4, wherein: The global attention mechanism GAM stores the image patches in a three-dimensional arrangement of the three-dimensional information, and uses a multi-layer perceptron MLP with a two-layer encoder-decoder structure to enhance the spectral-spatial correlation, and outputs spectral attention features using an inverse arrangement and a sigmoid activation function , and the spectral attention features are weighted with the image patches to obtain a feature map : ; Among them, represents element-wise multiplication; Use two layers of convolutional operations to fuse the feature map and output the spatial attention feature . Then, weight the spatial attention feature with the feature map to obtain the feature map : ; Among them, represents element-wise multiplication.
6. The image classification method according to claim 5, wherein: The classification module includes a convolutional tokenization, a linear projection layer, an EA 2 T module, and a classification multi-layer perceptron module, which are connected in sequence; The convolutional tokenization and the linear projection layer map the feature map such that the neighborhood around the central pixel is divided into different spatial tokens: , These spatial tokens are concatenated with the classification tokens and the position information is embedded into the above spatial tokens to generate a token sequence suitable for processing by the EA 2 T module, which is represented by the following formula: as follows: ; EA 2 The T module first performs layer normalization on the token sequence to generate , where , using two matrices and to transform the input feature into and , where , , is the dimension of the feature map, and the query matrix multiplies the learnable parameter vector to learn the query attention weights, obtaining the global attention query vector which is represented by the following formula: ; Global attention query vector Multiply by the query matrix And pool to produce the global query vector Which is represented by the following formula: ; Multiply the global query vector and the matrix element-wise to construct the global context, introduce a linear layer to process the interaction between to learn the hidden representation of the tokens, and obtain the output of the EA 2 The matrix output by the T module is given by the following formula: ; Among them, represents a normalized query matrix, represents a linear mapping; Input the matrix into the classification multi-layer perceptron for classification mapping, and output the classification category probability.
7. The image classification method according to claim 6, wherein: The dynamic threshold strategy includes the following calculation steps: By calculating the number of samples with classification probabilities greater than the class-adaptive threshold , to evaluate the classification difficulty of the th class ; Among them, represents the indicator function, which takes the value of 1 when the condition in the square brackets is satisfied and 0 otherwise. is the total number of unlabeled samples. is a preset threshold. s.t. is the abbreviation of the English phrase "subject to", which means "constrained by" or "satisfying the following conditions". indicates that the class predicted by the model is equal to ; Normalize to range: ; Among them, represents the normalized , represents maximizing the classification difficulty of the c-th class; Using mapping Construct weights to scale the initial threshold , the dynamic threshold is obtained using the following formula: ; For , the mapping is a non-linear convex function, and the mapping is performed using the following formula: ; Among them, for determining the value at 8. An image classification system, comprising a memory and a processor, wherein a computer program is stored on the memory, and characterized in that: When the computer program is executed by a processor, it implements the image classification method according to any one of claims 1-7 above.
Citation Information
Patent Citations
Hyperspectral image classification method and device, equipment and storage medium
CN113361481A
Hyperspectral image classification method and device based on spatial-spectral double-branch convolutional network
CN115249332A