Image classification method and system

By adopting a self-trained semi-supervised deep learning method and feature extraction network PCET-Net in hyperspectral image classification, the superpixel-guided label propagation and dynamic thresholding strategies are used to solve the problems of insufficient information capture and difficult to obtain labeled samples in hyperspectral image classification, and high-precision and robust image classification are achieved.

CN120047754AActive Publication Date: 2025-05-27QUANZHOU INST OF EQUIP MFG +1

Patent Information

Application Number
CN202510517559.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In hyperspectral image classification, the prior art is difficult to effectively capture medium- and long-term dependencies and global context information, and because labeled samples are difficult to obtain, deep learning models have low classification accuracy when training samples are insufficient.

Method used

A self-trained semi-supervised deep learning method is used to extract and classify features through the feature extraction network PCET-Net. The network includes an evaluation module, an SSMF module and a classification module, which utilizes superpixel-guided label propagation and dynamic thresholding strategies to select reliable unlabeled samples for weak and strong enhancement processing, improving the performance of the model.

Benefits of technology

It improves the accuracy and robustness of hyperspectral image classification, reduces the dependence on a large number of labeled samples, and can obtain excellent classification results when there are insufficient training samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047754A_ABST
    Figure CN120047754A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, in particular to an image classification method and system, and the method comprises the following steps: S1, obtaining a labeled sample in a hyperspectral image, constructing a labeled data set which comprises a training set and a test set, transmitting a label of the labeled sample of the training set to an unlabeled sample, unmarked samples with balanced data distribution are selected from the unmarked samples to serve as an unmarked data set; s2, constructing a feature extraction network PCET-Net, inputting the training set and the unmarked data set into the feature extraction network PCET-Net for training, and outputting a classification category probability; s3, reversely adjusting the feature extraction network PCET-Net by adopting a total loss function; S4, inputting the test set into the adjusted feature extraction network PCET-Net for image classification, and outputting a classification result; according to the method, self-training semi-supervised deep learning is adopted, and the feature extraction network PCET-Net is adopted to select the reliable unmarked sample to enhance the performance of the model, so that the classification precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to an image classification method and system. Background Art

[0002] Hyperspectral images (HSIs) record continuous and finely divided high-resolution spectral bands of dozens or even hundreds of targets. They can not only observe information such as the shape and size of substances like optical images, but also utilize the characteristics that the physical components of substances and the internal interactions within the components have specific responses to specific wavelengths of electromagnetic wave radiation to distinguish the physical and chemical compositions and even structural information inside the substances. Hyperspectral images contain rich spatial-spectral features, which are beneficial to the recognition and distribution evaluation of objects in the target area, providing information basis for geological resource exploration, agricultural management, environmental monitoring, medical assistance, and military security defense, etc. Due to the extremely rich information contained in hyperspectral images, how to effectively extract meaningful spectral and spatial features has become the key to improving the classification accuracy and application effect of HSIs.

[0003] With the rapid development of deep learning technology and the improvement of computing power, many deep learning network models have been successfully applied to hyperspectral image classification, which can automatically learn deep features from data, avoiding the process of manually designing feature extraction, such as Stacked Autoencoder (SAE), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Generative Adversarial Network (GAN), and so on. Among them, CNN performs excellently in hyperspectral image classification by virtue of its characteristics such as local connection, parameter sharing, and hierarchical structure. However, the fixed receptive field size of CNN limits its ability to capture more extensive context information and cannot effectively capture medium- and long-term dependencies. In addition, due to the high-dimensional spectral information of hyperspectral images, which can be regarded as continuous features, there are information losses and insufficient feature expressions when using CNN to process high-dimensional spectral data. In recent years, the Transformer model based on the attention mechanism has provided a more effective mechanism for processing long-distance dependencies and global context, but most current research focuses on improving the Transformer model from the perspective of feature extraction. In current research, the input samples are usually the neighborhoods centered on each labeled pixel. When the neighborhood contains category information inconsistent with the label of the central pixel, it is likely to affect the classification result of the Transformer for hyperspectral images.

[0004] In addition, in practical applications, it is difficult to obtain labeled samples of hyperspectral images. The cost of manual annotation is very high, time-consuming and laborious. A typical problem in applying deep learning networks is that a large number of labeled samples are required to train a high-precision classification model. Therefore, how to ensure that the model can still obtain excellent classification results under the condition of insufficient training samples is also an issue worthy of consideration. The commonly used solution is to use limited training labeled samples to enhance the training of the model through semi-supervised learning, thereby improving classification accuracy. However, this method depends on the quality of pseudo-label samples, making its performance limited by the accuracy and consistency of pseudo-labels. Summary of the Invention

[0005] The purpose of the present invention is to provide an image classification method and system for improving classification accuracy.

[0006] To achieve the above object, the present invention adopts the following technical solution: An image classification method, comprising the following steps executed in sequence: S1: Obtain labeled samples in the hyperspectral image, construct a labeled data set, the labeled data set includes a training set and a test set, perform superpixel segmentation on the hyperspectral image, propagate the labels of the labeled samples in the training set to unlabeled samples, and select unlabeled samples with balanced data distribution therefrom as the unlabeled data set; S2: Construct a feature extraction network PCET-Net. This feature extraction network PCET-Net includes an evaluation module for evaluating the quality of image patches and outputting the confidence of image patches, an SSMF module for fusing spectral information and spatial information, and a classification module. Input the training set into this feature extraction network PCET-Net for training, output the classification category probability, perform weak augmentation processing on the unlabeled data set to obtain and perform strong augmentation processing on the unlabeled data set to obtain , and and are respectively input into this feature extraction network PCET-Net for training, and the output category probability is compared with the dynamic threshold adjusted by the dynamic threshold strategy. If the output category probability is greater than the dynamic threshold, this category obtains the corresponding pseudo-label; otherwise, this category does not obtain a pseudo-label; S3: Calculate the loss functions of the training set and the unlabeled data set respectively according to the classification category probability output in step S2, and use the total loss function to reversely adjust the feature extraction network PCET-Net. The calculation method of the total loss function is as follows: S3-1: Calculate the loss function of the training set using the following formula : ; Where is the confidence level output by the evaluation module, represents the predicted value of the target category, represents for the category the predicted value, represents the exponential function; S3-2: Calculate the loss function of the unlabeled dataset using the following formula : ; where, represents the pseudo-label of the unlabeled sample, represents the corresponding dynamic threshold, is the weak augmentation for the th unlabeled sample, is the strong augmentation for the th unlabeled sample, is the total number of unlabeled samples, s.t. is the abbreviation of the English phrase "subject to", indicating "subject to" or "satisfying the following conditions", and max represents the maximization operation; S3-3: The total loss function is calculated using the following formula: ; S4: Input the test set into the adjusted feature extraction network PCET-Net for image classification and output the classification result.

[0007] Preferably, the label propagation in step S1 specifically includes the following steps executed in sequence: S1-1: Use simple linear iterative clustering to segment the hyperspectral image into a set of superpixels, and the number of superpixel segments is obtained by the following formula: ; where, represents the influence factor, and its value affects the number of pixels in each superpixel, and respectively represent the length and width of the hyperspectral image; S1-2: According to whether the superpixel contains labeled samples, divide these superpixels into two cases. One case is that the segmentation region consists of superpixels including both labeled and unlabeled samples. Calculate the number of labeled samples and propagate the label with the most corresponding labels to other unlabeled samples in the same superpixel. The other case is that the segmentation region only contains superpixels of unlabeled samples. Use the nearest neighbor algorithm to assign labels to the unlabeled samples in this superpixel. The distance between each superpixel and the labeled samples of each category Calculated through their average spectra, the distance between each superpixel and the labeled samples of each category is calculated by the following formula: ; wherein, represents the average spectrum of the -th superpixel containing only unlabeled samples, represents the norm of the average spectrum, represents the average spectrum of the labeled samples of the -th category; the category label corresponding to the maximum distance is assigned to all pixels in the -th superpixel containing only unlabeled samples; S1-3: According to the calculation in S1-2, all unlabeled superpixels have obtained propagated labels through label propagation, and unlabeled samples are uniformly selected from each category as the unlabeled dataset.

[0008] Preferably, the evaluation module in step S2 includes two branches. One branch fuses spectral information and spatial information of the image patch through the convolutional block attention module (CBAM) and outputs a fused feature map. The other branch sequentially performs feature extraction on the image patch through the first pointwise convolution (PW), depthwise convolution (DW), and the second pointwise convolution (PW) to obtain a preliminary feature map. The preliminary feature map and the fused feature map are added pointwise and then average pooled. The feature map after average pooling is sequentially flattened and batch-normalized, and some neurons in the feature map are randomly discarded. The features are mapped using a fully connected layer mapping, and the mapping result is processed by a sigmoid function to obtain the confidence , and the specific operation process is as follows: Assume the input image patch , represents the size of the image patch, represents the number of spectral dimensions. This convolutional block attention module (CBAM) includes a spectral attention module and a spatial attention module. The spectral attention module uses average pooling (AvgPool) and max pooling (MaxPool) operations to aggregate spatial information, thereby generating two context information feature maps with different spaces and . and are input into a shared multi-layer perceptron (MLP), and these two features are fused by element-wise summation to generate a spectral attention map : ; wherein, represents the sigmoid function, , respectively represent average pooling and max pooling operations; Spectral attention map and the input image patch The feature map obtained by weighting : ; This spatial attention module aggregates spectral information by applying average pooling and max pooling operations along the spectral dimension, generating two feature maps and , and and are concatenated in the spectral dimension to generate an effective feature description, and the concatenated feature description is processed using a convolutional layer to generate a spatial attention map : ; In the formula, represents the sigmoid function, represents 7 7 convolutional kernels, represents the concatenation operation of features; Spatial attention map and the feature map are weighted to obtain the fused feature map output by the convolutional attention module CBAM : ; The operations of the other branch are represented by the following formula: ; ; ; ; In the formula, and respectively represent the first and second pointwise convolution PW operations, represents the depth convolution operation, represents the result of the input passing through a series of and operations, represents the fused feature map output by the CBAM module, represents and are added pointwise and average pooled result, represents batch normalization and dropout processing on after flattening, represents the confidence level, Represents a fully connected layer mapping, represents the sigmoid activation function, represents randomly discarding some neurons, represents batch normalization, represents flattening.

[0009] Preferably, the SSMF module includes four branches. The first branch uses a convolutional kernel to extract the spectral features of the image patch to obtain a spectral feature map. The second branch uses a convolutional kernel to extract the spatial information of the image patch to obtain a spatial feature map. The third branch performs average pooling operation, convolutional operation and upsampling operation on the image patch in sequence to obtain a global description feature map. The fourth branch uses the global attention mechanism GAM to obtain the global context information of the image patch to obtain a global feature map. The spectral feature map, spatial feature map, global description feature map and global feature map are feature-connected, and after performing a convolutional operation on the concatenated feature map, the feature map is output. The operation of the SSMF module is represented by the following formula: ; In the formula, represents the connection operation of features, , respectively represent and convolution operations, represents that the features pass through average pooling, convolution operation and upsampling operation in sequence, represents the feature map output by the global attention mechanism.

[0010] Preferably, the global attention mechanism GAM stores the three-dimensional information of the image patch using a three-dimensional arrangement, and uses a multi-layer perceptron MLP with a two-layer encoder-decoder structure to enhance the spectral-spatial correlation, and outputs the spectral attention feature using an inverse arrangement and a sigmoid activation function. The spectral attention feature is weighted with the image patch to obtain the feature map : ; where, represents element-wise multiplication; uses two convolutional operations to fuse the feature map and outputs the spatial attention feature , the spatial attention features and the feature map are weighted to obtain the feature map : ; Among them, represents element-wise multiplication.

[0011] Preferably, the classification module includes a convolutional tokenization, a linear projection layer, an EA 2 T module, and a classification multi-layer perceptron module in sequence; The convolutional tokenization and the linear projection layer map the feature map such that the neighborhood around the central pixel is divided into different spatial tokens: , , and these spatial tokens are concatenated with the classification tokens , and the position information is embedded in the above spatial tokens to generate a token sequence 2 suitable for processing by the EA is represented by the following formula: ; The EA 2 T module first performs layer normalization on the token sequence to generate , where , and two matrices and are used to transform the input feature into and , where , , is the dimension of the feature map, and the query matrix multiplies the learnable parameter vector to learn the query attention weights to obtain the global attention query vector is represented by the following formula: ; The global attention query vector multiplies the query matrix and is pooled to generate the global query vector is represented by the following formula: ; The global query vector and the matrix are multiplied element-wise to construct the global context, introducing linear layer processing to learn the hidden representation of the tokens and obtain the EA2 The matrix output by the T module The output is shown by the following formula: ; Wherein, represents the normalized query matrix, represents a linear mapping; Input the matrix into the classification multi-layer perceptron for classification mapping, and output the classification category probability.

[0012] Preferably, the dynamic threshold strategy includes the following calculation steps: By calculating the number of samples whose classification probability is greater than the category adaptive threshold to evaluate the classification difficulty of the th class : ; Wherein, represents the indicator function, which takes the value of 1 when the condition in the square brackets is satisfied, otherwise it takes the value of 0, is the total number of pseudo-labeled samples, is a preset threshold, and s.t. is the abbreviation of the English phrase "subject to", indicating "constrained by" or "satisfying the following conditions", represents that the category predicted by the model is equal to ; Normalize to the range using the following formula: ; Wherein, represents the normalized , represents maximizing the classification difficulty of the Cth class; Construct weights using the mapping to scale the initial threshold , and the dynamic threshold is obtained using the following formula: ; For , the mapping is a non-linear convex function, and the mapping is performed using the following formula: ; Wherein, is used to determine the value when .

[0013] An image classification system includes a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, the image classification method described in any one of the above is implemented.

[0014] By adopting the foregoing design scheme, the beneficial effects of the present invention are as follows: 1. This application adopts self-training semi-supervised deep learning. The feature extraction network PCET-Net is used to select reliable unlabeled samples to enhance the performance of the model, thereby improving the classification accuracy; 2. This application adopts superpixel-guided label propagation and dynamic threshold strategies to alleviate the problems of unbalanced data distribution and large class differences, so as to better select reliable pseudo-label samples; 3. This application constructs a feature extraction network PCET-Net for feature extraction and classification, introduces an evaluation module to evaluate the quality of the input image patch and output the confidence as a training weight strategy, reduces the interference of heterogeneous pixels at the class boundary, and proposes an SSMF module to further enhance the feature expression ability of the feature extraction network PCET-Net, improving the classification accuracy and robustness. Description of the Drawings

[0015] Figure 1 It is a block diagram of the overall network model architecture of the present invention; Figure 2 It is a comparison diagram of the original image and SLIC superpixel segmentation of the present invention; Figure 3 It is an example diagram of the label propagation process of the present invention; Figure 4 It is a schematic diagram of the model structure of the feature extraction network PCET-Net of the present invention; Figure 5 It is a schematic diagram of the structure of the convolutional attention module CBAM of the present invention; Figure 6 It is a schematic diagram of the structure of the SSMF module of the present invention; Figure 7 It is a schematic diagram of the structure of the global attention mechanism GAM of the present invention; Figure 8 It is the EA 2 T module and a comparison diagram of different attention mechanisms. Detailed Embodiments

[0016] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0017] In the description and claims of the present invention and the above-mentioned drawings, the terms "first", "second", "third", etc. are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0018] An image classification method, as Figure 1 shown, includes the following steps executed in sequence: S1: Obtain labeled samples in the hyperspectral image, construct a labeled data set, the labeled data set includes a training set and a test set, perform superpixel segmentation on the hyperspectral image, propagate the labels of the labeled samples in the training set to the unlabeled samples, and select unlabeled samples with balanced data distribution therefrom as the unlabeled data set; As Figure 3 shown, the label propagation in step S1 specifically includes the following steps executed in sequence: S1-1: Use Simple Linear Iterative Clustering (SLIC) to segment the hyperspectral image (HSI) into a set of superpixels, these superpixels are composed of non-overlapping regions with different shapes and sizes, as Figure 2 shown, the number of superpixel segments is obtained by the following formula: ; where represents the influence factor, and its value affects the number of pixels in each superpixel, usually set to 50, and represent the length and width of the hyperspectral image respectively; S1-2: According to whether the superpixel contains labeled samples, the labeled samples in this step refer to the labeled samples in the training set, divide these superpixels into two cases. One case is that the segmentation region consists of superpixels including labeled samples and unlabeled samples, as Figure 3 in and , the pixels in the superpixel are highly similar and may belong to the same category. Therefore, calculate the number of known labeled samples, and propagate the label with the most corresponding labels to other unlabeled samples in the same superpixel.

[0019] Another type is that the segmented region only contains superpixels of unlabeled samples, such as Figure 3 in , and . The nearest neighbor algorithm is used to assign labels to the unlabeled samples of this superpixel. The distance between each superpixel and the labeled samples of each category is calculated through their average spectra. The distance between each superpixel and the labeled samples of each category is calculated by the following formula: ; where represents the average spectrum of the -th superpixel that only contains unlabeled samples, represents the modulus of the average spectrum, represents the average spectrum of the labeled samples of the -th category; the class label corresponding to the maximum distance is assigned to all pixels in the -th superpixel that only contains unlabeled samples, represents the cosine function of the included angle; the larger the distance , the higher the similarity between and , and vice versa. Therefore, the label c corresponding to the maximum distance is assigned to the superpixel , and this label is propagated to all pixels in this superpixel region.

[0020] S1-3: According to the calculation in S1-2, all unlabeled superpixels obtain propagated labels through label propagation, and unlabeled samples are uniformly selected from each category as the unlabeled dataset.

[0021] Since random selection may lead to unbalanced data distribution, reasonable selection of unlabeled samples is crucial for semi-supervised learning. To mitigate the negative impact of the unbalanced distribution, label propagation is adopted. Under the guidance of the superpixel segmentation technology based on spectral-spatial similarity, the limited known labeled samples are propagated to the entire hyperspectral image to achieve balanced sampling of samples in each category.

[0022] S2: Construct a feature extraction network PCET-Net, as shown in Figure 4 . This feature extraction network PCET-Net includes an evaluation module for evaluating the quality of image patches and outputting the confidence of image patches, an SSMF module for fusing spectral information and spatial information, and a classification module.

[0023] The training set is input into the feature extraction network PCET-Net for training to output the classification category probability. The unlabeled dataset is subjected to weak augmentation to obtain and strong augmentation is performed on the unlabeled dataset to obtain where weak augmentation is to randomly vertically and horizontally flip the samples, and strong augmentation adds random noise on the basis of weak augmentation. and are respectively input into the feature extraction network PCET-Net for training. The output category probability is compared with the dynamic threshold adjusted by the dynamic threshold strategy. If the output category probability is greater than the dynamic threshold, the category obtains the corresponding pseudo-label; otherwise, the category does not obtain a pseudo-label.

[0024] In this embodiment, the evaluation module in step S2 includes two branches. One branch fuses the spectral information and spatial information of the image patch through the convolutional block attention module CBAM to output a fused feature map. The other branch sequentially performs feature extraction on the image patch through the first pointwise convolution PW, depthwise convolution DW, and the second pointwise convolution PW to obtain a preliminary feature map. The preliminary feature map and the fused feature map are added pointwise and then average pooled. The feature map after average pooling is flattened and batch-normalized in sequence, and some neurons in the feature map are randomly discarded. The features are mapped using a fully connected layer mapping, and the mapping result is processed by the sigmoid function to obtain the confidence , and the specific operation process is as follows: Assume the input image patch , represents the size of the image patch, represents the number of spectral dimensions, as shown in Figure 5 . The convolutional block attention module CBAM includes a spectral attention module and a spatial attention module. The spectral attention module uses the inter-band relationship between features to generate a spectral attention map, and adopts average pooling AvgPool and max pooling MaxPool operations to aggregate the spatial information, thereby generating two context information feature maps in different spaces and . and are input into a shared multi-layer perceptron MLP, which has a hidden layer with a hidden activation size of , represents the compression ratio of the MLP, and these two features are fused by element-wise summation to generate a spectral attention map : ; where represents the sigmoid function, , respectively represent average pooling and max pooling operations.

[0025] Spectral attention map and the input image patch weighted feature map : ; This spatial attention module aggregates spectral information by applying average pooling and max pooling operations along the spectral dimension, generating two feature maps and , and and are concatenated in the spectral dimension to generate an effective feature description, and the concatenated feature description is processed using a convolutional layer to generate a spatial attention map : ; In the formula, represents the sigmoid function, represents 7 × 7 convolutional kernel, represents the concatenation operation of features; Spatial attention map and the feature map are weighted to obtain the fused feature map output by the convolutional attention module CBAM : ; By generating corresponding attention maps in the spectral and spatial dimensions, CBAM can dynamically adjust the weights of each part in the feature map. Subsequently, these attention maps are fused with the input feature map to make the model pay more attention to the important local information in the input feature map.

[0026] The operations of the other branch are represented by the following formula: ; ; ; ; In the formula, and respectively represent the first and second pointwise convolution PW operations, represents the depth convolution operation, represents the result of the input passing through a series of and operations, represents the fused feature map output by the CBAM module, represents the and perform point - by - point addition and average pooling The result of means that after flattening perform batch normalization and dropout processing represents the confidence represents the fully - connected layer mapping represents the sigmoid activation function represents randomly discarding some neurons represents batch normalization represents flattening

[0027] This evaluation module is used to evaluate the input image patch and output its confidence Confidence. A higher confidence represents a greater positive contribution of the image patch to training, and this confidence is used in the loss function for calculation, enabling the model to dynamically adjust the learning focus according to the quality of the input image patch

[0028] Depthwise convolution DW performs independent convolution on each spectral dimension using corresponding convolutional kernels, greatly reducing the number of parameters compared to standard convolution; Pointwise convolution PW then uses convolutional kernels to merge and remap the feature information of different channels, thus realizing the fusion of feature information of different spectral dimensions

[0029] In this embodiment, in the feature extraction stage, to make full use of the rich spectral and spatial information contained in the hyperspectral image and capture complementary and representative features from the image patch, an SSMF module is designed to fuse the spectral and spatial information, which can enhance the model's ability to distinguish different classes, thereby enabling the model to better capture the comprehensive features of substances

[0030] As Figure 6 shown, the SSMF module includes four branches, which perform different processes on the input image patch to generate 4 independent feature maps, aiming to extract feature information from different perspectives. The first branch uses convolutional kernels to extract the spectral features of the image patch to obtain the spectral feature map. The second branch uses convolutional kernels to extract the spatial information of the image patch to obtain the spatial feature map. The third branch sequentially performs average pooling operation, convolution operation, and upsampling operation on the image patch to obtain the global descriptive feature map. The fourth branch uses the global attention mechanism GAM to obtain the image patch The global context information is used to obtain the global feature map. The spectral feature map, spatial feature map, global description feature map, and global feature map are feature concatenated. After performing a convolution operation on the concatenated feature map, a feature map is output , and the operation of the SSMF module is represented by the following formula: ; In the formula, represents the feature concatenation operation, , respectively represent and convolution operations, represents that the feature sequentially passes through average pooling, convolution operation, and upsampling operation, represents the feature map output by the global attention mechanism.

[0031] The SSMF module integrates spectral and spatial information, and at the same time considers local details and global background, so as to be able to extract multi-scale feature representations of hyperspectral images and form a feature set with high resolution.

[0032] The global attention mechanism GAM more comprehensively considers the cross-dimensional information of features and enhances the interaction between global dimensions, as Figure 7 shown. The global attention mechanism GAM stores the three-dimensional information of the image patch in a three-dimensional arrangement, and uses a multi-layer perceptron MLP with a two-layer encoder-decoder structure to enhance the spectral-spatial correlation, and its hidden activation size is , represents its compression ratio, and uses inverse arrangement and sigmoid activation function to output spectral attention feature , and weights the spectral attention feature with the image patch to obtain the feature map : ; Among them, represents element-wise multiplication; uses two-layer convolution operations to fuse the feature map and outputs the spatial attention feature , and weights the spatial attention feature with the feature map to obtain the feature map : ; Among them, represents element-wise multiplication.

[0033] In this embodiment, the classification module includes a convolutional tokenization, a linear projection layer, an EA 2 T module, and a classification multi-layer perceptron module connected in sequence; The convolutional tokenization and the linear projection layer map the feature map such that the neighborhood around the central pixel is divided into different spatial tokens: , , and these spatial tokens are concatenated with a classification token for performing the classification task. To provide location information and embed the location information into the above spatial tokens, a token sequence 2 suitable for processing by the EA T module is generated and represented by the following formula: ; As Figure 8 shown, the EA 2 T module first performs layer normalization on the token sequence to generate , where , and uses as the input feature of the EA 2 T module. Two matrices and are used to transform the input feature into and , where , , is the dimension of the feature map. The query matrix multiplies the learnable parameter vector to learn the query attention weights, resulting in a global attention query vector represented by the following formula: ; The global attention query vector multiplies the query matrix and is pooled to produce a global query vector represented by the following formula: ; The global query vector is element-wise multiplied by the matrix to construct the global context, introducing the interaction between to learn the hidden representation of the tokens, resulting in the matrix 2 output by the EA whose output is shown by the following formula: ; Among them, represents the normalized query matrix, represents the linear mapping; The matrix is input into a classification multi-layer perceptron for classification mapping, and the classification category probability is output.

[0034] The dynamic threshold strategy includes the following calculation steps: By calculating the number of samples whose classification probability is greater than the category adaptive threshold to evaluate the classification difficulty of the th class : ; Among them, represents the indicator function, which takes the value of 1 when the condition in the square brackets is satisfied, and 0 otherwise. is the total number of unlabeled samples, is the preset threshold, and s.t. is the abbreviation of the English phrase "subject to", which means "constrained by" or "satisfies the following conditions". represents that the category predicted by the model is equal to ; The following formula is used to normalize to the range: ; Among them, represents the normalized , represents maximizing the classification difficulty of the Cth class; it can be seen that the class with fewer samples corresponds to a lower . To ensure that each class can obtain sufficient learnable labeled samples, for the class with fewer samples, its sample quantity should be increased; for the class with more samples, information redundancy should be reduced. This method can balance the sample quantity of each class and improve the discrimination and robustness of the model.

[0035] Use the mapping to construct weights to scale the initial threshold , and the dynamic threshold is obtained by the following formula: ; It can be seen that the threshold is positively correlated with the weight . Since the "easy" class with a large number of samples has a high threshold, and the "difficult" class with a small number of samples has a lower threshold, so when , the mapping should be a monotonically increasing function. In addition, for the "difficult" class, the weight decreases rapidly as decreases, so as to allocate more pseudo-labels for learning. For "easy" classes, the weight is close to the maximum value of 1, meaning that the weight increases slowly as increases. Therefore, for , the mapping should be a non-linear convex function.

[0036] Based on the above considerations, for , the mapping is a non-linear convex function, and the mapping is performed using the following formula: ; where is used to determine the value at , and the values of its parameters can be selected through ablation experiments.

[0037] S3: Calculate the loss functions of the training set and the unlabeled dataset respectively according to the classification category probabilities output in step S2, and use the total loss function to reversely adjust the feature extraction network PCET-Net. The calculation method of the total loss function is as follows: S3-1: Calculate the loss function of the training set using the following formula : ; where is the confidence level output by the evaluation module, represents the predicted value of the target category, represents the predicted value for the category , represents the exponential function; S3-2: Calculate the loss function of the unlabeled dataset using the following formula : ; where represents the pseudo-label of the unlabeled sample, represents the corresponding dynamic threshold, is the weak augmentation of the th unlabeled sample, is the strong augmentation of the th unlabeled sample, is the total number of unlabeled samples, s.t. is the abbreviation of the English phrase "subject to", indicating "subject to" or "satisfying the following conditions", and max represents the maximization operation; S3-3: The overall loss function consists of two parts: the supervised cross-entropy loss on the labeled sample training set and the unsupervised cross-entropy loss on the unlabeled samples , the total loss function is calculated using the following formula: ; S4: Input the test set into the adjusted feature extraction network PCET-Net for image classification and output the classification result.

[0038] In this embodiment, a system for implementing the above method is also provided.

[0039] An image classification system includes a memory and a processor. A computer program is stored on the memory. When the computer program is executed by the processor, the image classification method described in any one of the above is implemented.

[0040] In summary, the present application only requires a small number of labeled samples, alleviates the data distribution imbalance problem through superpixel-guided label propagation, constructs a high-precision and high-robustness feature extraction network PCET-Net based on consistency regularization, uses dynamic thresholds to select highly reliable unlabeled samples to enhance model training in a supervised manner, and improves the classification accuracy.

[0041] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. An image classification method, characterized in that: The process includes the following steps: S1: Obtain labeled samples in a hyperspectral image, construct a labeled data set, the labeled data set includes a training set and a test set, perform superpixel segmentation on the hyperspectral image, propagate the labels of the labeled samples in the training set to unlabeled samples, and select unlabeled samples with balanced data distribution as the unlabeled data set; S2: Construct a feature extraction network PCET-Net, which includes an evaluation module for evaluating the quality of image blocks and outputting the confidence of image blocks, an SSMF module for fusing spectral information and spatial information, and a classification module. Input the training set into the feature extraction network PCET-Net for training, output the classification category probability, perform weak enhancement processing on the unlabeled data set, and obtain , perform strong enhancement processing on the unlabeled data set to obtain ,Will and The output categories are respectively input into the feature extraction network PCET-Net for training, and the output category probability is compared with the dynamic threshold adjusted by the dynamic threshold strategy. If the output category probability is greater than the dynamic threshold, the category obtains the corresponding pseudo label, otherwise, the category does not obtain the pseudo label; S3: Calculate the loss functions of the training set and the unlabeled data set respectively according to the classification category probability output in step S2, and use the total loss function to reversely adjust the feature extraction network PCET-Net. The total loss function is calculated as follows: S3-1: Use the following formula to calculate the loss function of the training set : ; in, is the confidence of the evaluation module output, represents the predicted value of the target category, Indicates the category The predicted value of represents the exponential function; S3-2: Use the following formula to calculate the loss function of the unlabeled data set : ; in, represents the pseudo label of the unlabeled sample, represents the corresponding dynamic threshold, For the Weak enhancement of unlabeled samples The predicted probability of For the Strong enhancement of unlabeled samples The predicted probability of is the total number of unlabeled samples, st is the abbreviation of the English phrase subject to, which means "subject to" or "satisfies the following conditions", and max means maximization operation; S3-3: Total loss function The calculation is done using the following formula: ; S4: Input the test set into the adjusted feature extraction network PCET-Net for image classification, and output the classification results.

2. The image classification method according to claim 1, characterized in that: The label propagation in step S1 specifically includes the following steps performed in sequence: S1-1: Use simple linear iterative clustering to segment the hyperspectral image into a set of superpixels. The number of superpixel segmentations Obtained by the following formula: ; in, Represents the influencing factor, whose value affects the number of pixels in each superpixel. and represent the length and width of the hyperspectral image respectively; S1-2: Depending on whether the superpixel contains labeled samples, these superpixels are divided into two cases. One of them is that the segmented area consists of superpixels including labeled samples and unlabeled samples. The number of labeled samples is calculated, and the label with the most corresponding labels is propagated to other unlabeled samples in the same superpixel; the other is that the segmented area only contains superpixels with unlabeled samples. The nearest neighbor algorithm is used to assign labels to the unlabeled samples of the superpixel. The distance between each superpixel and the labeled samples of each category is The distance between each superpixel and the labeled sample of each class is calculated by their average spectrum Calculated by the following formula: ; in, Indicates The average spectrum of superpixels containing only unlabeled samples, represents the modulus of the average spectrum, Indicates The average spectrum of the labeled samples of the categories; corresponding to the maximum distance The class label is assigned to All pixels in a superpixel that contains only unlabeled samples; S1-3: According to the calculation of S1-2, all unlabeled superpixels obtain propagated labels through label propagation, and unlabeled samples are uniformly selected from each category as the unlabeled dataset.

3. The image classification method according to claim 1, characterized in that: The evaluation module in step S2 includes two branches, one of which fuses the spectral information and spatial information of the image block through the convolutional attention module CBAM and outputs a fused feature map, and the other branch extracts features of the image block through the first point-by-point convolution PW, the deep convolution DW and the second point-by-point convolution PW in turn to obtain a preliminary feature map, and adds the preliminary feature map to the fused feature map point by point and performs average pooling, flattens and batch normalizes the feature map after average pooling in turn, and randomly discards some neurons in the feature map, maps the features using the fully connected layer mapping, and the mapping result is processed by the sigmoid function to obtain the confidence The specific operation process is as follows: Assume that the input image block , represents the image block size, Represents the number of spectral dimensions. The convolutional attention module CBAM includes a spectral attention module and a spatial attention module. The spectral attention module uses average pooling AvgPool and maximum pooling MaxPool operations to aggregate spatial information to generate contextual information feature maps of two different spaces. and ,Will and The two features are input into a shared multi-layer perceptron MLP and fused by element-wise summation to generate a spectral attention map. : ; in, represents the sigmoid function, , Represent average pooling and maximum pooling operations respectively; Spectral Attention Map With the input image patch Weighted feature map : ; The spatial attention module applies average pooling and maximum pooling operations along the spectral dimension to aggregate the spectral information and generate two feature maps and , and and Splicing is performed on the spectral dimension to generate effective feature descriptions, and the convolutional layer is used to process the spliced ​​feature descriptions to generate a spatial attention map : ; In the formula, represents the sigmoid function, Indicates 7 7 convolution kernels, Represents the concatenation operation of features; Spatial Attention Map With feature map Weighted fusion feature map of the convolutional attention module CBAM output : ; The operation of the other branch is expressed as follows: ; ; ; ; In the formula, and denote the first and second point-wise convolution PW operations, respectively, Represents a deep convolution operate, Indicates that the input passes through a series of and The result of the operation, Represents the fused feature map output by the CBAM module, Express and Add point by point and perform average pooling As a result, Indicates the flattened Perform batch normalization and random dropout. Indicates confidence, represents the fully connected layer mapping, represents the sigmoid activation function, Indicates randomly discarding some neurons. represents batch normalization, Indicates flattening.

4. The image classification method according to claim 3, characterized in that: The SSMF module includes four branches. The first branch is used Convolution kernel extracts image patches Spectral features, obtain spectral feature map, the second branch uses Convolution kernel extracts image patches The spatial information of the image is obtained by the third branch. The average pooling operation, convolution operation and upsampling operation are performed in sequence to obtain the global description feature map. The fourth branch uses the global attention mechanism GAM to obtain the image block. The global context information is obtained, the global feature map is obtained, the spectral feature map, the spatial feature map, the global description feature map and the global feature map are feature-connected, and the feature map is output after the concatenated feature map is convolved. , the operation of the SSMF module is represented by the following formula: ; In the formula, represents the connection operation of the feature, , Respectively and The convolution operation, Indicates that the features are average pooled in sequence, Convolution and upsampling operations, Feature map representing the output of the global attention mechanism.

5. The image classification method according to claim 4, characterized in that: The global attention mechanism GAM stores image patches in a three-dimensional arrangement The three-dimensional information is obtained by using a two-layer encoder-decoder structured multi-layer perceptron MLP to enhance the spectral spatial correlation, and the spectral attention features are output using inverse permutation and sigmoid activation function. , the spectral attention feature With image blocks Weighted feature map : ; in, It means element-wise multiplication; Use two layers of convolution operation to transform the feature map Fusion and output of spatial attention features , the spatial attention feature With feature map Weighted feature map : ; in, Represents element-wise multiplication.

6. The image classification method according to claim 5, characterized in that: The classification module consists of sequentially connected convolutional tokenization, linear projection layer, EA 2 T module and classification multilayer perceptron module; The convolutional tokenization and the linear projection layer perform feature map Mapping is performed, and the neighborhood around the center pixel is divided into different spatial labels: , , these spatial labels and classification labels Splice and transfer the location information Embedded in the above space mark, generate suitable EA 2 The token sequence processed by the T module It is expressed by the following formula: ; EA 2 The T module first performs a token sequence Perform layer normalization to generate ,in , using two matrices and Input features Transformed to and ,in , , is the dimension of the feature map, the query matrix Multiply by the learnable parameter vector To learn the query attention weights, we get the global attention query vector It is expressed by the following formula: ; Global Attention Query Vector Multiply by the query matrix and pooled to produce a global query vector It is expressed by the following formula: ; The global query vector and matrix Multiply element by element to build a global context and introduce linear layer processing The interaction between them is used to learn the hidden representation of the token and obtain the EA 2 The matrix output by the T module The output is given by the following formula: ; in, represents the normalized query matrix, Represents a linear mapping; The matrix The classification multilayer perceptron is input for classification mapping, and the classification category probability is output.

7. The image classification method according to claim 6, characterized in that: The dynamic threshold strategy includes the following calculation steps: By calculating the classification probability greater than the class adaptation threshold The number of samples to evaluate the Classification difficulty : ; in, Represents an indicator function. When the condition in the square brackets is met, the value is 1, otherwise the value is 0. is the total number of unlabeled samples, is the preset threshold, and st is the abbreviation of the English phrase subject to, which means "subject to" or "satisfy the following conditions". Indicates that the category predicted by the model is equal to ; Use the following formula to Normalized to scope: ; in, Represents the normalized , Indicates maximizing the classification difficulty of the Cth category; Utilizing Mapping Construct weights to scale the initial threshold , dynamic threshold Use the following formula to obtain: ; for , mapping It is a nonlinear convex function and is mapped using the following formula: ; in, To determine The value at time.

8. An image classification system, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that: When the computer program is executed by a processor, the image classification method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Hyperspectral image classification method and device, equipment and storage medium

    CN113361481A

  • Hyperspectral image semi-supervised classification method and device based on improved active deep learning

    CN113723492A

  • Hyperspectral image classification method and device based on spatial-spectral double-branch convolutional network

    CN115249332A

  • Medical image processing method based on semi-supervised neural network

    CN116630299A

  • Multi-enhancement semi-supervised image classification method based on contrast loss

    CN117333709A

Cited By

  • Small sample image classification method based on parallel expert structure

    CN122156833A

  • Small sample image classification method based on parallel expert structure

    CN122156833B