Hyperspectral Image Classification Method Based on Spatial Pooling Transformer
By adopting spatially pooled Transformer network and noise-resistant learning algorithms in hyperspectral image classification, the problem of lack of global information processing and noise label processing capabilities in the prior art is solved, and more efficient and accurate image classification is achieved.
Patent Information
- Application Number
- CN202311221345.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-20
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-09-20
AI Technical Summary
The existing hyperspectral image classification methods lack global information processing in feature extraction and lack the processing capability of noise labels, resulting in limited classification performance.
The network structure based on space pooling Transformer is adopted, and the global spatial information is extracted through the self-attention mechanism, and the space pooling module is designed to remove redundant features. Combined with a noise-resistant learning algorithm, it distinguishes clean samples from noise samples, and uses a robust similar loss function to enhance the model performance.
More sufficient and accurate feature extraction is achieved, improving the accuracy of hyperspectral image classification, especially in the presence of noise labels, which significantly improves the robustness and performance of the model.
Smart Images

Figure CN117274691B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to a method for classifying remote sensing images based on spatial pooling Transformer in the field of image classification technology. The present invention can be used to classify the categories of pixels in hyperspectral images with noisy labels obtained in agricultural management, natural disaster assessment, and ecological monitoring. Background Art
[0002] Hyperspectral images have hundreds of narrow and continuous spectral bands, which contain rich spatial spectral information and provide favorable conditions for effectively analyzing ground objects. Among various hyperspectral image processing technologies, hyperspectral image classification is a basic technology, which aims to provide semantic labels for each pixel in a hyperspectral image. Reliable classification results can be applied to many remote sensing applications, such as agricultural management, ecological observation, and natural disaster assessment. In recent years, hyperspectral image classification methods have gradually shifted from traditional handcrafted feature extraction based on machine learning to deep feature extraction based on convolutional neural networks, and the classification performance of the models has been gradually improved. Although the above methods have achieved great success in hyperspectral image classification, they still have the following problems. On the one hand, most of the existing hyperspectral image classification methods rely on an assumption that the given handcrafted labels are correct. However, this assumption is not always correct in practice. Due to sensor errors and the complexity of hyperspectral images, incorrect annotations are inevitable. On the other hand, most of the existing hyperspectral image classification methods use convolutional neural networks as the feature extraction backbone. However, due to the limitation of the receptive field, the global information hidden in hyperspectral images is always ignored, which also limits the classification performance of the model to a certain extent.
[0003] Zhang Hongyan et al. proposed a hyperspectral image classification method based on sample selection and label correction in their published paper "A Superpixel Guided Sample Selection Neural Network for Handling Noisy Labels in Hyperspectral Image Classification" (IEEE Transactions on Geoscience and Remote Sensing, 2021). The implementation steps of this method are as follows: First, the samples are input into a convolutional neural network, and the samples with small loss values are regarded as clean samples. By selecting a fixed proportion of small-loss samples for training, the influence of noisy labels is alleviated. Then, the labels of the selected noisy data are corrected using superpixels, and samples are re-selected for training, reducing information loss. Although this method can effectively reduce the negative impact of noise, there are still two deficiencies. First, this method uses a convolutional neural network as the feature extraction backbone and does not consider the global pixel correlation information, resulting in insufficient and inaccurate feature extraction. Second, since sample selection cannot completely separate noisy samples, the subsequent label correction cannot guarantee the credibility of the new labels, limiting the improvement of classification performance.
[0004] Xidian University proposed a hyperspectral image classification method in its patent document "Hyperspectral Image Classification Method Based on Contrastive Generative Adversarial Network" (Patent Application No.: CN202110775163.7, Authorization Publication No.: CN113469084B). This method constructs a contrastive generative adversarial network, uses the generative network to generate false samples, uses the discriminative network to classify the training samples and false samples, obtains the contrastive loss by comparing the class features between samples, and constructs the loss functions of the generative network and the discriminative network, enhancing the network's feature extraction ability. However, the deficiency that still exists in this method is that due to an inaccurate positioning system and errors during ground surveys, the occurrence of mislabeled data in hyperspectral images is inevitable. Since this method does not set corresponding strategies for noisy labels, when the proportion of noisy labels gradually increases, the performance of the model will rapidly decline. Summary of the Invention
[0005] The purpose of the present invention is to address the above deficiencies of the prior art and propose a hyperspectral image classification method based on spatial pooling Transformer to solve the problems of insufficient and inaccurate network feature extraction and the negative impact when noisy labels appear in the prior art.
[0006] The technical idea for achieving the purpose of the present invention is: the present invention constructs a spatial pooling Transformer network, extracts global spatial information through a self-attention mechanism, and at the same time, based on the spatial characteristics of hyperspectral images, designs a spatial pooling module to perform redundancy operations on the obtained global spatial features, obtain more accurate feature representations, and thus perform more sufficient feature extraction, thereby solving the problem of insufficient and inaccurate network feature extraction in the prior art. The present invention designs a noise-resistant learning algorithm, distinguishes clean samples from noise samples by the loss value of the samples, and calculates the classification loss for the clean samples. In addition, a robust similarity loss function is designed, which is based on the spatial prior knowledge of hyperspectral images, closes the pixel representation of homogeneous areas, and uses the effective information in the noise samples, so that the model can obtain better performance even when containing noise labels, avoiding the problem of negative impact when noise labels appear.
[0007] According to the above technical ideas, the technical solutions adopted by the present invention include the following:
[0008] Step 1: Generate training set:
[0009] Obtain a hyperspectral image containing at least C pixel categories, each pixel category contains at least S pixels, and the category contains G pixels with noise labels; after normalizing the hyperspectral image, take each pixel containing the target as the center, and form a pixel block of the target with 9×9 pixels around it, and use the label of each target pixel as the label of the pixel block. All pixel blocks in the hyperspectral image and their corresponding pixel block labels form a training set, where C≥2, S≥1, G<S;
[0010] Step 2, respectively build a first spatial pooling Transformer network and a second spatial pooling Transformer network, the structures of the first and second spatial pooling Transformer networks are exactly the same; each spatial pooling Transformer network is composed of a spectral feature extraction module, a first spatial pooling Transformer module, a second spatial pooling Transformer module, a third spatial pooling Transformer module, a Transformer encoder, and a classifier connected in series in sequence;
[0011] Step 3: adopt a preheating strategy and use cross entropy loss to train the first and second spatial pooling Transformer networks respectively to obtain the preheated first and second spatial pooling Transformer networks;
[0012] Step 4: Use a noise-tolerant learning algorithm to train the pre-warmed first and second spatial pooling Transformer networks:
[0013] Step 4.1: Adopt a cross - selection strategy to calculate the cross - entropy losses of the first and second spatial pooling Transformer networks on the clean sample set respectively.
[0014] Step 4.2: Calculate the similarity losses of the first and second spatial pooling Transformer networks on the training set.
[0015] Step 4.3: Calculate the combined losses of the first and second spatial pooling Transformer networks respectively, and iteratively update the parameters of the two networks until the combined loss function converges, obtaining the trained first and second spatial pooling Transformer networks.
[0016] Step 5: Classify the hyperspectral image:
[0017] Adopt the same method as in Step 1 to extract the image patches to be classified from the normalized hyperspectral image. Input all the hyperspectral image pixel patches to be classified into the trained first and second spatial pooling Transformer networks. Take the average of the class probability vectors output by the two networks to obtain the class probability vector of all samples, and take the class corresponding to the maximum probability in the class probability vector as the classification result of the pixel.
[0018] Compared with the prior art, the present invention has the following advantages:
[0019] First, the spatial pooling Transformer network designed by the present invention can utilize the self - attention mechanism to extract the global inter - pixel relationship, fully considering the global information. At the same time, the spatial pooling module utilizes the spatial characteristics of the hyperspectral image to further remove the redundant information in the global spatial features, overcoming the problems in the prior art that the receptive field of the model is too small to capture the global spatial information, and the network fails to extract the feature representation sufficiently and accurately. The present invention enables full utilization of the spatial knowledge of the hyperspectral image and obtains a more accurate feature representation.
[0020] Second, the noise - tolerant learning algorithm designed by the present invention selects clean samples from the training set through a sample - selection strategy to calculate the classification loss. At the same time, based on the spatial prior knowledge of the hyperspectral image, a similarity loss is designed to utilize the effective information of all samples, overcoming the phenomenon that the performance drops sharply when there are noise labels in the hyperspectral image classification task, and the defect of insufficient credibility of label correction in the existing technologies for noise labels. The present invention alleviates the phenomenon that the network performance drops sharply when noise labels appear and improves the accuracy of hyperspectral image classification. Description of the Drawings
[0021] Figure 1 is the flow chart of the present invention;
[0022] Figure 2 It is a schematic diagram of the spatial pooling Transformer network structure constructed by the present invention;
[0023] Figure 3 It is a schematic diagram of the Transformer encoder structure constructed by the present invention;
[0024] Figure 4 It is a simulation result diagram of classifying the Indian Pines hyperspectral image dataset with noisy labels by the present invention and the prior art. Detailed implementation manners
[0025] In order to more clearly illustrate the specific implementation manners of the present invention or the technical solutions in the prior art, the following further describes the present invention in detail with reference to the drawings and embodiments.
[0026] Refer to Figure 1 to further describe in detail the implementation steps of the embodiments of the present invention.
[0027] Step 1, generate a training set.
[0028] Obtain a hyperspectral image that contains at least C pixel categories, each pixel category contains at least S pixel points, and there are G pixel points with noisy labels in this category.
[0029] The hyperspectral images used in the embodiments of the present invention are from the Indian Pines hyperspectral dataset. The hyperspectral images in this dataset are composed of 145×145 pixels containing 200 spectral bands. There are 10,249 labeled pixels in the image, which are divided into 16 pixel categories. 1,770 pixels are randomly selected from them to form a training set. The labels of 40% of the pixels in each category are randomly replaced with other categories to obtain noisy labels. After normalizing the hyperspectral image, with each pixel point containing a target as the center, the 9×9 pixels around it form the pixel block of this target, and the label of each target pixel point is used as the label of its pixel block. All the pixel blocks in the hyperspectral image and their corresponding pixel block labels form the training set.
[0030] Step 2: Build the first spatial pooling Transformer network and the second spatial pooling Transformer network respectively. The structures of the first and second spatial pooling Transformer networks are exactly the same. Each spatial pooling Transformer network is composed of a spectral feature extraction module, a first spatial pooling Transformer module, a second spatial pooling Transformer module, a third spatial pooling Transformer module, a Transformer encoder, and a classifier connected in series in sequence. The structure of the spatial pooling Transformer network is as Figure 2 shown.
[0031] The described spectral feature extraction module is composed of a 3D convolutional unit and a 2D convolutional unit connected in series, which extracts spectral information in the hyperspectral image from both global and local perspectives. The structures of the 3D convolutional unit and the 2D convolutional unit are the same, and both are composed of a convolutional layer, a batch normalization layer, and a non-linear activation layer connected in series. Set the convolutional kernel sizes of the convolutional layers in the 3D convolutional unit and the 2D convolutional unit to 5×1×1 and 1×1 respectively, and the numbers of convolutional kernels to 8 and 64 respectively. The non-linear activation layer is implemented using the LeakyReLU function.
[0032] Refer to Figure 3 , and the structure of the Transformer encoder is further described. It is used to further enhance the feature representation of the hyperspectral image, and it includes a position encoder, a first layer normalization, a multi-head self-attention unit, a second layer normalization, and a multi-layer perceptron. Among them, the feature matrix output by the position encoder is added to the feature matrix output after passing through the first layer normalization and the multi-head self-attention unit, and then added to the feature matrix output after passing through the second layer normalization and the multi-layer perceptron. Set the number of attention heads in the multi-head self-attention unit to 8.
[0033] The structures of the first spatial pooling Transformer module, the second spatial pooling Transformer module, and the third spatial pooling Transformer module are the same. They are all used to extract global pixel correlations and reduce redundant information in the feature map. Each spatial pooling Transformer module is composed of a Transformer encoder and a spatial pooling unit connected in series; the Transformer encoder is used to further enhance the feature representation of the hyperspectral image, which includes a position encoder, a first layer of normalization, a multi-head self-attention unit, a second layer of normalization, and a multi-layer perceptron. Among them, the feature matrix output by the position encoder is added to the feature matrix output after the first layer of normalization and the multi-head self-attention unit, and then added to the feature matrix output after the second layer of normalization and the multi-layer perceptron; the number of attention heads in the multi-head self-attention unit is set to 8; the spatial pooling unit is composed of a restoration layer, a spatial pooling layer, and a flattening layer connected in series. Among them, the restoration layer is used to transform the serialized features into two-dimensional features and then transmit them to the spatial pooling layer. The size of the two-dimensional features is equal to the square root of the length of the serialized features; the spatial pooling layer removes the outermost layer of features of the two-dimensional features obtained by the restoration layer and then transmits the refined feature map to the flattening layer; the flattening layer flattens the two-dimensional features back into one-dimensional features.
[0034] The classifier described above is composed of a global average pooling layer and a fully connected layer connected in series; the number of input nodes of the fully connected layer is set to 64, and the number of output nodes is set to be equal to the number of classes of the hyperspectral image. Since the number of classes of the images in the embodiments of the present invention is 16, the number of output nodes of the fully connected layer is set to 16 in the embodiments of the present invention.
[0035] Step 3: Adopt a warm-up strategy and use cross-entropy loss to train the first and second spatial pooling Transformer networks respectively to obtain the preheated first and second spatial pooling Transformer networks.
[0036] The warm-up strategy mentioned above means that the training set is input into the first and second spatial pooling Transformer networks respectively, the cross-entropy loss is calculated respectively, and the parameters of the two networks are iteratively updated by the gradient descent method until the cross-entropy loss function converges, so as to obtain the preheated first and second spatial pooling Transformer networks.
[0037] The cross-entropy loss is obtained by the following formula:
[0038]
[0039] Among them, represents the cross-entropy loss value, N represents the total number of pixel blocks in the training set, C represents the total number of pixel categories in the training set, ylk represents the sign function, where y is 1 when the class of the l-th sample is k, and 0 otherwise. p lk represents the probability that the i-th sample in the training set is predicted to be class k. lk
[0040] Step 4: Use a noise-resistant learning algorithm to train the preheated first and second spatial pooling Transformer networks.
[0041] Step 4.1: Adopt a cross-selection strategy to obtain the cross-entropy losses of the first and second spatial pooling Transformer networks on the clean sample set respectively. The calculation steps are as follows:
[0042] First step: Input the training set into the preliminarily trained first spatial pooling Transformer network. Use a two-component Gaussian mixture model to fit the loss values of all samples. Take the probability that each sample belongs to the Gaussian component with the minimum mean as the clean probability of the sample. All samples in the training set with a clean probability greater than or equal to the threshold τ form the clean sample set of the first spatial pooling Transformer network The rest form the noise sample set where τ is a hyperparameter, and its value range is: 0 ≤ τ < 1. In this example, the threshold τ is set to 0.3.
[0043] Second step: Use the same method as the first step to obtain the clean sample set and the noise sample set of the second spatial pooling Transformer network
[0044] Third step: Input the clean sample set selected by the second spatial pooling Transformer network into the first spatial pooling Transformer network to calculate the cross-entropy loss Input the clean sample set selected by the first spatial pooling Transformer network into the second spatial pooling Transformer network to calculate the cross-entropy loss By swapping the clean sample sets selected by the two networks, the cumulative error caused by incorrect selection can be effectively reduced.
[0045] The cross-entropy loss mentioned above is obtained by the following formula:
[0046]
[0047] where represents the cross-entropy loss value, N represents the total number of pixel blocks in the selected clean sample set, C represents the total number of pixel classes in the training set, and ylk represents the sign function, where y is 1 when the class of the l-th sample is k, and 0 otherwise, and p lk represents the probability that the i-th sample in the training set is predicted to be class k. lk
[0048] Step 4.2, calculate the similarity loss of the first and second spatial pooling Transformer networks on the training set according to the following method:
[0049] First step, use the simple linear iterative clustering algorithm to segment the original hyperspectral image obtained in Step 1 into homogeneous regions composed of M superpixels: {ξ1, ξ2, …, ξ m , …, ξ M}, where M is the total number of superpixels, and its value is determined by the classification accuracy. ξ m represents the m-th superpixel. In the embodiment of the present invention, the number of superpixels M is set to 800.
[0050] Second step, construct the following three sets of similarity pairs, and use the formula to obtain the final set of similarity pairs:
[0051]
[0052]
[0053]
[0054] where represent the first set of similarity pairs, the second set of similarity pairs, and the third set of similarity pairs respectively, and represent the probability vectors output by the first and second spatial pooling Transformer networks for the i-th sample x i respectively, and represent the probability vectors output by the first and second spatial pooling Transformer networks for the j-th sample x j respectively, (x i , x j ) ∈ ξ m means that the central pixels of the i-th and j-th samples in the training set belong to the same superpixel, and {·|·} represents the operation of incorporating the similarity pairs that meet the right condition into the corresponding similarity pair set on the left;
[0055] Third step, construct the similarity label matrix The value of each element in this matrix is obtained by the following formula:
[0056]
[0057] Among them, A uv represents the element value at coordinates u and v in the similarity label matrix A, and p u represents the u-th probability vector in the probability vectors output by the first and second spatial pooling Transformer networks for the training set samples, and p v represents the v-th probability vector in the probability vectors output by the first and second spatial pooling Transformer networks for the training set samples, represents the similarity pair included in the similarity pair set ;
[0058] Step 4: Calculate the similarity loss of all samples in the training set according to the following formula:
[0059]
[0060] Among them, represents the similarity loss of the first and second spatial pooling Transformer networks on the training set, and S uv represents the probability that two predicted probability vectors are similar, T represents the transpose operation.
[0061] Step 4.3: Calculate the combined losses of the first and second spatial pooling Transformer networks respectively, and iteratively update the parameters of the two networks until the combined loss function converges, obtaining the trained first and second spatial pooling Transformer networks.
[0062] The combined losses of the first and second spatial pooling Transformer networks described above are respectively:
[0063]
[0064]
[0065] Among them, and respectively represent the combined losses of the first and second spatial pooling Transformer networks, and λ s represents a hyperparameter, and its value range is: 0 ≤ λ s ≤ 1. In the embodiment of the present invention, the parameter λ s is set to 0.2.
[0066] Step 5: Classify the hyperspectral image:
[0067] Using the same method as in step 1, after normalizing the hyperspectral image, extract the pixel blocks to be classified, input all the hyperspectral image pixel blocks to be classified into the trained first and second spatial pooling Transformer networks, take the average of the class probability vectors output by the two networks to obtain the class probability vectors of all samples, and take the class corresponding to the maximum probability in the class probability vector as the classification result of the pixel.
[0068] The effect of the present invention can be further illustrated by the following simulation:
[0069] 1. Simulation experiment conditions
[0070] The simulation platform of the present invention selects python3.7 + pytorch and is completed on a workstation equipped with an Intel Xeon Silver 4210 CPU and 11G of memory and a GeForce RTX 2080Ti.
[0071] The dataset used in the simulation of the present invention is the Indian Pines hyperspectral image dataset, which was obtained by the AVIRIS sensor at the Indian Pine test site in northwestern Indiana. It consists of 145×145 pixels containing 220 spectral bands. Its band range covers 400 - 2500 nanometers. When using this dataset for experiments, its spectral bands will be reduced to 200 to eliminate the influence of the water absorption area. The image has 10249 labeled pixels, divided into 16 pixel classes, and 1770 pixels are randomly selected as the training set, and the rest are the test set.
[0072] 2. Simulation content and its result analysis
[0073] Under the above simulation conditions, the Indian Pines hyperspectral image dataset is classified using the present invention and four existing technologies RLPA, MSSG, DCRN, and S3Net respectively.
[0074] In the simulation experiment, the four existing technologies adopted refer to:
[0075] The prior art RLPA classification method refers to the method for classifying noisy hyperspectral images published by Jiang Junjun et al. on IEEE, that is: J. Jiang, J. Ma, Z. Wang, C. Chen, and X. Liu, “Hyperspectral image classification in the presence of noisy labels,” IEEE Transactions on Geoscience and Remote Sensing, vol. 57, no. 2, pp. 851-865, 2018;
[0076] The prior art MSSG classification method refers to the method for classifying noisy hyperspectral images published by Jiang Junjun et al. on IEEE, that is: J. Jiang, J. Ma, and X. Liu, “Multilayer spectral-spatial graphs for label-noisy robust hyperspectral image classification,” IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 2, pp. 839-852, 2022;
[0077] The prior art DCRN classification method refers to the method for classifying noisy hyperspectral images published by Li Zhaokui et al. on IEEE, that is: Y. Gao, F. Gao, J. Dong, and H.-C. Li, “SAR image change detection based on multiscale capsule network,” IEEE Geoscience and Remote Sensing Letters, vol. 18, no. 3, pp. 484–488, 2020;
[0078] The prior art S3Net classification method refers to the method for classifying noisy hyperspectral images published by Hongyan Zhang et al. in IEEE, namely: H. Xu, H. Zhang, and L. Zhang, “A superpixel guided sample selection neural network for handling noisy labels in hyperspectral image classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 11, pp. 9486 - 9503, 2020.
[0079] The following further describes the effect of the present invention in combination with Figure 4 the simulation diagram.
[0080] Figure 4 (a) is the true ground feature distribution map of the input Indian Pines hyperspectral image. Figure 4 (b) is the classification result map of classifying the Indian Pines hyperspectral image with a noise rate of 40% using the prior art RLPA classification method. Figure 4 (c) is the classification result map of classifying the Indian Pines hyperspectral image with a noise rate of 40% using the prior art MSSG classification method. Figure 4 (d) is the classification result map of classifying the Indian Pines hyperspectral image with a noise rate of 40% using the prior art DCRN classification method. Figure 4 (e) is the classification result map of classifying the Indian Pines hyperspectral image with a noise rate of 40% using the prior art S3Net classification method. Figure 4 (f) is the classification result map of classifying the Indian Pines hyperspectral image with a noise rate of 40% using the classification method of the present invention. Herein, the noise rate represents the proportion of incorrect labels in the entire training set.
[0081] It can be seen from Figure 4 that when noise labels appear, many misclassifications occur in each method, generating many noise points. However, relatively speaking, the present invention can still obtain a better classification result map in the case of containing noise labels, with fewer misclassification regions, which fully demonstrates the robustness of the method proposed by the present invention to noise labels.
[0082] To quantitatively illustrate the performance of the proposed method, three numerical performance indicators commonly used in hyperspectral image classification tasks were selected to measure the differences between the above-mentioned existing methods and the present invention. The performance indicators include overall accuracy (OA) and average accuracy (AA).
[0083] The overall accuracy OA represents the ratio of the number of correctly classified test samples to the total number of test samples;
[0084] The average accuracy AA represents the ratio of the number of correctly classified test samples to the total number of test samples in a certain category;
[0085] The classification performances of the above-mentioned existing hyperspectral image classification technologies and the present invention on the IndianPines dataset were compared using the evaluation indicators, and the results are shown in Table 1;
[0086] Table 1 Quantitative analysis table of the classification results of the present invention and four existing methods
[0087]
[0088] As can be seen from Table 1, compared with the existing four hyperspectral image classification methods, the present invention has higher accuracy in all three indicators, which proves that the present invention can obtain higher hyperspectral image classification accuracy in the case of noisy labels.
[0089] The above simulation experiments show that: the method of the present invention uses the constructed spatial pooling Transformer network to be able to remove redundant information in the feature map while extracting global pixel correlation information, combines a noise-resistant learning algorithm, fully considers the spatial characteristics of hyperspectral images, and at the same time uses clean samples and noise samples in the dataset to solve the problems existing in the existing technical methods, such as the lack of extraction of global information, insufficient and inaccurate feature extraction, and a significant reduction in the classification accuracy of the model when noisy labels appear. It is a very practical hyperspectral image classification method.
Claims
1. A hyperspectral image classification method based on spatial pooling Transformer, characterized in that, Build the first spatial pooling Transformer network and the second spatial pooling Transformer network respectively, and use a noise-resistant learning algorithm to train the first and second spatial pooling Transformer networks. The steps of this classification method are as follows: Step 1, generate a training set: Obtain a hyperspectral image containing at least C pixel categories, with at least S pixel points in each pixel category, and G pixel points with noisy labels in this category. After normalizing the hyperspectral image, with each target pixel point as the center, form a pixel block of 9×9 around it, and use the label of each target pixel point as the label of this pixel block. Combine all the pixel blocks in the hyperspectral image and their corresponding pixel block labels to form a training set, where C≥2, S≥1, and G<S; Step 2, build the first spatial pooling Transformer network and the second spatial pooling Transformer network respectively. The structures of the first and second spatial pooling Transformer networks are exactly the same. Each spatial pooling Transformer network is composed of a spectral feature extraction module, a first spatial pooling Transformer module, a second spatial pooling Transformer module, a third spatial pooling Transformer module, a Transformer encoder, and a classifier connected in series in sequence; Step 3, adopt a warm-up strategy and use cross-entropy loss to train the first and second spatial pooling Transformer networks respectively to obtain the pre-warmed first and second spatial pooling Transformer networks; Step 4, use a noise-resistant learning algorithm to train the pre-warmed first and second spatial pooling Transformer networks: Step 4.1, adopt a cross-selection strategy to calculate the cross-entropy loss of the first and second spatial pooling Transformer networks on the clean sample set respectively; Step 4.2, calculate the similarity loss of the first and second spatial pooling Transformer networks on the training set; Step 4.3, calculate the joint loss of the first and second spatial pooling Transformer networks respectively, and iteratively update the parameters of the two networks until the joint loss function converges to obtain the trained first and second spatial pooling Transformer networks; Step 5, classify the hyperspectral image: Adopt the same method as in Step 1 to extract the image blocks to be classified from the normalized hyperspectral image, input all the hyperspectral image pixel blocks to be classified into the trained first and second spatial pooling Transformer networks, take the average of the class probability vectors output by the two networks to obtain the class probability vectors of all samples, and take the class corresponding to the maximum probability in the class probability vector as the classification result of this pixel.
2. The hyperspectral image classification method based on spatial pooling Transformer according to claim 1, wherein The spectral feature extraction module described in step 2 is composed of a 3D convolutional unit and a 2D convolutional unit connected in series. The structures of the 3D convolutional unit and the 2D convolutional unit are the same, and both are composed of a convolutional layer, a batch normalization layer, and a non-linear activation layer connected in series. The kernel sizes of the convolutional layers in the 3D convolutional unit and the 2D convolutional unit are set to 5×1×1 and 1×1 respectively, and the numbers of kernels are set to 8 and 64 respectively. The non-linear activation layer is implemented using the LeakyReLU function.
3. The hyperspectral image classification method based on spatial pooling Transformer according to claim 1, wherein The Transformer encoder described in step 2 includes a position encoder, a first normalization layer, a multi-head self-attention unit, a second normalization layer, and a multi-layer perceptron. Among them, the position encoder, the first normalization layer, the multi-head self-attention unit, the second normalization layer, and the multi-layer perceptron are connected in series. After the feature matrix output by the position encoder is added to the feature matrix output after the first normalization layer and the multi-head self-attention unit, it is then added to the feature matrix output after the second normalization layer and the multi-layer perceptron. The number of attention heads in the multi-head self-attention unit is set to 8.
4. The hyperspectral image classification method based on spatial pooling Transformer according to claim 1, wherein The structures of the first spatial pooling Transformer module, the second spatial pooling Transformer module, and the third spatial pooling Transformer module described in step 2 are the same. Each spatial pooling Transformer module is composed of a Transformer encoder and a spatial pooling unit connected in series. The structure of the Transformer encoder is the same as that described in claim 3. The spatial pooling unit is composed of a restoration layer, a spatial pooling layer, and a flattening layer connected in series. The restoration layer is used to transform the serialized features into two-dimensional features and then transmit them to the spatial pooling layer. The size of the two-dimensional features is equal to the square root of the length of the serialized features. The spatial pooling layer removes the outermost layer of features of the two-dimensional features obtained by the restoration layer and then transmits them to the flattening layer. The flattening layer flattens the two-dimensional features back into one-dimensional features.
5. The hyperspectral image classification method based on spatial pooling Transformer according to claim 1, wherein The classifier described in step 2 is composed of a global average pooling layer and a fully connected layer connected in series. The number of input nodes of the fully connected layer is set to 64, and the number of output nodes is equal to the number of classes of the hyperspectral image.
6. The hyperspectral image classification method based on spatial pooling Transformer according to claim 1, wherein, The warm-up strategy described in step 3 means that the training set is input into the first and second spatial pooling Transformer networks respectively, the cross-entropy losses are calculated respectively, and the parameters of the two networks are iteratively updated by the gradient descent method until the cross-entropy loss function converges, so as to obtain the first and second spatial pooling Transformer networks after warm-up processing.
7. The hyperspectral image classification method based on spatial pooling Transformer according to claim 1, wherein The cross-entropy loss described in steps 3 and 4 is obtained by the following formula: Among them, represents the cross-entropy loss value. N represents the total number of pixel blocks in the training set in step 3 and the total number of pixel blocks in the selected clean sample set in step 4. C represents the total number of pixel categories in the training set, and y lk represents the sign function. When the category of the l-th sample is k, y lk = 1, otherwise 0, and p lk represents the probability that the i-th sample in the training set is predicted as category k.
8. The hyperspectral image classification method based on spatial pooling Transformer according to claim 7, wherein, The steps for calculating the cross-entropy losses of the first and second spatial pooling Transformer networks on the clean sample set described in step 4.1 are as follows: In the first step, the training set is input into the first spatially pooled Transformer network that has been preliminarily trained. The loss values of all samples are fitted using a two-component Gaussian mixture model. The probability that each sample belongs to the Gaussian component with the minimum mean is used as the clean probability of the sample. All samples in the training set with a clean probability greater than or equal to the threshold τ form the clean sample set of the first spatially pooled Transformer network The rest form the noise sample set Among them, τ is a hyperparameter, and its value range is: 0 ≤ τ < 1; In the second step, using the same method as in the first step, obtain the clean sample set of the second spatial pooling Transformer network and the noise sample set Step 3: Input the set of clean samples selected by the second spatial pooling Transformer network into the first spatial pooling Transformer network to calculate the cross-entropy loss Input the set of clean samples selected by the first spatial pooling Transformer network into the second spatial pooling Transformer network to calculate the cross-entropy loss 9. The hyperspectral image classification method based on spatial pooling Transformer according to claim 8, wherein, The steps for calculating the similarity losses of the first and second spatial pooling Transformer networks on the training set described in step 4.2 are as follows: First, the simple linear iterative clustering algorithm is used to segment the original hyperspectral image obtained in step 1 into a set of superpixels consisting of M superpixels: {ξ1, ξ2, …, ξ m , …, ξ M}, where M is the total number of superpixels, and its value is determined by the classification accuracy. ξ m represents the m-th superpixel; Step 2, construct the following three sets of similar pairs and use the formula to obtain the final set of similar pairs: Among them, respectively represent the first set of similar pairs, the second set of similar pairs, and the third set of similar pairs, and respectively represent the probability vectors output by the first and second spatial pooling Transformer networks for the \(i\)-th sample \(x\) i , and respectively represent the probability vectors output by the first and second spatial pooling Transformer networks for the \(j\)-th sample \(x\) j . \((x\) i , \(x\) j ) \(\in \xi\) m means that the central pixels of the \(i\)-th and \(j\)-th samples in the training set belong to the same superpixel. \(\{\cdot|\cdot\}\) represents the operation of incorporating similar pairs that satisfy the right condition into the corresponding set of similar pairs on the left; Step 3: Construct a similarity label matrix The value of each element in this matrix is obtained by the following formula: Among them, A uv represents the element value at coordinates u and v in the similarity label matrix A, and p u represents the u-th probability vector in the probability vectors output by the first and second spatial pooling Transformer networks for the training set samples, and p v represents the v-th probability vector in the probability vectors output by the first and second spatial pooling Transformer networks for the training set samples, represents being included in the set of similarity pairs in; In the fourth step, calculate the similarity losses of all samples in the training set according to the following formula: Among them, represents the similarity loss of the first and second spatial pooling Transformer networks on the training set, S uv represents the probability that two predicted probability vectors are similar, T represents the transpose operation.
10. The hyperspectral image classification method based on spatial pooling Transformer according to claim 9, characterized in that, The combined losses of the first and second spatial pooling Transformer networks calculated separately as described in step 4.3 are obtained by the following equations: wherein, and respectively represent the combined losses of the first and second spatial pooling Transformer networks, and λ s represents a hyperparameter, and its value range is: 0 ≤ λ s ≤ 1.
Citation Information
Patent Citations
Hyperspectral image classification method based on comparative generative adversarial network
CN113469084A
Hyperspectral Image Classification Method Based on Contrast Generative Adversarial Networks
CN113469084B
Multi-target garbage detection method based on improved YOLOv5 model
CN116452950A
Sports video motion identification method based on motion granularity grouping structure
CN116524596A