Hyperspectral image classification method based on frequency prompt and space spectrum Transform network
By adopting a method based on frequency prompt and null spectrum Transformer network in hyperspectral image classification, the problem of failure to fully utilize the spatial and spectral information of hyperspectral images in the prior art is solved, and a higher classification accuracy and information fusion effect are achieved.
Patent Information
- Application Number
- CN202510023366.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art has not fully mined the spatial information and spectral information of hyperspectral images, it is difficult to effectively capture the correlation between the two, and rarely pay attention to rich information in the frequency domain.
The hyperspectral image classification method based on frequency cues and null spectrum Transformer network is adopted to extract frequency cues through discrete cosine transformation, combining multi-stage strategy and null spectrum cross-attention design to promote effective interaction and enhancement of spatial information and spectral information.
The accuracy of hyperspectral image classification is improved, and the complementary effect of features at different stages is enhanced by effectively fusion of spatial information, spectral information and frequency domain information.
Smart Images

Figure CN119963898A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of hyperspectral image classification, and in particular relates to a hyperspectral image classification method based on frequency prompts and a spatial spectrum Transformer network. Background Art
[0002] Hyperspectral remote sensing images are a type of image data that contains a large number of bands, and each pixel usually has spectral information of hundreds of different bands. These hyperspectral data provide more spectral dimensions than traditional RGB images, making hyperspectral images widely used in areas such as land object classification, vegetation monitoring, and pollutant detection. However, due to the high-dimensional characteristics of hyperspectral images, how to effectively extract useful information from images and accurately classify them remains a challenge in hyperspectral image processing.
[0003] Traditional hyperspectral image classification methods are mostly based on a combination of feature extraction and classification algorithms. Early classification methods mainly rely on pixel-based techniques, such as support vector machines and K nearest neighbors. Although these methods can provide certain classification accuracy, they ignore the complex relationship between spatial and spectral information, resulting in poor classification performance in complex scenes. With the development of deep learning technology, methods based on convolutional neural networks and long short-term memory networks have gradually been applied to the field of hyperspectral image classification, and have achieved relatively remarkable results. However, deep learning methods still face some problems in hyperspectral image classification, such as high dimensionality of data, complexity of models, and high requirements for computing resources. Transformer networks, as an architecture based on self-attention mechanisms, have achieved remarkable results in image classification, natural language processing and other fields in recent years. In the hyperspectral image classification task, the Transformer model can capture the long-range dependencies between different bands in the image, which provides a new idea for improving classification performance. However, the traditional Transformer structure mainly focuses on global information modeling, and the capture of local information is not ideal. Therefore, how to effectively fuse spatial-spectral information and frequency domain information in the Transformer network has become a key issue in improving the accuracy of hyperspectral image classification. In addition, most of the existing deep learning-based methods pay little attention to the rich information in the frequency domain and fail to fully exploit the spatial and spectral information of hyperspectral images.
[0004] Hong et al. proposed a Transformer-based hyperspectral image classification network that uses grouped spectral embedding and cross-layer adaptive fusion to better model the sequence spectral information. This method only uses spectral information and ignores spatial information. Sun et al. proposed a spectral spatial feature tokenization converter method that combines convolutional neural networks and Transformers to represent features as semantic tags. This method extracts spectral features and spatial features at the same time, which may lead to insufficient fusion of spectral and spatial information and make it difficult to effectively capture the correlation between the two. Mei et al. designed a group-aware hierarchical Transformer network that uses a grouped pixel embedding module to simulate the local environment and uses a hierarchical architecture to extract multi-scale spatial spectral features. This method also extracts spectral features and spatial features at the same time and is difficult to effectively capture the correlation between the two. Qi et al. proposed a new method called global-local 3D convolutional Transformer network, in which 3D convolution is embedded in a dual-branch Transformer to simultaneously capture the global-local correlation of spectrum and space. This method processes spectral features and spatial features separately and fully mines the spatial information and spectral information of the image, but does not pay attention to the rich information in the frequency domain. Qiao et al. proposed a spectral-spatial-frequency Transformer network that uses discrete Fourier transform and learnable filters to extract features from the spectral-spatial and frequency domains. Qiao et al. further designed a dual-frequency Transformer network that extracts high-frequency and low-frequency features from hyperspectral images by using separate branches. Although these two methods utilize frequency domain information, they do not extract spectral features and spatial features separately.
[0005] In summary, the existing technology fails to fully explore the spatial and spectral information of hyperspectral images, does not process spectral features and spatial features separately, making it difficult to effectively capture the correlation between the two and pays little attention to the rich information in the frequency domain. Summary of the invention
[0006] In order to solve the above technical problems, the present invention provides a hyperspectral image classification method based on frequency prompts and spatial spectrum Transformer network, comprising:
[0007] S1: Obtain a hyperspectral image with a target category label and crop it into hyperspectral image patches centered on the pixels to be classified to form a training dataset;
[0008] S2: constructing a hyperspectral image classification model; the model includes: a spatial branch module, a spectral branch module and a fusion module;
[0009] S3: Input the hyperspectral image blocks in the training data set into the hyperspectral image classification model to perform model training;
[0010] S4: Input the hyperspectral image to be detected into the trained model for image classification to obtain the category label.
[0011] Beneficial effects of the present invention:
[0012] The present invention designs the network architecture based on frequency cues and spatial-spectral Transformer networks. First, the low-frequency part of the hyperspectral image is extracted through discrete cosine transform and the learnable high-frequency cues are used to form frequency cues to utilize the rich frequency domain information in the hyperspectral image. Secondly, a multi-stage strategy is used to fully explore the spatial information and spectral information contained in the hyperspectral image. At the same time, a spatial-spectral fusion Transformer and a multi-stage fusion Transformer are designed through a spatial-spectral cross-self-attention, thereby better promoting the effective interaction between the spatial information and spectral information of each stage and enhancing the complementary effect between the fusion features of different stages, thereby improving the classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a framework diagram of a hyperspectral image classification method based on frequency cues and spatial spectrum Transformer network of the present invention. DETAILED DESCRIPTION
[0014] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0015] A hyperspectral image classification method based on frequency cues and spatial spectrum Transformer network, such as Figure 1 As shown, including:
[0016] S1: Obtain a hyperspectral image with a target category label and crop it into hyperspectral image blocks centered on the pixels to be classified to form a training dataset for training a hyperspectral image classification model;
[0017] S2: constructing a hyperspectral image classification model; the model includes a spatial branch, a spectral branch and a fusion module;
[0018] S3: Input the hyperspectral image blocks in the training data set into the hyperspectral image classification model to perform model training;
[0019] S31: inputting the hyperspectral image block into the spatial convolution module and the spectral convolution module respectively for feature extraction, obtaining local spatial features and local spectral features, and flattening these features to form a fixed-size spatial feature vector and a spectral feature vector;
[0020] S32: performing a two-dimensional discrete cosine transform on the hyperspectral image block in the spatial dimension and a one-dimensional discrete cosine transform on the channel dimension to obtain a spatial frequency image block and a spectral frequency image block, then taking out a low-frequency part of the spatial frequency image block and a learnable high-frequency prompt, splicing and flattening them in the spatial dimension to obtain a spatial frequency prompt, and taking out a low-frequency part of the spectral frequency image block and a learnable high-frequency prompt, splicing and flattening them in the channel dimension to obtain a spectral frequency prompt;
[0021] S33: Send the spatial feature vector and the spectral feature vector to a series of spatial Transformers respectively to capture the global spatial information and the global spectral information;
[0022] The spatial Transformer consists of a multi-head spatial self-attention mechanism and a multi-layer perceptron, wherein each feature vector aggregates global spatial information through the spatial self-attention mechanism and outputs the global spatial information through the multi-layer perceptron; the spectral Transformer consists of a multi-head spectral self-attention mechanism and a multi-layer perceptron, wherein each feature vector aggregates global spectral information through the spectral self-attention mechanism and outputs the global spectral information through the multi-layer perceptron;
[0023] S34: spatial self-attention that feeds the spatial frequency cues into a series of spatial Transformers, and spectral self-attention that feeds the spectral frequency cues into a series of spectral Transformers;
[0024] S35: Send the output of the spatial Transformer and the output of the spectral Transformer to the spatial-spectral fusion Transformer to obtain refined spatial-spectral fusion features;
[0025] S36: reshape the output of the spatial transformer and the output of the spectral transformer into the original shape and size and perform the same second-stage operation as S31-S35, and then reshape the output of the second-stage spatial transformer and the output of the spectral transformer into the original shape and size and perform the same third-stage operation;
[0026] S37: sending the output of the first-stage spatial-spectral fusion Transformer, the output of the second-stage spatial-spectral fusion Transformer, and the output of the third-stage spatial-spectral fusion Transformer to the multi-stage fusion Transformer;
[0027] S38: The vector output by the multi-stage fusion Transformer is passed through a fully connected layer and a softmax activation function to predict the probability of the ground object coverage category;
[0028] S39: The classification loss is established by measuring the difference between the probability of the predicted ground object coverage category and the actual image label. The classification loss is used as the final loss function of the hyperspectral image classification model. During the model training process, the model training is completed by minimizing the loss function.
[0029] S4: Input the hyperspectral image to be detected into the trained model to obtain the category label.
[0030] In this embodiment, the hyperspectral image block X is input into the spatial convolution module and the spectral convolution module respectively for feature extraction to obtain local spatial features and local spectral features; these features are flattened to form a fixed-size spatial feature vector and a spectral feature vector, including:
[0031] The spatial convolution module includes a 3×3 two-dimensional convolution, batch normalization, and ReLU activation function, and the spectral convolution module includes two 1×1×7 three-dimensional convolutions, a 1×1×3 three-dimensional convolution, batch normalization, and ReLU activation function. The process can be expressed as:
[0032] T spa =Flatten(f(δ(Conv2D 3×3 (X))))
[0033] T spe =Flatten(f(δ(Conv3D 1×1×3 (Conv3D 1×1×7 (Conv3D 1×1×7 (X))))))
[0034] Among them, Conv2D 3×3 (·) represents a 3×3 two-dimensional convolution; δ(·) represents batch normalization; f(·) represents the ReLU activation function; Conv3D 1×1×7 (·) represents a 1×1×7 three-dimensional convolution; Conv3D 1×1×3 (·) represents a 1×1×3 three-dimensional convolution; Flatten(·) represents a flattening operation; T spa represents the spatial eigenvector; T spe represents the spectral feature vector.
[0035] In this embodiment, the hyperspectral image block is subjected to a two-dimensional discrete cosine transform in the spatial dimension and a one-dimensional discrete cosine transform in the channel dimension to obtain a spatial frequency image block and a spectral frequency image block, and then the low-frequency part F of the spatial frequency image block is taken out. spa_low and learnable high-frequency cues P spa_high Splicing and flattening in the spatial dimension to obtain the spatial frequency hint P spa , take out the low-frequency part F of the spectral frequency image block spe_low and learnable high-frequency cues P spe_high Splicing and flattening in the channel dimension to obtain the spectral frequency hint P spe .include:
[0036] For spatial frequency hint P spa and spectral frequency hint P spe It can be expressed as:
[0037] P spa =Flatten(Concat HW (F spa_low ,P spa_high ))
[0038] P spe =Flatten(Concat C (F spe_low ,P spe_high ))
[0039] Among them, P spa Indicates spatial frequency cue; P spe Indicates spectral frequency prompt; Concat HW (·) indicates concatenation of spatial dimensions; Flatten(·) indicates flattening operation; Concat C (·) represents the concatenation of channel dimensions; F spa_low Represents the low-frequency part of the spatial frequency image block; P spe_high Learnable high-frequency cues representing spectral branches.
[0040] In this embodiment, the spatial frequency cues are fed into the spatial self-attention of a series of spatial transformers, and the spectral frequency cues are fed into the spectral self-attention of a series of spectral transformers, including:
[0041] The output of the spatial self-attention includes:
[0042]
[0043] The output of the spectral self-attention includes:
[0044]
[0045] Among them, S pa SA(Q spa ,K spa ,V spa ) represents the output of spatial self-attention; S pe SA(Q spe , K spe , V spe ) represents the output of spectral self-attention; Q spa K represents the query matrix of spatial self-attention; spa V represents the key matrix of spatial self-attention; spa The value matrix representing spatial self-attention; d spa K spa Dimension; Q spe K represents the query matrix of spectral self-attention; spe V represents the bond matrix of spectral self-attention; spe The value matrix representing the spectral self-attention; d spe K spe dimension; softmax(·) represents the normalized exponential function; P spa Indicates spatial frequency cue; P spe represents the spectral frequency cue and T represents the matrix transpose.
[0046] In this embodiment, the output T of the spatial Transformer spa With the output T of the spectral Transformer spe Send it to the spatial-spectral fusion Transformer to obtain the refined spatial-spectral fusion feature T f ,include:
[0047] The spatial-spectral fusion Transformer consists of spatial-spectral cross attention, layer normalization, and a multi-layer perceptron. The spatial-spectral cross attention SSCA consists of spectral cross attention S pe CA and spatial cross attention S pa CA. Specifically, T spa Feed it into the normalized layer and generate the query matrix Q through linear transformation, T spe The input layer is normalized and linearly transformed to generate the key matrix K and the value matrix V. Then, the spectral cross attention is calculated by matrix multiplication using Q, K and V. After that, the output of the spectral cross attention is linearly transformed to generate a new key matrix K' and value matrix V'. Similarly, Q, K' and V' are used to calculate the spatial cross attention by matrix multiplication. The output of the spectral cross attention includes:
[0048]
[0049] The output of the spatial cross attention includes:
[0050]
[0051] Among them, S pe CA(Q,K,V) represents the output of spectral cross attention; S pa CA(Q,K′,V′) represents the output of spatial cross attention; softmax(·) represents the normalized exponential function; d pe represents the dimension of K; d pa represents the dimension of K′; T represents matrix transpose.
[0052] In this embodiment, the output T of the first stage spatial spectrum fusion Transformer is f1 , the output T of the second stage spatial-spectral fusion Transformer f2 And the output T of the third stage spatial spectrum fusion Transformer f3 Send to the multi-stage fusion Transformer, including:
[0053] The multi-stage fusion Transformer consists of two spatial spectral cross attention SSCA, layer normalization and multi-layer perceptron. f1 and T f3 It is sent to a SSCA to obtain an intermediate feature, which is then combined with T f2 It is sent to another SSCA to obtain more refined and rich features, and finally passes through a layer normalization and a multi-layer perceptron to obtain higher-level image features for final classification.
[0054] In this embodiment, the classification loss of the model includes:
[0055]
[0056] Where M represents the number of training samples; y i is the one-hot label vector of the i-th sample, is the predicted label distribution.
[0057] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A hyperspectral image classification method based on frequency cues and spatial spectrum Transformer network, characterized in that: include: S1: Obtain a hyperspectral image with a target category label and crop it into hyperspectral image patches centered on the pixels to be classified to form a training dataset; S2: construct a hyperspectral image classification model; the model includes: a spatial branch module, a spectral branch module and a fusion module; the fusion module includes: a spatial-spectral fusion Transformer and a multi-stage fusion Transformer; S3: Input the hyperspectral image blocks in the training data set into the hyperspectral image classification model to perform model training; S4: Input the hyperspectral image to be detected into the trained model for image classification to obtain the category label.
2. A hyperspectral image classification method based on frequency prompt and spatial spectrum Transformer network according to claim 1, characterized in that: The spatial convolution module includes a 3×3 two-dimensional convolution, batch normalization, and ReLU activation function; the spectral convolution module includes two 1×1×7 three-dimensional convolutions, a 1×1×3 three-dimensional convolution, batch normalization, and ReLU activation function.
3. The hyperspectral image classification method based on frequency prompt and spatial spectrum Transformer network according to claim 1 is characterized in that: The hyperspectral image blocks in the training data set are input into the hyperspectral image classification model to perform model training, including: S31: inputting the hyperspectral image block into the spatial convolution module and the spectral convolution module respectively for feature extraction, obtaining local spatial features and local spectral features, and flattening these features to form a fixed-size spatial feature vector and a spectral feature vector; S32: performing a two-dimensional discrete cosine transform on the hyperspectral image block in the spatial dimension and a one-dimensional discrete cosine transform on the channel dimension to obtain a spatial frequency image block and a spectral frequency image block, then taking out a low-frequency part of the spatial frequency image block and a learnable high-frequency prompt, splicing and flattening them in the spatial dimension to obtain a spatial frequency prompt, and taking out a low-frequency part of the spectral frequency image block and a learnable high-frequency prompt, splicing and flattening them in the channel dimension to obtain a spectral frequency prompt; S33: Send the spatial feature vector and the spectral feature vector to a series of spatial Transformers respectively to capture the global spatial information and the global spectral information; The spatial Transformer consists of a multi-head spatial self-attention mechanism and a multi-layer perceptron, wherein each feature vector aggregates global spatial information through the spatial self-attention mechanism and outputs the global spatial information through the multi-layer perceptron; the spectral Transformer consists of a multi-head spectral self-attention mechanism and a multi-layer perceptron, wherein each feature vector aggregates global spectral information through the spectral self-attention mechanism and outputs the global spectral information through the multi-layer perceptron; S34: spatial self-attention that feeds the spatial frequency cues into a series of spatial Transformers, and spectral self-attention that feeds the spectral frequency cues into a series of spectral Transformers; S35: Send the output of the spatial Transformer and the output of the spectral Transformer to the spatial-spectral fusion Transformer to obtain refined spatial-spectral fusion features; S36: reshape the output of the spatial transformer and the output of the spectral transformer into the original shape and size and perform the same second-stage operation as S31-S35, and then reshape the output of the second-stage spatial transformer and the output of the spectral transformer into the original shape and size and perform the same third-stage operation; S37: sending the output of the first-stage spatial-spectral fusion Transformer, the output of the second-stage spatial-spectral fusion Transformer, and the output of the third-stage spatial-spectral fusion Transformer to the multi-stage fusion Transformer; S38: The vector output by the multi-stage fusion Transformer is passed through a fully connected layer and a softmax activation function to predict the probability of the ground object coverage category; S39: The classification loss is established by measuring the difference between the probability of the predicted ground object coverage category and the actual image label. The classification loss is used as the final loss function of the hyperspectral image classification model. During the model training process, the model training is completed by minimizing the loss function.
4. A hyperspectral image classification method based on frequency prompt and spatial spectrum Transformer network according to claim 3, characterized in that: The spatial frequency prompt includes: P spa =Flatten(Concat HW (F spa_low ,P spa_high )) The spectral frequency prompts include: P spe =Flatten(Concat C (F spe_low ,P spe_high )) Among them, P spa Indicates spatial frequency cue; P spe Indicates spectral frequency prompt; Concat HW (·) indicates concatenation of spatial dimensions; Flatten(·) indicates flattening operation; Concat C (·) represents the concatenation of channel dimensions; F spa_low Represents the low-frequency part of the spatial frequency image block; P spe_high Learnable high-frequency cues representing spectral branches.
5. The method for hyperspectral image classification based on frequency prompt and spatial spectrum Transformer network according to claim 3 is characterized in that: The output of the spatial self-attention includes: The output of the spectral self-attention includes: Among them, S pa SA(Q spa ,K spa ,V spa ) represents the output of spatial self-attention; S pe SA(Q spe , K spe , V spe ) represents the output of spectral self-attention; Q spa K represents the query matrix of spatial self-attention; spa V represents the key matrix of spatial self-attention; spa The value matrix representing spatial self-attention; d spa K spa Dimension; Q spe K represents the query matrix of spectral self-attention; spe V represents the bond matrix of spectral self-attention; spe The value matrix representing the spectral self-attention; d spe K spe dimension; softmax(·) represents the normalized exponential function; P spa Indicates spatial frequency cue; P spe represents the spectral frequency cue and T represents the matrix transpose.
6. The method for hyperspectral image classification based on frequency cueing and spatial spectrum Transformer network according to claim 3, characterized in that: The output of the spatial transformer and the output of the spectral transformer are sent to the spatial-spectral fusion transformer to obtain refined spatial-spectral fusion features, including: The spatial-spectral fusion Transformer consists of spatial-spectral cross attention SSCA, layer normalization and multi-layer perceptron, where the spatial-spectral cross attention SSCA consists of spectral cross attention S pe CA and spatial cross attention S pa CA composition; The output T of the spatial Transformer spa The query matrix Q is generated by linear transformation and the output T of the spectral transformer is sent to the normalized layer. spe The input layer is normalized and linearly transformed to generate the key matrix K and the value matrix V; the spectral cross attention S is calculated by matrix multiplication using Q, K and V. pe The output of CA; the spectral cross attention S pe The output of CA generates a new key matrix K′ and value matrix V′ through linear transformation. Similarly, Q, K′ and V′ are used to calculate the spatial cross attention S through matrix multiplication. pa The output of CA; the spatial cross attention S pa CA obtains refined spatial-spectral fusion features through multi-layer perceptron processing.
7. The hyperspectral image classification method based on frequency prompt and spatial spectrum Transformer network according to claim 3 is characterized in that: The output of the first-stage spatial-spectral fusion Transformer, the output of the second-stage spatial-spectral fusion Transformer, and the output of the third-stage spatial-spectral fusion Transformer are sent to the multi-stage fusion Transformer, including: The multi-stage fusion Transformer consists of two spatial-spectral cross attention SSCA, layer normalization and multi-layer perceptron. The output T of the first stage spatial-spectral fusion Transformer is first f1 And the output T of the third stage spatial spectrum fusion Transformer f3 It is sent to an SSCA to obtain an intermediate feature, which is then fused with the output T of the Transformer in the second stage of spatial spectrum. f2 It is sent to another SSCA to obtain more refined and rich features, and finally passes through a layer normalization and a multi-layer perceptron to obtain higher-level image features for final classification.
8. The method for hyperspectral image classification based on frequency prompt and spatial spectrum Transformer network according to claim 3 is characterized in that: The classification loss of the model includes: Among them, L cls represents the classification loss; M represents the number of training samples; y i Represents the one-hot label vector of the i-th sample; represents the predicted label distribution; T represents the matrix transpose.
Citation Information
Cited By
Hyperspectral image reconstruction method based on spatial-spectral double-prior frequency domain enhancement
CN117764862A
A hyperspectral image reconstruction method based on spectral and spatial dual-prior frequency domain enhancement
CN117764862B