Satellite-borne hyperspectral image classification method based on lightweight spectrum-space joint feed-forward convolution Transform network

Through the lightweight spectroscopy-space combined feedforward convolutional Transformer network, the problems of large memory usage and slow speed in the classification of satellite hyperspectral images are solved, and high-precision and fast classification effects are achieved, which are suitable for real-time remote sensing tasks of satellite platforms.

CN120259872APending Publication Date: 2025-07-04SHANGHAI INSTITUTE OF TECHNICAL PHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510246422.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing hyperspectral remote sensing image classification algorithms have large memory usage on resource-constrained satellite platforms, slow prediction results, and low classification accuracy under small sample conditions, making it difficult to meet the needs of real-time remote sensing tasks.

Method used

The satellite-borne hyperspectral image classification method based on lightweight spectrospatial combined feedforward convolutional Transformer network is adopted. Feature extraction is performed through two-dimensional memory high-efficiency dimensionality reduction algorithm, combined with lightweight convolutional subnetwork and feedforward convolutional Transformer model, spectral and spatial, local and global information are captured, and adaptive average pooling layer and fully connected layer are used for classification.

Benefits of technology

It realizes high-precision classification in extremely low computing resources and shortest inference time, reduces the complexity of deep learning models, and maintains robust classification performance in sample scarce scenarios, and is suitable for resource-constrained satellite platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259872A_ABST
    Figure CN120259872A_ABST
Patent Text Reader

Abstract

The invention discloses a satellite-borne hyperspectral image classification method based on a lightweight spectrum-space joint feed-forward convolution Transform network (SSUFCT). The satellite-borne hyperspectral image classification method comprises the following steps of: selecting a satellite-borne hyperspectral image; according to the scheme, firstly, a novel two-dimensional memory efficient dimension reduction algorithm is adopted for feature extraction; in addition, the constructed SSUFCT model efficiently exerts the capability of capturing spectrum and space and local and global information of the designed lightweight convolution sub-network and the feed-forward convolution Transform model; and finally, according to the extracted high-level spectrum-spatial feature map, obtaining a classification result by using a self-adaptive average pooling layer and a full connection layer. According to the method, almost highest classification precision is realized by using extremely low computing resources and shortest reasoning time, the complexity of applying a deep learning model to the field of hyperspectral remote sensing image classification is greatly reduced, and stable classification performance can still be kept in various scenes with scarce samples; and an important reference is provided for on-orbit intelligent development of technologies such as satellite-borne hyperspectrum and fluorescence hyperspectrum.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hyperspectral remote sensing image processing, and in particular to a method for classifying spaceborne hyperspectral images based on a lightweight spectral-spatial joint feedforward convolutional Transformer network. Background Art

[0002] All kinds of surface anomalies caused by natural and human factors have significant suddenness and complexity, resulting in the fact that the timeliness of traditional remote sensing detection far cannot meet the needs of instant remote sensing tasks such as disaster relief, precision agriculture, and ecological supervision that require rapid response within hours. Instant detection of surface anomalies has become a high-tech strategic high point highly valued by various countries. Relying on the fine characterization of the spectral features of ground objects by hyperspectral technology, the classification of spaceborne hyperspectral images can clearly depict the disaster situation in the affected area, the distribution of vegetation fluorescence intensity, the composition of substances contained in soil and water, etc. on a large scale, providing an important basis for quickly discovering various anomalies.

[0003] Early hyperspectral image classification work mainly focused on mining spectral feature information, such as methods like spectral angle matching and spectral information divergence. However, in actual scenarios, a pixel often contains the mixed spectra of multiple ground objects. Simply relying on the pixel spectral information, it is very difficult to accurately distinguish these mixed components, and the classification accuracy is low and it does not have good noise robustness. In order to extract effective features and alleviate the Hughes phenomenon caused by high-dimensional redundant data, researchers combined dimensionality reduction techniques with machine learning-based classifiers for classification. However, most dimensionality reduction algorithms such as PCA and SVD algorithms need to first merge the length and width of each band image of the hyperspectral data into a one-dimensional vector, which destroys the spatial structure features of the image, and the memory occupation during the process of solving the eigenvector is very large. In addition, new machine learning classification methods such as kernel features and sparse representation have complex mathematical calculation processes and rely on manually designed features by domain experts, making it difficult to quickly and accurately implement on various types of hyperspectral data.

[0004] The deep learning algorithm based on big data and neural networks extracts deeper and more abstract features through step-by-step training, with high result accuracy, providing a new solution idea for hyperspectral image classification. However, when simply using convolutional networks, Transformers, long short-term memory networks, etc. to process high-dimensional data, the computational overhead is large and the efficiency is low. Although the LWNet and LiteDepthwiseNet proposed by Zhang, Cui, and others reduce the model complexity by using depthwise separable convolutions, stripping activation layers, and normalization layers, these methods will all lead to worse model stability and accuracy. Although the CNN-Swin network developed by Li, et al. reduces the size of the input block and the number of training parameters, operations such as shifting and rotation will increase the amount of computation, and the speed advantage is not obvious in a single-layer model. Therefore, it is necessary to develop a lighter, faster, and highly accurate hyperspectral image classification model to meet the requirements of resource-constrained on-orbit platforms and enhance the feasibility of actual deployment. Summary of the Invention

[0005] In view of the above problems, the present invention proposes a spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network. This method solves the problems of existing hyperspectral remote sensing classification algorithms based on intelligent learning, such as large memory occupancy, slow speed of obtaining prediction results, and low classification accuracy under small sample conditions. This model can be used as an efficient benchmark backbone architecture for hyperspectral image classification and has broad application prospects in the field of intelligent real-time remote sensing.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network, characterized in that a two-dimensional memory-efficient dimensionality reduction algorithm is used for feature extraction, and the constructed SSUFCT model realizes efficient capture of spectral, spatial, local, and global information through a lightweight convolutional subnet and a feedforward convolutional Transformer model; finally, according to the extracted high-level spectral-spatial feature map, an adaptive average pooling layer and a fully connected layer are used to obtain the classification result.

[0007] Furthermore, a spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network has the following specific implementation steps:

[0008] P1: Denote the original hyperspectral three-dimensional cube data as where l is the image length, w is the image width, and b is the number of spectral bands;

[0009] P2: The two-dimensional principal component analysis algorithm oriented to the spectral direction is used to reduce the dimensionality of the hyperspectral data. This process greatly saves the running memory and retains the spatial structure information. After reducing the number of bands from b to c, the dimensionality-reduced hyperspectral data is expressed as

[0010] P3: For each pixel s to be classified in the dimensionality-reduced hyperspectral data I i,j , where i = 1, 2, …, l and j = 1, 2, …, w, a three-dimensional image patch with a window size of d×d and a depth of c is extracted centered on the pixel (i, j), and they are stacked to form a new input space where p and q represent the length and width of the new space respectively;

[0011] P4: Each vector X i,j = [x 1,1,c , …, x i,j,c is used as the input of the spectral-spatial information joint network. The spectral and spatial features are extracted through a three-dimensional convolutional neural network (i.e., 3-D CNN) layer respectively, and the spatial features are extracted through a two-dimensional depthwise separable convolution (i.e., 2-D DSC) layer and the fused features are shaped into form, where m, n, and t are the height, width, and depth of the new feature space Z respectively;

[0012] P5: is used as the input of the feed-forward convolutional Transformer network. It undergoes linear projection, positional encoding embedding, normalization, and multi-head attention mechanism processing, and then is connected with the input through a residual connection and normalized to focus on the global features of each image patch;

[0013] P6: After the Transformer global feature extraction part, its output realizes information exchange within the local area through a two-dimensional inverted residual convolution block in the feed-forward convolutional layer;

[0014] P7: The output of the above feed-forward convolutional Transformer network aggregates the features at different positions through an adaptive average pooling layer, and outputs the final classification result through a fully connected layer.

[0015] Furthermore, the implementation process of the two-dimensional principal component analysis algorithm oriented to the spectral direction in step P2 is as follows:

[0016] A1: Combining the storage and operation methods of hyperspectral data in Python, after the data is loaded, the lengths l, widths w, and spectra b of the first, second, and third dimensions are regarded as the "spectrum", "length", and "width" dimensions respectively to achieve the purpose of data rotation, and the new hyperspectral data is obtained

[0017] A2: Calculate the mean of the w,b plane of S2 along the length l direction and centralize it to obtain the three-dimensional image Q w×b×l :

[0018] Q = S2 - E(S2)(1)

[0019] In the formula, E(·) represents the mean of the image.

[0020] A3: For the three-dimensional image Q w×b×l , it is considered that there are l images with a resolution of w×b. Each image is represented by q(i) (i = 1, 2, …, l). Therefore, the specific expression of E(Q) is Calculate the overall covariance matrix G of the image sample matrix r ;

[0021]

[0022] A4: Obtain the covariance matrix G by eigenvalue decomposition r The eigenvectors corresponding to the first c largest eigenvalues are used to form the optimal projection matrix W with a dimension of b×c. Its expression is:

[0023] W = [w1, w2, …, w c (3)

[0024] Project the image matrix S2 onto the obtained projection matrix W to obtain l hyperspectral data matrices with a spatial size of w×c after dimensionality reduction The expression of each principal component I0(i) = [I1, I2, …, I c of the image S2 starting from the l dimension is:

[0025] I0(i) = S2W(i)(4)

[0026] A5: Re-convert the first, second, and third dimensions of I0 into the initial hyperspectral data format, that is, the length l, width w, and spectral channel c dimensions, to obtain the finally dimension-reduced data I l×w×c .

[0027] Furthermore, the implementation process of extracting features by the spectral-spatial information joint network in step P4 is as follows:

[0028] B1: First, use a layer of 3-D CNN to perform spectral-spatial feature extraction on each hyperspectral data block X of the new input space X i,j to obtain the four-dimensional feature map Z1;

[0029] B2: After obtaining the spectral-spatial feature map through 3-D CNN, merge the spectral dimension and the channel dimension into to obtain the new three-dimensional feature map Z2;

[0030] B3: Use depthwise separable two-dimensional convolution to extract features from Z2, further enhancing the local spatial information of the extracted features.

[0031] Further, the implementation process of focusing on the global features between each input picture block in step P5 is as follows:

[0032] C1: The one-dimensional vector input x after linear projection is embedded through positional encoding to record the position information of each image block, and a learnable positional encoding is added to its embedding, as shown in formula (5);

[0033]

[0034] In the formula, E is the embedding projection process, x is the vector after linear projection, and E pos represents positional embedding.

[0035] C2: Perform normalization processing on the data; more comprehensively focus on the global features of each picture block by running the multi-head attention mechanism, as shown in formula (6).

[0036] MHSA = Concat(SA1, SA2, …, SA h )(6)

[0037] In the formula, MHSA is the output vector of the multi-head attention mechanism, and Concat(·) represents the concatenation operation.

[0038] C3: Add the result of MHSA to the input x and perform normalization to obtain the output features.

[0039] Further, the implementation process of the feed-forward convolutional Transformer in step P6 for information exchange within the local area through two-dimensional inverted residual convolutional blocks is as follows:

[0040] D1: Reshape the normalized multi-head attention output sequence L into a two-dimensional image feature map L 2D through Seqto2D, providing the possibility of introducing locality into the network, and its mathematical expression is:

[0041] L 2D = Seqto2D(L)(7)

[0042] D2: Use a convolutional layer with a kernel size of 1×1 to replace the first fully connected layer in the multi-layer perceptron of the traditional Transformer model;

[0043] D3: After the first convolutional layer, use a depth convolution with a kernel size of 3×3 to fuse the information between spatially adjacent pixels;

[0044] D4: After the depth convolution layer, a convolution layer with a kernel size of 1×1 is used again to replace the second fully connected layer in the multi-layer perceptron. The mathematical expressions of steps D2 to D4 are as follows:

[0045]

[0046] In the formula, σ is the RELU activation function, respectively represent the first 1×1 convolution layer, the depth convolution layer, and the second 1×1 convolution layer, represents the convolution operation, and Y 2d is the output through the two-dimensional inverted residual convolution block.

[0047] D5: Add the input L 2D and the output Y 2d to obtain the final output of the feed-forward convolutional Transformer.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. The adopted two-dimensional spectral dimensionality reduction algorithm directly calculates the covariance matrix from the perspective of the multi-layer space-spectral matrix, without performing destructive operations of merging the length and width of the image, avoiding memory overflow that may be caused by calculating the covariance of an extremely large matrix, and improving the average gradient of the image by retaining the spatial features of the original image.

[0050] 2. The designed novel 3-D CNN (i.e., three-dimensional convolutional neural network) and 2-D DSC (i.e., two-dimensional depthwise separable convolution) sub-networks with extremely low parameter and computational amounts efficiently extract the complex spectral-spatial joint features of ground objects, and the generated feature maps are also beneficial to reducing the complexity of the subsequent Transformer network.

[0051] 3. The designed feed-forward convolutional Transformer module can not only capture the long semantic information contained in the spectral sequence through the attention mechanism, but the feed-forward inverted residual convolution block also greatly reduces the complexity of the model and can extract the local spatial and spectral features between adjacent pixels of the reconstructed two-dimensional feature map. Using a single Encoder structure can fully focus on the high-level features of the hyperspectral image, and at the same time has the capabilities of fast speed and low computational consumption.

[0052] 4. The series hybrid lightweight model composed of the memory-saving two-dimensional spectral dimensionality reduction, CNN, and Transformer can be used as a benchmark general framework for spaceborne hyperspectral image classification. Subsequently, a detection model with stronger specialization and detail expression capabilities can be explored based on this architecture to meet more complex and diverse requirements.

[0053] 5. The present invention discloses a spaceborne hyperspectral image classification method based on a lightweight spectral-spatial unified feedforward convolutional transformer network (the full English name of "spectral-spatial unified feedforward convolutional transformer network" is Spectral-Spatial Unified Feedforward Convolutional Transformer, abbreviated as "SSUFCT"). This solution first uses a new type of two-dimensional memory-efficient dimensionality reduction algorithm for feature extraction, saving 75% of the memory compared with traditional methods and breaking the traditional idea that spectral dimensionality reduction must destroy the spatial structure of hyperspectral images. In addition, the constructed SSUFCT model effectively utilizes the designed lightweight convolutional sub-network and feedforward convolutional transformer model to capture the spectral, spatial, local, and global information. Finally, according to the extracted high-level spectral-spatial feature maps, an adaptive average pooling layer and a fully connected layer are used to obtain the classification results. The method of the present invention achieves almost the highest classification accuracy with extremely low computing resources and the shortest inference time, greatly reducing the complexity of applying deep learning models to the field of hyperspectral remote sensing image classification and ensuring robust classification performance in various scenarios with scarce samples, providing an important reference for the on-orbit intelligent development of spaceborne hyperspectral, fluorescence hyperspectral and other technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is the overall model architecture diagram implemented by the method of the present invention (i.e., the block diagram for implementing lightweight hyperspectral image classification based on spectral-spatial unified feedforward convolution);

[0055] Figure 2 It is a comparison diagram of the implementation processes of the traditional dimensionality reduction algorithm and the proposed dimensionality reduction algorithm;

[0056] Figure 3 It is the block diagram for implementing the feedforward convolutional transformer network;

[0057] Figure 4 It is the relationship diagram between the number of dimensionality reduction principal components and the average gradient of the image in Embodiment 1 of the present invention, where (a) is the PU dataset and (b) is the HHK dataset;

[0058] Figure 5 It is the result diagram obtained by classifying ground objects from the PU hyperspectral image data in Embodiment 2 of the present invention;

[0059] Figure 6 It is the result diagram obtained by classifying ground objects from the GF-5 HHK hyperspectral image data in Embodiment 3 of the present invention;

[0060] Figure 7 It is the speed comparison diagram between the method of the present invention and international advanced classification methods on different hyperspectral datasets. Specific Embodiment

[0061] To make the objectives, features, and advantages of the present invention clearer, a more detailed description of a specific embodiment of the present invention is given. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0062] The present invention provides a spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network. By using a cascaded hybrid model composed of two-dimensional spectral dimensionality reduction, CNN, and Transformer, it quickly and accurately realizes ground object classification. The model framework is as shown in Figure 1 . Verification experiments are completed on an airborne dataset and the spaceborne GF-5 dataset respectively, providing an important reference for realizing real-time classification of on-orbit hyperspectral images. Specifically, it includes the following steps:

[0063] Step 1 (i.e., P1): Denote the original hyperspectral three-dimensional cube data as where l is the image length, w is the image width, and b is the number of spectral bands. Each pixel point in S belongs to g land cover classes, and the class labels are represented as Y = y1, y2,..., y g , providing a standard basis for subsequent deep learning training and prediction processes;

[0064] Step 2 (i.e., P2): Compared with traditional PCA spectral dimensionality reduction and two-dimensional PCA spatial dimensionality reduction, a two-dimensional principal component analysis algorithm oriented to the spectral direction is used to perform dimensionality reduction on the hyperspectral data. This process greatly saves the running memory and retains the spatial structure information, Figure 2 clarifying the data change characteristics and memory consumption during the implementation of the three dimensionality reduction methods. This method reduces the number of bands from b to c, and the dimensionality-reduced hyperspectral data is represented as The specific process is as follows:

[0065] Step 2.1 (i.e., A1), combined with the storage and operation methods of hyperspectral data in Python, after loading the data, the lengths l, widths w, and spectral bands b of the first, second, and third dimensions are regarded as the "spectral", "length", and "width" dimensions respectively to achieve the purpose of data rotation, obtaining the new hyperspectral data

[0066] Step 2.2 (i.e., A2), then calculate the mean of the w, b plane of S2 along the length l direction and centralize it to obtain the decentralized three-dimensional image Q w×b×l :

[0067] Q = S2 - E(S2)(1)

[0068] In the formula, E(·) represents the mean value of the image.

[0069] Step 2.3 (i.e., A3), for the three-dimensional image Q w×b×l , it is considered that there are l images with a resolution of w×b. Each image is represented by q(i) (i = 1, 2, …, l). Therefore, the specific expression of E(Q) is Calculate the overall covariance matrix G of the image sample matrix r ;

[0070]

[0071] Step 2.4 (i.e., A4): Obtain the covariance matrix G by eigenvalue decomposition r The eigenvectors corresponding to the first c largest eigenvalues are used to form the optimal projection matrix W with a dimension of b×c. Its expression is:

[0072] W = [w1, w2, …, w c (3)

[0073] Project the image matrix S2 onto the obtained projection matrix W to obtain l hyperspectral data matrices with a reduced dimension of w×c in space The expression of each principal component I0(i) = [I1, I2, …, I c of the image S2 starting from the l dimension is:

[0074] I0(i) = S2W(i) (4)

[0075] Step 2.5 (i.e., A5): Re-convert the first, second, and third dimensions of I0 into the initial hyperspectral data format, that is, the length l, width w, and spectral channel c dimensions, to obtain the finally reduced-dimensional data I l×w×c .

[0076] Step 3 (i.e., P3): Each pixel s in the reduced-dimensional hyperspectral data I i,j is to be classified, where i = 1, 2, …, l and j = 1, 2, …, w. A three-dimensional image patch patch with a window size of d×d and a depth of c is extracted centered on the pixel (i, j) and stacked to form a new input space where p and q represent the length and width of the new space respectively.

[0077] Step 4 (i.e., P4): For each vector X i,j = [x 1,1,c , …, x i,j,cAs the input of the spectral-spatial information joint network, the spectral and spatial features are extracted through the 3-D CNN layer of the three-dimensional convolutional neural network and the spatial features are extracted through the 2-D DSC layer of the two-dimensional depthwise separable convolution, and the fused features are shaped into where m, n, and t are the height, width, and depth of the new feature space Z respectively. The specific implementation process is as follows:

[0078] Step 4.1 (i.e., B1): First, a 3-D CNN is used to extract spectral-spatial features from each hyperspectral data block X i,j of the new input space X. At this time, the number of input channels stride are all 1, and the padding are all 0. The size of the convolutional kernel is 8×3×3×3 (i.e., Figure 1 in ), and the output is a four-dimensional feature map Z1 with a shape of 8×(d - 2)×(d - 2)×(c - 2).

[0079] Step 4.2 (i.e., B2): After obtaining the spectral-spatial feature map through the 3-D CNN, the spectral dimension and the channel dimension are merged into to obtain a new three-dimensional feature map Z2 with a shape of 8(c - 2)×(d - 2)×(d - 2);

[0080] Step 4.3 (i.e., B3): The depthwise separable two-dimensional convolution (i.e., 2-D DSC) is used to extract features from Z2 to enhance the local spatial information of the extracted features. The two-dimensional depthwise separable convolution decomposes the standard two-dimensional convolution into two steps. The first step is the depthwise convolution, which independently performs convolution operations on each channel of the input feature map. The size of the two-dimensional convolutional kernel of the depthwise convolutional layer is designed to be 64×3×3 (i.e., Figure 1 in ), the stride and are set to 1, and the padding and are set to 0. The second step is the pointwise convolution, which uses a 1×1 convolutional kernel to perform a linear combination between channels on the output of the depthwise convolutional layer to fuse and combine the features of different channels. The shape of the output feature map Z3 after passing through the 2-D DSC layer is 64×(d - 4)×(d - 4).

[0081] Based on the above process, in the process of outputting the feature map Z1, the number of parameters and computational amount required for using 3-D CNN are 216 and 216×(d - 2)×(d - 2)×(c - 2) respectively. In the process of outputting the feature map Z3, the number of parameters and The computational amounts are 584×(c - 2) and 584×(c - 2)×(d - 4)×(d - 4) respectively. And if 3-D DSC is used to obtain the same output Z1, the number of parameters and are equal to 35 and 35×(d - 2)×(d - 2)×(c - 2) respectively. If 2-D CNN is used to obtain the same output Z3, the number of parameters and are equal to 4608×(c - 2) and 4608×(c - 2)×(d - 4)×(d - 4) respectively. Through the above comparison, it can be seen that under the condition that the number of channels in the first layer is only 1, the difference in the number of parameters and computational amounts between depthwise separable convolution and standard convolution is not obvious, and 3-D DSC will reduce the ability to capture spatio-spectral features and increase the complexity of using one convolutional layer. Therefore, 3-D CNN is adopted in the first layer here. And in the case of a relatively large number of channels in the second layer, using DSC operations can significantly reduce the number of training parameters and computational amounts. Therefore, 2-D DSC is adopted to further fuse local spatial features.

[0082] Step 5 (i.e., P5): As Figure 3 shown, take Z ∈ R m×n×t as the input of the feedforward convolutional Transformer network FCT. It is processed through linear projection, positional encoding embedding, normalization, and multi-head attention mechanism, and then residual connection is made with the input and normalized to focus on the global features of each image patch. The specific implementation process is as follows:

[0083] Step 5.1 (i.e., C1): The one-dimensional vector input x after linear projection records the position information of each image patch through positional encoding embedding, and a learnable positional encoding is added to its embedding, as shown in formula (5);

[0084]

[0085] In the formula, E is the embedding projection process, x is the vector after linear projection, and E pos represents the positional embedding.

[0086] Step 5.2 (i.e., C2): Normalize the data; more comprehensively focus on the global features among each image patch by running the multi-head attention mechanism, as shown in formula (6).

[0087] MHSA = Concat(SA1, SA2, …, SA h )(6)

[0088] In the formula, MHSA is the output vector of the multi-head attention mechanism, and Concat(·) represents the concatenation operation.

[0089] Step 5.3 (i.e., C3): Add the result of MHSA to the input x and perform normalization to obtain the output features.

[0090] Among them, each three-dimensional image patch can be mapped to a learnable one-dimensional vector through linear projection; the position encoding embedding is to record the position information of each image patch and add a learnable position encoding to its embedding, as shown in formula (5); using the Layer Normalization layer for normalization operations can accelerate the training speed and improve the training stability; the multi-head attention obtains the attention distribution of different subspaces of the input feature sequence by running multiple independent attention mechanisms in parallel, and captures more comprehensively the potential multiple semantic associations in the spectral sequence from the global perspective, as shown in formula (6).

[0091] Step 6 (i.e., P6): After passing through the Transformer global feature extraction part, its output realizes information exchange within the local area through a two-dimensional inverted residual convolution block in the feed-forward convolutional layer. This is because although the Transformer model uses two fully connected layers in the feed-forward layer to increase the nonlinearity of the model, the entire model has insufficient expression ability for local detail features. Therefore, after extracting the global features in step 5, a structure of a 1×1 convolution, a depth convolution, and a 1×1 convolution inverted residual is used in the feed-forward layer to realize information exchange within the local area, as Figure 3 shown. This structure not only avoids the high computational cost and the number of parameters brought by directly using multi-channel 2-D CNN, but also enhances the ability of information exchange within the local area. The specific implementation process of the feed-forward convolutional Transformer for information exchange within the local area through a two-dimensional inverted residual convolution block in step P6 is as follows:

[0092] Step 6.1 (i.e., D1): Reshape the normalized multi-head attention output sequence L into a two-dimensional image feature map L 2D , providing the possibility of introducing locality into the network, as shown in formula (7):

[0093] L 2D = Seqto2D(L) (7)

[0094] Step 6.2 (i.e., D2): Use a two-dimensional convolutional layer with a kernel size of 1×1 to replace the first fully connected layer in the traditional Transformer multi-layer perceptron;

[0095] Step 6.3 (i.e., D3): After passing through the first convolutional layer with a kernel size of 1×1, there is still a lack of information interaction between adjacent pixels. Therefore, a depth convolution with a kernel size of 3×3 is used to fuse the information between spatially adjacent pixels, and using depth convolution will not significantly increase the complexity of the model;

[0096] Step 6.4 (i.e., D4): After the depth convolution layer, a convolution layer with a kernel size of 1×1 is used again to replace the second fully connected layer in the traditional Transformer multi-layer perceptron, further enhancing the nonlinearity of the model. The mathematical expressions of steps 6.2 (i.e., D2) to 6.4 (i.e., D4) are as follows:

[0097]

[0098] In the formula, σ is the RELU activation function, respectively represent the first 1×1 convolution layer, the depth convolution layer, and the second 1×1 convolution layer, represents the convolution operation, and Y 2d is the output through the two-dimensional inverted residual convolution block.

[0099] Step 6.5 (i.e., D5): Add the input L 2D and the output Y 2d to obtain the final output of the feed-forward convolutional Transformer.

[0100] Step 7 (i.e., P7): The output of the above feed-forward convolutional Transformer network aggregates the features at different positions through an adaptive average pooling layer and outputs the final classification result through a fully connected layer.

[0101] The present invention uses a two-dimensional memory-efficient dimensionality reduction algorithm for feature extraction. The constructed SSUFCT model efficiently captures spectral, spatial, local, and global information through a lightweight convolutional sub-network and a feed-forward convolutional Transformer model; finally, according to the extracted high-level spectral-spatial feature map, an adaptive average pooling layer and a fully connected layer are used to obtain the classification result.

[0102] The effect of the method of the present invention can be further illustrated by the following experiments:

[0103] I. Experimental conditions and related parameter settings:

[0104] All experiments are implemented on CPU: Intel i7-12700H, GPU: NVIDIA GeForce RTX 3060, RAM: 16GB, GPU memory: 6GB, Python: 3.8.18, Pytorch: 1.12.1+cu113. Three-dimensional image patches with a size of 11×11 are used to sample the dimensionality-reduced hyperspectral data. The Adam optimizer and the cross-entropy loss function are used during training. The batch size of each input is set to 64, the learning rate lr is set to 0.001, 200 training epochs are used, and the weights with the highest accuracy on the validation set during the process are retained as the optimal weights.

[0105] To test the generalization ability of the invented method and save computational costs, we use a small number of samples for training and validation. Since the PaviaU University (PU) dataset (spatial dimensions length × width: 610 × 340, number of bands 103) has a balanced distribution of samples for each category and a large overall data volume, 3% of the data is used for training; for the GF-5 Yellow River Estuary (HHK) dataset (spatial dimensions length × width: 462 × 617, number of bands 150), due to the fragmented wetland patches, scattered vegetation distribution, and obvious changes in coverage, it is difficult to distinguish, so 15% of the data is selected for training.

[0106] In terms of evaluation metrics, three metrics, namely overall accuracy (OA), average accuracy (AA), and Kappa coefficient (KAPPA), are selected for accuracy, and the number of parameters (Parameters), computational volume (FLOPs), GPU memory occupancy (GPU memory), and prediction time are used to evaluate the lightweight performance of the model.

[0107] Comparison method: The ground truth map gt and four advanced hyperspectral image classification methods are used to compare with the method proposed in the present invention. The four comparison methods are as follows: the CNN-based classification method, from the article "Roy S K, Manna S, Song T, et al. Attention-based adaptive spectral–spatial kernel ResNet for hyperspectral image classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2020, 59(9): 7831-7843."; the Transformer-based method, from the article "Xu H, Zeng Z, Yao W, et al. CS2DT: Cross spatial–spectral dense transformer for hyperspectral image classification[J]. IEEE Geoscience and Remote Sensing Letters, 2023, 20: 1-5."; the CNN- and Transformer-based classification method, from the article "Xu R, Dong XM, Li W, et al. DBCTNet: Double branch convolution-transformer network for hyperspectral image classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024."; the classification method based on the combination of dimensionality reduction algorithm, CNN and Transformer, from the article "Liu B, Liu Y, Zhang W, et al. Fusing 3D-CNN and lightweight Swin Transformer networks for HSI[J]. 2023."

[0108] II. Comparison of Experimental Contents and Results

[0109] Example 1: On the PU and HHK datasets, the two-dimensional principal component analysis algorithm oriented to spectral direction in the present invention (i.e., the "proposed dimensionality reduction method" in Table 1) is compared with the widely used principal component analysis (PCA) and singular value decomposition (SVD) dimensionality reduction algorithms. The memory consumption of various dimensionality reduction algorithms when retaining 30 principal components is shown in Table 1, and the influence on the average gradient of the image when retaining different numbers of principal components is as Figure 4 shown.

[0110] Table 1 Comparison of memory consumption when dimensionality reduction algorithms are executed in PU and HHK datasets

[0111]

[0112] As can be seen from Table 1, the two-dimensional principal component analysis method oriented to spectral direction of the invention occupies the least memory when running on PU and HHK data, only requiring 52.9MB and 71.3MB respectively, while the SVD and PCA algorithms require 200MB to 500MB of memory space. Moreover, as the image size increases, the advantage of the proposed dimensionality reduction algorithm in small memory occupancy becomes more obvious. It should be noted that the time complexity of the proposed dimensionality reduction algorithm is similar to that of the PCA and SVD dimensionality reduction algorithms and does not take more running time.

[0113] From Figure 4 it can be seen that when retaining the same number of principal components as the SVD and PCA dimensionality reduction algorithms, the average gradient of the proposed dimensionality reduction algorithm is the highest, and the average gradient can reflect the clarity and detail contrast of the image, indicating that the two-dimensional principal component analysis method oriented to spectral direction retains more valuable information of the hyperspectral image. In addition, if the number of principal components is too small, the information retention is insufficient and the classification accuracy will be low; while too many principal components will not significantly improve the accuracy rate, but will increase the computational complexity and reduce the detection speed of the algorithm. Through experiments, the most suitable numbers of principal components for the PU and HHK datasets are 10 and 20.

[0114] Example 2: On the PU hyperspectral image data, the hyperspectral images are classified respectively using the method of the present invention and existing advanced classification methods (A2S2KNet, CS2DT, CNN-Swin, and DBCTNet), and the classification accuracy and lightweight performance evaluation indicators obtained are shown in Table 2, and the classification effect diagrams are as Figure 5 shown.

[0115] Table 2 Classification results of the PU dataset using 3% of the training samples

[0116]

[0117] As can be seen from Table 2, the OA, AA, and KAPPA accuracy values of the method of the present invention on PU data reach the international leading level. More importantly, the lightweight performance of this method is outstanding, with the lowest FLOPs and GPU memory, and the second lowest number of parameters. In addition, its whole-image prediction time is the shortest, only 6.37 s, and the running time of the fastest other prediction algorithms is about 5.1 times that of it. The speed comparison of each method is as Figure 7 shown. From Figure 5 it can be seen that the classification result map of the method of the present invention has stronger classification ability in the detail area. For example, in the area where the slender brick-paved road surface is mixed with trees in the lower left corner, other algorithms are prone to misclassify a large amount of the brick-paved road surface as asphalt road surface and trees, while this method has more accurate classification.

[0118] Example 3: On the HHK hyperspectral image data, the method of the present invention and existing advanced classification methods (A2S2KNet, CS2DT, CNN-Swin, and DBCTNet) are used to classify the hyperspectral images respectively. The classification accuracy and lightweight performance evaluation indexes obtained are shown in Table 3, and the classification effect diagrams are as Figure 6 shown.

[0119] As can be seen from Table 3, the three accuracy evaluation indexes OA, AA, and KAPPA of the method of the present invention on the HHK data reach the highest. However, the method of the invention only needs about 10 seconds to complete the whole-image prediction, while other methods need to spend 60 seconds or even hundreds of seconds. The speed comparison of each method is as Figure 7 shown. From Figure 6 it can be seen that the method of the invention classifies the low-coverage bare tidal flat in the upper left corner clearly and accurately, while other algorithms may misclassify a large area of it as water body, indicating that the method of the invention integrates local information through feedforward convolution on the basis of obtaining the global spatial-spectral information characteristics, which is beneficial to improving the classification accuracy of patchy and fragmented ground objects.

[0120] Table 3 Classification results of the HHK data set using 15% training samples

[0121]

[0122] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network, characterized in that, The two-dimensional memory-efficient dimensionality reduction algorithm is used for feature extraction. The constructed SSUFCT model realizes the efficient capture of spectral and spatial as well as local and global information through a lightweight convolutional sub-network and a feed-forward convolutional Transformer model. Finally, according to the extracted high-level spectral-spatial feature maps, an adaptive average pooling layer and a fully connected layer are used to obtain the classification result.

2. The spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network according to claim 1, characterized in that, It includes the following steps: P1: Denote the original hyperspectral three-dimensional cube data as where l is the image length, w is the image width, and b is the number of spectral bands; P2: The two-dimensional principal component analysis algorithm oriented to the spectral direction is used to perform dimensionality reduction on the hyperspectral data. This process greatly saves the running memory and retains the spatial structure information. After reducing the number of bands from b to c, the hyperspectral data after dimensionality reduction is expressed as P3: For each pixel s to be classified in the dimensionality-reduced hyperspectral data I i,j , where i = 1, 2, …, l and j = 1, 2, …, w, a three-dimensional image patch patch with a window size of d×d and a depth of c is extracted centered on the pixel (i, j), and they are stacked to form a new input space where p and q represent the length and width of the new space respectively; P4: Take each vector X i,j = [x 1,1,c , …, x i,j,c as the input of the spectral-spatial information joint network, extract spectral and spatial features through 3-D CNN layers respectively, and extract spatial features through 2-D DSC layers and shape the fused features into form, where m, n, and t are the height, width, and depth of the new feature space Z respectively; P5: Take as the input of the feed-forward convolutional Transformer network FCT, which is processed by linear projection, positional encoding embedding, normalization, and multi-head attention mechanism, and then undergoes residual connection with the input and normalization to focus on the global features of each image patch; P6: After the Transformer global feature extraction part, its output realizes information exchange within the local area through a two-dimensional inverted residual convolution block in the feed-forward convolutional layer; P7: The output of the above-mentioned feed-forward convolutional Transformer network aggregates features at different positions through an adaptive average pooling layer, and outputs the final classification result through a fully connected layer.

3. A spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network according to claim 2, characterized in that, The specific process of the two-dimensional principal component analysis algorithm oriented to the spectral direction in step P2 is as follows: A1: Combining the storage and operation methods of hyperspectral data in Python, after loading the data, the lengths l, widths w, and spectral bands b of the first, second, and third dimensions are regarded as the "spectral band", "length", and "width" dimensions respectively, to achieve the purpose of data rotation and obtain new hyperspectral data A2: Calculate the mean value of the w, b plane of S2 along the length l direction and centralize it to obtain the three-dimensional image Q w×b×l : Q = S2 - E(S2) (1) In the formula, E(·) represents the mean value of the image; A3: For the three-dimensional image Q w×b×l , it is considered that there are l images with a resolution of w×b, and each image is represented by q(i) (i = 1, 2,..., l). Therefore, the specific expression of E(Q) is Calculate the overall covariance matrix G of the image sample matrix r ; A4: Obtain the covariance matrix G by eigenvalue decomposition r The eigenvectors corresponding to the first c largest eigenvalues are used to form the optimal projection matrix W with dimensions b×c, and its expression is: W = [w1, w2, …, w c (3) Project the image matrix S2 onto the obtained projection matrix W to obtain l hyperspectral data matrices after dimensionality reduction with a spatial size of w×c. The expression of each principal component I0(i) = [I1, I2, …, I c starting from the l dimension is: I0(i) = S2W(i) (4) A5: Re-convert the first, second, and third dimensions of I0 into the initial hyperspectral initial data format, i.e., the length l, width w, and spectral channel c dimensions, to obtain the finally dimension-reduced data I l×w×c .

4. A spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network according to claim 2, characterized in that The specific process of the spectral-spatial information joint network extracting features in step P4 is as follows: B1: First, use a 3-D CNN layer to extract spectral-spatial features from each hyperspectral data block X in the new input space X i,j to obtain a four-dimensional feature map Z1; B2: After obtaining the spectral-spatial feature map through 3-D CNN, the spectral dimension and the channel dimension are merged into to obtain the new three-dimensional feature map Z2, B3: Use a 2-D DSC layer to extract features from Z2 to further enhance the local spatial information of the extracted features.

5. A spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network according to claim 2, characterized in that The specific process of paying attention to the global features of each input picture block in step P5 is as follows: C1: The one-dimensional vector input x after linear projection is embedded through position encoding to record the position information of each image block, and a learnable position encoding is added to its embedding, as shown in formula (5); In the formula, E is the embedding projection process, x is the vector after linear projection, and E pos represents the positional embedding; C2: Perform normalization processing on the data; more comprehensively pay attention to the global features of each picture block by running the multi-head attention mechanism, as shown in formula (6); MHSA = Concat(SA1, SA2, …, SA h )(6) In the formula, MHSA is the output vector of the multi-head attention mechanism, and Concat(·) represents the concatenation operation; C3: Add the result of MHSA to the input x and perform normalization to obtain the output features.

6. The spaceborne hyperspectral image classification method based on a lightweight spectral-spatial joint feedforward convolutional Transformer network according to claim 2, wherein The specific process of the feed-forward convolutional Transformer performing information exchange within the local area through a two-dimensional inverted residual convolution block in step P6 is as follows: D1: Reshape the normalized multi-head attention output sequence L into a two-dimensional image feature map L through Seqto2D, which provides the possibility of introducing locality into the network. Its mathematical expression is: 2D , which provides the possibility of introducing locality into the network. Its mathematical expression is: L 2D = Seqto2D(L)(7) D2: Use a convolutional layer with a kernel size of 1×1 to replace the first fully connected layer in the multi-layer perceptron of the traditional Transformer model; D3: After the first convolutional layer, use a depth convolution with a kernel size of 3×3 to fuse the information between spatially adjacent pixels; D4: After the depth convolutional layer, use a convolutional layer with a kernel size of 1×1 again to replace the second fully connected layer in the multi-layer perceptron. The mathematical expressions of steps D2 to D4 are as follows: where σ is the RELU activation function, represent the first 1×1 convolutional layer, the depth convolutional layer, and the second 1×1 convolutional layer respectively, represents the convolution operation, and Y 2d is the output through the two-dimensional inverted residual convolution block; D5: Add the input L 2D to the output Y 2d to obtain the final output of the feed-forward convolutional Transformer.