Hyperspectral remote sensing image classification method, system and device based on deep learning
Through the deep learning method of generating spectral-space tokens and multi-scale wavelet decomposition, the problem of insufficient spectral-space feature extraction in hyperspectral remote sensing image classification is solved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202510521209.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Traditional methods are difficult to effectively process three-dimensional data of hyperspectral remote sensing images, and cannot effectively extract spectral-spatial features, resulting in insufficient classification accuracy.
A deep learning-based method is adopted to generate a spectral-space token for feature extraction, combined with multi-scale wavelet decomposition and attention mechanism, adaptive fusion of low-frequency and high-frequency features is carried out, and classification scores are finally generated for classification.
The classification accuracy of hyperspectral remote sensing images is improved, especially on small-scale data sets.
Smart Images

Figure CN120451782A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing, and specifically relates to a hyperspectral remote sensing image classification method, system, and device based on deep learning. Background Art
[0002] Hyperspectral remote sensing image classification is an important research area in remote sensing, with broad application prospects. Unlike conventional RGB images, which contain only three bands: red, green, and blue, hyperspectral remote sensing images typically cover 200 bands ranging from 400 to 1000 nanometers. Hyperspectral data possesses both spatial resolution and spectral dimensionality, forming a three-dimensional data cube that contains rich spectral information at every pixel. Due to the high dimensionality, spectral-spatial correlation, and computational complexity of hyperspectral data, traditional monochrome, RGB, and multispectral image analysis techniques cannot be directly applied to hyperspectral remote sensing image classification. Instead, specialized adaptation methods are required to effectively extract spectral-spatial features. Summary of the Invention
[0003] The present invention provides a hyperspectral remote sensing image classification method, system, and device based on deep learning.
[0004] The technical solutions of the present invention are as follows:
[0005] The present invention provides a hyperspectral remote sensing image classification method based on deep learning, comprising:
[0006] S1: Acquire hyperspectral remote sensing image data, perform dimensionality conversion after 3D convolution processing, and generate spectral-spatial tokens. The spectral-spatial tokens are used to extract features through 3D convolution with different expansion rates. The extracted features are sequentially spliced and fused to obtain fused features.
[0007] S2: Based on the wavelet basis type, the fusion features are decomposed into two-dimensional wavelets to obtain the low-frequency features and high-frequency features of the corresponding wavelet basis type; using the learnable fusion weights, the low-frequency features and high-frequency features of different wavelet basis types are weighted summed to obtain low-frequency fusion features and high-frequency fusion features;
[0008] The low-frequency fusion features and high-frequency fusion features are processed by channel attention and spatial attention respectively to obtain low-frequency enhancement features and high-frequency enhancement features;
[0009] S3: The low-frequency enhanced feature is scanned along the spectral dimension to obtain a first scanning feature, and the first scanning feature is inverted to obtain a second scanning feature;
[0010] The high-frequency enhanced feature is scanned along the spatial dimension to obtain a third scanning feature, and the third scanning feature is inverted to obtain a fourth scanning feature;
[0011] Using learnable parameters, the first scanning feature and the second scanning feature, the third scanning feature and the fourth scanning feature are fused respectively to obtain a low-frequency scanning feature and a high-frequency scanning feature;
[0012] The high-frequency scanning features are processed by the dynamic gating mechanism and fused with the low-frequency scanning features to obtain the target features;
[0013] S4: After global average pooling and linear processing, the target features generate classification scores, and classification is performed based on the classification scores.
[0014] In S1, the spectral-spatial token extracts features through three-dimensional convolution with different expansion rates. The extracted features are sequentially spliced and fused to obtain fused features, specifically:
[0015] According to the formula: F s =σ(BN(Conv3D s (T))),s∈{1,2,4}, extract features;
[0016] Where, F s is the extracted feature, σ is the activation function, BN represents normalization, Conv3D s represents the three-dimensional convolution with different dilation rates, T represents the spectral-spatial token, and s is the dilation rate;
[0017] After the extracted features are spliced along the channel dimension, the spliced features are fused through convolution to obtain the fused features.
[0018] Said S2, based on the wavelet basis type, performs two-dimensional wavelet decomposition by integrating features to obtain low-frequency features and high-frequency features of the corresponding wavelet basis type, specifically:
[0019] According to the formula: cA i =X*W low , get the low-frequency features of the corresponding wavelet basis type;
[0020] According to the formula: cH i =X*W high-H , cV i =X*W high-V , cD i =X*W high-D ,and Obtain high-frequency features of the corresponding wavelet basis type;
[0021] Where cA i 、high_freq i Represent low-frequency features and high-frequency features respectively; X is the fusion feature; cH i 、cV i 、cD iRespectively represent horizontal high-frequency features, vertical high-frequency features, and diagonal high-frequency features; W low 、W high-H 、W high-V 、W high-D They represent low frequency, horizontal high frequency, vertical high frequency, and diagonal high frequency filters respectively.
[0022] After the fusion features are decomposed into two-dimensional wavelet, bilinear interpolation processing is also included.
[0023] The S2 uses the learnable fusion weights to perform weighted summation on the low-frequency features and high-frequency features of different wavelet basis types to obtain low-frequency fusion features and high-frequency fusion features, specifically:
[0024] According to the formula: accomplish;
[0025] Where, is the low-frequency fusion feature, is the low-frequency feature of wavelet basis type i, α i is the learnable fusion weight, M is the number of wavelet basis types, is the high-frequency fusion feature, is the high-frequency feature of wavelet basis type i.
[0026] The S3 uses learnable parameters to fuse the first scanning feature, the second scanning feature, the third scanning feature, and the fourth scanning feature to obtain low-frequency scanning features and high-frequency scanning features, specifically:
[0027] By the formula: Get low-frequency scanning features;
[0028] By the formula: Get high frequency sweep characteristics;
[0029] Where Y low 、Y high are low-frequency scanning features and high-frequency scanning features respectively; λ1, λ2, λ3, and λ4 represent the normalized weights; X low,spe 、 X high,spa 、 They represent the first scanning feature, the second scanning feature, the third scanning feature, and the fourth scanning feature respectively.
[0030] The present invention also provides a hyperspectral remote sensing image classification system based on deep learning, comprising:
[0031] Feature extraction module: used to obtain hyperspectral remote sensing image data, perform dimensionality conversion after 3D convolution processing, and generate spectral-spatial tokens. The spectral-spatial tokens are used to extract features through 3D convolution with different expansion rates. The extracted features are sequentially spliced and fused to obtain fused features.
[0032] Feature enhancement module: Based on the wavelet basis type, the fusion features are decomposed into two-dimensional wavelets to obtain the low-frequency features and high-frequency features of the corresponding wavelet basis type; using the learnable fusion weights, the low-frequency features and high-frequency features of different wavelet basis types are weighted summed to obtain low-frequency fusion features and high-frequency fusion features;
[0033] The low-frequency fusion features and high-frequency fusion features are processed by channel attention and spatial attention respectively to obtain low-frequency enhancement features and high-frequency enhancement features;
[0034] Scanning module: The low-frequency enhanced feature is scanned along the spectral dimension to obtain a first scanning feature, and the first scanning feature is inverted to obtain a second scanning feature;
[0035] The high-frequency enhanced feature is scanned along the spatial dimension to obtain a third scanning feature, and the third scanning feature is inverted to obtain a fourth scanning feature;
[0036] Using learnable parameters, the first scanning feature and the second scanning feature, the third scanning feature and the fourth scanning feature are fused respectively to obtain a low-frequency scanning feature and a high-frequency scanning feature;
[0037] The high-frequency scanning features are processed by the dynamic gating mechanism and fused with the low-frequency scanning features to obtain the target features;
[0038] Classification module: After the target features are globally averaged and linearly processed, a classification score is generated, and classification is performed based on the classification score.
[0039] The present invention also provides a hyperspectral remote sensing image classification device based on deep learning, comprising a processor and a memory, wherein the processor implements the hyperspectral remote sensing image classification method based on deep learning when executing a computer program stored in the memory.
[0040] Beneficial effects
[0041] The hyperspectral remote sensing image classification method based on deep learning of the present invention converts three-dimensional data into one dimension and performs preliminary feature extraction by generating spectral-spatial tokens; the features to be processed are decomposed into low-frequency features and high-frequency features through multi-scale wavelet decomposition, and then subjected to channel attention processing and spatial attention processing respectively to enhance the low-frequency and high-frequency features; the enhanced low-frequency features are expanded along the spectral dimension, and the enhanced high-frequency features are expanded along the spatial dimension, and dynamic path weights are introduced in the selective state space scanning mechanism to realize adaptive feature fusion of low-frequency features and high-frequency features; finally, the classification result is obtained according to the classification score. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 The following are the classification results of different hyperspectral remote sensing image classification methods on the Pavia University dataset, where: (a) is the false color image of the dataset; (b) is the ground truth label image of the dataset; (c) is the classification result of the 3D CNN method; (d) is the classification result of the SPRN method; (e) is the classification result of the CEGCN method; (f) is the classification result of the SSFTT method; (g) is the classification result of the MorphFormer method; (h) is the classification result of the SS-Mamba method; (i) is the classification result of the 3DSS-Mamba method; and (j) is the classification result of the method of the present invention.
[0043] Figure 2 The following are the classification results of different hyperspectral remote sensing image classification methods on the Indian Pines dataset, where: (a) is the false color image of the dataset; (b) is the ground truth label image of the dataset; (c) is the classification result of the 3D CNN method; (d) is the classification result of the SPRN method; (e) is the classification result of the CEGCN method; (f) is the classification result of the SSFTT method; (g) is the classification result of the MorphFormer method; (h) is the classification result of the SS-Mamba method; (i) is the classification result of the 3DSS-Mamba method; and (j) is the classification result of the method of the present invention.
[0044] Figure 3The following are the classification results of different hyperspectral remote sensing image classification methods on the Houston2013 dataset, where: (a) is the false color image of the dataset; (b) is the ground truth label image of the dataset; (c) is the classification result of the 3D CNN method; (d) is the classification result of the SPRN method; (e) is the classification result of the CEGCN method; (f) is the classification result of the SSFTT method; (g) is the classification result of the MorphFormer method; (h) is the classification result of the SS-Mamba method; (i) is the classification result of the 3DSS-Mamba method; and (j) is the classification result of the method of the present invention. DETAILED DESCRIPTION
[0045] The following examples are intended to illustrate the present invention rather than to further limit the present invention.
[0046] The present invention provides a hyperspectral remote sensing image classification method based on deep learning, comprising:
[0047] S1: Acquire hyperspectral remote sensing image data, perform dimensionality conversion after 3D convolution processing, and generate spectral-spatial tokens. The spectral-spatial tokens extract features through 3D convolution with different expansion rates. The extracted features are sequentially spliced and fused to obtain fused features.
[0048] The specific operations are as follows:
[0049] First, the hyperspectral remote sensing image data (3D cube data) is subjected to dimensionality reduction and divided into multiple pixel-level 3D patch cubes, denoted as x, x∈R B×B×d , where B is the size of the spatial dimensions (height and width) and d is the number of bands.
[0050] Then, it undergoes three-dimensional convolution processing, which includes a three-dimensional convolution layer, a normalization layer, and a ReLU activation function.
[0051] Then, the dimension is transformed through a linear layer of embedding operation to generate spectral-spatial tokens. The tokenization process can be expressed as:
[0052] T=Φ embed (φ 3DConv (x));
[0053] Where, φ embed represents the embedding operation; φ 3DConv represents a three-dimensional convolution process; T represents a spectral-spatial token; and T∈R N×P×P×K , N is the number of generated markers, each marker is a P×P two-dimensional spatial feature map. K is the feature dimension of each marker, and each marker has K eigenvalues at each spatial position, which capture the spectral and spatial information of the local area.
[0054] Afterwards, multi-scale features are extracted through three-dimensional convolution with different expansion rates to capture local and global information. Normalization is used to accelerate training and stabilize gradients, while activation functions are used to enhance the nonlinear expression ability of features. The formula for this process is:
[0055] F s =σ(BN(Conv3D s (T)),s∈{1,2,4};
[0056] Where, F s is the extracted feature, σ is the activation function, BN represents normalization, Conv3D s represents the three-dimensional convolution with different dilation rates, T represents the spectral-spatial token, and s is the dilation rate;
[0057] After the extracted features of different scales are spliced along the channel dimension, the spliced features are fused using convolution with a kernel of 1 to obtain the fused features.
[0058] S2: Based on the wavelet basis type, the fusion features are decomposed into two-dimensional wavelets to obtain the low-frequency features and high-frequency features of the corresponding wavelet basis type; using the learnable fusion weights, the low-frequency features and high-frequency features of different wavelet basis types are weighted summed to obtain low-frequency fusion features and high-frequency fusion features;
[0059] The low-frequency fusion features and high-frequency fusion features are processed by channel attention and spatial attention respectively to obtain low-frequency enhancement features and high-frequency enhancement features.
[0060] The specific operations are as follows:
[0061] First, three wavelet bases are defined: Haar, Daubechies-2, and Symlet-2. Based on the wavelet base type, the fusion features are used to perform two-dimensional wavelet decomposition to obtain the low-frequency and high-frequency features of the corresponding wavelet base type, specifically:
[0062] According to the formula: cA i =X*W low , get the low-frequency features of the corresponding wavelet basis type;
[0063] According to the formula: cH i =X*W high-H , cV i =X*W high-V , cD i =X*W high-D ,and Obtain high-frequency features of the corresponding wavelet basis type;
[0064] Where cA i、high_freq i Represent low-frequency features and high-frequency features respectively; X is the fusion feature; cH i 、cV i 、cD i Respectively represent horizontal high-frequency features, vertical high-frequency features, and diagonal high-frequency features; W low 、W high-H 、W high-V 、W high-D They represent low frequency, horizontal high frequency, vertical high frequency, and diagonal high frequency filters respectively.
[0065] Then, since wavelet decomposition will reduce the spatial size, the fusion features are decomposed into two-dimensional wavelet and then bilinear interpolation is performed to restore the original size. The formula for this process is:
[0066]
[0067] Where Upscale represents bilinear interpolation processing; H and W represent the height and width before two-dimensional wavelet decomposition respectively; cA i 、high_freq i Represent low-frequency features and high-frequency features respectively; They represent the low-frequency and high-frequency features of wavelet basis type i respectively.
[0068] Then, the decomposition results of different wavelets contribute differently to the final feature extraction, so the learnable fusion weight α is used i Perform a weighted sum on them. The formula for this process is:
[0069]
[0070] Where, is the low-frequency fusion feature, is the low-frequency feature of wavelet basis type i, α i is the learnable fusion weight, M is the number of wavelet basis types, is the high-frequency fusion feature, is the high-frequency feature of wavelet basis type i.
[0071] Afterwards, the low-frequency fusion features and high-frequency fusion features are processed by channel attention and spatial attention respectively to obtain low-frequency enhanced features and high-frequency enhanced features.
[0072] For example, for low-frequency fusion features, global average pooling and global maximum pooling are first used to extract global information in the channel dimension and calculate channel attention weights. Subsequently, the features are further enhanced through depthwise separable convolution operations, and finally the enhanced features are weighted fused in combination with the attention weights.
[0073] For high-frequency fusion features, spatial attention information is first extracted through channel average pooling and channel maximum pooling to capture key spatial structural features. Subsequently, the extracted high-frequency features are enhanced using depthwise separable convolution operations and weighted fused with the enhanced features using spatial attention weights.
[0074] Step S2 uses wavelet decomposition to perform multi-scale analysis on the hyperspectral remote sensing image, which can effectively capture the local and global features of the hyperspectral remote sensing image and enhance the expression ability of spectral-spatial information.
[0075] S3: The low-frequency enhanced feature is scanned along the spectral dimension to obtain a first scanning feature, and the first scanning feature is inverted to obtain a second scanning feature;
[0076] The high-frequency enhanced feature is scanned along the spatial dimension to obtain a third scanning feature, and the third scanning feature is inverted to obtain a fourth scanning feature;
[0077] Using learnable parameters, the first scanning feature and the second scanning feature, the third scanning feature and the fourth scanning feature are fused respectively to obtain a low-frequency scanning feature and a high-frequency scanning feature;
[0078] The high-frequency scanning features are processed by the dynamic gating mechanism and fused with the low-frequency scanning features to obtain the target features.
[0079] The specific operations are as follows:
[0080] First, the low-frequency enhancement features and high-frequency enhancement features are scanned using two methods: spectral priority (along the spectral dimension) and spatial priority (along the spatial dimension). Through the reshaping operation, each feature is expanded and rearranged in different dimensions, that is, four scanning paths. Specifically, the low-frequency enhancement features are expanded along the spectral dimension to obtain the first scanning feature X low,spe , the first scanning feature is reversed to obtain the second scanning feature
[0081] The high-frequency enhanced features are expanded along the spatial dimension to obtain the third scanning feature X high,spa , the third scanning feature is reversed to obtain the fourth scanning feature
[0082] X low,spe 、X high,spa After the above inversion operation, the direction perception ability is enhanced. Four scanning features are obtained to obtain a complete feature expression X scan ,
[0083] Then, a learnable parameter is set and the weights of the four scanning paths are calculated through Softmax normalization. The process of obtaining the low-frequency scanning features can be expressed as:
[0084]
[0085] Similarly, the process of obtaining high-frequency scanning features can be expressed as:
[0086]
[0087] Where Y low 、Y high are low-frequency scanning features and high-frequency scanning features respectively; λ1, λ2, λ3, and λ4 represent the normalized weights, where λ1 and λ2 correspond to the low-frequency scanning modes along the spectral dimension and its reverse direction, and λ3 and λ4 correspond to the high-frequency scanning modes along the spatial dimension and its reverse direction, respectively; X low,spe 、 X high,spa 、 They represent the first scanning feature, the second scanning feature, the third scanning feature, and the fourth scanning feature respectively.
[0088] Since low-frequency scanning features provide global structure and ensure spectral consistency, while high-frequency scanning features provide local details and make boundaries clearer, a dynamic gating mechanism is used to calculate the attention weights of high-frequency scanning features, so that high-frequency scanning features in different regions contribute differently to the final result. Finally, the low-frequency scanning features are fused with the weighted high-frequency scanning features. The process can be expressed as:
[0089] Y final =Y low +W high ·Y high ;
[0090] Where W high Calculated by the high-frequency gating mechanism, it ensures that the high-frequency sweep features of different regions can adaptively affect the final result, thereby improving the feature expression capability.
[0091] Step S3 not only adopts spectral priority (along the spectral dimension) and spatial priority (along the spatial dimension) scanning methods for low-frequency enhancement features and high-frequency enhancement features respectively, but also introduces dynamic path weights in the selective state space scanning mechanism to achieve adaptive feature fusion of low-frequency features and high-frequency features, and can perform efficient information dissemination, realize deep feature fusion across bands, and improve classification accuracy.
[0092] S4: After global average pooling and linear processing, the target features generate classification scores, and classification is performed based on the classification scores.
[0093] The specific process includes:
[0094] First, global average pooling is performed on the target features to reduce feature dimensionality while preserving global information. Next, a linear layer is used as the classification head. The pooled features are projected onto the num_classes dimension through the linear layer to generate a classification score. num_classes represents the number of classes in the classification task, such as 9 or 16 feature types. Classification is performed based on the classification score.
[0095] The hyperspectral remote sensing image classification method based on deep learning of the present invention converts three-dimensional data into one dimension and performs preliminary feature extraction by generating spectral-spatial tokens; the features to be processed are decomposed into low-frequency features and high-frequency features through multi-scale wavelet decomposition, and then subjected to channel attention processing and spatial attention processing respectively to enhance the low-frequency and high-frequency features; the enhanced low-frequency features are expanded along the spectral dimension, and the enhanced high-frequency features are expanded along the spatial dimension, and dynamic path weights are introduced in the selective state space scanning mechanism to realize adaptive feature fusion of low-frequency features and high-frequency features; finally, the classification result is obtained according to the classification score.
[0096] Experimental results
[0097] Different hyperspectral remote sensing image classification methods were used to perform classification on the Pavia University, Indian Pines, and Houston 2013 datasets. The classification methods included 3D CNN, SPRN, CEGCN, SSFTT, MorphFormer, SS-Mamba, 3DSS-Mamba, and the method of the present invention.
[0098] Among them, 3D CNN is a classic hyperspectral remote sensing image classification method that has significant advantages in datasets with large amounts of annotations. However, it is difficult to effectively extract spectral and spatial features in small datasets.
[0099] SPRN is a classification method based on convolutional neural networks and spectral segmentation. It has a high degree of extraction of spectral features but a low degree of extraction of spatial features, resulting in poor classification results.
[0100] CEGCN combines graph convolutional networks and reinforcement learning. It constructs a spectral-spatial feature cube and uses graph convolution modules to integrate the correlation information between samples.
[0101] SSFTT combines CNN and Transformer. The convolutional layer extracts low-level spectral-spatial features and converts them into semantic tokens. The Transformer encoder models its high-level semantic information.
[0102] MorphFormer combines morphological operations with the Transformer architecture, and performs better than traditional Transformer and CNN models in complex scenarios.
[0103] SS-Mamba mainly consists of a spectral-spatial token generation module and multiple superimposed spectral-spatial Mmaba blocks. Although it has lower computational complexity than the Transformer model, its overall classification effect is weaker;
[0104] 3DSS-Mamba captures the global spatial-spectral correlation of hyperspectral remote sensing images through a multi-directional scanning strategy, but it is weak in processing local features and has poor classification effect on small-scale datasets.
[0105] Depend on Figure 1 、 Figure 2 、 Figure 3 It can be seen that the classification method of the present invention has a higher classification accuracy than other classification methods. Figure 2 Shown performs better.
[0106] The present invention also provides a hyperspectral remote sensing image classification system based on deep learning, comprising:
[0107] Feature extraction module: used to obtain hyperspectral remote sensing image data, perform dimensionality conversion after 3D convolution processing, and generate spectral-spatial tokens. The spectral-spatial tokens are used to extract features through 3D convolution with different expansion rates. The extracted features are sequentially spliced and fused to obtain fused features.
[0108] Feature enhancement module: Based on the wavelet basis type, the fusion features are decomposed into two-dimensional wavelets to obtain the low-frequency features and high-frequency features of the corresponding wavelet basis type; using the learnable fusion weights, the low-frequency features and high-frequency features of different wavelet basis types are weighted summed to obtain low-frequency fusion features and high-frequency fusion features;
[0109] The low-frequency fusion features and high-frequency fusion features are processed by channel attention and spatial attention respectively to obtain low-frequency enhancement features and high-frequency enhancement features;
[0110] Scanning module: The low-frequency enhanced feature is scanned along the spectral dimension to obtain the first scanning feature and the second scanning feature;
[0111] The high-frequency enhanced feature is scanned along the spatial dimension to obtain a third scanning feature and a fourth scanning feature;
[0112] Using learnable parameters, the first scanning feature and the second scanning feature, the third scanning feature and the fourth scanning feature are fused respectively to obtain a low-frequency scanning feature and a high-frequency scanning feature;
[0113] The high-frequency scanning features are processed by the dynamic gating mechanism and fused with the low-frequency scanning features to obtain the target features;
[0114] Classification module: After the target features are globally averaged and linearly processed, a classification score is generated, and classification is performed based on the classification score.
[0115] The present invention also provides a hyperspectral remote sensing image classification device based on deep learning, comprising a processor and a memory, wherein the processor implements the hyperspectral remote sensing image classification method based on deep learning when executing a computer program stored in the memory.
Claims
1. A hyperspectral remote sensing image classification method based on deep learning, characterized in that: include: S1: Acquire hyperspectral remote sensing image data, perform dimensionality conversion after 3D convolution processing, and generate spectral-spatial tokens. The spectral-spatial tokens are used to extract features through 3D convolution with different expansion rates. The extracted features are sequentially spliced and fused to obtain fused features. S2: Based on the wavelet basis type, the fusion features are used to perform two-dimensional wavelet decomposition to obtain the low-frequency features and high-frequency features of the corresponding wavelet basis type; Using learnable fusion weights, the low-frequency features and high-frequency features of different wavelet basis types are weighted and summed to obtain low-frequency fusion features and high-frequency fusion features; The low-frequency fusion features and high-frequency fusion features are processed by channel attention and spatial attention respectively to obtain low-frequency enhancement features and high-frequency enhancement features; S3: The low-frequency enhanced feature is scanned along the spectral dimension to obtain a first scanning feature, and the first scanning feature is inverted to obtain a second scanning feature; The high-frequency enhanced feature is scanned along the spatial dimension to obtain a third scanning feature, and the third scanning feature is inverted to obtain a fourth scanning feature; Using learnable parameters, the first scanning feature and the second scanning feature, the third scanning feature and the fourth scanning feature are fused respectively to obtain a low-frequency scanning feature and a high-frequency scanning feature; The high-frequency scanning features are processed by the dynamic gating mechanism and fused with the low-frequency scanning features to obtain the target features; S4: After global average pooling and linear processing, the target features generate classification scores, and classification is performed based on the classification scores.
2. The hyperspectral remote sensing image classification method based on deep learning according to claim 1, characterized in that: In S1, the spectral-spatial token extracts features through three-dimensional convolution with different expansion rates. The extracted features are sequentially spliced and fused to obtain fused features, specifically: According to the formula: F s =σ(BN(Conv3D s (T))),s∈{1,2,4}, extract features; Where, F s is the extracted feature, σ is the activation function, BN represents normalization, Conv3D s represents the three-dimensional convolution with different dilation rates, T represents the spectral-spatial token, and s is the dilation rate; After the extracted features are spliced along the channel dimension, the spliced features are fused through convolution to obtain the fused features.
3. The hyperspectral remote sensing image classification method based on deep learning according to claim 1, characterized in that: Said S2, based on the wavelet basis type, performs two-dimensional wavelet decomposition by integrating features to obtain low-frequency features and high-frequency features of the corresponding wavelet basis type, specifically: According to the formula: cA i =X*W low , get the low-frequency features of the corresponding wavelet basis type; According to the formula: cH i =X*W high-H , cV i =X*W high-V , cD i =X*W high-D ,and Obtain high-frequency features of the corresponding wavelet basis type; Where cA i 、high_freq i Represent low-frequency features and high-frequency features respectively; X is the fusion feature; cH i 、cV i 、cD i Respectively represent horizontal high-frequency features, vertical high-frequency features, and diagonal high-frequency features; W low 、W high-H 、W high-V 、W high-D They represent low frequency, horizontal high frequency, vertical high frequency, and diagonal high frequency filters respectively.
4. The hyperspectral remote sensing image classification method based on deep learning according to claim 1, characterized in that: The S2 uses the learnable fusion weights to perform weighted summation on the low-frequency features and high-frequency features of different wavelet basis types to obtain low-frequency fusion features and high-frequency fusion features, specifically: According to the formula: accomplish; Where, is the low-frequency fusion feature, is the low-frequency feature of wavelet basis type i, α i is the learnable fusion weight, M is the number of wavelet basis types, is the high-frequency fusion feature, is the high-frequency feature of wavelet basis type i.
5. The hyperspectral remote sensing image classification method based on deep learning according to claim 1, characterized in that: The S3 uses learnable parameters to fuse the first scanning feature, the second scanning feature, the third scanning feature, and the fourth scanning feature to obtain low-frequency scanning features and high-frequency scanning features, specifically: By the formula: Get low-frequency scanning features; By the formula: Get high frequency sweep characteristics; Where Y low 、Y high are low-frequency scanning features and high-frequency scanning features respectively; λ1, λ2, λ3, and λ4 represent the normalized weights; X low,spe 、 X high,spa 、 They represent the first scanning feature, the second scanning feature, the third scanning feature, and the fourth scanning feature respectively.
6. The hyperspectral remote sensing image classification method based on deep learning according to claim 3, characterized in that: After the fusion features are decomposed into two-dimensional wavelet, bilinear interpolation processing is also included.
7. A hyperspectral remote sensing image classification system based on deep learning, characterized in that: include: Feature extraction module: used to obtain hyperspectral remote sensing image data, perform dimensionality conversion after 3D convolution processing, and generate spectral-spatial tokens. The spectral-spatial tokens are used to extract features through 3D convolution with different expansion rates. The extracted features are sequentially spliced and fused to obtain fused features. Feature enhancement module: Based on the wavelet basis type, the fusion features are decomposed into two-dimensional wavelets to obtain the low-frequency features and high-frequency features of the corresponding wavelet basis type; Using learnable fusion weights, the low-frequency features and high-frequency features of different wavelet basis types are weighted and summed to obtain low-frequency fusion features and high-frequency fusion features; The low-frequency fusion features and high-frequency fusion features are processed by channel attention and spatial attention respectively to obtain low-frequency enhancement features and high-frequency enhancement features; Scanning module: The low-frequency enhanced feature is scanned along the spectral dimension to obtain a first scanning feature, and the first scanning feature is inverted to obtain a second scanning feature; The high-frequency enhanced feature is scanned along the spatial dimension to obtain a third scanning feature, and the third scanning feature is inverted to obtain a fourth scanning feature; Using learnable parameters, the first scanning feature and the second scanning feature, the third scanning feature and the fourth scanning feature are fused respectively to obtain a low-frequency scanning feature and a high-frequency scanning feature; The high-frequency scanning features are processed by the dynamic gating mechanism and fused with the low-frequency scanning features to obtain the target features; Classification module: After the target features are globally averaged and linearly processed, a classification score is generated, and classification is performed based on the classification score.
8. A hyperspectral remote sensing image classification device based on deep learning, characterized in that: The method comprises a processor and a memory, wherein when the processor executes the computer program stored in the memory, the method for hyperspectral remote sensing image classification based on deep learning as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Drainage pipeline defect detection method and device capable of reducing false detection rate, equipment and medium
CN113129300A
Transform-based hyperspectral image classification method and system
CN117788889A
Image classification method and image classification system based on sparse feature fusion
CN117994579A
Multi-focus image fusion method based on multilayer semantics and multi-scale self-attention
CN119251062A