LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method
The LiDAR-guided multimodal fusion dynamic hyperspectral image band selection method utilizes wavelet deformable convolution and multi-scale feature alignment fusion modules to solve the problem of information loss in multimodal data fusion in existing methods. It achieves effective fusion of HSI and LiDAR data and dynamic band selection, improving classification performance and information extraction efficiency.
Patent Information
- Application Number
- CN202511744467.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-03-10
AI Technical Summary
Most existing band selection methods are designed for single-mode data, ignoring the potential advantages of elevation information in multi-mode data, which limits the comprehensiveness and accuracy of information mining.
We propose a LiDAR-guided multimodal fusion dynamic hyperspectral image band selection method. By constructing a dual-branch multimodal cross-feature fusion and dynamic band selection strategy model, and utilizing wavelet deformable convolution, multi-scale feature alignment fusion module and bidirectional cross-attention fusion module, we achieve effective fusion of HSI and LiDAR data and dynamic band selection.
It effectively captures multi-scale features, avoids the loss of key details, and achieves adaptive fusion of HSI and LiDAR data, thereby improving classification performance and information extraction efficiency.
Smart Images

Figure CN121640267A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a multi-modal fusion dynamic hyperspectral image band selection method. BACKGROUND
[0002] Remote sensing technology is a comprehensive technology for earth observation developed in the 1960s, which is a technology that uses the inherent characteristics of electromagnetic waves reflected or radiated by objects to identify, measure and analyze the properties of target objects without direct contact with the objects. It can reveal the spatial distribution characteristics and spatio-temporal variation rules of various elements on the earth's surface at the global level. Hyperspectral image (HSI) and LiDAR data are two important data types in the field of remote sensing, each with unique advantages, and when used together, they can provide more comprehensive information. HSI can capture continuous spectral information of target objects, covering multiple bands from visible light to infrared, providing strong support for land cover classification, vegetation health monitoring and environmental change analysis.In recent years, HSI has been widely applied in various fields, such as environmental monitoring [3] (Rajabi R, Zehtabian A, Singh KD, Tabatabaeenejad A, Ghamisi P and Homayouni S (2024) Editorial: Hyperspectral imaging in environmental monitoring and analysis. Front. Environ. Sci. 11: 1353447. doi: 10.3389 / fenvs.2023.1353447.), fine-grained classification [4] (J. Yuan, S. Wang, C. Wu and Y. Xu, "Fine-Grained Classification of Urban Functional Zones and Landscape Pattern Analysis Using Hyperspectral Satellite Imagery: A Case Study of Wuhan," in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 3972-3991, 2022.) and resource exploration [5] (X. Dong, F. Gan, N. Li, S. Zhang and T. Li, "Mineral mapping in the Duolong porphyry and epithermal ore district Tibet using the Gaofen-5 satellite hyperspectral remote sensing data", Ore Geol. Rev., vol. 151, 2022.), etc. In the context of land cover classification, HSI can accurately perceive and identify land cover [6] (C. Shi, D. Liao, T. Zhang and L. Wang, "Hyperspectral Image Classification Based on Expansion Convolution Network," in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-16, 2022.) due to its rich spectral bands.
[0003] However, HSI usually contains a large number of bands, which makes the data volume huge and brings great challenges to subsequent processing and analysis, such as high computational cost and large storage requirement[7](Hu, T; Guo, X; Gao, P. Hyperspectral Band Selection Method Based on Global Partition Clustering. Remote Sens. 2025, 17, 435. https: / / doi.org / 10.3390 / rs17030435.). Therefore, band selection becomes a key link in the processing of hyperspectral images, aiming to select the most representative and informative band combination from numerous bands to improve processing efficiency and analysis accuracy[8](W. Sun and Q. Du, "Hyperspectral Band Selection: A Review," in IEEE Geoscience and Remote Sensing Magazine, vol. 7, no. 2, pp. 118- 139, June 2019.). At present, the existing band selection methods are mainly divided into traditional band selection methods, band selection methods based on deep learning and the mixture of the two[9](R. N. Patro, S. Subudhi, P. K. Biswal and F. Dell’acqua, "A Review of Unsupervised Band Selection Techniques: Land Cover Classification for Hyperspectral Earth Observation Data," in IEEE Geoscience and Remote Sensing Magazine, vol. 9, no. 3, pp. 72-111, Sept. 2021.).
[0004] LiDAR data measures the distance of target objects through laser pulses, generating high-precision three-dimensional terrain and feature structure information, and is widely used in topographic mapping, forest structure analysis, and urban modeling
[10] (Shi, Y; Wang, T; Skidmore, A.K; Holzwarth, S; Heiden, U; Heurich, M. Mapping Individual Silver Fir Trees Using Hyperspectral and LiDAR Data in a Central European Mixed Forest. Int. J. Appl. Earth Obs. Geoinf. 2021, 98, 102311.). In addition, the independence of LiDAR data enables it to play a key role in multi-modal fusion. With the continuous development of sensor technology, the data of a single modality often cannot meet the information needs in complex scenarios. Multi-modal data fusion emerges as the times require, which integrates data from different sensors, different physical quantities, or different times to fully exploit the advantages of each modality and make up for the shortcomings of a single modality
[11] (Xie D, Zhang X, Gao X, et al. MAF-Net: A multimodal data fusion approach for human action recognition[J]. PloS one, 2025, 20(4): e0319656.). In recent years, the research on the fusion of LiDAR and HSI has gradually increased. LiDAR can provide high-precision terrain and feature height information, while HSI provides rich spectral information, and the fusion of the two can achieve complementary advantages
[12] (B. Yang, X. Wang, Y. Xing, C. Cheng, W. Jiang and Q. Feng, "Modality Fusion Vision Transformer for Hyperspectral and LiDAR Data Collaborative Classification," in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 17052-17065, 2024.). According to the existing research progress, the fusion method of HSI and LiDAR data can be classified into traditional methods and deep learning methods.
[0005] Current researches mainly focus on the features after fusion for feature classification and target recognition, and there are relatively few studies on hyperspectral image band selection guided by LiDAR. Yang et al.
[13] (Yang J X, Zhou J, Wang J, et al. Unsupervised Band Selection Using Fused HSI and LiDAR Attention Integrating With Autoencoder[J]. arXiv preprint arXiv:2404.05258, 2024.) proposed an unsupervised band selection method based on the attention mechanism and autoencoder of fused HSI and LiDAR data, which effectively improved the classification accuracy and information extraction efficiency of hyperspectral data by generating a fusion mask and using an autoencoder for reconstruction. In addition, Yang et al.
[14] (J. X. Yang, J. Zhou, J. Wang, H. Tian and A. W.-C. Liew, "LiDAR-Guided Cross-Attention Fusion for Hyperspectral Band Selection and Image Classification," in IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1-15, 2024.) proposed a hyperspectral band selection method based on cross-attention mechanism, which guided the selection of hyperspectral bands using LiDAR data. By fusing the elevation information of LiDAR and the spectral information of hyperspectral, the accuracy and efficiency of image classification were significantly improved, while the data redundancy and computational demand were reduced. However, different modal data often contain different features and information. On the one hand, when fusing data of different modalities, the existing fusion methods are prone to lose key detail features, especially the detail information of different frequency domains, affecting the final classification results. On the other hand, facing a large number of bands, many existing methods are difficult to adaptively select bands according to the characteristics of the data. SUMMARY
[0006] The present application is to solve the problem that most existing band selection methods are mainly for single modal data, often ignoring the potential advantages of elevation information in multi-modal data, resulting in limited comprehensiveness and accuracy of information mining. A multi-modal fusion dynamic hyperspectral image band selection method guided by LiDAR is proposed.
[0007] The specific process of the multi-modal fusion dynamic hyperspectral image band selection method guided by LiDAR is as follows:
[0008] Step one, constructing a double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS;
[0009] The double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS sequentially comprises a double-branch multi-modal cross-feature fusion module DBMCFM and a dynamic band selection strategy DBSS.
[0010] Step two, training the double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS based on a training set, to obtain a trained double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS;
[0011] Step three, collecting to-be-measured HSI data and LiDAR data, inputting the to-be-measured HSI data and LiDAR data into the trained double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS, and outputting selected bands by the double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS.
[0012] The LiDAR data is laser radar data; and the HSI data is hyperspectral data.
[0013] The present application has the following beneficial effects:
[0014] The application provides a LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method. First, a wavelet deformable convolution (WTDefconv) is proposed and used for downsampling, effectively capturing multi-scale features. Second, a multi-scale feature alignment fusion module (MSFAFM) is designed to adaptively align hyperspectral and LiDAR data features. Then, a bidirectional cross-attention fusion module (BCAFM) is constructed to effectively fuse hyperspectral and LiDAR data. Finally, a dynamic band selection module based on Top-k sparse attention (DBSM-TKSA) is designed to automatically learn band weights and select the optimal band. A large number of experiments show that the proposed method can achieve the best classification performance on three public datasets Houston 2013, Trento and MUUFL compared with other advanced methods, which fully demonstrates the effectiveness of the proposed band selection method. The application provides a LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method. First, a wavelet deformable convolution (WTDefconv) is proposed, which can utilize its processing capability in the frequency domain to extract multi-scale features of HSI and LiDAR modalities, avoiding the loss of key detail features. Second, through multi-scale feature alignment fusion and cross-attention fusion mechanisms, HSI and LiDAR features are effectively aligned and fused. Finally, band selection is adaptively performed according to data features, overcoming the limitations of existing methods that cannot dynamically select bands.
[0015] The application proposes a WTDefconv, which effectively captures key multi-scale features in a heterogeneous frequency domain by fusing frequency domain analysis and dynamic spatial perception technology, avoiding key information loss; the application designs a multi-scale feature alignment fusion module (MSFAFM), which realizes multi-scale feature alignment and adaptive fusion of HSI and LiDAR data by synergistically integrating the ability of dynamic receptive field adjustment and global-local feature extraction; the application constructs a bidirectional cross attention fusion module (BCAFM), which realizes global-local feature complementation and effective fusion of heterogeneous features of HSI and LiDAR data by synergistic design of bidirectional cross attention mechanism and dynamic convolution fusion; the application designs a dynamic band selection module based on Top-k sparse attention (DBSM-TKSA), which realizes dynamic selection of important bands by automatically learning the band weight of the fusion data through the sparse attention mechanism. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 FIG. 1 is a schematic diagram of the overall framework of the dual-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS of the application; Figure 2 FIG. 2 is a schematic diagram of the MSFAFM; Figure 3 FIG. 3 is a structural diagram of the BCAFM; Figure 4 FIG. 4 is a classification diagram of all methods on the Trento dataset; Figure 5a FIG. 5 is a location diagram of 20 bands selected by different methods on the Houston 2013 dataset; Figure 5b FIG. 6 is a schematic diagram of the entropy value corresponding to each spectral band of the Houston 2013 dataset, where Value of Entropy is the entropy value and Spectral Bands is the band. DETAILED DESCRIPTION
[0017] Embodiment 1: The specific process of the multi-modal fusion dynamic hyperspectral image band selection method under the guidance of LiDAR in this embodiment is as follows:
[0018] Step 1, constructing a dual-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS;
[0019] The dual-branch multi-modal cross feature fusion and dynamic band selection strategy model LGMF-DBS sequentially comprises a dual-branch multi-modal cross feature fusion module DBMCFM and a dynamic band selection strategy DBSS.
[0020] Step two, training the dual-branch multi-modal cross feature fusion and dynamic band selection strategy model LGMF-DBS based on the training set, to obtain the trained dual-branch multi-modal cross feature fusion and dynamic band selection strategy model LGMF-DBS;
[0021] Step three, collecting the to-be-tested HSI data and LiDAR data, inputting the to-be-tested HSI data and LiDAR data into the trained dual-branch multi-modal cross feature fusion and dynamic band selection strategy model LGMF-DBS, and outputting the selected band by the dual-branch multi-modal cross feature fusion and dynamic band selection strategy model LGMF-DBS;
[0022] The LiDAR data is laser radar data, and the HSI data is hyperspectral data.
[0023] Figure 1 The overall structure of the proposed method LGMF-DBS is shown. From Figure 1 It can be seen that LGMF-DBS includes two parts: a dual-branch multi-modal cross feature fusion module (DBMCFM) and a dynamic band selection strategy (DBSS). First, DBMCFM uses WTDefconv for downsampling to capture multi-scale features of HSI and LiDAR data respectively; then, the down-sampled multi-scale features are aligned through MSFAFM; finally, the aligned features are fused by BCAFM. DBSS automatically learns the band weight of the fused features through the Top-k sparse attention mechanism, and dynamically selects the optimal band subset.
[0024] Unlike existing hyperspectral image band selection methods, this invention proposes a LiDAR-guided multimodal fusion hyperspectral dynamic band selection method. First, to address the limitations of local receptive fields and the loss of detail information during multi-scale feature extraction, WTDefconv is proposed, which effectively expands the receptive field and captures multi-scale features across different frequency domains. Second, to solve the problems of multi-scale global-local feature extraction and adaptive alignment fusion, MSFAFM is proposed, enabling multi-scale global-local feature extraction and adaptive alignment fusion of HSI and LiDAR data. Then, to achieve complementary alignment fusion of HSI and LiDAR features, BCAFM is proposed, which achieves full feature complementarity alignment between HSI and LiDAR data through a bidirectional cross-attention mechanism. Finally, DBSM-TKSA is proposed, which automatically learns band weights through TKSA to dynamically select important bands. Extensive experiments demonstrate that the proposed method achieves better classification performance on three datasets compared to other state-of-the-art methods. In the future, we will also focus on exploring band selection strategies for HSI and LiDAR data fusion to further improve the effectiveness and reliability of the method.
[0025] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that: in step one, a dual-branch multimodal cross-feature fusion and dynamic band selection strategy model LGMF-DBS is constructed;
[0026] The dual-branch multi-modal cross feature fusion and dynamic band selection strategy model LGMF-DBS consists of a dual-branch multi-modal cross feature fusion module (DBMCFM) and a dynamic band selection strategy (DBSS); the specific process is as follows:
[0027] The dual-branch multi-modal cross feature fusion and dynamic band selection strategy model LGMF-DBS consists of a dual-branch multi-modal cross feature fusion module (DBMCFM) and a dynamic band selection strategy (DBSS).
[0028] The dual-branch multimodal cross-feature fusion module DBMCFM includes, in sequence, a first convolutional layer, a WTDefconv layer, an MSFAFM layer, and a BCAFM layer;
[0029] The dynamic band selection strategy DBSS includes, in sequence, a convolutional kernel size of... Two-dimensional convolutional layers and DBSM-TKSA;
[0030] The first convolutional layer is a two-dimensional convolution with a kernel size of [size missing]. ;
[0031] The MSFAFM layer includes a first group of PS blocks, a first max pooling layer, a second group of PS blocks, a second max pooling layer, a third group of PS blocks, a third max pooling layer, a fourth group of PS blocks, upsampling, a first group of DS blocks, a second group of DS blocks, a third group of DS blocks, a fourth group of DS blocks, upsampling, a first feature alignment block, a second feature alignment block, a third feature alignment block, and a fourth feature alignment block.
[0032] The BCAFM layer includes a layer normalization layer (LN), a flattening layer, an attention mechanism layer, a second convolutional layer, and a GELU layer. Other steps and parameters are the same as in Specific Implementation Method 1.
[0033] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method One or Two in that the working process of the dual-branch multimodal cross-feature fusion and dynamic band selection strategy model is as follows:
[0034] 1) Obtain the raw HSI data and LiDAR data ;
[0035] in, and These represent the height and width of the HSI and LiDAR data, respectively. Indicates the number of raw bands in the HSI data; This indicates the number of raw bands in the LiDAR data;
[0036] After filling the edge pixels (filling with 0 when edge information is missing), patches are extracted from each pixel of the HSI data and LiDAR data respectively to obtain the HSI data cube. and LiDAR data cube ;in, express ;
[0037] Then adopt and As input data for dual-branch downsampling.
[0038] 2) HSI data cube and LiDAR data cube Input the dual-branch multimodal cross-feature fusion module DBMCFM, and output the dual-branch multimodal cross-feature fusion module DBMCFM. The specific process is as follows:
[0039] 21) HSI data cube Input to the first convolutional layer, output features from the first convolutional layer ;
[0040] HSI data cube Input to the first convolutional layer, output features from the first convolutional layer ;
[0041] The first convolutional layer is a two-dimensional convolution with a kernel size of [size missing]. ;
[0042] 22) Output features of the first convolutional layer Input to the WTDefconv layer, output features from the WTDefconv layer. ;
[0043] The first convolutional layer outputs features Input to the WTDefconv layer, output features from the WTDefconv layer. ;
[0044] 23) Output features of WTDefconv layer Input to MSFAFM layer, output features of MSFAFM layer ;
[0045] WTDefconv layer output features Input to MSFAFM layer, output features of MSFAFM layer ;
[0046] 24) MSFAFM layer output characteristics and characteristics Input to BCAFM layer, output features of BCAFM layer ;
[0047] 3) Output features of BCAFM layer Input the Dynamic Band Selection Strategy (DBSS), and the DBSS will output the selected bands. Other steps and parameters are the same as in Specific Implementation Method 1 or 2.
[0048] Specific Implementation Method Four: This implementation method differs from Specific Implementation Methods One to Three in that: the output features of the first convolutional layer in step 22) are... Input to the WTDefconv layer, output features from the WTDefconv layer. The first convolutional layer outputs features. Input to the WTDefconv layer, output features from the WTDefconv layer. The specific process is as follows:
[0049] In feature extraction tasks, most previous methods directly used convolutional neural networks (CNNs) for downsampling. However, traditional CNN downsampling methods have limited receptive field expansion, limited ability to capture features of different scales and shapes, and are prone to losing detailed information. Considering that deformable convolutions can adaptively sample feature maps by learning offsets, thus better expanding the receptive field; and wavelet transforms can extract features at multiple scales in different frequency domains, we propose WTDefconv, a method that can effectively expand the receptive field and capture multi-scale detailed information in different frequency domains.
[0050] 221) Output features of the first convolutional layer Input to the WTDefconv layer, output features from the WTDefconv layer. The specific process is as follows:
[0051] 2211) Using wavelet transform on the features output of the first convolutional layer Downsampling is performed to obtain a low-frequency component feature. and three high-frequency component characteristics , and ; indicates as:
[0052]
[0053] In the formula, Indicates a low-pass filter. Indicates a high-pass filter; superscript This indicates the transpose;
[0054] 2212) Characteristics of high-frequency components , and Grouped convolutions were used separately to obtain the enhanced high-frequency component features. , and Enhance the local structural information of high-frequency features while avoiding cross-channel interference; represented as:
[0055]
[0056]
[0057]
[0058] In the formula, Indicates the kernel size as Grouped convolutions;
[0059] 2213) Enhanced high-frequency component characteristics , and Low-frequency component characteristics The cascaded features are then upsampled using inverse wavelet transform to obtain the reconstructed features. ; indicates as:
[0060]
[0061] In the formula, Indicates cascading; This indicates inverse wavelet transform upsampling;
[0062] 2214) Features output by the first convolutional layer After deformable convolution , to obtain features ;
[0063] 2215) Features and reconstructed features Element-wise summation is performed to obtain the output features of HSI after WTDefconv downsampling. ; indicates as:
[0064]
[0065] 222) Output features of the first convolutional layer Input to the WTDefconv layer, output features from the WTDefconv layer. The specific process is as follows:
[0066] 2221) Using wavelet transform on the features output of the first convolutional layer Downsampling is performed to obtain a low-frequency component feature. and three high-frequency component characteristics , and ; indicates as:
[0067]
[0068] In the formula, Indicates a low-pass filter. Indicates a high-pass filter; superscript This indicates the transpose;
[0069] 2222) Characteristics of high-frequency components , and Grouped convolutions were used separately to obtain the enhanced high-frequency component features. , and Enhance the local structural information of high-frequency features while avoiding cross-channel interference; represented as:
[0070]
[0071]
[0072]
[0073] In the formula, Indicates the kernel size as Grouped convolutions;
[0074] 2223) Enhanced high-frequency component characteristics , and Low-frequency component characteristics The cascaded features are then upsampled using inverse wavelet transform to obtain the reconstructed features. ; indicates as:
[0075]
[0076] In the formula, Indicates cascading; This indicates inverse wavelet transform upsampling;
[0077] 2224) Features output by the first convolutional layer After deformable convolution , to obtain features ;
[0078] 2225), Features and reconstructed features Element-wise summation is performed to obtain the output features of HSI after WTDefconv downsampling. ; indicates as:
[0079] .
[0080] The other steps and parameters are the same as those in one of the specific implementation methods one to three.
[0081] Specific Implementation Method Five: This implementation method differs from Specific Implementation Methods One to Four in that: the WTDefconv layer outputs features in 23) Input to MSFAFM layer, output features of MSFAFM layer ;WTDefconv layer output features Input to MSFAFM layer, output features of MSFAFM layer The specific process is as follows:
[0082] Figure 2 This is a schematic diagram of MSFAFM, whose core is global-local feature extraction and feature alignment and fusion. The encoder part performs long-range global-local feature extraction through SW-GCSA, while the decoder part achieves feature fusion through an Adaptive Feature Fusion Block (AFF Block). The intermediate attention gate achieves feature alignment between the encoder and decoder through skip connections.
[0083] 231) Output features of WTDefconv layer Input to MSFAFM layer, output features of MSFAFM layer The specific process is as follows:
[0084] 2311) Output features of WTDefconv layer Input the first set of PS blocks, and output the features of the first set of PS blocks. ;
[0085] The first group of PS block output features Input to the first max pooling layer, output features from the first max pooling layer. ;
[0086] The output features of the first max pooling layer Input the second set of PS blocks, and the second set of PS blocks outputs features. ;
[0087] Second group of PS block output features Input to the second max pooling layer, output features from the second max pooling layer. ;
[0088] Output features of the second max pooling layer Input the third PS block, and the third PS block outputs the features. ;
[0089] Third group of PS block output characteristics Input to the third max pooling layer, output features from the third max pooling layer. ;
[0090] Output features of the third max pooling layer Input the fourth PS block, and the fourth PS block outputs the features. ;
[0091] 2312), Output characteristics of the fourth group of PS blocks Features are obtained through upsampling ;
[0092] 2313) Features The first set of DS blocks is input, and the first set of DS blocks outputs features. ;
[0093] First group of DS block output features The second set of DS blocks is input, and the second set of DS blocks outputs features. ;
[0094] Second group of DS block output features Input the third DS block, and the third DS block outputs the features. ;
[0095] Third group of DS block output characteristics Input the fourth DS block, and the fourth DS block outputs the features. ;
[0096] 2314), Output characteristics of the fourth group of DS blocks Features are obtained through upsampling ;
[0097] 2315), Output characteristics of the first group of PS blocks and the output features of the fourth group of DS blocks Input the first feature alignment block, output the first feature alignment block features ;
[0098] The first feature alignment block includes, in sequence, a convolutional kernel size of... Two-dimensional convolution, with a kernel size of Two-dimensional convolution, ReLU activation function layer, convolution kernel size is Two-dimensional convolutional and sigmoid activation function layers;
[0099] 2316), Output characteristics of the second group of PS blocks and the output features of the third group of DS blocks Input the second feature alignment block, and the second feature alignment block outputs features. ;
[0100] The second feature alignment block includes, in sequence, a convolutional kernel size of... Two-dimensional convolution, with a kernel size of Two-dimensional convolution, ReLU activation function layer, convolution kernel size is Two-dimensional convolutional and sigmoid activation function layers;
[0101] 2317), Output characteristics of the third group of PS blocks Second group DS block output features Input the third feature alignment block, output the third feature alignment block features ;
[0102] The third feature alignment block includes, in sequence, a convolutional kernel size of... Two-dimensional convolution, with a kernel size of Two-dimensional convolution, ReLU activation function layer, convolution kernel size is Two-dimensional convolutional and sigmoid activation function layers;
[0103] 2318), Output characteristics of the fourth group of PS blocks and the output features of the first group of DS blocks Input the fourth feature alignment block, and output the features of the fourth feature alignment block. ;
[0104] The fourth feature alignment block includes, in sequence, a convolution kernel of size [size missing]. Two-dimensional convolution, with a kernel size of Two-dimensional convolution, ReLU activation function layer, convolution kernel size is Two-dimensional convolutional and sigmoid activation function layers;
[0105] 2319) Features and characteristics The summation is performed, and the summation result is used as the output feature of the MSFAFM layer. ;
[0106] 232) Output features of WTDefconv layer Input to MSFAFM layer, output features of MSFAFM layer The specific process is as follows:
[0107] 2321) Output features of WTDefconv layer Input the first set of PS blocks, and output the features of the first set of PS blocks. ;
[0108] The first group of PS block output features Input to the first max pooling layer, output features from the first max pooling layer. ;
[0109] The output features of the first max pooling layer Input the second set of PS blocks, and the second set of PS blocks outputs features. ;
[0110] Second group of PS block output features Input to the second max pooling layer, output features from the second max pooling layer. ;
[0111] Output features of the second max pooling layer Input the third PS block, and the third PS block outputs the features. ;
[0112] Third group of PS block output characteristics Input to the third max pooling layer, output features from the third max pooling layer. ;
[0113] Output features of the third max pooling layer Input the fourth PS block, and the fourth PS block outputs the features. ;
[0114] 2322), Output characteristics of the fourth group of PS blocks Features are obtained through upsampling ;
[0115] 2323), Features The first set of DS blocks is input, and the first set of DS blocks outputs features. ;
[0116] First group of DS block output features The second set of DS blocks is input, and the second set of DS blocks outputs features. ;
[0117] Second group of DS block output features Input the third DS block, and the third DS block outputs the features. ;
[0118] Third group of DS block output characteristics Input the fourth DS block, and the fourth DS block outputs the features. ;
[0119] 2324), Output characteristics of the fourth group of DS blocks Features are obtained through upsampling ;
[0120] 2325), Output characteristics of the first group of PS blocks and the output features of the fourth group of DS blocks Input the first feature alignment block, output the first feature alignment block features ;
[0121] The first feature alignment block includes, in sequence, a convolutional kernel size of... Two-dimensional convolution, with a kernel size of Two-dimensional convolution, ReLU activation function layer, convolution kernel size is Two-dimensional convolutional and sigmoid activation function layers;
[0122] 2326), Output characteristics of the second group of PS blocks and the output features of the third group of DS blocks Input the second feature alignment block, and the second feature alignment block outputs features. ;
[0123] The second feature alignment block includes, in sequence, a convolutional kernel size of... Two-dimensional convolution, with a kernel size of Two-dimensional convolution, ReLU activation function layer, convolution kernel size is Two-dimensional convolutional and sigmoid activation function layers;
[0124] 2327), Output characteristics of the third group of PS blocks Second group DS block output features Input the third feature alignment block, output the third feature alignment block features ;
[0125] The third feature alignment block includes, in sequence, a convolutional kernel size of... Two-dimensional convolution, with a kernel size of Two-dimensional convolution, ReLU activation function layer, convolution kernel size is Two-dimensional convolutional and sigmoid activation function layers;
[0126] 2328), Output characteristics of the fourth group of PS blocks and the output features of the first group of DS blocks Input the fourth feature alignment block, and output the features of the fourth feature alignment block. ;
[0127] The fourth feature alignment block includes, in sequence, a convolution kernel of size [size missing]. Two-dimensional convolution, with a kernel size of Two-dimensional convolution, ReLU activation function layer, convolution kernel size is Two-dimensional convolutional and sigmoid activation function layers.
[0128] 2329), Features and characteristics The summation is performed, and the summation result is used as the output feature of the MSFAFM layer. ;
[0129] The encoder section consists of It consists of PatchEmbedding and SW-GCSA. Since window attention computation cannot effectively capture long-distance dependencies across windows, Grouped Channel Self-Attention (GCSA) is used to enhance feature capture capabilities. GCSA computes attention independently within each channel group through a grouped channel attention mechanism.
[0130]
[0131] In the formula, , This refers to the number of channels after grouping. To prevent spatial information leakage, the number of channels after grouping must meet certain requirements. By independently calculating attention, local features can be captured more effectively, while global feature fusion is achieved through cross-group information interaction. Finally, by cascading and combining the weighted results of each group, the following output features are obtained.
[0132]
[0133] In the formula, For the first Group features, For channel attention functions.
[0134] The attention gate in the middle consists of three layers. Convolutional layers are used to align features between the encoder and decoder. The first layer... The convolutional layer performs a linear transformation on the output features of the decoder. (Second layer) The convolutional layer performs a linear transformation on the skip connection features. (Third layer) The convolutional layer further processes the features after ReLU activation, generating attention weights through the Sigmoid activation function. Finally, the output features are obtained after passing through the gated attention module.
[0135]
[0136]
[0137] Each stage of the decoder integrates an AFF Block. The AFF Block captures long-range dependencies through SW-GCSA; it also combines deformable convolution to dynamically adjust the local receptive field; and adaptively fuses multi-scale features.
[0138]
[0139]
[0140] .
[0141] The other steps and parameters are the same as in any of the specific implementation methods one to four.
[0142] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One to Five in that the working process of each group of PS blocks in the first group, the second group, the third group, and the fourth group is as follows:
[0143] Feature 1 is input to PatchEmbedding, and PatchEmbedding outputs Feature 2; Feature 2 is input to Layer Normalization, and Layer Normalization outputs Feature 3; Feature 3 is input to the first... Dilated convolution, first Dilated convolution output features ,feature The dimension size is , This represents the total number of channels. The number of groups. For height, Width; Feature 3 input second Dilated convolution, second Dilated convolution output features ,feature The dimension size is Feature 3 Input Third Dilated convolution, third Dilated convolution output features ,feature The dimension size is ;feature Transpose and features Element-wise multiplication yields the features. ,feature The dimension size is ;feature and characteristics Element-wise multiplication yields the features. ,feature Concatenate with feature 2 to obtain feature 2 ,feature Input to a multilayer perceptron (MLP), output features of the multilayer perceptron (MLP) ;feature This serves as the output for each group of PS blocks. Other steps and parameters are the same as in any of the specific implementation methods one through five.
[0144] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One through Six in that the working process of each group of DS blocks in the first group of DS blocks, the second group of DS blocks, the third group of DS blocks, and the fourth group of DS blocks is as follows:
[0145] Feature 1 is input to PatchEmbedding, and PatchEmbedding outputs Feature 2; Feature 2 is input to Layer Normalization, and Layer Normalization outputs Feature 3; Feature 3 is input to the first... Dilated convolution, first Dilated convolution output features ,feature The dimension size is , This represents the total number of channels. The number of groups. For height, Width; Feature 3 input second Dilated convolution, second Dilated convolution output features ,feature The dimension size is Feature 3 Input Third Dilated convolution, third Dilated convolution output features ,feature The dimension size is ;feature Transpose and features Element-wise multiplication yields the features. ,feature The dimension size is ;feature and characteristics Element-wise multiplication yields the features. ,feature Concatenate with feature 2 to obtain feature 2 ,feature Input to a multilayer perceptron (MLP), output features of the multilayer perceptron (MLP) ;feature This serves as the output for each group of DS blocks. Other steps and parameters are the same as in any of the specific implementation methods one through six.
[0146] Specific Implementation Method Eight: This implementation method differs from Specific Implementation Methods One through Seven in that: the MSFAFM layer output characteristics in 24) and characteristics Input to BCAFM layer, output features of BCAFM layer The specific process is as follows:
[0147] To map HSI and LiDAR features with different structures and distributions to the same representation space, such as Figure 3 As shown, a BCAFM is proposed, which mainly consists of a bidirectional cross-attention mechanism from HSI to LiDAR and from LiDAR to HSI, and a feature recombination and fusion module. Through the attention interaction between HSI to LiDAR and from LiDAR to HSI, the features of the two modalities adapt to each other, effectively achieving feature alignment and fusion.
[0148] 241) MSFAFM layer output characteristics Features are obtained by performing layer normalization. , will feature Flatten the vector into a sequence format; input the flattened sequence format into the attention mechanism to obtain the query vector. key vector value vector ;
[0149] 242) MSFAFM layer output characteristics Features are obtained by performing layer normalization. , will feature Flatten the vector into a sequence format; input the flattened sequence format into the attention mechanism to obtain the query vector. key vector value vector ;
[0150] 243) Query vector key vector value vector Input multi-head attention mechanism (8 heads), multi-head attention mechanism output sequence ; indicates as:
[0151]
[0152] In the formula, This represents the multi-head attention mechanism; Indicates batch size; Indicates the number of channels in the HSI; Represents the set of real numbers;
[0153] 244) Query vector key vector value vector Input multi-head attention mechanism (8 heads), multi-head attention mechanism output sequence ; indicates as:
[0154]
[0155] In the formula, This indicates the number of channels in LiDAR;
[0156] 245), the sequence and sequence Cascade After cascading, the features pass through sequentially 2D convolution And the activation function GELU, to obtain features ; indicates as:
[0157] .
[0158] The other steps and parameters are the same as those in any of the specific implementation methods one to seven.
[0159] Specific Implementation Method Nine: This implementation method differs from Specific Implementation Methods One through Eight in that: the output characteristics of the BCAFM layer in 3) Input the Dynamic Band Selection Strategy (DBSS), and the DBSS will output the selected bands; the specific process is as follows:
[0160] 31) Output characteristics of BCAFM layer The input convolution kernel size is Two-dimensional convolutional layers output features ;
[0161] 32) Features Input DBSM-TKSA, and DBSM-TKSA will output the selected band.
[0162] The other steps and parameters are the same as those in one of the specific implementation methods one to eight.
[0163] Specific Implementation Method Ten: This implementation method differs from Specific Implementation Methods One through Nine in that: the feature in 32) is... Input DBSM-TKSA, and DBSM-TKSA will output the selected band; the specific process is as follows:
[0164] 321) Initialize the global average pooling layer Initialize the linear layer for Conversion; Initialization of learnable temperature parameters , To create a matrix of all 1s; set The ratio is 0.5; an empty weight accumulation buffer is initialized. Initialize an empty sample count buffer. ;
[0165] 322) From the characteristics Extract batch size Number of bands ,high and width The sample;
[0166] 323) Input the samples extracted in 322) into the global average pooling layer. Global average pooling layer Output features ;
[0167] 324) Features The query vector is obtained by sequentially passing through a linear layer and then reshaping. ;
[0168] feature The bond vector is obtained by sequentially passing it through a linear layer and then reshaping it. ;
[0169] feature The value vector is obtained by sequentially passing through a linear layer and then reshaping. ;
[0170] Query vector transpose and key vector Perform element-wise multiplication and then combine with the learnable temperature parameter Multiply them to obtain the attention score matrix;
[0171] based on The ratio is 0.5, and the values less than 0.5 in the sparse matrix are set to 0 to obtain the sparse attention weight matrix;
[0172] The sparse attention weight matrix is processed by Softmax to obtain a matrix vector. Each row of the matrix vector obtained by Softmax represents the weight of each band 1 with other bands. The number of rows in the matrix vector obtained by Softmax is the number of batches, and the number of columns in the matrix vector obtained by Softmax is the number of bands.
[0173] For example, a matrix vector is Matrix vector 4 represents the band number, and 3 represents the batch.
[0174] The first row of the matrix vector represents the weights of band 1 and band 2 as 0.8, band 1 and band 3 as 0.7, band 1 and band 4 as 0.5, and band 1 and band 5 as 0.4. The second row of the matrix vector represents the weights of band 2 and band 1 as 0.2, band 2 and band 3 as 0.3, band 2 and band 4 as 0.7, and band 2 and band 5 as 0.4. The third row of the matrix vector represents the weight of band 3 to band 1 as 0.5, the weight of band 3 to band 2 as 0.1, the weight of band 3 to band 4 as 0.3, and the weight of band 3 to band 5 as 0.8.
[0175] The average of the channel dimensions of each column of the matrix vector obtained after Softmax is calculated to obtain... A matrix vector, where each value represents the average value for each band in each batch. ; Multiplying the matrix vector by the batch number (the number of rows in the matrix vector obtained after Softmax) yields the band weight matrix vector. );
[0176] The band weight matrix vector is placed into the weight accumulation buffer. ;
[0177] The batch corresponding to the band weight matrix vector is placed into the sample counting buffer. ;
[0178] Weight accumulation buffer The band weight matrix vectors in the sample are summed, and the summed band weight matrix vectors are divided by the total number of extracted samples (322) to obtain the average band weight. ; indicates as:
[0179]
[0180] in, This represents the summed band weight matrix vector (e.g., the band weight matrix vector). add add equal ); This represents the total number of samples extracted (322).
[0181] 325) Average band weights AND value vector Perform element-wise multiplication ( Each number in the equation is used as a weight and is respectively compared with the average band weight. Multiplying the corresponding numbers in the matrix yields a vector matrix and the feature vector matrix output from the BCAFM layer (31). Perform element-wise summation, then input the sum into the Sigmoid activation function, which outputs the features.
[0182] 326) Based on The algorithm processes the output features of the Sigmoid activation function and selects the first... Each index corresponds to a waveband; [Previous] The bands corresponding to each index are selected as the bands for DBSM-TKSA output.
[0183] In the Transformer architecture, self-attention is an important module for modeling the internal relationships of a sequence. This standard self-attention mechanism can capture the global dependencies between different positions in the sequence, but it also has a problem: it will sum the similarity scores of all query key pairs with weights, including those keys that are not relevant to the query or have low relevance. This may introduce noise and interfere with the subsequent feature aggregation process. Top-k Sparse Attention (TKSA) is to selectively retain the top k attention scores that are most relevant to the query, thereby removing attention values that are not useful for feature aggregation
[15] (Chen X, Li H, Li M, et al. Learning asparse transformer network for effective image deraining[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2023:5896-5905.). To select the most representative bands from the fused multimodal data, this invention proposes a DBSM-TKSA method. By introducing TKSA, self-attention calculation is performed at the band dimension, and Top-k sparsity is applied to retain only the bands most relevant to each band. The interaction of each band effectively filters redundant noise and enhances the ability to capture key band relationships. Simultaneously, during training, the automatic learning of global band weights makes band selection more adaptive.
[0184] The other steps and parameters are the same as those in any of the specific implementation methods one to nine.
[0185] Related work
[0186] A. Hyperspectral band selection: With the development of technology, deep learning has made great progress in various fields
[16] (C. Shi, T. Wang and L. Wang, "Branch Feature Fusion Convolution Network for Remote Sensing Scene Classification," in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 13, pp. 5194-5210,2020.). Convolutional neural networks, graph neural networks, Transformers, and autoencoders are all quite excellent in band selection methods
[17] (Y. Cai, X. Liu and Z. Cai, "BS-Nets: An End-to-End Framework for Band Selection of Hyperspectral Image," in IEEE Transactions on Geoscience and Remote Sensing, vol. 58, no. 3, pp. 1969-1984, March 2020.),
[18] (C. Yu, S. Zhou, M. Song, B. Gong, E. Zhao and C. -I. Chang, "Unsupervised Hyperspectral Band Selection via Hybrid Graph Convolutional Network," in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-15, 2022.).Depending on whether labeled samples are used, these methods can be further divided into supervised,
[19] (X. Cao, T. Xiong and L. Jiao, "Supervised Band Selection Using Local Spatial Information for Hyperspectral Image," in IEEE Geoscience and Remote Sensing Letters, vol. 13, no. 3, pp. 329-333, March 2016.), and semi-supervised,
[20] (J. Feng, L. Jiao, F. Liu, T. Sunand X. Zhang, "Mutual-Information-Based Semi-Supervised Hyperspectral Band Selection With High Discrimination, High Information, and Low Redundancy," in IEEE Transactions on Geoscience and Remote Sensing, vol. 53, no. 5, pp. 2956-2969, May 2015.)
[21] (X. Cao, C. Wei, Y. Ge, J. Feng, J. Zhao and L. Jiao, "Semi-Supervised Hyperspectral Band Selection Based on Dynamic ClassifierSelection," in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 12, no. 4, pp. 1289-1298, April 2019.
[22] (L. Jiao, J. Feng, F. Liu, T. Sun and X. Zhang, "Semisupervised Affinity PropagationBased on Normalized Trivariable Mutual Information for Hyperspectral Band Selection," in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 8, no. 6, pp. 2760-2773, June 2015.) and unsupervised methods. Compared with supervised and semi-supervised methods, unsupervised band selection methods are suitable for more downstream tasks because they do not require pre-labeling and are not constrained by labeled data. Cui et al.
[23] (Cui, C; Sun, X; Fu, B; Shang, X.SSANet-BS: Spectral–Spatial Cross-Dimensional Attention Network for Hyperspectral Band Selection. Remote Sens. 2024, 16, 2848.) proposed a novel unsupervised hyperspectral band selection method called SSANet-BS. By combining the spectral spatial cross-dimensional attention mechanism and the multi-scale reconstruction network, it can automatically identify and select the most representative bands in hyperspectral images, effectively improving the performance and stability of band selection. By segmenting the original image into multiple homogeneous regions and fusing the features of these regions into a low-dimensional latent space, Feng et al.
[24] (W. Feng et al., "Hyperspectral band selection via region-wise latent feature fusion and graph filter embedded subspace clustering", Eng. Appl. Artif. Intell., vol. 132, 2024.) proposed a band selection algorithm called region-wise latent feature fusion and graph filter embedded subspace clustering (RFGEC), which effectively captures spatial information.By utilizing the interaction between pixels and superpixels, spatial and structural information is embedded into the model. A new hyperspectral band selection method based on global local graph self-encoding (Tensorial Global-Local Graph Self-Representation, TGSR) was proposed
[25] (Y.Zhang, J. Qi, X. Wang, Z. Cai, J. Peng and Y. Zhou, "Tensorial Global-Local Graph Self-Representation for Hyperspectral Band Selection," in IEEE Transactions on Circuits and Systems for Video Technology, doi: 10.1109 / TCSVT.2024.), which improves the accuracy of band selection. Zhang et al.
[26] (W. Zhang, A. Yuan, J. Tangand X. Li, "Sparse Principal Component Analysis and Adaptive Multigraph Learning for Hyperspectral Band Selection," in IEEE Journal of SelectedTopics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 1419-1433, 2024.) proposed a hyperspectral band selection method based on sparse principal component analysis and adaptive multigraph learning (SPCA-AMGL). By combining sparse regularization and multigraph learning strategies, the low-dimensional manifold structure of the data is effectively preserved, and the quality and stability of band selection are significantly improved.
[0187] B. Fusion of HSI and LiDAR: In recent years, deep learning methods have overcome the limitations of manual feature extraction in traditional methods by automatically extracting and fusing features from multi-source data, thereby improving classification accuracy and efficiency. Zhang et al.
[27] (T.Zhang, S. Xiao, W. Dong, J. Qu and Y. Yang, "A Mutual Guidance Attention-Based Multi-Level Fusion Network for Hyperspectral and LiDAR Classification,"in IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1-5, 2022.) proposed a multi-branch convolutional neural network for HSI and LiDAR data classification, which improved classification accuracy through multi-level feature fusion and mutual guidance attention modules. Song et al.
[28] (L. Song, Z. Feng, S. Yang, X. Zhang and L. Jiao, "Discrepant Bi-Directional Interaction Fusion Network for Hyperspectral and LiDAR Data Classification," in IEEE Geoscience and RemoteSensing Letters, vol. 20, pp. 1-5, 2023.) proposed a novel inequality bi-directional interaction fusion network (DBIFNet) for classifying HSI and LiDAR data, which enhances feature extraction and fusion effects by designing different interaction modules. Feng et al.
[29] (Y. Feng, J. Jin, Y. Yin, C. Song and X. Wang, "MCFT: Multimodal Contrastive Fusion Transformer for Classification of Hyperspectral Image and LiDAR Data," in IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1-17, 2024.) proposed a multimodal contrastive fusion transformer (MCFT) for classifying HSI and LiDAR data, which improves feature fusion performance through contrastive learning and an improved Transformer architecture.Song et al.
[30] (T. Song, Z. Zeng, C. Gao, H. Chen and J. Li, "Joint Classification of Hyperspectral and LiDAR Data Using Height Information Guided Hierarchical Fusion-and-Separation Network," in IEEE Transactions on Geoscience and RemoteSensing, vol. 62, pp. 1-15, 2024.) proposed a highly information-guided hierarchical fusion-and-separation network (HFSNet) for the joint classification of HSI and LiDAR data, which improves classification performance through multi-level feature fusion and separation. A reinforcement learning-based Markov Edge Decoupled Fusion Network (MEDFN) was proposed by Wang et al.
[31] (H. Wang, Y. Cheng, X. Liu and X. Wang, "Reinforcement Learning Based Markov EdgeDecoupled Fusion Network for Fusion Classification of Hyperspectral and LiDAR," in IEEE Transactions on Multimedia, vol. 26, pp. 7174-7187, 2024.) for the fusion classification of HSI and LiDAR data. By intelligently constructing a graph structure and decoupling multimodal features, it makes full use of the complementary information of different modalities and improves the classification performance under complex spatial distribution. A joint classification method for HSI and LiDAR data based on the Mamba framework (HLMamba) was proposed by Liao et al.
[32] (D. Liao, Q. Wang, T. Lai and H. Huang, "Joint Classification of Hyperspectral and LiDAR Data Based on Mamba,"in IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1-15,2024.), which optimizes the feature fusion process by introducing the Mamba framework.Zhu et al.
[33] (F. Zhu, C. Shi, K. Shi and L. Wang, "Joint Classification of Hyperspectral and LiDAR Data Using Hierarchical Multimodal Feature Aggregation-Based Multihead Axial AttentionTransformer," in IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1-17, 2025.) proposed a multihead axial attention transformer (HMAT) based on hierarchical multimodal feature aggregation, which improves the classification performance of HSI and LiDAR data by fusing spatial spectral and elevation features.
[0188] The beneficial effects of the present invention are verified using the following embodiments:
[0189] Example 1:
[0190] The experiments used three commonly used public datasets: Houston 2013, Trento, and MUUFL. Table 1 shows the total number of classes, the number of samples for each class, and the total number of samples in the three datasets. The proposed method was implemented in the PyTorch 2.0 framework and validated using an NVIDIA RTX 4090D GPU with 24GB of RAM. The training parameters—batch size, learning rate, and patch size—were set to 32, ... And 16.
[0191] 1) Houston 2013: The Houston 2013 dataset includes information from multiple sensors. Furthermore, it incorporates radar data, increasing the diversity of the data.
[0192] 2) Trento: The Trento dataset consists of hyperspectral images of 63 spectral channels acquired by the AISA Eagle sensor, with a wavelength range of 402.89-989.09 nm and a spectral resolution of 9.2 nm. LiDAR data is also provided by the Optech ALTM 3100EA sensor.
[0193] 3) MUUFL: The MUUFL dataset contains hyperspectral images and LiDAR data.
[0194] Ablation experiments: The proposed LGMF-DBS method mainly consists of four modules: WTDefconv, MSFAFM, BCAFM, and DBSM-TKSA. Detailed ablation experiments were conducted in this section to verify the effectiveness of each module.
[0195] (1) Ablation Experiments on WTDefconv: The WTDefconv proposed in this invention employs wavelet transform and deformable convolution to effectively capture multi-scale features in different frequency domains. As shown in Case 2 of Table 2, the introduction of WTDefconv brings certain improvements compared to Case 1. Specifically, the OA values on the HT2013, Trento, and MUUFL datasets are improved by 0.56%, 0.25%, and 0.9%, respectively. As shown in Case 5 and Case 6, when WTDefconv is combined with the MSFAFM and BCAFM modules respectively, the OA values on the HT2013, Trento, and MUUFL datasets are improved compared to Case 1, by 0.34%, 0.22%, 0.81%, 0.45%, 0.17%, and 0.76%, respectively. The experimental results fully demonstrate the effectiveness of WTDefconv.
[0196] (2) Ablation Experiments on MSFAFM: The MSFAFM proposed in this invention effectively captures multi-scale global and local features and aligns and fuses these features. As shown in Case 3 of Table 2, the introduction of MSFAFM brings certain improvements compared to Case 1. Specifically, the OA values on the HT2013, Trento, and MUUFL datasets are improved by 0.28%, 0.13%, and 0.98%, respectively. As shown in Case 7, when MSFAFM is combined with the BCAFM module, the OA values on the HT2013, Trento, and MUUFL datasets are improved compared to Case 1, by 0.44%, 0.21%, and 0.75%, respectively. The experimental results demonstrate the effectiveness of MSFAFM.
[0197] (3) Ablation Experiments on BCAFM: The BCAFM proposed in this invention achieves complementary alignment and fusion of heterogeneous features of HSI and LiDAR data through a collaborative design of bidirectional cross-attention and dynamic convolution fusion. As shown in Case 4 of Table 1, compared with Case 1, using BCAFM alone improves classification performance to a certain extent. Specifically, the OA values on the HT2013, Trento, and MUUFL datasets are improved by 0.26%, 0.22%, and 1.09%, respectively. The experimental results fully demonstrate the effectiveness of BCAFM.
[0198]
[0199] Performance analysis: This experiment selected six advanced band selection methods and compared their performance with the LGMF-DBS proposed in this invention. These included five single-source data band selection methods, namely SSANet-BS, SPCA_AMGL, RFGEC, TGSR and MOBS-TD
[34] (X. Sun, P. Lin, X. Shang, H. Pang and X. Fu, "MOBS-TD: Multiobjective Band Selection With Ideal Solution Optimization Strategy for Hyperspectral Target Detection," in IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 10032-10050,2024.), and a multi-modal fusion band selection method, LGCAF. To better compare the classification performance of different methods, the Spectral–spatial Residual Network (SSRN)
[35] (Z. Zhong, J. Li, Z. Luo and M. Chapman, "Spectral–Spatial Residual Network for Hyperspectral Image Classification: A 3-D Deep Learning Framework," in IEEE Transactions on Geoscience and Remote Sensing, vol. 56, no. 2, pp. 847-858, Feb. 2018.) was used to classify and evaluate the selected band subsets of all methods. In addition, the classification metrics used included overall accuracy (OA), average accuracy (AA), and Kappa.
[0200] Tables 2-4 present the accuracy of the proposed method and all comparative methods on three datasets, with the same number of bands selected, for OA, AA, Kappa, and single-class classification. The number of bands selected is 20 for each dataset, and the training sample size is 5%. The results show that the proposed method achieves the highest classification results on all three datasets. As shown in Tables 2-4, the best results for each class in OA, AA, and Kappa are indicated in bold. For the Houston 2013 dataset, LGMF-DBS outperforms other methods in single-class accuracy. Compared to the suboptimal methods, the proposed method improves accuracy by 0.48%, 0.42%, and 0.53% in OA, AA, and Kappa, respectively. Figure 4 The classification maps further confirm these experimental results. As can be seen from the magnified local images, other methods exhibit large-scale misclassifications of "roads" and "buildings" compared to the proposed method. On the MUUFL dataset, LGMF-DBS continues to demonstrate superior classification performance, improving by 0.78%, 1.14%, and 1.02% on OA, AA, and Kappa datasets, respectively, compared to the suboptimal methods.
[0201]
[0202]
[0203]
[0204] Selected Band Analysis: This section analyzes the selected bands based on their location and entropy value. Table 5 shows the index values of the optimal bands obtained from the three datasets using different methods. Figure 5a , Figure 5b As shown, the upper part gives the spectral positions of the bands given in Table 5. Each row represents a band selection method and the spectral position of the selected band. The lower part shows the entropy values of the entire spectral band. As can be seen on the Houston 2013 dataset ( Figure 5a , Figure 5b The proposed method selects bands that are relatively dispersed, with most bands concentrated in areas with high entropy values. Compared to the proposed method, SSANet-BS selects bands that are more densely packed, mostly concentrated in the [35:55] interval, leading to significant redundancy in the selected bands. The method proposed in this invention selects bands that are more dispersed and have a better distribution even in areas with high entropy values, which greatly improves the classification performance of the proposed method.
[0205]
[0206] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A multi-modal fusion dynamic hyperspectral image band selection method under LiDAR guidance, characterized in that: The method specifically comprises the following steps: Step 1: constructing a double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS; The double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS comprises a double-branch multi-modal cross-feature fusion module DBMCFM and a dynamic band selection strategy DBSS in sequence; Step 2: training the double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS based on a training set to obtain a trained double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS; Step 3: collecting HSI data and LiDAR data to be measured, inputting the HSI data and LiDAR data to be measured into the trained double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS, and outputting selected bands by the double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS. The LiDAR data is laser radar data, and the HSI data is hyperspectral data.
2. The LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method according to claim 1, characterized in that: The double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS is constructed in step 1; The double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS comprises a double-branch multi-modal cross-feature fusion module DBMCFM and a dynamic band selection strategy DBSS in sequence; The double-branch multi-modal cross-feature fusion and dynamic band selection strategy model LGMF-DBS comprises a double-branch multi-modal cross-feature fusion module DBMCFM and a dynamic band selection strategy DBSS in sequence; The double-branch multi-modal cross-feature fusion module DBMCFM comprises a first convolutional layer, a WTDefconv layer, an MSFAFM layer and a BCAFM layer in sequence. The MSFAFM layer comprises a first group of PS blocks, a first maximum pooling layer, a second group of PS blocks, a second maximum pooling layer, a third group of PS blocks, a third maximum pooling layer, a fourth group of PS blocks, upsampling, a first group of DS blocks, a second group of DS blocks, a third group of DS blocks, a fourth group of DS blocks, upsampling, a first feature alignment block, a second feature alignment block, a third feature alignment block and a fourth feature alignment block. The dynamic band selection strategy DBSS sequentially comprises a two-dimensional convolution layer with a convolution kernel size of and a DBSM-TKSA. The first convolutional layer is a two-dimensional convolution, and the size of the convolution kernel is ; The BCAFM layer comprises a layer normalization layer LN, a flattening layer, an attention mechanism layer, a second convolutional layer and a GELU layer. The working process of the double-branch multi-modal cross-feature fusion and dynamic band selection strategy model is as follows:
3. The LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method of claim 2, wherein: The working process of each group of PS blocks in the first group of PS blocks, the second group of PS blocks, the third group of PS blocks and the fourth group of PS blocks is as follows: 1) Obtain original HSI data and LiDAR data ; wherein, and denote the height and width of the HSI and LiDAR data, respectively; denotes the original number of bands of the HSI data; denotes the original number of bands of the LiDAR data; patches are extracted from each pixel of the HSI data and the LiDAR data respectively, to obtain an HSI data cube and a LiDAR data cube respectively wherein represents ; 2) input the HSI data cube and the LiDAR data cube into the double-branch multi-modal cross-feature fusion module DBMCFM, and the double-branch multi-modal cross-feature fusion module DBMCFM outputs the feature The specific process is as follows: 21), the HSI datacube is transformed into a 2D image inputting the first convolutional layer, and outputting features ; cuboid of HSI data inputting a first convolutional layer, the first convolutional layer outputting features ; The first convolutional layer is a two-dimensional convolution, and the size of the convolution kernel is ; 22), first convolutional layer output feature input WTDefconv layer, WTDefconv layer output feature ; first convolutional layer output features input WTDefconv layer, WTDefconv layer output features ; 23) WTDefconv layer output features input MSFAFM layer, MSFAFM layer output features ; WTDefconv layer output features input MSFAFM layer, MSFAFM layer output features ; 24) MSF AFM layer output feature and feature input BCA FM layer, BCA FM layer output feature ; 3) BCAFM layer output characteristics An input dynamic band selection strategy DBSS, and the dynamic band selection strategy DBSS outputs the selected band.
4. The LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method of claim 3, wherein: The first convolutional layer output feature in 22) Input WTDefconv layer, WTDefconv layer output feature ; The first convolutional layer output feature Input WTDefconv layer, WTDefconv layer output feature ; The specific process is: 221), first convolutional layer output feature input WTDefconv layer, WTDefconv layer output feature The specific process is: 2211), using wavelet transform on the features output by the first convolutional layer down-sampling to obtain a low-frequency component feature and three high-frequency component features , and ; represented as: wherein denotes a low-pass filter, denotes a high-pass filter; upper index denotes transposition; 2212), high-frequency component features , and are respectively enhanced by using grouped convolutions, obtaining enhanced high-frequency component features , and ; denoted as: In the formula, denotes a grouped convolution with a kernel size of ; 2213), the enhanced high-frequency component feature 、 and is concatenated with the low-frequency component feature , and then is up-sampled by inverse wavelet transform to obtain the reconstructed feature ; which is represented as: wherein denotes concatenation; denotes up-sampling of the inverse wavelet transform; 2214), the features output by the first convolutional layer through deformable convolution , to obtain features ; 2215), the feature and reconstructing the feature are added element by element, resulting in the feature ; is represented as: 222), the first convolutional layer output feature input WTDefconv layer, the WTDefconv layer output feature The specific process is: 2221), using wavelet transform on the features output by the first convolutional layer down-sampling to obtain a low-frequency component feature and three high-frequency component features , and ; represented as: wherein denotes a low-pass filter, denotes a high-pass filter; upper index denotes transposition; 2222), high-frequency component features , and are respectively enhanced by using grouped convolutions, resulting in enhanced high-frequency component features , and ; denoted as: In the formula, denotes a grouped convolution with kernel size of 3 x 3. 2223), the enhanced high-frequency component feature 、 and is concatenated with the low-frequency component feature , and then is up-sampled by inverse wavelet transform to obtain the reconstructed feature ; which is represented as: wherein denotes concatenation; denotes up-sampling of the inverse wavelet transform; 2224), the features output by the first convolutional layer through deformable convolution , obtaining features ; 2225), the feature and reconstructing the feature are added element by element, resulting in the feature ; is represented as: 。 5. The LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method of claim 4, wherein: The WTDefconv layer output feature in 23) The input MSFAFM layer, the MSFAFM layer output feature The WTDefconv layer output feature The input MSFAFM layer, the MSFAFM layer output feature The specific process is: 231), WTDefconv layer output features input MSFAFM layer, MSFAFM layer output features The specific process is as follows: 2311), WTDefconv layer output features input first set of PS blocks, first set of PS blocks output features ; first group of PS block output features input a first max-pooling layer, the first max-pooling layer output features ; first max-pooling layer output feature input a second set of PS blocks, the second set of PS blocks output features ; Second set of PS block output features Input a second max-pooling layer, the second max-pooling layer outputs features ; second max-pooling layer output feature input a third set of PS blocks, the third set of PS blocks output features ; Third group of PS block output features input a third max-pooling layer, and output features of the third max-pooling layer ; third max-pooling layer output feature input a fourth set of PS blocks, fourth set of PS blocks output feature ; 2312), fourth group PS block output features up-sampled features ; 2313), feature inputting the first set of DS blocks, the first set of DS blocks outputting features ; first set of DS block output features input second set of DS blocks, second set of DS block output features ; second set of DS block output features input third set of DS blocks, third set of DS block output features ; third set of DS block output features input fourth set of DS blocks, fourth set of DS block output features ; 2314), fourth group of DS block output features features obtained by upsampling ; 2315), first set of PS block output features and fourth set of DS block output features input first feature alignment block, first feature alignment block outputs features ; The first feature alignment block sequentially comprises a two-dimensional convolution with a convolution kernel size of , a two-dimensional convolution with a convolution kernel size of , a ReLU activation function layer, a two-dimensional convolution with a convolution kernel size of , and a Sigmoid activation function layer. 2316), second set of PS block output features and third set of DS block output features inputting a second feature alignment block, the second feature alignment block outputting features ; The second feature alignment block sequentially comprises a two-dimensional convolution with a convolution kernel size of , a two-dimensional convolution with a convolution kernel size of , a ReLU activation function layer, a two-dimensional convolution with a convolution kernel size of , and a Sigmoid activation function layer. 2317), third set of PS block output features and second set of DS block output features inputting a third feature alignment block, the third feature alignment block outputting features ; The third feature alignment block sequentially comprises a two-dimensional convolution with a convolution kernel size of , a two-dimensional convolution with a convolution kernel size of , a ReLU activation function layer, a two-dimensional convolution with a convolution kernel size of , and a Sigmoid activation function layer. 2318), fourth set of PS block output features and first set of DS block output features input fourth feature alignment block, fourth feature alignment block outputs features ; The fourth feature alignment block sequentially comprises a two-dimensional convolution with a convolution kernel size of , a two-dimensional convolution with a convolution kernel size of , a ReLU activation function layer, a two-dimensional convolution with a convolution kernel size of , and a Sigmoid activation function layer. 2319), feature and features performing the summation, the result of the summation being output as a feature of the MSFAFM layer ; 232), WTDefconv layer output features input MSFAFM layer, MSFAFM layer output features The specific process is as follows: 2321), WTDefconv layer output features input first set of PS blocks, first set of PS blocks output features ; first group of PS block output features input a first max-pooling layer, the first max-pooling layer output features ; first max-pooling layer output features input second set of PS blocks, second set of PS blocks output features ; second set of PS block output features input a second max-pooling layer, the second max-pooling layer output features ; second max-pooling layer output feature inputting the third set of PS blocks, the third set of PS blocks outputting features ; Third group of PS block output features input a third max-pooling layer, the third max-pooling layer outputs features ; third max-pooling layer output feature input a fourth set of PS blocks, fourth set of PS blocks output feature ; 2322), fourth group PS block output features up-sampled features ; 2323), feature inputting a first set of DS blocks, the first set of DS blocks outputting a feature ; first set of DS block output features input second set of DS blocks, second set of DS block output features ; second set of DS block output features input third set of DS blocks, third set of DS block output features ; third set of DS block output features input fourth set of DS blocks, fourth set of DS block output features ; 2324), fourth group of DS block output features features obtained by upsampling ; 2325), first set of PS block output features and fourth set of DS block output features input first feature alignment block, first feature alignment block outputs features ; The first feature alignment block sequentially comprises a two-dimensional convolution with a convolution kernel size of , a two-dimensional convolution with a convolution kernel size of , a ReLU activation function layer, a two-dimensional convolution with a convolution kernel size of , and a Sigmoid activation function layer. 2326), second set of PS block output features and third set of DS block output features inputting a second feature alignment block, the second feature alignment block outputting features ; The second feature alignment block sequentially comprises a two-dimensional convolution with a convolution kernel size of , a two-dimensional convolution with a convolution kernel size of , a ReLU activation function layer, a two-dimensional convolution with a convolution kernel size of , and a Sigmoid activation function layer. 2327), third set of PS block output features and second set of DS block output features inputting a third feature alignment block, the third feature alignment block outputting features ; The third feature alignment block sequentially comprises a two-dimensional convolution with a convolution kernel size of , a two-dimensional convolution with a convolution kernel size of , a ReLU activation function layer, a two-dimensional convolution with a convolution kernel size of , and a Sigmoid activation function layer. 2328), fourth set of PS block output features and first set of DS block output features input fourth feature alignment block, fourth feature alignment block outputs features ; The fourth feature alignment block sequentially comprises a two-dimensional convolution with a convolution kernel size of , a two-dimensional convolution with a convolution kernel size of , a ReLU activation function layer, a two-dimensional convolution with a convolution kernel size of , and a Sigmoid activation function layer. 2329), features and features Summing is performed, and the summing result is output as a feature of the MSFAFM layer .
6. The LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method of claim 5, wherein: The working process of each group of DS blocks in the first group of DS blocks, the second group of DS blocks, the third group of DS blocks and the fourth group of DS blocks is as follows: The feature 1 inputs a PatchEmbedding, and the PatchEmbedding outputs a feature 2; the feature 2 inputs a layer normalization layer LayerNorm, and the layer normalization layer LayerNorm outputs a feature 3; the feature 3 inputs a first dilated convolution, and the first dilated convolution outputs a feature ; the feature 3 inputs a second dilated convolution, and the second dilated convolution outputs a feature ; the feature 3 inputs a third dilated convolution, and the third dilated convolution outputs a feature ; the feature is transposed and the feature is element-wise multiplied to obtain a feature ; the feature is element-wise multiplied with the feature to obtain a feature , the feature and the feature 2 are concatenated to obtain a feature , the feature inputs a multi-layer perception MLP, and the multi-layer perception MLP outputs a feature ; the feature is output as the output of each group of PS blocks.
7. The LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method of claim 6, wherein: The sparse attention weight matrix is subjected to Softmax to obtain a matrix vector; each row of the matrix vector obtained after Softmax represents the weight of each band 1 and other bands; the number of rows of the matrix vector obtained after Softmax is the batch number, and the number of columns of the matrix vector obtained after Softmax is the number of bands. Feature 1 is input into a PatchEmbedding, and the PatchEmbedding outputs Feature 2; Feature 2 is input into a layer normalization layer LayerNorm, and the layer normalization layer Layer Norm outputs Feature 3; Feature 3 is input into a first dilated convolution, and the first dilated convolution outputs Feature ; Feature 3 is input into a second dilated convolution, and the second dilated convolution outputs Feature ; Feature 3 is input into a third dilated convolution, and the third dilated convolution outputs Feature ; Feature is transposed and Feature is element-wise multiplied to obtain Feature ; Feature and Feature are element-wise multiplied to obtain Feature ; Feature and Feature 2 are concatenated to obtain Feature ; Feature is input into a multi-layer perceptron MLP, and the multi-layer perceptron MLP outputs Feature ; Feature is output as the output of each group of DS blocks.
8. The LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method of claim 7, wherein: The 24) MSFAFM layer output feature and features The input BCAFM layer, the BCAFM layer output feature The specific process is: 241), MSFAFM layer output feature Perform layer normalization to obtain features , the features Flatten into sequence format; input the flattened sequence format into the attention mechanism to obtain the query vector , the key vector , the value vector ; 242), MSFAFM layer output feature Perform layer normalization to obtain features , the features Flatten into sequence format; input the flattened sequence format into the attention mechanism to obtain the query vector , the key vector , the value vector ; 243), the query vector , the key vector , the value vector inputting the multi-head attention mechanism, and outputting a sequence of the multi-head attention mechanism ; represented as: wherein denotes a multi-head attention mechanism; denotes a batch size; denotes the number of channels of the HSI; denotes the set of real numbers; 244), the query vector , the key vector , the value vector inputting the multi-head attention mechanism, outputting a sequence of the multi-head attention mechanism ; represented as: In the formula, represents the number of channels of the LiDAR; 245), the sequence and the sequence are concatenated , and the features are sequentially passed through two-dimensional convolution and the activation function GELU to obtain the features ; represented as: 。 9. The LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method of claim 8, wherein: The 3) BCAFM layer output feature An input dynamic band selection strategy DBSS, the dynamic band selection strategy DBSS outputs a selected band; the specific process is as follows: 31), BCAFM layer output feature input convolution kernel size is two-dimensional convolution layer, output feature ; 32), feature The input DBSM-TKSA, and the DBSM-TKSA outputs the selected waveband.
10. The LiDAR-guided multi-modal fusion dynamic hyperspectral image band selection method of claim 9, wherein: The 32) feature The input DBSM-TKSA, and the DBSM-TKSA outputs the selected waveband; the specific process is: 321), initialize global average pooling layer ; initialize linear layer for conversion; initialize learnable temperature parameter , is all-ones matrix; set scale to 0.5; initialize empty weight accumulation buffer ; initialize empty sample count buffer ; 322), from the feature extracting batch size , number of wave bands , height and width of the sample; 323), inputting the sample extracted in 322) into a global average pooling layer global average pooling layer output feature ; 324)、 Features Passing through linear layer, Reshape in turn, get query vector ; Features Pass through linear layer, Reshape, get key vector ; Features Pass through linear layer, reshape, get value vector ; Query vector The transpose of the query vector is multiplied element-wise with the key vector and a learnable temperature parameter to obtain the attention score matrix. based on a ratio 0.5, the values less than 0.5 in the sparse matrix are set to 0 to obtain a sparse attention weight matrix; The channel dimension of each column of the matrix vector obtained after the Softmax is averaged to obtain a matrix vector, and each value in the matrix vector represents the average value of each wave band of each batch; the matrix vector is multiplied by the batch number to obtain a wave band weight matrix vector; Band weight matrix vector put into weight accumulation buffer ; The batch of waveband weight matrix vector pairs is placed into the sample count buffer ; Vector summing the band weight matrix vector in the weight accumulation buffer Vector summing the band weight matrix vector in the weight accumulation buffer ; is represented as: wherein denotes the summed waveband weight matrix vector; denotes the total number of samples extracted 322) 325), the average wave band weight element-wise multiplication with the value vector element-wise multiplication with the value vector element-wise addition, and then input into a Sigmoid activation function, and the Sigmoid activation function outputs features; 326), based on The algorithm processes the Sigmoid activation function output features, selects the wave band corresponding to the first The algorithm processes the Sigmoid activation function output features, selects the wave band corresponding to the first The algorithm processes the Sigmoid activation function output features, selects the wave band corresponding to the first