Multi-source Remote Sensing Image Classification Method Based on Key Band Retrieval Attention Mechanism

By building a neural network model based on the key band retrieval attention mechanism, the problem of redundant feature interference and insufficient modal interaction in multi-source remote sensing image classification is solved, and efficient feature extraction and classification accuracy are achieved.

CN120198818BActive Publication Date: 2025-07-25OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510668878.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-25
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

There are problems in the existing multi-source remote sensing image classification methods such as redundant feature interference, easy loss of key feature information, and lack of effective interaction mechanisms between multimodals, resulting in insufficient classification accuracy and robustness.

Method used

A neural network model based on the key band search attention mechanism is adopted, and the deep fusion of hyperspectral images and Lidar/SAR data is achieved through the combination of PCA dimensionality reduction, dual-domain fusion encoder, self-attention layer, key band search module and fully connected layer, and the deep fusion of hyperspectral images and Lidar/SAR data is improved, thereby improving feature extraction and modal interaction capabilities.

Benefits of technology

It significantly improves the accuracy and robustness of multi-source remote sensing image classification, enhances the discriminant and interpretable feature selection, and improves the classification performance of the model in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198818B_ABST
    Figure CN120198818B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-source remote sensing image classification method based on a key band retrieval attention mechanism, belonging to the technical field of remote sensing image processing. First, the present invention preprocesses multi-source remote sensing data and divides the data set. Subsequently, a key band retrieval attention network is constructed, and the divided training set and validation set are used to train and evaluate the model. Finally, the test set is used to conduct a visual analysis of the model results. By effectively extracting and fusing the features of hyperspectral images and lidar / synthetic aperture radar images, the present invention effectively reduces the interference of redundant information, improves the retention rate of hyperspectral key information, enhances the complementary expression ability between multi-source data, and significantly improves the accuracy and computational efficiency of remote sensing image classification. The present invention is applicable to the application scenarios of multi-source remote sensing data fusion and classification, can meet the needs of efficient and intelligent processing of complex surface information, and provides a precise and efficient remote sensing image classification solution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technology belongs to the field of remote sensing image processing and analysis, and specifically relates to a multi-source remote sensing image classification method based on key band retrieval attention mechanism. Background Art

[0002] Multi-source remote sensing image classification refers to the use of different types of remote sensing data, such as hyperspectral images (HSI), lidar, and synthetic aperture radar (SAR), etc., to accurately identify and classify surface coverings. This task has important application values in fields such as land cover identification, ecological environment monitoring, disaster warning, resource investigation, and urban planning. In recent years, with the improvement of remote sensing sensor performance, the availability of multi-source remote sensing data has been significantly enhanced, providing a richer information source for ground object classification in complex scenes.

[0003] Hyperspectral images have continuous narrow-band spectral resolution and can capture fine-grained spectral differences of ground objects in each band, which is an important data source for achieving high-precision classification. However, HSI usually contains hundreds of bands, with significant dimensional redundancy and computational burden. Traditional dimensionality reduction methods such as principal component analysis (PCA) alleviate the computational pressure to a certain extent, but there are problems of unstable information compression and easy loss of key discriminant features. In addition, HSI is sensitive to the atmosphere, illumination, and clouds in complex scenes, affecting its robustness in actual classification tasks. To make up for the limitations of HSI, data sources such as lidar and SAR are often introduced as complementary modalities. SAR has all-weather imaging capabilities, can penetrate clouds and fog, and provide terrain structure and backscattering information; lidar can obtain three-dimensional terrain and building information at high spatial resolution. These modalities have obvious advantages in spatial structure expression and can improve the classification accuracy in areas where HSI is difficult to distinguish. However, the spectral dimension of lidar / SAR is weak, making it difficult to effectively distinguish between ground objects with similar spectral characteristics.

[0004] Therefore, fusing HSI with Lidar / SAR has become an effective way to improve the classification performance of remote sensing images. Currently, multi-source remote sensing image classification methods based on deep learning can be mainly divided into two categories: local methods and global methods. Local methods usually rely on convolutional neural networks (CNNs) and leverage their advantages in local spatial feature extraction to fuse and model multi-source data. Global methods, represented by Transformer, model the long-range dependencies between different modalities through self-attention mechanisms, more effectively integrate multi-source information, and enhance the overall representation ability. Although existing deep learning methods have made significant progress in improving the accuracy of land cover classification, due to the differences in imaging mechanisms and information expression methods among various remote sensing data, achieving efficient fusion between multi-source data still faces the following two key challenges: (1) The problem of redundant feature interference and key feature loss: HSI has a high band dimension, and directly participating in modeling will lead to an excessive computational burden and introduce a large number of redundant features, interfering with model learning. Although the commonly used PCA dimensionality reduction can alleviate the redundancy problem, it lacks a selective retention mechanism for task-critical information and may cause compression and loss of discriminative bands; (2) The lack of an interaction mechanism between multi-modalities: Existing methods mainly focus on independent extraction and simple fusion, lacking explicit cross-modal interaction strategies, making it difficult to fully explore the complementary relationship between HSI and Lidar / SAR data, thereby affecting classification performance and result interpretability. Summary of the Invention

[0005] To address the problems commonly existing in existing deep learning-based multi-source remote sensing image classification methods, such as redundant feature interference, easy loss of key feature information, and the lack of an effective interaction mechanism between multi-modalities, it is urgent to construct a fusion method with efficient feature extraction capabilities and an explicit modality interaction mechanism to achieve the collaborative optimization of multi-source remote sensing data while retaining key spectral information, so as to improve the accuracy, robustness, and efficiency of multi-source remote sensing image classification.

[0006] The purpose of the present invention is to provide a multi-source remote sensing image classification method based on a key band retrieval attention mechanism, aiming to achieve the deep fusion of hyperspectral images and Lidar / SAR data, effectively improve the accuracy and robustness of multi-source remote sensing image classification, and make up for the deficiencies of the existing technology.

[0007] To achieve the above purpose, the technical solutions implemented in the present invention are as follows:

[0008] A multi-source remote sensing image classification method based on a key band retrieval attention mechanism, comprising the following steps:

[0009] S1: Collect multi-source remote sensing data of hyperspectral image HSI and Lidar / SAR, and preprocess the data;

[0010] S2: Construct a key band retrieval attention neural network model, which includes a dual-domain fusion encoder module, a self-attention layer, a key band retrieval module, a key band retrieval attention module, and a fully connected layer for classification; the data processing process of this network model is expressed as: input data → dual-domain fusion encoder module → self-attention layer → key band retrieval module → key band retrieval attention module → fully connected layer → output classification result.

[0011] S3: Input the preprocessed data into the key band retrieval attention neural network model for model training; during the training process, cross-entropy is used as the loss function, and the Adam optimizer is used for parameter optimization; after training is completed, use the trained key band retrieval attention neural network model for classification prediction and comprehensively evaluate the model performance;

[0012] S4: Use the trained key band retrieval attention neural network model to process the data to be measured, and visualize and analyze the classification results of the multi-source remote sensing images to assist in understanding the discrimination ability of the model.

[0013] Furthermore, the preprocessing in S1 includes: performing spectral dimension compression, data padding, unifying the numerical distribution of each modality data (i.e., normalization processing), constructing sample blocks, and introducing data augmentation operations on the hyperspectral image; and dividing the preprocessed data into a training set, a validation set, and a test set.

[0014] Even further, S1 specifically includes the following steps:

[0015] S1-1: Input the original hyperspectral image data and Lidar / SAR data, and use the principal component analysis (PCA) method to perform spectral dimension reduction on the hyperspectral data, so as to retain the main information features and reduce the redundant dimensions, and at the same time obtain the PCA weight matrix ;

[0016] S1-2: Perform edge padding (padding) operations on the original hyperspectral data, the hyperspectral data processed by PCA, and the Lidar / SAR data to ensure the integrity of the data in the edge area during the subsequent sliding window extraction process;

[0017] S1-3: Perform normalization processing on the processed hyperspectral data and Lidar / SAR data to make them conform to the standard normal distribution with a mean of 0 and a standard deviation of 1, so as to improve the stability and convergence speed of model training;

[0018] S1-4: Perform sliding sampling on the processed multi-source data with a window of a specified size, and divide the entire image into multiple data blocks of a fixed size, so as to construct a set of sample blocks for training and testing;

[0019] S1-5: Perform data augmentation operations on the constructed hyperspectral and Lidar / SAR sample blocks, and adopt strategies such as random horizontal flipping and vertical flipping to improve the generalization ability and robustness of the model;

[0020] S1-6: Annotate and partition the preprocessed samples according to the ground truth labels to construct a training set and a validation set. Since the ground truth labels do not annotate all sample points, an additional test data set containing all sample points is reserved for the subsequent visualization display and analysis of the classification results. Synchronously generate an index file to record the spatial coordinate information of the sample points in the training set, validation set, and test set to support the efficient loading and calling during the model training, evaluation, and visualization stages.

[0021] Further, in S2, the data processing flow of the key band retrieval attention neural network model is as follows:

[0022] S2-1: The inputs of the key band retrieval attention neural network model are the original HSI data, the hyperspectral data after PCA processing, and the Lidar / SAR data. The dimension of the HSI data is [B, C, N1, H, W], where B is the batch size, indicating the number of samples passed into the model in parallel, C is the number of channels (usually 1), N1 is the number of hyperspectral bands, and H and W respectively represent the height and width of the image block, which are equal to the size of the sliding window in the preprocessing stage. The dimension of the hyperspectral data after PCA processing is [B, C, N, H, W], and the meaning of each dimension is the same as that of the original HSI data. The dimension of the Lidar / SAR data is [B, C2, H, W], where B is the batch size, C2 is the number of channels, and H and W are the length and width of the samples. To achieve the consistent processing of multi-modal feature dimensions, the original HSI data and the HSI data after PCA processing are respectively reshaped. After reshaping, their dimensions are [B, C1, H, W] and [B, C, H, W] respectively, which are aligned with the dimension of the Lidar / SAR data, where C1 = C × N1, C pca , H, W], and the meaning of each dimension is the same as that of the original HSI data. The dimension of the Lidar / SAR data is [B, C2, H, W], where B is the batch size, C2 is the number of channels, and H and W are the length and width of the samples. To achieve the consistent processing of multi-modal feature dimensions, the original HSI data and the HSI data after PCA processing are respectively reshaped. After reshaping, their dimensions are [B, C1, H, W] and [B, C, H, W] respectively, which are aligned with the dimension of the Lidar / SAR data, where C1 = C × N1, C pca , H, W], and the meaning of each dimension is the same as that of the original HSI data. The dimension of the Lidar / SAR data is [B, C2, H, W], where B is the batch size, C2 is the number of channels, and H and W are the length and width of the samples. To achieve the consistent processing of multi-modal feature dimensions, the original HSI data and the HSI data after PCA processing are respectively reshaped. After reshaping, their dimensions are [B, C1, H, W] and [B, C, H, W] respectively, which are aligned with the dimension of the Lidar / SAR data, where C1 = C × N1, C pca = C × N pca .

[0023] S2-2: After the data enters the key band retrieval attention neural network model, first perform feature extraction on the HSI data. For the original HSI data and the HSI data after PCA processing Feature extraction is performed separately through the dual-domain fusion encoder module. The dual-domain fusion encoder module consists of two branches, the spatial domain and the frequency domain. The processing flow of the spatial domain branch is as follows: HSI data → Depthwise Convolution (DWConv) → BatchNorm2D → ReLU activation function → Spatial domain features In the frequency domain branch, a dynamic frequency filtering module is designed, and its processing flow is as follows: HSI data → Fast Fourier Transform (FFT) → Adaptive learnable mask → Hadamard product → Inverse Fast Fourier Transform (IFFT) → Frequency domain features Finally, the two are added together to obtain the dual-domain fusion features The entire calculation process is as follows:

[0024] ;

[0025] ;

[0026] ;

[0027] Among them, can represent the original HSI data or the HSI data after PCA processing , represents the learning process of the adaptive learnable mask, represents the Hadamard product. The learning process of the adaptive learnable mask can be expressed as: Data after Fast Fourier Transform → Linear layer → ReLU activation function → Linear layer → Mask. The original HSI data and the HSI data after PCA processing respectively obtain the dual-domain fusion features through the dual-domain fusion encoder module and .

[0028] S2-3: Perform simple feature extraction on the Lidar / SAR data This process uses a simple two-dimensional convolution structure to obtain its spatial structure information. The specific process is as follows: Lidar / SAR data → 2D Convolution (2DConv) → BatchNorm2D → ReLU activation function → Lidar / SAR data features The calculation process of this feature extraction operation is expressed as:

[0029] ;

[0030] S2-4: For the HSI data features and the HSI data features after PCA processing Perform a dimension reshaping operation to flatten its spatial dimensions H and W to obtain [B, C1, P] and [B, C pca , P], and at the same time, the Lidar / SAR data features also perform the same operation to reshape the dimensions to [B, C2, P], where P = H × W, indicating the number of pixels contained in each sample. After completing the flattening operation, and are concatenated along the channel dimension to obtain the intermediate state feature , whose dimension is [B, C pca +C2, P]. This fused feature combines the spectral information and spatial structure information after dimensionality reduction, providing cross-modal input support for the subsequent key band retrieval attention mechanism. The calculation process is as follows:

[0031] ;

[0032] S2-5: Enter the N SA layer self-attention layer to model long-range dependencies, which is used to model the long-range dependencies in the channel dimension and enhance the interaction representation ability between hyperspectral and Lidar / SAR. The specific process is as follows: Input data →Depth convolution→Self-attention mechanism→Residual connection→Feed-forward neural network→Residual connection→Processed feature . This module significantly enhances the and interaction ability through the long-range modeling mechanism, providing global dependency modeling ability for subsequent key band screening and classification. The entire calculation process is as follows:

[0033] ;

[0034] ;

[0035] ;

[0036] ;

[0037] Among them, Q, K, and V are the query, key, and value in the self-attention mechanism respectively, is the similarity matrix of the self-attention mechanism, and its dimension is [B, C pca +C 2, C pca+C2]. The FFN represents a feed-forward neural network, and its specific process is as follows: input data → 2D convolution (2DConv) → ReLU activation function → 2D convolution (2DConv) → Dropout → output data.

[0038] S2-6: Original HSI data features Enter the key band retrieval module to perform the retrieval of the first key band. This process is based on the attention similarity matrix generated in S2-5 to infer the important bands in the original HSI. First, extract the submatrix corresponding to the columns from the Cth pca+1 to the (C + C2)th column and the first C pca rows from the similarity matrix, denoted as pca , with a dimension of [B, C , C2]. The elements of the extracted similarity matrix represent the similarities between the HSI data features after PCA processing pca for each band and the Lidar / SAR data features. Next, perform hidden correlation Softmax on the similarity matrix to complete the band retrieval. The specific steps are as follows: input the similarity matrix and the HSI data features → perform Softmax normalization on the similarity matrix to obtain → perform Top-K retrieval on all elements in the similarity matrix to obtain the band indices of the top k% in similarity → retrieve obtained in S1-1 according to the band indices to obtain → perform Softmax normalization on to obtain → perform Top-K retrieval on to obtain the band indices of the top k% in similarity of the original HSI data → retrieve the key band features of the original HSI data through the indices → The key band features have a dimension of [B, k%C1, P]. The entire calculation process is as follows: The key band features ;

[0039] ;

[0040] ;

[0041] ;

[0042] ;

[0043] ;

[0044] ;

[0045] ;

[0046] Among them, represents rounding down. This module realizes the step-by-step screening of the most discriminative key bands from the full-band HSI features, providing information-compressed and highly discriminative input features for subsequent classification tasks.

[0047] S2-7: Key band features and intermediate state features are jointly input into the N KA -layer key band retrieval attention module for processing. This module designs a multi-source channel cross-sparse attention mechanism to further fuse the key band features of the original HSI with the global features extracted cross-modally, thereby improving the retrieval accuracy and discriminative ability. The specific process is as follows: input data and → and are concatenated along the channel dimension → depth convolution → cross-attention mechanism → residual connection → feed-forward neural network → residual connection → updated intermediate state features . The entire calculation process is as follows:

[0048] ;

[0049] ;

[0050] ;

[0051] ;

[0052] Among them, Q, K, and V are the query, key, and value in the multi-source channel cross-sparse attention mechanism respectively, is the similarity matrix, and its dimension is [B, C pca +C 2, C pca +C2+C1]. FFN represents the feed-forward neural network, and its specific process is: input data → 2D convolution (2DConv) → ReLU activation function → 2D convolution (2DConv) → Dropout → output data. The above similarity matrix and the original HSI features Send them into the aforementioned key band retrieval module (refer to step S2-6) together, and further perform spectral band screening based on the updated attention relationship. After passing through N KA layers of key band retrieval attention modules, the final key band features are obtained .

[0053] S2-8: Finally, process the key band features into classification results. The specific process is as follows: Input the key band features →max pooling layer→Dropout→fully connected layer→classification result . The entire calculation process is as follows:

[0054] ;

[0055] Among them, represents the max pooling layer, represents the fully connected layer.

[0056] Furthermore, in the aforementioned S3, model training and evaluation include:

[0057] S3-1: Input the training set obtained by partitioning in step S1 into the constructed key band retrieval attention neural network model. The inputs received by the model include original hyperspectral image data, hyperspectral data after PCA processing, and Lidar / SAR data, and perform feature extraction and fusion processing through each module of the neural network.

[0058] S3-2: During the training process, use cross entropy as the loss function to measure the difference between the predicted probability distribution of the model output categories and the true labels. The optimizer selects the Adam (Adaptive Moment Estimation) optimization algorithm, and effectively minimizes the loss function by adaptively updating the learning rate, thereby improving the learning effect of the model. The initial learning rate is set to 0.001, the batch size is 64, and the maximum number of iterations (Epochs) is 100. To prevent the model from overfitting during training, an early stopping mechanism is introduced. When the validation set loss does not decrease significantly within 10 consecutive epochs, the training is automatically terminated. The overall training strategy combines loss function optimization, adaptive scheduling, and regularization means, providing a strong guarantee for the improvement of model performance. The mathematical expression of the cross entropy loss function is as follows:

[0059] ;

[0060] Among them, N is the number of samples, C is the number of categories, Indicates the distribution of the true label of the i-th sample in the j-th class (usually 0 or 1). Indicates the predicted probability of the i-th sample in the j-th class.

[0061] S3-3: After training is completed, input the validation set into the trained key band retrieval attention neural network model for classification prediction. The model outputs the land cover class labels to which each test sample belongs.

[0062] S3-4: Evaluate the classification results on the validation set. Use the overall accuracy (OA), average accuracy (AA), and Kappa coefficient metrics to measure the model performance.

[0063] Furthermore, the S4 includes: Input the test data set into the trained key band retrieval attention neural network model, perform the forward propagation operation, and obtain the predicted classification results of each pixel or sample. Restore and splice the classification labels generated by the model for the test set to reconstruct the complete classification result image. Ensure that the output result is spatially consistent with the original remote sensing image for subsequent visualization and regional-level analysis and processing. Use the class-color mapping rule to visually display the classification map predicted by the model. Different classes are represented by different colors, making the distribution of land covers clear at a glance and facilitating the user to intuitively observe the classification accuracy and spatial continuity. By comparing the model classification results with the true label (Ground Truth) image, analyze the discrimination ability of the model between different land classes and identify the recognition accuracy of regional boundaries and the performance of easily confused classes.

[0064] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0065] 1. Improve the ability to retain key bands and enhance the discriminability and interpretability of feature selection. Existing methods usually only use principal component analysis (PCA) to compress the dimensions of hyperspectral images. Although it can reduce the computational cost, it is easy to cause key bands containing discriminant information to be compressed or even lost, seriously affecting the classification accuracy of the model and the interpretability of subsequent results. The present invention proposes a key band attention retrieval mechanism. Based on the initial PCA principal component analysis, a cross-modal attention-based band correlation analysis method is introduced. Using the similarity matrix of Lidar / SAR and PCA-HSI, calculate the high-correlation principal component index, and reverse infer it to the original HSI band based on the PCA weight. Finally, accurately retrieve the bands with the strongest discriminative power in the original spectral dimension. This mechanism not only has high selectivity and traceability but also significantly improves the model's utilization efficiency of key information, and has better spectral interpretability and engineering practicability.

[0066] 2. Explicit interaction modeling between multiple modalities is realized to fully exploit the complementarity between Lidar / SAR and HSI data. Most existing multi-source remote sensing image fusion methods mainly focus on feature-level stitching, stacking, or cascading, and fail to effectively model the interaction relationships between different modalities. The present invention constructs a dual-domain fusion encoder module and a key band retrieval attention module. The dual-domain fusion encoder module decouples and extracts features from the input data in the spatial domain and the frequency domain respectively to mine stable features from different spaces. The key band retrieval attention module adopts a multi-source cross-channel mechanism to guide the selection of key bands to align with Lidar / SAR, realizing the dynamic guidance and fusion of Lidar / SAR to HSI, thereby improving the information interaction quality. This structural modeling method not only enhances the coupled expression ability between features but also greatly improves the adaptability and fault tolerance of the model in heterogeneous modalities.

[0067] 3. An end-to-end trainable neural network framework is constructed, taking into account accuracy, efficiency, and transfer ability. The present invention constructs a complete end-to-end neural network architecture from data preprocessing, feature extraction, attention guidance, band selection to final classification output. Each module is designed with the advantages of high structure, trainable parameters, and transferable gradients. During the training process, the model uses cross-entropy loss, Adam optimizer, combined with an early stopping strategy to effectively suppress overfitting and improve training efficiency. In the testing stage, the model supports reasoning and visualization analysis of the full-image classification situation, and the output results are aligned with the original image spatially, facilitating on-site analysis and evaluation.

[0068] The present invention realizes the classification of multi-source remote sensing data by effectively extracting and fusing the features of hyperspectral images (HSI) and Light Detection and Ranging (LiDAR) / Synthetic Aperture Radar images (SAR). This technology is applicable to the fusion and classification scenarios of multi-source remote sensing data, can meet the remote sensing data processing requirements for complex and diverse surface information, and has significant advantages in improving classification accuracy and processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a schematic flowchart of an embodiment of the present invention.

[0070] Figure 2 It is a schematic principle flowchart of an embodiment of the present invention.

[0071] Figure 3 It is a schematic diagram of the neural network structure described in an embodiment of the present invention.

[0072] Figure 4 It is a detailed structural schematic diagram of the dual-domain fusion encoder module in an embodiment of the present invention.

[0073] Figure 5Schematic diagram of the detailed structure of the key band retrieval attention module in the embodiments of the present invention; wherein, (a) is the key band retrieval attention module, (b) is the multi-source channel cross sparse attention mechanism; (c) is the key band retrieval module.

[0074] Figure 6 Schematic diagram of the visualization images of multiple algorithms on the Augsburg dataset in the embodiments of the present invention.

[0075] Figure 7 Schematic diagram of the visualization images of multiple algorithms on the Berlin dataset in the embodiments of the present invention.

[0076] Figure 8 Schematic diagram of the visualization images of multiple algorithms on the Houston2013 dataset in the embodiments of the present invention. Specific implementation manners

[0077] The present invention will be further explained and illustrated below through specific embodiments in conjunction with the accompanying drawings.

[0078] Embodiment 1:

[0079] A multi-source remote sensing image classification method based on a key band retrieval attention mechanism, referring to Figure 1 , includes the following steps:

[0080] (1) Preprocess the multi-source remote sensing data and divide it into a training set and a test set.

[0081] (1.1) Data loading and principal component analysis (PCA) dimensionality reduction processing. Load the original hyperspectral image data and Lidar / SAR data into the memory. The dimension of the original HSI data is [H, W, N1], where H and W are the image height and width respectively, and N1 is the number of bands; the dimension of the Lidar / SAR data is [H, W, C2], where C2 represents its number of channels. To reduce the redundancy and computational burden of the hyperspectral data, apply the principal component analysis (PCA) method to the original HSI data, retain the first N PCA principal components to form the dimensionality-reduced hyperspectral image data, denoted as PCA-HSI, with a dimension of [H, W, N PCA . At the same time, record the weight matrix W PCA obtained by PCA dimensionality reduction for subsequent key band back-projection and retrieval.

[0082] (1.2)Edge Padding. Considering that the subsequent sample blocks are obtained by the sliding window extraction method, to avoid the loss of edge information during the sliding window process, mirror padding operations are performed on the original HSI data, PCA-HSI data, and Lidar / SAR data in the spatial dimensions (height and width). The width of the padded edge is set according to the window size S and the sliding step, usually half of the window size, to ensure the sample integrity of the image edge region.

[0083] (1.3)Normalization. To improve the convergence speed and stability of model training, the above-processed multi-source remote sensing data is normalized. The normalization method uses the standard normal distribution transformation to uniformly adjust the pixel values of each channel or band to a distribution with a mean of 0 and a standard deviation of 1. The formula is:

[0084] ;

[0085] where represents the original data, is the mean of this channel.

[0086] (1.4)Sample Block Construction. On the normalized multi-source data, sliding sampling is performed according to the preset sliding window size (e.g., 11×11 pixels) to construct a set of sample blocks covering the entire image. Each sample block is composed of an HSI block, a PCA-HSI block, and a Lidar / SAR block at the same position and is fed as input into the subsequent neural network model. Each sample block label is determined by the true ground object category corresponding to its central pixel.

[0087] (1.5)Data Augmentation Processing. To improve the robustness and generalization ability of the model, data augmentation strategies are applied to the constructed sample blocks, mainly including operations such as random horizontal flipping and vertical flipping, to enhance the diversity of samples and the model's tolerance to deformation.

[0088] (1.6)Data Partitioning and Index Generation. After the construction and enhancement of sample blocks of multi-source remote sensing data are completed, in order to facilitate the effective development of model training, validation, and testing, the labeled samples are reasonably partitioned. The specific partitioning strategy is as follows: First, according to the ground truth provided by each dataset, the samples are classified and labeled. Among them, the label of each sample is determined by the class number corresponding to the pixel position in the image where it is located. To ensure that the class distribution between the training set and the validation set is representative and balanced, the samples are partitioned using the method of stratified random sampling. Since there are unlabeled areas in remote sensing data, to avoid waste of samples, an additional full-image sample is reserved as the test set for the model to perform inference and visualization after training is completed. Spatial coordinate index files for the training set, validation set, and test set are generated synchronously to facilitate efficient positioning and visualization during subsequent model training and classification result restoration. The sample quantities of the Augsburg, Berlin, and Houston2013 datasets used are detailed in Tables 1, 2, and 3 respectively.

[0089] Table 1 Numbers of Training and Validation Samples in the Augsburg Dataset

[0090]

[0091] Table 2 Numbers of Training and Test Samples in the Berlin Dataset

[0092]

[0093] Table 3 Numbers of Training and Test Samples in the Houston2013 Dataset

[0094]

[0095] (2)Construct a Key Band Retrieval Attention Neural Network Model.

[0096] The overall structure of this model, as Figure 3 shown, is overall modularly designed and consists of an input layer, a dual-domain fusion encoder module, a self-attention layer, a key band retrieval module, a key band retrieval attention module, and a classification output module, etc., to achieve deep fusion and collaborative modeling of hyperspectral images (HSIs) and Lidar / SAR data.

[0097] (2.1) Model Input and Dimension Reconstruction. The neural network model constructed by the present invention simultaneously receives three types of input data, namely, original hyperspectral image data, hyperspectral image data after PCA dimensionality reduction processing, and Lidar / SAR data. The dimension of the HSI data is [B, C, N1, H, W], where B is the batch size, representing the number of samples passed into the model in parallel, C is the number of channels (usually 1), N1 is the number of hyperspectral bands, and H and W respectively represent the height and width of the image block, which are equal to the size of the sliding window in the preprocessing stage. The dimension of the hyperspectral data after PCA processing is [B, C, N pca , H, W], and the meaning of each dimension is the same as that of the original HSI data. The dimension of the Lidar / SAR data is [B, C2, H, W], where B is the batch size, C2 is the number of channels, and H and W are the length and width of the samples. To achieve consistent processing of multi-modal feature dimensions, the original HSI data and the HSI data after PCA processing are respectively dimensionally reshaped, and their dimensions after reshaping are [B, C1, H, W] and [B, C pca , H, W], which are aligned with the dimension of the Lidar / SAR data, where C1 = C × N1, C pca = C × N pca .

[0098] (2.2) Feature Extraction of the Dual-Domain Fusion Encoder Module. The structure of the dual-domain fusion encoder module is as Figure 4 shown. After the data enters the key band retrieval attention neural network model, the hyperspectral image data is first subjected to feature extraction processing. The present invention adopts a dual-domain fusion encoder module to respectively perform feature modeling on the original hyperspectral image data and the HSI data after PCA processing . This module consists of a spatial domain branch and a frequency domain branch, which can simultaneously capture the discriminant information of the image at two levels of spatial structure and spectral energy distribution, thereby enhancing the expression ability and robustness of the features.

[0099] In the spatial domain branch, the input hyperspectral image data first extracts its local spatial texture structure through depthwise separable convolution, and then sequentially passes through the BatchNorm2D normalization layer and the ReLU activation function to obtain the spatial domain features . Its calculation process can be expressed as:

[0100] ;

[0101] where, can represent the original hyperspectral image or the image after PCA processing .

[0102] In the frequency domain branch, first, the input data is applied with the fast Fourier transform (FFT) to obtain its spectral representation. To achieve adaptive modeling of frequency domain information and reduce computational costs, the present invention introduces a dynamic frequency filtering module based on the convolution theorem. According to this theorem, the convolution of functions corresponds to the Hadamard product of their Fourier transforms in the Fourier domain. The core of this module is the adaptive learnable mask . It is used to effectively dynamically modulate the spectral features, thereby enhancing the model's ability to extract key information. The mask generation process includes two layers of linear mapping and a non-linear activation function. The input is a frequency domain tensor, and the output is a response coefficient matrix of the same size, representing the importance of each frequency component. The mask learning process is expressed as:

[0103] ;

[0104] After obtaining the mask , it is subjected to the Hadamard product with the spectrum, and then restored to the image domain through the inverse Fourier transform (IFFT) to form the frequency domain feature , which is specifically expressed as:

[0105] ;

[0106] wherein, represents the Hadamard product operation.

[0107] Finally, the spatial domain feature and the frequency domain feature are fused by element-wise addition to form the final dual-domain fusion feature representation :

[0108] ;

[0109] This fusion feature combines the capabilities of local structure preservation and global frequency response modulation, and can effectively enhance the discriminative ability of subsequent models in key band selection and multi-modal interaction modeling. In actual operation, the original hyperspectral image data and the data after PCA dimensionality reduction respectively pass through the above dual-domain fusion encoder module independently, and finally obtain two corresponding fusion feature tensors and .

[0110] (2.3) Lidar / SAR data feature extraction. To make full use of the spatial structure information provided by Lidar or SAR remote sensing data, in the key band retrieval attention neural network model of the present invention, for SAR / Lidar data A lightweight feature extraction path is designed. The specific process is: Lidar / SAR data → 2D convolution (2DConv) → BatchNorm2D → ReLU activation function → Lidar / SAR data features . The calculation process of this feature extraction operation is expressed as:

[0111] .

[0112] (2.4) Feature flattening and self-attention modeling. After obtaining the original HSI features , PCA-HSI features and Lidar / SAR features , first, a flattening operation is performed on the spatial dimensions [H, W] of the three, and they are uniformly reshaped into a sequence form. Among them and are respectively transformed into dimensions of [B, C1, P] and [B, C pca , P], is reshaped into [B, C2, P], where P = H × W represents the number of pixels contained in each sample.

[0113] Subsequently, and are concatenated along the channel dimension to construct a fused feature , whose dimension is [B, C pca + C2, P]. This feature contains both spectral dimensionality reduction information and spatial structure information and is used as the input for cross-modal interaction modeling. The calculation expression is as follows:

[0114] .

[0115] To enhance the collaborative expression ability between multiple modalities, the fused feature is input into the N SA -layer self-attention layer for long-range dependence modeling in the channel dimension. This module adopts a self-attention mechanism with a query-key-value (QKV) structure and integrates a feed-forward network to enhance the feature expression ability. The specific process is as follows:

[0116] ;

[0117] ;

[0118] ;

[0119] ;

[0120] Among them, Q, K, and V are the query, key, and value in the self-attention mechanism respectively, is the similarity matrix of the self-attention mechanism, with dimensions [B, C pca +C 2, C pca +C2]. FFN represents the feed-forward neural network, and its specific process is: input data → 2D convolution (2DConv) → ReLU activation function → 2D convolution (2DConv) → Dropout → output data.

[0121] (2.5) Key band retrieval module (first key band screening). After completing the attention modeling, the model starts to perform the first retrieval task of key bands in the original hyperspectral image. The processing flow of the key band retrieval module is as shown in Figure 5 (c). Using the original hyperspectral image features as input, combined with the self-attention similarity matrix constructed in S2-5 , the original hyperspectral bands most relevant to Lidar / SAR are inferred from it.

[0122] First, extract the cross-modal similarity sub-matrix between PCA-HSI and Lidar / SAR data from the similarity matrix , denoted as , with dimensions [B,C pca ,C2], that is, select the sub-matrix corresponding to the C pca+1 th to the C pca +C2th columns, and the first C pca rows:

[0123] ;

[0124] Each element of this sub-matrix represents the attention weight between each principal component channel of PCA-HSI and the Lidar / SAR channel, used to measure their cross-modal correlation.

[0125] Subsequently, perform hidden correlation Softmax on and complete the band retrieval accordingly. First, perform the Softmax operation to obtain the normalized similarity matrix :

[0126] ;

[0127] Then, according to the global attention distribution in , adopt the Top-K strategy to select the top k% of PCA principal component channels with the highest cross-modal similarity, and obtain the corresponding channel index set :

[0128] ;

[0129] Next, call the PCA weight matrix generated in the preprocessing stage (1.1) , and extract the corresponding weight sub-matrix according to the above index :

[0130] .

[0131] Furthermore, perform hidden correlation Softmax on . First, perform softmax normalization:

[0132] ;

[0133] Subsequently, perform a Top-K operation on it to select the set of band indices with the most cross-modal correlation in the original HSI :

[0134] .

[0135] Finally, through the index select the corresponding key band channels from the original hyperspectral feature tensor to form the retrieved band feature tensor , that is:

[0136] ;

[0137] where represents rounding down.

[0138] Through the retrieval strategy of this module, the model can effectively screen out the discriminative bands strongly related to Lidar / SAR from the original high-dimensional hyperspectral image, significantly compress the input dimension while retaining the key semantic information, and provide an efficient and compact feature representation for subsequent classification prediction.

[0139] (2.6) Key band retrieval attention module. After completing the first key band selection, to further enhance the deep fusion expression between the original hyperspectral image and Lidar / SAR data, the present invention designs a key band retrieval attention module for jointly modeling the key band features of the original HSI and the globally fused features extracted cross-modally. The structure of the key band retrieval attention module is as shown in (a) in Figure 5 . This module takes the key band feature and the intermediate state feature output by self-attention as inputs, constructs a multi-source channel cross-sparse attention mechanism to improve the band selection accuracy and classification discrimination ability. First, splice the two input feature tensors along the channel dimension to form a joint fusion feature, and then perform attention calculation in combination with the query-key-value (QKV) structure. As shown inFigure 5 As shown in (b) of it, the specific processing flow is as follows: The query vector (Q) is generated by intermediate features through depthwise separable convolution; the key (K) and value (V) are generated from the fused features after concatenation. The calculation expression is as follows:

[0140] ;

[0141] Subsequently, the response relationship between channels is calculated through the dot-product attention mechanism to obtain the similarity matrix :

[0142] ;

[0143] Then, the value vector V is weighted and fused by the attention matrix and residual connection is performed to form an attention-enhanced feature representation:

[0144] .

[0145] Next, it is input into the feed-forward neural network (FFN) for non-linear mapping and channel reconstruction, and finally the updated intermediate features are output:

[0146] ;

[0147] Among them, the specific structure of the FFN is: input data → 2D convolution (2DConv) → ReLU activation function → 2D convolution (2DConv) → Dropout → output data.

[0148] Through the modeling process of N KA layers, the model can achieve dynamic alignment and mutual enhancement between the key band features of the original HSI and cross-modal information. The attention matrix calculated in this stage and the original hyperspectral features are input into the key band retrieval module together, and based on the updated cross-modal relationship, further fine screening of the key bands can be performed to obtain the final key band features .

[0149] This module realizes an explicit cross-modal attention mechanism in its structure, guides the information compression process of Lidar / SAR for HSI, significantly improves the accuracy and interpretability of spectral band selection, and provides a high-quality feature basis for the final classification task.

[0150] (2.7) Classification output module. After the final screening of the key band features, the obtained feature tensor The input classification output module is used to generate the final prediction results of ground object categories. This module aims to converge and classify the compressed high-dimensional spectral information, and adopts a lightweight structure to achieve efficient inference and stable output.

[0151] First, perform a max pooling operation (MaxPooling) on the key band features to compress the spatial dimension while keeping the channel dimension unchanged, enhancing the translational invariance of the features. Subsequently, a Dropout layer is introduced to reduce the risk of overfitting. Finally, the processed features are input into the fully connected layer to output the class probability distribution of each sample, obtaining the final classification result. The computational expression of the above process is as follows:

[0152] ;

[0153] where represents the max pooling layer, represents the fully connected layer.

[0154] (3) Train and evaluate the model to obtain the final classification result.

[0155] (3.1) Model training stage. First, input the training set obtained after preprocessing and division into the model. The input includes the original hyperspectral image data, the hyperspectral data after PCA dimensionality reduction, and the lidar / synthetic aperture radar data, which respectively correspond to different input channels of the model. Inside the network, it successively passes through the dual-domain fusion encoder module, the key band retrieval attention module, and the fully connected classification layer, thus completing the feature extraction, fusion, and classification of multi-source remote sensing images. During the training process, the cross-entropy loss function (Cross Entropy Loss) is used to measure the difference between the model prediction result and the true label, and the specific definition is as follows:

[0156] ;

[0157] where N is the number of samples, C is the number of categories, represents the distribution of the true label of the i-th sample in the j-th class (usually 0 or 1), represents the predicted probability of the i-th sample in the j-th class. The Adam (Adaptive Moment Estimation) algorithm is used as the optimizer to update the parameters. The initial learning rate is set to 0.001, and the exponential decay rate of the first-order moment estimation and the exponential decay rate Hyperparameter configurations such as these. To improve the efficiency and stability of the training process, an early stopping mechanism is introduced during model training. When the validation set loss does not decrease significantly within 10 consecutive training epochs, the training is automatically terminated. The batch size is set to 64, and the maximum number of training rounds is 100.

[0158] (3.2) Model validation phase. To comprehensively evaluate the performance of the proposed method in terms of classification accuracy, balance, and consistency, the validation set is input into the model of the current training round for inference, obtaining the predicted labels of each sample and comparing them with their true labels, thereby calculating multiple performance metrics. The specific evaluation metrics include overall classification accuracy (OverallAccuracy, OA), average classification accuracy (Average Accuracy, AA), and Kappa coefficient.

[0159] First, the overall classification accuracy (Overall Accuracy, OA) is the most commonly used basic evaluation metric, which is used to measure the overall recognition ability of the model for all validation samples. It is calculated as the ratio of the number of samples correctly classified by the model to the total number of samples in the validation set. Let the total number of samples in the validation set be , and the number of samples correctly classified be , then the overall classification accuracy is defined as:

[0160] ;

[0161] The higher the value of this metric, the stronger the classification ability of the model for the overall data. However, it should be noted that when the sample classes are severely imbalanced, OA may be dominated by the large classes, masking the classification effect of the model on the small classes. Therefore, its evaluation range mainly reflects the overall recognition accuracy of the model.

[0162] Secondly, the average classification accuracy (Average Accuracy, AA) focuses on the recognition ability of the model for each class and measures the classification balance between different classes. It is calculated as the average of the single-class classification accuracies of all classes. Let there be classes in the validation set, and the th class has samples, and the number of samples correctly classified is , then the single-class accuracy of this class is , and the calculation formula for AA is:

[0163] ;

[0164] This metric can overcome the deficiency of the overall accuracy in terms of bias towards large classes, and can better reflect the recognition performance of the model on minority classes or easily confused classes. The higher the AA value, the more balanced the classification performance of the model among various classes, which is particularly important for remote sensing image classification tasks with uneven distribution of ground object classes.

[0165] Finally, the Kappa coefficient is an evaluation metric that takes into account the accidental agreement between the classification results and the true labels, and is used to measure the degree of consistency between the model's prediction results and the ground truth labels. This metric incorporates the errors caused by random prediction and is commonly used to measure the robustness and reliability of a classifier. The formula for calculating the Kappa coefficient is:

[0166] ;

[0167] where, represents the actual agreement rate of the model (i.e., OA), while represents the theoretical agreement rate under complete random guessing. Specifically, 's calculation depends on the number of samples of each class predicted by the model and the class distribution in the true labels, and the formula is as follows:

[0168] ;

[0169] where, represents the number of true samples of the th class, represents the number of samples predicted by the model as the th class, is the total number of samples in the validation set. The value range of the Kappa value is usually [−1,1], where 1 represents perfect agreement, 0 represents agreement equivalent to random classification, and negative values represent a systematic deviation between the model's prediction and the true labels. Therefore, the Kappa coefficient can effectively reflect the prediction stability of the model and its ability to model the global data structure.

[0170] To ensure that the network model can accurately save its current state when achieving optimal performance and has good reusability and transfer ability, this embodiment introduces a stable and efficient model saving mechanism. This mechanism uses the classification performance on the validation set as the evaluation criterion and adopts a performance-driven strategy to dynamically save model parameters. After each training cycle, the system will compare the overall classification accuracy of the current model on the validation set with the historical best result. If the current metric is better than the historical record, it will immediately trigger the model saving operation and save all the trainable parameters of the current network as the "optimal model". In addition, to enhance the robustness and debugging flexibility of the model, the system will also automatically save the current model state every fixed number of training epochs (such as every 5 Epochs) to form an intermediate checkpoint, which is convenient for quick recovery in case of training interruption or parameter rollback. The model saving mechanism also operates in coordination with the early stopping strategy. When the system detects that the performance of the model on the validation set has not improved significantly for multiple consecutive training cycles (such as 10 rounds), it will automatically terminate the training process and retain the model parameters with the best evaluated performance, ensuring that the final output model has the best generalization ability. This mechanism not only improves the reliability and stability of the training process but also provides a solid technical foundation for application scenarios such as loading and calling, transfer learning, and deployment optimization of the model in the subsequent testing stage.

[0171] (3.3) Model testing phase. First, the preprocessed test set data is input into the saved optimal model. The inputs received by the model include the original hyperspectral image data, the hyperspectral data after PCA processing, and the registered Lidar / SAR data. The model runs in the inference mode without involving gradient backpropagation and finally outputs the probability distribution of the class to which each test sample belongs. For each pixel or image patch, the system uses the class index with the largest median in its classification probability distribution as the prediction label for that sample, thereby obtaining the classification results for the entire test set. Subsequently, the system rearranges and restores the prediction labels of all test samples according to their spatial coordinate indices and stitches them into a two-dimensional spatial distribution map with the same resolution as the original remote sensing image, that is, a complete remote sensing classification image is generated.

[0172] (4) Visualize and analyze the classification results.

[0173] To improve the readability of the classification results and the ability to express information, the skimage library is used to perform category-color mapping processing on the label map. In the specific operation, according to the category-color mapping rules preset for different datasets, a dedicated color with high contrast is assigned to each type of ground object, so that different categories show significant visual differences in the image, facilitating intuitive identification by the human eye and regional structure analysis. The generated color image is encoded in the standard RGB space and saved as a.png format file to ensure its good display adaptability and cross-platform compatibility. This visualized image can be directly used in various application scenarios such as classification result display, remote sensing image interpretation, and thematic map output, further enhancing the visualization effect and engineering practicality of the model output results.

[0174] Example 2:

[0175] This example further illustrates the effect of the present invention through simulation experiments:

[0176] The experimental platform for the simulation experiment in this embodiment is equipped with an NVIDIA RTX 4090 GPU (24GB video memory), an Intel Xeon Platinum 8352V processor (16 cores), and 120GB of memory. The software environment is based on the Ubuntu 20.04 operating system, and the training framework is built using PyTorch 1.11.0 and Python 3.8, with powerful parallel computing and deep learning model training capabilities, ensuring the stability and efficiency of the experimental results. This experiment was carried out on three publicly available multi-source remote sensing classification datasets, namely the Augsburg dataset, the Berlin dataset, and the Houston2013 dataset. The three datasets respectively integrate hyperspectral images with SAR or Lidar data, and have the characteristics of complex multi-modal structures, diverse class distributions, significant differences in spatial scales and spectral characteristics, etc., which can effectively evaluate the generalization ability of the method under different types of data fusion. The Augsburg dataset was collected from the Augsburg area in Germany and integrates a hyperspectral image composed of 180 spectral bands (0.4–2.5μm) collected by the HySpex sensor and SAR images (including VV, VH polarization intensities, and real and imaginary parts of the polarization covariance matrix features) obtained by the Sentinel-1 sensor. The image resolution is 30m, the spatial size is 332×485 pixels, and the ground truth labels include 7 types of land cover such as forests and industrial areas. The Berlin dataset covers urban and rural areas in the Berlin area of Germany. The HSI data is a simulated EnMAP hyperspectral image (constructed from HyMap data) and contains a total of 244 spectral bands; the SAR data is a Sentinel-1 dual-polarization SLC product provided by ESA, with a spatial resolution of 13.89m and a spatial size of 1723×476 pixels. To achieve registration, the two-modal data are aligned by nearest neighbor interpolation, and the labels cover 8 types of typical land cover, including commercial areas and water bodies. The Houston2013 dataset comes from the 2013 IEEE Geoscience and Remote Sensing Society Data Fusion Competition and was obtained by aerial survey of the Houston area and its surrounding areas in Texas, USA. It contains a total of 144-band HSI data (380–1050nm) and high-spatial-accuracy Lidar data, with a spatial resolution of 2.5m and an image size of 349×1905, including 15 types of land cover labels such as healthy grasslands, parking lots, and runways, which is suitable for fine-grained urban scene classification tasks.

[0177] To comprehensively evaluate the performance of the method of the present invention, it was systematically compared with a variety of advanced methods in the current field of multi-source classification of remote sensing images. These comparison methods are all derived from representative research work published in recent years, covering the mainstream fusion strategies based on convolutional neural networks and Transformers, and have strong representativeness and reference value. Specifically, the following methods are included: The TBCNN method was proposed in the article "Multisource remote sensing data classification based on convolutional neural network"; The FusAtNet method was proposed in the article "Fusatnet: Dual attention based spectrospatial multimodal fusion network for hyperspectral and lidar classification"; The method was proposed in the article "S2enet: Spatial–spectral cross-modal enhancement network for classification of hyperspectral and lidar data"; The DFINet method was proposed in the article "Hyperspectral and multispectral classification for coastal wetland using depthwise feature interaction network"; The ExViT method was proposed in the article "Extended Vision Transformer (ExViT) for Land Use and Land Cover Classification: A Multimodal Deep Learning Framework"; The HCT method was proposed in the article "Joint classification of hyperspectral and LiDAR data using a hierarchical CNN and Transformer"; The MACN method was proposed in the article "Mixing self-attention and convolution: A unified framework for multi-source remote sensing data classification";The MICF-Net method was proposed in the article "Multiple information collaborative fusion network for joint classification of hyperspectral and LiDAR data"; the M2FNet method was proposed in the article "Multiscale 3-D-2-D mixed CNN and lightweight attention-free transformer for hyperspectral and LiDAR classification".;

[0178] To more intuitively demonstrate the actual performance of the method of the present invention in the classification of multi-source remote sensing images, the experimental results of three typical datasets (Augsburg, Berlin, and Houston2013) were visually analyzed. As Figure 6 shown, on the Augsburg dataset, the method of the present invention has particularly significant classification effects in categories such as industrial areas, farmland, and water bodies. The visualization results show that the present invention can accurately identify the spatial structure of industrial areas and effectively avoid the common regional misjudgment and patchiness phenomena in other methods; in the classification of farmland areas, the model shows good regional integrity and boundary retention ability; for the water body category, the present method can identify clearer boundaries, and the classification results within the region are more consistent, significantly reducing the cases of misclassification and missed classification. These performances fully reflect the effectiveness and advantages of the method of the present invention in key band extraction and cross-modal collaborative modeling, especially having stronger capabilities in the identification of fine-grained ground objects and the maintenance of spatial structure. As Figure 7 shown, for the Berlin dataset, the classification results of the method of the present invention in the soil and water body categories are significantly better than other comparison methods. The visualization images show that the method of the present invention is more accurate in the classification of soil areas, can better maintain the consistency within the region, and effectively distinguish the boundaries of adjacent ground objects. In addition, for the classification of water body areas, the method of the present invention significantly improves the accuracy of the regional boundary and the internal consistency, demonstrating significant spectral-spatial collaborative modeling ability, and effectively reducing the phenomena of boundary blurring and misclassification in traditional methods. As Figure 8As shown in the experiments on the Houston2013 dataset, the method of the present invention is particularly outstanding in the recognition of categories such as healthy grassland, commercial area, soil, and parking lot 2. The visualization results show that the method can accurately capture the detailed texture features of the healthy grassland area, significantly reducing the common category confusion phenomenon in other methods; in the classification of commercial areas, the model exhibits higher spatial resolution ability, can clearly divide the regional boundaries, and accurately identify complex structures and subtle edge changes; in the categories of soil and parking lot 2, the method also significantly improves the classification accuracy, demonstrating good spatial consistency and category discrimination ability. These results fully verify the strong feature extraction and expression ability of the method of the present invention in fusing hyperspectral and Lidar / SAR data, and can effectively improve the overall performance and fine-grained recognition accuracy of the model in the multi-source remote sensing image classification task. Whether it is the accurate recognition of regional boundaries, the precise capture of detailed features, or the effectiveness of cross-modal information fusion, it is significantly better than the existing mainstream comparison methods, demonstrating its outstanding generalization ability and robust classification effect.

[0179] The present invention has been comprehensively and quantitatively evaluated on three typical multi-source remote sensing datasets in Augsburg, Berlin, and Houston 2013, and a comparative analysis has been carried out with various advanced methods in the current remote sensing classification field. From the statistical results of the three classic evaluation indicators, namely the overall classification accuracy (OA), average classification accuracy (AA), and Kappa coefficient, the method of the present invention shows significant advantages in all indicators. Specifically, as shown in Table 4, on the Augsburg dataset, the method of the present invention has achieved the best results in OA, AA, and Kappa coefficient, indicating that it not only has advantages in overall accuracy but also can maintain the balance and consistency of the classification performance of each category, fully demonstrating the significant performance advantages of the present invention in key band retrieval and cross-modal information fusion. As shown in Table 5, in the experimental results of the Berlin dataset, the present invention also obtained the highest OA and Kappa coefficient among all the compared methods. Especially in difficult-to-classify categories such as soil and water bodies, it significantly outperformed other advanced methods. This fully shows that the method of the present invention not only has excellent overall accuracy but also has excellent spectral-spatial collaborative modeling ability, can effectively extract clear features of the ground object boundaries, solves the problems of category confusion and boundary blur easily occurring in traditional methods, and significantly improves the robustness and reliability of the classification results. As shown in Table 6, on the Houston 2013 dataset, the method of the present invention once again achieved an overall lead in OA, AA, and Kappa coefficient. Compared with the current advanced multi-source fusion classification methods, the performance advantages of the present invention are more prominent. From the specific category performance, the method of the present invention has obvious accuracy advantages in the classification tasks of complex fine-grained categories such as healthy grassland, commercial area, soil, and parking lot 2, can effectively capture the subtle features of the internal structure of the ground object and the detailed changes of the boundaries, and greatly reduces the classification errors of easily confused categories.

[0180] Table 4 Comparison results of performance indicators of each algorithm on the Augsburg dataset

[0181]

[0182] Table 5 Comparison results of performance indicators of each algorithm on the Berlin dataset

[0183]

[0184] Table 6 Comparison results of performance indicators of each algorithm on the Houston 2013 dataset

[0185]

[0186] The multi-source remote sensing image classification method based on key band retrieval attention mechanism proposed by the present invention aims to solve problems such as information redundancy, feature mismatch, and insufficient classification accuracy in the process of hyperspectral image and Lidar / SAR data fusion. By constructing a dual-domain fusion encoder with frequency-space joint modeling ability and a key band retrieval attention module, it effectively extracts discriminative spectral information in remote sensing images and strengthens the deep interaction expression between different modalities, significantly improving the robustness and accuracy performance of the model in complex land cover category recognition tasks. In addition, the method proposed by the present invention is also applicable to other multi-source remote sensing classification scenarios with similar data characteristics, such as accurately classifying vegetation types in wetland areas to assist in ecological environment protection planning, finely classifying agricultural crop areas to achieve precision agriculture management, and accurately classifying urban areas to assist in urban fine governance and infrastructure planning, and other remote sensing application scenarios.

[0187] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A multi-source remote sensing image classification method based on a key band retrieval attention mechanism, characterized in that It includes the following steps: S1: Collect hyperspectral images HSI and Lidar / SAR multi-source remote sensing data, and preprocess the data; S2: Construct a key band retrieval attention neural network model, which includes a dual-domain fusion encoder module, a self-attention layer, a key band retrieval module, a key band retrieval attention module, and a fully connected layer for classification; The processing flow of the key band retrieval attention neural network model: S2-1: First, extract features from the HSI data. For the original HSI data and the HSI data after PCA processing perform feature extraction through the dual-domain fusion encoder module respectively. The dual-domain fusion encoder module consists of two branches: the spatial domain and the frequency domain. The processing flow of the spatial domain branch is: HSI data → depth convolution → BatchNorm2D → ReLU activation function → spatial domain features In the frequency domain branch, a dynamic frequency filtering module is designed, and its processing flow is: HSI data → fast Fourier transform → adaptive learnable mask → Hadamard product → inverse fast Fourier transform → frequency domain features Finally, add the two to obtain the dual-domain fusion features ; S2-2: Extract features from Lidar / SAR data The process is as follows: Lidar / SAR data → 2D convolution → BatchNorm2D → ReLU activation function → Lidar / SAR data features The calculation process of this feature extraction operation is expressed as: ; S2-3: HSI data features and the HSI data features after PCA processing perform a dimension reshaping operation, flatten the spatial dimensions H and W to obtain [B, C1, P] and [B, C pca , P], and at the same time, the Lidar / SAR data features also perform the same operation to reshape the dimensions to [B, C2, P], where P = H × W, indicating the number of pixels contained in each sample; after completing the flattening operation, and are concatenated along the channel dimension to obtain the intermediate state features , with the dimension of [B, C pca + C2, P]; the calculation process is as follows: ; S2-4: Enter N SA The N-layer self-attention layer models long-range dependencies. The specific process is as follows: The input data → Depth convolution → Self-attention mechanism → Residual connection → Feed-forward neural network → Residual connection → Processed features ; S2-5: Original HSI Data Features Enter the key band retrieval module to perform the retrieval of the first key band. This process is based on the attention similarity matrix generated in S2-5 to infer the important bands in the original HSI; First, extract the sub-matrix corresponding to the columns from the C pca+1 th to the C pca +C2 columns and the first C pca rows from the similarity matrix, denoted as , with a dimension of [B, C pca , C2]. The elements of the extracted similarity matrix represent the similarities between the HSI data features after PCA processing and the Lidar / SAR data features for each band; Next, perform hidden correlation Softmax on the similarity matrix to complete the band retrieval accordingly; S2-6: Key Band Features and intermediate state features are jointly input into the N KA layer key band retrieval attention module for processing. The specific process is as follows: Input data and → and are concatenated along the channel dimension → depth convolution → cross-attention mechanism → residual connection → feed-forward neural network → residual connection → updated intermediate state features ; S2-7: Finally, process the key band features into the classification result. The specific process is as follows: Input the key band features → Max pooling layer → Dropout → Fully connected layer → Classification result ; The entire calculation process is shown as follows: ; Among them, ; S3: Input the preprocessed data into the key band retrieval attention neural network model for model training; During the training process, cross-entropy is used as the loss function, and the Adam optimizer is used for parameter optimization; S4: Use the trained key band retrieval attention neural network model to process the data to be measured, and visually present and analyze the classification results of the multi-source remote sensing images.

2. The multi-source remote sensing image classification method according to claim 1, wherein The S1 includes: S1-1: Input the original hyperspectral image data and Lidar / SAR data, and use the principal component analysis method to reduce the spectral dimension of the hyperspectral data to obtain the PCA weight matrix ; S1-2: Perform edge padding operations on the original hyperspectral data, the hyperspectral data processed by PCA, and the Lidar / SAR data; S1-3: Normalize the processed hyperspectral data and Lidar / SAR data; S1-4: Slide and sample on the processed multi-source data with a window of a specified size, and divide the entire image into multiple data blocks of a fixed size; S1-5: Perform data augmentation operations on the constructed hyperspectral and Lidar / SAR sample blocks, using strategies such as random horizontal flipping and vertical flipping; S1-6: Label and divide the preprocessed samples according to the ground truth labels to construct a training set, a validation set, and a test set.

3. The multi-source remote sensing image classification method according to claim 1, characterized in that, In S2, the inputs of the key band retrieval attention neural network model are the original HSI data, the hyperspectral data after PCA processing, and the Lidar / SAR data. The dimension of the HSI data is [B, C, N1, H, W], where B is the batch size, representing the number of samples passed into the model in parallel, C is the number of channels, N1 is the number of hyperspectral bands, and H and W respectively represent the height and width of the image patch, which are equal to the size of the sliding window in the preprocessing stage; the dimension of the hyperspectral data after PCA processing is [B, C, N pca , H, W], and the meaning of each dimension is the same as that of the original HSI data; the dimension of the Lidar / SAR data is [B, C2, H, W], where B is the batch size, C2 is the number of channels, and H and W are the length and width of the sample; the original HSI data and the HSI data after PCA processing are respectively reshaped in dimension, and their dimensions after reshaping are [B, C1, H, W] and [B, C pca , H, W], which are aligned with the dimension of the Lidar / SAR data, where C1 = C × N1, C pca = C × N pca .

4. The multi-source remote sensing image classification method according to claim 1, wherein, The calculation process of the S2-1 is as follows: ; ; ; Among them, can represent the original HSI data or the HSI data after PCA processing , represents the learning process of the adaptive learnable mask, denotes the Hadamard product; the learning process of the adaptive learnable mask is expressed as: data after fast Fourier transform → linear layer → ReLU activation function → linear layer → mask; the original HSI data and the HSI data after PCA processing respectively obtain dual-domain fusion features through the dual-domain fusion encoder module and .

5. The multi-source remote sensing image classification method according to claim 1, wherein The calculation process of the S2-4 is as follows: ; ; ; ; Among them, Q, K, and V are the query, key, and value in the self-attention mechanism respectively. is the similarity matrix of the self-attention mechanism, and its dimension is [B, C pca +C 2, C pca +C2]; FFN represents the feed-forward neural network.

6. The multi-source remote sensing image classification method according to claim 1, wherein, The specific steps of S2-5 are as follows: Input similarity matrix and HSI data characteristics →Similarity matrix Perform Softmax normalization to obtain →Similarity matrix Perform Top-K search on all elements in to get the band index of the top k% of similarity →Based on band index Search S1-1 to obtain get →Yes Perform Softmax normalization to obtain →Yes Top-K search can obtain the band index of the top k% of the original HSI data similarity → By index The key band features of the original HSI data can be retrieved , key band characteristics The dimension is [B, k%C1, P]; the whole calculation process is as follows: ; ; ; ; ; ; ; Among them, represents rounding down.

7. The multi-source remote sensing image classification method according to claim 1, characterized in that The calculation process of the S2-6 is as follows: ; ; ; ; Among them, Q, K, and V are the query, key, and value in the multi-source channel cross-sparse attention mechanism, respectively. is the similarity matrix, and its dimension is [B, C pca +C 2, C pca +C2+C1]; FFN represents the feed-forward neural network, and its specific process is: input data → 2D convolution → ReLU activation function → 2D convolution → Dropout → output data; the above similarity matrix and the original HSI features are sent to the aforementioned key band retrieval module together to further perform spectral band screening based on the updated attention relationship. After passing through the N KA layers of key band retrieval attention modules, the final key band features .

8. The multi-source remote sensing image classification method according to claim 2, wherein In the S3, model training and evaluation include: S3-1: Input the training set obtained in step S1 into the constructed key band retrieval attention neural network model. The inputs received by the model include the original hyperspectral image data, the hyperspectral data processed by PCA, and the Lidar / SAR data, and perform feature extraction and fusion processing through each module of the neural network; S3-2: During the training process, the overall training strategy combines loss function optimization, adaptive scheduling, and regularization means. The mathematical expression of the cross-entropy loss function is as follows: ; where N is the number of samples, and C is the number of classes, represents the distribution of the true label of the i-th sample in the j-th class, represents the predicted probability of the i-th sample in the j-th class; S3-3: After training is completed, input the validation set into the trained key band retrieval attention neural network model for classification prediction; The model outputs the ground object category label to which each test sample belongs; S3-4: Evaluate the classification results on the validation set.

9. The multi-source remote sensing image classification method according to claim 2, characterized in that The S4 includes: Input the test data set into the trained key band retrieval attention neural network model, perform a forward propagation operation, and obtain the predicted classification results of each pixel or sample; Restore and splice the classification labels generated by the model for the test set to reconstruct a complete classification result image; Use the category-color mapping rule to visually display the classification map predicted by the model.

Citation Information

Patent Citations

  • Multi-source remote sensing image classification method based on spectrum adaptive feature fusion

    CN119723216A

  • A remote sensing image classification method based on data enhancement and knowledge distillation

    CN119741539A