Multi-modal fusion feature generation method and device and storage medium

Through wavelet transformation and adaptive frequency enhancement and sparse compensation models, the problems of insufficient fusion and feature imbalance in multimodal data fusion are solved, and standardized multimodal fusion features are generated, which improves the accuracy and robustness of land objects recognition and classification.

CN120495830AActive Publication Date: 2025-08-15NANJING UNIV OF INFORMATION SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510993463.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-15
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Traditional multimodal data fusion technology has problems such as insufficient fusion, loss of key frequency band information, noise interference and imbalance of high and low frequency characteristics, which affect the effect of land object identification and classification.

Method used

Wavelet transform is used to decompose hyperspectral imaging and lidar data, and high and low frequency characteristics are enhanced through adaptive frequency enhancement and sparse compensation models, and standardized multimodal fusion characteristics are generated by combining gating mechanisms and channel-space dual attention mechanisms.

Benefits of technology

It realizes full fusion of multimodal data, effectively suppresses noise interference, restores key frequency band information, and achieves 96.36% classification accuracy, improving the accuracy and robustness of land objects identification and classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495830A_ABST
    Figure CN120495830A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal fusion feature generation method and device and a storage medium in the technical field of data processing, and the method comprises the steps: carrying out the frequency feature decomposition of an image block through wavelet transform, and generating a corresponding high-frequency sub-band and a corresponding low-frequency sub-band, inputting the high-frequency sub-band and the low-frequency sub-band into a pre-constructed adaptive frequency enhancement and sparse compensation model; obtaining an enhanced high-frequency sub-band and an enhanced low-frequency sub-band; performing inverse wavelet transform based on the enhanced high-frequency sub-band and low-frequency sub-band to generate a reconstructed signal; and generating a standardized multi-modal fusion feature based on the image block and the reconstruction signal. According to the invention, the technical problems of insufficient fusion, key frequency band information loss, noise interference and imbalance of high and low frequency characteristics in traditional multi-modal data fusion can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a multimodal fusion feature generation method, device and storage medium. Background Art

[0002] Hyperspectral images (HSI) have rich spectral information and can accurately depict the spectral characteristics of objects, providing detailed information for object identification. Light-Detection-and-Ranging (LiDAR) has precise terrain and three-dimensional structure information, which can accurately depict the geometric shape and spatial distribution of objects, providing a reliable basis for the spatial positioning and height measurement of objects. The fusion of the two can improve the recognition and classification of objects, make comprehensive use of the spectral and spatial characteristics of objects, thereby improving the accuracy and robustness of classification, which is of great significance to the refined classification and understanding of complex objects.

[0003] However, traditional multimodal data fusion technology has some problems: 1. It only focuses on spatial or spectral features, and the fusion is insufficient, resulting in the ineffective utilization of some important feature information; 2. The rich inter-spectral features of HSI and the frequency band characteristics of LiDAR geometric elevation information are significantly different, and direct fusion can easily lead to the loss of key frequency band information; 3. Noise and cross-modal interference show a complex coupling relationship in the frequency domain, and traditional frequency band selection methods lack adaptive suppression capabilities; 4. There is a lack of targeted processing of low-frequency and high-frequency features, and the dynamic balance mechanism between high-frequency detail information and low-frequency contour information has not been effectively established, which restricts the improvement of classification accuracy.

[0004] Therefore, there is an urgent need for a multimodal fusion feature generation method, device and storage medium to solve the above technical problems. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a multimodal fusion feature generation method, device and storage medium that can solve the technical problems of insufficient fusion, loss of key frequency band information, noise interference and imbalance of high and low frequency features existing in traditional multimodal data fusion.

[0006] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0007] In a first aspect, the present invention provides a method for generating multimodal fusion features, comprising:

[0008] Acquire hyperspectral imaging and lidar data, and construct an image block centered on the pixel for each pixel sample in the hyperspectral imaging and lidar data;

[0009] Decomposing the image block by frequency features using wavelet transform to generate corresponding high-frequency sub-bands and low-frequency sub-bands;

[0010] Inputting the high frequency sub-band and the low frequency sub-band into a pre-built adaptive frequency enhancement and sparse compensation model, wherein the adaptive frequency enhancement and sparse compensation model includes an adaptive frequency enhancement unit and a sparse compensation unit;

[0011] The adaptive frequency enhancement unit enhances local detail information of the high-frequency sub-band and the low-frequency sub-band, and the sparse compensation unit restores global structural information of the high-frequency sub-band and the low-frequency sub-band that is weakened by frequency feature decomposition, thereby obtaining enhanced high-frequency sub-band and low-frequency sub-band;

[0012] Performing inverse wavelet transform based on the enhanced high-frequency sub-band and low-frequency sub-band to generate a reconstructed signal;

[0013] Based on the image blocks and the reconstructed signals, the gated fusion features of the hyperspectral imaging and lidar data are obtained through a gating mechanism, and the gated fusion features of the hyperspectral imaging and lidar data are spliced in the channel dimension. The channel-space dual attention mechanism is adopted to generate standardized multimodal fusion features.

[0014] Furthermore, the adaptive frequency enhancement unit utilizes task-driven dynamic weight learning and Laplace convolution to enhance local detail information of the high-frequency sub-band and the low-frequency sub-band, and obtains the high-frequency sub-band and the low-frequency sub-band after the local detail information is enhanced;

[0015] The sparse compensation unit captures the spatial correlation of the high-frequency subband and the low-frequency subband after the local detail information is enhanced, suppresses noise interference, and restores the global structural information weakened by the frequency feature decomposition to obtain the enhanced high-frequency subband and the low-frequency subband.

[0016] Furthermore, the adaptive frequency enhancement unit includes: an information quantity evaluation branch, a noise level evaluation branch, and a task relevance evaluation branch;

[0017] generating an information content evaluation component of an input subband by utilizing the information content evaluation branch;

[0018] generating a noise level estimation component of the input subband using a noise level estimation branch;

[0019] generating a task relevance evaluation component of the input subband using the task relevance evaluation branch;

[0020] generating a dynamic weight of the input subband according to the information quantity evaluation component, the noise level evaluation component, and the task relevance evaluation component, and performing adaptive frequency enhancement on the input subband by combining the dynamic weight and using a Laplace convolution form;

[0021] The input sub-band includes a high-frequency sub-band and a low-frequency sub-band.

[0022] In a second aspect, the present invention provides a multimodal fusion feature generation device, comprising:

[0023] A data acquisition module is used to acquire hyperspectral imaging and lidar data, and construct an image block centered on the pixel for each pixel sample in the hyperspectral imaging and lidar data;

[0024] A feature decomposition module, configured to perform frequency feature decomposition on the image block using wavelet transform to generate corresponding high-frequency sub-bands and low-frequency sub-bands;

[0025] An input module, configured to input the high frequency sub-band and the low frequency sub-band into a pre-built adaptive frequency enhancement and sparse compensation model, wherein the adaptive frequency enhancement and sparse compensation model includes an adaptive frequency enhancement unit and a sparse compensation unit;

[0026] an enhancement module, configured to enhance local detail information of the high-frequency sub-band and the low-frequency sub-band by the adaptive frequency enhancement unit, and restore global structural information of the high-frequency sub-band and the low-frequency sub-band that is weakened by frequency feature decomposition by the sparse compensation unit, thereby obtaining enhanced high-frequency sub-band and low-frequency sub-band;

[0027] A reconstruction module, configured to perform inverse wavelet transform based on the enhanced high-frequency sub-band and low-frequency sub-band to generate a reconstructed signal;

[0028] A fusion module is used to obtain the gated fusion features of hyperspectral imaging and lidar data based on the image blocks and the reconstructed signals through a gating mechanism, splice the gated fusion features of the hyperspectral imaging and lidar data in the channel dimension, and adopt a channel-space dual attention mechanism to generate standardized multimodal fusion features.

[0029] In a third aspect, the present invention provides an electronic terminal comprising a processor and a memory connected to the processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of any of the above methods are performed.

[0030] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the above methods when executed by a processor.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] The multimodal fusion feature generation method proposed in the present invention decomposes the HSI spectral features and LiDAR geometric elevation features into low-frequency sub-bands and high-frequency sub-bands based on the multi-scale decomposition of wavelet transform, and adaptively enhances the frequency of the high-frequency sub-bands through a dynamic weight generation mechanism combined with high-pass filter convolution, fully obtains high-frequency detail information, and effectively solves the problems of noise interference and information loss in multimodal data fusion; through the local spatial attention modeling of the sparse compensation unit and the gated low-frequency compensation network, while enhancing the high-frequency detail features, the over-suppressed low-frequency contour information is restored, and the dynamic balanced expression of high- and low-frequency features is achieved; combined with the residual gated fusion unit and the channel-space dual attention mechanism, the reconstructed frequency domain features are residually connected with the original features, and standardized multimodal fusion features are generated through global average pooling and layer normalization. Finally, the multimodal fusion features are input into the classifier and the classification results are output, which systematically solves the problems of significant differences in multimodal data fusion, noise interference and imbalance of high- and low-frequency features;

[0033] The present invention realizes the feature synergy between HSI and LiDAR data in the spectral, spatial and frequency domains through the cascade optimization of frequency domain feature decoupling and adaptive frequency enhancement and sparse compensation modules. Experiments show that the overall classification accuracy of this method on public datasets reaches 96.36%. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A first flow chart of a multimodal fusion feature generation method provided by an embodiment of the present invention;

[0035] Figure 2 A second flow chart of a multimodal fusion feature generation method provided by an embodiment of the present invention;

[0036] Figure 3 A model framework diagram of a multimodal fusion feature generation method provided by an embodiment of the present invention;

[0037] Figure 4 A pseudo-color image of a hyperspectral image of a certain region in an embodiment of the present invention;

[0038] Figure 5 A grayscale image of radar data of a certain area in an embodiment of the present invention;

[0039] Figure 6 A map of real landform types in a certain area in an embodiment of the present invention;

[0040] Figure 7 This is a diagram of the Coupled_CNN classification results in an embodiment of the present invention;

[0041] Figure 8 This is a diagram of the LSAF classification results in an embodiment of the present invention;

[0042] Figure 9 This is the ExViT classification result diagram in an embodiment of the present invention;

[0043] Figure 10 This is the FDNet classification result diagram in an embodiment of the present invention;

[0044] Figure 11 This is a Slice-Mamba classification result diagram in an embodiment of the present invention;

[0045] Figure 12 This is the AMSSE-Net classification result diagram in an embodiment of the present invention;

[0046] Figure 13 This is a CALC classification result diagram in an embodiment of the present invention;

[0047] Figure 14 This is a diagram of HCT classification results in an embodiment of the present invention;

[0048] Figure 15 This is the NCGLF2 classification result diagram in an embodiment of the present invention;

[0049] Figure 16 This is a classification result diagram of the multimodal fusion feature generation method in an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Unless there is a conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0051] The term "and / or" in this disclosure simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.

[0052] Example 1: Figure 1 This is a flow chart of the multimodal fusion feature generation method in an embodiment of the present invention. This flow chart only shows the logical sequence of the method described in this embodiment. In other possible embodiments of the present invention, different methods may be used without conflict. Figure 1 The steps shown or described are accomplished in the order shown.

[0053] The multimodal fusion feature generation method provided in this embodiment can be applied to a terminal and can be executed by a mechanical equipment fault identification device. The device can be implemented by software and / or hardware and can be integrated into a terminal, such as any smart phone, tablet computer or computer device with communication function. Figures 1 to 16 As shown, the method of this embodiment specifically includes the following steps:

[0054] Step 1: Obtain hyperspectral imaging and lidar data, and construct a pixel-centered image patch for each pixel sample in the hyperspectral imaging and lidar data:

[0055] The spatial shape and size of the patch is , the patch expression includes:

[0056] ,

[0057] ,

[0058] Where, For hyperspectral imaging data ( , ), For LiDAR data in ( , ), Indicates floor operation, is an odd number, Make sure the patch center is strictly aligned with the pixel ( , ), is the horizontal coordinate of the pixel point, is the vertical coordinate of the pixel.

[0059] Step 2: Use wavelet transform to decompose the frequency characteristics of the image block to generate corresponding high-frequency sub-bands and low-frequency sub-bands:

[0060] The image blocks of hyperspectral imaging data and lidar data are The convolution operation of the convolution kernel extracts the spatial and spectral features of hyperspectral imaging data and lidar data, as shown in the formula:

[0061] ,

[0062] ,

[0063] in, are the spatial and spectral characteristics of hyperspectral imaging data, The spatial and spectral characteristics of the lidar data;

[0064] The spatial and spectral features of the hyperspectral imaging and lidar data are decomposed into low-frequency sub-bands and multiple high-frequency sub-bands using wavelet transform, as shown in the formula:

[0065] ,

[0066] ,

[0067] Where, represents the low-frequency subband of HSI, Represents multiple different high-frequency sub-bands of HSI, represents the number of high frequency sub-bands. In this embodiment, It can be 1, 2, or 3. represents the scaling function, Indicates the The translation basis function of each high-frequency subband is represented by m and n, and p represents the size of the image block. Each high-frequency subband and low-frequency subband of HSI and LiDAR is input into the pre-built adaptive frequency enhancement and sparse compensation model separately, and the input is uniformly represented by x.

[0068] and inputting the high frequency sub-band and the low frequency sub-band into a pre-built adaptive frequency enhancement and sparse compensation model;

[0069] Step 3: Enhance the local detail information of the high-frequency subband and the low-frequency subband by the adaptive frequency enhancement unit of the adaptive frequency enhancement and sparse compensation model, and restore the global structural information of the high-frequency subband and the low-frequency subband weakened by the frequency feature decomposition by the sparse compensation unit of the adaptive frequency enhancement and sparse compensation model to obtain the enhanced high-frequency subband and the low-frequency subband:

[0070] The adaptive frequency enhancement unit utilizes task-driven dynamic weight learning and Laplace convolution to enhance local detail information of high-frequency sub-bands and low-frequency sub-bands, and obtains high-frequency sub-bands and low-frequency sub-bands after local detail information enhancement:

[0071] The adaptive frequency enhancement unit includes: an information quantity evaluation branch, a noise level evaluation branch and a task relevance evaluation branch;

[0072] generating an information content evaluation component of an input subband by utilizing the information content evaluation branch;

[0073] generating a noise level estimation component of the input subband using a noise level estimation branch;

[0074] generating a task relevance evaluation component of the input subband using the task relevance evaluation branch;

[0075] generating a dynamic weight of the input subband according to the information quantity evaluation component, the noise level evaluation component, and the task relevance evaluation component, and performing adaptive frequency enhancement on the input subband by combining the dynamic weight and using a Laplace convolution form;

[0076] The input sub-band includes a high-frequency sub-band and a low-frequency sub-band.

[0077] The expression of the information evaluation branch includes:

[0078] ,

[0079] Among them, S represents the information evaluation component of the input subband, x represents the input subband, Represents a 3×3 convolution operation, GELU represents Gaussian error linear unit, and AdaptiveAvgPool represents adaptive average pooling operation;

[0080] The expression of the noise level evaluation branch includes:

[0081] ,

[0082] Where N represents the noise level evaluation component of the input subband, Represents a 1×1 convolution operation;

[0083] The expression of the task dependency evaluation branch includes:

[0084] ,

[0085] Among them, R represents the task relevance evaluation component of the input subband, represents a 1×1 convolution operation that reduces the number of channels to half of the original number. represents a 1×1 convolution operation that restores the number of channels to the original number. Represents the Sigmoid activation function;

[0086] Generating a dynamic weight of the input subband according to the information quantity evaluation component, the noise level evaluation component, and the task relevance evaluation component, and then combining the dynamic weight and performing adaptive frequency enhancement on the input subband using a Laplace convolution form, including:

[0087] The information quantity evaluation component S, the noise level evaluation component N, and the task relevance evaluation component R are concatenated in the channel dimension. The channel dimension of the concatenated feature is then mapped to the number of channels of the input subband through a 1×1 convolution layer. The sigmoid function is then used to perform channel-by-channel normalization to generate dynamic weights in the range of [0, 1]. The expressions include:

[0088] ,

[0089] Where W is the dynamic weight, Concat represents the concatenation operation in the channel dimension;

[0090] The dynamic weight is combined with the Laplace convolution form to perform a convolution operation on the input subband, and the expression includes:

[0091] ,

[0092] in, represents the convolution form of Laplace, Indicates that the image block is , the vertical axis is The pixel at , r and s represent the index of the convolution kernel, indicating the offset of the convolution kernel in the horizontal and vertical directions, k represents the size of the convolution kernel in the row direction, l represents the size of the convolution kernel in the column direction, k and l determine the size of the convolution kernel;

[0093] The features after adaptive frequency enhancement are residually connected with the input subband. The expressions include:

[0094] ,

[0095] in, Represents the high-frequency sub-band and low-frequency sub-band after local detail information is enhanced.

[0096] The sparse compensation unit captures the spatial correlation of the high-frequency sub-band and the low-frequency sub-band after the local detail information is enhanced, suppresses noise interference, and restores the global structural information weakened by the frequency feature decomposition, thereby obtaining the enhanced high-frequency sub-band and the low-frequency sub-band:

[0097] The sparse attention mechanism of the sparse compensation unit is used to capture spatial correlation, the low-frequency feature compensation unit of the sparse compensation unit is used to restore the global structural information weakened by frequency decomposition, and the feedforward neural network unit of the sparse compensation unit is used to sparsely activate the features and suppress noise interference;

[0098] Specifically, the high-frequency sub-band and low-frequency sub-band after the input local detail information is enhanced are normalized through the normalization operation. The expressions include:

[0099] ,

[0100] in, Represents the high-frequency sub-band and low-frequency sub-band after local detail information enhancement, represents the first normalized feature, Representation layer normalization operation, represents global average pooling, represents the fully connected layer, represents the Sigmoid activation function,

[0101] The sparse attention mechanism constructs a local attention mask to force each spatial position to establish attention association only with the local area within its 3×3 neighborhood, thus constraining the range of feature interaction, as shown in the formula:

[0102] ,

[0103] in, Indicates that the horizontal axis in the first normalized feature map is , the vertical axis is The local attention mask at ;

[0104] A feature representation with rich contextual information is obtained through a multi-head attention mechanism, where the query vector, key vector, and value vector are derived from the mapping of the first normalized feature. The output features of the sparse attention mechanism are obtained through dynamic feature aggregation. The expression includes:

[0105] ,

[0106] ,

[0107] in, represents the multi-head attention mechanism, A represents the output feature of the sparse attention mechanism, , B, C, H and Represent the batch size, number of channels, height of feature map and width of feature map respectively, represents the activation function, Q, K, and V represent the query vector, key vector, and value vector, respectively. represents the transpose operation, is the dimension of the key vector, represents the Hadamard product operation, h=1,2,...H, where h represents the number of heads and H represents the maximum number of heads;

[0108] The output features of the sparse attention mechanism are transformed to restore the shape of the first normalized features, and the output features of the restored shape are obtained. The output features of the restored shape are subjected to a pre-built sparse gating mechanism and are residually connected with the high-frequency sub-band and the low-frequency sub-band after local detail information enhancement to obtain the features after residual connection. The expression includes:

[0109] ,

[0110] in, is the sparse threshold, Represents the features after residual connection, , represented as the output feature of the restored shape;

[0111] The low-pass filter convolution operation is performed on the residual connection feature to extract its low-frequency feature. The expression includes:

[0112] ,

[0113] in, represents the low-pass filter convolution operation, It is a low-frequency feature;

[0114] The low-frequency features are input into the low-frequency feature compensation unit, and the compensation features are constructed through a two-layer convolution structure with nonlinear activation. The expressions include:

[0115] ,

[0116] in, represents the compensation characteristic, Represents a 3×3 convolution kernel operation with a bias term, which is used to keep the feature map size unchanged. ReLU stands for rectified linear unit. Represents a 3×3 convolution kernel operation without a bias term, which is used to generate compensatory features that match the number of input channels while preserving the spatial resolution of the feature map;

[0117] Based on the compensation feature, a residual connection is performed with the high-frequency sub-band and low-frequency sub-band after the local detail information is enhanced after the sparse gating mechanism to achieve directional enhancement of the global structural information. The expression includes:

[0118] ,

[0119] in, Represents the features after global structural information enhancement;

[0120] The features after the global structural information enhancement are normalized by a normalization operation, and the expression includes:

[0121] ,

[0122] in, represents the second normalized feature;

[0123] Based on the second normalized feature, the feedforward neural network unit generates a channel description vector through adaptive average pooling, and converts it into an attention weight through a learnable weight matrix and a Sigmoid function. The attention weight is used to highlight strong channels and suppress weak channels, and the output features of the feedforward neural network unit are obtained. The expression includes:

[0124] ,

[0125] in, represents the output features of the feedforward neural network unit, Represents a learnable residual scaling factor, which is used to dynamically adjust the contribution ratio of the compensation feature. LeakyReLU represents a leaky linear rectifier unit. Represents a 3×3 convolution operation with an expanded dimension, which is used to expand the number of channels to twice the original dimension. Represents a 3×3 convolution operation to restore the number of channels to the original number. represents the weight matrix of the gating network;

[0126] The output features of the feedforward neural network unit are subjected to a sparse gating mechanism and then residually connected with the high-frequency sub-band and low-frequency sub-band after local detail information enhancement, as shown in the formula:

[0127] ,

[0128] in, Represents the high-frequency subband and low-frequency subband enhanced by the adaptive frequency enhancement and sparse compensation model, that is, the enhanced high-frequency subband and low-frequency subband.

[0129] Step 4: Perform inverse wavelet transform based on the enhanced high-frequency sub-band and low-frequency sub-band to generate a reconstructed signal:

[0130] ,

[0131] Where, Represents the reconstructed signal generated by the high-frequency sub-band and the low-frequency sub-band after hyperspectral imaging enhancement, LH represents the horizontal high-frequency sub-band, HL represents the vertical high-frequency sub-band, and HH represents the diagonal high-frequency sub-band. The reconstructed signal generated by the high-frequency sub-band and the low-frequency sub-band after lidar data enhancement is obtained using the same formula.

[0132] Step 5: Perform inverse wavelet transform based on the enhanced high-frequency subband and low-frequency subband to generate a reconstructed signal; based on the image block and the reconstructed signal, obtain the gated fusion features of hyperspectral imaging and lidar data through a gating mechanism, splice the gated fusion features of hyperspectral imaging and lidar data in the channel dimension, and use the channel-space dual attention mechanism to generate standardized multimodal fusion features. The expression includes:

[0133] ,

[0134] ,

[0135] ,

[0136] ,

[0137] in, and Respectively represent the features of hyperspectral imaging and lidar data after gated fusion, and represent image patches of hyperspectral imaging and lidar data, respectively, and represent the reconstructed signals of hyperspectral imaging and lidar data respectively, Represents a cascade of double-layer convolution and nonlinear activation functions, Tanh represents the hyperbolic tangent activation function, represents a 3×3 convolution operation, represents a 1×1 convolution operation aligned with the channel dimension, represents the Hadamard product operation;

[0138] The features after gated fusion of the hyperspectral imaging and lidar data are spliced using a channel dimension strategy to obtain their multimodal fusion features. The expression includes:

[0139] ,

[0140] in, Represents multimodal fusion features, and Concat represents the concatenation operation in the channel dimension;

[0141] The multimodal fusion features are passed through the channel-space dual attention mechanism to generate standardized multimodal fusion features, the expressions of which include:

[0142] ,

[0143] ,

[0144] in, represents the standardized multimodal fusion features, represents the standardized multimodal fusion features, represents the channel attention map, Represents the spatial attention map;

[0145] The Softmax function is used to perform probabilistic processing on the standardized multimodal fusion features, and the expressions include:

[0146] ,

[0147] ,

[0148] in, represents the probability distribution, Represents the predicted label of each category, C represents the total number of categories of the classification task, FC represents the fully connected layer, Indicates that within a given range Find the maximum value in is the index of the category;

[0149] In this embodiment, the processed features can be mapped to the output space through the classifier, and the ground object category recognition results can be obtained based on the probability distribution and the predicted label of each category, which is applied to the field of remote sensing classification. In addition, after obtaining the multimodal fusion features, it can also be applied to a variety of downstream tasks, including land cover classification, target detection, agricultural crop monitoring and urban planning, etc., which are all existing technologies and will not be repeated here.

[0150] Specific application example of this embodiment: The dataset used in this embodiment includes scenes captured by a compact airborne spectral imager (CASI) in a certain area and its surrounding urban areas. It includes digital surface model (DSM) data based on HSI and lidar, with a spatial resolution of 2.5 meters and a size of 349×1905 pixels. The HSI data consists of 144 spectral bands with a wavelength range of 0.38-1.35 μm. In addition, the dataset is annotated with 15,029 labeled samples representing 15 different land cover categories. Figure 4 A false-color composite image of the HSI data is shown. Figure 5 shows a grayscale image of the LiDAR data, Figure 6 The ground truth map is shown.

[0151] The methods used in the comparative experiments include Coupled-CNN (Coupled-Convolutional-Neural-Network), FDNet (Frequency-Domain-Network), ExViT (Extended-Vision-Transformer), LSAF (Linear-Self-attention-Fusion), Slice-Mamba (Block-based Mamba model), AMSSE-Net (Adaptive-Multiscale-Spatial-Spectral-Enhancement-Network), CALC (Coupled-Adversarial-Learning-based-Classification), HCT (Hierarchical-CNN-Transformer), and NCGLF² (Network-Combining-Global-and-Local-Features-for-Fusion). Among them, Coupled_CNN adopts a dual-branch CNN structure to simultaneously extract HSI spectral features and LiDAR spatial features, FDNet enhances the local feature expression ability through deep convolution stacking, ExViT uses global self-attention to model cross-modal long-range dependencies, LSAF improves computational efficiency through local self-attention windows, Slice-Mamba captures edge features based on the raster scanning mechanism of the state-space model, AMSSE-Net combines the multi-scale Mamba module to enhance texture representation, CALC realizes local context fusion through cross-modal cross-attention, HCT adopts a hierarchical Transformer architecture to encode multi-granularity semantic information, and NCGLF² introduces non-local operations and cross-granularity feature interactions to improve scene adaptability. By comparing representative methods such as CNN (Coupled-CNN, FDNet), Transformer (ExViT, LSAF, CALC, HCT), and Mamba (Slice-Mamba, AMSSE-Net), the comprehensive advantages of the multimodal fusion feature generation method (Adaptive-Frequency-Enhancement-and-Sparse-Compensation, AFESC) proposed in the embodiment in multimodal remote sensing data processing are verified.

[0152] The hyperparameter settings are as follows: Haar wavelet transform, image patch size of 7, PCA for dimensionality reduction to 20, 80 training epochs, an initial learning rate of 0.001, the Adam optimizer with a weight decay of 1e4, a cosine annealing learning rate scheduler with a maximum epoch of 80 and a minimum learning rate of 1e6; a validation set of 15% for model tuning, an encoder dimension of 64, and 20 samples per class in the training set to ensure class balance. These hyperparameters work together to optimize the model training process and improve classification performance.

[0153] Under this condition, all methods were tested 10 times, and their classification accuracy is shown in Table 1. The data in brackets represent the fluctuation range of the experimental results in the form of accuracy (standard deviation).

[0154] Table 1: Comparative experiment on classification accuracy of a certain region dataset

[0155]

[0156] As shown in Table 1, the multimodal fusion feature generation method (AFESC) proposed in this example performs best among all advanced algorithms. AFESC achieves the highest accuracy in most categories, such as health-grass (98.95%), highway (99.74%), and running-track (100.00%), significantly outperforming other methods. AFESC achieves 96.36% overall accuracy (OA), 96.82% average accuracy (AA), and 0.96 kappa coefficient (K) in overall accuracy, respectively, outperforming other compared methods. This demonstrates that AFESC effectively improves the feature representation capabilities of multimodal remote sensing data through adaptive frequency enhancement and sparsity compensation.

[0157] To visualize the classification results, Figures 7 to 15 The classification results of Coupled_CNN, FDNet, ExViT, LSAF, Slice-Mamba, AMSSE-Net, CALC, HCT, and NCGLF² methods are shown respectively. Figure 16 The following figure shows the classification results of the multimodal fusion feature generation method proposed in the examples of this application. It can be seen that the proposed method enhances the model's ability to distinguish different ground object categories, accurately identifying the ground object category to which the sample belongs, and performs well in the classification task, demonstrating its comprehensive advantages in multimodal remote sensing data processing.

[0158] Embodiment 2: This embodiment provides a multimodal fusion feature generation device, including:

[0159] A data acquisition module is used to acquire hyperspectral imaging and lidar data, and construct an image block centered on the pixel for each pixel sample in the hyperspectral imaging and lidar data;

[0160] A feature decomposition module, configured to perform frequency feature decomposition on the image block using wavelet transform to generate corresponding high-frequency sub-bands and low-frequency sub-bands;

[0161] An input module, configured to input the high frequency sub-band and the low frequency sub-band into a pre-built adaptive frequency enhancement and sparse compensation model, wherein the adaptive frequency enhancement and sparse compensation model includes an adaptive frequency enhancement unit and a sparse compensation unit;

[0162] an enhancement module, configured to enhance local detail information of the high-frequency sub-band and the low-frequency sub-band by the adaptive frequency enhancement unit, and restore global structural information of the high-frequency sub-band and the low-frequency sub-band that is weakened by frequency feature decomposition by the sparse compensation unit, thereby obtaining enhanced high-frequency sub-band and low-frequency sub-band;

[0163] A reconstruction module, configured to perform inverse wavelet transform based on the enhanced high-frequency sub-band and low-frequency sub-band to generate a reconstructed signal;

[0164] A fusion module is used to obtain the gated fusion features of hyperspectral imaging and lidar data based on the image blocks and the reconstructed signals through a gating mechanism, splice the gated fusion features of the hyperspectral imaging and lidar data in the channel dimension, and adopt a channel-space dual attention mechanism to generate standardized multimodal fusion features.

[0165] The multimodal fusion feature generation device provided in the second embodiment of the present invention can execute the multimodal fusion feature generation method provided in the first embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0166] Embodiment 3: This embodiment provides an electronic terminal, including a processor and a memory connected to the processor, wherein a computer program is stored in the memory, and the processor is configured to operate according to the instructions to execute the steps of the method described in Embodiment 1.

[0167] The electronic terminal provided in the third embodiment of the present invention can execute the adaptive frequency enhancement and sparse compensation method provided in the first embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0168] Embodiment 4: This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in Embodiment 1 are implemented, and the computer program has functional modules and beneficial effects corresponding to the execution method.

[0169] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, apparatuses, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0170] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (apparatus), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0171] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0173] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A multimodal fusion feature generation method, characterized in that: include: Acquire hyperspectral imaging and lidar data, and construct an image block centered on the pixel for each pixel sample in the hyperspectral imaging and lidar data; Decomposing the image block by frequency features using wavelet transform to generate corresponding high-frequency sub-bands and low-frequency sub-bands; Inputting the high frequency sub-band and the low frequency sub-band into a pre-built adaptive frequency enhancement and sparse compensation model, wherein the adaptive frequency enhancement and sparse compensation model includes an adaptive frequency enhancement unit and a sparse compensation unit; The adaptive frequency enhancement unit enhances local detail information of the high-frequency sub-band and the low-frequency sub-band, and the sparse compensation unit restores global structural information of the high-frequency sub-band and the low-frequency sub-band that is weakened by frequency feature decomposition, thereby obtaining enhanced high-frequency sub-band and low-frequency sub-band; Performing inverse wavelet transform based on the enhanced high-frequency sub-band and low-frequency sub-band to generate a reconstructed signal; Based on the image blocks and the reconstructed signals, the gated fusion features of the hyperspectral imaging and lidar data are obtained through a gating mechanism, and the gated fusion features of the hyperspectral imaging and lidar data are spliced in the channel dimension. The channel-space dual attention mechanism is adopted to generate standardized multimodal fusion features.

2. The multimodal fusion feature generation method according to claim 1, characterized in that: The adaptive frequency enhancement unit utilizes task-driven dynamic weight learning and Laplace convolution to enhance local detail information of the high-frequency sub-band and the low-frequency sub-band, and obtains the high-frequency sub-band and the low-frequency sub-band after the local detail information is enhanced; The sparse compensation unit captures the spatial correlation of the high-frequency subband and the low-frequency subband after the local detail information is enhanced, suppresses noise interference, and restores the global structural information weakened by the frequency feature decomposition to obtain the enhanced high-frequency subband and the low-frequency subband.

3. The multimodal fusion feature generation method according to claim 2, characterized in that: The adaptive frequency enhancement unit includes: an information quantity evaluation branch, a noise level evaluation branch and a task relevance evaluation branch; generating an information content evaluation component of an input subband by utilizing the information content evaluation branch; generating a noise level estimation component of the input subband using a noise level estimation branch; generating a task relevance evaluation component of the input subband using the task relevance evaluation branch; generating a dynamic weight of the input subband according to the information quantity evaluation component, the noise level evaluation component, and the task relevance evaluation component, and performing adaptive frequency enhancement on the input subband by combining the dynamic weight and using a Laplace convolution form; The input sub-band includes a high-frequency sub-band and a low-frequency sub-band.

4. The multimodal fusion feature generation method according to claim 3, characterized in that: The expression of the information evaluation branch includes: , Among them, S represents the information evaluation component of the input subband, x represents the input subband, Represents a 3×3 convolution operation, GELU represents Gaussian error linear unit, and AdaptiveAvgPool represents adaptive average pooling operation; The expression of the noise level evaluation branch includes: , Where N represents the noise level evaluation component of the input subband, Represents a 1×1 convolution operation; The expression of the task dependency evaluation branch includes: , Among them, R represents the task relevance evaluation component of the input subband, represents a 1×1 convolution operation that reduces the number of channels to half of the original number. represents a 1×1 convolution operation that restores the number of channels to the original number. Represents the Sigmoid activation function; Generating a dynamic weight of the input subband according to the information quantity evaluation component, the noise level evaluation component, and the task relevance evaluation component, and then combining the dynamic weight and performing adaptive frequency enhancement on the input subband using a Laplace convolution form, including: The information quantity evaluation component S, the noise level evaluation component N, and the task relevance evaluation component R are concatenated in the channel dimension. The channel dimension of the concatenated feature is then mapped to the number of channels of the input subband through a 1×1 convolution layer. The sigmoid function is then used to perform channel-by-channel normalization to generate dynamic weights in the range of [0, 1]. The expressions include: , Where W is the dynamic weight, Concat represents the concatenation operation in the channel dimension; The dynamic weight is combined with the Laplace convolution form to perform a convolution operation on the input subband, and the expression includes: , in, represents the convolution form of Laplace, Indicates that the image block is , the vertical axis is The pixel at , r and s represent the index of the convolution kernel, indicating the offset of the convolution kernel in the horizontal and vertical directions, k represents the size of the convolution kernel in the row direction, l represents the size of the convolution kernel in the column direction, k and l determine the size of the convolution kernel; The features after adaptive frequency enhancement are residually connected with the input subband. The expressions include: , in, Represents the high-frequency sub-band and low-frequency sub-band after local detail information is enhanced.

5. The multimodal fusion feature generation method according to claim 2, characterized in that: The sparse compensation unit includes: The spatial correlation is captured through the sparse attention mechanism of the sparse compensation unit, the global structural information weakened by frequency decomposition is restored by the low-frequency feature compensation unit of the sparse compensation unit, and the features are sparsely activated and noise interference is suppressed using the feedforward neural network unit of the sparse compensation unit.

6. The multimodal fusion feature generation method according to claim 5, characterized in that: include: The high-frequency sub-band and low-frequency sub-band after the input local detail information enhancement are normalized by the normalization operation. The expressions include: , in, Represents the high-frequency sub-band and low-frequency sub-band after local detail information enhancement, represents the first normalized feature, Representation layer normalization operation, represents global average pooling, represents the fully connected layer, Represents the Sigmoid activation function; The sparse attention mechanism constructs a local attention mask to force each spatial position to establish attention association only with the local area within its 3×3 neighborhood, thus constraining the range of feature interaction, as shown in the formula: , in, Indicates that the horizontal axis in the first normalized feature map is , the vertical axis is The local attention mask at ; A feature representation with rich contextual information is obtained through a multi-head attention mechanism, where the query vector, key vector, and value vector are derived from the mapping of the first normalized feature. The output features of the sparse attention mechanism are obtained through dynamic feature aggregation. The expression includes: , , in, represents the multi-head attention mechanism, A represents the output feature of the sparse attention mechanism, , B, C, H and Represent the batch size, number of channels, height of feature map and width of feature map respectively, represents the activation function, Q, K, and V represent the query vector, key vector, and value vector, respectively. represents the transpose operation, is the dimension of the key vector, represents the Hadamard product operation, h=1,2,...H, where h represents the number of heads and H represents the maximum number of heads; The output features of the sparse attention mechanism are transformed to restore the shape of the first normalized features, and the output features of the restored shape are obtained. The output features of the restored shape are subjected to a pre-built sparse gating mechanism and are residually connected with the high-frequency sub-band and the low-frequency sub-band after local detail information enhancement to obtain the features after residual connection. The expression includes: , in, is the sparse threshold, Represents the features after residual connection, , represented as the output feature of the restored shape; The low-pass filter convolution operation is performed on the residual connection feature to extract its low-frequency feature. The expression includes: , in, represents the low-pass filter convolution operation, It is a low-frequency feature; The low-frequency features are input into the low-frequency feature compensation unit, and the compensation features are constructed through a two-layer convolution structure with nonlinear activation. The expressions include: , in, represents the compensation characteristic, Represents a 3×3 convolution kernel operation with a bias term, which is used to keep the feature map size unchanged. ReLU stands for rectified linear unit. Represents a 3×3 convolution kernel operation without a bias term, which is used to generate compensatory features that match the number of input channels while preserving the spatial resolution of the feature map; Based on the compensation feature, a residual connection is performed with the high-frequency sub-band and low-frequency sub-band after the local detail information is enhanced after the sparse gating mechanism to achieve directional enhancement of the global structural information. The expression includes: , in, Represents the features after global structural information enhancement; The features after the global structural information enhancement are normalized by a normalization operation, and the expression includes: , in, represents the second normalized feature; Based on the second normalized feature, the feedforward neural network unit generates a channel description vector through adaptive average pooling, and converts it into an attention weight through a learnable weight matrix and a Sigmoid function. The attention weight is used to highlight strong channels and suppress weak channels, and the output features of the feedforward neural network unit are obtained. The expression includes: , in, represents the output features of the feedforward neural network unit, Represents a learnable residual scaling factor, which is used to dynamically adjust the contribution ratio of the compensation feature. LeakyReLU represents a leaky linear rectifier unit. Represents a 3×3 convolution operation with an expanded dimension, which is used to expand the number of channels to twice the original dimension. Represents a 3×3 convolution operation to restore the number of channels to the original number. represents the weight matrix of the gating network; The output features of the feedforward neural network unit are subjected to a sparse gating mechanism and then residually connected with the high-frequency sub-band and low-frequency sub-band after local detail information enhancement, as shown in the formula: , in, Represents the high-frequency subband and low-frequency subband enhanced by the adaptive frequency enhancement and sparse compensation model.

7. The multimodal fusion feature generation method according to claim 1, characterized in that: An inverse wavelet transform is performed based on the enhanced high-frequency subband and low-frequency subband to generate a reconstructed signal. Based on the image block and the reconstructed signal, a gating mechanism is used to obtain the gated fusion features of the hyperspectral imaging and lidar data. The gated fusion features of the hyperspectral imaging and lidar data are spliced in the channel dimension, and a channel-space dual attention mechanism is used to generate a standardized multimodal fusion feature. The expression includes: , , , , in, and Respectively represent the features of hyperspectral imaging and lidar data after gated fusion, and represent image patches of hyperspectral imaging and lidar data, respectively, and represent the reconstructed signals of hyperspectral imaging and lidar data respectively, Represents a cascade of double-layer convolution and nonlinear activation functions, Tanh represents the hyperbolic tangent activation function, represents a 3×3 convolution operation, represents a 1×1 convolution operation aligned with the channel dimension, represents the Hadamard product operation; The features after gated fusion of the hyperspectral imaging and lidar data are spliced using a channel dimension strategy to obtain their multimodal fusion features. The expression includes: , in, Represents multimodal fusion features, and Concat represents the concatenation operation in the channel dimension; The multimodal fusion features are passed through the channel-space dual attention mechanism to generate standardized multimodal fusion features, the expressions of which include: , , in, represents the standardized multimodal fusion features, represents the standardized multimodal fusion features, represents the channel attention map, represents the spatial attention map.

8. A multimodal fusion feature generation device, characterized in that: include: A data acquisition module is used to acquire hyperspectral imaging and lidar data, and construct an image block centered on the pixel for each pixel sample in the hyperspectral imaging and lidar data; A feature decomposition module, configured to perform frequency feature decomposition on the image block using wavelet transform to generate corresponding high-frequency sub-bands and low-frequency sub-bands; An input module, configured to input the high frequency sub-band and the low frequency sub-band into a pre-built adaptive frequency enhancement and sparse compensation model, wherein the adaptive frequency enhancement and sparse compensation model includes an adaptive frequency enhancement unit and a sparse compensation unit; an enhancement module, configured to enhance local detail information of the high-frequency sub-band and the low-frequency sub-band by the adaptive frequency enhancement unit, and restore global structural information of the high-frequency sub-band and the low-frequency sub-band that is weakened by frequency feature decomposition by the sparse compensation unit, thereby obtaining enhanced high-frequency sub-band and low-frequency sub-band; A reconstruction module, configured to perform inverse wavelet transform based on the enhanced high-frequency sub-band and low-frequency sub-band to generate a reconstructed signal; A fusion module is used to obtain the gated fusion features of hyperspectral imaging and lidar data based on the image blocks and the reconstructed signals through a gating mechanism, splice the gated fusion features of the hyperspectral imaging and lidar data in the channel dimension, and adopt a channel-space dual attention mechanism to generate standardized multimodal fusion features.

9. An electronic terminal, characterized in that: The method comprises a processor and a memory connected to the processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are executed.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-source remote sensing image classification method based on key band retrieval attention mechanism

    CN120198818A

  • Real-time face blind restoration method and device based on identity constraint and frequency domain enhancement

    CN120318122A

  • Two-order lightweight network panchromatic sharpening method combining guided filtering and nsct

    WO2023000505A1