Multispectral image and hyperspectral image fusion method and system
Through the spectral self-attention mechanism and the space-spectral fusion module, the problem of insufficient spectral feature extraction in the existing technology is solved, and high-quality multi-spectral and hyperspectral images are achieved, which improves the fusion performance and image quality.
Patent Information
- Application Number
- CN202510266257.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-18
AI Technical Summary
The existing multispectral and hyperspectral image fusion methods have shortcomings in spectral feature extraction, ignoring the global features in the spectral dimension, resulting in poor performance in spectral fidelity and failing to effectively utilize the complementary information between multispectral and hyperspectral images.
The spectral self-attention mechanism is used for feature extraction, and deep fusion is carried out through the space-spectral fusion module and the graph-based spectral perception module. Combined with the encoder and decoder of the spectral self-attention mechanism, deep fusion of spatial and spectral dimensions is achieved, and spectral feature extraction is enhanced.
The spatial and spectral resolution of the fusion image is improved, the fusion performance is improved, and the fidelity and quality of spectral characteristics is ensured. It is suitable for agricultural monitoring and military target recognition.
Smart Images

Figure CN120339016A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for fusing multi - spectral images and hyperspectral images, belonging to the fields of deep learning and signal processing. Background Art
[0002] Multi - spectral and hyperspectral image fusion (MHIF) is an important research direction in the field of image processing in recent years. Its main goal is to combine high - resolution multi - spectral images (HR - MSI) with low - resolution hyperspectral images (LR - HSI) to generate high - resolution hyperspectral images. High - resolution hyperspectral images play an important role in many application fields. For example, in the agricultural field, high - resolution hyperspectral images can help monitor the growth status of crops. By analyzing the spectral features in the hyperspectral images, the nutritional status, moisture content, and pest and disease conditions of crops can be accurately detected. In the military field, high - resolution hyperspectral images can be used to identify camouflaged targets, such as camouflaged military vehicles. By penetrating the camouflage materials through spectral features, the accuracy of target recognition can be improved. The continuous progress and optimization of multi - spectral and hyperspectral image fusion technology (MHIF) play an extremely important role in promoting the development of hyperspectral imaging technology, computer vision technology, and related application fields. It can not only provide higher - quality image data support for fields such as agriculture and military, but also show great application potential in many fields such as environmental protection, resource exploration, urban planning, and disaster monitoring, providing more effective technical means to solve practical problems.
[0003] Currently, the multi - spectral and hyperspectral image fusion technology based on early fusion is the mainstream method in this field. This method first performs upsampling on the low - resolution hyperspectral image (LR - HSI) to match the spatial resolution of the high - resolution multi - spectral image (HR - MSI). Subsequently, the upsampled hyperspectral image and the multi - spectral image are concatenated in the channel dimension to form a fused input data. This fused data is input into a neural network, and the network learns and extracts the joint representation of spatial and spectral features, and finally generates a fused image with high spatial resolution and high spectral resolution.
[0004] Specifically, the method based on early fusion concatenates the hyperspectral image and the multi - spectral image in the channel dimension, enabling the network to process data of both modalities simultaneously. In the neural network, through operations such as convolutional layers, pooling layers, and activation functions, the network can extract the spatial and spectral features of the fused data. To ensure the effective retention of spectral information, the upsampled hyperspectral image is added when outputting the result. This strategy ensures the performance of the fused image in spectral fidelity to a certain extent while improving the spatial resolution, thus realizing the generation of images with high spatial resolution and high spectral resolution.
[0005] The defects of existing multi - spectral and hyperspectral image fusion methods based on early fusion are as follows: Most existing multi - spectral and hyperspectral image fusion methods adopt the convolutional neural network (CNN) structure. The core of this structure lies in extracting the spatial features of images through operations such as convolutional layers, activation functions, and pooling layers. However, these methods often focus too much on the extraction of spatial local features and neglect the importance of global features. Hyperspectral images have the characteristics of spatial sparsity and spectral density, which means that the information is more dispersed in the spatial dimension while containing rich details in the spectral dimension. Most existing fusion methods only focus on the extraction of spatial features and neglect the global features in the spectral dimension, resulting in obvious deficiencies in spectral feature extraction.
[0006] In addition, some existing models simply regard the fusion of multi - spectral and hyperspectral images as a hyperspectral image super - resolution problem, mainly emphasizing the improvement of spatial resolution. This view neglects the spectral noise problem introduced by multi - spectral images, especially in Transformer models that only extract features in the spatial dimension. These models fail to fully consider the differences and connections between multi - spectral and hyperspectral images in the spectral dimension, resulting in poor performance of the fusion results in terms of spectral fidelity. Most models simply splice the data of the two modalities at the input end, lacking a customized fusion strategy based on the inherent features of the data. This simple splicing method cannot effectively utilize the complementary information between the two modalities and is difficult to fully exploit the potential of multi - spectral and hyperspectral image fusion. Summary of the Invention
[0007] To solve the problems existing in the above - mentioned background technology, the present invention provides a multi - spectral image and hyperspectral image fusion method and system, which adopts a spectral self - attention mechanism. Its core is to perform token partitioning and self - attention calculation in the spectral dimension. This method is designed specifically for the characteristics of hyperspectral images with spatial sparsity and spectral density, and can better solve the problem of insufficient spectral feature extraction in the fusion task of hyperspectral and multi - spectral images, thereby generating higher - quality hyperspectral images. To ensure the complementarity of information between multi - spectral and hyperspectral images in the spatial and spectral dimensions, the present invention proposes a spatial - spectral fusion module (S2FM), which can enhance the spatial - spectral features of the fusion results. In addition, the graph - based spectral perception module can further enhance the spectral features of the fusion results through graph convolution. Through the above optimization strategies, the full fusion and complementarity of spatial and spectral information of hyperspectral and multi - spectral images are ensured, which has important value for the research and application of MHIF.
[0008] Term Explanation:
[0009] Multispectral Image: A multispectral image refers to an image obtained by imaging a target scene through multiple specific bands (usually including visible light, infrared, ultraviolet, etc.). This image can capture information beyond the visible range of the human eye and is widely used in fields such as agriculture, healthcare, and environmental monitoring. The number of channels in a multispectral image is generally between 3 and 10.
[0010] Hyperspectral Image: A hyperspectral image is similar to a multispectral image. Through imaging technology with extremely high spectral resolution, it can record the reflection or radiation intensity of each pixel at multiple continuous spectral bands. Compared with a multispectral image, the number of channels in a hyperspectral image is usually in the dozens to hundreds, and can even reach thousands, depending on the application scenario and the design of the sensor. Due to the limitations of the inherent defects of the sensor, when imaging, a trade-off needs to be made between the spectral resolution and the spatial resolution of the image, resulting in the generated image generally being a high-spatial-resolution multispectral image or a low-spatial-resolution hyperspectral image.
[0011] The technical solution of the present invention is as follows:
[0012] The first aspect of the present invention provides a method for fusing multispectral images and hyperspectral images, including:
[0013] Input a high-resolution multispectral image and a low-resolution hyperspectral image into a multispectral image and hyperspectral image fusion model for training. After the model training is completed, test the model and evaluate the algorithm performance;
[0014] Step 1: Independently encode and extract features from the high-resolution multispectral image and the low-resolution hyperspectral image, and perform multi-scale feature extraction to obtain the hyperspectral image features and multispectral image features after feature extraction;
[0015] Step 2: Input the hyperspectral image features and multispectral image features obtained in Step 1 into a spatial-spectral fusion module for in-depth fusion in the spatial dimension and the spectral dimension to obtain a unified feature representation;
[0016] The spatial-spectral fusion module includes a spatial attention mechanism and a channel attention mechanism;
[0017] The spatial attention mechanism generates two two-dimensional feature maps through average pooling and max pooling operations along the channel axis, splices the two two-dimensional feature maps, and generates a spatial attention map through a convolutional operation;
[0018] The channel attention mechanism compresses the spatial dimension of the feature map into a real number through global average pooling and global max pooling operations, then learns the importance weights of each channel through a shared multi-layer perceptron (MLP), and finally generates a channel attention map through a Sigmoid function;
[0019] Step 3: Learn the correlation between channels of the hierarchical features to obtain the output features; the hierarchical features represent the hyperspectral image features and multispectral image features after feature extraction in Step 1.
[0020] After upsampling the output of the spatial-spectral fusion module, it is concatenated with the output features; then, layer-by-layer decoding is performed, and after decoding, it is added to the low-resolution hyperspectral image to obtain the final fusion result.
[0021] Preferably according to the present invention, an encoder based on a spectral self-attention mechanism is used to separately encode and extract features from the high-resolution multispectral image and the low-resolution hyperspectral image, and multi-scale feature extraction is performed to obtain the hyperspectral image features and multispectral image features after feature extraction, including:
[0022] The encoder based on the spectral self-attention mechanism includes a spectral self-attention module and a downsampling module.
[0023] The low-resolution hyperspectral image (UP-HIS) and the high-resolution multispectral image (HR-MSI) are respectively multi-scale encoded by the encoder based on the spectral self-attention mechanism, where the spectral self-attention module extracts features in the spectral dimension, and then the downsampling module performs downsampling in the spatial domain to reduce the resolution, thereby introducing multi-scale information.
[0024] Preferably according to the present invention, the spectral self-attention module extracts features in the spectral dimension, including:
[0025] The spectral self-attention module includes a multi-head attention layer, an aggregation layer, and a position encoding layer.
[0026] The position encoding layer includes two convolutional layers, a GELU activation function layer, and a reshaping operation.
[0027] The input resolution of the low-spatial-resolution hyperspectral image, that is, the low-resolution hyperspectral image (LR-HSI), is (32, 32, C), where C is the number of spectral channels of the hyperspectral image, and the input resolution of the high-spatial-resolution multispectral image, that is, the high-resolution multispectral image (HR-MSI), is (64, 64, c), where c is the number of spectral channels of the multispectral image. After upsampling the space of the low-spatial-resolution hyperspectral image (LR-HSI), the upsampled low-resolution hyperspectral image (UP-HSI) with a size of (64, 64, C) is obtained.
[0028] Input tensor That is, the low-resolution hyperspectral image and the high-resolution multispectral image are reshaped into form as the input of the spectral self-attention module, where H, W, and C are the length, width, and number of spectral channels of the input tensor respectively; subsequently, X is linearly projected into a query Key Sum value It is expressed as:
[0029] Q = XW Q , K = XW K , V = XW V ;
[0030] Where W Q , W K , W V All represent learnable parameters;
[0031] Subsequently, in the spectral dimension, Q, K, and V are divided into N groups and input into the multi - head attention layer. The multi - head attention layer performs self - attention calculation in the channels to learn different subspace representations. For the j - th head, it is as follows:
[0032]
[0033] Where, A j Represents the calculated attention weight, Represents the transposed matrix of K j , K j Represents the key of the j - th head, Q j Represents the query of the j - th head, softmax represents the normalized exponential function. Considering that the spectral density varies greatly with wavelength, the learnable parameter matrix σ j Is used to adjust the attention weight;
[0034] Finally, the multi - head attention layer is aggregated through the learnable parameter W and the learnable position encoding is added, which is expressed as:
[0035]
[0036] Where, SSA represents the spectral self - attention operation, V j Represents the value of the j - th head, concat represents the concatenation operation on the input in the channels, V represents the value matrix, A represents the attention matrix, f p (·) represents the position encoding layer, which is used to generate the learnable position encoding and specifically includes two depth - convolution 3×3 layers, one GELU activation function, and a reshape operation (Reshape); the input features are output after passing through the SSA module, that is, the spectral self - attention module Where, H, W, and C are the length, width, and number of channels of the input features respectively.
[0037] Preferably according to the present invention, downsampling is performed spatially through the downsampling module; it includes:
[0038] The downsampling module includes a Harr wavelet transform module and a convolutional layer;
[0039] To ensure maximum retention of high-frequency details during downsampling, a wavelet transform-based downsampling module is used for downsampling. The Harr wavelet transform module filters the output of the spectral self-attention module to perform filtering operations, respectively obtaining low-frequency information A, horizontal high-frequency information H, vertical high-frequency information V, and diagonal high-frequency information D; A, H, V, and D are concatenated in the channel dimension, and feature extraction is performed through the convolution operation of the convolutional layer to obtain the extracted hyperspectral image features and multispectral image features
[0040] According to the preference of the present invention, the hyperspectral image features and multispectral image features are input into the spatial-spectral fusion module for in-depth fusion in the spatial dimension and spectral dimension; including:
[0041] Among them, the operations in the spatial-spectral fusion module include: cross representation, spatial attention mechanism, channel attention mechanism, concatenation, convolution, non-linear activation, and batch normalization;
[0042] The cross representation includes convolution, non-linear activation, and batch normalization;
[0043] The spatial-spectral fusion module (S2FM) is used to capture the spectral and spatial correlation features of two modalities, namely high-spatial-resolution multispectral images and low-spatial-resolution hyperspectral images, and filter out redundant features, thereby obtaining a unified fusion representation;
[0044] Receive the encoded hyperspectral image features and multispectral image features as inputs, concatenate X HsI and X MSI in the channel dimension, and after convolution, non-linear activation, and batch normalization processing, obtain the cross representation
[0045] X cross includes the mixed information of X HSI and X MsI respectively perform spatial attention operations M cross on X s through the spatial attention mechanism, and perform channel attention operations M c through the channel attention mechanism to obtain the corresponding spatial attention map M s (F) and channel attention map M c (F), M s and M c are represented as follows:
[0046] M c(F) = σ(MLP(Avgpool(F)) + MLP(MaxPool(F)));
[0047] M s (F) = σ(f 7×7 ([Avgpool(F)); MLP(MaxPool(F)]));
[0048] Wherein, MLP represents a multi - layer perceptron (fully - connected layer), F represents the input feature, Avgpool represents the average pooling operation, MaxPool represents the max - pooling operation, f 7×7 represents performing a 7×7 convolution operation, and σ represents the sigmoid function;
[0049] To achieve selective filtering of information, multiply X HSI by the spatial attention map, multiply X MSI by the channel attention map, and after concatenation, convolution, non - linear activation, and batch normalization processing, obtain a unified feature representation Thereby realizing the deep - feature fusion of the two - modality data in the spatial dimension and the spectral dimension.
[0050] According to the preferred embodiment of the present invention, the graph - based spectral perception module learns the correlation between channels of hierarchical features to obtain output features. The hierarchical features represent the hyperspectral image features and multi - spectral image features after feature extraction in step 1, and include:
[0051] The graph - based spectral perception module includes an adaptive graph convolution module and a residual connection;
[0052] The graph - based spectral perception module (GS2M) includes two branches. One branch uses the adaptive graph convolution module (AGCM) to learn the correlation within the spectral dimension and effectively filter out irrelevant noise; the other branch adopts a residual connection to prevent the loss of feature representation. The core of the graph - based spectral perception module lies in the adaptive graph convolution module. The AGCM captures the adjacency relationship between different feature groups by constructing an adaptive graph and has been proven to have efficient feature extraction ability in the field of defect detection. The AGCM regards each channel as a feature vertex and then analyzes the relationship between channels in a graph - based manner to model the correlation between channels;
[0053] To further enhance the spectral features of the fusion result, the hierarchical features enter the graph-based spectral sensing module (GS2M). For the adaptive graph convolution module to obtain feature weights, it first processes the hierarchical features through convolution and a Softmax layer, and then uses the learned adaptive adjacency matrix to represent the relationship between features. After that, through convolution and ReLU processing, the output result is obtained. Finally, to prevent the problem of vanishing gradients caused by deepening the network, a residual connection is introduced, directly connecting the input hierarchical features to the output result. Specifically, the input hierarchical features are added to the output result to effectively transmit the information flow and obtain the output features.
[0054] Further preferably, the adaptive adjacency matrix A in the adaptive graph convolution module (AGCM) includes three parts: A0, A1, and A2.
[0055] Among them, A0 is an N×N identity matrix representing each feature vertex itself; A1 is an N×N diagonal matrix representing self-attention, where the weights of the feature vertices are obtained through a one-dimensional convolutional layer and a softmax layer and then arranged on the diagonal; A2 is an N×N adjacency matrix optimized through backpropagation to learn the relationship between any two feature vertices. The element values of A2 are not restricted, allowing for more flexible global information learning. It is expressed as follows:
[0056] AGCM(x,A) = σ(W·x·A);
[0057] A = A0×A1 + A2;
[0058] Among them, x represents the input feature, W represents the learnable matrix, and A represents the adjacency matrix.
[0059] From the above, the graph-based spectral sensing module (GS2M) is expressed as the following formula:
[0060] y = x·σ(F′ r (ReLU(AGCM(F r (GAP(x)), A)))) + x;
[0061] Among them, F r ′ and F r represent the MLP layer, GAP represents the global pooling operation, ReLU represents the activation function, x represents the input feature, y represents the output feature, AGCM represents the adaptive graph convolution module, and σ represents the Sigmoid activation function.
[0062] Preferably according to the present invention, the unified feature representation obtained by the spatial-spectral fusion module is input into the decoder based on the spectral self-attention mechanism, and the output features of the graph-based spectral perception module are input into the symmetric layer of the corresponding decoder based on the spectral self-attention mechanism. Since the overall framework follows the design of UNet, the network presents a symmetric structure in the encoder and decoder. Due to upsampling and downsampling, the network modules with the same input size in the encoder and decoder are called corresponding layers. Then, the decoder based on the spectral self-attention mechanism outputs a spectral image with high spatial resolution and high spectral resolution, including:
[0063] The decoder based on the spectral self-attention mechanism includes a spectral self-attention module and an upsampling module;
[0064] The upsampling module includes an inverse Harr wavelet transform module and a convolutional layer;
[0065] Among them, the spectral self-attention module extracts features in the spectral dimension, and the upsampling module improves the resolution in the spatial dimension, thereby introducing multi-scale information. The decoder based on the spectral self-attention mechanism receives the output of the spatial-spectral fusion module as input features for upsampling. In order to ensure that the high-frequency details are maximally retained during the upsampling process, an upsampling module based on inverse wavelet transform is used. Among them, the inverse Harr wavelet transform module processes the input features of the decoder, and the obtained low-frequency information A1, horizontal high-frequency information H1, vertical high-frequency information V1, and diagonal high-frequency information D1 are concatenated in the spatial dimension to restore the features After that, through the convolution operation, non-linear activation, and batch normalization operation of the convolutional layer, the output features
[0066] X dec_out are concatenated with y. Then, after two-dimensional convolution, they are input into the spectral self-attention module, decoded layer by layer, and then added to the low-resolution hyperspectral image (UP-HIS) to obtain the final fusion result.
[0067] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the multi-spectral image and hyperspectral image fusion method.
[0068] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the multi-spectral image and hyperspectral image fusion method.
[0069] The second aspect of the present invention provides a multi-spectral image and hyperspectral image fusion system, including:
[0070] The multi-scale feature extraction module is configured to: separately encode and extract features from the high-resolution hyperspectral image and the low-resolution hyperspectral image, and perform multi-scale feature extraction to obtain the hyperspectral image features and multi-spectral image features after feature extraction;
[0071] The deep fusion module is configured to: input the hyperspectral image features and multi-spectral image features obtained in step 1 into the spatial-spectral fusion module to perform deep fusion in the spatial dimension and spectral dimension, and obtain a unified feature representation;
[0072] The spatial-spectral fusion module includes a spatial attention mechanism and a channel attention mechanism;
[0073] The spatial attention mechanism generates two two-dimensional feature maps through average pooling and max pooling operations along the channel axis, and after splicing the two two-dimensional feature maps, generates a spatial attention map through a convolution operation;
[0074] The channel attention mechanism compresses the spatial dimension of the feature map into a real number through global average pooling and global max pooling operations, then learns the importance weights of each channel through a shared multi-layer perceptron (MLP), and finally generates a channel attention map through the Sigmoid function;
[0075] The spectral image generation module is configured to: learn the correlation between channels of hierarchical features to obtain output features; the hierarchical features represent the hyperspectral image features and multi-spectral image features after feature extraction by the multi-scale feature extraction module;
[0076] After upsampling the output of the spatial-spectral fusion module, it is spliced with the output features; then, layer-by-layer decoding is performed, and after decoding, an addition operation is performed with the low-resolution hyperspectral image to obtain the final fusion result.
[0077] The beneficial effects of the present invention are as follows:
[0078] 1. The present invention follows the technical route of multi-spectral and hyperspectral image fusion, combines the spectral self-attention mechanism to achieve feature extraction, and at the same time performs deep processing on the fused features using the spatial-spectral fusion module, effectively improving the feature matching performance. In addition, by introducing a graph-based spectral perception module, the spectral features are enhanced, forming a high-quality fused image, and finally significantly improving the overall fusion performance on the basis of the original architecture, which has important research and application value for fields such as agricultural monitoring and military target recognition.
[0079] 2. The present invention replaces the original feature extraction method with a spectral self-attention mechanism. By dividing the features of hyperspectral images and multispectral images into multiple spectral subsets and performing feature extraction and fusion sequentially according to the spectral dimension, it can effectively solve the problem of insufficient spectral feature extraction. At the same time, during the processing of each spectral subset, the complementarity of spatial and spectral information is considered simultaneously to ensure that the fused image can retain the optimal features of both modalities, thereby improving the fusion performance.
[0080] 3. The present invention processes the successfully matched features using a spatial-spectral fusion module, replacing the original simple stitching strategy, avoiding the interference of redundant information on the fusion result, and thus improving the quality of the fused image. This fusion method can effectively improve the spatial and spectral resolutions of the image, and has a reasonable computational cost, good robustness and scalability, which is of great significance for the research and application of multispectral and hyperspectral image fusion.
[0081] 4. The present invention uses a graph-based spectral perception module to enhance the spectral features of the fused image, further consolidating the spectral features, which can improve the quality of the fused image and significantly improve the overall performance of the fusion system. This module can be applied to any occasion where spectral features need to be enhanced and can be used for reference by other image fusion methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 is a framework diagram of the multispectral image and hyperspectral image fusion model proposed by the present invention;
[0083] Figure 2 is a schematic diagram of spectral self-attention calculation;
[0084] Figure 3 is a schematic diagram of the spatial-spectral fusion module;
[0085] Figure 4 is a schematic diagram of the graph-based spectral perception module;
[0086] Figure 5(a) is a visualization schematic diagram of the test image of the PAVIA dataset;
[0087] Figure 5(b) is a visualization schematic diagram of the fused image of the test image of the PAVIA dataset used in the present invention;
[0088] Figure 5(c) is a visualization schematic diagram of the error of the test image of the PAVIA dataset used in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0089] The present invention will be further described below by way of examples in conjunction with the accompanying drawings, but is not limited thereto.
[0090] Example 1
[0091] A method for fusing multi-spectral images and hyperspectral images, comprising:
[0092] Input the high-resolution multi-spectral image and the low-resolution hyperspectral image into the multi-spectral image and hyperspectral image fusion model for training. After the model training is completed, test the model and evaluate the algorithm performance;
[0093] Step 1: Separately encode and extract features from the high-resolution multi-spectral image and the low-resolution hyperspectral image, and perform multi-scale feature extraction to obtain the hyperspectral image features and multi-spectral image features after feature extraction;
[0094] Step 2: Input the hyperspectral image features and multi-spectral image features obtained in Step 1 into the spatial-spectral fusion module for in-depth fusion in the spatial dimension and spectral dimension to obtain a unified feature representation;
[0095] The spatial-spectral fusion module includes a spatial attention mechanism and a channel attention mechanism;
[0096] The spatial attention mechanism generates two two-dimensional feature maps through average pooling and max pooling operations along the channel axis, splices the two two-dimensional feature maps, and generates a spatial attention map through a convolutional operation;
[0097] The channel attention mechanism compresses the spatial dimension of the feature map into a real number through global average pooling and global max pooling operations, then learns the importance weights of each channel through a shared multi-layer perceptron (MLP), and finally generates a channel attention map through the Sigmoid function;
[0098] Step 3: Learn the correlation between the channels of the hierarchical features to obtain the output features; the hierarchical features represent the hyperspectral image features and multi-spectral image features after feature extraction in Step 1;
[0099] After upsampling the output of the spatial-spectral fusion module, splice it with the output features; then, perform layer-by-layer decoding, and after decoding, perform an addition operation with the low-resolution hyperspectral image to obtain the final fusion result.
[0100] Embodiment 2
[0101] A method for fusing multi-spectral images and hyperspectral images according to Embodiment 1, wherein:
[0102] Use an encoder based on the spectral self-attention mechanism to separately encode and extract features from the high-resolution multi-spectral image and the low-resolution hyperspectral image, and perform multi-scale feature extraction to obtain the hyperspectral image features and multi-spectral image features after feature extraction; including:
[0103] The encoder based on the spectral self-attention mechanism includes a spectral self-attention module and a downsampling module;
[0104] The low-resolution hyperspectral image (UP-HIS) and the high-resolution multispectral image (HR-MSI) are respectively encoded at multiple scales by the encoder based on the spectral self-attention mechanism. Among them, the spectral self-attention module extracts features in the spectral dimension, and then the downsampling module downsamples spatially to reduce the resolution, thereby introducing multi-scale information;
[0105] The spectral self-attention module extracts features in the spectral dimension, including:
[0106] The spectral self-attention module includes a multi-head attention layer, an aggregation layer, and a positional encoding layer;
[0107] The positional encoding layer includes two convolutional layers, a GELU activation function layer, and a reshaping operation;
[0108] The input resolution of the low-spatial-resolution hyperspectral image, that is, the low-resolution hyperspectral image (LR-HSI), is (32, 32, C), where C is the number of spectral channels of the hyperspectral image. The input resolution of the high-spatial-resolution multispectral image, that is, the high-resolution multispectral image (HR-MSI), is (64, 64, c), where c is the number of spectral channels of the multispectral image. After upsampling the space of the low-spatial-resolution hyperspectral image (LR-HSI), the upsampled low-resolution hyperspectral image (UP-HSI) of (64, 64, C) is obtained;
[0109] Input tensor That is, the low-resolution hyperspectral image and the high-resolution multispectral image are reshaped into form as the input of the spectral self-attention module, where H, W, and C are the length, width, and number of spectral channels of the input tensor respectively. Subsequently, X is linearly projected into a query key and value which are expressed as:
[0110] Q = XW Q , K = XW K , V = XW V ;
[0111] where W Q , W K , W V all represent learnable parameters;
[0112] Subsequently, in the spectral dimension, Q, K, and V are divided into N groups and input into the multi-head attention layer. The multi-head attention layer performs self-attention calculation in the channels to learn different subspace representations. For the j-th head, it is as follows:
[0113]
[0114] Among them, A j represents the calculated attention weight, represents the transposed matrix of K j , where K j represents the key of the j-th head, Q j represents the query of the j-th head, and softmax represents the normalized exponential function. Considering that the spectral density varies greatly with the wavelength, the learnable parameter matrix σ j is used to adjust the attention weight;
[0115] Finally, the multi-head attention layer is aggregated through the learnable parameter W and the learnable position encoding is added, which is expressed as:
[0116]
[0117] Among them, SSA represents the spectral self-attention operation, V j represents the value of the j-th head, concat represents the concatenation operation on the input in the channel dimension, V represents the value matrix, A represents the attention matrix, f p (·) represents the position encoding layer, which is used to generate the learnable position encoding, specifically including two depth convolution 3×3 layers, a GELU activation function and a reshape operation (Reshape); the input features are output after passing through the SSA module, that is, the spectral self-attention module Among them, H, W, and C are the length, width, and number of channels of the input features respectively.
[0118] Downsampling is performed spatially through the downsampling module; it includes:
[0119] The downsampling module includes a Harr wavelet transform module and a convolutional layer;
[0120] In order to ensure the maximum retention of high-frequency details during the downsampling process, a wavelet transform-based downsampling module is used during downsampling. The Harr wavelet transform module performs filtering operations on the output of the spectral self-attention module to obtain the low-frequency information A, the horizontal high-frequency information H, the vertical high-frequency information V, and the diagonal high-frequency information D respectively; A, H, V, and D are concatenated in the channel dimension, and feature extraction is performed through the convolution operation of the convolutional layer to obtain the extracted hyperspectral image features and the multispectral image features
[0121] The hyperspectral image features and the multispectral image features are input into the spatial-spectral fusion module for in-depth fusion in the spatial dimension and the spectral dimension; it includes:
[0122] Among them, the operations in the spatial-spectral fusion module include: cross representation, spatial attention mechanism, channel attention mechanism, concatenation, convolution, non-linear activation, and batch normalization;
[0123] The cross representation includes convolution, non-linear activation, and batch normalization;
[0124] The spatial-spectral fusion module (S2FM) is used to capture the spectral and spatial correlation features of two modalities, namely high spatial resolution multi-spectral images and low spatial resolution hyperspectral images, and filter out redundant features, thereby obtaining a unified fusion representation;
[0125] Receive the encoded hyperspectral image features and multi-spectral image features as inputs, concatenate X HSI and X MSI in the channel dimension. After convolution, non-linear activation, and batch normalization processing, a cross representation
[0126] X cross including the mixed information of X HSI and X MSO is obtained. Perform spatial attention operation M cross on X s respectively through the spatial attention mechanism, and perform channel attention operation M c through the channel attention mechanism, to obtain the corresponding spatial attention map M s (F) and channel attention map M c (F), M s and M c are represented as follows:
[0127] M c (F) = σ(MLP(Avgpool(F)) + MLP(MaxPool(F)));
[0128] M s (F) = σ(f 7×7 ([Avgpool(F)); MLP(MaxPool(F)]));
[0129] Among them, MLP represents a multi-layer perceptron (fully connected layer), F represents the input feature, Avgpool represents the average pooling operation, MaxPool represents the max pooling operation, f 7×7 represents the execution of a 7×7 convolution operation, and σ represents the sigmoid function;
[0130] To achieve selective filtering of information, multiply X HSI by the spatial attention map, and X MSIMultiply with the channel attention map, and after processes such as splicing, convolution, non-linear activation, and batch normalization, a unified feature representation is obtained. Thus, deep feature fusion of the two-modal data in the spatial dimension and the spectral dimension is realized.
[0131] The graph-based spectral perception module learns the correlation between channels of hierarchical features to obtain output features. The hierarchical feature representation in step 1 is the hyperspectral image feature and the multispectral image feature after feature extraction, including:
[0132] The graph-based spectral perception module includes an adaptive graph convolution module and a residual connection.
[0133] The graph-based spectral perception module (GS2M) includes two branches. One branch uses the adaptive graph convolution module (AGCM) to learn the correlation within the spectral dimension and effectively filter out irrelevant noise. The other branch adopts a residual connection to prevent the loss of feature representation. The core of the graph-based spectral perception module lies in the adaptive graph convolution module. AGCM captures the adjacency relationship between different feature groups by constructing an adaptive graph and has been proven to have efficient feature extraction capabilities in the defect detection field. AGCM regards each channel as a feature vertex and then analyzes the relationship between channels in a graph way to model the correlation between channels.
[0134] To further enhance the spectral features of the fusion result, the hierarchical features enter the graph-based spectral perception module (GS2M). In order to obtain feature weights, the adaptive graph convolution module first processes the hierarchical features through convolution and a Softmax layer, and then uses the learned adaptive adjacency matrix to represent the relationship between features. After that, after convolution and ReLU processing, the output result is obtained. Finally, to prevent the problem of gradient disappearance caused by the deepening of the network, a residual connection is introduced to directly connect the input hierarchical features to the output result. Specifically, the input hierarchical features are added to the output result to effectively transfer the information flow and obtain the output features.
[0135] The adaptive adjacency matrix A in the adaptive graph convolution module (AGCM) includes three parts: A0, A1, and A2.
[0136] Among them, A0 is an N×N identity matrix representing each feature vertex itself. A1 is an N×N diagonal matrix representing self-attention, where the weights of the feature vertices are obtained through a one-dimensional convolutional layer and a softmax layer and then arranged on the diagonal. A2 is an N×N adjacency matrix optimized through backpropagation to learn the relationship between any two feature vertices. The element values of A2 are not restricted, allowing for more flexible global information learning. The representation is as follows:
[0137] AGCM(x, A) = σ(W·x·A);
[0138] A = A0 × A1 + A2;
[0139] Where x represents the input feature, W represents the learnable matrix, and A represents the adjacency matrix;
[0140] From the above, the graph-based spectral sensing module (GS2M) is expressed as the following formula:
[0141] y = x·σ(F' r (ReLU(AGCM(F r (GAP(x)), A)))) + x;
[0142] Where F r r and F r represent the MLP layer, GAP represents the global pooling operation, ReLU represents the activation function, x represents the input feature, y represents the output feature, AGCM represents the adaptive graph convolution module, and σ represents the Sigmoid activation function;
[0143] Input the unified feature representation obtained by the spatial-spectral fusion module into the decoder based on the spectral self-attention mechanism, and input the output feature of the graph-based spectral sensing module into the symmetric layer of the corresponding decoder based on the spectral self-attention mechanism. Since the overall framework follows the design of UNet, the network presents a symmetric structure in the encoder and decoder. Due to upsampling and downsampling, the network modules with the same input size in the encoder and decoder are called corresponding layers; then, the decoder based on the spectral self-attention mechanism outputs a spectral image with high spatial resolution and high spectral resolution; including:
[0144] The decoder based on the spectral self-attention mechanism includes a spectral self-attention module and an upsampling module;
[0145] The upsampling module includes an inverse Harr wavelet transform module and a convolutional layer;
[0146] Where the spectral self-attention module extracts features in the spectral dimension, and the upsampling module improves the resolution in the spatial dimension, thereby introducing multi-scale information; the decoder based on the spectral self-attention mechanism receives the output of the spatial-spectral fusion module as the input feature for upsampling. To ensure that the high-frequency details are maximally retained during the upsampling process, an upsampling module based on the inverse wavelet transform is used; where the inverse Harr wavelet transform module processes the input feature of the decoder, and the obtained low-frequency information A1, horizontal high-frequency information H1, vertical high-frequency information V1, and diagonal high-frequency information D1 are concatenated in the spatial dimension to restore the feature After convolution operation, non-linear activation, and batch normalization operations in the convolutional layer, the output features are obtained.
[0147] X dec_out It is concatenated with y; then after two-dimensional convolution, it is input into the spectral self-attention module, decoded layer by layer, and then added to the low-resolution hyperspectral image (UP-HIS) to obtain the final fusion result.
[0148] The experiments selected the bilinear interpolation method, the CNN-based SSFCNN, ResTFNet, MSDCNN, and the Transformer-based PSRT model for comparative experiments on three commonly used datasets, CAVE, HARVARD, and PAVIA. The present invention achieved better fusion results in terms of both spatial and spectral metrics for the above datasets.
[0149] The descriptions of various benchmark methods are as follows:
[0150] Bilinear: It means that the low-spatial-resolution hyperspectral image is spatially enlarged by bilinear interpolation.
[0151] SSFCNN: A convolutional neural network architecture that utilizes the fusion of spatial and spectral information is proposed. By extracting and fusing features from the low-resolution hyperspectral image and the high-resolution multispectral image, super-resolution reconstruction of the hyperspectral image is achieved.
[0152] ResTFNet: Combines the residual mechanism and the two-stream architecture. The residual mechanism is used to avoid the problem of gradient vanishing or gradient explosion in the training of deep networks, thereby improving the training efficiency and performance of the network; the two-stream architecture can simultaneously process different types of input data and perform effective feature fusion.
[0153] MSDCNN: By introducing multi-scale feature extraction and densely connected convolutional layers, it can better capture different scale information in the image and make full use of the correlation between features.
[0154] PSRT: Combines the advantages of the pyramid structure and the Transformer architecture. The image is processed at multiple scales through the pyramid structure, and at the same time, the self-attention mechanism of the Transformer is used to establish long-range dependencies, thereby achieving efficient image fusion.
[0155] Ours (HST) is the multi-spectral - hyperspectral fusion method proposed in the present invention.
[0156] The experimental results of each benchmark method on the three datasets are shown in Table 1:
[0157] Table 1
[0158]
[0159]
[0160] The experimental results show that the fusion results of the HST model proposed in the present invention are superior to the existing model methods in both spatial and spectral metrics. On the CAVE and PAVIA datasets, the performance in terms of PSNR exceeds the existing best method by 1 - 2 points, and at the same time, excellent spectral fusion results are also demonstrated, far exceeding similar models.
[0161] Finally, the ablation experiment results of the three strategies proposed in the present invention on the PAVIA validation set are shown in Table 2:
[0162] Table 2
[0163] Method PSNR ERGAS SSIM RMSE SAM Baseline(Spectral Self Attention) 46.0132 0.9408 0.9959 0.0050 1.6673 Baseline+S2FM 46.2841 0.9261 0.9960 0.0049 1.6409 Baseline+GS2M 46.9248 0.8691 0.9967 0.0046 1.5076 Baseline+S2FM+GS2M 47.1525 0.8382 0.9969 0.0045 1.4415
[0164] Among them, S2FM refers to the proposed spatial - spectral fusion module, and GS2M refers to the graph - based spectral perception module. The present invention improves the PSNR by 1.14, reduces the SAM by 0.22, and reduces the ERGAS by 0.11 compared to the Baseline.
[0165] Example 3
[0166] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the multi - spectral image and hyperspectral image fusion method described in Example 1 or 2 are implemented.
[0167] Example 4
[0168] A computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the multi - spectral image and hyperspectral image fusion method described in Example 1 or 2 are implemented.
[0169] Example 5
[0170] A multi - spectral image and hyperspectral image fusion system includes:
[0171] A multi - scale feature extraction module configured to: separately encode and extract features from a high - resolution multi - spectral image and a low - resolution hyperspectral image, and perform multi - scale feature extraction to obtain the hyperspectral image features and multi - spectral image features after feature extraction;
[0172] A deep fusion module configured to: input the hyperspectral image features and multi - spectral image features obtained in step 1 into a spatial - spectral fusion module to perform deep fusion in the spatial dimension and spectral dimension to obtain a unified feature representation;
[0173] The spatial-spectral fusion module includes a spatial attention mechanism and a channel attention mechanism;
[0174] The spatial attention mechanism generates two two-dimensional feature maps through average pooling and max pooling operations along the channel axis, and after concatenating the two two-dimensional feature maps, generates a spatial attention map through a convolution operation;
[0175] The channel attention mechanism compresses the spatial dimension of the feature map into a real number through global average pooling and global max pooling operations, then learns the importance weights of each channel through a shared multi-layer perceptron (MLP), and finally generates a channel attention map through the Sigmoid function;
[0176] The spectral image generation module is configured to: learn the correlation between channels of hierarchical features to obtain output features; the hierarchical features represent the hyperspectral image features and multispectral image features after feature extraction by the multi-scale feature extraction module;
[0177] After upsampling the output of the spatial-spectral fusion module, it is concatenated with the output features; then, layer-by-layer decoding is performed, and after decoding, it is added to the low-resolution hyperspectral image to obtain the final fusion result.
Claims
1. A method for fusing multi-spectral images and hyperspectral images, characterized in that, Including: Step 1: Separately encode and extract features from the high-resolution hyperspectral image and the low-resolution hyperspectral image, and perform multi-scale feature extraction to obtain the hyperspectral image features and multi-spectral image features after feature extraction; Step 2: Input the hyperspectral image features and multi-spectral image features obtained in Step 1 into the spatial-spectral fusion module for in-depth fusion in the spatial dimension and spectral dimension to obtain a unified feature representation; The spatial-spectral fusion module includes a spatial attention mechanism and a channel attention mechanism; The spatial attention mechanism generates two two-dimensional feature maps through average pooling and max pooling operations along the channel axis, and after concatenating the two two-dimensional feature maps, generates a spatial attention map through a convolution operation; The channel attention mechanism compresses the spatial dimension of the feature map into a real number through global average pooling and global max pooling operations, then learns the importance weights of each channel through a shared multi-layer perceptron, and finally generates a channel attention map through the Sigmoid function; Step 3: Learn the correlation between the channels of the hierarchical features to obtain the output features; The hierarchical feature representation refers to the hyperspectral image features and multi-spectral image features after feature extraction in Step 1; After upsampling the output of the spatial-spectral fusion module, concatenate it with the output features; After that, perform layer-by-layer decoding, and add the decoded result to the low-resolution hyperspectral image to obtain the final fusion result.
2. The multispectral image and hyperspectral image fusion method according to claim 1, characterized in that, Use an encoder based on the spectral self-attention mechanism to separately encode and extract features from the high-resolution hyperspectral image and the low-resolution hyperspectral image, and perform multi-scale feature extraction to obtain the hyperspectral image features and multi-spectral image features after feature extraction; including: The encoder based on the spectral self-attention mechanism includes a spectral self-attention module and a downsampling module; The low-resolution hyperspectral image and the high-resolution multi-spectral image are respectively subjected to multi-scale encoding by the encoder based on the spectral self-attention mechanism, where the spectral self-attention module extracts features in the spectral dimension, and then the downsampling module performs downsampling in the space.
3. A method for fusing multi-spectral images and hyperspectral images according to claim 2, characterized in that, The spectral self-attention module extracts features in the spectral dimension; including: The spectral self-attention module includes a multi-head attention layer, an aggregation layer, and a position encoding layer; The position encoding layer includes two convolutional layers, a GELU activation function layer, and a reshaping operation; Input tensor That is, the low-resolution hyperspectral image and the high-resolution multispectral image are reshaped into form, as the input of the spectral self-attention module, where H, W, and C are the length, width, and spectral channel number of the input tensor respectively; subsequently, X is linearly projected into query key and value which are expressed as: Q = XW Q , K = XW K , V = XW V ; Among them, W Q , W K , W V all represent learnable parameters; Subsequently, in the spectral dimension, divide Q, K, and V into N groups and input them into the multi-head attention layer. The multi-head attention layer performs self-attention calculation in the channels. For the j-th head, as follows: Among them, A j represents the calculated attention weight, represents the transposed matrix of K j , and K j represents the key of the j-th head, Q j represents the query of the j-th head, and softmax represents the normalized exponential function, using the learnable parameter matrix σ j to adjust the attention weight; Finally, the multi-head attention layer is aggregated through the learnable parameter W and the learnable position encoding is added, expressed as: Among them, SSA represents the spectral self-attention operation, and V j represents the value of the j-th head, concat represents the concatenation operation on the input in the channel dimension, V represents the value matrix, A represents the attention matrix, and f p (·) represents the position encoding layer, which is used to generate learnable position encodings; the input features are output after passing through the SSA module, namely the spectral self-attention module Among them, H, W, and C are the length, width, and number of channels of the input features respectively.
4. A method for fusing multi-spectral images and hyperspectral images according to claim 3, characterized in that, Perform downsampling in the space through the downsampling module; including: The downsampling module includes a Harr wavelet transform module and a convolutional layer; The output of the spectral self-attention module by the Harr wavelet transform module Performs a filtering operation to obtain low-frequency information A, horizontal high-frequency information H, vertical high-frequency information V, and diagonal high-frequency information D respectively; concatenates A, H, V, and D in the channel dimension, and performs feature extraction through the convolution operation of the convolutional layer to obtain the extracted hyperspectral image features and multispectral image features 5. A method for fusing a multi-spectral image and a hyperspectral image according to claim 4, characterized in that, Input the hyperspectral image features and multi-spectral image features into the spatial-spectral fusion module for in-depth fusion in the spatial dimension and spectral dimension; Including: Among them, the operations in the spatial-spectral fusion module include: cross representation, spatial attention mechanism, channel attention mechanism, concatenation, convolution, non-linear activation, and batch normalization; The cross representation includes convolution, non-linear activation, and batch normalization; Receive the encoded hyperspectral image features and multispectral image features as inputs, concatenate X HSI and X MSI along the channel dimension. After convolution, non-linear activation, and batch normalization, obtain the cross representation X cross including X HSI and X MSO of the mixed information, for X cross respectively perform spatial attention operation M through the spatial attention mechanism s , and perform channel attention operation M through the channel attention mechanism c , to obtain the corresponding spatial attention map M s (F) and channel attention map M c (F), M s and M c are represented as follows: M c (F) = σ(MLP(Avgpool(F)) + MLP(MaxPool(F))); M s (F) = σ(f 7×7 ([Avgpool(F)); MLP(MaxPool(F)])); Among them, MLP represents a multi-layer perceptron, F represents input features, Avgpool represents an average pooling operation, MaxPool represents a max pooling operation, and f 7×7 represents the execution of a 7×7 convolutional operation, and σ represents the sigmoid function; Multiply X HSI by the spatial attention map, and X MSI by the channel attention map. After splicing, convolution, non-linear activation, and batch normalization processing, a unified feature representation is obtained 6. A method for fusing multi - spectral images and hyperspectral images according to claim 1, characterized in that, The graph-based spectral perception module learns the correlation between channels of hierarchical features to obtain output features, representing hierarchical features Step 1 performs high spectral image features and multi-spectral image features after feature extraction; including: The graph-based spectral perception module includes an adaptive graph convolution module and a residual connection; The graph-based spectral perception module includes two branches. One branch uses the adaptive graph convolution module to learn the correlation within the spectral dimension; the other branch adopts a residual connection to prevent the loss of feature representation; The hierarchical features enter the graph-based spectral perception module. The adaptive graph convolution module first performs convolution and Softmax layer processing on the hierarchical features, and then uses the learned adaptive adjacency matrix to represent the relationship between features; after that, after convolution and ReLU processing, the output result is obtained; finally, a residual connection is introduced to directly connect the input hierarchical features to the output result. Specifically, the input hierarchical features are added to the output result to obtain the output features; Further preferably, the adaptive adjacency matrix A in the adaptive graph convolution module includes three parts: A0, A1, and A2; Among them, A0 is an N×N identity matrix representing each feature vertex itself; A1 is an N×N diagonal matrix representing self-attention, where the weights of the feature vertices are obtained through a one-dimensional convolution layer and a softmax layer and then arranged on the diagonal; A2 is an N×N adjacency matrix optimized through backpropagation to learn the relationship between any two feature vertices; expressed as follows: AGCM(x,A) = σ(W·x·A); A = A0×A1 + A2; Among them, x represents the input feature, W represents the learnable matrix, and A represents the adjacency matrix; The graph-based spectral perception module is expressed as the following formula: y = x·σ(F′ r (ReLU(AGCM(F r (GAP(x)), A)))) + x; Among them, F' r and F r represent the MLP layer, GAP represents the global pooling operation, ReLU represents the activation function, x represents the input feature, y represents the output feature, AGCM represents the adaptive graph convolutional module, and σ represents the Sigmoid activation function.
7. A method for fusing multi-spectral images and hyperspectral images according to claim 6, wherein, The unified feature representation obtained by the spatial-spectral fusion module is input into the decoder based on the spectral self-attention mechanism, and the output features of the graph-based spectral perception module are input into the corresponding symmetric layer of the decoder based on the spectral self-attention mechanism. Then, the decoder based on the spectral self-attention mechanism outputs a spectral image with high spatial resolution and high spectral resolution; including: The decoder based on the spectral self-attention mechanism includes a spectral self-attention module and an upsampling module; The upsampling module includes an inverse Harr wavelet transform module and a convolution layer; Among them, the spectral self-attention module extracts features in the spectral dimension, and the upsampling module improves the resolution spatially; the decoder based on the spectral self-attention mechanism receives the output of the spatial-spectral fusion module as input features for upsampling, and uses the upsampling module based on the inverse wavelet transform; among them, the Harr inverse wavelet transform module processes the input features of the decoder, and splices the obtained low-frequency information A1, horizontal high-frequency information H1, vertical high-frequency information V1, and diagonal high-frequency information D1 in the spatial dimension to recover the features After convolution operation, non-linear activation, and batch normalization operation of the convolutional layer, the output features are obtained X dec_out Concatenate with y; then, after two-dimensional convolution, input it into the spectral self-attention module, decode layer by layer, and then perform an addition operation with the low-resolution hyperspectral image to obtain the final fusion result.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the multi-spectral image and hyperspectral image fusion method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-spectral image and hyperspectral image fusion method according to any one of claims 1-7.
10. A multi-spectral image and hyperspectral image fusion system, characterized in that, Including: The multi-scale feature extraction module is configured to: separately encode and extract features from the high-resolution multi-spectral image and the low-resolution hyperspectral image, and perform multi-scale feature extraction to obtain the hyperspectral image features and multi-spectral image features after feature extraction; The depth fusion module is configured to: input the hyperspectral image features and multispectral image features obtained in step 1 into the spatial-spectral fusion module to perform depth fusion in the spatial dimension and spectral dimension, and obtain a unified feature representation; The spatial-spectral fusion module includes a spatial attention mechanism and a channel attention mechanism; The spatial attention mechanism generates two two-dimensional feature maps through average pooling and max pooling operations along the channel axis, and after concatenating the two two-dimensional feature maps, generates a spatial attention map through a convolution operation; The channel attention mechanism compresses the spatial dimension of the feature map into a real number through global average pooling and global max pooling operations, then learns the importance weights of each channel through a shared multi-layer perceptron, and finally generates a channel attention map through the Sigmoid function; The spectral image generation module is configured to: learn the correlation between channels of hierarchical features to obtain output features; The hierarchical features represent the hyperspectral image features and multispectral image features after feature extraction by the multi-scale feature extraction module; After upsampling the output of the spatial-spectral fusion module, it is concatenated with the output features; After that, layer-by-layer decoding is performed, and after decoding, an addition operation is performed with the low-resolution hyperspectral image to obtain the final fusion result.
Citation Information
Cited By
Multispectral and hyperspectral imaging proportion balance optimization method based on satellite remote sensing
CN120746843A
Satellite remote sensing-based multispectral and hyperspectral imaging proportion balance optimization method
CN120746843B
Multispectral endoscope image intelligent fusion method
CN120746866A
A Smart Fusion Method for Multispectral Endoscopic Images
CN120746866B
Intelligent spectrum-integrated inspection system based on ultra-wide spectrum and physical evidence image
CN121068529A