Method and device for multi-scale fusion of hyperspectral image and multispectral image
Through the multi-scale fusion method of wavelet transformation and Mamba module, the spatial and spectral resolution of hyperspectral images is improved, the problem of insufficient fusion results in the prior art is solved, and high-quality image fusion is achieved.
Patent Information
- Application Number
- CN202510737033.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In the fusion of hyperspectral and multispectral images, it is difficult to retain spectral and spatial information at the same time, resulting in insufficient resolution of the fusion result.
Using a combination of wavelet transformation and Mamba module, multi-scale features are extracted by upsampling and stitching and fusing of hyperspectral images, and using a deep feature extraction module to capture spectral and spatial information to achieve multi-scale fusion.
The spatial resolution and spectral resolution of hyperspectral images are improved, and higher quality fusion results are obtained, solving the problem of insufficient spectral and spatial information.
Smart Images

Figure CN120259100A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for multi-scale fusion of hyperspectral images and multispectral images, and belongs to the technical field of hyperspectral image processing. Background Art
[0002] Hyperspectral images (HSIs) and multispectral images (MSIs) have extensive applications in fields such as remote sensing, environmental monitoring, agriculture, and geological exploration. Hyperspectral images can provide rich spectral information, but their spatial resolution is relatively low. In contrast, multispectral images have a higher spatial resolution but relatively less spectral information. Therefore, how to effectively fuse hyperspectral images and multispectral images to simultaneously obtain images with high spatial resolution and high spectral resolution has become an important research topic.
[0003] Traditional image fusion methods are mainly divided into two categories: spatial domain-based and transform domain-based. Spatial domain methods achieve fusion by directly operating on pixel values, but it is difficult to fully retain spectral and spatial information; transform domain methods perform fusion by transforming images into other domains. Although they can better retain details, they have high computational complexity when dealing with high-dimensional data and are difficult to handle non-linear relationships. In recent years, image fusion methods based on deep learning have made significant progress, but they rely on a large amount of training data, have a high model complexity, and are difficult to adapt to small-sample scenarios. In addition, existing methods often ignore the multi-scale characteristics when processing hyperspectral and multispectral images, resulting in insufficient retention of spectral and spatial information in the fusion results and poor spatial resolution. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method and device for multi-scale fusion of hyperspectral images and multispectral images. By performing multi-scale feature extraction through wavelet transform and adding a local feature extraction module to the Mamba module, spectral information and spatial information are fully obtained, realizing the multi-scale fusion of hyperspectral images and multispectral images, improving the accuracy of feature extraction, obtaining higher-quality high-resolution hyperspectral images, and solving the problem of low fusion accuracy caused by insufficient retention of spectral information and spatial information in the process of fusing multi-source remote sensing images.
[0005] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:
[0006] In a first aspect, the present invention provides a method for multi-scale fusion of hyperspectral images and multispectral images, including the following steps:
[0007] Obtain a hyperspectral image and a multispectral image;
[0008] Upsample the hyperspectral image to obtain the upsampled hyperspectral image, splice and fuse the upsampled hyperspectral image with the multispectral image to obtain the fused hyperspectral image;
[0009] Extract features from the fused hyperspectral image to obtain shallow spectral-spatial features;
[0010] Perform wavelet transform on the shallow spectral-spatial features to obtain multi-scale subband feature maps;
[0011] Input the multi-scale subband feature maps into the deep feature extraction module to obtain multi-scale deep spectral-spatial features;
[0012] Fuse the multi-scale deep spectral-spatial features to obtain the fused deep spectral-spatial features, and obtain the spectral-spatial features according to the fused deep spectral-spatial features and the shallow spectral-spatial features;
[0013] Obtain the high-resolution hyperspectral image according to the spectral-spatial features and the upsampled hyperspectral image.
[0014] Further, the upsampling of the hyperspectral image to obtain the upsampled hyperspectral image specifically includes:
[0015] Perform bilinear interpolation upsampling on the hyperspectral image to obtain the upsampled hyperspectral image;
[0016] The splicing and fusion of the upsampled hyperspectral image with the multispectral image to obtain the fused hyperspectral image specifically includes:
[0017] Use the splicing and fusion operation to splice and fuse the multispectral image with the upsampled hyperspectral image to obtain the fused hyperspectral image.
[0018] Further, the extraction of features from the fused hyperspectral image to obtain shallow spectral-spatial features, and its expression is:
[0019] ;
[0020] Among them, is the shallow spectral-spatial feature; is the fused hyperspectral image; is the shallow feature extraction module, which includes a convolution and multiple residual blocks. Among them, the convolution is used to perform preliminary feature extraction on the input initial features to obtain the first extracted features; the multiple residual blocks are used to layer by layer extract higher-level features from the first extracted features to obtain the second extracted features, and then perform residual connection on the initial features and the second extracted features to obtain the shallow spectral-spatial features.
[0021] Further, the wavelet transform of the shallow spectral-spatial features to obtain multi-scale subband feature maps specifically includes:
[0022] Select the Haar wavelet as the mother wavelet;
[0023] Perform low-pass and high-pass filtering on the input shallow spectral-spatial features in the horizontal and vertical directions respectively to generate multi-scale subband feature maps, and the multi-scale subband feature maps include subband feature maps, subband feature maps, subband feature maps, and subband feature maps, where the subband feature map represents the low-frequency component; the subband feature map represents the high-frequency component in the horizontal direction; the subband feature map represents the high-frequency component in the vertical direction; the subband feature map represents the high-frequency component in the diagonal direction.
[0024] Further, the deep feature extraction module includes a spectral Mamba block and a spatial Mamba block;
[0025] Inputting the multi-scale subband feature maps into the deep feature extraction module to obtain multi-scale deep spectral-spatial features specifically includes performing the following steps on the subband feature maps of each scale respectively:
[0026] Step A: Flatten the subband feature map into a vector form and vector serialize it to obtain a sequence;
[0027] Step B: Generate a position encoding through an encoder and add the position encoding to the sequence to retain the spatial information of the sequence, and obtain the sequence after obtaining the correct encoded position information;
[0028] Step C: Input the sequence after obtaining the correct encoded position information into the spectral Mamba block, and obtain deep spectral features through a decoder;
[0029] Replace the subband feature map in Step A with the deep spectral features, and then repeat Steps A and B to re-obtain the sequence after obtaining the correct encoded position information, and input the re-obtained sequence after obtaining the correct encoded position information into the spatial Mamba block, and obtain deep spectral-spatial features through a decoder.
[0030] Further, inputting the sequence after obtaining the correct encoded position information into the spectral Mamba block and obtaining deep spectral features through a decoder specifically includes:
[0031] Step a: Perform a linear projection on the sequence after obtaining the correct encoded position information The projected features are divided into two feature parts along the channel dimension And After q is processed by the activation function, p is processed by the activation function after convolution to extract spatial features, realizing the block processing of features and the enhancement of spatial features:
[0032] ;
[0033] ;
[0034] ;
[0035] Among them, is the projection operation, is the activation function, is the sequence after correctly encoding the position information, is the convolution operation, R is the scale, B is the batch, H is the height of the multispectral image, W is the width of the multispectral image, and C is the number of channels of the convolution in the shallow feature extraction module;
[0036] Step b: The feature part p generates the spectral token M or generates the spatial token N , focusing on the spectral features. The spectral token captures the global information through the state space model of Mamba and obtains the local information through the window scanning mechanism:
[0037] ;
[0038] ;
[0039] Among them, is all the information from different directions of each element mapped from the four corners of the feature map to the relative position, is the window scanning mechanism, capturing the local information between the sequence elements, represents the global information aggregated by the four-way scanning strategy, represents the local information aggregated by the window scanning. The and are input into the selective scanning spatial state sequence model in parallel for processing to obtain the intermediate deep spectral feature out_y;
[0040] Step c: Normalize the intermediate deep spectral feature and multiply it element-wise with the part to perform feature weighting and highlight the required important features to obtain the intermediate feature:
[0041] ;
[0042] Among them, is the intermediate feature, is the linear layer operation, is a layer normalization operation;
[0043] Step d: Obtain deep spectral features through the decoder based on the intermediate features ;
[0044] Inputting the sequence after re-obtaining the correct encoding position information into the spatial Mamba block and obtaining deep spectral-spatial features through the decoder specifically includes:
[0045] Based on the deep spectral features , re-obtain the sequence after the correct encoding position information;
[0046] Replace the sequence after the correct encoding position information in step a with the sequence after re-obtaining the correct encoding position information, and then repeat steps a, b, and c to re-obtain the intermediate features;
[0047] Obtain deep spectral-spatial features through the decoder based on the re-obtained intermediate features.
[0048] Furthermore, the fusion of the multi-scale deep spectral-spatial features specifically includes:
[0049] Reconstruct the multi-scale deep spectral-spatial features into the original signal through inverse two-dimensional discrete wavelet transform to obtain the fused deep spectral-spatial features, and the expression is as follows:
[0050] ;
[0051] where is the fused deep spectral-spatial feature; is the inverse two-dimensional discrete wavelet transform; represents the deep spectral-spatial feature obtained from the cA sub-band feature map, represents the deep spectral-spatial feature obtained from the cH sub-band feature map, represents the deep spectral-spatial feature obtained from the cV sub-band feature map, represents the deep spectral-spatial feature obtained from the cD sub-band feature map.
[0052] Furthermore, the obtaining of the spectral-spatial feature according to the fused deep spectral-spatial feature and the shallow spectral-spatial feature specifically includes:
[0053] Adopt skip connection to aggregate the shallow spectral-spatial feature and the fused deep spectral-spatial feature to obtain the spectral-spatial feature.
[0054] Furthermore, the obtaining of the high-resolution hyperspectral image according to the spectral-spatial feature and the upsampled hyperspectral image specifically includes:
[0055] Input the spectral-spatial features into the image reconstruction module to obtain a high-resolution reconstructed image;
[0056] Use residual connection between the high-resolution hyperspectral image and the upsampled hyperspectral image to fuse the low-resolution and high-resolution information, and obtain a high-resolution hyperspectral image.
[0057] In a second aspect, the present invention provides a device for multi-scale fusion of hyperspectral images and multispectral images, and the device includes:
[0058] Preliminary fusion module: Obtain a hyperspectral image and a multispectral image; upsample the hyperspectral image to obtain an upsampled hyperspectral image, splice and fuse the upsampled hyperspectral image and the multispectral image to obtain a fused hyperspectral image;
[0059] Shallow feature extraction module: Extract features from the fused hyperspectral image to obtain shallow spectral-spatial features;
[0060] Deep feature extraction module: Perform wavelet transform on the shallow spectral-spatial features to obtain multi-scale subband feature maps; input the multi-scale subband feature maps into the deep feature extraction module to obtain multi-scale deep spectral-spatial features;
[0061] Fusion and reconstruction module: Fuse the multi-scale deep spectral-spatial features to obtain fused deep spectral-spatial features, obtain spectral-spatial features according to the fused deep spectral-spatial features and the shallow spectral-spatial features; obtain a high-resolution hyperspectral image according to the spectral-spatial features and the upsampled hyperspectral image.
[0062] Compared with the prior art, the beneficial effects achieved by the present invention:
[0063] The present invention proposes a multi-scale fusion method and device based on the Mamba model and wavelet transform. The deep feature extraction module can efficiently process long sequence data. Through multi-scale analysis, wavelet transform can effectively capture the local features of the image, realizing the multi-scale fusion of hyperspectral images and multispectral images. While improving the spatial resolution, the spectral information is fully retained; the method proposed by the present invention realizes multi-scale feature extraction through wavelet transform, and fully obtains spectral information and spatial information according to the deep feature extraction module, realizing the multi-scale fusion of hyperspectral images and multispectral images, improving the accuracy of feature extraction, obtaining a higher-quality high-resolution hyperspectral image, and solving the problem of low fusion accuracy caused by insufficient retention of spectral information and spatial information in the process of fusing multi-source remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a schematic flowchart of a method for multi-scale fusion of hyperspectral images and multispectral images according to an embodiment of the present invention;
[0065] Figure 2 It is a schematic framework diagram of a method for fusing hyperspectral images and multispectral images according to an embodiment of the present invention;
[0066] Figure 3 It is a schematic framework diagram of the Mamba block provided according to an embodiment of the present invention;
[0067] Figure 4 It is the RGB heat map of the hyperspectral image of the CAVE dataset provided according to an embodiment of the present invention (using the 10th band, 20th band, and 30th band as the RGB image data);
[0068] Figure 5 It is the RGB heat map of the multispectral image of the CAVE dataset provided according to an embodiment of the present invention;
[0069] Figure 6 It is the fusion result map of the CAVE dataset provided according to an embodiment of the present invention;
[0070] Figure 7 It is a schematic structural diagram of a device for fusing hyperspectral images and multispectral images according to an embodiment of the present invention. Detailed implementation manners
[0071] The technical solution of the present invention will be described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0072] The term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.
[0073] Embodiment 1:
[0074] As Figure 1 shown, it is a schematic flow diagram of a method for fusing hyperspectral images and multispectral images according to an embodiment of the present invention, including the following steps:
[0075] Obtain hyperspectral images and multispectral images;
[0076] Specifically, obtaining the training set and validation set of the multispectral image specifically includes:
[0077] Obtain the CAVE dataset (Columbia Airborne Imaging and Photographic Experiment dataset), which contains 32 indoor hyperspectral images, each with a size of 512×512 pixels, covering a wavelength range of 400–700 nm in 31 bands;
[0078] Generate high-resolution multispectral images HR-MSI using the spectral response function of the camera;
[0079] Generate low-resolution hyperspectral images LR-HSI through Gaussian filtering and multiple downsamplings;
[0080] Randomly select 20 pairs of images from the dataset as the training set, and the remaining 12 pairs of images as the validation set;
[0081] Among them, in the training stage of this embodiment, randomly extract 64×64 patches from each 512×512 image and input them into the model. That is, in the training stage of this embodiment, take H = 64, W = 64, s = 3, h = 8, w = 8, S = 31, where h is the height of the hyperspectral image, w is the width of the hyperspectral image, S is the channel of the hyperspectral image, H is the height of the multispectral image, W is the width of the multispectral image, and s is the channel of the multispectral image.
[0082] As Figure 2 shown, it is a schematic framework diagram of a method for fusing hyperspectral images and multispectral images at multiple scales provided by an embodiment of the present invention;
[0083] Among them, the low frequency is the low-frequency component, the horizontal high frequency is the high-frequency component in the horizontal direction, the vertical high frequency is the high-frequency component in the vertical direction, and the diagonal high frequency is the high-frequency component in the diagonal direction.
[0084] In one embodiment, upsample the hyperspectral image to obtain the upsampled hyperspectral image, splice and fuse the upsampled hyperspectral image with the multispectral image to obtain the fused hyperspectral image;
[0085] The upsampling of the hyperspectral image to obtain the upsampled hyperspectral image specifically includes:
[0086] Perform bilinear interpolation upsampling on the hyperspectral image to obtain the upsampled hyperspectral image;
[0087] The splicing and fusing of the upsampled hyperspectral image with the multispectral image to obtain the fused hyperspectral image specifically includes:
[0088] Adopt splicing and fusion operations to splice and fuse the multispectral image with the upsampled hyperspectral image to obtain the fused hyperspectral image.
[0089] Specifically, the hyperspectral image is , and the multispectral image is Perform bilinear interpolation upsampling on the hyperspectral image to obtain the upsampled hyperspectral image, and splice and fuse the multispectral image with the upsampled hyperspectral image to obtain the fused hyperspectral image , where R represents the scale.
[0090] Specifically, for the hyperspectral image perform upsampling operation, and for the upsampled hyperspectral image and the multispectral image perform channel splicing and fusion, specifically including:
[0091] , ;
[0092] ;
[0093] Among them, is the upsampling operation, is the splicing and fusion operation, is the fused hyperspectral image.
[0094] In one embodiment, feature extraction is performed on the fused hyperspectral image to obtain shallow spectral-spatial features;
[0095] Specifically, the feature extraction of the fused hyperspectral image to obtain shallow spectral-spatial features , and its expression is:
[0096] ;
[0097] Among them, is the shallow spectral-spatial feature; is the fused hyperspectral image; is the shallow feature extraction module, which includes a convolution and multiple residual blocks. Among them, the convolution is used to perform preliminary feature extraction on the input initial features to obtain the first extracted features; the multiple residual blocks are used to gradually extract higher-level features from the first extracted features to obtain the second extracted features, and then perform residual connection on the initial features and the second extracted features to obtain the shallow spectral-spatial features.
[0098] Specifically, extract the shallow features of the fused hyperspectral image, and extract the initial features through convolution:
[0099] ;
[0100] Then, the residual blocks are used to gradually extract higher-level features and retain more detailed information, specifically including:
[0101] ;
[0102] ;
[0103] ;
[0104] Among them, is a convolution operation, is a convolution operation with a 3×3 convolution kernel, represents the continuous composition of functions, is the i-th residual block, where i takes 1, 2, 3; is the shallow spectral-spatial feature; is the number of residual blocks, The C in represents the number of channels of convolution in the shallow feature extraction module; in the present invention, C = 180 is taken, and = 3 is taken.
[0105] In one embodiment, the shallow spectral-spatial feature is subjected to wavelet transform to obtain a multi-scale sub-band feature map;
[0106] The step of subjecting the shallow spectral-spatial feature to wavelet transform to obtain a multi-scale sub-band feature map specifically includes:
[0107] Select the Haar wavelet as the mother wavelet;
[0108] Perform low-pass and high-pass filtering on the input shallow spectral-spatial feature in the horizontal and vertical directions respectively to generate a multi-scale sub-band feature map, and the multi-scale sub-band feature map includes sub-band feature map, sub-band feature map, sub-band feature map, and sub-band feature map, where sub-band feature map represents the low-frequency component; sub-band feature map represents the high-frequency component in the horizontal direction; sub-band feature map represents the high-frequency component in the vertical direction; sub-band feature map represents the high-frequency component in the diagonal direction.
[0109] Specifically, subjecting the extracted shallow spectral-spatial feature to wavelet transform to obtain sub-band feature maps of four scales, and obtaining more local detail information, specifically including:
[0110] ;
[0111] ;
[0112] Among them, It is a two-dimensional discrete wavelet transform. cA is the low-frequency component, cH is the high-frequency component in the horizontal direction, cV is the high-frequency component in the vertical direction, and cD is the high-frequency component in the diagonal direction. The mother wavelet function of the two-dimensional discrete wavelet transform is the Haar wavelet.
[0113] In one embodiment, the multi-scale sub-band feature maps are input into the deep feature extraction module to obtain multi-scale deep spectral-spatial features.
[0114] The deep feature extraction module includes a spectral Mamba block and a spatial Mamba block.
[0115] The input of the multi-scale sub-band feature maps into the deep feature extraction module to obtain multi-scale deep spectral-spatial features specifically includes performing the following steps on the sub-band feature maps of each scale:
[0116] Step A: Flatten the sub-band feature maps into a vector form and vector serialize them to obtain a sequence.
[0117] Step B: Generate positional encoding through an encoder and add the positional encoding to the sequence to retain the spatial information of the sequence, obtaining the sequence after obtaining the correct encoded position information.
[0118] Step C: Input the sequence after obtaining the correct encoded position information into the spectral Mamba (Mamba network) block, and obtain deep spectral features through a decoder. Among them, the spectral Mamba module consists of multiple Mamba blocks. The core feature of the Mamba block is to use a state space model to efficiently model long sequence data, capture global dependencies, enhance the global context modeling ability of sequence data through forward and backward scanning mechanisms, introduce window scanning to enhance the modeling ability of local features, enabling the spectral Mamba module to efficiently extract global and local features, and finally obtain a feature map containing spectral feature information through a decoder.
[0119] Replace the sub-band feature maps in Step A with the deep spectral features, and then repeat Step A and Step B to re-obtain the sequence after obtaining the correct encoded position information. Input the re-obtained sequence after obtaining the correct encoded position information into the spatial Mamba block, and obtain deep spectral-spatial features through a decoder. The deep spectral features are unfolded in the same way to obtain a spatial sequence and input into the spatial Mamba module to obtain spectral-spatial features. The Mamba module has the same structure as the spectral Mamba module, except that the spectral Mamba module focuses on spectral information while the spatial Mamba module focuses on spatial information.
[0120] In one embodiment, the input of the sequence after obtaining the correct encoded position information into the spectral Mamba block and obtaining deep spectral features through a decoder specifically includes:
[0121] Step a: Perform a linear projection on the sequence after correct encoding of the position information such that the projected features are divided into two feature parts along the channel dimension and . q is processed by an activation function, and p is subjected to convolution to extract spatial features and then processed by an activation function to achieve block processing of features and enhancement of spatial features:
[0122] ;
[0123] ;
[0124] ;
[0125] wherein, is the projection operation, is the activation function, is the sequence after correct encoding of the position information, is the convolution operation, R is the scale, B is the batch, H is the height of the multispectral image, W is the width of the multispectral image, and C is the number of channels of the convolution in the shallow feature extraction module;
[0126] Step b: The feature part p generates a spectral token M or generates a spatial token N , wherein, , represents the -th pixel point in the spatial direction; N = , represents the C-th spectrum; Focusing on the spectral features, the spectral token captures global information through the state space model of Mamba and obtains local information through the window scanning mechanism:
[0127] ;
[0128] ;
[0129] wherein, is all the information from the four corners of the feature map to each element of the relative position map in different directions, is the window scanning mechanism, capturing the local information between sequence elements, represents the global information aggregated by the four-way scanning strategy, represents the local information aggregated by the window scanning, and the and are input into the selective scanning spatial state sequence model in parallel for processing to obtain the intermediate deep spectral feature out_y;
[0130] Step c: Normalize the intermediate deep spectral feature and combine it with Element-wise multiplication is used to weight the features, highlighting the important features needed to obtain intermediate features:
[0131] ;
[0132] Among them, is the intermediate feature, is the linear layer operation, is the layer normalization operation;
[0133] Step d: According to the intermediate features, deep spectral features are obtained through the decoder ;
[0134] The sequence after re-obtaining the correct coding position information is input into the spatial Mamba block, and deep spectral-spatial features are obtained through the decoder, specifically including:
[0135] According to the deep spectral features , the sequence after re-obtaining the correct coding position information;
[0136] Replace the sequence after the correct coding position information in step a with the sequence after re-obtaining the correct coding position information, and then repeat steps a, b, and c to re-obtain the intermediate features;
[0137] According to the re-obtained intermediate features, deep spectral-spatial features are obtained through the decoder.
[0138] As Figure 3 shown, it is a schematic framework diagram of the Mamba block provided by the embodiment of the present invention; among them, the S6 module is a selective scanning spatial state sequence model.
[0139] Specifically, the four sub-band feature maps are sequentially input into the spectral Mamba block and the spatial Mamba block, the feature maps are unfolded to generate a 2D sequence, positional encoding is added to the 2D sequence to retain the spatial information of the sequence, and finally the encoded sequence is input into the Mamba network of the two Mamba blocks. Under the action of the self-cross-scanning mechanism of the Mamba network and the introduced window scanning mechanism, the local information sequence and the global information sequence obtained are selectively scanned by the spatial state sequence model to enhance the feature expression ability of the global features and local features. The specific operations are as in steps a, b, and c;
[0140] The above operations can obtain deep spectral features when input into the spectral Mamba block. Repeating the above operations to unfold the deep spectral features into a 2D sequence and then inputting them into the spatial Mamba block can obtain deep spectral-spatial features.
[0141] An embodiment is to fuse multi-scale deep spectral-spatial features to obtain the fused deep spectral-spatial features, and based on the fused deep spectral-spatial features and the shallow spectral-spatial features, obtain the spectral-spatial features.
[0142] The fusing of the multi-scale deep spectral-spatial features to obtain the fused deep spectral-spatial features specifically includes:
[0143] Reconstruct the multi-scale deep spectral-spatial features into the original signal through the inverse two-dimensional discrete wavelet transform to obtain the fused deep spectral-spatial features, and the expression is as follows:
[0144] ;
[0145] Where, is the fused deep spectral-spatial feature; is the inverse two-dimensional discrete wavelet transform; represents the deep spectral-spatial feature obtained from the cA sub-band feature map, represents the deep spectral-spatial feature obtained from the cH sub-band feature map, represents the deep spectral-spatial feature obtained from the cV sub-band feature map, represents the deep spectral-spatial feature obtained from the cD sub-band feature map; specifically, is the deep spectral-spatial feature obtained from the low-frequency component, is the deep spectral-spatial feature obtained from the high-frequency horizontal component, is the deep spectral-spatial feature obtained from the high-frequency vertical component, is the deep spectral-spatial feature obtained from the high-frequency diagonal component.
[0146] An embodiment, the obtaining of the spectral-spatial features based on the fused deep spectral-spatial features and the shallow spectral-spatial features specifically includes:
[0147] Adopt skip connection to aggregate the shallow spectral-spatial features and the fused deep spectral-spatial features to obtain the spectral-spatial features.
[0148] An embodiment is to obtain a high-resolution hyperspectral image based on the spectral-spatial features and the upsampled hyperspectral image.
[0149] The obtaining of the high-resolution hyperspectral image based on the spectral-spatial features and the upsampled hyperspectral image specifically includes:
[0150] Input the spectral-spatial features into an image reconstruction module to obtain a high-resolution reconstructed image;
[0151] Use residual connection between the high-resolution hyperspectral image and the upsampled hyperspectral image to fuse the low-resolution and high-resolution information to obtain the high-resolution hyperspectral image.
[0152] Inputting the spectral-spatial features into an image reconstruction module to obtain a high-resolution reconstructed image specifically includes:
[0153] Inputting the spectral-spatial features into an image reconstruction module to obtain a high-resolution reconstructed image, and the specific expression is as follows:
[0154] ;
[0155] ;
[0156] Wherein, is a convolution operation with a 3×3 convolution kernel, is an activation operation; is an addition operation; is the spectral-spatial feature; is the high-resolution reconstructed image;
[0157] Using a residual connection between the high-resolution reconstructed image and the upsampled hyperspectral image to fuse low-resolution and high-resolution information to obtain a high-resolution hyperspectral image, specifically includes:
[0158] ;
[0159] Where Y represents the final high-resolution hyperspectral image, is the input multispectral image, is the input hyperspectral image, is an upsampling operation, is a splicing operation, is a feature extraction operation, is an image reconstruction operation, is an operation of restricting the value within [0, 1].
[0160] As Figures 4-5 shown, it is the RGB heat map of the hyperspectral image of the CAVE dataset and the RGB heat map of the multispectral image of the CAVE dataset provided by the embodiment of the present invention; as Figure 6 shown, it is the fusion result map of the CAVE dataset provided by the embodiment of the present invention; in this embodiment, using the CAVE dataset to generate a low-resolution hyperspectral image as Figure 4 shown (using the 10th, 20th, and 30th bands as the RGB primary colors respectively), and a high-resolution multispectral image as Figure 5As shown; the CAVE dataset contains 32 indoor hyperspectral images, each with a size of 512×512 pixels, covering a wavelength range of 400–700 nm in 31 frequency bands; the existing TFNet (Remote sensing image fusion based on two-stream fusion network), CSSNet (Hyperspectral image super-resolution network based on cross-scale nonlocal attention), Fusformer (A transformer-based fusion network for hyperspectral image super-resolution), PSRT (Pyramid shuffle-and-reshuffle transformer for multispectral and hyper-spectral image fusion) and the fusion method of this application are respectively used to fuse hyperspectral images and multispectral images of the CAVE dataset example, as Figure 6 shown, including high-resolution hyperspectral images, true label images, absolute error difference maps and structural similarity index maps, and the fusion results are shown in Table 1:
[0161] Table 1. Comparison table of hyperspectral image and multispectral image fusion results;
[0162]
[0163] As can be seen from Table 1, SSIM is the structural similarity index and SAM is the spectral angle mapper; the PSNR (peak signal-to-noise ratio) of the hyperspectral image and multispectral image fusion result of this application reaches 47.52. Compared with other methods, it is 0.96 higher than the second-ranked Fusformer method.
[0164] The results of the method of this invention application are compared with the real pictures. The first row shows the fusion results and the real pictures, and the second row intuitively shows the results from two dimensions. The figure proves the feasibility of the method of this invention application and the fusion effect is good.
[0165] The above examples confirm that this application reduces the error value of hyperspectral image and multispectral image fusion.
[0166] Compared with the prior art, the method of the present invention for multi-scale fusion of hyperspectral images and multispectral images based on Mamba uses wavelet transform. First, the hyperspectral data is upsampled, then the upsampled hyperspectral image and the multispectral image are concatenated and fused in channels, and the shallow features of the fused image are extracted. Next, the shallow features are decomposed by wavelet transform to obtain feature maps at multiple scales. The small feature maps are serially input into the Mamba module to obtain deep features. The Mamba module uses a state space model to efficiently model long sequence data, capture global dependencies, and enhance the global context modeling ability for sequence data through a cross-scanning mechanism. A window scanning mode is introduced to enhance the modeling ability of local features and efficiently extract global and local features. Finally, the extracted deep features and shallow features are fused to obtain the final spectral-spatial features, and then the features are used for image reconstruction to obtain a high-resolution hyperspectral image.
[0167] The present invention proposes a multi-scale fusion method based on the Mamba model and wavelet transform. The deep feature extraction module (Mamba) can efficiently process long sequence data, and wavelet transform can effectively capture the local features of images through multi-scale analysis, realizing the multi-scale fusion of hyperspectral images and multispectral images. While improving the spatial resolution, the spectral information is fully retained. The method proposed by the present invention realizes multi-scale feature extraction through wavelet transform, and a local feature extraction module is added to the Mamba module to fully obtain spectral information and spatial information, realizing the multi-scale fusion of hyperspectral images and multispectral images, improving the accuracy of feature extraction, obtaining a higher-quality high-resolution hyperspectral image, and solving the problem of low fusion accuracy caused by insufficient retention of spectral information and spatial information in the process of fusing multi-source remote sensing images.
[0168] Embodiment 2:
[0169] As Figure 7 shown, it is a schematic structural diagram of a device for multi-scale fusion of hyperspectral images and multispectral images provided by an embodiment of the present invention. The device includes:
[0170] Preliminary fusion module: Obtain a hyperspectral image and a multispectral image; Upsample the hyperspectral image to obtain the upsampled hyperspectral image, and splice and fuse the upsampled hyperspectral image and the multispectral image to obtain the fused hyperspectral image;
[0171] Shallow feature extraction module: Extract features from the fused hyperspectral image to obtain shallow spectral-spatial features;
[0172] Deep feature extraction module: Perform wavelet transform on the shallow spectral-spatial features to obtain multi-scale subband feature maps; Input the multi-scale subband feature maps into the deep feature extraction module to obtain multi-scale deep spectral-spatial features;
[0173] Fusion and reconstruction module: fuse multi-scale deep spectral-spatial features to obtain fused deep spectral-spatial features, and obtain spectral-spatial features based on the fused deep spectral-spatial features and shallow spectral-spatial features; obtain a high-resolution hyperspectral image based on the spectral-spatial features and the upsampled hyperspectral image.
[0174] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0175] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0176] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0177] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0178] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for fusing hyperspectral images and multispectral images at multiple scales, characterized in that, It includes the following steps: Obtain hyperspectral images and multispectral images; Upsample the hyperspectral images to obtain the upsampled hyperspectral images, splice and fuse the upsampled hyperspectral images with the multispectral images to obtain the fused hyperspectral images; Extract features from the fused hyperspectral images to obtain shallow spectral-spatial features; Perform wavelet transform on the shallow spectral-spatial features to obtain multi-scale subband feature maps; Input the multi-scale subband feature maps into a deep feature extraction module to obtain multi-scale deep spectral-spatial features; Fuse the multi-scale deep spectral-spatial features to obtain the fused deep spectral-spatial features, and obtain spectral-spatial features based on the fused deep spectral-spatial features and the shallow spectral-spatial features; Obtain high-resolution hyperspectral images based on the spectral-spatial features and the upsampled hyperspectral images.
2. The method for fusing hyperspectral images and multispectral images at multiple scales according to claim 1, characterized in that, The upsampling of the hyperspectral images to obtain the upsampled hyperspectral images specifically includes: Perform bilinear interpolation upsampling on the hyperspectral images to obtain the upsampled hyperspectral images; The splicing and fusing of the upsampled hyperspectral images with the multispectral images to obtain the fused hyperspectral images specifically includes: Use splicing and fusion operations to splice and fuse the multispectral images with the upsampled hyperspectral images to obtain the fused hyperspectral images.
3. The method for multi-scale fusion of hyperspectral images and multispectral images according to claim 1, wherein, The feature extraction of the fused hyperspectral images to obtain shallow spectral-spatial features, and its expression is: ; Among them, is the shallow spectral spatial feature; is the fused hyperspectral image; is the shallow feature extraction module, which includes convolution and multiple residual blocks. Among them, convolution is used to perform preliminary feature extraction on the input initial features to obtain the first extracted features; multiple residual blocks are used to layer by layer extract higher-level features from the first extracted features to obtain the second extracted features, and then perform residual connection on the initial features and the second extracted features to obtain the shallow spectral spatial feature.
4. The method for fusing hyperspectral images and multispectral images at multiple scales according to claim 1, characterized in that, The wavelet transform of the shallow spectral-spatial features to obtain multi-scale subband feature maps specifically includes: Select the Haar wavelet as the mother wavelet; The shallow spectral spatial features of the input are respectively subjected to low-pass and high-pass filtering in the horizontal and vertical directions to generate multi-scale sub-band feature maps, and the multi-scale sub-band feature maps include sub-band feature maps, sub-band feature maps, sub-band feature maps, and sub-band feature maps, where the sub-band feature map represents the low-frequency component; the sub-band feature map represents the high-frequency component in the horizontal direction; the sub-band feature map represents the high-frequency component in the vertical direction; the sub-band feature map represents the high-frequency component in the diagonal direction.
5. The method for fusing hyperspectral images and multispectral images at multiple scales according to claim 3, wherein, The deep feature extraction module includes a spectral Mamba block and a spatial Mamba block; Inputting the multi-scale subband feature maps into the deep feature extraction module to obtain multi-scale deep spectral-spatial features specifically includes performing the following steps on the subband feature maps of each scale respectively: Step A: Flatten the subband feature maps into a vector form and vector serialize them to obtain a sequence; Step B: Generate position encoding through an encoder and add the position encoding to the sequence to retain the spatial information of the sequence, and obtain the sequence after obtaining the correct encoding position information; Step C: Input the sequence after obtaining the correct encoding position information into the spectral Mamba block, and obtain deep spectral features through a decoder; Replace the subband feature maps in Step A with the deep spectral features, and then repeat Steps A and B to re-obtain the sequence after obtaining the correct encoding position information. Input the re-obtained sequence after obtaining the correct encoding position information into the spatial Mamba block, and obtain deep spectral-spatial features through a decoder.
6. The method for fusing hyperspectral images and multispectral images at multiple scales according to claim 5, wherein Inputting the sequence after obtaining the correct encoding position information into the spectral Mamba block and obtaining deep spectral features through a decoder specifically includes: Step a: For the sequence after encoding the correct position information perform a linear projection, and the projected features are divided into two feature parts along the channel dimension and , After being processed by the activation function, p is convolved to extract spatial features and then processed by the activation function to achieve block processing of features and enhancement of spatial features: ; ; ; Among them, is the projection operation, is the activation function, is the sequence after correctly encoding the position information, is the convolution operation, R is the scale, B is the batch, H is the height of the multi-spectral image, W is the width of the multi-spectral image, and C is the number of channels of the convolution in the shallow feature extraction module; Step b: The feature part p generates a spectral token M or generates a spatial token N , focusing on spectral features. The spectral token captures global information through Mamba's state space model and obtains local information through a window scanning mechanism: ; ; Among them, All information from different directions for mapping each element from the four corners of the feature map to the relative position, Is the window scanning mechanism to capture local information between sequence elements, Represents the global information aggregated by the four-way scanning strategy, Represents the local information aggregated by window scanning, the And Are input into the selective scanning spatial state sequence model for parallel processing to obtain the intermediate deep spectral feature out_y; Step c: Normalize the intermediate deep spectral features and perform element-wise multiplication with to perform feature weighting for highlighting the required important features, thereby obtaining intermediate features: ; Among them, is an intermediate feature, is a linear layer operation, is a layer normalization operation; Step d: Obtain deep spectral features through a decoder based on the intermediate features ; Inputting the re-obtained sequence after obtaining the correct encoding position information into the spatial Mamba block and obtaining deep spectral-spatial features through a decoder specifically includes: According to the deep spectral features , the sequence after re-obtaining the correct coding position information; Replace the sequence after obtaining the correct encoding position information in Step a with the re-obtained sequence after obtaining the correct encoding position information, and then repeat Steps a, b, and c to re-obtain intermediate features; Obtain deep spectral-spatial features through a decoder according to the re-obtained intermediate features.
7. The method for fusing hyperspectral images and multispectral images with multi-scale as claimed in claim 4, wherein, The fusion of multi-scale deep spectral-spatial features to obtain the fused deep spectral-spatial features specifically includes: Reconstructing the multi-scale deep spectral-spatial features into the original signal through inverse two-dimensional discrete wavelet transform to obtain the fused deep spectral-spatial features, and the expression is as follows: ; Among them, is the fused deep spectral-spatial feature; is the inverse two-dimensional discrete wavelet transform; represents the deep spectral-spatial feature obtained from the cA sub-band feature map, represents the deep spectral-spatial feature obtained from the cH sub-band feature map, represents the deep spectral-spatial feature obtained from the cV sub-band feature map, represents the deep spectral-spatial feature obtained from the cD sub-band feature map.
8. The method for multi-scale fusion of hyperspectral images and multispectral images according to claim 7, wherein The obtaining of spectral-spatial features based on the fused deep spectral-spatial features and shallow spectral-spatial features specifically includes: Using skip connections to aggregate the shallow spectral-spatial features and the fused deep spectral-spatial features to obtain spectral-spatial features.
9. The method for fusing hyperspectral images and multispectral images with multi-scale as claimed in claim 2, wherein The obtaining of the high-resolution hyperspectral image based on the spectral-spatial features and the upsampled hyperspectral image specifically includes: Inputting the spectral-spatial features into an image reconstruction module to obtain a high-resolution reconstructed image; Using residual connections between the high-resolution hyperspectral image and the upsampled hyperspectral image to fuse the low-resolution and high-resolution information to obtain a high-resolution hyperspectral image.
10. An apparatus for fusing hyperspectral images and multispectral images at multiple scales, characterized in that, The device includes: Preliminary fusion module: Obtaining a hyperspectral image and a multi-spectral image; Upsampling the hyperspectral image to obtain the upsampled hyperspectral image, and splicing and fusing the upsampled hyperspectral image and the multi-spectral image to obtain the fused hyperspectral image; Shallow feature extraction module: Extracting features from the fused hyperspectral image to obtain shallow spectral-spatial features; Deep feature extraction module: Performing wavelet transform on the shallow spectral-spatial features to obtain multi-scale subband feature maps; Inputting the multi-scale subband feature maps into the deep feature extraction module to obtain multi-scale deep spectral-spatial features; Fusion and reconstruction module: Fusing the multi-scale deep spectral-spatial features to obtain the fused deep spectral-spatial features, obtaining spectral-spatial features based on the fused deep spectral-spatial features and shallow spectral-spatial features; Obtaining a high-resolution hyperspectral image based on the spectral-spatial features and the upsampled hyperspectral image.
Citation Information
Patent Citations
Hyperspectral and multispectral remote sensing image fusion method based on multi-level collaborative mapping
CN118898545A
Method for fusing infrared light and visible light images
WO2025103079A1
Cited By
Hyperspectral and multispectral image fusion method based on heterogeneous double-branch collaboration
CN120526278A
Hyperspectral and multispectral image fusion method based on heterogeneous dual-branch collaboration
CN120526278B
Image fusion method based on spectral space fusion state attention Mama model
CN120563988A
Method and device for image reconstruction of object
CN121563792A