Hyperspectral and Multispectral Image Fusion Method Based on Adaptive Multi-Scale Features

Through the image fusion method of adaptive multi-scale features, the spatial detail loss and spectral distortion problems in the fusion of hyperspectral and multispectral images are solved, and high-quality image fusion effect is achieved.

CN120047326BActive Publication Date: 2025-07-22NORTHEASTERN UNIV CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510517884.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-22
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing hyperspectral and multispectral image fusion technology faces key challenges such as loss of spatial details, insufficient cross-domain generalization capabilities and low spectral resolution resulting in spectral distortion.

Method used

The image fusion method based on adaptive multi-scale features is adopted, and the spatial resolution and spectral resolution of the image are improved through the channel interactive upsampling module, the multi-scale adaptive spatial reconstruction module and the spectral self-attention enhancement module to achieve full interaction and reconstruction of features.

Benefits of technology

It effectively avoids spatial detail loss and spectral distortion, improves the quality of image fusion and cross-domain generalization capabilities, and improves the overall performance of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047326B_ABST
    Figure CN120047326B_ABST
Patent Text Reader

Abstract

The present invention discloses a hyperspectral and multispectral image fusion method based on adaptive multi-scale features. First, a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution are used as the input of the network, and efficient feature expansion is achieved through upsampling by a channel interaction upsampling module. Then, they are fused into a super-multispectral image with the same size as the reference image. Next, diverse features are captured and spatial information is reconstructed via a multi-scale adaptive spatial reconstruction module. The optimized feature map enters a spectral self-attention enhancement module for spectral attention enhancement, calculates the correlation between feature maps to generate an attention map, and reconstructs spectral information. This avoids problems such as loss of spatial details, insufficient cross-domain generalization ability, and spectral distortion caused by low spectral resolution during the image fusion process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and particularly to a hyperspectral and multispectral image fusion method based on adaptive multi-scale features. Background Art

[0002] With the intensification of global climate change and the increasing urgency of sustainable development issues, high-precision environmental monitoring and ecological resource management have become the core needs of the international community to address challenges. Hyperspectral and multispectral image fusion technology can provide key data support for climate change research and ecosystem protection by extracting refined parameters such as surface cover, vegetation physiological status, and pollutant distribution. Therefore, it has become one of the important research directions in the field of remote sensing. The technology aims to integrate the fine spectral features of hyperspectral images (HSI) and the high spatial details of multispectral images (MSI), break through the physical performance limitations of a single imaging mode, make up for the deficiencies of a single sensor, and promote the leap of remote sensing from "data acquisition" to "information extraction". Compared with other remote sensing images, hyperspectral image HSI has a higher spectral resolution and can capture the spectral information of objects through hundreds of continuous narrow bands, providing unique advantages for ground object recognition. However, its spatial resolution is usually low due to the influence of sensor signal-to-noise ratio and energy dispersion effect. On the contrary, multispectral image MSI has a high spatial resolution, but the sparsity of spectral information limits its fine classification ability. However, the hyperspectral and multispectral image fusion task faces many challenges, and existing methods still face key challenges such as loss of spatial details, insufficient cross-domain generalization ability, and spectral distortion caused by low spectral resolution. Summary of the Invention

[0003] Based on this, it is necessary to address the above problems and propose a hyperspectral and multispectral image fusion method based on adaptive multi-scale features.

[0004] A hyperspectral and multispectral image fusion method based on adaptive multi-scale features, the method comprising:

[0005] Obtain a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution;

[0006] Convert the hyperspectral image with low spatial resolution and high spectral resolution into an upsampled feature through a channel interaction upsampling module, and the spatial resolution of the upsampled feature is the same as the spatial resolution of the multispectral image with high spatial resolution and low spectral resolution;

[0007] Determine an interpolation fusion feature according to the upsampled feature and the multispectral image with high spatial resolution and low spectral resolution through a multi-scale adaptive spatial reconstruction module;

[0008] Determine the spatial fusion feature according to the interpolation fusion feature, including: performing a convolution operation on the interpolation fusion feature to obtain a fusion weight and a comprehensive weight; determining an interaction feature according to the fusion weight and the interpolation fusion feature; determining a comprehensive feature according to the comprehensive weight and the interpolation fusion feature; determining the spatial fusion feature according to the interaction feature and the comprehensive feature;

[0009] Determine the residual fusion feature according to the spatial fusion feature, including: performing a convolution process on the spatial fusion feature to obtain a convolutional spatial fusion feature; performing a residual connection on the convolutional spatial fusion feature and the spatial fusion feature to obtain the residual fusion feature;

[0010] The spectral self-attention enhancement module determines a fused image according to the residual fusion feature and the spatial fusion feature.

[0011] In one embodiment, the determining the spatial fusion feature according to the interpolation fusion feature includes:

[0012] Performing a convolution operation on the interpolation fusion feature to obtain a fusion weight and a comprehensive weight;

[0013] Determining an interaction feature according to the fusion weight and the interpolation fusion feature;

[0014] Determining a comprehensive feature according to the comprehensive weight and the interpolation fusion feature;

[0015] Determining the spatial fusion feature according to the interaction feature and the comprehensive feature.

[0016] In one embodiment, determining the fused image according to the residual fusion feature and the spatial fusion feature includes:

[0017] Performing a convolution process on the residual fusion feature to obtain a query feature, a key feature, and a value feature;

[0018] Determining a spectral energy matrix according to the query feature and the key feature;

[0019] Determining an attention map according to the spectral energy matrix;

[0020] Determining the fused image according to the attention map, the value feature, and the residual fusion feature.

[0021] In one embodiment, the converting the hyperspectral image with low spatial resolution and high spectral resolution into an upsampled feature is implemented by the following expression:

[0022] CIUM(lr) = Conv 1×1 (CS(LeakyReLU(BN(DWC 3×3 (Up 4×(lr))))))

[0023] Among them, CIUM(lr) represents the upsampled feature, lr represents the hyperspectral image with low spatial resolution and high spectral resolution, Up 4× represents the four-fold upsampling operation, DWC 3×3 represents the depth convolution with a size of 3×3, BN represents the batch normalization layer, LeakyReLU represents the activation function, CS represents the Channel Shuffle function, Conv 1×1 represents the convolution with a size of 1.

[0024] In one embodiment, the convolution operation on the interpolated fusion feature to obtain the fusion weight and the comprehensive weight is implemented through the following expression:

[0025] W S =Conv S (x)

[0026] W C =Conv C (x)

[0027] W A =Conv A (x)

[0028] W H =Conv H (x)

[0029] W V =Conv V (x)

[0030] W concat =[W S ; W C ; W A ; W H ; W V

[0031] W1=Conv 1×1 (W concat )

[0032] W2=W S +W C +W A +W H +W V

[0033] Among them, x represents the interpolated fusion feature, W S represents the standard weight matrix output by the standard convolution, W C represents the central difference weight matrix output by the CDC convolution, W A ​Denote the angular difference weight matrix of the ADC convolution output as W H Denote the horizontal difference weight matrix of the HDC convolution output as W V Denote the vertical difference weight matrix of the VDC convolution output as W concat Denote the concatenated weight matrix after concatenating the channel dimensions; W1 represents the fusion weight and W2 represents the comprehensive weight.

[0034] In one embodiment, determining the interaction feature according to the fusion weight and the interpolation fusion feature is implemented through the following expression:

[0035] Feature1 = W1⊙ x

[0036] where Feature1 represents the interaction feature, x represents the interpolation fusion feature, W1 represents the fusion weight, and ⊙ represents element-wise multiplication;

[0037] Determining the comprehensive feature according to the comprehensive weight and the interpolation fusion feature is implemented through the following expression:

[0038] Feature2 = Conv (W2⊙x)

[0039] where Feature2 represents the comprehensive feature, W2 represents the comprehensive weight, and ⊙ represents element-wise multiplication.

[0040] In one embodiment, determining the spatial fusion feature according to the interaction feature and the comprehensive feature is implemented through the following expression:

[0041] α = σ(Feature1+ Feature2)

[0042] Feature merged = α·Feature1+ (1-α)·Feature2

[0043] Feature = Feature merged + x

[0044] where σ represents the sigmoid function, α∈[0,1] represents the adaptive feature perception coefficient, Feature1 represents the interaction feature, Feature2 represents the comprehensive feature, Feature merged represents the fusion feature, Feature represents the spatial fusion feature, and x represents the interpolation fusion feature.

[0045] In one embodiment, determining the fused image according to the residual fusion feature and the spatial fusion feature is implemented through the following expression:

[0046] A= Softmax (E)

[0047] E = Q T K

[0048] Y = β·A·V + X

[0049] Y = β·Softmax(Q T K)·V + X

[0050] Among them, Y represents the fused image, X represents the residual fusion feature, Q ∈ represents the query feature, Q T represents the transpose of the query feature, K ∈ represents the key feature, A ∈ represents the attention weight map, V ∈ represents the value feature, E ∈ represents the spectral energy matrix, Softmax represents the normalization operation, and β represents the learnable parameter.

[0051] In this application, a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution are obtained. The hyperspectral image with low spatial resolution and high spectral resolution is the hyperspectral image with low spatial resolution and high spectral resolution, and the multispectral image with high spatial resolution and low spectral resolution is the multispectral image with high spatial resolution and low spectral resolution; the hyperspectral image with low spatial resolution and high spectral resolution is converted into an upsampled feature, and the spatial resolution of the upsampled feature is the same as that of the multispectral image with high spatial resolution and low spectral resolution; an interpolation fusion feature is determined according to the upsampled feature and the multispectral image with high spatial resolution and low spectral resolution; a spatial fusion feature is determined according to the interpolation fusion feature; a residual fusion feature is determined according to the spatial fusion feature; a fused image is determined according to the residual fusion feature and the spatial fusion feature. This avoids problems such as loss of spatial details, insufficient cross-domain generalization ability, and spectral distortion caused by low spectral resolution during the image fusion process. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0053] Among them:

[0054] Figure 1 is an application environment diagram of the hyperspectral and multispectral image fusion method based on adaptive multi-scale features in one embodiment;

[0055] Figure 2 It is a flowchart of a hyperspectral and multispectral image fusion method based on adaptive multi-scale features in an embodiment;

[0056] Figure 3 It is a flowchart of upsampling feature acquisition;

[0057] Figure 4 It is a flowchart of the implementation of the Channel Shuffle function;

[0058] Figure 5 It is a network structure diagram of the MASR Module;

[0059] Figure 6 It is a structure diagram of the spectral self-attention enhancement module;

[0060] Figure 7 It is a comparison chart of fusion effects;

[0061] Figure 8 It is a structural block diagram of a hyperspectral and multispectral image fusion system based on adaptive multi-scale features in an embodiment;

[0062] Figure 9 It is a structural block diagram of a computer device in an embodiment. Specific implementation manners

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.

[0064] With the intensification of global climate change and the increasing urgency of sustainable development issues, high-precision environmental monitoring and ecological resource management have become the core needs of the international community to address challenges. Hyperspectral and multispectral image fusion technology can provide key data support for climate change research and ecosystem protection by extracting refined parameters such as land cover, vegetation physiological status, and pollutant distribution. Therefore, it has become one of the important research directions in the field of remote sensing. This technology aims to break through the physical performance limitations of a single imaging mode, make up for the deficiencies of a single sensor, and promote the leap of remote sensing from "data acquisition" to "information extraction" by integrating the fine spectral features of hyperspectral images (HSI) and the high spatial details of multispectral images (MSI). Compared with other remote sensing images, hyperspectral image HSI has a higher spectral resolution and can capture the spectral information of objects through hundreds of continuous narrow bands, providing unique advantages for ground object recognition. However, its spatial resolution is usually low due to the influence of sensor signal-to-noise ratio and energy dispersion effect. On the contrary, multispectral image MSI has a high spatial resolution, but the sparsity of spectral information limits its fine classification ability. However, the hyperspectral and multispectral image fusion task faces many challenges, and existing methods still face key challenges such as loss of spatial details, insufficient cross-domain generalization ability, and spectral distortion caused by low spectral resolution. To solve the above technical problems, this application provides a spectral and multispectral image fusion method, Figure 1 is an application environment diagram for hyperspectral and multispectral image fusion in an embodiment. Refer to Figure 1 , this hyperspectral and multispectral image fusion method is applied to a hyperspectral and multispectral image fusion system. The hyperspectral and multispectral image fusion system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 can specifically be a desktop terminal or a mobile terminal, and the mobile terminal can specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 is used to obtain a hyperspectral image with a target low spatial resolution and high spectral resolution and a multispectral image with a high spatial resolution and low spectral resolution. The hyperspectral image with a low spatial resolution and high spectral resolution is a hyperspectral image with a low spatial resolution and high spectral resolution, and the multispectral image with a high spatial resolution and low spectral resolution is a multispectral image with a high spatial resolution and low spectral resolution; the server 120 is used to convert the hyperspectral image with a low spatial resolution and high spectral resolution into an upsampled feature, and the spatial resolution of the upsampled feature is the same as the spatial resolution of the multispectral image with a high spatial resolution and low spectral resolution; determine an interpolation fusion feature according to the upsampled feature and the multispectral image with a high spatial resolution and low spectral resolution; determine a spatial fusion feature according to the interpolation fusion feature; determine a residual fusion feature according to the spatial fusion feature; determine a fusion image according to the residual fusion feature and the spatial fusion feature.

[0065] As Figure 2 shown, in one embodiment, a hyperspectral and multispectral image fusion method based on adaptive multi-scale features is provided. This method can be applied to both terminals and servers. In this embodiment, an example of applying it to a terminal is given. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features specifically includes the following steps:

[0066] S10: Obtain a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution.

[0067] S20: Convert the hyperspectral image with low spatial resolution and high spectral resolution into upsampled features through a channel interaction upsampling module, and the spatial resolution of the upsampled features is the same as that of the multispectral image with high spatial resolution and low spectral resolution.

[0068] S30: Determine interpolation fusion features according to the upsampled features and the multispectral image with high spatial resolution and low spectral resolution through a multi-scale adaptive spatial reconstruction module.

[0069] S40: Determine spatial fusion features according to the interpolation fusion features.

[0070] S50: Determine residual fusion features according to the spatial fusion features.

[0071] S60: Determine a fused image according to the residual fusion features and the spatial fusion features through a spectral self-attention enhancement module.

[0072] This application avoids problems such as loss of spatial details, insufficient cross-domain generalization ability, and spectral distortion caused by low spectral resolution during the image fusion process.

[0073] In step S20, the conversion of the hyperspectral image with low spatial resolution and high spectral resolution into upsampled features is realized through the following expression:

[0074] CIUM(lr)=Conv 1×1 (CS(LeakyReLU(BN(DWC 3×3 (Up 4× (lr)))))) (1)

[0075] where CIUM(lr) represents the upsampled features, lr represents the hyperspectral image with low spatial resolution and high spectral resolution, Up 4× represents a four-fold upsampling operation, DWC 3×3Conv represents a depth convolution with a size of 3×3, BN represents a batch normalization layer, LeakyReLU represents an activation function, CS represents a Channel Shuffle function, and Conv 1×1 represents a convolution with a size of 1.

[0076] Specifically, in the hyperspectral and multispectral image fusion task, due to the different spatial resolutions of the input images, it is usually necessary to first upsample the low-resolution hyperspectral image to match its high spatial resolution with that of the multispectral image, and then perform subsequent feature extraction and fusion operations. Traditional upsampling methods mainly expand the feature map size based on pixel value interpolation, paying more attention to the expansion of the spatial dimension, but there are deficiencies in retaining the context information of the feature map and the information interaction between channels. Since upsampling is a crucial step in this task and is usually applied at the beginning of the forward propagation of the network, the spatial information and channel interaction ignored by it often have a more obvious impact on the performance of the model and the image fusion effect. To solve the problems of spatial information loss and edge feature blurring caused by traditional upsampling methods, this paper proposes a channel interaction upsampling module (CIUM), as Figure 3 shown. At the same time, the multi-scale adaptive spatial reconstruction module is represented by MASR, and the spectral self-attention enhancement module is represented by SAE.

[0077] By gradually upsampling the features, the spatial resolution of the upsampled features is increased to the spatial resolution of the multispectral image with high spatial resolution and low spectral resolution; specifically, the low-spatial-resolution and high-spectral-resolution hyperspectral image LR-HSI is used as the input of the module, and 4-fold upsampling is performed to obtain a feature map with the same spatial resolution as the multispectral image with high spatial resolution and low spectral resolution, that is, the multispectral image HR-MIS with high spatial resolution and low spectral resolution. Then, the spatial details are enhanced through depth convolution, BatchNorm, and the LeakyReLU activation function, and then the channels of the low-spatial-resolution and high-spectral-resolution hyperspectral image are rearranged via the Channel Shuffle function to fully fuse different channel information, thereby increasing the information interaction between channels. Finally, pointwise convolution is used to adjust the number of channels to obtain the final output upsampled features.

[0078] Performing depth convolution with a convolution kernel size of 3×3 on each channel can more effectively capture local spatial features within the channel. Taking the edge information in an image as an example, depth convolution can accurately extract detailed features such as the direction and intensity of the edge through convolution calculations on the pixel values at the edge of each channel. Compared with traditional upsampling methods, it has a better effect on extracting and retaining spatial detail information such as edges and textures in the image. However, in the prior art, it is proposed that depth convolution is performed independently within each channel, which easily ignores the relationship between different channels. Therefore, we introduce the ChannelShuffle function operation to enhance the interaction ability of the model between different channels. The implementation process of the Channel Shuffle function is as Figure 4 shown. First, reshape the input hyperspectral image lr with low spatial resolution and high spectral resolution in R C×H×W (where C represents the total number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map), expanding 1 dimension of the input feature channels into 2 dimensions, namely the number of convolution groups G and the number of channels included in each convolution N, so as to obtain the feature matrix Q ∈ R G×N×H×W , where G×N = C. Secondly, transpose the first dimension and the second dimension of the matrix Q to obtain the matrix S ∈ R N×G×H×W . Finally, flatten the 2 dimensions in the matrix S into 1 dimension to obtain the final upsampled feature map CIUM(lr) ∈ R C ×H×W , thus realizing the information flow between channels.

[0079] In step S30, determining the interpolation fusion feature according to the upsampled feature and the multispectral image with high spatial resolution and low spectral resolution includes: constructing the interpolation fusion feature by means of a gap interpolation strategy for the multispectral image with high spatial resolution and low spectral resolution and the upsampled feature.

[0080] In one embodiment, determining the spatial fusion feature according to the interpolation fusion feature includes:

[0081] S401: Performing a convolution operation on the interpolation fusion feature to obtain a fusion weight and a comprehensive weight.

[0082] S402: Determining an interaction feature according to the fusion weight and the interpolation fusion feature.

[0083] S403: Determining a comprehensive feature according to the comprehensive weight and the interpolation fusion feature.

[0084] S404: Determining the spatial fusion feature according to the interaction feature and the comprehensive feature.

[0085] Specifically, in step S401, the convolution operation on the interpolation fusion feature to obtain the fusion weight and the comprehensive weight is realized through the following expression:

[0086] W S =Conv S (x) (2)

[0087] W C =Conv C (x) (3)

[0088] W A =Conv A (x) (4)

[0089] W H =Conv H (x) (5)

[0090] W V =Conv V (x) (6)

[0091] W concat =[W S ; W C ; W A ; W H ; W V (7)

[0092] W1=Conv 1×1 (W concat ) (8)

[0093] W2=W S +W C +W A +W H +W V (9)

[0094] Among them, x represents the interpolation fusion feature, W S represents the standard weight matrix output by the standard convolution, W C represents the central difference weight matrix output by the CDC convolution, W A represents the angular difference weight matrix output by the ADC convolution, W H represents the horizontal difference weight matrix output by the HDC convolution, W V represents the vertical difference weight matrix output by the VDC convolution, W concat represents the concatenated weight matrix after concatenating the channel dimensions; W1 represents the fusion weight, and W2 represents the comprehensive weight.

[0095] In step S402, the determination of the interaction feature according to the fusion weight and the interpolation fusion feature is realized through the following expression:

[0096] Feature1 = W1⊙ x (10)

[0097] Among them, Feature1 represents the interaction feature, x represents the interpolation fusion feature, W1 represents the fusion weight, and ⊙ represents element-wise multiplication.

[0098] In step S403, determining the comprehensive feature according to the comprehensive weight and the interpolation fusion feature is implemented through the following expression:

[0099] Feature2 = Conv (W2⊙ x) (11)

[0100] Among them, Feature2 represents the comprehensive feature, W2 represents the comprehensive weight, and ⊙ represents element-wise multiplication.

[0101] In step S404, determining the spatial fusion feature according to the interaction feature and the comprehensive feature is implemented through the following expression:

[0102] α = σ(Feature1+ Feature2) (12)

[0103] Feature merged = α·Feature1+ (1-α)·Feature2 (13)

[0104] Feature = Feature merged + x (14)

[0105] Among them, σ represents the sigmoid function, α∈[0,1] represents the adaptive feature perception coefficient, Feature1 represents the interaction feature, Feature2 represents the comprehensive feature, Feature merged represents the fusion feature, Feature represents the spatial fusion feature, and x represents the interpolation fusion feature.

[0106] Specifically, in this embodiment, spatial reconstruction is completed through a multi-scale convolution combination, including 4 differential convolutions and 1 standard convolution. Among them, the central difference convolution (CDC) enhances the edge feature response through the central difference convolution, the horizontal difference convolution (HDC) and the vertical difference convolution (VDC) extract direction-sensitive features using horizontal and vertical gradient convolutions respectively, the angular difference convolution (ADC) can adapt to different angle changes and capture rotation-invariant features, and the standard convolution retains the original spatial context information. This application uses a dual-branch adaptive feature enhancement mechanism and residual connection to further improve the spatial reconstruction efficiency. The network structure of the MASR Module is as Figure 5 shown. As can be seen from the attached drawings, first, the input interpolation fusion feature x ∈ R C×H×W is input in parallel to 4 differential convolutions and 1 standard convolution. After the weight matrices output by each branch are concatenated along the channel dimension, 1×1 convolution is used to achieve cross-channel feature compression and interaction, generating a fusion weight W1 that fuses multi-scale features. Subsequently, dynamic enhancement is achieved through a dual-branch feature fusion mechanism: In branch one, the fusion weight W1 is applied to the input feature to generate an interaction feature Feature1 with channel interaction characteristics. In branch two, the five weight matrices generated by the multi-scale convolution combination are directly added to obtain a comprehensive weight W2, and a standard convolution operation is performed to generate a comprehensive feature Feature2. The sigmoid function is used to perform a non-linear transformation on Feature1 + Feature2, outputting an adaptive feature perception coefficient α ∈ [0, 1]. Through weighting, dynamic fusion of the two features is achieved, and high-frequency detail information is enhanced through multi-scale feature enhancement. At the same time, the gating mechanism can automatically adjust the contribution weights of each component according to the local characteristics of the input feature. Subsequent experiments show that this method has achieved the best results on datasets in different regions and scenarios, significantly improving the model generalization performance.

[0107] In the above embodiment, in step S50, the determining the residual fusion feature according to the spatial fusion feature includes: performing convolution processing on the spatial fusion feature to obtain a convolution spatial fusion feature; and performing a residual connection between the convolution spatial fusion feature and the spatial fusion feature to obtain the residual fusion feature.

[0108] Meanwhile, in step S60, determining the fused image according to the residual fusion feature and the spatial fusion feature includes:

[0109] S601: Performing convolution processing on the residual fusion feature to obtain a query feature, a key feature, and a value feature.

[0110] S602: Determine a spectral energy matrix according to the query feature and the key feature.

[0111] S603: Determine an attention map according to the spectral energy matrix.

[0112] S604: Determine the fused image according to the attention map, the value feature, and the residual fusion feature.

[0113] Specifically, the above steps are implemented through the following expressions:

[0114] A = Softmax(E) (15)

[0115] E = Q T K (16)

[0116] Y = β·A·V + X (17)

[0117] Y = β·Softmax(Q T K)·V + X (18)

[0118] Where Y represents the fused image, X represents the residual fusion feature, Q ∈ represents the query feature, Q T represents the transpose of the query feature, K ∈ represents the key feature, A ∈ represents the attention weight map, V ∈ represents the value feature, E ∈ represents the spectral energy matrix, Softmax represents the normalization operation, and β represents a learnable scaling factor.

[0119] Specifically, in the hyperspectral and multispectral image fusion task, due to its local receptive field characteristics, CNN often has difficulty in fully capturing the global dependencies and cross-band spectral information in the image. When traditional convolution operations are passed layer by layer, they can only capture the spectral feature interactions within the local neighborhood through filters of a fixed size, resulting in limited ability to express the non-linear coupling relationship between consecutive bands in the hyperspectral image. This limitation is prone to cause spectral distortion during the fusion process, that is, the spectral curve of the reconstructed image deviates from the true ground object reflectance, specifically manifested as the blurring and confusion of spectral features of different materials. The root cause lies in the lack of explicit modeling of the global correlation in the channel dimension during the feature extraction process, and the complex correlations between hundreds of bands of the spectral image cannot be effectively coordinated. To address the above problems, this application reconstructs the spectral information lost in the high-spatial-resolution and low-spectral-resolution multispectral image HR-MSI through an attention mechanism. Adopting a Query-Key-Value attention paradigm similar to Transformer, first, the input interpolated fusion features are respectively mapped to the query space and the key space, and the spectral energy matrix is calculated through matrix multiplication to generate an attention weight map with global perception ability. Subsequently, the weights are applied to the value space to reconstruct the features, and finally, the residual fusion of the attention features and the original features is realized through a learnable scaling factor β, as Figure 6 shown.

[0120] Specifically, first, the input residual fusion feature X ∈ passes through three independent 1×1 convolutional layers to respectively generate the query feature Q ∈ , the key feature K ∈ and the value feature V ∈ . It should be noted that we compress the number of channels in the query feature Q and the key feature K to one-eighth to reduce the subsequent attention calculation complexity. Secondly, the query feature Q is transposed and multiplied by the key feature K to calculate the spectral energy matrix E ∈ , and then through the Softmax normalization operation, the attention weight map A ∈ is obtained, thereby quantifying the spectral influence weight of each band group on other band groups. Then, the attention map A is multiplied by the value feature V to achieve feature recombination, so that the features of each band group are fused with the weighted spectral information of other band groups. Finally, the fusion ratio of the reconstructed feature and the original feature is controlled by the learnable parameter β. By establishing a global attention mapping, this application breaks through the local limitation of traditional convolution, enabling the model to dynamically focus on the regions related to the spectral characteristics of the current position, thereby effectively alleviating the spectral distortion problem caused by the local convolution receptive field limitation.

[0121] In this paper, Mean Squared Error (MSE) is used as the loss function to verify the obtained fusion image, and the specific expression is as follows:

[0122] (19)

[0123] Among them, ∈ represents the fused image generated by fusion, ∈ represents the ground truth. Among them, represents the ground truth pixel value at the spatial position in the k - band, represents the estimated value at this position. The property of its convex function means that there is only one global minimum and no local optimal solution problem. Therefore, it can keep the model relatively stable and more likely to converge during the training and optimization process. In the actual training process, we update the model parameters by minimizing the MSE loss, so that it can learn how to better fuse hyperspectral and multispectral images to generate a fused image closer to the real situation.

[0124] In the experiments of this paper, four widely used metrics are adopted to comprehensively evaluate from three aspects: the overall quality, the preservation degree of spatial information, and spectral information:

[0125] Metric 1: Root - Mean - Squared Error (RMSE) is used to estimate the difference between the reference HR - HSI and the estimated HR - HSI. The smaller the RMSE, the better the quality of the reconstructed image, as shown in formula (20).

[0126] (20)

[0127] Metric 2: Erreur Relative Globale Adimensionnelle de Synthèse (ERGAS) comprehensively examines the performance in both spatial and spectral aspects and is used to reflect the overall quality of image fusion. The lower the value of ERGAS, the higher the similarity between the fused image and the original image, and the better the fusion effect, as shown in formula (21). Among them, represents the spatial scaling ratio, represents the mean value of the pixel values in the k - th band of the original image.

[0128] (21)

[0129] Metric 3: Peak Signal - to - Noise Ratio (PSNR) is used to measure the fusion quality in the spatial domain and can evaluate the ability of the model to retain and reconstruct spatial details. The larger the PSNR value, the smaller the difference between the fused image and the original image, as shown in formula (22). Among them, represents the maximum value of the pixel values in the k - th band of the original image.

[0130] (22)

[0131] Indicator 4: The Spectral Angle Mapper (SAM) can focus on the relative relationship of spectral vectors, measure their similarity by calculating the angle between the spectra of the fused image and the reference image, and is used to evaluate the preservation degree of spectral information at each pixel. The smaller the SAM value, the better the spectral quality of the fused image, as shown in formula (23). Among them, represents the inner product operation.

[0132] (23)

[0133] Through the above four evaluation indicators, the proposed AMHF-Net is comprehensively compared with 10 other advanced algorithms in different datasets, and the results are elaborated and analyzed. The comparison algorithms used include both traditional algorithms, such as CNMF, SSE, LTTR, and deep learning-based methods (SSFCNN, ConSSFCNN, MSDCNN, TFNet, ResTFNet, SSR-NET, MCT-NET).

[0134]

[0135] The experimental results on the Washington DC dataset are shown in Table 1. The proposed method achieves the best results in terms of the overall image quality, spatial and spectral quality. This dataset covers complex mixed areas of cities and nature, and has high requirements for the model's ability to understand spatial context information. The performance of CNN-based methods is not outstanding compared to the dual-branch Transformer structure (MCT-Net). However, the method we proposed comprehensively outperforms MCT-Net, which benefits from the fact that AMHF-Net fully considers global dependencies and cross-channel information interaction.

[0136]

[0137]

[0138]

[0139] The experimental results on the Pavia Center, PaviaU, and Urban datasets are shown in Tables 2 to 4. These datasets include both the core areas of modern cities, a large number of historical buildings, and mixed areas with agriculture and forestry. Different lighting conditions have a significant impact on the spectral information of different ground objects. Thanks to the excellent module cooperation, the method we proposed shows the most prominent performance on the three datasets.

[0140]

[0141]

[0142] As shown in Tables 5 and 6, AMHF-Net ranks first in all indicators on the Botswana dataset, has the most prominent overall quality of the fused image in the IndianP dataset, and also achieves competitive results in the evaluation of spatial and spectral quality. Different from the datasets of urban scenes, the Botswana and IndianP datasets contain more complex ground objects such as farmland, forest, and wetland, and there are large spectral differences between different ground objects. Due to the MASR module we proposed, the model can adaptively adjust the feature weights according to the distribution of different ground objects and reconstruct HR-HSI efficiently and accurately. To further visually prove the effectiveness of AMHF-Net, we show the results of different method fusions in the form of pictures. As Figure 7 shown, the first row in each dataset is the HR-HSIs generated by the model, and the second row is the difference image between the estimated image processed by the pseudo-color technology and the reference R-G-B image. (a) SSFCNN. (b) ConSSFCNN. (c) MSDCNN. (d) TFNet. (e)ResTFNet. (f) SSR-NET. (g) MCT-NET. (h) AMHF-Net. (i) GT.

[0143] In summary, the experiments on 6 datasets with different characteristics show that the method not only has the highest quality image fusion ability but also has prominent generalization ability. It is worth noting that in the Washington DC Mall, urban, and Botswana datasets, there is almost a breakaway lead. From the perspectives of spatial reconstruction quality (ERGAS), spectral reconstruction quality (SAM), or element reconstruction quality (RMSE and PSNR), the proposed AMHF-Net has obvious advantages.

[0144] To verify the effectiveness of each module, ablation experiments on the CIUM, MASR, and SAE modules are carried out on the Washington DC Mall and Urban datasets, as shown in Tables 7 and 8.

[0145]

[0146]

[0147] The experimental results show that the roles of different modules vary in different datasets. In the Washington DC dataset, the CIUM and SAE modules perform better. When CIUM and SAE are enabled simultaneously, the RMSE is as low as 0.6009, the ERGAS is as low as 0.1041, the PSNR is as high as 49.8910, and the SAM is as low as 0.1997. This benefits from the channel interaction design of CIUM and the global attention mechanism of SAE, which can better handle the global feature correlation and channel information fusion of images. In the Urban dataset with complex landforms, the MASR module has more advantages. When MASR is enabled in combination with other modules, the fusion effect is better. Due to its multi-scale feature extraction method, the model can more accurately capture the diverse detailed features in the dataset with complex landforms. Generally speaking, the roles of the modules vary due to the characteristics of different datasets. CIUM and SAE are more suitable for global and channel modeling of large-size images, while MASR has more advantages in feature extraction of diverse landform data. When the three modules are applied simultaneously, the model performance is optimal.

[0148] This application also provides a hyperspectral and multispectral image fusion system based on adaptive multi-scale features, as Figure 8 shown, the system includes:

[0149] An acquisition module 10 for acquiring a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution;

[0150] A conversion module 20 for converting the hyperspectral image with low spatial resolution and high spectral resolution into upsampled features through a channel interaction upsampling module, where the spatial resolution of the upsampled features is the same as that of the multispectral image with high spatial resolution and low spectral resolution;

[0151] A first determination module 30 for determining interpolation fusion features based on the upsampled features and the multispectral image with high spatial resolution and low spectral resolution through a multi-scale adaptive spatial reconstruction module;

[0152] A second determination module 40 for determining spatial fusion features based on the interpolation fusion features;

[0153] A third determination module 50 for determining residual fusion features based on the spatial fusion features;

[0154] The fourth determination module 60, and the spectral self-attention enhancement module determines a fused image according to the residual fusion feature and the spatial fusion feature.

[0155] A computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the following method:

[0156] S10: Obtain a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution;

[0157] S20: Convert the hyperspectral image with low spatial resolution and high spectral resolution into an upsampled feature through a channel interaction upsampling module, where the spatial resolution of the upsampled feature is the same as the spatial resolution of the multispectral image with high spatial resolution and low spectral resolution;

[0158] S30: Determine an interpolation fusion feature according to the upsampled feature and the multispectral image with high spatial resolution and low spectral resolution through a multi-scale adaptive spatial reconstruction module;

[0159] S40: Determine a spatial fusion feature according to the interpolation fusion feature;

[0160] S50: Determine a residual fusion feature according to the spatial fusion feature;

[0161] S60: The spectral self-attention enhancement module determines a fused image according to the residual fusion feature and the spatial fusion feature.

[0162] Figure 9 The internal structure diagram of a computer device in an embodiment is shown. The computer device can specifically be a terminal or a server. As Figure 9 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and can also store a computer program. When the computer program is executed by the processor, the processor can implement a hyperspectral and multispectral image fusion method. The internal memory can also store a computer program. When the computer program is executed by the processor, the processor can execute the hyperspectral and multispectral image fusion method. Those skilled in the art can understand that Figure 9 the structure shown in

[0163] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0164] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0165] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A hyperspectral and multispectral image fusion method based on adaptive multi-scale features, characterized in that The method includes: Obtaining a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution; Converting the hyperspectral image with low spatial resolution and high spectral resolution into upsampled features through a channel interaction upsampling module, where the spatial resolution of the upsampled features is the same as that of the multispectral image with high spatial resolution and low spectral resolution; Determining interpolation fusion features based on the upsampled features and the multispectral image with high spatial resolution and low spectral resolution; Determining spatial fusion features based on the interpolation fusion features, including: performing a convolution operation on the interpolation fusion features to obtain a fusion weight and a comprehensive weight; determining interaction features based on the fusion weight and the interpolation fusion features; determining comprehensive features based on the comprehensive weight and the interpolation fusion features; determining spatial fusion features based on the interaction features and the comprehensive features; the performing a convolution operation on the interpolation fusion features to obtain a fusion weight and a comprehensive weight is implemented through the following expression: W S =Conv S (x) W C =Conv C (x) W A =Conv A (x) W H =Conv H (x) W V =Conv V (x) W concat =[W S ;W C ;W A ;W H ;W V ​ W1=Conv 1×1 (W concat ) W2 = W S + W C + W A + W H + W V Among them, x represents the interpolated fusion feature, W S represents the standard weight matrix of the standard convolution output, W C represents the central difference weight matrix of the CDC convolution output, W A represents the angular difference weight matrix of the ADC convolution output, W H represents the horizontal difference weight matrix of the HDC convolution output, W V represents the vertical difference weight matrix of the VDC convolution output, W concat represents the concatenated weight matrix after concatenating the channel dimensions; W1 represents the fusion weight, and W2 represents the comprehensive weight; Determining residual fusion features based on the spatial fusion features, including: performing a convolution process on the spatial fusion features to obtain a convolution spatial fusion feature; performing a residual connection between the convolution spatial fusion feature and the spatial fusion feature to obtain the residual fusion features; The spectral self-attention enhancement module determines a fusion image based on the residual fusion features and the spatial fusion features, where determining a fusion image based on the residual fusion features and the spatial fusion features includes: performing a convolution process on the residual fusion features to obtain a query feature, a key feature, and a value feature; determining a spectral energy matrix based on the query feature and the key feature; determining an attention map based on the spectral energy matrix; determining the fusion image based on the attention map, the value feature, and the residual fusion features; implemented through the following expressions: A = Softmax (E) E = Q T K Y = β·A·V + X Y = β·Softmax(Q T K)·V + X Among them, Y represents the fused image, X represents the residual fusion feature, Q ∈ represents the query feature, and Q T represents the transpose of the query feature, K ∈ represents the key feature, A ∈ represents the attention weight map, V ∈ represents the value feature, E ∈ represents the spectral energy matrix, Softmax represents the normalization operation, and β represents the learnable parameter.

2. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features according to claim 1, wherein The converting the hyperspectral image with low spatial resolution and high spectral resolution into upsampled features is implemented through the following expression: CIUM(lr)=Conv 1×1 (CS(LeakyReLU(BN(DWC 3×3 (Up 4× (lr)))))) Among them, CIUM(lr) represents the upsampled feature, lr represents the hyperspectral image with low spatial resolution and high spectral resolution, and Up 4× represents a four-fold upsampling operation, and DWC 3×3 represents a depth convolution with a size of 3×3, BN represents the batch normalization layer, LeakyReLU represents the activation function, CS represents the Channel Shuffle function, and Conv 1×1 represents a convolution with a size of 1.

3. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features according to claim 1, wherein The determining interaction features based on the fusion weight and the interpolation fusion features is implemented through the following expression: Feature1 = W1⊙ x where Feature1 represents the interaction features, x represents the interpolation fusion features, W1 represents the fusion weight, and ⊙ represents element-wise multiplication; The determining comprehensive features based on the comprehensive weight and the interpolation fusion features is implemented through the following expression: Feature2 = Conv (W2⊙x) where Feature2 represents the comprehensive features, W2 represents the comprehensive weight, and ⊙ represents element-wise multiplication.

4. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features according to claim 3, wherein Determining spatial fusion features based on the interaction features and the comprehensive features is implemented through the following expression: α = σ(Feature1 + Feature2) Feature merged = α·Feature1+ (1-α)·Feature2 Feature = Feature merged + x Among them, σ represents the sigmoid function, α ∈ [0, 1] represents the adaptive feature perception coefficient, Feature1 represents the interaction feature, Feature2 represents the comprehensive feature, Feature merged represents the fusion feature, Feature represents the spatial fusion feature, and x represents the interpolation fusion feature.

Citation Information

Patent Citations

  • Hyperspectral image and multispectral image fusion method based on multi-scale visual attention

    CN118918016A

  • Hyperspectral and multispectral image fusion method and system based on improved Transform

    CN119559066A