Hyperspectral and multispectral image fusion method based on adaptive multi-scale features

By adopting an adaptive multi-scale feature method in the fusion of hyperspectral and multispectral images, the problems of spatial details loss, insufficient cross-domain generalization capabilities and spectral distortion in image fusion are solved, and high-quality image fusion effect is achieved.

CN120047326AActive Publication Date: 2025-05-27NORTHEASTERN UNIV CHINA +1

Patent Information

Application Number
CN202510517884.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The hyperspectral and multispectral image fusion task faces spectral distortion problems caused by loss of spatial details, insufficient cross-domain generalization capabilities and low spectral resolution.

Method used

The image fusion method based on adaptive multi-scale features is adopted, and the hyperspectral and multi-spectral images are converted and fused through the channel interactive upsampling module, the multi-scale adaptive spatial reconstruction module and the spectral self-attention enhancement module to ensure the effective fusion of spatial and spectral information.

Benefits of technology

It effectively avoids spatial details loss, insufficient cross-domain generalization capabilities and spectral distortion problems, and improves the quality and accuracy of image fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047326A_ABST
    Figure CN120047326A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral and multispectral image fusion method based on adaptive multi-scale features. According to the method, firstly, a low-spatial-resolution high-spectral-resolution hyperspectral image and a high-spatial-resolution low-spectral-resolution multispectral image serve as input of a network and are subjected to up-sampling through a sampling module on channel interaction to achieve efficient feature expansion, and then the hyperspectral image and the multispectral image are fused into a super-multispectral image with the same size as a reference image; and a multi-scale adaptive spatial reconstruction module captures diversified features and reconstructs spatial information, optimized feature maps enter a spectrum self-attention enhancement module to perform spectrum attention enhancement, correlation between the feature maps is calculated to generate an attention map, and spectrum information is reconstructed. In the image fusion process, the problems of spatial detail loss, insufficient cross-domain generalization ability, spectral distortion caused by low spectral resolution and the like are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image fusion, and in particular to a hyperspectral and multispectral image fusion method based on adaptive multi-scale features. Background Art

[0002] With the intensification of global climate change and the increasing urgency of sustainable development issues, high-precision environmental monitoring and ecological resource management have become the core needs of the international community to address challenges. Hyperspectral and multispectral image fusion technology can provide key data support for climate change research and ecosystem protection by extracting refined parameters such as land cover, vegetation physiological status, and pollutant distribution. Therefore, it has become one of the important research directions in the field of remote sensing. The technology aims to integrate the fine spectral features of hyperspectral images (HSI) and the high spatial details of multispectral images (MSI), break through the physical performance limitations of a single imaging mode, make up for the deficiencies of a single sensor, and promote the leap of remote sensing from "data acquisition" to "information extraction". Compared with other remote sensing images, hyperspectral image HSI has higher spectral resolution and can capture the spectral information of objects through hundreds of continuous narrow bands, providing unique advantages for ground object recognition. However, its spatial resolution is usually low due to the influence of sensor signal-to-noise ratio and energy dispersion effect. On the contrary, multispectral image MSI has high spatial resolution, but the sparsity of spectral information limits its fine classification ability. However, the hyperspectral and multispectral image fusion task faces many challenges, and existing methods still face key challenges such as loss of spatial details, insufficient cross-domain generalization ability, and spectral distortion caused by low spectral resolution. Summary of the Invention

[0003] Based on this, it is necessary to address the above problems and propose a hyperspectral and multispectral image fusion method based on adaptive multi-scale features.

[0004] A hyperspectral and multispectral image fusion method based on adaptive multi-scale features, the method comprising: Obtaining a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution; Converting the hyperspectral image with low spatial resolution and high spectral resolution into an upsampled feature through a channel interaction upsampling module, the spatial resolution of the upsampled feature being the same as the spatial resolution of the multispectral image with high spatial resolution and low spectral resolution; Determining an interpolation fusion feature according to the upsampled feature and the multispectral image with high spatial resolution and low spectral resolution through a multi-scale adaptive spatial reconstruction module; Determine the spatial fusion feature according to the interpolation fusion feature, including: performing a convolution operation on the interpolation fusion feature to obtain a fusion weight and a comprehensive weight; determining an interaction feature according to the fusion weight and the interpolation fusion feature; determining a comprehensive feature according to the comprehensive weight and the interpolation fusion feature; determining the spatial fusion feature according to the interaction feature and the comprehensive feature; Determine the residual fusion feature according to the spatial fusion feature, including: performing a convolution process on the spatial fusion feature to obtain a convolutional spatial fusion feature; performing a residual connection on the convolutional spatial fusion feature and the spatial fusion feature to obtain the residual fusion feature; The spectral self-attention enhancement module determines a fused image according to the residual fusion feature and the spatial fusion feature.

[0005] In one embodiment, the determining the spatial fusion feature according to the interpolation fusion feature includes: Performing a convolution operation on the interpolation fusion feature to obtain a fusion weight and a comprehensive weight; Determining an interaction feature according to the fusion weight and the interpolation fusion feature; Determining a comprehensive feature according to the comprehensive weight and the interpolation fusion feature; Determining the spatial fusion feature according to the interaction feature and the comprehensive feature.

[0006] In one embodiment, determining the fused image according to the residual fusion feature and the spatial fusion feature includes: Performing a convolution process on the residual fusion feature to obtain a query feature, a key feature, and a value feature; Determining a spectral energy matrix according to the query feature and the key feature; Determining an attention map according to the spectral energy matrix; Determining the fused image according to the attention map, the value feature, and the residual fusion feature.

[0007] In one embodiment, the converting the hyperspectral image with low spatial resolution and high spectral resolution into an upsampled feature is implemented by the following expression: CIUM(lr) = Conv 1×1 (CS(LeakyReLU(BN(DWC 3×3 (Up 4× (lr)))))) where CIUM(lr) represents the upsampled feature, lr represents the hyperspectral image with low spatial resolution and high spectral resolution, Up 4× represents a four-fold upsampling operation, DWC 3×3Depth convolution with a size of 3×3 is denoted as, BN represents the batch normalization layer, LeakyReLU represents the activation function, CS represents the Channel Shuffle function, and Conv 1×1 denotes a convolution with a size of 1.

[0008] In one embodiment, the convolution operation on the interpolated fusion feature to obtain the fusion weight and the comprehensive weight is implemented by the following expression: W S =Conv S (x) W C =Conv C (x) W A =Conv A (x) W H =Conv H (x) W V =Conv V (x) W concat =[W S ; W C ; W A ; W H ; W V W 1 =Conv 1×1 (W concat ) W 2 =W S +W C +W A +W H +W V where x represents the interpolated fusion feature, W S represents the standard weight matrix output by the standard convolution, W C represents the central difference weight matrix output by the CDC convolution, W A represents the angular difference weight matrix output by the ADC convolution, W H represents the horizontal difference weight matrix output by the HDC convolution, W V represents the vertical difference weight matrix output by the VDC convolution, W concat represents the concatenated weight matrix after concatenation in the channel dimension; W 1 represents the fusion weight, and W 2 represents the comprehensive weight.

[0009] ​In one embodiment, determining the interaction feature according to the fusion weight and the interpolation fusion feature is implemented by the following expression: Feature 1 = W 1 ⊙ x where Feature 1 represents the interaction feature, x represents the interpolation fusion feature, W 1 represents the fusion weight, and ⊙ represents element-wise multiplication; Determining the comprehensive feature according to the comprehensive weight and the interpolation fusion feature is implemented by the following expression: Feature 2 = Conv (W 2 ⊙x) where Feature 2 represents the comprehensive feature, W 2 represents the comprehensive weight, and ⊙ represents element-wise multiplication.

[0010] In one embodiment, determining the spatial fusion feature according to the interaction feature and the comprehensive feature is implemented by the following expression: α = σ(Feature 1 + Feature 2 ) Feature merged = α·Feature 1 + (1-α)·Feature 2 Feature = Feature merged + x where σ represents the sigmoid function, α∈[0,1] represents the adaptive feature perception coefficient, Feature 1 represents the interaction feature, Feature 2 represents the comprehensive feature, Feature merged represents the fusion feature, Feature represents the spatial fusion feature, and x represents the interpolation fusion feature.

[0011] In one embodiment, determining the fused image according to the residual fusion feature and the spatial fusion feature is implemented by the following expression: A = Softmax (E) E = Q T K Y = β·A·V + X Y = β·Softmax (Q T K)·V + X Among them, Y represents the fused image, X represents the residual fusion feature, Q ∈ represents the query feature, and Q T represents the transpose of the query feature, K ∈ represents the key feature, A ∈ represents the attention weight map, V ∈ represents the value feature, E ∈ represents the spectral energy matrix, Softmax represents the normalization operation, and β represents the learnable parameter.

[0012] In this application, a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution are obtained. The hyperspectral image with low spatial resolution and high spectral resolution is the hyperspectral image with low spatial resolution and high spectral resolution, and the multispectral image with high spatial resolution and low spectral resolution is the multispectral image with high spatial resolution and low spectral resolution. The hyperspectral image with low spatial resolution and high spectral resolution is converted into an upsampled feature, and the spatial resolution of the upsampled feature is the same as that of the multispectral image with high spatial resolution and low spectral resolution. An interpolation fusion feature is determined according to the upsampled feature and the multispectral image with high spatial resolution and low spectral resolution. A spatial fusion feature is determined according to the interpolation fusion feature. A residual fusion feature is determined according to the spatial fusion feature. A fused image is determined according to the residual fusion feature and the spatial fusion feature. During the fusion process of the image, problems such as loss of spatial details, insufficient cross-domain generalization ability, and spectral distortion caused by low spectral resolution are avoided. Description of the Drawings

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0014] Among them: Figure 1 is an application environment diagram of a hyperspectral and multispectral image fusion method based on adaptive multi-scale features in an embodiment; Figure 2 is a flowchart of a hyperspectral and multispectral image fusion method based on adaptive multi-scale features in an embodiment; Figure 3 is a flowchart for obtaining the upsampled feature; Figure 4 is a flowchart for implementing the Channel Shuffle function; Figure 5It is the network structure diagram of the MASR Module; Figure 6 It is the structure diagram of the spectral self-attention enhancement module; Figure 7 It is the comparison diagram of the fusion effect; Figure 8 It is the structural block diagram of the hyperspectral and multispectral image fusion system based on adaptive multi-scale features in one embodiment; Figure 9 It is the structural block diagram of a computer device in one embodiment. Detailed implementation manners

[0015] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0016] With the intensification of global climate change and the increasing urgency of sustainable development issues, high-precision environmental monitoring and ecological resource management have become the core needs of the international community to address challenges. The hyperspectral and multispectral image fusion technology can provide key data support for climate change research and ecosystem protection by extracting refined parameters such as surface cover, vegetation physiological state, and pollutant distribution. Therefore, it has become one of the important research directions in the field of remote sensing. This technology aims to integrate the fine spectral features of hyperspectral images (HSIs) and the high spatial details of multispectral images (MSIs), break through the physical performance limitations of a single imaging mode, make up for the deficiencies of a single sensor, and promote the leap of remote sensing from "data acquisition" to "information extraction". Compared with other remote sensing images, the hyperspectral image HSI has a higher spectral resolution and can capture the spectral information of objects through hundreds of continuous narrow bands, providing unique advantages for ground object recognition. However, its spatial resolution is usually low due to the influence of the sensor signal-to-noise ratio and energy dispersion effect. On the contrary, the multispectral image MSI has a high spatial resolution, but the sparsity of spectral information limits its fine classification ability. However, the hyperspectral and multispectral image fusion task faces many challenges, and existing methods still face key challenges such as loss of spatial details, insufficient cross-domain generalization ability, and spectral distortion caused by low spectral resolution. To solve the above technical problems, this application provides a spectral and multispectral image fusion method, Figure 1 It is the application environment diagram of the hyperspectral and multispectral image fusion in one embodiment. Refer to Figure 1, the hyperspectral and multispectral image fusion method is applied to a hyperspectral and multispectral image fusion system. The hyperspectral and multispectral image fusion system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 can specifically be a desktop terminal or a mobile terminal, and the mobile terminal can specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 is used to obtain a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution. The hyperspectral image with low spatial resolution and high spectral resolution is a hyperspectral image with low spatial resolution and high spectral resolution, and the multispectral image with high spatial resolution and low spectral resolution is a multispectral image with high spatial resolution and low spectral resolution; the server 120 is used to convert the hyperspectral image with low spatial resolution and high spectral resolution into an upsampled feature, and the spatial resolution of the upsampled feature is the same as the spatial resolution of the multispectral image with high spatial resolution and low spectral resolution; determine an interpolation fusion feature according to the upsampled feature and the multispectral image with high spatial resolution and low spectral resolution; determine a spatial fusion feature according to the interpolation fusion feature; determine a residual fusion feature according to the spatial fusion feature; determine a fusion image according to the residual fusion feature and the spatial fusion feature.

[0017] As Figure 2 shown, in one embodiment, a hyperspectral and multispectral image fusion method based on adaptive multi-scale features is provided. This method can be applied to both the terminal and the server. In this embodiment, taking the application to the terminal as an example, the hyperspectral and multispectral image fusion method based on adaptive multi-scale features specifically includes the following steps: S10: Obtain a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution.

[0018] S20: Convert the hyperspectral image with low spatial resolution and high spectral resolution into an upsampled feature through a channel interaction upsampling module, and the spatial resolution of the upsampled feature is the same as the spatial resolution of the multispectral image with high spatial resolution and low spectral resolution.

[0019] S30: Determine an interpolation fusion feature according to the upsampled feature and the multispectral image with high spatial resolution and low spectral resolution through a multi-scale adaptive spatial reconstruction module.

[0020] S40: Determine a spatial fusion feature according to the interpolation fusion feature.

[0021] S50: Determine a residual fusion feature according to the spatial fusion feature.

[0022] S60: The spectral self-attention enhancement module determines a fused image based on the residual fusion feature and the spatial fusion feature.

[0023] This application avoids problems such as loss of spatial details, insufficient cross-domain generalization ability, and spectral distortion caused by low spectral resolution during the image fusion process.

[0024] In step S20, the conversion of the hyperspectral image with low spatial resolution and high spectral resolution into an upsampled feature is achieved through the following expression: CIUM(lr) = Conv 1×1 (CS(LeakyReLU(BN(DWC 3×3 (Up 4× (lr)))))) (1) where CIUM(lr) represents the upsampled feature, lr represents the hyperspectral image with low spatial resolution and high spectral resolution, Up 4× represents a four-fold upsampling operation, DWC 3×3 represents a depth convolution of size 3×3, BN represents a batch normalization layer, LeakyReLU represents an activation function, CS represents a Channel Shuffle function, and Conv 1×1 represents a convolution of size 1.

[0025] Specifically, in the hyperspectral and multispectral image fusion task, due to the different spatial resolutions of the input images, it is usually necessary to first upsample the low-resolution hyperspectral image to match the high spatial resolution of the multispectral image, and then perform subsequent feature extraction and fusion operations. Traditional upsampling methods mainly expand the feature map size based on pixel value interpolation, paying more attention to the expansion of the spatial dimension, but there are deficiencies in retaining the context information of the feature map and the information interaction between channels. Since upsampling is a crucial step in this task and is usually applied at the beginning of the forward propagation of the network, the spatial information and channel interaction ignored by it often have a more obvious impact on the performance of the model and the image fusion effect. To solve the problems of spatial information loss and edge feature blurring caused by traditional upsampling methods, this paper proposes a channel interaction upsampling module (CIUM), as Figure 3 shown. At the same time, the multi-scale adaptive spatial reconstruction module is represented by MASR, and the spectral self-attention enhancement module is represented by SAE.

[0026] By gradually upsampling the features, the spatial resolution of the upsampled features is increased to the spatial resolution of the multi-spectral image with high spatial resolution and low spectral resolution. Specifically, the high-spectral image with low spatial resolution and high spectral resolution, LR-HSI, is used as the input of the module, and 4-fold upsampling is performed to obtain a feature map with the same spatial resolution as the multi-spectral image with high spatial resolution and low spectral resolution, i.e., the multi-spectral image HR-MIS with high spatial resolution and low spectral resolution. Then, the spatial details are enhanced through depth convolution, BatchNorm, and the LeakyReLU activation function. Subsequently, the Channel Shuffle function is used to rearrange the channels of the high-spectral image with low spatial resolution and high spectral resolution, enabling full fusion of information from different channels and thus increasing the information interaction between channels. Finally, pointwise convolution is used to adjust the number of channels to obtain the final output upsampled features.

[0027] Depth convolution with a convolution kernel size of 3×3 is performed on each channel separately, which can more effectively capture local spatial features within the channels. Taking the edge information in the image as an example, depth convolution can accurately extract detailed features such as the direction and intensity of the edge through the convolution calculation of pixel values at the edge of each channel. Compared with traditional upsampling methods, it has a better effect of extracting and retaining spatial detail information such as edges and textures in the image. However, in the prior art, it is proposed that depth convolution is performed independently within each channel, which easily ignores the relationship between different channels. Therefore, we introduce the ChannelShuffle function operation to enhance the interaction ability of the model between different channels. The implementation process of the Channel Shuffle function is as Figure 4 shown. First, the input high-spectral image with low spatial resolution and high spectral resolution, lr ∈ R C×H×W is reshaped (where C represents the total number of channels of the feature map, H represents the height of the feature map, and W represents the width of the feature map), expanding 1 dimension of the input feature channels into 2 dimensions, namely the number of convolution groups G and the number of channels included in each convolution N, thereby obtaining the feature matrix Q ∈ R G×N×H×W , where G × N = C. Secondly, the first and second dimensions of the matrix Q are transposed to obtain the matrix S ∈ R N×G×H×W . Finally, the 2 dimensions in the matrix S are flattened into 1 dimension to obtain the final upsampled feature map CIUM(lr) ∈ R C ×H×W , thus realizing the information flow between channels.

[0028] In step S30, determining the interpolation fusion feature according to the upsampled feature and the multi-spectral image with high spatial resolution and low spectral resolution includes: constructing the interpolation fusion feature by means of a gap interpolation strategy for the multi-spectral image with high spatial resolution and low spectral resolution and the upsampled feature.

[0029] In one embodiment, determining the spatial fusion feature according to the interpolation fusion feature includes: S401: Performing a convolution operation on the interpolation fusion feature to obtain a fusion weight and a comprehensive weight.

[0030] S402: Determining an interaction feature according to the fusion weight and the interpolation fusion feature.

[0031] S403: Determining a comprehensive feature according to the comprehensive weight and the interpolation fusion feature.

[0032] S404: Determining the spatial fusion feature according to the interaction feature and the comprehensive feature.

[0033] Specifically, in step S401, performing a convolution operation on the interpolation fusion feature to obtain a fusion weight and a comprehensive weight is implemented through the following expressions: W S =Conv S (x) (2) W C =Conv C (x) (3) W A =Conv A (x) (4) W H =Conv H (x) (5) W V =Conv V (x) (6) W concat =[W S ; W C ; W A ; W H ; W V (7) W 1 =Conv 1×1 (W concat ) (8) W 2 =W S +W C +W A +W H +W V (9) wherein, x represents the interpolation fusion feature, W S represents the standard weight matrix output by the standard convolution, W C represents the central difference weight matrix output by the CDC convolution, W ARepresents the angular difference weight matrix of the ADC convolution output, W H Represents the horizontal difference weight matrix of the HDC convolution output, W V Represents the vertical difference weight matrix of the VDC convolution output, W concat Represents the concatenated weight matrix after concatenating the channel dimensions; W 1 Represents the fusion weight, W 2 Represents the comprehensive weight.

[0034] In step S402, the interaction feature is determined based on the fusion weight and the interpolated fusion feature, and is implemented through the following expression: Feature 1 = W 1 ⊙ x (10) where Feature 1 represents the interaction feature, x represents the interpolated fusion feature, W 1 represents the fusion weight, and ⊙ represents element-wise multiplication.

[0035] In step S403, the comprehensive feature is determined based on the comprehensive weight and the interpolated fusion feature, and is implemented through the following expression: Feature 2 = Conv (W 2 ⊙ x) (11) where Feature 2 represents the comprehensive feature, W 2 represents the comprehensive weight, and ⊙ represents element-wise multiplication.

[0036] In step S404, the spatial fusion feature is determined based on the interaction feature and the comprehensive feature, and is implemented through the following expression: α = σ(Feature 1 + Feature 2 ) (12) Feature merged = α·Feature 1 + (1-α)·Feature 2 (13) Feature = Feature merged + x (14) where σ represents the sigmoid function, α∈[0,1] represents the adaptive feature perception coefficient, Feature 1 represents the interaction feature, Feature 2 represents the comprehensive feature, Feature mergedFusion feature is denoted as, spatial fusion feature is denoted as Feature, and interpolation fusion feature is denoted as x.

[0037] Specifically, in this embodiment, spatial reconstruction is completed through a multi-scale convolution combination, including 4 differential convolutions and 1 standard convolution. Among them, the central difference convolution (CDC) enhances the edge feature response through the central difference convolution, the horizontal difference convolution (HDC) and the vertical difference convolution (VDC) extract direction-sensitive features by using horizontal and vertical gradient convolutions respectively, the angular difference convolution (ADC) can adapt to different angle changes and capture rotation-invariant features, and the standard convolution retains the original spatial context information. This application uses a dual-branch adaptive feature enhancement mechanism and a residual connection to further improve the spatial reconstruction efficiency. The network structure of the MASR Module is as Figure 5 shown. It can be seen from the attached drawings that first, the input interpolation fusion feature x ∈ R C×H×W is input in parallel to 4 differential convolutions and 1 standard convolution. After the weight matrices output by each branch are concatenated in the channel dimension, cross-channel feature compression and interaction are realized through a 1×1 convolution to generate the fusion weight W 1 for the fused multi-scale features. Subsequently, dynamic enhancement is achieved through a dual-branch feature fusion mechanism: In the first branch, the fusion weight W 1 acts on the input feature to generate an interaction feature Feature 1 with channel interaction characteristics. In the second branch, the five weight matrices generated by the multi-scale convolution combination are directly added to obtain the comprehensive weight W 2 , and a standard convolution operation is performed to generate the comprehensive feature Feature 2 . The sigmoid function is used to perform a non-linear transformation on Feature 1 +Feature 2 to output the adaptive feature perception coefficient α ∈ [0, 1]. Dynamic fusion of the two features is achieved through weighting. High-frequency detail information is enhanced through multi-scale feature enhancement, and at the same time, the gating mechanism can automatically adjust the contribution weights of each component according to the local characteristics of the input feature. Subsequent experiments show that this method has achieved the best results on datasets in different regions and scenarios, significantly improving the model generalization performance.

[0038] In the above embodiments, in step S50, the determining the residual fusion feature according to the spatial fusion feature includes: performing convolution processing on the spatial fusion feature to obtain a convolutional spatial fusion feature; performing residual connection on the convolutional spatial fusion feature and the spatial fusion feature to obtain the residual fusion feature.

[0039] Meanwhile, in step S60, determining the fused image according to the residual fusion feature and the spatial fusion feature includes: S601: Performing convolution processing on the residual fusion feature to obtain a query feature, a key feature, and a value feature.

[0040] S602: Determining a spectral energy matrix according to the query feature and the key feature.

[0041] S603: Determining an attention map according to the spectral energy matrix.

[0042] S604: Determining the fused image according to the attention map, the value feature, and the residual fusion feature.

[0043] Specifically, the above steps are implemented through the following expressions: A = Softmax(E) (15) E = Q T K (16) Y = β·A·V + X (17) Y = β·Softmax(Q T K)·V + X (18) where Y represents the fused image, X represents the residual fusion feature, Q ∈ represents the query feature, Q T represents the transpose of the query feature, K ∈ represents the key feature, A ∈ represents the attention weight map, V ∈ represents the value feature, E ∈ represents the spectral energy matrix, Softmax represents the normalization operation, and β represents a learnable scaling factor.

[0044] Specifically, in the hyperspectral and multispectral image fusion task, due to its local receptive field characteristics, CNN often has difficulty fully capturing the global dependencies and cross-band spectral information in images. When traditional convolution operations are passed layer by layer, they can only capture the spectral feature interactions within the local neighborhood through filters of a fixed size, resulting in limited expression ability for the non-linear coupling relationships between consecutive bands in hyperspectral images. This limitation is prone to cause spectral distortion during the fusion process, that is, the spectral curve of the reconstructed image deviates from the true ground object reflectance, specifically manifested as the blurring and confusion of spectral features of different materials. The root cause lies in the lack of explicit modeling of the global correlation in the channel dimension during the feature extraction process, and the complex correlations between hundreds of bands in spectral images cannot be effectively coordinated. To address the above problems, this application reconstructs the spectral information lost in the high-spatial-resolution and low-spectral-resolution multispectral image HR-MSI through an attention mechanism. Adopting a Query-Key-Value attention paradigm similar to Transformer, first, the input interpolated fusion features are respectively mapped to the query space and the key space, and the spectral energy matrix is calculated through matrix multiplication to generate an attention weight map with global perception ability. Subsequently, the weights are applied to the value space to reconstruct the features, and finally, the residual fusion of the attention features and the original features is achieved through a learnable scaling factor β, as Figure 6 shown.

[0045] Specifically, first, the input residual fusion feature X ∈ passes through three independent 1×1 convolutional layers to respectively generate the query feature Q ∈ , the key feature K ∈ and the value feature V ∈ . It should be noted that we compress the number of channels in the query feature Q and the key feature K to one-eighth to reduce the subsequent attention calculation complexity. Secondly, the query feature Q is transposed and matrix-multiplied with the key feature K to calculate the spectral energy matrix E ∈ , and then through the Softmax normalization operation, the attention weight map A ∈ is obtained, thereby quantifying the spectral influence weights of each band group on other band groups. Then, the attention map A is multiplied by the value feature V to achieve feature recombination, so that the features of each band group are fused with the weighted spectral information of other band groups. Finally, the fusion ratio of the reconstructed feature and the original feature is controlled by the learnable parameter β. This application breaks through the local limitation of traditional convolution by establishing a global attention mapping, enabling the model to dynamically focus on the regions related to the spectral characteristics of the current position, thereby effectively alleviating the spectral distortion problem caused by the local convolution receptive field limitation.

[0046] In this paper, Mean Squared Error (MSE) is used as the loss function to verify the obtained fusion image, and the specific expression is as follows: (19) Among them, ∈ represents the fused image generated by fusion, ∈ represents the ground truth. Among them, represents the ground truth pixel value at the spatial position in the k - band, represents the estimated value at this position. The property of its convex function means that there is only one global minimum and no local optimal solution problem. Therefore, it can make the model remain relatively stable and converge more easily during the training and optimization process. In the actual training process, we update the model parameters by minimizing the MSE loss, so that it can learn how to better fuse hyperspectral and multispectral images to generate a fused image closer to the real situation.

[0047] In this paper's experiments, four widely used metrics are adopted to comprehensively evaluate from three aspects: overall quality, preservation degree of spatial information, and spectral information: Metric 1: Root - Mean - Squared Error (RMSE) is used to estimate the difference between the reference HR - HSI and the estimated HR - HSI. The smaller the RMSE, the better the quality of the reconstructed image, as shown in formula (20).

[0048] (20) Metric 2: Erreur Relative Globale Adimensionnelle de Synthèse (ERGAS) comprehensively examines the performance in both spatial and spectral aspects and is used to reflect the overall quality of image fusion. The lower the value of ERGAS, the higher the similarity between the fused image and the original image, and the better the fusion effect, as shown in formula (21). Among them, represents the spatial scaling ratio, represents the mean value of the pixel values in the k - band of the original image.

[0049] (21) Metric 3: Peak Signal - to - Noise Ratio (PSNR) is used to measure the fusion quality in the spatial domain and can evaluate the model's ability to retain and reconstruct spatial details. The larger the PSNR value, the smaller the difference between the fused image and the original image, as shown in formula (22). Among them, represents the maximum value of the pixel values in the k - band of the original image.

[0050] (22) Indicator 4: The Spectral Angle Mapper (SAM) can focus on the relative relationship of spectral vectors and measure their similarity by calculating the angle between the spectral of the fused image and the reference image, which is used to evaluate the preservation degree of spectral information at each pixel. The smaller the SAM value, the better the spectral quality of the fused image, as shown in formula (23). Among them, represents the inner product operation.

[0051] (23) Through the above four evaluation indicators, the proposed AMHF-Net is comprehensively compared with 10 other advanced algorithms in different datasets, and the results are elaborated and analyzed. The comparison algorithms used include both traditional algorithms, such as CNMF, SSE, LTTR, and deep learning-based methods (SSFCNN, ConSSFCNN, MSDCNN, TFNet, ResTFNet, SSR-NET, MCT-NET).

[0052] The experimental results on the Washington DC dataset are shown in Table 1. The proposed method achieves the best results in terms of the overall image quality, spatial and spectral quality. This dataset covers complex mixed areas of cities and nature, and has high requirements for the model's ability to understand spatial context information. The performance of CNN-based methods is not outstanding compared to the dual-branch Transformer structure (MCT-Net). However, the proposed method comprehensively outperforms MCT-Net, which benefits from the fact that AMHF-Net fully considers global dependencies and cross-channel information interaction.

[0053] The experimental results on the Pavia Center, PaviaU, and Urban datasets are shown in Tables 2 to 4. These datasets include both the core areas of modern cities, a large number of historical buildings, and mixed areas with agriculture and forestry. Different lighting levels have a great impact on the spectral information of different ground objects. Thanks to the excellent module cooperation, the proposed method has the most prominent performance on the three datasets.

[0054] As shown in Table 5 and Table 6, AMHF-Net ranked first in all indicators on the Botswana dataset. The overall quality of the fused images in the IndianP dataset was the most prominent, and competitive results were also obtained in the evaluation of spatial and spectral quality. Different from the datasets of urban scenes, the Botswana and IndianP datasets contain more complex ground objects such as farmland, forests, and wetlands, and the spectral differences between different ground objects are relatively large. Due to the proposed MASR module, the model can adaptively adjust the feature weights according to the distribution of different ground objects, and reconstruct HR-HSI efficiently and accurately. To further intuitively demonstrate the effectiveness of AMHF-Net, we present the results of different methods in the form of pictures. As Figure 7 shown, the first row in each dataset is the HR-HSIs generated by the model, and the second row is the difference image between the estimated image processed by the pseudo-color technology and the reference R-G-B image. (a) SSFCNN. (b) ConSSFCNN. (c) MSDCNN. (d) TFNet. (e)ResTFNet. (f) SSR-NET. (g) MCT-NET. (h) AMHF-Net. (i) GT。

[0055] In summary, experiments on 6 datasets with different characteristics show that the method not only has the highest-quality image fusion ability but also has outstanding generalization ability. It is worth noting that in the Washington DC Mall, urban, and Botswana datasets, an almost overwhelming lead was achieved. From the perspectives of spatial reconstruction quality (ERGAS), spectral reconstruction quality (SAM), or element reconstruction quality (RMSE and PSNR), the proposed AMHF-Net has obvious advantages.

[0056] To verify the effectiveness of each module, ablation experiments were conducted on the CIUM, MASR, and SAE modules on the Washington DC Mall and Urban datasets, as shown in Table 7 and Table 8.

[0057] The experimental results show that the roles of different modules vary in different datasets. In the Washington DC dataset, the CIUM and SAE modules perform better. When CIUM and SAE are enabled simultaneously, the RMSE is as low as 0.6009, the ERGAS is as low as 0.1041, the PSNR is as high as 49.8910, and the SAM is as low as 0.1997. This benefits from the channel interaction design of CIUM and the global attention mechanism of SAE, which can better handle the global feature correlation and channel information fusion of images. In the Urban dataset with complex landforms, the MASR module has more advantages. When MASR is enabled in combination with other modules, the fusion effect is better. Due to its multi-scale feature extraction method, the model can more accurately capture the diverse detailed features in the dataset with complex landforms. Generally speaking, the roles of the modules vary due to the characteristics of different datasets. CIUM and SAE are more suitable for global and channel modeling of large-sized images, while MASR has more advantages in feature extraction of diverse landform data. When the three modules are applied simultaneously, the model performance is optimal.

[0058] This application also provides a hyperspectral and multispectral image fusion system based on adaptive multi-scale features, as Figure 8 shown, the system includes: An acquisition module 10, which acquires a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution; A conversion module 20, which converts the hyperspectral image with low spatial resolution and high spectral resolution into upsampled features through a channel interaction upsampling module, and the spatial resolution of the upsampled features is the same as that of the multispectral image with high spatial resolution and low spectral resolution; A first determination module 30, which determines an interpolation fusion feature according to the upsampled features and the multispectral image with high spatial resolution and low spectral resolution through a multi-scale adaptive spatial reconstruction module; A second determination module 40, which determines a spatial fusion feature according to the interpolation fusion feature; A third determination module 50, which determines a residual fusion feature according to the spatial fusion feature; A fourth determination module 60, and a spectral self-attention enhancement module determines a fused image according to the residual fusion feature and the spatial fusion feature.

[0059] A computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the following method: S10: Acquire a hyperspectral image with low spatial resolution and high spectral resolution and a multispectral image with high spatial resolution and low spectral resolution; S20: Convert the hyperspectral image with low spatial resolution and high spectral resolution into an upsampled feature through a channel interaction upsampling module, where the spatial resolution of the upsampled feature is the same as that of the multispectral image with high spatial resolution and low spectral resolution; S30: Determine an interpolation fusion feature according to the upsampled feature and the multispectral image with high spatial resolution and low spectral resolution through a multi-scale adaptive spatial reconstruction module; S40: Determine a spatial fusion feature according to the interpolation fusion feature; S50: Determine a residual fusion feature according to the spatial fusion feature; S60: The spectral self-attention enhancement module determines a fused image according to the residual fusion feature and the spatial fusion feature.

[0060] Figure 9 The internal structure diagram of a computer device in an embodiment is shown. The computer device can specifically be a terminal or a server. As Figure 9 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and can also store a computer program. When the computer program is executed by the processor, the processor can implement the hyperspectral and multispectral image fusion method. The internal memory can also store a computer program. When the computer program is executed by the processor, the processor can execute the hyperspectral and multispectral image fusion method. Those skilled in the art can understand that Figure 9 the structure shown in

[0061] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0062] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0063] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A hyperspectral and multispectral image fusion method based on adaptive multi-scale features, characterized in that: The method comprises: Acquire hyperspectral images with low spatial resolution and high spectral resolution and multispectral images with high spatial resolution and low spectral resolution; The high-spectral image with low spatial resolution and high spectral resolution is converted into up-sampled features through a channel interactive up-sampling module, wherein the spatial resolution of the up-sampled features is the same as the spatial resolution of the multi-spectral image with high spatial resolution and low spectral resolution; Determining interpolation fusion features according to the up-sampling features and the multispectral image with high spatial resolution and low spectral resolution through a multiscale adaptive spatial reconstruction module; Determining a spatial fusion feature according to the interpolation fusion feature includes: performing a convolution operation on the interpolation fusion feature to obtain a fusion weight and a comprehensive weight; determining an interaction feature according to the fusion weight and the interpolation fusion feature; determining a comprehensive feature according to the comprehensive weight and the interpolation fusion feature; and determining a spatial fusion feature according to the interaction feature and the comprehensive feature. Determining a residual fusion feature according to the spatial fusion feature includes: performing convolution processing on the spatial fusion feature to obtain a convolution spatial fusion feature; performing residual connection on the convolution spatial fusion feature and the spatial fusion feature to obtain the residual fusion feature; The spectral self-attention enhancement module determines a fused image according to the residual fusion feature and the spatial fusion feature.

2. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features according to claim 1, characterized in that: Determining a fused image according to the residual fusion feature and the spatial fusion feature includes: Convolution processing is performed on the residual fusion features to obtain query features, key features and value features; determining a spectral energy matrix based on the query features and the key features; determining an attention map according to the spectral energy matrix; The fused image is determined according to the attention map, the value feature and the residual fusion feature.

3. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features according to claim 1, characterized in that: The conversion of the hyperspectral image with low spatial resolution and high spectral resolution into up-sampled features is achieved by the following expression: CIUM(lr)=Conv 1×1 (CS(LeakyReLU(BN(DWC 3×3 (Up 4× (lr)))))) Among them, CIUM (lr) represents the upsampling feature, lr represents the hyperspectral image with low spatial resolution and high spectral resolution, Up 4× Indicates a four-fold upsampling operation, DWC 3×3 represents a 3×3 depth convolution, BN represents a batch normalization layer, LeakyReLU represents an activation function, CS represents a Channel Shuffle function, Conv 1×1 represents a convolution of size 1.

4. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features according to claim 1, characterized in that: The convolution operation on the interpolation fusion feature to obtain the fusion weight and the comprehensive weight is implemented by the following expression: W S =Conv S (x) W C =Conv C (x) W A =Conv A (x) W H =Conv H (x) W V =Conv V (x) IN concat =[In S ;IN C ;IN A ;IN H ;IN V ] W1=Conv 1×1 (W concat ) W2=W S +W C +W A +W H +W V Among them, x represents the interpolation fusion feature, W S represents the standard weight matrix of the standard convolution output, W C Represents the central difference weight matrix of the CDC convolution output, W A Represents the angle differential weight matrix of the ADC convolution output, W H Denotes the horizontal differential weight matrix of the HDC convolution output, W V Represents the vertical difference weight matrix of the VDC convolution output, W concat Represents the splicing weight matrix after channel dimension splicing; W1 represents the fusion weight, and W2 represents the comprehensive weight.

5. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features according to claim 4 is characterized in that: The determining of the interactive feature according to the fusion weight and the interpolation fusion feature is implemented by the following expression: Feature1 = W1⊙ x Among them, Feature1 represents the interactive feature, x represents the interpolation fusion feature, W1 represents the fusion weight, and ⊙ represents element-by-element multiplication; The comprehensive feature is determined according to the comprehensive weight and the interpolation fusion feature, and is implemented by the following expression: Feature2 = Conv (W2⊙x) Among them, Feature2 represents the comprehensive feature, W2 represents the comprehensive weight, and ⊙ represents element-by-element multiplication.

6. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features according to claim 5, characterized in that: The spatial fusion feature is determined according to the interaction feature and the comprehensive feature, and is implemented by the following expression: α = σ(Feature1+ Feature2) Feature merged = α·Feature1+ (1-α)·Feature2 Feature = Feature merged + x Among them, σ represents the sigmoid function, α∈[0,1] represents the adaptive feature perception coefficient, Feature1 represents the interactive feature, Feature2 represents the comprehensive feature, and Feature merged represents fusion features, Feature represents spatial fusion features, and x represents interpolation fusion features.

7. The hyperspectral and multispectral image fusion method based on adaptive multi-scale features according to claim 2, characterized in that: This is achieved through the following expression: A = Softmax (E) E=Q T K Y = β·A·V + X Y=β·Softmax (Q T K)·V + X Among them, Y represents the fused image, X represents the residual fusion feature, and Q∈ represents the query feature, Q T represents the transpose of the query feature, K∈ represents the key feature, A∈ represents the attention weight map, V∈ Represents the value feature, E∈ represents the spectral energy matrix, Softmax represents the normalization operation, and β represents the learnable parameter.

Citation Information

Patent Citations

  • Transform deformation-based multi-scale hyperspectral and multispectral image fusion method and system

    CN118537233A

  • Hyperspectral image and multispectral image fusion method based on multi-scale visual attention

    CN118918016A

  • Hyperspectral and multispectral image fusion method and system based on improved Transform

    CN119559066A

  • Retina vessel segmentation method based on difference feature fusion

    CN119559190A

  • Change detection and change monitoring of natural and man-made features in multispectral and hyperspectral satellite imagery

    US20160307073A1

Cited By

  • Hyperspectral image classification method based on spectral feature reconstruction

    CN120356015A