Hyperspectral and multispectral image fusion method based on dynamic double-branch interaction network
Patent Information
- Application Number
- CN202611080472.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-07-21
AI Technical Summary
[0004]尽管上述深度学习方法在实验指标上取得了长足进步,但在处理HR-MSI与LR-HSI之间显著的模态异构性时,依然面临着空间结构一致性与光谱保真度难以兼顾的技术瓶颈
本发明针对高、多光谱图像模态差异分别设计专用特征提取支路,借助哈尔小波频域模块强化多光谱图像空间纹理边缘,依靠双池化通道注意力完整保留高光谱连续光谱信息;并通过全局/局部窗口双分支交叉注意力同步实现全局长程依赖建模与局部像素精细对齐,搭配可学习动态平衡头自适应调节两类交互特征权重,解决了现有融合网络无法兼顾全局结构与局部细节、空间与光谱特征难以均衡融合的缺陷,可大幅提升复杂遥感场景下特征表达能力。
Smart Images

Figure CN122597204B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image fusion technology, specifically relating to a hyperspectral and multispectral image fusion method based on a dynamic bi-branch interactive network. Background Technology
[0002] Traditional methods for hyperspectral and multispectral image fusion (HMIF) are primarily based on rigorous mathematical frameworks and physical imaging models. Their technological evolution can be clearly divided into two main branches: spatial detail injection and low-rank physical modeling. One branch, represented by component substitution (CS) and multi-resolution analysis (MRA), involves spatial detail injection. Its core logic is to map the low-resolution hyperspectral image (LR-HSI) to a specific feature space such as principal components or brightness components. Then, high-frequency spatial details are extracted from the high-resolution multispectral image (HR-MSI) using methods such as wavelet decomposition and forcibly injected. A representative algorithm is the fusion architecture combining principal component analysis (PCA) and wavelet transform. The other branch, physical modeling methods, such as matrix / tensor decomposition (MF / TF) and spectral unmixing techniques, tend to treat LR-HSI as a product of low-rank matrices or tensors. They estimate the shared endmember spectra and high-resolution abundance maps through an alternating iterative process (such as coupled nonnegative matrix factorization (CNMF)). However, these methods are inherently highly dependent on the researchers' hand-designed prior knowledge and highly idealized degradation model assumptions. This makes them often exhibit inherent drawbacks such as insufficient feature characterization ability and limited model robustness when facing extremely complex spatial heterogeneity and nonlinear spectral distortion in real remote sensing scenarios.
[0003] With the rapid development of deep learning technology, the ability to model nonlinear features based on large-scale data has become the mainstream technical paradigm in the current HMIF field, undergoing a significant transformation from a single convolutional stream to a multi-branch collaborative architecture. Convolutional Neural Networks (CNNs), with their inherent local connectivity and weight-sharing mechanisms, have demonstrated superior performance in extracting fine-scale spatial features such as edges and textures from HR-MSI. Their development has evolved from early single-stream networks (such as SSFCNN and MSDCNN) that performed regression mapping through simple channel concatenation to dual-stream interactive architectures (such as TFNet and SSRNet) capable of independently encoding features for spectral and spatial branches. In recent years, researchers have further explored the possibilities of hybrid architectures, attempting to combine the powerful local perception capabilities of CNNs with the excellent long-range global context modeling capabilities of Transformers. By introducing self-attention mechanisms to ensure the consistency of the fused image in macroscopic structure, for example, multi-level cross-Transformers (such as the MCT architecture) can achieve deep interaction of cross-modal features at the same scale, thereby compensating to some extent for the missing spatial details or spectral information between different modalities.
[0004] Despite significant advancements in experimental metrics achieved by the aforementioned deep learning methods, they still face a technical bottleneck in addressing the substantial modal heterogeneity between HR-MSI and LR-HSI, where achieving both spatial structure consistency and spectral fidelity remains challenging. Specifically, the local receptive field of traditional convolutional operations limits their ability to capture global spatial dependencies in HR-MSI, risking the loss of fine-grained edge information. Meanwhile, the standard Transformer mechanism often lacks targeted discriminative modeling strategies when dealing with HSI, a modality characterized by continuous, dense, and narrow-band distribution. Existing methods often fail to achieve a dynamic balance between macroscopic global background perception and microscopic local detail alignment during the fusion process, leading to edge blurring, jagged edges, or significant spectral color shifts in complex scenes (such as urban buildings and areas with non-uniform vegetation cover). Summary of the Invention
[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a hyperspectral and multispectral image fusion method based on a dynamic bi-branch interactive network. Through frequency domain spatial enhancement, spectral channel attention and global-local bi-branch adaptive fusion, the spatial detail clarity and spectral fidelity of the fused image can be guaranteed simultaneously.
[0006] To achieve the above objectives, this invention provides a hyperspectral and multispectral image fusion method based on a dynamic bi-branch interactive network, comprising the following steps: S1. For the input high-resolution multispectral image and low-resolution hyperspectral image, the multi-scale feature extraction module extracts the multi-level initial multi-scale features of the two images in the multi-scale dimension. S2. Input the initial multi-scale features of the corresponding level of the high-resolution multispectral image into the frequency-driven structure extraction module, realize frequency domain decomposition based on discrete Haar wavelet transform, extract high-frequency spatial texture and boundary features, and obtain the corresponding level optimized spatial feature map; The initial multi-scale features of the corresponding level of the low-resolution hyperspectral image are input into the spectral channel feature extraction module. The channel attention is constructed by dual pooling joint shared multilayer perceptron, the complete continuous spectral features are extracted, and the optimized spectral feature map of the corresponding level is output. S3. The optimized spatial feature map and spectral feature map of each level are fed into the dynamic dual-branch fusion module. The global cross-attention branch and the local window cross-attention branch are used to complete the bidirectional cross-modal feature interaction. The adaptive fusion weight is obtained by normalizing the two sets of implicit learnable scalar parameters of the built-in dynamic balancing head through Softmax. The output features of the two attention branches are balanced, and the multi-scale fusion features corresponding to each level are generated layer by layer. S4. Construct an image reconstruction module to concatenate and stitch the multi-scale fusion features corresponding to all levels along the channel dimension. Gradually restore the original size of the image through multi-layer convolution, activation functions and progressive upsampling operations, and reconstruct and output the fused target image. S5. The fusion model composed of the spatial edge gradient loss, spectral angle mapping loss, and mean square error loss is used to train the fusion model. Based on the trained fusion model, the high-resolution multispectral image and the low-resolution hyperspectral image of the actual input are fused.
[0007] As a preferred embodiment of the present invention, the processing procedure of the multi-scale feature extraction module in S1 is as follows: Shallow basic features are extracted by using 3×3 convolution with ReLU activation function on the input high-resolution multispectral image and low-resolution hyperspectral image respectively. The shallow basic features are downsampled layer by layer to generate three different resolution layered feature maps. After normalization, all layered feature maps are used to obtain three-level initial multi-scale features. The initial multi-scale features corresponding to each level of the high-resolution multispectral image are respectively sent to the frequency-driven structure extraction module, while the initial multi-scale features corresponding to each level of the low-resolution hyperspectral image are respectively sent to the spectral channel feature extraction module.
[0008] As a preferred embodiment of the present invention, in step S2, the input to the frequency-driven structure extraction module is the initial multi-scale features corresponding to a single-level, high-resolution multispectral image, and the processing procedure of the frequency-driven structure extraction module is as follows: The input single-level initial multi-scale features are subjected to discrete Haar wavelet transform to decompose them into low-frequency structural components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. A spatial guidance map is generated using low-frequency structural components, and the texture and edge details of the high-frequency components in three directions are enhanced by relying on an attention mechanism. The inverse Haar wavelet transform is performed on the low-frequency structural components and the enhanced high-frequency components in each direction. After fusion and reconstruction, the optimized spatial feature map of this level is output and sent to the dynamic dual-branch fusion module.
[0009] As a preferred embodiment of the present invention, in step S2, the input to the spectral channel feature extraction module is the initial multi-scale features corresponding to a single-level, low-resolution hyperspectral image, and the processing procedure of the spectral channel feature extraction module is as follows: Global average pooling and global max pooling are performed simultaneously on the input single-level initial multi-scale features to extract two sets of global statistical features of spectral bands. The two sets of global statistical features of spectral bands are input into a shared multilayer perceptron. The output features of the two perceptrons are summed and then processed by the Sigmoid activation function to calculate and normalize the attention weights corresponding to each spectral channel. The attention weights corresponding to each spectral channel are multiplied element-wise with the initial multi-scale features of the single-level input, and then the multi-band information is integrated through a 3×3 convolution layer to output the optimized spectral feature map of this level. The optimized spectral feature map of this level is then sent to the dynamic dual-branch fusion module.
[0010] As a preferred embodiment of the present invention, in S3, the dynamic dual-branch fusion module comprises three parts: a global cross-attention branch, a local window cross-attention branch, and a dynamic balancing head. The input to the dynamic dual-branch fusion module is the optimized spatial feature map and the optimized spectral feature map paired at the same level. The processing procedure is as follows: The global cross-attention branch uses the optimized spatial feature map at the same level as the query matrix and the spectral feature map as the key matrix and value matrix to calculate the long-range dependency of the global range across modalities and output the enhanced spatial interaction features. At the same time, it uses the spectral feature map as the query matrix and the spatial feature map as the key matrix and value matrix to output the enhanced spectral interaction features. The enhanced spatial interaction features and spectral interaction features together constitute the global interaction features at this level. The local window cross-attention branch divides the two types of single-level feature maps into non-overlapping windows of equal size, performs cross-attention calculation only within the window, achieves fine alignment of local pixels, and outputs the local interaction features of that level. The dynamic balancing head is configured with two sets of independent, unconstrained, implicitly learnable scalar parameters. First, the global and local interactive features corresponding to the high-resolution multispectral images within the same level are summed and then normalized by layer. At the same time, the global and local interactive features corresponding to the low-resolution hyperspectral images are summed and then normalized by layer. Then, the two sets of independent, unconstrained, implicitly learnable scalar parameters are normalized by two-dimensional Softmax to obtain two sets of constrained fusion weights. The normalized two-mode fusion features are weighted and summed by the two sets of constrained fusion weights to generate the multi-scale fusion features corresponding to that level. The multi-scale fusion features corresponding to that level are then sent to the image reconstruction module.
[0011] As a preferred embodiment of the present invention, in S4, the process of the image reconstruction module reconstructing and outputting the fused target image is as follows: the multi-scale fusion features corresponding to the three levels are concatenated and spliced in the channel dimension to obtain spliced fusion features. The spliced fusion features are then integrated with depth information by multi-layer 3×3 convolution and ReLU activation function in sequence. The original width and height of the image are gradually restored by stepwise upsampling operation, and the fused target image is obtained by decoding and reconstruction.
[0012] In a preferred embodiment of the present invention, in step S5, the combination loss function L is: ; In the formula, , , For learnable scalar noise scale parameters; For spatial edge gradient loss, Represents the spatial edge gradient of the real reference image. This represents the spatial edge gradient of the fused target image. The L1 norm is represented by H, W, and C, which represent the height, width, and number of channels of the image, respectively. HWC represents the height × width × number of channels of the image, i.e., the total number of pixels. For spectral angle mapping loss, This represents the spectral vector of the target image at a specific pixel location. This represents the spectral vector of the real reference image at the corresponding pixel location. HW represents the height × width of the image, and i and j are the indices of the height and width of the image, respectively; the superscript T indicates transpose. For mean square error loss, Represents a real reference image. This indicates the target image to be fused.
[0013] As a preferred embodiment of the present invention, in S5, the training process is as follows: a training sample set is formed by pairing and matching low-resolution hyperspectral images, high-resolution multispectral images, and corresponding real reference images; the training samples are input in batches into a fusion model composed of a multi-scale feature extraction module, a frequency-driven structure extraction module, a spectral channel feature extraction module, a dynamic dual-branch fusion module, and an image reconstruction module. The fusion model sequentially executes steps S1 to S4 to output the fused target image and calculates the combined loss function value between the fused target image and the real reference image. The backpropagation algorithm combined with the gradient descent optimizer is used to update the network parameters of all modules in the fusion model layer by layer in reverse based on the combined loss function value. The process of iteratively executing sample forward inference, loss calculation, and parameter update continues until the combined loss function converges to a preset threshold, thus completing the training of the fusion model and obtaining the trained fusion model.
[0014] The beneficial effects of this invention are: This invention designs dedicated feature extraction branches for modal differences in hyperspectral and multispectral images, enhances the spatial texture edges of multispectral images using a Haar wavelet frequency domain module, and fully preserves continuous hyperspectral spectral information through dual-pooling channel attention. Furthermore, it achieves global long-range dependency modeling and fine local pixel alignment through global / local window dual-branch cross-attention synchronization, and adaptively adjusts the weights of two types of interactive features using a learnable dynamic balancing head. This solves the shortcomings of existing fusion networks that cannot take into account both global structure and local details, and that spatial and spectral features are difficult to fuse in a balanced way, and can significantly improve the feature representation capability in complex remote sensing scenarios.
[0015] This invention constructs a combined loss function that includes gradient edge loss, spectral angle mapping loss, and mean square error loss. The model training is simultaneously constrained from three dimensions: image contour sharpness, spectral vector fidelity, and overall pixel fidelity, correcting the edge blurring and spectral drift problems caused by single loss functions. The entire modular network architecture, which integrates multi-scale extraction, frequency domain / spectral optimization, dual-branch fusion, and image reconstruction, is end-to-end trainable. It does not require manual priors or complex pre-decomposition operations, resulting in stronger model generalization. The final output fused image possesses both clear and complete spatial details and spectral features that closely match real-world features. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the principle of this invention; Figure 2 This is a flowchart illustrating the process of obtaining the fused target image based on the trained fusion model. Detailed Implementation
[0017] The embodiments of the present invention will be further described below with reference to the accompanying drawings: Example 1: As Figure 1 As shown, the hyperspectral and multispectral image fusion method based on a dynamic dual-branch interactive network includes the following steps: S1. For the input high-resolution multispectral image and low-resolution hyperspectral image, the multi-scale feature extraction module extracts the multi-level initial multi-scale features of the two images in the multi-scale dimension. S2. Input the initial multi-scale features of the corresponding level of the high-resolution multispectral image into the frequency-driven structure extraction module, realize frequency domain decomposition based on discrete Haar wavelet transform, extract high-frequency spatial texture and boundary features, and obtain the corresponding level optimized spatial feature map; The initial multi-scale features of the corresponding level of the low-resolution hyperspectral image are input into the spectral channel feature extraction module. The channel attention is constructed by dual pooling joint shared multilayer perceptron, the complete continuous spectral features are extracted, and the optimized spectral feature map of the corresponding level is output. S3. The optimized spatial feature map and spectral feature map of each level are fed into the dynamic dual-branch fusion module. The global cross-attention branch and the local window cross-attention branch are used to complete the bidirectional cross-modal feature interaction. The adaptive fusion weight is obtained by normalizing the two sets of implicit learnable scalar parameters of the built-in dynamic balancing head through Softmax. The output features of the two attention branches are balanced, and the multi-scale fusion features corresponding to each level are generated layer by layer. S4. Construct an image reconstruction module to concatenate and stitch the multi-scale fusion features corresponding to all levels along the channel dimension. Gradually restore the original size of the image through multi-layer convolution, activation functions and progressive upsampling operations, and reconstruct and output the fused target image. S5. The fusion model composed of the spatial edge gradient loss, spectral angle mapping loss, and mean square error loss is used to train the fusion model. Based on the trained fusion model, the high-resolution multispectral image and the low-resolution hyperspectral image of the actual input are fused.
[0018] Based on the trained fusion model, a schematic diagram is shown below of fusing the actual input high-resolution multispectral image (HR-MSI) and low-resolution hyperspectral image (LR-HSI) to obtain the fused target image (HR-HSI). Figure 2 As shown.
[0019] In S1, the processing procedure of the multi-scale feature extraction module is as follows: Shallow basic features are extracted by using 3×3 convolution with ReLU activation function on the input high-resolution multispectral image and low-resolution hyperspectral image respectively. The shallow basic features are downsampled layer by layer to generate three different resolution layered feature maps. After normalization, all layered feature maps are used to obtain three-level initial multi-scale features. The three different resolutions are the original shallow resolution, medium resolution, and lowest resolution. The original shallow resolution is the highest resolution, and the corresponding layered feature map has not undergone downsampling and its size is exactly the same as the input image, which is used to preserve the complete original spatial details. The layered feature map corresponding to the medium resolution is obtained by performing one downsampling on the shallow basic features, and the width and height are simultaneously reduced to 1 / 2 of the original image, which is used to extract medium-scale ground feature contour features. The layered feature map corresponding to the lowest resolution is obtained by performing another downsampling on the medium resolution feature map, and the width and height are reduced to 1 / 4 of the original image, which is used to capture global large-scale ground feature structure and long-distance contextual dependencies.
[0020] The initial multi-scale features corresponding to each level of the high-resolution multispectral image are respectively sent to the frequency-driven structure extraction module, while the initial multi-scale features corresponding to each level of the low-resolution hyperspectral image are respectively sent to the spectral channel feature extraction module.
[0021] In S2, the input to the frequency-driven structure extraction module is the initial multi-scale features corresponding to a single-level, high-resolution multispectral image. The processing procedure of the frequency-driven structure extraction module is as follows: The input single-level initial multi-scale features are subjected to discrete Haar wavelet transform to decompose them into low-frequency structural components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. A spatial guidance map is generated using low-frequency structural components, and the texture and edge details of the high-frequency components in three directions are enhanced by relying on attention mechanisms (such as single-channel global spatial attention). The process of generating the spatial guide map is as follows: first, global average pooling and global max pooling are performed on the low-frequency structural components to extract the global spatial statistics of the low-frequency structural components. After 1×1 convolution to compress the channels, a single-channel spatial guide map is generated through the Sigmoid activation function.
[0022] The inverse Haar wavelet transform is performed on the low-frequency structural components and the enhanced high-frequency components in each direction. After fusion and reconstruction, the optimized spatial feature map of this level is output and sent to the dynamic dual-branch fusion module.
[0023] The input to the spectral channel feature extraction module is the initial multi-scale features corresponding to a single-level, low-resolution hyperspectral image. The processing procedure of the spectral channel feature extraction module is as follows: Global average pooling and global max pooling are simultaneously performed on the input single-level initial multi-scale features to extract two sets of global statistical features for spectral bands. These two sets of global statistical features are then input into a shared multilayer perceptron. The output features of the two perceptrons are summed and processed using a sigmoid activation function. The attention weights for each spectral channel are then calculated and normalized. ; In the formula, This represents the initial multi-scale features of the input single-level layer; express Attention weights corresponding to the spectral channels; The sigmoid activation function is represented by MLP; MLP stands for shared multilayer perceptron. , They represent respectively to The pooling results after performing global average pooling and global max pooling; The attention weights corresponding to each spectral channel are multiplied element-wise with the initial multi-scale features of the single-level input, and then the multi-band information is integrated through a 3×3 convolution layer to output the optimized spectral feature map of this level. The optimized spectral feature map of this level is then sent to the dynamic dual-branch fusion module.
[0024] In S3, the dynamic dual-branch fusion module consists of three parts: a global cross-attention branch, a local window cross-attention branch, and a dynamic balancing head. The input to the dynamic dual-branch fusion module is the optimized spatial feature map and the optimized spectral feature map paired at the same level. The processing procedure is as follows: The global cross-attention branch uses the optimized spatial feature map at the same level as the query matrix and the spectral feature map as the key matrix and value matrix to calculate the long-range dependency of the global range across modalities and output the enhanced spatial interaction features. At the same time, it uses the spectral feature map as the query matrix and the spatial feature map as the key matrix and value matrix to output the enhanced spectral interaction features. The enhanced spatial interaction features and spectral interaction features together constitute the global interaction features at this level. The local window cross-attention branch divides the two types of single-level feature maps (i.e., the optimized spatial feature map and the optimized spectral feature map paired at the same level) into non-overlapping windows of equal size. Cross-attention calculation is only performed within the window to achieve fine alignment of local pixels and output the local interaction features of that level. The dynamic balancing head is configured with two sets of independent, unconstrained, implicitly learnable scalar parameters. and First, the global and local interaction features corresponding to the high-resolution multispectral images within the same layer are summed and then normalized. Simultaneously, the global and local interaction features corresponding to the low-resolution hyperspectral images are summed and then normalized. Then... and Two sets of constraint fusion weights were obtained after two-dimensional Softmax normalization. and ,satisfy and ,pass and The normalized two types of modal fusion features are weighted and summed to generate the multi-scale fusion features corresponding to the level, and the multi-scale fusion features corresponding to the level are sent to the image reconstruction module.
[0025] Generate multi-scale fusion features corresponding to a certain level Represented as: ; In the formula, LN represents the layer normalization operation; This represents the global interactive features output from a high-resolution multispectral image after a global cross-attention branch. This represents the local interactive features output from a high-resolution multispectral image through a local window cross-attention branch; This represents the global interactive features output from a low-resolution hyperspectral image after a global cross-attention branch. This represents the local interactive features output from a low-resolution hyperspectral image through a local window cross-attention branch. The model is automatically updated during training using backpropagation. and Indirect adjustment and , Modulate the overall contribution of spatial texture information in fused images. The overall contribution of spectral information is regulated, and the fusion weights of global long-distance ground object context association and local fine pixel matching information are dynamically balanced in a coordinated manner, without the need for manual weight fixing.
[0026] For example, for large, contiguous land features (woodlands, water bodies) with simple spatial structures and significant spectral characteristics, fusion models can improve performance. This amplifies the overall contribution of spectral information, ensuring accurate reconstruction of ground feature spectral characteristics. For small, fragmented targets and textured areas with rich spatial details and dramatic spectral variations, the fusion model will improve... This amplifies the overall contribution of spatial texture information and ensures the clear preservation of edge structures.
[0027] In S4, the process of the image reconstruction module reconstructing the output fused target image is as follows: the multi-scale fusion features corresponding to the three levels are concatenated and stitched together in the channel dimension to obtain the stitched fusion features. The stitched fusion features are then integrated with depth information by passing multiple layers of 3×3 convolution and ReLU activation function. The original width and height of the image are gradually restored by step-by-step upsampling operation, and the fused target image is obtained by decoding and reconstruction.
[0028] In S5, the combination loss function L is: ; In the formula, , , The learnable scalar noise scale parameter corresponds to the spatial edge gradient loss, spectral angle mapping loss, and mean square error loss, respectively. The initial value is set to 1, and it is updated synchronously with all weights of the network during the training phase through backpropagation. , , These are adaptive weighting coefficients; the greater the uncertainty of the corresponding loss term, the smaller the weight. , , It is an uncertainty regularization constraint term to prevent the uncertainty from increasing indefinitely and causing the weight to return to zero.
[0029] For spatial edge gradient loss, Represents the spatial edge gradient of the real reference image. This represents the spatial edge gradient of the fused target image. The L1 norm is represented by H, W, and C, which represent the height, width, and number of channels of the image, respectively. HWC represents the height × width × number of channels of the image, i.e., the total number of pixels. For spectral angle mapping loss, This represents the spectral vector of the target image at a specific pixel location. This represents the spectral vector of the real reference image at the corresponding pixel location. HW represents the height × width of the image, and i and j are the indices of the height and width of the image, respectively; the superscript T indicates transpose. For mean square error loss, Represents a real reference image. This indicates the target image to be fused.
[0030] The training process is as follows: a training sample set is constructed by pairing low-resolution hyperspectral images, high-resolution multispectral images, and corresponding real reference images. The training samples are then input in batches into a fusion model consisting of a multi-scale feature extraction module, a frequency-driven structure extraction module, a spectral channel feature extraction module, a dynamic dual-branch fusion module, and an image reconstruction module. The fusion model sequentially executes steps S1 to S4 to output the fused target image and calculates the combined loss function value between the fused target image and the real reference image. The backpropagation algorithm combined with the gradient descent optimizer is used to update the network parameters of all modules in the fusion model layer by layer in reverse based on the combined loss function value. The process of iteratively executing sample forward inference, loss calculation, and parameter update continues until the combined loss function converges to a preset threshold, thus completing the training of the fusion model and obtaining the trained fusion model.
[0031] In terms of fusion performance metrics, the method in this embodiment has achieved industry-leading levels on multiple public datasets such as Salinas and Houston. Its performance in multiple core evaluation criteria, such as root mean square error (RMSE), peak signal-to-noise ratio (PSNR), and spectral angle mapping (SAM), is comprehensively superior to the current mainstream state-of-the-art methods. The fused image not only has extremely high edge sharpness in visual appearance, but also maintains a very high degree of fit with the real reference image in spectral curves, which has excellent engineering application and scientific research value.
[0032] Example 2: The difference between this example and Example 1 is that a preprocessing step is added to the multi-scale feature extraction module: The input low-resolution hyperspectral image is upsampled to the same spatial size as the high-resolution multispectral image to obtain the spectral reference image. Shallow basic features are extracted from both the high-resolution multispectral image and the spectral reference image using 3×3 convolution with ReLU activation function. The shallow basic features are downsampled layer by layer to generate three different resolution layered feature maps. All layered feature maps are normalized to obtain three levels of initial multiscale features. The initial multiscale features corresponding to each level of the high-resolution multispectral image are sent to the frequency-driven structure extraction module, and the initial multiscale features corresponding to each level of the spectral reference image are sent to the spectral channel feature extraction module.
[0033] In subsequent steps, the initial multi-scale features corresponding to each level of the spectral reference image are used to replace the initial multi-scale features corresponding to each level of the low-resolution hyperspectral image.
[0034] By pre-sampling low-resolution hyperspectral images to the same spatial size as high-resolution multispectral images to obtain spectral reference images, spatial scale alignment of two features can be achieved without losing spectral channel information. This enables subsequent multi-scale feature extraction modules and dynamic dual-branch fusion modules to perform layer-by-layer pairing interactions at the same spatial resolution level, avoiding the complex operation of cross-scale feature alignment.
[0035] In the image reconstruction module, the output fused target image is reconstructed using the following method: The multi-scale fusion features corresponding to the three levels are reconstructed step by step from bottom to top. Starting from the lowest resolution level, the reconstructed features of the current level are upsampled to the spatial size of the adjacent higher-level multi-scale fusion features and added to them element by element. Then, the deep information is integrated through the reconstruction sub-network of the corresponding level. This process is iterated until the highest resolution level is reached to obtain the highest-level reconstructed features. The highest-level reconstructed features are upsampled to the original spatial size of the high-resolution multispectral image and mapped to the number of hyperspectral channels to obtain the initial fused image. The initial fused image, the high-resolution multispectral image, and the spectral reference image are concatenated along the channel dimension to obtain the residual stitching features. A 3×3 convolution operation is performed on the residual stitching features and added element by element to the spectral reference image to obtain the basic fused image. The basic fused image is input into the spatial residual branch and the spectral residual branch respectively to extract the spatial thinning residual and the spectral thinning residual. The basic fused image is added element by element to the weighted spatial thinning residual and the spectral thinning residual, and the fused target image is decoded and reconstructed.
[0036] Compared to the reconstruction method in Example 1, which concatenates all multi-scale fusion features across all levels and then performs unified convolutional upsampling, this example adopts a bottom-up, step-by-step upsampling fusion strategy. By upsampling the reconstruction features of low-resolution levels layer by layer and adding them element by element to the fusion features of adjacent higher-level levels, global semantic information is gradually transmitted and preserved during the process of progressively restoring spatial resolution. This avoids spatial detail distortion and multi-scale information aliasing caused by a single large-scale upsampling. At the same time, each level is configured with an independent reconstruction sub-network, which can perform differentiated deep integration for fusion features of different scales. Combined with the residual connection based on the spectral reference image at the end and the spatial and spectral dual-branch refinement structure, it can adaptively compensate for spatial texture and spectral information respectively, and can more effectively ensure the high-frequency spatial detail clarity and spectral fidelity of the fused target image simultaneously.
[0037] In spectral angle mapping loss, adding a minimum positive number To prevent the denominator from being zero, thus avoiding training crashes in extreme cases: ; Example 3: A hyperspectral and multispectral image fusion device based on a dynamic dual-branch interactive network, comprising: One or more processors; Memory, used to store one or more computer programs; When one or more programs are executed by one or more processors, the one or more processors perform the method in Embodiment 1 or Embodiment 2.
[0038] Example 4: A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method in Example 1 or Example 2.
[0039] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can make equivalent substitutions or modifications based on the technical solution and concept of the present invention within the scope of the technology disclosed in the present invention, and such modifications should also be considered to fall within the scope of protection of the present invention.
Claims
1. A hyperspectral and multispectral image fusion method based on a dynamic double-branch interaction network, characterized in that Includes the following steps: S1. For the input high-resolution multispectral image and low-resolution hyperspectral image, the multi-scale feature extraction module extracts the multi-level initial multi-scale features of the two images in the multi-scale dimension. S2. Input the initial multi-scale features of the corresponding level of the high-resolution multispectral image into the frequency-driven structure extraction module, realize frequency domain decomposition based on discrete Haar wavelet transform, extract high-frequency spatial texture and boundary features, and obtain the corresponding level optimized spatial feature map; The initial multi-scale features of the corresponding level of the low-resolution hyperspectral image are input into the spectral channel feature extraction module. The channel attention is constructed by dual pooling joint shared multilayer perceptron, the complete continuous spectral features are extracted, and the optimized spectral feature map of the corresponding level is output. S3. The optimized spatial feature map and spectral feature map of each level are fed into the dynamic dual-branch fusion module. The global cross-attention branch and the local window cross-attention branch are used to complete the bidirectional cross-modal feature interaction. The adaptive fusion weight is obtained by normalizing the two sets of implicit learnable scalar parameters of the built-in dynamic balancing head through Softmax. The output features of the two attention branches are balanced, and the multi-scale fusion features corresponding to each level are generated layer by layer. S4. Construct an image reconstruction module to concatenate and stitch the multi-scale fusion features corresponding to all levels along the channel dimension. Gradually restore the original size of the image through multi-layer convolution, activation functions and progressive upsampling operations, and reconstruct and output the fused target image. S5. The fusion model composed of the spatial edge gradient loss, spectral angle mapping loss, and mean square error loss is used to train the fusion model. Based on the trained fusion model, the high-resolution multispectral image and the low-resolution hyperspectral image of the actual input are fused.
2. The hyperspectral and multispectral image fusion method based on a dynamic dual-branch interactive network according to claim 1, characterized in that, In S1, the processing procedure of the multi-scale feature extraction module is as follows: Shallow basic features are extracted by using 3×3 convolution with ReLU activation function on the input high-resolution multispectral image and low-resolution hyperspectral image respectively. The shallow basic features are downsampled layer by layer to generate three different resolution layered feature maps. After normalization, all layered feature maps are used to obtain three-level initial multi-scale features. The initial multi-scale features corresponding to each level of the high-resolution multispectral image are respectively sent to the frequency-driven structure extraction module, while the initial multi-scale features corresponding to each level of the low-resolution hyperspectral image are respectively sent to the spectral channel feature extraction module.
3. The hyperspectral and multispectral image fusion method based on a dynamic dual-branch interactive network according to claim 1, characterized in that, In S2, the input to the frequency-driven structure extraction module is the initial multi-scale features corresponding to a single-level, high-resolution multispectral image. The processing procedure of the frequency-driven structure extraction module is as follows: The input single-level initial multi-scale features are subjected to discrete Haar wavelet transform to decompose them into low-frequency structural components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. A spatial guidance map is generated using low-frequency structural components, and the texture and edge details of the high-frequency components in three directions are enhanced by relying on an attention mechanism. The inverse Haar wavelet transform is performed on the low-frequency structural components and the enhanced high-frequency components in each direction. After fusion and reconstruction, the optimized spatial feature map of this level is output and sent to the dynamic dual-branch fusion module.
4. The hyperspectral and multispectral image fusion method based on a dynamic dual-branch interactive network according to claim 1, characterized in that, In S2, the input to the spectral channel feature extraction module is the initial multi-scale features corresponding to a single-level, low-resolution hyperspectral image. The processing procedure of the spectral channel feature extraction module is as follows: Global average pooling and global max pooling are performed simultaneously on the input single-level initial multi-scale features to extract two sets of global statistical features of spectral bands. The two sets of global statistical features of spectral bands are input into a shared multilayer perceptron. The output features of the two perceptrons are summed and then processed by the Sigmoid activation function to calculate and normalize the attention weights corresponding to each spectral channel. The attention weights corresponding to each spectral channel are multiplied element-wise with the initial multi-scale features of the single-level input, and then the multi-band information is integrated through a 3×3 convolution layer to output the optimized spectral feature map of this level. The optimized spectral feature map of this level is then sent to the dynamic dual-branch fusion module.
5. The hyperspectral and multispectral image fusion method based on a dynamic dual-branch interactive network according to claim 1, characterized in that, In S3, the dynamic dual-branch fusion module comprises three parts: a global cross-attention branch, a local window cross-attention branch, and a dynamic balancing head. The input to the dynamic dual-branch fusion module is the optimized spatial feature map and the optimized spectral feature map paired at the same level. The processing procedure is as follows: The global cross-attention branch uses the optimized spatial feature map at the same level as the query matrix and the spectral feature map as the key matrix and value matrix to calculate the long-range dependency of the global range across modalities and output the enhanced spatial interaction features. At the same time, it uses the spectral feature map as the query matrix and the spatial feature map as the key matrix and value matrix to output the enhanced spectral interaction features. The enhanced spatial interaction features and spectral interaction features together constitute the global interaction features at this level. The local window cross-attention branch divides the two types of single-level feature maps into non-overlapping windows of equal size, performs cross-attention calculation only within the window, achieves fine alignment of local pixels, and outputs the local interaction features of that level. The dynamic balancing head is configured with two sets of independent, unconstrained, implicitly learnable scalar parameters. First, the global and local interactive features corresponding to the high-resolution multispectral images within the same level are summed and then normalized by layer. At the same time, the global and local interactive features corresponding to the low-resolution hyperspectral images are summed and then normalized by layer. Then, the two sets of independent, unconstrained, implicitly learnable scalar parameters are normalized by two-dimensional Softmax to obtain two sets of constrained fusion weights. The normalized two-mode fusion features are weighted and summed by the two sets of constrained fusion weights to generate the multi-scale fusion features corresponding to that level. The multi-scale fusion features corresponding to that level are then sent to the image reconstruction module.
6. The hyperspectral and multispectral image fusion method based on a dynamic dual-branch interactive network according to claim 2, characterized in that, In S4, the process of the image reconstruction module reconstructing the output fused target image is as follows: the multi-scale fusion features corresponding to the three levels are concatenated and stitched together in the channel dimension to obtain the stitched fusion features. The stitched fusion features are then integrated with depth information by passing multiple layers of 3×3 convolution and ReLU activation function in sequence. The original width and height of the image are gradually restored by stepwise upsampling operation, and the fused target image is obtained by decoding and reconstruction.
7. The hyperspectral and multispectral image fusion method based on a dynamic dual-branch interactive network according to claim 1, characterized in that, In S5, the combined loss function L is: ; In the formula, , , For learnable scalar noise scale parameters; For spatial edge gradient loss, Represents the spatial edge gradient of the real reference image. This represents the spatial edge gradient of the fused target image. The L1 norm is represented by H, W, and C, which represent the height, width, and number of channels of the image, respectively. HWC represents the height × width × number of channels of the image, i.e., the total number of pixels. For spectral angle mapping loss, This represents the spectral vector of the target image at a specific pixel location. This represents the spectral vector of the real reference image at the corresponding pixel location. HW represents the height × width of the image, and i and j are the indices of the height and width of the image, respectively; the superscript T indicates transpose. For mean square error loss, Represents a real reference image. This indicates the target image to be fused.
8. The hyperspectral and multispectral image fusion method based on a dynamic dual-branch interactive network according to claim 1, characterized in that, In S5, the training process is as follows: a training sample set is formed by pairing low-resolution hyperspectral images, high-resolution multispectral images, and corresponding real reference images. The training samples are then input in batches into a fusion model consisting of a multi-scale feature extraction module, a frequency-driven structure extraction module, a spectral channel feature extraction module, a dynamic dual-branch fusion module, and an image reconstruction module. The fusion model sequentially executes steps S1 to S4 to output the fused target image and calculates the combined loss function value between the fused target image and the real reference image. The backpropagation algorithm combined with the gradient descent optimizer is used to update the network parameters of all modules in the fusion model layer by layer in reverse based on the combined loss function value. The process of iteratively executing sample forward inference, loss calculation, and parameter update continues until the combined loss function converges to a preset threshold, thus completing the training of the fusion model and obtaining the trained fusion model.
Citation Information
Patent Citations
Hyperspectral and multispectral image fusion method based on attention mechanism
CN117474781A
Hyperspectral image super-resolution reconstruction method based on multi-scale cavity convolution guidance
CN120852162A