Image fusion method based on multi-dimensional complementary features and multi-scale detail enhancement
By using the parallel multi-dimensional complementary fusion network PMCFusion and the high-frequency detail enhancement module HFDEM, the problem of high-frequency detail loss in infrared and visible light image fusion is solved, achieving a clearer texture and higher contrast fusion effect.
Patent Information
- Application Number
- CN202511137834.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-12-12
AI Technical Summary
Existing infrared and visible light image fusion methods have limitations in feature interaction and multimodal information fusion, resulting in the loss of high-frequency details and difficulty in generating fusion results with clear texture and high contrast.
The parallel multi-dimensional complementary fusion network PMCFusion is adopted. Through the parallel three-branch fusion module PTFM and the high-frequency detail enhancement module HFDEM, spatial, channel and frequency domain information are combined for collaborative sensing and efficient integration. A high-frequency loss function is introduced for constraint, and a multi-scale architecture is constructed to improve the fusion quality.
It significantly improves the clarity and detail fidelity of the fused image, overcomes the information loss and high-frequency detail loss problems in existing technologies, and generates higher quality fusion results.
Smart Images

Figure CN121120408A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of infrared and visible light image fusion, and particularly relates to an image fusion method based on multi-dimensional complementary features and multi-scale detail enhancement. BACKGROUND
[0002] As a key image processing technology, the core goal of image fusion is to integrate multi-source image information from different sensors or imaging conditions to generate a single enhanced image containing richer scene interpretation clues. With its unique advantages in information enhancement and complementation, image fusion technology has shown far-reaching application value in many frontier fields such as geological exploration, intelligent security, clinical medical auxiliary diagnosis and robot navigation. Among the many multi-modal fusion tasks, the fusion of infrared and visible light images is particularly eye-catching. Visible light images can capture the fine texture and color levels of the scene in the manner of human visual habits, but their imaging quality is easily disturbed by environmental factors such as light changes, smoke and haze, resulting in serious information loss in adverse conditions. In sharp contrast, infrared images can highlight the target outline stably in low light or even complete darkness by sensing the heat radiation of objects, but their inherent defects are often lack of sufficient background details and delicate texture description. Therefore, how to effectively combine the images with different imaging mechanisms, make full use of the rich details of visible light images and the prominent target characteristics of infrared images to generate a fusion result containing clear texture and high-contrast targets has become the core issue of continuous exploration in this field, and has crucial significance for improving the environmental perception and intelligent analysis capability in complex scenes. Although there are a large number of methods for infrared and visible light image fusion, there are still some deficiencies in the details and textures of the fusion results.
[0003] Limitations of multi-dimensional complementary relationship: Under the impetus of the deep learning wave, various neural network-based infrared and visible image fusion algorithms have emerged, and have significantly outperformed traditional methods. Convolutional neural networks, with their excellent local perception ability, perform well in texture and structure information extraction; generative adversarial networks, by introducing an adversarial mechanism, strive to generate more natural and realistic fusion images. However, a deep analysis of these mainstream deep fusion frameworks reveals that they still have limitations in feature interaction and multi-modal information fusion. Most current models mainly focus on spatial domain and channel domain feature learning and fusion decision-making. Even the research that introduces frequency domain information often uses a simple serial processing method. Such a non-parallel process may cause information loss during the conversion between different domains, or fail to effectively utilize the guidance from another domain during feature extraction in a certain domain, making it difficult to achieve true cross-domain collaboration and feature enhancement. Especially for infrared and visible light modalities, which have significant differences in physical properties and information distribution, how to build a unified framework that can parallelly fuse spatial structure, channel characteristics, and frequency domain representation, fully exploiting their potential cross-dimensional complementarity, has become a key to improving fusion quality and an important direction for current research that needs to be broken through.
[0004] Limitations of image high-frequency detail loss: Current deep learning fusion networks generally have the inherent limitation of image high-frequency detail loss, which is due to the internal defects in their network architecture design. In order to extract high-level semantic features and expand the receptive field, mainstream architectures such as encoder-decoder extensively use pooling, step convolution, and other downsampling operations. These operations, while reducing the size of the feature map, will forcibly discard the spatial information within the local region, causing irreversible damage to high-frequency signals representing object edges and textures. For example, U-Net, which is representative and used by many fusion models. At the same time, the deep stacking of convolutional layers produces a cumulative low-pass filtering effect, systematically attenuating high-frequency components during feature forward propagation, resulting in a "bottleneck effect" in information transmission. More critically, existing technical frameworks generally lack a dedicated, explicit high-frequency compensation mechanism, relying solely on global loss functions for end-to-end implicit optimization. This indirect supervision method is weak and uncontrollable for restoring fine, local high-frequency detail signals. For example, many advanced fusion models such as DDcGAN and SwinFusion. Therefore, these inherent limitations collectively result in the final generated fusion image, although the macro structure is preserved, the micro visual quality is degraded, the edges are blurred, and the texture details are flat, failing to fully retain the rich detail advantages of the source images (especially visible light images), severely affecting the visual fidelity and practical application value of the fusion results. SUMMARY
[0005] The present application aims at the above-mentioned problems existing at present, and provides an image fusion method based on multi-dimensional complementary features and multi-scale detail enhancement, a parallel multi-dimensional complementary fusion network PMCFusion is arranged to realize collaborative perception and efficient integration of spatial features, channel correlation and frequency domain information through a parallel three-branch fusion module PTFM, and the PMCFusion can selectively enhance multi-dimensional features and achieve deep complementation through cross-dimensional attention interaction. In order to further improve the detail definition and structural integrity of the fused image, a high-frequency detail enhancement module HFDEM is integrated into the model, and a corresponding high-frequency loss function is used for constraint. The whole network adopts a multi-scale architecture, which guarantees efficient fusion and robustness of the model under different resolutions. A large number of experimental results show that the method of the present application is significantly better than the current mainstream fusion algorithms in subjective visual effect and objective evaluation index.
[0006] The technical scheme of the present application is as follows: An image fusion method based on multi-dimensional complementary features and multi-scale detail enhancement comprises the following steps: An edge map of the source image is calculated through an edge attention module, and the edge map and the source image are spliced in the channel dimension; The spliced features pass through a preliminary convolution layer and enter a high-frequency detail enhancement module to enhance feature texture and edge details; The enhanced high-resolution features are generated into three feature maps of different scales through successive downsampling operations, and each feature map is subjected to feature depth optimization and refinement through a residual feature distillation module to obtain a feature representation with higher expression capability; From the deepest scale, the infrared and visible light features of this level are fused through a parallel three-branch fusion module; The fused features are used to reconstruct a low-resolution fusion result, and are transmitted to the previous layer through an upsampling operation; in the previous scale, the upsampled features are spliced with the features of the same scale in the encoder path to form new features to be fused, and the spliced features are fused through a parallel three-branch fusion module; the process is repeated step by step until the final full-resolution fusion image is generated.
[0007] Through the above method, the PMCFusion constitutes an end-to-end trainable system which performs deep feature extraction, cross-modal fusion and reconstruction supervision at each scale.
[0008] Further, the parallel three-branch fusion module fully mines and fuses complementary information of multi-modal data from various aspects by introducing a three-branch parallel network and subsequent high-low frequency filtering operations; the parallel three-branch fusion module includes a spatial branch, a channel branch and a frequency domain branch, and the multi-dimensional feature representation is further optimized by high-low frequency filtering through the integration of the outputs of the branches, so as to complete cross-modal information complementation and fusion.
[0009] Further, the spatial branch is a spatial irrelevance branch, which generates spatial attention weights by calculating the difference of feature maps in spatial positions, focusing on and retaining the spatial region with the most information in the modal; the spatial irrelevance branch includes the following steps: Calculate the pixel-by-pixel absolute difference between the input features and to obtain a difference map ; average pooling and maximum pooling are performed on along the channel dimension, and the results are spliced and then passed through a multi-scale convolution spatial attention network to generate the final spatial weight map : ,
[0010] wherein, represents a splicing operation, and represent average pooling and maximum pooling respectively, is a Sigmoid activation function; the weight map can highlight the regions where the two modalities have significant differences in space; the output feature of the spatial branch is obtained by : .
[0011] Further, the channel branch is a channel irrelevance branch, which generates weights by calculating the difference between the two modal features in the channel dimension; specifically including the following steps: An enhanced channel attention module is used to generate a channel attention map for each modality; the correlation between the attention maps is measured by calculating the cosine similarity, and the final channel irrelevance weight is calculated by combining the difference : , , wherein, and are the channel attention maps of the query and key features respectively, represents cosine similarity calculation, represents element-wise multiplication, represents normalization operation; the weight amplifies the channels with significant differences between the two modalities; the output feature of the channel branch is : .
[0012] Further, the frequency domain branch is a frequency domain complementary branch, which directly fuses features in the frequency domain space; and specifically includes the following steps: The input and are converted into the frequency domain by fast Fourier transform to obtain respective amplitude spectrum and phase spectrum; The amplitude spectrum and the phase spectrum of the two modalities are spliced respectively, and two independent convolutional networks are used for fusion to learn and generate fused amplitude spectrum and phase spectrum ; The amplitude spectrum and the phase spectrum of the two different modalities are spliced, and then respectively enter two combinations of convolutional layers with a convolution kernel of 3*3 and LeakyReLU filled with 1, to obtain features and ; The fused frequency domain representation is converted back to the spatial domain by inverse fast Fourier transform to obtain the output feature of the frequency domain branch ; The specific process can be summarized as: , , , Among them, represents a convolutional network for fusing frequency domain spectrum, indicates reconstruction of complex frequency spectrum according to Euler's formula.
[0013] Further, the parallel three-branch fusion module integrates the complementary information of the three parallel branches through a dynamic weight adjustment module; the dynamic weight adjustment module includes the following steps: The original and are taken as inputs, and a small convolutional network is used to dynamically generate weights of the three branches for each spatial position; The fused features are weighted and fused as follows: , The channel is reduced to 64 using a convolution kernel size of 1 and GELU, and then the channel size is reduced to 32 using a convolution kernel size of 3 filled with 1 and BatchNorm : .
[0014] Further, the parallel three-branch fusion module further refines the features, and performs frequency-selective reorganization on the output feature : An independent Gaussian low-pass filter and a Laplacian high-pass filter are used to and decomposed into low-frequency components and high-frequency components ; a network composed of a convolutional layer with a convolution kernel of 3 filled with 1 and a ReLU activation function and a convolutional layer with a convolution kernel of 1 and a Softmax activation function learns the fusion weights of low-frequency and high-frequency according to the input features ; The frequency components from different modalities are fused by weighting, and the final output of the parallel three-branch fusion module is obtained by residual connection with the output of the three branches : , , .
[0015] Through the above method, the PTFM module not only extracts irrelevant and complementary features from three dimensions in parallel, but also realizes the deep integration and optimization of multi-modal information through dynamic weighting and subsequent frequency-selective fusion, providing rich feature representation for generating high-quality fused images.
[0016] Further, the high-frequency detail enhancement module enhances the texture and edge details of the features through residual connection; comprising the following steps: using a set of Laplacian high-pass filters with different sizes on the input feature map , so as to separate the edge and texture information of fine, medium and coarse granularity in parallel: , wherein, represents a high-pass filter with a size of , is the high-frequency feature map under the corresponding scale; the high-frequency feature of each scale is sent into a respective independent convolutional network for enhancement to learn and strengthen the specific detail pattern under that scale: .
[0017] Further, the high-frequency detail enhancement module focuses on the detail scale through an adaptive weight module; the adaptive weight module takes the original input feature as a condition, generates three weights corresponding to the three high-frequency scales through a small convolutional network and a Softmax function; the weighted enhanced features are calculated by the following formula: , , wherein, is a weight generation network, represents element-wise multiplication.
[0018] Further, the high-frequency detail enhancement module enhances the details of the edge region through an edge-preserving attention mechanism; the edge-preserving attention mechanism enhances the details of the edge region through an attention network analyzes the original feature to generate a spatial attention map whose value is higher in the edge region and lower in the smooth region: , wherein is a Sigmoid activation function; The weighted multi-scale enhanced features are spliced and modulated through the attention map, and finally pass through a fusion network composed of a convolution layer with a convolution kernel size of 3 and a padding of 1, a BN normalization layer, a ReLU activation function, and a convolution layer with a convolution kernel size of 1 to obtain the enhanced high-frequency feature : , , The enhanced high-frequency feature is added back to the original input feature to obtain the final output of the module : .
[0019] Through the above method, through the processing of the HFDEM, the network can obtain a feature representation rich in fine details in the early stage, providing high-quality input for subsequent cross-modal deep fusion, thereby significantly improving the clarity and detail fidelity of the final fused image.
[0020] Compared with the prior art, the beneficial effects of the present application are: 1. By leveraging the complementary information from the spatial domain, channel domain, and frequency domain, this approach fundamentally overcomes the "dimensional limitations" of existing fusion methods in the feature fusion stage. It decouples the complex fusion problem from a single spatial domain to three parallel processing branches: channel, spatial, and frequency. The channel branch generates differentiated weights by calculating the difference in response intensity between two modes in each channel; the spatial branch focuses on regions of complementary information by capturing local spatial feature differences at different scales; and as a key innovation, it introduces a frequency branch, generally overlooked in existing technologies. This branch uses Fourier transform to explicitly map features to the frequency domain, enabling lower-level selection and fusion of advantageous information at the amplitude and phase levels. Finally, the outputs of these three parallel branches are not simply added together, but are "arbitrated" by a lightweight dynamic weight network. This network adaptively assigns weights to the results of each branch based on the original input, achieving intelligent weighted fusion. Through this closed-loop design of "parallel decoupling - irrelevance measurement - dynamic arbitration," this module upgrades the originally blind fusion process into a multi-dimensional, adaptive, and intelligent information integration process, thereby maximizing the extraction and preservation of cross-modal complementary information. 2. Utilizing the rich high-frequency information contained in the early stages of feature fusion to enhance later feature information, this approach aims to precisely combat the inherent "high-frequency information loss" problem in feature extraction of deep networks. The design philosophy shifts from "passive recovery" to "active enhancement," constructing a dedicated, explicit "extraction-enhancement-injection" high-frequency compensation pathway. Recognizing the multi-scale characteristics of image details, a filter bank composed of high-pass filters of various sizes is first employed to accurately extract full-spectrum high-frequency components from features in parallel, ranging from fine textures to coarse edges. For these separated multi-scale high-frequency information, a "divide and conquer" strategy is adopted, designing independent convolutional networks for differentiated enhancement at each scale to learn the optimal enhancement strategy for details at different scales. To make the enhancement process more intelligent, two adaptive mechanisms are further introduced: an adaptive enhancement weight module that dynamically determines the fusion ratio of high-frequency components at different scales based on image content; and an edge-preserving attention module that primarily applies the enhancement effect to the edges and textured regions of features, effectively avoiding the amplification of noise in smooth areas. Ultimately, the carefully processed enhanced high-frequency features are directly "injected" back into the backbone feature stream through a residual connection, returning the lost details to the main network in the most direct and powerful way, fundamentally improving the clarity, detail richness, and overall visual fidelity of the final fused image. Attached Figure Description
[0021] Figure 1 This is a diagram illustrating the overall framework of the method described in this application.
[0022] Figure 2 This is a schematic diagram of the HFDEM module in this application.
[0023] Figure 3 Schematic diagram of the PTFM module part of the present application.
[0024] Figure 4 Schematic diagram of the remaining part of the PTFM module of the present application. DETAILED DESCRIPTION
[0025] It should be noted that the relational terms herein, such as first and second, and the like, are used solely to distinguish one from another entity or action without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0026] The features and nature of the present application will become more apparent from the detailed description set forth below, taken in conjunction with the accompanying drawings.
[0027] Referring to Figures 1-4 , an image fusion method based on multi-dimensional complementary features and multi-scale detail enhancement, comprising: The overall network architecture of PMCFusion is shown in Figure 1 , which adopts a carefully designed multi-scale encoder-decoder structure, aiming to comprehensively extract and integrate complementary information from the input infrared image and visible light image . In the encoder (feature extraction) path, the network follows the process from high resolution to low resolution.
[0028] First, an edge attention module is used to calculate the edge map of the source image, which is then concatenated with the original image in the channel dimension to enhance the initial perception of structural information. The concatenated features are processed by a preliminary convolution layer, and then immediately sent to a high-frequency detail enhancement module (HFDEM) for processing. The HFDEM enhances the texture and edge details of the features through a residual connection.
[0029] As shown in Figure 2 , in order to capture detailed information of different granularities, the HFDEM first uses a multi-scale strategy to extract high-frequency components. A set of Laplacian high-pass filters (HPF) with different sizes (3x3, 5x5, 7x7) are used to act on the input feature map , thus separating the edge and texture information of fine, medium and coarse granularity in parallel. This process can be represented as: , where, represents a high-pass filter of size k x k, is the high-frequency feature map at the corresponding scale. In this way, the module can comprehensively perceive various high-frequency information from fine texture to prominent contour.
[0030] The high-frequency features extracted at different scales have different physical meanings, and therefore need to be processed specifically. The high-frequency features at each scale are fed into a separate convolutional network for enhancement, to learn and strengthen the specific detail patterns at that scale: , To enable the module to dynamically focus on the most important detail scale according to the content of the input image, an adaptive weight module is set up. This module, conditioned on the original input feature , generates three weights , corresponding to the three high-frequency scales, through a small convolutional network and a Softmax function. This allows the network to adaptively allocate the contribution of high-frequency information at different scales. The weighted enhanced features are calculated by: , , where, is the weight generation network, represents element-wise multiplication.
[0031] Before fusing the multi-scale high-frequency information, to ensure that the enhancement process does not introduce noise in smooth areas, while further strengthening the details in edge areas, we introduce an edge-preserving attention mechanism. This mechanism analyzes the original feature through an attention network , generating a spatial attention map whose value is higher in edge areas and lower in smooth areas: , where is the Sigmoid activation function. Subsequently, the weighted multi-scale enhanced features are concatenated and modulated by the attention map, and finally pass through a fusion network composed of a convolution layer with a kernel size of 3 and a padding of 1, a BN normalization layer, a ReLU activation function, and a convolution layer with a kernel size of 1, to obtain the enhanced high-frequency features : , , Finally, following the idea of residual learning, the enhanced high-frequency features are added back to the original input features to obtain the final output of the module . This additive fusion ensures that the module retains the original low-frequency information while enhancing the high-frequency details.
[0032] , Through the processing of HFDEM, the network can obtain feature representations rich in fine details at an early stage, providing high-quality inputs for subsequent cross-modal deep fusion, thereby significantly improving the clarity and detail fidelity of the final fused image.
[0033] The high-resolution features enhanced by HFDEM will be progressively generated into three feature maps of different scales (level 1, level 2, level 3) through successive downsampling operations. At each scale, a residual feature distillation module (RFDB) is used to optimize and refine the features at that level to obtain more expressive feature representations. In the decoder (feature fusion and reconstruction) path, the network follows a "bottom-up" process from low resolution to high resolution. The fusion process starts from the deepest scale (level 3), where the infrared and visible light features of this level are fed into the first parallel three-branch fusion module (PTFM) for fusion.
[0034] This module is not simply a stack of components, but a complete and precise fusion unit designed to extract and integrate complementary information from multiple dimensions through a coherent process. Unlike strategies that focus on a single dimension or use serial processing, PTFM introduces a three-branch parallel network and subsequent high-low frequency filtering operations to fully extract and fuse complementary information from various aspects of multi-modal data. This unit not only focuses on feature modeling within a single modality, but also emphasizes information interaction and selective enhancement between cross-modalities, thereby improving the network's performance in extracting and fusing discriminative information from multiple modalities and dimensions.
[0035] The overall structure of PTFM and the interaction between its key components are shown in Figure 3 and Figure 4 . PTFM contains three core parallel branches: spatial branch, channel branch, and frequency domain branch, and how the outputs of these branches are integrated and further optimized through subsequent high-low frequency filtering to achieve deep complementarity and efficient fusion of cross-modal information.
[0036] Information in spatial dimensions, such as edges, contours, and textures, is the key to image fusion. Infrared images are good at capturing the prominent contours of targets, while visible images are rich in fine texture details. To adaptively preserve these complementary spatial information, a spatial irrelevance branch is set up. This branch generates spatial attention weights by calculating the difference in spatial positions of feature maps, guiding the model to focus on and preserve the most informative spatial regions in each modality.
[0037] First, the pixel-wise absolute difference between input features and is calculated to obtain the difference map . Then, the average pooling and max pooling are performed along the channel dimension, and the results are concatenated and sent to a spatial attention network containing multi-scale convolution (using 3x3 and 7x7 convolution kernels in parallel) to generate the final spatial weight map :
[0038] where denotes the concatenation operation, and represent average pooling and max pooling, respectively, is the Sigmoid activation function. This weight map can highlight the regions where the two modalities have significant differences in space. The output feature of the spatial branch is obtained by: Feature responses in the channel dimension represent different abstract semantic information of the image. In infrared and visible light images, some channels may have stronger responses to specific information (such as heat radiation or texture details). To highlight the unique features of each modality, a channel irrelevance branch is set up. This branch generates weights by calculating the difference in channel dimension between the features of the two modalities, thereby enhancing those features that are prominent in one modality but not in the other.
[0039] First, an enhanced channel attention module (ECA) is used to generate a channel attention map for each modality. Then, the cosine similarity between the attention maps is calculated to measure their relevance, and the final channel irrelevance weight is calculated by combining their differences: where, and are the channel attention maps of query and key features respectively, represents cosine similarity calculation, denotes element-wise multiplication, denotes normalization operation. The weight will amplify the channels that have significant differences between two modalities. Finally, the output feature of the channel branch is defined as: , Through the above operations, the model can pay more attention to and enhance the channel features in that show greater complementarity with in the channel response.
[0040] Unlike seeking differences in spatial and channel domains, the frequency domain provides another unique perspective for feature fusion. The low-frequency components of an image usually represent its overall contour and background, while the high-frequency components contain edge and detail information. The information distribution of infrared and visible light images in different frequency bands has natural complementarity.
[0041] The frequency domain complementary branch aims to directly fuse features in the frequency domain space. The input and are converted to the frequency domain by fast Fourier transform (FFT) to obtain their respective amplitude spectrum and phase spectrum. Subsequently, the amplitude spectrum and phase spectrum of the two modalities are spliced respectively, and two independent convolutional networks are used for fusion to learn the fused amplitude spectrum and phase spectrum . Then the amplitude spectrum and phase spectrum of the two different modalities are spliced, and then enter two combinations of convolution layers with a convolution kernel of 3x3 and LeakyReLU filled with 1, respectively, to obtain features and . Finally, the inverse fast Fourier transform (IFFT) is used to convert the fused frequency domain representation back to the spatial domain to obtain the output feature of the frequency domain branch. The whole process can be summarized as: , , , where, represents the convolutional network used to fuse the frequency domain spectrum, denotes the reconstruction of complex frequency spectrum according to Euler's formula.
[0042] In order to optimally integrate the complementary information from the three parallel branches, a dynamic weight adjustment module is introduced. This module will adjust the original and As input, weights for each spatial location are dynamically generated for the three branches (spatial, channel, frequency) by a small convolutional network . This enables the network to adaptively determine the contribution of each branch according to the characteristics of the local region. The weighted fused features are: , Then, in order to prevent information loss caused by directly reducing the dimension to 32, the channel is first reduced to 64 using a convolution kernel size of 1 and GELU, and then a convolution kernel size of 3 is filled with 1 and BatchNorm to obtain an output feature with a channel size of 32 : , On this basis, in order to further refine the features, the PTFM module performs frequency-selective reorganization on . Using independent Gaussian low-pass filters and Laplacian high-pass filters, the and are decomposed into low-frequency components and high-frequency components . At the same time, another network composed of a convolution layer with a convolution kernel size of 3 filled with 1 and a ReLU activation function and a convolution layer with a convolution kernel size of 1 and a Softmax activation function learns the fusion weights of low-frequency and high-frequency from the input features . Finally, by weighting and fusing the frequency components from different modalities, and performing residual connection with the output of the three branches, the final output of the PTFM is obtained : , , , In this way, the PTFM module not only extracts irrelevant and complementary features from three dimensions in parallel, but also realizes the deep integration and optimization of multi-modal information through dynamic weighting and subsequent frequency-selective fusion, providing rich feature representation for generating high-quality fused images.
[0043] Then, the fused features of this layer are used to reconstruct the low-resolution fusion result on the one hand, and are passed to the previous layer (level 2) through the up-sampling operation on the other hand. In level 2, the up-sampled features will be concatenated (skip-connection) with the same scale features from the encoder path to form new features to be fused. These concatenated features are then sent to the PTFM module of the current scale for fusion. This "up-sampling- concatenation-fusion" process is repeated level by level in the decoder path until the final full-resolution fused image is generated In this way, the PMCFusion constitutes an end-to-end trainable system that performs deep feature extraction, cross-modal fusion and reconstruction supervision at each scale.
[0044] The above-described embodiments are merely illustrative of the present application and do not limit the scope of the present application. It should be pointed out that, for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application.
Claims
1. An image fusion method based on multi-dimension complementary features and multi-scale details enhancement, characterized in that, The method comprises the following steps: An edge map of the source image is calculated by an edge attention module, and the edge map is spliced with the source image in a channel dimension; The spliced features pass through a preliminary convolution layer to enter a high-frequency detail enhancement module to enhance feature texture and edge details; The enhanced high-resolution features pass through successive downsampling operations to generate three feature maps of different scales, and each feature map is optimized and refined in feature depth by a residual feature distillation module to obtain a more expressive feature representation; Starting from the deepest scale, the infrared and visible light features of this level are fused by a parallel three-branch fusion module; The fused features are used to reconstruct a low-resolution fusion result and are passed to the previous layer through upsampling operations; in the previous layer scale, the upsampled features are spliced with the same scale features of the encoder path to form new features to be fused, and the spliced features are fused by a parallel three-branch fusion module; the process is repeated step by step until the final full-resolution fusion image is generated. 2.The image fusion method based on multi-dimension complementary features and multi-scale details enhancement of claim 1, characterized in that, The parallel three-branch fusion module fully excavates and fuses the complementary information of multi-modal data from various aspects by introducing a three-branch parallel network and subsequent high-low frequency filtering operations; the parallel three-branch fusion module includes a spatial branch, a channel branch, and a frequency domain branch, which further optimizes multi-dimensional feature representation through high-low frequency filtering of the outputs of the branches, and completes cross-modal information complementation and fusion. 3.The image fusion method based on multi-dimension complementary features and multi-scale details enhancement of claim 2, characterized in that, The spatial branch is a spatial irrelevance branch that generates spatial attention weights by calculating the differences of the feature maps in the spatial position, and focuses on and retains the spatial regions with the most information in the modal; the spatial irrelevance branch comprises the following steps: Calculate input features and The difference map is obtained by taking the absolute difference between each pixel. ;Will Average pooling and max pooling are performed along the channel dimension, and the results are concatenated and then passed through a multi-scale convolutional spatial attention network to generate the final spatial weight map. : , wherein, denotes concatenation operation, and represent average pooling and max pooling, respectively, is a sigmoid activation function; weight map can highlight the regions where the two modalities have significant differences in space; output features of the spatial branch is obtained by the following formula: 。 4. The image fusion method based on multi-dimension complementary features and multi-scale details enhancement according to claim 3, characterized in that, The channel branch is a channel irrelevance branch that generates weights by calculating the differences of the two modal features in the channel dimension; Specifically, the method comprises the following steps: The enhanced channel attention module is used to generate a channel attention map for each modality; the correlation is measured by calculating the cosine similarity between the attention maps, and the final channel irrelevance weight is calculated by combining the difference : , , in, and These are the channel attention maps for the query and key features, respectively. Represents cosine similarity calculation. This represents element-wise multiplication. Indicates the normalization operation; weight It amplifies channels with significant differences between the two modes; the output characteristics of channel branches. for: 。 5. The image fusion method based on multi-dimension complementary features and multi-scale details enhancement according to claim 4, characterized in that, The frequency domain branch is a frequency domain complementary branch that directly fuses features in the frequency domain space; specifically, the method comprises the following steps: The input is and Transformed into the frequency domain by fast Fourier transform, the respective amplitude spectrum and phase spectrum are obtained; The amplitude spectrum and the phase spectrum of the two modalities are spliced respectively, and two independent convolution networks are used for fusion to learn and generate a fused amplitude spectrum and a phase spectrum The amplitude spectrum and the phase spectrum of two different modalities are spliced, and then respectively enter a combination of a convolutional layer with a convolution kernel of 3x3 filled with 1 and LeakyReLU to obtain features and ; The fused frequency domain representation is converted back to the spatial domain by an inverse fast Fourier transform to obtain the output features of the frequency domain branch The specific process can be summarized as follows: , , , wherein, represents a convolutional network for fusing frequency domain spectra, denotes the reconstruction of complex spectra according to Euler's formula.
6. The image fusion method based on multi-dimension complementary features and multi-scale details enhancement according to claim 5, characterized in that, The parallel three-branch fusion module integrates the complementary information of the three parallel branches through a dynamic weight adjustment module; the dynamic weight adjustment module comprises the following steps: The original and As input, weights for the three branches are dynamically generated for each spatial location by a small convolutional network ; Weighted fused features is: , Using a convolution kernel size of 1 and GELU to reduce the channels to 64, then using a convolution kernel size of 3 with padding of 1 and BatchNorm to get an output feature with a channel size of 32 : 。 7. The image fusion method based on multi-dimension complementary features and multi-scale details enhancement according to claim 6, characterized in that, The parallel three-branch fusion module is used for further refining features, and the output features are frequency-selectively reorganized: using independent Gaussian low-pass filters and Laplacian high-pass filters, the image is decomposed into low-frequency components and and high-frequency components and ; A network composed of a convolutional layer with a convolution kernel of 3 filled with 1 and a ReLU activation function and a convolutional layer with a convolution kernel of 1 and a Softmax activation function learns the fusion weights of low frequency and high frequency according to the input features ; The final output of the parallel three-branch fusion module is obtained by weighting and fusing the frequency components from different modalities and performing residual connection with the outputs of the three branches : , , 。 8.The image fusion method based on multi-dimension complementary features and multi-scale details enhancement of claim 1, characterized in that, The high-frequency detail enhancement module enhances the texture and edge details of the features through a residual connection; the method comprises the following steps: applying a set of laplacian high-pass filters with different sizes to the input feature maps thereby separating fine, medium and coarse edge and texture information in parallel: , wherein, represents a high-pass filter with a size of , is a high-frequency feature map at a corresponding scale; the high-frequency features at each scale are fed into a respective independent convolutional network for enhancement to learn and strengthen specific detail patterns at that scale: 。 9. The image fusion method based on multi-dimension complementary features and multi-scale details enhancement according to claim 8, characterized in that, In the high-frequency detail enhancement module, an adaptive weight module focuses on the detail scale; In the high-frequency detail enhancement module, an adaptive weight module focuses on the detail scale; The adaptive weight module takes the original input features as conditions, generates three weights through a small convolutional network and a Softmax function , respectively corresponding to three high-frequency scales; the weighted enhanced features are calculated by the following formula: , , wherein, is a weight generation network, represents an element-wise multiplication.
10. The image fusion method based on multi-dimension complementary features and multi-scale details enhancement according to claim 9, characterized in that, The high-frequency detail enhancement module enhances the details of the edge region through an edge-preserving attention mechanism; the edge-preserving attention mechanism passes an attention network analyzing the original features to generate a spatial attention map whose value is higher in the edge region and lower in the smooth region: , wherein is a Sigmoid activation function; The weighted multi-scale enhanced features are spliced and modulated through the attention map, and finally pass through a fusion network composed of a convolution layer with a convolution kernel size of 3 and padding of 1, a BN normalization layer, a ReLU activation function, and a convolution layer with a convolution kernel size of 1 obtain enhanced high-frequency features : , , The enhanced high-frequency features are added back to the original input features , to obtain the final output of the module : 。
Citation Information
Cited By
Multi-modal image fusion method and system based on dynamic frequency domain fusion
CN122415356A
A Multimodal Image Fusion Method and System Based on Dynamic Frequency Domain Fusion
CN122415356B