Prior knowledge guided dual-domain feature network multi-modal image fusion method and system

CN122222836BActive Publication Date: 2026-09-15XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610706361.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-09-15
Estimated Expiration
2046-05-21

AI Technical Summary

Benefits of technology

[0053] 1. The prior knowledge-guided dual-domain feature network multimodal image fusion method of this invention introduces prior knowledge to guide feature learning, performs shallow feature extraction and fusion, strengthens the structural alignment and consistency between cross-modal images, and lays a solid foundation for subsequent deep processing. Subsequently, the fusion process is decoupled, and the complementary advantages of the spatial and frequency domain branches are synergistically utilized to achieve balanced optimization of structure and details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122222836B_ABST
    Figure CN122222836B_ABST
Patent Text Reader

Abstract

The present application relates to image fusion method, specifically to prior knowledge guided dual-domain feature network multi-modal image fusion method and system. In order to solve the deficiency that the deep learning model in the prior art is not enough for the specific extraction of high frequency details or does not effectively utilize the inherent prior structure information between cross modal, resulting in the deficiency that the structural consistency of the fused image is damaged, the present application constructs the space-frequency dual branch cross modal fusion network PGDFusion by prior knowledge guided dual-domain feature network multi-modal image fusion method for fusing infrared image and visible light image, PGDFusion learns features guided by prior knowledge, extracts and fuses shallow features, then obtains space domain features by fully modeling cross scale space structure features through the space domain branch, obtains frequency domain features by acquiring high frequency texture details and structure edge information through the frequency domain branch, and finally reconstructs the final fusion image based on the space domain features and the frequency domain features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to image fusion methods, specifically to a prior knowledge-guided dual-domain feature network multimodal image fusion method and system. Background Technology

[0002] With the rapid development of multimedia sensor technology, multimodal image acquisition equipment has been widely used in fields such as security monitoring, autonomous driving, medical diagnosis, and remote sensing. Due to differences in imaging mechanisms, different modal sensors often have unique imaging characteristics and limitations. For example, visible light sensors can capture rich texture details and color information, but are highly susceptible to lighting conditions and inclement weather; while infrared sensors, based on thermal radiation imaging, can highlight prominent targets in all weather conditions, but often lack clear background structure and texture details. Multimodal image fusion aims to effectively integrate complementary information from different source image pairs into a single image, enabling it to simultaneously possess high-contrast target features and high-fidelity background details, thereby providing richer and more accurate data support for subsequent advanced vision tasks (such as object detection and semantic segmentation).

[0003] Despite significant progress in multimodal image fusion research, achieving both high-fidelity structural reconstruction and refined detail preservation within a unified framework remains a core challenge. Traditional fusion methods often rely on hand-designed feature extraction and fusion rules, which have limited representational capabilities and struggle to model complex cross-modal relationships. This often leads to issues such as texture blurring, edge artifacts, or contrast imbalances in the fusion results, making it difficult to balance structural fidelity with the fine preservation of high-frequency details. In recent years, deep learning methods, represented by convolutional neural networks, have significantly improved fusion performance through their powerful feature learning capabilities. However, many existing deep learning models either focus on global or local feature modeling in the spatial domain, neglecting targeted extraction of high-frequency details, or fail to effectively utilize inherent prior structural information across modalities, resulting in compromised structural consistency in the fused image. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing deep learning models in extracting high-frequency details effectively or failing to utilize the inherent prior structural information across modalities, which leads to impaired structural consistency of fused images. This invention provides a prior knowledge-guided method and system for multimodal image fusion using dual-domain feature networks.

[0005] The inventive concept of this invention is as follows: To solve the aforementioned problems, this invention establishes a prior knowledge-guided spatial-frequency dual-branch cross-modal fusion network. First, prior knowledge is used to constrain the extraction of shallow cross-modal features, which are then fused with these shallow features to build a robust foundation for structural consistency. Subsequently, the processing flows in the spatial and frequency domains are decoupled: the spatial branch utilizes a Unet integrating global and local receptive fields to accurately capture spatial dependencies between long and short distances; the frequency branch introduces a learnable Butterworth high-frequency filtering dense network, dynamically adjusting filter parameters to selectively extract and retain fine-grained high-frequency texture details. Finally, a high-quality fused image is reconstructed through the fusion decoding of spatial and frequency features. Extensive experiments demonstrate that this invention outperforms traditional methods in multiple objective metrics and subjective visual perception.

[0006] To achieve the above objectives, the technical solution provided by this invention is as follows:

[0007] A prior knowledge-guided multimodal image fusion method using a dual-domain feature network, characterized by the following steps:

[0008] Step 1, acquire multiple sets of infrared images I ir and visible light image I vi As source image pairs, as training sets, each group of infrared images I ir and visible light image I vi Targeting the same objective;

[0009] Step 2: Construct the space-frequency dual-branch cross-modal fusion network PGDFusion, which includes a cross-modal encoder, a spatial encoder, a frequency encoder, and a space-frequency fusion decoder.

[0010] Step 3: Input the training set into PGDFusion, and extract the infrared image I using the cross-modal encoder. ir Visible light image I vi and prior knowledge The shallow features are then fused to obtain the initial fused feature C;

[0011] Step 4: The spatial encoder acquires the contour, region brightness structure, and local texture information from the initial fused feature C to obtain the spatial feature. The frequency domain encoder transforms the initial fused feature C into the frequency domain, performs high-frequency enhancement, and obtains the frequency domain feature. ;

[0012] Step 5: Spatial-frequency domain fusion decoder fuses spatial features. Frequency domain characteristics The final fused image is reconstructed and output.

[0013] Step 6: Update the PGDFusion network parameters by backpropagating the loss function, return to step 3, and continue until the loss function reaches a local minimum to obtain the trained PGDFusion.

[0014] Step 7: Input the infrared image and visible light image to be fused into PGDFusion, process them to obtain the final fused image. At this time, the prior knowledge takes the maximum value of the corresponding source image pair.

[0015] Furthermore, in step 4, the spatial encoder acquires the contour, region brightness structure, and local texture information from the initial fusion feature C to obtain the spatial feature. Specifically:

[0016] a1, the initial fused feature C is alternately passed through multiple levels of hybrid Transformer-convolution modules and multiple levels of downsampling to output deep feature L3;

[0017] The specific process of the hybrid Transformer-convolutional module is as follows: the input features are processed through local branches and global branches, and the corresponding features are output respectively. The features are then fused and output.

[0018] a2, after downsampling the deep feature L3, multi-scale feature enhancement is performed to obtain the enhanced deep feature;

[0019] a3, the enhanced deep features are alternately processed through multiple levels of upsampling and multiple levels of hybrid Transformer-convolution modules to obtain spatial domain features. In this process, the output features of each upsampled level are fused with the features obtained from the downsampled level at the corresponding level before entering the hybrid Transformer-convolution module.

[0020] The number of upsampling and downsampling levels are the same.

[0021] Furthermore, step a1 specifically includes the following steps:

[0022] a11. Input the initial fused feature C, and after passing through the hybrid Transformer-convolution module l1, obtain the shallow feature L1;

[0023] a12. Downsample the shallow feature L1 by d1, and then pass it through the hybrid Transformer-convolution module l2 to obtain feature L2;

[0024] a13. Subtract feature L2 from feature L2 and then pass it through a hybrid Transformer-convolution module l3 to obtain deep feature L3.

[0025] Furthermore, step a3 specifically includes the following steps:

[0026] a31. Upsample the enhanced deep features u3 to obtain features U3 that match the L3 scale of the deep features;

[0027] a32. The feature U3 is concatenated with the deep feature L3 located in the same layer during downsampling, and then processed by the hybrid Transformer-convolution module l4 after fusion.

[0028] a33. Upsample the features output by the hybrid Transformer-convolution module l4 by u2 to obtain feature U2, concatenate it with feature L2, and then process it through the hybrid Transformer-convolution module l5;

[0029] a34. The features output by the hybrid Transformer-convolutional module l5 are upsampled by u1 to obtain feature U1, which is then concatenated with the shallow feature L1 and passed through the hybrid Transformer-convolutional module l6 to output the spatial domain feature. .

[0030] Furthermore, in step a1, the downsampling employs the PixelUnShuffle operation;

[0031] In step a3, the upsampling is performed using the PixelShuffle operation;

[0032] Step 5 specifically involves: [Illegible text - likely related to spatial features] Frequency domain characteristics Fusion is achieved through channel splicing:

[0033] ,

[0034] in, Indicates splicing along the channel. Features after splicing;

[0035] The final fused image is then output through a hybrid decoder consisting of alternating Transformer and convolutional layers.

[0036] Furthermore, in step 4, the frequency domain encoder transforms the initial fused feature C to the frequency domain, performs high-frequency enhancement, and obtains the frequency domain feature. Specifically:

[0037] b1, transform the initial fused feature C into the frequency domain to obtain the intermediate frequency domain feature. ;

[0038] b2, using a Butterworth high-pass modulation filter to analyze the intermediate frequency domain characteristics Filtering is performed to obtain the filtered frequency domain characteristics. ;

[0039] b3, using inverse Fourier transform to extract the frequency domain characteristics of the filter. Reverting to the spatial domain yields the initial frequency domain features. ;

[0040] b4, initial frequency domain features Enhancement is performed sequentially through two serial enhancement paths, and the frequency domain enhancement features F are output respectively. T1 and frequency domain enhancement features F T2 ;

[0041] b5, the initial fusion feature C and the frequency domain enhancement feature F T1 and F T2 The data is then spliced ​​together and restored to the original number of channels to output the frequency domain feature F. f .

[0042] Furthermore, in step b1, the initial fused feature C undergoes channel adjustment via 1×1 convolution, followed by a fast Fourier transform to obtain the intermediate frequency domain feature. ;

[0043] Step b4 specifically involves processing the initial frequency domain features. Feature extraction is performed sequentially through two convolutional transformers: initial frequency domain features. After one 3×3 convolution layer, followed by a Transformer module, the frequency domain enhanced feature F is obtained. T1 The frequency domain enhancement feature FT1 is processed through a 3×3 convolution layer, and then through a Transformer module to obtain the frequency domain enhancement feature F. T2 ;

[0044] In step b5, restoring to the original number of channels specifically involves performing channel integration and dimensionality reduction through 1×1 convolution to restore to the original number of channels.

[0045] Furthermore, after step 1, the process also includes processing the infrared image I. ir and visible light image I vi Preprocessing is performed; the preprocessing includes normalization and image cropping.

[0046] In step 3, the candidate fusion image with the best performance is determined according to a comprehensive evaluation function. The comprehensive evaluation function integrates multiple complementary objective indicators, including structural similarity loss that measures structural fidelity and / or mutual information that measures information content.

[0047] In step 3, the prior knowledge mentioned during training Obtain it through the following methods:

[0048] Infrared and visible light images of the same target are acquired and fused using several different existing image fusion methods to obtain multiple candidate fused images. The candidate fused image with the best performance is selected as prior knowledge.

[0049] In step 7, the infrared image and the visible light image to be fused are input into PGDFusion after the preprocessing.

[0050] This invention also provides a prior knowledge-guided dual-domain feature network multimodal image fusion system, characterized in that it includes a cross-modal encoder, a spatial encoder, a frequency encoder, and a spatial-frequency fusion decoder; the cross-modal encoder is used to extract infrared image I. ir Visible light image I vi and prior knowledge The shallow features are collected and fused to obtain the initial fused feature C. The input of the spatial encoder is connected to the output of the cross-modal encoder and is used to collect the contour, regional brightness structure and local texture information in the initial fused feature C. The input of the frequency encoder is connected to the output of the cross-modal encoder and is used to convert the initial fused feature C to the frequency domain for high-frequency enhancement. The input of the spatial-frequency fusion decoder is connected to the outputs of the spatial encoder and the frequency encoder respectively and is used to fuse the spatial features and frequency features and reconstruct the final fused image.

[0051] Meanwhile, the present invention also provides a computer program product, including a computer program, which is characterized in that: when the program is executed by a processor, it implements the steps of the above-mentioned prior knowledge-guided dual-domain feature network multimodal image fusion method.

[0052] Compared with the prior art, the present invention has the following beneficial technical effects:

[0053] 1. The prior knowledge-guided dual-domain feature network multimodal image fusion method of this invention introduces prior knowledge to guide feature learning, performs shallow feature extraction and fusion, strengthens the structural alignment and consistency between cross-modal images, and lays a solid foundation for subsequent deep processing. Subsequently, the fusion process is decoupled, and the complementary advantages of the spatial and frequency domain branches are synergistically utilized to achieve balanced optimization of structure and details.

[0054] 2. The prior knowledge-guided dual-domain feature network multimodal image fusion method of this invention adopts a UNet with global-local receptive fields in the spatial domain branch, including downsampling modules, upsampling modules, and multi-level hybrid Transformer-convolutional modules, to comprehensively capture the long and short distance spatial dependencies from local details to global context in the image, ensuring the coherence and integrity of the main structure.

[0055] 3. The prior knowledge-guided dual-domain feature network multimodal image fusion method of this invention introduces a learnable Butterworth high-pass modulation filter in the frequency domain branch for the first time. It can dynamically and adaptively extract and enhance fine-grained high-frequency components in the source image pair, thereby focusing on solving the problem of easy loss of details.

[0056] 4. The present invention also provides a computer program product capable of executing the above method steps, which can extend and apply the method of the present invention to realize multimodal image fusion on corresponding hardware devices. Attached Figure Description

[0057] Figure 1 This is a flowchart of an embodiment of the prior knowledge-guided dual-domain feature network multimodal image fusion method of the present invention;

[0058] Figure 2 This is a schematic diagram of the structure of PGDFusion in an embodiment of the prior knowledge-guided dual-domain feature network multimodal image fusion method of the present invention;

[0059] Figure 3 This is a flowchart of the cross-modal encoder in an embodiment of the prior knowledge-guided dual-domain feature network multimodal image fusion method of the present invention;

[0060] Figure 4 This is a schematic diagram of the spatial encoder structure in an embodiment of the prior knowledge-guided dual-domain feature network multimodal image fusion method of the present invention;

[0061] Figure 5 This is a schematic diagram of the frequency domain encoder in an embodiment of the prior knowledge-guided dual-domain feature network multimodal image fusion method of the present invention;

[0062] Figure 6 This is a schematic diagram of the spatial-frequency domain fusion decoder in an embodiment of the prior knowledge-guided dual-domain feature network multimodal image fusion method of the present invention. Detailed Implementation

[0063] This invention establishes a space-frequency dual-branch cross-modal fusion network (PGDFusion) guided by prior knowledge, such as... Figure 2 As shown, its core idea lies in: guiding feature learning by introducing prior knowledge and synergistically utilizing the complementary advantages of the spatial and frequency domain branches to achieve balanced optimization of structure and details. Based on the above network, this invention provides a cross-modal spatial-frequency joint image fusion system—a prior knowledge-guided dual-domain feature network multimodal image fusion system. This system consists of four key components: a cross-modal encoder, a spatial encoder, a frequency encoder, and a spatial-frequency fusion decoder.

[0064] In this invention, the prior knowledge-guided dual-domain feature network multimodal image fusion method first utilizes a CNN-Transformer-based cross-modal encoder to extract images from visible light images (I... vi Infrared images (I) ir ) and prior knowledge (I pr The method extracts complementary shallow features and fuses them to obtain initial fused features. Then, the extracted initial fused features are fed to the spatial domain encoder and frequency domain encoder respectively to capture cross-scale spatial structure features and adaptive frequency texture. Finally, the decoder jointly models the spatial and frequency domain features to generate a final fused image that combines brightness structure, texture details, and thermal features. Based on the above system, the prior knowledge-guided dual-domain feature network multimodal image fusion method of this invention mainly includes the following:

[0065] (1) Acquisition of prior knowledge;

[0066] To construct reliable and informative prior knowledge to guide the fusion process, this invention employs a prior generation strategy based on multi-method evaluation and selection. This method utilizes several existing representative fusion methods to generate candidate fused images, quantitatively evaluates them through an objective evaluation system, and then filters and aggregates the relatively optimal results in terms of structure, detail, and visual perception, using these as prior knowledge to guide subsequent network learning.

[0067] The multimodal images of this invention include infrared images and visible light images. First, a set of infrared and visible light images targeting the same target are acquired. Then, n representative image fusion methods with different fusion principles are selected, where n≥3. The infrared and visible light images are fused separately to obtain t candidate fused images. Where t ≥ 3, and i is the index of t candidate fusion images, taking values ​​from 1 to t. The image fusion method can be DIDFuse, CDDFuse, PIAFusion, DDFM, etc. Subsequently, a comprehensive evaluation function is defined. This function integrates multiple complementary objective metrics, including Structural Similarity Loss (SSIM), which measures structural fidelity, and Mutual Information (MI), which measures information content, through a comprehensive evaluation function. Each candidate image is given a comprehensive score. Finally, the candidate fused image with the highest score is selected as the prior knowledge for that sample. ,Right now:

[0068] ,

[0069] Here, argmax is used to find the maximum value.

[0070] The method described above produces a candidate fusion image that performs best among existing methods. It implicitly encodes a fusion decision that is relatively balanced between structure and detail for the current scene.

[0071] (2) Cross-modal encoder;

[0072] Significant differences exist between multimodal images resulting from different imaging mechanisms: visible light images contain rich texture details, while infrared images highlight thermal radiation information. To fully utilize the complementarity between the two, this invention designs a cross-modal encoder that integrates convolution and Transformer as a shallow feature extraction and fusion module, used to extract features from visible light images I... vi Infrared image I ir and prior knowledge I pr Extracting a globally consistent low-level representation, where the visible light image I vi Infrared image I ir The source image is a pair of images. Before entering the cross-modal encoder for processing, the visible light image and the infrared image are preprocessed separately. In this embodiment, the preprocessing includes normalization and image cropping. In other embodiments, other preprocessing methods may also be used.

[0073] A cross-modal encoder consists of multiple levels of convolutional layers and a Transformer module arranged sequentially. Each modality is processed sequentially through these multiple levels of convolutional layers and the Transformer module, such as... Figure 3 As shown in the diagram. Multi-level convolutional layers are used to capture local edge and texture features, resulting in shallow features. The Transformer module then uses a multi-head self-attention mechanism to establish long-range dependencies and achieves stronger cross-modal correlation modeling through a cross-modal encoder. It encodes the shallow features and outputs the modally encoded shallow features, as shown below:

[0074] ,

[0075] in, For cross-modal encoders, Shallow features encoded from infrared images. Shallow features encoded from visible light images. Shallow features are encoded from prior knowledge.

[0076] To further obtain a unified shallow feature representation, the Transformer module uses a learnable weighting method to perform trimodal fusion on the shallow features encoded by the above modalities, outputting the initial fused feature C:

[0077] ,

[0078] In the formula, Visible light image I vi Infrared image I ir and prior knowledge I pr The fusion weights, in this embodiment, are set to values ​​of 0.5, 0.5, and 1, respectively, to adjust the importance of each mode. The initial fusion feature C serves as the input to the subsequent spatial domain encoder and frequency domain encoder.

[0079] (3) Spatial encoder;

[0080] Spatial branch settings spatial encoder Spatial encoder A UNet architecture with a global-local receptive field is adopted to fully model cross-scale spatial structural features, especially the contours, regional brightness structures, and local texture information of the initial fused feature C. The spatial encoder includes downsampling modules, upsampling modules, and multi-level hybrid Transformer-convolutional modules, such as... Figure 4 As shown. The spatial structure features across scales, i.e., spatial features, output by the spatial encoder. Recorded as:

[0081] ,

[0082] The downsampling module and upsampling module are configured with multi-level downsampling and multi-level upsampling respectively. The processing of the spatial encoder mainly includes top-down downsampling and bottom-up upsampling, specifically including the following steps:

[0083] 3.1. Input the initial fused feature C, and after passing through the hybrid Transformer-convolution module l1, obtain the shallow feature L1.

[0084] 3.2. The shallow feature L1 is downsampled by d1 and then passed through the hybrid Transformer-convolution module l2 to obtain feature L2.

[0085] 3.3. Feature L2 is obtained by downsampling d2 and hybrid Transformer-convolution module l3 to obtain deep feature L3.

[0086] 3.4. After downsampling the deep feature L3 by d3, multi-scale feature enhancement is performed using the SPPF module to obtain the enhanced deep feature. Then, subsequent steps upsample the enhanced deep feature from bottom to top.

[0087] 3.5. Upsample the enhanced deep features u3 to obtain features U3 that match the scale of the deep features L3.

[0088] 3.6. The feature U3 is concatenated with the deep feature L3 located in the same layer during downsampling, and then fused and passed through the hybrid Transformer-convolution module l4.

[0089] 3.7. Upsample the features output by the hybrid Transformer-convolution module l4 by u2 to obtain feature U2, concatenate it with feature L2, and then process it through the hybrid Transformer-convolution module l5.

[0090] 3.8. The features output by the hybrid Transformer-convolution module l5 are upsampled by u1 to obtain feature U1, which is then concatenated with the shallow feature L1 and passed through the hybrid Transformer-convolution module l6 to output the spatial domain feature. .

[0091] To fully model the spatial structure information of the input image, this invention sets up a hybrid Transformer-convolution module li in each downsampling and upsampling stage, where i = 1, 2, 3, 4, 5, 6, as shown below. Figure 4 As shown, the hybrid Transformer-convolutional module consists of two complementary feature extraction paths: a local branch responsible for capturing short-range texture and edge details, and a global branch utilizing a multi-head self-attention mechanism to introduce long-range context modeling capabilities. Through residual fusion, the features of the two paths are effectively integrated in a unified spatial domain, enabling the network to simultaneously possess local sensitivity and global structural understanding. The computation process of this module can be formally represented as follows:

[0092] ,

[0093] ,

[0094] ,

[0095] ,

[0096] Where x is the input feature, For the activation function, in this embodiment, the ReLU activation function is used, and TF stands for Transformer operation. This is for point-by-point operations; + indicates point-by-point addition. As an intermediate variable, For the output features of local branches, For the output features of the global branch, This is a 3×3 convolution operation. For a 1×1 convolution operation, l represents the output feature of the hybrid Transformer-convolution module.

[0097] Downsampling employs the PixelUnShuffle operation, which has the advantage of increasing the channel dimension while preserving resolution information, enabling the encoder to learn features from a higher-dimensional space. The process can be represented as follows:

[0098] ,

[0099] Where x is the input feature and d is the downsampling operation.

[0100] Upsampling uses the PixelShuffle operation to restore spatial resolution using the following formula, in order to avoid the artifact problem caused by traditional deconvolution.

[0101] ,

[0102] Where x is the input feature and u is the upsampling operation.

[0103] (4) Frequency domain encoder;

[0104] To further extract high-frequency texture details and structural edge information from the input image, this invention introduces a Butterworth high-frequency filter dense network into the frequency domain encoder. Its core is a learnable, channel-independent Butterworth high-pass modulation filter. The frequency domain encoder uses the Butterworth high-pass modulation filter to perform frequency domain modulation, achieving interpretable and adaptive high-frequency enhancement, filtering out low-frequency components and amplifying high-frequency information, such as... Figure 5 As shown.

[0105] The first step of the frequency domain encoder is to input the initial fused feature C, perform channel adjustment through a 1×1 convolution, and then perform a fast Fourier transform (FFT). The intermediate frequency domain characteristics composed of amplitude and phase are obtained. .

[0106] ,

[0107] Subsequently, learnable Butterworth high-pass modulation is applied to the intermediate frequency domain features using a Butterworth high-pass modulation filter. (The core operation is matrix multiplication) is defined as:

[0108] ,

[0109] Where D(u,v) is the frequency radius. Let n be the learnable cutoff frequency and the order, respectively. and These are the high-frequency enhancement component and the high-frequency transfer function, respectively. Let (u,v) be the c-th channel Butterworth high-pass modulation filter with coordinates (u,v), where (u,v) are frequency domain coordinates.

[0110] The frequency domain characteristics of the filter obtained after filtering by the Butterworth high-pass modulation filter , represented as:

[0111] ,

[0112] Finally, the filtered frequency domain features are restored to the spatial domain using the inverse Fourier transform (IFFT) to obtain the initial frequency domain features. :

[0113] ,

[0114] Then the initial frequency domain features The input is fed into two sequential enhancement paths, that is, feature extraction is performed alternately through convolution and transformer: after one 3×3 convolution layer, and then through the Transformer module, the frequency domain enhanced feature F is obtained. T1 The FT1 feature is processed through a 3×3 convolution layer, and then through a Transformer module to obtain the frequency domain enhanced feature F. T2 .

[0115] The initial fusion feature C and the frequency domain enhancement feature F are combined. T1 and F T2 The features are stitched together to achieve complementarity and fusion of spatial and frequency domain information. The stitched features are then subjected to 1×1 convolution for channel integration and dimensionality reduction, restoring the original number of channels, and finally outputting the frequency domain feature F. f .

[0116] This frequency domain encoder can effectively enhance fine textures, edge details, and the contour areas of infrared thermal targets in images.

[0117] (5) Spatial-frequency domain fusion decoder;

[0118] Spatial-frequency domain fusion decoder is used to fuse spatial features Frequency domain characteristics The final fused image is then reconstructed, preserving both the detailed information of the visible light image and the structural characteristics of the infrared image. First, the two features are fused through channel stitching:

[0119] ,

[0120] in, Indicates splicing along the channel. These are the features after splicing.

[0121] The spliced ​​features The input is a hybrid decoder consisting of Transformers and convolutional layers, such as... Figure 6 As shown, this fully integrates cross-scale spatial structural features with high-frequency texture information, and then is processed by a hybrid decoder. Output the final fused image :

[0122] ,

[0123] Hybrid decoder It can effectively integrate structured information from the spatial domain and detailed information from the frequency domain, so that the output image is comprehensively optimized in terms of brightness distribution, edge sharpness, texture details and overall structure, and reconstructs a final fused image that performs well in both subjective visual perception and multiple objective metrics.

[0124] (6) Establishment of the loss function;

[0125] The loss function is mainly used for training the space-frequency dual-branch cross-modal fusion network (PGDFusion). To achieve a comprehensive balance between preserving structural details, taking into account significant infrared and visible light information, and stabilizing training, this invention designs a loss function consisting of three parts. Including encoder consistency loss Basis / details decomposition constraint loss and the losses from integration and reconstruction These correspond to the feature consistency constraint of the cross-modal encoder, the structure preservation constraint of the two branches of the spatial domain encoder and the frequency domain encoder, and the supervision constraint of the final fused image output. See below for details.

[0126] ,

[0127] The visible light images, infrared images, and prior knowledge from the training set are respectively input into the last two layers of the corresponding Transformer modules in the cross-modal encoder, and the shallow features of the visible light images are output respectively. Shallow features of infrared images Shallow features of prior knowledge and deep features of visible light images. Deep features of infrared images Deep characteristics of prior knowledge , denoted as:

[0128] ,

[0129] Then as Figure 2 As shown, the encoder consistency loss can be expressed as:

[0130] ,

[0131] ,

[0132] Among them, L v L i L pThe L1 loss is calculated for visible light images, infrared images, and prior knowledge, respectively. Indicates L1 loss, , , These are the weight parameters for visible light images, infrared images, and prior knowledge, respectively, with a value of 1.

[0133] To fully leverage the advantages of the dual branches in the spatial and frequency domains, this invention decomposes the initial fused features into spatial features. Frequency domain characteristics The former primarily characterizes large-scale brightness and contour structure, while the latter emphasizes texture and high-frequency details. To explicitly constrain these two components to cover the salient information in the source image pair, this invention employs a basis / details decomposition constraint loss. We set up a combined loss of structural similarity and mean squared error, and used the maximum response per pixel as the supervision target to approximate the most salient information, as shown below:

[0134] ,

[0135] ,

[0136] ,

[0137] in, To address the decomposition constraint loss for spatial domain features, For the decomposition constraint loss targeting frequency domain features, SSIM is the structural similarity loss, and MSE is the mean squared error loss. This is the balance coefficient.

[0138] This invention obtains the final fused image F by constraining the fusion reconstruction loss, while preserving significant information from both the visible light and infrared images. Specifically, it is achieved through the following formula:

[0139] ,

[0140] ,

[0141] ,

[0142] in, L1 loss is the pixel loss. To adjust the parameters, For gradient loss, This represents the Sobel gradient operator. It outputs the supervision constraints of the fused image. This is used to encourage the fusion results to simultaneously approximate both modalities in terms of visual effects and structural fidelity.

[0143] Prior knowledge during the training phase Instead of being a hard target, it serves as a soft guiding signal, projected into the feature space by a lightweight cross-modal encoder, and combined with shallow features encoded from infrared images directly extracted from the source image pair. Shallow features after visible light image encoding Interacting with the network allows for the injection of enhanced structural consistency and detail saliency information in the early stages of network training, guiding the space-frequency dual-branch cross-modal fusion network to learn complementary features across modalities more efficiently and bridging the uncertainty in data-driven learning.

[0144] The training process in this embodiment is as follows: Figure 1 As shown, the image dataset, including infrared images and their corresponding visible light images, is divided into a training set and a test set. At the same time, a set of infrared images and their corresponding visible light images from the training set are used to obtain prior knowledge through the above step (1). After preprocessing the training set, it is processed through the above steps (2) to (5) to output the final fused image. Then, the loss function is calculated by comparing the final fused image with the infrared images and their corresponding visible light images in the training set. The network parameters of PGDFusion are updated by backpropagation. This process is repeated until the loss function reaches a local minimum point, the training ends, and the optimal network parameters are output. Then, PGDFusion is tested using the test set to obtain the corresponding final fused image. During the test, the maximum value of the corresponding source image pair in the test set is taken as the prior knowledge. That is, the source image pair is compared pixel by pixel, and the larger pixel value at each position is taken to generate a new input source to replace the prior knowledge during training. In general, during multimodal image fusion, the prior knowledge is taken as the maximum value of the corresponding source image pair in the test set.

[0145] Specifically, this embodiment selected 266 pairs of medical images from the Harvard Medical School website as a medical image dataset, of which 130 pairs were used for training and the remainder as a test set, including 21 pairs of MRI-CT images, 42 pairs of MRI-PET images, and 73 pairs of MRI-SPECT images. The training method followed... Figure 1 The method shown is used. In other embodiments of the present invention, other visible light and infrared fusion experiments were also conducted, using four popular benchmark datasets and medical image datasets to verify the effectiveness of PGDFusion. The benchmark datasets include M3FD, MSRS, RoadScene, and TNO. The present invention uses the MSRS training set (1083 pairs) and M3FD (240 pairs) for network training, and the MSRS test set (361 pairs), RoadScene (50 pairs), M3FD (60 pairs), and TNO (25 pairs) as test sets to comprehensively verify the fusion performance.

[0146] This embodiment uses six metrics to quantitatively evaluate the fusion results: entropy (EN), standard deviation (SD), visual information fidelity (VIF), edge information preservation (Qabf), spatial frequency (SF), and average gradient (AG). Higher metric values ​​indicate better fused images. The fusion results are then compared with methods including DID, U2F, FGAN, ITF, PIA, DDFM, and UMF.

[0147] The experiment was conducted using machines equipped with two NVIDIA GeForce RTX 3090 GPUs. Images were first preprocessed: training samples were randomly cropped into 128×128 image patches. The number of training epochs was set to 80. The batch size was set to 24. This invention uses the Adam optimizer with an initial learning rate of 10. 4 The decay rate was 0.5 every 20 rounds, and the experimental results are shown in Tables 1-7.

[0148] Table 1: Quantitative results of different methods on the TNO dataset.

[0149]

[0150] Table 2: Quantitative results of different methods on the RoadScene dataset.

[0151]

[0152] Table 3: Quantitative results of different methods on the M3FD dataset.

[0153]

[0154] Table 4: Quantitative results of different methods on the MSRS dataset.

[0155]

[0156] Table 5: Quantitative results of different methods on MRI-CT datasets.

[0157]

[0158] Table 6: Quantitative results of different methods on MRI-PET datasets.

[0159]

[0160] Table 7: Quantitative results of different methods on the MRI-SPECT dataset.

[0161]

[0162] Tables 1 to 4 present the quantitative evaluation results of different methods on the four datasets TNO, RoadScene, M3FD, and MSRS. It can be seen that PGDFusion achieves the best or second-best performance on most metrics.

[0163] In terms of detail preservation, PGDFusion significantly outperforms other methods in both SF and AG metrics, demonstrating its clear advantage in high-frequency details and edge sharpness. For example, on the M3FD dataset, PGDFusion's SF value reaches 18.42, far exceeding the second-place PIA's 11.49. EN and SD metrics reflect the information richness and contrast of the fused image. PGDFusion achieves the highest values ​​on multiple datasets, indicating its effective integration of bimodal information and avoidance of information loss. VIF and Qabf metrics evaluate the visual consistency between the fused image and the source image pair. PGDFusion performs stably and excellently on VIF, especially ranking first on the TNO and MSRS datasets, indicating good visual naturalness and structural fidelity.

[0164] In the MRI-CT fusion task, PGDFusion achieved SF and SD scores of 33.42 and 79.56, respectively, far exceeding other contrast methods, indicating that the fused images possess extremely high contrast and rich bone and tissue details. In Tables 6 and 7, PGDFusion also performed excellently, with Qabf scores of 0.72 and 0.77, and VIF scores of 0.67 and 0.81, respectively. These leading objective metrics demonstrate that PGDFusion successfully transferred edge information from functional images to anatomical images while preserving the salient features of the source image pairs to the greatest extent possible, providing high-quality fused image support.

[0165] The prior knowledge-guided dual-domain feature network multimodal image fusion method of the present invention can also be formed into a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the prior knowledge-guided dual-domain feature network multimodal image fusion method.

Claims

1. A prior knowledge-guided dual-domain feature network multimodal image fusion method, characterized in that, Includes the following steps: Step 1, acquire multiple sets of infrared images I ir and visible light image I vi As source image pairs, as training sets, each group of infrared images I ir and visible light image I vi Targeting the same objective; Step 2: Construct the space-frequency dual-branch cross-modal fusion network PGDFusion, which includes a cross-modal encoder, a spatial encoder, a frequency encoder, and a space-frequency fusion decoder. Step 3: Input the training set into PGDFusion, and extract the infrared image I using the cross-modal encoder. ir Visible light image I vi and prior knowledge The shallow features are then fused to obtain the initial fused feature C; The prior knowledge Obtain it through the following methods: Infrared and visible light images of the same target are acquired and fused using several different existing image fusion methods to obtain multiple candidate fused images. The candidate fused image with the best performance is selected as prior knowledge. Step 4: The spatial encoder acquires the contour, region brightness structure, and local texture information from the initial fused feature C to obtain the spatial feature. The frequency domain encoder transforms the initial fused feature C into the frequency domain, performs high-frequency enhancement, and obtains the frequency domain feature. ; Step 5: Spatial-frequency domain fusion decoder fuses spatial features. Frequency domain characteristics The final fused image is reconstructed and output. Step 6: Update the PGDFusion network parameters by backpropagating the loss function, return to step 3, and continue until the loss function reaches a local minimum to obtain the trained PGDFusion. Step 7: Input the infrared image and visible light image to be fused into PGDFusion, process them to obtain the final fused image. At this time, the prior knowledge takes the maximum value of the corresponding source image pair.

2. The prior knowledge-guided dual-domain feature network multimodal image fusion method according to claim 1, characterized in that, In step 4, the spatial encoder acquires the contour, region brightness structure, and local texture information from the initial fusion feature C to obtain the spatial feature. Specifically: a1, the initial fused feature C is alternately passed through multiple levels of hybrid Transformer-convolution modules and multiple levels of downsampling to output deep feature L3; The specific process of the hybrid Transformer-convolutional module is as follows: the input features are processed through local branches and global branches, and the corresponding features are output respectively. The features are then fused and output. a2, after downsampling the deep feature L3, multi-scale feature enhancement is performed to obtain the enhanced deep feature; a3, the enhanced deep features are alternately processed through multiple levels of upsampling and multiple levels of hybrid Transformer-convolution modules to obtain spatial domain features. In this process, the output features of each upsampled level are fused with the features obtained from the downsampled level at the corresponding level before entering the hybrid Transformer-convolution module. The number of upsampling and downsampling levels are the same.

3. The prior knowledge-guided dual-domain feature network multimodal image fusion method according to claim 2, characterized in that, Step a1 specifically includes the following steps: a11. Input the initial fused feature C, and after passing through the hybrid Transformer-convolution module l1, obtain the shallow feature L1; a12. Downsample the shallow feature L1 by d1, and then pass it through the hybrid Transformer-convolution module l2 to obtain feature L2; a13. Subtract feature L2 from feature L2 and then pass it through a hybrid Transformer-convolution module l3 to obtain deep feature L3.

4. The prior knowledge-guided dual-domain feature network multimodal image fusion method according to claim 3, characterized in that, Step a3 specifically includes the following steps: a31. Upsample the enhanced deep features u3 to obtain features U3 that match the L3 scale of the deep features; a32. The feature U3 is concatenated with the deep feature L3 located in the same layer during downsampling, and then processed by the hybrid Transformer-convolution module l4 after fusion. a33. Upsample the features output by the hybrid Transformer-convolution module l4 by u2 to obtain feature U2, concatenate it with feature L2, and then process it through the hybrid Transformer-convolution module l5; a34. The features output by the hybrid Transformer-convolutional module l5 are upsampled by u1 to obtain feature U1, which is then concatenated with the shallow feature L1 and passed through the hybrid Transformer-convolutional module l6 to output the spatial domain feature. .

5. The prior knowledge-guided dual-domain feature network multimodal image fusion method according to claim 2, characterized in that: In step a1, the downsampling is performed using the PixelUnShuffle operation; In step a3, the upsampling is performed using the PixelShuffle operation; Step 5 specifically involves: [Illegible text - likely related to spatial features] Frequency domain characteristics Fusion is achieved through channel splicing: ; in, Indicates splicing along the channel. Features after splicing; The final fused image is then output through a hybrid decoder consisting of alternating Transformer and convolutional layers.

6. The prior knowledge-guided dual-domain feature network multimodal image fusion method according to any one of claims 1-5, characterized in that, In step 4, the frequency domain encoder transforms the initial fused feature C to the frequency domain, performs high-frequency enhancement, and obtains the frequency domain feature. Specifically: b1, transform the initial fused feature C into the frequency domain to obtain the intermediate frequency domain feature. ; b2, using a Butterworth high-pass modulation filter to analyze the intermediate frequency domain characteristics Filtering is performed to obtain the filtered frequency domain characteristics. ; b3, using inverse Fourier transform to extract the frequency domain characteristics of the filter. Reverting to the spatial domain yields the initial frequency domain features. ; b4, the initial frequency domain features Enhancement is performed sequentially through two serial enhancement paths, and the frequency domain enhancement features F are output respectively. T1 and frequency domain enhancement features F T2 ; b5, the initial fusion feature C and the frequency domain enhancement feature F T1 and F T2 The components are then spliced ​​together and restored to their original number of channels to output the frequency domain feature F. f .

7. The prior knowledge-guided dual-domain feature network multimodal image fusion method according to claim 6, characterized in that: In step b1, the initial fused feature C undergoes channel adjustment via 1×1 convolution, followed by a Fast Fourier Transform to obtain the intermediate frequency domain feature. ; Step b4 specifically involves processing the initial frequency domain features. Feature extraction is performed sequentially through two convolutional transformers: initial frequency domain features. After one 3×3 convolution layer, followed by a Transformer module, the frequency domain enhanced feature F is obtained. T1 Frequency domain enhancement feature F T1 After one 3×3 convolution layer, followed by a Transformer module, the frequency domain enhanced feature F is obtained. T2 ; In step b5, restoring to the original number of channels specifically involves performing channel integration and dimensionality reduction through 1×1 convolution to restore to the original number of channels.

8. The prior knowledge-guided dual-domain feature network multimodal image fusion method according to claim 1, characterized in that: Step 1 is followed by processing the infrared image I. ir and visible light image I vi Preprocessing is performed; the preprocessing includes normalization and image cropping. In step 3, the candidate fusion image with the best performance is determined according to a comprehensive evaluation function. The comprehensive evaluation function integrates multiple complementary objective indicators, including structural similarity loss that measures structural fidelity and / or mutual information that measures information content. In step 7, the infrared image and the visible light image to be fused are input into PGDFusion after the preprocessing.

9. A prior knowledge-guided dual-domain feature network multimodal image fusion system, used to implement the prior knowledge-guided dual-domain feature network multimodal image fusion method of claim 1, characterized in that: This includes cross-modal encoders, spatial encoders, frequency encoders, and spatial-frequency fusion decoders; The cross-modal encoder is used to extract infrared image I. ir Visible light image I vi and prior knowledge The shallow features are then fused to obtain the initial fused feature C; The input end of the spatial encoder is connected to the output end of the cross-modal encoder and is used to acquire contour, regional brightness structure and local texture information in the initial fusion feature C; The input end of the frequency domain encoder is connected to the output end of the cross-modal encoder, and is used to convert the initial fused feature C to the frequency domain for high-frequency enhancement; The input of the spatial-frequency domain fusion decoder is connected to the output of the spatial domain encoder and the frequency domain encoder, respectively, for fusing spatial and frequency domain features and reconstructing the final fused image.

10. A computer program product, comprising a computer program, characterized in that: When executed by a processor, the program implements the steps of the prior knowledge-guided dual-domain feature network multimodal image fusion method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Parallel U-Net-based dual-domain collaborative infrared and visible light image fusion method

    CN121599852A

  • Image high-quality harmonization model training and device

    WO2024187901A1