Multi-modal image fusion method based on space-frequency interaction
By constructing a space-frequency interactive multimodal image fusion model and using wavelet transform for multi-scale decomposition and enhancement, the problem of underutilization of frequency domain information in existing technologies is solved, achieving efficient fusion of infrared and visible light images and improving image quality and information richness.
Patent Information
- Application Number
- CN202510980602.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-11-25
AI Technical Summary
Existing multimodal image fusion methods ignore frequency domain information during spatial domain fusion, resulting in insufficient extraction of modality-specific features. Furthermore, the interaction mechanism between the spatial and frequency domains is not well studied, affecting the fusion quality and information richness.
A multimodal image fusion method based on space-frequency interaction is adopted. Spatial and frequency features are extracted from infrared and visible light images through a dual-branch model. Wavelet transform is used for multi-scale decomposition and enhancement to achieve complementary fusion of frequency features. Image quality is improved through cross-domain feature fusion.
It improves the spatial quality and information richness of the fused images, and enhances the overall fusion effect.
Smart Images

Figure CN121010858A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of computer vision, and particularly relates to a multi-modal image fusion method based on space-frequency interaction. BACKGROUND
[0002] Multi-modal image fusion is a low-level vision processing technology that enhances image quality by integrating complementary and redundant information from heterogeneous sensors into a unified representation. Among them, the fusion of infrared images and visible light images is the most representative, which can cooperatively integrate the thermal radiation data of infrared images and the texture details of visible light images to compensate for the noise interference of infrared images and the defects of insufficient imaging conditions of visible light images, and has important application value in many scenes such as intelligent monitoring and automatic driving.
[0003] Recent studies have proposed various technical solutions to address the challenges of multi-modal image fusion, which can be systematically divided into traditional methods and deep learning methods. Traditional methods usually extract artificial design features through sparse representation, multi-scale transformation, subspace decomposition, and saliency analysis, but their generalization ability is obviously insufficient due to the inherent limitations of artificial design. Deep learning methods can be roughly divided into non-generative and generative methods, among which non-generative methods use convolutional neural networks, Transformer architecture, and autoencoders to efficiently model feature extraction and fusion processes. Methods based on convolutional neural networks aim to capture local spatial features of dual modalities, but often ignore long-range dependencies that play a key role in representing global structures. Techniques based on Transformer model long-range spatial dependencies in image sequences through self-attention mechanisms, but they have the inherent limitation of high computational complexity. Methods based on autoencoders usually pre-train the encoder on large datasets, but the fusion process relies on artificial design rules, which can easily cause modality preference problems and distort the saliency image. Generative methods mainly include generative adversarial networks and diffusion models. Techniques based on generative adversarial networks model the source data image distribution to generate fusion images by establishing an adversarial framework between the generator and the discriminator, but the adversarial training can easily lead to modality imbalance and constraint conflicts, often causing model collapse and degradation of fusion quality. Denoising diffusion probability models exhibit excellent high-quality image generation capabilities by modeling the diffusion process of gradually restoring noisy images to clear images, but excessive fine control of local details can compromise the spatial fidelity of the fusion image.
[0004] Although existing methods perform well in the multi-modal image fusion task, there are still several limitations that need to be broken through: first, most methods only focus on spatial domain image fusion, often ignoring the unique characteristics of source image frequency domain information, leading to insufficient modal-specific feature extraction; second, although some studies explore the potential of frequency domain in multi-modal fusion, they usually analyze the frequency domain of each source image independently and manually fuse the frequency components, failing to fully exploit the inter-modal and intra-modal frequency domain correlation across modalities. Taking infrared image and visible light image fusion as an example, visible light images mainly capture high-frequency details such as edges and textures, while infrared images focus on low-frequency thermal radiation information, so effectively exploring the inter-modal / intra-modal frequency domain correlation and integrating it into the fusion architecture is still a key challenge. In addition, the interaction mechanism between spatial domain and frequency domain is not sufficient, and this interaction plays a decisive role in realizing cross-domain complementary feature fusion and enhancing information transmission. SUMMARY
[0005] To overcome the shortcomings of the prior art, the present application provides a multi-modal image fusion method, system, medium and equipment based on spatial-frequency interaction, which systematically excavates frequency domain features and realizes adaptive interaction between frequency domain and spatial domain through spatial-frequency interaction and collaborative fusion, so that the model can efficiently process infrared image and visible light image fusion tasks, and improve the spatial quality and information richness of the fused image.
[0006] The technical scheme adopted by the present application to solve its technical problems is:
[0007] A multi-modal image fusion method based on spatial-frequency interaction, comprising the following steps:
[0008] Step 1, input infrared image and visible light image;
[0009] Step 2, use a basic convolutional neural network to preliminarily encode and extract features from the infrared image and the visible light image;
[0010] Step 3, perform deep spatial feature extraction on the preliminarily extracted spatial features of the infrared image and the visible light image, and then fuse the spatial features of the two;
[0011] Step 4, perform two-stage frequency feature extraction on the preliminarily fused spatial encoding features, where the first stage performs inter-modal feature interaction fusion, and the second stage performs intra-modal frequency enhancement;
[0012] Step 5, perform cross-domain feature fusion on the spatial features and frequency features of the fused infrared image and visible light image;
[0013] Step 6, input the cross-domain fused features into the decoder in turn for feature decoding;
[0014] Step 7, the decoded features are image reconstructed to obtain an output image fused from the infrared image and the visible light image.
[0015] Further, in step 3, the preliminary spatial encoding features of the infrared image and the visible light image are respectively applied to a deep feature extraction network, a state space model is used to extract deep spatial features, and a residual connection is combined to avoid overfitting and information loss in training; a convolution network, an activation function and other operations are applied to fuse the deep spatial features of the two.
[0016] Further, in step 4, the two-stage frequency feature extraction process for the preliminary fused spatial encoding features is as follows: in the first stage, wavelet transform is used to decompose the frequency subband to obtain high and low frequency components, then targeted high and low frequency variable enhancement operations are sequentially performed to improve the expression of high and low frequency features, then adaptive frequency fusion operations are performed to fuse the high and low frequency components of the two modalities, thereby obtaining fused high and low frequency components, and finally an inverse wavelet transform is performed to obtain a fused frequency feature; in the second stage, wavelet transform is also used to decompose the frequency subband to obtain high and low frequency components, then a selective scanning mechanism of a state space model is used to adaptively enhance and aggregate the information of each frequency band, thereby further enhancing the expression of frequency information.
[0017] Further, in step 5, the spatial features and the frequency features are first applied to feature addition to obtain preliminary fused features; a state space model is applied to capture global spatial correlation, and an activation function is used to obtain a feature weight map; the original spatial features and the frequency features are weighted and fused, and a residual connection is combined to obtain the final cross-domain fused features.
[0018] In the decoding stage of step 6, the features of the two scales are decoded to form unified scale decoded features.
[0019] In step 7, image reconstruction is applied to a convolutional neural network and an activation function to obtain a final fused image.
[0020] The preliminary encoding, deep spatial encoding, frequency feature extraction, cross-domain feature fusion, decoding and image reconstruction form a multi-modal image fusion model, which adopts a structural similarity loss, a pixel intensity loss, a spatial gradient loss and a frequency loss based on wavelet transform.
[0021] The technical concept of the present application is: a dual-branch space-frequency interactive multi-modal image fusion model is constructed, and the spatial and frequency features of infrared images and visible light images are extracted, which makes up for the problem of insufficient image information caused by fusion only in the spatial domain; a frequency feature extraction branch based on wavelet transform is proposed, the multi-scale decomposition capability of wavelet transform is used to analyze and enhance the high and low frequency components in each modal image, and complementary fusion of the frequency features of two modal images is realized, which saves the parameter quantity of the model while retaining the key frequency information of two modal images; through the space-frequency interactive fusion, the complementary fusion of space-frequency is realized, which efficiently improves the quality of the fusion image and better enables the downstream tasks.
[0022] The beneficial effects of the present application mainly manifest in: efficiently improving the spatial quality and information richness of the fusion image. BRIEF DESCRIPTION OF DRAWINGS
[0023] Fig. 1 is a flowchart of a multi-modal image fusion method based on space-frequency interaction;
[0024] Fig. 2 is a model architecture diagram of a multi-modal image fusion method based on space-frequency interaction;
[0025] Fig. 3 is a first-stage frequency feature extraction model detail diagram;
[0026] Fig. 4 is a second-stage frequency feature extraction model detail diagram;
[0027] Fig. 5 is a multi-modal image fusion system diagram based on space-frequency interaction. DETAILED DESCRIPTION
[0028] The present application will be further described below with reference to the accompanying drawings.
[0029] Referring to Figs. 1-4 A multi-modal image fusion method based on space-frequency interaction, comprising the following steps:
[0030] Step 1, input infrared images and visible light images;
[0031] First, use multi-modal sensors to collect image data, including infrared images and visible light images, to ensure that the sizes and resolutions of the paired images are consistent, and then normalize the collected images before inputting.
[0032] In this embodiment, the method is used to normalize the above data, which scales the numerical value to the interval [0, 1], which can be expressed as the formula:
[0033]
[0034] where X max and X min respectively represent the maximum and minimum pixel values of the image;
[0035] Step 2, using a basic convolutional neural network to preliminarily encode and extract features of the infrared image and the visible light image;
[0036] This embodiment uses a convolutional neural network and an activation function to preliminarily encode features of the preprocessed infrared image and the visible light image, which can be represented by the formula:
[0037] F = ReLU(Conv(X i )), X i ∈ {X ir ,X vi} (2)
[0038] where X ir and X vi respectively represent the infrared image and the visible light image, Conv(·) represents a convolutional neural network, and ReLU(·) represents an activation function;
[0039] Step 3, performing deep spatial feature extraction on the preliminarily extracted spatial features of the infrared image and the visible light image, and then fusing the spatial features of the two;
[0040] For deep spatial feature extraction, this embodiment introduces a state space model to perform efficient global spatial modeling, and then combines a gating structure to maintain spatial structure information. The deep spatial feature extraction can be represented by the formula:
[0041] F1 = FeedForward(SS2D(F)) (3)
[0042] where F represents the infrared image or the preliminarily encoded features of the visible light image, SS2D(·) represents a state space model, and FeedForward(·) represents a gating structure operation;
[0043] Step 4, performing two-stage frequency feature extraction on the preliminarily fused spatial encoded features, where the first stage performs inter-modal feature interaction fusion, and the second stage performs intra-modal frequency enhancement;
[0044] This embodiment introduces frequency feature learning in multi-modal image fusion, which makes up for the shortcomings of single spatial learning features, and simultaneously introduces an adaptive spatial-frequency fusion mechanism to effectively integrate spatial and frequency features, solving the integration problem caused by the domain difference between the spatial domain and the frequency domain.
[0045] For frequency feature extraction, the extraction is divided into two stages. The first stage is inter-modal frequency feature aggregation. The stage first uses discrete wavelet transform to decompose the input features into four sub-bands according to frequency components, which can be expressed as formula:
[0046] I LL ,I LH ,I HL ,I HH = DWT(I) (4)
[0047] where I represents the infrared image or the visible light preliminary encoded feature, I LL ,I LH ,I HL ,I HH represent the low-frequency component, the horizontal high-frequency component, the vertical high-frequency component and the diagonal high-frequency component respectively, and DWT(·) represents the discrete wavelet transform.
[0048] Then, in order to enhance the expression of each frequency component, specific high-frequency and low-frequency component enhancement processing is designed for the difference of information contained in the high-frequency component and the low-frequency component. The high-frequency component often has rich boundary and texture information. A state space model is introduced to model the global correlation of the high-frequency component, and the global high-frequency information expression is improved. The process can be expressed as formula:
[0049] X1 = SS2D(X0) + X0 (5)
[0050] where X0∈[I LH ,I HL ,I HH ] represents the stacked high-frequency component feature, and SS2D(·) represents the space state model. The low-frequency component is basically composed of smooth information such as background and sky, and only convolution processing is used to realize the enhancement of local low-frequency information. The process can be expressed as formula:
[0051] Y1 = Conv 1×1 (LeakyReLU(Conv 3×3 (Y0)) (6)
[0052] where Y0 represents the low-frequency component feature, Conv 3×3 (·), LeakyReLU(·), and Conv 1×1 (·) represent the convolutional neural network with kernel size 3, the activation function, and the convolutional neural network with kernel size 1, respectively.
[0053] Further, an attention mechanism-based fusion module is introduced to correspondingly integrate the low and high frequency features from the infrared image and the visible light image, and the overall process is as follows:
[0054]
[0055] where X i ,X v represent the high-frequency or low-frequency features of the infrared and visible light images respectively, Cat(·) represents the feature concatenation processing, Conv 3×3 (·) represents the convolutional neural network with a kernel size of 3, Sigmoid(·) is the activation function, weight represents the generated feature weight map, Fused0 represents the preliminary weighted fused feature, Fused2 represents the final fused feature, ChanAtten(·) and SpaAtten(·) represent the channel attention calculation and spatial attention calculation respectively.
[0056] Thus, F LL ,F LH ,F HL ,F HH are obtained. Finally, the features are converted from the frequency domain to the spatial domain using the inverse discrete wavelet transform, and the specific process is as follows:
[0057] F fused = IDWT([F LL ,F LH ,F HL ,F HH ]) (8)
[0058] where IDWT(·) represents the inverse discrete wavelet transform, and the obtained fused feature is F fused .
[0059] The second stage is the modal internal frequency enhancement processing. The fused feature Fused obtained in the first stage is further subjected to the discrete wavelet transform, and then the four frequency bands after decomposition are spliced. Then, by virtue of the selective scanning ability of the state space model, the feature enhancement is realized from low frequency to high frequency. The overall process is shown in the following formula:
[0060] Freq1 = IDWT(aX + FreqSS2D(DWT(F fused ))) (9)
[0061] where FreqSS2D(·) represents the frequency-based state space model, and a is a learnable scale parameter.
[0062] Step 5: Cross-domain feature fusion of the spatial features and the frequency features of the fused infrared image and visible light image;
[0063] The embodiment introduces a cross-domain feature fusion module to realize the complementary integration of spatial features and frequency features. The operation adopts a coarse-to-fine progressive fusion strategy, first uses a simple tensor addition operation to coarsely fuse the spatial domain and frequency domain features, then uses a state space model for global modeling, then uses an activation function to obtain a feature weight map, and then the spatial domain and frequency domain features are finely weighted and fused according to the weight size. The overall process is shown in the following formula:
[0064]
[0065] Wherein, X and Y represent spatial domain features and frequency domain features respectively.
[0066] Step 6, the cross-domain fused features are sent into the decoder for feature decoding;
[0067] In order to realize multi-scale analysis, the above-mentioned encoding stage divides the features into three scales for extraction, which are (H, W), Depth spatial feature extraction is performed under each scale. In the (H, W) scale, frequency feature extraction is performed in the first stage to aggregate frequency features from two modalities into one frequency feature, and then in the Second stage frequency feature extraction is only performed in two scales. The spatial-frequency fusion features obtained in three scales are The spatial-frequency fused features are decoded, specifically, a state space model is used for decoding, which is expressed in the formula as follows:
[0068]
[0069] Wherein, Upsample ×2 (·) represents 2 times up-sampling processing.
[0070] Step 7, the decoded features are reconstructed to obtain the output image of the fused infrared image and visible light image.
[0071] The embodiment uses convolution and activation function to reconstruct the final fused image, which is expressed in the formula as follows:
[0072] F out =Tanh(Conv 3×3 (Decode2)) (12)
[0073] Wherein, Decode2 is the final decoding feature, Tanh(·) is the activation function, and F out is the final fused image.
[0074] In this embodiment, preliminary feature encoding, deep spatial feature-frequency feature extraction and fusion, feature decoding, and image reconstruction constitute a complete multimodal image fusion model. Model training is constrained by structural similarity loss, pixel intensity loss, spatial gradient loss, and frequency loss based on wavelet transform. The structural similarity loss is expressed as follows:
[0075]
[0076] Among them, F out ,F i ,F v These represent the fused image, infrared image, and visible light image, respectively. `ssim(·)` represents structural similarity calculation. The formula for calculating pixel intensity loss is as follows:
[0077] L int =||F out -max (F i ,F v )‖1 (14)
[0078] Where max(·) is the maximum value function, and ||·||1 represents the F-1 norm calculation. The formula for calculating the spatial gradient loss is as follows:
[0079]
[0080] in, Let || represent the gradients of the fused image, infrared image, and visible light image, respectively, and |·| be the absolute value operation. The frequency loss formula based on wavelet transform is as follows:
[0081] L freq =‖|Freq out |-max(|Freq i |,|Freq v |)‖1,Freq∈{LL,LH,HL,HH} (16)
[0082] Among them, Freq out Freq i Freq v These represent the frequency components of the fused image, infrared image, and visible light image, respectively, including four types: LL, LH, HL, and HH.
[0083] The total loss can be expressed as:
[0084] L=αL ssim +βL int +γL grad +δL freq (17)
[0085] Wherein, a, b, g, d represent the weight of each loss term respectively.
[0086] Gradient clipping is applied in the model training process to prevent gradient explosion, the initial learning rate is set to 1e-4, and a learning rate decay strategy is designed, and the model parameters are saved regularly.
[0087] The embodiment also provides a multi-modal image fusion system based on space-frequency interaction, comprising a data acquisition module, a data processing module, a feature encoding module, a feature decoding module and an image reconstruction module, and specific descriptions are as follows:
[0088] The data acquisition module is used to acquire infrared image data and visible light image data.
[0089] The data processing module is used to pre-process the images, including image pairing, size cropping and normalization.
[0090] The feature encoding module is used to sequentially perform primary feature extraction, spatial feature extraction, frequency domain feature extraction and space-frequency feature fusion on the image data.
[0091] The feature decoding module is used to perform feature decoding on the encoded features.
[0092] The image reconstruction module is used to reconstruct the decoded features to obtain a final fusion image.
[0093] The embodiment also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the processing processes of the data acquisition module, the data processing module, the feature encoding module, the feature decoding module and the image reconstruction module.
[0094] The embodiment also provides an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor realizes the processing processes of the data acquisition module, the data processing module, the feature encoding module, the feature decoding module and the image reconstruction module when executing the computer program.
[0095] The content described in the embodiments of the present specification is only a list of implementation forms of the inventive concept, and is only for the purpose of description. The protection scope of the present application should not be regarded as being limited to the specific forms described in the embodiments, and the protection scope of the present application also extends to equivalent technical means that can be thought of by those skilled in the art according to the inventive concept.
Claims
1. A multimodal image fusion method based on space-frequency interaction, characterized in that, The method includes the following steps: Step 1: Input the infrared image and the visible light image; Step 2: Use basic convolutional neural networks to perform preliminary encoding and feature extraction on infrared and visible light images; Step 3: Extract depth spatial features from the spatial features of the initially extracted infrared and visible light images, and then fuse the spatial features of the two. Step 4: Perform two-stage frequency feature extraction on the initially fused spatial coding features. The first stage involves feature interaction fusion between modalities, and the second stage involves frequency enhancement within modalities. Step 5: Perform cross-domain feature fusion on the spatial and frequency features of the fused infrared and visible light images; Step 6: The cross-domain fused features are then fed into the decoder for feature decoding. Step 7: Reconstruct the image using the decoded features to obtain the output image after fusing the infrared image and the visible light image.
2. The multimodal image fusion method based on space-frequency interaction as described in claim 1, characterized in that, In step 3, the preliminary spatial coding features of infrared and visible light images are respectively applied to a deep feature extraction network, and the deep spatial features are extracted using a state space model. Residual connections are used to avoid overfitting and information loss during training. Convolutional networks and activation function operations are applied to fuse the deep spatial features of the two images.
3. A multimodal image fusion method based on space-frequency interaction as described in claim 1 or 2, characterized in that, In step 4, the process of extracting frequency features from the preliminary fused spatial coding features in two stages is as follows: In the first stage, the frequency subband is decomposed using wavelet transform to obtain high-frequency and low-frequency components. Then, targeted high-frequency and low-frequency variable enhancement operations are performed to improve the expression of high-frequency and low-frequency features. Next, the high-frequency and low-frequency components of the two modes are fused through adaptive frequency fusion operation to obtain fused high-frequency and low-frequency components. Finally, a fused frequency feature is obtained through inverse wavelet transform. The second stage also uses wavelet transform to decompose the frequency subbands to obtain high and low frequency components. Then, the selective scanning mechanism of the state-space model is used to adaptively enhance and aggregate the information of each frequency band, further enhancing the expression of frequency information.
4. A multimodal image fusion method based on space-frequency interaction as described in claim 1 or 2, characterized in that, In step 5, feature addition is first applied to spatial features and frequency features to obtain preliminary fusion features; A state-space model is applied to capture global spatial correlations, and an activation function is used to obtain a feature weight map. The original spatial features and frequency features are weighted and fused, and combined with residual connections, to obtain the final cross-domain fused features.
5. A multimodal image fusion method based on space-frequency interaction as described in claim 1 or 2, characterized in that, In the decoding stage of step 6, features from both scales are combined for decoding and unified into decoding features of a single scale.
6. A multimodal image fusion method based on space-frequency interaction as described in claim 1 or 2, characterized in that, The image reconstruction in step 7 involves applying a convolutional neural network and activation functions to obtain the final fused image.
7. A multimodal image fusion method based on space-frequency interaction as described in claim 1 or 2, characterized in that, The initial encoding, deep spatial encoding, frequency feature extraction, cross-domain feature fusion, decoding, and image reconstruction form a multimodal image fusion model. This model employs structural similarity loss, pixel intensity loss, spatial gradient loss, and frequency loss based on wavelet transform.