Multi-modal image fusion deep network model based on multiple filters and application
By adopting a deep network model based on multi-modal image fusion, using LEF, Gabor filter and INN for feature extraction and fusion, the problems of high computational complexity and poor interpretability in the prior art are solved, and the efficient and high-quality image fusion effect is achieved.
Patent Information
- Application Number
- CN202510100632.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
Existing deep learning models have high computational complexity in multimodal image fusion, prolonged training and inference time, and lack interpretability, making it difficult to achieve efficient image fusion.
Multimodal image fusion deep network model (MFF model) based on multi-filters is used to extract image features using CNN technology and multi-filter technology, including low-frequency enhancement filter (LEF), Gabor filter and reversible neural network model (INN), high-frequency and low-frequency features are fused through addition strategies, and feature decoding is used using Transformer model.
It improves the quality and speed of image fusion, reduces the computational complexity, and achieves more efficient image feature extraction and fusion, which has strong universality and interpretability.
Smart Images

Figure CN120013785A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine vision, and specifically relates to a multi-filter based multimodal image fusion deep network model and application. Background Art
[0002] For many years, multimodal image fusion technology has been a hot topic in machine vision research. According to the type of image to be fused, it can be divided into: medical image fusion, out-of-focus image fusion, infrared-visible light image fusion, multi-exposure image fusion, etc. This technology has a wide range of applications in improving image quality, achieving accurate image recognition and tracking, promoting remote sensing capabilities, and advancing medical imaging technology. Very important in this field is the integration of infrared and visible light images. Due to their different wavelength characteristics, visible light captures reflected light, while infrared captures thermal radiation, thereby capturing unique data. Therefore, the combined output provides more comprehensive insights compared to the individual single modality inputs. In addition, they also exhibit complementary characteristics. Although visible light images are closely consistent with human visual perception by providing high resolution and clear brightness contrast, they are easily affected by the environment, which significantly reduces the imaging effect. The fusion of these two modes effectively compensates for each other's limitations, thereby maximizing the information content in the composite image. Visible-infrared image fusion (VIF) has practical significance across fields, such as advanced target detection and tracking systems, remote sensing technology for accurate data acquisition, and intelligent computing-driven car navigation systems.
[0003] Compared with traditional machine learning models, deep learning models have less stringent requirements on image feature values and higher image quality after fusion. Therefore, there are related deep learning model solutions at home and abroad to meet the above requirements. For example, the image fusion method using CNN extracts features from images through convolutional neural networks to improve the accuracy and diversity of feature fusion. This is achieved by integrating residual connections, dense connections, attention mechanisms, and multi-scale extraction techniques. The image fusion method based on autoencoders uses autoencoders to pre-train the encoder using input images and extracts features through the decoder. Then, the final extracted features are fused according to the established fusion rules. The famous infrared visible image fusion method DenseFuse is based on autoencoders. Fu et al. also proposed a method based on a dual-branch encoder, which uses acoustic emission to perform feature fusion through addition strategy and channel selection strategy. The Transformer model has made significant contributions in the field of image fusion due to its excellent performance in long-range dependencies. A notable example is the CDDFuse model, which combines the Transformer with the CNN feature extractor. LiteTransformer (LT) uses a long-range attention mechanism to process low-frequency global features.
[0004] In summary, the image fusion methods using deep learning can be divided into CNN-based methods, autoencoder (AE)-based methods, GAN-based methods, and the recently popular Transformer-based methods. Most deep learning models for image fusion follow a series of steps, including image feature extraction, feature fusion, and image reconstruction. Among them, models based on the Transformer structure show strong performance in image fusion. However, they have high requirements on computing power and data volume, resulting in prolonged training and inference time. In addition, CNN-based image fusion models lack interpretability. Although deep models based on decoder-encoder structures provide faster computing speeds, their comprehensive indicators do not match the performance of Transformer. In order to enhance the feature extraction of the images to be fused to achieve the purpose of fusion, it is very necessary to invent a multimodal image fusion model with universality and low time computational complexity. Summary of the invention
[0005] In view of this, the present invention aims to solve the technical problems in the prior art and provide a multi-filter based multimodal image fusion deep network model and application. The multi-filter based multimodal image fusion deep network model (i.e., MFF model) of the present invention is established on the basis of the automatic encoding model, and uses CNN technology and multi-filter technology to extract image features. The MFF model of the present invention has greatly improved image fusion quality, image fusion speed, and image fusion universality compared with the prior art.
[0006] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0007] A multi-modal image fusion deep network model based on multiple filters, comprising an encoder module, a merging strategy module and a decoder module connected in sequence;
[0008] The encoder module includes two sets of feature extraction modules;
[0009] Each set of feature extraction modules includes LEF module, GaborCNN layer and INN module;
[0010] The LEF module is used for extracting low-frequency features of images;
[0011] The GaborCNN layer is used to extract texture features from high-frequency features of the image;
[0012] The INN module is used for edge feature extraction in high-frequency features of images;
[0013] The merging strategy module adopts an addition strategy to fuse the high-frequency features and low-frequency features extracted from the multimodal image pair by the feature extraction module;
[0014] The decoder module uses the Transformer model to perform feature decoding and combines the convolution structure to achieve fused image restoration.
[0015] In the above technical solution, the steps of using the LEF module to extract low-frequency features of an image are as follows:
[0016] First, a 1×1 convolution block is used to adjust the number of channels, transforming the input tensor shape from R h×w×1 Convert to R h×w×32 ; A multi-scale structured adaptive average pooling layer is used to implement low-frequency filtering, and upsampling technology is used to restore the uniformity of feature scales; the input dimension is divided into four equal parts, and the feature F obtained by the first convolution is expressed as:
[0017] F={f1,f2,f3,f4} (1)
[0018] f i ∈F(i=1,2,3,4) have the same shape; for each sub-feature f of feature F i Perform multi-scale pooling operations and use a bilinear interpolation sampling algorithm to upsample the pooling results for each sub-feature f i Generate the corresponding feature η i ; The process formula is as follows:
[0019] η i =Up(β s (f i )) (2)
[0020] s represents the size of the pooling layer. i Four pooling operations are required, and the size of the pooling layer in each operation is s∈{1,2,3,6}; β s represents a pooling operation with a size of s∈{1,2,3,6}; Up represents an upsampling operation;
[0021] Finally, the four sub-features f i The extracted features are connected and the ReLU activation function is applied; then, they are adjusted through the convolution layer Conv to meet the shape requirements of subsequent high-frequency features; finally, the extracted features are connected and then the ReLU activation function is applied; then, they are refined through the convolution layer Conv to meet the size requirements of high-frequency features; the low-frequency features finally obtained are It is expressed as:
[0022]
[0023] In the above technical solution, the steps of using the GaborCNN layer to extract texture features from high-frequency features of an image are as follows:
[0024] The two-dimensional Gabor filter is a Gaussian kernel function modulated by a sine wave. The function formula is as follows:
[0025]
[0026] Where x′ and y′ are expressed as:
[0027] x′=xcosθ+ysinθ (5)
[0028] y′=-xcosθ+ycosθ (6)
[0029] In formula (4), σ is the Gaussian variance, is the phase, ω is the length of the sine wave; x, y are the pixel coordinate positions, x′, y′ represent the x, y pixel coordinate positions after the parallelogram rule respectively;
[0030] In formula (5) and formula (6), θ is the angle between two intersecting lines in the parallelogram rule;
[0031] The real part of the two-dimensional Gabor filter is expressed as:
[0032]
[0033] The real part of the two-dimensional Gabor filter is an effective method to extract image texture. By inputting an image containing texture information into the GaborCNN layer, the texture features of the image can be obtained. I is the input image;
[0034]
[0035] In the above technical solution, the steps of extracting edge features from high-frequency features of an image using the INN module are as follows:
[0036] First, the input of each reversible module is divided into two parts along the channel axis, and a mapping function is introduced, using the Bottleneck Residual Block (BRB) in MobileNetV2 as its mapping function;
[0037] I=I[1:a]+I[a+1:A] (9)
[0038]
[0039] In formula (9) to formula (12), I is the input image, A is the total number of channels, BRB is the bottleneck residual block, ⊙ is the Hadamard product, is the extracted edge feature; since the edge feature extraction of the INN module is carried out in multiple times, represents the edge feature of the k+1th iteration; a is the channel value corresponding to the channel axis;
[0040] The final high-frequency features It is expressed as:
[0041]
[0042] In the above technical solution, when the multi-filter based multimodal image fusion deep network model of the present invention is used for the fusion of infrared and visible light two-modal images, the merging strategy module adopts an addition strategy to fuse the high-frequency features and low-frequency features extracted from the multimodal image pair, which is expressed as:
[0043]
[0044] in, represents the operation of the merge strategy, is the high-frequency feature of the visible light image, is the high-frequency feature of the infrared image, is the low-frequency feature of the visible light image, is the low-frequency feature of the infrared image, Fuse High Fuse is the combined high-frequency features of the visible light image and the infrared image. Low It is the combined low-frequency feature obtained by fusing the low-frequency features of the visible light image with the low-frequency features of the infrared image.
[0045] In the above technical solution, the decoder module uses the Transformer model to perform feature decoding as follows:
[0046] The decoder module uses the Transformer model to perform feature decoding in two stages;
[0047] In the first stage, the high- and low-frequency features of the visible light image and the high- and low-frequency features of the infrared image are decoded respectively;
[0048] In the second stage, the high-frequency fusion features of the infrared image and the visible light image are decoded, and the low-frequency fusion features of the visible light image and the infrared image are decoded.
[0049] In the above technical solution, the multi-filter based multimodal image fusion deep network model of the present invention is trained using the following loss function;
[0050] The loss function is divided into two stages, expressed as follows:
[0051] L MFF =L Ⅰ +L Ⅱ (15)
[0052] L MFF is the overall loss function of the MFF model, L Ⅰ is the loss function of the first stage, L Ⅱ is the loss function of the second stage;
[0053] The loss function L in the first stage Ⅰ It is expressed as follows:
[0054] L Ⅰ =L i +αL v +βL cc (16)
[0055] In formula (16), L i represents the reconstruction loss of infrared image, L v represents the reconstruction loss of the visible light image, L cc Represents the related loss, α and β are weight parameters;
[0056] The reconstruction loss function includes the mean square error loss Lmse, the structural similarity index measurement loss Lssim and the gradient loss Lgrad; combined with formula (16), the above three losses are expressed as:
[0057]
[0058]
[0059] In formula (17), Ir represents the original visible light image, represents the reconstructed visible light image; Vi and represent the original and reconstructed infrared images respectively;
[0060] In formula (16), Lcc is expressed as:
[0061]
[0062] In formula (18), CC is the correlation coefficient operator and ε is the bias;
[0063] The loss function L in the second stage Ⅱ :
[0064] L Ⅱ =L in +ωL max_grad +τLcc (19)
[0065] L in Indicates strength loss, L max_grad represents the maximum gradient loss, L cc Indicates the relevant loss;
[0066] L in =‖Fuse Image -Max(Ir,Vi)‖
[0067]
[0068] In formula (20), Fuse Image is the final synthetic infrared and visible light image fusion result. represents the Sobel gradient operator, CC is the correlation coefficient operator, and Max is the maximum value operation;
[0069] In formula (19), ω and τ both represent weight parameters.
[0070] The multi-filter based multimodal image fusion deep network model of the present invention is applied to target tracking or medical image fusion.
[0071] The beneficial effects of the present invention are:
[0072] The multi-filter based multimodal image fusion deep network model of the present invention has the following advantages:
[0073] 1. In the process of multimodal image fusion, a multi-filter mechanism is introduced to extract high-frequency and low-frequency features respectively using various types of filters. By effectively integrating these features, this method not only improves the quality of image fusion but also reduces the computational complexity.
[0074] 2. Considering that high-frequency features contain richer information, we classify high-frequency features more finely. Specifically, we distinguish between high-frequency edge information and high-frequency texture information. When extracting high-frequency edge information, an invertible neural network model (INN model) is used; while for extracting high-frequency texture features, we use Gabor filters. We use a low-frequency enhancement filter (LEF) to extract low-frequency features. Multiple groups with different kernel sizes and step sizes are used to create a dynamic low-pass filter. The LEF model not only performs well in low-frequency feature extraction, but also significantly reduces the computational complexity.
[0075] 3. A two-stage training mode is adopted. Taking infrared and visible light images as an example, the high- and low-frequency features of visible light images and infrared images are decoded in the first stage. In the second stage, the high-frequency fusion features of infrared images and visible light images are decoded, and the low-frequency fusion features of visible light images and infrared images are decoded. This allows the fusion results to have more feature information.
[0076] 4. The MFF model of the present invention, as an alternative to the Transformer model, produces a lightweight image fusion model that reduces the required computing power. The present invention uses three VIF datasets to verify the ability of the MFF model to fuse infrared and visible light images, and further demonstrates its advantages in image fusion quality through downstream machine vision tasks. In addition, three medical image datasets are used to verify the versatility and superior performance of the MFF model in multimodal data fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0078] Figure 1 This is an architecture diagram of the multi-filter based multimodal image fusion deep network model of the present invention.
[0079] Figure 2 This is the LEF module architecture diagram.
[0080] Figure 3 This is the visible light infrared image fusion image of data No. 11 in the TNO dataset.
[0081] Figure 4 This is a visible light infrared image fusion image of the FLIR 04602 data in the Roadscene dataset.
[0082] Figure 5 This is the visible light infrared image fusion image of the 00706N data in the MSRS dataset.
[0083] Figure 6 Detailed comparison of the fusion results of the MSRS dataset.
[0084] Figure 7 Detailed comparison of the fusion results of the RoadScene dataset.
[0085] Figure 8 Detailed comparison of the fusion results of the TNO dataset.
[0086] Fig. 9 This is a comparison result diagram of target detection for the fused results.
[0087] Fig.10This is a comparison chart of MRI-CT image fusion results.
[0088] Fig.11 This is a comparison chart of MRI-PET image fusion results.
[0089] Fig.12 This is a comparison chart of MRI-SPECT image fusion results. DETAILED DESCRIPTION
[0090] The inventive concept of the present invention is: the present invention provides a method for improving image texture detail information and image feature dimension by using a multi-filter based multimodal image fusion deep network model (MFF model). When using the traditional fusion method, it is generally necessary to perform feature analysis of limited dimensions on the image, such as brightness value, grayscale value and RGB parameters, etc., but it is impossible to perform universal fusion for two images of different modalities. When using a deep learning model, the more popular transformer model architecture (Transformer) is usually used to extract dimensions, but the Transformer only has good performance in extracting low-frequency features, and the performance of extracting high-frequency features is average. The present invention uses multiple filter networks to extract corresponding high-frequency and low-frequency features, thereby improving the training and reasoning time of the model. The more important high-frequency features are refined, the high-frequency texture features are extracted using Gabor filters, the high-frequency edge features are extracted using reversible neural networks, and the low-frequency enhancement filter (Low-Frequency Enhancement Filter) is used to enhance and extract low-frequency features, and the fused image is obtained through fusion strategy and two-stage model training. After verification by multiple experimental data sets, the model method proposed in the present invention has good performance in the fusion of infrared images and visible light images for multi-type image fusion, and has strong universality, high fusion speed, high fusion accuracy, low fusion time complexity, and can be directly migrated to medical image fusion, such as CT and MRI, MRI and PET, PET and SPECT image fusion also has good performance, which can also prove the universality of the present invention.
[0091] The present invention discloses a multi-modal image fusion deep network model based on multiple filters, which encodes, fuses and decodes the multi-modal image pairs to be fused through a deep learning model to form a final fusion result.
[0092] (1) Encoding, fusion, and decoding process
[0093] (a) Encoding process: INN module is used for edge feature extraction, Gabor filter is used for high-frequency texture feature extraction, and LEF module is used for low-frequency feature extraction.
[0094] (b) Fusion: Using the addition strategy, the high-frequency features and low-frequency features extracted from the multimodal image pairs are fused.
[0095] (c) Decoding process: The fused features are decoded using a two-stage process, and the model is trained using a two-stage training process. In the first stage, the feature extraction unit extracts high-frequency and low-frequency features from the visible light image (or infrared image) and inputs them into the decoding unit to obtain the final reconstructed image. In the second stage, a pair of visible light and infrared images are input into the encoder module to extract the corresponding high-frequency and low-frequency features. The comprehensive features of the visible light image (or infrared image) are obtained through the fusion layer, and then the fused features are sent to the decoding module for image synthesis and reconstruction.
[0096] (d) Training process: The loss function is also divided into two stages. Taking the input images of the two modalities of infrared and visible light as an example, the first stage extracts high-frequency and low-frequency features from the visible light image and the infrared image respectively. These features are then fused in their respective domains to reconstruct the corresponding infrared and visible light images. This stage mainly focuses on training the encoder and decoder. In the second stage, the high-frequency features of the infrared and visible light images are combined to obtain fused high-frequency features, and the low-frequency features also undergo a similar fusion process. Then, the high-frequency and low-frequency features are fused to enhance the image quality. This stage emphasizes the fusion ability of the training model. Therefore, the loss function is also divided into two components, each of which corresponds to one of the above training stages.
[0097] (2) According to (a) in (1), use the following formulas (1) to (3) to extract low-frequency features from the input image.
[0098] (3) According to (a) in (1), high-frequency texture features are extracted from the input image using the following formulas (4) to (8).
[0099] (4) According to (a) in (1), high-frequency edge features are extracted from the input image using the following formulas (9) to (12).
[0100] (5) According to (3) and (4), the high-frequency features of the image are obtained using the following formula (13).
[0101] (6) According to (b) in (1), the high-frequency features and low-frequency features of the image pair are fused at different frequencies using the following formula (14).
[0102] (7) According to (d) in (1), the loss function is set for the first stage of the training process using the following formulas (16) to (18).
[0103] (8) According to (d) in (1), use the following formulas (19) to (20) to set the loss function for the second stage of the training process.
[0104] The multi-filter based multimodal image fusion deep network model of the present invention can also be used for target tracking or medical image fusion.
[0105] The present invention is described in detail below with reference to the accompanying drawings.
[0106] The multi-filter based multimodal image fusion deep network model (MFF model) of the present invention comprises an encoder module, a merging strategy module and a decoder module connected in sequence;
[0107] The encoder module includes two groups of feature extraction modules; each group of feature extraction modules includes a LEF module (low frequency enhancement filter), a GaborCNN layer (Gabor filter) and an INN module (reversible network model); the LEF module is used for image low frequency feature extraction; the GaborCNN layer is used for texture feature extraction in image high frequency features; the INN module is used for edge feature extraction in image high frequency features;
[0108] The merging strategy module adopts an addition strategy to fuse the high-frequency features and low-frequency features extracted from the multimodal image pair by the feature extraction module;
[0109] The decoder module uses the Transformer model to perform feature decoding and combines the convolution structure to achieve fused image restoration.
[0110] The various parts of the MFF model of the present invention are described in more detail below. The overall architecture of the MFF model is as follows: Figure 1 As shown, the fusion of multimodal images is performed according to the encoder module, the merging strategy module, and the decoder module.
[0111] First, let’s introduce the encoder module:
[0112] (1) Low-frequency feature extraction:
[0113] The present invention uses a low frequency enhancement filter (LEF) to extract low frequency features. The framework diagram of the LEF module is as follows: Figure 2 shown.
[0114] First, a 1×1 convolution block is used to adjust the number of channels, transforming the input tensor shape from R h×w×1 Convert to R h ×w×32A multi-scale structured adaptive average pooling layer is used to implement low-frequency filtering, and upsampling technology is used to restore the uniformity of feature scales. The input dimension is divided into four equal parts. The feature F obtained by the first convolution can be expressed as:
[0115] F={f1,f2,f3,f4} (1)
[0116] f i ∈F(i=1,2,3,4) have the same shape. For each sub-feature f of feature F i Perform multi-scale pooling operations and use a bilinear interpolation sampling algorithm to upsample the pooling results for each sub-feature f i Generate the corresponding feature η i The process can refer to formula (2).
[0117] η i =Up(β s (f i )) (2)
[0118] s represents the size of the pooling layer. i Four pooling operations are required, and the size of the pooling layer in each operation is s∈{1,2,3,6}. s Represents a pooling operation with a size of s∈{1,2,3,6}. Up represents an upsampling operation. Finally, the four sub-features are concatenated and the ReLU activation function is applied. Then, they are adjusted through the convolution layer Conv to meet the shape requirements of the subsequent high-frequency features. Finally, the extracted features are concatenated and then ReLU activated. Then, they are refined through the convolution layer Conv to meet the size requirements of the high-frequency features. The low-frequency features finally obtained It can be expressed as:
[0119]
[0120] (2) Texture feature extraction:
[0121] In the context of texture feature extraction, Gabor filters are used because they are very effective in capturing frequency and orientation features. As mentioned earlier, Gabor filters closely match the characteristics of the human visual system, which makes them particularly effective in tasks involving texture representation and discrimination. A two-dimensional Gabor filter can be mathematically described as a Gaussian kernel function modulated by a sine wave. The process is shown in the following formula (4):
[0122]
[0123] As shown in formula (4), x′ and y′ can be expressed as:
[0124] x′=xcosθ+ysinθ (5)
[0125] y′=-xcosθ+ycosθ (6)
[0126] In formula (4), σ is the Gaussian variance, is the phase, ω is the length of the sine wave, x, y are the pixel coordinate positions, and x′, y′ represent the x, y pixel coordinate positions after the parallelogram rule respectively;
[0127] θ in formula (5) and formula (6) is the angle between two intersecting lines in the parallelogram rule. The above parameters are all parameters that can be directly learned by the system. The real part of the Gabor filter can be expressed as:
[0128]
[0129] The real part of the Gabor filter is an effective way to extract image texture. By inputting an image containing texture information into the Gabor layer, we can obtain the texture features of the image. I is the input image.
[0130]
[0131] (3) Edge feature extraction:
[0132] The reversible network model (INN module) can effectively preserve the original features of the image during the training process. The model performs well not only in extracting edge features, but also in capturing other high-frequency features. Combining the features extracted by the INN module with the texture features obtained by the GaborCNN layer can preserve more detailed information in the synthesized image.
[0133] First, the input of each reversible network module is divided into two parts along the channel axis, and a mapping function is introduced to ensure lossless transmission of information in the reversible layer. The MFF model uses the Bottleneck Residual Block (BRB) in MobileNetV2 as its mapping function.
[0134] I=I[1:a]+I[a+1:A] (9)
[0135]
[0136] Formula (9) to Formula (12) describe the extraction process of high-frequency edge features of an image in the present invention, where I is the input image, A is the total number of channels, BRB is the bottleneck residual block, ⊙ is the Hadamard product, is the extracted edge feature. Since the edge feature extraction of the INN module is divided into multiple times, represents the edge feature of the k+1th iteration; a is the channel value corresponding to the channel axis;
[0137] The final high-frequency features It can be expressed as:
[0138]
[0139] In the present invention, the extracted high-frequency features and low-frequency features are merged by merging the features of the two multimodal images to be fused according to the corresponding high and low frequencies. Infrared and visible light modes are used as examples for introduction.
[0140] Secondly, the merge strategy module is introduced;
[0141] The merge strategy can be expressed as follows:
[0142] The merging strategy combines the features of visible light image and infrared image extracted by the encoding module.
[0143]
[0144] To represent the operation of the merging strategy, the present invention adopts the calculation form of addition. is the high-frequency feature of the visible light image, is the high-frequency feature of the infrared image, is the low-frequency feature of the visible light image, Fuse is the low-frequency feature of infrared image. High Fuse is the combined high-frequency features of the visible light image and the infrared image. Low It is the combined low-frequency feature obtained by fusing the low-frequency features of the visible light image with the low-frequency features of the infrared image.
[0145] Finally, the decoder module is introduced:
[0146] In the decoder part, the Transformer model is used for feature decoding, and the convolution structure is combined to realize the restoration of the fused image. Due to the two-stage training method, the decoder module is also divided into two parts. Taking the input images of infrared and visible light as examples, in the first stage, the high-frequency and low-frequency features of the visible light image and the high-frequency and low-frequency features of the infrared image are decoded respectively. In the second stage, the high-frequency fusion features of the infrared image and the visible light image are decoded, and the low-frequency fusion features of the visible light image and the infrared image are decoded. In the first stage, the feature extraction unit extracts high-frequency and low-frequency features from the visible light image (or infrared image), inputs them into the decoding unit, and obtains the final reconstructed image. In the second stage, a pair of visible light and infrared images are input into the encoder module to extract the corresponding high-frequency and low-frequency features. The comprehensive features of the visible light image (or infrared image) are obtained through the fusion layer, and then the fused features are sent to the decoding module for image synthesis and reconstruction.
[0147] The loss function of the model proposed in this invention is defined as follows:
[0148] The loss function is also divided into two stages. Taking the input images of the two modalities of infrared and visible light as an example, the first stage extracts high-frequency and low-frequency features from the visible light image and the infrared image respectively. These features are then fused in their respective domains to reconstruct the corresponding infrared and visible light images. This stage focuses on training the encoder and decoder. In the second stage, the high-frequency features of the infrared and visible light images are combined to obtain fused high-frequency features, and the low-frequency features also undergo a similar fusion process. Then, the high-frequency and low-frequency features are fused to enhance the image quality. This stage emphasizes the fusion ability of the training model. Therefore, the loss function is also divided into two components, each of which corresponds to one of the above training stages.
[0149] L MFF =L Ⅰ +L Ⅱ (15)
[0150] L MFF is the overall loss function of the MFF model, L Ⅰ is the loss function of the first stage, L Ⅱ is the loss function of the second stage. For the loss function L of the first stage Ⅰ It can be expressed as follows:
[0151] L Ⅰ =L i +αL v +βL cc (16)
[0152] In formula (16), the first two terms are the reconstruction losses of infrared images and visible light images, and the third term is the feature decomposition loss. irepresents the reconstruction loss of infrared image, L v represents the reconstruction loss of the visible light image, L cc represents the relevant loss. α and β are weight parameters.
[0153] The reconstruction loss function includes the mean square error loss Lmse, the structural similarity index measurement loss Lssim and the gradient loss Lgrad. Combined with formula (16), the above three losses can be expressed as:
[0154]
[0155] In formula (17), Ir represents the original visible light image, Represents the reconstructed visible light image. Similarly, Vi and Represent the original and reconstructed infrared images respectively. In formula (16), Lcc can be expressed as:
[0156]
[0157] In formula (18), CC is the correlation coefficient operator and ε is the bias.
[0158] For the loss function of the second stage:
[0159] L Ⅱ =L in +ωL max_grad +τL cc (19)
[0160] L in Indicates strength loss, L max_grad represents the maximum gradient loss, L cc Indicates the associated loss.
[0161] L in =‖Fuse Image -Max(Ir,Vi)‖
[0162]
[0163] In formula (20), Fuse Image is the final synthetic infrared and visible light image fusion result. represents the Sobel gradient operator, CC is the correlation coefficient operator, and Max is the maximum value operation; in formula (19), ω and τ represent weight parameters.
[0164] Fusion result analysis:
[0165] The present invention uses the MSRS dataset to train the model, which is divided into a training set and a test set in a ratio of 3:1. The training set consists of 1083 pairs of images, each pair of images consists of an infrared image and a visible light image. In addition, we use the RoadScene dataset (containing 50 pairs of images) and the TNO dataset (containing 25 pairs of images) as part of our test set. The present invention analyzes the fusion results in two parts. The first part is a subjective evaluation of the quality of the fused image, and the second part is a quantitative analysis of the experimental results. Seven quality indicators, including entropy (EN), standard deviation (SD), spatial frequency (SF), mutual information (MI), difference correlation sum (SCD), QAB / F and structural similarity index (SSIM), are used to quantitatively compare our fusion method with other existing fusion methods. The visual fusion quality comparison is a subjective evaluation of the image fusion quality through visual evaluation, which mainly analyzes the fusion quality of the color distribution, noise points, image details and other aspects after fusion. The comparison models are MFEIF, DenseFuse, CDDFuse, U2Fusion, NestFuse, DIDFuse and ReCoNet. Figures 3 to 8 The fusion results and details of several images in the three datasets are shown.
[0166] The quantitative analysis is to compare the proposed MFF model with the baseline model through the validation indicators. The evaluation indicators of TNO, RoadScene and MSRS datasets are shown in Tables 1 to 3. Regarding running a fusion, the time consumption comparison results of each model are shown in Table 4.
[0167] Table 1 Quantitative analysis results of TNO dataset
[0168]
[0169] It can be seen from Table 1 that when the comparison indicators are entropy (EN), standard deviation (SD), spatial frequency (SF), mutual information (MI), correlation sum of differences (SCD), QAB / F and structural similarity index (SSIM), the best results of the MFF model of the present invention are 7.13, 46.87, 14.15, 2.20, 1.84, 0.55, and 0.67 respectively; the next better results are 7.12, 46.02, 13.15, 2.19, 1.78, 0.54, and 0.66 respectively.
[0170] Table 2 Quantitative analysis results of RoadScene dataset
[0171]
[0172] It can be seen from Table 2 that when the comparison indicators are entropy (En), standard deviation (SD), spatial frequency (SF), mutual information (MI), correlation sum of differences (SCD), QAB / F and structural similarity index (SSIM), the best results of the MFF model of the present invention are 7.44, 60.84, 16.65, 2.35, 1.87, 0.54, and 0.67 respectively; the next better results are 7.42, 54.63, 16.34, 2.31, 1.84, 0.53, and 0.66 respectively.
[0173] Table 3 Quantitative analysis results of MSRS dataset
[0174]
[0175] It can be seen from Table 3 that when the comparison indicators are entropy (En), standard deviation (SD), spatial frequency (SF), mutual information (MI), correlation sum of differences (SCD), QAB / F and structural similarity index (SSIM), the best results of the MFF model of the present invention are 6.71, 43.45, 11.59, 3.43, 1.68, 0.69, and 0.7 respectively; the next better results are 6.70, 43.38, 11.55, 3.21, 1.62, 0.55, and 0.69 respectively.
[0176] Table 4 Comparison of computational complexity
[0177]
[0178] As shown in Table 4, in terms of complexity comparison, the best parameter value result of the MFF model of the present invention is 0.007, and the best time result is 0.024; the next better result is: the parameter value result is 0.074, and the best time result is 0.053. The next better result is: the parameter value result is 0.151, and the best time result is 0.055.
[0179] The multi-filter-based multimodal image fusion deep network model of the present invention can not only be used to improve the image fusion effect, but also have an improvement effect on downstream operations after fusion, such as target tracking. Fig. 9 shown.
[0180] The multi-modal image fusion deep network model based on multiple filters of the present invention is also universal for medical image fusion, such as CT-MRI image fusion, MRI-PET image fusion, and PET-SPECT image fusion. Figures 10 to 12 shown.
[0181] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.
Claims
1. A multi-modal image fusion deep network model based on multiple filters, comprising an encoder module, a merging strategy module and a decoder module connected in sequence; It is characterized in that The encoder module includes two sets of feature extraction modules; Each set of feature extraction modules includes LEF module, GaborCNN layer and INN module; The LEF module is used for extracting low-frequency features of images; The GaborCNN layer is used to extract texture features from high-frequency features of the image; The INN module is used for edge feature extraction in high-frequency features of images; The merging strategy module adopts an addition strategy to fuse the high-frequency features and low-frequency features extracted from the multimodal image pair by the feature extraction module; The decoder module uses the Transformer model to perform feature decoding and combines the convolution structure to achieve fused image restoration.
2. The multi-filter based multimodal image fusion deep network model according to claim 1, characterized in that: The steps of using the LEF module to extract low-frequency features of an image are as follows: First, a 1×1 convolution block is used to adjust the number of channels, transforming the input tensor shape from R h×w×1 Convert to R h×w×32 ; Adopt multi-scale structure adaptive average pooling layer to achieve low-frequency filtering, and use upsampling technology to restore the uniformity of feature scale; The input dimension is divided into four equal parts, and the feature F obtained by the first convolution is expressed as: F={f1,f2,f3,f4} (1) f i ∈F, i = 1, 2, 3, 4, with the same shape; for each sub-feature f of feature F i Perform multi-scale pooling operations and use a bilinear interpolation sampling algorithm to upsample the pooling results for each sub-feature f i Generate the corresponding feature η i ; The process formula is as follows: or i =Up(β s (f i )) (2) s represents the size of the pooling layer. i Four pooling operations are required, and the size of the pooling layer in each operation is s∈{1,2,3,6}; β s represents a pooling operation with a size of s∈{1,2,3,6}; Up represents an upsampling operation; Finally, the four sub-features f i The extracted features are connected and the ReLU activation function is applied; then, they are adjusted through the convolution layer Conv to meet the shape requirements of subsequent high-frequency features; finally, the extracted features are connected and then the ReLU activation function is applied; then, they are refined through the convolution layer Conv to meet the size requirements of high-frequency features; the low-frequency features finally obtained are It is expressed as:
3. The multi-filter based multimodal image fusion deep network model according to claim 1, characterized in that: The steps of using the GaborCNN layer to extract texture features from high-frequency features of an image are as follows: The two-dimensional Gabor filter is a Gaussian kernel function modulated by a sine wave. The function formula is as follows: Where x′ and y′ are expressed as: x′=xcosθ+ysinθ (5) y′=-xcosθ+ycosθ (6) In formula (4), σ is the Gaussian variance, is the phase, ω is the length of the sine wave; x, y are the pixel coordinate positions, x′, y′ represent the x, y pixel coordinate positions after the parallelogram rule respectively; In formula (5) and formula (6), θ is the angle between two intersecting lines in the parallelogram rule; The real part of the two-dimensional Gabor filter is expressed as: The real part of the two-dimensional Gabor filter is an effective method to extract image texture. By inputting an image containing texture information into the GaborCNN layer, the texture features of the image can be obtained. I is the input image; 4. The multi-filter based multimodal image fusion deep network model according to claim 1, characterized in that: The steps of the INN module for edge feature extraction in high-frequency features of an image are as follows: First, the input of each reversible module is divided into two parts along the channel axis, and a mapping function is introduced, using the Bottleneck Residual Block (BRB) in MobileNetV2 as its mapping function; I=I[1:a]+I[a+1:A] (9) In formula (9) to formula (12), I is the input image, A is the total number of channels, BRB is the bottleneck residual block, ⊙ is the Hadamard product, is the extracted edge feature; since the edge feature extraction of the INN module is carried out in multiple times, Represents the edge features of the k+1th iteration; a is the channel value corresponding to the channel axis; The final high-frequency features It is expressed as:
5. The multi-filter based multimodal image fusion deep network model according to claim 1, characterized in that: When the model is used for the fusion of infrared and visible light images, the merging strategy module uses an addition strategy to fuse the high-frequency features and low-frequency features extracted from the multi-modal image pair, which is expressed as: in, represents the operation of the merge strategy, is the high-frequency feature of the visible light image, is the high-frequency feature of the infrared image, is the low-frequency feature of the visible light image, is the low-frequency feature of the infrared image, Fuse High Fuse is the combined high-frequency features of the visible light image and the infrared image. Low It is the combined low-frequency feature obtained by fusing the low-frequency features of the visible light image with the low-frequency features of the infrared image.
6. The multi-filter based multimodal image fusion deep network model according to claim 5, characterized in that: The decoder module uses the Transformer model to perform feature decoding as follows: The decoder module uses the Transformer model to perform feature decoding in two stages; In the first stage, the high- and low-frequency features of the visible light image and the high- and low-frequency features of the infrared image are decoded respectively; In the second stage, the high-frequency fusion features of the infrared image and the visible light image are decoded, and the low-frequency fusion features of the visible light image and the infrared image are decoded.
7. The multi-filter based multimodal image fusion deep network model according to claim 6, characterized in that: The model is trained using the following loss function; The loss function is divided into two stages, expressed as follows: L MFF =L Ⅰ +L Ⅱ (15) L MFF is the overall loss function of the MFF model, L Ⅰ is the loss function of the first stage, L Ⅱ is the loss function of the second stage; The loss function L in the first stage Ⅰ It is expressed as follows: L Ⅰ =L i +αL v +βL cc (16) In formula (16), L i represents the reconstruction loss of infrared image, L v represents the reconstruction loss of the visible light image, L cc Represents the related loss, α and β are weight parameters; The reconstruction loss function includes the mean square error loss Lmse, the structural similarity index measurement loss Lssim and the gradient loss Lgrad; combined with formula (16), the above three losses are expressed as: In formula (17), Ir represents the original visible light image, represents the reconstructed visible light image; Vi and represent the original and reconstructed infrared images respectively; In formula (16), Lcc is expressed as: In formula (18), CC is the correlation coefficient operator and ε is the bias; The loss function L in the second stage Ⅱ : L Ⅱ =L in +ωL max_grad +τL cc (19) L in Indicates strength loss, L max_grad represents the maximum gradient loss, L cc Indicates the relevant loss; In formula (20), Fuse Image is the final synthetic infrared and visible light image fusion result. represents the Sobel gradient operator, CC is the correlation coefficient operator, and Max is the maximum value operation; In formula (19), ω and τ both represent weight parameters.
8. An application of the multi-filter based multimodal image fusion deep network model described in any one of claims 1 to 4 in target tracking or medical image fusion.