Infrared and visible light image fusion method and system based on wavelet transform feature decoupling
Through the wavelet transform feature decoupling method, the features of infrared and visible light images are extracted using the Restormer block and WTConv block, which solves the problem of insufficient feature extraction and decomposition in the prior art, and generates efficient image fusion results, which are suitable for application scenarios with resource limitations.
Patent Information
- Application Number
- CN202510357899.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-25
AI Technical Summary
The existing infrared and visible image fusion algorithms have shortcomings in feature extraction and decomposition, making it difficult to effectively extract and fuse multimodal features, and the deep learning model has high computational complexity in resource-constrained scenarios.
Using a method based on wavelet transform feature decoupling, shared features are extracted through the Restormer block, the DenseConv block and the WTConv block extract the low-frequency basic features and high-frequency detailed features respectively, and a multiplication and multiplication fusion module and two-stage training loss function are designed to improve the feature extraction and fusion effect.
It significantly improves the network's feature extraction ability, generates more information-rich and accurate fusion images, suitable for practical application scenarios with limited resources.
Smart Images

Figure CN120374413A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an infrared and visible light image fusion method, specifically an infrared and visible light image fusion method and system based on wavelet transform feature decoupling. Background Art
[0002] Infrared and visible light image fusion is a research topic with broad application prospects, involving multiple fields such as security monitoring, target recognition, and medical image processing. Different from traditional single-modal images, infrared images and visible light images each have unique information dimensions and characteristics: infrared images mainly reflect thermal radiation information, are suitable for low-light or completely dark environments, and can highlight thermal targets; while visible light images provide rich detail and texture information and are important data sources in daytime and normal lighting environments. However, in multi-modal fusion tasks, how to effectively extract, decompose, and fuse this information has always been a complex and challenging research problem. And current mainstream infrared and visible light image fusion algorithms often face the problem of being difficult to effectively extract and decompose multi-modal features when dealing with multi-modal features. Therefore, better extracting the complementary characteristics and global shared information of infrared and visible light images and improving the fusion effect is an extremely important and meaningful research topic. And with the significant progress of deep learning technology in image fusion tasks, its high computational cost limits the application of the model in actual scenarios (such as embedded devices or real-time processing scenarios). Therefore, how to compress the knowledge of complex network models into lightweight models, significantly reduce the computational complexity, and make the fusion algorithm more suitable for resource-constrained scenarios is also a problem with broad application scenarios and great significance.
[0003] In recent years, deep learning has dominated the development of the computer vision field with its powerful feature extraction and expression capabilities and demonstrated significant performance advantages in visual tasks such as image classification, object detection, and semantic segmentation. To overcome the deficiencies of traditional algorithms, researchers in the field of image fusion have also explored a large number of deep learning-based image fusion algorithms. Existing deep learning-based image fusion algorithms mainly focus on solving three key problems in image fusion: feature extraction, feature fusion, and image reconstruction. According to the adopted network architecture, deep learning-based image fusion algorithms can be divided into three categories: image fusion frameworks based on autoencoders (AE), image fusion frameworks based on convolutional neural networks (CNN), and image fusion frameworks based on generative adversarial networks (GAN).
[0004] 1. The image fusion framework based on the autoencoder (AE) first pre-trains an autoencoder on a large dataset to achieve feature extraction and image reconstruction. For example, the MS-COCO (Microsoft common objects in context) dataset, the ImageNet dataset. Then, a manually designed fusion strategy is adopted to integrate the deep features extracted from different source images to achieve image fusion. However, these manually designed fusion strategies may not be applicable to deep features, thus limiting the performance of the AE-based fusion framework.
[0005] 2. The image fusion framework based on the convolutional neural network (CNN) realizes end-to-end feature extraction, feature fusion, and image reconstruction by designing the network structure and loss function, thus avoiding the tediousness of manually designing fusion rules. The mainstream CNN-based image fusion framework constructs a loss function by measuring the similarity between the fused image and the source images to guide the end-to-end training of the network. There are also methods that use prior knowledge to design a pseudo-label image and the fused image to construct a loss function. In addition, some CNN-based methods use the convolutional neural network as part of the overall method for feature extraction or activity level measurement.
[0006] 3. The image fusion framework based on the generative adversarial network (GAN) models the image fusion problem as an adversarial game problem between the generator and the discriminator. The GAN-based image fusion framework forces the fused result generated by the generator to be consistent with the target distribution in terms of probability distribution through the discriminator, thus implicitly realizing feature extraction, fusion, and image reconstruction. Existing GAN-based fusion methods construct the target distribution through the source images or pseudo-label images. According to the supervised paradigm used in the training process, deep learning-based image fusion algorithms can also be divided into unsupervised image fusion frameworks, self-supervised image fusion frameworks, and supervised image fusion frameworks.
[0007] In summary, in the infrared and visible light image fusion, the feature extraction ability of the network plays a very important decisive role in the subsequent fusion effect. Summary of the Invention
[0008] In order to address the problems of cross-modal feature modeling and decomposing ideal specific-modal features and modal-shared features, and to improve the network feature extraction ability, the present invention proposes an infrared and visible light image fusion method based on wavelet transform feature decoupling, which alleviates the problem of incomplete frequency feature fusion in existing fusion networks, thereby obtaining a more informative and accurate fused image, and promoting the application of infrared and visible light image fusion technology in downstream advanced vision tasks.
[0009] According to one aspect of the specification of the present invention, there is provided an infrared and visible light image fusion method based on wavelet transform feature decoupling, comprising:
[0010] Obtain an infrared image and a visible light image;
[0011] Input the obtained infrared image and visible light image into the trained image fusion model to output an image fusion result; wherein, the training of the image fusion model includes:
[0012] Construct a dataset of infrared images and visible light images;
[0013] Construct an image fusion model based on wavelet transform feature decoupling, including: an encoder module for extracting shallow features from the input infrared image and visible light image, and respectively extracting the low-frequency basic features and high-frequency detail features of the two modalities from the shallow features based on wavelet transform feature decoupling; a fusion module for fusing the low-frequency basic features and high-frequency detail features of the two modalities; a decoder module for obtaining the image fusion result based on the fused features;
[0014] Design a loss function, and perform two-stage training on the model based on the constructed dataset to output the trained image fusion model.
[0015] As a further technical solution, the construction of the encoder module further includes:
[0016] Construct a shared feature encoder based on the Restormer block for extracting shallow features from the input infrared image and visible light image;
[0017] Construct a Base Transformer encoder based on the DenseConv block for respectively extracting the low-frequency basic features of the two modalities from the shallow features;
[0018] Construct a Detail Transformer encoder based on the WTConv block for respectively extracting the high-frequency detail features of the two modalities from the shallow features.
[0019] As a further technical solution, when respectively extracting the low-frequency basic features or high-frequency detail features of the two modalities from the shallow features, it further includes:
[0020] Use wavelet transform to filter and downsample the output low-frequency and high-frequency contents;
[0021] Before constructing the output using wavelet inverse transform, perform small kernel depth convolution on different frequency maps, and use the cascade principle to combine multiple single convolution operations performed in the wavelet domain.
[0022] As a further technical solution, the construction of the fusion module includes:
[0023] Construct a fusion module based on the multiplication-addition fusion idea.
[0024] As a further technical solution, when the decoder module performs decoding, it includes:
[0025] Concatenate the decomposed features as input in the channel dimension, and concatenate the original image as input in the channel dimension.
[0026] As a further technical solution, the structure of the decoder module is the same as that of the shared feature encoder.
[0027] As a further technical solution, the two-stage training of the image fusion model includes:
[0028] In training stage 1, the infrared image and the visible light image are input into the shared feature encoder to extract shallow features; the Base Transformer encoder based on the DenseConv block and the Detail Transformer encoder based on the WTConv block are used to extract the low-frequency basic features and high-frequency detail features of the two different modalities respectively; the basic features of the infrared image or the basic features and detail features of the visible light image are incorporated into the decoder to reconstruct the original infrared image or visible light image;
[0029] In training stage 2, the paired infrared image and visible light image are input into the trained encoder to obtain decomposed features; the decomposed basic features and detail features are respectively input into the fusion layer; the fused features are input into the decoder to obtain the fused image.
[0030] According to one aspect of the specification of the present invention, there is provided an infrared and visible light image fusion system based on wavelet transform feature decoupling, including:
[0031] An image input module for acquiring an infrared image and a visible light image;
[0032] An image fusion module for inputting the acquired infrared image and visible light image into the trained image fusion model and outputting an image fusion result; wherein, the training of the image fusion model includes:
[0033] Construct a data set of infrared images and visible light images;
[0034] Construct an image fusion model based on wavelet transform feature decoupling, including: an encoder module for extracting shallow features from the input infrared image and visible light image, and respectively extracting the low-frequency basic features and high-frequency detail features of the two modalities from the shallow features based on wavelet transform feature decoupling; a fusion module for performing feature fusion on the low-frequency basic features and high-frequency detail features of the two modalities; a decoder module for obtaining an image fusion result based on the fused features;
[0035] Design a loss function, perform two-stage training on the model based on the constructed dataset, and output a trained image fusion model.
[0036] According to one aspect of the specification of the present invention, there is provided an infrared and visible light image fusion device based on wavelet transform feature decoupling, including a memory and a processor. The memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the infrared and visible light image fusion method based on wavelet transform feature decoupling.
[0037] According to one aspect of the specification of the present invention, there is provided a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the infrared and visible light image fusion method based on wavelet transform feature decoupling.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0039] The present invention proposes an infrared and visible light image fusion method based on wavelet transform feature decoupling. First, a shared feature encoder (ResBlock) based on the Restormer block is designed to obtain the shallow features of infrared and visible light images. Secondly, a Base Transformer of the DenseConv block is designed, and this module extracts features from infrared and visible light images respectively to ensure that the extracted features contain the main features of the source images. Then, a DetailTransformer frequency feature generation module based on the WTConv block is used to focus on extracting the detail features in the source images, thereby effectively improving the network feature extraction ability and having a better fusion effect. Finally, an image fusion module based on the addition and multiplication idea is designed to improve the effect of image fusion. Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings used in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a schematic flowchart of the infrared and visible light image fusion method based on wavelet transform feature decoupling provided by the embodiment of the present invention.
[0042] Figure 2 It is a schematic network structure diagram of the infrared and visible light image fusion method based on wavelet transform feature decoupling provided by the embodiment of the present invention. Detailed Embodiments
[0043] In the description, claims and above-mentioned drawings of the present invention, the terms "comprising", "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0044] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. In addition, the technical features in each embodiment or individual embodiment provided by the present invention can be combined arbitrarily with each other to form a new technical solution. This combination is not restricted by the order of steps and / or the mode of structural composition, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0045] An infrared and visible light image fusion method based on wavelet transform feature decoupling is provided in an embodiment of the present invention. As Figure 1 shown, first, an infrared image and a visible light image are obtained; then, the obtained infrared image and visible light image are input into a trained image fusion model, and an image fusion result is output.
[0046] Among them, the training of the image fusion model specifically includes the following steps:
[0047] Step 1: Using the infrared image and the visible light image as inputs respectively, dividing the data into a training set and a test set.
[0048] Specifically, two common benchmarks are used in the experiment to verify the described image fusion model, namely the MSRS dataset and the TNO dataset. The image fusion model is trained on the MSRS training set (1083 data pairs), and 50 data pairs in TNO are used for verification. The test dataset includes the MSRS test set (361 pairs) and the TNO test set (25 pairs), which can comprehensively verify the fusion performance. Eight metrics are used in the embodiments of the present invention to quantitatively measure the fusion result: entropy (EN), standard deviation (SD), spatial frequency (SF), mutual information (MI), sum of differential correlations (SCD), visual information fidelity (VIF), QAB / F, and structural similarity index (SSIM). The higher the values of these metrics, the better the fusion effect.
[0049] Step 2: As shown in Figure 2 , design an infrared and visible light image fusion network with wavelet transform feature decoupling to achieve image fusion, including:
[0050] 1) Design an encoder module: The encoder module consists of three parts: a shared feature encoder (ResBlock) based on the Restormer block, a Base Transformer based on the DenseConv block, and a DetailTransformer based on the WTConv block.
[0051] For a clearer presentation, some symbols are defined below. The shared feature encoder, the Base Transformer encoder based on the DenseConv block, and the Detail Transformer encoder based on the WTConv block are denoted as S(·), B(·), and D(·), respectively.
[0052] Shared feature encoder (ResBlock) based on the Restormer block: The purpose of the shared feature encoder is to extract shallow features from the input infrared and visible light images {I ir , I vi}, that is: Namely:
[0053]
[0054] In terms of network hyperparameter settings, the number of Restormer blocks in the shared feature encoder is 4, with 8 attention heads and 64 dimensions. The Restormer block can extract global features from high-resolution input images through self-attention across feature dimensions. Therefore, it can be used to extract shallow cross-modal features.
[0055] Base Transformer encoder based on the DenseConv block: The purpose of the Base Transformer encoder is to extract low-frequency basic features from the shared features, that is, shallow features:
[0056]
[0057] Where respectively represent the basic features of the infrared image I ir and the basic features of the visible light image I vi . It can be seen from Figure 2 that two 3×3 convolutional layers and a common convolutional layer with a kernel size of 1×1 are deployed in the main dense stream of the Base Transformer encoder.
[0058] It should be emphasized that in this part, dense connections are introduced to make full use of the features extracted by each convolutional layer. The residual stream is the input. Then, the outputs of the main dense stream and the residual stream are added through element-wise addition to integrate the infrared image I ir with the visible light image I vi of the basic features. And placing the self-attention module after the feature extraction module can enhance the global feature modeling ability, effectively reduce the computational complexity, reduce resource occupancy, and improve the running efficiency of the model. The self-attention module can enhance important features and suppress irrelevant features through the weight mechanism, thereby improving the distinguishability of features and further enhancing the stability of the model.
[0059] Detail Transformer Encoder Based on WTConv Block: The Detail Transformer encoder extracts high-frequency detail information from the shared features, and its formula is:
[0060]
[0061] Considering that the edge and texture information in the detail features are very important for the image fusion task, we hope that the WTConv block in the Detail Transformer encoder can retain as much detail information as possible. Therefore, wavelet transform is used to expand the receptive field of the Transformer, so as to effectively extract the high-frequency information of the image.
[0062] The present invention uses convolution to perform wavelet transform. Here, Haar wavelet transform is used, which has less computational cost compared to other wavelet bases. For the shallow features of infrared and visible light images A single-level Haar wavelet transform in one spatial dimension (width or height) is given by a depth convolution with kernels and plus a standard downsampling operator with a coefficient of 2. To perform a two-dimensional Haar wavelet transform, we perform a combined operation in two dimensions and use the following four filters for depth convolution with a stride of 2:
[0063]
[0064] where f LL is a low-pass filter, while f LH , f HL , f HH are a set of high-pass filters. For each input channel, the output of the convolutional filter is
[0065] [Φ LL , Φ LH , Φ HL , Φ HH = Conv([fLL , f LH , f HL , f HH , Φ)
[0066] where Φ LL is the low-frequency component of the input Φ, while Φ LH , Φ HL , Φ HH are the horizontal, vertical, and diagonal high-frequency components of the input Φ. Since the kernels of the above four filters form an orthogonal basis, when applying the inverse wavelet transform (IWT), it can be obtained by transposed convolution:
[0067] Φ = Conv-transposed([f LL , f LH , f HL , f HH , [Φ LL , Φ LH , Φ HL , Φ HH )
[0068] Then, the cascaded wavelet decomposition is given by recursively decomposing the low-frequency classification. The decomposition formula for each level is as follows:
[0069]
[0070] where and i is the current level number. This results in an increase in frequency resolution while a decrease in the spatial resolution of the low frequency.
[0071] Increasing the size of the convolution layer kernel causes the number of parameters and degrees of freedom to increase quadratically. Therefore, the wavelet transform is first used to filter and downsample the low-frequency and high-frequency content of the output. Then, before using the inverse wavelet transform to construct the output, small kernel depth convolution is performed on different frequency maps. The entire process is as follows:
[0072] Φ out = IWT(Conv(W, WT(Φ in )))
[0073] where Φ in are the base features of the infrared image I ir and the base features of the visible light image I vi respectively. W is the weight tensor of the k×k depth kernel, and the number of its input channels is four times that of Φ in . This operation can not only separate the convolution between frequency components but also allow smaller kernels to work over a larger range of the original input, that is, increase their receptive field for the input.
[0074] And by using the cascading principle in the above content, this combined operation at this level is further improved. The process is as follows:
[0075]
[0076] Among them, represents the input of this layer, represents all three high-frequency mappings of the i-th layer.
[0077] To merge the outputs of different frequencies, the fact that wavelet transform and inverse wavelet transform are linear operations is utilized. That is, IWT(X + Y) = IWT(X) + IWT(Y). Therefore, the result of the execution is the convolution addition of different layers.
[0078]
[0079] Among them, is the summary output after the i-th layer.
[0080] 2) Design an image fusion module based on the idea of addition-multiplication fusion: The goal of the fusion module is to obtain a fused feature map with the same feature size. For the cross-modal complementary information and cross-modal common information obtained by the feature extraction module, a fusion method that can efficiently retain its effective information is required. Usually, addition operation is used to fuse the cross-modal complementary information, while multiplication operation is used to fuse the cross-modal common information.
[0081] The present invention integrates using the idea of addition-multiplication fusion and reduces the number of channels through 1×1 convolution. It is specifically represented by the following formula:
[0082]
[0083] where Cat represents the direct merging operation. And channel attention is used to better select the common features and unique features.
[0084] 3) Design the decoder module: In the decoder DC(·), the decomposed features are concatenated in the channel dimension as the input, and then the original image is concatenated in the channel dimension as the input, which can be represented by the following formula:
[0085]
[0086] Since the input here involves cross-modal and multi-frequency characteristics, the present invention keeps the structure of the decoder consistent with the design of the shared feature encoder, that is, using the Restormer module as the basic unit of the decoder.
[0087] In terms of network hyperparameter settings, it is kept consistent with the shared feature encoder. The number of Restormer blocks is 4, with 8 attention heads and 64 dimensions.
[0088] Step 3: Perform two-stage training on the image fusion network in Step 2.
[0089] In training stage 1, the infrared image and the visible light image {I ir , I vi} are input into the shared feature encoder to extract shallow features Then, the Base Transformer based on the DenseConv block and the Detail Transformer based on the WTConv block are used to extract the low-frequency basic features of the two different modalities and the high-frequency detail features Then, the basic features of the infrared image (or the basic features of the visible light image ) and the detail features are incorporated into the decoder to reconstruct the original infrared image (or the visible light image ).
[0090] In training stage 2, the paired infrared image and visible light image {I ir , I vi} are input into the trained encoder to obtain the decomposed features. Then the decomposed basic features and the detail features are respectively input into the fusion layers F B and F D . Finally, the fused features {Φ B , Φ D} are input into the decoder to obtain the fused image F.
[0091] It should be noted that the embodiment of the present invention adopts two-stage training and trains with the original images in both stages, which has the following advantages:
[0092] (1) Information integrity: The original infrared image and visible light image contain complete modal information. Directly inputting the original images can ensure the integrity and accuracy of information in the feature extraction process, enabling the encoder to capture richer and more real image features.
[0093] (2) Avoiding error accumulation: The reconstructed images and in the first stage may have certain reconstruction errors. If these reconstructed images are used as inputs in the second stage, the errors may be propagated to the fusion layer, affecting the fusion effect, while directly using the original images avoids this accumulation of errors.
[0094] (3) Preserve the original features: The features of the original image are not processed by the decoder, which can more truly reflect the original content of the image. This helps the fusion layer better learn how to effectively fuse the features of different modalities, thereby generating a higher-quality fused image.
[0095] Loss function:
[0096] In training stage 1, the total loss function is:
[0097]
[0098] where and are the reconstruction losses of the infrared image and the visible light image, is the feature decomposition loss, and α1 and α2 are adjustment parameters. The reconstruction loss mainly ensures that the information contained in the image is not lost during the encoding and decoding processes, that is
[0099]
[0100] where SSIM(·,·) is the structural similarity index.
[0101]
[0102] where CC(·,·) is the correlation coefficient operator, and ∈ is set to 1.01 here to ensure that this term is always positive.
[0103] The reason for generating this loss term is that the decomposed features will contain more modality-shared information, such as the background and large-scale environment, so they tend to be highly correlated. In contrast, represents texture and detail information in the visible light image and thermal radiation and clear edge information in the infrared image, which are all modality-specific information. Empirically, under the guidance of in gradient descent, gradually approaches 0, while becomes larger and larger, which conforms to the intuition of feature decomposition.
[0104] Subsequently, in training stage 2, the total loss becomes:
[0105]
[0106] where represents the Sobel gradient operator, and α3 and α4 are adjustment parameters.
[0107] During the network training process, the parameter settings are shown in Table 1.
[0108] Table 1 Training Parameters
[0109] Name Parameter CPU Intel Core i5-14600KF GPU GeForce RTX 2080Ti 11GB Deep learning framework PyTorch 2.1 Training set MSRS (6482 pairs) Training set image size 120×120 Network optimizer Adam Learning rate <![CDATA[1×10 -4 > Number of training epochs 40 epochs in stage I, 60 epochs in stage II Loss function weight <![CDATA[λ = 5, μ = 5, α1 = 2, α2 = 5, α3 = 5, α4 = 1, α5 = 10]]> Test set MSRS (30 pairs), TNO (42 pairs), LLVIP (50 pairs)
[0110] In the test stage, the trained image fusion network model is used to fuse the image data of the test set.
[0111] The implementation basis of each embodiment of the present invention is realized through programmed processing by a device with processor functions. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, an embodiment of the present invention provides an infrared and visible light image fusion system based on wavelet transform feature decoupling, which is used to execute the infrared and visible light image fusion method based on wavelet transform feature decoupling in the above method embodiments.
[0112] The system includes: an image input module for acquiring infrared images and visible light images; an image fusion module for inputting the acquired infrared images and visible light images into the trained image fusion model and outputting an image fusion result; wherein, the training of the image fusion model includes: constructing a data set of infrared images and visible light images; constructing an image fusion model based on wavelet transform feature decoupling, including: an encoder module for extracting shallow features from the input infrared images and visible light images, and respectively extracting low-frequency basic features and high-frequency detail features of the two modalities from the shallow features based on wavelet transform feature decoupling; a fusion module for performing feature fusion on the low-frequency basic features and high-frequency detail features of the two modalities; a decoder module for obtaining an image fusion result based on the fused features; designing a loss function, and performing two-stage training on the model based on the constructed data set to output a trained image fusion model.
[0113] The infrared and visible light image fusion system based on wavelet transform feature decoupling provided by the embodiment of the present invention faces the problem of incomplete frequency feature fusion in the existing fusion network. By using the foregoing several modules, it effectively extracts local features of time and frequency from the signal by combining wavelet transform, significantly improving the network feature extraction ability.
[0114] It should be noted that the system embodiments provided by the present invention are used not only to implement the methods in the above method embodiments, but also to implement the methods in other method embodiments provided by the present invention. The difference lies only in setting corresponding functional modules. The principle is basically the same as that of the above system embodiments provided by the present invention. As long as those skilled in the art, on the basis of the above system embodiments, refer to the specific technical solutions in other method embodiments, obtain corresponding technical means by combining technical features, and the technical solutions constituted by these technical means, and improve the modules in the above system embodiments on the premise of ensuring the practicability of the technical solutions, corresponding system-like embodiments can be obtained to implement the methods in other method-like embodiments.
[0115] Based on the same inventive concept as the foregoing embodiments, the embodiments of the present invention further provide an infrared and visible light image fusion device based on wavelet transform feature decoupling, including a memory and a processor. The memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the infrared and visible light image fusion method based on wavelet transform feature decoupling.
[0116] Based on the same inventive concept as the foregoing embodiments, the embodiments of the present invention further provide a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the infrared and visible light image fusion method based on wavelet transform feature decoupling.
[0117] In summary of the above embodiments, the present invention relates to the field of computer vision technology. In order to address the problems of cross-modal feature modeling and decomposing ideal specific-modal features and modality-shared features, the present invention discloses an infrared and visible light image fusion method based on wavelet transform feature decoupling. First, the infrared image and the visible light image are respectively used as inputs, and the data is divided into a training set and a test set. The Restormer block is used to extract cross-modal shallow features. Then, a dual-branch feature extractor is used to extract features from the images. Among them, the Base Transformer block uses DenseConv to process low-frequency global features, while the Detail Transformer block uses the relevant characteristics of wavelet transform to extract high-frequency local features. A loss function is designed to supervise the training process of the network model; in the test stage, the infrared and visible light source images are input, and the network will output the final image fusion result. The present invention combines wavelet transform to effectively extract local features of time and frequency from signals, significantly improving the network feature extraction ability.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. An infrared and visible light image fusion method based on wavelet transform feature decoupling, characterized in that, Including: Obtain infrared images and visible light images; Input the obtained infrared images and visible light images into the trained image fusion model, and output the image fusion result; wherein, the training of the image fusion model includes: Construct a dataset of infrared images and visible light images; Construct an image fusion model based on wavelet transform feature decoupling, including: an encoder module for extracting shallow features from the input infrared images and visible light images, and respectively extracting the low-frequency basic features and high-frequency detail features of the two modalities from the shallow features based on wavelet transform feature decoupling; a fusion module for performing feature fusion on the low-frequency basic features and high-frequency detail features of the two modalities; a decoder module for obtaining the image fusion result based on the fusion features; Design a loss function, and perform two-stage training on the model based on the constructed dataset to output the trained image fusion model.
2. The infrared and visible light image fusion method based on wavelet transform feature decoupling according to claim 1, wherein The construction of the encoder module further includes: Construct a shared feature encoder based on the Restormer block for extracting shallow features from the input infrared images and visible light images; Construct a Base Transformer encoder based on the DenseConv block for respectively extracting the low-frequency basic features of the two modalities from the shallow features; Construct a Detail Transformer encoder based on the WTConv block for respectively extracting the high-frequency detail features of the two modalities from the shallow features.
3. The infrared and visible light image fusion method based on wavelet transform feature decoupling according to claim 1, wherein When respectively extracting the low-frequency basic features or high-frequency detail features of the two modalities from the shallow features, it further includes: Use wavelet transform to filter and downsample the output low-frequency and high-frequency contents; Before constructing the output using inverse wavelet transform, perform small kernel depth convolution on different frequency maps, and use the cascade principle to combine multiple single convolution operations performed in the wavelet domain.
4. The infrared and visible light image fusion method based on wavelet transform feature decoupling according to claim 1, wherein The construction of the fusion module includes: Construct a fusion module based on the addition and multiplication fusion idea.
5. The infrared and visible light image fusion method based on wavelet transform feature decoupling according to claim 1, characterized in that When the decoder module performs decoding, it includes: Concatenate the decomposed features as inputs in the channel dimension, and concatenate the original images as inputs in the channel dimension.
6. The infrared and visible light image fusion method based on wavelet transform feature decoupling according to claim 2, wherein The structure of the decoder module is the same as that of the shared feature encoder.
7. The infrared and visible light image fusion method based on wavelet transform feature decoupling according to claim 2, wherein The two-stage training of the image fusion model includes: In training stage 1, the infrared images and visible light images are input into the shared feature encoder to extract shallow features; use the Base Transformer encoder based on the DenseConv block and the Detail Transformer encoder based on the WTConv block to respectively extract the low-frequency basic features and high-frequency detail features of the two different modalities; incorporate the basic features of the infrared images or the basic features and detail features of the visible light images into the decoder to reconstruct the original infrared images or visible light images; In training stage 2, input the paired infrared images and visible light images into the trained encoder to obtain the decomposed features; input the decomposed basic features and detail features into the fusion layer respectively; input the fused features into the decoder to obtain the fused image.
8. Infrared and visible light image fusion system based on wavelet transform feature decoupling, characterized in that Including: An image input module for obtaining infrared images and visible light images; An image fusion module, configured to input the acquired infrared image and visible light image into a trained image fusion model, and output an image fusion result; wherein the training of the image fusion model includes: Constructing a data set of infrared images and visible light images; Constructing an image fusion model based on wavelet transform feature decoupling, including: an encoder module, configured to extract shallow features from the input infrared image and visible light image, and respectively extract the low-frequency basic features and high-frequency detail features of the two modalities from the shallow features based on wavelet transform feature decoupling; a fusion module, performing feature fusion on the low-frequency basic features and high-frequency detail features of the two modalities; a decoder module, configured to obtain an image fusion result based on the fused features; Designing a loss function, and performing two-stage training on the model based on the constructed data set to output a trained image fusion model.
9. An infrared and visible light image fusion device based on wavelet transform feature decoupling, characterized in that It includes a memory and a processor, the memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the infrared and visible light image fusion method based on wavelet transform feature decoupling according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the infrared and visible light image fusion method based on wavelet transform feature decoupling according to any one of claims 1 to 7.
Citation Information
Cited By
Visible light-thermal infrared scene understanding method in intelligent automobile based on display frequency decoupling
CN121438264A
A visible-thermal infrared scene understanding method based on explicit frequency decoupling in intelligent vehicles
CN121438264B