Infrared light and visible light image fusion method based on diffusion model
Through the image fusion method based on the diffusion model, the fusion prior image of infrared light and visible light images is solved, and the problems of instability and pattern collapse in image fusion are generated, and high-quality fusion images are suitable for advanced visual tasks.
Patent Information
- Application Number
- CN202411729195.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-28
AI Technical Summary
The prior art has unstable training process and pattern crash problems in image fusion, resulting in unreasonable distribution of the generated fusion images and low quality, affecting practicality.
Using an image fusion method based on the diffusion model, the data set includes registered infrared images, visible light images and fusion prior images, using forward noise addition and reverse denoising diffusion processing, combined infrared light and visible light images as input conditions for the diffusion model, the fusion prior images are reconstructed to generate high-quality fusion images.
The effective fusion of infrared light and visible light images is achieved. The generated fusion image combines the advantages of two modal information, improves the robustness and information volume of the image, and is suitable for advanced visual tasks.
Smart Images

Figure CN119963954A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a method for fusing an infrared light image with a visible light image. Background Art
[0002] Due to the hardware limitations of imaging devices, sensors of a single type or in a single setting are usually unable to fully characterize the imaging scene. The images collected by visible light cameras are consistent with human visual perception, with rich texture details and spatial resolution, but in harsh environmental conditions such as insufficient light, rainy and foggy days, the image quality will be significantly reduced, making the target information difficult to identify. Thermal infrared camera imaging relies on the thermal radiation information of the target, has better smoke penetration ability, is not affected by lighting conditions, and is suitable for working under all-weather conditions. For example, when performing nighttime mountain forests or fire scene personnel rescue missions, thermal infrared cameras have great advantages, but the image itself lacks details and cannot reflect scene information. Therefore, there are great limitations in relying solely on a single type of image to perform visual tasks. How to use image fusion technology to organically integrate image information collected by different sensors to generate images that are more robust, richer in information, and conducive to human eye perception has become a hot topic in the current image processing field.
[0003] In order to solve the challenges in image fusion and avoid manual and tedious fusion scheme design. Among them, methods based on generative adversarial networks (GANs) stand out as end-to-end methods with excellent performance. The workflow of GAN-based fusion methods mainly involves a generator that generates fused images by integrating information from different fields, and a discriminator responsible for evaluating the similarity of probability distribution between the generated image and the source image. Although GAN-based image fusion models have produced satisfactory fused images, they are plagued by several problems, including unstable training processes and mode collapse. These challenges lead to unreasonable distribution and low quality of the generated fused images, which seriously affect the practicality of GAN-based methods. In recent years, the denoising diffusion probability model (DDPM) has attracted much attention. It has good mathematical interpretability and high-quality generation results, and has achieved great success in image generation. However, in multimodal image fusion, the diffusion model cannot achieve direct image fusion due to the lack of true value. This is a huge challenge for the current application of diffusion models in multimodal image fusion. Summary of the invention
[0004] The purpose of the present invention is to provide an image fusion method for infrared light images and visible light images based on a diffusion model.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A method for fusing infrared light and visible light images, specifically comprising:
[0007] Construct a dataset, including the registered infrared image I ir , visible light image I vi and fused prior image F0, where the fused prior image represents the general result of image fusion, which is determined by the target search function f ts (F j ) is obtained and used in the model training process.
[0008] Perform forward denoising and reverse denoising diffusion processing on the fused prior image F0:
[0009] In the forward process, the variance table [β1,β2,β3,...,β T ] Repeatedly add Gaussian noise to the fused prior image F0, and F0 will gradually degenerate into a Gaussian noise image F within T time steps. T , where T is defined as the maximum time step, t∈[1,T], T=2000; the forward diffusion process q is formulated as:
[0010]
[0011] where α t =1-β t , N represents Gaussian normal distribution; through small variance Gaussian noise α t Repeat the superposition steps, F0 gradually becomes a Gaussian image F similar to the standard Gaussian form t ;
[0012] In the reverse diffusion process, by introducing the infrared light image I ir and visible light image I vi As the input condition of the diffusion model, the Gaussian image F t In the reconstruction of the fused prior image F0, each back diffusion step uses a mean value of μ θ The neural network p θ Written as:
[0013]
[0014] in Mean μ θ Formulated as:
[0015]
[0016] where ε θ is the neural network p θ Estimating the noise of the output; Neural network p θ The role of is to learn the modeling of Gaussian normal distribution N at each time step; at the same time, the neural network is based on infrared image Iir and visible light image I vi As a condition, it is used to reconstruct the fused prior image F0 and establish the learning objective as follows:
[0017] L t =||ε t -ε θ (F t ,I ir ,I vi ,t)|| 2
[0018] Through continuous training, the learning target L is reduced t The value of .
[0019] The infrared and visible light image pairs of arbitrary resolution are input into the trained diffusion model. The diffusion model automatically outputs a fused image that retains the rich texture and color information of the visible light image and the prominent target information of the infrared image through continuous iteration.
[0020] The size of the input image sample is (H, W, C), where H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image. ir , visible light image I vi and Gaussian image F t The channels are concatenated together and used as the input to the neural network noise predictor.
[0021] Noise predictor ε θ It is a U-net type neural network that converts the number of channels of the feature map to C through the convolution module. θ It includes a down-sampling module, a mid-step module, and an up-sampling module. The down-sampling and up-sampling modules contain four steps, consisting of several residual blocks, two down-sampling layers, and two up-sampling layers. This paper uses a two-dimensional convolution with a step size of 2 to achieve down-sampling, and uses a two-dimensional transposed convolution to achieve up-sampling. In addition, the output of each down-sampling step is added to the input of the corresponding up-sampling step to make full use of the image features. For a time step t, the position embedding module in the Transformer is used to convert the time step t into the time embedding t e , and t e is added to each residual block. It is worth noting that the noise predictor in this paper is lightweight compared to the noise predictor in the original paper of the diffusion model, which can significantly reduce the training cost.
[0022] Compared with the prior art, the present invention has the following advantages:
[0023] 1. This invention proposes a method for fusion of infrared and visible light images based on a diffusion model. Aiming at the lack of true value in image fusion tasks, this method creatively introduces a fusion prior image to guide training. The fusion knowledge prior represents the general distribution of the fusion results of infrared and visible light images, which contains fusion results from various strategies, and the fusion prior image is sampled from this distribution.
[0024] 2. The diffusion model network used in the present invention can fuse infrared images and visible light images to obtain a fused image that combines information from the two modalities, which can help improve the performance of advanced visual tasks (such as target tracking, target detection, and semantic segmentation).
[0025] 3. The training framework proposed in the present invention can be extended to other related image fusion fields to solve similar problems and is universal. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is the dataset of this method.
[0027] Figure 2 It is the forward noise adding process of this method.
[0028] Figure 3 is the architecture of the conditional noise predictor of our method.
[0029] Figure 4 This is the schematic diagram of the image fusion process of this method.
[0030] Figure 5 It is a schematic diagram of the training process of this method.
[0031] Figure 6 It is a schematic diagram of the reasoning process of this method.
[0032] Figure 7 It is a schematic diagram of generating a fused image by this method. DETAILED DESCRIPTION
[0033] The technical solution of the present invention is further described below in conjunction with the accompanying drawings. The meanings of the parameters are shown in Table 1.
[0034] Table 1
[0035]
[0036]
[0037] This embodiment takes infrared light, visible light and fused priori images in a data set as examples to illustrate the image fusion method based on a diffusion model.
[0038] like Figure 1As shown, the data set used in the present invention is composed of registered infrared light images, visible light images and fused prior images. The infrared and visible light images are used as inputs of the noise predictor, and the fused prior images are used in the training process of the noise predictor.
[0039] like Figure 2 As shown, the fused prior image F0 is continuously noised, and the formula Get the noise image F at time t t , we can see that as the time step increases, the original fusion prior image F0 is constantly "destroyed" and eventually becomes a Gaussian noise image F t Our ultimate goal is to use the infrared image and the visible light image as conditions in the inference stage to predict the noise added to the fused image at time t, thereby reconstructing the fused image. For the infrared and visible light image pairs used for inference, the reconstruction process is the image fusion process.
[0040] like Figure 3 As shown in Figure 1, infrared, visible light, and Gaussian noise images are spliced together as the input of the noise predictor. The noise ε added at time t predicted by the U-net network output θ (F t ,I ir ,I vi ,t), and the noise ε added at time t t For comparison, the noise predictor is trained by minimizing the following loss function:
[0041] L t =||ε t -ε θ (F t ,I ir ,I vi ,t)|| 2 .
[0042] like Figure 4 As shown, in the inference stage, the infrared and visible light images and a noise image F of the same size as them are combined. t-1 spliced together as the input of the trained noise predictor, the noise predictor will output the predicted noise ε at this time t pred , through the formula Get the Gaussian noise image F at the previous moment t-1 . Continue to iterate until the fused image F0 is obtained.
[0043] The training process of the present invention is as follows Figure 5 As shown, the input is: the fusion prior image F0 passes through the forward process of the diffusion model The resulting noisy image F t , infrared image Iir , visible light image I vi , concatenate these three images at the channel level. The denoising module (DM) is a conditional noise predictor ε θ (F t ,I ir ,I vi ,t), the purpose is to predict the noise ε “added” in the forward process t , establish learning objectives L t =||ε t -ε θ (F t ,I ir ,I vi ,t)|| 2 .
[0044] The reasoning stage of the present invention is as follows Figure 6 As shown, input: infrared image I ir , visible light image I vi and the noise image f T , F T ~N(0,I). The trained model outputs the predicted noise ε at time T pred , through the formula Get the denoised image f T-1 The model is expressed by the formula Continuously iterate and finally output the fused image This enables the reconstruction of the fused prior image f0.
[0045] like Figure 7 As shown in the figure, under bad lighting conditions, visible light images are degraded and it is difficult to capture pedestrian targets, while infrared images can reflect thermal radiation information and clearly display the location information of pedestrians, but the texture information of the background is incomplete. The fusion method adopted by the present invention can better combine the information of the two modal images, generate images with stronger robustness, richer information, and conducive to human eye perception, and improve the performance of other advanced visual tasks.
[0046] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. The scope of the present invention is defined by the appended claims rather than the above description, and therefore it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims be included in the present invention.
Claims
1. A method for fusing infrared light and visible light based on a diffusion model, characterized in that: Step 1: Build a data set, including a pair of registered infrared images I ir and visible light image I vi And their corresponding fusion prior image F0, where the fusion prior image represents the general result of image fusion and is used in the model training process; the fusion prior image F0 is a pre-generated fusion image, which is used as the true value of model training, so that the image fusion task changes from an unsupervised task to a supervised task; the distribution of the fusion image F0 is obtained by the target search function Get, where ω n is the weight of different indicators, T n represents the evaluation function, T n (F j ) represents the evaluation score of the given j-th sample, F j are sample images generated by multiple methods; Step 2: Perform forward denoising and reverse denoising diffusion processing on the fused prior image F0: In the forward process, the variance table [β1,β2,...,β t ,...,β T ] controls the amount of noise added at each time step, β t It gradually increases with time step t, indicating that more noise is added at each time step; Gaussian noise is repeatedly added to the fused prior image F0, and f0 will gradually degenerate into a Gaussian noise image F within t time steps. t , where T is defined as the maximum time step length, t∈[1,T], T=1000; the forward diffusion process q is formulated as: where α t =1-β t , N represents Gaussian normal distribution, I is the unit matrix, is the coefficient in the diffusion process, which controls the noise intensity at each time step. Represents the noise variance in the forward diffusion process, according to α t Calculated; In the back diffusion process, by gradually denoising the Gaussian image F t Reconstruct the fusion prior image F0 and use the infrared image I ir and visible light image I vi As the input condition of the diffusion model, each back diffusion step uses the mean μ θ The neural network p θ Written as: in Mean μ θ Formulated as: where ε θ is the neural network p θ Estimating the noise of the output; Neural network p θ Learn to model the Gaussian normal distribution N at each time step; at the same time, the neural network is based on the infrared image I ir and visible light image I vi As a condition, it is used to reconstruct the fused prior image F0 and establish the learning objective as follows: L t =E(||e t -e θ (F t ,I ir ,I vi ,t)|| 2 ) Through continuous training, the learning target L is reduced t The value of Step 3: Output of the diffusion model network: The infrared and visible light image pairs of arbitrary resolution are input into the trained diffusion model. The diffusion model automatically outputs a fused image that retains the rich texture and color information of the visible light image and the prominent target information of the infrared image through continuous iteration.
2. The method according to claim 1, characterized in that The size of the input F0 image sample is (H, W, C), where H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image; the input images are spliced together in the channel dimension as the input of the neural network noise predictor.
3. The method according to claim 2, characterized in that Noise predictor ε θ It is a U-net type neural network that converts the number of channels of the feature map into C through the convolution module; ε θ It includes a downsampling module, an intermediate module, and an upsampling module; the downsampling and upsampling modules contain four steps, consisting of several residual blocks, two downsampling layers, and two upsampling layers; downsampling is achieved using a two-dimensional convolution with a step size of 2, and upsampling is achieved using a two-dimensional transposed convolution; in addition, the output of each downsampling step is added to the input of the corresponding upsampling step to make full use of image features; for a time step t, the position embedding module in the Transformer is used to convert the time step t into a time embedding t e , and t e is added to each residual block.
Citation Information
Patent Citations
Multi-source information fusion uncertainty discrimination method
CN115457351A
Multi-source information fusion vehicle collision airbag control system and vehicle
CN116461512A
Infrared visible light image fusion method and device, electronic equipment and storage medium
CN118096578A
Multiband image fusion method based on diffusion model
CN118154435A
Line scanning image super-resolution method and device based on de-noising diffusion fusion model
CN118297803A
Cited By
Intelligent system evaluation method based on multi-model cross identification data credibility
CN120123890A
Method for converting visible light into infrared image based on diffusion model
CN120318060A