Infrared and Visible Image Fusion Target Detection Method Based on Diffusion Generation Driving
Through a diffusion generation-driven method, infrared and visible light images are fused to generate a collaborative diffusion generation prior, solving the problem of object detection accuracy in low-quality images and complex environments in the prior art, and achieving efficient object detection in all weather and all scenes.
Patent Information
- Application Number
- CN202510323992.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The existing infrared and visible light image fusion object detection methods do not perform well when processing detection technologies in low-quality images and complex environments, and it is difficult to achieve accurate object detection in complex environments caused by multiple factors such as noise, low contrast, low resolution, and low light.
Using an infrared and visible image fusion object detection method based on diffusion generation drive, the conditional diffusion model is pre-trained by constructing low-quality and high-quality image data sets to generate a collaborative diffusion generation prior, and then fusion features are obtained and object detection is performed.
This method can effectively reduce the noise impact caused by low-quality images, improve the accuracy of object detection in complex environments, and achieve efficient object detection in all weather and all scenarios.
Smart Images

Figure CN119850935B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection in computer vision, and particularly to an infrared and visible light image fusion object detection method driven by diffusion generation. Background Art
[0002] Object detection, as an important research direction in the field of computer vision, has extensive applications in many fields of real life. However, in many actual scenarios such as security monitoring and intelligent driving, it is often necessary to maintain high-efficiency object detection capabilities day and night and under different meteorological conditions. At the same time, due to limitations of the environment, equipment, etc., the quality of the collected images may be low. Facing such challenges, object detection methods based on high-quality images collected by a single sensor can no longer meet the requirements. Therefore, a large number of object detection methods based on multi-sensor data have become new research hotspots, which to a certain extent make up for the shortcomings of limited information collected by a single sensor and improve the accuracy of object detection.
[0003] Currently, the most typical and common multi-sensor fusion source is the combination of infrared and visible light images. By combining the heat source target information sensed by the infrared sensor with the object reflection light imaging captured by the visible light sensor, the advantages of the two sensors can be mutually complemented, the feature information of the two modal images can be effectively integrated, synthesized, and explored, the object can be highlighted while enhancing the understanding of the scene, and object detection in all-weather and full-scene can be realized.
[0004] Although the existing infrared and visible light image fusion object detection methods have achieved certain success in the fusion of multi-sensor image data, these methods mainly rely on data under ideal scenarios and acquisition conditions for training. Although some methods have paid attention to image fusion object detection under low-light conditions, they still face challenges in detection technologies affected by multiple factors. When dealing with challenging application scenarios, such as complex environments caused by multiple factors such as noise, low contrast, low resolution, and low light, they often perform poorly. Summary of the Invention
[0005] The purpose of the present invention is to provide an infrared and visible light image fusion object detection method driven by diffusion generation, which can effectively process low-quality infrared and visible light images and perform accurate object detection.
[0006] To achieve the above purpose, the present invention provides an infrared and visible light image fusion object detection method driven by diffusion generation, including the following steps:
[0007] Step S1: Based on the publicly available infrared and visible light image dataset M3FD and the low-light and normal-light image dataset LOL, construct a low-quality infrared image dataset, a high-quality infrared image dataset, a low-quality visible light image dataset, and a high-quality visible light image dataset to obtain a dataset for conditional diffusion model training;
[0008] Step S2: Pre-train a conditional diffusion model for low-quality image data to obtain a co-diffusion generation prior for low-quality images;
[0009] Step S3: Obtain fusion features based on the fused images of the co-diffusion generation prior;
[0010] Step S4: Perform object detection based on the generation-driven fused images to produce the final object detection results.
[0011] Preferably, in step S1, use the publicly available infrared and visible light image dataset M3FD and the low-light and normal-light image dataset LOL to construct low-quality infrared and visible light image datasets for training;
[0012] Among them, perform simulation degradation processing on all infrared images in the M3FD dataset to construct a low-quality infrared image dataset with noise and low-resolution characteristics. The specific process is as follows:
[0013] By using a downsampling factor of 0.25, reduce the resolution of the image to one-fourth of the original image. At the same time, add stripe-shaped noise to the downsampled image, and set the noise intensity to , where is an index to measure the noise intensity;
[0014] The original infrared images in the M3FD dataset correspond to the high-quality infrared image dataset;
[0015] The low-light images in the LOL dataset are directly used as low-quality visible light image datasets with low resolution and low light characteristics;
[0016] The normal-light images in the LOL dataset correspond to the high-quality infrared image dataset.
[0017] Preferably, in step S2, pre-train a conditional diffusion model for low-quality image data to obtain a co-diffusion generation prior for low-quality images. The specific process is as follows:
[0018] Step S21: First, define low-quality images as follows:
[0019] ;
[0020] Among them, represents a low-quality image sample; including infrared images and visible light images ;
[0021] Secondly, define high-quality images as follows:
[0022] ;
[0023] Among them, represents a high-quality image sample;
[0024] Step S22. In the forward diffusion stage, define the distribution of data as follows:
[0025] ;
[0026] Among them, obeys the distribution ; represents an original data; represents the probability distribution of the original data ;
[0027] Step S23. By gradually injecting Gaussian noise in the forward diffusion stage, the original data is gradually destroyed in diffusion time steps; is the total number of steps in the entire diffusion process, set to 1000;
[0028] Step S24. Pre-define that the Markov chain satisfies the following distribution:
[0029] ;
[0030] Among them, is the state of the sample at time step t; represents the normal distribution; represents the probability distribution; and are noise adjustment parameters in the diffusion process, controlling the intensity of adding noise at the th time step, and the relationship between the two is ; represents the identity matrix;
[0031] At the same time, the marginal distribution is derived as follows:
[0032] ;
[0033] Among them, represents the original data sample; , ; When tends to the time step When tends to 0, it follows a normal distribution , and at the same time, the forward process ends;
[0034] Step S25: In the reverse process, denoising is gradually performed from the Gaussian noise through the Markov chain. The probability distribution of the generated image is as follows:
[0035] ;
[0036] Among them, represents the probability distribution of the random variable and at time step under the conditions ; are the parameters of the model; the variance is a time - related constant; is the predicted mean at time step ;
[0037] and are calculated as follows:
[0038] ;
[0039] ;
[0040] Among them, the state of the sample at time step t is concatenated with the low - quality image and input into the noise estimation network with parameters . The degraded image will provide semantic and structural information for sample generation;
[0041] Step S26: The optimization objective is as follows:
[0042] ;
[0043] Among them, represents the expectation; represents the noise vector added at time step ; represents the square of the norm;
[0044] Step S27: Finally, in the inference stage of low - quality image restoration, the low - quality image will be used as the diffusion condition to guide the model to gradually iterate starting from the standard Gaussian noise to generate the high - quality source image.
[0045] Preferably, in step S3, a prior fusion image is generated based on collaborative diffusion to obtain fusion features. The specific process is as follows:
[0046] Step S31: First, initialize the fusion image samples from Gaussian noise , define , representing the fusion image; the mean in the sampling process is calculated and updated as follows:
[0047] ;
[0048] where, represents the low-quality images including infrared and visible light modalities; represents the state of the fusion image at time step t; represents the high-quality sample related to the th modality at time step ; is the combination weight in the reverse process and satisfies ;
[0049] By combining the current fusion sample , the high-quality sample , and the weight of the previous step and inputting them into the weight prediction network , it is calculated as follows:
[0050] ;
[0051] where, is the unnormalized weight of the th modality at time step ;
[0052] Step S32: Normalize the unnormalized weight of the th modality at time step to obtain the final combination weight , as follows:
[0053] ;
[0054] where, represents the e-exponential form of the unnormalized weight of the th modality at time step ; At the same time, in this image task, since there are only two modalities, the weight of the corresponding other modality is directly defined as ;
[0055] Step S33, Weight Prediction Network The loss function is as follows:
[0056] ;
[0057] wherein, represents the fused image predicted at time step ; represents the maximum aggregation strategy; represents the gradient operator; represents norm; is the high-quality source image restored from the diffusion model for low-quality images obtained through pre-training in the previous stage.
[0058] Preferably, in step S4, based on the generated-driven fused image, object detection is performed to generate the final object detection result. The specific process is as follows:
[0059] Through the image fusion step based on the combined prior until the network converges, clear fused features are finally obtained. The fused features are input into the task head to extract task-specific information. The task head uses the YOLOv5 architecture;
[0060] Finally, through post-processing, the final prediction result is obtained, as follows:
[0061] ;
[0062] wherein, represents the learnable parameters of the task head ; The post-processing process includes the class label, corresponding bounding box coordinates, and confidence score of each detected object.
[0063] Therefore, the present invention adopts the above-mentioned diffusion generation-driven infrared and visible light image fusion object detection method, weakens the noise influence brought by low-quality images, can effectively process low-quality infrared and visible light images, and performs accurate object detection.
[0064] Next, through the drawings and embodiments, the technical solution of the present invention will be further described in detail. Brief Description of the Drawings
[0065] Figure 1 is the flowchart of the diffusion generation-driven infrared and visible light image fusion object detection method of the present invention;
[0066] Figure 2It is the flowchart for pre-training the conditional diffusion model for low-quality image data in the present invention;
[0067] Figure 3 It is the flowchart for image fusion and object detection based on co-diffusion generation prior in the present invention. Detailed implementation manners
[0068] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0069] As Figure 1 shown, the infrared and visible light image fusion object detection method based on diffusion generation driving includes the following steps:
[0070] Step S1: Based on the publicly available infrared and visible light image dataset M3FD and the low-light and normal-light image dataset LOL, construct a low-quality infrared image dataset, a high-quality infrared image dataset, a low-quality visible light image dataset, and a high-quality visible light image dataset to obtain a dataset for conditional diffusion model training;
[0071] Step S2: Pre-train the conditional diffusion model for low-quality image data to obtain the co-diffusion generation prior of low-quality images;
[0072] Step S3: Obtain fusion features based on the fusion images with co-diffusion generation prior;
[0073] Step S4: Perform object detection based on the generation-driven fusion images to generate the final object detection results.
[0074] Embodiment
[0075] Step S1: Based on the publicly available infrared and visible light image dataset M3FD and the low-light and normal-light image dataset LOL, construct a low-quality infrared image dataset, a high-quality infrared image dataset, a low-quality visible light image dataset, and a high-quality visible light image dataset to obtain a dataset for conditional diffusion model training.
[0076] Use the publicly available infrared and visible light image dataset M3FD and the low-light and normal-light image dataset LOL. Both of these datasets include various scenes in a variety of real environments. In order to enable the method proposed in this method to more effectively reduce the noise impact brought by low-quality images, therefore, in the image preprocessing process, a low-quality infrared and visible light image dataset for training will be constructed first.
[0077] Among them, all infrared images in the M3FD dataset are subjected to simulated degradation processing to construct a low-quality infrared image dataset mainly featuring noise and low resolution; the original infrared images correspond to the high-quality infrared image dataset. Specifically, by using a downsampling factor of 0.25, the resolution of the image is reduced to one-fourth of the original image, and at the same time, stripe-shaped noise is added to the downsampled image, and the noise intensity is set to , where is an index for measuring the noise intensity, and the larger the value, the stronger the noise.
[0078] The low-light images in the LOL dataset have the characteristics of low contrast and low illumination, and can be directly used as a low-quality visible light image dataset mainly featuring low resolution and low illumination; the normal light images in the LOL dataset correspond to the high-quality visible light image dataset.
[0079] Step S2, pre-train the conditional diffusion model for low-quality image data to obtain the co-diffusion generation prior of low-quality images.
[0080] In the training stage of low-quality image restoration, the two infrared and visible light image datasets constructed in the previous stage are used to train a denoising diffusion probability model with image restoration ability, as Figure 2 shown.
[0081] Step S21, first, define the low-quality image as follows:
[0082] ;
[0083] Among them, represents a low-quality image sample; includes infrared images and visible light images .
[0084] Secondly, define the high-quality image as follows:
[0085] ;
[0086] Among them, represents a high-quality image sample.
[0087] Step S22, in the forward diffusion stage, define the distribution of the data as follows:
[0088] ;
[0089] Among them, follows the distribution ; represents an original data; represents the original data Probability distribution
[0090] Step S23: By gradually injecting Gaussian noise in the forward diffusion stage, the original data is gradually destroyed at diffusion time steps ; is the total number of steps in the entire diffusion process, set to 1000
[0091] Step S24: The predefined Markov chain satisfies the following distribution:
[0092] ;
[0093] where, is the state of the sample at time step t; represents the normal distribution; represents the probability distribution; and are the noise adjustment parameters in the diffusion process, controlling the intensity of the noise added at the th time step. The relationship between the two is ; represents the identity matrix.
[0094] At the same time, the marginal distribution is derived as follows:
[0095] ;
[0096] where, represents the original data sample; , ; when tends to the time step , tends to 0, obeys the normal distribution , and at the same time the forward process ends.
[0097] Step S25: In the reverse process, starting from the Gaussian noise denoising is gradually performed through the Markov chain. The probability distribution of the generated image is as follows:
[0098] ;
[0099] where, represents the probability distribution of the random variable and at time step under the conditions ; are the parameters of the model; the variance is a time-related constant; is at time step The predicted mean value at .
[0100] and The calculation formula is as follows:
[0101] ;
[0102] ;
[0103] Among them, the state of the sample at time step t With low quality images After splicing, the input parameters are Noise estimation network In this paper, the degraded image will provide semantic and structural information for sample generation.
[0104] Step S26, optimizing the target, as shown below:
[0105] ;
[0106] in, express expectations; Indicates that at time step the added noise vector; express The square of the norm.
[0107] Step S27. Finally, in the inference stage of low-quality image restoration, the low-quality image will be used as a diffusion condition to guide the model to iterate step by step starting from standard Gaussian noise and finally generate a high-quality source image.
[0108] Step S3: Generate a priori fusion image based on cooperative diffusion to obtain fusion features.
[0109] In the image fusion generation stage, after pre-training the diffusion model for low-quality images, the prior distribution of unimodal high-quality images is obtained. The prior is determined by the noise predicted at each sampling step, and combining the noise from multiple diffusion models can generate a result that contains their respective prior distributions.
[0110] Therefore, by flexibly combining the noise generated by the pre-trained diffusion model, image fusion based on the generative model is achieved, such as Figure 3 shown.
[0111] Step S31: First, initialize the fused image sample from Gaussian noise ,definition , represents the fused image. The mean value during sampling Compute the update as follows:
[0112] ;
[0113] Among them, represents a low-quality image including infrared and visible light modalities; represents the state of the fused image at time step t; represents the high-quality sample related to the modality at time step ; is the combination weight in the reverse process and satisfies .
[0114] By combining the current fused sample , the high-quality sample , and the weight of the previous step and inputting them into the weight prediction network , it is calculated as follows:
[0115] ;
[0116] Among them, is the unnormalized weight of the th modality at time step .
[0117] Step S32. Then, normalize the unnormalized weight of the th modality at time step to obtain the final combination weight , as follows:
[0118] ;
[0119] Among them, represents the exponential form of the unnormalized weight of the th modality at time step ; at the same time, in this image task, since there are only two modalities, the weight of the corresponding other modality can be directly defined as .
[0120] Step S33. The loss function of the weight prediction network is as follows:
[0121] ;
[0122] Among them, represents the fused image predicted at time step ; Represents the maximum aggregation strategy; Represents the gradient operator; Represents Norm; is a high-quality source image restored from a diffusion model for low-quality images based on the pre-training in the previous stage.
[0123] Step S4: Based on the generated driving fusion image, perform object detection to generate the final object detection result.
[0124] Through the image fusion step based on the combined prior until the network converges, clear fusion features can be finally obtained , and the fusion features are passed into the task head to extract task-specific information, and the task head uses the YOLOv5 architecture.
[0125] Finally, through post-processing, the final prediction result is obtained , as follows:
[0126] ;
[0127] Among them, represents the learnable parameters of the task head ; the post-processing process includes the class label, corresponding bounding box coordinates, and confidence score of each detected object.
[0128] Therefore, the present invention adopts the above-mentioned infrared and visible light image fusion object detection method based on diffusion generation driving, weakens the noise influence brought by low-quality images, can effectively process low-quality infrared and visible light images, and perform accurate object detection.
[0129] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A target detection method based on infrared and visible light image fusion driven by diffusion generation, characterized in that: The following steps are involved: Step S1, based on the public infrared and visible light image dataset M3FD and the low-light and normal light image dataset LOL, construct a low-quality infrared image dataset, a high-quality infrared image dataset, a low-quality visible light image dataset, and a high-quality visible light image dataset to obtain a dataset for conditional diffusion model training; Among them, all infrared images of the M3FD dataset are subjected to simulated degradation processing to construct a low-quality infrared image dataset with noise and low resolution characteristics. The specific process is as follows: By using a downsampling factor of 0.25, the image resolution is reduced to a quarter of the original image, and stripe noise is added to the downsampled image. The noise intensity is set to ,in It is a measure of noise intensity; The original infrared images of the M3FD dataset correspond to the high-quality infrared image dataset; The low-light images of the LOL dataset are directly used as a low-quality visible light image dataset with low resolution and low-light characteristics; The normal light images of the LOL dataset correspond to the high-quality infrared image dataset; Step S2, pre-training the conditional diffusion model for low-quality image data to obtain a collaborative diffusion generation prior for low-quality images; Step S3, generating a priori fusion image based on cooperative diffusion to obtain fusion features; Step S4: Based on the generated driving fusion image, target detection is performed to generate a final target detection result.
2. The infrared and visible light image fusion target detection method based on diffusion generation drive according to claim 1 is characterized in that: In step S2, the conditional diffusion model for low-quality image data is pre-trained to obtain the co-diffusion generation prior of low-quality images. The specific process is as follows: Step S21: First, define a low-quality image as follows: ; in, Represents low-quality image samples; Including infrared images and visible light images ; Second, define a high-quality image as follows: ; in, Represents a high-quality image sample; Step S22: In the forward diffusion stage, define the distribution of data as follows: ; in, Follow the distribution ; Represents a raw data; Represents the original data The probability distribution of Step S23: gradually inject Gaussian noise in the forward diffusion stage to achieve diffusion time steps gradually destroy the original data ; is the total number of steps in the entire diffusion process, which is set to 1000; Step S24: predefine the Markov chain to satisfy the following distribution: ; in, is the state of the sample at time step t; represents normal distribution; represents a probability distribution; and is the noise adjustment parameter in the diffusion process, which is controlled in the The intensity of adding noise in each time step is related to ; represents the identity matrix; At the same time, the marginal distribution is derived as follows: ; in, represents the original data sample; , ;when Trending time step hour, tends to 0, Normal distribution , and the forward process ends at the same time; Step S25, reverse process from Gaussian noise The probability distribution of the generated image is shown below after Markov chain denoising step by step: ; in, Indicates that the condition and Next, time step A random variable The probability distribution of are the parameters of the model; variance is a time-dependent constant; is at the time step The predicted mean value at ; and The calculation formula is as follows: ; ; Among them, the state of the sample at time step t With low quality images After splicing, the input parameters are Noise estimation network In , the degraded image will provide semantic and structural information for sample generation; Step S26, optimizing the target, as shown below: ; in, express expectations; Indicates that at time step the added noise vector; express The square of the norm; Step S27. Finally, in the inference stage of low-quality image restoration, the low-quality image will be used as a diffusion condition to guide the model to iterate step by step starting from standard Gaussian noise to generate a high-quality source image.
3. The infrared and visible light image fusion target detection method based on diffusion generation drive according to claim 2 is characterized in that: In step S3, a priori fusion image is generated based on cooperative diffusion to obtain fusion features. The specific process is as follows: Step S31: First, initialize the fused image sample from Gaussian noise ,definition , represents the fused image; the mean value during sampling Compute the update as follows: ; in, Representation includes low-quality images of infrared and visible light modalities; Represents the state of the fused image at time step t; Representatives and The mode-dependent High-quality samples; is the combined weight in the reverse process and satisfies ; By taking the current fusion sample , high quality samples , and the weight of the previous step Combined and input into the weight prediction network It is calculated as follows: ; in, For the The mode at the time step The unnormalized weights of Step S32, then, The mode at the time step The unnormalized weights Standardize and get the final combined weight , as shown below: ; in, Indicates The mode at the time step The unnormalized weights e exponential form; at the same time, since there are only two modes, the weight of the corresponding other mode is directly defined as ; Step S33: Weight prediction network The loss function , as shown below: ; in, Represents the time step Predicted fused image; represents the maximum aggregation strategy; represents the gradient operator; express norm; It is a high-quality source image restored based on the diffusion model for low-quality images obtained through pre-training in the previous stage.
4. The infrared and visible light image fusion target detection method based on diffusion generation drive according to claim 1 is characterized in that: In step S4, target detection is performed based on the generated drive fusion image to generate the final target detection result. The specific process is as follows: After the image fusion step based on the combined prior until the network converges, the fusion feature is finally obtained. , the fusion features Incoming task header To extract task-specific information, the task head uses the YOLOv5 architecture; Finally, after post-processing, the final prediction result is obtained. , as shown below: ; in, Represents the task head The post-processing process includes the category label, corresponding bounding box coordinates and confidence score of each detected object.
Citation Information
Patent Citations
Multi-modal image fusion and super-resolution method based on conditional diffusion probability model
CN118314022A
Infrared and visible light image fusion method based on conditional diffusion model
CN119540070A