A method for converting visible light to infrared images based on diffusion model
Through a conversion method based on diffusion model, combined with retinal cortex theory and adaptive noise control, multi-scale discriminator and physical constraints are used to solve the noise suppression and detail recovery problems generated by infrared images under low light, and the generation and recognition of high-quality infrared images are achieved.
Patent Information
- Application Number
- CN202510808734.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-17
AI Technical Summary
In the visible light to infrared image conversion, it is difficult to effectively suppress noise, restore details and ensure the physical consistency of the generated images in low light and complex environments. Especially in dynamic backgrounds, it is easy to cause details loss and artifacts, and it is difficult to meet the needs of high-precision infrared image generation.
Using a conversion method based on diffusion model, reflection and illumination information are separated through retinal cortex theory, combined with adaptive noise control and multi-scale discriminator, physical constraint loss function is introduced, and the model is trained using a joint optimized total loss function to achieve high-quality conversion of images.
Generate infrared images with clear details and physically real in complex environments, with strong environmental adaptability and generalization capabilities, significantly improving infrared imaging quality and target recognition performance.
Smart Images

Figure CN120318060B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for converting visible light into infrared images based on a diffusion model, and belongs to the technical fields of image processing and computer vision. Background Art
[0002] Image processing and computer vision technologies face numerous challenges in the field of visible-to-infrared image conversion. This conversion task is particularly complex due to significant differences between visible-light and infrared images in imaging mechanisms, information representation, and application scenarios, particularly in low-light and complex environments. Infrared images primarily reflect the thermal radiation characteristics of objects, while visible-light images capture the reflected light information from the object's surface. This leads to significant differences in visual representation, detail preservation, and feature extraction. Therefore, achieving high-quality conversion from visible-light images to infrared images, especially accurately restoring the thermal radiation characteristics of infrared images under low-light conditions, has become an important and challenging research topic.
[0003] While existing image conversion methods, such as those based on generative adversarial networks (GANs) and traditional image translation techniques, have advanced this field to some extent, they still have significant limitations. Firstly, these methods typically rely on fixed network structures and predefined training data, making them difficult to adapt to changing environmental conditions. This is particularly true in low light or dynamic backgrounds, where they can easily lead to loss of detail and artifacts. Secondly, most methods fail to adequately address the global structure and local details of the image, making it difficult to meet the requirements for generating high-precision, physically consistent infrared images, limiting their practicality.
[0004] Furthermore, current methods perform poorly in noise suppression, contrast enhancement, and detail recovery. Specifically, in the process of converting visible-light images to infrared images, effectively suppressing noise, restoring image detail, and ensuring the physical consistency of the generated image remain pressing challenges. While some studies have attempted to improve image quality by enhancing contrast or employing denoising techniques, these often result in the loss of target details or the introduction of artifacts under complex background conditions, reducing the accuracy and application value of the generated image.
[0005] Given these shortcomings, a novel method is urgently needed that can effectively suppress noise and enhance image contrast and structure while preserving realistic image details, thus meeting the requirements for high-quality infrared image generation. Furthermore, this method should possess strong adaptability, automatically adjusting the generation strategy based on environmental changes to ensure the efficiency and reliability of the image conversion process. Summary of the Invention
[0006] The present invention aims to propose a method for converting visible light to infrared images based on a diffusion model, targeting the challenges of image generation in complex lighting and low-light environments, and achieving high-quality infrared image reconstruction while taking into account physical consistency and detail preservation.
[0007] To achieve the above object, the present invention provides a method for converting visible light to infrared images based on a diffusion model, comprising the following steps:
[0008] S1, collect visible light images of a certain scene;
[0009] S2, performing data input operation on the visible light image collected in S1;
[0010] S3, using retinal cortex theory to perform separation operation on the image after the data input operation in S2;
[0011] S4, using the diffusion model to perform a forward diffusion operation on the image after the retinal cortex theoretical separation operation in S3;
[0012] S5, using an adaptive noise control mechanism to perform dynamic noise removal on the image after the forward diffusion operation in S4;
[0013] S6, using a diffusion model to perform a reverse diffusion operation on the image after the dynamic denoising operation in S5 and generate an infrared image;
[0014] S7, using physical formulas to perform physical consistency constraint operations on the infrared image generated after the back diffusion operation in S6;
[0015] S8, using a multi-scale discriminator to perform quality assessment on the infrared image after the physical constraint operation in S7 at different resolution scales;
[0016] S9. Use the jointly optimized total loss function to perform model parameter optimization training operations on the infrared image after multi-scale discrimination in S8.
[0017] Furthermore, the step of using the retinal cortex theory to perform separation operation on the image after the data input operation in S2 in S3 is:
[0018] S3.1 introduces retinal cortex theory, which decomposes the input visible light image into a reflected image and an illuminated image, thereby eliminating the influence of lighting conditions on infrared image generation and ensuring that the generated image reflects the true surface characteristics of the object. The retinal cortex theory can be expressed by the following mathematical formula:
[0019] ;
[0020] in, Represents the spatial position coordinates of a pixel in the image, represents the input visible light image, represents the reflected image, Represents an illumination image.
[0021] Furthermore, the step of using the diffusion model in S4 to perform a forward diffusion operation on the image after the retinal cortex theoretical separation operation in S3 is:
[0022] S4.1 performs a forward diffusion process on the input visible light image, gradually adds noise, and generates a noise image. The mathematical formula of the forward diffusion process is as follows:
[0023] ;
[0024] in, is the image at step t, is the image of the previous step, is noise sampled from a standard normal distribution, It is a parameter that controls the intensity of noise addition and gradually decreases as the number of steps increases.
[0025] Furthermore, the step of using the adaptive noise control mechanism in S5 to perform a dynamic noise removal operation on the image after the retinal cortex theoretical separation operation in S4 is:
[0026] S5.1 In the reverse denoising process, the noise removal strength is dynamically adjusted according to the local features of the image (such as gradients and edges). Adaptive noise control ensures that the image details are preserved by adjusting the denoising strength in each area. In particular, the denoising strength needs to be reduced in edge and high-frequency areas to avoid losing details. The adaptive noise control mechanism can be implemented by the following formula:
[0027] ;
[0028] in, is the adaptive noise mapping, is the control coefficient, It is the image gradient, which is used to represent the intensity of image changes.
[0029] Furthermore, the step of using the diffusion model in S6 to perform a reverse diffusion operation on the image after the dynamic denoising operation in S5 and generate an infrared image is:
[0030] S6.1 uses a denoising neural network to gradually remove noise from the noisy image, restore the original image, and generate an infrared image. The goal of the reverse denoising process is to gradually restore a clear image from the noisy image. Each step removes noise until the image is restored to its original state. This can be achieved using the following formula:
[0031] ;
[0032] in, represents the denoising network, represents the parameters of the denoising network, represents the current noise image, represents the restored image after denoising.
[0033] Furthermore, the step of using a physical formula in S7 to perform a physical consistency constraint operation on the infrared image generated after the back diffusion operation in S6 is:
[0034] S7.1 introduces physical constraints to ensure that the generated infrared image complies with physical laws such as thermal radiation. Physical constraints ensure that the generated infrared image is not only visually consistent but also complies with the physical laws of thermal radiation and temperature changes, thereby ensuring the authenticity of the image. The physical loss function is shown in the formula:
[0035] ;
[0036] in, and represent the generated and real infrared images respectively, and Represent the generated and real image temperature information respectively.
[0037] Furthermore, the step of using a multi-scale discriminator in S8 to perform quality assessment on the infrared image generated after the back diffusion operation in S7 at different resolution scales is as follows:
[0038] S8.1 uses a multi-scale discriminator to evaluate the generated infrared image, ensuring that the image has good details and global consistency at different resolutions. The multi-scale discriminator optimizes the image details to ensure that the generated infrared image maintains good quality at different scales. The discriminator output at each scale is:
[0039] ;
[0040] The final multi-scale loss function is:
[0041] ;in, Indicates the The discriminator output of the scale, Indicates the The discriminative loss on scale, represents the generated infrared image, Represents a real infrared image.
[0042] Furthermore, the step of using the jointly optimized total loss function in S9 to perform model parameter optimization training on the infrared image after multi-scale discrimination in S8 is as follows:
[0043] S9.1 uses a total loss function to uniformly train and optimize the diffusion model. The total loss function comprehensively considers the generative adversarial loss, diffusion loss, physical constraint loss, multi-scale discriminant loss, and adaptive noise control loss. It can comprehensively optimize the quality, physical realism, visual detail, and noise suppression of the generated infrared image. The mathematical formulas for each loss function are as follows:
[0044] Generative Adversarial Loss (GAN Loss), used to ensure the authenticity of generated images, is defined as:
[0045] ;
[0046] Diffusion loss, which constrains the ability of the diffusion model to gradually recover the noise-free image, is defined as:
[0047] ;
[0048] The physical constraint loss ensures that the generated image conforms to the physical properties of thermal radiation and temperature distribution and is defined as:
[0049] ;
[0050] Multi-scale discriminant loss evaluates and optimizes the details of the image at different scales and is defined as:
[0051] ;
[0052] Adaptive noise control loss, which dynamically adjusts the denoising strength and optimizes local image details, is defined as:
[0053] ;
[0054] The combination of the above loss functions forms the total loss function, whose mathematical expression is:
[0055] ;
[0056] Among them, the is the weight coefficient, which is used to balance the contribution of each loss in the overall optimization. By optimizing the total loss function and adopting the gradient descent algorithm to back-propagate and update the model parameters, the quality of the generated infrared image is gradually improved, and finally an infrared image result with clear details, physical reality and visual realism is obtained.
[0057] Beneficial effects: This method effectively captures and restores the thermal radiation characteristics of the image by fusing the forward diffusion and reverse denoising mechanisms of the diffusion model; combines the multi-scale discriminator to perform fine evaluation and optimization of the details and structures of the generated image at different spatial scales, thereby improving the visual realism and detail integrity of the image. The method further introduces the retinal cortex theory to systematically separate the reflection and illumination information, enhancing the lighting robustness and physical rationality of the generated image; and through the physical constraint loss function, ensures the high consistency of the temperature distribution and thermal radiation characteristics of the infrared image. This method has the advantage of not requiring a large amount of labeled data for training, exhibits excellent generalization ability and environmental adaptability, and can automatically adjust the generation strategy to adapt to changing real-world scenarios. It is widely used in high-precision infrared imaging applications such as military reconnaissance, night monitoring, and traffic monitoring under low-light conditions, and significantly improves the detection and recognition performance of space targets in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic flow diagram of the present invention;
[0059] Figure 2 This is a schematic diagram of visible light image separation based on retinal cortex theory;
[0060] Figure 3 It is a schematic diagram of the forward and reverse diffusion process of the diffusion model;
[0061] Figure 4 This is a schematic diagram of the evaluation infrared image of the multi-scale discriminator;
[0062] Figure 5 It is a schematic diagram of the total loss function. DETAILED DESCRIPTION
[0063] The present invention will be further described below with reference to the accompanying drawings.
[0064] like Figure 1 As shown, a method for converting visible light to infrared images based on a diffusion model includes the following steps:
[0065] S1, collect visible light images of a certain scene;
[0066] S2, performing data input operation on the visible light image collected in S1;
[0067] S3, using retinal cortex theory to perform separation operation on the image after the data input operation in S2;
[0068] S4, using the diffusion model to perform a forward diffusion operation on the image after the retinal cortex theoretical separation operation in S3;
[0069] S5, using an adaptive noise control mechanism to perform dynamic noise removal on the image after the forward diffusion operation in S4;
[0070] S6, using a diffusion model to perform a reverse diffusion operation on the image after the dynamic denoising operation in S5 and generate an infrared image;
[0071] S7, using physical formulas to perform physical consistency constraint operations on the infrared image generated after the back diffusion operation in S6;
[0072] S8, using a multi-scale discriminator to perform quality assessment on the infrared image after the physical constraint operation in S7 at different resolution scales;
[0073] S9. Use the jointly optimized total loss function to perform model parameter optimization training operations on the infrared image after multi-scale discrimination in S8.
[0074] As a preferred embodiment, the step of using the retinal cortex theory to separate the image after the data input operation in S2 in S3 is:
[0075] S3.1 introduces retinal cortex theory, which decomposes the input visible light image into a reflected image and an illuminated image, thereby eliminating the influence of lighting conditions on infrared image generation and ensuring that the generated image reflects the true surface characteristics of the object. The retinal cortex theory can be expressed by the following mathematical formula:
[0076] ;
[0077] in, Represents the spatial position coordinates of a pixel in the image, represents the input visible light image, represents the reflected image, represents the lighting image, such as Figure 2 shown.
[0078] Furthermore, the step of using the diffusion model in S4 to perform a forward diffusion operation on the image after the retinal cortex theoretical separation operation in S3 is:
[0079] S4.1 performs a forward diffusion process on the input visible light image, gradually adds noise, and generates a noise image. The mathematical formula of the forward diffusion process is as follows:
[0080] ;
[0081] in, is the image at step t, is the image of the previous step, is noise sampled from a standard normal distribution, is a parameter that controls the intensity of noise addition and gradually decreases as the number of steps increases, such as Figure 3 shown.
[0082] As a preferred embodiment, the step of using the adaptive noise control mechanism in S5 to perform dynamic noise removal on the image after the retinal cortex theoretical separation operation in S4 is:
[0083] S5.1 In the reverse denoising process, the noise removal strength is dynamically adjusted according to the local features of the image (such as gradients and edges). Adaptive noise control ensures that the image details are preserved by adjusting the denoising strength in each area. In particular, the denoising strength needs to be reduced in edge and high-frequency areas to avoid losing details. The adaptive noise control mechanism can be implemented by the following formula:
[0084] ;
[0085] in, is the adaptive noise mapping, is the control coefficient, It is the image gradient, which is used to represent the intensity of image changes.
[0086] As a preferred embodiment, the step of using the diffusion model in S6 to perform a reverse diffusion operation on the image after the dynamic denoising operation in S5 and generate an infrared image is:
[0087] S6.1 uses a denoising neural network to gradually remove noise from the noisy image, restore the original image, and generate an infrared image. The goal of the reverse denoising process is to gradually restore a clear image from the noisy image. Each step removes noise until the image is restored to its original state. This can be achieved using the following formula:
[0088] ;
[0089] in, represents the denoising network, represents the parameters of the denoising network, represents the current noise image, represents the restored image after denoising, such as Figure 3 shown.
[0090] As a preferred embodiment, the step of using a physical formula in S7 to perform a physical consistency constraint operation on the infrared image generated after the back diffusion operation in S6 is:
[0091] S7.1 introduces physical constraints to ensure that the generated infrared image complies with physical laws such as thermal radiation. Physical constraints ensure that the generated infrared image is not only visually consistent but also complies with the physical laws of thermal radiation and temperature changes, thereby ensuring the authenticity of the image. The physical loss function is shown in the formula:
[0092] ;
[0093] in, and represent the generated and real infrared images respectively, and Represent the generated and real image temperature information respectively.
[0094] As a preferred embodiment, the step of using a multi-scale discriminator in S8 to perform quality assessment on the infrared image generated after the back diffusion operation in S7 at different resolution scales is as follows:
[0095] S8.1 uses a multi-scale discriminator to evaluate the generated infrared image, ensuring that the image has good details and global consistency at different resolutions. The multi-scale discriminator optimizes the image details to ensure that the generated infrared image maintains good quality at different scales. The discriminator output at each scale is:
[0096] ;
[0097] The final multi-scale loss function is:
[0098] ;in, Indicates the The discriminator output of the scale, Indicates the The discriminative loss on scale, represents the generated infrared image, Represents a real infrared image, such as Figure 4 shown.
[0099] As a preferred embodiment, the step of using the jointly optimized total loss function in S9 to perform model parameter optimization training on the infrared image after multi-scale discrimination in S8 is as follows:
[0100] S9.1 uses a total loss function to uniformly train and optimize the diffusion model. The total loss function comprehensively considers the generative adversarial loss, diffusion loss, physical constraint loss, multi-scale discriminant loss, and adaptive noise control loss. It can comprehensively optimize the quality, physical authenticity, visual details, and noise suppression effect of the generated infrared image. The mathematical formulas of each loss function are expressed as follows:
[0101] Generative Adversarial Loss (GAN Loss), used to ensure the authenticity of generated images, is defined as:
[0102] ;
[0103] Diffusion loss, which constrains the ability of the diffusion model to gradually recover the noise-free image, is defined as:
[0104] ;
[0105] The physical constraint loss ensures that the generated image conforms to the physical properties of thermal radiation and temperature distribution and is defined as:
[0106] ;
[0107] Multi-scale discriminant loss evaluates and optimizes the details of the image at different scales and is defined as:
[0108] ;
[0109] Adaptive noise control loss, which dynamically adjusts the denoising strength and optimizes local image details, is defined as:
[0110] ;
[0111] The combination of the above loss functions forms the total loss function, whose mathematical expression is:
[0112] ;
[0113] Among them, the is the weight coefficient, which is used to balance the contribution of each loss in the overall optimization. By optimizing the total loss function, the gradient descent algorithm is used to back-propagate the model parameters to update the generated infrared image quality, so that the infrared image results with clear details, physical reality and visual realism are finally obtained. Figure 5 shown.
[0114] In summary, a diffusion-based method for converting visible light to infrared images achieves high-quality infrared image generation in complex environments by integrating the forward noise addition and backward progressive denoising mechanisms of the diffusion model with the precise identification of image details at different resolutions by a multi-scale discriminator. While ensuring physical image consistency, this method effectively separates illumination and reflectance information based on retinal cortex theory, further enhancing the realism and detail of images under low-light and complex lighting conditions. Furthermore, adaptive noise control dynamically adjusts the denoising intensity to effectively suppress the impact of noise on image quality, avoiding detail loss and artifacts. This method not only significantly improves the detail recovery and structural integrity of infrared images, but also exhibits strong generalization and adaptability, adapting to diverse scenarios and environments without requiring extensive annotated data. It has broad application potential in military reconnaissance, nighttime surveillance, intelligent transportation, and other fields, and possesses significant application value and innovative significance for improving infrared imaging quality and target recognition accuracy in low-light and complex environments.
Claims
1. A method for converting visible light to infrared images based on a diffusion model, characterized in that: The steps include: S1, collect visible light images of a certain scene; S2, performing data input operation on the visible light image collected in S1; S3, using retinal cortex theory to perform separation operation on the image after the data input operation in S2, the specific steps are: S3.1 introduces the retinal cortex theory, which decomposes the input visible light image into a reflection image and an illumination image. The retinal cortex theory is expressed by the following mathematical formula: ; in, Represents the spatial position coordinates of a pixel in the image, represents the input visible light image, represents the reflected image, represents an illumination image; S4. Use the diffusion model to perform a forward diffusion operation on the image after the retinal cortex theoretical separation operation in S3. The specific steps are as follows: S4.1 performs a forward diffusion process on the input visible light image, gradually adds noise, and generates a noise image. The mathematical formula of the forward diffusion process is as follows: ; in, is the image at step t, is the image of the previous step, is the noise sampled from a standard normal distribution, It is a parameter that controls the intensity of noise addition and gradually decreases as the number of steps increases; S5, using an adaptive noise control mechanism to perform dynamic noise removal on the image after the forward diffusion operation in S4, the specific steps are as follows: S5.1 In the reverse denoising process, the noise removal strength is dynamically adjusted according to the local characteristics of the image. The adaptive noise control mechanism is implemented by the following formula: ; in, is the adaptive noise mapping, is the control coefficient, is the image gradient, which is used to represent the intensity of image changes; S6, using a diffusion model to perform a reverse diffusion operation on the image after the dynamic denoising operation in S5 and generate an infrared image; S7, using physical formulas to perform physical consistency constraint operations on the infrared image generated after the back diffusion operation in S6; S8, using a multi-scale discriminator to perform quality assessment on the infrared image after the physical constraint operation in S7 at different resolution scales; S9. Use the jointly optimized total loss function to perform model parameter optimization training operations on the infrared image after multi-scale discrimination in S8.
2. The method for converting visible light to infrared images based on a diffusion model according to claim 1, characterized in that: The steps in S6 of using the diffusion model to perform a reverse diffusion operation on the image after the dynamic denoising operation in S5 and generating an infrared image are as follows: S6.1 uses a denoising neural network to gradually remove noise from the noisy image, restore the original image, and generate an infrared image. Each step removes the noise until the image is restored to its original state. This is achieved through the following formula: ; in, represents the denoising network, represents the parameters of the denoising network, represents the current noise image, represents the restored image after denoising.
3. The method for converting visible light to infrared images based on a diffusion model according to claim 1, characterized in that: The step of using physical formulas in S7 to perform physical consistency constraint operation on the infrared image generated after the back diffusion operation in S6 is: S7.1 introduces physical constraints to ensure that the generated infrared image complies with the physical laws of thermal radiation. The physical loss function is shown in the formula: ; in, and represent the generated and real infrared images respectively, and Represent the generated and real image temperature information respectively.
4. The method for converting visible light to infrared images based on a diffusion model according to claim 1, characterized in that: The steps of using a multi-scale discriminator in S8 to perform quality assessment on the infrared image generated after the back diffusion operation in S7 at different resolution scales are as follows: S8.1 uses a multi-scale discriminator to evaluate the generated infrared image. The discriminator output at each scale is: ;; The final multi-scale loss function is: ; in, Indicates the The discriminator output of the scale, Indicates the The discriminative loss on scale, represents the generated infrared image, Represents a real infrared image.
5. The method for converting visible light to infrared images based on a diffusion model according to claim 1, wherein: The steps of using the jointly optimized total loss function in S9 to perform model parameter optimization training on the infrared image after multi-scale discrimination in S8 are as follows: S9.1 uses a total loss function to uniformly train and optimize the diffusion model. The total loss function comprehensively considers the generative adversarial loss, diffusion loss, physical constraint loss, multi-scale discriminant loss, and adaptive noise control loss. The mathematical formulas of each loss function are as follows: Generative adversarial loss, also known as GAN loss, is used to ensure the authenticity of generated images and is defined as: ; Diffusion loss, which constrains the ability of the diffusion model to gradually recover the noise-free image, is defined as: ; The physical constraint loss ensures that the generated image conforms to the physical properties of thermal radiation and temperature distribution and is defined as: ; Multi-scale discriminant loss evaluates and optimizes the details of the image at different scales and is defined as: ; Adaptive noise control loss, which dynamically adjusts the denoising strength and optimizes local image details, is defined as: ; The combination of the above loss functions forms the total loss function, whose mathematical expression is: ; Among them, the is the weight coefficient.
Citation Information
Patent Citations
Method, diffusion model and system for generating infrared-visible light data in pairs
CN119091100A
Denoising diffusion model driven texture enhanced infrared and visible light image fusion method and system
CN119540071A