A visible-to-infrared image conversion system and method based on a diffusion model

By using a diffusion-based visible-to-infrared image conversion system, and leveraging a visual language understanding module and multi-stage progressive learning, the system addresses the issues of insufficient semantic information utilization and style control in the conversion of visible images to infrared images, thereby achieving high-quality infrared image generation.

CN120912421BActive Publication Date: 2026-04-03NINGBO INST OF NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for converting visible light images to infrared images have shortcomings in terms of semantic information utilization and style control, making it difficult to address the issues of infrared image imaging being greatly affected by the semantics of the surrounding environment and style differences between different datasets.

Method used

A visible-to-infrared image conversion system based on a diffusion model is adopted. The system generates control flags by fusing text descriptions through a visual language understanding module. Combined with the diffusion model framework, the system uses visible light images, segmentation images, and infrared single-light data to perform multi-stage progressive learning and optimize the image conversion process.

Benefits of technology

It improves the semantic consistency and style control accuracy of infrared image generation, reduces error accumulation, enhances image conversion speed and structural similarity, and solves the problem of insufficient utilization of semantic information in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912421B_ABST
    Figure CN120912421B_ABST
Patent Text Reader

Abstract

This invention relates to a visible-to-infrared image conversion system and method based on a diffusion model. By fusing text descriptions through a visual language understanding module to generate image text descriptions, and using the image text descriptions as control flags for the diffusion model algorithm, the diffusion model algorithm executed by the diffusion model framework can precisely adjust the style features of the infrared image according to semantic instructions. This ultimately overcomes the technical defects of traditional single-style image conversion in terms of insufficient style control, and makes full use of semantic information and controllable flags to significantly improve the semantic consistency of key areas of the infrared image generated by the entire system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to a visible-to-infrared image conversion system and method based on a diffusion model. Background Technology

[0002] Converting visible light images into infrared images is a core task in the field of image analysis. It aims to convert objects such as targets and scenes in visible light images into infrared modes. This technology has important application value in fields such as video surveillance, moving target detection and tracking, and ecological environment assessment.

[0003] However, the process of converting visible light images to infrared images faces complex technical challenges. First, the imaging of infrared images is highly dependent on the surrounding environment and the semantic information of the target object. For example, different parts of the same building may have similar color distributions in the visible light modality, but due to differences in heat absorption leading to varying temperature radiation, the infrared modal pixel values ​​of different parts will differ significantly. Therefore, it is necessary to analyze environmental information such as illumination and season shown in the visible light image to ensure the quality of the generated infrared image. Second, differences in imaging conditions and sensor types between different infrared datasets result in significant differences in the style of infrared images from different datasets, making it difficult to directly apply a model trained on one dataset to another. Finally, because infrared images are much harder to obtain than visible light images, publicly available infrared image datasets are relatively small. These phenomena significantly limit the usability of visible light to infrared image conversion technology.

[0004] For the task of converting visible light images to infrared images, some existing methods directly define it as a simple pixel-by-pixel translation problem and use common image generation models, such as variational autoencoders, generative adversarial networks and diffusion models, to complete this task. However, they are difficult to deal with the problem that infrared image imaging is greatly affected by the semantics of the surrounding environment.

[0005] To address the issue of infrared image imaging being significantly influenced by the semantics of the surrounding environment, some existing methods incorporate low-level semantic information, such as structural graphs and geometric consistency, into the generative model. However, the role of this low-level semantic information is limited, failing to fully utilize the abundant semantic information revealed in the visible light image. To address the problem of infrared image imaging being heavily influenced by imaging conditions and sensor type, some existing methods introduce constraints such as physical information into the generative model, but these still struggle to compensate for style differences between different datasets. Summary of the Invention

[0006] The technical problem to be solved by this invention is how to overcome the technical defects of existing visible light image to infrared image conversion methods in terms of insufficient semantic information utilization and style control. In order to overcome this technical defect, this invention provides a visible and infrared light image conversion system and method based on a diffusion model, specifically including a visible and infrared light image conversion system based on a diffusion model and a visible and infrared light image conversion method based on a diffusion model.

[0007] This invention provides a visible-to-infrared image conversion system based on a diffusion model, comprising:

[0008] The visual language understanding module is configured to perform image encoding based on self-attention mapping on the visible light image to obtain image features, then integrate a preset text description into the image features to generate control flag bits, and then decode to obtain the image text description;

[0009] The diffusion model framework, which communicates with the visual language understanding module, is configured to convert a preset noise map, the visible light map, the segmentation map of the visible light map, the infrared single-light data of the visible light map, and the image text description into three styles of infrared images through the diffusion model algorithm.

[0010] The segmentation map of the visible light image is a semantic segmentation map, an instance segmentation map, or a panoramic segmentation map of the visible light image, and the infrared single-light data of the visible light image is an infrared light image obtained by converting the visible light image into an infrared light image.

[0011] The visible-to-infrared image conversion system based on a diffusion model disclosed in this invention integrates text descriptions into an image text description through a visual language understanding module. This image text description serves as a control flag for the diffusion model algorithm, enabling the algorithm to precisely control the style features of the infrared image according to semantic instructions. This overcomes the technical shortcomings of traditional single-style conversion in terms of style control. Furthermore, the diffusion model framework combines visible light images, segmentation images, and infrared single-light data to fully utilize semantic information, significantly reducing the error accumulation of traditional multi-stage conversions. Under the same hardware conditions, both image conversion speed and structural similarity (SSIM) metrics are improved. Moreover, because the visual language understanding module also performs self-attention mapping-based image encoding on the visible light image, this self-attention mapping image encoding technique deeply binds text descriptions with visual features to fully utilize semantic information. The controllable flag significantly improves the semantic consistency of key regions (such as human outlines and high-temperature objects) in the generated infrared image, avoiding the text-image disconnect problem found in traditional methods.

[0012] In one possible implementation, the visual language understanding module includes:

[0013] An image encoder is configured to perform self-attention mapping-based image encoding on the visible light image to obtain image features;

[0014] The text decoder, which communicates with both the image encoder and the diffusion model framework, is configured to first incorporate a preset text description generation control flag into the image features, and then decode the image features incorporating the text description generation control flag to obtain an image text description.

[0015] The visual language understanding module, equipped with the above structure and functions, performs image encoding on the visible light image based on self-attention mapping and incorporates preset text description generation control flags, further ensuring that the diffusion model algorithm can accurately adjust the style features of the infrared image according to semantic instructions, thereby improving conversion efficiency.

[0016] In one possible implementation, the image encoder includes a self-attention module, an encoding controller, and a feedforward module 1 connected in series. The encoding controller also communicates with the text decoder. The encoding controller is configured to call the self-attention module to perform self-attention mapping processing on the visible light image to obtain self-attention features, and sum the self-attention features with the visible light image to obtain a summed image. Subsequently, the feedforward module 1 is called to perform feedforward network model mapping processing on the summed image to obtain a feedforward feature map, and the feedforward feature map is summed with the summed image to obtain the image features. The image encoder with the above structure forms a collaborative strategy of self-attention mechanism and residual connection. Through this collaborative optimization of self-attention mechanism and residual connection, the problems of traditional encoders ignoring long-distance dependencies, feature loss, and modal fragmentation can be solved.

[0017] In one possible implementation, the text decoder includes a causal attention module, a cross-attention module, a second feedforward module, and a decoding controller. The decoding controller communicates simultaneously with the causal attention module, the cross-attention module, the second feedforward module, and the diffusion model framework. The cross-attention module also communicates with the encoding controller.

[0018] The decoding controller is configured to call the causal attention module to perform causal attention mapping on the text description generation control flag to obtain causal attention features, and sum the causal attention features with the text description generation control flag to obtain fusion feature one; then call the cross attention module to perform cross attention mapping on the fusion feature one and the image features to obtain cross features, and sum the cross features with the fusion feature one to obtain fusion feature two; finally, call the feedforward module two to perform feedforward network model mapping processing on the fusion feature two to obtain feedforward text, and sum the feedforward text with the fusion feature two to obtain the image text description.

[0019] By using a cascaded processing mechanism that coordinates multiple modules with a decoding controller, the generation quality and semantic accuracy of image text descriptions are significantly improved, specifically through feature fusion optimization, cross-modal alignment enhancement, and improved generation efficiency.

[0020] In one possible implementation, the diffusion model framework includes:

[0021] The latent space coding module is configured to perform latent space coding on the noise map, the visible light map, the segmentation map of the visible light map, and the infrared single-light data of the visible light map to obtain code one, code two, and code three;

[0022] The denoising module communicates with both the decoding controller and the latent space coding module, and is configured to run in parallel to denoise the first code to obtain denoising result one, denoise the second code to obtain denoising result two, and denoise the fusion result of the image text description and the third code to obtain denoising result three.

[0023] The latent space decoding module communicates with the denoising module and is configured to perform latent space decoding on the denoising result one, the denoising result two, and the denoising result three respectively to obtain infrared images of three styles.

[0024] The diffusion model framework with the above structure and functions combines visible light image, segmentation image and infrared single light data with latent space coding module. With the support of denoising module and latent space decoding module, it further reduces the error accumulation of traditional multi-stage conversion, and further improves image conversion speed and structural similarity index.

[0025] In one possible implementation, the latent space coding module is configured to operate as follows:

[0026] The infrared single-light data of the visible light image is mapped using a latent space coding algorithm to obtain the first code;

[0027] The result of the mapping operation between the noise map and the latent space coding algorithm performed on the visible light map is summed to obtain code two;

[0028] The result of the mapping operation of the noise map, the result of the latent space coding algorithm performed on the visible light map, and the result of the mapping operation of the latent space coding algorithm performed on the segmentation map of the visible light map are summed to obtain code three.

[0029] The latent space coding module, operating in the manner described above, can efficiently and accurately perform latent space coding on noise maps, visible light maps, segmented maps of visible light maps, and infrared single-light data of visible light maps to obtain Code 1, Code 2, and Code 3, ensuring the orderly operation of the diffusion model algorithm.

[0030] In one possible implementation, the noise reduction module includes:

[0031] A text encoder, communicating with the decoding controller, is configured to perform text encoding on the image text description to obtain text features;

[0032] A cross-attention layer model, which communicates with the text encoder, is configured to perform cross-attention mapping on the text features to obtain cross-text features;

[0033] The first-stage denoising network communicates with the latent space coding module and is configured to denoise the first code to obtain the first denoising result.

[0034] The second-stage denoising network communicates with the latent space coding module and is configured to denoise the second code to obtain the second denoising result.

[0035] The third-stage denoising network, which communicates with both the cross-attention layer model and the latent space encoding module, is configured to denoise the fusion result of the cross-text features and the encoding three to obtain the denoised result three.

[0036] The denoising module, equipped with the aforementioned structure and functions, extracts image descriptive features through a text encoder, maps them through a cross-attention layer, and then fuses them with the image encoding. This ensures that the denoising process strictly follows the text semantic guidance, resolving the detail distortion problem caused by semantic bias in traditional denoising, and particularly improving the detail restoration accuracy in complex scenes (such as medical images and artistic creations). Furthermore, it employs a phased approach: processing basic noise in the first stage (encoding one), refining high-frequency details in the second stage (encoding two), and fusing text features in the third stage to achieve semantic calibration. This phased, progressive denoising reduces the computational load per iteration, enables dynamic resource allocation, and enhances noise robustness.

[0037] In one possible implementation, the first-stage denoising network, the second-stage denoising network, and the third-stage denoising network all include a denoising encoder and a denoising decoder. The denoising encoder is a network structure formed by sequentially connecting an input convolutional layer model, multiple cross-attention downsampling modules, a downsampling module, and a bottleneck module. The denoising decoder is a network structure formed by sequentially connecting an upsampling module, multiple cross-attention upsampling modules, and an output convolutional layer model.

[0038] The input convolutional layer models in the first-stage denoising network and the second-stage denoising network communicate with the latent space coding module. The input convolutional layer model in the third-stage denoising network communicates with both the cross-attention layer model and the latent space coding module. The upsampling module communicates with the bottleneck module. The output convolutional layer model communicates with the latent space decoding module.

[0039] The number of cross-attention downsampling modules is the same as the number of cross-attention upsampling modules, and they communicate in a one-to-one manner.

[0040] This modular design supports adaptation to different modal inputs, and can solve the problem of universality in cross-modal data collaborative denoising while achieving denoising.

[0041] In one possible implementation, the latent space decoding module includes three latent space decoders that communicate one-to-one with the output convolutional layer model to perform latent space decoding on each denoising result obtained by the denoising module, thereby further improving the acquisition of three different styles of infrared images.

[0042] Another technical solution of the present invention is to provide a visible and infrared light image conversion method based on a diffusion model, comprising the following steps:

[0043] S1: Construct an infrared single-light dataset, a visible light-infrared image pair dataset, and a style-specific visible light-infrared image pair dataset;

[0044] S2: Construct a diffusion model training framework based on progressive learning. Optimize the parameters of the diffusion model training framework using the infrared single-light dataset, visible-infrared image pair dataset, and specific-style visible-infrared image pair dataset constructed in step S1 to obtain the diffusion model framework.

[0045] S3: The visible light image is encoded using a self-attention mapping-based image encoding method by the visual language understanding module to obtain image features. The image features are then decoded by incorporating a preset text description into the image features to generate control flag bits, thereby obtaining an image text description.

[0046] S4: The preset noise map, the visible light map, the segmentation map of the visible light map, the infrared single-light data of the visible light map, and the image text description are converted into three styles of infrared images through the diffusion model framework.

[0047] The visible-to-infrared image conversion method based on a diffusion model disclosed in this invention first constructs an infrared single-light dataset, a visible-infrared image pair dataset, and a specific-style visible-infrared image pair dataset. These datasets are then used to perform three stages of parameter optimization on the diffusion model training framework to obtain the diffusion model framework. This effectively combines the visual language information provided by a multimodal visual language model, achieving semantically driven visible-to-infrared image conversion. Furthermore, this three-stage progressive learning framework addresses the problem of excessive differences in infrared imaging styles between different datasets. The collection of infrared single-light datasets, visible-infrared image pair datasets, and specific-style visible-infrared image pair datasets supports the entire training process, achieving high-quality visible-to-infrared image conversion and overcoming the technical shortcomings of existing visible-to-infrared image conversion methods in terms of semantic information utilization and style control. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of a visible-to-infrared light image conversion system based on a diffusion model disclosed in an embodiment of the present invention;

[0049] Figure 2 This is a schematic diagram of the operation of the visual language understanding module disclosed in the embodiments of the present invention;

[0050] Figure 3 This is a schematic diagram illustrating the operation of the diffusion model framework disclosed in this embodiment of the invention;

[0051] Figure 4 This is a schematic diagram of the first-stage denoising network, the second-stage denoising network, or the third-stage denoising network structure disclosed in the embodiments of the present invention.

[0052] Figure 5 This is a flowchart of the method disclosed in the embodiments of the present invention. Detailed Implementation

[0053] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of this application and are not intended to limit the scope of protection of the embodiments of this application. Those skilled in the art can make adjustments as needed to adapt to specific application scenarios.

[0054] In the description of the embodiments of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "communication connection" and "communication" should be interpreted broadly, that is, a connection method that enables the transmission of information between each other. For example, it can be a circuit connection achieved through conductive lines, an electrical connection achieved through a radio signal channel, or a combination of both.

[0055] In the embodiments of this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0056] The present application will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0057] See Figure 1 This application discloses a visible-to-infrared light image conversion system based on a diffusion model. Figure 1 The diagram shows the structure of the conversion system, which includes a visual language understanding module and a diffusion model framework, wherein the diffusion model framework communicates with the visual language understanding module.

[0058] See Figure 1 and Figure 2 In this conversion system, the visual language understanding module is configured to perform image encoding based on self-attention mapping on the visible light image to obtain image features, then incorporate preset text description generation control flags into the image features, and finally decode to obtain the image text description. In this embodiment, the visual language understanding module includes an image encoder and a text decoder, wherein the text decoder communicates with both the image encoder and the diffusion model framework.

[0059] See Figure 2 In the visual language understanding module, the image encoder is configured to perform self-attention mapping-based image encoding on the visible light image to obtain image features. In this embodiment, the image encoder includes a self-attention module and an encoding controller connected in series. Figure 2 The diagram illustrates its operating principle and feedforward module one. The encoder controller also communicates with the text decoder. Figure 2As shown, the encoding controller is configured to call the self-attention module to perform self-attention mapping on the visible light image to obtain self-attention features, and sum the self-attention features with the visible light image to obtain a summed image; then call the feedforward module to perform feedforward network model mapping on the summed image to obtain a feedforward feature map, and sum the feedforward feature map with the summed image to obtain image features.

[0060] See Figure 2 In the visual language understanding module, the text decoder is configured to generate control flags by incorporating preset text descriptions into image features. Figure 2 The image features are decoded to obtain image text descriptions by using a preset text description generation control flag (hereinafter referred to as text description generation control). In this embodiment, the text decoder is an image feature-based text decoder, including a causal attention module, a cross-attention module, a feedforward module, and a decoding controller (…). Figure 2 The diagram illustrates the operating principle of the decoding controller. The decoding controller communicates simultaneously with the causal attention module, the cross-attention module, feedforward module two, and the diffusion model framework. The cross-attention module also communicates with the encoding controller. In this embodiment, feedforward module one and feedforward module two are feedforward modules of the same specification. It should be noted that if resource consumption is to be reduced, the text decoder and image encoder could share a single feedforward module. However, to avoid crosstalk, this embodiment uses a method where the text decoder and image encoder each have a feedforward module of the same specification, such as... Figure 2 As shown.

[0061] Please continue reading Figure 2 In the text decoder, the decoding controller is configured to call the causal attention module to perform causal attention mapping on the text description generation control flag to obtain causal attention features, and sum the causal attention features with the text description generation control flag to obtain fusion feature one; then call the cross attention module to perform cross attention mapping on fusion feature one and image features to obtain cross features, and sum the cross features with fusion feature one to obtain fusion feature two; finally call the feedforward module two to perform the mapping processing of fusion feature two on the feedforward network model to obtain feedforward text, and sum the feedforward text with fusion feature two to obtain image text description.

[0062] See Figure 1 , Figure 2 , Figure 3 and Figure 4 In this conversion system, the diffusion model framework is set up to use a diffusion model algorithm to convert the preset noise map, visible light map, visible light map segmentation map, and infrared single-light data of the visible light map (…). Figure 3The infrared single-light data of the visible light image (referred to as infrared image) and the image text description are converted into three styles of infrared images, namely... Figure 3 Examples of the first, second, and third stages of generation are provided. The segmentation map of the visible light image can be a semantic segmentation map, an instance segmentation map, or a panoramic segmentation map. Figure 2 The leftmost part of the diagram shows the origin of the visible light image segmentation map used in this embodiment, which is obtained by transforming the visible light image through a pre-trained segmentation model. The infrared single-light data of the visible light image is any infrared image obtained by performing visible-infrared image conversion on the visible light image. This can be the result of performing visible-infrared image conversion on the visible light image using existing variational autoencoders, generative adversarial networks, or diffusion models.

[0063] See Figure 1 and Figure 3 In this embodiment, the diffusion model framework includes a latent space encoding module, a denoising module, and a latent space decoding module. The denoising module communicates with both the decoding controller and the latent space encoding module, and the latent space decoding module communicates with the denoising module.

[0064] In the diffusion model framework, the latent space coding module is configured to perform latent space coding on the noise map, visible light map, visible light map segmentation map, and infrared single-light data of the visible light map to obtain Code 1, Code 2, and Code 3. Figure 3 The diagram illustrates the operating principle of the latent space coding module within the diffusion model framework. In this embodiment, the latent space coding module is configured to operate as follows: First, the infrared single-light data of the visible light image is mapped using a latent space coding algorithm to obtain code one; second, the noise image is summed with the mapping result of the latent space coding algorithm performed on the visible light image to obtain code two; third, the noise image, the mapping result of the latent space coding algorithm performed on the visible light image, and the mapping result of the latent space coding algorithm performed on the segmented image of the visible light image are summed to obtain code three.

[0065] Specifically, the latent space coding module operates according to the following calculation formula:

[0066] ,

[0067] ,

[0068] ,

[0069] In the formula,

[0070] Infrared single-light data representing the visible light pattern;

[0071] Represents a noise graph;

[0072] A segmented image representing the visible light spectrum;

[0073] Represents a visible light diagram;

[0074] The mapping operation represents the latent space coding algorithm;

[0075] This represents the summation operation;

[0076] Represents code one;

[0077] Representative code two;

[0078] This represents code three.

[0079] See Figure 3 and Figure 4 In the diffusion model framework, the denoising module is configured to run in parallel: denoising code 1 to obtain denoised result 1; denoising code 2 to obtain denoised result 2; and denoising the fusion result of the image text description and code 3 to obtain denoised result 3. See also... Figure 3 In this embodiment, the denoising module includes a text encoder and a cross-attention layer model (…). Figure 3 The denoising network consists of three stages: a cross-attention layer (referred to as the cross-attention layer), a first-stage denoising network, a second-stage denoising network, and a third-stage denoising network. The text encoder communicates with the decoding controller, the cross-attention layer model communicates with the text encoder, the first-stage denoising network communicates with the latent space coding module, the second-stage denoising network communicates with the latent space coding module, and the third-stage denoising network communicates with both the cross-attention layer model and the latent space coding module.

[0080] In the denoising module, the text encoder is configured to encode the text description of the image to obtain text features. The cross-attention layer model is configured to perform cross-attention mapping on the text features to obtain cross-text features. The first-stage denoising network is configured to denoise the first encoding to obtain denoised result one. The second-stage denoising network is configured to denoise the second encoding to obtain denoised result two. The third-stage denoising network is configured to denoise the fusion result of the cross-text features and the third encoding to obtain denoised result three.

[0081] See Figure 4 In the denoising module, the first-stage denoising network, the second-stage denoising network, and the third-stage denoising network all contain a denoising encoder and a denoising decoder. The denoising encoder is a model consisting of an input convolutional layer (…). Figure 4 The network structure consists of an input convolutional layer (referred to as the input convolutional layer), multiple cross-attention downsampling modules, a downsampling module, and a bottleneck module, connected in series. The denoising decoder is a model consisting of an upsampling module, multiple cross-attention upsampling modules, and an output convolutional layer. Figure 4 The network structure consists of concatenated output convolutional layers (referred to as output convolutional layers in the original text). In the first and second stage denoising networks, the input convolutional layer models communicate with the latent space coding module. In the third stage denoising network, the input convolutional layer models communicate with both the cross-attention layer model and the latent space coding module. The upsampling module communicates with the bottleneck module, and all output convolutional layer models communicate with the latent space decoding module. The number of cross-attention downsampling modules is the same as the number of cross-attention upsampling modules, and they communicate one-to-one. In this embodiment, the bottleneck module is the two-dimensional cross-attention UNet bottleneck module. Figure 4 In this context, N represents the number of channels, and H and W are the height and width of the input image after passing through the latent space encoder.

[0082] See Figure 3 In the diffusion model framework, the latent space decoding module is configured to perform latent space decoding on denoised result one, denoised result two, and denoised result three respectively to obtain infrared images of three styles. Figure 3 The diagram illustrates the operating principle of the latent space decoding module. The mapping operation represents the latent space decoder model. In this embodiment, the latent space decoding module includes three latent space decoders, which communicate with the output convolutional layer model in a one-to-one manner.

[0083] See Figure 5 The conversion method corresponding to the visible and infrared light image conversion system based on the diffusion model in this embodiment will be further disclosed below. Figure 5 The overall flowchart of this conversion method is as follows:

[0084] S1: Constructing an infrared single-light dataset, a visible light-infrared image pair dataset, and a style-specific visible light-infrared image pair dataset. In this embodiment, a large-scale infrared single-light dataset, a large-scale visible light-infrared image pair dataset, and a style-specific visible light-infrared image pair dataset were constructed. Specifically, the infrared single-light dataset contains 500,000 infrared single-light images, the visible light-infrared image pair dataset contains 70,000 visible light-infrared image pairs, and the style-specific visible light-infrared image pair dataset contains 4,300 image pairs. To ensure experimental fairness, the style-specific visible light-infrared image pair dataset is not included in the visible light-infrared image pair dataset.

[0085] S2: Construct a diffusion model training framework based on progressive learning. Optimize the parameters of the diffusion model training framework using the infrared single-light dataset, visible-infrared image pair dataset, and specific-style visible-infrared image pair dataset constructed in step S1 to obtain the diffusion model framework. Figure 3 The dashed arrows in the diagram represent the training order. In this embodiment, a diffusion model training framework based on progressive learning is constructed. The parameter optimization process includes two processes: training and testing. The training consists of three stages. Through the three stages of training, the visual language features of the visible light image and the style features of the specific infrared image are received, and corresponding high-quality infrared images are generated.

[0086] The established diffusion model training framework is as follows: Figure 3 As shown, it includes a latent space encoding module, a latent space decoding module, a text encoder, and a cross-attention layer model. Figure 3 The denoising network consists of three stages: a cross-attention layer (referred to as the cross-attention layer), a first-stage denoising network, a second-stage denoising network, and a third-stage denoising network. The overall architecture of the three stages is the same. During training, the latent space encoding module, the latent space decoding module, the text encoder, and the cross-attention layer model do not participate in parameter updates; only the parameters of the three denoising networks are updated during training.

[0087] Specifically, step S2 includes the following steps:

[0088] S21: The first stage of training for the diffusion model training framework is performed based on a large-scale infrared single-light dataset; the objective function for the first stage of training is first constructed as follows:

[0089] ,

[0090] In the formula, This represents the objective function for fine-tuning the training framework of the diffusion model in the first stage. This represents the mean. This refers to the input image (generally referring to infrared single-light data of the visible light image, noise image, segmentation image of the visible light image, or visible light image). Represents real noise. Indicating in guiding conditions Next The noise predicted by the step-by-step denoising network This represents a denoising network. express noise map of the step, Indicates the number of denoising steps. Indicates guiding conditions. The representative standard is the Zhengtai distribution.

[0091] Subsequently, the random noise map and the keyword "infrared" are input into the diffusion model training framework. At this point, the guiding condition in the objective function update of the diffusion model training framework is the keyword "infrared," and the infrared image is used as the training target. The diffusion model training framework is then fine-tuned using low-rank matrix factorization, enabling it to generate infrared images based on the keyword "infrared." The formula for low-rank matrix factorization is as follows:

[0092] ,

[0093] ,

[0094] in, This represents the model parameters of the first-stage denoising network in the updated diffusion model training framework. This represents the model parameters of the original first-stage denoising network. This indicates the updated parameters of the first-stage denoising network before and after the update. and This represents the matrix obtained after performing low-rank decomposition on the updated parameter matrix. ,and This indicates the dimension of the parameters in the first-stage denoising network model. Indicates the magnitude of the intrinsic rank. .

[0095] S22: The diffusion model training framework is trained in the second stage based on a large-scale visible light-infrared image dataset. The noise map and the visible light map are combined by adding channels. At this time, the objective function of the diffusion model training framework is the same as the objective function in step S21, but the guiding condition is the visible light map and the corresponding infrared image is used as the training target. The diffusion model training framework is used to extract the general mapping relationship between the visible light map and the corresponding infrared image, so that the diffusion model training framework has the ability to convert the visible light image into the infrared image.

[0096] S23: Construct a visual language understanding module to obtain image-text descriptions of the visible light image, and use a pre-trained segmentation model to obtain a segmentation map. Combine the segmentation map, visible light image, and random noise image by adding channels together, and introduce the image-text description in the form of cross-attention. At this point, the objective function of the diffusion model training framework follows the objective function in step S21, but the guiding conditions are the visual language information and the visible light image. Infrared images are used as the training target, and style information is obtained based on a specific style of visible light-infrared image dataset. The diffusion model training framework is used to fuse visual language information and style information to obtain infrared images of a specific style. The introduction of various conditions adopts a classifier-free guided loss function, described as follows:

[0097] ,

[0098] in, This represents the loss function for a classifier-free guided policy. This indicates the prediction of conditionally guided diffusion by the training framework of the diffusion model. express noise map of the step, This represents the visible light map in the guiding conditions. This represents the segmentation graph in the guiding conditions. This refers to the image and text description in the guiding conditions. This indicates the weight of the visible light map in the guiding conditions. This indicates the weight of the segmentation graph in the guiding conditions. This indicates the weight of the image and text descriptions in the guiding conditions. This indicates that the guiding condition is empty.

[0099] S24: The diffusion model training framework was tested on a dataset of visible-infrared image pairs with specific styles. Specifically, Frechessel perceptual distance, peak signal-to-noise ratio, and structural similarity were used as image generation quality evaluation metrics to assess the generation performance of visible-infrared images based on progressive learning and visual language understanding. The Frechessel perceptual distance formula is:

[0100] ,

[0101] ,

[0102] ,

[0103] in, Characteristics representing a true infrared image, This represents the mean output of the Inception feature extractor. Represents a true infrared image. Features representing the generation of infrared images, This indicates the generation of an infrared image. This indicates that Reichert sensed distance. This represents the mean of the features of a true infrared image. This represents the mean of the features generated in the infrared image. Represents the trace of a matrix. The covariance representing the features of a true infrared image. This represents the covariance of the features generated in the infrared image.

[0104] The formula for peak signal-to-noise ratio is:

[0105] ,

[0106] ,

[0107] in, This represents the mean square error between the actual infrared image and the generated infrared image. Indicates the length of the infrared image. Indicates the width of the infrared image. Pixels representing a real infrared image Represents the pixels that generate the infrared image. This represents the maximum value of the image color. This represents the peak signal-to-noise ratio.

[0108] The formula for structural similarity is:

[0109] ,

[0110] in, Indicating the structural similarity, Represents a true infrared image. This indicates the generation of an infrared image. This represents the mean of the true infrared image. This represents the mean of the generated infrared image. Represents the variance of the true infrared image. This represents the variance of the generated infrared image. This represents the covariance between the real infrared image and the generated infrared image. , This represents two constants.

[0111] S3: The visible light image is encoded using a self-attention mapping-based method by the visual language understanding module to obtain image features. The image features are then decoded by incorporating a preset text description into the image features to generate control flags, thereby obtaining the image text description.

[0112] S4: Using a diffusion model framework, the preset noise map, visible light map, visible light map segmentation map, infrared single-light data of the visible light map, and image text description are converted into infrared images of three styles.

[0113] The aforementioned method first constructs an infrared single-light dataset, a visible light-infrared image pair dataset, and a style-specific visible light-infrared image pair dataset. It then uses these datasets to perform three stages of parameter optimization on the diffusion model training framework to obtain the diffusion model framework. This effectively combines the visual language information provided by the multimodal visual language large model, achieving semantically driven visible light image to infrared image conversion. Furthermore, this three-stage progressive learning framework addresses the problem of excessive differences in infrared imaging styles between different datasets. The collection of large-scale infrared single-light datasets, large-scale visible light-infrared image pair datasets, and style-specific visible light-infrared image pair datasets supports the entire training process, achieving high-quality visible light image to infrared image conversion. This overcomes the technical shortcomings of existing visible light image to infrared image conversion methods in terms of insufficient semantic information utilization and style control.

[0114] The visible-to-infrared image conversion system based on a diffusion model disclosed in this embodiment generates image-text descriptions by fusing text descriptions through a visual language understanding module. These image-text descriptions serve as control flags for the diffusion model algorithm, enabling the algorithm to precisely control the style features of the infrared image according to semantic instructions. This overcomes the technical shortcomings of traditional single-style conversion in terms of style control. Furthermore, by combining visible light images, segmentation images, and single-light infrared data, the diffusion model framework significantly reduces the error accumulation of traditional multi-stage conversions. Under the same hardware conditions, both image conversion speed and structural similarity indicators are improved. Moreover, by employing self-attention mapping image encoding technology to deeply bind text descriptions with visual features, and fully utilizing semantic information, the controllable flags significantly improve the semantic consistency of key regions in the generated infrared image, avoiding the text-image disconnect problem found in traditional methods.

[0115] In the description of the embodiments of this application, it should be noted that the terms "inner" and "outer" and other terms indicating direction or positional relationship are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this application.

[0116] In the description of this application, the references to terms such as "an embodiment," "some embodiments," "in this embodiment," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0117] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A visible-to-infrared image conversion system based on a diffusion model, characterized in that, include: The visual language understanding module is configured to perform image encoding based on self-attention mapping on the visible light image to obtain image features, then integrate a preset text description into the image features to generate control flag bits, and then decode to obtain the image text description; The diffusion model framework, which communicates with the visual language understanding module, is configured to convert a preset noise map, the visible light map, a segmentation map of the visible light map, infrared single-light data, and the image text description into infrared images of three styles through a diffusion model algorithm. The segmented image of the visible light image is obtained by transforming the visible light image using a pre-trained segmentation model; The diffusion model framework comprises three training phases: the first phase fine-tunes the framework using low-rank matrix factorization, enabling it to generate infrared images based on the keyword "infrared"; the second phase extracts a universal mapping relationship between visible light images and corresponding infrared images, allowing the framework to convert visible light images into infrared images; and the third phase obtains style information from a dataset of visible light-infrared images with specific styles, then fuses visual language information and style information using the diffusion model framework to obtain infrared images with specific styles. The visual language understanding module includes an image encoder and a text decoder. The image encoder includes a self-attention module, an encoding controller, and a feedforward module connected in series. The encoding controller also communicates with the text decoder. The encoding controller is configured to call the self-attention module to perform self-attention mapping processing on the visible light image to obtain self-attention features, and sum the self-attention features with the visible light image to obtain a summed image. Subsequently, it calls the feedforward module to perform feedforward network model mapping processing on the summed image to obtain a feedforward feature map, and sum the feedforward feature map with the summed image to obtain the image features. The text decoder includes a causal attention module, a cross attention module, a second feedforward module, and a decoding controller. The decoding controller communicates with the causal attention module, the cross attention module, the second feedforward module, and the diffusion model framework. The cross attention module also communicates with the encoding controller. The decoding controller is configured to call the causal attention module to perform causal attention mapping on the text description generation control flag to obtain causal attention features, and sum the causal attention features with the text description generation control flag to obtain fusion feature one; then call the cross attention module to perform cross attention mapping on the fusion feature one and the image features to obtain cross features, and sum the cross features with the fusion feature one to obtain fusion feature two; finally, call the feedforward module two to perform feedforward network model mapping processing on the fusion feature two to obtain feedforward text, and sum the feedforward text with the fusion feature two to obtain the image text description; The diffusion model framework includes: The latent space coding module is configured to perform latent space coding on the noise map, the visible light map, the segmentation map of the visible light map, and the infrared single light data to obtain code one, code two, and code three; The denoising module communicates with both the decoding controller and the latent space coding module, and is configured to run in parallel to denoise the first code to obtain denoising result one, denoise the second code to obtain denoising result two, and denoise the fusion result of the image text description and the third code to obtain denoising result three. The latent space decoding module communicates with the denoising module and is configured to perform latent space decoding on the denoising result one, the denoising result two, and the denoising result three respectively to obtain infrared images of three styles; The noise reduction module includes: A text encoder, communicating with the decoding controller, is configured to perform text encoding on the image text description to obtain text features; A cross-attention layer model, which communicates with the text encoder, is configured to perform cross-attention mapping on the text features to obtain cross-text features; The first-stage denoising network communicates with the latent space coding module and is configured to denoise the first code to obtain the first denoising result. The second-stage denoising network communicates with the latent space coding module and is configured to denoise the second code to obtain the second denoising result. The third-stage denoising network, which communicates with both the cross-attention layer model and the latent space encoding module, is configured to denoise the fusion result of the cross-text features and the encoding three to obtain the denoised result three.

2. The visible-to-infrared image conversion system based on a diffusion model according to claim 1, characterized in that, The latent space coding module is configured to operate as follows: The infrared single-light data is mapped using a latent space coding algorithm to obtain the first code; The result of the mapping operation between the noise map and the latent space coding algorithm performed on the visible light map is summed to obtain code two; The result of the mapping operation of the noise map, the result of the latent space coding algorithm performed on the visible light map, and the result of the mapping operation of the latent space coding algorithm performed on the segmentation map of the visible light map are summed to obtain code three.

3. The visible-to-infrared image conversion system based on a diffusion model according to claim 2, characterized in that, The first-stage denoising network, the second-stage denoising network, and the third-stage denoising network all include a denoising encoder and a denoising decoder. The denoising encoder is a network structure formed by sequentially connecting an input convolutional layer model, multiple cross-attention downsampling modules, a downsampling module, and a bottleneck module. The denoising decoder is a network structure formed by sequentially connecting an upsampling module, multiple cross-attention upsampling modules, and an output convolutional layer model. The input convolutional layer models in the first-stage denoising network and the second-stage denoising network communicate with the latent space coding module. The input convolutional layer model in the third-stage denoising network communicates with both the cross-attention layer model and the latent space coding module. The upsampling module communicates with the bottleneck module. The output convolutional layer model communicates with the latent space decoding module. The number of cross-attention downsampling modules is the same as the number of cross-attention upsampling modules, and they communicate in a one-to-one manner.

4. The visible-to-infrared image conversion system based on a diffusion model according to claim 3, characterized in that, The latent space decoding module contains three latent space decoders, which communicate with the output convolutional layer model in a one-to-one manner.

5. A method for converting visible and infrared light images based on a diffusion model, characterized in that, The visible-to-infrared image conversion system based on the diffusion model according to any one of claims 1-4 includes the following steps: S1: Construct an infrared single-light dataset, a visible light-infrared image pair dataset, and a style-specific visible light-infrared image pair dataset; S2: Construct a diffusion model training framework based on progressive learning. Optimize the parameters of the diffusion model training framework using the infrared single-light dataset, visible-infrared image pair dataset, and specific-style visible-infrared image pair dataset constructed in step S1 to obtain the diffusion model framework. S3: The visible light image is encoded using a self-attention mapping-based image encoding method by the visual language understanding module to obtain image features. The image features are then decoded by incorporating a preset text description into the image features to generate control flag bits, thereby obtaining an image text description. S4: The preset noise map, the visible light map, the segmentation map of the visible light map, the infrared single light data, and the image text description are converted into three styles of infrared images through the diffusion model framework.

Citation Information

Patent Citations

  • Method and device for converting visible light image into infrared image based on feature fusion

    CN119863378A