Small sample infrared data generation method
By building a diffusion model for small sample infrared data generation and using a small amount of visible light and infrared image data for training, the problem of inability to generate infrared images in the existing technology is solved, efficient and economical infrared data generation is achieved, and the robustness of the model and actual task performance are improved.
Patent Information
- Application Number
- CN202411783486.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The existing diffusion model cannot effectively generate infrared images, especially under small sample data conditions, and the existing visible light to infrared conversion model has limited performance, making it difficult to meet image processing requirements.
A diffusion model is constructed for small sample infrared data generation, and a diffusion encoder, a diffusion middleware and a diffusion decoder are designed, combined with a VAE decoder, and a small amount of visible and infrared image data are used for training to generate high-resolution infrared images corresponding to the visible light image.
The large amount of infrared data generation in the case of small sample infrared data is realized, providing an efficient and economical solution for the generation of small sample infrared data, improving the robustness of the model and its performance stability in actual visual tasks.
Smart Images

Figure CN119941548A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of small sample multimodal image generation algorithms, and in particular relates to a small sample infrared data generation method. Background Art
[0002] In recent studies, diffusion models have attracted much attention for their excellent performance in data generation. The design of these models is inspired by the diffusion process in physics. In the field of deep learning, diffusion models provide a rich data resource for tasks that are expensive to collect and involve privacy. The model is able to generate multiple data modalities including images, videos, and audio. The working principle of the diffusion model includes two stages: the forward stage, which gradually adds noise to the real data until it is transformed into a Gaussian noise distribution; the reverse stage, which reversely removes noise from the Gaussian noise distribution through neural network learning to generate samples close to the original data distribution. At present, diffusion models have been widely used in fields such as image generation and restoration, and have become an important tool for data expansion.
[0003] Current research on diffusion models mainly focuses on the generation of RGB images in the visible light band, but models and systems for computer vision tasks can also utilize other modal information, such as temperature information, image depth information, and texture perception information. For example, at night or in low light conditions, RGB images provide limited information, while infrared images are able to detect objects with temperatures above absolute zero by capturing thermal radiation with a wavelength of 0.75-15 microns. Therefore, the richer the modal data obtained, the more robust the model is, and the more stable its performance in actual visual tasks. However, current infrared image acquisition equipment is usually expensive, which poses an economic burden on the research team; and manual proofreading is required to ensure data quality; in addition, the existing visible light to infrared conversion models have limited performance and there are bottlenecks in conversion quality.
[0004] In view of this, there is an urgent need to develop a diffusion model that can generate infrared data. Summary of the invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art.
[0006] The present invention provides a method for generating small sample infrared data, the method comprising: constructing a diffusion model for generating small sample infrared data, generating an infrared image corresponding to a visible light image based on the constructed diffusion model for generating small sample infrared data;
[0007] Among them, building a diffusion model for small sample infrared data generation specifically includes: preparing data, designing a diffusion model, designing a loss function and optimizing the diffusion model;
[0008] Prepare data: Collect several pairs of representative visible light images and corresponding infrared images as small sample reference images; also collect several pairs of visible light images and corresponding infrared images. The visible light images are used to generate image condition images. c , the infrared image is used as the real infrared image when calculating the loss function Image gt ;
[0009] Design diffusion model: The diffusion model architecture includes diffusion encoder, diffusion middleware and diffusion decoder;
[0010] Optimized diffusion model: The image condition image formed by any original noise data n0, all representative samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function, and perform back propagation to continuously optimize the diffusion model.
[0011] By applying the technical solution of the present invention, a method for generating small sample infrared data is provided. The method drives a mature visible light diffusion data generation model through small sample infrared data, uses a small amount of visible light and infrared image data for training, and generates a high-resolution infrared image corresponding to the visible light image. The present invention is based on the visible light data diffusion model and guided by a small amount of infrared data to achieve the generation of a large amount of infrared data in the case of small sample infrared data, providing an efficient and economical solution for the generation of small sample infrared data. Compared with the prior art, the technical solution of the present invention can solve the technical problem that the diffusion model in the prior art cannot meet the image processing requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The included drawings are used to provide a further understanding of the embodiments of the present invention, which constitute a part of the specification, are used to illustrate the embodiments of the present invention, and together with the text description, explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0013] Figure 1 A schematic diagram showing the principle of generating small sample infrared data provided according to a specific embodiment of the present invention is shown. DETAILED DESCRIPTION
[0014] It should be noted that, in the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present invention and its application or use. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0015] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0016] Unless otherwise specifically stated, the relative arrangement, numerical expressions and numerical values of the parts and steps set forth in these embodiments do not limit the scope of the present invention. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the accompanying drawings are not drawn according to actual proportional relationships. The technology, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, the technology, methods and equipment should be considered as a part of the specification. In all examples shown and discussed here, any specific value should be interpreted as being merely exemplary, rather than as a limitation. Therefore, other examples of exemplary embodiments may have different values.
[0017] like Figure 1 As shown, according to a specific embodiment of the present invention, a method for generating small sample infrared data is provided, the method comprising: constructing a diffusion model for generating small sample infrared data, and generating an infrared image corresponding to a visible light image based on the constructed diffusion model for generating small sample infrared data;
[0018] Among them, building a diffusion model for small sample infrared data generation specifically includes: preparing data, designing a diffusion model, designing a loss function and optimizing the diffusion model;
[0019] Prepare data: Collect several pairs of representative visible light images and corresponding infrared images as small sample reference images; also collect several pairs of visible light images and corresponding infrared images. The visible light images are used to generate image condition images. c , the infrared image is used as the real infrared image when calculating the loss function Imagegt ;
[0020] Design diffusion model: The diffusion model architecture includes diffusion encoder, diffusion middleware and diffusion decoder;
[0021] Optimized diffusion model: The image condition image formed by any original noise data n0, all representative samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function, and perform back propagation to continuously optimize the diffusion model.
[0022] By applying this configuration, a method for generating small sample infrared data is provided. The method drives a mature visible light diffusion data generation model through small sample infrared data, uses a small amount of visible light and infrared image data for training, and generates a high-resolution infrared image corresponding to the visible light image. The present invention is based on the visible light data diffusion model and guided by a small amount of infrared data to achieve the generation of a large amount of infrared data in the case of small sample infrared data, providing an efficient and economical solution for the generation of small sample infrared data.
[0023] The method of the present invention is divided into two stages: a training stage and an inference stage. The training stage is to adjust a model that can only generate images randomly to a model that can generate infrared images, that is, to build a diffusion model for generating small sample infrared data; the inference stage is to apply the diffusion model built for generating small sample infrared data to generate an infrared image corresponding to the visible light image.
[0024] Firstly, in the present invention, it is necessary to construct a diffusion model for generating small sample infrared data.
[0025] Constructing a diffusion model for small sample infrared data generation specifically includes: preparing data, designing the diffusion model architecture, designing the loss function, and optimizing the diffusion model. Each step is introduced in detail below.
[0026] Prepare data: Collect several pairs of representative visible light images and corresponding infrared images as small sample reference images; also collect several pairs of visible light images and corresponding infrared images. The visible light images are used to generate image condition images. c , the infrared image is used as the real infrared image when calculating the loss function Image gt ;
[0027] As a specific embodiment of the present invention, the number of reference images may be appropriately expanded.
[0028] (2) Design of diffusion model: The diffusion model architecture includes diffusion encoder, diffusion middleware and diffusion decoder.
[0029] The diffusion encoder receives the original noise data n0 and the text condition text c and time code time c As input conditions. The original noise data n0 represents the initial state of the model; the text condition text c Used to describe the language information related to noise data; time code time c Used to provide the model with an understanding of the time step, ensuring that temporal information is appropriately incorporated into the diffusion process.
[0030] As a specific embodiment of the present invention, the present invention can use CLIP technology to encode text to form text conditions. c It can be expressed as:
[0031] text c = CLIP(text description)
[0032] The text description is a text description of the image to be generated.
[0033] The task of the diffusion encoder is to map these input conditions to a preliminary noise code z0. The function of the diffusion encoder can be expressed as:
[0034] z0=f encoder (n0,text c ,time c )
[0035] Among them, f encoder Represents the nonlinear mapping function of the encoder, which is responsible for combining the original noise data with the conditional input to generate the noise code z0. The noise code is used as the input for the subsequent diffusion process.
[0036] The diffusion middleware receives two main inputs: the noise code z0 and the image condition image c , where the image condition image c Represents visual information related to the noise data, such as image features or scene description. Through the diffusion middleware processing, the noise code z0 is transformed into the intermediate code z. The key to this step is how to adjust the distribution of noise according to the image conditions.
[0037] As a specific embodiment of the present invention, the present invention can use the ControlNet model and convolution technology to encode the image to form an image condition. The image condition formed by using the convolution technology to encode the image can be expressed as:
[0038] image c =ControlNet(Conv(concat(image_pair)),Conv(image_v))
[0039] Among them, image c is the image condition, Conv is the convolution operation, ControlNet is the public model, concat is the dimension splicing operation, which splices the infrared image and the visible light image together, image_pair is a representative small sample data, and image_v is the corresponding visible light image of the infrared image to be generated.
[0040] The functionality of diffusion middleware can be expressed as:
[0041] z=f middleware (z0,image c )
[0042] Among them, f middleware Represents the mapping function of the middleware, which is used to encode z0 and image condition according to the input noise c Generate an intermediate code z.
[0043] The diffusion decoder is the last step in the diffusion model. The diffusion decoder receives the intermediate code z and the image condition image c As input, the denoised noise code z1 is output, and the denoised noise code z1 is decoded. The core goal of this process is to remove useless noise information through the denoising operation and restore a clearer and qualified image or signal according to the image conditions.
[0044] The denoised output of the diffusion decoder can be expressed as:
[0045] z1=f decoder (z,image c )
[0046] Among them, f decoder is the mapping function of the decoder, used to combine the intermediate code z and the image condition image c De-noising is performed to generate the de-noised noise code z1.
[0047] The present invention uses a VAE decoder to decode the denoised noise code, which can be expressed as:
[0048] Image gen =VAE decoder (z1)
[0049] Among them, VAE decoder is the VAE decoder, Image genThe final infrared image is generated.
[0050] (3) Design loss function:
[0051] In the model of the present invention, the design of the loss function is crucial and directly affects the optimization process of model training and the final generation effect. In order to effectively evaluate the difference between the generated infrared image and the real infrared image, the present invention adopts Mean Squared Error (MSE) as the loss function.
[0052] As a specific embodiment of the present invention, it is assumed that the infrared image generated by the model is Image gen , the real infrared image is Image gt , then the mean square error (MSE) loss function L MES It can be expressed as:
[0053]
[0054] Where N is the number of pixels in the image or the total number of elements contained in the image; Image Gen,i is the value of the generated infrared image at the i-th pixel; Image gt,i is the value of the real infrared image at the i-th pixel.
[0055] (4) Optimize the diffusion model:
[0056] The image condition image formed by any original noise data n0, all representative small samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function L MES , back propagation is performed to continuously optimize the diffusion model.
[0057] Furthermore, in the present invention, after the construction of the diffusion model for generating small sample infrared data is completed, an infrared image corresponding to the visible light image is generated based on the constructed diffusion model for generating small sample infrared data.
[0058] Specifically, the steps of the reasoning phase include:
[0059] (1) First, the paired representative small sample data are spliced in dimension, and the features are extracted through convolution operation to obtain the preliminary features of the small sample pairs;
[0060] (2) Through convolution operation, feature extraction is performed on the reference image to obtain preliminary reference features of the visible light reference image;
[0061] (3) The preliminary features of the small sample and the reference preliminary features are fused by summing and input into ControlNet for encoding to form the image condition image c ;
[0062] (4) For text conditions, use CLIP's text encoder to encode the input text instructions to form text conditions text c ;
[0063] (5) The original noise data n0 and text condition text c and time code time c Input diffusion encoder to form noise code z0;
[0064] (6) The noise code z0 and the image condition image c Input diffusion middleware to form intermediate code z;
[0065] (7) The intermediate code z and image condition image c The input diffusion decoder forms the denoised noise code z1;
[0066] (8) Finally, the denoised noise code z1 is input into the decoder to form the infrared image Image corresponding to the visible light image gen .
[0067] The diffusion model for generating small sample infrared data constructed by the present invention takes into account that a large amount of data is required to train a diffusion model for data generation, while there is less infrared data and it is difficult to train the diffusion model. Therefore, a mature visible light diffusion data generation model is driven by a small amount of infrared data to achieve diffusion data generation driven by small sample infrared data. The diffusion model of the present invention will provide a large amount of data for military, night driving and other fields, and promote the development of these fields.
[0068] The present invention introduces a combination of text prompts, ControlNet control signals and diffusion models, and uses a small amount of visible light and infrared image data for training to generate high-resolution infrared images corresponding to visible light images. The system generates with small sample data, which not only reduces the dependence on a large amount of annotated data, but also generates infrared images with wide applicability in various application scenarios while maintaining image quality, thus achieving high-quality infrared image generation and data enhancement.
[0069] In order to further understand the present invention, the following Figure 1 The small sample infrared data generating method of the present invention is described in detail.
[0070] like Figure 1As shown, according to a specific embodiment of the present invention, a method for generating small sample infrared data is provided, which specifically includes the following steps.
[0071] Step 1: Construct a diffusion model for small sample infrared data generation.
[0072] (1) Prepare data:
[0073] Collect 10 pairs of visible light images and infrared image samples as small samples to ensure that the sample size is small but representative. Also collect 100,000 visible light images and their corresponding infrared images. The visible light images are used to generate the image condition image. c , the infrared image is used as the real infrared image when calculating the loss function Image gt , and provide a text description of the generated infrared image.
[0074] (2) Design of diffusion model: The diffusion model architecture includes diffusion encoder, diffusion middleware and diffusion decoder.
[0075] Use CLIP technology to encode text to form text conditions: text c =CLIP(text description).
[0076] The diffusion encoder receives the original noise data n0 and the text condition text c and time code time c As input conditions, these input conditions are mapped to a preliminary noise code z0: z0 = f encoder (n0,text c ,time c ).
[0077] Use the ControlNet model and convolution technology to encode the image to form image conditions: image c =ControlNet(Conv(concat(image_pair)),Conv(image_v)).
[0078] The diffusion middleware receives the noise code z0 and the image condition image c , generate the intermediate code z: z = f middleware (z0,image c ).
[0079] The diffusion decoder receives the intermediate code z and the image condition image c As input, output the denoised noise code z1: z1 = f decoder (z,image c ) and decode the denoised noise code z1: Imagegen =VAE decoder (z1).
[0080] (3) Design loss function:
[0081] (4) Optimize the diffusion model:
[0082] The image condition image formed by any original noise data n0, all representative small samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function L MES , back propagation is performed to continuously optimize the diffusion model.
[0083] Step 2: Generate an infrared image corresponding to the visible light image based on the diffusion model constructed for small sample infrared data generation:
[0084] (1) First, the paired representative small sample data are spliced in dimension, and the features are extracted through convolution operation to obtain the preliminary features of the small sample pairs;
[0085] (2) Through convolution operation, feature extraction is performed on the reference image to obtain preliminary reference features of the visible light reference image;
[0086] (3) The preliminary features of the small sample and the reference preliminary features are fused by summing and input into ControlNet for encoding to form the image condition image c ;
[0087] (4) For text conditions, use CLIP's text encoder to encode the input text instructions to form text conditions text c ;
[0088] (5) The original noise data n0 and text condition text c and time code time c Input diffusion encoder to form noise code z0;
[0089] (6) The noise code z0 and the image condition image c Input diffusion middleware to form intermediate code z;
[0090] (7) The intermediate code z and image condition image c The input diffusion decoder forms the denoised noise code z1;
[0091] (8) Finally, the denoised noise code z1 is input into the decoder to form the infrared image Image corresponding to the visible light image gen .
[0092] In summary, the present invention provides a small sample infrared data generation method, which inputs an infrared-visible light small sample pair, a text instruction, and a visible light reference image, compares the infrared-visible light small sample pair, and outputs an infrared image of the corresponding visible light reference image.
[0093] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for generating small sample infrared data, characterized in that: The small sample infrared data generation method comprises: constructing a diffusion model for generating small sample infrared data, and generating an infrared image corresponding to a visible light image based on the constructed diffusion model for generating small sample infrared data; Among them, building a diffusion model for small sample infrared data generation specifically includes: preparing data, designing a diffusion model, designing a loss function and optimizing the diffusion model; Prepare data: Collect several pairs of representative visible light images and corresponding infrared images as small sample reference images; also collect several pairs of visible light images and corresponding infrared images. The visible light images are used to generate image condition images. c , the infrared image is used as the real infrared image when calculating the loss function Image gt ; Design diffusion model: The diffusion model architecture includes diffusion encoder, diffusion middleware and diffusion decoder; Optimized diffusion model: The image condition image formed by any original noise data n0, all representative samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function, and perform back propagation to continuously optimize the diffusion model.
2. The method for generating small sample infrared data according to claim 1, characterized in that: The diffusion encoder receives the original noise data n0 and the text condition text c and time code time c As input conditions, the input conditions are mapped to a preliminary noise code z0: z0=f encoder (n0,text c ,time c ) Among them, f encoder Represents the nonlinear mapping function of the encoder, which is responsible for combining the original noisy data with the conditional input to generate the noise code z0.
3. The small sample infrared data generation method according to claim 2, characterized in that: Use CLIP technology to encode text to form text conditional text c : text c = CLIP(text description) The text description is a text description of the image to be generated.
4. The small sample infrared data generation method according to claim 2, characterized in that: The diffusion middleware receives the noise code z0 and the image condition image c As input, generate the intermediate code z: z=f middleware (z0,image c ) Among them, f middleware Represents the mapping function of the middleware, which is used to encode z0 and image condition according to the input noise c Generate an intermediate code z.
5. The method for generating small sample infrared data according to claim 4, characterized in that: The image is encoded using convolutional technology to form image conditions: image c =ControlNet(Conv(concat(image_pair)),Conv(image_v)) Among them, image c is the image condition, Conv is the convolution operation, ControlNet is the public model, concat is the dimension concatenation operation used to concatenate infrared images and visible light images, image_pair is a representative small sample data, and image_v is the corresponding visible light image of the infrared image to be generated.
6. The method for generating small sample infrared data according to claim 4, characterized in that: The diffusion decoder receives the intermediate code z and the image condition image c As input, the denoised noise code z1 is output, and the denoised noise code z1 is decoded; the denoised output of the diffusion decoder is expressed as: z1=f decoder (z,image c ) Among them, f decoder is the mapping function of the decoder, used to combine the intermediate code z and the image condition image c Perform denoising.
7. The method for generating small sample infrared data according to claim 6, characterized in that: Use VAE decoder to decode the denoised noise code: Picture gen =VAE decoder (z1) Among them, VAE decoder is the VAE decoder, Image gen The final infrared image is generated.
8. The method for generating small sample infrared data according to claim 1, characterized in that: Loss function L MES for: Where N is the number of pixels in the image or the total number of elements contained in the image; Image Gen,i is the value of the generated infrared image at the i-th pixel; Image gt,i is the value of the real infrared image at the i-th pixel.
9. The method for generating small sample infrared data according to claim 1, characterized in that: The infrared image corresponding to the visible light image is generated based on the diffusion model constructed for small sample infrared data generation, which specifically includes: (1) First, the paired representative small sample data are spliced in dimension, and the features are extracted through convolution operation to obtain the preliminary features of the small sample pairs; (2) Through convolution operation, feature extraction is performed on the reference image to obtain preliminary reference features of the visible light reference image; (3) The preliminary features of the small sample and the reference preliminary features are fused by summing and input into ControlNet for encoding to form the image condition image c ; (4) For text conditions, use CLIP's text encoder to encode the input text instructions to form text conditions text c ; (5) The original noise data n0 and text condition text c and time code time c Input diffusion encoder to form noise code z0; (6) The noise code z0 and the image condition image c Input diffusion middleware to form intermediate code z; (7) The intermediate code z and image condition image c The input diffusion decoder forms the denoised noise code z1; (8) Finally, the denoised noise code z1 is input into the decoder to form the infrared image Image corresponding to the visible light image gen .
Citation Information
Patent Citations
Small sample font generation method based on diffusion model
CN118898549A
Personalized text-to-image generation
US20240355022A1