A small-sample infrared data generation method

By constructing a diffusion model, using a small amount of visible light and infrared image data, a diffusion encoder, diffusion middleware and decoder are designed to optimize the loss function, and the problems of high cost and limited quality of infrared image generation in the existing technology are solved, and efficient and economical infrared image generation is achieved.

CN119941548BActive Publication Date: 2025-08-05AEROSPACE SCI & IND GRP INTELLIGENT TECH RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411783486.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-08-05
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

The existing diffusion models have problems such as high cost in generating infrared images, difficulty in data acquisition and limited conversion quality, which is difficult to meet the needs of computer vision tasks.

Method used

A diffusion model is constructed for small sample infrared data generation. By collecting a small amount of visible light and infrared image data, a diffusion encoder, diffusion middleware and diffusion decoder are designed, and a mean square error loss function optimization model is used to generate high-resolution infrared images.

Benefits of technology

It realizes the generation of a large number of high-quality infrared images by driving a small amount of infrared data, which reduces data acquisition costs, improves the robustness of the model and image generation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941548B_ABST
    Figure CN119941548B_ABST
Patent Text Reader

Abstract

The present invention provides a method for generating small-sample infrared data. The method includes: constructing a diffusion model for generating small-sample infrared data; and generating an infrared image corresponding to a visible light image based on the diffusion model constructed for generating small-sample infrared data. Constructing the diffusion model for generating small-sample infrared data specifically includes: preparing data, designing a diffusion model, designing a loss function, and optimizing the diffusion model. The data preparation includes collecting several pairs of representative visible light images and corresponding infrared images to serve as small samples; collecting several pairs of visible light images and corresponding infrared images to optimize the diffusion model; and designing the diffusion model, wherein the diffusion model architecture includes a diffusion encoder, diffusion middleware, and diffusion decoder. The application of the technical solution of the present invention can solve the technical problem that diffusion models in the prior art cannot meet image processing requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of small sample multimodal image generation algorithms, and in particular relates to a small sample infrared data generation method. Background Art

[0002] In recent research, diffusion models have garnered significant attention for their exceptional performance in data generation. These models are inspired by the diffusion process in physics. In deep learning, they provide a rich data resource for tasks where data acquisition is expensive and privacy-sensitive. These models are capable of generating a variety of data modalities, including images, video, and audio. Diffusion models operate in two phases: a forward phase, which gradually adds noise to real data until it transforms into a Gaussian noise distribution; and a reverse phase, which uses neural network learning to remove noise from the Gaussian noise distribution to generate samples that approximate the original data distribution. Diffusion models have been widely used in fields such as image generation and restoration, becoming a crucial tool for data augmentation.

[0003] Current research on diffusion models mainly focuses on the generation of RGB images in the visible light band, but models and systems for computer vision tasks can also utilize other modal information, such as temperature information, image depth information, and texture perception information. For example, at night or in low light conditions, RGB images provide limited information, while infrared images can detect objects with temperatures above absolute zero by capturing thermal radiation with a wavelength of 0.75-15 microns. Therefore, the richer the modal data obtained, the more robust the model and the more stable its performance in actual visual tasks. However, current infrared image acquisition equipment is usually expensive, which poses an economic burden on the research team; manual proofreading is required to ensure data quality; in addition, the existing visible light to infrared conversion model has limited performance and there is a bottleneck in the conversion quality.

[0004] In view of this, there is an urgent need to develop a diffusion model that can generate infrared data. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art.

[0006] The present invention provides a method for generating small sample infrared data, the method comprising: constructing a diffusion model for generating small sample infrared data, generating an infrared image corresponding to a visible light image based on the constructed diffusion model for generating small sample infrared data;

[0007] Among them, building a diffusion model for small sample infrared data generation specifically includes: preparing data, designing a diffusion model, designing a loss function and optimizing the diffusion model;

[0008] Prepare data: Collect several pairs of representative visible light images and corresponding infrared images to use as small sample reference images; also collect several pairs of visible light images and corresponding infrared images, and the visible light images are used to generate image condition images. c , the real infrared image when the infrared image is used to calculate the loss function Image gt ;

[0009] Design the diffusion model: The diffusion model architecture includes a diffusion encoder, diffusion middleware, and diffusion decoder;

[0010] Optimized diffusion model: The image condition image formed by any original noise data n0, all representative samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function, and perform back propagation to continuously optimize the diffusion model.

[0011] The technical solution of the present invention provides a method for generating small-sample infrared data. This method uses small-sample infrared data to drive a mature visible light diffusion data generation model, trained using a small amount of visible light and infrared image data, to generate a high-resolution infrared image corresponding to the visible light image. Based on the visible light data diffusion model and guided by a small amount of infrared data, the present invention achieves the generation of large amounts of infrared data from small-sample infrared data, providing an efficient and economical solution for generating small-sample infrared data. Compared with the prior art, the technical solution of the present invention can solve the technical problem that the diffusion model in the prior art cannot meet the requirements of image processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are included to provide a further understanding of the embodiments of the present invention, constitute a part of the specification, illustrate the embodiments of the present invention, and together with the description, explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0013] Figure 1 A schematic diagram showing the principle of generating small sample infrared data provided according to a specific embodiment of the present invention is shown. DETAILED DESCRIPTION

[0014] It should be noted that, in the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is in no way intended to limit the present invention and its application or use. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0015] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0016] Unless otherwise specifically stated, the relative arrangement of the parts and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention. Meanwhile, it should be understood that, for ease of description, the sizes of the various parts shown in the accompanying drawings are not drawn according to actual proportional relationships. Technology, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, the technology, methods and equipment should be considered as a part of the specification. In all examples shown and discussed here, any specific value should be interpreted as being merely exemplary, rather than as a limitation. Therefore, other examples of exemplary embodiments may have different values.

[0017] like Figure 1 As shown, according to a specific embodiment of the present invention, a method for generating small sample infrared data is provided, the method comprising: constructing a diffusion model for generating small sample infrared data, and generating an infrared image corresponding to a visible light image based on the constructed diffusion model for generating small sample infrared data;

[0018] Among them, building a diffusion model for small sample infrared data generation specifically includes: preparing data, designing a diffusion model, designing a loss function and optimizing the diffusion model;

[0019] Prepare data: Collect several pairs of representative visible light images and corresponding infrared images to use as small sample reference images; also collect several pairs of visible light images and corresponding infrared images, and the visible light images are used to generate image condition images. c , the real infrared image when the infrared image is used to calculate the loss function Imagegt ;

[0020] Design the diffusion model: The diffusion model architecture includes a diffusion encoder, diffusion middleware, and diffusion decoder;

[0021] Optimized diffusion model: The image condition image formed by any original noise data n0, all representative samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function, and perform back propagation to continuously optimize the diffusion model.

[0022] This configuration provides a method for generating small-sample infrared data. This method uses small-sample infrared data to drive a mature visible light diffusion data generation model, training it with a small amount of visible light and infrared image data to generate a high-resolution infrared image corresponding to the visible light image. Based on the visible light data diffusion model and guided by a small amount of infrared data, this method achieves the generation of large amounts of infrared data from small-sample infrared data, providing an efficient and economical solution for generating small-sample infrared data.

[0023] The method of the present invention consists of two phases: a training phase and an inference phase. The training phase aims to adjust a model that can only generate random images into a model that can generate infrared images, that is, to construct a diffusion model for generating small-sample infrared data. The inference phase uses the diffusion model constructed for generating small-sample infrared data to generate an infrared image corresponding to the visible light image.

[0024] First, in the present invention, it is necessary to construct a diffusion model for generating small sample infrared data.

[0025] Building a diffusion model for small-sample infrared data generation involves preparing the data, designing the diffusion model architecture, designing the loss function, and optimizing the diffusion model. Each step is described in detail below.

[0026] Prepare data: Collect several pairs of representative visible light images and corresponding infrared images to use as small sample reference images; also collect several pairs of visible light images and corresponding infrared images, and the visible light images are used to generate image condition images. c , the real infrared image when the infrared image is used to calculate the loss function Image gt ;

[0027] As a specific embodiment of the present invention, the number of reference images can be appropriately expanded.

[0028] (2) Design of diffusion model: The diffusion model architecture includes diffusion encoder, diffusion middleware and diffusion decoder.

[0029] The diffusion encoder receives the original noise data n0 and the text condition text c and time code time c As input conditions. The original noise data n0 represents the initial state of the model; the text condition text c Used to describe language information related to noise data; time code time c It is used to provide the model with an understanding of the time step, ensuring that temporal information is properly incorporated into the diffusion process.

[0030] As a specific embodiment of the present invention, the present invention can use CLIP technology to encode text to form text conditions. c It can be expressed as:

[0031] text c =CLIP(text description)

[0032] The text description is a text description of the image to be generated.

[0033] The task of the diffusion encoder is to map these input conditions to a preliminary noise code z0. The function of the diffusion encoder can be expressed as:

[0034] z0=f encoder (n0,text c ,time c )

[0035] Among them, f encoder Represents the nonlinear mapping function of the encoder, which is responsible for combining the original noise data with the conditional input to generate the noise code z0. The noise code is used as the input for the subsequent diffusion process.

[0036] The diffusion middleware receives two main inputs: the noise code z0 and the image condition image c , where the image condition image c Represents visual information related to the noise data, such as image features or scene description. Through the diffusion middleware processing, the noise code z0 is converted into an intermediate code z. The key to this step is how to adjust the distribution of noise according to the image conditions.

[0037] As a specific embodiment of the present invention, the present invention can use the ControlNet model and convolution technology to encode the image to form an image condition. The image condition formed by using the convolution technology to encode the image can be expressed as:

[0038] image c =ControlNet(Conv(concat(image_pair)),Conv(image_v))

[0039] Among them, image c is the image condition, Conv is the convolution operation, ControlNet is the public model, concat is the dimension splicing operation, which splices the infrared image and the visible light image together, image_pair is a representative small sample data, and image_v is the corresponding visible light image of the infrared image to be generated.

[0040] The functionality of the diffusion middleware can be expressed as:

[0041] z=f middleware (z0,image c )

[0042] Among them, f middleware Represents the mapping function of the middleware, which is used to encode z0 and image condition according to the input noise c Generate the intermediate code z.

[0043] The diffusion decoder is the last step in the diffusion model. The diffusion decoder receives the intermediate code z and the image condition image c As input, the denoised noise code z1 is output, and the denoised noise code z1 is decoded. The core goal of this process is to remove useless noise information through the denoising operation and restore a clearer and qualified image or signal according to the image conditions.

[0044] The denoised output of the diffusion decoder can be expressed as:

[0045] z1=f decoder (z,image c )

[0046] Among them, f decoder is the mapping function of the decoder, which is used to combine the intermediate code z and the image condition image c De-noising is performed to generate the de-noised noise code z1.

[0047] The present invention uses a VAE decoder to decode the denoised noise code, which can be expressed as:

[0048] Image gen =VAE decoder (z1)

[0049] Among them, VAE decoder is the VAE decoder, Image genThe final infrared image is generated.

[0050] (3) Design loss function:

[0051] In the model of this invention, the design of the loss function is crucial, directly affecting the optimization process of model training and the final generation effect. To effectively evaluate the difference between the generated infrared image and the real infrared image, this invention uses the mean squared error (MSE) as the loss function.

[0052] As a specific embodiment of the present invention, it is assumed that the infrared image generated by the model is Image gen , the real infrared image is Image gt , then the mean square error (MSE) loss function L MES It can be expressed as:

[0053]

[0054] Where N is the number of pixels in the image or the total number of elements contained in the image; Image Gen,i is the value of the generated infrared image at the i-th pixel; Image gt,i is the value of the real infrared image at the i-th pixel.

[0055] (4) Optimize the diffusion model:

[0056] The image condition image formed by any original noise data n0, all representative small samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function L MES , perform back propagation to continuously optimize the diffusion model.

[0057] Furthermore, in the present invention, after the diffusion model for generating small sample infrared data is constructed, an infrared image corresponding to the visible light image is generated based on the constructed diffusion model for generating small sample infrared data.

[0058] Specifically, the steps of the reasoning phase include:

[0059] (1) First, the paired representative small sample data are spliced in dimension, and the features are extracted through convolution operation to obtain the preliminary features of the small sample pairs;

[0060] (2) Through convolution operation, the reference image is subjected to feature extraction to obtain the reference preliminary features of the visible light reference image;

[0061] (3) The preliminary features of the small sample and the reference preliminary features are fused by summing and input into ControlNet for encoding to form the image condition image c ;

[0062] (4) For text conditions, use CLIP's text encoder to encode the input text instructions to form text conditions text c ;

[0063] (5) The original noise data n0, text condition text c and time code time c Input diffusion encoder to form noise code z0;

[0064] (6) The noise code z0 and the image condition image c Input diffusion middleware to form intermediate code z;

[0065] (7) The intermediate code z and image condition image c Input diffusion decoder to form denoised noise code z1;

[0066] (8) Finally, the denoised noise code z1 is input into the decoder to form the infrared image Image corresponding to the visible light image gen .

[0067] The diffusion model constructed in this paper for generating small-sample infrared data takes into account the large amount of data required to train a diffusion model. However, infrared data is relatively scarce, making it difficult to train a diffusion model. Therefore, a small amount of infrared data is used to drive a well-established visible light diffusion data generation model, thereby achieving diffusion data generation driven by small-sample infrared data. This diffusion model will provide a large amount of data for applications such as military and nighttime driving, promoting their development.

[0068] This paper combines text prompts, ControlNet control signals, and a diffusion model, using a small amount of visible light and infrared image data for training to generate high-resolution infrared images corresponding to visible light images. This system generates images from small sample data, reducing its reliance on large amounts of annotated data. It also generates infrared images that are widely applicable across a variety of application scenarios while maintaining image quality, achieving high-quality infrared image generation and data enhancement.

[0069] In order to have a further understanding of the present invention, the following Figure 1 The small sample infrared data generation method of the present invention is described in detail.

[0070] like Figure 1As shown, according to a specific embodiment of the present invention, a method for generating small sample infrared data is provided, which specifically includes the following steps.

[0071] Step 1: Construct a diffusion model for small sample infrared data generation.

[0072] (1) Prepare data:

[0073] Collect 10 pairs of visible light images and infrared image samples as small samples to ensure that the sample size is small but representative. Also collect 100,000 visible light images and their corresponding infrared images. The visible light images are used to generate the image condition image. c , the real infrared image when the infrared image is used to calculate the loss function Image gt , and provide a text description of the generated infrared image.

[0074] (2) Design of diffusion model: The diffusion model architecture includes diffusion encoder, diffusion middleware and diffusion decoder.

[0075] Use CLIP technology to encode text to form text conditions: text c =CLIP(text description).

[0076] The diffusion encoder receives the original noise data n0 and the text condition text c and time code time c As input conditions, these input conditions are mapped to a preliminary noise code z0: z0 = f encoder (n0,text c ,time c ).

[0077] Use the ControlNet model and convolution technology to encode the image to form image conditions: image c =ControlNet(Conv(concat(image_pair)),Conv(image_v)).

[0078] The diffusion middleware receives the noise code z0 and the image condition image c , generate the intermediate code z: z = f middleware (z0,image c ).

[0079] The diffusion decoder receives the intermediate code z and the image condition image c As input, output the denoised noise code z1: z1 = f decoder (z,image c ) and decode the denoised noise code z1: Imagegen =VAE decoder (z1).

[0080] (3) Design loss function:

[0081] (4) Optimize the diffusion model:

[0082] The image condition image formed by any original noise data n0, all representative small samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function L MES , perform back propagation to continuously optimize the diffusion model.

[0083] Step 2: Generate an infrared image corresponding to the visible light image based on the diffusion model constructed for small sample infrared data generation:

[0084] (1) First, the paired representative small sample data are spliced in dimension, and the features are extracted through convolution operation to obtain the preliminary features of the small sample pairs;

[0085] (2) Through convolution operation, the reference image is subjected to feature extraction to obtain the reference preliminary features of the visible light reference image;

[0086] (3) The preliminary features of the small sample and the reference preliminary features are fused by summing and input into ControlNet for encoding to form the image condition image c ;

[0087] (4) For text conditions, use CLIP's text encoder to encode the input text instructions to form text conditions text c ;

[0088] (5) The original noise data n0, text condition text c and time code time c Input diffusion encoder to form noise code z0;

[0089] (6) The noise code z0 and the image condition image c Input diffusion middleware to form intermediate code z;

[0090] (7) The intermediate code z and image condition image c Input diffusion decoder to form denoised noise code z1;

[0091] (8) Finally, the denoised noise code z1 is input into the decoder to form the infrared image Image corresponding to the visible light image gen .

[0092] In summary, the present invention provides a small sample infrared data generation method, which inputs an infrared-visible light small sample pair, a text instruction, and a visible light reference image, compares the infrared-visible light small sample pair, and outputs an infrared image of the corresponding visible light reference image.

[0093] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for generating small sample infrared data, characterized in that: The small sample infrared data generation method includes: constructing a diffusion model for generating small sample infrared data, and generating an infrared image corresponding to a visible light image based on the constructed diffusion model for generating small sample infrared data; Among them, building a diffusion model for small sample infrared data generation specifically includes: preparing data, designing a diffusion model, designing a loss function and optimizing the diffusion model; Prepare data: Collect several pairs of representative visible light images and corresponding infrared images to use as small sample reference images; also collect several pairs of visible light images and corresponding infrared images, and the visible light images are used to generate image condition images. c , the real infrared image when the infrared image is used to calculate the loss function Image gt ; Design the diffusion model: The diffusion model architecture includes a diffusion encoder, diffusion middleware, and diffusion decoder; Optimized diffusion model: The image condition image formed by any original noise data n0, all representative samples and a visible light image c , and the corresponding text condition text c Input the diffusion model to generate the denoised noise code z1, and generate the infrared image Image through the VAE decoder gen , and calculate the loss function, and perform back propagation to continuously optimize the diffusion model; The infrared image corresponding to the visible light image is generated based on the diffusion model constructed for small sample infrared data generation, specifically including: (1) First, the paired representative small sample data are spliced in dimension, and the features are extracted through convolution operation to obtain the preliminary features of the small sample pairs; (2) Through convolution operation, the reference image is subjected to feature extraction to obtain the reference preliminary features of the visible light reference image; (3) The preliminary features of the small sample and the reference preliminary features are fused by summing and input into ControlNet for encoding to form the image condition image c ; (4) For text conditions, use CLIP's text encoder to encode the input text instructions to form text conditions text c ; (5) The original noise data n0 and text condition text c and time code time c Input diffusion encoder to form noise code z0; (6) The noise code z0 and the image condition image c Input diffusion middleware to form intermediate code z; (7) Combine the intermediate code z and the image condition image c Input diffusion decoder to form denoised noise code z1; (8) Finally, the denoised noise code z1 is input into the VAE decoder to form the infrared image Image corresponding to the visible light image gen .

2. The method for generating small sample infrared data according to claim 1, characterized in that: The diffusion encoder receives the original noise data n0 and the text condition text c and time code time c As input conditions, the input conditions are mapped to a preliminary noise code z0: , Among them, f encoder Represents the nonlinear mapping function of the encoder, which is responsible for combining the original noisy data with the conditional input to generate the noisy code z0.

3. The method for generating small sample infrared data according to claim 2, characterized in that: Use CLIP technology to encode text to form text conditional text c : text c =CLIP(text description) The text description is a text description of the image to be generated.

4. The method for generating small sample infrared data according to claim 2, characterized in that: The diffusion middleware receives the noise code z0 and the image condition image c As input, generate the intermediate code z: , Among them, f middleware Represents the mapping function of the middleware, which is used to encode z0 and image condition according to the input noise c Generate the intermediate code z.

5. The method for generating small sample infrared data according to claim 4, characterized in that: The convolution technique is used to encode the image to form the image condition: image c =ControlNet(Conv(concat(image_pair)),Conv(image_v)) Among them, image c is the image condition, Conv is the convolution operation, ControlNet is the public model, concat is the dimension splicing operation used to stitch infrared images and visible light images, image_pair is a representative small sample data, and image_v is the corresponding visible light image of the infrared image to be generated.

6. The method for generating small sample infrared data according to claim 4, characterized in that: The diffusion decoder receives the intermediate code z and the image condition image c As input, the denoised noise code z1 is output and the denoised noise code z1 is decoded; the denoised output of the diffusion decoder is expressed as: z1=f decoder (z,image c ) Among them, f decoder is the mapping function of the decoder, which is used to combine the intermediate code z and the image condition image c Perform denoising.

7. The method for generating small sample infrared data according to claim 6, characterized in that: Use VAE decoder to decode the denoised noise code: Picture gen =VAE decoder (z1) Among them, VAE decoder is the VAE decoder, Image gen The final infrared image is generated.

8. The method for generating small sample infrared data according to claim 1, wherein: The method for generating small sample infrared data according to claim 1 is characterized in that the loss function for: , in, N is the number of pixels in the image or the total number of elements contained in the image; is the value of the generated infrared image at the i-th pixel; is the value of the real infrared image at the i-th pixel.

Citation Information

Patent Citations

  • Small sample font generation method based on diffusion model

    CN118898549A

  • Personalized text-to-image generation

    US20240355022A1