Multi-source prompt infrared target image customization generation method

Through the customized generation method of infrared target image with multi-source prompts, multi-element text and indication information encoding, combined with Transformer network for denoising processing, the problem of infrared image generation relying on visible light images in the prior art is solved, and high-quality and customized infrared image generation is achieved.

CN120451702APending Publication Date: 2025-08-08XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510432556.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing infrared image generation methods rely on visible light image guidance and cannot flexibly generate infrared images required for specific tasks. The generated infrared images often have deviations in target characteristics, and the generation quality is poor.

Method used

The multi-source prompted infrared target image customization generation method is adopted, and the multi-factor text and indication information are obtained for encoding processing, and the infrared image generation backbone network based on Transformer is used for the denoising process, and the layered supervised image is customized to achieve high-quality infrared image generation without the guidance of visible light images.

Benefits of technology

It realizes high-quality infrared image generation without the guidance of visible light images, with accurate target characteristics and controllable generation process, and can accurately control the environmental characteristics and object characteristics of infrared images based on multi-factor text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451702A_ABST
    Figure CN120451702A_ABST
Patent Text Reader

Abstract

The invention particularly relates to a multi-source prompt infrared target image customization generation method, which adopts a multi-element text encoding module to encode a multi-element text, and inputs the multi-element text into a Transform-based infrared image generation backbone network to perform perceptual association mapping, thereby improving the consistency of a generated image and a given multi-element text. When an infrared image is generated, customized generation is carried out on an input specific target image through a layered redrawing method. According to the method, the infrared image of the corresponding target in the background corresponding to the multi-element text can be generated from the target infrared image and the multi-element text, and customized generation of the infrared image is realized. A layered redrawing method is used, different denoising steps are input in the image generation process, and the generation effect of the infrared image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of infrared image technology, and in particular to a customized generation method of infrared target images with multi-source prompts. Background Art

[0002] In the related art, infrared images are obtained by capturing infrared radiation emitted by objects using infrared cameras. Because they can display temperature differences under various environmental conditions and perform well at night and inclement weather, they are widely used in tasks such as border surveillance, military reconnaissance, and environmental monitoring. However, when actually shooting scenes with infrared cameras, factors such as uncooperative targets and high human and material costs often make it difficult and costly to obtain infrared images under specific conditions. Simulation software is currently often used as an alternative method for acquiring infrared images. However, existing infrared image production methods based on physical simulation software require modeling each target in the scene and assigning factors such as material and emissivity. The scene heat and target heat are then calculated ray by ray, resulting in a long infrared image production cycle.

[0003] Currently, mainstream generative model-based infrared image generation methods focus on image style transfer techniques. These methods typically take visible light images as input and generate corresponding infrared images using generative adversarial networks or diffusion models. However, these methods rely heavily on visible light images for guidance, making them inflexible in generating infrared images required for specific tasks. Furthermore, the generated infrared images often exhibit deviations from target characteristics, resulting in poor quality. Directly fine-tuning existing visible light generative models to generate infrared images also presents challenges such as inaccurate infrared target characteristics and an inability to accurately control the target.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0005] The present invention provides a method for customized generation of infrared target images with multi-source prompts, a computer program product, and an electronic device, which realize a method for customized generation of high-quality infrared images without the need for visible light image guidance, thereby overcoming the defects existing in the prior art to a certain extent.

[0006] Other features and advantages of the present invention will become apparent from the following detailed description, or may be learned in part by practice of the present invention.

[0007] According to a first aspect of the present invention, a method for generating a customized infrared target image with multi-source prompts is provided, the method comprising:

[0008] Acquire a multi-element text, an initial infrared image, and indication information; wherein the multi-element text is used to describe environmental features of a target infrared image to be generated, and the indication information is used to describe object features of an object in the initial infrared image in the target infrared image;

[0009] Encode multi-element text to obtain corresponding condition information and text features;

[0010] Performing image segmentation processing on the initial infrared image to obtain a foreground image, transforming the foreground image according to the instruction information; combining the transformed foreground image with the target grayscale image to obtain an intermediate image; performing grayscale assignment processing on the intermediate image based on a preconfigured redraw retention intensity to obtain a layered supervision image;

[0011] The initial infrared image is compressed and encoded to obtain the corresponding latent space code; the latent space code is randomly sampled in a normal distribution to obtain the latent noise;

[0012] The conditional information, text features, and potential noise are input into a Transformer-based infrared image generation backbone network, and a hierarchical supervision image is used for supervision to perform a denoising process with a total number of steps n to form a latent space coded image customized based on the multi-element text and indicative information; and the customized latent space coded image is decoded to obtain a customized target infrared image.

[0013] In some exemplary embodiments, the multi-element text includes: a combination of any multiple of objects, scenes, weather, seasons, time periods, viewing angles, and lighting conditions;

[0014] The indication information includes the size and position of the object in the target infrared image.

[0015] In some exemplary embodiments,

[0016] The performing image segmentation processing on the initial infrared image to obtain a foreground image, transforming the foreground image according to the instruction information; and combining the transformed foreground image with the target grayscale image to obtain an intermediate image, including:

[0017] Performing image segmentation processing on the initial infrared image using an image segmentation model to obtain a foreground area corresponding to the object, and generating a foreground image with a transparent background based on the foreground area;

[0018] After scaling and / or position adjustment are performed on the foreground area according to the instruction information, the foreground area is combined with the grayscale image with a preset grayscale value to obtain the intermediate image.

[0019] In some exemplary embodiments, processing the intermediate image based on a preconfigured redrawing preservation strength to obtain a layered supervision image includes:

[0020] Assign grayscale values to the foreground area based on the preconfigured object redraw retention strength;

[0021] Dilate the foreground area and assign grayscale values to the boundary area according to the pre-configured object boundary redraw retention intensity;

[0022] Based on the background redrawing retention intensity, grayscale assignment is performed on the background area;

[0023] Image fusion processing is performed on the foreground area, boundary area, and background area after grayscale assignment to obtain the layered supervision image; wherein the values of the redrawing retention strength of the preconfigured object, object boundary, and background are different.

[0024] In some exemplary embodiments, encoding the multi-element text to obtain corresponding condition information includes:

[0025] Inputting the multi-element text into a text encoder; wherein the text encoder includes: a CLIP-G / 14 encoder, a CLIP-L / 14 encoder, and a T5XXL encoder;

[0026] Encoding the multi-element text using a CLIP-G / 14 encoder and a CLIP-L / 14 encoder respectively to obtain first encoded data and second encoded data, and combining the first encoded data and the second encoded data to obtain third encoded data for representing global semantic features;

[0027] Performing feature extraction on the third encoded data using a multi-layer perceptron to obtain first feature data;

[0028] The time step is encoded as follows to obtain the corresponding current step number code; the current step number code is extracted using a multi-layer perceptron to obtain the second feature data;

[0029] The first feature data and the second feature data are concatenated to obtain condition information.

[0030] In some exemplary embodiments, encoding the multi-element text to obtain corresponding text features includes:

[0031] Encode the multi-element text using a T5XXL encoder to obtain corresponding fourth encoded data;

[0032] The CLIP-G / 14 encoder and CLIP-L / 14 encoder are used to encode the multi-factor text and obtain the corresponding intermediate layer features.

[0033] The intermediate layer features are concatenated with the fourth encoded data and the zero matrix to obtain the text features.

[0034] In some exemplary embodiments, a denoising process with a total number of steps n is performed using a layered supervision image for supervision to form a latent space encoding image customized based on the multi-element text and the indication information, including:

[0035] When the current time step is t, the hierarchical supervision image is downsampled to obtain the hierarchical supervision image of the latent space;

[0036] Perform pixel-by-pixel comparison of the hierarchical supervision image with the current time step parameters to obtain a hierarchical mask; and, based on the supervision of the hierarchical mask, gradually blend and de-noise the latent space code;

[0037] After the latent noise is denoised in n steps, a customized latent space encoding image is formed; where n and t are positive integers respectively.

[0038] In some exemplary embodiments, the Transformer-based infrared image generation backbone network includes several denoising blocks, modulation layers, and linear layers arranged in sequence; wherein the text features and potential noise are used to input the first denoising block; and the conditional information is used to input each denoising block and modulation layer.

[0039] In some exemplary embodiments, the method further comprises:

[0040] The denoising block performs normalization processing on the text features and potential noise respectively, performs modulation processing and linearization processing based on the conditional information, and obtains the corresponding first independent features and second independent features;

[0041] Use the self-attention layer to extract the global feature relationship and interact the feature information of the first independent feature and the second independent feature to obtain the corresponding third feature and fourth feature;

[0042] The third feature and the fourth feature are linearized, normalized, modulated with conditional information, and feature extracted using a multi-layer perceptron to obtain the corresponding first deep semantic information and second deep semantic information.

[0043] According to a second aspect of the present invention, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the above-mentioned customized generation method of infrared target images with multi-source prompts.

[0044] According to a third aspect of the present invention, there is provided an electronic device, comprising:

[0045] processor; and

[0046] A memory for storing executable instructions of the processor; wherein the memory is used to store executable instructions of the processor; the processor is configured to implement the above-mentioned customized generation method of infrared target images with multi-source prompts when executing instructions by executing the executable instructions.

[0047] According to a fourth aspect of the present invention, there is provided a storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned method for customized generation of infrared target images with multi-source prompts.

[0048] The embodiment of the present invention provides a customized generation method for infrared target images with multi-source prompts. By encoding multi-element text and instruction information, it can control multiple image factors, and then input them into the Transformer-based infrared image generation backbone network for perceptual association mapping, thereby improving the consistency of the generated image with the given multi-element text. When generating the infrared image, a layered redrawing method is used to customize the input specific target image. Different shapes can be given to different target areas. For target areas where infrared characteristics need to be retained, a higher redrawing retention strength is given; for background areas where multi-element text generation is required, a lower redrawing retention strength is given; and for boundary areas where thermal interaction occurs between the target and the background, a general redrawing retention strength is given to achieve coupling between the target and the scene, making the infrared characteristics more accurate. The size of the boundary area can also be freely selected according to the target radiation intensity. The layered redrawing method realizes customized generation of infrared images with a high degree of freedom, ensuring that the target characteristics of the generated infrared image are accurate and the generation process is controllable.

[0049] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present invention, and together with the description, serve to explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and it is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0051] Figure 1 A schematic diagram schematically illustrates a method for customized generation of infrared target images with multi-source prompts according to an exemplary embodiment of the present invention;

[0052] Figure 2 A schematic diagram schematically illustrates a method flow for multi-element text encoding according to an exemplary embodiment of the present invention;

[0053] Figure 3 A schematic diagram schematically illustrates an infrared image generation backbone network structure according to an exemplary embodiment of the present invention;

[0054] Figure 4 A schematic diagram schematically illustrates an input initial infrared image according to an exemplary embodiment of the present invention;

[0055] Figure 5 A schematic diagram schematically illustrates a hierarchical supervision image in an exemplary embodiment of the present invention;

[0056] Figure 6 A schematic diagram schematically illustrates a customized infrared image generated by a model in an exemplary embodiment of the present invention;

[0057] Figure 7 A schematic diagram schematically illustrates an electronic device according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0058] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0059] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0060] In the related art, mainstream generative model-based infrared image generation methods primarily focus on image style transfer techniques. These methods typically take visible light images as input and generate corresponding infrared images using generative adversarial networks or diffusion models. However, these methods rely heavily on visible light images for guidance, making them inflexible in generating infrared images required for specific tasks. Furthermore, the generated infrared images often exhibit deviations in target characteristics, resulting in poor quality. Directly fine-tuning existing visible light generative models to generate infrared images also presents challenges such as inaccurate infrared target characteristics and an inability to accurately control the target.

[0061] In view of the shortcomings and deficiencies of the existing technology, this example embodiment provides a customized generation method of infrared target images with multi-source prompts, which can achieve a generation method of high-quality infrared images that can be customized without the need for visible light image guidance. Figure 1 As shown, the method may include the following steps:

[0062] Step S11, obtaining a multi-element text, an initial infrared image, and instruction information; wherein the multi-element text is used to describe the environmental features of the target infrared image to be generated, and the instruction information is used to describe the object features of the object in the initial infrared image in the target infrared image;

[0063] Step S12: Encode the multi-element text to obtain corresponding condition information and text features;

[0064] Step S13, performing image segmentation processing on the initial infrared image to obtain a foreground image, transforming the foreground image according to the instruction information; combining the transformed foreground image with the target grayscale image to obtain an intermediate image; performing grayscale assignment processing on the intermediate image based on a preconfigured redraw retention intensity to obtain a layered supervision image;

[0065] Step S14, compressing and encoding the initial infrared image to obtain the corresponding latent space code; randomly sampling the latent space code in a normal distribution to obtain latent noise;

[0066] Step S15: input the conditional information, text features, and potential noise into a Transformer-based infrared image generation backbone network, use a hierarchical supervision image for supervision, and perform a denoising process with a total number of steps n to form a latent space coded image customized based on the multi-element text and indication information; and decode the customized latent space coded image to obtain a customized target infrared image.

[0067] Hereinafter, each step of the customized generation method of infrared target images with multi-source prompts in this exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.

[0068] In step S11, a multi-element text, an initial infrared image, and indication information are obtained; wherein the multi-element text is used to describe the environmental features of the target infrared image to be generated, and the indication information is used to describe the object features of the object in the initial infrared image in the target infrared image.

[0069] Exemplarily, the multi-element text includes but is not limited to: a combination of any multiple of objects, scenes, weather, seasons, time periods, viewing angles, and lighting conditions;

[0070] The indication information includes the size and position of the object in the target infrared image.

[0071] Specifically, the user can create an image processing task for generating infrared images on the terminal device. An initial infrared image to be processed can be provided as a data basis. Secondly, the user can also configure multi-element text and / or instruction information. Through multi-element text and instruction information, the user can customize the content of the target infrared image to be generated, so that the target infrared image finally output includes the characteristics of the scene, weather, season, time period, viewing angle, and lighting conditions specified in the above-mentioned multi-element text, as well as the position and size of the object in the target infrared image in the instruction information. Among them, the scene can be an urban scene, an outdoor scene, an indoor scene, etc.; the weather can be sunny, cloudy, rainy, etc.; the time period can be divided into a time period every two hours; and the lighting condition can be the light intensity. Combining weather, season, time period, lighting conditions, and scene, the environmental conditions can be accurately defined from multiple dimensions.

[0072] In step S12, the multi-element text is encoded to obtain corresponding condition information and text features.

[0073] Exemplarily, encoding the multi-element text to obtain corresponding condition information includes:

[0074] Step S211, inputting the multi-element text into a text encoder; wherein the text encoder includes: a CLIP-G / 14 encoder, a CLIP-L / 14 encoder, and a T5XXL encoder;

[0075] Step S212: Encode the multi-element text using a CLIP-G / 14 encoder and a CLIP-L / 14 encoder respectively to obtain first encoded data and second encoded data, and combine the first encoded data and the second encoded data to obtain third encoded data for representing global semantic features;

[0076] Step S213, using a multi-layer perceptron to perform feature extraction on the third encoded data to obtain first feature data;

[0077] Step S214: Encode the time step as follows to obtain the corresponding current step number code; perform feature extraction on the current step number code using a multi-layer perceptron to obtain second feature data;

[0078] Step S215: concatenate the first feature data and the second feature data to obtain condition information.

[0079] Exemplarily, encoding the multi-element text to obtain corresponding text features includes:

[0080] Step S221, using the T5XXL encoder to encode the multi-element text to obtain corresponding fourth encoded data;

[0081] Step S222: Use CLIP-G / 14 encoder and CLIP-L / 14 encoder to encode the multi-element text and obtain the corresponding intermediate layer features.

[0082] Step S223: Concatenate the intermediate layer features with the fourth encoded data and the zero matrix to obtain the text features.

[0083] Specifically, a customized infrared image generation model can be provided, including infrared image encoding and decoding modules, multi-element text encoding modules and infrared image generation backbone networks. For the currently input multi-element text, the multi-element text encoding module can be used to encode it. Figure 2 As shown, the multi-factor text encoding module is implemented using three encoders pre-trained on large-scale text: CLIP-G / 14, CLIP-L / 14, and T5XXL.

[0084] Specifically, first input the original multi-factor text to obtain the text encoding of two CLIP text encoders as the global semantic features of the text, and then add the current step encoding t encoded by the following formula enc The conditional information is obtained by passing through the multi-layer perceptron and splicing. The formula may include:

[0085]

[0086] Where t is the current time step and dim is the dimension of the encoded output.

[0087] In order to obtain fine-grained features, the penultimate layer features of the two CLIP text encoders are extracted respectively, and then zero-padded and concatenated with the last layer features of the T5-XXL text encoder to obtain text features.

[0088] In step S13, the initial infrared image is subjected to image segmentation processing to obtain a foreground image, and the foreground image is transformed according to the instruction information; and the transformed foreground image is combined with the target grayscale image to obtain an intermediate image; based on the preconfigured redrawing retention intensity, the intermediate image is subjected to grayscale assignment processing to obtain a layered supervision image.

[0089] Exemplarily, performing image segmentation processing on the initial infrared image to obtain a foreground image, transforming the foreground image according to the instruction information; and combining the transformed foreground image with the target grayscale image to obtain an intermediate image includes:

[0090] Step S31, performing image segmentation processing on the initial infrared image using an image segmentation model, obtaining a foreground area corresponding to the object, and generating a foreground image with a transparent background based on the foreground area;

[0091] Step S32: After scaling and / or position adjustment of the foreground area according to the instruction information, the foreground area is combined with the grayscale image of the preset grayscale value to obtain the intermediate image.

[0092] Specifically, for a specific initial infrared image x tgt , use the segmentation model to segment the target in the initial infrared image to form the first infrared image x with transparent background tp , by specifying the target size and generating position x tp Scale and paste into the grayscale image with a value of 127 to get the intermediate image x i .

[0093] Exemplarily, the processing of the intermediate image based on the preconfigured redrawing retention strength to obtain the layered supervision image includes:

[0094] Step S41, assigning grayscale values to the foreground area based on the pre-configured object redraw retention strength;

[0095] Step S42, dilating the foreground area and assigning grayscale values to the boundary area according to a pre-configured object boundary redrawing retention intensity;

[0096] Step S43, assigning grayscale to the background area based on the background redraw retention intensity;

[0097] Step S44, performing image fusion processing on the foreground area, boundary area, and background area after grayscale assignment to obtain the layered supervision image; wherein the values of the redrawing retention strength of the preconfigured objects, object boundaries, and background are different.

[0098] Specifically, you can pre-configure the object, object boundary, and background to have three different redrawing retention intensities, k tgt 、k srd 、k gnd The corresponding value range is 0-255. tp The target part in (object, e.g. Figure 4 The car in the image) is assigned a grayscale value of k tgt , the target is subjected to a dilation algorithm with a structure element of k×k, where the dilated part (i.e. the object boundary part) is assigned a grayscale value of k srd , assign gray value k to the background gnd , and finally get the layered supervision image x spv .

[0099] In step S14, the initial infrared image is compressed and encoded to obtain the corresponding latent space code; the latent space code is randomly sampled in a normal distribution to obtain latent noise.

[0100] For example, the initial infrared image can be compressed and encoded using a pre-trained infrared image encoding / decoding module. Specifically, the initial infrared image x with a shape of 1×h×w can be i The latent space encoded image z is compressed to 16×n / 8×w / 8 i .

[0101] Afterwards, the potential noise z of size 16×n / 8×w / 8 can be randomly sampled from a normal distribution n .

[0102] At the same time, we can also perform layer-wise supervision on the shape of 1×h×w spv Downsample to form a latent space hierarchical supervision image z with a shape of 1×h / 8×w / 8 spv .

[0103] Exemplarily, the infrared image encoding / decoding module can be pre-trained. For example, the infrared image encoding / decoding module is implemented using a variational auto-encoder (VAE), which includes an encoder E and a decoder D. The encoder in the VAE compresses, identifies, and separates the different features and factors that affect the process of generating infrared images from noise data, forming a latent space encoding of the infrared image, which serves as the input of the subsequent image generation backbone network. For an infrared image with a shape of 1×h×w, the encoder E compresses it into a latent space encoding with a shape of 16×h / 8×w / 8. During image generation, the decoder is used to reconstruct the low-dimensional spatial features into an infrared image.

[0104] In step S15, the conditional information, text features, and potential noise are input into a Transformer-based infrared image generation backbone network, and a hierarchical supervision image is used for supervision to perform a denoising process with a total number of steps n to form a latent space coded image customized based on the multi-element text and indication information; and the customized latent space coded image is decoded to obtain a customized target infrared image.

[0105] Exemplarily, a layered supervision image is used for supervision, and a denoising process with a total number of steps n is performed to form a latent space coding image customized based on the multi-element text and the indication information, including:

[0106] When the current time step is t, the hierarchical supervision image is downsampled to obtain the hierarchical supervision image of the latent space;

[0107] Perform pixel-by-pixel comparison of the hierarchical supervision image with the current time step parameters to obtain a hierarchical mask; and, based on the supervision of the hierarchical mask, gradually blend and de-noise the latent space code;

[0108] After the latent noise is denoised in n steps, a customized latent space encoding image is formed; where n and t are positive integers respectively.

[0109] Specifically, the number of generation steps can be pre-configured as n. After obtaining the condition information, text features, and potential noise, the condition information, text features, and potential noise z n At the same time, it is input into the infrared image generation backbone network for denoising with a total number of steps of n.

[0110] Let the current time step be t, for the layered supervision image x spv and (Current time step parameter) performs pixel-by-pixel traversal comparison to obtain a hierarchical mask containing only 0 or 1

[0111] By layering mask m t Supervision, the latent space encoding image z i During the gradual blending and denoising process:

[0112] z t =z t+1 ⊙m t +noise(z i , t)⊙(1-m t )

[0113] Here, ⊙ represents pixel-level multiplication, and noise(·, t) represents the noise addition operation based on the noise level at the current time step t.

[0114] Potential noise z n After n steps of denoising, a customized latent space coded image z0 is formed. This image can then be decoded into a customized infrared image x0, i.e., the target infrared image, by the infrared image decoding module.

[0115] Exemplarily, the Transformer-based infrared image generation backbone network includes several denoising blocks, modulation layers, and linear layers arranged in sequence; wherein the text features and potential noise are used to input the first denoising block; and the conditional information is used to input each denoising block and modulation layer.

[0116] Exemplarily, the method further includes:

[0117] The denoising block performs normalization processing on the text features and potential noise respectively, performs modulation processing and linearization processing based on the conditional information, and obtains the corresponding first independent features and second independent features;

[0118] Use the self-attention layer to extract the global feature relationship and interact the feature information of the first independent feature and the second independent feature to obtain the corresponding third feature and fourth feature;

[0119] The third feature and the fourth feature are linearized, normalized, modulated with conditional information, and feature extracted using a multi-layer perceptron to obtain the corresponding first deep semantic information and second deep semantic information.

[0120] Specifically, refer to Figure 3 As shown in the figure, the conditional information, text features and noise in the latent space are input into the infrared image generation backbone network together to obtain the customized target infrared image output by the model. Since the encoding of text and infrared images is conceptually completely different, two completely different sets of weights are used to process different modalities. The text features output by the multi-factor text encoding module are used as the text modality, and the latent space noise output by the infrared image encoding module is used as the infrared image modality and input into the infrared image generation backbone network at the same time. The structure of the infrared image generation backbone network is shown in the figure. Figure 3 As shown in the figure, it contains several denoising blocks with the same structure. Each denoising block encodes the text modality and infrared image modality into independent feature representations through a linear layer after normalization and modulation modules. The two feature representations are then simultaneously passed through a self-attention module to capture global feature relationships. This allows the feature representations of text and infrared images to operate in their respective spaces while interacting with each other, achieving bidirectional information flow. Subsequently, the features are further extracted through linear layers, modulation, and multi-layer perceptrons, and the original features are retained through residual connections to obtain the output. The conditional information output by the multi-factor text encoding module is also injected into the denoising block through a modulation module and direct splicing after passing through a linear layer. Finally, after passing through several denoising blocks, the output of the infrared image modality is again passed through a modulation module and linear layer to obtain the final customized latent space denoised image.

[0121] For example, a customized infrared image generation model can be pre-built, including an infrared image encoding / decoding module, a multi-element text encoding module, and an infrared image generation backbone network. The customized infrared image generation model can then be trained. Specifically, the following steps can be performed during model training:

[0122] 1) Use the annotation model to annotate sample infrared images with multi-element text, including the target, scene, weather, season, and time of day, to generate text-infrared image pairs. For a given infrared image, the annotation model will output the corresponding annotations for the target, scene, weather, season, and time of day, generating text-infrared image pairs.

[0123] 2) Build a customized infrared image generation model, including an infrared image encoding / decoding module, a multi-element text encoding module, and an infrared image generation backbone network;

[0124] A customized generative model for infrared images is constructed. First, the image encoding module is used to convert infrared images into a low-dimensional latent space, compressing image features. Next, based on the diffusion model, a Transformer-based infrared image generation backbone network is used in the latent space for image denoising. Combined with the text information encoded by the multi-element text encoding module, this network controls the changes in image attributes, thereby achieving conditional generation of associations between text and infrared images. Finally, the image decoding module reconstructs the generated data in the latent space into an infrared image.

[0125] The infrared image encoding / decoding module is implemented using a variational auto-encoder (VAE), consisting of an encoder E and a decoder D. The VAE's encoder compresses, identifies, and separates the various features and factors that influence the generation of infrared images from noise data, forming a latent space encoding of the infrared image. This encoding serves as input for subsequent training of the image generation backbone network. For infrared images of shape 1 × h × w, the encoder E compresses them into a latent space encoding of shape 16 × h / 8 × w / 8. During image generation, the decoder reconstructs the low-dimensional spatial features into an infrared image.

[0126] Specifically, the training loss function of the encoder E and decoder D in VAE is:

[0127] L=L VAE (E, D) + L EMB (E, D, X)

[0128] Among them, L VAE is the reconstruction loss, L EMB is the embedding space loss.

[0129] For the reconstruction loss, we can use the learned perceptual image patch similarity to measure the difference between the compressed and reconstructed infrared image and the original infrared image. Specifically, we can use a pre-trained deep network to extract infrared image features, and then calculate the distance between these features to evaluate the perceptual similarity between images.

[0130] For embedding space loss, an adversarial training process can be introduced. Through the additional training of a discriminator X, the goal is to distinguish between the original infrared image and the compressed and reconstructed infrared image as much as possible. The discriminator, the encoder and decoder compete with each other and are continuously optimized. The final equilibrium state is that the distribution of the compressed and reconstructed infrared image of the encoder and decoder is consistent with the distribution of the original infrared image, and the discriminator cannot distinguish the compressed and reconstructed infrared image from the original infrared image.

[0131] The multi-factor text encoding module is implemented using three encoders pre-trained on large-scale text: CLIP-G / 14, CLIP-L / 14, and T5XXL. The specific structure of the module is as follows: Figure 2 During training, we first input the original multi-factor text to obtain the text encoding of the two CLIP text encoders as the global semantic features of the text, and then add the current step encoding t encoded by the following formula enc The conditional information is obtained by passing through the multi-layer perceptron and splicing.

[0132] The conditional information, text features and noise in the latent space are input into the infrared image generation backbone network for training and generation. Since the encoding of text and infrared images is conceptually completely different, two completely different sets of weights are used to process different modalities. The text features output by the multi-factor text encoding module are used as the text modality, and the latent space noise output by the infrared image encoding module is used as the infrared image modality and input into the infrared image generation backbone network at the same time. The structure of the infrared image generation backbone network is as follows: Figure 3 As shown in the figure, it contains several denoising blocks with the same structure. Each denoising block can encode the text modality and infrared image modality into independent feature representations through a linear layer after normalization and modulation modules. The two feature representations are then simultaneously passed through a self-attention module to capture global feature relationships. This allows the feature representations of text and infrared images to operate in their respective spaces while interacting with each other, achieving bidirectional information flow. Subsequently, the features are further extracted through linear layers, modulation, and multi-layer perceptrons, and the original features are retained through residual connections to obtain the output. The conditional information output by the multi-factor text encoding module is also injected into the denoising block through a modulation module and direct splicing after passing through a linear layer. Finally, after passing through several denoising blocks, the output of the infrared image modality is again passed through a modulation module and linear layers to obtain the final latent space denoised image.

[0133] 3) inputting the text-infrared image pair into the infrared image customized generation model for training to obtain a trained infrared image customized generation model;

[0134] The model is trained using the concept of a diffusion model, which consists of two phases: a forward process and a backward process. The customized infrared image generation model constructed in step 2 is trained using the text-infrared image pairs obtained in step 1. For infrared images x0 to q(x), the diffusion model defines the forward process as a Markov chain whose stationary distribution is a Gaussian distribution, where T is the end time and t is the current time step:

[0135]

[0136] In this process, random Gaussian noise is slowly diffused into the infrared image. The transition probability of each step is constructed to add Gaussian noise:

[0137]

[0138] Among them, β t ∈(0,1) is the variance of the noise, β t With the value of the current time step t, the variance table is obtained. As the time step t increases, the weight of the data gradually decreases, while the weight of the noise gradually increases. When t is large enough, the conditional distribution q(x t |x0) will converge to a standard Gaussian distribution that is independent of x0, that is:

[0139] lim t→∞ q(x t )=N(0,I)

[0140] As long as a sufficiently large end time T is taken, a distribution from the infrared image data distribution x0 to the standard Gaussian distribution x T The mapping:

[0141]

[0142] Among them, α t =1-β t , The reverse process of the diffusion model is to learn from the standard Gaussian distribution p(x T )=N(x T ; 0, 1) start of the denoising process:

[0143]

[0144] According to Bayes' theorem, the transition probability of the reverse process can be obtained:

[0145]

[0146] Among them, the mean and variance

[0147]

[0148] in,

[0149] Since the noise ò is unknown, noise ò is added during training by minimizing t Customized generative model with infrared image prediction noise ∈ θ (x t ; t) to optimize the loss function:

[0150]

[0151] Among them, θ is the model parameter, For expectation.

[0152] 4) Input multi-element text, specific target image, target size and its generation position, and use the infrared image customized generation model through the layered redrawing method to perform customized generation of infrared images.

[0153] For the input specific target infrared image x tgt , use the segmentation model to segment the target in the image and form an infrared image of the target with transparent background x tp , by specifying the target size and generating position x tp Scale and paste into the grayscale image with a value of 127 to get the input image x i .

[0154] Set the target, target boundary and background three different redrawing retention strengths as k tgt 、k srd 、k gnd , the range is 0-255, and x tp The target part in the grayscale value is k tgt , the target is subjected to a dilation algorithm with a structure element of k×k, where the dilated part (i.e. the target boundary part) is assigned a grayscale value of k srd , assign gray value k to the background gnd , and finally get the layered supervision image x spv .

[0155] Use the infrared image encoding module in the infrared image customized generation model to convert the input image x with a shape of 1×h×w i The latent space coded image z is compressed to 16×h / 8×w / 8 i . At the same time, for the layered supervision image x of shape 1×h×w spv Downsample to form a latent space hierarchical supervision image z with a shape of 1×h / 8×w / 8 spv .

[0156] Input the multi-factor text into the multi-factor text encoding module to obtain conditional information and text features; set the number of generation steps to n, and randomly sample the potential noise z with a size of 16×h / 8×w / 8 in the normal distribution. n . Combine conditional information, text features and potential noise z n At the same time, it is input into the infrared image generation backbone network for denoising with a total number of steps of n.

[0157] Let the current time step be t, for the layered supervision image x spv and Perform pixel-by-pixel comparison to obtain a hierarchical mask containing only 0 or 1 By layering mask m t Supervision, the latent space encoding image z i Gradually blended into the denoising process.

[0158] For example, during model training, the annotation model can be deployed on a GPU server. For a given infrared image, the annotation model will output the corresponding annotations: "Object: vehicle, tree; Scene: road, sidewalk; Weather: cloudy; Season: summer; Time: afternoon" (for example). The output annotations are saved in a .txt file with the same name as the image, resulting in a dataset of text-infrared image pairs.

[0159] Specify the generated image size as 512, load the target SUV image, specify the target size as 200×200, and the generated position as (187, 173), and obtain the input image of the infrared image customized generation model as follows Figure 4 As shown. The redrawing retention strength of the target, target boundary and background is set to 255, 41 and 0 respectively, and the structural element of the expansion algorithm is 132×132, and the hierarchical supervision image is obtained as shown in Figure 5 As shown. Load the trained infrared image customized generation model, enter the text prompt as: "Target: vehicle, tree; Scene: city, road; Weather: sunny; Season: summer; Time: morning", set the number of generation steps to 30, and the corresponding customized infrared image generated by the infrared image customized generation model is as follows Figure 6 shown.

[0160] The method provided by the present invention adopts a multi-element text encoding module to encode the multi-element text, and inputs it into the Transformer-based infrared image generation backbone network for perceptual association mapping, thereby improving the consistency of the generated image with the given multi-element text; when generating the infrared image, the input specific target image is customized and generated through a layered redrawing method.

[0161] Specifically, the multi-element text encoding module adopted by the present invention can realize the encoding of multi-element texts including targets, scenes, weather, seasons, time periods, etc., so as to carry out subsequent controllable multi-element generation. The layered redrawing method adopted by the present invention can give different shapes to different target areas. For target areas that need to retain infrared characteristics, a higher redrawing retention intensity is given; for background areas that need multi-element text generation, a lower redrawing retention intensity is given; and for boundary areas where thermal interaction occurs between the target and the background, a general redrawing retention intensity is given to achieve coupling between the target and the scene, making the infrared characteristics more accurate. The size of the boundary area can also be freely selected according to the level of the target radiation intensity. The layered redrawing method realizes customized generation of infrared images with a high degree of freedom, ensuring that the target characteristics of the generated infrared image are accurate and the generation process is controllable.

[0162] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0163] Furthermore, an electronic device is provided in an embodiment of this example, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to implement the above-mentioned three-dimensional infrared radiation field modeling method based on physical mechanism constraints by executing the executable instructions.

[0164] Figure 7 A schematic diagram of an electronic device suitable for implementing an embodiment of the present invention is shown.

[0165] It should be noted that Figure 7 The electronic device 1000 shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0166] like Figure 7As shown, electronic device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in read-only memory (ROM) 1002 or the program loaded from storage portion 1008 into random access memory (RAM) 1003. Various programs and data required for system operation are also stored in RAM 1003. CPU 1001, ROM 1002 and RAM 1003 are connected to each other via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0167] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read from the removable media can be installed in the storage section 1008 as needed.

[0168] In particular, according to an embodiment of the present invention, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a storage medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009 and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the various functions defined in the system of the present application are performed.

[0169] It should be noted that the storage medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any storage medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0171] The units involved in the embodiments of the present invention may be implemented in software or hardware, and the units described may also be provided in a processor. In some cases, the names of these units do not limit the units themselves.

[0172] It should be noted that, as another aspect, the present application also provides a storage medium, which can be included in an electronic device; or it can exist independently without being installed in the electronic device. The above storage medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments. For example, the electronic device can implement the following Figure 1 、 Figure 2 The individual steps of the method are shown.

[0173] In one embodiment, the present application provides a computer program product, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0174] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0175] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to an embodiment of the present invention, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0176] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.

[0177] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof, which is limited only by the appended claims.

Claims

1. A customized generation method of infrared target images with multi-source prompts, characterized in that: The method comprises: Acquire a multi-element text, an initial infrared image, and indication information; wherein the multi-element text is used to describe environmental features of a target infrared image to be generated, and the indication information is used to describe object features of an object in the initial infrared image in the target infrared image; Encode multi-element text to obtain corresponding condition information and text features; Performing image segmentation processing on the initial infrared image to obtain a foreground image, transforming the foreground image according to the instruction information; combining the transformed foreground image with the target grayscale image to obtain an intermediate image; performing grayscale assignment processing on the intermediate image based on a preconfigured redraw retention intensity to obtain a layered supervision image; The initial infrared image is compressed and encoded to obtain the corresponding latent space code; the latent space code is randomly sampled in a normal distribution to obtain the latent noise; The conditional information, text features, and potential noise are input into a Transformer-based infrared image generation backbone network, and a hierarchical supervision image is used for supervision to perform a denoising process with a total number of steps n to form a latent space coded image customized based on the multi-element text and indicative information; and the customized latent space coded image is decoded to obtain a customized target infrared image.

2. The method according to claim 1, characterized in that The multi-element text includes any combination of objects, scenes, weather, seasons, time periods, viewing angles, and lighting conditions; The indication information includes the size and position of the object in the target infrared image.

3. The method according to claim 2, characterized in that performing image segmentation processing on the initial infrared image to obtain a foreground image, and transforming the foreground image according to the instruction information; The transformed foreground image is combined with the target grayscale image to obtain an intermediate image, including: Performing image segmentation processing on the initial infrared image using an image segmentation model to obtain a foreground area corresponding to the object, and generating a foreground image with a transparent background based on the foreground area; After scaling and / or position adjustment are performed on the foreground area according to the instruction information, the foreground area is combined with the grayscale image with a preset grayscale value to obtain the intermediate image.

4. The method according to claim 3, characterized in that The processing of the intermediate image based on the preconfigured redrawing retention strength to obtain the layered supervision image includes: Assign grayscale values to the foreground area based on the preconfigured object redraw retention strength; Dilate the foreground area and assign grayscale values to the boundary area according to the pre-configured object boundary redraw retention intensity; Based on the background redrawing retention intensity, grayscale assignment is performed on the background area; Image fusion processing is performed on the foreground area, boundary area, and background area after grayscale assignment to obtain the layered supervision image; wherein the values of the redrawing retention strength of the preconfigured object, object boundary, and background are different.

5. The method according to claim 1, wherein The encoding process of the multi-element text to obtain the corresponding condition information includes: Inputting the multi-element text into a text encoder; wherein the text encoder includes: a CLIP-G / 14 encoder, a CLIP-L / 14 encoder, and a T5 XXL encoder; Encoding the multi-element text using a CLIP-G / 14 encoder and a CLIP-L / 14 encoder respectively to obtain first encoded data and second encoded data, and combining the first encoded data and the second encoded data to obtain third encoded data for representing global semantic features; Performing feature extraction on the third encoded data using a multi-layer perceptron to obtain first feature data; The time step is encoded as follows to obtain the corresponding current step number code; the current step number code is extracted using a multi-layer perceptron to obtain the second feature data; The first feature data and the second feature data are concatenated to obtain condition information.

6. The method according to claim 5, characterized in that The encoding process of the multi-element text to obtain corresponding text features includes: Encode the multi-element text using a T5 XXL encoder to obtain corresponding fourth encoded data; The CLIP-G / 14 encoder and CLIP-L / 14 encoder are used to encode the multi-factor text and obtain the corresponding intermediate layer features. The intermediate layer features are concatenated with the fourth encoded data and the zero matrix to obtain the text features.

7. The method according to claim 1, characterized in that Using the layered supervision image for supervision, a denoising process with a total number of steps n is performed to form a latent space encoding image customized based on the multi-element text and the indication information, including: When the current time step is t, the hierarchical supervision image is downsampled to obtain the hierarchical supervision image of the latent space; Perform pixel-by-pixel comparison of the hierarchical supervision image with the current time step parameters to obtain a hierarchical mask; and, based on the supervision of the hierarchical mask, gradually blend and de-noise the latent space code; After the latent noise is denoised in n steps, a customized latent space encoding image is formed; where n and t are positive integers respectively.

8. The method according to claim 1 or 7, characterized in that The Transformer-based infrared image generation backbone network includes several denoising blocks, modulation layers, and linear layers arranged in sequence; wherein the text features and potential noise are used to input the first denoising block; and the conditional information is used to input each denoising block and modulation layer.

9. The method according to claim 8, characterized in that The method further comprises: The denoising block performs normalization processing on the text features and potential noise respectively, performs modulation processing and linearization processing based on the conditional information, and obtains the corresponding first independent features and second independent features; Use the self-attention layer to extract the global feature relationship and interact the feature information of the first independent feature and the second independent feature to obtain the corresponding third feature and fourth feature; The third feature and the fourth feature are linearized, normalized, modulated with conditional information, and feature extracted using a multi-layer perceptron to obtain the corresponding first deep semantic information and second deep semantic information.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating customized infrared target images with multi-source prompts according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Cross-modal text-to-infrared image generation method based on thermal mask constraint

    CN120807718A