Target and background coupling editing diffusion generation method of infrared image

Through a method of coupled editing and diffusion generation for target and background for infrared images, the controllability and quality problems of infrared image editing generation in the prior art are solved, and efficient and low-cost diversified infrared image data generation is achieved.

CN120107415APending Publication Date: 2025-06-06XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510171823.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When used in infrared images, the existing visible light image coupled editing and diffusion generation method has problems such as inaccurate target characteristics and poor controllability, resulting in low quality of infrared images and the inability to edit and generate diversified data at low cost and high efficiency.

Method used

A method for coupling editing and diffusion generation of target and background of infrared images is proposed. By acquiring the original infrared image, background image, target coupling position and parameters, image segmentation, feature extraction and model generation are performed to achieve efficient and controllable coupled editing of infrared objects and backgrounds.

Benefits of technology

It realizes efficient and high-quality generation of infrared images under zero sample conditions, improves the production efficiency of infrared images, and has high fidelity and diversity in generated images, reducing costs and time expenses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107415A_ABST
    Figure CN120107415A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of computer vision, in particular to an infrared image target and background coupling editing diffusion generation method, which comprises the following steps: acquiring an original infrared image containing an infrared target, a background image not containing the infrared target, an infrared target coupling position and an infrared target coupling parameter; segmenting the original infrared image to obtain a segmented pure infrared target image; performing feature extraction on the pure infrared target image to obtain overall features; simply fitting the pure infrared target image into the background image according to the infrared target coupling position, and performing feature extraction to obtain detail features; and inputting the overall features and the detail features into a pre-trained coupling editing diffusion generation model, and generating an infrared target and background coupling editing image by the coupling editing diffusion generation model. According to the method, the required infrared image can be generated through efficient and high-quality coupling editing diffusion under the zero-sample condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer vision technology, and in particular to a method for generating target and background coupled editing diffusion of infrared images. Background Art

[0002] The role of generative models is to learn the latent distribution of data and generate new samples similar to the training samples. The origin of generative models can be traced back to the probabilistic model in the field of statistics. Most of the early generative models were based on traditional probability distribution and classical statistical learning methods, such as Gaussian mixture models and hidden Markov models. These models can generate samples similar to real data by learning the latent distribution of data. However, due to the limitations of traditional computing power, they cannot effectively process high-dimensional complex data, especially images, videos, and voice. In recent years, image generation models based on deep learning have become a research hotspot. Domestic and foreign research teams have successively proposed VAE (Variational Auto Encoder), GAN (Generative Adversial Network), and diffusion models.

[0003] Stability AI has proposed a stable diffusion model. Unlike VAE and GAN, the stable diffusion model uses a conditional generation strategy to generate corresponding images by controlling the input prompt text. Compared with traditional diffusion models, it has higher generation efficiency and image quality, and has a wider range of application scenarios. The innovation of the stable diffusion model lies in its use of the diffusion process of the latent space. It compresses the input image into a latent variable through a specific image encoder, rather than directly calculating in the pixel space of the image, which greatly improves the efficiency of image generation. The stable diffusion model also combines the encoded text information with the model through the mutual attention mechanism, so that the generated image can be generated according to the given text prompt.

[0004] However, simply relying on the text-based image method for data production will result in poor image interpretability and lack of human controllability. There is also a phenomenon that the generation effect is not good and the user's intention cannot be accurately understood. Therefore, it is necessary to introduce editing methods based on the generated image. At present, the target and background coupled editing diffusion generation methods of visible light images are mainly divided into the following categories: methods based entirely on text descriptions, methods with text descriptions and control conditions (such as mask maps, object structure maps, segmentation maps, etc.), and methods with image control conditions. These methods can be used to add, delete, replace, change targets, and transform and transfer styles of backgrounds. In order to meet people's requirements for controllability and refinement of edited images, more image editing methods with indicator control information have been proposed.

[0005] However, the coupled edit diffusion generation method for visible light images that has been proposed requires the use of image data, semantic descriptions, and other complex parameter information (such as segmentation maps). When editing, the semantic description of the editing method or the mask map of the target must also be provided, which has poor controllability of the generated image. In addition, there is currently no effective research on the coupled edit diffusion generation method for infrared images. Directly applying the coupled edit diffusion generation method for visible light images to infrared images will cause problems such as inaccurate target characteristics and poor controllability of the target, resulting in low quality of infrared images generated by coupled edit diffusion, and it is impossible to achieve low-cost and high-efficiency editing and generation of diversified data. Summary of the invention

[0006] In order to solve the above technical problems, an embodiment of the present application proposes a method for generating coupled editing diffusion of targets and backgrounds in infrared images, which allows users to interactively place targets at specific positions in the background, and use optional morphological controls to control the shape and posture of the targets, ultimately achieving efficient and high-quality coupled editing diffusion generation of the required infrared images under zero-sample conditions.

[0007] In order to achieve the above-mentioned purpose, an embodiment of the present application proposes a method for generating coupled edited diffusion of targets and backgrounds of infrared images, the method comprising: obtaining an original infrared image containing an infrared target, a background image not containing an infrared target, an infrared target coupling position and an infrared target coupling parameter; segmenting the original infrared image, removing the background of the infrared target, and aligning the extracted infrared target to the center of the image to obtain a segmented pure infrared target image; extracting features from the pure infrared target image to obtain overall features; simply fitting the pure infrared target image to the background image according to the infrared target coupling position to obtain a priori image as a priori of a pre-trained coupled edited diffusion generation model, and extracting features from the priori image by high-pass filtering, wavelet transform, and discrete cosine transform to obtain detail features; inputting the overall features and detail features into the coupled edited diffusion generation model, and the coupled edited diffusion generation model generates an infrared target and background coupled edited image according to the overall features, detail features, infrared target coupling position, and infrared target coupling parameters.

[0008] In order to achieve the above-mentioned purpose, an embodiment of the present application also proposes a target and background coupled editing diffusion generation system of an infrared image, the system comprising: an acquisition module, used to acquire an original infrared image containing an infrared target, a background image not containing an infrared target, an infrared target coupling position and an infrared target coupling parameter; a segmentation module, used to segment the original infrared image, remove the background of the infrared target, and align the extracted infrared target to the center of the image to obtain a segmented pure infrared target image; an overall feature extraction module, used to extract features from the pure infrared target image to obtain overall features; a detail feature extraction module, used to simply fit the pure infrared target image to the background image according to the infrared target coupling position to obtain a priori image as a priori of a pre-trained coupled editing diffusion generation model, and extract features from the priori image by high-pass filtering, wavelet transform, and discrete cosine transform to obtain detail features; a coupled editing diffusion generation model, used to generate an infrared target and background coupled editing image based on overall features, detail features, infrared target coupling position, and infrared target coupling parameters.

[0009] In order to achieve the above-mentioned purpose, an embodiment of the present application also proposes an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a method for generating target and background coupled editing diffusion of infrared images as described above.

[0010] In order to achieve the above objectives, an embodiment of the present application further proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement a target and background coupled editing diffusion generation method for infrared images as described above.

[0011] The embodiment of the present application proposes a method for generating infrared images with coupled editing and diffusion of targets and backgrounds. The pre-trained coupled editing and diffusion generation model is used to perform image generation tasks. A variety of infrared images containing infrared targets can be generated at low cost and high efficiency. The generation speed is fast and the target types are diverse, which improves the production efficiency of infrared images. In view of the user's control requirements for specific infrared targets, the present application can achieve high-fidelity coupled generation of infrared targets and backgrounds. On the basis of ensuring the image fidelity, the size, rotation angle, location, etc. of the infrared target can be arbitrarily edited to achieve the purpose of high-value data amplification. The user interaction operation is simple, and the image can be visually edited. The experimental results show that the infrared target and background coupled editing and diffusion generation method proposed in the present application has fast generation speed, low cost, high fidelity and strong diversity of the generated image, and the generation speed can be as high as 10 seconds / sheet. Compared with the current method of obtaining data by actual measurement and simulation, the time, manpower and material costs required to obtain the same amount of data are greatly reduced when the amount of data processed is similar. At the same time, the present application belongs to an intelligent generative large model, the result can be simply operated by the user, and the image can be visually edited, which well meets the diverse needs of users.

[0012] In some optional embodiments, the original infrared image containing the infrared target, the background image not containing the infrared target, the infrared target coupling position and the infrared target coupling parameters are obtained through a preset user visualization interaction interface, and the pre-trained coupling editing diffusion generation model is pre-loaded into the kernel of the user visualization interaction interface and called when in use.

[0013] In some optional embodiments, the original infrared image containing the infrared target carries the coordinate information of the infrared target and the reference mask corresponding to the infrared target. The infrared target coupling position characterization requires coupling the infrared target to the position of the background image that does not contain the infrared target. The infrared target coupling parameters include control strength, reasoning steps, guidance scale, seed code, whether to perform DDE image enhancement, whether to refine the reference mask, and whether to perform shape controllability; wherein, the higher the guidance scale, the higher the fidelity of the infrared target, and vice versa, the better the fusion with the background image. If DDE image enhancement is selected, DDE image enhancement will be performed on the original infrared image and the background image. If the reference mask is refined, a pre-loaded pre-trained improvement model will be called to refine the reference mask corresponding to the infrared target. If shape controllability is selected, the user is allowed to manually adjust the shape and posture of the infrared target in the infrared target and background coupled edited image. If shape controllability is not selected, the shape and posture of the infrared target in the infrared target and background coupled edited image will be automatically adjusted.

[0014] In some optional embodiments, the original infrared image is segmented, the background of the infrared target is removed, and the extracted infrared target is aligned to the center of the image to obtain a segmented pure infrared target image, including: finding the infrared target in the original infrared image based on the coordinate information of the infrared target; segmenting and extracting the infrared target from the original infrared image according to the coordinate information of the infrared target and a reference mask corresponding to the infrared target to remove the background of the infrared target; aligning the extracted infrared target to the center of the image to obtain a segmented pure infrared target image.

[0015] In some optional embodiments, after obtaining the segmented pure infrared target image, the method further includes: performing feature extraction on the pure infrared target image, and before obtaining the overall feature, performing image enhancement on the pure infrared target image by means of histogram equalization, noise filtering, and DDE to obtain an enhanced pure infrared target image.

[0016] In some optional embodiments, feature extraction is performed on a pure infrared target image to obtain overall features, including: encoding the enhanced pure infrared target image into global tokens and patch tokens respectively by means of tokens; splicing the global tokens and patch tokens, and then using a single linear layer as a projection to align the spliced ​​tokens into the latent space of a coupled edited diffusion generative model to obtain overall features.

[0017] In some optional embodiments, the pre-trained coupled-editing diffusion generation model is trained by the following steps: collecting sample infrared target images and sample background images that do not contain sample infrared targets by actual shooting, simulation generation, and intelligent generation; obtaining target anchor points and shape masks, generating real labels corresponding to the sample infrared target images and the sample background images, and performing label annotation; wherein the target anchor points indicate the positions where the sample infrared targets are coupled to the sample background images; performing feature extraction on the sample infrared target images to obtain overall features of the sample infrared targets; and simply fitting the sample infrared target images to the sample background images according to the infrared target coupling positions to obtain sample prior images as prior images for the coupled-editing diffusion generation model, and performing label annotation on the coupled-editing diffusion generation model. The features of the sample prior target image are extracted by high-pass filtering, wavelet transform and discrete cosine transform to obtain the detailed features of the sample infrared target; the overall features and detailed features of the sample infrared target are input into the coupled edited diffusion generation model for iterative training, and the target anchor point and the shape mask are aligned with each other, and the shape mask is used to indicate the pose of the sample infrared target image. At the same time, the shape mask is sampled at different scales, and random dilation and erosion are applied to remove details until the coupled edited diffusion generation model converges, so that the coupled edited diffusion generation model can obtain the ability of coupled edited diffusion generation of the target and background of the infrared image, and the ability to control the shape of the infrared target based on the obtained rough shape mask. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related technologies, the drawings required for use in the embodiments of the present application or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 It is a flow chart of a method for generating target and background coupled editing diffusion of infrared images provided in one embodiment of the present application;

[0020] Figure 2 It is a schematic diagram of the principle of a method for generating target and background coupled editing diffusion of infrared images provided in one embodiment of the present application;

[0021] Figure 3 is a schematic diagram of regulating infrared target coupling parameters provided in one embodiment of the present application;

[0022] Figure 4 is a schematic diagram of an infrared target coupling position provided in one embodiment of the present application;

[0023] Figure 5 The generated infrared target and background coupled editing image provided in one embodiment of the present application

[0024] Figure 6 is a structural schematic diagram of a target and background coupled editing diffusion generation system for infrared images provided in another embodiment of the present application;

[0025] Figure 7 It is a structural schematic diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings. In the various embodiments of the present application, many technical details are proposed in order to make the reader better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present application can also be implemented. The division of the following embodiments is only for the convenience of description, and the specific implementation mode of the present application should not constitute any limitation. The various embodiments can be combined with each other and referenced to each other under the premise of no contradiction.

[0027] In recent years, image generation models based on deep learning have become a research hotspot, and research teams at home and abroad have successively proposed VAE, GAN, and diffusion models.

[0028] In 2013, Kingma and Welling proposed a VAE generative model based on deep learning. The core idea of ​​VAE is a deep learning model based on the variational Bayesian method, which is used to learn the potential representation of data distribution. It learns the data generation process by maximizing the lower bound of the log-likelihood of the data, and then uses this lower bound to approximate the posterior distribution of the latent variables to generate new data samples. The advantage of VAE is that it combines variational inference and autoencoder frameworks, and performs well when processing high-dimensional data. However, there are problems such as insufficient quality and diversity of VAE's generated samples. The generated samples may be accompanied by distortion, and are not flexible enough when processing complex data distributions. In addition, the generation process of VAE involves complex probability calculations and optimization problems, and the time and computational costs are high during the training process.

[0029] In 2014, Goodfellow et al. proposed the generative adversarial network (GAN), which trains the generator and discriminator through game theory, in which the generator is responsible for generating fake data, and the discriminator tries to distinguish between real data and generated data. In the end, the generator and the discriminator are continuously optimized through game theory to achieve the goal of generating high-fidelity data. GAN has achieved good results in tasks such as image generation and image super-resolution. For example, the Pix2Pix and CycleGAN models can convert input images into output images of specific styles or specific tasks, while the StyleGAN model can control the style and content of generated images, making the generated images more diverse and controllable. However, when processing multiple different categories of image data, the training and optimization process is unstable, and the quality and effect of the generated images may be poor.

[0030] In 2020, Ho et al. proposed DDPM (Denoising Diffusion Probabilistic Models) to simulate the physical diffusion process. DDPM includes a forward diffusion process and a reverse denoising process. The forward diffusion process gradually adds Gaussian noise to the real image data to finally obtain pure noise. The forward diffusion process generates a series of images with different noise levels, providing a large number of data samples for model training. In the reverse denoising process, a neural network model is constructed and the data samples provided for the forward diffusion process are input. The training goal of the model is to predict the noise added to a given noisy image. In this way, image data can be generated through a completely given Gaussian noise and several steps of reverse denoising process. Since the training process of the diffusion model is very stable and can generate high-quality images, it has gradually become a hot topic in the field of generative models in recent years. In 2022, based on DDPM, Stability AI proposed a stable diffusion model, which is a large text-to-image generative model based on a potential diffusion model. Unlike VAE and GAN, the stable diffusion model uses a conditional generation strategy to generate corresponding images by controlling the input prompt text. Compared with the traditional diffusion model, it has higher generation efficiency and image quality, and has a wider range of application scenarios. The innovation of the stable diffusion model lies in its use of the diffusion process of the latent space. It compresses the input image into a latent variable through a specific image encoder, rather than directly calculating in the pixel space of the image, which greatly improves the efficiency of image generation. The stable diffusion model also combines the encoded text information with the model through the mutual attention mechanism, so that the generated image can be generated according to the given text prompt.

[0031] However, simply relying on the text-based image method for data production will result in poor image interpretability and lack of human controllability. There is also a phenomenon that the generation effect is not good and the user's intention cannot be accurately understood. Therefore, it is necessary to introduce editing methods based on the generated image. At present, the target and background coupled editing diffusion generation methods of visible light images are mainly divided into the following categories: methods based entirely on text descriptions, methods with text descriptions and control conditions (such as mask maps, object structure maps, segmentation maps, etc.), and methods with image control conditions. These methods can be used to add, delete, replace, change targets, and transform and transfer styles of backgrounds. In order to meet people's requirements for controllability and refinement of edited images, more image editing methods with indicator control information have been proposed. From the classic Dreambooth, which can generate an edited image of a target by specifying 3 to 5 images and text descriptions of the target, to the later Instruct-pix2pix, which is completely driven by text training and can achieve style editing of the target, and Break-a-scene, which can achieve text-based editing and generation of multiple targets in a single image, to the recent Pair-diffusion, which uses structural diagrams to guide image editing, and the HIVE method, which can achieve target addition, deletion, and transformation, more and more editing and generation schemes have been proposed.

[0032] However, the coupled edit diffusion generation method for visible light images that has been proposed requires the use of image data, semantic description, and other complex parameter information (such as segmentation map). When editing, it is also necessary to provide semantic description of the editing method or information such as the mask map of the target, which has poor controllability of the generated image. In addition, there is no effective research on the coupled edit diffusion generation method for infrared images. Directly applying the coupled edit diffusion generation method for visible light images to infrared images will cause problems such as inaccurate target characteristics and poor controllability of the target, resulting in low quality of infrared images generated by coupled edit diffusion, and it is impossible to realize the editing and generation of diversified data at low cost and high efficiency. In order to realize the coupled generation of data for special scenes and specific infrared targets in infrared images, it is urgent to study a generation method that can control the position, shape, posture, etc. of infrared targets in infrared images and can accurately couple infrared targets and background images.

[0033] An embodiment of the present application proposes a method for generating a target and background coupled editing diffusion of an infrared image, which is applied to an electronic device, wherein the electronic device may be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described using a server as an example. The implementation details of the method for generating a target and background coupled editing diffusion of an infrared image proposed in this embodiment are described in detail below. The following content is only the implementation details provided for ease of understanding and is not necessary for implementing this solution.

[0034] The specific process of the infrared image target and background coupled editing diffusion generation method proposed in this embodiment can be as follows: Figure 1 As shown, its working principle is as follows Figure 2 As shown, the method specifically comprises the following steps:

[0035] Step 101, obtaining an original infrared image containing an infrared target, a background image not containing an infrared target, an infrared target coupling position, and an infrared target coupling parameter.

[0036] In the specific implementation, the server first needs to obtain the original infrared image containing the infrared target, the background image excluding the infrared target, the infrared target coupling position and the infrared target coupling parameters. The infrared target coupling position characterizes the position where the infrared target needs to be coupled to the background image excluding the infrared target, and the infrared target coupling parameters are used to control the posture, shape, quality, etc. of the infrared target coupled to the background image excluding the infrared target.

[0037] In an example, the server can obtain the original infrared image containing the infrared target, the background image not containing the infrared target, the infrared target coupling position and the infrared target coupling parameters input by the user through a preset user visualization interaction interface, and the pre-trained coupling editing diffusion generation model is pre-loaded into the kernel of the user visualization interaction interface and called when used.

[0038] In one example, the original infrared image containing the infrared target carries the coordinate information of the infrared target and the reference mask corresponding to the infrared target. The infrared target coupling position characterization requires coupling the infrared target to the position of the background image that does not contain the infrared target. The interface for inputting and selecting the infrared target coupling parameters can be as follows: Figure 3 As shown, the infrared target coupling parameters include control strength, reasoning steps, guidance scale, seed code, whether to perform DDE image enhancement, whether to refine the reference mask, and whether to perform shape controllability. Among them, the higher the guidance scale, the higher the fidelity of the infrared target, and vice versa, the better the fusion with the background image. If DDE image enhancement is selected, the original infrared image and background image will be enhanced by DDE. If the reference mask is refined, the pre-loaded pre-trained improvement model will be called to refine the reference mask corresponding to the infrared target. If shape controllability is selected, the user is allowed to manually adjust the shape and posture of the infrared target in the infrared target and background coupling edited image. If shape controllability is not selected, the shape and posture of the infrared target in the infrared target and background coupling edited image will be automatically adjusted.

[0039] Step 102, segmenting the original infrared image, removing the background of the infrared target, and aligning the extracted infrared target to the center of the image to obtain a segmented pure infrared target image.

[0040] In a specific implementation, after obtaining the original infrared image containing the infrared target, the background image not containing the infrared target, the infrared target coupling position and the infrared target coupling parameters, the server can segment the original infrared image, remove the background of the infrared target, and align the extracted infrared target to the center of the image, thereby obtaining a segmented pure infrared target image.

[0041] In one example, when the server performs image segmentation, it needs to first find the infrared target in the original infrared image based on the coordinate information of the infrared target, and then segment and extract the infrared target from the original infrared image according to the coordinate information of the infrared target and the reference mask corresponding to the infrared target to remove the background of the infrared target, and finally align the extracted infrared target to the center of the image to obtain a segmented pure infrared target image.

[0042] It is worth noting that the image segmentation task can be performed by a pre-trained image segmentation model. The server only needs to input the original infrared image, the coordinate information of the infrared target, and the reference mask corresponding to the infrared target into the image segmentation model to obtain the pure infrared target image output by the image segmentation model.

[0043] In one example, considering that the infrared target is likely to be only a small part of the original infrared image and the quality may not be high, after obtaining the segmented pure infrared target image, the server needs to enhance the pure infrared target image through histogram equalization, noise filtering, and DDE to obtain an enhanced pure infrared target image.

[0044] Step 103, extracting features from the pure infrared target image to obtain overall features.

[0045] In a specific implementation, after obtaining the pure infrared target image, the server can extract features from the pure infrared target image to obtain overall features.

[0046] In one example, the server can encode the enhanced clean infrared target image into global tokens and patch tokens respectively by tokens, then concatenate the global tokens and patch tokens, and then use a single linear layer as a projection to align the concatenated tokens into the latent space of the coupled edit diffusion generative model to obtain the overall features. Concatenating the global tokens and patch tokens can well preserve more feature information.

[0047] It is worth noting that the overall feature extraction task can be performed by a pre-trained overall feature extractor. The server only needs to input the enhanced pure infrared target image into the overall feature extractor to obtain the overall features output by the overall feature extractor.

[0048] Step 104, simply fit the pure infrared target image to the background image according to the infrared target coupling position to obtain a prior image, which is used as the prior of the pre-trained coupled editing diffusion generation model, and extracts features from the prior image by high-pass filtering, wavelet transform, and discrete cosine transform to obtain detail features.

[0049] In the specific implementation, the server needs to extract detailed features while extracting overall features. The server simply fits the pure infrared target image to the background image according to the infrared target coupling position to obtain a prior image, which is used as the prior of the pre-trained coupled editing diffusion generation model. Then, the server extracts features from the prior image through high-pass filtering, wavelet transformation, and discrete cosine transform to obtain detailed features.

[0050] It is understandable that step 103 may be performed before step 104 , may be performed after step 104 , or may be performed simultaneously with step 104 .

[0051] In one example, the server needs to simply fit the enhanced pure infrared target image to the background image according to the infrared target coupling position to obtain a prior image as the prior of the coupled edit diffusion generation model, and then extract the features of the prior image through high-pass filtering, wavelet transformation, and discrete cosine transform to obtain detail features. Detail features can well supplement the detail information ignored during the overall feature extraction.

[0052] It is worth noting that the detail feature extraction task can be performed by a pre-trained detail feature extractor. The server only needs to input the prior image into the detail feature extractor to obtain the detail features output by the detail feature extractor.

[0053] Step 105, inputting the overall features and the detail features into the coupled edited diffusion generation model, and the coupled edited diffusion generation model generates an infrared target and background coupled edited image according to the overall features, the detail features, the infrared target coupling position, and the infrared target coupling parameters.

[0054] In a specific implementation, after obtaining the overall features and detail features, the server can input the overall features and detail features into a coupled editing diffusion generation model, and the coupled editing diffusion generation model generates an infrared target and background coupled editing image based on the overall features, detail features, infrared target coupling position, and infrared target coupling parameters.

[0055] In one example, the coupled edited diffusion generation model is trained by the following steps. First, the server needs to collect sample infrared target images and sample background images that do not contain sample infrared targets by actual shooting, simulation generation, and intelligent generation. During the collection, preliminary screening is required to remove data that do not meet the training and reasoning requirements of the infrared target and background coupling task, such as poor clarity and overly blurred targets. Subsequently, the server obtains the target anchor point and shape mask, generates the real label corresponding to the sample infrared target image and the sample background image, and performs label annotation. The target anchor point indicates the position where the sample infrared target is coupled to the sample background image. Next, the server needs to perform feature extraction on the sample infrared target image to obtain the overall features of the sample infrared target, and simply fit the sample infrared target image to the sample background image according to the infrared target coupling position to obtain the sample prior image as the prior of the coupled edited diffusion generation model. The sample prior target image is feature extracted by high-pass filtering, wavelet transformation, and discrete cosine transform to obtain the detailed features of the sample infrared target. Finally, the server inputs the overall features and detailed features of the sample infrared target into the coupled edit diffusion generation model for iterative training, and aligns the target anchor point with the shape mask, uses the shape mask to indicate the pose of the sample infrared target image, and samples the shape mask at different scales, and applies random dilation and erosion to remove details until the coupled edit diffusion generation model converges, so that the coupled edit diffusion generation model can obtain the ability of coupled edit diffusion generation of the target and background of the infrared image, and the ability to control the shape of the infrared target based on the obtained rough shape mask.

[0056] In one example, the server may also train the segmentation model, the overall feature extractor, the detail feature extractor, and the coupled editing diffusion generation model as a whole large model.

[0057] The present embodiment proposes a method for generating infrared images by coupling editing and diffusion of targets and backgrounds. The pre-trained coupled editing and diffusion generation model is used to perform image generation tasks. A variety of infrared images containing infrared targets can be generated at low cost and high efficiency. The generation speed is fast and the target types are diverse, which improves the production efficiency of infrared images. In view of the user's control requirements for specific infrared targets, the present application can achieve high-fidelity coupled generation of infrared targets and backgrounds. On the basis of ensuring the image fidelity, the size, rotation angle, location, etc. of the infrared target can be arbitrarily edited to achieve the purpose of high-value data amplification. The user interaction operation is simple, and the image can be visually edited. The experimental results show that the infrared target and background coupled editing and diffusion generation method proposed in the present application has fast generation speed, low cost, high fidelity and strong diversity of the generated image, and the generation speed can be as high as 10 seconds / sheet. Compared with the current method of obtaining data by actual measurement and simulation, the time, manpower and material costs required to obtain the same amount of data are greatly reduced when the amount of data processed is similar. At the same time, the present application belongs to an intelligent generative large model, the result can be simply operated by the user, and the image can be visually edited, which well meets the diverse needs of users.

[0058] The step division of the above methods is only for clear description. When implemented, they can be combined into one step, or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this application.

[0059] In one embodiment, we conducted relevant simulation experiments to verify the superiority of a target and background coupled edited diffusion generation method for infrared images (hereinafter referred to as this method) proposed in this application.

[0060] In the simulation experiment, we used a simulation data set containing 21041 image pairs to train the coupled editing diffusion generation model. The image resolution was set to 512px×512px, the optimizer used Adam, the learning rate was set to 0.0001, the number of training samples per batch on the GPU was set to 12, and the total number of training steps was set to 11856. The experimental environment used Python 3.8.5, the CPU model used was Intel(R) Xeon(R) Platinum 8352V, the GPU was NVIDIARTX A6000×4, and the operating system was Ubuntu 20.04.06LTS. In the coupled generation stage of infrared target and background images, infrared image data obtained by various methods including simulation platform simulation, instrument measurement, and large model generation were used.

[0061] The infrared target coupling position can be as follows Figure 4As shown, the generated infrared target and background coupled edited image can be Figure 5 As shown, input the background image in the measured data set and the infrared car target image generated by simulation, draw the location of the car in the target image, draw the target placement position in the background image, adjust the target size, click Run, the model generates the result of coupling the car with the infrared street background, and the location of the car is consistent with the specified location.

[0062] Another embodiment of the present application proposes a target and background coupled editing and diffusion generation system for infrared images. The implementation details of the target and background coupled editing and diffusion generation system for infrared images proposed in this embodiment are described in detail below. The following content is only the implementation details provided for the convenience of understanding and is not necessary for the implementation of this example.

[0063] Figure 6 It is a structural diagram of a coupled editing diffusion generation system for a target and background of an infrared image proposed in this embodiment. The system includes: an acquisition module 201, a segmentation module 202, an overall feature extraction module 203, a detail feature extraction module 204 and a coupled editing diffusion generation model 205.

[0064] The acquisition module 201 is used to acquire an original infrared image containing an infrared target, a background image not containing an infrared target, an infrared target coupling position and an infrared target coupling parameter.

[0065] The segmentation module 202 is used to segment the original infrared image, remove the background of the infrared target, and align the extracted infrared target to the center of the image to obtain a segmented pure infrared target image.

[0066] The overall feature extraction module 203 is used to extract features from the pure infrared target image to obtain overall features.

[0067] The detail feature extraction module 204 is used to simply fit the pure infrared target image to the background image according to the infrared target coupling position to obtain a prior image as the prior of the pre-trained coupled editing diffusion generation model, and extract features from the prior image by high-pass filtering, wavelet transformation, and discrete cosine transform to obtain detail features.

[0068] The coupled edit diffusion generation model 205 is used to generate an infrared target and background coupled edit image according to the overall features, detail features, infrared target coupling positions, and infrared target coupling parameters.

[0069] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application, but this does not mean that there are no other units in this embodiment.

[0070] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in conjunction with the above method embodiment. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment, and in order to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiments.

[0071] Another embodiment of the present application provides an electronic device, whose specific structure is as follows: Figure 7 As shown, it includes: at least one processor 301; and a memory 302 that is communicatively connected to the at least one processor 301; wherein the memory 302 stores instructions that can be executed by the at least one processor 301, and the instructions are executed by the at least one processor 301 so that the at least one processor 301 can execute a method for generating target and background coupled editing diffusion of infrared images as described in the above-mentioned method embodiments.

[0072] Among them, the memory and the processor can be connected in a bus manner, and the bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are all well known in the art and will not be further described in this article. The bus interface is responsible for providing an interface between the bus and the transceiver. The transceiver can be one component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor is transmitted on a wireless medium through an antenna, and further, the antenna also receives data and transmits the data to the processor.

[0073] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.

[0074] Another embodiment of the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement a method for generating target and background coupled editing diffusion of infrared images as described in the above method embodiments.

[0075] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (such as a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, ROM (Read-Only Memory), RAM (Random Access Memory), disk or optical disk and other media that can store program codes.

[0076] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.

Claims

1. A method for generating target and background coupled editing diffusion of infrared images, characterized in that: include: Acquire an original infrared image containing an infrared target, a background image not containing an infrared target, an infrared target coupling position, and an infrared target coupling parameter; Segment the original infrared image, remove the background of the infrared target, and align the extracted infrared target to the center of the image to obtain a segmented pure infrared target image; Extract features from pure infrared target images to obtain overall features; The pure infrared target image is simply fitted to the background image according to the infrared target coupling position to obtain a prior image, which is used as the prior of the pre-trained coupled editing diffusion generation model. The feature of the prior image is extracted by high-pass filtering, wavelet transform, and discrete cosine transform to obtain detail features. The overall features and detail features are input into the coupled edited diffusion generation model, and the coupled edited diffusion generation model generates an infrared target and background coupled edited image according to the overall features, detail features, infrared target coupling position, and infrared target coupling parameters.

2. The target and background coupled editing diffusion generation method of infrared image according to claim 1 is characterized in that: The original infrared image containing the infrared target, the background image excluding the infrared target, the infrared target coupling position and the infrared target coupling parameters are obtained through the preset user visualization interaction interface. The pre-trained coupling editing diffusion generation model is pre-loaded into the kernel of the user visualization interaction interface and is called when used.

3. The target and background coupled editing diffusion generation method of infrared image according to claim 2 is characterized in that: The original infrared image containing the infrared target carries the coordinate information of the infrared target and the reference mask corresponding to the infrared target. The infrared target coupling position characterization requires coupling the infrared target to the position of the background image that does not contain the infrared target. The infrared target coupling parameters include control strength, number of inference steps, guidance scale, seed code, whether to perform DDE image enhancement, whether to refine the reference mask, and whether to perform shape controllability. Among them, the higher the guidance scale, the higher the fidelity of the infrared target, and vice versa, the better the fusion with the background image. If you choose to perform DDE image enhancement, DDE image enhancement will be performed on the original infrared image and the background image. If you choose to refine the reference mask, the pre-loaded pre-trained improvement model will be called to refine the reference mask corresponding to the infrared target. If you choose to perform shape controllability, the user is allowed to manually adjust the shape and posture of the infrared target in the infrared target and background coupled editing image. If you choose not to perform shape controllability, the shape and posture of the infrared target in the infrared target and background coupled editing image will be automatically adjusted.

4. The target and background coupled editing diffusion generation method of infrared image according to claim 3 is characterized in that: Segment the original infrared image, remove the background of the infrared target, and align the extracted infrared target to the center of the image to obtain a segmented pure infrared target image, including: Find the infrared target in the original infrared image based on the coordinate information of the infrared target; According to the coordinate information of the infrared target and the reference mask corresponding to the infrared target, the infrared target is segmented and extracted from the original infrared image to remove the background of the infrared target; Align the extracted infrared target to the center of the image to obtain the segmented pure infrared target image.

5. The target and background coupled editing diffusion generation method of infrared image according to claim 4 is characterized in that: After obtaining the segmented pure infrared target image, extracting features from the pure infrared target image, and before obtaining the overall features, the method further includes: The pure infrared target image is enhanced by histogram equalization, noise filtering and DDE to obtain an enhanced pure infrared target image.

6. The target and background coupled editing diffusion generation method of infrared image according to claim 5 is characterized in that: Perform feature extraction on the pure infrared target image to obtain the overall features, including: By tokenization, the enhanced clean infrared target image is encoded into global tokens and patch tokens respectively; The global token and the patch token are concatenated, and a single linear layer is used as a projection to align the concatenated tokens into the latent space of the coupled edit-diffusion generative model to obtain the overall features.

7. A method for generating target and background coupled editing diffusion of infrared images according to any one of claims 1 to 6, characterized in that: The pre-trained coupled edit diffusion generative model is trained by the following steps: Collect sample infrared target images and sample background images that do not contain sample infrared targets by actual shooting, simulation generation, and intelligent generation; Obtaining target anchor points and shape masks, generating true labels corresponding to the sample infrared target image and the sample background image, and labeling the labels; wherein the target anchor points indicate the position where the sample infrared target is coupled to the sample background image; Extract features of sample infrared target images to obtain overall features of the sample infrared target; The sample infrared target image is simply fitted to the sample background image according to the infrared target coupling position to obtain the sample prior image, which is used as the prior of the coupled editing diffusion generation model. The sample prior target image is feature extracted by high-pass filtering, wavelet transform, and discrete cosine transform to obtain the detailed features of the sample infrared target. The overall features and detailed features of the sample infrared target are input into the coupled edited diffusion generation model for iterative training, and the target anchor point and the shape mask are aligned with each other. The shape mask is used to indicate the pose of the sample infrared target image. The shape mask is sampled at different scales, and random dilation and erosion are applied to remove details until the coupled edited diffusion generation model converges. The coupled edited diffusion generation model obtains the ability of coupled edited diffusion generation of targets and backgrounds of infrared images, as well as the ability to control the shape of infrared targets based on the obtained rough shape mask.

8. A target and background coupled editing diffusion generation system for infrared images, characterized in that: include: An acquisition module, used for acquiring an original infrared image containing an infrared target, a background image not containing an infrared target, an infrared target coupling position and an infrared target coupling parameter; The segmentation module is used to segment the original infrared image, remove the background of the infrared target, and align the extracted infrared target to the center of the image to obtain a segmented pure infrared target image; The overall feature extraction module is used to extract features from pure infrared target images to obtain overall features; The detail feature extraction module is used to simply fit the pure infrared target image to the background image according to the infrared target coupling position to obtain a prior image as the prior of the pre-trained coupled editing diffusion generation model, and extract features from the prior image by high-pass filtering, wavelet transform, and discrete cosine transform to obtain detail features; The coupled editing diffusion generation model is used to generate the coupled editing image of the infrared target and the background according to the overall features, the detailed features, the infrared target coupling position and the infrared target coupling parameters.

9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; Wherein, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a target and background coupled editing diffusion generation method for infrared images as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement a target and background coupled editing diffusion generation method for infrared images as described in any one of claims 1 to 7.