Method for training and testing controllable image generation model capable of reflecting fine-grained instance layout and learning device and testing device using the same
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-08-12
Smart Images

Figure R1020250135582_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a method for training and testing a controllable image generation model capable of reflecting a fine-grained instance layout, and a training device and a testing device using the same. More specifically, the invention relates to a method for training and testing a diffusion model, which is a controllable image generation model, using a semantic mask and a corresponding caption to reflect a fine-grained instance layout, and a training device and a testing device using the same. Background Technology
[0002] With the advancement of AI technology, image generation models that create new images based on inputs such as text or noise are being widely used. Representative examples include the Generative Adversarial Network (GAN) and the Diffusion model. The GAN model is a method in which a generator that creates images and a discriminator that determines whether an image is real or fake learn through competition, while the Diffusion model is a method that gradually generates images from noise.
[0003] Among these, the diffusion model utilizes a forward process that adds noise to the data to turn it into complete noise, and a reverse process that, conversely, generates data by gradually restoring it from the noise.
[0004] Recently, research in layout-to-image series for controlling instance layouts within images has been conducted, but these mainly rely on box-based layout representations. However, boxes are inaccurate for representing thin or slanted objects, and there is a problem in that the precision of instance-level control is reduced as the background is included inside the box.
[0005] Therefore, research is being conducted to utilize semantic masks that enable precise layout control by providing more accurate instance geometry and position information.
[0006] In the recent study 3DIS (Depth-Driven Decoupled Instance Synthesis), a method is proposed to more accurately reflect the instance layout during the denoising process of the inference step of the diffusion model by utilizing a mask obtained from a control image through a segmentation model.
[0007] However, 3DIS relies on depth-based control images, and since it is difficult to reflect small objects or fine instance masks, it has the problem of being limited to coarse layout controls.
[0008] Accordingly, the applicant intends to propose a diffusion model capable of generating an image by reflecting not only coars but also fine-grained instance layouts. The problem to be solved
[0009] The present invention aims to solve all of the aforementioned problems.
[0010] In addition, the present invention has another objective of providing a controllable image generation model capable of reflecting a fine-grain instance layout.
[0011] In addition, the present invention has another objective of improving image generation speed by eliminating dependency on an external segmentation model.
[0012] In addition, another objective of the present invention is to enable the expansion of the degree of control freedom for the image generation process by adjusting the balance between image fidelity and control information based on the time step during the denoising process. means of solving the problem
[0013] The characteristic configuration of the present invention for achieving the objectives of the present invention as described above and realizing the characteristic effects of the present invention described below is as follows.
[0014] According to one embodiment of the present invention, in a method for training a controllable image generation model capable of reflecting a fine-grained instance layout, the method comprises: (a) when a training image and a training time step are acquired, a training device inputs the training image into an encoder of a Variational Autoencoder (VAE) to cause the encoder of the VAE to generate a training image latent, and generates a training noisy image latent by repeatedly adding noise to the training image latent according to the training time step through a scheduler;(b) When the learning device acquires a learning semantic mask corresponding to the learning image and a learning image caption—the learning image caption includes a learning instance-level text for a learning class corresponding to the learning semantic mask and a learning global-level text corresponding to the learning image—the learning semantic mask is input to the encoder of the VAE to cause the encoder of the VAE to generate a learning semantic mask latent, and the learning control information is generated using the learning noisy image latent and the learning semantic mask latent; the learning image caption is input to a text encoder to cause the text encoder to generate a learning text embedding including a learning instance-level text embedding and a learning global-level text embedding; and the learning A step of performing a process of inputting a time step into a time step encoder to cause the time step encoder to encode the training time step and generate a training time step embedding;(c) A step in which the learning device inputs the learning time-step embedding, the learning text embedding, and the learning control information into a denoising network, causing the denoising network to generate learning prediction noise by referencing the learning control information, and causes the scheduler to generate a learning synthetic image latency by removing noise from the learning control information by referencing the learning prediction noise, wherein the denoising network generates learning intermediate prediction noise by referencing the learning intermediate synthetic image latency according to the learning time-step embedding, and the denoising process in which the scheduler generates the learning intermediate synthetic image latency from which noise has been removed by referencing the learning intermediate prediction noise is repeated, thereby generating the learning synthetic image latency; A method is disclosed comprising: (d) a step in which the learning device inputs the learning synthetic image latency to the decoder of the VAE to cause the decoder of the VAE to decode the learning synthetic image latency to generate a learning synthetic image, and trains the denoising network to minimize the loss generated by referencing the learning synthetic image and the learning image.
[0015] In the above embodiment, in step (b), the learning device concatenates the learning noisy image latency and the learning semantic mask latency by channel to generate the learning control information.
[0016] In the above embodiment, in step (c), the learning device performs the prediction noise generation process by adding zero-initial weights corresponding to the increased number of channels to the first layer of the denoising network in correspondence with the increased number of channels as the learning noisy image latency and the learning semantic mask latency are concatenated.
[0017] In the above embodiment, in step (b), the learning device inputs the learning noisy image latency, the learning semantic mask latency, the learning text embedding, and the learning time step embedding into a ControlNet to cause the ControlNet to generate a Control Signal, and generates learning control information including the Control Signal and the learning noisy image latency.
[0018] In the above embodiment, the control net is created by copying some of the layers of a pre-trained diffusion model and adding zero convolution layers to the first and last layers of the control net.
[0019] In the above embodiment, the learning semantic mask is an RGB image in which unique colors are assigned to each learning class included in the learning image.
[0020] According to another embodiment of the present invention, a method for testing a controllable image generation model capable of reflecting a fine-grained instance layout comprises: (a) a subprocess in which, by a training device, (i) when a training image and a training time step are acquired, the training image is input into an encoder of a Variational Autoencoder (VAE) to cause the encoder of the VAE to generate a training image latent, and a scheduler is used to repeatedly add noise to the training image latent according to the training time step to generate a training noisy image latent; and (ii) a training semantic mask corresponding to the training image and a training image caption, wherein the training image caption comprises training instance-level text for a training class corresponding to the training semantic mask and training global corresponding to the training image. Including Global-level Text - when this is obtained, the process of inputting the training semantic mask into the encoder of the VAE to cause the encoder of the VAE to generate a training semantic mask latency, and generating training control information using the training noisy image latency and the training semantic mask latency; the process of inputting the training image caption into a text encoder to cause the text encoder to generate a training text embedding including a training instance-level text embedding and a training global-level text embedding.and a subprocess including a process of inputting the training time step into a time step encoder to cause the time step encoder to encode the training time step to generate a training time step embedding, (iii) a subprocess of inputting the training time step embedding, the training text embedding, and the training control information into a denoising network to cause the denoising network to generate training prediction noise by referencing the training control information, and causing the scheduler to generate a training synthetic image latency by removing noise from the training control information by referencing the training prediction noise, wherein the prediction noise generation process in which the denoising network generates training intermediate prediction noise according to the time step embedding and the denoising process in which the scheduler generates the training intermediate synthetic image latency from which noise has been removed by referencing the training intermediate prediction noise are repeated, thereby generating the training synthetic image latency. and (iv) a sub-process is performed to input the training synthetic image latency into the decoder of the VAE so that the decoder of the VAE decodes the training synthetic image latency to generate a training synthetic image, and to train the denoising network to minimize the loss generated by referencing the training synthetic image and the training image, so that the denoising model is trained, and the test device, a test noisy latency, a test semantic mask,A step of obtaining a test image caption corresponding thereto—the test image caption includes test instance-level text for a test class corresponding to the test semantic mask and test global-level text corresponding to the test semantic mask—and a test time step; (b) The test device performs the steps of: generating a test semantic noisy latency using the test semantic mask and the test noisy latency; inputting the test image caption into the text encoder to cause the text encoder to generate a test text embedding including a test instance-level text embedding and a test global-level text embedding; and inputting the test time step into the time step encoder to cause the time step encoder to encode the test time step to generate a test time step embedding; (c) The test device inputs the test semantic noisy latency, the test text embedding, and the test timestep embedding into the denoising network to cause the denoising network to generate test prediction noise by referencing the test semantic noisy latency, and causes the scheduler to generate test synthetic image latency by removing noise from the test semantic noisy latency by referencing the test prediction noise, whereinA method is disclosed comprising: a step of generating a test synthetic image latency by repeating a prediction noise generation process in which the denoising network generates test intermediate prediction noise by referencing a test intermediate synthetic image latency according to the test time step embedding, and a denoising process in which the scheduler generates the test intermediate synthetic image latency from which noise has been removed by referencing the test intermediate prediction noise; and (d) a step in which the test device inputs the test synthetic image latency to a decoder of the VAE and causes the decoder of the VAE to decode the test synthetic image latency to generate a test synthetic image.
[0021] In the above embodiment, in step (b), the test device converts the test semantic mask to the same resolution as the test noisy latency, and then performs a 1:1 mapping operation with the test noisy latency to generate the test semantic noisy latency.
[0022] In the above embodiment, in step (c), the test device repeats the predictive noise generation process and the denoising process, wherein, according to the test time step, (i) in the initial denoising process in which the predictive noise generation process and the denoising process are repeated up to k - where k is a preset integer greater than or equal to 1 - the output calculation operation of each layer included in the denoising network includes (i-1) an attention operation for parts corresponding to an instance-level mask generated by mapping the test instance-level text embedding and the test noisy semantic mask latency or the test instance-level text embedding and the test intermediate synthetic image latency, and (i-2) an attention operation for the test global-level text embedding and the test noisy semantic mask latency or the test global-level text embedding and the test intermediate synthetic image latency, and (ii) a subsequent denoising process in which the predictive noise generation process and the denoising process are repeated after k times. In the process, the output calculation operation of each layer included in the denoising network includes an attention operation for the test global level text embedding and the test noisy semantic mask latency, or for the test global level text embedding and the test intermediate synthetic image latency.
[0023] In the above embodiment, the test semantic mask is an RGB image in which unique colors are assigned to each test class to be generated.
[0024] According to one embodiment of the present invention, a training device for a controllable image generation model capable of reflecting a fine-grained instance layout comprises: one or more memories for storing instructions; and includes one or more processors configured to execute the above instructions, wherein the processor comprises: (I) a process in which, when a training image and a training time step are acquired, the training image is input into an encoder of a Variational Autoencoder (VAE) to cause the encoder of the VAE to generate a training image latency, and a process in which noise is repeatedly added to the training image latency according to the training time step through a scheduler to generate a training noisy image latency; (II) a training semantic mask corresponding to the training image and a training image caption—the training image caption includes training instance-level text for a training class corresponding to the training semantic mask and training global-level text corresponding to the training image—when these are acquired, the training semantic mask is input into the encoder of the VAE to generate the VAE A process of causing an encoder to generate a learning semantic mask latency, and generating learning control information using the learning noisy image latency and the learning semantic mask latency,A process of inputting the above-mentioned training image caption into a text encoder to cause the text encoder to generate a training text embedding including a training instance-level text embedding and a training global-level text embedding, and a process of inputting the above-mentioned training time step into a time step encoder to cause the time step encoder to encode the above-mentioned training time step to generate a training time step embedding, (III) inputting the above-mentioned training time step embedding, the above-mentioned training text embedding, and the above-mentioned training control information into a denoising network to cause the denoising network to generate training prediction noise by referencing the above-mentioned training control information, and causing the scheduler to generate a training synthetic image latency by removing noise from the above-mentioned training control information by referencing the above-mentioned training prediction noise, wherein the denoising according to the above-mentioned training time step embedding A learning device is disclosed that performs a process of generating prediction noise in which a network references a learning intermediate synthetic image latency to generate learning intermediate prediction noise, and a process of generating a learning intermediate synthetic image latency by repeating a denoising process in which a scheduler references the learning intermediate prediction noise to generate the learning intermediate synthetic image latency from which noise has been removed, thereby generating the learning synthetic image latency; and (IV) input the learning synthetic image latency to a decoder of the VAE to cause the decoder of the VAE to decode the learning synthetic image latency to generate a learning synthetic image, and training the denoising network to minimize the loss generated by referencing the learning synthetic image and the learning image.
[0025] In the above embodiment, the processor generates the learning control information by concatenating the learning noisy image latency and the learning semantic mask latency by channel in the process (II).
[0026] In the above embodiment, the processor performs the prediction noise generation process by adding zero-initial weights corresponding to the increased number of channels to the first layer of the denoising network in correspondence with the increased number of channels as the learning noisy image latency and the learning semantic mask latency are concatenated in the process (III).
[0027] In the above embodiment, the processor inputs the learning noisy image latency, the learning semantic mask latency, the learning text embedding, and the learning time step embedding into the ControlNet in the process (II) to cause the ControlNet to generate a Control Signal, and generates the learning control information including the Control Signal and the learning noisy image latency.
[0028] In the above embodiment, the control net is created by copying some of the layers of a pre-trained diffusion model and adding zero convolution layers to the first and last layers of the control net.
[0029] In the above embodiment, the learning semantic mask is an RGB image in which unique colors are assigned to each learning class included in the learning image.
[0030] According to another embodiment of the present invention, in a test device for a controllable image generation model capable of reflecting a fine-grained instance layout, one or more memories for storing instructions; and includes one or more processors configured to execute the above instructions, wherein the processor comprises: (I) a subprocess that, when a training image and a training timestep are acquired by a training device, inputs the training image into an encoder of a Variational Autoencoder (VAE) to cause the encoder of the VAE to generate a training image latency, and generates a training noisy image latency by repeatedly adding noise to the training image latency according to the training timestep through a scheduler; (ii) when a training semantic mask corresponding to the training image and a training image caption—the training image caption includes training instance-level text for a training class corresponding to the training semantic mask and training global-level text corresponding to the training image—are acquired, the training semantic mask of the VAE A process of inputting into an encoder to cause the encoder of the VAE to generate a learning semantic mask latency, and generating learning control information using the learning noisy image latency and the learning semantic mask latency,A subprocess comprising: a process of inputting the above-mentioned training image caption into a text encoder to cause the text encoder to generate a training text embedding including a training instance-level text embedding and a training global-level text embedding; and a process of inputting the above-mentioned training time step into a time step encoder to cause the time step encoder to encode the above-mentioned training time step to generate a training time step embedding; (iii) inputting the above-mentioned training time step embedding, the above-mentioned training text embedding, and the above-mentioned training control information into a denoising network to cause the denoising network to generate training prediction noise by referencing the above-mentioned training control information, and causing the scheduler to generate a training synthetic image latency by removing noise from the above-mentioned training control information by referencing the above-mentioned training prediction noise, wherein the denoising network according to the above-mentioned time step embedding A sub-process that generates a prediction noise generation process for generating training intermediate prediction noise and a denoising process in which the scheduler refers to the training intermediate prediction noise to generate the training intermediate synthetic image latency from which noise has been removed, thereby generating the training synthetic image latency; and (iv) a sub-process that inputs the training synthetic image latency to the decoder of the VAE to cause the decoder of the VAE to decode the training synthetic image latency to generate a training synthetic image, and trains the denoising network to minimize the loss generated by referring to the training synthetic image and the training image, so that the denoising model is trained, and a test noisy latency,A test semantic mask, a test image caption corresponding thereto—the test image caption includes test instance-level text for a test class corresponding to the test semantic mask and test global-level text corresponding to the test semantic mask—and a processor acquiring a test timestep, (II) a process of generating a test semantic noisy latency using the test semantic mask and the test noisy latency, a process of inputting the test image caption into the text encoder to cause the text encoder to generate a test text embedding including a test instance-level text embedding and a test global-level text embedding, and a process of inputting the test timestep into the timestep encoder to cause the timestep encoder to encode the test timestep to generate a test timestep embedding. A process for performing the process, (III) inputting the test semantic noisy latency, the test text embedding, and the test timestep embedding into the denoising network to cause the denoising network to generate test prediction noise by referencing the test semantic noisy latency, and causing the scheduler to generate test synthetic image latency by removing noise from the test semantic noisy latency by referencing the test prediction noise,A test device is disclosed that performs the following processes: a prediction noise generation process in which the denoising network generates test intermediate prediction noise by referencing the test intermediate synthetic image latency according to the test time step embedding, and a denoising process in which the scheduler generates the test intermediate synthetic image latency from which noise has been removed by referencing the test intermediate prediction noise, thereby generating the test synthetic image latency; and (IV) input the test synthetic image latency to the decoder of the VAE and cause the decoder of the VAE to decode the test synthetic image latency to generate a test synthetic image.
[0031] In the above embodiment, the processor, in the process (II), converts the test semantic mask to the same resolution as the test noisy latency, and then performs a 1:1 mapping operation with the test noisy latency to generate the test semantic noisy latency.
[0032] In the above embodiment, the processor, in step (III), repeats the predictive noise generation process and the denoising process, wherein, according to the test time step, (i) in the initial denoising process in which the predictive noise generation process and the denoising process are repeated up to k - where k is a preset integer greater than or equal to 1 - the output calculation operation of each layer included in the denoising network includes (i-1) an attention operation for parts corresponding to an instance-level mask generated by mapping the test instance-level text embedding and the test noisy semantic mask latency or the test instance-level text embedding and the test intermediate synthetic image latency, and (i-2) an attention operation for the test global-level text embedding and the test noisy semantic mask latency or the test global-level text embedding and the test intermediate synthetic image latency, and (ii) a subsequent process in which the predictive noise generation process and the denoising process are repeated after k times. In the denoising process, the output calculation operation of each layer included in the denoising network includes an attention operation for the test global level text embedding, the test noisy semantic mask latency, and the test intermediate synthetic image latency.
[0033] In the above embodiment, the test semantic mask is an RGB image in which unique colors are assigned to each test class to be generated. Effects of the invention
[0034] The present invention has the effect of providing a controllable image generation model capable of reflecting a fine-grain instance layout.
[0035] In addition, the present invention has the effect of improving image generation speed by eliminating dependency on external segmentation models.
[0036] In addition, the present invention has the effect of expanding the degree of control freedom for the image generation process by adjusting the balance between image fidelity and control information based on the time step during the denoising process. Brief explanation of the drawing
[0037] The drawings attached below for use in describing embodiments of the present invention are merely some of the embodiments of the present invention, and other drawings can be obtained based on these drawings without inventive work by a person skilled in the art to which the present invention pertains (hereinafter "person skilled in the art"). FIG. 1 schematically illustrates a learning device for learning a diffusion model, which is a controllable image generation model capable of reflecting a fine-grained instance layout according to an embodiment of the present invention. FIG. 2 schematically illustrates a diffusion model, which is a controllable image generation model according to one embodiment of the present invention, and FIG. 3a illustrates an exemplary method for training a diffusion model according to an embodiment of the present invention, and FIG. 3b schematically illustrates another method of training a diffusion model according to one embodiment of the present invention, and FIG. 4 schematically illustrates a test apparatus for testing a diffusion model, which is a controllable image generation model capable of reflecting a fine-grain instance layout according to an embodiment of the present invention. FIG. 5 schematically illustrates a diffusion model learned according to one embodiment of the present invention, and FIG. 6 illustrates an exemplary method for testing a diffusion model according to one embodiment of the present invention, and FIG. 7 exemplarily illustrates the process of an initial denoising step and a later denoising step according to the progression of a time step in accordance with an embodiment of the present invention, and FIG. 8 schematically illustrates a semantic mask, which is a control image input according to one embodiment of the present invention. Specific details for implementing the invention
[0038] The following detailed description of the invention refers to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention. It should be understood that various embodiments of the invention are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be modified from one embodiment to another without departing from the spirit and scope of the invention. It should also be understood that the location or arrangement of individual components within each embodiment may be modified without departing from the spirit and scope of the invention. Accordingly, the following detailed description is not intended to be limited in meaning, and the scope of the invention should be understood to encompass the scope claimed by the claims and all equivalents thereof. Similar reference numerals in the drawings indicate identical or similar components across various aspects.
[0039] For reference, throughout this specification, "for learning" or "learning" has been added to terms related to the learning process, and "for testing" or "test" has been added to terms related to the testing process to avoid possible confusion.
[0040] Hereinafter, in order to enable a person skilled in the art to easily practice the present invention, various preferred embodiments of the present invention will be described in detail with reference to the attached drawings.
[0041] FIG. 1 schematically illustrates a learning device for learning a controllable image generation model capable of reflecting a fine-grain instance layout according to an embodiment of the present invention. Referring to FIG. 1, the learning device (100) may include a memory (110) in which instructions for learning a diffusion model (300), which is a controllable image generation model capable of reflecting a fine-grain instance layout, are stored, and a processor (120) that performs an operation to learn the diffusion model (300) according to the instructions stored in the memory (110). At this time, although the diffusion model (300) is shown as being installed in the learning device (100), alternatively, the diffusion model (300) may be installed in a cloud environment or installed in a computing device different from the learning device (100).
[0042] Specifically, the learning device (100) may achieve desired system performance by utilizing a combination of a computing device (e.g., a device that may include components of a computer processor, memory, storage, input device and output device, and other conventional computing devices; an electronic communication device such as a router, switch, etc.; an electronic information storage system such as a Network Attached Storage (NAS) and a Storage Area Network (SAN)) and computer software (i.e., instructions that cause the computing device to function in a specific way), but is not limited thereto.
[0043] Additionally, the processor (120) of the learning device (100) may include hardware configurations such as an MPU (Micro Processing Unit) or CPU (Central Processing Unit), cache memory, and data bus. Additionally, the computing device may further include software configurations such as an operating system and an application for a specific purpose.
[0044] However, this does not exclude the case where the learning device (100) includes an integrated processor in which a medium, a processor, and a memory are integrated for implementing the present invention.
[0045] Additionally, referring to FIG. 2, a diffusion model (300), which is a controllable image generation model according to one embodiment of the present invention, may include a denoising network (310) that generates predicted noise that predicts input noise will be denoised at the next time step, a scheduler (320) that performs a forward process that adds noise to a training image to make the training image safe noise and a reverse process that generates a training image by gradually restoring it from the noise, a Variational AutoEncoder (VAE) (330) including an encoder that encodes an input image into an image latency and a decoder that decodes the image latency into an image, a text encoder (340) that encodes input text, and a time step encoder (350) that encodes an input time step. However, the present invention is not limited thereto, and the diffusion model (300) may be configured with various architectures depending on the user. In this case, the time step may be expressed as an integer greater than or equal to 1.
[0046] A method for learning a diffusion model according to an embodiment of the present invention using a learning device (100) configured as described above is explained below with reference to FIGS. 2 and FIGS. 3a. First, when a learning image (10) and a learning time step are obtained according to a forward process, the learning device (100) inputs the learning image (10) into the encoder of the VAE (320) to cause the encoder of the VAE (330) to encode the learning image (10) to generate a learning image latent, and can generate a learning noisy image latent (14) by repeatedly adding noise to the learning image latent according to the learning time step through a scheduler (320). At this time, the learning time step indicates the number of iterations for adding or removing noise, but is not limited thereto. In addition, at this time, the learning noisy image latency (14) is complete noise generated by the repetitive addition of noise, and may be Gaussian noise, which is random noise following a normal distribution, but is not limited thereto.
[0047] For example, if noise is applied using 1,000 steps as training time steps, the scheduler (320) can generate a schedule sequence containing 1,000 numbers such as [0, 1, 2, 999] and the scheduler can generate a training noisy image latency by applying Gaussian noise to each noise step according to a predefined value included in a variance schedule. Here, the variance schedule is a fixed schedule and refers to small amounts of variance values that determine how much noise to add at each time step of the schedule sequence. Thus, the values of the variance schedule are defined to start from a small number close to 0 and gradually increase so that the data becomes increasingly corrupted over time. As another example, if noise is applied using 50 steps as the training time step, the scheduler can generate a training noisy image latency by generating a schedule sequence containing 50 numbers such as [0, 20, 940, 960, 980] and applying Gaussian noise according to the variance schedule. The numbers in the schedule sequence generated here may be generated as [0, 940, 960, 980] following uniform intervals, or they may follow non-uniform intervals using other optimization sampling methods.
[0048] Here, the learning time step may be determined randomly by the learning device (100), but the present invention is not limited thereto and may be determined by various methods such as settings by a user.
[0049] Next, a reverse process for removing noise from the learning noisy image latency (14) can be performed, and first, the input required for the denoising network (310) of the diffusion model (300) can be obtained through the following process.
[0050] First, a learning device (100) can obtain a learning semantic mask (11) and a learning image caption (12) corresponding to a learning image (10). Here, the learning semantic mask (11) may be an RGB image in which unique colors are assigned to each learning class included in the learning image (10), as shown in (a), (b), and (c) of FIG. 8. Additionally, the learning image caption (12) may include a learning instance-level text for the learning class corresponding to the learning semantic mask (11) and a learning global-level text corresponding to the learning image (10). For example, the learning instance-level text may be "car", "dog", "person", and the learning global-level text may be "a dog and a person are walking next to a car".
[0051] And, the learning device (100) inputs a learning semantic mask (11) into the encoder of the VAE (330) to cause the encoder of the VAE (330) to encode the learning semantic mask (11) to generate a learning semantic mask latency (13), and generates learning control information using the learning noisy image latency (15) and the learning semantic mask latency (13); inputs a learning image caption (12) into a text encoder (340) to cause the text encoder (340) to encode the learning image caption (12) to generate a learning text embedding (16) including a learning instance-level text embedding and a learning global-level text embedding; and inputs a learning time step into a time step encoder (350) to cause the time step encoder (350) to encode the learning time step to generate a learning time step A process can be performed to generate an embedding (17).
[0052] Meanwhile, although the above description explains that the learning device (100) acquires a learning image (10) and a learning time step for a forward process and acquires a learning semantic mask (11) and a learning image caption (12) for a reverse process, the present invention is not limited thereto, and the learning device (100) may acquire the learning image (10), the learning time step, the learning semantic mask (11), and the learning image caption (12), and then perform a forward process and a reverse process.
[0053] Afterwards, the learning device (100) can generate learning control information (15) by concatenating the learning noisy image latency (14) and the learning semantic mask latency (13) by channel.
[0054] Here, the learning control information (15) may be a latency in which the learning noisy image latency (14) and the learning semantic mask latency (13) are concatenated.
[0055] Additionally, by concatenating the learning noisy image latency (14) and the learning semantic mask latency (13) to correspond to the increased number of channels, zero-initial weights equal to the increased number of channels are added to the first layer of the denoising network (310), thereby enabling the denoising network (310) to process the learning control information (15).
[0056] And, the learning device (100) inputs the learning time step embedding (17), the learning text embedding (16), and the learning control information (15) into the denoising network (310) so that the denoising network (310) generates the learning prediction noise (18) by referring to the learning control information (15) according to the learning time step embedding (17), and the scheduler (320) can generate the learning synthetic image latency (19) by removing noise from the learning control information (15) by referring to the learning prediction noise. At this time, the learning prediction noise may be the prediction of the restored noise, that is, the noise that is restored according to the learning time step embedding, that is, the learning control information (15) or the learning intermediate synthetic image latency.
[0057] Specifically, the learning device (100) causes the denoising network (310) to generate learning intermediate prediction noise by referring to learning control information (15), inputs the learning intermediate prediction noise into the scheduler (320) so that the scheduler (320) removes noise from the learning control information (15) by referring to the learning intermediate prediction noise to generate learning intermediate synthetic image latency, and inputs the learning intermediate synthetic image latency again as learning control information into the denoising network (310) so that the denoising network (310) generates learning intermediate prediction noise by referring to the learning intermediate synthetic image latency, and inputs the learning intermediate prediction noise into the scheduler (320) so that the scheduler (320) removes noise by referring to the learning intermediate prediction noise to generate learning intermediate By repeating a denoising process that removes noise from intermediate synthetic image latency to generate a learning intermediate synthetic image latency corresponding to a learning time step, a learning synthetic image latency (19) can be generated. Here, when the learning device (100) inputs the learning intermediate synthetic image latency into the denoising network (310), it may concatenate the learning intermediate synthetic image latency and the learning semantic mask latency (13) to generate updated learning control information (15) and input the updated learning control information (15) into the denoising network (310).
[0058] For reference, since the reverse process is a process that removes noise unlike the forward process, if the scheduler (320) uses 50 steps as a time step for learning, the scheduler (320) can generate a schedule sequence containing 50 numbers such as [980, 960, 20, 0] and remove noise according to the predefined values included in the variation schedule.
[0059] When a synthetic image latency (19) for training is generated by the reverse process as described above, the training device (100) inputs the synthetic image latency (19) for training into the decoder of the VAE (330) to cause the decoder of the VAE (330) to decode the synthetic image latency (19) for training to generate a synthetic image for training (30), and can train a denoising network (310) to minimize the loss generated by referencing the synthetic image for training (30) and the training image (10). At this time, unlike training the denoising network (310) using the loss generated by referencing the synthetic image for training (30) and the training image (10), the denoising network (310) may also be trained using the loss generated by referencing the noise added at each time step and the predicted noise. However, the present invention is not limited thereto, and by using various loss functions, the denoising network (310) can be trained to predict noise from the previous time step from the noise at each time step. Through such training of the denoising network (310), the diffusion model (300) becomes able to recognize a semantic mask having fine-grained information.
[0060] Meanwhile, in the above, learning control information (15) was generated by concatenating the learning noisy image latency (14) and the learning semantic mask latency (13). Alternatively, learning control information (15) may be generated using a ControlNet, and this is explained as follows with reference to FIG. 3b. Below, detailed explanations will be omitted for parts that can be easily understood from the explanation with reference to FIG. 3a.
[0061] First, the learning device (100) inputs the learning noisy image latency (14), the learning semantic mask latency (13), the learning text embedding (16), and the learning time step embedding (17) into the control net (360) to cause the control net (360) to generate a control signal (not shown), and can generate learning control information including the control signal (not shown) and the learning noisy image latency (14). At this time, the control net (360) may be created by copying some of the layers of the pre-trained diffusion model (300) and adding zero convolution layers to the first and last layers of the control net (360).
[0062] As described above, the learning device (100) inputs the learning time step embedding (17), the learning noisy image latency (14), the learning semantic mask latency (13), the learning text embedding (16), and the learning time step embedding (17) into the control net (360) to cause the control net (360) to generate a control signal (not shown), and then inputs the learning time step embedding (17), the learning text embedding (16), and the control signal (not shown) into the denoising network (310) to cause the denoising network (310) to generate learning prediction noise by referring to the learning control information, namely the control signal (not shown) and the learning noisy image latency (14), according to the learning time step embedding (17), and the scheduler (320) removes noise from the noisy image latency (14) by referring to the learning prediction noise to produce a learning synthetic image Latency (19) can be generated. In this case, the learning prediction noise may be a prediction of the restored noise that is restored according to the time step of the input noise, that is, the learning noisy image latency (14) or the learning intermediate synthetic image latency.
[0063] Specifically, the learning device (100) causes the denoising network (310) to generate learning intermediate prediction noise by referencing learning control information, namely (i) learning noisy image latency (14) and (ii) a control signal (not shown) output from the control network (360); inputs the learning intermediate prediction noise into the scheduler (320) so that the scheduler (320) generates a learning intermediate synthetic image latency (19) by removing noise from the learning noisy image latency (14) included in the learning control information by referencing the learning intermediate prediction noise; and inputs the learning intermediate synthetic image latency again as learning control information into the denoising network (310) so that the denoising network (310) generates learning intermediate prediction noise by referencing the learning intermediate synthetic image latency. By inputting the process and the learning intermediate prediction noise into the scheduler (320) and repeating the denoising process so that the scheduler (320) removes noise from the learning intermediate synthetic image latency by referencing the learning intermediate prediction noise to generate a learning intermediate synthetic image latency corresponding to the learning time step, a learning synthetic image latency (19) can be generated. Here, when the learning device (100) inputs the learning intermediate synthetic image latency into the denoising network (310), it may generate updated learning control information including the learning intermediate synthetic image latency and the updated control signal output by inputting the learning intermediate synthetic image and the learning semantic mask latency (13) into the control network (360), and input the updated learning control information into the denoising network (310).
[0064] When a synthetic image latency (19) for training is generated by the reverse process as described above, a denoising network (310) can be trained as described with reference to FIG. 3a.
[0065] For reference, the control net (360) can be trained in a manner similar to the denoising network (310). That is, it can be trained to minimize the loss generated by referencing the training synthetic image (30) and the training image (10). Additionally, using various loss functions, the control net (360) can be trained to generate a control signal (not shown) by predicting the noise from the previous time step from the noise at each time step. Through such training, the control net (360) can generate a control signal (not shown) that follows both the training text embedding (16) and the training semantic mask latency (13) by minimizing the error in the predicted noise.
[0066] With the diffusion model (300) trained by the learning method described above, a test can be performed to generate a synthetic image reflecting a fine-grain instance layout using the diffusion model (300).
[0067] FIG. 4 illustrates a schematic configuration of a test device for testing a controllable image generation model capable of reflecting a fine-grain instance layout using a diffusion model (300) learned according to an embodiment of the present invention. Referring to FIG. 4, the test device (200) may include a memory (210) in which instructions for testing the diffusion model (300), which is a controllable image generation model capable of reflecting a fine-grain instance layout, are stored, and a processor (220) that performs an operation to test the diffusion model (300) according to the instructions stored in the memory (210). At this time, although the diffusion model (300) is shown as being installed in the test device (200), alternatively, the diffusion model (300) may be installed in a cloud environment or installed in a computing device different from the test device (200).
[0068] Specifically, the test device (200) may achieve desired system performance by utilizing a combination of a computing device (e.g., a device that may include components of a computer processor, memory, storage, input device and output device, and other conventional computing devices; an electronic communication device such as a router, switch, etc.; an electronic information storage system such as a Network Attached Storage (NAS) and a Storage Area Network (SAN)) and computer software (i.e., instructions that cause the computing device to function in a specific way), but is not limited thereto.
[0069] Additionally, the processor (220) of the test device (200) may include hardware configurations such as an MPU (Micro Processing Unit) or CPU (Central Processing Unit), cache memory, and data bus. Additionally, the computing device may further include software configurations such as an operating system and an application for a specific purpose.
[0070] However, this does not exclude the case where the test device (200) includes an integrated processor in which a medium, a processor, and a memory are integrated for carrying out the present invention.
[0071] A method for testing a controllable image generation model capable of reflecting a fine-grain instance layout according to one embodiment of the present invention using a test device configured as described above is explained below with reference to FIGS. 5 and 6.
[0072] With the diffusion model (300) trained to recognize a fine-grain instance layout, i.e., a fine-grain semantic mask, by the learning method described with reference to FIGS. 3a and 3b, the test device (200) can acquire a test noisy latency (20), a test semantic mask (21), a corresponding test image caption (22), and a test time step. Here, the test semantic mask (21) may be an RGB image in which unique colors are assigned to each test class included in the test synthetic image (31) to be generated as an image such as (a), (b), and (c) shown in FIG. 8. Additionally, the test image caption (22) may include test instance-level text for the test class corresponding to the test semantic mask (21) and test global-level text corresponding to the test semantic mask (21). For example, the test instance-level text may be "car", "dog", or "person", and the test global-level text may be "a dog and a person are walking next to a car". Additionally, the test noisy latency (20) may be generated using a random seed, but is not limited thereto. Furthermore, the test time step may be randomly determined by the test device (200), but the present invention is not limited thereto and may be determined by various methods such as user settings. At this time, the test time step represents the maximum number of iterations for removing noise, but is not limited thereto.
[0073] Subsequently, the test device (200) may perform the following steps: generating a test semantic noisy latency (25) using a test semantic mask (21) and a test noisy latency (20); inputting a test image caption (22) into a text encoder (340) to cause the text encoder (340) to encode the test image caption (22) to generate a test text embedding (26) including a test instance-level text embedding and a test global-level text embedding; and inputting a test time step into a time step encoder (350) to cause the time step encoder (350) to encode the test time step to generate a test time step embedding (27).
[0074] Here, the test device (200) can convert the test semantic mask to the same resolution as the test noisy latency (20), and then perform a 1:1 mapping operation with the test noisy latency (20) to generate the test semantic noisy latency (25).
[0075] Next, the test device (200) inputs the test semantic noisy latency (25), the test text embedding (26), and the test time step embedding (27) into the denoising network (310) to cause the denoising network (310) to generate test prediction noise (28) by referencing the test semantic noisy latency (25), and the scheduler (320) may generate test synthetic image latency (29) by removing noise from the test semantic noisy latency (25) by referencing the test prediction noise (28). At this time, the test prediction noise may be a prediction of the restored noise, i.e., the test semantic noisy latency (25) or the test intermediate synthetic image latency, which is restored according to the test time step embedding.
[0076] Specifically, the test device (200) has a denoising network (310) to generate test intermediate prediction noise by referencing the test semantic noisy latency (25), input the test intermediate prediction noise into a scheduler (320) to have the scheduler (320) remove noise from the test semantic noisy latency (25) by referencing the test intermediate prediction noise to generate a test intermediate synthetic image latency, and input the test intermediate synthetic image latency as the test semantic noisy latency (25) into the denoising network (310) to have the denoising network (310) generate test intermediate prediction noise by referencing the test intermediate synthetic image latency, and input the test intermediate prediction noise into the scheduler (320). By having the scheduler (320) repeat a denoising process to remove noise from the test intermediate synthetic image latency by referring to the test intermediate prediction noise and to generate a test intermediate synthetic image latency corresponding to the test time step, a test synthetic image latency (31) can be generated. Here, when the test device (200) inputs the test intermediate synthetic image latency into the denoising network (310), it may input the updated test semantic noisy latency (25) into the denoising network (310) by mapping the test intermediate synthetic image latency and the test semantic mask latency (23) in a 1:1 ratio.
[0077] At this time, an attention mask can be applied by referencing the test instance-level text embedding and the test global-level text embedding in the attention operation included in the denoising network (310) so that the layout shown in the test semantic mask, such as (a), (b), and (c) shown in FIG. 8, provided as control information used to generate a test synthetic image, is well reflected.
[0078] Specifically, referring to FIG. 7, the test device (200) repeats the prediction noise generation process and the denoising process as described above, wherein, according to the test time step, (i) in the initial denoising process in which the prediction noise generation process and the denoising process are repeated up to the k-th time, the output calculation operation of each layer included in the denoising network includes (i-1) an attention operation for parts corresponding to the instance level mask generated by mapping the test instance level text embedding and the test noisy semantic mask latency or the test instance level text embedding and the test intermediate synthetic image latency, and (i-2) an attention operation for the test global level text embedding and the test noisy semantic mask latency or the test global level text embedding and the test intermediate synthetic image latency, and (ii) in the subsequent denoising process in which the prediction noise generation process and the denoising process are repeated after the k-th time, the output calculation operation of each layer included in the denoising network includes the test global level text embedding and By including attention operations for test noisy semantic mask latency or test global-level text embeddings and test intermediate synthetic image latency, the reflection of the semantic mask at the instance level can also be finely grained.
[0079] That is, the denoising network (310) can perform various attention operations. It can organize how each part of the image is connected to each other during the process of generating a synthetic image from noise by performing a self-attention operation, and can control it according to external conditions such as text captions by performing a cross-attention operation. FIG. 7 exemplarily illustrates the representation of the process of the self-attention operation included among the various operations performed by the denoising network (310) to predict noise. Specifically, FIG. 7 shows a mapped noisy latency space expressed in a 2D form generated according to the method of the present invention. As shown in FIG. 7 (a) initial denoising step, initially, a self-attention operation is performed by applying an attention mask to parts corresponding to the instance-level mask in the noisy semantic mask latency obtained by mapping the test semantic mask latency and the test noisy latency, and a test global-level text embedding and the test noisy semantic mask latency or a test global-level text embedding and the test intermediate synthetic image. Although a self-attention operation that applies an attention mask to the latency is included, in the later stage, as illustrated in the later denoising step (b) of FIG. 7, only a self-attention operation that applies an attention mask to the test global level text embedding, the test noisy semantic mask latency, and the test intermediate synthetic image latency may be performed.
[0080] Here, k, which divides the initial denoising phase and the later denoising phase, may be an integer greater than or equal to 1 or less than or equal to a test time step representing a repeating process and may be determined by the user, but is not limited thereto.
[0081] After generating a test synthetic image latency in the manner described above, the test device (200) can input the test synthetic image latency (29) into the decoder of the VAE (330) to cause the decoder of the VAE (330) to decode the test synthetic image latency (29) and generate a test synthetic image (31).
[0082] The embodiments according to the present invention described above may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the computer-readable recording medium may be those specifically designed and configured for the present invention, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware device may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.
[0083] Although the present invention has been described above with specific details such as specific components, limited embodiments, and drawings, this is provided only to aid in a more comprehensive understanding of the invention, and the invention is not limited to the above embodiments, and a person skilled in the art to which the invention belongs can make various modifications and variations from this description.
[0084] Accordingly, the scope of the present invention should not be limited to the embodiments described above, and all modifications equivalent to or equivalent to the claims set forth below, as well as the claims described below, shall be considered to fall within the scope of the concept of the present invention.
Claims
Claim 1 A method for training a controllable image generation model capable of reflecting a fine-grained instance layout, comprising: (a) when a training image and a training time step are acquired, a training device inputs the training image into an encoder of a Variational Autoencoder (VAE) to cause the encoder of the VAE to generate a training image latent, and generates a training noisy image latent by repeatedly adding noise to the training image latent according to the training time step through a scheduler;(b) When the learning device acquires a learning semantic mask corresponding to the learning image and a learning image caption—the learning image caption includes a learning instance-level text for a learning class corresponding to the learning semantic mask and a learning global-level text corresponding to the learning image—the learning semantic mask is input to the encoder of the VAE to cause the encoder of the VAE to generate a learning semantic mask latent, and the learning control information is generated using the learning noisy image latent and the learning semantic mask latent; the learning image caption is input to a text encoder to cause the text encoder to generate a learning text embedding including a learning instance-level text embedding and a learning global-level text embedding; and the learning A step of performing a process of inputting a time step into a time step encoder to cause the time step encoder to encode the training time step and generate a training time step embedding;(c) A step in which the learning device inputs the learning time-step embedding, the learning text embedding, and the learning control information into a denoising network, causing the denoising network to generate learning prediction noise by referencing the learning control information, and causes the scheduler to generate a learning synthetic image latency by removing noise from the learning control information by referencing the learning prediction noise, wherein the denoising network generates learning intermediate prediction noise by referencing the learning intermediate synthetic image latency according to the learning time-step embedding, and the denoising process in which the scheduler generates the learning intermediate synthetic image latency from which noise has been removed by referencing the learning intermediate prediction noise are repeated, thereby generating the learning synthetic image latency; and (d) a step comprising: the learning device inputting the learning synthetic image latency to the decoder of the VAE to cause the decoder of the VAE to decode the learning synthetic image latency to generate a learning synthetic image, and training the denoising network to minimize the loss generated by referencing the learning synthetic image and the learning image; Claim 2 In claim 1, in step (b), the learning device concatenates the learning noisy image latency and the learning semantic mask latency channel by channel to generate the learning control information. Claim 3 In claim 2, in step (c), the learning device performs the prediction noise generation process by adding zero-initial weights corresponding to the increased number of channels to the first layer of the denoising network in correspondence with the increased number of channels as the learning noisy image latency and the learning semantic mask latency are concatenated. Claim 4 In claim 1, in step (b), the learning device inputs the learning noisy image latency, the learning semantic mask latency, the learning text embedding, and the learning timestep embedding into a ControlNet to cause the ControlNet to generate a Control Signal, and a method for generating learning control information including the Control Signal and the learning noisy image latency. Claim 5 In claim 4, the control net is created by copying some of the layers of a pre-trained diffusion model and adding zero convolution layers to the first and last layers of the control net. Claim 6 In claim 1, the method wherein the learning semantic mask is an RGB image in which unique colors are assigned to each learning class included in the learning image. Claim 7 A method for testing a controllable image generation model capable of reflecting a fine-grained instance layout, wherein (a) by a training device, (i) when a training image and a training time step are acquired, the training image is input into an encoder of a Variational Autoencoder (VAE) to cause the encoder of the VAE to generate a training image latent, and a scheduler is used to repeatedly add noise to the training image latent according to the training time step to generate a training noisy image latent; (ii) a training semantic mask corresponding to the training image and a training image caption—the training image caption includes training instance-level text for a training class corresponding to the training semantic mask and training global-level text corresponding to the training image. - When this is obtained, the process of inputting the training semantic mask into the encoder of the VAE to cause the encoder of the VAE to generate a training semantic mask latency, and generating training control information using the training noisy image latency and the training semantic mask latency; the process of inputting the training image caption into a text encoder to cause the text encoder to generate a training text embedding including a training instance-level text embedding and a training global-level text embedding;and a subprocess including a process of inputting the training time step into a time step encoder to cause the time step encoder to encode the training time step to generate a training time step embedding; (iii) a subprocess including inputting the training time step embedding, the training text embedding, and the training control information into a denoising network to cause the denoising network to generate training prediction noise by referencing the training control information, and causing the scheduler to generate a training synthetic image latency by removing noise from the training control information by referencing the training prediction noise, wherein the prediction noise generation process in which the denoising network generates training intermediate prediction noise according to the time step embedding and the denoising process in which the scheduler generates a training intermediate synthetic image latency from which noise has been removed by referencing the training intermediate prediction noise are repeated, thereby generating the training synthetic image latency; and (iv) A sub-process is performed to input the above-mentioned training synthetic image latency into the decoder of the VAE so that the decoder of the VAE decodes the above-mentioned training synthetic image latency to generate a training synthetic image, and to train the denoising network to minimize the loss generated by referencing the training synthetic image and the training image, so that the denoising model is trained, and the test device, test noisy latency, test semantic mask,(b) a step of obtaining a test image caption corresponding thereto—the test image caption includes test instance-level text for a test class corresponding to the test semantic mask and test global-level text corresponding to the test semantic mask—and a test time step; (b) a step in which the test device generates a test semantic noisy latency using the test semantic mask and the test noisy latency, inputs the test image caption to the text encoder to cause the text encoder to generate a test text embedding including a test instance-level text embedding and a test global-level text embedding, and inputs the test time step to the time step encoder to cause the time step encoder to encode the test time step to generate a test time step embedding. A step of performing a process; (c) the test device inputs the test semantic noisy latency, the test text embedding, and the test timestep embedding into the denoising network to cause the denoising network to generate test prediction noise by referencing the test semantic noisy latency, and causes the scheduler to generate test synthetic image latency by removing noise from the test semantic noisy latency by referencing the test prediction noise, whereinA method comprising: a step of generating a test synthetic image latency by repeating the prediction noise generation process in which the denoising network generates test intermediate prediction noise by referencing the test intermediate synthetic image latency according to the test timestep embedding, and the denoising process in which the scheduler generates the test intermediate synthetic image latency from which noise has been removed by referencing the test intermediate prediction noise; and (d) a step in which the test device inputs the test synthetic image latency to the decoder of the VAE, causing the decoder of the VAE to decode the test synthetic image latency to generate a test synthetic image. Claim 8 In claim 7, in step (b) above, the test device converts the test semantic mask to the same resolution as the test noisy latency, and then performs a 1:1 mapping operation with the test noisy latency to generate the test semantic noisy latency. Claim 9 In claim 8, in step (c) above, the test device repeats the predictive noise generation process and the denoising process, wherein, according to the test time step, (i) in the initial denoising process in which the predictive noise generation process and the denoising process are repeated up to k - where k is a preset integer greater than or equal to 1 - the output calculation operation of each layer included in the denoising network includes (i-1) an attention operation for parts corresponding to an instance-level mask generated by mapping the test instance-level text embedding and the test noisy semantic mask latency or the test instance-level text embedding and the test intermediate synthetic image latency, and (i-2) an attention operation for the test global-level text embedding and the test noisy semantic mask latency or the test global-level text embedding and the test intermediate synthetic image latency, and (ii) in the subsequent denoising process in which the predictive noise generation process and the denoising process are repeated after k times A method in which the output calculation operation of each layer included in the above denoising network includes an attention operation for the test global level text embedding and the test noisy semantic mask latency or the test global level text embedding and the test intermediate synthetic image latency. Claim 10 In claim 7, the above test semantic mask is an RGB image in which unique colors are assigned to each test class to be generated. Claim 11 A training device for a controllable image generation model capable of reflecting a fine-grained instance layout, comprising: one or more memories for storing instructions; and includes one or more processors configured to execute the above instructions, wherein the processor comprises: (I) a process in which, when a training image and a training time step are acquired, the training image is input into an encoder of a Variational Autoencoder (VAE) to cause the encoder of the VAE to generate a training image latency, and a process in which noise is repeatedly added to the training image latency according to the training time step through a scheduler to generate a training noisy image latency; (II) a training semantic mask corresponding to the training image and a training image caption—the training image caption includes training instance-level text for a training class corresponding to the training semantic mask and training global-level text corresponding to the training image—when these are acquired, the training semantic mask is input into the encoder of the VAE to the VAE A process of causing an encoder to generate a learning semantic mask latency, and generating learning control information using the learning noisy image latency and the learning semantic mask latency,A process of inputting the above-mentioned training image caption into a text encoder to cause the text encoder to generate a training text embedding including a training instance-level text embedding and a training global-level text embedding, and a process of inputting the above-mentioned training time step into a time step encoder to cause the time step encoder to encode the above-mentioned training time step to generate a training time step embedding, (III) inputting the above-mentioned training time step embedding, the above-mentioned training text embedding, and the above-mentioned training control information into a denoising network to cause the denoising network to generate training prediction noise by referencing the above-mentioned training control information, and causing the scheduler to generate a training synthetic image latency by removing noise from the above-mentioned training control information by referencing the above-mentioned training prediction noise, wherein the denoising according to the above-mentioned training time step embedding A process for generating a learning synthetic image latency by repeating a prediction noise generation process in which a network references a learning intermediate synthetic image latency to generate learning intermediate prediction noise, and a denoising process in which a scheduler references the learning intermediate prediction noise to generate the learning intermediate synthetic image latency from which noise has been removed; and (IV) a learning device that inputs the learning synthetic image latency to a decoder of the VAE to cause the decoder of the VAE to decode the learning synthetic image latency to generate a learning synthetic image, and trains the denoising network to minimize the loss generated by referencing the learning synthetic image and the learning image. Claim 12 In claim 11, the processor is a learning device that generates learning control information by concatenating the learning noisy image latency and the learning semantic mask latency by channel in the process (II). Claim 13 In claim 12, the processor is a learning device that, in the process (III), adds zero-initial weights corresponding to the increased number of channels to the first layer of the denoising network in correspondence with the increased number of channels as the learning noisy image latency and the learning semantic mask latency are concatenated, thereby enabling the prediction noise generation process to be performed. Claim 14 In claim 11, the processor inputs the learning noisy image latency, the learning semantic mask latency, the learning text embedding, and the learning timestep embedding into a ControlNet in the process (II) to cause the ControlNet to generate a Control Signal, and a learning device that generates learning control information including the Control Signal and the learning noisy image latency. Claim 15 In claim 14, the control net is a learning device created by copying some of the layers of a pre-trained diffusion model and adding zero convolution layers to the first and last layers of the control net. Claim 16 In claim 11, the above-mentioned learning semantic mask is a learning device in which each unique color is assigned to each of the above-mentioned learning classes included in the above-mentioned learning image. Claim 17 A test device for a controllable image generation model capable of reflecting a fine-grained instance layout comprises: one or more memories storing instructions; and one or more processors configured to execute said instructions, wherein the processor comprises: (I) a sub-process in which, when a training image and a training time step are acquired by a training device, (i) input the training image into an encoder of a Variational Autoencoder (VAE) to cause the encoder of the VAE to generate a training image latent, and in which noise is repeatedly added to the training image latent according to the training time step through a scheduler to generate a training noisy image latent; (ii) a training semantic mask corresponding to the training image and a training When an image caption—the above training image caption includes training instance-level text for a training class corresponding to the above training semantic mask and training global-level text corresponding to the above training image—is obtained, the training semantic mask is input into the encoder of the above VAE to cause the encoder of the above VAE to generate a training semantic mask latency, and a process of generating training control information using the training noisy image latency and the training semantic mask latency,A subprocess comprising: a process of inputting the above-mentioned training image caption into a text encoder to cause the text encoder to generate a training text embedding including a training instance-level text embedding and a training global-level text embedding; and a process of inputting the above-mentioned training time step into a time step encoder to cause the time step encoder to encode the above-mentioned training time step to generate a training time step embedding; (iii) inputting the above-mentioned training time step embedding, the above-mentioned training text embedding, and the above-mentioned training control information into a denoising network to cause the denoising network to generate training prediction noise by referencing the above-mentioned training control information, and causing the scheduler to generate a training synthetic image latency by removing noise from the above-mentioned training control information by referencing the above-mentioned training prediction noise, wherein the denoising network according to the above-mentioned time step embedding A sub-process that generates a prediction noise generation process for generating training intermediate prediction noise and a denoising process in which the scheduler refers to the training intermediate prediction noise to generate a noise-removed training intermediate synthetic image latency, thereby generating the training synthetic image latency; and (iv) a sub-process that inputs the training synthetic image latency to the decoder of the VAE to cause the decoder of the VAE to decode the training synthetic image latency to generate a training synthetic image, and trains the denoising network to minimize the loss generated by referring to the training synthetic image and the training image, so that the denoising model is trained, and therein thereis a test noisy latency, a test semantic mask,A processor that obtains a test image caption corresponding thereto—the test image caption includes test instance-level text for a test class corresponding to the test semantic mask and test global-level text corresponding to the test semantic mask—and a test time step; (II) a process of generating a test semantic noisy latency using the test semantic mask and the test noisy latency; a process of inputting the test image caption into the text encoder to cause the text encoder to generate a test text embedding including a test instance-level text embedding and a test global-level text embedding; and a process of inputting the test time step into the time step encoder to cause the time step encoder to encode the test time step to generate a test time step embedding. (III) Input the above test semantic noisy latency, the above test text embedding, and the above test time step embedding into the denoising network so that the denoising network generates test prediction noise by referencing the above test semantic noisy latency, and the scheduler generates test synthetic image latency by removing noise from the above test semantic noisy latency by referencing the above test prediction noise, whereinA process for generating a test synthetic image latency by repeating the prediction noise generation process in which the denoising network generates test intermediate prediction noise by referencing the test intermediate synthetic image latency according to the test time step embedding, and the denoising process in which the scheduler generates the test intermediate synthetic image latency from which noise has been removed by referencing the test intermediate prediction noise; and (IV) a test device that performs a process of inputting the test synthetic image latency into the decoder of the VAE and causing the decoder of the VAE to decode the test synthetic image latency to generate a test synthetic image. Claim 18 In claim 17, the processor is a test device that, in the process (II), converts the test semantic mask to the same resolution as the test noisy latency and then performs a 1:1 mapping operation with the test noisy latency to generate the test semantic noisy latency. Claim 19 In claim 18, the processor, in the process (III), repeats the predictive noise generation process and the denoising process, wherein, according to the test time step, (i) in the initial denoising process in which the predictive noise generation process and the denoising process are repeated up to k - where k is a preset integer greater than or equal to 1 - the output calculation operation of each layer included in the denoising network comprises (i-1) an attention operation for parts corresponding to an instance-level mask generated by mapping the test instance-level text embedding and the test noisy semantic mask latency or the test instance-level text embedding and the test intermediate synthetic image latency, and (i-2) an attention operation for the test global-level text embedding and the test noisy semantic mask latency or the test global-level text embedding and the test intermediate synthetic image latency, and (ii) a subsequent process in which the predictive noise generation process and the denoising process are repeated after k. A test device in which the output calculation operation of each layer included in the denoising network during the denoising process includes an attention operation for the test global level text embedding and the test noisy semantic mask latency or the test global level text embedding and the test intermediate synthetic image latency. Claim 20 In claim 17, the above test semantic mask is a test device that is an RGB image in which unique colors are assigned to each test class to be generated.
Citation Information
Patent Citations
Method for semantic image synthesis using condition diffusion and apparatus for same
KR1020240134643A
Method and apparatus for image processing based on neural diffusion
KR1020250042426A
System and method for controllable text-to-3d room mesh generation with layout constraints
US20250131656A1
Segmentation free guidance in diffusion models
US20250166236A1