A method and system for generative imaging
The method integrates image and text data processing using a single trained model to encode and denoise features, addressing inefficiencies in conventional multi-model systems by maintaining context and reducing complexity in image editing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-11-05
- Publication Date
- 2026-05-15
AI Technical Summary
Conventional image editing methods require multiple models to analyze image and text data separately, leading to a loss of context and increased complexity, making the process inefficient and time-consuming.
A method and system that uses a trained model to encode input images with predefined noise, identify alteration and non-alteration features based on text prompts, and generate output images by denoising these features, integrating image and text data processing in a single model.
Enhances efficiency and robustness in image editing by maintaining context and reducing complexity through a unified model that processes both image and text data simultaneously.
Smart Images

Figure KR2025018027_15052026_PF_FP_ABST
Abstract
Description
A METHOD AND SYSTEM FOR GENERATIVE IMAGING
[0001] The present disclosure relates to imaging. More particularly, the present disclosure relates to a method for generative imaging and a system thereof.
[0002] Recently, various methods for image editing are emerging as the need for image editing is increasing. Image editing has various advantages such as brand building, photo-intensive tasks becoming easier, robust social media strategy, visual story telling and the like.
[0003] Generally, for editing images, an image and text data are provided as inputs from a user. The image is edited based on the text data. However, in conventional techniques for editing images, multiple models are used. For example, a model may be used to mask regions of interest i.e., the regions to be edited in the image. In an embodiment, the user may mask the regions to be edited and provide it to the system. Thus, making the process of generating an output image time consuming. Another model may be used to analyze the region of interest based on the text data. Editing images using multiple models is not robust or efficient, as the image and the text data being analyzed by different models separately may lead to a loss of the context of the text data. Using multiple models also increases the complexity in integrating the models to obtain the edited image. Therefore, there is a need for a method and a system to overcome the aforementioned problems.
[0004] The information disclosed in this background of the disclosure section is only for enhancement of understanding of the general background of the invention and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art.
[0005] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.
[0006] In an embodiment, a method for generative imaging is disclosed. The method includes receiving an input image and text data. In one exemplary aspect, the text data comprises one or more prompts for generating an output image. The method includes identifying the one or more prompts in the text data. The method includes encoding the identified one or more prompts for generating text embeddings, using a trained model. The method includes encoding the input image by including a predefined noise, using the trained model. The method includes identifying, using the trained model, one or more alteration features and one or more non-alteration features based on the identified one or more prompts and the input image. The method includes generating, using the trained model, the output image by denoising the encoded input image based on the one or more alteration features and the one or more non-alteration features.
[0007] In an embodiment, a system for generative imaging is disclosed. The system includes a memory that stores processor-executable instructions. The system includes a processor configured to execute the processor-executable instructions stored in the memory and thereby configured to receive an input image and text data. In one exemplary aspect, the text data comprises one or more prompts for generating an output image. The processor is configured to identify the one or more prompts in the text data. The processor is configured to encode the identified one or more prompts for generating text embeddings, using a trained model. The processor is configured to encode the input image by including a predefined noise, using the trained model. The processor is configured to identify, using the trained model, one or more alteration features and one or more non-alteration features based on the identified one or more prompts and the input image. The processor is configured to generate, using the trained model, the output image by denoising the encoded input image based on the one or more alteration features and the one or more non-alteration features.
[0008] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles. The same numbers are used throughout the figures to reference features and components. Some embodiments of at least one of device and methods in accordance with embodiments of the present subject matter are now described, by way of example only, and with reference to the accompanying figures, in which:
[0009] Fig. 1a-b illustrate an environment in which some embodiments of the present disclosure may be practiced;
[0010] Fig. 2 illustrates a system for generative imaging, in accordance with an embodiment of the present disclosure;
[0011] Fig. 3 illustrates an exemplary depiction of generating text embeddings, in accordance with an embodiment of the present disclosure;
[0012] Fig. 4 illustrates an exemplary depiction of encoding an input image, in accordance with an embodiment of the present disclosure;
[0013] Fig. 5 illustrates an exemplary depiction of identifying one or more alteration features and one or more non-alteration features, in accordance with an embodiment of the present disclosure;
[0014] Fig. 6 illustrates identifying one or more alteration features and one or more non-alteration features, in accordance with another embodiment of the present disclosure;
[0015] Fig. 7 illustrates an exemplary depiction of generative imaging, in accordance with an embodiment of the present disclosure;
[0016] Fig. 8 illustrates an exemplary depiction of generative imaging, in accordance with another embodiment of the present disclosure;
[0017] Fig. 9 illustrates a trained model for generative imaging, in accordance with an embodiment of the present disclosure;
[0018] Fig. 10a-c illustrate exemplary depictions of generative imaging, in accordance with an embodiment of the present disclosure;
[0019] Fig. 11 illustrates a flow chart of a method of generative imaging, in accordance with an embodiment of the present disclosure; and
[0020] Fig. 12 illustrates a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.
[0021] Fig. 13 illustrates an exemplary depiction of generative imaging, in accordance with an embodiment of the present disclosure.
[0022] Fig. 14 illustrates an exemplary depiction of generative imaging, in accordance with an embodiment of the present disclosure.
[0023] The figures depict embodiments of the disclosure for purposes of illustration only. One skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles of the disclosure described herein.
[0024] In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.
[0025] While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the spirit and the scope of the disclosure.
[0026] The terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a device or system or apparatus proceeded by "comprises... a" does not, without more constraints, preclude the existence of other elements or additional elements in the device or system or apparatus.
[0027] In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense.
[0028] Fig. 1a illustrates an environment 100a in which some embodiments of the present disclosure may be practiced.
[0029] The environment 100a exemplarily depicts a system 102, an input image 104, text data 106 and an output image 108. The input image 104 and the text data 106 are provided to the system 102. The input image 104 is an image to be edited and the text data 106 comprises information essential to edit the image. The text data 106 may comprise one or more prompts to edit the input image 104. The system 102 analyzes the text data 106 and identifies regions to be edited in the input image 104. In an embodiment, a user device may provide the input image 104 and the text data 106 to the system 102.
[0030] Fig. 1b illustrates an exemplary environment 100b in which some embodiments of the present disclosure may be practiced. In Fig. 1b, the system 102, a user device 110, and a communication network 112 are disclosed. The user device 110 may establish a communication with the system 102 to generate images. Some examples of the user device 110 may include, but not limited to, electronic devices such as, a smartphone, a laptop, a desktop, a personal computer, or any spatial computing device capable of performing wireless communication. For example, the user device 110 may work on multiple platforms and / or Operating Systems to perform different operations related to wireless communication. The user device 110 transmits the input image 104 and text data 106 to the system 102, to generate the output image 108. Conventionally, the user device 110 transmits the input image 104 and the text data 106 to separate models and / or systems. The text data 106 is analyzed separately and the input image 104 is analyzed separately and later combined together to generate the output image 108. However, the conventional technique is complex and time consuming. Therefore, the system 102 is deployed to generate the output image 108 directly, by receiving both the input image 104 and the text data 106 from the user device 110.
[0031] The user device 110 may establish a connection with the system 102 via a communication network 112. It is understood that the user device 110 may be in operative communication with the communication network 112, such as the Internet, enabled by a network provider, also known as an Internet Service Provider (ISP). The user device 110 may be connected to the communication network 112 using a wireless network. Some non-limiting examples of wireless networks may include the Wireless LAN (WLAN), cellular networks, Bluetooth or ZigBee networks, and the like.
[0032] Various embodiments of the present disclosure disclose a method performed by the system 102 for generative imaging. The operations performed by the system 102 are explained in detail next with reference to Fig. 2.
[0033] Fig. 2 illustrates a system 102 for generative imaging, in accordance with an embodiment of the present disclosure. In an embodiment, the system 102 comprises a processor 202, a memory 204, an input / output module 206 and a communication interface 208. It shall be noted that, in some embodiments, the system 102 may include more or fewer components than those depicted herein. The various components of the system 102 may be implemented using hardware, software, firmware, or any combinations thereof. Further, the various components of the system 102 may be operably coupled with each other. More specifically, various components of the system 102 may be capable of communicating with each other using communication channel media (such as buses, interconnects, etc.).
[0034] In one embodiment, the processor 202 may be embodied as a multi-core processor, a single core processor, or a combination of one or more multi-core processors and one or more single core processors. For example, the processor 202 may be embodied as one or more of various processing devices, such as a coprocessor, a microprocessor, a controller, a digital signal processor (DSP), a processing circuitry with or without an accompanying DSP, or various other processing devices including, a microcontroller unit (MCU), a hardware accelerator, a special-purpose computer chip, or the like.
[0035] In an embodiment, the processor 202 may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU).
[0036] In an embodiment, the processor 202 is configured to: (1) receive the input image 104 and text data 106; the text data 106 comprises one or more prompts for generating the output image 108, (2) identify at least one of: one or more static prompts 302 (refer Fig. 3) and one or more dynamic prompts 304 from the one or more prompts in the text data, (3) encode the identified prompts for generating text embeddings 306, using a trained model, (4) encode the input image 104 (refer Fig. 4) by including a predefined noise, using the trained model, (5) identify, using the trained model, at least one of: one or more alteration features in the encoded input image 104 based on the one or more dynamic prompts 304, and one or more non-alteration features in the encoded input image 104 based on the one or more static prompts 302, and (6) generate, using the trained model, the output image 108 by transforming the one or more alteration features in the encoded input image 104 based on the one or more dynamic prompts 304.
[0037] In one embodiment, the memory 204 is capable of storing machine executable instructions, referred to herein as instructions 205 and a trained model 212. In an embodiment, the processor 202 is embodied as an executor of software instructions. As such, the processor 202 is capable of executing the instructions 205 stored in the memory 204 to perform one or more operations described herein.
[0038] In an embodiment, at least one of a plurality of modules of the system 102 may be implemented through an Artificial Intelligence (AI) model. A function associated with the AI may be performed through the non-volatile memory, the volatile memory, and the processor 202.
[0039] The term "trained model(s)" used herein refers to a model that has been trained on a set of data to recognize certain patterns and / or make certain decisions without further human intervention. The trained model(s) apply different algorithms to relevant data inputs to achieve the tasks and / or output they have been trained for. The "trained model(s)" may be a continuous learning model in some embodiments. The trained model(s) may be a deep learning model such as a deep attention model. An attention technique used in the deep attention model enhances the deep learning models by selectively focusing on important input elements, improving prediction accuracy and computational efficiency. The deep attention model prioritizes and emphasizes relevant information, acting as a spotlight to enhance overall model performance. In an embodiment, the model is trained using a large data set of images to identify various features in the images. In an embodiment, the model is trained based on a supervised learning. The trained model is trained based on a loss function.
[0040] The processor 202 controls the processing of the input data in accordance with a predefined operating rule or the AI model stored in the non-volatile memory and the volatile memory. The predefined operating rule or trained model 212 is provided through training or learning. Herein, being provided through learning means that, by applying a learning technique to learn data, a predefined operating rule or trained model of a desired characteristic is made. The learning may be performed in a device itself in which AI according to an embodiment is performed, and / or may be implemented through a separate server / system.
[0041] The trained model 212 may consist of a plurality of models (not shown). Each model may comprise a plurality of neural network layers. Each layer has a plurality of weight values and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Network (RNN), restricted Boltzmann Machine (RBM), Deep Belief Network (DBN), Bidirectional Recurrent Deep Neural Network (BRDNN), Generative Adversarial Networks (GAN), and deep Q-networks. The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning.
[0042] The memory 204 can be any type of storage accessible to the processor 202 to perform respective functionalities. For example, the memory 204 may include one or more volatile or non-volatile memories, or a combination thereof. For example, the memory 204 may be embodied as semiconductor memories, such as flash memory, mask ROM, PROM (programmable ROM), EPROM (erasable PROM), RAM (random access memory), etc. and the like.
[0043] In an embodiment, the I / O module 206 may include mechanisms configured to receive inputs from and provide outputs to operator of the system 102. To enable reception of inputs and provide outputs from the system 102, the I / O module 206 may include at least one input interface and / or at least one output interface. Some examples of the input interface may include, but are not limited to, a keyboard, a mouse, a joystick, a keypad, a touch screen, soft keys, a microphone, and the like. Examples of the output interface may include, but are not limited to, a display such as a light emitting diode display, a thin-film transistor (TFT) display, a liquid crystal display, an active-matrix organic light-emitting diode (AMOLED) display, a microphone, a speaker, a ringer, and the like.
[0044] In an embodiment, the communication interface 208 may include mechanisms configured to communicate with other entities to receive the input image 104 and the text data 106. In an embodiment, the communication interface 208 of the system 102 receives the input image 104 and the text data 106 and generates the output image 108.
[0045] In an embodiment, the input image 104 and the text data 106 may be stored in a database 210. The system 102 is depicted to be in operative communication with a database 210. In one embodiment, the database 210 is configured to store the input image 104 and the text data 106 received by the user device 110 over a period of time.
[0046] The database 210 may include multiple storage units such as hard disks and / or solid-state disks in a redundant array of inexpensive disks (RAID) configuration. In some embodiments, the database 210 may include a storage area network (SAN) and / or a network attached storage (NAS) system. In one embodiment, the database 210 may correspond to a distributed storage system, wherein individual databases are configured to store custom information, such as correlation information, data of a plurality of domains, and the like.
[0047] In some embodiments, the database 210 is integrated within the system 102. For example, the system 102 may include one or more hard disk drives like database 210. In other embodiments, the database 210 is external to the system 102 and may be accessed by the system 102 using a storage interface (not shown in Fig. 2). The storage interface is any component capable of providing the processor 202 with access to the database 210. The storage interface may include, for example, an Advanced Technology Attachment (ATA) adapter, a Serial ATA (SATA) adapter, a Small Computer System Interface (SCSI) adapter, a RAID controller, a SAN adapter, a network adapter, and / or any component providing the processor 202 with access to the database 210.
[0048] Fig. 3 illustrates an exemplary depiction of generating text embeddings, in accordance with an embodiment of the present disclosure. The text data 106 received from the user device 110 comprise one or more prompts for generating the output image 108 (illustrated in Fig. 7). The one or more prompts comprises at least one of: one or more static prompts 302 and one or more dynamic prompts 304. The one or more static prompts 302 include information associated with features in the input image that should not be altered. The one or more dynamic prompts 304 include information associated with features in the input image that may be altered or edited. For example, if the text data is "modify t-shirt color to white keeping the same brand logo, hands down, jeans color to sky-blue with one leg crossed", the one or more prompts are: modify t-shirt color to white, keep the same brand logo, hands down, jeans color to sky-blue, and one leg crossed. The one or more static prompts 302 are: keep the same brand logo. The one or more dynamic prompts 304 are: modify t-shirt color to white, hands down, jeans color to sky-blue, and one leg crossed.
[0049] The one or more static prompts 302 and the one or more dynamic prompts 304 are encoded using an encoder to obtain text embeddings 306.
[0050] Fig. 4 illustrates an exemplary depiction of encoding an input image 104, in accordance with an embodiment of the present disclosure. The input image 104 is encoded using an image encoder. The encoder transforms the input image 104 to an encoded input image 402 by adding a predefined noise. The encoder transforms the input image 104 from a pixel space to a latent space. The encoded image 402 is a noisy latent. The noisy latent comprises timesteps, the timesteps are indicators of depth of the predefined noise. The timesteps may be used to denoise the noisy latent.
[0051] Fig. 5 illustrates an exemplary depiction of identifying one or more alteration features and one or more non-alteration features, in accordance with an embodiment of the present disclosure. An input 502 comprising the text embeddings 306 and the encoded input image 402 is provided to the trained model 212. The trained model 212 identifies one or more alteration features in the encoded input image 402 based on the one or more dynamic prompts 304 in the text embeddings 306. For example, consider the text data associated with the text embeddings 306 is "modify t-shirt color to white keeping the same brand logo, hands down, jeans color to sky-blue with one leg crossed". The one or more static prompts 302 are: keep the same brand logo. The one or more dynamic prompts 304 are: modify t-shirt color to white, hands down, jeans color to sky-blue, and one leg crossed. The one or more alteration features are identified as the t-shirt 504, the jeans 508, the hands 510 and the legs 512. The one or more non-alteration features are identified as the brand logo 506. In an embodiment, the one or more alteration features are the regions of the input image 104 that may be edited.
[0052] Fig. 6 illustrates identifying one or more alteration features and one or more non-alteration features, in accordance with another embodiment of the present disclosure. Fig. 6 illustrates a self attention based trained model for identifying the one or more alteration features and one or more non-alteration features. A query and key vector based attention technique is depicted. The input to the trained model 212 is the input image 104 and the text data 106 comprising at least one of: the one or more static prompts 302 and the one or more dynamic prompts 304. The one or more prompts comprise a query and a key vector pair. For example, consider the one or more prompts are "modify t-shirt color to white keeping the same brand logo, hands down, jeans color to sky-blue with one leg crossed". For the prompt "modify t-shirt color to white", the query vector is the "t-shirt" and the key vector is the "color to white". Similarly, for "keeping the same brand logo", the query vector is "brand logo" and the key vector is "keeping the same". Similarly, for each prompt, there exists a query and a key pair. The one or more alteration and non-alteration features are determined by the query and key vector pair. Each query vector is analyzed using a separate layer of query layers 606. Each key vector determined by the trained model 212 is provided to a corresponding key layer in key layers 608. The input image 104 is provided to a multilayer attention model at 610 to obtain the input image 104 in latent space by transforming the input image 104 from a pixel space to the latent space. The query layers 606 and the corresponding key layers 608 and input image in latent space 610 is provided to a corresponding layer of an attention model 612. The attention model 612 determines one or more alteration features and one or more non-alteration features based on the query and the corresponding key vector pair. The input image 104 is edited based on the query and key to generate an output image 108. The generation of the output image 108 is illustrated in Fig. 7.
[0053] Fig. 7 illustrates an exemplary depiction of a generative imaging, in accordance with an embodiment of the present disclosure. The input image 104 and the text data 106 are provided to the trained model 212. The trained model 212 transforms the input image 104 to an encoded input image 402 by adding a predefined noise. The trained model 212 encodes the text data 106 to obtain the text embeddings 306. The encoded input image 402 and the text embeddings 306 are provided to the trained model 212. The trained model 212 identifies one or more alteration features and the one or more non-alteration features from the text embeddings 306 (explained in Fig. 5 and 6). The trained model 212 edits the one or more alteration features based on the one or more dynamic prompts 304. The trained model 212 retains the one or more non-alteration features based on the one or more static prompts 302 from the text embeddings 306. The trained model 212 generates the output image 108 based on the text embeddings 306 and the encoded input image 402.
[0054] Fig. 8 illustrates an exemplary depiction of a generative imaging, in accordance with another embodiment of the present disclosure. The input image 104 and the text data 106 is provided to the trained model 212, from the user device 110.
[0055] The text data 106 received from the user device 110 comprise one or more prompts for generating the output image 108. The one or more prompts comprises at least one of: one or more static prompts 302 and one or more dynamic prompts 304. The one or more static prompts 302 include information associated with features in the input image that may be altered or edited. The one or more dynamic prompts 304 include information associated with features in the input image that may not be altered. For example, if the text data is "modify t-shirt color to white keeping the same brand logo, hands down, jeans color to sky-blue with one leg crossed", the one or more prompts are: modify t-shirt color to white, keep the same brand logo, hands down, jeans color to sky-blue, and one leg crossed. The one or more static prompts 302 are: keep the same brand logo. The one or more dynamic prompts 304 are: modify t-shirt color to white, hands down, jeans color to sky-blue, and one leg crossed. The one or more static prompts 302 and the one or more dynamic prompts 304 are encoded using an encoder to obtain text embeddings 306.
[0056] The input image 104 is encoded using the image encoder. The image encoder transforms the input image 104 to the encoded input image 402 by adding a predefined noise. The predefined noise may be defined by a user, stored in the memory 204 or the database 210, may be provided by the user device 110, and the like. The encoder transforms the input image 104 from the pixel space to the latent space. The encoded image 404 is the noisy latent. The noisy latent comprises timesteps, the timesteps are indicators of depth of the predefined noise. The timesteps may be used to denoise the noisy latent.
[0057] The trained model 212 identifies one or more alteration features in the encoded input image 402 based on the one or more dynamic prompts 304 in the text embeddings 306. For example, consider the text data associated with the text embeddings 306 is "modify t-shirt color to white keeping the same brand logo, hands down, jeans color to sky-blue with one leg crossed". The one or more static prompts 302 are: keep the same brand logo. The one or more dynamic prompts 304 are: modify t-shirt color to white, hands down, jeans color to sky-blue, and one leg crossed. The one or more alteration features are identified as the t-shirt 504, the jeans 508, the hands 510 and the legs 512. The one or more non-alteration features are identified as the brand logo 506. In an embodiment, the one or more alteration features are the regions of the input image 104 that may be altered or edited.
[0058] The trained model 212 alters and / or edits the one or more alteration features and retains the one or more non-alteration features based on the one or more prompts in the text embeddings 306, to obtain an edited image. The trained model 212 determines a noise to be removed from the edited image based on the predefined noise and the timestep. For example, consider 1000 timesteps and the trained model 212 is trained for 20 inference steps, the image is encoded from the input image 104 having timestep 0 to pure Gaussian noise i.e., the encoded input image 402 having timestep 999. Therefore, for denoising the edited image, the model may determine the timestep as 999-20=979 (i.e., [t-20]).
[0059] In an embodiment, the trained model 212 removes the noise from the edited image to obtain a denoised edited image i.e., the output image 108. The trained model 212 determines if the timestep in the edited image 108 is 0. If the timestep is not 0, the edited image 108 is given back to the trained model 212 to remove the noise. This step is repeated until the timestep is 0. The output of the trained model 212 when the timestep is 0 is the output image 802.
[0060] Fig. 9 illustrates the trained model 212 for generative imaging, in accordance with an embodiment of the present disclosure. In an embodiment, the trained model 212 is a self attention based model. The input image 104, the text data 106, the predefined noise and the timestep based on the predefined noise is provided to Multi-Layer Perceptron (MLP) layers 902. A multilayer Perceptron (MLP) is a feedforward artificial neural network, consisting of fully connected neurons with a nonlinear activation function, organized in at least three layers. The MLP may distinguish data that is not linearly separable. The MLP1 902 encodes the input image 104 by adding the predefined noise. The text embeddings 306 and the encoded input image 402 are provided to a Self Attention (SA) block-1 904. The SA1 904 performs attention mechanism to determine the one or more alteration features and the one or more non-alteration features in the input image 104 based on the text embeddings 306. The output of SA1 904 is provided to another SA model viz., SA2 906 along with the predefined noise. The SA2 906 performs attention mechanism to determine what part of the predefined noise to be modified and generates the edited image. The output from SA2 906 and the time steps are provided to the MLP2 908. The MLP2 908 determines how the noise may be removed from the edited image. The edited image may be provided to MLP3 910 to provide the output image 108.
[0061] The MLP1 block comprises of an input layer, one or more hidden layers, and an output layer. The input layer takes in the input features i.e., the input image 104, the text embeddings 306 and the predefined noise. The hidden layers process these features using an activation function. The output layer produces the final output. For example, the MLP1 block may encode the input image 104 to generate the encoded input image 402.
[0062] Fig. 10a-c illustrate exemplary depictions of generative imaging, in accordance with an embodiment of the present disclosure. In Fig. 10a, the text data 106 is "a girl with red curly hair" and the input image 104 is provided to the system 102. The system 102 generates an output image 108. In the output image 108, the girl has red curly hair. In Fig. 10b, the text data 106 is "add a teddy bear sitting on the bench" and the input image 104 is provided to the system 102. The system 102 generates an output image 108. In the output image 108, the teddy is placed on the bench. In Fig. 10c, an image 1008 with a different style is provided as a directive image with the input image 104 to the system 102. The system 102 generates an output image 108 such that the style of the image 1008 is adapted to the input image 104. In an embodiment, the input image 104 may be edited for changing one or more features of the input image 104, changing style of input image 104, adding one or more objects to the input image 104, removing one or more objects from the input image 104, and the like.
[0063] Fig. 11 illustrates a flow chart of a method 1100 of generative imaging, in accordance with an embodiment of the present disclosure.
[0064] At 1102, the system 102 receives the input image 104 and text data 106. The text data 106 comprises one or more prompts for generating the output image 108.
[0065] At 1104, the system 102 identifies the one or more prompts in the text data. For example, the trained model 212 in the system 102 identifies at least one of: the one or more static prompts 302 and the one or more dynamic prompts 304 from the one or more prompts in the text data 106. For example, if the text data 106 is "modify t-shirt color to white keeping the same brand logo, hands down, jeans color to sky-blue with one leg crossed", the one or more prompts are: modify t-shirt color to white, keep the same brand logo, hands down, jeans color to sky-blue, and one leg crossed. The one or more static prompts 302 are: keep the same brand logo. The one or more dynamic prompts 304 are: modify t-shirt color to white, hands down, jeans color to sky-blue, and one leg crossed.
[0066] At 1106, the system 102 encodes the identified one or more prompts for generating text embeddings. For example, the trained model 212 encodes the identified prompts for generating text embeddings 306.
[0067] At 1108, the system 102 encodes the input image 104 by including the predefined noise, For example, the trained model 212 encodes the input image 104 by including the predefined noise to obtain the encoded input image 402.
[0068] At 1110, the system 102 identifies one or more alteration features and one or more non-alteration features based on the identified one or more prompts and the input image 402. For example, the trained model 212 identifies at least one of: the one or more alteration features in the encoded input image 402 based on the one or more dynamic prompts 304, and the one or more non-alteration features in the encoded input image 402 based on the one or more static prompts 302. For example, consider the text data associated with the text embeddings 306 is "modify t-shirt color to white keeping the same brand logo, hands down, jeans color to sky-blue with one leg crossed". The one or more static prompts 302 are: keep the same brand logo. The one or more dynamic prompts 304 are: modify t-shirt color to white, hands down, jeans color to sky-blue, and one leg crossed. The one or more alteration features are identified as the t-shirt 504, the jeans 508, the hands 510 and the legs 512. The one or more non-alteration features are identified as the brand logo 506. In an embodiment, the one or more non-alteration features is a subset of the one or more alteration features for example, the t-shirt is an alteration feature and the brand logo on the t-shirt is a non-alteration feature. In another embodiment, the one or more alteration features is a subset of the one or more non-alteration features. The system 102 may identify the one or more alteration features and the one or more non-alteration features based on the identified one or more prompts, the input image 402 and the directive image 1008. The system 102 may identify the one or more alteration features and the one or more non-alteration features based on the identified one or more prompts, the input image 402, the directive image 1008 and the text data.
[0069] At 1112, the system 102 generates the output image by denoising the encoded input image based on the one or more alteration features and the one or more non-alteration features. For example, the system 102 determines a noise to be removed from the encoded input image based on the one or more alteration features and the one or more non-alteration features. The system 102 removes the noise from the encoded input image 402 to obtain the output image 108.
[0070] According to an embodiment of the present disclosure, the trained model 212 generates the output image 108 by transforming the one or more alteration features in the encoded input image 402 based on the one or more dynamic prompts 304. The trained model 212 determines a noise to be removed from the output image 108 based on the predefined noise and a timestep. The trained model 212 removes the noise from the output image 108 to obtain a denoised output image 802.
[0071] The disclosed method 1100 with reference to Fig. 11, may be implemented using software including computer-executable instructions stored on one or more computer-readable media (e.g., non-transitory computer-readable media, such as one or more optical media discs, volatile memory components (e.g., DRAM or SRAM), or non-volatile memory or storage components (e.g., hard drives or solid-state non-volatile memory components, such as Flash memory components) and executed on a computer (e.g., any suitable computer, such as a laptop computer, net book, Web book, tablet computing device, smart phone, or other mobile computing device). Such software may be executed, for example, on a single local computer.
[0072] The order in which the method 1100 is described is not intended to be construed as a limitation, and any number of the described method blocks can be combined in any order to implement the method. Additionally, individual blocks may be deleted from the methods without departing from the scope of the subject matter described herein. Furthermore, the method can be implemented in any suitable hardware, software, firmware, or combination thereof.
[0073] Fig. 12 illustrates a block diagram of an exemplary computer system 1200, for implementing embodiments consistent with the present disclosure. The computer system 1200 may be, without limitation to, the system 102. The computer system 1200 may include a central processing unit ("CPU" or "processor") 1201. The processor 1201 may include at least one data processor for executing processes. The processor 1201 may include specialized processing units such as, integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc.
[0074] The processor 1201 may be disposed in communication with one or more input / output (I / O) devices 1208 and 1209 via I / O interface 1207. The I / O interface 1207 may employ communication protocols / methods such as, without limitation, audio, analog, digital, monaural, RCA, stereo, IEEE-1394, serial bus, universal serial bus (USB), infrared, PS / 2, BNC, coaxial, component, composite, digital visual interface (DVI), high-definition multimedia interface (HDMI), RF antennas, S-Video, VGA, IEEE 902.n / b / g / n / x, Bluetooth, cellular (e.g., code-division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, or the like), etc.
[0075] Using the I / O interface 1207, the computer system 1200 may communicate with one or more I / O devices 1208 and 1209. For example, the input devices 1208 may be an antenna, keyboard, mouse, joystick, (infrared) remote control, camera, card reader, fax machine, dongle, biometric reader, microphone, touch screen, touchpad, trackball, stylus, scanner, storage device, transceiver, video device / source, etc. The output devices 1209 may be a printer, fax machine, video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, Plasma display panel (PDP), Organic light-emitting diode display (OLED) or the like), audio speaker, etc.
[0076] In some embodiments, the processor 1201 may be disposed in communication with external elements such as external computer systems, servers, network elements. The network interface 1210 may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc.
[0077] In some embodiments, the processor 1201 may be disposed in communication with a memory 1203 (e.g., RAM, ROM, etc.) via a storage interface 1202. The storage interface 1202 may connect to memory 1203 including, without limitation, memory drives, removable disc drives, etc., employing connection protocols such as, serial advanced technology attachment (SATA), Integrated Drive Electronics (IDE), IEEE-1394, Universal Serial Bus (USB), fibre channel, Small Computer Systems Interface (SCSI), etc. The memory drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory devices, solid-state drives, etc.
[0078] The memory 1203 may store a collection of program or database components, including, without limitation, user interface 1204, an operating system 1205, a web browser 1206 etc. In some embodiments, computer system 1200 may store user / application data, such as, the data, variables, records, etc., as described in this disclosure. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle® or Sybase®.
[0079] The operating system 1205 may facilitate resource management and operation of the computer system 1200. Examples of operating systems include, without limitation, APPLE MACINTOSH® OS X, UNIX®, UNIX-like system distributions (E.G., BERKELEY SOFTWARE DISTRIBUTIONTM (BSD), FREEBSDTM, NETBSDTM, OPENBSDTM, etc.), LINUX DISTRIBUTIONSTM (E.G., RED HATTM, UBUNTUTM, KUBUNTUTM, etc.), IBMTM OS / 2, MICROSOFTTM WINDOWSTM (XPTM, VISTATM / 7 / 8, 10 etc.), APPLE® IOSTM, GOOGLE® ANDROIDTM, BLACKBERRY® OS, or the like.
[0080] In some embodiments, the computer system 1200 may implement the web browser 1206 stored program components. The web browser 1206 may be a hypertext viewing application, such as MICROSOFT® INTERNET EXPLORER®, GOOGLETM CHROMETM, MOZILLA® FIREFOX®, APPLE® SAFARI®, etc. Secure web browsing may be provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. Web browsers 1206 may utilize facilities such as AJAX, DHTML, ADOBE® FLASH®, JAVASCRIPT®, JAVA®, Application Programming Interfaces (APIs), etc. In some embodiments, the computer system 1200 may implement a mail server stored program component. The mail server may be an Internet mail server such as Microsoft Exchange, or the like. The mail server may utilize facilities such as Active Server Pages (ASP), ACTIVEX®, ANSI® C++ / C#, MICROSOFT®, .NET, CGI SCRIPTS, JAVA®, JAVASCRIPT®, PERL®, PHP, PYTHON®, WEBOBJECTS®, etc. The mail server may utilize communication protocols such as Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), MICROSOFT® exchange, Post Office Protocol (POP), Simple Mail Transfer Protocol (SMTP), or the like. In some embodiments, the computer system 1200 may implement a mail client stored program component. The mail client may be a mail viewing application, such as APPLE® MAIL, MICROSOFT® ENTOURAGE®, MICROSOFT® OUTLOOK®, MOZILLA® THUNDERBIRD®, etc.
[0081] Fig. 13 illustrates an exemplary depiction of generative imaging, in accordance with an embodiment of the present disclosure. The system 102 receives the input image 104 and the text data 106 and generates the output image 108. The system 102 encodes the input image 104 and the text data 106. The trained model 212 combines the encoded image and the encoded text to obtain a combined contextual embedding. The trained model 212 performs a first level of self-attention step 1310. In the first level of self-attention step 1310, the trained model 212 identifies alteration features that need to be edited and non-alteration features that do not need to be edited, based on the combined contextual embedding. The alteration features include mask areas (for example, t-shirt area) and attributes (for example, color to white). The non-alteration features include mask areas (for example, brand logo area).
[0082] The system 102 transforms the input image 104 from a pixel space to a latent space. The system 102 obtains an encoded input image 402 by adding a predefined noise to the transformed input image. The trained model 212 performs a second level of self-attention step 1320. In the second level of self-attention step 1320, the trained model 212 iteratively removes noises from the encoded input image 402 based on the alteration features and the non-alteration features to obtain the output image 108.
[0083] Fig. 14 illustrates an exemplary depiction of generative imaging, in accordance with an embodiment of the present disclosure. The system 102 receives the input image 104, a directive image 1410 and the text data 106. An image encoder 1420 transforms the input image 104 into the latent space as an image embedding. An image encoder 1420 transforms the directive image 1410 into the latent space as a directive image embedding. The text encoder 1440 processes the text data 106 and transforms the text data 106 into the text embedding. The deep attention denoising module 1450 identifies the alteration features and the non-alteration features base on the text embedding, the image embedding and the directive image embedding. The system 102 transforms the input image 104 from a pixel space to a latent space. The system 102 obtains an encoded input image 402 by adding a predefined noise to the transformed input image. The deep attention denoising module 1450 iteratively denoises the encoded input image 402 based on the alteration features and the non-alteration features to obtain the output image 108.
[0084] The present disclosure discloses a method and system for generative imaging. The present disclosure provides a system which can take both the image and the text data as input and generate an output image.
[0085] An embodiment of the present system eliminates the need for multiple models to edit an image.
[0086] The present system is efficient as the system performs context aware editing of images. The system is less complex and easy to deploy, as there is no need for integration of multiple models.
[0087] The described operations may be implemented as a method, system or article of manufacture using standard programming and / or engineering techniques to produce software, firmware, hardware, or any combination thereof. The described operations may be implemented as code maintained in a "non-transitory computer readable medium", where a processor may read and execute the code from the computer readable medium. The processor is at least one of a microprocessor and a processor capable of processing and executing the queries. A non-transitory computer readable medium may include media such as magnetic storage medium (e.g., hard disk drives, floppy disks, tape, etc.), optical storage (CD-ROMs, DVDs, optical disks, etc.), volatile and non-volatile memory devices (e.g., EEPROMs, ROMs, PROMs, RAMs, DRAMs, SRAMs, Flash Memory, firmware, programmable logic, etc.), etc. Further, non-transitory computer-readable media may include all computer-readable media except for a transitory. The code implementing the described operations may further be implemented in hardware logic (e.g., an integrated circuit chip, Programmable Gate Array (PGA), Application Specific Integrated Circuit (ASIC), etc.).
[0088] The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that ongoing technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein. Such alternatives fall within the scope and spirit of the disclosed embodiments. Also, the words "comprising," "having," "containing," and "including," and other similar forms are intended to be equivalent in meaning and be open ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items or meant to be limited to only the listed item or items. It must also be noted that as used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.
[0089] Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present disclosure. A computer readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term "computer readable medium" should be understood to include tangible items and exclude carrier waves and transient signals, i.e., are non-transitory. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, and any other known physical storage media.
[0090] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. Accordingly, the disclosure of the embodiments of the disclosure is intended to be illustrative, but not limiting, of the scope of the disclosure.
[0091] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.
Claims
1.A method of generative imaging, comprising:receiving an input image (104) and text data (106), wherein the text data (106) comprises one or more prompts for generating an output image (108);identifying the one or more prompts in the text data (106);encoding the identified one or more prompts for generating text embeddings (306), using a trained model (212);encoding the input image (104) by including a predefined noise, using the trained model (212);identifying, using the trained model (212), one or more alteration features and one or more non-alteration features based on the identified one or more prompts and the input image (104); andgenerating, using the trained model (212), the output image (108) by denoising the encoded input image (402) based on the one or more alteration features and the one or more non-alteration features.2.The method of claim 1, wherein encoding the input image (104) comprising:transforming the input image (104) from a pixel space to a latent space.3.The method of claim 1 or claim 2, wherein identifying the one or more prompts in the text data (106) comprising:identifying one or more static prompts (302) and one or more dynamic prompts (304) in the text data (106),wherein identifying, using the trained model (212), at least one of one or more alteration features and one or more non-alteration features based on the identified one or more prompts and the input image (104) comprising:identifying the one or more alteration features in the encoded input image (402) based on the one or more dynamic prompts (304), andidentifying one or more non-alteration features in the encoded input image (402) based on the one or more static prompts (302).4.The method of any one of claims 1 to 3, further comprising:determining a noise to be removed from the encoded input image (402) based on the one or more alteration features and the one or more non-alteration features; andremoving the noise from the encoded input image (402) to obtain the output image (108).5.The method of any one of claims 1 to 4, wherein the one or more alteration features is a subset of the one or more non-alteration features.6.The method of any one of claims 1 to 5, wherein the one or more non-alteration features is a subset of the one or more alteration features.7.The method of any one of claims 1 to 6, further comprising:receiving a directive image (1410) for generating the output image (108),wherein the identifying, using the trained model (212), the at least one of one or more alteration features and the one or more non-alteration features based on the identified one or more prompts and the input image (104) comprise identifying, using the trained model (212), the at least one of one or more alteration features and the one or more non-alteration features based on the identified one or more prompts, the input image (104) and the directive image (1410).8.A system for generative imaging, comprises:a memory configured to store instructions; anda processor configured to execute the instructions stored in the memory and thereby configured to:receive an input image (104) and text data (106), wherein the text data (106) comprises one or more prompts for generating an output image (108);identify the one or more prompts in the text data (106);encode the identified one or more prompts for generating text embeddings (306), using a trained model;encode the input image (104) by including a predefined noise, using the trained model; andidentify, using the trained model, one or more alteration features and one or more non-alteration features based on the identified one or more prompts and the input image (104);generate, using the trained model, the output image (108) by denoising the encoded input image (402) based on the one or more alteration features and the one or more non-alteration features.9.The system of claim 8, wherein to encode the input image (104), the processor is configured to:transform the input image (104) from a pixel space to a latent space.10.The system of claim 8 or claim 9, wherein the processor is configured to:identify one or more static prompts (302) and one or more dynamic prompts (304) in the text data (106),identify the one or more alteration features in the encoded input image (402) based on the one or more dynamic prompts (304), andidentify the one or more non-alteration features in the encoded input image (402) based on the one or more static prompts (302).11.The system of any one of claims 8 to 10, wherein the processor is configured to:determine a noise to be removed from the encoded input image (402) based on the one or more alteration features and the one or more non-alteration features; andremove the noise from the encoded input image (402) to obtain the output image (108).12.The system of any one of claims 8 to 11, wherein the one or more alteration features is a subset of the one or more non-alteration features.13.The system of any one of claims 8 to 12, wherein the one or more non-alteration features is a subset of the one or more alteration features.14.The system of any one of claims 8 to 13, wherein the processor is configured to:receive a directive image (1410) for generating the output image (108); andidentify, using the trained model, the at least one of one or more alteration features and the one or more non-alteration features based on the identified one or more prompts, the input image (104) and the directive image (1410).15.A non-transitory computer-readable recording medium having recorded thereon computer-readable codes as a program for executing the method of claim 1.