Multi-attribute conversion for text-to-image synthesis
By providing specific attribute lexicons in different layers and time-step sets of the image generation model, and using the multi-attribute transformation algorithm (MATTE) to untangle complex attributes in the reference image, the problem of difficulty in maintaining user control and independently selecting image attributes in the prior art is solved, and more refined and creative control of the image generation process is achieved.
Patent Information
- Application Number
- CN202411175430.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-08-26
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art is difficult to maintain user control when synthesising images with complex attributes from reference images, and it is impossible to independently select and modify specific attributes such as colors, styles, layouts, and objects.
By providing specific attribute lexicons in different layers and time step sets of image generation models, using a multi-attribute transformation algorithm (MATTE) to unwrap complex attributes in reference images, and operating across layers and time step dimensions during the generation process, allowing users to independently select and modify image attributes.
More refined and creative control of the image generation process is achieved, and the generated synthetic images can more accurately reflect the user's intentions and preferences.
Smart Images

Figure CN120020883A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computers, and more particularly, to image processing. Background Art
[0002] The following generally relates to image processing, and more particularly to image generation using diffusion models. Diffusion models are increasingly being used in the field of image generation, particularly because they are capable of performing conditional image synthesis. By guiding the generative process with specific inputs, these models facilitate the creation of images that adhere to predefined attributes or themes. The adaptability of diffusion models to such tasks is useful in a wide range of applications, from design to automated content generation. Summary of the Invention
[0003] A method, apparatus, and non-transitory computer-readable medium for image processing are described. One or more aspects of the method, apparatus, and non-transitory computer-readable medium include: obtaining a text prompt, a first attribute token, and a second attribute token; identifying a first set of layers and a first set of time steps of an image generation model for the first attribute token, and identifying a second set of layers and a second set of time steps of the image generation model for the second attribute token; and generating a synthetic image using the image generation model based on the text prompt, the first attribute token, and the second attribute token by providing the first attribute token to the first set of layers of the image generation model during the first set of time steps and providing the second attribute token to the second set of layers of the image generation model during the second set of time steps.
[0004] A method, apparatus, and non-transitory computer-readable medium for image processing are described. One or more aspects of the method, apparatus, and non-transitory computer-readable medium include: obtaining training data that includes a first attribute token representing a first attribute and a second attribute token representing a second attribute; identifying a first set of layers and a first set of time steps of an image generation model for the first attribute token, and identifying a second set of layers and a second set of time steps of the image generation model for the second attribute token; and training the image generation model to generate a synthetic image that includes the first attribute and the second attribute by providing the first attribute token to the first set of layers of the image generation model during the first set of time steps and providing the second attribute token to the second set of layers of the image generation model during the second set of time steps.
[0005] A device and method for image processing are described. One or more aspects of the device and method include: at least one processor; at least one memory storing instructions executable by the at least one processor; and an image generation model including parameters stored in the at least one memory and trained to generate a synthetic image including a first attribute and a second attribute by providing a first attribute token to a first set of layers of the image generation model during a first set of time steps and providing a second attribute token to a second set of layers of the image generation model during a second set of time steps. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 An example of an image processing system in accordance with aspects of the present disclosure is shown.
[0007] Figure 2 An example of an image generation application in accordance with aspects of the present disclosure is shown.
[0008] Figure 3 An example of a synthetic image using attribute-guided image synthesis in accordance with aspects of the present disclosure is shown.
[0009] Figure 4 An example of multi-prompt conditioning jointly across U-Net layers and denoising time steps in accordance with aspects of the present disclosure is shown.
[0010] Figure 5 An example of an image processing device in accordance with aspects of the present disclosure is shown.
[0011] Figure 6 An example of a machine learning model for generating a synthetic image in accordance with an embodiment of the present disclosure is shown.
[0012] Figure 7 An example of a guided diffusion architecture in accordance with aspects of the present disclosure is shown.
[0013] Figure 8 An example of a U-Net in accordance with aspects of the present disclosure is shown.
[0014] Figure 9 An example of an example diffusion process in accordance with aspects of the present disclosure is shown.
[0015] Figure 10 An example of a machine learning model for computing input conditioning in accordance with aspects of the present disclosure is shown.
[0016] Figure 11 An example of an image processing method in accordance with aspects of the present disclosure is shown.
[0017] Figure 12 An example of an image processing method in accordance with aspects of the present disclosure is shown.
[0018] Figure 13 Illustrates an example of a method for training a machine learning model according to aspects of the present disclosure.
[0019] Figure 14 Illustrates an example of training a machine learning model according to aspects of the present disclosure.
[0020] Figure 15 Illustrates an example of a computing device according to aspects of the present disclosure. Detailed Description
[0021] Recent developments in machine learning have broadened the scope of image processing by introducing generative models, which can synthesize realistic images from various inputs. Diffusion models stand out for their ability to generate synthetic images formed by text descriptions and visual references.
[0022] Traditional techniques for generating personalized outputs typically use a single conditioning vector extracted from a reference image to produce variations. These methods rely on data where each token is represented by a vector to drive the customization process. Additionally, some other methods involve adjusting the weights of the diffusion model for more personalization and performing transformation tasks via latent optimization and model fine-tuning using generative adversarial networks (GANs). However, these methods have limitations in maintaining user control while synthesizing images with complex attributes from reference images.
[0023] Embodiments of the present disclosure include an improved image generation model that generates more accurate synthetic images (e.g., a synthetic image depicting an object from text in the style of another image) based on attributes indicated by a reference image and a text prompt. Conventional image generation devices do not untangle and extract attributes from a reference image and do not apply the attributes extracted from the reference image to the generated image. In contrast, embodiments of the present disclosure enable a user to independently select and modify specific attributes (such as color, style, layout, and objects), thereby allowing for more refined and creative control over the image generation process.
[0024] For example, embodiments of the present disclosure provide a method that untangles complex attributes in a reference image by considering the layers and time step dimensions of a denoising diffusion probability model (DDPM). This method enables the identification of specific aspects and stages of the generation process that capture various attributes. Some embodiments include a multi-attribute transformation algorithm (MATTE) that includes an orientation regularization loss for enhancing attribute untangling. MATTE operates across layer and time step dimensions, effectively allowing the extraction and subsequent synthesis of images based on multiple attributes (including color, object, layout, and style) from a single reference image.
[0025] The applications of the present disclosure can be integrated with image processing applications, where the MATTE algorithm enables enhanced user control over the image generation process without retraining the model. The algorithm facilitates the generation of images that comply with the property constraints derived from a reference image, thereby enabling the generation of variations that match user preferences. Additionally, the systems and methods described herein are designed to be integrated into systems with clear UI components, which helps detect infringement.
[0026] Accordingly, embodiments of the present disclosure provide a technological advancement in the field of image synthesis, thus providing a method for disentangling and personalized generation of images. This can be achieved by integrating the MATTE algorithm, thereby enhancing the ability to guide image synthesis using multiple attributes extracted from a reference image provided by the user.
[0027] Image Processing System
[0028] In Figures 1 - 4 , an image processing method is described. One or more aspects of the method include: obtaining a text prompt, a first attribute token, and a second attribute token; identifying a first set of layers and a first set of time steps of an image generation model for the first attribute token, and identifying a second set of layers and a second set of time steps of the image generation model for the second attribute token; and generating a synthetic image using the image generation model based on the text prompt, the first attribute token, and the second attribute token by providing the first attribute token to the first set of layers of the image generation model during the first set of time steps and providing the second attribute token to the second set of layers of the image generation model during the second set of time steps.
[0029] In some aspects, the first attribute token includes a first token type, and the second attribute token includes a second token type, where the first token type and the second token type are selected from a set of token types including a color token type, an object token type, a style token type, and a layout token type. In some aspects, the synthetic image includes elements described by the text prompt, a first attribute represented by the first attribute token, and a second attribute represented by the second attribute token. The first attribute token and the second attribute token each include a learnable token corresponding to the first attribute and the second attribute, respectively. Some examples of the method, apparatus, and non-transitory computer-readable medium also include obtaining the first attribute token and the second attribute token, including receiving user input indicating the first attribute and the second attribute. In some aspects, the first set of layers does not overlap with the second set of layers, and the first set of time steps does not overlap with the second set of time steps.
[0030] Some examples of the method, apparatus, and non-transitory computer-readable medium further include: generating a synthetic image by performing a reverse diffusion process on a noisy input image, where the reverse diffusion process is based on a plurality of time steps including a first set of time steps and a second set of time steps. Some examples of the method, apparatus, and non-transitory computer-readable medium further include encoding a text prompt to obtain a text embedding, where the synthetic image is generated based on the text embedding.
[0031] Figure 1 An example of an image processing system in accordance with aspects of the present disclosure is shown. The example shown includes a user 100, a user device 105, an image processing apparatus 110, a cloud 115, and a database 120. The image processing apparatus 110 is an example of the corresponding element described with reference to Figures 5 - 8 or includes aspects of the corresponding element described with reference to Figures 5 - 8 described.
[0032] In Figure 1 the example shown, a reference image and a text prompt are provided to the image processing apparatus 110, for example, via the user device 105 and the cloud 115. The reference image features a cat with a specific color, style, and layout, and the object of interest is the cat itself. The text prompt specifies an element from the text, namely, "dog". The user also indicates that the synthetic image conforms to the attributes among the color, style, layout, and object attributes of the reference image. For example, the user can indicate conforming to the style of the reference image. Then, the image processing apparatus 110 processes the reference image to capture its style attributes and obtains an image embedding encoding these characteristics. In addition, the image processing apparatus 110 processes the text prompt to generate a text embedding, which includes a requested subject transformation from the cat in the reference image to the dog indicated by the text prompt.
[0033] In this example, via a generative model such as a diffusion model, the image processing apparatus 110 synthesizes an output image based on the reference image and the text embedding. The output image depicts a dog characterized by the style of the reference cat image. For example, if the reference cat is depicted in the style of Vincent van Gogh and has a specific layout and color scheme, the output image is a dog presented in these same style characteristics. Using the reference cat as a style template, the image processing apparatus 110 also ensures that the style, color, layout, and presence of the specified object (now a dog) are coherently integrated. In this example, the synthesized dog not only replicates the physical form of the cat but also adopts the aesthetic and compositional attributes of the reference image. Then, the resulting output image is returned to the user 100 via the cloud 115 and the user device 105.
[0034] The user device 105 can be a personal computer, laptop computer, mainframe computer, handheld computer, personal assistant, mobile device, or any other suitable processing device. In some examples, the user device 105 includes software that incorporates image processing applications (e.g., query answering, image editing, relationship detection). In some examples, the image editing application on the user device 105 can include the functionality of the image processing apparatus 110.
[0035] The user interface enables the user 100 to interact with the user device 105. In some embodiments, the user interface can include an audio device, such as an external speaker system, an external display device such as a display screen, or an input device (e.g., a remote control device that engages with the user interface directly or through an I / O controller module). In some cases, the user interface can be a graphical user interface (GUI). In some examples, the user interface can be represented by code that is sent to the user device 105 and locally rendered by a browser. Refer to Figure 2 A process of using the image processing apparatus 110 is further described.
[0036] The image processing apparatus 110 includes a computer-implemented network that includes an image encoder, a text encoder, a multimodal encoder, and a decoder. The image processing apparatus 110 can also include a processor unit, a memory unit, an I / O module, and a training component. The training component is used to train a machine learning model (or an image processing network). Additionally, the image processing apparatus 110 can communicate with the database 120 via the cloud 115. In some cases, the architecture of the image processing network is also referred to as a network, a machine learning model, or a network model. Refer to Figures 5 - 8 More details about the architecture of the image processing apparatus 110 are provided. Refer to Figures 5 - 8 More details about the operation of the image processing apparatus 110 are provided.
[0037] In some cases, the image processing apparatus 110 is implemented on a server. The server provides one or more functions to users via one or more various networks. In some cases, the server includes a single microprocessor board that includes a microprocessor responsible for controlling all aspects of the server. In some cases, the server uses the microprocessor and protocols to exchange data with other devices / users on one or more networks via the Hypertext Transfer Protocol (HTTP) and the Simple Mail Transfer Protocol (SMTP), although other protocols such as the File Transfer Protocol (FTP) and the Simple Network Management Protocol (SNMP) can also be used. In some cases, the server is configured to send and receive files in Hypertext Markup Language (HTML) format (e.g., for displaying web pages). In various embodiments, the server includes a general-purpose computing device, a personal computer, a laptop computer, a mainframe computer, a supercomputer, or any other suitable processing device.
[0038] The cloud 115 is a computer network that is configured to provide on-demand availability of computer system resources such as data storage and computing power. In some examples, the cloud 115 provides resources without the user actively managing them. The term "cloud" is sometimes used to describe a data center that is available via the Internet to many users. Some large cloud networks have functions distributed across multiple locations from a central server. If a server has a direct or close connection to a user, it is designated as an edge server. In some cases, the cloud 115 is limited to a single organization. In other examples, the cloud 115 is available to many organizations. In one example, the cloud 115 includes a multi-layer communication network consisting of multiple edge routers and core routers. In another example, the cloud 115 is based on a collection of local switches in a single physical location.
[0039] The database 120 is an organized collection of data. For example, the database 120 stores data in a specified format called a schema. The database 120 can be constructed as a single database, a distributed database, multiple distributed databases, or an emergency backup database. In some cases, a database controller can manage the data storage and processing in the database 120. In some cases, a user interacts with the database controller. In other cases, the database controller can operate automatically without user interaction.
[0040] Figure 2 An example of an image generation application according to aspects of the present disclosure is shown. In some examples, these operations are performed by a system including a processor that executes a set of codes to control the functional elements of the apparatus. Additionally or alternatively, certain processes are performed using dedicated hardware. Generally, these operations are performed in accordance with the methods and processes described according to aspects of the present disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.
[0041] At operation 205, the user provides an image and a text prompt. In some cases, the image includes a first attribute token and a second attribute token. In some cases, the operation of this step refers to the user as described in reference Figure 1 or can be performed by the user.
[0042] For example, at operation 205, the user initiates an image generation process by providing a reference image and a text prompt. The reference image depicts a cat characterized by specific style elements that form the first attribute token, such as the unique color palette and brushstroke techniques of Vincent van Gogh. The second attribute token can be the layout or composition of the image, which includes the positioning and pose of the cat. Along with the reference image, the user inputs the text prompt "of a dog", setting the system task to replicate the attributes of the reference image but change the subject to a dog.
[0043] At operation 210, the system identifies the corresponding layers and time steps of the image generation model for the first attribute token and the second attribute token. In some cases, the operation of this step refers to the image processing device as described in reference Figures 5 - 8 or can be performed by the image processing device as described in reference Figures 5 - 8 described.
[0044] For example, at operation 210, the system analyzes the provided reference image to identify the distinct layers within the image generation model that correspond to the first attribute token and the second attribute token. For example, this step involves mapping the style and object attributes of the reference image to the specific layers and time steps within the model that are responsible for generating those attributes. For example, the system can identify that some layers are crucial in re - creating the style while other layers affect the object.
[0045] At operation 215, the system generates a synthetic image based on the text prompt, the first attribute token, and the second attribute token based on the corresponding layers and time steps. In some cases, the operation of this step refers to the image processing device described in reference Figures 5 - 8 or can be performed by the image processing device described in reference Figures 5 - 8 described.
[0046] For example, at operation 215, the system continues to generate the synthetic image, which incorporates the user - specified attributes into the new subject indicated by the text prompt and the reference image. The system utilizes the identified layers and time steps to ensure that the first attribute (e.g., the style as indicated in the reference image) is mirrored when rendering the synthetic image. Additionally, the system ensures that the second attribute (e.g., the object indicated by the text prompt, which is a dog in this example) is reflected in the synthetic image.
[0047] At operation 220, the system displays the synthesized image to the user. In some cases, the operation of this step refers to the image processing device described by reference Figures 5 - 8 or can be performed by the image processing device described by reference Figures 5 - 8 the image processing device.
[0048] For example, at operation 220, the system displays the generated synthesized image to the user. The generated image is a synthesized image of a dog, which is generated to display the distinguishable style attributes initially depicted in the reference image of a cat. The presentation of the image allows the user to evaluate the effectiveness of the attribute transfer and the overall quality of the synthesized image. For example, when the user wants to make more adjustments, the user can choose to provide different text prompts or select different attribute tokens to generate different synthesized images.
[0049] Figure 3 An example of a synthesized image generated using attribute-guided image synthesis according to aspects of the present disclosure is shown. Refer to Figure 3 A reference image 300 is provided to generate the synthesized image. The synthesized image is generated based on one or more attributes of the reference image (e.g., color, style, layout, or object).
[0050] In some examples, the generated synthesized image complies with one attribute among the attributes (i.e., color, style, layout, and object) of the reference image. The generated synthesized image can also comply with a combination of attributes. For example, when the user chooses to comply with the color of the reference image 300 and provides the text prompt "of a vase", the synthesized image 325 that complies with the color of the reference image is generated. For example, when the user chooses to comply with the style of the reference image 300 and provides the text prompt "dog", the synthesized image 330 that complies with the style of the reference image 300 is generated. For example, when the user chooses to comply with the layout of the reference image 300 and provides the text prompt "origami-style panda", the synthesized image 335 that complies with the layout of the reference image 300 is generated. For example, when the user chooses to comply with the object of the reference image 300 and provides the text prompt "doodle style", the synthesized image 340 with one or more objects that are the same or similar to one or more objects of the reference image 300 is generated.
[0051] Figure 4 An example of a method for jointly conditioning across U-Net layers and denoising time steps according to aspects of the present disclosure is shown. The present disclosure provides a method for conditional guidance of a text-to-image diffusion model. According to some embodiments, the attribute distribution during the generation process is analyzed, considering both the layer and time step dimensions. For example, the analysis can be used to identify which layers (e.g., within the DDPM model) and time steps (e.g., during the reverse process) collaborate to capture attributes (e.g., color, style, layout, or object) during image generation.
[0052] For example, the U-Net model includes 16 cross-attention layers with different resolutions. These layers are classified into rough layers, medium layers, and fine layers. Denoising the time steps is divided into four stages: (800 - 1000), (600 - 800), (200 - 600), and (0 - 200). In forward diffusion, it corresponds to (the notation previously used for the time step denoising stage) of the backward denoising process. Therefore, the nature of the backward denoising stage is directly related to the corresponding forward diffusion stage.
[0053] Embodiments of the present disclosure identify specific layers and time steps corresponding to four attributes, such that modifications identifying specific attributes are mainly determined during specific layers and during specific time steps, with no significant impact during other layers or during other time steps.
[0054] According to embodiments of the present disclosure, to identify the layers and time steps corresponding to an attribute, conditioning is added or removed in the time steps and layers, and the resulting output is analyzed. See Figure 4 , which shows the results of image generation based on joint cues across layers and denoising stages. For the finally generated image (i.e., a standing red cat in an oil painting style), text conditioning corresponding to each key attribute (red, standing, cat, and oil painting) in the prompt is specified only for a subset of the layers and only along specific time steps. For example, although blue is specified in some layers, a red cat is generated, indicating that some attributes are not uniformly affected by the modifications across layers and time stages.
[0055] In Figure 4 's example, tokens for four attributes are provided. Regarding the color attribute, specifying colors such as green and blue in the conditioning of the fine and rough layers respectively does not affect the generated red image. Similarly, in the later denoising stage of the medium layer, colors such as white do not affect the generated image, and an image of a red cat is generated in the example of Figure 4 . This indicates that the color is captured in the initial denoising stage across the medium layer.
[0056] See Figure 4 , the U-Net includes 16 cross-attention layers with resolutions of 8, 16, 32, and 64. These layers are divided into three sets: rough (L 6 -L 9 ), medium (L 3 -L 5 and L 10 -L 13 ) and fine (L 1 -L 2 and L 14 -L 16 ). Denoising the time steps is divided into four stages: t1 , t 2 , t 3 and t 4 . Analyze the layers and time steps and layers to determine where these four attributes are captured during generation. For example, embodiments of the present disclosure add or remove conditioning from the time steps and layers and analyze the output.
[0057] In Figure 4 , the synthetic image 405 is generated. The synthetic image 405 is a red standing cat in an oil painting style. See Figure 4 , to generate the synthetic image 405, text conditioning corresponding to each key attribute of red, standing, cat, and oil painting is specified only for a subset of the layers and only along specific time steps. For example, although blue is specified in layers L 1 -L 2 and L 14 -L 16 (including in dimension 410 ("blue, lizard, sitting, graffiti style")), a red cat is still generated, indicating that some patterns exist in how these attributes are distributed across layers and time phases. Specifically, regarding the color attribute in the synthetic image 405, specifying colors such as green and blue in the conditioning of the fine layers (L 1 -L 2 and L 14 -L 16 ) and the rough layers (L 6 -L 9 ) has no effect on the generated red image. Similarly, colors such as white in the later denoising phases (t 3 -L 5 and L 10 -L 13 ) of the medium layers (L 3 , t 4 ) have no effect on the final generation, and we indeed get a red cat. This indicates that the color is captured in the initial denoising phases (t 3 -L 5 and L 10 -L 13 ) of the medium layers (L 1 , t 2 ).
[0058] For example, regarding the style of the synthetic image 405, graffiti is specified across the rough (L 6 -L 9 ) and fine (L 1 -L 2 and L 14 -L 16 ) layers and to the later denoising phases (t 3 , t 4)(Including in dimension 415 (“Blue, Lizard, Sitting Position, Graffiti Style”)) It is specified that the graffiti has no effect on the generated image. The composite image 405 is still based on the cross (L 3 -L 5 and L 10 -L 13 ) of the (t 1 、t 2 ) oil painting.
[0059] For example, regarding the object of the composite image 405, although a cow is specified in the initial stage and the later stage (t 1 、t 4 ), a cat is generated, indicating that the object is captured in the intermediate stage (t 2 、t 3 ). The rough layer (L 6 -L 9 ) is the reason because specifying other types such as lizards in other layers has no effect.
[0060] For example, regarding the layout of the composite image 405, changing the layout aspect from standing to sitting after the first denoising stage (including in dimension 420 (“Blue, Lizard, Sitting Position, Graffiti Style”)) has no effect on the posture of the generated cat. This indicates that the layout is captured in the initial stage (t 1 ). In addition, in this example, the layer with a resolution of 16 is responsible for the layout attribute. In particular, the layout property is mainly captured in the initial few time steps across the layers L 6 、L 8 and L 9 .
[0061] Image Generation Device
[0062] In Figures 5 - 9 , an apparatus for image processing is described. One or more aspects of the apparatus include at least one processor; at least one memory storing instructions executable by at least one processor; and an image generation model including parameters stored in at least one memory and trained to generate a composite image including a first attribute and a second attribute by providing a first attribute token to a first set of layers of the image generation model during a first set of time steps and a second attribute token to a second set of layers of the image generation model during a second set of time steps.
[0063] In some aspects, the image generation model includes a U-Net architecture, and wherein the first set of layers and the second set of layers include different layers of the U-Net architecture.
[0064] In some aspects, the first attribute includes a color attribute or a style attribute, and the first set of layers includes the medium-resolution layers of the image generation model.
[0065] In some aspects, the second attribute includes an object attribute or a layout attribute, and the second layer set includes a rough resolution layer of an image generation model.
[0066] In some aspects, the first attribute includes a color attribute, a style attribute, or a layout attribute, and the first time step set includes an initial time step set.
[0067] In some aspects, the second attribute includes an object attribute, and the second time step set includes an intermediate time step set.
[0068] Figure 5 An example of an image generation device 500 according to aspects of the present disclosure is shown. The image generation device 500 is an example of a corresponding element described with reference to Figure 1 or includes aspects of a corresponding element described with reference to Figure 1 In one aspect, the image generation device 500 includes a processor unit 505, an I / O module 510, a training component 515, a memory unit 520, and a machine learning model 525 including an image generation model 530 and a text encoder 535.
[0069] The processor unit 505 includes one or more processors. A processor is an intelligent hardware device such as a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, discrete gate or transistor logic components, discrete hardware components, or any combination thereof.
[0070] In some cases, the processor unit 505 is configured to operate a memory array using a memory controller. In other cases, the memory controller is integrated into the processor unit 505. In some cases, the processor unit 505 is configured to execute computer-readable instructions stored in the memory unit 520 to perform various functions. In some aspects, the processor unit 505 includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing. According to some aspects, the processor unit 505 includes one or more processors described with reference to Figure 15
[0071] The memory unit 520 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory devices include solid-state memory and hard disk drives. In some examples, the memory is used to store computer-readable and computer-executable software that includes instructions that, when executed, cause at least one processor of the processor unit 505 to perform the various functions described herein.
[0072] In some cases, the memory unit 520 includes a basic input / output system (BIOS) that controls basic hardware or software operations such as interacting with peripheral components or devices. In some cases, the memory unit 520 includes a memory controller that operates the storage cells of the memory unit 520. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, the storage cells within the memory unit 520 store information in the form of logical states. According to some aspects, the memory unit 520 includes the memory subsystem described Figure 15 herein.
[0073] According to some aspects, the image generation device 500 uses one or more processors of the processor unit 505 to execute instructions stored in the memory unit 520 to perform the functions described herein. For example, in some cases, the image generation device 500 obtains a prompt. In some cases, the prompt includes a text prompt.
[0074] Machine learning parameters (also known as model parameters or weights) are variables that provide the behavior and characteristics of a machine learning model. Machine learning parameters can be learned or estimated from training data and are used to make predictions or perform tasks based on the patterns and relationships learned in the data.
[0075] Machine learning parameters are typically adjusted during the training process to minimize a loss function or maximize a performance metric. The goal of the training process is to find the optimal values of the parameters that allow the machine learning model to make accurate predictions or perform well on a given task.
[0076] For example, during the training process, an algorithm adjusts the machine learning parameters according to optimization techniques such as gradient descent, stochastic gradient descent, or other optimization algorithms to minimize the error or loss between the predicted output and the actual target. Once the machine learning parameters are learned from the training data, they are used to make predictions on new, unseen data.
[0077] An artificial neural network (ANN) has many parameters, including weights and biases associated with each neuron in the network, which control the degree of connection between neurons and affect the ability of the neural network to capture complex patterns in the data.
[0078] An ANN is a hardware component or a software component that includes a number of connected nodes (i.e., artificial neurons) that roughly correspond to the neurons in the human brain. Each connection (or edge) transmits a signal from one node to another (like the physical synapses in the brain). When a node receives a signal, it processes the signal and then transmits the processed signal to other connected nodes.
[0079] In some cases, the signals between nodes include real numbers, and the output of each node is calculated as a function of the sum of its inputs. In some examples, a node may use other mathematical algorithms to determine its output, such as selecting the maximum value from the inputs as the output, or any other suitable algorithm for activating the node. Each node and edge is associated with one or more node weights that determine how the signals are processed and transmitted.
[0080] In an ANN, the hidden layer (or intermediate layer) includes hidden nodes and is located between the input layer and the output layer. The hidden layer performs a non - linear transformation on the inputs fed into the network. Each hidden layer is trained to produce a defined output that contributes to the combined output of the ANN's output layer. A hidden representation is a machine - readable data representation of the inputs learned from the hidden layers of the ANN and produced by the output layer. As the ANN is trained, the understanding of the inputs to the ANN improves, and the hidden representation gradually differentiates from earlier iterations.
[0081] During the training process of an ANN, the node weights are adjusted to improve the accuracy of the results (i.e., by minimizing a loss that in some way corresponds to the difference between the current result and the target result). The weights of the edges increase or decrease the strength of the signals transmitted between nodes. In some cases, a node has a threshold below which no signal is transmitted at all. In some examples, nodes are aggregated into layers. Different layers perform different transformations on their inputs. The initial layer is called the input layer, and the last layer is called the output layer. In some cases, the signals pass through some layers multiple times.
[0082] According to some aspects, the machine - learning model 525 identifies a first set of layers and a first set of time steps of the image - generation model for a first attribute token, and identifies a second set of layers and a second set of time steps of the image - generation model for a second attribute token. In some examples, the machine - learning model 525 performs a forward diffusion process on the training images to obtain a noisy input image. In some examples, the machine - learning model 525 performs a reverse diffusion process on the noisy image to obtain a predicted image.
[0083] According to some aspects, the training component 515 obtains training data that includes a first attribute token representing a first attribute and a second attribute token representing a second attribute. In some examples, the training component 515 trains an image generation model to generate a synthetic image including the first attribute and the second attribute by providing the first attribute token to a first set of layers of the image generation model during a first set of time steps and providing the second attribute token to a second set of layers of the image generation model during a second set of time steps. In some examples, the training component 515 compares a predicted image with a training image. In some examples, the training component 515 training the image generation model includes training the image generation model to generate a synthetic image including a color attribute, an object attribute, a style attribute, and a layout attribute. In some examples, the training component 515 training the image generation model includes calculating a color-style disentanglement loss. In some examples, the training component 515 obtaining the training data further includes optimizing the first attribute token to represent the first attribute and optimizing the second attribute token to represent the second attribute. In some examples, the training component 515 obtaining the training data further includes obtaining a training image and a text prompt describing the training image, wherein the image generation is trained based on the training image and the text prompt.
[0084] Figure 6 An example of a machine learning model for generating synthetic images according to an embodiment of the present disclosure is shown. Refer to Figure 6 , the machine learning model 615 generates a synthetic image 630 based on a reference image 650 and a text prompt 610. The machine learning model 615 includes an image generation model 620 and a text encoder 625. The text encoder 625 takes the text prompt 610 as input and generates a text embedding 630. The image generation model 620 generates a synthetic image 635 based on the text embedding 630 and the reference image 605.
[0085] The image generation model 620 uses a diffusion process to synthesize the image 635. Initially, the model 620 introduces a series of noise variables into the reference image 605 through multiple steps, creating a series of images with increasing noise. This process of adding noise is the forward diffusion process, which transforms the original reference image into a high-entropy state. Each step corresponds to a time step in the diffusion model and is quantified by the parameters of the model.
[0086] After the forward diffusion, the image generation model 620 proceeds with the reverse diffusion process. Starting from the most noisy state, the model 620 gradually reduces the noise from the image in the reverse diffusion steps. In each reverse time step, the model consults the text embedding 630, which encapsulates the characteristics to be embodied in the synthetic image as indicated by the text prompt 610. The reverse diffusion process is guided by the text embedding 630, ensuring that the final image is consistent with the user's input while preserving the style and content attributes of the reference image 605.
[0087] During this reverse diffusion, the model 620 systematically refines the image, enhancing the fidelity to the target attributes and the consistency with the input text prompt 610. The model 620 integrates the text embedding 630 at the policy layer and time steps, effectively shaping the synthetic image to highlight the theme of the dog in the style conveyed by the reference image 605. When the image is fully transformed, the reverse diffusion ends, resulting in the generation of the synthetic image 635, which demonstrates the fusion of the text input and the visual cues from the reference image.
[0088] Figure 7 An example of a guided diffusion architecture 700 in accordance with aspects of the present disclosure is shown. Diffusion models are a class of generative ANNs that can be trained to generate new data having features similar to those found in the training data. In particular, diffusion models can be used to generate novel images. Diffusion models can be used for various image generation tasks, including image super-resolution, image generation with perceptual metrics, conditional generation (e.g., text-guided generation), inpainting, and image processing.
[0089] Diffusion models learn to recover data by iteratively adding noise to the data during the forward diffusion process and then denoising the data during the reverse diffusion process. Examples of diffusion models include denoising diffusion probabilistic models (DDPMs) and denoising diffusion implicit models (DDIMs). In a DDPM, the generative process involves reversing a stochastic Markov diffusion process. On the other hand, a DDIM uses a deterministic process such that the same input produces the same output. Diffusion models can also be characterized by whether the noise is added to the image itself (such as pixel diffusion) or to the image features generated by the encoder (such as latent diffusion).
[0090] For example, according to some aspects, the forward diffusion process 715 gradually adds noise to the original image 705 to obtain noise images 720 at various noise levels. In some cases, the forward diffusion process 715 is implemented by a forward diffusion component (such as the forward diffusion component described with reference Figure 5 and Figure 8 to).
[0091] According to some aspects, the first reverse diffusion process 725 gradually removes noise from the noise images 720 at various noise levels in various diffusion steps to obtain predicted denoised images 730. In some cases, a predicted denoised image 730 is created from each of the various noise levels. For example, in some cases, in each diffusion step of the first reverse diffusion process 725, a first diffusion model (such as the first diffusion model described with reference Figure 5 and Figure 8The described first diffusion model) predicts a partially denoised image, where the partially denoised image is a combination of the predicted denoised image (e.g., the predicted final output) of the diffusion step and the noise. Thus, in some cases, each predicted denoised image can be considered as the first diffusion model's prediction of the final noise-free output for each diffusion step, and each predicted denoised image 730 can thus be considered as an "early" prediction of the final output of the corresponding diffusion step of the first reverse diffusion process 725.
[0092] According to some aspects, the predicted denoised image 730 is provided to an upsampling component 735 (such as the Figure 5 and Figure 8 described upsampling component). In some cases, the upsampling component 735 upsamples the predicted denoised image 730 to output an upsampled denoised image 740 at a higher resolution. In some cases, the forward diffusion process 715 gradually adds isotropic noise to the upsampled denoised image 740 at various noise levels to obtain an intermediate input image 745. In some cases, the intermediate input image 745 can be considered as an enlarged version of the partially denoised image at the time step of the first reverse diffusion process 725 corresponding to the predicted denoised image 730, where the intermediate input image 745 includes a Gaussian distribution of noise.
[0093] According to some aspects, the second reverse diffusion process 750 gradually removes noise from the intermediate noisy image 745 to obtain a higher resolution output image 755. In some cases, the output image 755 is created from each of the various noise levels.
[0094] In some cases, each of the first reverse diffusion process 725 and the second reverse diffusion process 750 is implemented via a U-Net ANN (such as the Figure 7 described U-Net architecture). The forward diffusion process 715, the first reverse diffusion process 725, and the second reverse diffusion process 750 are examples of the corresponding elements described in Figure 10 or include aspects of the corresponding elements described in Figure 10 described.
[0095] In some cases, each of the first reverse diffusion process 725 and the second reverse diffusion process 750 is guided based on a prompt 760 (such as a text prompt, an image, a layout, a segmentation map, etc.). The prompt 760 can be encoded using an encoder 765 (in some cases a multimodal encoder) to obtain guidance features 770 (e.g., prompt embeddings) in a guidance space 775.
[0096] According to some aspects, the guidance feature 775 is combined with the noisy image 720 and the intermediate input image 745 at one or more layers of the first reverse diffusion process 720 and the second reverse diffusion process 750 respectively to guide the predicted denoised image 730 and the output image 755 towards including what is described in the prompt 760. For example, cross-attention blocks within the first reverse diffusion process 725 and the second reverse diffusion process 750 can be used to combine the guidance feature 770 with the noisy image 720 and the intermediate input image 745 respectively. In some cases, the guidance feature 770 can be weighted such that the guidance feature 770 has a greater or lesser representation in the predicted denoised image 730 and the output image 755.
[0097] Cross-attention (also known as multi-head attention) is an extension of the attention mechanism used in some ANNs for NLP tasks. In some cases, cross-attention enables each reverse diffusion process in the first reverse diffusion process 725 and the second reverse diffusion process 750 to simultaneously attend to multiple parts of the input sequence, thereby capturing the interactions and dependencies between different elements. In cross-attention, there are typically two input sequences: the query sequence and the key-value sequence. The query sequence represents the elements to be attended to, while the key-value sequence contains the elements to be focused on. In some cases, to compute cross-attention, the cross-attention block transforms each element in the query sequence (e.g., using a linear projection) into a "query" representation, and transforms the elements in the key-value sequence into "key" and "value" representations.
[0098] The cross-attention block computes attention scores by measuring the similarity between each query representation and key representation, where a higher similarity indicates a higher degree of attention to the key element. The attention scores indicate the importance or relevance of each key element to the corresponding query element.
[0099] Then, the cross-attention block normalizes the attention scores to obtain attention weights (e.g., using the softmax function), where the attention weights determine how much information in each value element is incorporated into the final attended representation. By simultaneously attending to different parts of the key-value sequence, the cross-attention block captures the relationships and dependencies across the input sequence, thereby allowing each reverse diffusion process in the first reverse diffusion process 725 and the second reverse diffusion process 750 to better understand the context and generate more accurate and contextually relevant outputs.
[0100] As Figure 7 shown, the guidance diffusion architecture 700 is implemented according to a pixel diffusion model. According to some aspects, the guidance diffusion architecture 700 is implemented according to a latent diffusion model. In the latent diffusion model, the forward diffusion process and the reverse diffusion process occur in the latent space rather than in the pixel space.
[0101] For example, in some cases, an image encoder encodes an original image 705 into image features in a latent space. In some cases, a forward diffusion process 715 adds noise to the image features rather than the original image 705 to obtain noisy image features. In some cases, a first reverse diffusion process 725 gradually removes noise from the noisy image features (guided by a guidance feature 770 in some cases) to obtain predicted denoised image features at intermediate steps of the first reverse diffusion process 725. In some cases, an upsampling component upsamples the predicted denoised image features to obtain upsampled image features. In some cases, the forward diffusion process 715 gradually adds noise to the upsampled image features to obtain intermediate image features. In some cases, a second reverse diffusion process 750 gradually removes noise from the intermediate image features to obtain output image features.
[0102] In some cases, an image decoder decodes the output image features to obtain an output image 755 in a pixel space 710. In some cases, encoding the original image 705 to obtain image features can significantly reduce the inference time because the size of the image features in the latent space may be significantly smaller than the resolution of the image in the pixel space (e.g., 32, 64, etc. compared to 256, 512, etc.).
[0103] Figure 8 An example of a U-Net 800 according to aspects of the present disclosure is shown. According to some aspects, a diffusion model includes an ANN architecture referred to as a U-Net. In some cases, the U-Net 800 implements the reverse diffusion process described in Figure 7 reference.
[0104] According to some aspects, the U-Net 800 receives an input feature 805, where the input feature 805 includes an initial resolution and an initial number of channels, and processes the input feature 805 using an initial neural network layer 810 (e.g., a convolutional neural network layer) to produce an intermediate feature 815.
[0105] In some cases, the intermediate feature 815 is then downsampled using a downsampling layer 820 such that the downsampled feature 825 has a resolution less than the initial resolution and a number of channels greater than the initial number of channels.
[0106] In some cases, the process is repeated multiple times and then reversed. For example, the downsampled features 825 are upsampled using the upsampling process 830 to obtain upsampled features 835. In some cases, the upsampled features 835 are combined with intermediate features 815 having the same resolution and number of channels via a skip connection 840. In some cases, the combination of the intermediate features 815 and the upsampled features 835 is processed using a final neural network layer 845 to produce output features 850. In some cases, the output features 850 have the same resolution as the initial resolution and the same number of channels as the initial number of channels.
[0107] According to some aspects, the U-Net 800 receives additional input features to produce a conditional generated output. In some cases, the additional input features include a vector representation of an input prompt. In some cases, the additional input features are combined with the intermediate features 815 in the U-Net 800 at one or more layers. For example, in some cases, a cross-attention module is used to combine the additional input features with the intermediate features 815.
[0108] Image Generation Process
[0109] Figure 9 An example of an image processing method 900 according to aspects of the present disclosure is shown. In some examples, these operations are performed by a system including a processor that executes a set of codes to control the functional elements of the device. Additionally or alternatively, certain processes are performed using dedicated hardware. Generally, these operations are performed according to the methods and processes described in aspects of the present disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.
[0110] At operation 905, the system obtains a text prompt, a first attribute token, and a second attribute token. In some cases, the operation of this step refers to Figures 5 - 8 the image generation model described, or can be performed by the image generation model described in Figures 5 - 8 reference.
[0111] For example, at operation 905, the system retrieves the text prompt and selects two attribute tokens, namely, the first attribute token and the second attribute token. These two attribute tokens are used to guide the image synthesis process, and each attribute token affects a different attribute that the output image is to reflect.
[0112] At operation 910, the system identifies a first set of layers and a first set of time steps of the image generation model for the first attribute token, and identifies a second set of layers and a second set of time steps of the image generation model for the second attribute token. In some cases, the operation of this step refers to Figures 5 - 8 the image generation model described, or can be performed by the image generation model described inFigures 5 - 8 performed by the described image generation model.
[0113] For example, at operation 910, the system selects a first set of layers in the image generation model, the first set of layers being used to capture and interpret characteristics associated with the first attribute token. This can involve layers that specifically handle texture details or color patterns. Additionally, when these specific layers are effective in processing the first attribute, the system identifies a first set of time steps. Similarly, the system selects a second set of layers based on the ability of the second set of layers to reflect attributes related to the second attribute token, which can involve the depiction and arrangement of objects within the scene. It also determines a second set of time steps for these layers to operate effectively.
[0114] At operation 915, by providing the first attribute token to the first set of layers of the image generation model during the first set of time steps and the second attribute token to the second set of layers of the image generation model during the second set of time steps, the system uses the image generation model to generate a synthetic image based on the text prompt, the first attribute token, and the second attribute token. In some cases, the operation of this step refers to Figures 5 - 8 the described image generation model, or can be performed by Figures 5 - 8 the described image generation model.
[0115] For example, at operation 915, the system guides the image generation model to generate a synthetic image, focusing on the first attribute token processed by the first set of layers during the specified first set of time steps. Thus, the first attribute is integrated into the image with high fidelity. Additionally, the system enables the second set of layers to process the second attribute token with the corresponding second set of time steps, where the modification of the second attribute depends on the corresponding second set of time steps. By providing the first attribute token and the second attribute token to their respective layer sets and corresponding time steps, the image generation model can construct a synthetic image that not only reflects the specified attributes but also closely aligns with the text narrative provided in the prompt.
[0116] In some examples, the system can select attribute tokens from a set including color, object, style, and layout type. The first attribute token can represent a color attribute from a reference image, while the second attribute token may correspond to an object attribute depicting a specific item within the image.
[0117] In some examples, the generated synthetic image includes elements from the text prompt, where the first attribute is represented by the first attribute token and the second attribute is represented by the second attribute token. For example, if the text prompt specifies "garden", the system applies the first attribute token for color to capture greenery and the second token for object to incorporate elements such as plants or flowers.
[0118] In some examples, the method assigns a first attribute token and a second attribute token to different sets of layers within the model. This non-overlapping assignment allows for clear and controlled modulation of the resulting image, where one set of layers affects the style attribute and another set of layers modulates the color attribute.
[0119] In some examples, generating a synthetic image involves a reverse diffusion process applied to a noise image. This reverse process occurs over several time steps, guided first by a first set of attribute tokens and then by a second set of attribute tokens, thereby building up the image details in a controlled order.
[0120] In some examples, a text prompt undergoes an encoding process to produce a text embedding. Then, based on this text embedding, a synthetic image is generated, aligning the final product with the text description and the visual attributes indicated by the attribute tokens.
[0121] Figure 10 Examples of an image processing method according to aspects of the present disclosure are shown. In some examples, these operations are performed by a system including a processor that executes a generation set of code to control the functional elements of the device. Additionally or alternatively, certain processes are performed using dedicated hardware. Generally, these operations are performed in accordance with the methods and processes described in aspects of the present disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.
[0122] At operation 1005, the system obtains a text prompt, a first attribute token, and a second attribute token, including receiving user input indicating the first and second attributes. In some cases, the operation of this step refers to Figures 5 - 8 the image generation model described, or can be performed by the image generation model described in Figures 5 - 8 reference.
[0123] For example, at operation 1005, the system begins by obtaining inputs for image synthesis. This includes receiving a text prompt that serves as a narrative guide for the desired image and a reference image including the first and second attribute tokens. For example, these attribute tokens are identified through user input, where the user specifies the desired attributes of the image. The system captures these user-defined attributes as tokens that will guide the synthesis process to generate an image that embodies these specified characteristics.
[0124] At operation 1010, the system uses the image generation model to generate a synthetic image based on the text prompt. In some cases, the operation of this step refers to Figures 5 - 8 the image generation model described, or can be performed by the image generation model described in Figures 5 - 8 reference.
[0125] For example, at operation 1010, the system continues to use an image generation model to create a synthetic image. This process is anchored by a previously received text prompt. For example, the generation model utilizes the correlation between learned text and visual content to initiate image generation.
[0126] At operation 1015, the system includes in the synthetic image the elements described by the text prompt, the first attribute represented by the first attribute token, and the second attribute represented by the second attribute token. In some cases, the operation of this step refers to Figures 5 - 8 the image generation model described, or can be performed by Figures 5 - 8 the image generation model described.
[0127] For example, at operation 1015, the descriptive elements of the text prompt are visibly embodied in the image. Additionally, both the first attribute encapsulated by the first attribute token and the second attribute encapsulated by the second attribute token have prominent features. The system integrates these attributes into the image such that the final synthesized output visually embodies the elements described by the text and the first and second attributes selected by the user. Thus, the system generates a synthetic image that is a coherent visual representation containing the distinctive elements indicated by the input.
[0128] Training
[0129] Figure 11 An example of a diffusion process according to aspects of the present disclosure is shown. The example shown includes a forward diffusion process 1105 (such as Figure 7 the forward diffusion process described) and a reverse diffusion process 1110 (such as the first and second reverse diffusion processes described with reference to Figure 7 ). In some cases, the forward diffusion process 1105 adds noise to an image (or image features in the latent space). In some cases, the reverse diffusion process 1110 denoises an image (or image features in the latent space) to obtain a denoised image.
[0130] According to some aspects, according to a known variance schedule 0 < β 1 < β 2 < … < β T < 1, the forward diffusion component iteratively adds Gaussian noise to the input at each diffusion step t using the forward diffusion process 1005:
[0131]
[0132] According to some aspects, the Gaussian noise is extracted from a Gaussian distribution with mean and variance σ 2 = β t ≥ 1, by Sampling is performed and settings are made Therefore, starting from the initial input x 0 the forward diffusion process 1005 generates x 1 ,…, x t ,… x T , where x T is pure Gaussian noise.
[0133] In some cases, a Markov chain is used to map the observed variable x 0 (such as, the original image 1130) to intermediate variables x 1 、…、x T , where the intermediate variables x 1 、…、x T have the same dimension as the observed variable x 0 . In some cases, the Markov chain gradually adds Gaussian noise to the observed variable x 0 or the intermediate variables x 1 、…、x T to obtain an approximate posterior q(x 1:T ∣x 0 ).
[0134] According to some aspects, during the reverse diffusion process 1110, the diffusion model gradually removes noise from x T to obtain a prediction of the observed variable x 0 (e.g., a representation of what the diffusion model thinks the original image 1130 should be). In some cases, this prediction is influenced by a guidance prompt or a guidance vector (e.g., referring to Figure 6 the prompts or prompt embeddings described). However, the diffusion model does not know the conditional distribution p(x 0 ∣x t-1 ) of the observed variable x t , because calculating the conditional distribution requires knowing the distribution of all possible images. Therefore, the diffusion model is trained to approximately calculate (e.g., learn) the conditional probability distribution p t-1 (x t ∣x θ ) of the conditional distribution p(x t-1 ∣x t ):
[0135]
[0136] In some cases, the mean of the conditional probability distribution p θ (x t-1 ∣x t ) is parameterized by μ θ , and the conditional probability distribution p θ (xt-1 |x t ) has a variance parameterized by Σ θ In some cases, the mean and variance are conditioned on a noise level t (e.g., the amount of noise corresponding to a diffusion step t). According to some aspects, the diffusion model is trained to learn the mean and / or variance.
[0137] According to some aspects, the diffusion model initiates a reverse diffusion process 1110 with noise data x T (such as, noise image 1115). According to some aspects, the diffusion model iteratively denoises the noise data x T to obtain a conditional probability distribution p θ (x t-1 |x t ). For example, in some cases, at each step t - 1 of the reverse diffusion process 1110, the diffusion model takes x t (such as, a first intermediate image 1120) and t as inputs (where t represents a step in a series of transformations associated with different noise levels), and iteratively outputs a prediction of x t-1 (such as a second intermediate image 1125) until the noise data x T is recovered as a prediction of the observed variable x 0 (e.g., a predicted image of the original image 1130). According to some aspects, the joint probability of a sequence of samples in a Markov chain is determined by the product of conditional and marginal probabilities:
[0138]
[0139] In some cases, is a pure noise distribution because the reverse diffusion process 1110 takes the result of the forward diffusion process 1105 (e.g., samples of pure noise x T ) as input, and represents a sequence of Gaussian transformations corresponding to a sequence of adding Gaussian noise to the samples.
[0140] Figure 12 illustrates an example of a method for computing input conditioning according to aspects of the present disclosure. See Figure 12 , the input prompt is transformed into a conditioning vector. For example, i ∈ [1, 5] corresponds to Figure 12 a five - layer subset in ij , while j ∈ [1, 4] corresponds to four time - step phases. Thus, P
[0141] See Figure 12 , the learnable token<c> 、 <o> 、 <s>And <l>is included in the input prompts across multiple layers and time-step phases. For example, object properties are captured at the coarse layer (L 6 -L 9 ) and during the intermediate $t 2 , $t 3 denoising-backward phase. Thus, the token <o>can be designed to be conditioned only on the rough layer and the forward t′ 2 ,t′ 3 Diffusion phase. For example, during the sampling time step t ∈ t′3 in the forward diffusion process, <o>It may affect the final conditioning vector across the coarse U-Net layers. It can be derived from Figure 12 that <c> 、 <s>And <l>Similar observations for individual tokens. Subsequently, during the conversion, for a particular backpropagation, depending on the sampled time step, only the embeddings corresponding to the active tokens from Figure 12 can be optimized.
[0142] Embodiments of the present disclosure provide the MATTE algorithm. The learning objective of MATTE consists of three parts. The initial component is the standard reconstruction loss, where the function represents the learnable embeddings of a subset of tokens, which depends on the time step sampled during the forward diffusion process:
[0143]
[0144] Here, p j includes for <c> 、 <o> 、 <s>And <l>Learnable embeddings of subsets of tokens, which depend specifically on the time step \(t\in[0,1000]\) sampled during the forward diffusion process. Additionally, since color and style attributes are captured across similar layers and time step phases, an additional color-style disentanglement loss is provided to facilitate the disentanglement of these tokens:
[0145]
[0146] For the encoding vector of token \(c\), it is paired with \(s\), which is the encoding vector of a style randomly selected from a set of latent styles (such as watercolor, graffiti, oil painting, etc.). Additionally, \(c\) is also associated with the CLIP embedding 1205 of some ground truth colors in the reference image. The encoding 1210 is generated based on \(P\) ij′ where \(i = [1,5]\) and \(j = j'\), which is according to one of the 4 phases of the sampled \(t\).
[0147] Here, \(c\) is the token <c>The encoded vector, s is the encoded vector of any style randomly selected from a set of 30 styles (such as watercolor painting, graffiti, oil painting), and c gt is the CLIP embedding 1205 of all ground-truth colors in the reference image, and this embedding can be extracted using some datasets or libraries. The basic intuition is to align the learned embedding cc of the token c with the token c by ensuring that both are equidistant from s. This process makes the embedding of c closer to the CLIP color feature space, making it distinguishable from the embedding for s.
[0148] In addition, embodiments of the present disclosure ascertain from analysis and visualization that object and layout information is captured by the same set of rough U-Net layers. To further distinguish object and layout tokens, a regularization method for learning layout tokens is proposed. This regularization ensures that layout tokens respect the classes of objects depicted in the reference image. This is achieved by calculating the CLIP vector of the ground-truth class label and making it closely aligned with the vector of the layout token. This is represented by the following formula:
[0149]
[0150] Here, o is the token <o>The learning vector, and o gt is the true value. The overall loss function of MATTE is expressed as follows, where λ CS = λ O = 0.1:
[0151] L inv = L R + λ CS L CS + λ O L O (7)
[0152] In Figures 13 - 14 a method for training a machine learning model is described. One or more aspects of the method include: obtaining training data, the training data including a first attribute token representing a first attribute and a second attribute token representing a second attribute; identifying a first set of layers of an image generation model and a first set of time steps for the first attribute token, and identifying a second set of layers of the image generation model and a second set of time steps for the second attribute token; and training the image generation model to generate a synthetic image including the first attribute and the second attribute by providing the first attribute token to the first set of layers of the image generation model during the first set of time steps and providing the second attribute token to the second set of layers of the image generation model during the second set of time steps.
[0153] Some examples of the method, apparatus, and non-transitory computer-readable medium further include performing a forward diffusion process on the training image to obtain a noisy input image. Some examples further include performing a reverse diffusion process on the noisy image to obtain a predicted image. Some examples further include comparing the predicted image with the training image.
[0154] Some examples of the method, apparatus, and non-transitory computer-readable medium further include: training the image generation model includes training the image generation model to generate a synthetic image including a color attribute, an object attribute, a style attribute, and a layout attribute.
[0155] Some examples of the method, apparatus, and non-transitory computer-readable medium further include: training the image generation model includes calculating a color style disentanglement loss.
[0156] Some examples of the method, apparatus, and non-transitory computer-readable medium further include: obtaining the training data further includes optimizing the first attribute token to represent the first attribute and optimizing the second attribute token to represent the second attribute.
[0157] Some examples of the method, apparatus, and non-transitory computer-readable medium further include: obtaining the training data further includes obtaining a training image and a text prompt describing the training image, wherein image generation is trained based on the training image and the text prompt.
[0158] Figure 13 An example of method 1300 for training a machine learning model including an image generation model in accordance with aspects of the present disclosure is shown. In some examples, these operations are performed by a system including a processor that executes a set of code to control the functional elements of the device. Additionally or alternatively, certain processes are performed using dedicated hardware. Generally, these operations are performed in accordance with the methods and processes described in aspects of the present disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.
[0159] At operation 1305, the system obtains training data including a first attribute token representing a first attribute and a second attribute token representing a second attribute. In some cases, the operation of this step refers to a machine learning model described in Figures 5 - 8 or may be performed by a machine learning model described in Figures 5 - 8
[0160] For example, at operation 1305, the system obtains a dataset including distinct images in which the first attribute token and the second attribute token are distinctively represented. The first attribute token corresponds to a particular attribute observed across the individual images in the dataset, such as color. The second attribute token similarly represents another distinct attribute, such as style. These tokens can be used as proxies for their respective attributes, enabling the system to identify and manipulate these attributes separately during the image generation process.
[0161] At operation 1310, the system identifies a first set of layers and a first set of time steps of the image generation model for the first attribute token, and identifies a second set of layers and a second set of time steps of the image generation model for the second attribute token. In some cases, the operation of this step refers to a machine learning model described in Figures 5 - 8 or may be performed by a machine learning model described in Figures 5 - 8
[0162] For example, at operation 1310, the system identifies the specific layers within the image generation model that are most effective for the attributes represented by the first and second tokens. The system identifies a first set of layers responsible for the type of information encoded by the first attribute token and also identifies a first set of time steps. Similarly, the system determines a second set of layers customized for the attribute of the second token and a second set of time steps that ensure optimal engagement of these layers in the attribute representation.
[0163] At operation 1315, the system trains an image generation model to generate a synthetic image including a first attribute and a second attribute by providing a first attribute token to a first set of layers of the image generation model during a first set of time steps and a second attribute token to a second set of layers of the image generation model during a second set of time steps. In some cases, the operation of this step refers to the training component described in Figures 5 - 8 or can be performed by the training component described in Figures 5 - 8 .
[0164] For example, at operation 1315, the system trains the image generation model with a specified set of layers and time steps. This involves inputting a first attribute token into the first set of layers at a specified first set of time steps to enable the model to have the ability to render an image including the first attribute. At the same time, a second attribute token is introduced into the second set of layers at a determined second set of time steps. This dual-input method allows the model to learn to accurately describe both attributes simultaneously. Through repeated training iterations, the model becomes good at generating synthetic images that truly display the first and second attributes defined by the corresponding tokens.
[0165] According to some embodiments, the image generation model is trained to generate synthetic images based on more than one attribute. The model is trained not only to replicate a single attribute but also to combine a series of attributes. For example, these attributes include color attributes that specify the chromatic composition of an image, object attributes that specify the central element within a scene, style attributes that specify the artistic and style rendering, and layout attributes that arrange the spatial distribution of elements within an image. Training the generation model involves training the model to generate complex images with the interactions of these attributes.
[0166] According to some embodiments, the image generation model is trained to minimize a color-style disentanglement loss. This particular loss function is calibrated so that the model can independently distinguish and manipulate color and style attributes. The disentanglement loss quantifies the ability of the model to change one attribute (such as color) without inadvertently affecting the style of the generated image, thus maintaining the integrity of the style while changing the color scheme.
[0167] According to some embodiments, the acquisition of training data involves a refinement step for the attribute tokens. The first attribute token is specifically optimized to capture the essence of the first attribute with higher fidelity. Similarly, the second attribute token undergoes a fine-tuning process to enhance its representation of the second attribute. This optimization ensures that each token is a more precise manifestation of its corresponding attribute, thereby facilitating the image generation model to learn clearer and more distinct attributes.
[0168] According to some embodiments, the process of obtaining training data includes collecting training images and their corresponding text prompts. The image generation model uses this paired data to learn the correlation between text descriptions and visual attributes. This training scheme combines reference images and descriptive prompts, allowing the model to understand and replicate the depicted attributes when generating new images based on text input.
[0169] Figure 14 An example of training a machine learning model according to aspects of the present disclosure is shown. In some examples, these operations are performed by a system including a processor that executes a set of codes to control the functional elements of the device. Additionally or alternatively, dedicated hardware is used to perform certain processes. Generally, these operations are performed according to the methods and processes described in aspects of the present disclosure. In some cases, the operations described herein consist of various sub-steps or are performed in combination with other operations.
[0170] At operation 1405, the system performs a forward diffusion process on the training image to obtain a noisy input image. In some cases, the operation of this step refers to the machine learning model described in Figures 5 - 8 or can be performed by the machine learning model described in Figures 5 - 8
[0171] For example, at operation 1405, the system applies the forward diffusion process to the training image, thereby increasing the noise level within the image. This step can generate a series of gradually noisier images, resulting in a fully noisy representation. These noisy images provide a basis for the system to learn the complex patterns for converting the noisy image into a denoised image.
[0172] At operation 1410, the system performs a reverse diffusion process on the noisy image to obtain a predicted image. In some cases, the operation of this step refers to the machine learning model described in Figures 5 - 8 or can be performed by the machine learning model described in Figures 5 - 8
[0173] For example, at operation 1410, the system performs a reverse diffusion process on the noisy image generated from operation 1405. This process iteratively denoises the image, using the training parameters of the image generation model to restore the image to its original state or generate a new image sharing characteristics with the original image. This step is used to train the model to produce clear and detailed synthetic images from a noisy starting point.
[0174] At operation 1415, the system compares the predicted image with the training image. In some cases, the operation of this step refers to the training component described in Figures 5 - 8 or can be performed by the training component described in Figures 5 - 8
[0175] For example, at operation 1415, the system compares the resulting image of the reverse diffusion process with the original training image. The system uses this comparison to calculate the difference between the predicted image and the original image, thereby informing the optimization process of the image generation model.
[0176] Figure 15 An example of a computing device 1500 in accordance with aspects of the present disclosure is shown. The computing device 1500 includes one or more processors 1505, a memory subsystem 1510, a communication interface 1515, an I / O interface 1520, a user interface component 1525, and a channel 1530.
[0177] In some embodiments, the computing device 1500 is an example of the image generation device described with reference to Figure 1 and Figure 5 or includes aspects of the image generation device described with reference to Figure 1 and Figure 5 In some embodiments, the computing device 1500 includes one or more processors 1505 that may execute instructions stored in the memory subsystem 1510 to generate a synthetic image including a first attribute and a second attribute by providing a first attribute token to a first set of layers of the image generation model during a first set of time steps and a second attribute token to a second set of layers of the image generation model during a second set of time steps.
[0178] According to some aspects, the computing device 1500 includes one or more processors 1505. The processor 1505 is an example of the processor unit described with reference to Figure 5 or includes aspects of the processor unit described with reference to Figure 5 In some cases, the processor is an intelligent hardware device, e.g., a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device, discrete gate or transistor logic components, discrete hardware components, or a combination thereof.
[0179] In some cases, the processor is configured to operate a memory array using a memory controller. In other cases, the memory controller is integrated into the processor. In some cases, the processor is configured to execute computer-readable instructions stored in the memory to perform various functions. In some embodiments, the processor includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing.
[0180] According to some aspects, the memory subsystem 1510 includes one or more memory devices. The memory subsystem 1510 is the reference Figure 5 Examples of the described memory cells, or aspects including reference Figure 5 Examples of the described memory cells. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or hard disks. Examples of memory devices include solid-state memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software that includes instructions that, when executed, cause a processor to perform the various functions described herein. In some cases, the memory includes, among other things, a basic input / output system (BIOS) that controls basic hardware or software operations such as interacting with peripheral components or devices. In some cases, a memory controller operates the memory cells. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, the memory cells within the memory store information in the form of logical states.
[0181] According to some aspects, the communication interface 1515 operates at the boundary between communication entities such as the computing device 1500, one or more user devices, the cloud, and one or more databases and the channel 1530, and may record and process communications. In some cases, the communication interface 1515 is provided to enable the processing system to couple to a transceiver (e.g., a transmitter and / or a receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for the communication device via an antenna.
[0182] According to some aspects, the I / O interface 1520 is controlled by an I / O controller to manage the input and output signals of the computing device 1500. In some cases, the I / O interface 1520 manages peripheral devices that are not integrated into the computing device 1500. In some cases, the I / O interface 1520 represents a physical connection or port to an external peripheral device. In some cases, the I / O controller uses an operating system such as or other known operating systems. In some cases, the I / O controller represents a modem, a keyboard, a mouse, a touch screen, or a similar device, or interacts with a modem, a keyboard, a mouse, a touch screen, or a similar device. In some cases, the I / O controller is implemented as a component of a processor. In some cases, a user interacts with the device via the I / O interface 1520 or via a hardware component controlled by the I / O controller.
[0183] According to some aspects, the user interface component(s) 1525 enables a user to interact with the computing device 1500. In some cases, the user interface component(s) 1525 includes an audio device, such as an external speaker system, an external display device (such as a display screen), an input device (e.g., a remote control device that directly or through an I / O controller interacts with the user interface), or a combination thereof. In some cases, the user interface component 1525 includes a GUI.
[0184] Evaluation
[0185] Embodiments of the present disclosure demonstrate multi-attribute transfer from reference images. The proposed formula (7) is used to learn specific embeddings for attributes such as α, β, γ, and δ. By specifying α and β in the input, embodiments of the present disclosure generate an image of a cat in a watercolor style that follows the reference color. Similarly, an image of a bottle in a watercolor style is generated in γ color. Embodiments of the present disclosure correctly infer the layout and objects from the reference, resulting in an image of pebbles stacked together. Embodiments of the present disclosure identify the pencil sketch style (η) of a bird (θ) from the reference, resulting in an image that combines the two attributes.
[0186] Embodiments of the present disclosure provide a user study on the generated images. In an example, survey respondents were asked to select which set of images best represents the input constraints. The survey respondents were presented with a reference image, a text prompt, and a set of attributes from the reference image that ideally should be transferred to the final generated image. The results of this study indicate that people prefer the images generated by the method proposed in the present disclosure, highlighting the effectiveness of the proposed transformation technique in constrained text-to-image generation.
[0187] Thus, embodiments of the present disclosure provide an algorithm that learns attributes such as color, style, layout, and objects from a reference image. Then, the algorithm is used for attribute-guided text-to-image synthesis. The present disclosure conditions both the layer and the time-step dimensions. This approach results in a novel transformation algorithm that includes an explicit disentanglement enhancement regularizer. Evaluations show that the method according to embodiments of the present disclosure effectively extracts attributes from the reference image and successfully transfers these attributes to the new generation.
[0188] The descriptions and figures presented herein represent example configurations and do not represent all implementations within the scope of the claims. For example, operations and steps may be rearranged, combined, or otherwise modified. Additionally, structures and devices may be represented in block diagram form to show the relationships between components and to avoid obscuring the described concepts. Similar components or features may have the same name but may have different reference numerals corresponding to different figures.
[0189] Those skilled in the art can easily think of making some modifications to the present disclosure, and the principles defined herein can be applied to other variations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features described herein.
[0190] The described method can be implemented or executed by a device including a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof. The general-purpose processor can be a microprocessor, a conventional processor, a controller, a microcontroller, or a state machine. The processor can also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors with a DSP core, or any other such configuration). Thus, the functions described herein can be implemented in hardware or software and can be executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, these functions can be stored in the form of instructions or code on a computer-readable medium.
[0191] Computer-readable media includes non-transitory computer storage media and communication media, including any medium that facilitates the transfer of code or data. The non-transitory storage media can be any available media that can be accessed by a computer. For example, non-transitory computer-readable media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact discs (CDs) or other optical disc memories, magnetic disk memories, or any other non-transitory media for carrying or storing data or code.
[0192] In addition, a connecting component can be appropriately referred to as a computer-readable medium. For example, if code or data is transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, or microwave signals) from a website, server, or other remote source, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology is included in the definition of the medium. Combinations of media are also included within the scope of computer-readable media.
[0193] In the present disclosure and the following claims, the word "or" indicates an inclusive list, e.g., such that a list of X, Y, or Z represents X or Y or Z, or XY or XZ or YZ, or XYZ. In addition, the phrase "based on" is not used to represent a closed set of conditions. For example, a step described as "based on condition A" can be based on condition A and condition B. In other words, the phrase "based on" should be interpreted as "at least partially based on". In addition, "a" or "an" means "at least one".< / o> < / c> < / l> < / s> < / o> < / c> < / l> < / s> < / c> < / o> < / o> < / l> < / s> < / o> < / c>
Claims
1. A method comprising: Obtaining a text prompt, a first attribute word unit, and a second attribute word unit; Identifying a first set of layers and a first set of time steps of an image generation model for the first attribute word-gram, and identifying a second set of layers and a second set of time steps of the image generation model for the second attribute word-gram; as well as The image generation model is used to generate a synthetic image based on the text prompt, the first attribute word-gram, and the second attribute word-gram by providing the first attribute word-gram to the first set of layers of the image generation model during the first set of time steps and providing the second attribute word-gram to the second set of layers of the image generation model during the second set of time steps.
2. The method according to claim 1, wherein: The first attribute word-gram includes a first word-gram type, and the second attribute word-gram includes a second word-gram type, and wherein the first word-gram type and the second word-gram type are selected from a word-gram type set including a color word-gram type, an object word-gram type, a style word-gram type, and a layout word-gram type.
3. The method according to claim 1, wherein: The composite image includes the element described by the text hint, a first attribute represented by the first attribute word-gram, and a second attribute represented by the second attribute word-gram.
4. The method according to claim 3, wherein: The first attribute word-gram and the second attribute word-gram respectively include learnable word-grams corresponding to the first attribute and the second attribute.
5. The method according to claim 3, wherein obtaining the first attribute word-gram and the second attribute word-gram comprises: User input is received indicating the first attribute and the second attribute.
6. The method according to claim 1, wherein: The first set of layers does not overlap with the second set of layers, and the first set of time steps does not overlap with the second set of time steps.
7. The method of claim 1 , wherein generating the composite image comprises: A back diffusion process is performed on the noisy input image, wherein the back diffusion process is based on a plurality of time steps including the first set of time steps and the second set of time steps.
8. The method according to claim 1, further comprising: The textual cue is encoded to obtain a text embedding, wherein the composite image is generated based on the text embedding.
9. A method for training a machine learning model, comprising: Obtaining training data, wherein the training data includes a first attribute word-gram representing a first attribute and a second attribute word-gram representing a second attribute; Identifying a first set of layers and a first set of time steps of an image generation model for the first attribute word-gram, and identifying a second set of layers and a second set of time steps of the image generation model for the second attribute word-gram; as well as The image generation model is trained to generate a composite image including the first attribute and the second attribute by providing the first attribute tokens to the first set of layers of the image generation model during the first set of time steps and providing the second attribute tokens to the second set of layers of the image generation model during the second set of time steps.
10. The method of claim 9, wherein training the image generation model comprises: Perform a forward diffusion process on the training image to obtain a noisy input image; performing a back diffusion process on the noise image to obtain a predicted image; as well as The predicted image is compared to the training image.
11. The method of claim 9, wherein training the image generation model comprises: The image generation model is trained to generate a synthetic image including color attributes, object attributes, style attributes, and layout attributes.
12. The method of claim 9, wherein training the image generation model comprises: Training the image generation model includes computing a color-style disentanglement loss.
13. The method according to claim 9, wherein obtaining the training data further comprises: The first attribute word-gram is optimized to represent the first attribute and the second attribute word-gram is optimized to represent the second attribute.
14. The method according to claim 9, wherein obtaining the training data further comprises: A training image and a textual cue describing the training image are obtained, wherein the image generation is trained based on the training image and the textual cue.
15. An apparatus for image processing, comprising: at least one processor; at least one memory storing instructions executable by the at least one processor; as well as The device also includes an image generation model, which includes parameters stored in the at least one memory and is trained to generate a composite image including a first attribute and a second attribute by providing a first attribute word-gram to a first layer set of the image generation model during a first set of time steps and providing a second attribute word-gram to a second layer set of the image generation model during a second set of time steps.
16. The device according to claim 15, wherein: The image generation model comprises a U-Net architecture, and wherein the first set of layers and the second set of layers comprise different layers of the U-Net architecture.
17. The device according to claim 15, wherein: The first attribute comprises a color attribute or a style attribute, and the first set of layers comprises a medium-resolution layer of the image generation model.
18. The apparatus of claim 15, wherein: The second attribute comprises an object attribute or a layout attribute, and the second set of layers comprises a coarse resolution layer of the image generation model.
19. The apparatus of claim 15, wherein: The first attribute comprises a color attribute, a style attribute, or a layout attribute, and the first set of time steps comprises an initial set of time steps.
20. The apparatus of claim 15, wherein: The second attributes include object attributes, and the second set of time steps includes a set of intermediate time steps.