Image re-illumination using machine learning

By combining the low-rank adaptive layer (LoRA) and the color transformation function, accurate re-illuminated images are generated, which solves the problems of low efficiency and multi-model training in the existing technology and achieves efficient and accurate image re-illumination and background generation.

CN120707659APending Publication Date: 2025-09-26ADOBE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146625.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-15
Filing Date
2025-02-10
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing image rendering process is inefficient and expensive. Conventional machine learning models require training multiple models to generate accurate re-illuminated images, and background generation is easily affected by the deviation of foreground objects, resulting in insufficient background diversity.

Method used

A low-rank adaptive layer (LoRA) is used to generate image features, and combined with a color transformation function, a single model is used to generate accurate re-illuminated images, including objects and backgrounds, avoiding the cost and inefficiency of training multiple models and reducing generation bias.

Benefits of technology

The efficiency and accuracy of image relighting are improved, and the generated images are more consistent with the expected lighting conditions and have more diverse backgrounds, avoiding the inefficiency and bias problems of conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707659A_ABST
    Figure CN120707659A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to image relighting using machine learning. A method, apparatus, non-transitory computer readable medium and system for image generation includes acquiring an input image and an input cue, wherein the input image depicts an object and the input cue describes a lighting condition for the object; generating a relighted image feature based on the input image and the input cue, where the relighted image feature represents an object having an illumination condition; and generating a composite image based on the relighted image features, where the composite image depicts the object having the illumination condition.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 569,897, filed in the U.S. Patent and Trademark Office on March 26, 2024, and corresponding U.S. Non-Provisional Application No. 18 / 949,023, filed on November 15, 2024, the disclosures of which are incorporated herein by reference in their entireties. Background Art

[0003] The following content relates generally to image generation, and more specifically to image relighting. Machine learning algorithms build models based on sample data, called training data, to make predictions or decisions in response to input without being explicitly programmed. One application area for machine learning is image generation.

[0004] Machine learning models can be used to generate images based on input guidance provided by text or images. Image relighting refers to the process of replacing the lighting conditions of an input image with new lighting conditions in the relighted image. Summary of the Invention

[0005] Systems and methods are described for generating a relighted image using a low-rank adaptive layer of an image generation model. In one example, the low-rank adaptive layer effectively adapts the weights of a trained image generation model to perform a relighting task of generating relighted image features for image elements of an input image in a latent space. The relighted image features can be generated based on an input prompt (such as a text prompt or an image prompt) that describes desired lighting conditions for the image elements. The image generation model then decodes the relighted image features to obtain a relighted image that includes the image elements illuminated according to the desired lighting conditions. Furthermore, in some embodiments, the image generation model computes a color transformation function based on the relighted image features, and the image generation system obtains the relighted image by applying a color transformation predicted by the color transformation function to the input image.

[0006] By generating images based on relighted image features and, in some embodiments, based on color transformation functions, the image generation model provides relighted images that more efficiently and accurately depict relighted image elements than conventional machine learning models.

[0007] Additionally, in some embodiments, the relighted image includes a background generated by an image generation model based on an input prompt. By using a model to generate both the relighted image element and the background based on the same input prompt, aspects of the present disclosure avoid the expense and inefficiency of using at least two machine learning models to complete similar composite tasks.

[0008] This summary introduces a selection of concepts in a simplified form that will be further described in the detailed description below. Therefore, this summary is not intended to identify essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The detailed description is described with reference to the accompanying drawings. Entities shown in the drawings refer to one or more entities and thus the discussion may refer to the singular or plural of the entities interchangeably.

[0010] Figure 1 An example of an image generation system according to aspects of the present disclosure is shown.

[0011] Figure 2 An example of a method for image relighting according to aspects of the present disclosure is shown.

[0012] Figure 3 An example of an image generation system for generating a relighted image according to aspects of the present disclosure is shown.

[0013] Figure 4 An example of an image generation system for generating a composite image using a color transform function in accordance with aspects of the present disclosure is shown.

[0014] Figure 5 An example of an image generation system for generating a composite image background according to aspects of the present disclosure is shown.

[0015] Figure 6 An example of a guided diffusion model according to aspects of the present disclosure is shown.

[0016] Figure 7 An example of a U-Net according to aspects of the present disclosure is shown.

[0017] Figure 8 An example of a method for generating a composite image according to aspects of the present disclosure is shown.

[0018] Figure 9 An example of a method for conditional image generation according to aspects of the present disclosure is shown.

[0019] Figure 10 An example of a diffusion process according to aspects of the present disclosure is shown.

[0020] Figure 11 An example of a method for training an image generation model according to aspects of the present disclosure is shown.

[0021] Figure 12An example of a flowchart depicting an algorithm as a step-by-step process for training a machine learning model is shown, in accordance with aspects of the present disclosure.

[0022] Figure 13 An example of a method for training a diffusion model according to aspects of the present disclosure is shown.

[0023] Figure 14 An example of a computing device according to aspects of the present disclosure is shown.

[0024] Figure 15 An example of an image generating device according to aspects of the present disclosure is shown. DETAILED DESCRIPTION

[0025] Overview

[0026] The following relates to image relighting using machine learning. Image relighting refers to the process of replacing the lighting conditions of an input image (e.g., the visual characteristics of the lighting included in the image) with new lighting conditions in the relighted image. Image relighting can be accomplished using image rendering or a machine learning process. However, conventional image rendering processes are unable to generate accurate relighted images or are inefficient due to the use of specialized and expensive image capture and rendering hardware and software, while conventional machine learning processes require training a full model to relight foreground objects, and require training at least two models to provide images that include relighted foreground and background objects. In addition, conventional machine learning models for generating image backgrounds are often affected by biases caused by foreground objects, which leads to a lack of diversity among the generated backgrounds.

[0027] Thus, aspects of the present disclosure generate a relighted image by encoding an input image depicting image elements to obtain image features in a latent space, and generating relighted image features using a low-rank adaptation (LoRA) layer of an image generation model based on the image features and lighting conditions described by an input prompt (e.g., a text prompt or an image prompt). The LoRA layer effectively adapts the weights of the trained image generation model to generate relighted image features according to the lighting conditions. The image generation model then decodes the relighted image features or, in some embodiments, applies a color transform computed based on the relighted image features by an additional color transform layer to the input image to obtain a relighted image comprising image elements having lighting according to the desired lighting conditions.

[0028] By generating relighted image features using low-rank adaptive layers and, in some embodiments, using color transformation functions, the image generation model provides a relighted image that more efficiently and accurately depicts relighted image elements than conventional machine learning models.

[0029] Additionally, in some embodiments, the relighted image includes a background generated by an image generation model based on an input prompt (e.g., using a diffusion process). In some embodiments, the image generation model generates the background using a second LoRA layer that is trained to generate images using a reverse diffusion process. By generating the relighted image element and the background based on the same input prompt, aspects of the present disclosure avoid the expense and inefficiency of training at least two machine learning models to perform similar tasks and avoid biasing the background in the relighted image element.

[0030] One aspect of the present disclosure is used in the context of image composition. For example, a user provides an input image depicting an object (e.g., a basketball) and a text prompt describing an expected setting for the object (e.g., "a crowded arena floor") to an image generation device of an image generation system. The image generation system generates a composite image by relighting the basketball included in the input image according to the lighting conditions implied by the text prompt and generating a composite background that at least partially surrounds the object based on the content described by the text prompt. As an example result, the composite image depicts the basketball resting on the floor of a crowded arena, and the lighting of the basketball makes the basketball appear to blend harmoniously with the background scene.

[0031] refer to Figure 1 and Figure 2 Further example applications of the present disclosure in the context of image synthesis are provided. Figure 1 、 Figure 3-Figure 7 and Figure 14-15 To provide an example reference for the process of image generation Figures 8-10 To provide an example reference for the process of training a machine learning model Figure 11-13 To provide.

[0032] Embodiments of the present disclosure improve upon conventional image generation systems by making the image relighting process more efficient and accurate. For example, some embodiments achieve this efficiency and accuracy by generating relighted image features for image objects using a LoRA layer of an image generation model, wherein the LoRA layer effectively adapts the weights of a trained image generation model to perform the feature generation process at a higher speed, and generates a composite image based on the relighted image features, wherein the composite image depicts the image object according to the input lighting conditions. In contrast, conventional image rendering processes for image relighting fail to generate accurate relighted images or are inefficient due to the use of specialized and expensive image capture and rendering hardware and software, while conventional machine learning processes for image relighting require training of a full model to relight foreground objects.

[0033] In addition, some embodiments of the present disclosure improve upon conventional image generation systems by using image generation processes performed by other layers of the image generation model to effectively generate backgrounds for synthetic images based on lighting conditions. In some embodiments, the other layers include one or more additional LoRA layers. In contrast, conventional machine learning processes for image relighting require training at least two models to provide images that include relighted foreground and background objects or require training each layer of the machine learning model to perform relighting and background functions. Furthermore, in some cases, one or more low-rank adaptive layers allow for avoiding generation biases that may be caused by image objects, thereby producing synthetic images that depict background scenes that are more diverse than those provided by conventional image generation systems.

[0034] Image Generation System

[0035] Figure 1 An example of an image generation system 100 according to aspects of the present disclosure is shown. The example shown includes the image generation system 100, an input image 140, an input prompt 145, and a composite image 150. The image generation system 100 is a reference Figure 3-Figure 5 Examples of corresponding elements described or including references Figure 3-Figure 5 In one aspect, the image generation system 100 includes an image generation device 105 , a cloud 120 , a database 125 , a user device 130 , and a user 135 . In one aspect, the image generation device 105 includes an image generation model 110 and a user interface 115 .

[0036] refer to Figure 1 According to some aspects, image generation device 105 obtains an input image (e.g., input image 140) and an input prompt (e.g., input prompt 145). The input image depicts an object and the input prompt describes lighting conditions for the object. In some aspects, the lighting conditions describe or imply at least one of color, brightness, shadow, and reflection properties.

[0037] For example, input image 140 depicts a butterfly and input hint 145 is the text hint "golden hour," suggesting lighting conditions that are characteristic (e.g., color, brightness, shadows, etc.) associated with the golden hour shortly after sunrise or before sunset. Figure 1 In the example of FIG. 1 , the user 135 provides the image generating apparatus 105 with an input image 140 and an input prompt 145 through the user interface 115 displayed by the image generating apparatus 105 on the user device 130 .

[0038] The image generation model 110 generates relighted image features based on the input image and the input prompt using a low-rank adaptive layer of the image generation model 110. The relighted image features represent an object with lighting conditions. In some aspects, the low-rank adaptive layer is included in a memory component (e.g., a reference Figure 15 In some aspects, the image generation model 110 also includes color transformation parameters that are trained to perform a color transformation function based on the relighted image features.

[0039] Image generation model 110 generates a composite image (e.g., composite image 150) based on the re-illuminated image features. The composite image depicts an object with the lighting conditions. In this example, composite image 150 depicts a butterfly from input image 140 according to the "golden hour" lighting conditions described by input prompt 145.

[0040] "Input prompt" refers to a text prompt (e.g., a text string) or an image prompt (e.g., an image) used to provide guiding information to a machine learning model. For example, in some cases, the prompt describes the expected lighting conditions and / or the content of the image to be generated.

[0041] "Lighting conditions" refers to information used to depict the lighting of an object included in an image. For example, in some cases, the apparent lighting of an object is affected by one or more of the object's color, reflectivity, shadows, etc. An object "illuminating with lighting conditions" refers to an object depicted as being illuminated according to the lighting conditions.

[0042] "Embedding" or "feature" refers to the representation of an object (e.g., an element) in a low-dimensional space (embedding space) such that semantic information about the object can be more easily captured and analyzed by machine learning models. For example, an embedding is a numerical representation of an object in a continuous vector space (embedding space) in which objects that include similar semantic information to one another correspond to numerically similar vectors and are therefore "closer" to one another, allowing similarities between different objects corresponding to different embeddings to be easily determined. "Embedding space" (or "vector space") refers to a mathematical set with embeddings (or vectors) as components and is characterized by dimensions that specify several independent directions in the embedding space.

[0043] A “synthetic image” refers to an image generated by an image generation model. A “synthetic background” refers to a background generated by an image generation model. A “background” may refer to a scene that at least partially surrounds or is covered by an object.

[0044] The image generating device 105 is a reference Figure 3-Figure 5 、 Figure 14 and Figure 15Examples of corresponding elements described or including references Figure 3-Figure 5 、 Figure 14 and Figure 15 According to some aspects, the image generation device 105 includes a computer-implemented network. In some embodiments, the computer-implemented network includes a machine learning model (such as a reference Figure 3-Figure 7 and Figure 15 The image generation model 110 described in further detail). The image generation device 105 may also include the image generation model 110 as shown in FIG. Figure 14 The one or more processors, memory subsystems, communication interfaces, I / O interfaces, one or more user interface components, and buses described herein are described. Additionally, the image generation apparatus 105 can communicate with a user device 130 and a database 125 via the cloud 120 .

[0045] According to some aspects, the image generation device 105 is implemented on a server. The server provides one or more functions to users connected via one or more networks in various networks, such as the cloud 120. The server may include a microprocessor board that includes a microprocessor responsible for controlling all aspects of the server. The server uses a microprocessor and protocols such as Hypertext Transfer Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), File Transfer Protocol (FTP), and Simple Network Management Protocol (SNMP) to exchange data with other devices or users on one or more networks in the network. The server can be configured to send and receive files in Hypertext Markup Language (HTML) format (e.g., for displaying web pages). In various embodiments, the server includes a general-purpose computing device, a personal computer, a laptop computer, a mainframe computer, a supercomputer, or any other suitable processing device.

[0046] The image generation model 110 is a reference Figure 3-Figure 7 and Figure 15 Examples of corresponding elements described or including references Figure 3-Figure 7 and Figure 15 According to some aspects, the image generation model 110 includes a memory unit (e.g., reference to Figure 15

[00106] In some aspects, the image generation model includes an artificial neural network (ANN) that is trained to generate the synthesized image.

[0047] For more details on the architecture of the image generation system, see Figure 3-Figure 7 and Figure 14-15 For more details about the image generation process, refer to Figure 2 and Figures 8-10For more details on the process used to train machine learning models, refer to Figure 11-13 To provide.

[0048] The cloud 120 is a computer network configured to provide computer system resource availability (such as data storage and computing power) on demand. The cloud 120 can provide resources without the need for active management by the user. The term "cloud" is sometimes used to describe a data center that is available to many users via the Internet. Some large cloud networks have functionality distributed across multiple locations from a central server. A server is designated as an edge server if it has a direct or close connection to a user. The cloud 120 can be limited to a single organization or available to multiple organizations. In one example, the cloud 120 includes a multi-layer communication network with multiple edge routers and core routers. In another example, the cloud 120 is based on a collection of local switches in a single physical location. According to some aspects, the cloud 120 provides communication between the image generation device 105, the database 125, and the user device 130.

[0049] Database 125 is an organized collection of data. In an example, database 125 stores data in a specified format, called a schema. According to some aspects, database 125 is structured as a single database, a distributed database, multiple distributed databases, or an emergency backup database. A database controller can manage the storage and processing of data in database 125. A user can interact with the database controller, or the database controller can operate automatically without user interaction. According to some aspects, database 125 is included in image generation device 105. According to some aspects, database 125 is external to image generation device 105 and communicates with image generation device 105 via cloud 120.

[0050] According to some aspects, user device 130 is a personal computer, a laptop computer, a mainframe computer, a palmtop computer, a personal assistant, a mobile device, or any other suitable processing device. User device 130 may include software that displays a user interface 115 (e.g., a graphical user interface) provided by image generation device 105. User interface 115 allows information (such as images, prompts, etc.) to be transferred between user 135 and image generation device 105.

[0051] According to some aspects, the user interface of the user device enables user 135 to interact with user device 130. In some embodiments, the user interface of the user device may include an audio device, such as an external speaker system, an external display device such as a display screen, or an input device (e.g., a remote control device that interfaces directly with the user interface or interfaces with an I / O controller module). In some cases, the user interface of the user device may be a graphical user interface.

[0052] The input image 140 is a reference Figure 3-Figure 5 Examples of corresponding elements described or including references Figure 3-Figure 5 Describes aspects of the corresponding element. Input prompt 145 is a reference Figure 3-Figure 5 Examples of corresponding elements described or including references Figure 3-Figure 5 The composite image 150 is a reference to the corresponding elements of the description. Figure 3-Figure 5 Examples of corresponding elements described or including references Figure 3-Figure 5 Describes aspects of the corresponding element.

[0053] Figure 2 An example of a method 200 for image relighting according to aspects of the present disclosure is shown. According to some aspects, an image generation system performs method 200 to generate a composite image depicting an object that is relighted according to lighting conditions described by an input prompt.

[0054] LoRA layer (e.g., reference Figure 3 and Figure 4 The LoRA layer described in the previous section) adapts the weights of the base model to perform the task in a parameter-efficient manner so that the base model does not need to be trained to perform the task. In some embodiments, the LoRA layer borrows the weights of the pre-trained image generation model (e.g., reference Figure 1 、 Figure 3-Figure 7 and Figure 15 The weights of the image generation model described in

[15] , such as U-Net, are borrowed and the relighted image features are generated in the latent space based on the input image and input prompt in the pixel space. Once the LoRA layer is trained (e.g., as described in

[15] ), the weights of the image generation model described in

[15] , such as U-Net, are borrowed and the relighted image features are generated in the latent space based on the input image and input prompt in the pixel space. Figure 11-13 As described above, an image generation model including a trained LoRA layer uses one time step (e.g., T=0) to predict the relighted image features. The image generation model can then generate a synthetic image depicting the relighted object based on the relighted image features by decoding the relighted image features from the latent space to the pixel space.

[0055] Additionally, in some embodiments, the image generation model uses a LoRA layer included in the encoder portion of the image generation model to generate relighted image features and uses a color transformation function to predict color transformation parameters based on the relighted image features. The image generation model can then generate a synthetic image depicting the relighted object by applying the predicted color transformation parameters to the input image.

[0056] In addition, in some embodiments, the image generation model uses image generation processes performed by other layers of the image generation model, such as a diffusion process, to generate a background for the composite image based on input prompts in parallel with or after generating the relighted object. In some embodiments, the other layers include one or more additional LoRA layers that adapt the weights of the image generation model to generate the background for the composite image using a diffusion process. Thus, the image generation system obtains a composite image in which the relighted object is accurately composited with the background, wherein the relighted object and the background are generated according to the same lighting conditions. Thus, in some embodiments, one or more LoRA layers allow the image generation model to function as a relighting module and a diffusion external filler.

[0057] At operation 205, the user provides an input image depicting an object and a prompt describing the lighting conditions. In some cases, the operation of this step refers to or can be performed by referring to Figure 1 For example, the user generates an image on a user device (such as a reference Figure 1 A user interface provided on a user device described in Figure 1 The user interface 115 described above) is provided to the image generating device (such as the reference Figure 1 The image generation device 105 described above provides an input image and a prompt. The prompt can be a text prompt (for example, "golden hour"), an image prompt, or other forms of prompts (such as audio).

[0058] At operation 210, the system generates a composite image that depicts the object with the lighting conditions. In some cases, the operation of this step refers to or can be represented by, for example, reference to Figure 1 、 Figure 3-Figure 5 、 Figure 14 and Figure 15 In the example, the image generating device generates a reference Figure 3 、 Figure 4 or Figure 5 Composite image as described.

[0059] At operation 215, the system provides the composite image to the user. In some cases, the operation of this step refers to or can be represented by reference to Figure 1 、 Figure 3-Figure 5 、 Figure 14 and Figure 15 In an example, the image generating device provides the composite image to the user via the user interface.

[0060] Figure 3An example of an image generation system 300 for generating a composite image 380 according to aspects of the present disclosure is shown. The illustrated example includes the image generation system 300, an input image 355, input image features 360, an input prompt 365, an input prompt embedding 370, relighted image features 375, and a composite image 380. In one aspect, the image generation system 300 includes an image generation device 305. In one aspect, the image generation device 305 includes an image generation model 310 and an input prompt encoder 350. In one aspect, the image generation model 310 includes a variational autoencoder (VAE) 315 and a U-Net 330. In one aspect, the VAE 315 includes a VAE encoder 320 and a VAE decoder 325. In one aspect, the U-Net 330 includes a U-Net encoder 335, a U-Net decoder 340, and a LoRA layer 345.

[0061] refer to Figure 3 , image generation system 300 generates a composite image (e.g., composite image 380) in pixel space based on relighted image features (e.g., relighted image features 375), where the relighted image features are generated based on an input image depicting an object (e.g., input image 355 depicting a butterfly) and an input hint that explicitly or implicitly describes the lighting conditions (e.g., input hint 365 describing the lighting condition "golden hour").

[0062] For example, the VAE encoder 320 of the VAE 315 generates image features (e.g., input image features 360) based on the input image. The image features are embeddings of the input image in a latent space (e.g., an embedding space). The input prompt encoder 350 (e.g., a text encoder, an image encoder, an encoder for another modality, or a multimodal encoder) generates input prompt embeddings (e.g., input prompt embeddings 370) based on the input prompts.

[0063] U-Net 330 includes a LoRA layer 345 with weights borrowed from U-Net encoder 335 and U-Net decoder 340. LoRA layer 345 generates relighted image features 375 based on the image features and the input hint embedding. Figure 3 T=0 indicates that the LoRA layer 345 predicts or generates the re-illuminated image features within one time step. The VAE decoder 325 decodes the re-illuminated image features from the latent space to the pixel space, thereby obtaining a synthesized image.

[0064] In some embodiments, the input image describes the object independently of other image elements. In some embodiments, the image generation device 305 generates the input image by extracting the object from another image (eg, using object detection).

[0065] The image generation system 300 is a reference Figure 1 、 Figure 4 and Figure 5 Examples of corresponding elements described or including references Figure 1 、 Figure 4 and Figure 5 The image generating device 305 is a reference to the corresponding elements of the description. Figure 1 、 Figure 4 、 Figure 5 、 Figure 14 and Figure 15 Examples of corresponding elements described or including references Figure 1 、 Figure 4 、 Figure 5 、 Figure 14 and Figure 15 The image generation model 310 is a reference to the corresponding elements of the Figure 1 、 Figure 4-Figure 7 and Figure 15 Examples of corresponding elements described or including references Figure 1 、 Figure 4-Figure 7 and Figure 15 Describes aspects of the corresponding element.

[0066] VAE 315 and VAE Encoder 320 are reference Figure 4 Examples of corresponding elements described or including references Figure 4 According to some aspects, the VAE 315 includes a memory unit (such as a memory unit of the image generation device 305) in the image generation device 305. Figure 15 Autoencoder parameters (e.g., machine learning parameters) stored in the memory unit 1510 described above.

[0067] A variational autoencoder (VAE) includes an ANN that is trained to encode input data into a low-dimensional latent space and then decode the encoded input data back into the original input space. In some cases, a VAE differs from other autoencoder implementations in that it imposes a probabilistic structure on the latent space.

[0068] According to some aspects, VAEs can generate new data samples by sampling from a learned latent space distribution, thereby generating new data points that are similar to the training data. VAEs are widely used in various applications, including image generation, data compression, and representation learning, due to their ability to learn rich probabilistic representations of high-dimensional data. VAEs provide a principled framework for generative modeling and have successfully generated realistic samples across different domains.

[0069] The VAE encoder 320 receives input data and outputs a mean vector and a variance vector, which represent the parameters of a probability distribution (such as a Gaussian distribution) of the input data in a latent space. In some cases, the VAE encoder 320 samples a latent vector from the mean vector and the variance vector using a reparameterization technique, where the latent vector is obtained by sampling from a standard normal distribution and then scaling and shifting the samples according to the mean vector and the variance vector. According to some aspects, the VAE decoder 325 reconstructs the original input data based on the latent vector.

[0070] According to some aspects, the VAE 315 is trained by optimizing a loss function that includes a reconstruction loss and a regularization term. In some cases, the reconstruction loss forces the decoder network to generate outputs similar to the original input, while the regularization term, such as the Kullback-Leibler (KL) divergence between the learned distribution and the prior distribution, forces the latent space to follow a specific distribution, such as a standard normal distribution. In some cases, the regularization term helps ensure that the latent space is well-structured and continuous.

[0071] U-Net 330 is a reference Figure 4 and Figure 7 Examples of corresponding elements described or including references Figure 4 and Figure 7 The U-Net encoder 335 is a reference to the corresponding elements of Figure 4 Examples of corresponding elements described or including references Figure 4 Describes aspects of the corresponding element.

[0072] According to some aspects, a U-Net is a type of ANN characterized by a U-shaped structure comprising a contracting path (e.g., encoder) and an expanding path (e.g., decoder). The contracting path comprises a series of convolutional and pooling layers that progressively reduce the spatial dimensions of the input image while increasing the number of feature channels. Each convolutional block may be followed by a rectified linear unit (ReLU) activation function, which introduces nonlinearity into the U-Net.

[0073] The expansion path consists of a series of upsampling and concatenation operations that gradually increase the spatial dimension of the feature map while reducing the number of feature channels. Each upsampling operation can be followed by a convolution block with ReLU activation. The output of each block in the expansion path is concatenated with the corresponding feature map from the contraction path, allowing the U-Net to retain detailed spatial information from earlier layers.

[0074] The U-Net architecture can include skip connections connecting corresponding layers between the contracting path and the expanding path, enabling the U-Net to bypass the loss of spatial information during downsampling and promote accurate localization of objects in the segmentation output. Depending on the number of classes in the segmentation task, the last layer of the U-Net can include a convolutional layer followed by a softmax or sigmoid activation function.

[0075] LoRA layer 345 is reference Figure 4 Examples of corresponding elements described or including references Figure 4 Aspects of the corresponding elements described. According to some aspects, a low-rank adaptation (LoRA) layer is an ANN component designed to adapt the weights of a pre-trained model (e.g., U-Net 330) to a new task or domain, thereby reducing computational complexity and memory requirements. In transfer learning, a model trained on a source domain is fine-tuned on a target domain (e.g., a specific dataset related to a specific task) to leverage knowledge learned from the source domain. However, due to differences in data distribution between the source and target domains, transferring the entire parameter set from a pre-trained model may not be optimal, resulting in poor performance or overfitting.

[0076] The low-rank adaptive layer addresses this problem by factorizing the weight matrix of a neural network layer into a low-rank matrix. This decomposition reduces the number of parameters in the low-rank adaptive layer, making it more adaptable to the target domain while retaining important features learned from the source domain. By reducing the rank of the weight matrix, the low-rank adaptive layer can more effectively capture the underlying structure of the data.

[0077] Methods for low-rank adaptation include singular value decomposition (SVD) and truncated SVD, which factorize the weight matrix into two or more low-rank matrices. These low-rank matrices are then used to initialize the weights of the adapted model.

[0078] By reducing the number of parameters in the model that includes low-rank adaptation, low-rank adaptation layers can speed up training and inference, thereby helping to achieve tasks in resource-constrained environments. By adapting the model's parameters to the target domain while retaining important features from the source domain, low-rank adaptation can lead to more efficient generalization performance on the target task. Low-rank adaptation can act as a form of regularization, mitigating overfitting by imposing constraints on the model's parameter space.

[0079] According to some aspects, LoRA layers are added to one or more layers of U-Net 330. According to some aspects, LoRA layers are not added to one or more criss-cross attention layers of U-Net 330.

[0080] Therefore, low-rank adaptation is a parameter-efficient fine-tuning technique for generating images with photorealistic quality and high diversity. In some embodiments, the LoRA layer 345 does not modify other layers included in the image generation model 310. In some embodiments, the LoRA layer 345 helps maintain text compliance and content understanding of the image generation model 310.

[0081] According to some aspects, U-Net 330 includes a first LoRA module and a second LoRA module, each of which includes one or more LoRA layers (e.g., LoRA layer 345). In some embodiments, the first LoRA module is initialized using the weights of U-Net 330 and is trained to predict re-illuminated image features using a single time step (e.g., T=0) without performing diffusion denoising. In some embodiments, at inference time, the weights of the first LoRA module are combined into U-Net 330, and U-Net 330 uses the weights of the first LoRA module and a single time step (e.g., without performing diffusion denoising) to predict re-illuminated image features.

[0082] In some embodiments, the second LoRa module is initialized using the weights of U-Net 330 and trained to perform a diffusion external filling process (e.g., a diffusion process for generating an image background). In some embodiments, at inference time, the weights of the second LoRa module are combined into U-Net 330 independently of the weights of the first LoRa module, and U-Net 330 generates the background of the image using the weights of the second LoRa module and the diffusion process.

[0083] Input prompt encoder 350 is reference Figure 4 and Figure 5 Examples of corresponding elements described or including references Figure 4 and Figure 5 According to some aspects, the input prompt encoder 350 includes encoding parameters (eg, machine learning parameters) stored in a memory unit of the image generation device 305 .

[0084] According to some aspects, the input prompt encoder 350 includes an ANN configured to generate text embeddings based on text input, such as a recurrent neural network (RNN) or a transformer. According to some aspects, the input prompt encoder 350 includes an ANN configured to generate image embeddings based on image embeddings, such as a CNN or a visual transformer (ViT). According to some aspects, the input prompt encoder 350 includes an ANN configured to generate multimodal embeddings for text input or image input in a multimodal embedding space.

[0085] Input image 355, input prompt 365 and composite image 380 are reference Figure 1 、 Figure 4 and Figure 5 Examples of corresponding elements described or including references Figure 1 、 Figure 4 and Figure 5 The input image features 360 and the relighted image features 375 are referenced Figure 4 Examples of corresponding elements described or including references Figure 4 The input hint embedding 370 is referenced Figure 4 and Figure 5 Examples of corresponding elements described or including references Figure 4 and Figure 5 Describes the aspect of the corresponding element.

[0086] Figure 4 An example of an image generation system 400 for generating a composite image 485 using a color transformation function in accordance with aspects of the present disclosure is shown. The illustrated example includes the image generation system 400, an input image 455, input image features 460, an input hint 465, an input hint embedding 470, relighted image features 475, color parameters 480, and a composite image 485.

[0087] In one aspect, the image generation system 400 includes an image generation device 405. In one aspect, the image generation device 405 includes an image generation model 410, an input hint encoder 445, and an image relighting component 450. In one aspect, the image generation model 410 includes a variational autoencoder (VAE) 415, a U-Net 425, and (multiple) color transformation layers. In one aspect, the VAE 415 includes a VAE encoder 420. In one aspect, the U-Net 425 includes a U-Net encoder 430 and a LoRA layer 435.

[0088] refer to Figure 4 According to some aspects, image generation system 400 generates a composite image (eg, composite image 485 ) by applying color parameters (eg, color parameters 480 ) to an input image (eg, input image 455 ).

[0089] For example, the VAE encoder 420 of the VAE 415 generates image features (e.g., input image features 460) based on the input image. The input prompt encoder 445 generates input prompt embeddings (e.g., input prompt embeddings 470) based on the input prompts.

[0090] U-Net 425 includes a LoRA layer 435 with weights borrowed from U-Net encoder 430. LoRA layer 435 generates relighted image features (e.g., relighted image features 475) based on the image features and the input hint embedding. Figure 4 T=0 indicates that the LoRA layer 435 generates the re-illuminated image features within one time step.

[0091] One or more color transform layers 440 (e.g., a parameterized head) use a color transform function to predict color parameters 480. For example, in some cases, the input image is an RGB image comprising three color channels, e.g., red, green, and blue channels, where each color channel is associated with the same number of color parameters (e.g., 32 for a total of 96 color parameters associated with the input image). The color transform layer 440 predicts the color parameters 480 according to the color transform function based on the relighted image features. For example, the color transform function can be viewed as a lookup table, such that the one or more color transform layers 440 maps color information (e.g., color intensity, brightness, etc.) to the input image for the color information provided by the relighted image features, thereby obtaining a composite image.

[0092] In some embodiments, the image relighting component 450 applies a color transform to each pixel of the input image using parameter coordination. In an example, the color parameters are three piecewise linear curves with 32 control points, and the image relighting component 450 applies these three linear curves independently to each color channel of the input image. Application of the three linear curves is a per-pixel operation that can be efficiently computed at any resolution. In some embodiments, the image relighting component 450 performs histogram smoothing on the composite image to remove artifacts in the composite image.

[0093] In some embodiments, the color transform function operates independently of the resolution of the input image, allowing the color transform function to scale to any output resolution so that the resolution of the composite image is independent of the resolution of the input image. In some embodiments, the user selects the resolution of the composite image using a user interface.

[0094] Therefore, Figure 3 and Figure 4 In comparison, U-Net 425 omits the U-Net decoder, and VAE 415 omits the VAE encoder. Instead of decoding the re-illuminated image features to obtain a synthesized image, the image generation device 405 predicts color parameters based on the re-illuminated image features and applies the predicted color parameters to the input image to obtain a synthesized image.

[0095] Image generation system 400 is a reference Figure 1、 Figure 3 and Figure 5 Examples of corresponding elements described or including references Figure 1 、 Figure 3 and Figure 5 The image generating device 405 is a reference to the corresponding elements of the description. Figure 1 、 Figure 3 、 Figure 5 、 Figure 14 and Figure 15 Examples of corresponding elements described or including references Figure 1 、 Figure 3 、 Figure 5 、 Figure 14 and Figure 15 The image generation model 410 is a reference to the corresponding elements of the description. Figure 1 、 Figure 3 、 Figure 5-Figure 7 and Figure 15 Examples of corresponding elements described or including references Figure 1 、 Figure 3 、 Figure 5-Figure 7 and Figure 15 Describes aspects of the corresponding element.

[0096] VAE 415, VAE encoder 420, U-Net encoder 430 and LoRA layer 435 are reference Figure 3 Examples of corresponding elements described or including references Figure 3 The U-Net 425 is a reference to the corresponding elements. Figure 3 and Figure 7 Examples of corresponding elements described or including references Figure 3 and Figure 7 Describes aspects of the corresponding element.

[0097] Color conversion layer(s) 440 are included in a memory unit (such as reference Figure 15 Color transformation parameters (e.g., machine learning parameters) are stored in memory unit 1510 (described in detail in the preceding text). According to some aspects, color transformation layer 440 comprises a feedforward network. A feedforward network is an ANN in which the connections between ANN nodes do not form loops, so that information moves in one direction without any feedback loops, from the input nodes of the ANN, through the hidden layers (if any), and finally to the output nodes of the ANN. According to some aspects, color transformation layer 440 comprises one or more convolutional layers. According to some aspects, color transformation layer 440 comprises one or more linear layers.

[0098] According to some aspects, U-Net 425 and color transform layer(s) 440 comprise a parameterized U-Net that borrows layers from U-Net 425. According to some aspects, the output block of U-Net 425 is omitted. Input hint encoder 445 is referenced Figure 3 and Figure 5 Examples of corresponding elements described or including references Figure 3 and Figure 5 Describes aspects of the corresponding element.

[0099] Input image 455, input prompt 465 and composite image 485 are reference Figure 1 、 Figure 3 and Figure 5 Examples of corresponding elements described or including references Figure 1 、 Figure 3 and Figure 5 The input image features 460 and the relighted image features 475 are referenced Figure 3 Examples of corresponding elements described or including references Figure 3 The input prompt embed 470 is a reference to the corresponding elements of the description. Figure 3 and Figure 5 Examples of corresponding elements described or including references Figure 3 and Figure 5 Describes aspects of the corresponding element.

[0100] Thus, methods, systems, and non-transitory computer-readable media storing code for image processing include: obtaining an input image and an input prompt describing a lighting condition; generating an image generation model, wherein relighted image features are based on the input image and the input prompt; calculating a color transform function based on the relighted image features; and using the image generation model, generating a composite image based on the input image and the color transform, wherein the composite image depicts the input image with the lighting condition.

[0101] Figure 5 An example of an image generation system 500 for generating a composite image background according to aspects of the present disclosure is shown. The illustrated example includes the image generation system 500, an input image 520, an edit mask 525, an input hint 530, an input hint embedding 535, and a composite image 540. In one aspect, the image generation system 500 includes an image generation device 505. In one aspect, the image generation device 505 includes an image generation model 510 and an input hint encoder 515.

[0102] refer to Figure 5, the image generation model 510 generates a synthetic image (eg, synthetic image 540) based on an input image (eg, input image 520), the synthetic image including a re-illuminated object and a background. Figure 3 or Figure 4 Generated by the description.

[0103] The image generation model 510 receives as input an input image generated by the input prompt encoder 515 based on an input prompt (e.g., input prompt 530), an edit mask (e.g., edit mask 525), and an input prompt embedding (e.g., input prompt embedding 535), and uses an image generation process (such as reference image generation) to generate an image. Figure 6 and Figure 9-10 ), generating a composite image 540. The input prompt describes the content of the background.

[0104] The edit mask restricts the image generation process to generate new content in the area corresponding to the masked region, and thus the objects in the input image corresponding to the masked region are not affected by the image generation process. Figure 3 or Figure 4 The obtained re-illuminated object is composited with a synthetic background generated by the image generation process in the synthetic image. In some embodiments, the image generation device generates an edit mask based on the input image. In an example, the image generation device uses an object detection or bounding box algorithm or an ANN (such as Mask-R-CNN) to generate the mask.

[0105] In some embodiments, the input prompt used to generate the relighted object is also used in the image generation process. The input prompt can imply the lighting conditions (e.g., "a forest with green trees") or can directly describe the lighting conditions (e.g., "sneakers in orange lighting with a solid background"). Because the same input prompt is used to generate the relighted object and the background of the composite image, the relighted object and the background share the same lighting conditions.

[0106] Image generation system 500 is a reference Figure 1 、 Figure 3 and Figure 4 Examples of corresponding elements described or including references Figure 1 、 Figure 3 and Figure 4 The image generating device 505 is a reference to the corresponding elements of the description. Figure 1 、 Figure 3 、 Figure 4 、 Figure 14 and Figure 15 Examples of corresponding elements described or including references Figure 1 、 Figure 3 、 Figure 4 、 Figure 14 and Figure 15 The image generation model 510 is a reference to the corresponding elements of the description. Figure 1 、 Figure 3-Figure 4 、 Figure 6-Figure 7 and Figure 15 Examples of corresponding elements described or including references Figure 1 、 Figure 3-Figure 4 、 Figure 6-Figure 7 and Figure 15 The input prompt encoder 515 is a reference to the corresponding elements of the description. Figure 3 and Figure 4 Examples of corresponding elements described or including references Figure 3 and Figure 4 Describes aspects of the corresponding element.

[0107] The input image 520, the input prompt 530 and the synthesized image 540 are reference images. Figure 1 、 Figure 3 and Figure 4 Examples of corresponding elements described or including references Figure 1 、 Figure 3 and Figure 4 The input hint embed 535 is a reference to the corresponding element. Figure 3 and Figure 4 Examples of corresponding elements described or including references Figure 3 and Figure 4 Describes aspects of the corresponding element.

[0108] Figure 6 An example of a guided diffusion model 600 according to aspects of the present disclosure is shown. In some examples, the guided diffusion model 600 describes a reference Figure 15 The operation and architecture of the image generation model 1515 are described.

[0109] Diffusion models are a type of generative neural network that is trained to generate new data with features similar to those found in the training data. Specifically, diffusion models can be used to generate new images. Diffusion models can be used for a variety of image generation tasks, including image super-resolution, image generation using perceptual metrics, conditional generation (e.g., based on textual guidance), image interior filling, and image manipulation.

[0110] The diffusion model works by iteratively adding noise to the data during the forward process and then learning to recover the data by denoising the data during the backward process. For example, during training, the guided diffusion model 600 can take as input an original image 605 in pixel space 610 and apply an image encoder 615 to convert the original image 605 into an original image feature 620 in a latent space 625. Then, a forward diffusion process 630 gradually adds noise to the original image feature 620 to obtain noise features 635 (also in the latent space 625) at different noise levels.

[0111] Next, a back-diffusion process 640 (e.g., a U-Net ANN, such as that described in reference Figure 7 The U-Net described above gradually removes noise from the noise features 635 at various noise levels to obtain denoised image features 645 in the latent space 625. In some examples, the denoised image features 645 are compared with the original image features 620 at each of the various noise levels and the parameters of the back-diffusion process 640 of the diffusion model are updated based on the comparison. Finally, the image decoder 650 decodes the denoised image features 645 to obtain an output image 655 in the pixel space 610. In some cases, the output image 655 is created at each different noise level. The output image 655 can be compared with the original image 605 to train the back-diffusion process 640.

[0112] In some cases, the image encoder 615 and the image decoder 650 are pre-trained before training the back diffusion process 640. In some examples, they are jointly trained or the image encoder 615 and the image decoder 650 are fine-tuned jointly with the back diffusion process 640.

[0113] The back diffusion process 640 can also be guided based on the textual hint 660 or other guiding hints (such as images, layouts, segmentation maps, masks, etc.). The textual hint 660 can be encoded using an encoder 665 (e.g., a multimodal encoder) to obtain guided features 670 in a guided space 675. The guided features 670 can be combined with the noise features 635 at one or more layers of the back diffusion process 640 to ensure that the output image 655 includes the content described by the textual hint 660 or other guiding hint. For example, the guided features 670 can be combined with the noise features 635 using a cross-attention block within the back diffusion process 640.

[0114] Cross-attention, also known as multi-head attention, is an extension of the attention mechanism. In some cases, cross-attention enables the back-diffusion process 640 to focus on multiple parts of the input sequence simultaneously, thereby capturing interactions and correlations between different elements. In cross-attention, there are typically two input sequences: a query sequence and a key-value sequence. The query sequence represents the elements to receive attention, while the key-value sequence contains the elements to be attended to. In some cases, to compute cross-attention, the cross-attention block transforms (e.g., using linear projection) each element in the query sequence into a "query" representation, while the elements in the key-value sequence are transformed into "key" and "value" representations.

[0115] The crisscross attention block computes an attention score by measuring the similarity between each query representation and the key representation, where a higher similarity indicates more attention is paid to the key element. The attention score indicates the importance or relevance of each key element to the corresponding query element.

[0116] The crisscross attention block then normalizes the attention scores to obtain attention weights (e.g., using a softmax function), where the attention weights determine how much information from each value element is incorporated into the final attention representation. By simultaneously paying attention to different parts of the key-value sequence, the crisscross attention block captures relationships and correlations across the input sequence, allowing the back-diffusion process 1040 to better understand the context and generate more accurate and contextually relevant outputs.

[0117] Methods of operating diffusion models include denoising diffusion probabilistic models (DDPMs) and denoising diffusion implicit models (DDIMs). In DDPMs, the generation process includes inverting a random Markov diffusion process. On the other hand, DDIMs use a deterministic process so that the same input produces the same output. In some cases, DDIMs can reduce the number of time steps during image generation. Diffusion models can also be characterized by whether noise is added to the image itself or to image features generated by the encoder (i.e., latent diffusion). In a pixel diffusion model, noise is added and removed in pixel space. In a potential diffusion model, noise is added (and removed) in the latent space of image features rather than in pixel space. Therefore, the potential diffusion model uses reverse diffusion to generate image features and these image features can be decoded to obtain a synthetic image. In some embodiments, the guided diffusion model 600 is implemented as a guided pixel diffusion model.

[0118] Figure 7 An example of a U-Net 700 according to aspects of the present disclosure is shown. In some examples, the U-Net 700 is a reference Figure 6 Examples of components of the reverse diffusion process 640 of the guided diffusion model 600 are described and include reference to Figure 15Architectural elements of the image generation model 1515 are described. Figure 7 The U-Net 700 depicted in the reference Figure 6 Examples of architectures used within the described backdiffusion process or including references Figure 6 Describes various aspects of the architecture used within the back-diffusion process. U-Net 700 is a reference Figure 3 and Figure 4 Examples of corresponding elements described or including references Figure 3 and Figure 4 Describes aspects of the corresponding element.

[0119] In some examples, the diffusion model is based on a neural network architecture called U-Net. U-Net 700 takes input features 705 having an initial resolution and an initial number of channels and processes the input features 705 using an initial neural network layer 710 (e.g., a convolutional network layer) to produce intermediate features 715. The intermediate features 715 are then downsampled using a downsampling layer 720 so that the downsampled features 725 have a resolution smaller than the initial resolution and a number of channels greater than the initial number of channels.

[0120] This process is repeated multiple times and then the process is reversed. That is, the downsampled features 725 are upsampled using the upsampling process 730 to obtain upsampled features 735. The upsampled features 735 can be combined with the intermediate features 715 of the same resolution and number of channels via a skip connection 740. These inputs are processed using the final neural network layer 745 to produce output features 750. In some cases, the output features 750 have the same resolution as the initial resolution and the same number of channels as the initial number of channels.

[0121] In some cases, U-Net 700 requires additional input features to produce conditionally generated outputs. For example, the additional input elements may include vector representations of input prompts. The additional input features may be combined with intermediate features 715 at one or more layers within the neural network. For example, a cross-attention module may be used to combine the additional input features with the intermediate features 715.

[0122] Image Generation

[0123] Figure 8 An example of a method 800 for generating a composite image according to aspects of the present disclosure is shown. Figure 8 , image generation systems (such as reference Figure 1 and Figure 3-Figure 5The image generation system described herein executes method 800 to generate a composite image based on an input image and re-illuminated image features, such that objects included in the input image are included in the composite image and are re-illuminated (e.g., including one or more of different colors, reflective properties, shadows, brightness, etc.) according to lighting conditions included in an input prompt (such as a text prompt or an image prompt).

[0124] In some cases, the composite image includes a background that at least partially surrounds the object, where the content of the background is described by an input cue (such as a text cue or an image cue). Thus, in some cases, aspects of the present disclosure provide a unified model for diffusion-based exterior fill, relighting using a single base model with task-specific layers, end-to-end text conditioning for more diverse background generation and synthesis, and image- or text-guided harmony and exterior fill. In some cases, aspects of the present disclosure provide scalable parameterized harmony effects based on low-rank adaptation.

[0125] At operation 805, the system obtains an input image and an input prompt, wherein the input image depicts an object and the input prompt describes lighting conditions for the object. In some cases, the operation of this step refers to or can be performed by referring to Figure 1 、 Figure 3-Figure 5 、 Figure 14 and Figure 15 In an example, a user provides an input image and an input prompt to a user interface provided by the image generation device on the user device. In some cases, the input prompt is a text prompt including a text string.

[0126] In some cases, the input image depicts only the object and omits other content. In some cases, the input prompt is an image prompt that includes an image. In some cases, the input prompt directly describes the lighting conditions (e.g., "sunlit object"). In some cases, the input prompt indirectly implies the lighting conditions (e.g., "tree-lined forest"). In some embodiments, the lighting conditions include at least one of color, brightness, shadow, and reflection properties.

[0127] At operation 810, the system generates relighted image features based on the input image and the input prompt using a low-rank adaptive layer of the image generation model, wherein the relighted image features represent an object with lighting conditions. In some cases, the operation of this step refers to or can be performed by referring to Figure 1 、 Figure 3-Figure 7 and Figure 15 In the example, the image generation model generates the image as described in reference Figure 3 or Figure 4 The described reilluminated image characteristics.

[0128] According to some aspects, generating the relighted image features includes encoding an input prompt to obtain a prompt embedding, wherein the relighted image features are based on the prompt embedding. In some embodiments, the input prompt includes a text prompt. In some embodiments, the input prompt includes an image prompt. In some embodiments, generating the relighted image features includes encoding an input image to obtain an image embedding, wherein the relighted image features are based on the image embedding.

[0129] At operation 815, the system generates a composite image based on the re-illuminated image features using the image generation model, wherein the composite image depicts the object with the lighting conditions. In some cases, the operation of this step refers to or can be performed by referring to Figure 1 、 Figure 3-Figure 7 and Figure 15 In the example, the image generation model generates the image as described in reference Figure 3 、 Figure 4 or Figure 5 Composite image as described.

[0130] In an example, generating a composite image includes decoding features of the relighted image to obtain a composite image. In an example, generating the composite image includes computing a color transformation function based on the features of the relighted image, wherein the composite image is based on the color transformation function. In an example, generating the composite image includes generating a background for the composite image, wherein content of the background is described by an input prompt. In an example, the background is generated based on an edit mask. In an example, generating the composite image includes obtaining a noise map and denoising the noise map based on the input prompt to obtain the composite image.

[0131] In some cases, the image generation model uses the image generation parameters to perform a diffusion process (such as the reference Figure 6 and Figure 10 In some cases, the back diffusion process described in the preceding paragraph is used to obtain a composite image, wherein the composite image includes an object and a background that at least partially surrounds the object. In some cases, the generation of the background is guided by a hint and an edit mask so that diffusion does not occur in obscured areas corresponding to the object and occurs in obscured areas corresponding to the background. In some cases, the user provides the edit mask via a user interface. In some cases, the mask component generates the edit mask based on the input image so that the obscured area corresponds to the outline of the object and the unobscured area surrounds the obscured area.

[0132] According to some aspects, the image generation model generates a synthetic image in two stages. In some cases, the image generation model generates a relighted object for the synthetic image in a first stage and generates a background for the synthetic image in a second stage after the first stage. In some embodiments, the first stage includes using a first LoRA module to predict relighted image features using a single time step, and the second stage includes using a second LoRA module to generate a background for the synthetic image using a diffusion process. According to some aspects, the image generation model generates the synthetic image in a single image generation stage.

[0133] Figure 9 An example of a method 900 for conditional image generation according to aspects of the present disclosure is shown. In some examples, the method 900 describes a method of generating a conditional image according to aspects of the present disclosure. Figure 15 The operation of the image generation model 1515 described, such as applying reference Figure 6 The described guided diffusion model 600. In some examples, these operations are performed by a system that includes a processor that executes a code set to control functional elements of a device, such as Figure 6 The image generation model described in .

[0134] Additionally or alternatively, the steps of method 900 can be performed using dedicated hardware. Typically, these operations are performed according to the methods and processes described in various aspects of this disclosure. In some cases, the operations described herein are composed of various sub-steps or are performed in conjunction with other operations.

[0135] At operation 905, the user provides an input image and an input prompt that describes the content to be included in the generated image. For example, the user may provide the prompt "sneakers in orange lighting conditions with a solid background." In some examples, guidance may be provided in a form other than text, such as via an image, sketch, layout, or edit mask.

[0136] At operation 910, the system converts the text prompt (or other guidance) into a conditional guidance vector or other multidimensional representation. For example, the text can be converted into a vector or a series of vectors using a transformer model or a multimodal encoder. In some cases, the encoder for the conditional guidance is trained independently of the diffusion model.

[0137] At operation 915, a noise map comprising random noise is initialized. The noise map can be in pixel space or latent space. By initializing an image with random noise, different variations of the image comprising content described by the conditional guidance can be generated.

[0138] At operation 920, the system generates an image based on the noise map and the conditional guidance vector. For example, the image can be generated using a reference Figure 10is generated by the reverse diffusion process described in .

[0139] Figure 10 An example of a diffusion process 1000 is shown in accordance with aspects of the present disclosure. In some examples, the diffusion process 1000 describes a reference Figure 15 The operation of the image generation model 1515 is described.

[0140] As mentioned above Figure 6 As described above, using a diffusion model may involve a forward diffusion process 1005 for adding noise to an image (or a feature in a latent space) and a backward diffusion process 1010 for denoising the image (or feature) to obtain a denoised image. The forward diffusion process 1005 may be represented as q(x t ∣x t-1 ), and the reverse diffusion process 1010 can be expressed as p(x t-1 ∣x t In some cases, the forward diffusion process 1005 is used during training to generate images with continuously increasing noise, and the neural network is trained to perform the backward diffusion process 1010 (ie, to continuously remove noise).

[0141] In the example forward pass for the latent diffusion model, the model uses a Markov chain to map the observed variable x0 (in pixel space or latent space) to intermediate variables x1,…,x T As the latent variables are passed through a neural network such as U-Net, the Markov chain gradually adds Gaussian noise to the data to obtain an approximate posterior q(x 1:T |x0), where x1,…,x T has the same dimensions as x0.

[0142] The neural network can be trained to perform the reverse process. During the back diffusion process 1010, the model is trained from the noise data x T (such as the noisy image 1015) and denoise the data to obtain p(x t-1 ∣x t ). At each step t-1, the back diffusion process 1010 takes x t (such as the first intermediate image 1020) and t as input. Here, t represents the step in the transformation sequence associated with different noise levels. The back diffusion process 1010 iteratively outputs x t-1 , such as the second intermediate image 1025, until x T Restore to x0, i.e., the original image 1030. The reverse process can be expressed as:

[0143] p θ (x t-1 ∣xt ):=N(x t-1 ;μ θ (x t ,t),∑ θ (x t ,t)), (1)

[0144] The joint probability of a sequence of samples in a Markov chain can be written as the product of the conditional probability and the marginal probability:

[0145]

[0146] where p(x T )=N(x T ; 0,I) is a pure noise distribution, because the reverse process takes the result of the forward process (pure noise samples) as input, and Represents the Gaussian transformed sequence corresponding to a sequence with Gaussian noise added to the samples.

[0147] At inference time, the observation data x0 in pixel space can be mapped into the latent space as input and the generated data As the output, it is mapped from the latent space back to the pixel space. In some examples, x0 represents the original input image with low image quality, and the latent variables x1,…,x T represents a noisy image and Indicates the generated image has high image quality.

[0148] Thus, a method for relighting an image using machine learning is described. One or more aspects of the method include obtaining an input image and an input prompt, wherein the input image depicts an object and the input prompt describes lighting conditions for the object; generating, using a low-rank adaptive layer of an image generative model, relighted image features based on the input image and the input prompt, wherein the relighted image features represent the object with the lighting conditions; and generating, using the image generative model, a composite image based on the relighted image features, wherein the composite image depicts the object with the lighting conditions. In some aspects, the lighting conditions include at least one of color, brightness, shadows, and reflectance properties.

[0149] In some examples, generating the relighted image features includes encoding an input prompt to obtain a prompt embedding, wherein the relighted image features are based on the prompt embedding. In some aspects, the input prompt includes a text prompt. In some aspects, the input prompt includes an image prompt.

[0150] In some examples, generating the relighted image features includes encoding the input image to obtain an image embedding, wherein the relighted image features are based on the image embedding. In some examples, generating the relighted image features includes computing a color transformation function, wherein the relighted image features are based on the color transformation function. Some examples further include predicting one or more color parameters based on the color transformation function, wherein the synthesized image is based on the one or more color parameters.

[0151] In some examples, generating the composite image includes generating a background for the composite image, wherein content of the background is described by an input prompt. Some examples of the method also include generating the background for the composite image based on the edit mask.

[0152] In some examples, generating the composite image includes obtaining a noise map. Some examples also include denoising the noise map based on the input prompt to obtain the composite image.

[0153] In some examples, these operations are performed by a system that includes a processor that executes a code set to control the functional elements of the device. Additionally, certain processes are performed using dedicated hardware. Generally, these operations are performed according to the methods and processes described in various aspects of this disclosure. In some cases, the operations described herein are composed of various sub-steps or are performed in combination with other operations.

[0154] train

[0155] Figure 11 An example of a method 1100 for training an image generation model according to aspects of the present disclosure is shown. Figure 11 , image generation models (such as reference Figure 1 、 Figure 3-Figure 7 and Figure 15 An image generation model described herein is trained to generate synthetic images based on training images and a lighting input, where the synthetic images depict foreground objects under lighting conditions from the lighting input.

[0156] According to some aspects, a U-Net (such as the one in reference Figure 3-Figure 4 and Figure 7 The U-Net described in

[15] is implemented as a backbone for an image generation model and the U-Net is reused for an image-to-image generation training task to train an image generation model to apply lighting conditions including one or more of shading, brightness, shadow, and reflection properties to a given foreground object based on an input prompt (such as a text prompt or an image prompt).

[0157] In one example, the low-rank adaptive layer of the image generation model is derived from a U-Net (e.g., as shown in reference Figure 13In some embodiments, the low-rank adaptive layer operates at one time step during training, which is different from the iterative denoising process of the diffusion network (e.g., T=0).

[0158] At operation 1105, the system obtains a training set comprising a real image depicting an object, a lighting input indicating a lighting condition of the object, and a training image depicting the object under a second lighting condition. In some cases, the operation of this step refers to or can be performed by referring to Figure 15 The training components described are executed.

[0159] In some embodiments, the training component retrieves data from a database (such as a reference Figure 1 In some examples, the training component augments the colors of real images (e.g., randomly) to obtain training images. In some embodiments, the lighting input includes text, images, or other media items describing lighting conditions.

[0160] At operation 1110, the system trains an image generation model using the training set to generate a composite image based on the training image and the lighting input, wherein the composite image depicts the foreground object under the lighting conditions from the lighting input. In some cases, the operation of this step refers to or can be performed by referring to Figure 15 The training components described are executed.

[0161] In an example, the image generation model generates a predicted image based on a training image and a lighting input. In some embodiments, the training component generates a predicted image based on a pre-trained image generation model (e.g., a reference image). Figure 13 The image generation model is initialized by using the pre-trained image generation model described in

[15] and the LoRA adaptive layer is added to the pre-trained image generation model. The image generation model is based on the image generation model described in

[15] Figure 3 The training image and illumination input described in

[15] are used to generate re-illuminated image features using the LoRA layer, and based on the reference

[15] Figure 3 The training component calculates the reconstruction loss function based on the predicted image and the real image and updates the parameters of the LoRA layer based on the reconstruction loss function.

[0162] In some embodiments, the image generation model is based on a reference Figure 4 The training image and illumination input described above are used to generate re-illuminated image features using the LoRA layer and based on the reference Figure 4 The color transformation function of the color transformation layer described above generates a predicted image. The training component calculates a reconstruction loss function based on the predicted image and the real image and updates the parameters of one or more layers in the LoRA layer and the color transformation layer based on the reconstruction loss function.

[0163] Figure 12 An example of a flowchart depicting an algorithm as a step-by-step process 1200 for training a machine learning model according to aspects of the present disclosure is shown. In some embodiments, the process 1200 describes the operation of a training component 1525, which is described for configuring as described in reference Figure 15 Described image generation model 1515. Process 1200 provides one or more examples of generating training data, using the training data to train a machine learning model, and using the trained machine learning model to perform a task.

[0164] To get started in this example, the machine learning system collects training data (block 1202), which is used as the basis for training the machine learning model, i.e., it defines what is being modeled. Training data can be collected by the machine learning system from a variety of sources. Examples of training data sources include public datasets, service provider system platforms that expose application programming interfaces (e.g., social media platforms), user data collection systems (e.g., digital surveys and online crowdsourcing systems), etc. Training data collection can also include data augmentation and synthetic data generation techniques to expand and diversify the available training data, balancing techniques to balance the number of positive and negative examples, etc.

[0165] The machine learning system can also be configured to identify features related to the type of task for which the machine learning model is to be trained (block 1204). Examples of tasks include classification, natural language processing, generative artificial intelligence, recommendation engines, reinforcement learning, clustering, and the like. To this end, the machine learning system collects training data based on the identified features and / or, after collection, filters the training data based on the identified features. The training data is then used to train the machine learning model.

[0166] To train the machine learning model in the illustrated example, the machine learning model is first initialized (block 1206). Initialization of the machine learning model includes selecting a model architecture to be trained (block 1208). Examples of model architectures include neural networks, convolutional neural networks (CNNs), long short-term memory (LSTM) neural networks, generative adversarial networks (GANs), decision trees, support vector machines, linear regression, logistic regression, Bayesian networks, random forest learning, dimensionality reduction algorithms, boosting algorithms, deep learning neural networks, and the like.

[0167] A loss function is also selected (block 1210). The loss function is used to measure the difference between the output (i.e., prediction) of the machine learning model and the target value (e.g., represented by the training data) used to train the machine learning model. Additionally, an optimization algorithm is selected (block 1212) that is used in conjunction with the loss function to optimize the parameters of the machine learning model during training. Examples of optimization algorithms include gradient descent, stochastic gradient descent (SGD), and the like.

[0168] Initialization of the machine learning model also includes setting initial values ​​for the machine learning model (block 1214). Examples of initialization include initializing the weights and biases of the nodes to improve training efficiency and the computational resources consumed as part of the training. Hyperparameters for controlling the training of the machine learning model are also set. Examples of hyperparameters include regularization parameters, model parameters (e.g., the number of layers in a neural network), learning rates, batch sizes selected from the training data, and the like. Hyperparameters are set using various techniques, including randomization techniques, heuristics learned from other training scenarios, and the like.

[0169] The machine learning model is then trained by the machine learning system using the training data (block 1218). A machine learning model refers to a computer representation that can be tuned (e.g., trained and retrained) to approximate an unknown function based on the input of the training data. Specifically, the term machine learning model can include models that learn and predict from known data using an algorithm (e.g., using the model architecture described above) to learn and relearn by analyzing the training data to generate outputs that reflect the patterns and properties expressed by the training data.

[0170] Examples of training types include supervised learning using labeled data, unsupervised learning involving finding underlying structures or patterns in training data, reinforcement learning based on optimization functions (e.g., rewards and / or penalties), using nodes as part of "deep learning," and the like. A machine learning model can be configured, for example, to include multiple nodes that collectively form multiple layers. A layer can be configured, for example, to include an input layer, an output layer, and one or more hidden layers. Computations are performed by the nodes within the layer using hidden states through a system of weighted connections that are "learned" during training, for example, by using a selected loss function and backpropagation to optimize the performance of the machine learning model to perform the associated task.

[0171] As part of training the machine learning model, a determination is made as to whether a stopping criterion is met (decision block 1220), i.e., the stopping criterion is used to validate the machine learning model. The stopping criterion can be used to reduce overfitting of the machine learning model, reduce consumption of computational resources, and improve the ability of the machine learning model to handle previously unseen data (i.e., data not specifically included as examples in the training data). Examples of stopping criteria include, but are not limited to, a predefined number of epochs, verifying loss stability, achieving a performance improvement threshold, whether a threshold level of accuracy is met, or based on performance metrics such as precision and recall. If the stopping criterion is not met ("No" from decision block 1220), process 1200 continues to train the machine learning model using the training data in the example (block 1218).

[0172] If the stopping criteria is met ("yes" from decision block 1620), the trained machine learning model is then used to generate an output based on subsequent data (block 1222). The trained machine learning model is, for example, trained to perform the task described above and thus, once trained, is configured to perform the task based on subsequent data received as input and processed by the machine learning model.

[0173] Figure 13 An example of a method 1300 for training a diffusion model according to aspects of the present disclosure is shown. In some embodiments, the method 1300 describes the operation of a training component 1525, which is described as being configured as described in reference Figure 15 The image generation model 1515. According to some aspects, Figure 13 Operations are described for training one or more LoRA layers of an image generation model using weights of pretrained layers of the image generation model to perform a backdiffusion process.

[0174] In some examples, these operations are performed by a system that includes a processor that executes a set of codes to control functional elements of a device, such as Figure 6 The guided diffusion model described in .

[0175] Additionally or alternatively, some processes of method 1300 can be performed using dedicated hardware. Typically, these operations are performed according to the methods and processes described in various aspects of this disclosure. In some cases, the operations described herein are composed of various sub-steps or are performed in combination with other operations.

[0176] At operation 1305, the user initializes the untrained model. Initialization can include defining the architecture of the model and establishing initial values ​​for the model parameters. In some cases, initialization can include defining hyperparameters such as the number of layers, the resolution and channels of each layer block, the location of skip connections, etc.

[0177] At operation 1310, the system adds noise to the training image using an N-stage forward diffusion process. In some cases, the forward diffusion process is a fixed process in which Gaussian noise is continuously added to the image. In a latent diffusion model, Gaussian noise can be continuously added to features in the latent space.

[0178] At operation 1315, the system, at each stage n, starting from stage N, uses a back diffusion process to predict an image or image features at stage n-1. For example, the back diffusion process can predict the noise added by the forward diffusion process, and the predicted noise can be removed from the image to obtain a predicted image. In some cases, the original image is predicted at each stage of the training process.

[0179] At operation 1320, the system compares the image (or image features) predicted at stage n-1 (such as the image at stage n-1 or the original input image) with the actual image (or image features). For example, given observation data x, a diffusion model can be trained to compare the negative log-likelihood of the training data to -logp θ Minimize the variational upper bound of (x).

[0180] At operation 1325, the system updates the parameters of the model based on the comparison. For example, the parameters of the U-Net can be updated using gradient descent. The time-dependent parameters of the Gaussian transformation can also be learned.

[0181] Thus, a method for training a machine learning model is described. One or more aspects of the method include obtaining a training set comprising a real image depicting an object, a lighting input indicating a lighting condition of the object, and a training image depicting the object under a second lighting condition; and using the training set to train an image generation model to generate a composite image based on the training image and the lighting input, wherein the composite image depicts a foreground object under the lighting condition from the lighting input. In some examples, obtaining the training set includes augmenting the colors of the real image to obtain the training image.

[0182] In some examples, training the image generation model includes generating a predicted image based on the training image and the illumination input. Some examples also include calculating a reconstruction loss function based on the predicted image and the real image. Some examples also include updating parameters of the image generation model based on the reconstruction loss function.

[0183] In some examples, updating the parameters of the image generative model includes updating parameters of a low-rank adaptive layer. Some examples of the method also include initializing the image generative model based on a pretrained image generative model. Some examples also include adding the low-rank adaptive layer to the pretrained image generative model. In some examples, updating the parameters of the image generative model includes updating parameters of the image generative model corresponding to a color transformation function.

[0184] In some examples, these operations are performed by a system that includes a processor that executes a code set to control functional elements of the device. Additionally or alternatively, some processes are performed using dedicated hardware. Generally, these operations are performed according to the methods and processes described in various aspects of this disclosure. In some cases, the operations described herein are composed of various sub-steps or are performed in combination with other operations.

[0185] Image generating device

[0186] Figure 14 1 shows an example of a computing device 1400 according to aspects of the present disclosure. The computing device 1400 is a reference Figure 1 、 Figure 3-Figure 5 and Figure 15 Examples of image generation devices described herein include reference Figure 1 、 Figure 3-Figure 5 and Figure 15 In one aspect, computing device 1400 includes processor(s) 1405 , memory subsystem 1410 , communication interface 1415 , I / O interface 1420 , user interface component(s) 1425 , and channel 1430 .

[0187] In some embodiments, computing device 1400 is Figure 6 Examples of image generation models include Figure 6 In some embodiments, computing device 1400 includes one or more processors 1405 that can execute instructions stored in memory subsystem 1410 to perform image generation.

[0188] According to some aspects, computing device 1400 includes one or more processors 1405. In some cases, the processor is an intelligent hardware device, for example, a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or a combination thereof. In some cases, the processor is configured to operate the memory array using a memory controller. In other cases, the memory controller is integrated into the processor. In some cases, the processor is configured to execute computer-readable instructions stored in the memory to perform various functions. In some embodiments, the processor includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing.

[0189] According to some aspects, the memory subsystem 1410 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory devices include solid-state memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software, which includes instructions that cause the processor to perform the various functions described herein when executed. In some cases, the memory includes a basic input / output system (BIOS), etc., which controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, a memory controller operates the memory cells. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, the memory cells within the memory store information in the form of logical states.

[0190] According to some aspects, the communication interface 1415 operates at the boundary between the communication entities (such as the computing device 1400, one or more user devices, the cloud, and one or more databases) and the channel 1430 and can record and process communications. In some cases, the communication interface 1415 is provided to enable a processing system coupled to a transceiver (e.g., a transmitter and / or receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for the communication device via an antenna.

[0191] According to some aspects, I / O interface 1420 is controlled by an I / O controller to manage input and output signals for computing device 1400. In some cases, I / O interface 1420 manages peripheral devices that are not integrated into computing device 1400. In some cases, I / O interface 1420 represents a physical connection or port to an external peripheral device. In some cases, an I / O controller uses a controller such as 1420 or other known operating systems. In some cases, an I / O controller represents or interacts with a modem, keyboard, mouse, touch screen, or similar device. In some cases, an I / O controller is implemented as a component of a processor. In some cases, a user interacts with the device via I / O interface 1420 or via hardware components controlled by the I / O controller.

[0192] According to some aspects, user interface component(s) 1425 enable a user to interact with computing device 1400. In some cases, user interface component(s) 1425 include an audio device (such as an external speaker system), an external display device (such as a display screen), an input device (e.g., a remote control device that interfaces with the user interface directly or via an I / O controller), or a combination thereof. In some cases, user interface component(s) 1425 include a GUI.

[0193] Figure 15 An example of an image generating apparatus 1500 according to aspects of the present disclosure is shown. The image generating apparatus 1500 is a reference Figure 1 、 Figure 3-Figure 5 and Figure 14 Examples of corresponding elements described or including references Figure 1 、 Figure 3-Figure 5 and Figure 14 The image generating apparatus 1500 may include reference Figure 6 Examples or aspects of the guided diffusion model described and references Figure 7 In some embodiments, image generation device 1500 includes a processor unit 1505, a memory unit 1510, an image generation model 1515, an I / O module 1520, and a training component 1525. Training component 1525 updates the parameters of image generation model 1515 stored in memory unit 1510. In some examples, training component 1525 is located outside image generation device 1500.

[0194] The processor unit 1505 includes one or more processors. A processor is an intelligent hardware device, such as a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device, discrete gate or transistor logic components, discrete hardware components, or any combination thereof.

[0195] In some cases, processor unit 2505 is configured to operate a memory array using a memory controller. In other cases, the memory controller is integrated into processor unit 1505. In some cases, processor unit 1505 is configured to execute computer-readable instructions stored in memory unit 1510 to perform various functions. In some aspects, processor unit 1505 includes dedicated components for modem processing, baseband processing, digital signal processing, or transport processing. In some aspects, processor unit 1505 includes reference Figure 14 One or more processors 1405 are described.

[0196] Memory unit 1510 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory devices include solid-state memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software, which includes instructions that, when executed, cause at least one processor of processor unit 1505 to perform the various functions described herein.

[0197] In some cases, memory unit 1510 includes a basic input / output system (BIOS) that controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, memory unit 1510 includes a memory controller that operates the memory cells of memory unit 1510. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, the memory cells within memory unit 1510 store information in the form of logical states. According to some aspects, memory unit 1510 is a reference to Figure 14 An example of a memory subsystem 1410 is depicted.

[0198] According to some aspects, the image generation device 1500 uses one or more processors of the processor unit 1505 to execute instructions stored in the memory unit 1510 to perform the functions described herein. For example, the image generation device 1500 can perform operations including acquiring an input image and an input prompt, wherein the input image depicts an object and the input prompt describes a lighting condition for the object; using a low-rank adaptive layer of an image generation model to generate re-illuminated image features based on the input image and the input prompt, wherein the re-illuminated image features represent the object under the lighting condition; and using the image generation model to generate a synthetic image based on the re-illuminated image features, wherein the synthetic image depicts the object with the lighting condition.

[0199] The memory unit 1510 may include an image generation model 1515 that is trained to generate a synthetic image based on a training image and a lighting input. For example, after training, the image generation model 1515 may perform a reference Figures 8-10 The inference operation is performed to generate a synthetic image based on the re-illuminated image features, wherein the synthetic image depicts the object with the lighting conditions. The image generation model 1515 is a reference Figure 1 、 Figure 3 and Figure 5 Examples of corresponding elements described or including references Figure 1 、 Figure 3 and Figure 5 Describes aspects of the corresponding element.

[0200] In some embodiments, the image generation model 1515 is an artificial neural network (ANN), such as Figure 6 Guided diffusion model described and referenced Figure 7 The U-Net described above. An artificial neural network (ANN) can be a hardware or software component consisting of connected nodes (i.e., artificial neurons) that loosely correspond to neurons in the human brain. Each connection, or edge, transmits a signal from one node to another (just like a physical synapse in the brain). When a node receives a signal, it processes it and then transmits the processed signal to other connected nodes.

[0201] ANNs have many parameters, including weights and biases associated with each neuron in the network. These parameters control the degree of connectivity between neurons and influence the neural network's ability to capture complex patterns in data. These parameters, also known as model parameters or model weights, are variables that determine the behavior and characteristics of a machine learning model.

[0202] In some cases, the signals between nodes consist of real numbers, and the output of each node is computed as a function of its inputs. For example, a node may use other mathematical algorithms to determine its output, such as selecting the maximum value from its inputs, or any other suitable algorithm to activate a node. Each node and edge is associated with one or more node weights that determine how the signal is processed and transmitted. In some cases, nodes have thresholds below which the signal is not transmitted at all. In some examples, nodes are aggregated into layers.

[0203] The parameters of the image generation model 1515 can be organized into layers. Different layers perform different transformations on their inputs. The initial layer is called the input layer and the last layer is called the output layer. In some cases, the signal traverses certain layers multiple times. The hidden (or intermediate) layer includes hidden nodes and is located between the input layer and the output layer. The hidden layer performs nonlinear transformations on the inputs into the network. Each hidden layer is trained to produce a defined output that contributes to the joint output of the ANN's output layer. The hidden representation is a machine-readable data representation of the input that is learned from the ANN's hidden layers and generated by the output layer. As the ANN is trained, the ANN's understanding of the input improves, and the hidden representation gradually distinguishes itself from earlier iterations.

[0204] The training component 1525 can train the image generation model 1515. For example, the parameters of the image generation model 1515 can be learned or estimated from the training data and then used to make predictions or perform tasks based on the patterns and relationships learned in the data. In some examples, the parameters are adjusted during the training process to minimize a loss function or maximize a performance metric (e.g., as shown in FIG. Figure 11-13The goal of the training process can be to find optimal values ​​for the parameters that allow the image generation model 1515 to make accurate predictions or perform well at a given task.

[0205] Accordingly, the node weights can be adjusted to improve the accuracy of the output (i.e., by minimizing a loss that corresponds in some way to the difference between the current result and the target result). The weights of the edges increase or decrease the strength of the signal transmitted between the nodes. For example, during the training process, the algorithm adjusts the machine learning parameters according to an optimization technique such as gradient descent, stochastic gradient descent, or other optimization algorithm to minimize the error or loss between the predicted output and the actual target. Once the machine learning parameters are learned from the training data, the image generation model 1515 can be used to make predictions on new, unseen data (i.e., during inference).

[0206] I / O module 1520 receives input from image generating device 1500 and transmits output of image generating device 1500 to other devices or users. For example, I / O module 1520 receives input for image generating model 1515 and transmits output of image generating model 1515. According to some aspects, I / O module 1520 is a reference Figure 14 An example of I / O interface 1420 is described.

[0207] According to some aspects, training component 1525 comprises executable code (eg, software) stored in memory unit 1510, firmware, one or more hardware circuits, or a combination thereof.

[0208] Thus, systems and apparatus for image relighting using machine learning are described. One or more aspects of the systems and apparatus include a memory component and a processing device coupled to the memory component, the processing device configured to perform operations including: acquiring an input image and an input prompt, wherein the input image depicts an object and the input prompt describes a lighting condition for the object; generating, using a low-rank adaptive layer of an image generative model, relighted image features based on the input image and the input prompt, wherein the relighted image features represent the object with the lighting condition; and generating, using the image generative model, a synthetic image based on the relighted image features, wherein the synthetic image depicts the object with the lighting condition.

[0209] In some aspects, the low-rank adaptation layer includes image relighting parameters stored in the memory component. In some aspects, the image generation model also includes color transformation parameters trained to perform a color transformation function.

[0210] The descriptions and drawings described herein represent example configurations and do not represent all implementations within the scope of the claims. For example, operations and steps may be rearranged, combined, or otherwise modified. In addition, structures and devices may be represented in block diagram form to illustrate the relationships between components and avoid obscuring the concepts being described. Similar components or features may have the same name but may have different reference numerals corresponding to different figures.

[0211] Some modifications of the present disclosure will be apparent to those skilled in the art and the principles defined herein may be applied to other variations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0212] The methods described may be implemented or performed by a device including a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general purpose processor may be a microprocessor, a conventional processor, a controller, a microcontroller, or a state machine. A processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration). Thus, the functions described herein may be implemented in hardware or software and may be performed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored on a computer-readable medium in the form of instructions or code.

[0213] Computer-readable media include non-transient computer storage media and communication media, and communication media include any media that facilitate the transmission of code or data. Non-transient storage media can be any available media that a computer can access. For example, non-transient computer-readable media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc (CD) or other optical disc storage device, magnetic disk storage device or any other non-transient medium for carrying or storing data or code.

[0214] Additionally, a connecting component may be properly referred to as a computer-readable medium. For example, if code or data is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, or microwave signals), the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology is included in the definition of medium. Combinations of media are also included within the scope of computer-readable media.

[0215] In this disclosure and the appended claims, the word "or" indicates an inclusive list, so that, for example, a list of X, Y, or Z means X or Y or Z or XY or XZ or YZ or XYZ. Furthermore, the phrase "based on" is not used to indicate a closed set of conditions. For example, a step described as "based on condition A" can be based on both condition A and condition B. In other words, the phrase "based on" should be interpreted as "based at least in part on." Furthermore, the words "a" or "an" indicate "at least one."

Claims

1. A method for image generation, comprising: Obtaining an input image and an input cue, wherein the input image depicts an object and the input cue describes lighting conditions for the object; generating, using a low-rank adaptive layer of an image generation model, relighted image features based on the input image and the input prompt, wherein the relighted image features represent the object having the lighting conditions; and A composite image is generated based on the relighted image features using the image generation model, wherein the composite image depicts the object with the lighting conditions.

2. The method according to claim 1, wherein: The lighting condition includes at least one of the following: color, brightness, shadow and reflection properties.

3. The method of claim 1 , wherein generating the relighted image features comprises: The input hint is encoded to obtain a hint embedding, wherein the relighted image features are based on the hint embedding.

4. The method according to claim 3, wherein: The input prompt includes a text prompt or an image prompt.

5. The method according to claim 1, wherein: By adding the low-rank adaptive layer to a pre-trained image generation model, the image generation model is trained based on the pre-trained image generation model.

6. The method of claim 1 , wherein generating the relighted image features comprises: The input image is encoded to obtain an image embedding, wherein the relighted image features are based on the image embedding.

7. The method of claim 1 , wherein generating the composite image comprises: Based on the relighted image features, a color transformation function is calculated, wherein the composite image is based on the color transformation function.

8. The method according to claim 7, further comprising: Based on the color transformation function, one or more color parameters are predicted, wherein the composite image is based on the one or more color parameters.

9. The method of claim 1 , wherein generating the composite image comprises: A background of the composite image is generated, wherein content of the background is described by the input prompt.

10. The method according to claim 9, further comprising: Based on an edit mask, the background of the composite image is generated.

11. The method of claim 1 , wherein generating the composite image comprises: Get the noise map; as well as Based on the input prompt, the noise map is denoised to obtain the composite image.

12. A non-transitory computer-readable medium storing code for image processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: Get an input image and an input prompt describing the lighting conditions; generating, by an image generation model, relighted image features based on the input image and the input prompt; Calculating a color transformation function based on the re-illuminated image features; as well as A composite image is generated based on the input image and the color transform using the image generation model, wherein the composite image depicts the input image with the lighting conditions.

13. The non-transitory computer-readable medium of claim 12, the operations further comprising: Based on the color transformation function, one or more color parameters are predicted, wherein the composite image is based on the one or more color parameters.

14. The non-transitory computer-readable medium of claim 12, wherein: The relighted image features are generated using a low-rank adaptive layer of the image generation model.

15. The non-transitory computer readable medium of claim 14, wherein: By adding the low-rank adaptive layer to a pre-trained image generation model, the image generation model is trained based on the pre-trained image generation model.

16. The non-transitory computer readable medium of claim 15, wherein: The image generation model is trained based on a reconstruction loss.

17. The non-transitory computer-readable medium of claim 12, wherein generating the relighted image features comprises: The input image is encoded to obtain an image embedding, wherein the relighted image features are based on the image embedding.

18. A system for image generation, comprising: Memory components; as well as a processing device coupled to the memory component, the processing device configured to perform operations comprising: Obtaining an input image and an input cue, wherein the input image depicts an object and the input cue describes lighting conditions for the object; generating, using a low-rank adaptive layer of an image generation model, relighted image features based on the input image and the input prompt, wherein the relighted image features represent the object having the lighting conditions; and A composite image is generated based on the relighted image features using the image generation model, wherein the composite image depicts the object with the lighting conditions.

19. The system of claim 18, wherein: The low-rank adaptation layer includes image relighting parameters stored in the memory component.

20. The system of claim 18, wherein: The image generation model also includes color transformation parameters trained to perform a color transformation function.