Text-based Realistic Image Editing Using Diffusion Models
The method employs a text-image diffusion model to optimize text embeddings and fine-tune a diffusion model, addressing limitations in existing image editing technologies by enabling sophisticated, high-resolution semantic edits within the described framework.
Patent Information
- Application Number
- JP2024067535
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-04-18
- Filing Date
- 2024-04-18
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2044-04-18
AI Technical Summary
Existing methods for applying non-trivial semantic edits to real photographs are limited by the need for specific edit types, domain restrictions, or auxiliary inputs such as image masks or detailed text descriptions.
A computer-implemented method using a text-image diffusion model that receives an input image and target text, optimizes text embeddings, fine-tunes a diffusion model, and interpolates between embeddings to generate an edited image with desired semantic edits.
Enables sophisticated, high-resolution image editing that aligns with target text while maintaining the original image's background, structure, and composition, allowing for a wide variety of edits and semantically significant linear interpolation.
Smart Images

Figure 0007691548000004 
Figure 0007691548000005 
Figure 0007691548000006
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to digital image editing. More specifically, the present disclosure relates to applying text-based semantic edits to an image using only an input image and target text by leveraging a text-image diffusion model.
Background Art
[0002] Applying non-trivial semantic edits to real photographs has been a challenge in image processing. Many methods for text-based image editing suffer from drawbacks such as being limited to a specific set of edits (such as drawing on top of the image, adding objects, or style transfer), working only on images from a specific domain or images generated by synthesis, or requiring auxiliary inputs such as an image mask indicating the desired edit location, multiple images of the same object, or text describing the original image in addition to the input image.
Summary of the Invention
Means for Solving the Problems
[0003] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the description that follows, or may be learned from the description, or may be learned through practice of the embodiments.
[0004] One exemplary aspect of the present disclosure is directed to a computer-implemented method for editing an image. The method includes receiving, by a computing system, an input image and a target text, where the target text indicates a desired edit for the input image, and obtaining, by the computing system, a target text embedding based on the target text. The method also includes obtaining, by the computing system, an optimized text embedding based on the target text embedding and the input image, and fine-tuning, by the computing system, a diffusion model based on the optimized text embedding. The method includes interpolating, by the computing system, the target text embedding and the optimized text embedding to obtain an interpolated embedding, and generating, by the computing system, an edited image including the desired edit using the diffusion model based on the input image and the interpolated embedding.
[0005] Another exemplary aspect of the present disclosure is directed to a computing system for editing an image. The computing system includes one or more processors and a non-transitory computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations. The operations include receiving an input image and target text, where the target text indicates a desired edit for the input image, and obtaining a target text embedding based on the target text. The operations also include obtaining an optimized text embedding based on the target text embedding and the input image, and fine-tuning a diffusion model based on the optimized text embedding. The operations further include interpolating the target text embedding and the optimized text embedding to obtain an interpolated embedding, and generating an edited image including the desired edit using the diffusion model based on the input image and the interpolated embedding.
[0006] Another exemplary aspect of the present disclosure is directed to a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations include receiving an input image and target text, where the target text indicates a desired edit for the input image, and obtaining a target text embedding based on the target text. The operations also include obtaining an optimized text embedding based on the target text embedding and the input image, and fine-tuning a diffusion model based on the optimized text embedding. The operations further include interpolating the target text embedding and the optimized text embedding to obtain an interpolated embedding, and generating an edited image including the desired edit using the diffusion model based on the input image and the interpolated embedding.
[0007] Other aspects of the present disclosure are directed to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0008] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate exemplary embodiments of the present disclosure and, together with the description, serve to explain the relevant principles.
[0009] A detailed discussion of embodiments directed to those of ordinary skill in the art is described herein with reference to the accompanying drawings.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Best Mode for Carrying Out the Invention
[0011] Reference numerals repeated across multiple figures identify the same features in various implementations.
[0012] Overview Generally, the present disclosure is directed to semantic image editing given only a single text prompt that describes an input image to be edited and a target edit.
[0013] The text-to-image diffusion model uses a diffusion model capable of high-quality image synthesis. When conditioned on a natural language text prompt, the diffusion model can generate an image that sufficiently matches the requested text. The present invention adapts the diffusion model to edit real-world images rather than synthesize new images.
[0014] This is performed in a three-step process. First, the text embedding is optimized, resulting in a text embedding that best matches the given input image, which is near the target text embedding. Second, the diffusion model is fine-tuned to better match the given input image or reconstruct the given input image. Finally, the optimized embedding and the target text embedding are linearly interpolated to find a point that achieves both fidelity to the image and the target text alignment.
[0015] The method of editing an image described above enables sophisticated, yet not strict, editing of high-resolution, realistic images. The resulting image output is well-aligned with the target text while maintaining the overall background, structure, and composition of the original image. For example, an input image of a standing dog accompanied by the text "lying dog" results in an output of a dog in a new pose (lying instead of standing), with the same coat color, same background, same expression, same features, etc. Multiple objects in the image can be edited, and a very wide variety of edits can be performed, all while maintaining the overall structure and composition of the original image.
[0016] Furthermore, the present invention enables semantically significant linear interpolation between two text-embedding sequences to be generated, which demonstrates the powerful compositional ability of the text-image diffusion model and allows the user to gradually edit the objects in the image.
[0017] Referring now to the drawings, exemplary embodiments of the present disclosure will be discussed in more detail.
[0018] Exemplary Model Configuration FIG. 1 shows a block diagram of an exemplary image model 100 according to an exemplary embodiment of the present disclosure.
[0019] The image editing model 100 includes a pre-trained diffusion model 105. The premise underlying the pre-trained diffusion model is to initialize with a randomly sampled noise image (such as the randomly sampled noise image 110), and then repeatedly refine the image in a controlled manner until the image is synthesized into a "clean", i.e., a realistic image without artifacts from the generation process. Each refinement step may include the application of a neural network to the current sample followed by a random Gaussian noise perturbation, which results in a more defined sample. The network is trained for noise removal, which leads to a learned image distribution that is highly faithful to the target distribution. These models may be generalized to learn conditional distributions by conditioning the noise removal network on an auxiliary input, and the resulting diffusion process can sample from the data distribution conditioned on the auxiliary input. In some embodiments, the auxiliary input may be a text sequence describing the desired image. By incorporating knowledge from large language models or hybrid vision-language models, these text-to-image diffusion models can generate realistic high-resolution images using only a text prompt describing the scene. A low-resolution image is first synthesized using generative diffusion and then converted to a high-resolution image using an additional auxiliary model.
[0020] The present invention receives, as input, an input image (such as input image 115) representing an actual object, animal, person, etc., and a target text describing a desired edit. The ultimate goal of the pre-trained model 105 is to edit the input image so as to satisfy a given text while maintaining the maximum amount of detail from the input image (e.g., background details and the identity of the objects within the image). To accomplish this, the text embedding layer of the pre-trained diffusion model 105 may be used to perform semantic operations. This is done by finding important representations, which, when fed through the generation process, produce an image similar to the input image 115. The pre-trained model 105 is then fine-tuned to better reconstruct the input image 115, and then, ultimately, the latent representation is manipulated to obtain the edited result.
[0021] First, text embedding optimization may be performed. Text embedding optimization includes passing the target text through a text encoder, which outputs a corresponding target text embedding 120. The output target text embedding indicates the number of tokens in the target text and may have a token embedding dimension.
[0022] The parameters of the pre-trained diffusion model 105 are then frozen, and the target text embedding 120 is optimized using a denoising diffusion objective such as the objective shown in Equation 1.
[0023]
Equation
[0024] In Equation 1, x tis the noisy version of the input image, and theta is the set of pre-trained diffusion model weights. The resulting optimized text embedding 125 is a text embedding that matches the input image as closely as possible (e.g., as indicated by the embedding distance 130). In some embodiments, this process is carried out in relatively few steps in order to remain close to the initial target text embedding 120 when obtaining the optimized test embedding 125. This optimized text embedding does not necessarily lead directly to the input image x when passed through the generative diffusion process, because the optimization operates for only a few steps. This proximity enables important interpolation in the embedding space and does not exhibit linear behavior for distant embeddings.
[0025] Next, the pre-trained diffusion model 105 is fine-tuned. Since the optimization for obtaining the optimized text embedding 125 operates for only a few steps, the model parameters theta of the pre-trained diffusion model 105 are optimized using the same loss function as shown in Equation 1 and the optimized text embedding 125 is frozen. This process shifts the pre-trained diffusion model 105 to fit the input image 115 at the point of the optimized text embedding 125. In parallel, an auxiliary diffusion model existing within the base generative model may be fine-tuned with the same reconstruction loss as the pre-trained diffusion model 105, but conditioned on the target text embedding 120 rather than the optimized text embedding 125. The optimization of the auxiliary model ensures the preservation of high-frequency details from the input image 115 that do not exist at the base resolution.
[0026] After the pre-trained diffusion model 105 is trained to sufficiently reproduce the input image 115 in the optimized embedding 125, a newly created fine-tuned diffusion model 150 (e.g., the pre-trained diffusion model 105 that has been trained and fine-tuned) is used to apply the desired edits to the input image 115 by moving in the direction of the target text embedding 120. In other words, this stage of the process is a simple linear interpolation 155 between the target text embedding 120 and the optimized text embedding 125. For a given hyperparameter η ∈ [0, 1], Equation 2 can be used to obtain the embedding representing the desired edited image.
[0027]
Number
[0028] In Equation 2, e tgt is the target text embedding 120, and e opt is the optimized text embedding 125. The output of Equation 2 is the embedding representing the desired edited image. The fine-tuned diffusion model 150 applies the base generation diffusion process (
Number
[0029] In some embodiments, the frameworks of the pre-trained diffusion model 105 and the fine-tuned diffusion model 150 may include three different text-conditioned diffusion models, namely, a generative diffusion model for 64×64 pixel images, a super-resolution diffusion model that changes a 64×64 pixel image to a 256×256 pixel image, and a second super-resolution diffusion model that converts a 256×256 pixel image to a 1024×1024 pixel image. By cascading these three models and using classifier-free guidance, the pre-trained diffusion model 105 and the fine-tuned diffusion model 150 can generate image edits based on text guidance from the target text.
[0030] An example of the image editing process can be seen in FIG. 2. For example, based on an input image (such as any of the input images 205), the fine-tuned diffusion model 150 may be fine-tuned for a specific input image, and as a result, various edited images (such as any of the edited images 210) are output based on the target text input associated with the desired image edit. In one example, an input dog image 215 of a dog standing in a field may be provided along with one of the desired image edit target texts 220 such as "a sitting dog", "a jumping dog", "a lying dog". Note that for some target texts such as "a dog playing with a toy" or "a dog jumping and holding a frisbee", new objects as well as edits may be added to the currently existing objects on which the edits are being performed. Other examples of the input image 205 and the resulting edited images 210 are also shown to illustrate other types of edits that can be performed on various input images.
[0031] Returning to FIG. 1 here, in some embodiments, the final output image 160 may vary based on the random noise 165 input into the fine-tuning diffusion model 150. For example, FIG. 3 shows the results 305 when different random noise samples are input into the fine-tuning diffusion model 150. Given a first input image 310 of a bird resting with its wings down having the target text "photo of a bird spreading its wings", different noise samples can result in different output results 305. A second example of a second input image 315 of a forest associated with the target text "drawing of a forest by a child" is also provided to show how random noise samples 165 can result in different final output images 160.
[0032] In some embodiments, the value of the hyperparameter η can result in different edited images. For example, FIG. 4 shows how different edited images 415 can result based on a given input image 405 when the value 410 of the hyperparameter η is changed. In a second example, certain different η values 420 can lead to different generated results based on an input image of a cake and the target text "photo of a pistachio cake".
[0033] Exemplary Method FIG. 5 shows a flowchart diagram of an exemplary method 500 for performing image editing according to an exemplary embodiment of the present disclosure. FIG. 5 shows steps performed in a specific order for purposes of illustration and explanation, but the methods of the present disclosure are not limited to the specifically shown order or sequence. The various steps of method 500 may be variously omitted, reordered, combined, and / or adapted without departing from the scope of the present disclosure.
[0034] At 502, the computing system receives an input image and a target text, where the target text indicates the desired edit for the input image. The target text is a text presentation of the desired edit, such as "a dog sitting" provided with an input image of a dog standing in a field. The input image is one from which an edited image is generated while performing the desired edit indicated by the target text (e.g., while preserving details from the input image).
[0035] At 504, the computing system obtains a target text embedding based on the target text. The target text embedding is a text embedding that represents the target text among embeddings usable by one or more models. In some embodiments, the computing system obtains the target text embedding by providing the target text to a text encoder and receiving the target text embedding as an output from the text encoder. In some embodiments, the target text embedding includes the number of tokens in the target text and has a token embedding dimension.
[0036] In 506, the computing system obtains an optimized text embedding based on the target text embedding and the input image. To obtain the optimized text embedding, the parameters of the diffusion model are frozen, and the target text embedding is optimized using a denoising decoding objective, which may be a loss function such as the function described in Equation 2 above. The resulting text embedding represents the input image as accurately as possible. This optimization process is run for a predetermined number of steps (relatively few) in order to remain close to the initial target text embedding. The embedding obtained after several steps are run is the optimized text embedding. The optimized text embedding enables important linear interpolation in the embedding space, which does not exhibit linear behavior for distant embeddings. The optimized text embedding does not necessarily lead directly to the input image when passed through the diffusion model, because the optimization runs for only a few steps.
[0037] In 508, the computing system fine-tunes the diffusion model based on the optimized text embedding. To fine-tune the diffusion model, the optimal text embedding is frozen, and the model parameters of the diffusion model are optimized using the frozen optimal text embedding. A loss function is used to optimize the model parameters. This process shifts the model to fit the input image x at the point of the optimized text embedding. In some embodiments, optimizing the model parameters may include conditioning the diffusion model on the optimized text embedding.
[0038] In some embodiments, fine-tuning the diffusion model may also include fine-tuning at least one auxiliary diffusion model, such as a super-resolution generation diffusion model. Fine-tuning the auxiliary diffusion model may be similar to fine-tuning the diffusion model. For example, fine-tuning the auxiliary diffusion model may include using the same loss function as the diffusion model to optimize one or more parameters of the auxiliary diffusion model, but instead of freezing the optimized text embedding and conditioning the auxiliary diffusion model with the optimized text embedding, the target text embedding is frozen instead, and the auxiliary diffusion model is conditioned on the target text embedding. Optimizing the auxiliary diffusion model with the target text embedding ensures that high-frequency details from the input image that do not exist at the base resolution generated by the diffusion model are maintained.
[0039] At 510, the computing system interpolates between the target text embedding and the optimized text embedding to obtain an interpolated embedding. The diffusion model is trained sufficiently and precisely (e.g., overfitted) to adequately reproduce the input image in the optimized embedding. Thus, to apply the desired edit, the computing system starts at the optimized embedding and moves through the embedding space towards the target text embedding to generate the edit. To "move in the direction" of the target text embedding, simple linear interpolation (described by Equation 2 disclosed above) is performed between the optimized text embedding and the target text embedding. In some embodiments, this interpolation is represented by a hyperparameter |. Based on the linear interpolation of the optimized text embedding, the target text embedding, and the hyperparameter, an interpolated embedding is obtained.
[0040] At 512, the computing system generates an edited image including a desired edit using a diffusion model based on the input image and the interpolated embedding. The fine-tuned diffusion model then performs a base generation diffusion process conditioned on the interpolated embedding to obtain a low-resolution image (e.g., a 64×64 pixel image) including the desired edit. In some embodiments, one or more auxiliary diffusion models, such as a super-resolution model, may then be used to super-resolve the low-resolution image into a high-resolution final image (e.g., a 256×256 and / or 1024×1024 pixel image), which is then output as the final edited image.
[0041] Exemplary Devices and Systems FIG. 6A shows a block diagram of an exemplary computing system 600 for performing image editing according to an exemplary embodiment of the present disclosure. System 600 includes a user computing device 602, a server computing system 630, and a training computing system 650 communicatively coupled via a network 680.
[0042] The user computing device 602 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0043] The user computing device 602 includes one or more processors 612 and a memory 614. The one or more processors 612 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 614 can include one or more non-transitory computer-readable recording media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 614 can store data 616 and instructions 618 that are executed by the processor 612 to cause the user computing device 602 to perform operations.
[0044] In some implementations, the user computing device 602 can store or include one or more image editing models 620. For example, the image editing model 620 can be various machine learning models such as a neural network (e.g., a deep neural network) or other types of machine learning models including non-linear models and / or linear models, or can include those machine learning models. The neural network can include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Some exemplary machine learning models can utilize attention mechanisms such as self-attention. For example, some exemplary machine learning models can include a multi-head self-attention model (e.g., a transformer model). The exemplary image editing model 620 is discussed with reference to FIG. 1.
[0045] In some implementations, one or more image editing models 620 are received from a server computing system 630 via a network 680, stored in a user computing device memory 614, and then used by, or otherwise implemented by, one or more processors 612. In some implementations, a user computing device 602 can implement multiple parallel instances of a single image editing model 620 (e.g., to perform parallel image editing across multiple instances of image editing).
[0046] More specifically, given an input image and target text indicating a desired edit to the image, one or more image editing models 620 can edit the image to satisfy the given text while maintaining the maximum amount of detail from the original input image, such as details in the background, details about unedited portions of one or more objects, and the like. One or more models 620 then use a text embedding layer of the one or more models to perform semantic operations on the input image to obtain a latent representation of the input image, and this representation can then be manipulated to obtain an output image in which the desired edit is specified by the given text.
[0047] Additionally or alternatively, one or more image editing models 640 can be included in a server computing system 630 that communicates with a user computing device 602 according to a client - server relationship, or otherwise stored and implemented by the server computing system 630. For example, an image editing model 640 can be implemented by a server computing system 630 as part of a web service (e.g., an image editing service). Thus, one or more models 620 can be stored and implemented in a user computing device 602 and / or one or more models 640 can be stored and implemented in a server computing system 630.
[0048] The user computing device 602 may also include one or more user input components 622 that receive user input. For example, the user input component 622 may be a touch-sensitive component (such as a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (such as a finger or a stylus). The touch-sensitive component can be useful for implementing a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which a user can provide user input.
[0049] The server computing system 630 includes one or more processors 632 and a memory 634. The one or more processors 632 may be any suitable processing device (such as a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or multiple processors operably connected. The memory 634 can include one or more non-transitory computer-readable recording media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 634 can store data 636 and instructions 638 that are executed by the processor 632 to cause the server computing system 630 to perform operations.
[0050] In some implementations, the server computing system 630 includes or is otherwise implemented by one or more server computing devices. In cases where the server computing system 630 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0051] As described above, the server computing system 630 can store or otherwise include one or more image editing models 640. For example, the model 640 can be or otherwise include various machine learning models. Exemplary machine learning models include neural networks or other multi-layer non-linear models. Exemplary neural networks include feed-forward neural networks, deep neural networks, regression neural networks, and convolutional neural networks. Some exemplary machine learning models can utilize attention mechanisms such as self-attention. For example, some exemplary machine learning models can include multi-head self-attention models (e.g., transformer models). Exemplary model 640 is discussed with reference to FIG. 1.
[0052] The user computing device 602 and / or the server computing system 630 can train the models 620 and / or 640 through interaction with a training computing system 650 communicatively coupled via a network 680. The training computing system 650 can be separate from the server computing system 630 or can be a part of the server computing system 630.
[0053] The training computing system 650 includes one or more processors 652 and a memory 654. The one or more processors 652 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 654 can include one or more non-transitory computer-readable recording media such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 654 can store data 656 and instructions 658 that are executed by the processor 652 to cause the training computing system 650 to perform operations. In some implementations, the training computing system 650 includes one or more server computing devices or, alternatively, is implemented by one or more server computing devices.
[0054] The training computing system 650 can include a model trainer 660 that trains a machine learning model 620 and / or 640 stored in the user computing device 602 and / or the server computing system 630 using various training or learning techniques such as, for example, backpropagation of error. For example, a loss function can be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or other various loss functions. Gradient descent techniques can be used to iteratively update the parameters over several training iterations.
[0055] In some implementations, performing backpropagation of errors may include performing abbreviated backpropagation over time. The model trainer 660 can implement some generalization techniques (such as weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0056] In particular, the model trainer 660 can train the image editing models 620 and / or 640 based on a set of training data 662. The training data 662 can include, for example, input images for one or more of the editing models 620 and 640. In some implementations, the editing models 620 and 640 can include a generative diffusion model for 64×64 pixel images, a first super-resolution diffusion model for converting 64×64 pixel images to 256×256 pixel images, and a second super-resolution diffusion model for converting 256×256 pixel images to 1024×1024 pixel images. The text embeddings for the image editing models 620 and 640 can be optimized during training using the 64×64 diffusion model. In some implementations, an Adam optimizer with 100 sets and a fixed learning rate of 1e-3 can be used. The 64×64 diffusion model can then be fine-tuned by continuing training for 1500 steps on the input images, conditioned on the optimized embeddings. In parallel, the first super-resolution diffusion model can be fine-tuned for 1500 steps using the target text embedding and the original image to capture high-frequency details from the input image.
[0057] In another implementation, the image editing models 620 and 640 can be trained by applying a diffusion process in a latent space of size 4×64×64 of a pre-trained autoencoder that acts on 512×512 images with an optimization of text embeddings taking 1000 steps at a learning rate of 2e-3 using Adam.
[0058] Based on this training, the text embedding can be interpolated according to Equation 3 described above with respect to FIG. 1. With the fine-tuning process, as η increases, the output image aligns with the target text. For example, when η is 0, the original image is generated, and when η is 1, a completely new image (e.g., not sharing characteristics with the input image) is generated. In some embodiments, η may be intermediate between 0.6 and 0.8. The output image can then be generated using image editing model 620 or 640.
[0059] In some implementations, when the user gives consent, the training examples may be provided by user computing device 602. Thus, in such implementations, model 620 provided to user computing device 602 can be trained by training computing system 650 with respect to user-specific data received from user computing device 602. In some cases, this process may be referred to as personalizing the model.
[0060] Model trainer 660 includes computer logic used to provide the desired functionality. Model trainer 660 can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, model trainer 660 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, model trainer 660 includes one or more sets of computer-executable instructions stored on a tangible computer-readable recording medium such as RAM, a hard disk, or an optical or magnetic medium.
[0061] Network 680 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication via network 680 can be carried over any type of wired and / or wireless connection using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0062] The machine learning models described herein may be used in a variety of tasks, applications, and / or use cases.
[0063] In some implementations, the input to the machine learning models of the present disclosure may be image data. The machine learning models may process the image data to generate an output. By way of example, the machine learning models may process the image data to generate an image recognition output (e.g., recognition of the image data, latent embedding of the image data, encoded representation of the image data, hash of the image data, etc.). As another example, the machine learning models may process the image data to generate an image segmentation output. As another example, the machine learning models may process the image data to generate an image classification output. As another example, the machine learning models may process the image data to generate an image data modification output (e.g., modification of the image data, etc.). As another example, the machine learning models may process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, the machine learning models may process the image data to generate an upscaled image data output. As another example, the machine learning models may process the image data to generate a prediction output.
[0064] In some implementations, the input to the machine learning model of the present disclosure may be text or natural language data. The machine learning model may process the text or natural language data to generate an output. By way of example, the machine learning model may process natural language data to generate a language encoding output. As another example, the machine learning model may process text or natural language data to generate a latent text embedding output. As another example, the machine learning model may process text or natural language data to generate a conversion output. As another example, the machine learning model may process text or natural language data to generate a classification output. As another example, the machine learning model may process text or natural language data to generate a text segmentation output. As another example, the machine learning model may process text or natural language data to generate a semantic intent output. As another example, the machine learning model may process text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is of higher quality than the input text or natural language, etc.). As another example, the machine learning model may process text or natural language data to generate a prediction output.
[0065] In some implementations, the input to the machine learning model of the present disclosure may be audio data. The machine learning model may process the audio data to generate an output. As an example, the machine learning model may process the audio data to generate an audio recognition output. As another example, the machine learning model may process the audio data to generate an audio translation output. As another example, the machine learning model may process the audio data to generate a latent embedding output. As another example, the machine learning model may process the audio data to generate an encoded audio output (e.g., an encoded and / or compressed representation of the audio data, etc.). As another example, the machine learning model may process the audio data to generate an upscaled audio output (e.g., audio data of higher quality than the input audio data, etc.). As another example, the machine learning model may process the audio data to generate a text representation output (e.g., a text representation of the input audio data, etc.). As another example, the machine learning model may process the audio data to generate a prediction output.
[0066] In some implementations, the input to the machine learning model of the present disclosure may be latent encoding data (e.g., a latent space representation of the input, etc.). The machine learning model may process the latent encoding data to generate an output. As an example, the machine learning model may process the latent encoding data to generate a recognition output. As another example, the machine learning model may process the latent encoding data to generate a reconstruction output. As another example, the machine learning model may process the latent encoding data to generate a search output. As another example, the machine learning model may process the latent encoding data to generate a reclustering output. As another example, the machine learning model may process the latent encoding data to generate a prediction output.
[0067] In some cases, the machine learning model can be configured to perform tasks that include encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task may be an audio compression task. The input may include audio data, and the output may include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task may include generating an embedding for the input data (e.g., input audio or visual data).
[0068] FIG. 6A shows one exemplary computing system that can be used to implement the present disclosure. Other computing systems may be used. For example, in some implementations, the user computing device 602 can include a model trainer 660 and a training dataset 662. In such implementations, the model 620 can be both locally trained and used on the user computing device 602. In some of such implementations, the user computing device 602 can implement the model trainer 660 to individualize the model 620 based on user-specific data.
[0069] FIG. 6B shows a block diagram of an exemplary computing device 700 implemented in accordance with an exemplary embodiment of the present disclosure. The computing device 700 can be a user computing device or a server computing device.
[0070] Computing device 700 includes several applications (e.g., applications 1 to N). Each application includes its own machine learning library and machine learning model. For example, each application can include a machine learning model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.
[0071] As shown in FIG. 6B, each application can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0072] FIG. 6C shows a block diagram of an exemplary computing device 800 implemented according to an exemplary embodiment of the present disclosure. Computing device 800 can be a user computing device or a server computing device.
[0073] Computing device 800 includes several applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).
[0074] The central intelligence layer includes several machine learning models. For example, as shown in FIG. 6C, each machine learning model can be provided to each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model to all applications. In some implementations, the central intelligence layer is included in or otherwise implemented by the operating system of the computing device 800.
[0075] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized repository of data for the computing device 800. As shown in FIG. 6C, the central device data layer can communicate with some other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0076] Additional disclosure The technology described in this specification refers to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent between such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionality between components. For example, the processes described in this specification can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate continuously or in parallel.
[0077] Although the present subject matter has been described in detail with respect to its various specific exemplary embodiments, each example is provided by way of illustration and not limitation of the disclosure. Those skilled in the art will readily appreciate that upon understanding the foregoing, modifications, variations, and equivalents to such embodiments can easily occur. Accordingly, the disclosure is not intended to exclude such modifications, variations, and / or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. For example, features illustrated or described as part of one embodiment can be used with another embodiment to further create additional embodiments. Accordingly, the disclosure is intended to embrace such alternatives, variations, and equivalents.
Description of Reference Numerals
[0078] 100 Image model, Image editing model 105 Pre-trained diffusion model, Pre-trained model 600 Computing system, System 602 User computing device 612 Processor 614 Memory, User computing device memory 622 User input component 630 Server computing system 632 Processor 634 Memory 650 Training Computing System 652 Processor 654 Memory 660 Model Trainer 680 Network 700 Computing Device 800 Computing Device
Claims
1. 1. A computer-implemented method for editing an image, comprising: receiving, by a computing system, an input image and target text, the target text indicating a desired edit for the input image; obtaining, by the computing system, a target text embedding based on the target text; obtaining, by the computing system, optimized text embeddings based on the target text embeddings and the input image; fine-tuning, by the computing system, a diffusion model based on the optimized text embeddings; interpolating, by the computing system, the target text embedding and the optimized text embedding to obtain an interpolated embedding; generating, by the computing system, an edited image including the desired edits based on the input image and the interpolated embedding using the diffusion model.
2. The step of obtaining the target text embedding comprises: providing, by the computing system, the target text to a text encoder; and receiving, by the computing system, the target text embeddings from the text encoder.
3. The step of obtaining the optimized text embeddings comprises: freezing, by the computing system, parameters of the diffusion model; optimizing, by the computing system, the target text embedding using a denoising diffusion objective to obtain the optimized text embedding; and outputting the optimized text embedding, the optimized text embedding being a text embedding that matches the input image.
4. The step of fine-tuning the diffusion model comprises: freezing, by the computing system, the optimized text embeddings; and optimizing, by the computing system, at least one model parameter of the diffusion model using a loss function.
5. 5. The computer-implemented method of claim 4, wherein optimizing at least one model parameter comprises conditioning the diffusion model on the optimized text embeddings.
6. The step of fine-tuning the diffusion model comprises: The method further includes fine-tuning, by the computing system, at least one auxiliary diffusion model, the fine-tuning comprising: freezing, by the computing system, the target text embeddings; and optimizing, by the computing system, the at least one auxiliary diffusion model using the loss function.
7. 7. The computer-implemented method of claim 6, wherein optimizing the at least one auxiliary diffusion model comprises conditioning the at least one auxiliary diffusion model on the target text embeddings.
8. The step of generating an edited image includes: generating, by the computing system, a low-resolution version of the edited image using the diffusion model; and super-resolving, by the computing system, the low-resolution version of the edited image into a final high-resolution version of the edited image using the at least one auxiliary diffusion model.
9. The step of generating an edited image includes:
2. The computer-implemented method of claim 1, comprising generating, by the computing system, the edited image using the diffusion model conditioned on the interpolated embedding.
10. 1. A computing system for editing images, the computing system comprising: one or more processors; and a computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: receiving an input image and target text, the target text indicating a desired edit for the input image; obtaining a target text embedding based on the target text; and obtaining optimized text embeddings based on the target text embeddings and the input image; and fine-tuning a diffusion model based on the optimized text embeddings; and interpolating the target text embedding and the optimized text embedding to obtain an interpolated embedding; generating an edited image including the desired edits using the diffusion model based on the input image and the interpolated embedding.
11. Obtaining the target text embeddings includes: providing the target text to a text encoder; and receiving the target text embeddings from the text encoder.
12. Obtaining the optimized text embeddings includes: freezing, by the computing system, parameters of the diffusion model; optimizing, by the computing system, the target text embeddings using a denoising diffusion objective to obtain the optimized text embeddings; and outputting the optimized text embedding, the optimized text embedding being a text embedding that matches the input image.
13. Fine-tuning the diffusion model includes: Freezing the optimized text embeddings; and and optimizing at least one model parameter of the diffusion model using a loss function.
14. 14. The computing system of claim 13, wherein optimizing the at least one model parameter comprises conditioning the diffusion model with the optimized text embeddings.
15. Fine-tuning the diffusion model includes: and further comprising fine-tuning at least one auxiliary diffusion model, wherein the fine-tuning of the at least one auxiliary diffusion model comprises: Freezing the target text embeddings; and and optimizing the at least one auxiliary diffusion model using the loss function.
16. 16. The computing system of claim 15, wherein optimizing the at least one auxiliary diffusion model comprises conditioning the at least one auxiliary diffusion model on the target text embedding.
17. generating the edited image includes: generating a low resolution version of the edited image using the diffusion model; and and super-resolving the low-resolution version of the edited image into a final high-resolution version of the edited image using the at least one auxiliary diffusion model.
18. generating the edited image includes: The computing system of claim 10 further comprising generating the edited image using the diffusion model conditioned on the interpolated embedding.
19. 1. A computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: receiving an input image and target text, the target text indicating a desired edit for the input image; obtaining a target text embedding based on the target text; and obtaining optimized text embeddings based on the target text embeddings and the input image; and fine-tuning a diffusion model based on the optimized text embeddings; and interpolating the target text embedding and the optimized text embedding to obtain an interpolated embedding; generating an edited image including the desired edits using the diffusion model based on the input image and the interpolated embedding.
20. generating the edited image includes:
20. The computer-readable medium of claim 19, further comprising generating the edited image using the diffusion model conditioned on the interpolated embedding.
Citation Information
Patent Citations
Image processing method, image processing model training method, device, and storage medium
JP2022180519A
Semantic image manipulation using visual-semantic joint embeddings
US20220036127A1