Text-driven image editing via image-specific trimming of diffusion model

By using a fine-tuned machine learning diffusion model to process natural language editing prompts and basic images, high-fidelity output images are generated, which solves the fidelity and complexity problems of text-driven image editing in the prior art, and achieves the effect of maskless high-fidelity editing.

CN120051803APending Publication Date: 2025-05-27GOOGLE LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380073286.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2023-10-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to achieve high fidelity text-driven image editing, especially when processing the editing of the image dependent portion of the masked image, it is difficult to effectively process the model.

Method used

Generate output images by using a fine-tuned machine learning diffusion model, combining editing tips and basic images of natural language descriptions. The method includes splicing fine-tuning prompts with edit prompts into combination prompts, and processing these prompts using a diffusion model, performing classifier-free boots or prompt boots to ensure that the output image is consistent with the desired edit.

Benefits of technology

It realizes high fidelity to perform image editing without masking, maintaining the integrity of visual and semantic details of the input image, and being able to handle complex cross-domain scenes and local editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051803A_ABST
    Figure CN120051803A_ABST
Patent Text Reader

Abstract

Systems and methods are provided for generic text-driven image editing, and example implementations of these systems and methods may be referred to as' UniTune '. The UniTune may receive any image and text edit description as input, and may perform edit while maintaining high semantics and visual fidelity of the input image. The UniTune does not require any additional input such as a mask or sketch. According to one aspect of the present disclosure, by properly selecting parameters, an example system described herein may fine-tune a large diffusion model (e.g., Image) on a single image, causing the model to visually and semantically remain fidelity to the input image while still allowing expressive operation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 416,838, filed Oct. 17, 2022. U.S. Provisional Patent Application No. 63 / 416,838 is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to machine learning. More specifically, the present disclosure relates to text-driven image editing via image-specific fine-tuning of diffusion models. Background Art

[0004] Performing high-fidelity image manipulation via text commands is a long-standing problem in computer graphics research. Using free-form commands such as "men wearing tuxedos", "pixelart", or "blue house" to describe the desired edits is significantly easier than manually implementing the modifications in image editing software. Language-based interfaces have the potential to make experts more efficient and unlock graphic design capabilities for casual users. Despite the remarkable progress of image generation methods, high-fidelity image editing in the general domain remains an unsolved problem.

[0005] Specifically, revolutionary text-to-image models such as Dall-E, Imagen, and StableDiffusion are good at creating images from scratch or filling in regions manually removed from existing images in a context-aware manner. However, for editing operations, these models typically require the user to specify masks and often struggle to handle edits that depend on the masked parts of the image. Summary of the Invention

[0006] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.

[0007] One general aspect includes a computer-implemented method for performing text-driven image editing. The computer-implemented method includes obtaining a base image and an editing prompt by a computing system that may include one or more computing devices, where the editing prompt may include or be derived from a natural language description of a desired edit to the base image. The method further includes accessing, by the computing system, a machine learning diffusion model that has been fine-tuned on one or more fine-tuning tuples, each fine-tuning tuple may include a base image and a fine-tuning prompt. The method further includes processing, by the computing system, the fine-tuning prompt and the editing prompt using the machine learning diffusion model to generate an output image. The method further includes providing, by the computing system, output data as an output. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0008] The implementation may include one or more of the following features. The computer-implemented method described above, wherein the fine-tuning prompt may include or be derived from one or more rare tokens. The fine-tuning prompt may include or be derived from a base prompt, which may include a natural language description of the base image. Processing the fine-tuning prompt and the editing prompt by the computing system using a machine learning diffusion model may include: concatenating, by the computing system, the fine-tuning prompt and the editing prompt to generate a combined prompt; and processing, by the computing system, the combined prompt using a machine learning diffusion model. Accessing, by the computing system, a machine learning diffusion model may include: fine-tuning, by the computing system, the machine learning diffusion model using the one or more fine-tuning tuples. The machine learning diffusion model may include a text-to-image diffusion model and one or more super-resolution diffusion models that sequentially follow the text-to-image diffusion model, and wherein at least one of the text-to-image diffusion model and the one or more super-resolution diffusion models has been fine-tuned using the one or more fine-tuning tuples. Processing the fine-tuning prompt and the editing prompt by the computing system using a machine learning diffusion model may include: performing, by the computing system, classifier-free guidance to bias the machine learning diffusion model towards the editing prompt and away from the unconditional setting. Processing the fine-tuning prompt and the editing prompt by the computing system using a machine learning diffusion model may include: performing, by the computing system, prompt guidance to bias the machine learning diffusion model towards the editing prompt and away from the base prompt. Processing the fine-tuning prompt and the editing prompt by the computing system using a machine learning diffusion model may include: processing, by the computing system, the fine-tuning prompt and the editing prompt and a set of noise to generate an output image. Processing the fine-tuning prompt and the editing prompt by the computing system using a machine learning diffusion model may include: processing, by the computing system, the fine-tuning prompt and the editing prompt and a noisy image generated from the base image to generate an output image. The computer-implemented method described above may include: interpolating, by the computing system, the base image and the output image to generate a second output image. The output image may depict the desired edit performed on the base image. Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0009] Another general aspect includes a computer-implemented method for training a diffusion model to perform text-driven image editing. The computer-implemented method includes obtaining, by a computing system that may include one or more computing devices, a base image and an edit prompt, where the edit prompt may include or be derived from a natural language description of a desired edit to the base image. The method further includes accessing, by the computing system, a machine learning diffusion model. The method further includes fine-tuning, by the computing system, the machine learning diffusion model on one or more fine-tuning tuples, each fine-tuning tuple may include the base image and a fine-tuning prompt. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0010] Other aspects of the present disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0011] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate

[0012] example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Reference is made to the following drawings in which detailed discussions of embodiments for those of ordinary skill in the art are set forth:

[0014] Figures 1A to 1C An example process for text-driven image editing with image-specific fine-tuning via a diffusion model is shown. Specifically, Figure 1A A block diagram depicting an example method for training a diffusion model to perform text-conditioned image generation in accordance with an example embodiment of the present disclosure. Figure 1B A block diagram depicting an example method for fine-tuning a diffusion model on a specific base image in accordance with an example embodiment of the present disclosure. Figure 1C A block diagram depicting an example method for generating an output image depicting a desired edit to a base image in accordance with an example embodiment of the present disclosure.

[0015] Figures 2A to 2C Example computing systems and devices are shown. Specifically, Figure 2A A block diagram depicting an example computing system in accordance with an example embodiment of the present disclosure. Figure 2B A block diagram depicting an example computing device in accordance with an example embodiment of the present disclosure. Figure 2C A block diagram depicting an example computing device in accordance with an example embodiment of the present disclosure.

[0016] Reference numerals that are repeated in multiple figures are intended to identify like features in various implementations. Detailed Description

[0017] Overview

[0018] Generally, the present disclosure relates to systems and methods for general text-driven image editing, and example implementations of these systems and methods may be referred to as "UniTune". Example implementations of the present disclosure may receive any image and a text editing description as input and may perform the editing while maintaining a high semantic and visual fidelity to the input image. Example implementations of the present disclosure do not require any additional input such as masks or sketches. According to one aspect of the present disclosure, by correctly selecting parameters, the example systems described herein may fine-tune a large diffusion model (e.g., Imagen) on a single image, thereby prompting the model to maintain fidelity to the input image both visually and semantically while still allowing for expressive operations. The proposed method has been tested in a series of different use cases, and the experimental results demonstrate its broad applicability.

[0019] An example aspect of the present disclosure relates to a computer-implemented method for performing text-driven image editing using a machine learning diffusion model. This technique can be used to edit a base image according to a natural language description provided in an edit prompt. For example, the edit prompt may instruct the model to change the color of an object in the image, add a new element, change the lighting of the image, and / or perform other edits on the image. The model can then generate an output image that incorporates these edits while maintaining the overall fidelity of the base image.

[0020] The diffusion model used in the present disclosure may be fine-tuned on one or more fine-tuning tuples. Each of these tuples may include a base image and a fine-tuning prompt. The fine-tuning prompt may include or be derived from one or more rare word elements, which helps the model focus on specific features of the base image that need to be retained.

[0021] Another aspect of the present disclosure relates to using a combined prompt during the image editing process. This combined prompt may be created by concatenating the fine-tuning prompt with the edit prompt. This approach ensures that the model considers both the original features of the base image and the desired edits when generating the output image. For example, if the base image is a beach scene and the edit prompt requests a sunset, the combined prompt will guide the model to retain the beach elements while adding a sunset to the scene.

[0022] In some implementations, the machine learning diffusion model can include a text-to-image diffusion model and one or more super-resolution diffusion models. The text-to-image model can be responsible for generating a low-resolution version of the output image, while the super-resolution model can refine this image to produce a high-resolution result. At least one of the super-resolution models in the text-to-image model and the super-resolution model can be fine-tuned using a fine-tuning tuple. This multi-stage approach allows the technique to generate high-quality edited images even when the editing prompt involves complex or detailed changes.

[0023] Some example methods described in this disclosure also involve using classifier-free guidance. This technique causes the diffusion model to be adjusted towards the editing prompt and away from the unconditional setting, thus ensuring that the output of the model is closely consistent with the desired edit. For example, if the editing prompt requires adding a sunset to a beach scene, the classifier-free guidance will direct the model to add the sunset while preserving the remaining beach elements.

[0024] Prompt guidance is another technique used in this disclosure. Prompt guidance can cause the diffusion model to be adjusted towards the editing prompt and away from the base prompt. This allows the model to focus on the desired edit without being overly influenced by the original features of the base image. For example, if the base image is a beach scene during the day and the editing prompt requires a sunset, the prompt guidance will help the model add the sunset without overly preserving the daylight.

[0025] Some example implementations of this disclosure also incorporate a set of noise into the image editing process. This noise can be processed by the diffusion model together with the fine-tuning prompt and the editing prompt to generate the output image. Incorporating noise into the process can help create a more natural and realistic edited image. For example, the noise can cause subtle changes in color or texture, which make the edited elements blend more seamlessly with the rest of the image, or can alternatively be used as a random position to initialize the model.

[0026] In some implementations, the method involves using the diffusion model to process a noisy image generated from the base image to create the output image. This method can increase the visual fidelity of the output image because the method starts from a version of the base image that has incorporated a certain level of noise. For example, if the base image is a beach scene and the editing prompt requires a sunset, the noisy image may include a slightly altered version of the beach scene that makes the addition of the sunset look more natural.

[0027] Some example methods provided herein also include interpolating a base image and an output image to generate a second output image. This interpolation process can further enhance the visual fidelity of the edited image as it ensures that the final result retains some of the original characteristics of the base image. For example, if the base image is a beach scene and the output image includes a newly added sunset, the interpolation will blend the two images to create a final result that looks like a natural beach scene at sunset.

[0028] The present disclosure provides a novel and efficient method for text-driven image editing using a machine learning diffusion model. This technique can significantly enhance the performance of a computing system in several ways. For example, using a fine-tuned diffusion model can enable faster and more accurate image editing tasks. The machine learning diffusion model can quickly process editing cues to generate an output image, thus saving significant computational time and resources.

[0029] Furthermore, the present disclosure can increase the energy efficiency of a computing system. The method uses a fine-tuned diffusion model that requires less energy to be maintained in memory and to perform computations within the model. This results in less energy being consumed to perform a given task such as maintaining the model in memory or performing computations within the model. This increased energy efficiency can also allow for more tasks to be completed within a given energy budget. For example, a greater number of tasks, more complex tasks, or the same tasks to be completed with higher accuracy or precision can be achieved.

[0030] The present disclosure not only enhances the performance and energy efficiency of a computing system but also enables new functionality. For example, the method allows for fine-tuning of the diffusion model using rare tokens. This unique functionality allows the model to focus on specific characteristics of the base image that need to be preserved during the editing process.

[0031] Accordingly, the present disclosure provides systems and methods, example implementations of which may be referred to as "UniTune", which represent a meaningful step towards the goal of high-fidelity image editing in the general domain. Specifically, the present disclosure provides a novel method for editing an image without a mask by simply providing a text description of the desired result. Additionally, example implementations of the present disclosure retain high fidelity to the overall input image including the edited portion. The fidelity can preserve both visual details (e.g., shape, color, and texture) and semantic details (e.g., objects, poses, and actions).

[0032] In addition, some example implementations of the present disclosure are capable of editing any image in complex cross-domain scenarios. The example implementations demonstrate the ability to perform both local editing and extensive global editing. What is unique about the proposed technique is that it can precisely locate complex local edits without an editing mask and can make image-wide style changes that only preserve semantic details.

[0033] More specifically, some example implementations of the present disclosure perform expressive image editing by leveraging the capabilities of large-scale text-to-image diffusion models. The proposed method uses a simple yet powerful technique to transfer the visual and semantic capabilities of such diffusion models to the domain of image editing.

[0034] A key observation is that, with the right parameters, fine-tuning a large diffusion model on a single (image, prompt) pair does not result in complete catastrophic forgetting. The fine-tuned model will strongly prefer to associate the provided image and prompt together and will strongly prefer to draw samples that are nearly identical to the provided image when given other prompts. However, the visual and semantic knowledge that the model acquired during its original training is still available for a wide variety of editing operations (e.g., can be exploited by performing classifier-free guidance). The fidelity-expression balance can be tuned by controlling the number of training steps and the learning rate, or the amount of classifier-free guidance and SDEdit.

[0035] Fine-tuning of diffusion models is a powerful technique relevant to many use cases, including, for example, image-to-image translation and subject-driven image generation. These methods attempt to mitigate overfitting during training time by data augmentation, using large datasets, or restricting fine-tuning to the embeddings of specific tokens. This allows these techniques to learn, for example, the essence of a subject without learning transient image-specific attributes such as pose, camera angle, background, etc. For our image editing use case, some amount of overfitting is beneficial as our actual goal is to maintain high fidelity to the source image. The present disclosure represents the first method to use fine-tuning of large diffusion models for an image editing use case.

[0036] Referring now to the accompanying drawings, example embodiments of the present disclosure will be discussed in more detail.

[0037] Example text-driven image editing

[0038] Some example implementations are configured to convert the text-to-image diffusion model f θ into a text-conditioned image editor for a base image x provided for a specific user (b) The new model should be able to accept an edit prompt c that describes the edited image and output an edited image x that satisfies the condition c and maintains fidelity to x (b) 0 ​​。

[0039] Some example methods consist of two stages: (1) fine-tuning the model solely on the base image x (b) ; (2) using a modified sampling process to balance the fidelity to the base image x (b) and the alignment with the edit prompt C.

[0040] Example fine-tuning

[0041] Some example implementations fine-tune the model on x (b) for a fixed number of steps to encourage the model to produce images close to the base image. For example, some example implementations use a text condition c (b) in the fine-tuning stage, which consists of a certain number (e.g., 3) of rare tokens to create rare words not found in the original training data of f θ . Some example implementations utilize the fixed condition and the image to use the diffusion model denoising loss:

[0042]

[0043] where and w t , α t , σ t are functions of t determined by the noise schedule of the diffusion model.

[0044] Example sampling

[0045] To perform the edit operation, some example implementations sample the fine-tuned model by concatenating c and c (b) (i.e., the string "[rare_tokens|edit_prompt"). Through naive sampling, the fine-tuned model tends to bias towards x (b) more than the provided prompt c, and the model produces images very similar to x (b) . Classifier-free guidance can be used to direct the model towards the concatenated prompt, producing images that maintain fidelity to x (b) while satisfying c. Since some example implementations use high values and / or apply oscillatory guidance and / or dynamic thresholding for classifier-free guidance. To increase visual fidelity, some example implementations start sampling at a lower step t (instead of starting from t = 1), and use the diffusion forward process to initialize the sampling with an appropriately noisy version of x (b) (instead of random Gaussian noise):

[0046] z t = α t x (b) + σ tε(2)

[0047] z t is the initialization value, and α t , σ t are functions of t determined by the noise schedule of the diffusion model. Finally, to further preserve the fine details from the source image x (b) linear interpolation is performed on the pixels of the generated image and the pixels of x (b) . The interpolation weights can be determined by the similarity of the pixel neighborhoods.

[0048] Example implementation details

[0049] Some example implementations use Imagen as the text-to-image model, where a frozen T5-XXL encoder is used for text embeddings. Imagen consists of a text-to-image model and two super-resolution models. The text-to-image model generates a 64x64 pixel output, and the two super-resolution models convert the 64x64 image to a 256x256 image and then to a 1024x1024 image. Some example implementations fine-tune these two models and use the default 1024 model. Some example implementations use Adafactor to train the 64x64 model and Adamw to train the 256x256 model with a learning rate of 0.0001 (the same settings as used when training Imagen). Some example implementations use a batch size of 4 and emit weights at 16, 32, 64, 128 training steps (some example implementations use fewer training steps when expressiveness is needed and more training steps when fidelity is needed). In this setup, an example fine-tuned model can reproduce x (b) .

[0050] It is observed that when fine-tuning with a very small number of steps (16 - 128) and using classifier-free guidance, the model takes into account the edit prompt and then returns an image x (b) similar to x and satisfying the desired c 0 . Surprisingly, despite the fact that some example implementations use a pixel-level MSE loss during fine-tuning, the similarity between x 0 and x (b) is usually semantic, and their MSE can be very different, indicating that fine-tuning changes the bias in the model's internal semantic representation.

[0051] Some example implementations demonstrate that when using a fine-tuned model, starting sampling with a very noisy version of x (b) is sufficient to maintain high visual fidelity (since the model has converged to x (b)Bias). Thus, to increase visual fidelity, some example implementations start sampling at step 0.8 ≤ t ≤ 0.98 and appropriately add noise to the base image. Some example implementations also use various values for classifier-free guidance, and the value 32 has proven beneficial.

[0052] Example data flow visualization

[0053] Figures 1A to 1C illustrates an example process for text-driven image editing for image-specific fine-tuning via a diffusion model. Specifically, Figure 1A depicts a block diagram of an example method for training a diffusion model to perform text-conditioned image generation according to an example embodiment of the present disclosure. Figure 1B depicts a block diagram of an example method for fine-tuning a diffusion model on a specific base image according to an example embodiment of the present disclosure. Figure 1C depicts a block diagram of an example method for generating an output image that depicts a desired edit to a base image according to an example embodiment of the present disclosure.

[0054] First, referring to Figure 1A , the training method can be applied to a large number of training tuples. Figure 1A An example training tuple 12 is shown in. The training tuple 12 includes a training image 16 and a set of training texts 14 that describe the content of the training image 16.

[0055] The training image 16 can be processed by the diffusion model 18 in the noise-adding direction to generate a set of latent noises 20. The training texts 14 can be processed by the text encoder 22 to generate conditional prompts. The text encoder 22 can be a transformer model (e.g., BERT, T5, etc.), a CLIP model, and / or other encoders.

[0056] The latent noises 20 and the conditional prompts 24 can be processed by the diffusion model 26 in the denoising direction to generate a reconstructed image 28. The reconstructed image 28 can be an attempt by the diffusion model 26 to reconstruct the training image 16.

[0057] One or more loss functions 30 can be used to evaluate the reconstructed image 28. For example, one loss function 30 can compare the reconstructed image 28 with the training image 16. The diffusion models 26 and 18 can be trained based on the loss function 30. For example, the loss function 30 can be backpropagated through the models 26 and 18.

[0058] By performing the training method shown in Figure 1A on a large number of training examples, the diffusion model 26 can be trained to generate images based on input conditional prompts generated from text.

[0059] Now referring to Figure 1BAccording to one aspect of the present disclosure, diffusion models 26 and 18 may be fine-tuned on a particular image so as to be able to later generate an output image corresponding to an edited version of the particular image.

[0060] Specifically, Figure 1B As shown in , the fine-tuning tuple 212 may include a base image 216 and a set of fine-tuning texts 214. In one example, the fine-tuning text 214 may include a set of one or more rare word-grams. An example is the rare word-gram "beikkpic." In another example, the fine-tuning text 214 may be a set of base texts that describe the content of the base image 216. For example, for Figure 1B For the example image 216 shown in FIG. 2 , the example base text may state: “Two men sit in a restaurant booth in front of plates of food. The man in the red hoodie has his arm around the man in the blue striped shirt and glasses.”

[0061] The base image 216 may be processed by the diffusion model 18 along the noise addition direction to generate a set of potential noises 220. The fine-tuned text 214 may be processed by the text encoder 22 to generate a fine-tuned hint 224. The text encoder 22 may be a transformer model (e.g., BERT, T5, etc.), a CLIP model, and / or other encoders.

[0062] Potential noise 220 and fine-tuning hints 224 may be processed by diffusion model 26 in a denoising direction to generate reconstructed base image 228. Reconstructed base image 228 may be an attempt by diffusion model 26 to reconstruct base image 216.

[0063] The reconstructed base image 228 may be evaluated using one or more loss functions 30. For example, one loss function 30 may compare the reconstructed base image 228 to the base image 216. The diffusion models 26 and 18 may be trained based on the loss function 30. For example, the loss function 30 may be back-propagated through the models 26 and 18.

[0064] In some implementations, the machine learning diffusion model 26 may include a text-to-image diffusion model and one or more super-resolution diffusion models sequentially following the text-to-image diffusion model. In at least some such implementations, the text-to-image diffusion model and at least one of the one or more super-resolution diffusion models may be as follows:Figure 1B perform fine-tuning as shown.

[0065] By performing multiple iterations of the fine-tuning method shown in Figure 1B on the same base image 216 and the same fine-tuning text 214, the diffusion models 26 and 18 can learn to associate specific visual and semantic features of a particular base image 216 with the fine-tuning text 214. Thereafter, by providing a combination of the fine-tuning text 214 and an edit text that describes the desired edits to the base image 216, the diffusion model 26 can generate an output image that reflects the desired edits while still retaining the visual and semantic features of the base image 216.

[0066] Specifically, now referring to Figure 1C , a set of edit texts 314 can be received. The edit texts 314 can be natural language descriptions of the desired edits to the base image shown in Figure 1B . In the example shown in Figure 1C , the edit text 314 is "Teddy Bears". The text encoder 22 can process the edit text 314 to generate an edit prompt 324.

[0067] The diffusion model 26 along the denoising direction can process the edit prompt 324, the fine-tuning prompt 224 from Figure 1B and a set of latent noise 320 to generate an output image 328. The output image 328 can depict the desired edits performed on the base image 216.

[0068] In some implementations, the fine-tuning prompt 224 and the edit prompt 324 can be concatenated to generate a combined prompt, and the diffusion model 26 can process the combined prompt.

[0069] In some implementations, the latent noise 320 can be the same as the latent noise 220 from Figure 1B , or can be a different set of latent noise. Alternatively, in some implementations, the diffusion model 26 can process a noisy image generated from the base image 216 instead of processing the latent noise 220. For example, the noisy image can be obtained from a later layer of the diffusion model 18 along the noise addition direction.

[0070] In some implementations, processing the fine-tuning prompt 224 and the edit prompt 324 using the diffusion model 26 can include performing classifier-free guidance to bias the diffusion model towards the edit prompt 324 and away from the unconditional setting for adjustment.

[0071] In some implementations, processing the fine-tuning prompt 224 and the edit prompt 324 using the diffusion model 26 can include performing prompt guidance to bias the diffusion model 26 towards the edit prompt 324 and away from the base prompt that describes the base image for adjustment.

[0072] In some implementations, the output image 328 can be interpolated with the base image 226 to generate a second output image that may have improved visual consistency.

[0073] Example apparatus and system

[0074] Figure 2A FIG. 1 depicts a block diagram of an example computing system 100 in accordance with example embodiments of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.

[0075] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or a desktop computer), a mobile computing device (e.g., a smart phone or a tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0076] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple processors operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0077] In some implementations, the user computing device 102 can store or include one or more machine learning models 120. For example, the machine learning model 120 can be or otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including non-linear models and / or linear models. Neural networks can include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some example machine learning models can utilize an attention mechanism, such as self-attention. For example, some example machine learning models can include a multi-head self-attention model (e.g., a transformer model). Refer to Figures 1A to 1C for a discussion of example machine learning models 120.

[0078] In some implementations, one or more machine learning models 120 can be received via network 180 from server computing system 130, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing device 102 can implement multiple parallel instances of a single machine learning model 120 (e.g., to perform parallel image editing across multiple instances of an input image).

[0079] Additionally or alternatively, one or more machine learning models 140 can be included in or otherwise stored and implemented by server computing system 130, which communicates with user computing device 102 according to a client-server relationship. For example, machine learning model 140 can be implemented by server computing system 140 as part of a web service (e.g., an image editing service). Thus, one or more models 120 can be stored and implemented at user computing device 102, and / or one or more models 140 can be stored and implemented at server computing system 130.

[0080] User computing device 102 can also include one or more user input components 122 that receive user input. For example, user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to a user input object (e.g., a finger or a stylus). The touch-sensitive component can be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices by which a user can provide user input.

[0081] Server computing system 130 includes one or more processors 132 and memory 134. One or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 can store data 136 and instructions 138 that are executed by processor 132 to cause server computing system 130 to perform operations.

[0082] In some implementations, server computing system 130 includes or is otherwise implemented by one or more server computing devices. In the case where server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0083] As described above, the server computing system 130 may store or otherwise include one or more machine learning models 140. For example, the model 140 may be or otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models (e.g., transformer models). Some example models include diffusion models. Refer to Figures 1A to 1C Discuss example model 140.

[0084] The user computing device 102 and / or the server computing system 130 may train the models 120 and / or 140 via interaction with a training computing system 150 communicatively coupled via a network 180. The training computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130.

[0085] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or multiple processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 may store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or is otherwise implemented by the one or more server computing devices.

[0086] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, error backpropagation. For example, a loss function may be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters over multiple training iterations.

[0087] In some implementations, performing backpropagation of error can include performing backpropagation through truncated time. The model trainer 160 can perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0088] Specifically, the model trainer 160 can train the machine learning models 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, image and text pairs and specific base images.

[0089] In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some cases, this process can be referred to as personalizing the model.

[0090] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium (such as RAM, a hard disk, or optical or magnetic media).

[0091] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over the network 180 can be performed using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL), via any type of wired and / or wireless connection.

[0092] In some implementations, the input to the machine learning model of the present disclosure can be image data. The machine learning model can process the image data to generate an output. As an example, the machine learning model can process the image data to generate an image recognition output (e.g., recognition of the image data, latent embedding of the image data, encoded representation of the image data, hash of the image data, etc.). As another example, the machine learning model can process the image data to generate an image segmentation output. As another example, the machine learning model can process the image data to generate an image classification output. As another example, the machine learning model can process the image data to generate an image data modification output (e.g., change of the image data, etc.). As another example, the machine learning model can process the image data to generate an encoded image data output (e.g., encoded representation and / or compressed representation of the image data, etc.). As another example, the machine learning model can process the image data to generate an enlarged image data output. As another example, the machine learning model can process the image data to generate a prediction output.

[0093] In some implementations, the input to the machine learning model of the present disclosure can be text or natural language data. The machine learning model can process the text or natural language data to generate an output. As an example, the machine learning model can process the natural language data to generate a language encoding output. As another example, the machine learning model can process the text or natural language data to generate a latent text embedding output. As another example, the machine learning model can process the text or natural language data to generate a translation output. As another example, the machine learning model can process the text or natural language data to generate a classification output. As another example, the machine learning model can process the text or natural language data to generate a text segmentation output. As another example, the machine learning model can process the text or natural language data to generate a semantic intent output. As another example, the machine learning model can process the text or natural language data to generate an enlarged text or natural language output (e.g., text or natural language data with higher quality than the input text or natural language, etc.). As another example, the machine learning model can process the text or natural language data to generate a prediction output.

[0094] In some implementations, the input to the machine learning model of the present disclosure can be speech data. The machine learning model can process the speech data to generate an output. As an example, the machine learning model can process the speech data to generate a speech recognition output. As another example, the machine learning model can process the speech data to generate a speech conversion output. As another example, the machine learning model can process the speech data to generate a latent embedding output. As another example, the machine learning model can process the speech data to generate an encoded speech output (e.g., an encoded representation and / or a compressed representation of the speech data, etc.). As another example, the machine learning model can process the speech data to generate an amplified speech output (e.g., speech data with higher quality than the input speech data, etc.). As another example, the machine learning model can process the speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine learning model can process the speech data to generate a prediction output.

[0095] In some implementations, the input to the machine learning model of the present disclosure can be latent encoded data (e.g., an input latent space representation, etc.). The machine learning model can process the latent encoded data to generate an output. As an example, the machine learning model can process the latent encoded data to generate a recognition output. As another example, the machine learning model can process the latent encoded data to generate a reconstruction output. As another example, the machine learning model can process the latent encoded data to generate a search output. As another example, the machine learning model can process the latent encoded data to generate a reclustering output. As another example, the machine learning model can process the latent encoded data to generate a prediction output.

[0096] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data for one or more images and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images depict an object belonging to that object class. The image processing task can be object detection, where the image processing output identifies one or more regions in one or more images and, for each region, identifies the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in one or more images, the corresponding likelihood of each category in a predetermined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines, for each pixel in one or more images, a corresponding depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images and the image processing output defines, for each pixel in one of the input images, the motion of the scene depicted at the pixel between the images in the network input.

[0097] In some cases, the input includes audio data representing spoken utterances and the task is a speech recognition task. The output can include a text output mapped to the spoken utterance. In some cases, the task includes encrypting or decrypting the input data. In some cases, the task includes microprocessor performance tasks such as branch prediction or memory address translation.

[0098] Figure 2A An example computing system that can be used to implement the present disclosure is shown. Other computing systems can also be used. For example, in some implementations, the user computing device 102 can include a model trainer 160 and a training data set 162. In such implementations, the model 120 can be trained and used locally on the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the model 120 based on user-specific data.

[0099] Figure 2B A block diagram of an example computing device 10 performing in accordance with an example embodiment of the present disclosure is depicted. The computing device 10 can be a user computing device or a server computing device.

[0100] The computing device 10 includes a plurality of applications (e.g., Application 1 to Application N). Each application contains its own machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.

[0101] As Figure 2B shown, each application may communicate with a plurality of other components of the computing device (such as, for example, one or more sensors, a context manager, a device status component, and / or additional components). In some implementations, each application may use an API (e.g., a common API) to communicate with each device component. In some implementations, the API used by each application is specific to that application.

[0102] Figure 2C A block diagram of an example computing device 50 performing in accordance with an example embodiment of the present disclosure is depicted. The computing device 50 may be a user computing device or a server computing device.

[0103] The computing device 50 includes a plurality of applications (e.g., Application 1 to Application N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some implementations, each application may use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).

[0104] The central intelligence layer includes a plurality of machine learning models. For example, as Figure 2C shown, a corresponding machine learning model may be provided for each application, and the corresponding machine learning model may be managed by the central intelligence layer. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer may provide a single model for all applications. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of the computing device 50.

[0105] The central intelligence layer may communicate with a central device data layer. The central device data layer may be a centralized data repository for the computing device 50. As Figure 2C shown, the central device data layer may communicate with a plurality of other components of the computing device (e.g., such as one or more sensors, a context manager, a device status component, and / or additional components). In some implementations, the central device data layer may use an API (e.g., a private API) to communicate with each device component.

[0106] Additional disclosure

[0107] The technology discussed in this document relates to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a variety of possible configurations, combinations, and divisions of tasks and functions among and within components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0108] Although the subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation and not limitation of the disclosure. Those skilled in the art will readily conceive of changes, variations, and equivalents to such embodiments upon understanding the foregoing. Accordingly, the disclosure does not exclude including such modifications, variations, and / or additions to the subject matter that would be readily apparent to a person of ordinary skill in the art. For example, features shown or described as part of one embodiment can be used with another embodiment to yield yet another embodiment. Accordingly, the disclosure is intended to cover such changes, variations, and equivalents.

Claims

1. A computer-implemented method for performing text-driven image editing, the method include: Obtaining, by a computing system including one or more computing devices, a base image and an edit hint, wherein the edit hint includes or is derived from a natural language description of a desired edit to the base image; accessing, by the computing system, a machine-learned diffusion model, wherein the machine-learned diffusion model has been fine-tuned on one or more fine-tuning tuples, each fine-tuning tuple comprising the base image and a fine-tuning hint; processing, by the computing system, the fine-tuning hint and the editing hint using the machine learning diffusion model to generate an output image; and The output image is provided as output by the computing system.

2. A computer-implemented method according to any preceding claim, in, The fine-tuning hint includes or is derived from one or more rare word-grams.

3. A computer-implemented method according to any preceding claim, in, The fine-tuning hint includes or is derived from a base hint including a natural language description of the base image.

4. A computer-implemented method according to any preceding claim, in, Processing the fine-tuning hint and the editing hint by the computing system using the machine learning diffusion model includes: splicing, by the computing system, the fine-tuning hint and the editing hint to generate a combined hint; and The combined prompt is processed by the computing system using the machine learning diffusion model.

5. The computer-implemented method of any preceding claim, further comprising fine-tuning, by the computing system, the machine learning diffusion model using the one or more fine-tuning tuples.

6. A computer-implemented method according to any preceding claim, in, The machine learning diffusion model comprises a text-to-image diffusion model and one or more super-resolution diffusion models sequentially following the text-to-image diffusion model, and wherein at least one of the text-to-image diffusion model and the one or more super-resolution diffusion models has been fine-tuned using the one or more fine-tuning tuples.

7. A computer-implemented method according to any preceding claim, in, Processing the fine-tuning hint and the editing hint by the computing system using the machine learning diffusion model includes: performing classifier-free guidance by the computing system to adjust the machine learning diffusion model toward the editing hint and away from the unconditional setting.

8. A computer-implemented method according to any preceding claim, in, Processing the fine-tuning prompt and the editing prompt by the computing system using the machine learning diffusion model includes: the computing system performs prompt guidance to make the machine learning diffusion model adjust toward the editing prompt and away from the base prompt.

9. A computer-implemented method according to any preceding claim, in, Processing the fine-tuning hint and the editing hint by the computing system using the machine learning diffusion model includes: processing the fine-tuning hint and the editing hint and the noise set by the computing system using the machine learning diffusion model to generate the output image.

10. A computer-implemented method according to any one of claims 1 to 8, in, Processing the fine-tuning hint and the editing hint by the computing system using the machine learning diffusion model includes: processing the fine-tuning hint and the editing hint and the noisy image generated from the base image by the computing system using the machine learning diffusion model to generate the output image.

11. The computer-implemented method of any preceding claim, further comprising interpolating, by the computing system, the base image and the output image to generate a second output image.

12. A computer-implemented method according to any preceding claim, in, The output image depicts the desired edit performed on the base image.

13. A computer-implemented method for training a diffusion model to perform text-driven image editing, the method include: Obtaining, by a computing system including one or more computing devices, a base image and an edit hint, wherein the edit hint includes or is derived from a natural language description of a desired edit to the base image; accessing, by the computing system, a machine learning diffusion model; and The computing system fine-tunes the machine learning diffusion model on one or more fine-tuning tuples, each fine-tuning tuple comprising the base image and a fine-tuning hint.

14. A computing system configured to perform a method according to any preceding claim.

15. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform the method of any one of claims 1-13.

Citation Information

Cited By

  • Semantic image editing method and system based on physical perception

    CN120876669A