Proxy-guided image editing

By employing a two-stage image generation model, utilizing low-resolution agent guidance and teacher knowledge distillation, the problem of generating unrealistic pixels during image element deletion in existing systems is solved, achieving high-quality and efficient image editing.

CN120833413APending Publication Date: 2025-10-24ADOBE INC
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510135784.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-22
Filing Date
2025-02-07
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing image generation systems often generate unrealistic or incorrect pixels when deleting image elements, leading to a decrease in image quality and requiring a large amount of computing resources.

Method used

A two-stage image generation model is adopted. First, the first image generation model generates a proxy guide based on low-resolution input. Then, the second image generation model generates a high-resolution synthetic image based on the proxy guide. The first model is trained by knowledge distillation of the teacher's image generation model to ensure accurate deletion of image elements.

Benefits of technology

It generates more accurate synthetic images, avoids unwanted artifacts, improves image quality, and reduces computational requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833413A_ABST
    Figure CN120833413A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to agent-guided image editing. A method, apparatus, non-transitory computer readable medium and system for image processing includes obtaining an input image and an input mask, wherein the input mask indicates a region in the input image to be modified; and generating an intermediate result based on the input image and the input mask using the first image generation model, where the intermediate result modifies a region indicated by the input mask in the input image. The second image generation model generates a composite image based on the input image and the intermediate result, where the composite image depicts the input image at a higher level of detail with content from the modified region than the intermediate result.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 637,748, filed April 23, 2024, in the U.S. Patent and Trademark Office, and corresponding U.S. Non-Provisional Application No. 18 / 956,284, filed November 22, 2024, the disclosures of which are incorporated by reference herein in their entireties. TECHNICAL FIELD

[0003] The following relates generally to image processing, and more specifically to image editing using machine learning models. Image processing refers to the use of computers to edit images using algorithms or processing networks. In some cases, image processing software can be used for various image processing tasks, such as image inpainting, image detection, image generation, image compositing, and image editing. BACKGROUND

[0004] In some cases, image editing includes using a machine learning model to edit an input image based on an adjustment to generate an output image. For example, a machine learning model is trained to generate an edited image based on a textual prompt, a mask input, and / or an input image. In some cases, the edited image can depict a modification to the input image, such as a deletion of an element from the input image. SUMMARY

[0005] Aspects of the present disclosure provide a method and system for image generation. In one aspect, a system receives an input image and an input mask and generates an edited image based on the input image and the input mask. In one aspect, the system includes a first image generation model trained to generate a proxy guide based on a lower resolution input. The proxy guide is used as an input to a second image generation model to guide an image generation process. In one aspect, the system includes a teacher image generation model trained to delete an element from an input image. In one aspect, the first image generation model is trained using knowledge distilled from the teacher image generation model to generate a proxy guide. In one aspect, the second image generation model generates a composite image based on the input image, the input mask, and the proxy guide. In one aspect, one or more elements are deleted from the input image and the result is depicted in the composite image. In one aspect, the input image and the composite image are high resolution images.

[0006] A method, apparatus, non-transitory computer-readable medium, and system for image processing include obtaining an input image and an input mask, where the input mask indicates a region in the input image to be modified; generating, using a first image generation model, an intermediate result based on the input image and the input mask, where the intermediate result modifies the region in the input image indicated by the input mask; and generating, using a second image generation model, a composite image based on the input image and the intermediate result, where the composite image depicts the input image with a higher level of detail from content in the modified region than the intermediate result.

[0007] A method, apparatus, non-transitory computer-readable medium, and system for image processing include obtaining an input image and an input mask, where the input mask indicates an element in the input image to be deleted; generating, using a first image generation model, an intermediate result based on the input image and the input mask, where the intermediate result deletes the element in the input image indicated by the input mask; and generating, using a second image generation model, a composite image based on the input image and the intermediate result, where the composite image depicts the input image without the element deleted in the intermediate result.

[0008] A method, apparatus, non-transitory computer-readable medium, and system for training a machine learning model include obtaining a training set, the training set including an input image, the input image including an element; generating, using a teacher image generation model, a predicted image, the predicted image replacing the element from the input image with generated content; and training, using the training set and the predicted image, a first image generation model to replace the element of the input image with the generated content.

[0009] An apparatus and system for image processing include a memory component and a processing device coupled with the memory component, the processing device configured to perform operations including obtaining an input image and an input mask, where the input image depicts an element and the input mask indicates a region of the element in the input image; generating, using a first image generation model, an intermediate result based on the input image and the input mask, where the intermediate result includes first generated content to replace the element within the region indicated by the input mask; and generating, using a second image generation model, a composite image based on the input image and the intermediate result, where the composite image includes second generated content to replace the element within the region indicated by the input mask. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 An example of an image processing system is shown in accordance with aspects of the present disclosure.

[0011] Figure 2 An example of a method for image editing is shown in accordance with aspects of the present disclosure.

[0012] Figure 3 、 Figure 4 and Figure 5 An example of object deletion using proxy guidance is shown in accordance with aspects of the present disclosure.

[0013] Figure 6 An example of a method for image editing based on proxy guidance is shown in accordance with aspects of the present disclosure.

[0014] Figure 7 An example of an image processing device is shown in accordance with aspects of the present disclosure.

[0015] Figure 8 An example of a machine learning model is shown in accordance with aspects of the present disclosure.

[0016] Figure 9 An example of an object deletion system is shown in accordance with aspects of the present disclosure.

[0017] Figure 10 An example of an image generation model is shown in accordance with aspects of the present disclosure.

[0018] Figure 11 An example of a U-Net architecture is shown in accordance with aspects of the present disclosure.

[0019] Figure 12 An example of a diffusion process is shown in accordance with aspects of the present disclosure.

[0020] Figure 13 An embodiment of a method of training a machine learning model in accordance with various aspects of the present disclosure is shown.

[0021] Figure 14 An example of a method for training a machine learning model in accordance with aspects of the present disclosure is shown.

[0022] Figure 15 An example of a flowchart depicting an algorithm as a distributed program in an example implementation of operations executable for training a machine learning model in accordance with aspects of the present disclosure is shown.

[0023] Figure 16 An example of a method for training a diffusion model in accordance with aspects of the present disclosure is shown.

[0024] Figure 17 An example of a computing device in accordance with aspects of the present disclosure is shown. DETAILED DESCRIPTION

[0025] Aspects of the present disclosure relate to image editing using generative machine learning. Some embodiments of the present disclosure relate to image generation systems that accurately and efficiently generate synthetic images that depict modifications (e.g., deletions) to image elements from input images. In some cases, the system includes a first image generation model that is trained using a teacher image generation model to generate a proxy guide (e.g., an intermediate result). The system also includes a second image generation model that is trained to generate a synthetic image based on an input image, an input mask, and the proxy guide. The proxy guide generated by the first image generation model is provided to the second image generation model to ensure that the synthetic image accurately depicts the deletion of an object indicated by the input mask.

[0026] In the field of image editing, particularly in object deletion, machine learning systems are used to delete one or more elements from an input image. For example, these systems can identify and segment one or more objects within an input image and then fill or inpaint missing pixels in the area of the deleted one or more objects. In some cases, these systems are trained on large image datasets to understand patterns, textures, and context. However, in some cases, these systems can generate unrealistic or incorrect pixels in the area of the missing object that is deleted. In high resolution object deletion, these systems can introduce additional artifacts or affect the image quality of the generated image. In some cases, these systems require powerful computing capabilities.

[0027] In some cases, when an input is provided to delete an object from an image, conventional systems can generate a different object to replace the object to be deleted. For example, when an object mask indicating an object to be deleted is provided to a conventional system, the conventional system can generate a different object instead of deleting the target object indicated by the object mask. Thus, conventional systems are unable to accurately generate a synthetic image indicating the deletion of an object from an input image.

[0028] Accordingly, the present disclosure provides systems and methods that improve conventional image generation systems by accurately and efficiently generating synthetic images that depict the deletion of image elements from input images. This is achieved using a system that includes a first image generation model trained to generate a proxy guide and a second image generation model trained to generate a synthetic image based on the proxy guide.

[0029] According to some aspects, a system receives an input image and an input mask and generates an edited image (e.g., a synthetic image) based on the input image and the input mask. In one aspect, the system includes a first image generation model trained to generate a proxy guide based on lower resolution inputs (e.g., a low resolution input image and a low resolution input mask). In one aspect, the system includes a teacher image generation model trained to remove elements from an input image. In one aspect, the first image generation model is trained using knowledge distilled from the teacher image generation model to generate the proxy guide.

[0030] According to some aspects, the proxy guide is used as an input to a second image generation model to guide an image generation process. In one aspect, the second image generation model generates a synthetic image based on the input image, the input mask, and the proxy guide. In one aspect, one or more elements are removed from the input image and the result is depicted in the synthetic image. In one aspect, the input image and the synthetic image are high resolution images.

[0031] Reference is made to Figure 1 and Figure 17 Example systems are provided that embody inventive concepts in image processing. Reference is made to Figure 2-Figure 5 Example applications of inventive concepts in image processing are provided. Reference is made to Figure 7-11 Details relating to the architecture of an image processing device are provided. Reference is made to Figure 6 and Figure 12 Examples of processes for image processing are provided. Reference is made to Figure 13-16 A description of example training processes is provided.

[0032] Accordingly, embodiments of the present disclosure improve upon conventional image generation models by generating more accurate synthetic images. For example, embodiments generate images that accurately depict the removal of elements from an input image without unwanted artifacts (e.g., such as replacing objects with unwanted replacement objects). Some embodiments include a first image generation model trained to generate a proxy guide based on a low resolution input image. The proxy guide is provided to a second image generation model to generate a high resolution synthetic image based on an input image. Accordingly, by using the proxy guide to guide the diffusion process of the second image generation model, the system is able to accurately generate images that depict the removal of elements from an input image.

[0033] Image editing

[0034] In Figures 1-6 and Figure 12In some aspects, methods, apparatus, non-transitory computer-readable media, and systems for image processing include obtaining an input image and an input mask, where the input image depicts an element and the input mask indicates a region of the element in the input image; generating, using a first image generation model, an intermediate result based on the input image and the input mask, where the intermediate result includes first generated content to replace the element within the region indicated by the input mask; and generating, using a second image generation model, a composite image based on the input image and the intermediate result, where the composite image includes second generated content to replace the element within the region indicated by the input mask.

[0035] Some embodiments include obtaining an input image and an input mask, where the input mask indicates a region of the input image to be modified; generating, using a first image generation model, an intermediate result based on the input image and the input mask, where the intermediate result modifies the region of the input image indicated by the input mask; and generating, using a second image generation model, a composite image based on the input image and the intermediate result, where the composite image depicts the input image with a higher level of detail from content in the modified region compared to the intermediate result. The composite image can include content with higher resolution or additional texture detail compared to the intermediate result. In some embodiments, the first image generation model is a smaller model (e.g., it can have fewer layers or fewer parameters) than the second image generation model, but it is specifically trained for an object removal task.

[0036] Some examples of methods, apparatus, non-transitory computer-readable media, and systems further include segmenting the input image to identify the region of the element in the input image. Some examples of methods, apparatus, non-transitory computer-readable media, and systems further include receiving a position input. Some examples further include generating the input mask based on the position input.

[0037] Some examples of methods, apparatus, non-transitory computer-readable media, and systems further include receiving a deletion prompt, where the deletion prompt includes a command to delete the element from the input image. Some examples further include selecting a deletion mode based on the deletion prompt, where the intermediate result is generated based on the deletion mode.

[0038] In some aspects, the intermediate result includes an intermediate image with lower resolution compared to the composite image. In some aspects, the first image generation model has fewer parameters than the second image generation model. In some aspects, the first image generation model is trained to replace the element using a predicted image generated by a teacher image generation model.

[0039] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include obtaining a first noisy input. Some examples further include denoising the first noisy input to obtain an intermediate result. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include obtaining a second noisy input. Some examples further include denoising the second noisy input to generate a synthetic image.

[0040] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include inpainting an area indicated by the input mask with content consistent with the input image. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include generating a plurality of synthetic images, the plurality of synthetic images including different generated content in place of the element.

[0041] According to some aspects, a method, apparatus, non-transitory computer-readable medium, and system for image processing include obtaining an input image and an input mask, where the input image depicts an element and the input mask indicates a location of the element in the input image; generating, using a first image generation model, an intermediate result based on the input image and the input mask, where the intermediate result includes the element removed from the input image; and generating, using a second image generation model, a synthetic image based on the intermediate result, where the synthetic image includes the element removed from the input image.

[0042] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include segmenting the input image based on the element. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include receiving a location input from a user. Some examples further include generating the input mask based on the location input.

[0043] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include receiving a deletion prompt from a user. In some cases, the deletion prompt includes a command to remove the element from the input image. Some examples further include selecting a deletion mode based on the deletion prompt. In some cases, the intermediate result is generated based on the deletion mode.

[0044] In some aspects, the intermediate result includes an intermediate image having a lower resolution than the synthetic image. In some aspects, the first image generation model has fewer parameters than the second image generation model. In some aspects, the first image generation model is trained based on an object removal task. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include inpainting an area indicated by the input mask with content consistent with the input image.

[0045] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include obtaining a first noise input. Some examples further include performing a first diffusion process on the first noise input. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include obtaining a second noise input. Some examples further include performing a second diffusion process on the second noise input and the intermediate result.

[0046] Figure 1 An example of an image processing system is shown in accordance with aspects of the present disclosure. The shown example includes a user 100, a user device 105, an image processing apparatus 110, a cloud 115, and a database 120. The image processing apparatus 110 is described with reference to Figure 7 described with reference to Figure 7 described with reference to

[0047] described with reference to Figure 1 The user 100 provides an input image and an input mask to the image processing apparatus 110 via the user device 105 and the cloud 115. For example, the input image depicts two cups of coffee on a table. The input mask indicates a rough area on the table where the first cup of iced coffee is located, for example. In some cases, the user 100 can provide an additional command, such as “delete,” to the image processing apparatus 110 to delete the object indicated by the input mask. The machine learning model of the image processing apparatus 110 then generates an intermediate image from a student agent image generation model (e.g., a first image generation model) based on the input image and the input mask. For example, in some cases, the intermediate image is a low resolution edited image depicting one cup of iced coffee (e.g., the first cup of iced coffee on the lower left side of the input image is deleted). The intermediate image is used as a proxy guide for a second image generation model to generate a synthetic image (high resolution) based on the input image and the input mask. For example, the synthetic image depicts an edited image of the input image with no coffee cup indicated by the input mask. The image processing apparatus 110 displays the synthetic image to the user 100 via the user device 105 and the cloud 115.

[0048] The user device 105 can be a personal computer, a laptop, a mainframe computer, a palmtop computer, a personal assistant, a mobile device, or any other suitable processing apparatus. In some examples, the user device 105 includes software that incorporates an image processing application. In some examples, the image processing application on the user device 105 can include the functionality of the image processing apparatus 110.

[0049] The user interface can enable the user 100 to interact with the user device 105. In some embodiments, the user interface can include an audio device such as an external speaker system, an external display device such as a display screen, or an input device (e.g., a remote control device that interfaces with the user interface directly or via an I / O controller module). In some cases, the user interface can be a graphical user interface (GUI). In some examples, the user interface can be represented in code, where the code is sent to the user device 105 and rendered locally by a browser. The process using the image processing apparatus 110 is described with reference to Figure 2 Further described.

[0050] According to some aspects, the image processing apparatus 110 comprises a computer-implemented network comprising a machine learning model, a first image generation model, and a second image generation model. The image processing apparatus 110 further comprises a processor unit, a memory unit, an I / O module, a training component, and a teacher image generation model. In some embodiments, the image processing apparatus 110 further comprises the communication interface, the user interface component, and the bus described with reference to Figure 17 Additionally or alternatively, the image processing apparatus 110 communicates with the user device 105 and the database 120 via the cloud 115. The image processing apparatus 110 is described with reference to Figure 7 described examples of elements or comprises aspects of elements described with reference to Figure 7 described examples of elements or comprises aspects of elements described with reference to Figure 2 described.

[0051] In some cases, the image processing apparatus 110 is implemented on a server. A server provides one or more functions for users linked through one or more of a variety of networks. In some cases, a server includes a single microprocessor board that includes a microprocessor responsible for controlling aspects of the server. In some cases, a server uses microprocessors and protocols to exchange data with other devices / users on one or more networks via the hypertext transfer protocol (HTTP) and the simple mail transfer protocol (SMTP), although other protocols can be used, such as the file transfer protocol (FTP) and the simple network management protocol (SNMP). In some cases, a server is configured to send and receive files in hypertext markup language (HTML) format (e.g., for displaying web pages). In various embodiments, a server includes a general computing device, a personal computer, a notebook computer, a mainframe computer, a supercomputer, or any other suitable processing apparatus.

[0052] The cloud 115 is a computer network configured to provide on-demand availability of computer system resources, such as data storage and computing power. In some examples, the cloud 115 provides resources without active management by a user (e.g., the user 100). The term cloud is sometimes used to describe a data center that is available to many users over the Internet. Some large cloud networks have functionality distributed over multiple locations from a central server. If a server has a direct or close connection to a user, the server is designated as an edge server. In some cases, the cloud 115 is limited to a single organization. In other examples, the cloud 115 is available to many organizations. In one example, the cloud 115 includes a multi-tiered communication network with multiple edge routers and core routers. In another example, the cloud 115 is based on a set of local switches in a single physical location.

[0053] According to some aspects, the database 120 stores training data (or training set), which includes input images that include elements. The database 120 is a collection of organized data. For example, the database 120 stores data in a specified format known as a schema. The database 120 can be architected as a single database, a distributed database, multiple distributed databases, or an emergency backup database. In some cases, a database controller can manage data storage and processing in the database 120. In some cases, a user (e.g., the user 100) interacts with the database controller. In other cases, the database controller can operate automatically without user interaction.

[0054] Figure 2 An example of a method 200 for image editing according to aspects of the present disclosure is shown. In some examples, the operations are performed by a system including a processor executing a set of codes to control functional elements of a device. Additionally or alternatively, some processes are performed using special-purpose hardware. Generally, the operations are performed in accordance with the methods and processes described according to aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps, or performed in conjunction with other operations.

[0055] At operation 205, the system provides an input image and an input mask. In some cases, the operations of this step refer to or can be performed by a user as described with reference to Figure 1 For example, the input image depicts two cups of iced coffee on a table. For example, the input mask describes a region of elements to be removed from the input image. For example, in some cases, a command can also be provided to the system in addition to the input image and the input mask. For example, the command can be “remove.”

[0056] At operation 210, the system generates a conditional guidance result. In some cases, the operations of this step refer to or can be performed by a user as described with reference to Figure 1and Figure 7 The described image processing apparatus to perform. In some cases, the operations of this step refer to or can be performed by an image processing apparatus as described with reference to Figure 7-Figure 9 and Figure 14 The described first image generation model to perform. In some cases, the system generates a low resolution proxy result based on the input image and the input mask. In some cases, the conditional guidance result includes the low resolution proxy result generated by the first image generation model. The low resolution proxy result is used as guidance to guide the image generation process of the second image generation model to generate the synthetic image.

[0057] At operation 215, the system initializes a noise input. In some cases, the operations of this step refer to or can be performed by an image processing apparatus as described with reference to Figure 1 and Figure 7 The described image processing apparatus to perform. In some cases, the operations of this step refer to or can be performed by an image processing apparatus as described with reference to Figure 7-Figure 9 and Figure 14 The described second image generation model to perform. In some cases, the noise input including random noise is initialized. The noise input can be in a latent space. By initializing the image generation model with random noise, different variations of the synthetic image can be generated. In some cases, a conditional embedding, such as a text encoding or text embedding, can be combined with the noise features to guide the image generation process using cross-attention blocks within the image generation model. More details regarding the image generation process are described with reference to Figure 10 .

[0058] At operation 220, the system generates media content. In some cases, the operations of this step refer to or can be performed by an image processing apparatus as described with reference to Figure 1 and Figure 7 The described image processing apparatus to perform. In some cases, the operations of this step refer to or can be performed by an image processing apparatus as described with reference to Figure 7-Figure 9 and Figure 14 The described second image generation model to perform. For example, in some cases, the media content includes synthetic images or modified images each depicting the input image with a cup of iced coffee removed. For example, the synthetic images include image pixels generated by the image generation model. For example, the modified images include image pixels from the input image and image pixels generated by the image generation model. In some cases, the media content is displayed to a user via a user device.

[0059] Figure 3 An example of object removal using proxy guidance is shown in accordance with aspects of the present disclosure. The shown example includes an image editing system 300, an input image 305, an input mask 310, a machine learning model 315, a synthetic image 320, a regular model 325, and a regular synthetic image 330.

[0060] refer to Figure 3 , the machine learning model 315 receives the input image 305 and the input mask 310 to generate a composite image 320. For example, the input image 305 depicts a person holding a writing tablet and a pen. For example, the input mask 310 indicates an area (or multiple areas) representing an object to be deleted (e.g., a writing tablet and a pen). In some cases, the user can provide a rough sketch indicating the location of the object to be deleted. The machine learning model 315 can then generate a precise mask based on the rough sketch. For example, the machine learning model can segment the input image 305 to obtain multiple segmented objects, and each of the multiple segmented objects represents an object in the input image 305.

[0061] In some embodiments, the machine learning model 315 generates an intermediate result based on the input image 305 and the input mask 310. For example, the intermediate result is a low-resolution image that depicts the input image 305 without the elements indicated by the input mask 310. In some cases, the intermediate result is used as guidance to direct the image generation process of the image generation model of the machine learning model 315 to generate the synthesized image 320. For example, the synthesized image 320 is a high-resolution image that depicts the input image 305 without the elements indicated by the input mask 310.

[0062] In contrast to the composite image, conventional composite image 330 generated by conventional model 325 depicts the replacement of the object indicated by input mask 310, rather than the deletion of the object. For example, the left image of conventional composite image 330 depicts the deletion of the pen and the replacement of the writing board with a mobile phone. For example, the center image of conventional composite image 330 depicts the deletion of the pen and the replacement of the writing board with a different writing board. For example, the right image of conventional composite image 330 depicts the deletion of the pen and the replacement of the writing board with a stack of paper towels.

[0063] The image editing system 300 is a reference Figure 4 、 Figure 5 and Figure 9 Examples of corresponding elements described or including references Figure 4 、 Figure 5 and Figure 9 The input image 305 is a reference to the corresponding elements. Figure 4 、 Figure 5 、 Figure 8 、 Figure 9 and Figure 14 Examples of corresponding elements described or including references Figure 4 、 Figure 5 、 Figure 8 、 Figure 9 and Figure 14 The input mask 310 is a reference to the corresponding elements of the description. Figure 4 、 Figure 5 , Figure 8 and Figure 9 described with reference to Figure 4 , Figure 5 , Figure 8 and Figure 9 described with reference to

[0064] The machine learning model 315 is an example of, or includes aspects of, the corresponding element described with reference to Figure 4 , Figure 5 and Figure 7 described with reference to Figure 4 , Figure 5 and Figure 7 described with reference to Figure 4 , Figure 5 and Figure 8 described with reference to Figure 4 , Figure 5 and Figure 8 described with reference to Figure 4 and Figure 5 described with reference to Figure 4 and Figure 5 described with reference to Figure 4 and Figure 5 described with reference to Figure 4 and Figure 5 described with reference to

[0065] Figure 4 An example of object removal using proxy guidance is shown in accordance with aspects of the present disclosure. The shown example includes an image editing system 400, an input image 405, an input mask 410, a machine learning model 415, a synthetic image 420, a canonical model 425, and a canonical synthetic image 430.

[0066] The machine learning model 415 receives the input image 405 and the input mask 410 to generate the synthetic image 420, with reference to Figure 4 For example, the input image 405 depicts a sculpture on a sofa. For example, the input mask 410 indicates a region representing an object (e.g., the sculpture) to be removed. In some cases, a user can provide a rough sketch indicating a location of the object to be removed. The machine learning model 415 can then generate a precise mask representing the object based on the rough sketch. For example, the machine learning model can segment the input image 405 to obtain a plurality of segmented objects, and each of the plurality of segmented objects represents an object in the input image 405.

[0067] In some embodiments, machine learning model 415 generates an intermediate result based on input image 405 and input mask 410. For example, the intermediate result is a low-resolution image that depicts input image 405 without the elements indicated by input mask 410. In some cases, the intermediate result is used as guidance to direct the image generation process of the image generation model of machine learning model 415 to generate synthetic image 420. For example, synthetic image 420 is a high-resolution image that depicts input image 405 without the elements indicated by input mask 410. In contrast, regular model 425 generates regular synthetic image 430 that has replaced objects instead of the deleted objects indicated by input mask 410.

[0068] Image editing system 400 is a reference Figure 3 、 Figure 5 and Figure 9 Examples of corresponding elements described or including references Figure 3 、 Figure 5 and Figure 9 The input image 405 is a reference to the corresponding elements. Figure 3 、 Figure 5 、 Figure 8 、 Figure 9 and Figure 14 Examples of corresponding elements described or including references Figure 3 、 Figure 5 、 Figure 8 、 Figure 9 and Figure 14 The input mask 410 is a reference to the corresponding elements of the description. Figure 3 、 Figure 5 、 Figure 8 and Figure 9 Examples of corresponding elements described or including references Figure 3 、 Figure 5 、 Figure 8 and Figure 9 Describes aspects of the corresponding element.

[0069] Machine Learning Model 415 is a reference Figure 3 、 Figure 5 and Figure 7 Examples of corresponding elements described or including references Figure 3 、 Figure 5 and Figure 7 The composite image 420 is a reference to the corresponding elements of the description. Figure 3 、 Figure 5 and Figure 8 Examples of corresponding elements described or including references Figure 3 、 Figure 5 and Figure 8 The conventional model 425 is a reference to the corresponding elements of the Figure 3 and Figure 5 described corresponding elements or include references to Figure 3 and Figure 5 aspects of the corresponding elements described. The regular synthetic image 430 is an example of or includes references to Figure 3 and Figure 5 described corresponding elements or include references to Figure 3 and Figure 5 aspects of the corresponding elements described.

[0070] Figure 5 An example of object removal using proxy guidance is shown in accordance with aspects of the present disclosure. The shown example includes an image editing system 500, an input image 505, an input mask 510, a machine learning model 515, a synthetic image 520, a regular model 525, and a regular synthetic image 530.

[0071] The image editing system 500 is an example of or includes references to Figure 5 , the machine learning model 515 receives the input image 505 and the input mask 510 to generate the synthetic image 520. For example, the input image 505 depicts two cups of coffee on a table. The input mask 510 indicates a region representing an object to be removed (e.g., the first cup of coffee in the front), for example. In some cases, a user can provide a rough sketch indicating a location of the object to be removed. The machine learning model 515 can then generate a precise mask representing the object based on the rough sketch. For example, the machine learning model 515 can segment the input image 505 to obtain a plurality of segmented objects, and each of the plurality of segmented objects represents an object in the input image 505.

[0072] In some embodiments, the machine learning model 515 generates an intermediate result based on the input image 505 and the input mask 510. The intermediate result is a low resolution image depicting the input image 505 without the element indicated by the input mask 510, for example. In some cases, the intermediate result is used as guidance to guide an image generation process of an image generation model of the machine learning model 515 to generate the synthetic image 520. The synthetic image 520 is a high resolution image depicting the input image 505 without the element indicated by the input mask 510, for example. In contrast, the regular model 525 generates the regular synthetic image 530, which has an object replaced instead of the removed object indicated by the input mask 510.

[0073] The image editing system 500 is an example of or includes references to Figure 3 , Figure 4 and Figure 9 described corresponding elements or include references to Figure 3 , Figure 4 and Figure 9Aspects of the described corresponding elements. Input image 505 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , 4 , 8, 9, and 14. Input mask 510 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , 4 , 8, 9, and 14. Input mask 510 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , Figure 4 , Figure 8 , and Figure 9 . Synthetic image 520 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , Figure 4 , Figure 8 , and Figure 9 .

[0074] Machine learning model 515 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , Figure 4 , and Figure 7 . Synthetic image 520 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , Figure 4 , and Figure 7 . Synthetic image 520 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , Figure 4 , and Figure 8 . Regular model 525 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , Figure 4 , and Figure 8 . Regular synthetic image 530 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , and Figure 4 . Regular synthetic image 530 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , and Figure 4 . Regular synthetic image 530 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , and Figure 4 . Regular synthetic image 530 is an example of, or includes aspects of, the corresponding elements described with reference to Figure 3 , and Figure 4 .

[0075] Figure 6 An example of a method 600 for image editing based on agent guidance, in accordance with aspects of the present disclosure, is shown. In some examples, these operations are performed by a system including a processor that executes a set of codes to control the functional elements of a device. Additionally or alternatively, some processes are performed using special-purpose hardware. Generally, these operations are performed in accordance with the methods and processes described with respect to aspects of the present disclosure. In some instances, the operations described herein are performed by various sub-steps or in conjunction with other operations.

[0076] At operation 605, the system obtains an input image and an input mask, where the input image depicts an element and the input mask indicates a region of the element in the input image. In some instances, the operations of this step are performed with reference to or can be executed by a machine learning model as described with reference to Figure 3-Figure 5 and Figure 7 In some instances, the element can be substantially the same as an image element. For example, an image element is an image component or image feature that makes up the whole composition of an image (e.g., the input image), such as an object, entity, subject, shape, color, texture, pattern, background scene, visual attribute, and / or style. For example, an image element can be an animal (such as a cat or dog), a person, an object (such as a hat or table), a scene (such as a beach or mountaintop), or a combination thereof. For example, an image element or element can be a cup of iced coffee as depicted in Figure 5

[0077] At operation 610, the system generates, using a first image generation model, an intermediate result based on the input image and the input mask, where the intermediate result includes first generated content that replaces the element within the region indicated by the input mask. In some instances, the operations of this step are performed with reference to or can be executed by a first image generation model as described with reference to Figure 7-Figure 9 and Figure 14 In some instances, the first image generation model is a student agent model that is trained to generate a low-resolution output based on a low-resolution input. In some instances, a teacher image generation model (or agent teacher model) distills knowledge in the teacher image generation model to the student agent model. For example, the student agent model is a smaller, simpler, and faster model than the teacher agent model. For example, in some instances, the teacher agent model is a high-capacity model that is trained on a large dataset. By distilling knowledge from the teacher agent model to the student agent model, the student agent model is able to accurately generate a low-resolution output while reducing computational costs. In some instances, the intermediate result includes the low-resolution output. In some instances, the intermediate result is used as a guide for a second image generation model to guide the image generation process.

[0078] At operation 615, the system generates, using a second image generation model, a synthetic image based on the input image and the intermediate result, where the synthetic image includes second generated content that replaces the element within the region indicated by the input mask. In some instances, the operations of this step are performed with reference to or can be executed by a second image generation model as described with reference to Figure 7-Figure 9 and Figure 14 ​The described second image generation model is executed. In some cases, the second image generation model conditionally processes the reference image (e.g., the intermediate result) to generate the synthetic image. For example, the intermediate result is a low resolution image that depicts the input image without the objects indicated by the input mask. The second image generation model is guided based on the low resolution image to generate the synthetic image. In some cases, the second image generation model is a larger model compared to the first image generation model. For example, the second image generation model can receive a high resolution input image and generate a high resolution synthetic image.

[0079] In some cases, the second image generation model can generate a synthetic image or a modified image. For example, the synthetic image includes image pixels generated by the image generation model. For example, the modified image includes image pixels from the input image and image pixels generated by the image generation model.

[0080] System Architecture

[0081] In Figure 7-11 and Figure 17 In some cases, the second image generation model can generate a synthetic image or a modified image. For example, the synthetic image includes image pixels generated by the image generation model. For example, the modified image includes image pixels from the input image and image pixels generated by the image generation model.

[0082] Some examples of the apparatus and system further include a teacher image generation model, where the teacher image generation model is trained to delete the element from the input image. In some aspects, the first image generation model and the second image generation model are diffusion models. In some aspects, the first image generation model has fewer parameters than the second image generation model.

[0083] Figure 7 An example of an image processing apparatus 700 is shown in accordance with aspects of the present disclosure. The illustrated example includes an image processing apparatus 700, a processor unit 705, an I / O module 710, a memory unit 715, a training component 735, and a teacher image generation model 740. In one aspect, the memory unit 715 includes a machine learning model 720, a first image generation model 725, and a second image generation model 730.

[0084] According to some embodiments of the present disclosure, the image processing device 700 includes a computer-implemented artificial neural network (ANN). ANNs are hardware or software components that include several connected nodes (e.g., artificial neurons) that loosely correspond to neurons in the human brain. Each connection or edge transmits a signal from one node to another (just like a physical synapse in the brain). When a node receives a signal, the node processes the signal and then transmits the processed signal to other connected nodes. In some cases, the signals between nodes include real numbers and the output of each node is computed by the sum of its inputs. In some examples, the nodes can use other mathematical algorithms (e.g., select the maximum value from the inputs as the output) or any other suitable algorithm to determine the output to activate the node. Each node and edge is associated with one or more node weights that determine how the signal is processed and transmitted. The image processing device 700 is an example of a processor described with reference to Figure 1 described corresponding elements, or includes aspects of the described corresponding elements. Figure 1 described corresponding elements, or includes aspects of the described corresponding elements.

[0085] The processor unit 705 is a smart hardware device (e.g., a general- purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a separate hardware component, or any combination thereof). In some cases, the processor unit 705 is configured to operate a memory array using a memory controller. In other cases, a memory controller is integrated into the processor. In some cases, the processor unit 705 is configured to execute computer-readable instructions stored in a memory to perform various Figure 17 described corresponding elements, or includes aspects of the described corresponding elements. Figure 17 described corresponding elements, or includes aspects of the described corresponding elements.

[0086] The I / O module 710 (e.g., an input / output interface) can include an I / O controller. The I / O controller can manage input and output signals for the device. The I / O controller can also manage peripherals not integrated to the device. In some cases, the I / O controller can represent a physical connection or port to an external peripheral. In some cases, the I / O controller can utilize an operating system such as or other known operating systems. In other cases, the I / O controller can represent or interact with a modem, a keyboard, a mouse, a touch screen, or similar device. In some cases, the I / O controller can be implemented as part of a processor. In some cases, a user can interact with the device via the I / O controller or via hardware components controlled by the I / O controller.

[0087] In some examples, the I / O module 710 includes a user interface. The user interface can enable a user to interact with the device. In some embodiments, the user interface can include an audio device, such as an external speaker system, an external display device, such as a display screen, or an input device (e.g., a remote control device that interfaces with the user interface directly or with the aid of the I / O controller module). In some cases, the user interface can be a graphical user interface (GUI). In some examples, the communication interface operates at the boundary between a communication entity and a channel and can also record and process communications. A communication interface is provided herein to enable a processing system coupled with a transceiver (e.g., a transmitter and / or a receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for the communication device via an antenna. The I / O module 710 is described with reference to Figure 17 Examples of the I / O interface described or aspects of the I / O interface described are included with reference to Figure 17 Examples of the I / O interface described or aspects of the I / O interface described are included with reference to

[0088] Examples of the memory unit 715 include random access memory (RAM), read only memory (ROM), or hard drives. Examples of the memory unit 715 include solid state memory and hard disk drives. In some examples, the memory unit 715 is used to store computer-readable, computer-executable software including instructions that, when executed, cause the processor to perform various functions described herein.

[0089] In some cases, the memory unit 715 includes, among other things, a basic input / output system (BIOS) that controls basic hardware or software operation such as interaction with peripheral components or devices. In some cases, a memory controller operates the memory unit. For example, the memory controller can include a row decoder, a column decoder, or both. In some cases, memory units within the memory unit 715 store information in the form of logical states.

[0090] In one aspect, the memory unit 715 includes a machine learning model 720, a first image generation model 725, and a second image generation model 730. In some aspects, the machine learning model 720 includes the first image generation model 725 and the second image generation model 730. The memory unit 715 is described with reference to Figure 17Examples of the described memory subsystem or include references Figure 17 Aspects of the described memory subsystem.

[0091] In some cases, the machine learning model 720 is a computational algorithm, model, or system designed to recognize patterns, make predictions, or perform specific tasks (e.g., image processing) without being explicitly programmed. According to some aspects, the machine learning model 720 is implemented as software stored in the memory unit 715 and can be executed by the processor unit 705 as firmware, as one or more hardware circuits, or as a combination thereof.

[0092] According to some embodiments of the present disclosure, the machine learning model 720 includes an ANN, which is a hardware or software component that includes several connected nodes (e.g., artificial neurons) that loosely correspond to neurons in the human brain. Each connection or edge transmits a signal from one node to another (just like a physical synapse in the brain). When a node receives a signal, the node processes the signal and then transmits the processed signal to other connected nodes. In some cases, the signals between nodes include real numbers and the output of each node is computed by the sum of its inputs. In some examples, the nodes can use other mathematical algorithms (e.g., select the maximum value from the inputs as the output) or any other suitable algorithm to determine the output to activate the node. Each node and edge is associated with one or more node weights that determine how the signal is processed and transmitted.

[0093] During the training process, the one or more node weights are adjusted to improve the accuracy of the results (e.g., by minimizing a loss function that corresponds in some way to the difference between the current results and target results). The weights of the edges increase or decrease the strength of the signal transmitted between nodes. In some cases, the nodes have a threshold below which the signal is not transmitted at all. In some examples, the nodes are aggregated into layers. Different layers perform different transformations on the respective inputs. The initial layer is called the input layer and the last layer is called the output layer. In some cases, the signals pass through certain layers multiple times.

[0094] According to some embodiments, the machine learning model 720 includes a computer- implemented convolutional neural network (CNN). CNNs are a class of neural networks commonly used in computer vision or image classification systems. In some cases, CNNs can enable the processing of digital images with minimal pre-processing. CNNs are characterized by the use of convolutional (or cross-correlation) hidden layers. These layers apply a convolution operation to the input before signaling the result to the next layer. Each convolutional node can process data for a limited input field (e.g., a receptive field). During forward pass of the CNN, the filters at each layer can be convolved across the input volume, computing the dot product between the filter and the input. During the training process, the filters can be modified such that the filters are activated when they detect a particular feature within the input.

[0095] In one aspect, the machine learning model 720 includes machine learning parameters. Also referred to as model parameters or weights, the machine learning parameters are the variables that provide the behavior and characteristics of the machine learning model 720. The machine learning parameters can be learned or estimated from training data and used to make predictions or perform tasks based on the patterns and relationships learned in the data.

[0096] The machine learning parameters are adjusted during a training process to minimize a loss function or maximize a performance metric. The goal of the training process is to find the best values for the parameters that allow the machine learning model 720 to make accurate predictions or perform well in a given task.

[0097] For example, during the training process, the algorithm adjusts the machine learning parameters according to an optimization technique such as gradient descent, stochastic gradient descent, or other optimization algorithm to minimize the error or loss between the predicted output and the actual target. Once the machine learning parameters are learned from the training data, the machine learning parameters can be used to make predictions on new, unseen data.

[0098] According to some embodiments, the machine learning model 720 includes a computer- implemented recurrent neural network (RNN). RNNs are a class of ANN in which the connections between nodes form a directed graph along an ordered (e.g., temporal) sequence. This enables RNNs to model temporal dynamic behavior, such as predicting which element in a sequence will come next. Thus, RNNs are suitable for tasks involving ordered sequences, such as text recognition (ordering of words in a sentence). In some cases, the RNN includes one or more feed-forward networks (characterized by nodes forming a directed acyclic graph), one or more recurrent networks (characterized by nodes forming a directed cyclic graph), or a combination thereof.

[0099] According to some embodiments, the machine learning model 720 includes a transformer (or transformer model, or transformer network), where the transformer is a type of neural network model used for natural language processing tasks. The transformer network transforms one sequence into another sequence using an encoder and a decoder. The encoder and decoder include modules that can be stacked on top of each other multiple times. The modules include multi-headed attention and feed forward layers. The input and output (target sentence) are first embedded into an n-dimensional space. Positional encodings of different words (e.g., giving a relative position for each word / portion in the sequence, as the sequence is related to the order of its elements) are added to the embedding representation (n-dimensional vector) of each word. In some examples, the transformer network includes an attention mechanism, where the attention looks at the input sequence and decides at each step which other parts of the sequence are important. The attention mechanism involves queries, keys, and values, denoted by Q, K, and V, respectively. Q is a matrix containing the queries (vector representations of one word in the sequence), K is the keys (vector representations of words in the sequence), and V is the values, which are again vector representations of words in the sequence. For the encoder and decoder, multi-headed attention modules, V consists of the same sequence of words as Q. However, for the attention modules that consider both the encoder and decoder sequences, V is different from the sequence represented by Q. In some cases, the values in V are multiplied by some attention weights a and summed.

[0100] In the field of machine learning, an attention mechanism (e.g., implemented in one or more ANNs) is a method of placing different levels of importance on different input elements. Computing attention can involve three basic steps. First, a similarity between queries and key vectors obtained from the input is computed to generate attention weights. The similarity function used for this process can include dot product, concatenation, detectors, etc. Next, a softmax function is used to normalize the attention weights. Finally, the attention weights are weighted together with corresponding values. In the context of attention networks, the keys and values are vectors or matrices used to represent the input data. The keys are used to determine which parts of the input the attention mechanism should focus on, while the values are used to represent the actual data being processed.

[0101] Attention mechanisms are a key component in some ANN architectures, particularly ANNs used in natural language processing (NLP) and sequence-to-sequence tasks, that allow the ANN to focus on different parts of the input sequence when making predictions or generating outputs. Some sequence models, such as RNNs, process the input sequence in order, maintaining an internal hidden state to capture information from previous steps. However, in some cases, this sequential processing makes it difficult to capture long-range dependencies or focus on specific parts of the input sequence.

[0102] Attention mechanisms address these difficulties by enabling ANNs to selectively focus on different parts of the input sequence, assigning different degrees of importance or attention to each part. Attention mechanisms enable selective focus by considering the relevance of each input element with respect to the current state of the ANN.

[0103] The term“self-attention” refers to a machine learning model 720 in which the representations of the inputs interact with each other to determine the attention weights for the inputs. Because the attention weights are determined at least in part by the inputs, self-attention can be distinguished from other attention models.

[0104] According to some aspects, the machine learning model 720 takes an input image and an input mask, where the input image depicts an element and the input mask indicates a region of the input image in which the element is located. In some examples, the machine learning model 720 segments the input image to identify the region of the input image in which the element is located. In some examples, the machine learning model 720 receives a position input. In some examples, the machine learning model 720 generates the input mask based on the position input.

[0105] In some examples, the machine learning model 720 receives a deletion cue, where the deletion cue includes a command to delete the element from the input image. In some examples, the machine learning model 720 selects a deletion mode based on the deletion cue, where the intermediate result is generated based on the deletion mode. The machine learning model 720 is described with reference to Figure 3-Figure 5 Examples of corresponding elements described or include aspects of Figure 3-Figure 5 Examples of corresponding elements described or include aspects of

[0106] According to some aspects, the first image generation model 725 is implemented as software stored in the memory unit 715 and executable by the processor unit 705, as firmware, one or more hardware circuits, or a combination thereof. According to some aspects, the first image generation model 725 generates the intermediate result based on the input image and the input mask, where the intermediate result includes first generated content that replaces the element within the region indicated by the input mask. In some aspects, the intermediate result includes an intermediate image that has a lower resolution than the composite image. In some aspects, the first image generation model 725 has fewer parameters than the second image generation model 730.

[0107] In some examples, the first image generation model 725 takes a first noise input. In some examples, the first image generation model 725 denoises the first noise input to obtain the intermediate result. In some aspects, the first image generation model 725 is trained to replace the element using a predicted image generated by the teacher image generation model 740. In some aspects, the first image generation model 725 has fewer parameters than the teacher image generation model 740.

[0108] According to some aspects, the first image generation model 725 generates, based on the input image and the input mask, an intermediate result, where the intermediate result includes first generated content that replaces elements within the region indicated by the input mask. In some aspects, the first image generation model 725 and the second image generation model 730 are diffusion models. In some aspects, the first image generation model 725 has fewer parameters than the second image generation model 730. The first image generation model 725 is an example of, or includes aspects of, the corresponding elements described in reference to Figure 8 、 Figure 9 and Figure 14 . Figure 8 、 Figure 9 and Figure 14 .

[0109] According to some aspects, the second image generation model 730 is implemented as software stored in the memory unit 715 and executable by the processor unit 705, as firmware, one or more hardware circuits, or a combination thereof. According to some aspects, the second image generation model 730 generates, based on the input image and the intermediate result, a composite image, where the composite image includes second generated content that replaces elements within the region indicated by the input mask. In some examples, the second image generation model 730 takes a second noise input. In some examples, the second image generation model 730 denoises the second noise input to generate the composite image.

[0110] In some examples, the second image generation model 730 internally fills the region indicated by the input mask with content consistent with the input image. In some examples, the second image generation model 730 generates a set of composite images, the set of composite images including different generated content that replaces the elements. According to some aspects, the second image generation model 730 generates, based on the input image and the intermediate result, a composite image, where the composite image includes second generated content that replaces elements within the region indicated by the input mask. The second image generation model 730 is an example of, or includes aspects of, the corresponding elements described in reference to Figure 8 、 Figure 9 and Figure 14 . Figure 8 、 Figure 9 and Figure 14 .

[0111] According to some aspects, the training component 735 is implemented as software stored in the memory unit 715 and executable by the processor unit 705, as firmware, as one or more hardware circuits, or a combination thereof. According to some embodiments, the training component 735 is implemented as software stored in the memory unit and executable by a processor of a processor unit of a separate computing device, as firmware in the separate computing device, as one or more hardware circuits of the separate computing device, or a combination thereof. In some examples, the training component 735 is part of another device other than the image processing apparatus 700 and is in communication with the image processing apparatus 700. In some examples, the training component 735 is part of the image processing apparatus 700.

[0112] According to some aspects, the training component 735 obtains a training set, the training set comprising input images, the input images including elements. In some examples, the training component 735 trains the first image generation model 725 to replace the elements from the input images with generated content using the training set and the predicted images. In some examples, the training component 735 trains the teacher image generation model 740 to replace the elements from the input images. In some examples, the training component 735 computes a diffusion loss. In some examples, the training component 735 updates parameters of the first image generation model 725 based on the diffusion loss. In some examples, the training component 735 trains the second image generation model 730 to replace the elements from the input images based on the output of the first image generation model 725.

[0113] According to some aspects, the training component 735 obtains a training set, the training set comprising input images, the input images including elements. In some examples, the training component 735 obtains a teacher image generation model 740, wherein the teacher image generation model 740 is trained to delete the elements from the input images. In some examples, the training component 735 trains the first image generation model 725 to delete the elements from the input images using the training set and the teacher image generation model 740. In some examples, the training component 735 trains the teacher image generation model 740 to delete the elements from the input images.

[0114] In some examples, the training component 735 obtains a real image, the training component 735 obtaining the real image comprising deleting the elements from the input images, wherein the teacher image generation model 740 is trained based on the real image. In some examples, the training component 735 computes a diffusion loss. In some examples, the training component 735 updates parameters of the first image generation model 725 based on the diffusion loss. In some examples, the training component 735 trains the second image generation model 730 to delete the elements from the input images based on the output of the first image generation model 725.

[0115] According to some aspects, the teacher image generation model 740 is implemented as software stored in the memory unit 715 and executable by the processor unit 705, as firmware, one or more hardware circuits, or a combination thereof. According to some embodiments, the teacher image generation model 740 is implemented as software stored in the memory unit and executable by a processor in a processor unit of a separate computing device, as firmware in the separate computing device, one or more hardware circuits of the separate computing device, or a combination thereof. In some examples, the teacher image generation model 740 is part of another device other than the image processing apparatus 700 and is in communication with the image processing apparatus 700. In some examples, the teacher image generation model 740 is part of the image processing apparatus 700.

[0116] According to some aspects, the teacher image generation model 740 generates a predicted image that replaces an element from the input image with generated content. According to some aspects, the teacher image generation model 740 is trained to delete an element from an input image. The teacher image generation model 740 is a reference Figure 9 and Figure 14 described corresponding element or includes aspects of the corresponding element described Figure 9 and Figure 14 described corresponding element.

[0117] Figure 8 An example of a machine learning model according to aspects of the present disclosure is shown. The shown example includes a machine learning system 800, an input image 805, an input mask 810, a first image generation model 815, an agent result 820, an agent guidance 825, a second image generation model 830, and a synthetic image 835.

[0118] Referring to Figure 8 , the machine learning system 800 receives an input image 805 and an input mask 810 to generate a synthetic image 835. For example, a first image generation model 815 receives the input image 805 and the input mask 810 to generate an agent result 820. In one aspect, the first image generation model 815 is a student agent diffusion model that is trained to generate a low resolution output image based on a low resolution input image. In some embodiments, the first image generation model 815 is trained on an object deletion task. For example, the object deletion task involves identifying and deleting an object from a region of an image while seamlessly reconstructing missing image pixels of the deleted object to maintain a natural appearance. In some cases, the process uses an intra-painting technique to fill in the deleted region or image pixels that are typically guided by context, texture, and image elements from surrounding regions of the image.

[0119] For example, the teacher agent image generation model distills knowledge from the teacher agent image generation model (or teacher image generation model) to the student image generation model (e.g., the first image generation model 815) during training. In one aspect, the teacher agent image generation model is a complex, high-capacity image generation model trained on a large dataset. The teacher agent image generation model is trained to remove objects / elements from an input image (e.g., the input image 805) to generate an edited image without the elements. When the distillation process is performed, knowledge from the teacher agent image generation model is migrated to the student agent model (or the first image generation model 815) such that the first image generation model 815 is able to perform the same function as the teacher agent image generation model in a low-capacity manner. For example, in some cases, the teacher agent image generation model is distilled from 40 steps to 5 steps. In one aspect, the first image generation model 815 is a smaller, simpler, faster model compared to the teacher agent image generation model. In one aspect, the first image generation model 815 has fewer parameters and uses fewer computational resources. As a result, the first image generation model 815 is able to generate low-resolution output images faster and more accurately.

[0120] In some embodiments, the second image generation model 830 receives the input image 805, the input mask 810, and the agent guidance 825 to generate one or more synthetic images 835. In some cases, the second image generation model 830 is a high-capacity (or high-resolution) diffusion model. For example, the second image generation model 830 has more parameters than the first image generation model 815. In some embodiments, the first image generation model 815 has the same number of parameters as the second image generation model 830. For example, the agent result 820 is used as the agent guidance 825 to the second image generation model 830. In some embodiments, the second image generation model 830 is trained to generate synthetic images 835 that are conditioned on a reference image (e.g., the agent result 820 or the agent guidance 825). For example, the agent guidance 825 is used as guidance to guide a diffusion process in the second image generation model 830. In some cases, the synthetic images 835 depict various versions of the input image 805 without the objects indicated by the input mask 810. In some cases, the synthetic images 835 are high-quality images.

[0121] In some embodiments, the second image generation model 830 is trained for a hybrid task of image editing and image generation. For example, the second image generation model 830 can be trained for tasks such as in-painting, super-resolution, object removal, style transfer, and image-to-image translation. In some cases, the second image generation model 830 is able to efficiently perform various image editing operations while reducing the need for separate models for each specific task.

[0122] The input image 805 is a reference Figure 3-Figure 5 、 Figure 9 and Figure 14 Examples of corresponding elements described or including references Figure 3-Figure 5 、 Figure 9 and Figure 14 The input mask 810 is a reference to the corresponding elements of the description. Figure 3-Figure 5 and Figure 9 Examples of corresponding elements described or including references Figure 3-Figure 5 and Figure 9 The first image generation model 815 is a reference to the corresponding elements of Figure 7 、 Figure 9 and Figure 14 Examples of corresponding elements described or including references Figure 7 、 Figure 9 and Figure 14 Describes aspects of the corresponding element.

[0123] Agent Guidance 825 is a reference Figure 9 and Figure 14 Examples of corresponding elements described or including references Figure 9 and Figure 14 The second image generation model 830 is a reference to the corresponding elements of Figure 7 、 Figure 9 and Figure 14 Examples of corresponding elements described or including references Figure 7 、 Figure 9 and Figure 14 The composite image 835 is a reference to the corresponding elements of the description. Figure 3-Figure 5 Examples of corresponding elements described or including references Figure 3-Figure 5 Describes aspects of the corresponding element.

[0124] Figure 9 An example of an object removal system according to aspects of the present disclosure is shown. The example shown includes an image editing system 900, a low-resolution input image 905, a low-resolution input mask 910, a first image generation model 915, a teacher image generation model 920, an agent guidance 925, an input image 930, an input mask 935, a second image generation model 940, and a composite image 945.

[0125] refer to Figure 9The first image generation model 915 receives a low resolution input image 905 and a low resolution input mask 910 to generate an agent guidance 925. In some cases, the teacher image generation model 920 performs distillation that migrates knowledge to the first image generation model 915 during training time. In some cases, the first image generation model 915 is a low capacity diffusion model that can generate output (e.g., agent guidance 925) more quickly. By performing knowledge distillation, knowledge from the teacher image generation model 920 is migrated to the first image generation model 915 such that the first image generation model 915 is able to perform the same functionality as the teacher image generation model 920 in a low capacity manner. In some cases, the teacher image generation model 920 is a larger model and has more parameters compared to the first image generation model 915.

[0126] In some embodiments, the second image generation model 940 receives an input image 930 and an input mask 935 to generate a synthetic image 945. For example, the input image 930 is a high resolution image. For example, in some aspects, the second image generation model 940 is a large capacity diffusion model, where the second image generation model 940 has more parameters than the first image generation model 915. In some cases, the synthetic image 945 is a high resolution image. In some embodiments, the second image generation model 940 receives the agent guidance 925 as guidance to guide an image generation process (e.g., a diffusion process). For example, the agent guidance 925 is an image that depicts the input image 930 (or low resolution input image 905) without the objects indicated by the input mask 935 (or low resolution input mask 910). When the agent guidance 925 is used to guide the image generation process, image features that closely represent the agent guidance 925 are input into the second image generation model 940. As a result, the synthetic image 945 includes features of the agent guidance 925.

[0127] In some embodiments, the second image generation model 940 performs interior filling on the area indicated by the input mask 935 in the input image 930. For example, the second image generation model 940 generates additional pixels with content consistent with the input image 930. In some cases, the second image generation model 940 identifies the image area (e.g., the input image 930) to be filled with content. In one aspect, the area is indicated by the input mask 935. The second image generation model 940 then uses pixel information from nearby pixels of the input image 930 to interior fill the area. In some cases, the second image generation model 940 uses texture synthesis to perform interior filling. In some cases, the second image generation model 940 can perform blending between the synthetically generated pixels in the interior filled area of ​​the input image 930 and the rest of the input image 930. In some cases, the interior filled image (or the synthesized image 945) is thinned or filtered.

[0128] Image editing system 900 is a reference Figure 3-Figure 5 Examples of corresponding elements described or including references Figure 3-Figure 5 The first image generation model 915 is a reference to the corresponding elements of Figure 7 、 Figure 8 and Figure 14 Examples of corresponding elements described or including references Figure 7 、 Figure 8 and Figure 14 The teacher image generation model 920 is a reference to the corresponding elements of the description. Figure 7 and Figure 14 Examples of corresponding elements described or including references Figure 7 and Figure 14 The agent guide 925 is a reference to the corresponding elements of the description. Figure 8 and Figure 14 Examples of corresponding elements described or including references Figure 8 and Figure 14 Describes aspects of the corresponding element.

[0129] The input image 930 is a reference Figure 3-Figure 5 、 Figure 8 and Figure 14 Examples of corresponding elements described or including references Figure 3-Figure 5 、 Figure 8 and Figure 14 The input mask 935 is a reference to the corresponding element. Figure 3-Figure 5 and Figure 8 Examples of corresponding elements described or including references Figure 3-Figure 5 and Figure 8 The second image generation model 940 is a reference to the corresponding elements of Figure 7 、 Figure 8 and Figure 14 described corresponding elements or include aspects of the corresponding elements described Figure 7 , Figure 8 and Figure 14 described corresponding elements or include aspects of the corresponding elements described Figure 14 described corresponding elements or include aspects of the corresponding elements described Figure 14 described corresponding elements or include aspects of the corresponding elements described

[0130] Figure 10 An example of an image generation model is shown in accordance with aspects of the present disclosure. The shown example includes a diffusion model 1000, an original image 1005, a pixel space 1010, an image encoder 1015, an original image feature 1020, a latent space 1025, a forward diffusion process 1030, a noise feature 1035, a reverse diffusion process 1040, a denoised image feature 1045, an image decoder 1050, an output image 1055, a text prompt 1060, a text encoder 1065, a guidance feature 1070, and a guidance space 1075.

[0131] Diffusion models are a class of generative neural networks that can be trained to generate new data with similar characteristics to those found in training data. Specifically, diffusion models can be used to generate new images. Diffusion models can be used for various image generation tasks, including image super-resolution, generating images using perceptual metrics, conditional generation (e.g., generation based on text guidance, color guidance, style guidance, and image guidance), image inpainting, and image manipulation.

[0132] Types of diffusion models include denoising diffusion probabilistic models (DDPM) and denoising diffusion implicit models (DDIM). In DDPM, the generation process includes reversing a stochastic Markov diffusion process. In another aspect, DDIM uses a deterministic process such that the same input produces the same output. Diffusion models can also be characterized by whether noise is added to an image or to an image feature (e.g., latent diffusion) generated by an encoder.

[0133] Diffusion models work by iteratively adding noise to data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, the diffusion model 1000 can take an original image 1005 in a pixel space 1010 as input and apply an image encoder 1015 to convert the original image 1005 to an original image feature 1020 in a latent space 1025. A forward diffusion process 1030 then gradually adds noise to the original image feature 1020 to obtain noise features 1035 at different noise levels (also in the latent space 1025).

[0134] Next, a reverse diffusion process 1040 (e.g., a U-Net ANN) progressively removes noise from the noise features 1035 at various noise levels to obtain denoised image features 1045 in the latent space 1025. In some examples, the denoised image features 1045 are compared to the original image features 1020 at each of the various noise levels and parameters of the reverse diffusion process 1040 of the diffusion model are updated based on the comparison. Finally, an image decoder 1050 decodes the denoised image features 1045 to obtain an output image 1055 in the pixel space 1010. In some cases, the output image 1055 is created at each of the various noise levels. The output image 1055 can be compared to the original image 1005 to train the reverse diffusion process 1040. In some cases, the output image 1055 refers to a synthetic image (e.g., a reference Figure 3-Figure 5 and Figure 8-Figure 9 as described

[0135] In some cases, the image encoder 1015 and the image decoder 1050 are pre-trained prior to training the reverse diffusion process 1040. In some examples, the image encoder 1015 and the image decoder 1050 are jointly trained or the image encoder 1015 and the image decoder 1050 are jointly fine-tuned with the reverse diffusion process 1040.

[0136] The reverse diffusion process 1040 can also be guided based on a text prompt 1060 or other guiding prompts (such as images, layouts, styles, colors, segmentation maps, etc.). The text prompt 1060 can be encoded using a text encoder 1065 (e.g., a multi-modal encoder) to obtain guiding features 1070 in a guiding space 1075. The guiding features 1070 can be combined with the noise features 1035 at one or more layers of the reverse diffusion process 1040 to ensure that the output image 1055 includes content described by the text prompt 1060. For example, the guiding features 1070 can be combined with the noise features 1035 within the reverse diffusion process 1040 using cross-attention blocks.

[0137] Cross-attention, also known as multi-head attention, is an extension of the attention mechanism used in some ANNs, for example, for NLP tasks. In some cases, cross-attention attends to multiple parts of an input sequence simultaneously, thereby capturing interactions and correlations between different elements. In cross-attention, there are two input sequences: a query sequence and a key-value sequence. The query sequence represents the elements that need to be attended to, while the key-value sequence contains the elements to be attended to. In some cases, to compute cross-attention, a cross-attention block transforms (e.g., using a linear projection) each element in the query sequence into a “query” representation, while the elements in the key-value sequence are transformed into “key” and “value” representations.

[0138] The cross-attention block computes attention scores by measuring the similarity between each query representation and key representation, where a higher similarity indicates that more attention is given to the key element. The attention scores indicate the importance or relevance of each key element to the corresponding query element.

[0139] The cross-attention block then normalizes the attention scores to obtain attention weights (e.g., using a softmax function), where the attention weights determine how much information from each value element is incorporated into the final attention representation. By simultaneously focusing on different parts of the key-value sequence, the cross-attention block captures relationships and correlations across the input sequence, enabling the machine learning model to understand context and generate more accurate and contextually relevant outputs.

[0140] In some examples, the diffusion model is based on a neural network architecture known as a U-Net. The U-Net takes an input feature with an initial resolution and an initial number of channels and processes the input feature using an initial neural network layer (e.g., a convolutional network layer) to generate an intermediate feature. The intermediate feature is then downsampled using a downsample layer such that the downsampled feature has a smaller resolution than the initial resolution and a larger number of channels than the initial number of channels.

[0141] The process is repeated multiple times and then the process is reversed. For example, the downsampled feature is upsampled using an upsample process to obtain an upsampled feature. The upsampled feature can be combined with the intermediate feature having the same resolution and number of channels via a skip connection. These inputs are processed using a final neural network layer to produce an output feature. In some cases, the output feature has the same resolution as the initial resolution and the same number of channels as the initial number of channels.

[0142] In some cases, the U-Net requires additional input features to produce a conditioned generated output. For example, the additional input features can include a vector representation of an input prompt. The additional input features can be combined with the intermediate features at one or more layers within the neural network. For example, a cross-attention module can be used to combine the additional input features and the intermediate features. Further details regarding the U-Net are described with reference to Figure 11

[0143] ​The diffusion process can also be modified based on a conditional guidance. In some cases, a user provides a textual prompt (e.g., textual prompt 1060) that describes content to include in the generated image. In some examples, the guidance can be provided in a form other than text, such as via an image, sketch, color, style, or layout. The system converts the textual prompt 1060 (or other guidance) into a conditional guidance vector or other multi-dimensional representation. For example, the text can be converted into a vector or series of vectors using a transformer model or multi-modal encoder. In some cases, the encoder of the conditional guidance is trained independently of the diffusion model.

[0144] A noise map including random noise is initialized. The noise map can be in pixel space or latent space. By initializing the image using random noise, different variations of the image including the content described by the conditional guidance can be generated. The diffusion model 1000 then generates an image based on the noise map and the conditional guidance vector.

[0145] The diffusion process can include a forward diffusion process 1030 for adding noise to an image (e.g., original image 1005) or feature (e.g., original image feature 1020) in latent space 1025 and a reverse diffusion process 1040 for denoising the image (or feature) to obtain a denoised image (e.g., output image 1055). The forward diffusion process 1030 can be represented as q(x t | x t-1 ), and the reverse diffusion process 1040 can be represented as p θ (x t-1 | x t ). Additional details regarding the diffusion process are described with reference to Figure 12

[0146] The diffusion model 1000 can be trained using both the forward diffusion process 1030 and the reverse diffusion process 1040. In one example, a user initializes an untrained model. The initialization can include defining the architecture of the model and establishing initial values for the model parameters. In some cases, the initialization can include defining hyperparameters, such as the number of layers, the resolution and channels of each layer block, the location of skip connections, etc.

[0147] The system then uses an N-stage forward diffusion process 1030 to add noise to a training image. In some cases, the forward diffusion process 1030 is a fixed process in which Gaussian noise is added to the image successively. In latent diffusion models, the Gaussian noise can be added to a feature (e.g., original image feature 1020) in latent space 1025 successively.

[0148] ​At each stage n, starting from stage N, a reverse diffusion process 1040 is used to predict the image or image features at stage n-1. For example, the reverse diffusion process 1040 can predict the noise added by the forward diffusion process 1030, and the predicted noise can be removed from the image to obtain a predicted image. In some cases, the original image 1005 is predicted at each stage of the training process.

[0149] The training component (e.g., the reference Figure 7 described training component) compares the predicted image (or image features) at stage n-1 to the actual image (or image features), such as the image at stage n-1 or the original input image. For example, given the observation data x, the diffusion model 1000 can be trained to minimize the variational upper bound on the negative log-likelihood value -logp θ (x) of the training data. The training component then updates the parameters of the diffusion model 1000 based on the comparison. For example, the parameters of the U-Net can be updated using gradient descent. The time-dependent parameters of the Gaussian transform can also be learned. Further details related to training the diffusion model are described with reference to Figure 16 .

[0150] The original image 1005 is an example of, or includes aspects of, the corresponding element described with reference to Figure 12 . The forward diffusion process 1030 is an example of, or includes aspects of, the corresponding element described with reference to Figure 12 . The reverse diffusion process 1040 is an example of, or includes aspects of, the corresponding element described with reference to Figure 12 . The forward diffusion process 1030 is an example of, or includes aspects of, the corresponding element described with reference to Figure 12 . The reverse diffusion process 1040 is an example of, or includes aspects of, the corresponding element described with reference to Figure 12 . The reverse diffusion process 1040 is an example of, or includes aspects of, the corresponding element described with reference to Figure 12 . The reverse diffusion process 1040 is an example of, or includes aspects of, the corresponding element described with reference to

[0151] Figure 11 An example of a U-Net 1100 architecture is shown in accordance with aspects of the present disclosure. The example shown includes a U-Net 1100, input features 1105, initial neural network layers 1110, intermediate features 1115, down-sampling layers 1120, down-sampled features 1125, up-sampling process 1130, up-sampled features 1135, skip connections 1140, final neural network layers 1145, and output features 1150.

[0152] In some examples, the U-Net 1100 is an example of, and includes aspects of, the components of the reverse diffusion process 1040 of the diffusion model 1000 described with reference to Figure 10 . The U-Net 1100 is an example of, and includes aspects of, the architecture elements of the image generation model (e.g., the first image generation model 725 and the second image generation model 730) described with reference to Figure 7 . The U-Net 1100 is an example of, and includes aspects of, the architecture elements of the image generation model (e.g., the first image generation model 725 and the second image generation model 730) described with reference to Figure 11The depicted U-Net 1100 is a reference Figure 10 Examples of architectures used in the described reverse diffusion process or including Figure 10 Aspects of the described reverse diffusion process.

[0153] In some examples, the diffusion model is based on a neural network architecture known as a U-Net. The U-Net 1100 takes an input feature 1105 having an initial resolution and an initial number of channels and processes the input feature 1105 using an initial neural network layer 1110 (e.g., a convolutional network layer) to produce an intermediate feature 1115. The intermediate feature 1115 is then downsampled using a downsample layer 1120 such that the downsampled feature 1125 has a smaller resolution than the initial resolution and a larger number of channels than the initial number of channels.

[0154] The process is repeated multiple times and then the process is reversed. For example, the downsampled feature 1125 is upsampled using an upsample process 1130 to obtain an upsampled feature 1135. The upsampled feature 1135 can be combined with the intermediate feature 1115 having the same resolution and number of channels via a skip connection 1140. These inputs are processed using a final neural network layer 1145 to produce an output feature 1150. In some cases, the output feature 1150 has the same resolution as the initial resolution and has the same number of channels as the initial number of channels.

[0155] In some cases, the U-Net 1100 takes additional input features to produce a conditionally generated output. For example, the additional input features can include a vector representation of an input prompt. The additional input features can be combined with the intermediate feature 1115 at one or more layers within the neural network. For example, a cross-attention module can be used to combine the additional input features and the intermediate feature 1115.

[0156] Diffusion process

[0157] Figure 12 An example of a diffusion process 1200 is shown in accordance with aspects of the present disclosure. The shown example includes the diffusion process 1200, a forward diffusion process 1205, a reverse diffusion process 1210, a noise image 1215, a first intermediate image 1220, a second intermediate image 1225, and an original image 1230.

[0158] The diffusion process 1200 can include the forward diffusion process 1205, which is used to diffuse an original image 1230 (e.g., reference Figure 10 the described original image 1005) or a feature (e.g., reference Figure 10In some aspects, the diffusion process 1200 includes a back diffusion process 1210 for denoising the noisy image 1215 (or image features) to obtain a denoised image (or original image 1230). The forward diffusion process 1205 can be represented as q(x t ∣x t-1 ) and the back diffusion process 1210 can be represented as p θ (x t-1 ∣x t In some cases, the forward diffusion process 1205 is used during training to generate images with continuously increasing noise, and the neural network is trained to perform the backward diffusion process 1210 (eg, to continuously remove the noise).

[0159] In the context of potential diffusion models (e.g., ref. Figure 10 In the forward diffusion process 1205 of the diffusion model 1000 described above, the diffusion model uses a Markov chain to map the observed variable x0 (in pixel space or latent space) to obtain intermediate variables x1, ..., x T When the latent variable is passed through a neural network such as U-Net, the Markov chain gradually adds Gaussian noise to the data to obtain an approximate posterior q(x 1:T |x0), where x1,…,x T has the same dimensions as x0.

[0160] A neural network can be trained to perform the back diffusion process 1210. During the back diffusion process 1210, the diffusion model utilizes the noise data x T (such as the noisy image 1215), and denoise the data to obtain p μ (x t-1 ∣x t ). At each step t-1, the back diffusion process 1210 uses x t (such as the first intermediate image 1220) and t as input. Here, t represents a step in the transformation sequence associated with different noise levels. The back diffusion process 1210 iteratively outputs x t-1 , such as the second intermediate image 1225, until x T Restore to x0, that is, the original image 1230. The back diffusion process 1210 can be expressed as:

[0161] p μ (x t-1 ∣x t ):=N(x t-1 ;μ θ (x t ,t),∑ θ (xt ,t)),(1)

[0162] The joint probability of a sequence of samples in a Markov chain can be written as the product of the conditional probability and the marginal probability:

[0163]

[0164] where p(x T )=N(x T 0, I) is the pure noise distribution when the reverse diffusion process 1210 uses the result of the forward diffusion process 1205 (pure noise sample) as input and Represents the Gaussian transformed sequence corresponding to a sequence with Gaussian noise added to the samples.

[0165] At inference time, the observation data x0 in pixel space can be mapped into the latent space as input and the generated data It can be mapped back from the latent space to the pixel space as output. In some examples, x0 represents the original input image with low image quality, and the latent variables x1,…,x T represents a noisy image and Indicates the generated image has high image quality.

[0166] Forward diffusion process 1205 is reference Figure 10 Examples of corresponding elements described or including references Figure 10 The reverse diffusion process 1210 is a reference to the corresponding elements of the description. Figure 10 Examples of corresponding elements described or including references Figure 10 The original image 1230 is a reference to the corresponding elements of the Figure 10 Examples of corresponding elements described or including references Figure 10 Describes aspects of the corresponding element.

[0167] Training and evaluation

[0168] exist Figure 13-16 In the present invention, a method, apparatus, non-transitory computer-readable medium, and system for training a machine learning model include obtaining a training set, the training set including an input image, the input image including elements; using a teacher image generation model to generate a predicted image, the predicted image using generated content to replace elements from the input image; and using the training set and the predicted image to train a first image generation model to replace elements from the input image with the generated content.

[0169] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include training the teacher image generation model to replace an element from the input image. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include training the second image generation model to replace an element from the input image based on the output of the first image generation model.

[0170] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include computing a diffusion loss. Some examples further include updating the parameters of the first image generation model based on the diffusion loss. In some aspects, the first image generation model has fewer parameters than the teacher image generation model.

[0171] According to some aspects, some examples of the method, apparatus, non-transitory computer-readable medium, and system further include training the teacher image generation model to delete an element from the input image. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include obtaining a real image, obtaining the real image including deleting the element from the input image, wherein the teacher image generation model is trained based on the real image.

[0172] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include computing a diffusion loss. Some examples further include updating the parameters of the first image generation model based on the diffusion loss. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include training the second image generation model to delete an element from the input image based on the output of the first image generation model. In some aspects, the first image generation model has fewer parameters than the teacher image generation model.

[0173] Figure 13 An example of a method 1300 for training a machine learning model is shown in accordance with aspects of the present disclosure. In some examples, the operations are performed by a system including a processor executing a set of codes to control functional elements of a device. Additionally or alternatively, some processes are performed using special-purpose hardware. In general, the operations are performed in accordance with the methods and processes described with respect to aspects of the present disclosure. In some cases, the operations described herein are performed with various sub-steps or in conjunction with other operations.

[0174] At operation 1305, the system obtains a training set including an input image including an element. In some cases, the operations of this step are performed with reference to or can be performed by a training component as described with reference to Figure 7 In some cases, the training set is stored in a database as described with reference to Figure 1

[0175] ​At operation 1310, the system generates a predicted image using the teacher image generation model, the predicted image replacing elements from the input image with generated content. In some cases, the operations of this step are performed with reference to or can be performed by a teacher image generation model as described with reference to Figure 7 、 Figure 9 and Figure 14 In some cases, the teacher image generation model is a complex, high-capacity image generation model trained on a large dataset. The teacher image generation model is trained to delete one or more elements from an input image to generate a predicted image without the elements. In some cases, the teacher image generation model is trained to internally fill in a deletion region using pixels that are consistent with pixels surrounding the deletion region.

[0176] At operation 1315, the system trains the first image generation model to delete elements from input images using the training set and the teacher image generation model. In some cases, the operations of this step are performed with reference to or can be performed by a training component as described with reference to Figure 7 In some cases, a diffusion loss is computed based on the predicted image and the real image. In some cases, the parameters of the first image generation model are updated based on the diffusion loss. Further details related to training the first image generation model are described with reference to Figure 14 .

[0177] Figure 14 An example of training a machine learning model is shown in accordance with aspects of the present disclosure. The shown example includes a training system 1400, a low resolution input 1405, a first image generation model 1410, an agent guide 1415, a teacher image generation model 1420, an input image 1425, a second image generation model 1430, and a synthetic image 1435.

[0178] With reference to Figure 14 The first image generation model 1410 receives a low resolution input 1405 (e.g., a low resolution image) to generate a low resolution output image (e.g., the agent guide 1415). The teacher image generation model 1420 distills knowledge from the teacher image generation model 1420 to the first image generation model 1410. In some cases, the first image generation model 1410 is a smaller model (and has fewer parameters) compared to the teacher image generation model 1420. In some cases, a diffusion loss is computed based on the low resolution output image and a real low resolution image to fine-tune the first image generation model 1410.

[0179] In some embodiments, the second image generation model 1430 receives the input image 1425 and the agent guidance 1415 to generate a synthetic image 1435. In some cases, the input image 1425 is a high resolution image. In some cases, the synthetic image 1435 is a high resolution image that depicts the input image 1425 without the objects indicated, for example, by the input mask. In one aspect, the second image generation model 1430 has more parameters than the first image generation model 1410. For example, a diffusion loss is computed based on the synthetic image 1435 and a real image that depicts the deleted objects, and the diffusion loss is used to fine-tune the second image generation model 1430.

[0180] The first image generation model 1410 is an example of, or includes reference to, the corresponding element described in Figure 7-Figure 9 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in Figure 7-Figure 9 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in Figure 8 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in Figure 9 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in Figure 8 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in Figure 9 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in Figure 7 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in Figure 9 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in Figure 7 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in Figure 9 The agent guidance 1415 is an example of, or includes reference to, the corresponding element described in

[0181] The input image 1425 is an example of, or includes reference to, the corresponding element described in Figure 3-5 Figure 8 The input image 1425 is an example of, or includes reference to, the corresponding element described in Figure 9 The input image 1425 is an example of, or includes reference to, the corresponding element described in Figure 3-5 The input image 1425 is an example of, or includes reference to, the corresponding element described in Figure 8 The input image 1425 is an example of, or includes reference to, the corresponding element described in Figure 9 The second image generation model 1430 is an example of, or includes reference to, the corresponding element described in Figure 7-Figure 9 The second image generation model 1430 is an example of, or includes reference to, the corresponding element described in Figure 7-Figure 9 The second image generation model 1430 is an example of, or includes reference to, the corresponding element described in Figure 9 The second image generation model 1430 is an example of, or includes reference to, the corresponding element described in Figure 9 The second image generation model 1430 is an example of, or includes reference to, the corresponding element described in

[0182] Figure 15 An example illustrates a flow diagram that depicts a step-by-step procedure as an example operational implementation of an algorithm executable for training a machine learning model, in accordance with aspects of the present disclosure. In some embodiments, the procedure 1500 describes operations of the training component 735 described for configuring a machine learning model as reference Figure 7 ​The first image generation model 725 and / or the second image generation model 730 are described. The program 1500 provides one or more examples of generating training data, training a machine learning model using the training data, and using the trained machine learning model to perform a task.

[0183] To begin in this example, the machine learning system collects training data (block 1502), which is used as a basis for training the machine learning model, i.e., it defines what is being modeled. The training data can be collected by the machine learning system from a variety of sources. Examples of training data sources include public datasets, service provider system platforms that expose application programming interfaces (e.g., social media platforms), user data collection systems (e.g., digital surveys and online crowdsourcing systems), etc. Training data collection can also include data augmentation and synthetic data generation techniques for expanding and diversifying available training data, balancing techniques for balancing the number of positive and negative examples, etc.

[0184] The machine learning system can also be configured to identify features related to the type of task for the machine learning model to be trained (block 1504). Examples of tasks include classification, natural language processing, generative artificial intelligence, recommendation engines, reinforcement learning, clustering, etc. To this end, the machine learning system collects training data based on the identified features and / or filters the training data after collection based on the identified features. The training data is then used to train the machine learning model.

[0185] To train the machine learning model in the illustrated example, the machine learning model is first initialized (block 1506). The initialization of the machine learning model includes selecting a model architecture to be trained (block 1508). Examples of model architectures include neural networks, convolutional neural networks (CNNs), long short-term memory (LSTM) neural networks, generative adversarial networks (GANs), decision trees, support vector machines, linear regression, logistic regression, Bayesian networks, random forest learning, dimensionality reduction algorithms, boosting algorithms, deep learning neural networks, U-Net architectures, etc.

[0186] A loss function is also selected (block 1510). The loss function is used to measure the difference between the output of the machine learning model (e.g., model predictions) and the target values (e.g., represented by the training data) used to train the machine learning model. Additionally, an optimization algorithm is selected (1512), which is used in conjunction with the loss function to optimize the parameters of the machine learning model during training, examples of optimization algorithms include gradient descent, stochastic gradient descent (SGD), etc.

[0187] Initialization of the machine learning model also includes setting initial values for the machine learning model (block 1516), examples of initialization include initializing weights and biases of nodes to increase training efficiency and computational resources consumed as part of training. Hyperparameters for controlling training of the machine learning model are also set (block 1514), examples of hyperparameters include regularization parameters, model parameters (e.g., number of layers in a neural network), learning rate, batch size selected from training data, etc. Hyperparameters are set using various techniques, including using randomization techniques, by using heuristics learned from other training scenarios, etc.

[0188] The machine learning model is then trained by the machine learning system using the training data (block 1518). A machine learning model refers to a computer representation that can be tuned (e.g., trained and retrained) to approximate an unknown function based on inputs of the training data. In particular, the term machine learning model can include a model that learns and relearns by analyzing training data to generate outputs that reflect patterns and properties expressed by the training data, using algorithms (e.g., using model architectures described above), to learn and predict from known data.

[0189] Examples of training types include supervised learning that employs labeled data, unsupervised learning that involves finding underlying structure or patterns in training data, reinforcement learning that is based on an optimization function (e.g., rewards and / or penalties), using nodes as part of “deep learning,” etc. The machine learning model may, for example, be configured to include a plurality of nodes that collectively form a plurality of layers. The layers may, for example, be configured to include an input layer, an output layer, and one or more hidden layers. The computations are performed by the nodes within the layers by way of hidden states, through a system of weighted connections that are “learned” during training, e.g., by using a selected loss function and backpropagation to optimize performance of the machine learning model to perform an associated task.

[0190] As part of training the machine learning model, it is determined whether a stopping criterion is satisfied (decision block 1520), i.e., the stopping criterion is used to validate the machine learning model. The stopping criterion can be used to reduce overfitting of the machine learning model, reduce computational resource consumption, and improve the ability of the machine learning model to handle unseen data (data not included as an example in the training data). Examples of stopping criteria include, but are not limited to, a predefined number of epochs, validation loss stability, achieving a performance improvement threshold, whether a threshold level of accuracy is satisfied, or based on performance metrics such as precision and recall. If the stopping criterion is not satisfied (“No” from decision block 1520), the program 1500 continues to train the machine learning model using the training data in the example (block 1518).

[0191] If the stopping criterion is satisfied ("Yes" from decision block 1520), the trained machine learning model is then used to generate an output based on subsequent data (block 1522). The trained machine learning model, for example, is trained to perform the task described above and is thus configured to perform that task based on subsequent data received as input and processed by the machine learning model once trained.

[0192] Figure 16 An example of a method 1600 for training a diffusion model is shown in accordance with aspects of the present disclosure. In some examples, the operations are performed by a system including a processor that executes a set of codes to control the functional elements of a device. Additionally or alternatively, some processes are performed using specialized hardware. Generally, these operations are performed in accordance with the methods and processes described in accordance with aspects of the present disclosure. In some cases, the operations described herein are performed by various sub-steps or in conjunction with other operations.

[0193] In some embodiments, the method 1600 describes operations for training a training component 735 of a machine learning model 720 as described with reference to Figure 7 The method 1600 represents an example for training a reverse diffusion process as described above with reference to Figure 12 In some examples, the operations are performed by a system including a processor that executes a set of codes to control the functional elements of a device, such as the first and second image generation models described in Figure 7

[0194] At operation 1605, the system initializes an untrained model. In some cases, the operations of this step are performed with reference to or can be performed by a training component as described with reference to Figure 7 The initialization can include defining the architecture of the model and establishing initial values for the model parameters. In some cases, the initialization can include defining hyperparameters, such as the number of layers, the resolution and channels of each layer block, the location of skip connections, and the like.

[0195] At operation 1610, the system uses an N-stage forward diffusion process to add noise to a media item. In some cases, the operations of this step are performed with reference to or can be performed by a training component as described with reference to Figure 7 For example, in some cases, the media item is a training image. In some cases, the forward diffusion process is a fixed process in which Gaussian noise is continuously added to the media item, such as an original image. In latent diffusion models, Gaussian noise can be continuously added to features in the latent space.

[0196] ​At operation 1615, the system predicts, at each stage n, starting from stage N, a media item for stage n-1. In some cases, the operations of this step are performed with reference to or can be performed by a training component as described with reference to Figure 7 In some cases, the media item is a synthetic image generated using an image generation model. For example, a reverse diffusion process can predict the noise added by a forward diffusion process and the predicted noise can be removed from a noise input to obtain a predicted output. In some cases, the original media item is predicted at each stage of the training process.

[0197] At operation 1620, the system compares the predicted media item (or feature) at stage n-1 to the media item at stage n-1. For example, in some cases, the system compares a synthetic image (or predicted image feature) at stage n-1 to a real image (or real feature) at stage n-1. In some cases, the operations of this step are performed with reference to or can be performed by a training component as described with reference to Figure 7 For example, given observation data x, a diffusion model can be trained to minimize the variational upper bound on the negative log-likelihood value -logp μ (x) of the training data.

[0198] At operation 1625, the system updates the parameters of the model based on the comparison. In some cases, the operations of this step are performed with reference to or can be performed by a training component as described with reference to Figure 7 For example, the parameters of a U-Net can be updated using gradient descent. The time-dependent parameters of a Gaussian transform can also be learned.

[0199] Computing device

[0200] Figure 17 An example of a computing device 1700 is shown in accordance with aspects of the present disclosure. The illustrated example includes a computing device 1700, a processor 1705, a memory subsystem 1710, a communication interface 1715, an I / O interface 1720, a user interface component 1725, and a channel 1730.

[0201] In some embodiments, the computing device 1700 is an example of or includes aspects of the image processing apparatus described with reference to Figure 1 and Figure 7 In some embodiments, the computing device 1700 is an example of or includes aspects of the image processing apparatus described with reference to Figure 1 and Figure 7 In some embodiments, the computing device 1700 includes a processor 1705 that can execute instructions stored in the memory subsystem 1710 to take an input image and an input mask, generate an intermediate result based on the input image and the input mask, and generate a synthetic image based on the input image and the intermediate result.

[0202] According to some embodiments, the processor 1705 includes one or more processors. In some cases, the processor 1705 is a smart hardware device, e.g., a general- purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, a specialized application integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, discrete gate or transistor logic components, discrete hardware components, or a combination thereof. In some cases, the processor 1705 is configured to operate a memory array using a memory controller. In other cases, a memory controller is integrated into the processor 1705. In some cases, the processor 1705 is configured to execute computer-readable instructions stored in a memory to perform various functions. In some embodiments, the processor 1705 includes specialized components to perform modem processing, baseband processing, digital signal processing, or transmit processing. The processor 1705 is an example of a processor unit described with reference to Figure 7 described processor units or includes aspects of the processor units described with reference to Figure 7 described processor units or includes aspects of the processor units described with reference to

[0203] According to some embodiments, the memory subsystem 1710 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read only memory (ROM), or hard drives. Examples of memory devices include solid state memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause the processor to perform various functions described herein. In some cases, the memory contains a basic input / output system (BIOS), etc., that Figure 7 described memory units or includes aspects of the memory units described with reference to Figure 7 described memory units or includes aspects of the memory units described with reference to

[0204] According to some embodiments, the communication interface 1715 operates at the boundary between a communication entity, such as the computing device 1700, one or more user devices, a cloud, and one or more databases, and a channel 1730 and can record and process communications. In some cases, the communication interface 1715 is provided for enabling the processing system coupled with a transceiver (e.g., a transmitter and / or a receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for the communication device via an antenna. In some cases, a bus is used in the communication interface 1715.

[0205] According to some embodiments, I / O interface 1720 is controlled by an I / O controller to manage input and output signals for computing device 1700. In some cases, I / O interface 1720 manages peripherals that are not integrated into computing device 1700. In some cases, I / O interface 1720 represents a physical connection or port to external peripherals. In some cases, the I / O controller utilizes an operating system such as iOS®, ANDROID®, PHOTOS®, or other known operating systems. In some cases, the I / O controller represents a modem, a keyboard, a mouse, a touchscreen, or similar device or interacts with a modem, a keyboard, a mouse, a touchscreen, or similar device. In some cases, the I / O controller is implemented as a component of the processor. In some cases, a user interacts with the device via I / O interface 1720 or a hardware component controlled by the I / O controller. I / O interface 1720 is an example of an I / O module as described in reference to In some cases, I / O interface 1720 is an example of an I / O module as described in reference to Figure 7 In some cases, I / O interface 1720 includes aspects of an I / O module as described in reference to Figure 7

[0206] According to some embodiments, user interface component 1725 enables user interaction with computing device 1700. In some cases, user interface component 1725 includes an audio device such as an external speaker system, an external display device such as a display screen, an input device (e.g., a remote control device that interfaces with the user interface directly or via an I / O controller), or a combination thereof.

[0207] The performance of the apparatuses, systems, and methods of the present disclosure have been evaluated and the results indicate that embodiments of the present disclosure have achieved higher performance than conventional techniques (e.g., conventional image generation models). Example experiments demonstrate that image processing apparatuses based on embodiments of the present disclosure outperform conventional image generation models. Details related to example use cases based on embodiments of the present disclosure are described in reference to Figure 3-Figure 5

[0208] The descriptions and drawings herein illustrate example configurations and do not represent all implementations within the scope of the claims. For example, operations and steps can be rearranged, combined, or otherwise modified. Additionally, structures and devices can be represented in block diagram form in order to indicate the relationship between components and to avoid obscuring the described concepts. Similar components or features can have similar names but can have different reference numerals corresponding to different figures.

[0209] ​​Some modifications of the present disclosure will be apparent to those skilled in the art and the principles defined herein may be applied to other variations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0210] The methods described may be implemented or performed by a device including a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general purpose processor may be a microprocessor, a conventional processor, a controller, a microcontroller, or a state machine. A processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration). Thus, the functions described herein may be implemented in hardware or software and may be performed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored on a computer-readable medium in the form of instructions or code.

[0211] Computer-readable media include non-transitory computer storage media and communication media, including any media that facilitate the transmission of code or data. Non-transitory storage media can be any available media that can be accessed by a computer. For example, non-transitory computer-readable media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact discs (CDs) or other optical disc storage devices, magnetic disk storage devices, or any other non-transitory media for carrying or storing data or code.

[0212] Additionally, a connecting component may be properly referred to as a computer-readable medium. For example, if code or data is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, or microwave signals), the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology is included in the definition of medium. Combinations of media are also included within the scope of computer-readable media.

[0213] In this disclosure and the accompanying claims, the word "or" is meant to indicate a inclusive list such that, for example, the list of X, Y, or Z indicates that X or Y or Z or XY or XZ or YZ or XYZ. Furthermore, the phrase "based on" is not used to denote closed set of conditions. For example, a step described as "based on condition A" can be based on both condition A and condition B. In other words, the phrase "based on" should be interpreted as "based at least in part on." Also, the word "a" or "an" indicates "at least one."

Claims

1. A method comprising: obtaining an input image and an input mask, wherein the input mask indicates a region of the input image to be modified; generating, using a first image generation model, an intermediate result based on the input image and the input mask, wherein the intermediate result modifies the region of the input image indicated by the input mask; and generating, using a second image generation model, a composite image based on the input image and the intermediate result, wherein the composite image depicts the input image with a higher level of detail from content of the modified region compared to the intermediate result.

2. The method of claim 1, wherein obtaining the input mask comprises: segmenting the input image to identify elements of the input image.

3. The method of claim 1, wherein obtaining the input mask comprises: receiving a position input; and generating the input mask based on the position input.

4. The method of claim 1, further comprising: receiving a deletion cue, wherein the deletion cue comprises a command to delete an element from the input image; and selecting a deletion mode based on the deletion cue, wherein the intermediate result is based on the deletion mode.

5. The method of claim 1, wherein: the intermediate result comprises an intermediate image having a lower resolution than the composite image.

6. The method of claim 1, wherein: the input mask indicates an element of the input image and the composite image deletes the element from the input image.

7. The method of claim 1, wherein generating the intermediate result comprises: obtaining a first noise input; and denoising the first noise input to obtain the intermediate result.

8. The method of claim 1, wherein generating the composite image comprises: obtaining a second noise input; and denoising the second noise input to generate the composite image.

9. The method of claim 1, wherein generating the composite image comprises: internally filling the region indicated by the input mask using content consistent with the input image.

10. The method of claim 1, wherein: the first image generation model is trained to delete an image element using a predicted image generated by a teacher image generation model.

11. The method of claim 10, wherein: the second image generation model is trained to replace the element from the input image based on an output of the first image generation model.

12. A non-transitory computer-readable medium storing code for image processing, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: obtaining an input image and an input mask, wherein the input mask indicates an element of the input image to be deleted; generating, using a first image generation model, an intermediate result based on the input image and the input mask, wherein the intermediate result removes the element in the input image indicated by the input mask; and generating, using a second image generation model, a composite image based on the input image and the intermediate result, wherein the composite image depicts the input image without the element removed in the intermediate result.

13. The non-transitory computer-readable medium of claim 12, the operations further comprising: receiving a deletion prompt, wherein the deletion prompt comprises a command to delete an element from the input image; and selecting a deletion mode based on the deletion prompt, wherein the intermediate result is based on the deletion mode.

14. The non-transitory computer-readable medium of claim 12, wherein: the first image generation model is trained to remove the element using a predicted image generated by a teacher image generation model.

15. The non-transitory computer-readable medium of claim 12, wherein the first image generation model is trained by computing a diffusion loss and updating parameters of the first image generation model based on the diffusion loss.

16. The non-transitory computer-readable medium of claim 12, wherein the second image generation model is trained to replace the element from the input image based on an output of the first image generation model.

17. A system comprising: a memory component; a processing device coupled to the memory component, the processing device configured to perform operations comprising: obtaining an input image and an input mask, wherein the input image depicts an element and the input mask indicates a region of the input image where the element is located; generating, using a first image generation model, an intermediate result based on the input image and the input mask, wherein the intermediate result includes first generated content that replaces the element within the region indicated by the input mask; and generating, using a second image generation model, a composite image based on the input image and the intermediate result, wherein the composite image includes second generated content that replaces the element within the region indicated by the input mask.

18. The system of claim 17, further comprising: a teacher image generation model, wherein the teacher image generation model is trained to delete the element from the input image.

19. The system of claim 17, wherein: the first image generation model and the second image generation model are diffusion models.

20. The system of claim 17, wherein: the first image generation model has fewer parameters than the second image generation model.

Citation Information

Cited By

  • Method and system for removing object in image editing

    CN120599096A

  • Methods and systems for removing objects in image editing

    CN120599096B