Prompt-driven image editing using machine learning

The diffusion model-based method addresses the challenge of accurately transforming images by preserving the face and enhancing detailed features, resulting in improved image generation quality.

JP7795043B2Active Publication Date: 2026-01-06GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025500058
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-05-09
Filing Date
2024-05-09
Publication Date
2026-01-06
Estimated Expiration
2044-05-09

AI Technical Summary

Technical Problem

Generative AI struggles to accurately capture fine details in images, particularly when people are involved, leading to inadequate representation of features like fingers, eyes, and mouths.

Method used

A computer-implemented method using a diffusion model to modify images by generating a storage mask for the face, performing text conditioning and forward diffusion, and blending the denoised initial and transformed images to preserve the face while satisfying text requests, with optional convolutional neural networks for de-diffusion and self-attention maps.

Benefits of technology

The method effectively maintains the structure of the initial image while modifying other elements, ensuring accurate and detailed image transformations without altering the face, thus improving the quality of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007795043000001
    Figure 0007795043000001
  • Figure 0007795043000002
    Figure 0007795043000002
  • Figure 0007795043000003
    Figure 0007795043000003
Patent Text Reader

Abstract

The media application receives an initial image and a text request to modify the initial image, the initial image including a subject with a face. The media application generates a storage mask from the initial image corresponding to the subject's face. The media application provides the text request, the initial image, and the storage mask as inputs to a diffusion model. The diffusion model outputs a denoised initial image based on the initial image, performs text conditioning and forward diffusion of the text request to generate a noisy transformed image that satisfies the text request, and outputs the denoised transformed image based on the noisy transformed image, the extracted features, and the self-attention map. The media application blends the denoised initial image, the storage mask, and the denoised transformed image to form an output image, the storage mask preventing modifications to the face from the initial image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 465,226, entitled "Prompt-Driven Image Editing Using Machine Learning," filed May 9, 2023, which is incorporated herein in its entirety. [Background technology]

[0002] Generative artificial intelligence (AI) is sometimes used to generate images from text prompts. For example, a user can request an image of an avocado chair, which is then created by the generative AI. The results are often problematic, especially when the image contains people, as more detailed aspects may be inadequately represented. For example, generative AI is still developing when it comes to capturing the fine details of features such as fingers, eyes, and mouths.

[0003] The background art provided herein is described for purposes of generally presenting the context for the present disclosure. The work of the inventors named herein is not admitted expressly or impliedly as prior art to the present disclosure to the extent described in this background art section, including aspects of the description that may not qualify as prior art at the time of filing. Summary of the Invention

[0004] A computer-implemented method includes receiving an initial image and a text request to modify the initial image, where the initial image includes a subject with a face. The method further includes generating a storage mask corresponding to the subject's face. The method further includes providing the text request, the initial image, and the storage mask as inputs to a diffusion model. The method further includes outputting a denoised initial image based on the initial image using the diffusion model. The method further includes performing text conditioning and forward diffusion of the text request using the diffusion model to generate a noisy transformed image that satisfies the text request. The method further includes outputting a denoised transformed image based on the noisy transformed image, the extracted features, and the self-attention map using the diffusion model. The method further includes blending the denoised initial image, the storage mask, and the denoised transformed image to form an output image, where the storage mask prevents modifications to the face from the initial image.

[0005] In some embodiments, outputting the denoised initial image includes performing de-diffusion of the initial image using a diffusion model to generate a noisy initial image based on the initial image, providing the noisy initial image to a first convolutional neural network (CNN) and outputting the de-noised initial image, and outputting the de-noised transformed image includes providing the noisy transformed image to a second CNN, injecting the extracted features and self-attention map into the diffusion, and outputting the de-noised transformed image. In some embodiments, the de-diffusion is a de-noised diffusion implicit model (DDIM) inversion.

[0006] In some embodiments, the method further includes receiving a selection of a first object in the initial image, and the text request includes a comment to replace the first object in the initial image with a second object. In some embodiments, the method further includes identifying, from the initial image, an area in the background to replace or modify and providing a suggestion to replace or modify the background, and the text request is associated with the suggestion. In some embodiments, the method further includes identifying, from the initial image, an object in the background to remove and providing a suggestion to remove the object from the background. In some embodiments, the method further includes identifying, from the initial image, one or more objects to replace and providing a suggestion to replace the object.

[0007] In some embodiments, the text request is to change the background of the initial image, and the saved mask further includes one or more parts of the subject in addition to the subject's face. In some embodiments, the text request further includes at least one selection from the group of a global preset, a menu of options, a library of pre-made prompts, and combinations thereof.

[0008] A non-transitory computer-readable medium includes stored instructions that, when executed by one or more processors, cause the one or more processors to perform operations including receiving an initial image and a text request to modify the initial image, the initial image including a subject with a face, generating a storage mask corresponding to the subject's face, providing the text request, the initial image, and the storage mask as inputs to a diffusion model, outputting a denoised initial image based on the initial image using the diffusion model, performing text conditioning and forward diffusion of the text request using the diffusion model to generate a noisy transformed image that satisfies the text request, using the diffusion model to output the denoised transformed image based on the noisy transformed image, the extracted features, and the self-attention map, and blending the denoised initial image, the storage mask, and the denoised transformed image to form an output image, the storage mask preventing modifications to the face from the initial image.

[0009] In some embodiments, outputting the denoised initial image includes performing de-diffusion of the initial image using a diffusion model to generate a noisy initial image based on the initial image, providing the noisy initial image to a first convolutional neural network (CNN) and outputting the denoised initial image, and outputting the denoised transformed image includes providing the noisy transformed image to a second CNN, injecting the extracted features and self-attention map into the diffusion, and outputting the denoised transformed image. In some embodiments, the de-diffusion is DDIM inversion. In some embodiments, the operations further include receiving a selection of a first object in the initial image, and the text request includes a comment for replacing the first object in the initial image with a second object. In some embodiments, the operations further include identifying an area in the background to replace or modify from the initial image, and providing a suggestion to replace or modify the background, and the text request is associated with the suggestion. In some embodiments, the operations further include identifying an area in the background to replace or modify from the initial image and providing a suggestion to replace or modify the background, wherein a text request is associated with the suggestion.

[0010] The system includes a processor and a memory coupled to the processor having stored thereon instructions that, when executed by the processor, cause the processor to perform operations including receiving an initial image and a text request to modify the initial image, the initial image including a subject with a face, generating a storage mask corresponding to the subject's face, providing the text request, the initial image, and the storage mask as inputs to a diffusion model, outputting a denoised initial image based on the initial image using the diffusion model, performing text conditioning and forward diffusion of the text request using the diffusion model to generate a noisy transformed image that satisfies the text request, using the diffusion model to output the denoised transformed image based on the noisy transformed image, the extracted features, and the self-attention map, and blending the denoised initial image, the storage mask, and the denoised transformed image to form an output image, the storage mask preventing modifications to the face from the initial image.

[0011] In some embodiments, outputting the denoised initial image includes performing de-diffusion of the initial image using a diffusion model to generate a noisy initial image based on the initial image, providing the noisy initial image to a first convolutional neural network (CNN) and outputting the denoised initial image, and outputting the denoised transformed image includes providing the noisy transformed image to a second CNN, injecting the extracted features and self-attention map into the diffusion, and outputting the denoised transformed image. In some embodiments, the de-diffusion is DDIM inversion. In some embodiments, the operations further include receiving a selection of a first object in the initial image, and the text request includes a comment for replacing the first object in the initial image with a second object. In some embodiments, the operations further include identifying an area in the background to replace or modify from the initial image, and providing a suggestion to replace or modify the background, and the text request is associated with the suggestion. In some embodiments, the operations further include identifying an area in the background to replace or modify from the initial image, and providing a suggestion to replace or modify the background, wherein a text request is associated with the suggestion. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram of an exemplary network environment, according to some embodiments described herein. [Figure 2] FIG. 1 is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3A] FIG. 1 illustrates an exemplary initial image, according to certain embodiments described herein. [Figure 3B] 10A-10C illustrate exemplary output images in which shirts are swapped and objects are replaced with different objects, according to some embodiments described herein. [Figure 3C]10A-10C illustrate exemplary output images in which a user's clothing, hair, and skin are modified, according to some embodiments described herein. [Figure 3D] 10A-10C illustrate exemplary output images with modified backgrounds according to certain embodiments described herein. [Figure 4] FIG. 1 illustrates an exemplary user interface including options for selecting different areas of an image to modify, a global preset to apply, a field for providing text, and an exemplary output image, according to some embodiments described herein. [Figure 5] FIG. 10 illustrates an example user interface including options for receiving user input to select an object to be replaced with a text request, according to some embodiments described herein. [Figure 6] FIG. 1 illustrates an exemplary user interface including a menu of options and a library of pre-made prompts, according to some embodiments described herein. [Figure 7] FIG. 1 illustrates an example architecture for generating and using self-attention maps during diffusion, according to some embodiments described herein. [Figure 8] FIG. 1 illustrates an example architecture for generating an image using an inversion and text conditional diffusion process, according to some embodiments described herein. [Figure 9] FIG. 1 is a block diagram of an example architecture for generating an image incorporating a text request, according to some embodiments described herein. [Figure 10] FIG. 2 illustrates an example method for generating an output image from a text request, according to some embodiments described herein. [Figure 11] FIG. 10 illustrates another exemplary method for generating an output image from a text request, according to some embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION

[0013] Generative artificial intelligence (AI) is sometimes used to generate images from text prompts. For example, a user can request an image of an avocado chair, which is then created by the generative AI. The results are often problematic, especially when the image contains a person, as more detailed aspects may be inadequately represented. For example, generative AI is still developing when it comes to capturing the fine details of features such as fingers, eyes, and mouths.

[0014] The techniques described below include a media application receiving an initial image and a text request to modify the initial image. The initial image includes subjects with faces, such as an initial image of a family on a mountain, and a text request to replace the family's clothing, including heavy jackets, with summer clothing. The text request may be received directly from a user or may be selected from a library of pre-made prompts or a menu of options. The media application generates a stored mask corresponding to the face.

[0015] The text request, the initial image, and the preservation mask are provided as inputs to a diffusion model. The diffusion model outputs a denoised initial image based on the initial image, performs text conditioning and forward diffusion of the text request to generate a noisy transformed image that satisfies the text request, and outputs the denoised transformed image based on the noisy transformed image, the extracted features, and the self-attention map. The denoised initial image, the preservation mask, and the denoised transformed image are blended to form an output image, where the preservation mask prevents modifications to the face from the initial image. The extracted features and the self-attention map are used to maintain the structure of the initial image and modify the initial image in a process that is faster than if the output image were generated without reference to the initial image.

[0016] The media application may perform additional steps, such as identifying areas to replace or modify. Continuing with the example above, the media application may suggest reducing clouds in the background. The media application may also identify objects to remove. For example, the media application may suggest removing other people and hiking equipment from the background of the initial image.

[0017] Exemplary Environment 100

[0018] FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, environment 100 includes a media server 101, user device 115a, and user device 115n coupled to a network 105. Users 125a, 125n may be associated with respective user devices 115a, 115n. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1. In FIG. 1 and the remaining figures, a letter following a reference number, e.g., "115a," denotes a reference to the element with that specific reference number. A reference number in text without a following letter, e.g., "115," denotes a general reference to an embodiment of the element bearing that reference number.

[0019] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively coupled to a network 105 via signal line 102. The signal line 102 may be a wired connection, such as Ethernet, coaxial cable, or fiber optic cable, or a wireless connection, such as Wi-Fi, Bluetooth, or other wireless technology. In some embodiments, the media server 101 transmits and receives data to and from one or more of the user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0020] The database 199 may store machine learning models, training data sets, images, etc. The database 199 may also store social network data associated with the user 125, the user preferences of the user 125, etc.

[0021] The user device 115 may be a computing device that includes a memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device that can access the network 105.

[0022] In the illustrated implementation, user device 115a is coupled to network 105 via signal line 108, and user device 115n is coupled to network 105 via signal line 110. Media application 103 may be stored as media application 103b on user device 115a and / or as media application 103c on user device 115n. Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections, such as Wi-Fi, Bluetooth, or other wireless technologies. User devices 115a and 115n are accessed by users 125a and 125n, respectively. The user devices 115a and 115n in FIG. 1 are used as an example. While FIG. 1 shows two user devices 115a and 115n, the present disclosure applies to system architectures having one or more user devices 115.

[0023] The media application 103 may be stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are executed on the media server 101 or the user device 115. In some embodiments, some operations may be executed on the media server 101 and some may be executed on the user device 115. Execution of operations is subject to user settings. For example, the user 125a may specify that operations be executed on their respective user device 115a and not on the media server 101. Such settings result in operations described herein being executed entirely on the user device 115a and no operations being executed on the media server 101. Furthermore, the user 125a may specify that user images and / or other data be stored only locally on the user device 115a, and not on the media server 101. Such settings result in user data not being transmitted to or stored on the media server 101. The transmission of user data to the media server 101, any temporary or permanent storage of such data by the media server 101, and the performance of actions on such data by the media server 101 are performed only if the user consents to the transmission, storage, and performance of actions by the media server 101. The user is provided with the option to change settings at any time, for example, so that the user can enable or disable use of the media server 101.

[0024] Machine learning models (e.g., neural networks or other types of models) are stored locally on the user device 115 and utilized for one or more operations with specific user permission. Server-side models are used only if authorized by the user. Additionally, trained models may be provided for use on the user device 115. During such use, on-device training of the model may be performed if authorized by the user 125. Updated model parameters may be sent to the media server 101 if authorized by the user 125, for example, to enable federated learning. The model parameters do not include any user data.

[0025] The media application 103 receives an initial image and a text request to modify the initial image, where the initial image includes a subject with a face. For example, the media application 103 receives the initial image from a camera that is part of the user device 115, or the media application 103 receives the initial image over the network 105. From the initial image, the media application 103 generates a storage mask corresponding to the subject's face. The media application 103 provides the text request, the initial image, and the storage mask as inputs to a diffusion model. The diffusion model outputs a denoised initial image based on the initial image, performs text conditioning and forward diffusion of the text request to generate a noisy transformed image that satisfies the text request, and outputs the denoised transformed image based on the noisy transformed image, the extracted features, and the self-attention map. The media application 103 blends the denoised initial image, the storage mask, and the denoised transformed image to form an output image, where the storage mask prevents modifications to the face from the initial image. The output image that satisfies the text requirements corresponds to the initial image, and in the output image, the pixels of the subject's face are the pixels of the subject's face in the initial image, and the pixels that do not define a face are the pixels of the initial image modified by the diffusion process according to the text requirements.

[0026] In some embodiments, media application 103 may be implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, media application 103a may be implemented using a combination of hardware and software.

[0027] Exemplary Computing Device 200

[0028] 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. The computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a media server 101 used to implement a media application 103a. In another example, the computing device 200 is a user device 115.

[0029] In some embodiments, computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, a display 241, a camera 243, and a storage device 245, all coupled via bus 218. Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, display 241 may be coupled to bus 218 via signal line 228, camera 243 may be coupled to bus 218 via signal line 230, and storage device 245 may be coupled to bus 218 via signal line 232.

[0030] Processor 235 may be one or more processors and / or processing circuits that execute program code and control basic operations of computing device 200. A “processor” includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuits for achieving functionality, dedicated processors for performing neural network model-based processing, neural circuits, systems with processors optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, processor 235 may include one or more coprocessors that perform neural network processing. In some embodiments, processor 235 may be a processor that processes data to produce a probabilistic output; for example, the output produced by processor 235 may be inaccurate or accurate within a range from an expected output. Processing need not be limited to a particular geographic location or have time limitations. For example, a processor may perform its functions in real time, offline, in batch mode, etc. Portions of processing may be performed at different times and in different locations by different (or the same) processing systems. A computer may be any processor in communication with a memory.

[0031] Memory 237 is typically provided in computing device 200 for access by processor 235 and may be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., located separately from and / or integrated with processor 235, suitable for storing instructions for execution by the processor or set of processors. Memory 237 may store software operated on computing device 200 by processor 235, including media application 103.

[0032] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, etc. One or more methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application (“app”) running on a mobile computing device, etc.

[0033] Application data 266 may be data generated by other applications 264 or by hardware of computing device 200. For example, application data 266 may include images used by an image library application, user actions identified by other applications 264 (e.g., social networking applications), etc.

[0034] I / O interface 239 may provide functionality that allows computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of computing device 200 or may be separate and in communication with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices may communicate through I / O interface 239. In some embodiments, I / O interface 239 may connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.).

[0035] Some examples of interfaced devices that can be connected to I / O interface 239 may include a display 241 that can be used to display content, such as images, video, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from a user. For example, display 241 may be utilized to display a user interface, including a graphical guide, on a viewfinder. Display 241 may include any suitable display device, such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in an eyeglass form factor or headset device, or a monitor screen of a computing device.

[0036] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that I / O interface 239 sends to media application 103.

[0037] The storage device 245 stores data related to the media application 103. For example, the storage device 245 may store training data sets including labeled images, machine learning models, output from machine learning models, etc.

[0038] FIG. 2 illustrates an exemplary media application 103 stored in memory 237 that includes a user interface module 202 , a segmenter 204 , a repair module 206 , and a diffusion module 208 .

[0039] The user interface module 202 generates graphic data for displaying a user interface including an image. The user interface module 202 receives an initial image. The initial image may be received from the camera 243 of the computing device 200 or from the media server 101 via the I / O interface 239. The initial image includes a subject having a face, such as a person or an animal.

[0040] The user interface includes an option for providing a text request associated with the initial image. For example, the user interface may include a text field in which the user directly enters the text request, an audio button for providing audio input that is converted into the text request, etc. In some embodiments, the user interface may update with autocomplete suggestions while the user is providing the text request. For example, for an outdoor scene in which the text field includes "change to m," the user interface module 202 may add "mountain" as an autocomplete suggestion. The user interface also includes an option for receiving user input to select objects and regions to modify and / or replace.

[0041] The user interface module 202 may identify one or more objects in the initial image to replace. In some embodiments, the user interface module 202 performs object recognition to identify objects in the initial image and provide suggestions for replacing the objects. For example, the user interface may include a menu of options or a library of pre-built prompts for highlighting and selecting objects. The user interface may include a library of pre-built prompts regarding how to change the scene, change the clothing of a person in the image, etc. For example, a user may select a pre-built prompt to change an initial image of a beach scene with a house in the background to an output image in which the house is replaced with a sandcastle.

[0042] In some embodiments, the user interface includes options for selecting various people, objects, parts of people or objects, or backgrounds in the initial image to modify or change based on a request. For example, a user may select a person by tapping on a person, circling an object, etc. In some embodiments, the user interface generates recommendations for modifying the image, such as highlighting the person's boots in the image, and asking if the user would like to change the person's boots.

[0043] In some embodiments, the text request relates to replacing a first object in the initial image with a second object. For example, a user may select a first object and provide text input related to replacing the selected object with a second object that corresponds to the text request.

[0044] The user interface module 202 may identify areas in the background of the initial image to replace or modify and provide suggestions for replacing or modifying the background. For example, the suggestions may include global presets, such as a list of themes from which to select to change the initial image.

[0045] The user interface module 202 may identify an object in the background of the initial image to remove. For example, the object may be identified by the user interface module 202 in response to performing object recognition. The user interface may provide suggestions for removing the object from the background. In some embodiments, the user interface includes a text field for receiving a text request from the user to replace the object based on the text request.

[0046] In some embodiments, the user interface module 202 generates graphical data for displaying the output image. The user interface may also include options for editing the output image, sharing the output image, adding the output image to a photo album, etc.

[0047] 3A , an exemplary initial image 300 is shown in accordance with some embodiments described herein. The initial image 300 includes a woman 302 as a subject, a shirt 307 worn by the woman 302, and a suitcase 304. The initial image 300 is displayed as part of a user interface (not shown) that includes a way for a user to provide user input. For example, the user interface may include a text field for receiving text input, and the user interface may identify a finger or mouse (or other indicator) for selecting an object by clicking on the object, circling the object, moving back and forth over the object to highlight it, and so on.

[0048] Figure 3B shows an example output image 310 in which shirt 312 has been changed from shirt 307 in Figure 3A and suitcase 304 in Figure 3A has been replaced with turtle 314 in Figure 3B, according to some embodiments described herein. To achieve output image 310 in Figure 3B, a request provided for initial image 300 in Figure 3A can include science shirt 312 and replace suitcase 304 in Figure 3A with turtle 314.

[0049] 3C shows an exemplary output image 320 in which a subject's shirt 322 and hair 324 have been altered and a tattoo 326 has been added to the subject's arm, according to some embodiments described herein. To achieve the output image 320 in FIG. 3C, a user may have provided a text request to the initial image 300 in FIG. 3A to make the subject more punk rock, selected a punk rock theme from a library of pre-made prompts, and individually altered each object in the input image 300.

[0050] 3D shows an example output image 330 with a changed background, according to some embodiments described herein. To achieve the output image 330 in FIG. 3D, the text request provided to the initial image 300 in FIG. 3A may include a request to add a waterfall background. In some embodiments, the save mask in FIG. 3D encompasses all of the subject (e.g., not just the face) because the save mask prevents the subject from being modified while the background is replaced.

[0051] As discussed in more detail below, the diffusion module 208 uses the conservation masks of Figures 3B, 3C, and 3D, which include at least the subject's face, to prevent the subject's face from being modified during blending with the composite image.

[0052] 4 illustrates exemplary user interfaces 400, 425, 450 including options for selecting different regions of an image to modify, a global preset to apply, a field for providing text, and an exemplary output image, according to some embodiments described herein. Specifically, the first user interface 400 automatically provides a global preset 405 for a user to select modifications to an input image 401, such as an oil painting, a surreal world, or a nostalgic scene. The user can temporarily modify the input image by selecting an option from the global preset 405, such as an oil painting, a surreal world, or a nostalgic scene, and can return to the original input image by selecting no option from the global preset 405.

[0053] The first user interface 400 also includes circles 410, 411, 412 that represent the identification of different regions in the initial image 401. The user can specify changes to make to the sky by tapping the first circle 410, to the bridge by tapping the second circle 411, and to the people by tapping the third circle 412.

[0054] In response to a user selecting one of the circles 410, 411, 412, the user interface may update the display to provide a menu of options (not shown). For example, selecting the first circle 410 may cause the user interface to display suggestions such as changing a cloudy sky to a clear sky. Selecting the second circle 411 may cause the user interface to display suggestions such as an option to remove the bridge associated with the second circle 411, or to replace the bridge with a different type of bridge or a boat. Selecting the third circle 412 may cause the user interface to display a suggestion to remove a person.

[0055] The second user interface 425 includes an input image 426 and a text entry field 430 where the user can specify the changes they want to make. The user can include a description specific enough to encompass the object they want to change (e.g., changing boots to colorful, sparkly boots), or the user can select the object they want to change in the second user interface 425 and then describe the specific changes to be made. For example, the user may select the object by tapping on it, circling it, marking it, etc. In this case, the user selects the boots 427 on the subject.

[0056] A third user interface 450 includes an output image 451 in which the text request 452 for "colorful sparkly boots" is fulfilled. The boots 453 are changed to sparkly, colorful stars. The user interface also includes an option that allows the user to undo 454 the changes made to the initial image.

[0057] 5 shows exemplary user interfaces 500, 525, 550 including options for receiving user input for selecting objects to be replaced based on a text request, according to some embodiments described herein. The first user interface 500 includes an initial image 501 and suggested global presets 505. In some embodiments, the user interface module 202 performs object recognition on the initial image 501 to determine objects in the initial image 501 and provides suggested global presets 505 based on the determined objects. For example, the user interface module 202 may suggest stylized, sketched, and vintage global presets 505 for an outdoor scene, as these presets are particularly suited to outdoor scenes.

[0058] The second user interface 525 includes an initial image 526 in which the user has circled a specific object in the initial image 526. The second user interface 525 highlights the selected object 527 with an outline. A text entry field 530 includes the text the user has entered, "rock," after the initial prompt "change to," to indicate that the user wants to replace the highlighted section with a rock. The second user interface 525 also includes suggested objects 532, such as rocks, water, and shrubs. In some embodiments, the suggested objects are provided based on the user interface module 202 performing object recognition and suggesting objects that would commonly be found along with other objects found in the image, such as streams, rocks, and mountains.

[0059] The third user interface 550 includes an output image 551 generated by the user interface module 202 to replace the selected object 527 in the second user interface 525 with a rock specified by the user. The third user interface 550 also includes options to save a copy of the output image 552 and to undo changes to the output image 553.

[0060] 6 shows an example user interface 600 including a menu of options 605 and a library of pre-made prompts 610, according to some embodiments described herein. In this example, the menu of options 605 includes options for modifying the subject's clothing and scenery in the input image 601, as well as free text options for providing more specific changes. The library of pre-made prompts 610 includes various themes that can be applied to the input image 601, including ocean adventurer, ancient warrior, space crusader, sage, nobleman, and space mission.

[0061] In some embodiments, the user interface module 202 generates a user interface that includes options for modifying user preferences. For example, the user interface may include user preferences for specifying a level of probability, including how much noise the user wants to see in the output image (e.g., the degree to which the output image differs from the initial image) and the degree to which realistic seeds are used in the output image (e.g., the degree to which the output image differs from reality). For example, if the input image shows a boy drinking liquid from a straw from a mug, a spectrum of increasing probability would result in output images with small changes, such as a different type of mug and a slightly different-looking background, to more widespread changes, such as a different mug, a completely different background, different clothing for the boy, and a different table on which the mug is sitting. In another example, as the degree of seed type increases, a spectrum of increasing differences in seed type would result in output images with an unrecognizable mug and the background changed from a countertop to another room in a house, to an unrecognizable mug and the background changed to a room on a spaceship.

[0062] In some embodiments, the user interface module 202 generates a graphical user interface that includes options for receiving user input from a user that can be converted into an image. For example, a user may sketch the outline of a dinosaur, and the diffusion module 208 may generate an output image that includes the dinosaur based on the sketch. In some embodiments, the user interface receives user input for an initial image. For example, a user may sketch a hat on an initial image of a child, and the user interface module 202 updates the user interface to include the output image generated by the diffusion module 208, which includes a rendering of the sketched hat.

[0063] The segmenter 204 segments one or more objects, including the subject's face, from the initial image. A face segment includes pixels corresponding to the location of the face in the initial image. The segmenter 204 segments the subject's face to generate a storage mask that the diffusion module 208 uses to prevent modification of the face during generation of the output image. Face segmentation may be used to prevent modification to the subject's face while changing aspects of the subject's hair, clothing, etc.

[0064] The segmenter 204 may also not only segment the face, such as the entire body, in which case the entire body is prevented from being modified. A body segment includes pixels corresponding to the location of the body in the initial image. Body segmentation may be used to prevent modifications to the entire body while the rest of the image is modified, such as changes to the background of the initial image. In some embodiments, the saved mask includes all aspects of the initial image except for the portions that are being modified. For example, the saved mask may encompass the face, hair, and background, while the subject's clothing is modified.

[0065] The segmenter 204 may segment other objects in the initial image automatically or in response to user input. For example, if the user interface module 202 generates suggestions to modify, remove, and / or replace an object in the initial image, the segmenter 204 segments the object. In another example, the user interface receives user input identifying an object to be modified, removed, and / or replaced, and the segmenter 204 segments the object in response to the object being selected. In some embodiments, the segmenter 204 generates a segmentation map that associates identification with each pixel in the initial image as belonging to a face, body, object, etc.

[0066] In some embodiments, segmenter 204 uses an alpha map as part of a technique for distinguishing between the foreground and background of the initial image during segmentation. Segmenter 204 may also identify the texture of selected objects in the foreground of the initial image.

[0067] The segmenter 204 may perform segmentation by detecting objects in the initial image. The objects may be people, animals, cars, buildings, etc. A person may be a subject in the initial image or may not be a subject in the initial image (i.e., a bystander). A bystander may include a person walking, running, bicycling, standing, or in another situation behind the subject in the initial image. In another example, a bystander may be in the foreground (e.g., a person passing in front of the camera), at the same depth as the subject (e.g., a person standing next to the subject), or in the background. In some examples, there may be multiple bystanders in the initial image. A bystander may be a human in any pose, e.g., standing, sitting, crouching, lying down, jumping, etc. A bystander may be facing the camera, at an angle to the camera, or facing away from the camera.

[0068] The segmenter 204 may perform object recognition and detect the type of object by comparing the object to prior information about objects such as people, vehicles, buildings, etc., to identify the object's expected shape and determine whether a pixel is associated with the selected object or the background. The segmenter 204 may generate a region of interest for the selected object, such as a bounding box with x-coordinates, y-coordinates, and scale.

[0069] The segmenter 204 generates a storage mask encompassing at least the face of the subject. The face storage mask may include pixels corresponding to pixels of the face segment in the initial image. In some embodiments, the storage mask includes additional or different body parts, such as the subject's entire head, hands, or body. In some embodiments, the storage mask is generated based on generating superpixels of the image and matching the centers of the superpixels to depth map values ​​(e.g., obtained by the camera 243 using a depth sensor or by deriving depth from pixel values) to perform depth-based cluster detection. More specifically, the depth values ​​of the masked area may be used to determine a depth range, and superpixels that fall within the depth range may be identified. Another technique for generating a mask includes weighting depth values ​​based on how close the depth values ​​are to the mask, where the weights are represented by a distance transform map.

[0070] In some embodiments, segmenter 204 may specify circuitry (e.g., for a programmable processor, for a field programmable gate array (FPGA), etc.) that enables processor 235 to apply the machine learning model. In some embodiments, segmenter 204 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, segmenter 204 may provide an application programming interface (API) that can be used by operating system 262 and / or other applications 264 to invoke segmenter 204, for example, to apply the machine learning model to application data 266 and output a saved mask.

[0071] The segmenter 204 uses training data to generate a trained machine learning model. For example, the training data may include pairs of an initial image with one or more objects and an output image with one or more storage masks.

[0072] The training data may come from any source, e.g., a data repository specifically marked for training, data that has been given permission to be used as training data for machine learning, etc. In some embodiments, training may occur on the media server 101, which provides the training data directly to the user device 115, training may occur locally on the user device 115, or a combination of both.

[0073] In some embodiments, segmenter 204 uses weights obtained and unedited / transferred from another application. For example, in these embodiments, a trained model may be generated, e.g., on a different device, and provided as part of segmenter 204. In various embodiments, the trained model may be provided as a data file that includes a model structure or format (e.g., defining the number and type of neural network nodes, the connectivity between the nodes, and the organization of the nodes into multiple layers) and associated weights. Segmenter 204 may read the trained model data file and implement a neural network with node connectivity, layers, and weights based on the model structure or format specified in the trained model.

[0074] The trained machine learning model may include one or more model formats or structures. For example, the model format or structure may include any type of neural network, such as a linear network, a deep learning neural network implementing multiple layers (e.g., "hidden layers" between an input layer and an output layer, each layer being a linear network), a convolutional neural network (e.g., a network that divides or partitions input data into multiple portions or tiles, processes each tile separately using one or more neural network layers, and aggregates the results from processing each tile), a sequence-to-sequence neural network (e.g., a network that receives sequential data as input, such as words in a sentence or frames in a video, and produces a sequence of results as output), etc.

[0075] The model format or structure may specify the connectivity between various nodes and their organization into layers. For example, nodes in a first layer (e.g., an input layer) may receive data as input or application data. Such data may include, for example, one or more pixels per node, for example, when the trained model is used to analyze an initial image. Subsequent intermediate layers may receive as input the output of nodes in the previous layer according to the connectivity specified in the model format or structure. These layers may also be referred to as hidden layers. For example, the first layer may output a segmentation between foreground and background. The final layer (e.g., an output layer) produces the output of the machine learning model. For example, the output layer may receive the segmentation of the initial image into foreground and background and output whether a pixel is part of a saved mask. In some embodiments, the model format or structure also specifies the number and / or type of nodes in each layer.

[0076] In another embodiment, the trained model may include one or more models. One or more of the models may include multiple nodes arranged in layers according to a model structure or format. In some embodiments, a node may be a memoryless computational node configured, for example, to process a unit of input and generate a unit of output. The computation performed by the node may include, for example, multiplying each of multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce the node output. In some embodiments, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, computations by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, individual processing units of a graphics processing unit (GPU), or dedicated neural circuitry. In some embodiments, a node may include memory, for example, capable of storing and using one or more previous inputs when processing a subsequent input. For example, a node with memory may include a long short-term memory (LSTM) node, which may use memory to maintain a "state" that allows the node to behave like a finite state machine (FSM).

[0077] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may begin as multiple nodes organized into layers, as specified by the model format or model structure. At initialization, corresponding weights may be applied to nodes connected according to the model format, e.g., to connections between each pair of nodes in successive layers of a neural network. For example, each weight may be randomly assigned or initialized to a default value. The model may then be trained, e.g., using training data, to produce results.

[0078] Training can include applying supervised learning techniques. In supervised learning, the training data can include multiple inputs (e.g., images, storage masks, etc.) and corresponding ground truth outputs for each input (e.g., ground truth masks that correctly identify portions of the subject, such as the subject's face, in each image). Based on a comparison of the model's outputs and the ground truth outputs, the values ​​of the weights can be automatically adjusted, for example, in a manner that increases the probability that the model will produce the ground truth output for the image.

[0079] In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In some embodiments, the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights. In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In embodiments where data is omitted, the segmenter 204 may generate a trained model that is based on prior training, e.g., by a developer of the segmenter 204, by a third party, etc. In some embodiments, the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights.

[0080] In some embodiments, the trained machine learning model receives an initial image having one or more objects. In some embodiments, the trained machine learning model outputs one or more saved masks corresponding to the one or more objects. For example, the one or more saved masks may be for the faces of the one or more objects.

[0081] In situations where an object is removed from the initial image, the inpainting module 206 generates an inpainting image that replaces object pixels corresponding to one or more objects with background pixels. The background pixels may be based on pixels from a reference image of the same location without the object. Alternatively, the inpainting module 206 may identify background pixels to replace the removed object based on the proximity of the background pixels to other pixels surrounding the object. The inpainting module 206 may use gradients of nearby pixels to determine the characteristics of the background pixel. For example, if a bystander was standing on the ground, the inpainting module 206 would replace the background pixel with a ground pixel. Other inpainting techniques are possible, including machine learning-based inpainting techniques that output background pixels based on training data containing images of similar structure.

[0082] In embodiments in which the user chooses to erase the selected object, the user interface module 202 may display an inpainted image in which the selected object has been removed and the selected object pixels have been replaced with background pixels.

[0083] The diffusion model receives as input a request (e.g., a text request provided directly by the user, a selection of a pre-made prompt, a selection of a global preset, a selection of an option from a menu, etc.), an initial image, and a storage mask. The diffusion model encodes the image in latent space, performs diffusion, and decodes it to pixel space.

[0084] The diffusion module 208 performs text conditioning of the request. Text conditioning describes the process of generating an image that is conditioned on (e.g., aligned with) a text prompt. For example, if the text request is to replace the red shirt worn by the subject in the initial image with a blue shirt, the diffusion model 208 performs text conditioning by generating an output image of the blue shirt.

[0085] In some embodiments, the diffusion module 208 trains the diffusion model using two types of training data. The first type of training data includes pairs of images, which may include synthetic pairs generated through an inter-prompt generation machine learning model. The inter-prompt generation machine learning model is a diffusion model that receives a text prompt, uses self-attention to extract keys and values ​​from the text prompt, switches portions of the attention map previously generated for the input image based on the input text prompt, and outputs an output image that matches the text prompt.

[0086] The inter-prompt generative machine learning model generates a self-attention map. Self-attention calculates the interactions between different elements of an input sequence (e.g., different words in a text request). This contrasts with cross-attention, where the interactions are between two different input sequences (e.g., how a text request relates to the original prompt).

[0087] A self-attention map describes the structure and different semantic domains in an image. For example, an image described in a self-attention map as "pepperoni pizza next to orange juice" includes how certain pixels on the pizza's crust are attentive to other pixels on the crust. Conversely, in a cross-attention map, pixels on the pizza's crust are attentive to the orange juice.

[0088] Turning to Figure 7, an exemplary architecture 700 for generating and using self-attention maps during diffusion is shown, according to some embodiments described herein. Figure 7 includes two ways of describing the process: text-to-image self-attention 705 and self-attention control 740.

[0089] During text-to-image self-attention 705, the diffusion module 208 generates an attention map by extracting pixel features 710 from the noisy input image. The pixel features 710 from the noisy input image are projected onto a pixel query 715 matrix. The text embeddings are projected onto a key matrix, which takes the form of token keys from an initial image prompt 720. The initial image prompt may be "The cat in the hat is lying on a beach chair."

[0090] The pixel query 715 is multiplied with the token key from the initial image prompt 720 to produce different layer self-attention maps 725. Different layers are used to focus on different features in the input image at different levels of abstraction. The self-attention map 725 contains rich semantic relationships that greatly influence the generated image. In the self-attention map 725, each cell defines the weight of the value of a particular token for a particular pixel based on the latent projection dimension of the token key from the initial image prompt 720 and the pixel query 715.

[0091] The self-attention map 725 is used to create an output image that is a modification of the initial image. Continuing with the example above, the initial image prompt could be modified to "A cat wearing a Musketeer hat is lying on a beach chair." The diffusion model creates token values ​​from the request prompt 730. The self-attention map 725 is multiplied by the token values ​​from the request prompt 730 to obtain a self-attention output 735, where the weights are an attention map that correlates with the similarity between the pixel query 714 and the token keys from the initial image prompt 720. The self-attention output 735 is used to create the output image.

[0092] The self-attention control 740 shows how the self-attention map is revised based on differences between the initial pictorial prompt and the request prompt. For example, "cat on chair" may be replaced with "dog on chair." The main challenge is to address the content of the new prompt while preserving the original structure. If a word in the initial pictorial prompt is replaced, the diffusion model performs a word swap 745 from the self-attention map 725 to the revised self-attention map 750. The revised self-attention map 750 replaces the "cat" token with a "dog" token. If a word is replaced using a different number of tokens, the diffusion model may replicate or average the self-attention map 725 to obtain the revised self-attention map 750.

[0093] Additional words may be added to the initial picture prompt. For example, "cat on a chair" may be replaced with "black cat on a white chair." The diffusion model performs prompt refinement 750 by adding self-attention maps for the new words and creating an elaborated self-attention map 755 that incorporates the new self-attention maps for the new words. To preserve common details, the diffusion model applies attention injection to common tokens from both prompts (e.g., cat on a chair). In some embodiments, the diffusion model uses an alignment function that receives token indexes from the target prompt and outputs corresponding token indexes.

[0094] The self-attention map is used in a text-conditional diffusion model to use structure and different semantic regions in the input image to change one or more token values ​​while keeping the self-attention map fixed and preserving the composition of the scene. In some embodiments, the diffusion model adds a new word to the prompt, keeping attention to previous tokens static while allowing new attention to flow to the new token. This results in specific objects in the input image being globally edited or modified to match the text request.

[0095] Each diffusion step predicts the noise from the noisy image and the text embedding. In the final step, the process results in a generated image. The interaction between the text prompt and the image occurs during noise prediction, where visual features and text feature embeddings are fused using a self-attention layer, resulting in a spatial attention map for each text token.

[0096] The second type of training data includes pairs of real and synthetic images. The real images are received by a diffusion model, such as a denoising diffusion implicit model (DDIM). The diffusion model outputs a synthetic image using an inversion method based on the real image and instructions on how to edit the input image. The diffusion module 208 trains the diffusion model to generate output images from the requests using a forward process, in which the diffusion model adds noise to the data, and an inverse process, in which the diffusion model learns to recover the data from the noise.

[0097] 8 shows an example transformation 800 including an inversion 805 and a text conditional diffusion process 815, according to some embodiments described herein. The inversion 805 starts with an input image 810 with a text prompt that reads, "The cat in the hat is lying in a beach chair." The inversion gradually adds noise to the input image 810.

[0098] A text-conditional diffusion process 815 conditions images based on text requests. In this example, "The Musketeers" is added to the text request associated with the input image 810, resulting in a text request for images containing "A cat wearing a Musketeer hat lying on a beach chair." The diffusion model is forward diffusion, combining text conditioning of the text request with the noisy image.

[0099] The diffusion module 208 trains the diffusion model to maintain photorealism and preserve the identities of people shown in the image. During training, the diffusion model receives editing instructions and modifies the editing instructions to create corresponding prompts based on a language model, such as a large-scale language model. For example, the diffusion module 208 uses a language model to turn the editing instruction "Make a person look like an astronaut" into a prompt that describes various aspects of what clothing to apply to make it look like a space suit.

[0100] The diffusion model creates a set of input and output image pairs from the generated prompt pairs, where each prompt can generate N images (using a different seed). The diffusion module 208 filters certain images from the image pairs, such as image transformations that do not match the given editing instructions, image transformations that do not result in sufficiently aligned images, and mismatched pairs. In some embodiments, the diffusion module 208 also filters images based on an edit alignment score that reflects the alignment between the image-to-image transformation and the original edited caption, and an image-text alignment score that reflects the alignment between the input / output image and the corresponding input / output prompt. In some embodiments, the diffusion module 208 trains the diffusion model by generating one or more loss functions based on the filtered images from the image pairs.

[0101] The diffusion model is trained to generate images by gradually adding noise to the image, and then the diffusion model learns how to gradually remove the noise. The diffusion model applies a denoising process to random seeds to generate realistic images. By simulating diffusion, the diffusion model generates one or more noisy images.

[0102] Once the diffusion model is trained, it receives an input image and performs a de-diffusion process on the initial image to generate a noisy image based on the initial image. In some embodiments, the diffusion model 208 performs de-diffusion using DDIM inversion.

[0103] The diffusion model provides a noisy image to a first CNN equipped with features and a self-attention mechanism. The first CNN samples the input image and extracts features from the input image. The first CNN directly injects the extracted features and self-attention map into a second CNN. The first CNN performs forward diffusion of the noisy initial image, which is a process of gradually denoising the noisy image using sampling to output a denoised initial image.

[0104] The text request and the noisy image are provided as inputs to a second CNN, which uses a self-attention map to align semantic features of the text request with the structure of the noisy image to generate a noisy transformed image, and performs forward diffusion on the noisy transformed image to output a denoised transformed image.

[0105] The denoised initial image is combined with the denoised transformed image and the preserved mask, which advantageously prevents modifications to the face that might otherwise be modified in a way that results in unrealistic features. In some embodiments, the diffusion module 208 performs the blending by using a mask smoothing algorithm and Poisson blending.

[0106] In some embodiments, the saved mask includes other parts of the subject, such as the subject's hair if the user wants to keep their hair intact, the subject's fingers since fingers are often retouched in unrealistic ways by machine learning models, the subject's entire body if the subject is a pet to prevent the pet from being over-retouched, etc. In some embodiments where the output image retouches the subject's clothing, the saved mask may include everything but the subject's clothing, so that the body (excluding clothing) and background of the initial image are preserved.

[0107] The combined denoised image and preservation mask is blended with the denoised transformed image to form an output image that meets the text requirements.

[0108] 9, a block diagram of an exemplary architecture 900 for generating an output image incorporating a text request is shown. An initial image 905 is provided to a diffusion model, where a denoising diffusion implicit model (DDIM) inversion 910 is performed on the input image 705 to output a noisy image 915 obtained by inverting the input image. During DDIM inversion, features are extracted from the input image.

[0109] The noisy image 915 is provided as input to a first CNN 920 with a self-attention mechanism. The CNN utilizes an aggregation function over local receptive fields according to the convolutional filter weights shared by the feature map. The self-attention map applies a weighted averaging operation based on the context of the input features, where the attention weights are dynamically calculated using a similarity function between related pixel pairs. The exemplary architecture 900 uses a combination of both types of feature extraction to output a denoised initial image. The first CNN 920 outputs a denoised initial image 925.

[0110] Text conditioning is performed on the text input 935, specifically "add a hat to the photo," to predict the class that best matches the output image corresponding to the text input 935. Using the text conditioning, a noisy text-guided transformed image 940 is generated and provided to a second CNN 945. The second CNN 945 receives the extracted features and self-attention map from the first CNN 920 and uses the extracted features and self-attention map to align the noisy text-guided transformed image 940 with the structure of the initial image 905.

[0111] The second CNN 945 outputs the denoised transformed image blended with the denoised initial image 925 combined with the preservation mask 930 to form an output image 950 that satisfies the text requirement of a subject wearing a hat. In this example, the preservation mask 930 encompasses the face to prevent modification of the face from the initial image 905 during blending. At each blending step, the original latent image inside the mask is used instead of the denoised transformed image to preserve the original features.

[0112] Exemplary Methods

[0113] 10 shows an example flowchart of a method 1000 for generating an output image. Method 1000 may be performed by computing device 200 in FIG. 2. In some embodiments, method 1000 is performed by user device 115, media server 101, or partially on user device 115 and partially on media server 101.

[0114] 10 may begin at block 1002. In block 1002, an initial image and a text request to modify the initial image are received, the initial image including a subject with a face. The text request may include a request to modify the subject of the initial image, a background of the initial image, and / or an attribute of a selection of a first object, and a request to modify the first object of the initial image with a second object. In some embodiments, the request is at least one selection from a group of global presets, a menu of options, and / or a library of pre-made prompts.

[0115] The method may further include receiving a selection of a first object in the initial image, the text request including a comment to replace the first object in the initial image with a second object. The method further includes identifying an area in a background to replace or modify from the initial image and providing a suggestion to replace or modify the background, the text request being associated with the suggestion. The text request may include a request to change the background of the initial object.

[0116] In addition to the text request, the method may further include identifying from the initial image an object in the background to remove and providing a suggestion to remove the object from the background. In some embodiments, the method includes identifying from the initial image one or more objects to replace and providing a suggestion to replace the object. Block 1002 may be followed by block 1004.

[0117] In block 1004, a saved mask corresponding to at least the subject's face is generated. The segmenter 204 may also be applied to other portions of the initial image depending on where the modification is intended to occur. For example, if the text request includes a request to change the background of the initial image, the saved mask may include one or more portions of the subject in addition to the subject's face. Block 1004 may be followed by block 1006.

[0118] In block 1006, the text request, the initial image, and the storage mask are provided as inputs to a diffusion model. Block 1006 may be followed by block 1008.

[0119] In block 1008, the diffusion model performs dediffusion of the initial image to generate a noisy initial image based on the initial image. Block 1008 may be followed by block 1010.

[0120] In block 1010, a first CNN is provided with a noisy initial image and outputs a denoised initial image. Block 1010 may be followed by block 1012.

[0121] In block 1012, the diffusion model performs text conditioning of the text request and forward diffusion to generate a noisy transformed image that satisfies the text request. Block 1012 may be followed by block 1014.

[0122] In block 1014, a second CNN is provided with the noisy transformed image and outputs a denoised transformed image, where the second CNN injects the extracted features and self-attention map to output the denoised transformed image. Block 1014 may be followed by block 1016.

[0123] At block 1016, the denoised initial image, the preservation mask, and the denoised transformed image are blended to form an output image, where the preservation mask prevents modifications to the face from the initial image.

[0124] 11 shows an example flowchart of a method 1100 for generating an output image. Method 1100 may be performed by computing device 200 in FIG. 2. In some embodiments, method 1100 is performed by user device 115, media server 101, or partially on user device 115 and partially on media server 101.

[0125] 11 may begin at block 1102. In block 1102, an initial image and a text request to modify the initial image are received, the initial image including a subject with a face. Block 1102 may be followed by block 1104.

[0126] A storage mask corresponding to the subject's face is generated in block 1104. Block 1104 may be followed by block 1106.

[0127] In block 1106, the text request, the initial image, and the storage mask are provided as inputs to a diffusion model. Block 1106 may be followed by block 1108.

[0128] In block 1108, the diffusion model outputs a denoised initial image based on the initial image. Block 1108 may be followed by block 1110.

[0129] In block 1110, the diffusion model performs text conditioning and forward diffusion of the text request to generate a noisy transformed image that satisfies the text request. Block 1110 may be followed by block 1112.

[0130] In block 1112, the diffusion model outputs a denoised transformed image based on the noisy transformed image, the extracted features, and the self-attention map. Block 1112 may be followed by block 1114.

[0131] In block 1114, the denoised initial image, the preservation mask, and the denoised transformed image are blended to form an output image, where the preservation mask prevents modifications to the face from the initial image.

[0132] In addition to the above, a user may be provided with controls that allow the user to choose both whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information about the user's social networks, social actions, or activities, occupation, user preferences, or the user's current location) and whether content or communications are sent from the server to the user. Furthermore, certain data may be processed in one or more ways to remove personally identifiable information before being stored or used. For example, a user's identifying information may be processed so that personally identifiable information about the user cannot be determined, or if location information is obtained (such as to the city, zip code, or state level), the user's geographic location may be generalized so that the user's specific location cannot be determined. Thus, a user may control what information is collected about them, how that information is used, and what information is provided to them.

[0133] In the foregoing description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the specification. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form to avoid obscuring the specification. For example, embodiments may be described above primarily with reference to user interfaces and specific hardware. However, embodiments are applicable to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.

[0134] A reference herein to "some embodiments" or "some instances" means that a particular feature, structure, or characteristic described in connection with an embodiment or instance may be included in at least one implementation herein. The appearances of the phrase "in some embodiments" in various places in this specification are not necessarily all referring to the same embodiments.

[0135] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. These steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It is convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0136] It should be recognized, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specified, as will be apparent from the discussion that follows, throughout this specification, discussions utilizing terms including "processing," "calculating," "figuring out," "determining," or "displaying," etc., will be understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulates and converts data represented as physical (electronic) quantities in the computer system's registers and memory into other data that is also represented as physical quantities in the computer system's memory or registers, or other such information storage, transmission, or display devices.

[0137]

[0013] Embodiments herein may also relate to a processor for performing one or more steps of the above-described methods. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored on a non-transitory computer-readable storage medium, including, but not limited to, an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory including a USB key with non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0138] This specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, this specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.

[0139] Furthermore, this specification may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any apparatus that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, instruction execution apparatus, or instruction execution device.

[0140] A data processing system suitable for storing or executing program code will include at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.

Claims

1. 1. A computer-implemented method comprising: receiving an initial image and a text request to modify the initial image, the initial image including a subject with a face, the method comprising: generating a storage mask corresponding to the face of the subject; providing the text request, the initial image, and the saved mask as inputs to a diffusion model; outputting a denoised initial image based on the initial image using the diffusion model; performing text conditioning and forward diffusion of the text request using the diffusion model to generate a noisy transformed image that satisfies the text request; outputting a denoised transformed image based on the noisy transformed image, the extracted features, and the self-attention map using the diffusion model; blending the denoised initial image, the preservation mask, and the denoised transformed image to form an output image, wherein the preservation mask prevents modifications to the face from the initial image.

2. outputting the denoised initial image performing dediffusion of the initial image using the diffusion model to generate a noisy initial image based on the initial image; providing the noisy initial image to a first convolutional neural network (CNN) and outputting the denoised initial image; outputting the denoised transformed image providing the noisy transformed image to a second CNN; injecting the extracted features and the self-attention map during diffusion; and outputting the denoised transformed image.

3. The method of claim 2 , wherein the despreading is a denoising diffusion implicit model (DDIM) inversion.

4. 2. The method of claim 1, further comprising receiving a selection of a first object in the initial image, wherein the text request includes a comment to replace the first object in the initial image with a second object.

5. identifying an area in the background to replace or modify from the initial image; The method of claim 1 , further comprising providing a suggestion to replace or modify the background, wherein the text request is associated with the suggestion.

6. identifying an object in the background from the initial image to be removed; The method of claim 1 , further comprising: providing suggestions to remove the object from the background.

7. identifying one or more objects to replace from the initial image; The method of claim 1 , further comprising: providing suggestions to replace the object.

8. the text request is to change the background of the initial image; The method of claim 1 , wherein the preservation mask further includes one or more portions of the subject in addition to the face of the subject.

9. The method of claim 1 , wherein the text request further comprises at least one selection from the group of a global preset, a menu of options, a library of pre-made prompts, and combinations thereof.

10. A program causing one or more processors to carry out the method according to any one of claims 1 to 9.

11. 1. A system comprising: a processor; A system comprising: a memory coupled to the processor, the memory, when executed by the processor, causing the processor to perform the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Generating training data for machine learning

    DE102022202223A1

  • Appearance editing for participants during video conferences

    JP2015516625A

  • Image-to-Image Mapping by Iterative De-Noising

    US20230103638A1