Segmentation of objects in images

By preprocessing images to identify and segment objects using convolutional neural networks and diffusion models, the media application addresses delays and inaccuracies in conventional image editing, achieving accurate and efficient object manipulation.

JP2025525285AInactive Publication Date: 2025-08-05GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024565993
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-07
Filing Date
2024-05-09
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Conventional image editing software experiences delays and inaccuracies in object segmentation, leading to improper identification of object boundaries and unnatural object movement during editing.

Method used

A media application performs preprocessing on images to identify objects and segments them based on user likelihood, reducing processing time and improving segmentation accuracy by using convolutional neural networks and diffusion models to enhance or modify selected objects.

Benefits of technology

The solution reduces processing delays and enhances segmentation accuracy, ensuring background pixels are not misclassified with selected objects, resulting in natural object movement and improved editing quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025525285000001_ABST
    Figure 2025525285000001_ABST
Patent Text Reader

Abstract

The media application performs object recognition on the initial image to identify a set of objects in the initial image. The media application determines whether the initial image is an outdoor scene. In response to the initial image being an outdoor scene, the media application identifies a sky segment from the initial image. The media application determines whether the initial image includes an object that is a human or an animal. In response to the initial image including the object, the media application identifies an object segment from the initial image. The media application receives user input at a user interface that includes the initial image, corresponding to selecting a selected object from the set of objects. The media application updates the user interface to include an indication that the selected object has been selected.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 465,232, entitled "Selecting a Region of an Image," filed May 9, 2023; U.S. Provisional Patent Application No. 63 / 465,224, entitled "Relighting of Outdoor Images Using Machine Learning," filed May 9, 2023; U.S. Provisional Patent Application No. 63 / 465,226, entitled "Prompt-Drive Image Editing Using Machine Learning," filed May 9, 2023; U.S. Provisional Patent Application No. 63 / 465,230, entitled "Repositioning Objects in an Image," filed May 9, 2023; and U.S. Provisional Patent Application No. 63 / 562,634, entitled "Performing Scene Impact Editing Tasks Using Diffusion Neural Networks," filed March 7, 2024, each of which is incorporated herein in its entirety.

[0002] As image editing technology improves, it becomes increasingly important to intuitively translate user intent into object selection in a user interface.

[0003] The discussion of the background art provided herein is intended to provide a general background to the present disclosure. The inventor's work, to the extent described in this background art section, is not admitted, expressly or impliedly, as prior art to the present disclosure, and any portions of the description that may not qualify as prior art at the time of filing are likewise not admitted, expressly or impliedly, as prior art to the present disclosure. Summary of the Invention

[0004] The computer-implemented method includes performing object recognition on an initial image to identify a set of objects in the initial image. The method further includes determining whether the initial image is an outdoor scene. In response to the initial image being an outdoor scene, the method identifies a sky segment from the initial image. The method further includes determining whether the initial image includes an object, the object being a human or an animal. In response to the initial image including the object, the method identifies an object segment from the initial image. The method determines whether the initial image includes one or more obtrusive objects. In response to the initial image including one or more obtrusive objects, the method identifies one or more obtrusive segments from the initial image. The method further includes user input at a user interface including the initial image, the user input corresponding to selecting a selected object from the set of objects. The method further includes updating the user interface to include an indication that the selected object has been selected.

[0005] In some embodiments, the user input includes tapping the selection object multiple times, and the method further includes determining a number of taps from the user input and determining the selection object based on the number of taps, where a first tap is associated with a different region than a second tap. In some embodiments, the method further includes generating a background segment in response to the initial image including the object, where the object segment is associated with a foreground region and the background segment is associated with a background region, and pixels in the initial image are associated with the foreground region or the background region, and the method further includes determining that the user input corresponds to the foreground region based on the user input contacting pixels associated with the foreground region. In some embodiments, performing object recognition to identify objects in the initial image includes determining an object bounding box for each of the objects, and the method further includes determining that the user input corresponds to the selection object based on a proximity of the user input to the nearest object bounding box.

[0006] In some embodiments, a convolutional neural network (CNN) performs the segmentation, and the method further includes providing the initial image and a heat map of keypoints as inputs to the CNN and outputting, using the convolutional neural network, a segmentation mask corresponding to a sky segment, an object segment, and one or more obtrusive segments. In some embodiments, the user input includes selecting sky, and the method further includes receiving a request from the user to modify the illumination of the initial image, providing the initial image and the request to modify the illumination of the initial image as inputs to a diffusion model, and using the diffusion model to output an output image that meets the request. In some embodiments, the user input includes selecting one or more background objects for removal, and the method further includes removing the one or more obtrusive objects from the initial image based on object recognition and generating a modified image that includes inpainting pixels associated with the one or more obtrusive segments.

[0007] In some embodiments, the selected object is an incomplete object, where a missing portion of the incomplete object is clipped by a boundary of the initial image or obscured by another object, and the method includes generating a segmentation mask including the incomplete object, removing the incomplete object from the initial image, generating an inpainting image in which pixels of the incomplete object corresponding to the incomplete object are replaced with background pixels that match a background of the initial image, providing the segmentation mask, the incomplete object, and the inpainting image as inputs to a diffusion model, outputting the complete object using the diffusion model, and generating a modified image by blending one or more versions of the complete object with one or more versions of the inpainting image using the preservation mask. In some embodiments, the method further includes receiving a text request to modify the selected object in the initial image, identifying face segments for a face of the subject from the initial image based on the subject segments, generating a preservation mask corresponding to the face segments, providing the text request, the initial image, and the preservation mask as inputs to a diffusion model, and outputting an output image that satisfies the text request using the diffusion model.

[0008] A non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: performing object recognition on an initial image to identify a set of objects in the initial image; determining whether the initial image is an outdoor scene; responsive to the initial image being an outdoor scene, identifying a sky segment from the initial image; determining whether the initial image includes an object that is a human or an animal; responsive to the initial image including the object, identifying an object segment from the initial image; determining whether the initial image includes one or more offending objects; responsive to the initial image including the one or more offending objects, identifying one or more offending segments from the initial image; receiving, at a user interface including the initial image, user input corresponding to selecting a selected object from the set of objects; and updating the user interface to include an indication that the selected object was selected.

[0009] In some embodiments, the user input includes tapping the selection object multiple times, and the operations further include determining a number of taps from the user input and determining the selection object based on the number of taps, where a first tap is associated with a different region than a second tap. In some embodiments, the operations further include generating a background segment in response to the initial image including the object, where the object segment is associated with a foreground region and the background segment is associated with a background region, and pixels in the initial image are associated with the foreground region or the background region, and the operations further include determining that the user input corresponds to the foreground region based on the user input contacting pixels associated with the foreground region. In some embodiments, performing object recognition to identify objects in the initial image includes determining an object bounding box for each of the objects, and the operations further include determining that the user input corresponds to the selection object based on a proximity of the user input to a nearest object bounding box. In some embodiments, the CNN performs the segmentation, and the operations further include providing the initial image and a heat map of keypoints as inputs to the CNN, and outputting, using a convolutional neural network, a segmentation mask corresponding to a sky segment, an object segment, and one or more obtrusive segments. In some embodiments, the user input includes selecting a sky, and the operations further include receiving a request from a user to modify the lighting of the initial image, providing the initial image and the request to modify the lighting of the initial image as inputs to a diffusion model, and outputting, using the diffusion model, an output image that meets the request.

[0010] The system includes a processor and a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform operations including: performing object recognition on an initial image to identify a set of objects in the initial image; determining whether the initial image is an outdoor scene; responsive to the initial image being an outdoor scene, identifying a sky segment from the initial image; determining whether the initial image includes an object that is a human or an animal; responsive to the initial image including the object, identifying an object segment from the initial image; determining whether the initial image includes one or more offending objects; responsive to the initial image including the one or more offending objects, identifying one or more offending segments from the initial image; receiving, at a user interface including the initial image, user input corresponding to selecting a selected object from the set of objects; and updating the user interface to include an indication that the selected object has been selected.

[0011] In some embodiments, the user input includes tapping the selection object multiple times, and the operations further include determining a number of taps from the user input and determining the selection object based on the number of taps, where a first tap is associated with a different region than a second tap. In some embodiments, the operations further include generating a background segment in response to the initial image including the object, where the object segment is associated with a foreground region and the background segment is associated with a background region, and pixels in the initial image are associated with the foreground region or the background region, and the operations further include determining that the user input corresponds to the foreground region based on the user input contacting pixels associated with the foreground region. In some embodiments, performing object recognition to identify objects in the initial image includes determining an object bounding box for each of the objects, and the operations further include determining that the user input corresponds to the selection object based on a proximity of the user input to a nearest object bounding box. In some embodiments, the CNN performs the segmentation, and the operations further include providing the initial image and the heat map of keypoints as input to the CNN, and outputting, using a convolutional neural network, segmentation masks corresponding to sky segments, object segments, and one or more distracting segments. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram of an exemplary network environment according to some embodiments described herein. [Figure 2] FIG. 1 is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3] FIG. 1 is a block diagram of an example architecture of a trained tap-to-segment machine learning model according to some embodiments described herein. [Figure 4A]1A-1C illustrate exemplary user interfaces for selecting regions of an image, according to some embodiments described herein. [Figure 4B] 1A-1C illustrate exemplary user interfaces for selecting regions of an image, according to some embodiments described herein. [Figure 4C] 1A-1C illustrate exemplary user interfaces for selecting regions of an image, according to some embodiments described herein. [Figure 5A] FIG. 1 illustrates an exemplary initial image of a child sitting on a bench and holding a balloon (the balloon is partially cut off by the boundary of the initial image), according to some embodiments described herein. [Figure 5B] FIG. 10 illustrates an exemplary modified image in which the child, bench, and balloon have been moved to a second position, according to certain embodiments described herein. [Figure 6] 1 illustrates an exemplary user interface, which includes options for selecting different regions of an image to modify, a global preset to apply, a field for providing text, and an exemplary output image, in accordance with some embodiments described herein. [Figure 7] 1 is an exemplary flowchart illustrating a method of modification to an initial image according to some embodiments described herein. [Figure 8A] 1 is an exemplary flowchart illustrating a method for segmenting an initial image according to some embodiments described herein. [Figure 8B] 1 is an exemplary flowchart illustrating a method for segmenting an initial image according to some embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION

[0013] As image editing technology improves, it becomes increasingly important to intuitively translate user intent into object selection in a user interface. Conventional software applications for editing images perform segmentation of objects within an image after a user selects an object. This can result in unnecessary delays while the software application performs the segmentation. Furthermore, software applications may attempt to compensate for delays caused by performing the segmentation using a less accurate but faster process. Inaccurate segmentation can result in improper identification of object boundaries, resulting in pixels associated with the background or other objects being misclassified as belonging to the object. When editing involves moving an object, the object may appear unnatural in its new location if it contains pixels associated with the background from its previous location.

[0014] The media application performs preprocessing on the initial image before user interaction to identify a set of objects in the initial image. For example, the media application performs object recognition to identify subjects (e.g., people, dogs, children, etc.), trees, bystanders, sky, etc. The media application performs segmentation of different objects based on the likelihood of the object being selected by the user. For example, if the initial image is of an outdoor scene, the user may select the sky, change the sky color, remove clouds, etc.

[0015] The media application determines whether the initial image is an outdoor scene based on object recognition. In response to the initial image being an outdoor scene, the media application identifies a sky segment from the initial image, where pixels corresponding to the sky are identified as sky pixels. The media application determines whether the initial image includes an object, the object being a human or an animal, based on object recognition. In response to the initial image including the object, the media application identifies an object segment from the initial image, where pixels corresponding to the object are identified as object pixels. The media application determines whether the initial image includes one or more obtrusive objects. In response to the initial image including one or more obtrusive objects, the media application identifies one or more obtrusive segments from the initial image, where pixels corresponding to the sky are identified as obtrusive object pixels. In some embodiments, the obtrusive objects are identified based on being a type of object that is frequently removed from the initial image.

[0016] The media application receives, at a user interface including the initial image, user input corresponding to selecting a selected object from a set of objects identified based on performing object recognition. For example, the user can select a subject and provide a text request to add a hat to the subject, select a bystander and request that the bystander be removed from the image, select an incomplete object clipped by a boundary of the initial image, and move the incomplete object to a new location so that the media application can generate the complete object in the new location. The media application updates the user interface to include an indication that the selected object has been selected. For example, the indication may include a highlighted object, an outline around the selected object, etc.

[0017] By performing segmentation before receiving user input, the media application advantageously reduces the processing time that the user must wait for the segmentation to occur and improves the quality of the segmentation, resulting in an output image in which background pixels are not inappropriately associated with the selected object.

[0018] Exemplary Environment 100 FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, environment 100 includes a media server 101, a user device 115a, and a user device 115n coupled to a network 105. Users 125a, 125n may be associated with each user device 115a, 115n. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1. In FIG. 1 and other figures, a reference number followed by a letter, e.g., "115a," denotes a reference to the element with that particular reference number. A reference number in text without a following letter, e.g., "115," denotes a general reference to an embodiment of the element with that reference number.

[0019] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively coupled to a network 105 via signal line 102. The signal line 102 may be a wired connection, such as Ethernet, coaxial cable, or fiber optic cable, or a wireless connection, such as Wi-Fi, Bluetooth, or other wireless technology. In some embodiments, the media server 101 transmits data to and receives data from one or more of the user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0020] The database 199 may store machine learning models, training data sets, images, etc. The database 199 may also store social network data associated with the user 125, the user preferences of the user 125, etc.

[0021] The user device 115 may be a computing device that includes a memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device that can access the network 105.

[0022] In the illustrated embodiment, user device 115a is coupled to network 105 via signal line 108, and user device 115n is coupled to network 105 via signal line 110. Media application 103 may be stored on user device 115a as media application 103b and / or on user device 115n as media application 103c. Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections, such as Wi-Fi, Bluetooth, or other wireless technologies. User devices 115a and 115n are accessed by users 125a and 125n, respectively. In FIG. 1, user devices 115a and 115n are used as an example. While FIG. 1 shows two user devices, 115a and 115n, the present disclosure applies to system architectures having one or more user devices 115.

[0023] The media application 103 may be stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are performed on the media server 101 or the user device 115. In some embodiments, some operations may be performed on the media server 101 and some operations may be performed on the user device 115. The execution of the operations is subject to user settings. For example, the user 125a may specify that operations be performed on each user device 115a and not on the media server 101. With such settings, the operations described herein are performed entirely on the user device 115a and not on the media server 101. Furthermore, the user 125a may specify that user images and / or other data be stored locally only on the user device 115a and not on the media server 101. With such settings, user data is not sent to or stored on the media server 101. The transmission of user data to the media server 101, any temporary or permanent storage of such data by the media server 101, and the performance of operations on such data by the media server 101, occurs only if the user consents to the transmission, storage, and performance of operations by the media server 101. The user has the option to change settings at any time, for example, so that the user can enable or disable use of the media server 101.

[0024] Machine learning models (e.g., neural networks or other types of models) are stored locally on the user device 115 and used with specific user permission when utilized for one or more operations. Server-side models are used only with user permission. Additionally, trained models may be provided for use on the user device 115. In such use, on-device training of the model may be performed if the user 125 permits it. Updated model parameters may be sent to the media server 101 if the user 125 permits it, for example, to enable federated learning. The model parameters do not include user data.

[0025] The media application 103 performs object recognition on the initial image to identify a set of objects in the initial image. The media application 103 determines whether the initial image is an outdoor scene. In response to the initial image being an outdoor scene, the media application 103 identifies a sky segment from the initial image. The media application 103 determines whether the initial image includes an object that is a human or an animal. In response to the initial image including the object, the media application 103 identifies an object segment from the initial image. The media application 103 determines whether the initial image includes one or more offending objects. In response to the initial image including one or more offending objects, the media application 103 identifies one or more offending segments from the initial image.

[0026] The media application 103 receives user input corresponding to selecting a selected object from the set of objects in a user interface that includes the initial image. The media application 103 updates the user interface to include an indication that the selected object has been selected.

[0027] In some embodiments, media application 103 may be implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, media application 103a may be implemented using a combination of hardware and software.

[0028] Exemplary Computing Device 200 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. Computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, computing device 200 is media server 101 used to implement media application 103a. In another example, computing device 200 is user device 115.

[0029] In some embodiments, computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, a display 241, a camera 243, and a storage device 245, all coupled via bus 218. Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, display 241 may be coupled to bus 218 via signal line 228, camera 243 may be coupled to bus 218 via signal line 230, and storage device 245 may be coupled to bus 218 via signal line 232.

[0030] Processor 235 may be one or more processors and / or processing circuits that execute program code and control basic operations of computing device 200. A “processor” includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) with one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuitry to achieve a function, a special-purpose processor that performs neural network model-based processing, a neural circuit, a system with a processor optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, processor 235 may include one or more coprocessors that perform neural network processing. In some embodiments, processor 235 may be a processor that processes data to generate a probabilistic output; for example, the output generated by processor 235 may be inaccurate or accurate within a range from an expected output. Processing need not be limited to a particular geographic location, nor need it have time limitations. For example, a processor may perform its functions in real time, offline, in batch mode, etc. Portions of processing may be performed by different (or the same) processing systems at different times and in different locations. A computer may be any processor in communication with a memory.

[0031] Memory 237 is typically provided within computing device 200 for access by processor 235 and may be any suitable processor-readable storage medium suitable for storing instructions for execution by a processor or set of processors, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., located separately from and / or integrated with processor 235. Memory 237 may store software operated on computing device 200 by processor 235 and including media application 103.

[0032] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, etc. One or more methods disclosed herein may be implemented in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, or as a mobile application (“app”) that runs on a mobile computing device.

[0033] Application data 266 may be data generated by other applications 264 or by hardware of computing device 200. For example, application data 266 may include images used by an image library application, user actions identified by other applications 264 (e.g., social networking applications), etc.

[0034] I / O interface 239 may provide functionality that allows computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of computing device 200 or may be separate and in communication with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices may communicate through I / O interface 239. In some embodiments, I / O interface 239 may connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.).

[0035] Some examples of interface devices that may be connected to I / O interface 239 include display 241, which may be used to display content (e.g., images, video, and / or user interfaces of output applications described herein) and receive touch (or gesture) input from a user. For example, display 241 may be used to display a user interface, including a graphical guide, in a viewfinder. Display 241 may include any suitable display device, such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in a component of glasses or a headset device, or a monitor screen on a computing device.

[0036] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that I / O interface 239 sends to media application 103.

[0037] The storage device 245 stores data related to the media application 103. For example, the storage device 245 can store training data sets including labeled images, machine learning models, output from machine learning models, etc.

[0038] FIG. 2 illustrates an exemplary media application 103 including a segmenter 202, a user interface module 204, a repair module 206, and a diffusion module 208 stored in memory 237.

[0039] Segmentation is the process of labeling pixels in an initial image that are associated with a particular class. Segmentation can be used for a variety of reasons. For example, segmentation can be used to identify objects in an image that a user wants to remove, such as bystanders, power lines, or scooters. Segmentation can also be used to select objects that a user wants to enhance. For example, a user may want to change the background of an image or change the clothing of a subject in an image. Segmentation can also be used to identify regions of the initial image that should be preserved by generating a preservation mask that includes pixels associated with objects that will be prevented from being altered by blending with the synthetically generated image.

[0040] Once the pixels are labeled, the output of the segmentation is one or more segmentation masks. The one or more segmentation masks include pixels associated with segmented objects or regions in the initial image. The segmentation masks can be used as groups of pixels associated with the objects or regions, so that when the user interface receives user input, the user interface module 204 determines whether the user input corresponds to a particular segmentation mask based on the location of the user input. For example, the user interface module 204 can identify that the user input touched some pixels associated with a background segmentation mask. The segmentation masks can be used as retention masks to prevent modification of pixels associated with the retention mask while modifying pixels not associated with the retention mask. For example, during the process of generating an output image using a diffusion model, a retention mask can be used on a subject's face to prevent the face from being distorted during generation of the output image.

[0041] The segmenter 202 receives an initial image. The initial image may be captured by a camera 243 associated with the computing device 200, received from another application 264, or the like. The segmenter 202 performs object recognition on the initial image to identify a set of objects in the initial image. The object recognition may be performed by a machine learning model or other algorithm. In some embodiments, the segmenter 202 determines an object bounding box for each of the objects in the set of objects. The object bounding box may include pixels associated with a particular object and may be associated with metadata describing the object bounding box (e.g., (x, y) coordinates describing the edges of the object bounding box).

[0042] The segmenter 202 performs segmentation of the initial image. For example, the segmenter 202 identifies pixels associated with a subset of a set of objects in the initial image based on object recognition and a likelihood that the subset of objects will be selected by a user. The likelihood that the subset of objects will be selected by a user may be based on anonymous information about what people select in an image. For example, a user is most likely to select the subject of an image, more likely to select an obtrusive object in an image, and least likely to select an aesthetic background object, such as a tree, a building in a cityscape, or a boat on the water.

[0043] In some embodiments, segmenter 202 determines whether the initial image has a particular type of object and performs segmentation in response to the initial image including the particular type of object. For example, segmenter 202 determines whether the initial image is an outdoor scene based on object recognition that identifies the presence of sky. Outdoor scenes are characterized by images that include sky. In some embodiments, segmenter 202 determines that the initial image is an outdoor scene based on the initial image including particular colors associated with outdoor scenes and / or based on the initial image including particular colors located in areas where sky is expected. Outdoor scenes may include additional objects, such as buildings, trees, beaches, etc. If the initial image is an outdoor scene, segmenter 202 identifies a sky segment for the initial image.

[0044] The segmenter 202 determines whether the initial image includes a human or animal object based on object recognition that identifies objects associated with human and / or animal categories. For example, the object may be a cat, a chicken, a human, etc. If the initial image includes a human or animal object, the segmenter 202 identifies an object segment from the initial image.

[0045] The segmenter 202 determines whether the initial image includes one or more obtrusive objects. Obtrusive objects can be based on types of objects that are frequently removed from the initial image (e.g., people, cars, power lines, etc. that are not subjects of the initial image). Conversely, the segmenter 202 may not segment objects such as trees because trees are not frequently removed from the initial image. In some embodiments, classifying objects as obtrusive objects is based on ranking the types of objects to be removed from the initial image using a cutoff value (e.g., classifying the top 20 most frequently removed objects as obtrusive object types, likelihood of an object type being removed from the initial image exceeding a threshold likelihood, etc.). If the initial image includes one or more obtrusive objects, the segmenter identifies one or more obtrusive segments from the initial image.

[0046] In some embodiments, the segmentation further includes foreground / background segmentation, sky segmentation, and / or panoramic segmentation (e.g., segmenting an image into semantically significant parts or regions). The foreground / background segmentation may be used by the media application 103 to perform selective tone mapping. Tone mapping is used to modify the tonal values of pixels. Tone mapping may be used to adjust the tonal values of an initial image with a high dynamic range for applications such as display on a digital display.

[0047] The segmenter 202 can use different approaches to segment a subset of objects in an image. In some embodiments, the segmenter 202 segments objects into regions. In some embodiments, the segmenter 202 divides the image into foreground and background and segments objects based on whether they are in the foreground or background.

[0048] In some embodiments, segmenter 202 generates different types of segmentation masks for segmentation performed on an image. For example, segmenter 202 may generate a subject mask that preserves the subject's face or includes more of the subject, such as the subject's entire head, hands, body, etc. In some embodiments, the segmentation mask is generated based on generating superpixels of the image and matching the centroids of the superpixels to depth map values for depth-based cluster detection (e.g., values obtained by camera 243 using a depth sensor or by deriving depth from pixel values). More specifically, the depth values of the masked region can be used to determine a depth range, and superpixels that fall within the depth range can be identified.

[0049] Another technique for generating a segmentation mask involves weighting depth values based on how close they are to the mask, where the weights are represented in a distance transform map.

[0050] In some embodiments, segmenter 202 can specify circuitry (e.g., for a programmable processor, for a field programmable gate array (FPGA), etc.) that enables application of the machine learning model by processor 235. In some embodiments, segmenter 202 can include software instructions, hardware instructions, or a combination thereof. In some embodiments, segmenter 202 can provide an application programming interface (API) that can be used by operating system 262 and / or other applications 264 to invoke segmenter 202, for example, to apply the machine learning model to application data 266 and output a segmentation mask.

[0051] The segmenter 202 uses training data to generate a trained machine learning model. In some embodiments, the training data includes images (e.g., red, green, blue (RGB) images) and a heat map of keypoints in the images. Keypoints are feature or salient points in the initial image that are used to identify, describe, or match objects or features in a scene. For example, keypoints may be determined using a scale-invariant feature transform (SIFT). In some embodiments, the training data further includes a corresponding segmentation mask.

[0052] The training data may be obtained from any source (e.g., a data repository marked for training purposes, data with permission to be used as training data for machine learning, etc.) In some embodiments, training may occur on the media server 101 providing the training data directly to the user device 115, may occur locally on the user device 115, or may occur as a combination of both.

[0053] In some embodiments, segmenter 202 uses unedited / untransferred weights obtained from another application. For example, in these embodiments, a trained model may be generated, e.g., on a different device, and provided as part of segmenter 202. In various embodiments, the trained model may be provided as a data file that includes a model structure or format (e.g., a model structure or format that defines the number and type of neural network nodes, the connectivity between the nodes, and the organization of the nodes into multiple layers) and associated weights. Segmenter 202 may read the trained model data file and implement a neural network with node connectivity, layers, and weights based on the model structure or format specified in the trained model.

[0054] A trained machine learning model may include one or more model forms or structures, such as any type of neural network, e.g., a linear network, a deep learning neural network that implements multiple layers (e.g., "hidden layers" between the input and output layers, where each layer is a linear network), a convolutional neural network (CNN) (e.g., a network that divides or partitions input data into multiple portions or tiles, processes each tile separately using one or more neural network layers, and assembles the results from processing each tile), a sequence-to-sequence neural network (e.g., a network that receives as input sequential data, such as words in a sentence or frames in a video, and produces as output a sequence of results), etc.

[0055] The model format or structure may specify the connectivity between various nodes and their organization into layers. For example, nodes in a first layer (e.g., an input layer) may receive data as input or application data. Such data may include, for example, one or more pixels per node, for example, when the trained model is used to analyze an initial image. Subsequent intermediate layers may receive as input the output of nodes in the previous layer according to the connectivity specified in the model format or structure. These layers may also be referred to as hidden layers. For example, the first layer may output a segmentation between foreground and background. The final layer (e.g., an output layer) generates the output of the machine learning model. For example, the output layer may receive the segmentation of the initial image into foreground and background and output whether a pixel is part of the segmentation mask. In some embodiments, the model format or structure further specifies the number and / or type of nodes in each layer.

[0056] 3 is a block diagram of an example architecture 300 of a trained tap-to-segment machine learning model according to some embodiments described herein. The example architecture includes a CNN that receives input and generates output. The CNN includes a convolutional layer that applies filters to the input data to extract features. The convolutional layer may be followed by a pooling layer that reduces spatial dimensionality and improves computational efficiency.

[0057] The CNN includes a CNN encoder 315 and a CNN decoder 320. The encoder receives an image and encodes the image into a vector or matrix representation of the image. The CNN encoder 315 receives an RGB image 305 and a corresponding heat map of keypoints 310. The RGB image 305 is an image containing pixels that include one of three color channels (red, green, and blue). The keypoints 310 include locations within the initial image where the user touches. The keypoints 310 may be defined as locations where the user input exceeds a user input threshold.

[0058] The CNN encodes the RGB image 305 into increasingly abstract information, with each convolutional layer representing a different level of abstraction. The CNN decoder 320 decodes the abstract information and outputs a segmentation mask 325. The segmentation mask 325 identifies pixels associated with one or more objects in the RGB image 305. For example, the RGB image 305 may be an image of a coffee mug on a table, and the heat map of keypoints 310 has a keypoint in the center of the coffee mug to indicate that a user typically selects the coffee mug and nothing else in the image. Because it is likely that a user will tap the coffee mug and not other objects in the image, the CNN decoder 320 outputs a segmentation mask that segments the coffee mug from the rest of the image.

[0059] In different embodiments, the trained model may include one or more models. One or more of the models may include multiple nodes arranged in layers according to a model structure or model format. In some embodiments, a node may be, for example, a memoryless computational node configured to process a unit of input and generate a unit of output. The computation performed by the node may include, for example, multiplying each of multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias value or intercept value to generate the node output. In some embodiments, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, computation by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, using individual processing units of a graphics processing unit (GPU), or dedicated neural circuitry. In some embodiments, a node may include memory, for example, capable of storing and using one or more previous inputs when processing a subsequent input. For example, a node with memory may include a long short-term memory (LSTM) node, which may use memory to maintain a "state" that allows the node to operate like a finite state machine (FSM).

[0060] In some embodiments, the trained model may include embedding values or weights for individual nodes. For example, the model may begin as a plurality of nodes organized into layers, as specified by the model format or model structure. At initialization, a respective weight may be applied to the connection between each pair of nodes (e.g., nodes in successive layers of a neural network) connected according to the model format. For example, each weight may be randomly assigned or initialized to a default value. The model may then be trained to generate results, for example, using training data.

[0061] The training may include applying a supervised learning method. In supervised learning, the training data may include multiple inputs (e.g., images, segmentation masks, etc.) and a corresponding ground truth output for each input (e.g., a ground truth mask that accurately identifies a portion of a subject, such as the subject's face, in each image). Based on a comparison of the model's output and the ground truth output, the values of the weights are automatically adjusted, for example, to increase the probability that the model generates the ground truth output for the image.

[0062] In various embodiments, the trained model includes a set of embedding values or weights corresponding to the model structure. In some embodiments, the trained model may include a set of weights that are fixed (e.g., downloaded from a server that provides the weights). In various embodiments, the trained model includes a set of embedding values or weights corresponding to the model structure. In embodiments where data is missing, segmenter 202 may generate a trained model based on pre-training, for example, by a developer of segmenter 202, a third party, etc. In some embodiments, the trained model may include a set of weights that are fixed (e.g., downloaded from a server that provides the weights).

[0063] In some embodiments, the trained machine learning model receives an initial image having objects identified by object recognition. In some embodiments, the trained machine learning model outputs one or more segmentation masks corresponding to one or more of the objects. For example, the trained machine learning model outputs segmentation masks for the sky, the subject, and one or more obtrusive objects. In another example, the trained machine learning model outputs segmentation masks for the background and the foreground.

[0064] The user interface module 204 generates graphical data for displaying a user interface including the image. The user interface displays various options for associating user input with a corresponding region in the image. Figures 4A-4C show exemplary user interfaces for selecting a region of an image, according to some embodiments described herein.

[0065] FIG. 4A includes a first user interface 400 instructing a user to circle any object the user wants to select, according to some embodiments described herein. This may be referred to as stroke selection. In this example, the user is circling a subject in an image with a circle 402. In a second user interface 405, the user is instructed to tap one of the circles to select an object. For example, the user may select circle 406 to select the sky, circle 407 to select a tree, and circle 408 to select a user. In a third user interface 410, the user is instructed to select one of the following regions / objects from a list 412: sky, person, car, sign, background, and clothing. "Person" in list 412 is highlighted as an indication that the user has selected a person.

[0066] 4B includes a fourth user interface 415, in which a single circle is associated with multiple regions / objects, according to some embodiments described herein. In this example, the circle 416 may be selected a first time to select the sky, and the circle 416 may be selected a second time to select the background. In response to the user selecting the circle 416 once, the user interface module 204 may update the user interface to display a segment mask indicating pixels associated with the sky segment. In response to the user selecting the circle 416 a second time, the user interface module 204 may update the user interface to display a segment mask indicating pixels associated with the background segment.

[0067] In the fifth user interface 420, the segmenter 202 has segmented the image into foreground and background segments. The person is in the foreground and everything else is in the background. As a result, selecting any area within the foreground region results in a selection of the person, as indicated by the exemplary foreground arrow 422. Selecting any area within the background region results in a selection of the background, as indicated by the background arrow 424.

[0068] In the sixth user interface 423, the user has selected a pixel corresponding to the foreground segment in the fifth user interface 420, and the user interface module 204 has updated the user interface to include a representation that is a segmentation mask 425 associated with the foreground segment.

[0069] FIG. 4C includes a seventh user interface 430. In the seventh user interface 430, a user can tap an object to select the corresponding object. Objects are associated with object bounding boxes. When a user taps within a bounding box, the object is selected. For example, tapping within bounding box 426 selects a car. Tapping within bounding box 427 selects a stop sign. This scenario can lead to confusion when a user taps a section that is within two bounding boxes, such as area 429 where bounding box 427 and bounding box 428 overlap. In some embodiments, if the selection is ambiguous, the user interface may display text asking the user for confirmation as to which object the user intended to select, or the user interface may update the display to provide an indication of which object the user more likely intended to select, allowing the user to change it if they disagree.

[0070] Once the user interface module 204 determines which object / area the user input corresponds to, the user interface module 204 generates graphical data to display an indicator that the object has been selected. For example, the user interface may add an outline around the selected object, highlight the selected object, etc.

[0071] 5A shows an exemplary initial image 500 of a child 505 sitting on a bench 510 and holding a balloon 515 (the balloon 515 is partially cut off by the boundary of the initial image 500), according to some embodiments described herein. In this example, the user interface module 204 provides a user interface with options for the user to select objects segmented by the segmenter 202. The user selects the child 505, the bench 510, and the balloon 515 in a first position. Here, the balloon 515 is in an incomplete image.

[0072] The user interface module 204 includes an option for moving the selected object to another location. The user selects a second location. The segmenter 202 removes the selected object from the initial image. The inpainting module 206 generates an inpainting image. The inpainting image replaces object pixels corresponding to the removed object with background pixels that match the background of the initial image.

[0073] The diffusion module 208 receives the coordinates of the selected object and the second location as input and outputs the complete object: the balloon and the longer bench. Figure 5B shows an example retouched image 550 in which the child 555, bench 560, and balloon 565 have been moved to a second location, according to some embodiments described herein. In this example, the diffusion module 208 outputs the retouched image. The retouched image uses a segmentation mask to blend one or more versions of the child 555, bench 560, and balloon 565 with one or more versions from the inpainting image.

[0074] 6 illustrates exemplary user interfaces 600, 625, 650. The exemplary user interfaces 600, 625, 650 include options for selecting different regions of an image to modify, a global preset to apply, a field for providing text, and an exemplary output image, according to some embodiments described herein. Specifically, the first user interface 600 automatically provides a global preset 605. Through the global preset 605, a user selects to modify an input image 601 to look like an oil painting, a surreal world, or a nostalgic scene.

[0075] The first user interface 600 also includes circles 610, 611, 612 that represent the identification of different regions within the initial image 601. The user can specify changes to be made to the sky by tapping the first circle 610, to the bridge by tapping the second circle 611, and to the people by tapping the third circle 612.

[0076] In response to a user selecting one of the circles 610, 611, 612, the user interface may update the display to provide a menu of options (not shown). For example, selecting the first circle 610 may cause the user interface to display suggestions such as changing a cloudy sky to a sunny sky. Selecting the second circle 611 may cause the user interface to display suggestions such as removing the bridge associated with the second circle 611, replacing the bridge with a different type of bridge or a boat, etc. Selecting the third circle 612 may cause the user interface to display suggestions to remove the person.

[0077] The second user interface 625 includes an input image 626 and a text entry field 630. In the text entry field 630, the user can specify the changes they want to make. The user can include a description that is specific enough to cover the objects they want to change (e.g., change boots to colorful, shiny boots), or the user can select the objects they want to change in the second user interface 625 and then describe the specific changes to be made. For example, the user may select the object by tapping on it, circling it, scribbling on it, etc. In this case, the user selects the subject boots 627.

[0078] A third user interface 650 includes an output image 651 in which a text request 652 for "colorful, shiny boots" has been fulfilled. The boots 653 have been modified to resemble sparkly, colorful stars.

[0079] In situations where an object is removed from the initial image, the inpainting module 206 generates an inpainting image. The inpainting image replaces object pixels corresponding to one or more objects with background pixels. The background pixels can be based on pixels from a reference image at the same location without the object. Alternatively, the inpainting module 206 may identify background pixels to replace the removed object based on the proximity of the background pixels to other pixels surrounding the object. The inpainting module 206 can use a gradient of nearby pixels to determine the properties of the background pixels. For example, if a bystander was standing on the ground, the inpainting module 206 replaces the background pixels with ground pixels. Other inpainting methods are possible, including machine learning-based inpainting methods that output background pixels based on training data containing images with similar configurations.

[0080] In embodiments where the user chooses to erase the selected object, the user interface module 204 may display a repair image in which the selected object has been removed and the pixels of the selected object have been replaced with background pixels.

[0081] The diffusion model includes a forward process in which the diffusion model adds noise to the data and a backward process in which the diffusion model learns to recover the data from the noise. For example, when a selected object is moved from a first position to a second position, the diffusion module 208 applies the diffusion model by blending the selected object with progressively noisier versions of the inpainting image and then with progressively denoised versions of the inpainting image. In some embodiments, an object stitching diffusion model is used to move the object from the first position to the second position. In some embodiments, a generative diffusion model is used when the object is an incomplete object and portions of the object are generated and / or for new objects generated from text prompts.

[0082] Object Stitching Diffusion Model In some embodiments, an object stitching diffusion model is used when moving an object from a first position to a second position. In some embodiments, the diffusion module 208 includes an object image encoder that extracts semantic features from a selected object, a diffusion model that blends the object with the image, and a content adapter that converts a sequence of visual tokens into a sequence of text tokens to overcome the domain gap between images and text. In some embodiments, the diffusion module 208 trains the diffusion model using self-supervised learning based on training data, where the training data includes image and text pairs. In some embodiments, the diffusion model is trained with synthetic data that simulates real-world scenarios. The diffusion model may also be trained using data augmentation, which is generated by introducing random shifts and crop augmentations during training while ensuring that foreground objects are contained within a crop window.

[0083] During the first stage, the content adapter is trained to preserve the high-level semantics of the object using image and text pairs, and during the second stage, the content adapter is trained in the context of a diffusion model to encode important discriminative features of the object by facilitating the visual reconstruction of the object in the original image. The diffusion model may be trained with the embedding values generated by the content adapter via the cross-attention block.

[0084] The diffusion model blends the inpainted image with the object using a preservation mask. The diffusion model may remove noise from the masked region. The content adapter may convert visual features from the object image encoder into text features (tokens) for use as conditioning for the diffusion model.

[0085] Generative Diffusion Model In some embodiments, a generative diffusion model is used to output a complete object based on an incomplete object or to generate a new object. The diffusion module 208 trains the generative diffusion model based on training data. The training data may include image and text pairs. The image and text pairs are used to create an embedding space for the image and text. The image and text pair may include an image and its associated text (e.g., an image of a dog and text containing "pitbull"). The diffusion module 208 may be trained with a loss that reflects the cosine distance between the embedding value of the text prompt and the embedding value of the estimated clean image (i.e., without the text-generating object).

[0086] The diffusion module 208 can use the training data to perform text conditioning, where text conditioning describes the process of outputting an object conditioned by a text prompt. The diffusion module 208 can train a neural network to output an object based on a text prompt provided by a user or by a media application. For example, the text prompt can be a suggestion generated by the media application based on the context of an initial image (e.g., if the initial image is a beach, the text prompt can be about a beach ball, a turtle, etc.).

[0087] In some embodiments, the diffusion module 208 can output at least a portion of the missing portion of the object based on receiving as input data an incomplete object, a location in the image to which the incomplete object is moved, and output dimensions (e.g., original or modified dimensions if the object is resized). For example, if a user selects an object in a user interface that is partially cut off by a boundary and moves the object from a first position to a second position, where the second position also cuts off a portion of the object, the diffusion module 208 can output a modified object that includes more of the object that became visible based on moving the object in the image. In some embodiments, the diffusion module 208 can output a complete object based on an incomplete object selected by a user. For example, if a user selects a beach ball that is partially obscured by another object, a diffusion model can be trained to output the complete beach ball.

[0088] In some embodiments, the diffusion module 208 generates progressively noisier versions of the complete object (compared to previous versions) and progressively noisier versions of the inpainted image (compared to previous versions). For example, a forward Markov noise process generates a series of noisy inpainted images by gradually adding Gaussian noise until approximately isotropic Gaussian noise samples are obtained. The forward noise process defines a progression of image manifolds, where each manifold consists of noise images.

[0089] The diffusion module 208 can use a preservation mask to spatially blend a noisy version of the complete object with a corresponding noisy version of the inpainted image. For example, the diffusion module 208 can use a preservation mask to blend each noisy version of the complete object with each corresponding noisy version of the inpainted image, where the preservation mask defines the boundary of the complete object and, therefore, the preservation mask defines the area to be modified during the blending process. In some embodiments, the diffusion process can include local complete object guided diffusion, where the image generation loss determined during the training process is used under the preservation mask during local object generation diffusion.

[0090] The diffusion module 208 can perform a diffusion step to denoise the latent space in a direction dependent on the text prompt. The diffusion module 208 generates an incrementally denoised version of the complete object (compared to the previous version) and an incrementally denoised version of the inpainted image (compared to the previous version). For example, an inverse Markov process transforms Gaussian noise samples by iteratively denoising the inpainted image using the learned posterior distribution. Each step of the denoising diffusion process projects the noisy image onto the next less noisy manifold.

[0091] The diffusion module 208 performs a denoising diffusion step after each blend to restore consistency by projecting onto the next manifold. Once spatial blending is complete, the diffusion module 208 preserves the background by replacing areas outside the preservation mask with corresponding areas from the inpainted image.

[0092] In some embodiments, the diffusion module 208 applies an iterative refinement scheme using cross-domain compositing to inject contextual information into the object to match the style of the inpainted image. For example, if an object is generated for an indoor setting and added to an outdoor inpainted image, the object may be modified to be brighter to match the inpainted image. In another example, if an object is in a first position in shadow and a second position in sunlight, the object may be modified to match the brightness of the second position.

[0093] Object Removal Model In some embodiments, instead of using the segmenter 202 to remove objects and the inpainting module 206 to add pixels to the removed regions in the initial image, a diffusion model is trained to include an object removal model.

[0094] The diffusion module 208 generates counterfactual training data to train the diffusion model, including the object removal model. For each counterfactual image pair, the diffusion module 208 captures a real image containing the object in the scene, physically removes the object while avoiding camera movement, lighting changes, or other object movement, captures a counterfactual image of the scene without the object, and segments the real image to create a storage mask. Segmenting the real image generates a segmentation map (M) of the object O removed from the real image X. o )

[0095] For each image pair, the diffusion module 208 creates a composite image that includes the real image, a preservation mask, and a counterfactual image. The preservation mask is a binary preservation mask (M o (X)), and the counterfactual image pair is the input pair (X,M o (X)) and the output counterfactual image (X cf )

[0096] Given a real image x and a binary preserving mask, the diffusion module 208 trains a diffusion model based on the use of counterfactual image pairs to generate a distribution P(X cf |X=x,M o (X)). The diffusion module 208 determines the estimate by minimizing the loss function ζ(θ) using the following equation:

[0097]

number

[0098]

number

[0099]

number

[0100] where x represents the image without the object (counterfactual), and α t and σ t is determined by the noise schedule, ε~N(O,I).

[0101] Once the diffusion model, including the object removal model, has been trained, the user interface module 204 can receive a request to remove a selected object from a first modified image. The initial image and the request are provided as inputs to the object removal model, which outputs a modified image that does not include the selected object.

[0102] Object Insertion Model In some embodiments, instead of removing the object using the segmenter 202, the inpainting module 206 is used to add pixels to the removed region of the initial image, and the diffusion model is used to blend the object with the pixels in its new location, including the object insertion model, so that the diffusion model is trained.

[0103] In some embodiments, the object insertion model is trained on a number of image pairs that exceeds the number of available counterfactual image pairs. As a result, the diffusion module 208 generates synthetic training data. For each synthetic image pair, the diffusion module 208 selects an original image that contains the object, uses the object removal model to output a modified image from the original image that is free of the object, generates an input image by inserting the object into the modified image, and segments the original image to create a conservation mask. The modified image that lacks the object is denoted by z using the following formula: i Let's say.

[0104] z i ~P(X cf |x i ,M o (x i )) Equation 3 where the original images are x1, x2, …, x n and the corresponding storage mask is M o (x1), M o (x2),…,M o (x n ) The diffusion module 208 calculates the diffusion coefficient for an object-free scene z using the following formula: i The input image is generated by inserting objects into the resulting image without shadows and reflections.

[0105]

number

[0106] The composite image pair is (y i ,M o (x i)) and the corresponding target is the original image x i While both the input image and the output image contain the object o, the input image does not contain the effect of the object on the scene, while the output image does. In some embodiments, the diffusion module 208 trains the object insertion model using the diffusion object presented in Equation 1.

[0107] For each composite image pair, the diffusion module 208 creates a second composite image that includes the original image, the preservation mask, and the input image. The diffusion module 208 pre-trains the diffusion model to include an object insertion model based on using the composite image pair, and fine-tunes the diffusion model to include the object insertion model based on using the counterfactual image pair used to train the object removal model.

[0108] In some embodiments, the user interface module 204 generates graphics data for displaying a user interface that provides the user with options for specifying the object's location and resizing the object. The diffusion module 208 adds the selected object removed from the initial image to a new location. In some embodiments, the diffusion module 208 provides the selected object as input to a diffusion model, further providing the location where the selected object will be placed in the retouched image, and outputs the retouched image blending the selected object with the inpainted image. For example, the diffusion module 208 can spatially blend a noisy version of the inpainted image with a noisy version of the selected object.

[0109] In some embodiments, the diffusion module 208 may add a shadow to the selected object at the new position. The shadow may match the direction of the light in the image. For example, if the sun casts light from the upper left corner of the image, the shadow may appear to the right of the person and / or object. In some embodiments, the diffusion module 208 uses a machine learning model to output a shadow mask that is used to generate the shadow that falls on the object.

[0110] A diffusion model for text requests. In some embodiments, a user can select an object or region and provide a request to modify the selected object or region. For example, a user can select a subject to change the subject's clothing, or select a sky to change the sky's lighting. The diffusion model receives as input a request (e.g., a text request provided directly by the user, a selection of a pre-made prompt, a selection of a global preset, a selection of an option from a menu, etc.), an initial image, and a save mask. The diffusion model encodes the image in latent space, performs diffusion, and decodes back to pixel space.

[0111] The diffusion module 208 performs text conditioning of the request. Text conditioning describes the process of generating an image that is conditioned on (e.g., aligned with) a text prompt. For example, if the text request is to replace the red shirt worn by the subject in the initial image with a blue shirt, the diffusion module 208 performs the text conditioning by generating an output image of the blue shirt.

[0112] In some embodiments, the diffusion module 208 trains the diffusion model using two types of training data: The first type of training data includes pairs of images, which may include synthetic pairs generated via a prompt-to-prompt generation machine learning model. The prompt-to-prompt generation machine learning model is a diffusion model that receives a text prompt, uses self-attention to extract keys and values from the text prompt, swaps a portion of the attention map previously generated for the input image based on the input text prompt, and outputs an output image to match the text prompt.

[0113] The prompt-to-prompt generation machine learning model generates a self-attention map. Self-attention calculates the interactions between different elements of an input sequence (e.g., different words in a text request). This contrasts with cross-attention, where the interactions are between two different input sequences (e.g., how the text request relates to the original prompt).

[0114] Self-attention maps describe the structure and different semantic regions within an image. For example, an image described in a self-attention map as "pepperoni pizza next to orange juice" incorporates how certain pixels on the pizza's crust attend to other pixels on the crust. Conversely, in a cross-attention map, pixels on the pizza's crust attend to the orange juice.

[0115] The self-attention map is used in the text-conditional diffusion model to modify one or more token values using structure and different semantic regions in the input image, but the self-attention map is fixed to preserve the scene's configuration. In some embodiments, the diffusion model adds a new word to the prompt and fixes attention on the previous token, allowing new attention to flow to the new token. This results in a global edit or modification of specific objects in the input image to match the text request.

[0116] Each diffusion step predicts noise from the noisy image and text embeddings. In the final step, the process results in a generated image. The interaction between the text prompt and the image occurs during noise prediction, where visual and text feature embeddings are fused using a self-attention layer to generate a spatial attention map for each text token.

[0117] The second type of training data includes pairs of real and synthetic images. The real images are received by a diffusion model, such as a denoising diffusion implicit model (DDIM). The diffusion model uses an inverse method to output a synthetic image based on the real image and instructions on how to edit the input image. The diffusion module 208 trains the diffusion model to generate output images from requests using a forward process, in which the diffusion model adds noise to the data, and a backward process, in which the diffusion model learns to recover the data from the noise.

[0118] The diffusion module 208 trains the diffusion model to maintain photorealism and preserve the identity of objects shown in the image. During training, the diffusion model receives editing instructions and modifies the editing instructions to create corresponding prompts based on a language model, such as a large-scale language model. For example, the diffusion module 208 uses a language model to convert the editing instruction "Make the person look like an astronaut" into prompts that describe various aspects of what a space suit looks like.

[0119] The diffusion model creates a set of input and output image pairs from the generated prompt pairs, where each prompt can generate N images (using a different seed). The diffusion module 208 filters certain images from the image pairs (e.g., image transformations that do not match the given editing instructions, image transformations that do not produce sufficiently aligned images, and mismatched pairs). In some embodiments, the diffusion module 208 also filters images based on an edit alignment score and an image-text alignment score. The edit alignment score reflects the alignment between the image-to-image transformation and the original edited caption. The image-text alignment score reflects the alignment between the input / output image and the corresponding input / output prompt. In some embodiments, the diffusion module 208 trains the diffusion model by generating one or more loss functions based on the filtered images from the image pairs.

[0120] The diffusion model is trained to generate images by gradually adding noise to the image. The diffusion model then learns how to gradually remove the noise. The diffusion model applies the denoising process to random seeds to generate realistic images. By simulating diffusion, the diffusion model generates one or more noisy images.

[0121] Once the diffusion model is trained, it receives an input image and performs a de-diffusion process on the initial image to generate a noisy image based on the initial image. In some embodiments, the diffusion module 208 performs de-diffusion using DDIM inversion.

[0122] The diffusion model provides a noisy image to a first CNN equipped with features and a self-attention mechanism. The first CNN samples the input image and extracts features from the input image. The first CNN directly injects the extracted features and self-attention map into a second CNN. The first CNN performs forward diffusion of the noisy initial image, which is a process of gradually denoising the noisy image using sampling to output a denoised initial image.

[0123] The text request and the noisy image are provided as input to a second CNN, which uses a self-attention map to align semantic features of the text request with the structure of the noisy image to generate a noisy transformed image, and performs forward diffusion on the noisy transformed image to output a denoised transformed image.

[0124] The denoised initial image is combined with the denoised transformed image and the storage mask. This advantageously prevents modification of the face, which may otherwise be modified in ways that result in unrealistic features. In some embodiments, the diffusion module 208 performs the blending using a mask smoothing algorithm and Poisson blending.

[0125] In some embodiments, the preserved mask includes other parts of the subject (e.g., the subject's hair if the user wants the hair to remain the same, the subject's fingers (because fingers are often unrealistically modified by machine learning models), the entire subject if the subject is a pet that should be protected from excessive modification, etc.). In some embodiments in which the output image modifies the subject's clothing, the preserved mask may include everything except the subject's clothing, thereby preserving the torso (excluding the clothing) and background of the initial image.

[0126] The combined denoised image and preservation mask are blended with the denoised transformed image to form an output image that meets the text requirements.

[0127] Exemplary Flowchart The media application 103 may include different methods for editing the initial image. Figure 7 shows an example flowchart of a method 700 of making modifications to an initial image according to some embodiments described herein. The method 700 may be performed by the computing device 200 of Figure 2. In some embodiments, the method 700 is performed by the user device 115, by the media server 101, or partially on the user device 115 and partially on the media server 101.

[0128] 7 may begin at block 705. At block 705, it is determined whether the user grants permission to access the initial image. If the user does not grant permission, the method 700 ends. If the user grants permission, block 705 may be followed by block 710.

[0129] Block 710 receives a request to modify the entire initial image, a request to modify a portion of the initial image, or a text request. Modifications to the entire initial image may include, for example, a request to change the style of the initial image to resemble an impressionist painting. Modifications to a portion of the initial image may include, for example, a request to move an object from one location to another, a request to remove power lines, etc. Modifications, including text requests, may be directed to a specific object in the image (e.g., a request to replace a subject's shirt with a jacket), creating a new object (e.g., a request to add a turtle to an initial image of a beach), or a change to the entire image (e.g., a request to change an outdoor scene from a daytime image to a nighttime image). Block 710 may be followed by block 715 for modifying the entire image, block 720 for modifying a portion of the image, or block 725 for a text request.

[0130] At block 715, in response to the request being to modify the entire image, a selection of a preset is received. The preset may include changing an outdoor scene to an evening, night, or cloudy scene, etc., changing the initial image to an oil painting, a surreal image, a nostalgic image, etc., and changing the theme to a sea adventurer, an ancient warrior, a space crusader, a wise wizard, an aristocrat, a space exploration mission, etc. Block 715 may be followed by block 730.

[0131] At block 720, a selection of a region is received in response to modifying a portion of the image. The region may include a group of objects (e.g., sky with clouds) or a single object. The region may be selected by clicking a circle in a user interface, encircling the region, tapping the region until the desired region is highlighted with an indicator, etc. Block 720 may be followed by block 730.

[0132] In block 725, an open text prompt is used in response to a request to modify using a text prompt. Block 725 may be followed by block 730.

[0133] A corrected image is generated at block 730. Block 730 may be followed by block 735.

[0134] Block 735 determines whether the user is satisfied with the modified image. If the user is not satisfied with the modified image, block 735 may be followed by block 740.

[0135] The modified image is modified or refreshed in response to the user providing further user input at block 740. The cycle from block 735 to block 740 is repeated until the user is satisfied with the modified image, at which point block 735 may be followed by block 745.

[0136] At block 745, the modified image is saved. 8A-8B show an example flowchart of a method 800 for segmenting an initial image according to some embodiments described herein. Method 800 may be performed by computing device 200 of FIG. 2. In some embodiments, method 800 is performed by user device 115, by media server 101, or partially on user device 115 and partially on media server 101.

[0137] 8 may begin at block 802. At block 802, it is determined whether a user grants permission to access an initial image. If the user does not grant permission, method 800 ends. If the user grants permission, block 802 may be followed by block 804.

[0138] At block 810, object recognition is performed on the initial image to identify objects in the input image. In some embodiments, performing object recognition to identify objects in the initial image includes determining an object bounding box for each of the objects, and the method further includes determining that the user input corresponds to the selected object based on a proximity of the user input to the nearest object bounding box. Block 810 is followed by block 815.

[0139] Block 815 determines whether the initial image is an indoor scene. If the initial image is an indoor scene, block 815 may be followed by block 820. If the initial image is not an indoor scene, block 815 may be followed by block 825.

[0140] In block 820, empty segments are identified from the initial image. Block 820 may be followed by block 825.

[0141] At block 825, it is determined whether the initial image has a subject that is a human or an animal. If the initial image has a subject that is a human or an animal, block 825 may be followed by block 830. In some embodiments, the method further includes generating a background segment in response to the initial image including the subject, where the subject segment is associated with a foreground region and the background segment is associated with a background region, and pixels in the initial image are associated with the foreground region or the background region, and the method further includes determining that the user input corresponds to a foreground region based on the user input contacting a pixel associated with the foreground region. If the initial image does not have a subject that is a human or an animal, block 825 may be followed by block 835.

[0142] Object segments are identified from the initial image in block 830. Block 830 may be followed by block 835.

[0143] At block 835, it is determined whether the initial image has one or more obtrusive objects. If the image does not have one or more obtrusive objects, block 835 may be followed by block 840.

[0144] At block 840, in response to receiving the user input, the selected object is segmented.

[0145] If the initial image has one or more obtrusive objects, block 835 may be followed by block 845 of FIG. 8B.

[0146] At block 845, in response to the initial image including one or more offending objects, one or more offending segments are identified from the initial image. In some embodiments, a convolutional neural network (CNN) performs the segmentation, and method 800 further includes providing the initial image and the heat map of keypoints as input to the CNN, and outputting, using the convolutional neural network, a segmentation mask corresponding to sky segments, object segments, and one or more offending segments. Block 845 may be followed by block 850.

[0147] At block 850, the user interface including the initial image receives user input corresponding to a selected object from the set of objects. The user input may include tapping the selected object multiple times. In this case, method 800 may further include determining a number of taps from the user input and determining the selected object based on the number of taps, where a first tap is associated with a different region than a second tap.

[0148] In some embodiments, the user input includes selecting a sky, and the method further includes receiving a request from the user to modify the lighting of the initial image; providing the initial image and the request to modify the lighting of the initial image as inputs to a diffusion model; and using the diffusion model to output an output image that meets the request.

[0149] In some embodiments, the user input includes selecting one or more background objects for removal, and the method further includes removing the one or more offending objects from the initial image based on object recognition, and generating a modified image including repairing pixels associated with the one or more offending segments.

[0150] In some embodiments, the selected object is an incomplete object, where a missing portion of the incomplete object is clipped by a boundary of the initial image or obscured by another object, and the method further includes generating a segmentation mask including the incomplete object, removing the incomplete object from the initial image, generating an inpainting image in which pixels of the incomplete object corresponding to the incomplete object are replaced with background pixels that match a background of the initial image, providing the segmentation mask, the incomplete object, and the inpainting image as inputs to a diffusion model, outputting the complete object using the diffusion model, and generating a modified image by blending one or more versions of the complete object with one or more versions of the inpainting image using the conservation mask. Block 845 may be followed by block 855.

[0151] At block 855, the user interface is updated to include an indication that the selected object has been selected.

[0152] In some embodiments, the method further includes receiving a text request to modify a selected object in the initial image; identifying a face segment from the initial image for a face of the subject based on the subject segment; generating a saved mask corresponding to the face segment; providing the text request, the initial image, and the saved mask as inputs to a diffusion model; and outputting an output image that satisfies the text request using the diffusion model.

[0153] In addition to the above, the system, program, or functionality described herein may provide users with controls that allow them to choose both when and if user information (e.g., information about the user's social network, social actions, or activities, occupation, user preferences, or the user's current location) may be collected and when content or information is sent from the server to the user. Furthermore, certain data may be processed in one or more ways to remove personally identifiable information before storage or use. For example, the user's identity may be processed so that personally identifiable information about the user cannot be determined, or if location information (such as to the city, zip code, or state level) is obtained, the user's geographic location may be generalized so that the user's specific location cannot be determined. Thus, users have control over what information is collected about them, how that information is used, and what information is provided to them.

[0154] In accordance with the above, the media application performs object recognition on the initial image to identify a set of objects in the initial image. The media application determines whether the initial image is an outdoor scene. In response to the initial image being an outdoor scene, the media application identifies a sky segment from the initial image. The media application determines whether the initial image includes an object that is a human or an animal. In response to the initial image including an object that is a human or an animal, the media application identifies an object segment from the initial image. The media application receives user input corresponding to selecting a selected object from the set of objects at a user interface that includes the initial image. The media application updates the user interface to include an indication that the selected object has been selected.

[0155] In the foregoing description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without such specific details. In some instances, structures and devices are shown in block diagram form to avoid obscuring the description. For example, embodiments may be described above primarily with reference to user interfaces and specific hardware. However, embodiments may apply to any type of computing device capable of receiving data and instructions, and any peripheral device that provides services.

[0156] References herein to "some embodiments" or "some examples" mean that a particular feature, structure, or characteristic described in connection with an embodiment or example may be included in at least one implementation of the description. The appearances of the phrase "in some embodiments" in various places in this specification do not necessarily all refer to the same embodiments.

[0157] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, such quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It will be apparent that it is sometimes convenient, principally for reasons of common usage, to refer to such data as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0158] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. As will be apparent from the description that follows, unless otherwise indicated, descriptions using terms such as "processing" or "computing" or "calculating" or "determining" or "displaying" throughout this specification will naturally refer to the actions and processes of a computer system or similar electronic computing device that manipulates and converts data, which are represented as physical (electronic) quantities in the computer system's registers and memory, into other data, which are also represented as physical quantities in the computer system's memory or registers or other such information storage, transmission, or display device.

[0159]

[0013] Embodiments herein may also relate to a processor for executing one or more steps of the above-described methods. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory computer-readable storage medium. Such storage media include, but are not limited to, any type of disk, such as an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory, such as a USB key, with non-volatile memory, or any type of medium suitable for storing electronic instructions, each of which is coupled to a computer system bus.

[0160] The specifications may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, the specifications are implemented in software, including but not limited to firmware, resident software, microcode, etc.

[0161] Furthermore, the subject matter may take the form of a computer program product. Such a computer program product is accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For purposes of the subject matter, a computer-usable or computer-readable medium may be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0162] A data processing system suitable for storing or executing program code includes at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.

Claims

1. performing object recognition on an initial image to identify a set of objects within the initial image; determining whether the initial image is an outdoor scene; responsive to the initial image being an outdoor scene, identifying a sky segment from the initial image; determining whether the initial image includes a subject that is a human or an animal; identifying an object segment from the initial image in response to the initial image including the object; determining whether the initial image includes one or more obtrusive objects; identifying one or more offending segments from the initial image in response to the initial image including one or more offending objects; receiving user input at a user interface including the initial image corresponding to selecting a selected object from the set of objects; updating the user interface to include an indication that the selected object has been selected.

2. The user input includes tapping the selected object multiple times, and the method further comprises: determining a number of taps from the user input; and determining the selected object based on the number of taps, wherein a first tap is associated with a different region than a second tap.

3. generating a background segment in response to the initial image including the object, the object segment being associated with a foreground region and the background segment being associated with a background region, and pixels in the initial image being associated with the foreground region or the background region, the method further comprising: The method of claim 1 , further comprising: determining that the user input corresponds to the foreground region based on the user input contacting a pixel associated with the foreground region.

4. Performing object recognition to identify objects in the initial image includes determining an object bounding box for each of the objects, the method further comprising: The method of claim 1 , comprising determining that the user input corresponds to the selected object based on a proximity of the user input to a nearest object bounding box.

5. A convolutional neural network (CNN) performs the segmentation, and the method further comprises: providing the initial image and a heatmap of keypoints as input to the CNN; and using the convolutional neural network to output segmentation masks corresponding to the sky segment, the object segment, and the one or more distracting segments.

6. The user input includes selecting sky, and the method further comprises: receiving a request from a user to change the lighting of the initial image; providing an initial image and a request to modify the illumination of said initial image as input to a diffusion model; and using the diffusion model to output an output image that satisfies the requirement.

7. The user input includes selecting one or more background objects for removal, and the method further comprises: removing the one or more offending objects from the initial image based on object recognition; and generating a modified image including repairing pixels associated with the one or more offending objects.

8. The selected object is an incomplete object, a missing portion of the incomplete object being clipped by a boundary of the initial image or obscured by another object, and the method further comprises: generating a storage mask including the incomplete object; removing the incomplete object from the initial image; generating an inpainted image in which pixels of the incomplete object corresponding to the incomplete object are replaced with background pixels that match a background of the initial image; providing the stored mask, the incomplete object, and the inpainted image as inputs to a diffusion model; outputting a complete object using said diffusion model; and generating a corrected image by blending one or more versions of the complete object with one or more versions of the inpainted image using the preservation mask.

9. receiving a text request to modify the selected object in the initial image; identifying face segments for the face of the subject from the initial image based on the subject segments; generating a storage mask corresponding to the face segment; providing the text request, the initial image, and the saved mask as inputs to a diffusion model; The method of claim 1 , further comprising: using the diffusion model to output an output image that satisfies the text request.

10. 1. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: performing object recognition on an initial image to identify a set of objects within the initial image; determining whether the initial image is an outdoor scene; responsive to the initial image being an outdoor scene, identifying a sky segment from the initial image; determining whether the initial image includes a subject that is a human or an animal; identifying an object segment from the initial image in response to the initial image including the object; determining whether the initial image includes one or more obtrusive objects; identifying one or more offending segments from the initial image in response to the initial image including one or more offending objects; receiving user input at a user interface including the initial image corresponding to selecting a selected object from the set of objects; and updating the user interface to include an indication that the selected object has been selected.

11. The user input includes tapping the selected object multiple times, and the action further comprises: determining a number of taps from the user input; and determining the selected object based on the number of taps, wherein a first tap is associated with a different region than a second tap.

12. The operation further comprises: generating a background segment in response to the initial image including the object, the object segment being associated with a foreground region and the background segment being associated with a background region, and pixels in the initial image being associated with the foreground region or the background region, the operations further comprising: and determining that the user input corresponds to the foreground region based on the user input contacting a pixel associated with the foreground region.

13. performing object recognition to identify objects in the initial image includes determining an object bounding box for each of the objects, the operations further comprising:

11. The non-transitory computer-readable medium of claim 10, comprising determining that the user input corresponds to the selected object based on a proximity of the user input to a nearest object bounding box.

14. A convolutional neural network (CNN) performs the segmentation, said operation further comprising: providing the initial image and a heatmap of keypoints as input to the CNN; and using the convolutional neural network to output segmentation masks corresponding to the sky segment, the object segment, and the one or more obtrusive segments.

15. The user input includes selecting sky, and the operation further comprises: receiving a request from a user to change the lighting of the initial image; providing an initial image and a request to modify the illumination of said initial image as input to a diffusion model; and using the diffusion model to output an output image that satisfies the requirement.

16. a processor; a memory coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the processor to perform operations, the operations including: performing object recognition on an initial image to identify a set of objects within the initial image; determining whether the initial image is an outdoor scene; responsive to the initial image being an outdoor scene, identifying a sky segment from the initial image; determining whether the initial image includes a subject that is a human or an animal; identifying an object segment from the initial image in response to the initial image including the object; determining whether the initial image includes one or more obtrusive objects; identifying one or more offending segments from the initial image in response to the initial image including one or more offending objects; receiving user input at a user interface including the initial image corresponding to selecting a selected object from the set of objects; and updating the user interface to include an indication that the selected object has been selected.

17. The user input includes tapping the selected object multiple times, and the action further comprises: determining a number of taps from the user input; and determining the selected object based on the number of taps, wherein a first tap is associated with a different region than a second tap.

18. The operation further comprises: generating a background segment in response to the initial image including the object, the object segment being associated with a foreground region and the background segment being associated with a background region, and pixels in the initial image being associated with the foreground region or the background region, the operations further comprising: and determining that the user input corresponds to the foreground region based on the user input contacting a pixel associated with the foreground region.

19. performing object recognition to identify objects in the initial image includes determining an object bounding box for each of the objects, the operations further comprising: The system of claim 16 , further comprising determining that the user input corresponds to the selected object based on a proximity of the user input to a nearest object bounding box.

20. A convolutional neural network (CNN) performs the segmentation, said operation further comprising: providing the initial image and a heatmap of keypoints as input to the CNN; and using the convolutional neural network to output segmentation masks corresponding to the sky segment, the object segment, and the one or more distracting segments.

Citation Information

Patent Citations

  • Makeup simulation system, makeup simulation method, makeup simulation program, and makeup simulation device

    JP2022133792A

  • User input based distraction removal in media items - Patents.com

    JP2024518695A

  • User input based distraction removal in media items

    US20230118361A1

  • Information processing device, information processing method, and program

    WO2010090106A1