Rearrange, replace, and generate objects in an image
A diffusion model-based method inpaints and blends objects within images to achieve realistic and high-quality repositioning and editing, addressing the limitations of existing image editing techniques.
Patent Information
- Application Number
- JP2025503412
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2024-05-09
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-05-09
AI Technical Summary
Existing techniques for moving objects within images often result in poor editing outcomes, such as improperly identified pixels, out-of-place filling of empty spaces, or mismatched background pixels, leading to unrealistic image edits.
A computer-implemented method using a diffusion model to generate a complete object by inpainting missing portions and blending it with the initial image, allowing for seamless repositioning and editing of objects within the image.
The method effectively corrects flawed images by ensuring realistic and high-quality object repositioning and editing, avoiding common errors associated with traditional image editing techniques.
Smart Images

Figure 2025530976000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 465,230, filed May 9, 2023, entitled "Repositioning Objects in an Image," and U.S. Provisional Patent Application No. 63 / 562,634, filed March 7, 2024, entitled "Performing Scene Impact Editing Tasks Using Diffusion Neural Networks," each of which is incorporated herein in its entirety. [Background technology]
[0002] A user may capture an image in which an object is in an undesired position. For example, the object may be clipped by the image boundary or by other objects. Techniques exist for moving objects within an image. However, attempts to move an object to a different position within an image can produce disastrous results. For example, pixels associated with an object may be improperly identified such that part of the object remains in its original position while the rest of the object is moved to a different position (e.g., a chicken's body is moved while the chicken's legs remain unmoved). In other examples, empty space created by removing pixels associated with a moved object may be filled with pixels that appear out of place. In yet other examples, pixels surrounding a moved object may appear different from the background, resulting in an image that appears poorly edited.
[0003] The discussion of the background art provided herein is intended to generally present the context for the present disclosure. The work of the presently named inventors, to the extent described in this background art section, as well as aspects of the present disclosure that may not specifically qualify as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure. Summary of the Invention
[0004] A computer-implemented method includes receiving a selection of an incomplete object in an initial image, the incomplete object being associated with a first location in the initial image, and an omitted portion of the incomplete object being clipped by a boundary of the initial image or obscured by another object. The method further includes generating an object mask including incomplete object pixels associated with the incomplete object. The method further includes removing the incomplete object pixels associated with the incomplete object from the initial image. The method further includes generating an inpainting image that replaces the incomplete object pixels corresponding to the incomplete object with inpainting pixels. The method further includes providing the object mask, the incomplete object, and the inpainting image as inputs to a diffusion model. The method further includes outputting the complete object using the diffusion model. The method further includes generating a modified image by blending one or more versions of the complete object with one or more versions of the inpainting image using the object mask, the complete object being located at a second location in the modified image that is different from its first location in the initial image.
[0005] In some embodiments, the modified image is a first modified image, and the method further includes receiving a request to remove a selected object from the first modified image, providing the first modified image and the request as input to an object removal model associated with the diffusion model, and using the object removal model to output a second modified image without the selected object. In some embodiments, the modified image is the first modified image, and the method further includes receiving a request to move the selected object from a third position to a fourth position, providing the first modified image and the request as input to an object insertion model associated with the diffusion model, and using the object insertion model to output a second modified image with the selected object at the fourth position based on the request.
[0006] In some embodiments, the modified image is a first modified image, and the method further includes receiving a request to add an additional object to the initial image, outputting the additional object using a diffusion model, and outputting a second modified image by blending one or more versions of the additional object with one or more versions of the inpainted image using the diffusion model. In some embodiments, the request to add the additional object includes a text prompt describing the additional object.
[0007] In some embodiments, the whole object is resized based on a change from a first position in the initial image to a second position in the modified image. In some embodiments, the method further includes receiving a command to uncrop the modified image to extend an uncrop boundary of the modified image to the extended boundary, and outputting an uncropped image including repair pixels between the uncrop boundary and the extended boundary of the modified image based on the command. In some embodiments, the command to uncrop the repair image includes selecting an uncrop button and either a command to directly extend the uncrop boundary of the modified image to the extended boundary or a movement of the whole object that extends the uncrop boundary of the modified image to the extended boundary. In some embodiments, the method further includes modifying lighting of the modified image and adding a shadow to the whole object based on a direction of lighting in the modified image.
[0008] In some embodiments, a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including receiving a selection of an incomplete object in an initial image, the incomplete object being associated with a first location in the initial image, with an omitted portion of the incomplete object being clipped by a boundary of the initial image or obscured by another object, generating an object mask including incomplete object pixels associated with the incomplete object, removing the incomplete object pixels associated with the incomplete object from the initial image, generating an inpainting image that replaces the incomplete object pixels corresponding to the incomplete object with inpainting pixels, providing the object mask, the incomplete object, and the inpainting image as inputs to a diffusion model, outputting the complete object using the diffusion model, and generating a modified image by blending one or more versions of the complete object with one or more versions of the inpainting image using the object mask, the complete object being located at a second location in the modified image that is different from its first location in the initial image.
[0009] In some embodiments, the modified image is a first modified image, and the operations further include receiving a request to remove a selected object from the first modified image, providing the first modified image and the request as input to an object removal model associated with the diffusion model, and using the object removal model to output a second modified image without the selected object. In some embodiments, the modified image is the first modified image, and the operations further include receiving a request to move the selected object from a third position to a fourth position, providing the first modified image and the request as input to an object insertion model associated with the diffusion model, and using the object insertion model to output the second modified image with the selected object at the fourth position based on the request.
[0010] In some embodiments, the modified image is a first modified image, and the operations further include receiving a request to add an additional object to the initial image, the request including a text prompt describing the additional object; outputting the additional object using a diffusion model; and outputting a second modified image by blending one or more versions of the additional object with one or more versions of the inpainted image using the diffusion model. In some embodiments, the complete object is resized based on a change from a first position in the initial image to a second position in the modified image. In some embodiments, the operations further include receiving a command to uncrop the modified image to extend an uncrop boundary of the modified image to an extended boundary; and outputting an uncropped image including inpainted pixels between the uncrop boundary and the extended boundary of the modified image based on the command.
[0011] In some embodiments, a system includes a processor and a memory coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to perform operations including receiving a selection of an imperfect object in an initial image, the imperfect object being associated with a first location in the initial image and an omitted portion of the imperfect object being clipped by a boundary of the initial image or obscured by other objects, generating an object mask including imperfect object pixels associated with the imperfect object, removing the imperfect object pixels associated with the imperfect object from the initial image, generating an inpainting image that replaces the imperfect object pixels corresponding to the incomplete object with inpainting pixels, providing the object mask, the imperfect object, and the inpainting image as inputs to a diffusion model, outputting the complete object using the diffusion model, and generating a modified image by blending one or more versions of the complete object with one or more versions of the inpainting image using the object mask, the complete object being located at a second location in the modified image that is different from its first location in the initial image.
[0012] In some embodiments, the modified image is a first modified image, and the operations further include receiving a request to remove a selected object from the first modified image, providing the first modified image and the request as input to an object removal model associated with the diffusion model, and using the object removal model to output a second modified image without the selected object. In some embodiments, the modified image is the first modified image, and the operations further include receiving a request to move the selected object from a third position to a fourth position, providing the first modified image and the request as input to an object insertion model associated with the diffusion model, and using the object insertion model to output the second modified image with the selected object at the fourth position based on the request.
[0013] In some embodiments, the modified image is a first modified image, and the operations further include receiving a request to add an additional object to the initial image, the request including a text prompt describing the additional object; outputting the additional object using a diffusion model; and outputting a second modified image by blending one or more versions of the additional object with one or more versions of the inpainted image using the diffusion model. In some embodiments, the complete object is resized based on a change from a first position in the initial image to a second position in the modified image. In some embodiments, the operations further include receiving a command to uncrop the modified image to extend an uncrop boundary of the modified image to an extended boundary; and outputting an uncropped image including inpainted pixels between the uncrop boundary and the extended boundary of the modified image based on the command. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a block diagram of an exemplary network environment, according to some embodiments described herein. [Figure 2]FIG. 1 is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3A] 1 illustrates an exemplary initial image according to certain embodiments described herein. [Figure 3B] 3B illustrates an example initial image in which two objects from FIG. 3A are selected for modification, according to certain embodiments described herein. [Figure 3C] 1 illustrates an exemplary initial image in which a mask surrounds an object, according to certain embodiments described herein. [Figure 3D] 1 illustrates an exemplary inpainted image in which bystander objects have been removed and subject objects have been shifted and resized, according to some embodiments described herein. [Figure 3E] 1 illustrates an exemplary restoration image in which roads have been replaced with grass, trees of a first type have been replaced with trees of a second type, and cloudy skies have been replaced with clear skies, according to some embodiments described herein. [Figure 4A] 1 shows an exemplary initial image of a child sitting on a bench and holding a balloon that is partially cut off by the boundary of the initial image, according to some embodiments described herein. [Figure 4B] 10 illustrates an exemplary modified image in which the child, bench, and balloon have been moved to a second position, according to certain embodiments described herein. [Figure 5A] 10 illustrates an example user interface for an initial image including a button for modifying the border of the initial image, according to certain embodiments described herein. [Figure 5B] 10 illustrates an exemplary user interface with indicators used to expand the boundaries of an initial image, according to some embodiments described herein. [Figure 5C] 10 illustrates an example user interface for an uncropped image output based on an initial image, according to some embodiments described herein. [Figure 5D]10 illustrates an alternative exemplary user interface in which a selected object is used to extend the boundary of an initial image, according to some embodiments described herein. [Figure 6] 1 shows an example flowchart of a method for generating a corrected image of a complete object from an incomplete object, according to some embodiments described herein. [Figure 7] 1 shows an example flowchart of a method for outputting an uncropped image from an initial image according to some embodiments described herein. [Figure 8] 1 illustrates an example flowchart of a method for training an object removal model according to some embodiments described herein. [Figure 9] 1 illustrates an example flowchart of a method for training an object insertion model according to some embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION
[0015] A user may capture an image in which an object is in an undesired position. For example, the object may be clipped by the image boundary or by other objects. Techniques exist for moving objects within an image. However, attempts to move an object to a different position within an image can produce disastrous results. For example, pixels associated with an object may be improperly identified such that part of the object remains in its original position while the rest of the object is moved to a different position (e.g., a chicken's body is moved while the chicken's legs remain unmoved). In other examples, empty space created by removing pixels associated with a moved object may be filled with pixels that appear out of place. In yet other examples, pixels surrounding a moved object may appear different from the background, resulting in an image that appears poorly edited.
[0016] The techniques described below advantageously solve these problems by providing an incomplete object as input to a diffusion machine learning model, referred to herein as a diffusion model, and outputting a complete object. An incomplete object is a partial representation of an object in an image. The object is partially present (not completely present) in the image. The portion of the object that is not present in the image is referred to herein as the "omitted portion" of the incomplete object. A user can select the incomplete object in the initial image and move the position of the incomplete object.
[0017] The space left by the incomplete object is inpainted with inpainting pixels to form an inpainting image. The inpainting image is a complete representation of the object, including the incomplete object and omitted portions of the incomplete object. The inpainting image is an image that differs from the initial image in that the incomplete object pixels associated with the incomplete object are removed from the inpainting image and the incomplete object pixels are replaced with inpainting pixels that may be selected based on their proximity to surrounding pixels selected from a reference image, including background pixels, etc.
[0018] The diffusion model ensures that the complete object fits into the new position. For example, if an object is moved from the background to the foreground, the diffusion model increases the size of the moved object. The diffusion model uses an object mask to blend one or more versions of the complete object with one or more versions of the inpainted image, thereby outputting a modified image in which the complete object is seamlessly merged with the inpainted image. For example, the diffusion model can blend progressively noisier versions of the complete object with corresponding noisier versions of the inpainted image, but it can also generate a denoised version of the complete object and a corresponding denoised version of the inpainted image. The noisy version of the complete object is created by increasing the entropy of the image; the more noise, the less discernible the details of the complete object in the image. Similarly, the noisy version of the inpainted image is created by increasing the entropy of the inpainted image; the more noise, the less discernible the details of the inpainted image.
[0019] By using a diffusion model instead of other machine learning models, media applications maintain a realistic appearance of the modified image under a wide variety of circumstances. The techniques described below enable the correction of flawed images, i.e., images containing imperfect objects, in an efficient manner. The perfect objects created by utilizing a diffusion model have high quality and are free of the above-mentioned errors. The image processing described herein effectively and efficiently corrects images for imperfect objects present in the image.
[0020] Exemplary Environment 100 FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, environment 100 includes a media server 101, user device 115a, and user device 115n coupled to a network 105. Users 125a, 125n may be associated with respective user devices 115a, 115n. In some embodiments, environment 100 may include other servers or devices not shown in FIG. 1. In FIG. 1 and the remaining figures, a letter following a reference number, e.g., "115a," denotes a reference to the element with that specific reference number. A reference number in the text without a following letter, e.g., "115," denotes a general reference to an embodiment of the element bearing that reference number.
[0021] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively coupled to a network 105 via signal line 102. The signal line 102 may be a wired connection, such as Ethernet, coaxial cable, or fiber optic cable, or a wireless connection, such as Wi-Fi, Bluetooth, or other wireless technology. In some embodiments, the media server 101 transmits and receives data to and from one or more of the user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.
[0022] The database 199 may store machine learning models, training data sets, images, etc. The database 199 may also store social network data associated with the user 125, the user preferences of the user 125, etc.
[0023] The user device 115 may be a computing device that includes a memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or any other electronic device that can access the network 105.
[0024] In the illustrated embodiment, user device 115a is coupled to network 105 via signal line 108, and user device 115n is coupled to network 105 via signal line 110. Media application 103 may be stored on user device 115a as media application 103b and / or on user device 115n as media application 103c. Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections, such as Wi-Fi, Bluetooth, or other wireless technologies. User devices 115a and 115n are accessed by users 125a and 125n, respectively. The user devices 115a and 115n in FIG. 1 are used as an example. While FIG. 1 shows two user devices, 115a and 115n, this disclosure applies to system architectures having one or more user devices 115.
[0025] The media application 103 may be stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are executed on the media server 101 or the user device 115. In some embodiments, some operations may be executed on the media server 101 and some may be executed on the user device 115. Execution of the operations is subject to user settings. For example, the user 125a may specify that operations be executed on each device 115a and not on the media server 101. Such settings result in the operations described herein being executed entirely on the user device 115a and not on the media server 101. Furthermore, the user 125a may specify that the user's images and / or other data be stored only locally on the user device 115a and not on the media server 101. Such settings result in user data not being transmitted to or stored on the media server 101. The transmission of user data to the media server 101, any temporary or permanent storage of such data by the media server 101, and the performance of actions on such data by the media server 101 are performed only if the user consents to the transmission, storage, and performance of actions by the media server 101. The user is provided with the option to change settings at any time, for example, so that the user can enable or disable use of the media server 101.
[0026] Machine learning models (e.g., diffusion models, neural networks, or other types of models) are stored locally on the user device 115 and utilized with specific user permission when utilized for one or more operations. Server-side models are utilized only with user permission. Additionally, trained models may be provided for use on the user device 115. During such utilization, on-device training of the model may be performed if authorized by the user 125. Updated model parameters may be transmitted to the media server 101 if authorized by the user 125, for example, to enable federated learning. The model parameters do not include any user data.
[0027] The media application 103 receives an initial image. For example, the media application 103 receives the initial image from a camera that is part of the user device 115, or the media application 103 receives the initial image over the network 105. The media application 103 receives a selection of an incomplete object in the initial image. The incomplete object is associated with a first position in the initial image, and an omitted portion of the incomplete object is clipped by a boundary of the initial image or obscured by another object. The incomplete object may be selected when the user 125 taps the object, draws a shape (e.g., a circle) around the object, confirms a suggestion by the media application 103 to modify the object, etc.
[0028] The media application 103 generates an object mask including the incomplete object pixels associated with the incomplete object, removes the incomplete object pixels associated with the incomplete object from the initial image, and generates an inpainting image that replaces the incomplete object pixels corresponding to the incomplete object with inpainting pixels.
[0029] The media application 103 uses a diffusion model to output the complete object. For example, if an incomplete object is clipped by an edge of the initial image and the incomplete object is moved to the center of the initial image, the diffusion model outputs the complete object that fills in the missing portion of the incomplete object. The media application 103 uses the diffusion model to output the modified image by blending one or more versions of the complete object with one or more versions of the inpainted image using an object mask. The complete object is placed at a second position in the modified image that is different from its first position in the initial image. In some embodiments, the modified image may include a watermark or other indicator to identify that the modified image was generated using a machine learning model.
[0030] In some embodiments, the media application 103 may be implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, the media application 103a may be implemented using a combination of hardware and software.
[0031] Exemplary Computing Device 200 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. Computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, computing device 200 is media server 101 used to implement media application 103a. In another example, computing device 200 is user device 115.
[0032] In some embodiments, computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, a display 241, a camera 243, and a storage device 245, all coupled via bus 218. Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, display 241 may be coupled to bus 218 via signal line 228, camera 243 may be coupled to bus 218 via signal line 230, and storage device 245 may be coupled to bus 218 via signal line 232.
[0033] Processor 235 may be one or more processors and / or processing circuits that execute program code and control basic operations of computing device 200. A “processor” includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) with one or more cores (e.g., single-core, dual-core, or multi-core configurations), a system having multiple processing units (e.g., multiprocessor configurations), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuits for realizing functions, dedicated processors for implementing processing based on neural network models, neural circuits, processors optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, processor 235 may include one or more coprocessors that perform neural network processing. In some embodiments, processor 235 may be a processor that processes data to generate probabilistic outputs; for example, the outputs generated by processor 235 may be inaccurate or accurate within a range from an expected output. Processing need not be limited to a particular geographic location or have time limitations. For example, a processor may perform its functions in real time, offline, in batch mode, etc. Portions of processing may be performed at different times and in different locations by different (or the same) processing systems. A computer may be any processor in communication with a memory.
[0034] Memory 237 is provided within computing device 200 for access by processor 235 and may be any suitable processor-readable storage medium, located separate from and / or integral with processor 235, suitable for storing instructions for execution by a processor or set of processors, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc. Memory 237 may store software operated on computing device 200 by processor 235, including media application 103.
[0035] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, etc. One or more methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application (“app”) running on a mobile computing device, etc.
[0036] Application data 266 may be data generated by other applications 264 or by hardware of computing device 200. For example, application data 266 may include images used by an image library application, user actions identified by other applications 264 (e.g., social networking applications), etc.
[0037] I / O interface 239 may provide functionality that allows computing device 200 to interface with other systems and devices. The interfaced devices may be included as part of computing device 200 or may be separate and in communication with computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices may communicate through I / O interface 239. In some embodiments, I / O interface 239 may connect to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.).
[0038] Some examples of interface devices that can be connected to I / O interface 239 can include display 241, which can be used to display content, e.g., images, video, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from a user. For example, display 241 can be utilized to display a user interface, including a graphical guide, on a viewfinder. Display 241 can include any suitable display device, such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touchscreen, a three-dimensional display screen, or other visual display device. For example, display 241 can be a flat display screen provided on a mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen of a computing device.
[0039] Camera 243 may be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that I / O interface 239 sends to media application 103.
[0040] The storage device 245 stores data related to the media application 103. For example, the storage device 245 can store training data sets including labeled images, machine learning models, output from machine learning models, etc.
[0041] FIG. 2 illustrates an exemplary media application 103 including a user interface module 202 , a segmenter 204 , a restorer module 206 , and a diffusion module 208 stored in memory 237 .
[0042] The user interface module 202 generates graphical data for displaying a user interface including an image. In some embodiments, the user interface module 202 receives an initial image. The initial image may be received from the camera 243 of the computing device 200 or from the media server 101 via the I / O interface 239. The initial image includes a subject, such as a person. In some embodiments, the user interface includes options for selecting various people and other objects within the initial image. For example, a user may select a person by tapping on the person, circling the object, brushing on the object, etc. In some embodiments, the user interface generates recommendations for modifying the image, such as displaying text asking if the user wants to remove bystanders from the object.
[0043] In some embodiments, when a user selects an object, the user interface module 202 updates the graphical data to include a highlighted version of the selected object. The user can modify the selected object. For example, the user may drag and drop the selected object from a first location to a second location, the user may resize an image, the user may select a button to clear the image, etc.
[0044] In some embodiments, the segmenter 204 generates a segmentation score that reflects the quality of identification of pixels associated with the selected object in the initial image. The user interface may include various options for modifying the selected object based on the segmentation score. For example, if the segmentation score exceeds a threshold, the user interface module 202 provides options to move the selected object, replace the selected object with a different object, or delete the selected object. In other examples, if the segmentation score does not exceed a threshold, the user interface module 202 does not provide an option to move the selected object, but provides options to replace the selected object or delete the selected object.
[0045] In some embodiments, the user interface module 202 generates graphical data for displaying the repaired image, moving a selected object from a first position to a second position within the image, resizing the selected image, adding additional objects, etc. The user interface may also include options for editing the repaired image, sharing the repaired image, adding the repaired image to a photo album, etc.
[0046] The segmenter 204 segments the selected object from the initial image by identifying pixels that correspond to the selected object. In some embodiments, the segmenter 204 uses an alpha map as part of a technique for distinguishing between the foreground and background of the initial image during segmentation. The segmenter 204 may also identify the texture of the selected object in the foreground of the initial image. In some embodiments, the segmenter 204 generates a segmentation map that identifies pixels associated with one or more objects in the initial image. For example, the segmentation map may include identification of pixels associated with the selected object.
[0047] The segmenter 204 can perform segmentation by detecting objects in the initial image. The objects can be people, animals, cars, buildings, etc. A person can be a subject of the initial image or can be a bystander (i.e., a person who is not a subject of the initial image). Bystanders can include people walking, running, bicycling, standing behind the subject, or otherwise entering the initial image. In different examples, bystanders can be in the foreground (e.g., a person passing in front of the camera), at the same depth as the subject (e.g., a person standing next to the subject), or in the background. In some examples, there can be multiple bystanders in the initial image. A bystander can be a human in any pose, such as standing, sitting, crouching, lying down, jumping, etc. A bystander can be facing the camera, angled relative to the camera, or facing away from the camera.
[0048] The segmenter 204 may perform object recognition to identify the likely shape of the object to determine whether the pixel is associated with the selected object or the background, and may detect the type of object by comparing the object to previous objects such as people, vehicles, buildings, etc. The segmenter 204 may generate a region of interest for the selected object, such as a bounding box having x, y coordinates and a scale.
[0049] The segmenter 204 generates one or more object masks for one or more selected objects in the initial image. The object masks represent regions of interest. Object masks are described in more detail below with reference to the diffusion model.
[0050] In some embodiments, one or more object masks are generated based on generating superpixels of the image and matching the centroids of the superpixels to depth map values (e.g., values obtained by the camera 243 using a depth sensor or by deriving depth from pixel values) with depth-based cluster detection. More specifically, the depth values of the masked region may be used to determine a depth range, and superpixels that fall within the depth range may be identified. Other techniques for generating masks include weighting depth values based on how close the depth values are to the object mask, where the weights are represented by a distance transform map.
[0051] In some embodiments, segmenter 204 uses a machine learning algorithm, such as a neural network, or more specifically, a convolutional neural network, to segment the initial image and generate an object mask. Segmenter 204 can specify circuitry (e.g., for a programmable processor, for a field programmable gate array (FPGA), etc.) that enables processor 235 to apply the machine learning model. In some embodiments, segmenter 204 can include software instructions, hardware instructions, or a combination thereof. In some embodiments, segmenter 204 can provide an application programming interface (API) that can be used by operating system 262 and / or other applications 264 to invoke segmenter 204, for example, to apply a machine learning model to application data 266 and output the object mask.
[0052] The segmenter 204 uses training data to generate a trained machine learning model. For example, the training data may include pairs of an initial image with one or more objects and an output image with one or more corresponding object masks.
[0053] The training data may be obtained from any source, e.g., a data repository marked for training purposes, data that has been given permission to be used as training data for machine learning, etc. In some embodiments, training may occur on the media server 101 providing the training data directly to the user device 115, training occurs locally on the user device 115, or a combination of both.
[0054] In some embodiments, segmenter 204 uses unedited / untransferred weights obtained from other applications. For example, in these embodiments, a trained model may be generated, e.g., on a different device, and provided as part of segmenter 204. In various embodiments, the trained model may be provided as a data file that includes a model structure or format (e.g., defining the number and type of neural network nodes, the connectivity between the nodes, and the organization of the nodes into multiple layers) and associated weights. Segmenter 204 may read the trained model data file and implement a neural network with node connectivity, layers, and weights based on the model structure or format specified in the trained model.
[0055] The trained machine learning model may include one or more model formats or structures, such as a linear network, a deep learning neural network that implements multiple layers (e.g., each layer is a linear network with "hidden layers" between the input and output layers), a convolutional neural network (e.g., a network that divides or partitions input data into multiple portions or tiles, processes each tile separately using one or more neural network layers, and aggregates the results of processing each tile), a sequence-to-sequence neural network (e.g., a network that receives sequential data as input, such as words in a sentence or frames of a video, and outputs a sequence of results), or any other type of neural network.
[0056] The model format or structure may specify the connectivity between various nodes and their organization into layers. For example, nodes in a first layer (e.g., an input layer) may receive data as input or application data. Such data may include, for example, one or more pixels per node, for example, when the trained model is used to analyze an initial image. Subsequent intermediate layers may receive as input the output of nodes in the previous layer according to the connectivity specified in the model format or structure. These layers may also be referred to as hidden layers. For example, the first layer may output a segmentation between foreground and background. The final layer (e.g., an output layer) generates the output of the machine learning model. For example, the output layer may receive the segmentation of the initial image into foreground and background and output whether the pixel is part of an object mask or the remainder of the initial image. In some embodiments, the model format or structure also specifies the number and / or type of nodes in each layer.
[0057] In different embodiments, the trained model may include one or more models. One or more of the models may include multiple nodes arranged in layers per model structure or format. In some embodiments, a node may be a memoryless computational node configured, for example, to process a unit of input and generate a unit of output. The computation performed by the node may include, for example, multiplying each of multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to generate the node output. In some embodiments, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some implementations, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, computations by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, individual processing units of a graphics processing unit (GPU), or dedicated neural circuitry. In some embodiments, a node may include memory, for example, capable of storing and using one or more previous inputs when processing a subsequent input. For example, a node with memory may include a long short-term memory (LSTM) node, which may use memory to maintain a "state" that allows the node to operate like a finite state machine (FSM).
[0058] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may begin as multiple nodes organized into layers, as specified by the model format or model structure. At initialization, a respective weight may be applied to the connection between each pair of nodes connected according to the model format, e.g., nodes in successive layers of a neural network. For example, each weight may be randomly assigned or initialized to a default value. The model may then be trained to produce results, e.g., using training data.
[0059] The training may include applying supervised learning techniques. In supervised learning, the training data may include multiple inputs (e.g., an initial image, an object mask, an object mask, etc.) and a corresponding ground truth output for each input (e.g., a ground truth mask that correctly identifies the object in each image). Based on a comparison of the model's output and the ground truth output, the values of the weights are automatically adjusted, for example, to increase the probability that the model will generate the ground truth output for the image.
[0060] In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In some embodiments, the trained model may include a fixed set of weights, for example, downloaded from a server that provides the weights. In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In embodiments in which data is omitted, the segmenter 204 may generate a trained model based on pre-training, for example, pre-training by a developer of the segmenter 204, pre-training by a third party, etc. In some embodiments, the trained model may include a fixed set of weights, for example, downloaded from a server that provides the weights.
[0061] In some embodiments, the trained machine learning model receives an initial image including one or more selected objects, hi some embodiments, the trained machine learning model outputs one or more object masks including the one or more objects.
[0062] After one or more object masks are output by the segmenter 204 (eg, from a machine learning model), the segmenter 204 removes the one or more selected objects from the initial image.
[0063] The restorer module 206 generates an inpainting image that replaces object pixels corresponding to one or more objects with inpainting pixels. The inpainting pixels may be based on pixels from a reference image at the same location without the object. Alternatively, the restorer module 206 may identify inpainting pixels to replace the removed object based on the proximity of the inpainting pixels to other pixels surrounding the object. The restorer module 206 may use gradients of nearby pixels to determine the characteristics of the inpainting pixel. For example, if a bystander was standing on the ground, the restorer module 206 replaces the inpainting pixel with a ground pixel. Other inpainting techniques are possible, including machine learning-based inpainting techniques that output inpainting pixels based on training data containing images of similar structures.
[0064] In embodiments where the user chooses to erase the selected object, the user interface module 202 may display a repair image in which the selected object has been removed and the pixels of the selected object have been replaced with repair pixels.
[0065] In embodiments where the user chooses to move a selected object, replace a selected object, or add an object to the image, the diffusion module 208 uses a diffusion model to perform blending of the object with the object mask and the inpainting image. The diffusion model may receive the object mask, the incomplete object, and the inpainting image as inputs. The diffusion model may receive additional inputs such as a text request, the number of pixels to fill to output a complete object based on the incomplete object, and the dimensions of the inpainting image.
[0066] The diffusion model includes a forward process, in which the diffusion model adds noise to the data, and a backward process, in which the diffusion model learns to recover the data from the noise. For example, if a selected object is moved from a first position to a second position, the diffusion module 208 applies the diffusion model by blending the selected object with progressively noisier versions of the inpainted image, and then with progressively denoised versions. In some embodiments, an object stitching diffusion model is used to move the object from the first position to the second position. In some embodiments, a generative diffusion model is used when the object is an incomplete object and portions of the object are generated, and / or for new objects generated from text prompts.
[0067] Object Stitching Diffusion Model In some embodiments, an object stitching diffusion model is used when an object is moved from a first position to a second position. In some embodiments, the diffusion module 208 includes an object image encoder that extracts semantic features from a selected object, a diffusion model that blends the object with the image, and a content adapter that converts a sequence of visual tokens into a sequence of text tokens to bridge the domain gap between the image and the text. In some embodiments, the diffusion module 208 trains the diffusion model using self-supervision based on training data including image and text pairs. In some embodiments, the diffusion model is trained with synthetic data that simulates real-world scenarios. The diffusion model may also be trained using data augmentation, which is generated by introducing random shifts and crop augmentations during training, while ensuring that foreground objects are contained within the crop window.
[0068] During the first stage, the content adapter is trained to preserve the high-level semantics of the object using image and text pairs, and during the second stage, the content adapter is trained in the context of a diffusion model to encode important identity features of the object by facilitating the visual reconstruction of the object in the original image. The diffusion model may be trained with the embeddings generated by the content adapter via the cross-attention block.
[0069] The diffusion model uses the object mask to blend the inpainted image with the object. The diffusion model may denoise the masked region. The content adapter may convert visual features from the object image encoder into text features (tokens) to use as conditioning for the diffusion model.
[0070] Generative Diffusion Model In some embodiments, the generative diffusion model is used to output a complete object based on an incomplete object, or to generate a new object. The diffusion module 208 trains the generative diffusion model based on training data. The training data may include image and text pairs used to create an embedding space for the image and text. The image and text pair may include an image associated with corresponding text, such as an image of a dog and text containing "pit bull." The diffusion module 208 may be trained with a loss that reflects the cosine distance between the embedding of the text prompt and the embedding of the estimated clean image (i.e., without the text-generated object).
[0071] The diffusion module 208 can use the training data to perform text conditioning, which describes the process of outputting an object conditioned by a text prompt. The diffusion module 208 can train a neural network to output an object based on a text prompt provided by a user or by a media application. For example, the text prompt can be a suggestion generated by the media application based on the context of an initial image (e.g., if the initial image is a beach, the text prompt can be about a beach ball, a turtle, etc.).
[0072] In some embodiments, the diffusion module 208 may output at least a portion of the missing portion of the object, a location in the image to which the incomplete object is moved, and output dimensions (e.g., original or modified dimensions if the object is resized) based on receiving an incomplete object as input data. For example, if a user selects an object in the user interface that is partially cropped by a boundary and moves the object from a first position to a second position, where the second position also crops out a portion of the object, the diffusion module 208 may output a modified object that includes more of the object that is visible based on moving the object in the image. In some embodiments, the diffusion module 208 may output a complete object based on an incomplete object selected by a user. For example, if a user selects a beach ball that is partially obscured by another object, a diffusion model may be trained to output the complete beach ball.
[0073] In some embodiments, the diffusion module 208 generates progressively noisier versions of the complete object compared to previous versions, and generates progressively noisier versions of the inpainted image compared to previous versions. For example, a forward Markov noise process generates a series of noisy inpainted images by gradually adding Gaussian noise until nearly isotropic Gaussian noise samples are obtained. The forward noise process defines a progression of image manifolds, each manifold consisting of noise images.
[0074] The diffusion module 208 can use an object mask to spatially blend a noisy version of the complete object with a corresponding noisy version of the inpainted image. For example, the diffusion module 208 can use an object mask to blend each noisy version of the complete object with each corresponding noisy version of the inpainted image, where the object mask defines the boundary of the complete object such that the object mask defines the region to be modified during the blending process. In some embodiments, the diffusion process can include local complete-object guided diffusion, where the image generation loss determined during the training process is used under the object mask during local object generation diffusion.
[0075] The diffusion module 208 can perform a diffusion step to denoise the latent space in a direction dependent on the text prompt. The diffusion module 208 generates a version of the complete object that is progressively denoised compared to the previous version, and a version of the inpainted image that is progressively denoised compared to the previous version. For example, an inverse Markov process transforms Gaussian noise samples by iteratively denoising the inpainted image using the learned posterior. Each step of the denoising diffusion process projects the noisy image onto the next, less noisy manifold.
[0076] The diffusion module 208 performs a denoising diffusion step after each blend to restore consistency by projecting onto the next manifold. Once spatial blending is complete, the diffusion module 208 preserves the background by replacing regions outside the object mask with corresponding regions from the inpainting image.
[0077] In some embodiments, the diffusion module 208 applies an iterative refinement scheme to inject contextual information into the object to match the style of the inpainted image using cross-domain compositing. For example, if an object is generated for an indoor setting and added to an outdoor inpainted image, the object may be modified to be brighter to match the inpainted image. In another example, if an object is in a first position in the shadow and a second position is in full sun, the object may be modified to match the brightness of the second position.
[0078] Object Removal Model In some embodiments, instead of using the segmenter 204 to remove objects and the restorer module 206 to add pixels to the removed regions of the initial image, the diffusion model 208 is trained to include an object removal model.
[0079] The diffusion module 208 generates counterfactual training data to train the diffusion model, including the object removal model. For each counterfactual image pair, the diffusion module 208 captures a factual image containing the object in the scene, physically removes the object while avoiding camera movement, lighting changes, or other object movement, captures a counterfactual image of the scene without the object, and segments the factual image to create an object mask. Segmenting the factual image generates a segmentation map (M) of the object O removed from the factual image X. o )
[0080] For each image pair, the diffusion module 208 creates a combined image that includes the factual image, the object mask, and the counterfactual image. The object mask is a binary object mask (M o (X)), and the counterfactual image pair is the factual image and the binary object mask (X,M o (X)) and the output counterfactual image (X cf )
[0081]
number
[0082] Once the diffusion model, including the object removal model, has been trained, the user interface module 202 can receive a request to remove a selected object from a first modified image, where the initial image and the request are provided as inputs to the object removal model, which outputs a modified image that does not include the selected object.
[0083] Object Insertion Model In some embodiments, instead of removing the object using the segmenter 204, a restorer module 206 is used to add pixels to the removed region of the initial image, and a diffusion model 208 is used to blend the object with the pixels in its new location, where the diffusion model 208 is trained to include an object insertion model.
[0084] In some embodiments, the object insertion model is trained on a number of image pairs that exceeds the number of available counterfactual image pairs. As a result, the diffusion module 208 generates synthetic training data. For each synthetic image pair, the diffusion module 208 selects an original image that contains the object, uses the object removal model to output a modified image from the original image without the object, generates an input image by inserting the object into the modified image, and segments the original image to create an object mask. The modified image lacking the object is computed using the following formula: i It is called.
[0085]
number
[0086] The composite image pair is (y i ,M o (x i )) and the corresponding target is the original image x i Both the input image and the output image contain the object o, but the input image does not contain the effect of the object on the scene, while the output image contains the effect of the object on the scene. In some embodiments, the diffusion module 208 uses the diffusion target presented in Equation 1 to train an object insertion model.
[0087] For each synthetic image pair, the diffusion module 208 creates a second combined image that includes the original image, the object mask, and the input image. The diffusion module 208 pre-trains the diffusion model to include an object insertion model based on using the synthetic image pair, and fine-tunes the diffusion model to include the object insertion model based on using the counterfactual image pair used to train the object removal model.
[0088] In some embodiments, the user interface module 202 generates graphical data for displaying a user interface that provides a user with options for specifying the location of the object and for resizing the object. The diffusion module 208 adds the selected object removed from the initial image to a new location. In some embodiments, the diffusion module 208 provides the selected object and the location where the selected object is located in the corrected image as input to a diffusion model, and outputs a corrected image that blends the selected object with the inpainted image. For example, the diffusion module 208 can spatially blend a noisy version of the inpainted image with a noisy version of the selected object.
[0089] In some embodiments, the diffusion module 208 may add a shadow to the selected object at its new position. The shadow may match the direction of the light in the image. For example, if the sun casts light from the upper left corner of the image, a shadow may appear to the right of the person and / or object. In some embodiments, the diffusion module 208 uses a machine learning model to output a shadow mask that is used to generate the shadow that is applied to the object.
[0090] Once the selected object has been added to the inpainted image, the user interface module 202 may include additional functionality for modifying the inpainted image, such as an option to change the lighting of the inpainted image.
[0091] 3A shows an example initial image 300. The initial image includes a person 301, which is the subject of the initial image 300, a bystander 302 in the foreground of the initial image 300, grass 303, a road 304, a tree 305, and an overcast sky 306.
[0092] 3B shows an exemplary initial image 310 in which two objects from FIG. 3A have been selected for modification. The two objects are displayed with outlines 311, 312 of how the user selected the two objects. The user may have selected the two objects using a finger, mouse, or other object by circling, brushing, double-tapping, etc., around the two objects. The user interface module 202 may identify the objects in the initial image 310 using an object selection tool, a lasso tool, an artificial intelligence segmentation tool, etc. In some embodiments, the user interface module 202 may suggest selecting the two objects, and in response to the user confirming the selection, the two objects are highlighted.
[0093] Once the two objects are selected, the segmenter 204 segments the objects from the initial image 310 and generates object masks. Figure 3C shows an example initial image 320 with object masks 321, 322 surrounding two objects. The person is surrounded by the first object mask 321, and the bystander is surrounded by the second object mask 322. The segmenter 204 removes the person and the bystander from the initial image 320.
[0094] Once the person and bystanders are removed from the initial image 320 of Figure 3C, the restorer module 206 generates an inpainting image that replaces object pixels corresponding to the removed objects with inpainting pixels. Figure 3D shows an example inpainting image 330 in which the bystanders have been removed from the initial image 320 and the person 331 has been moved and resized.
[0095] In some embodiments, the diffusion module 208 resizes the object to be larger or smaller than the object in the initial image. For example, an object may be resized to be larger when moved forward and smaller when moved backward. In this example, the user provided input to resize the person 331 smaller, and the diffusion module 208 resized the person 331 and blended the person into the new position. In some embodiments, moving the person 331 may be a separate action from the resizing, or both moving and resizing may be part of the same action.
[0096] Additional modifications may be made to the inpainted image. Figure 3E shows an example of an inpainted image 340 in which the road 332 from Figure 3D has been replaced with grass 341. The inpainted image 340 also includes a first type of tree replaced with a second type of tree 342. The inpainted image 340 also includes the cloudy sky from Figure 3D replaced with a sunny sky with sunlight 343 emanating from the upper right corner of the sky. The direction of the sun 343 causes the diffusion module 208 to output a shadow 344 that matches the person 345. In some embodiments, the road is replaced with grass 341 by using the diffusion module 208 to output grass and blending the grass with the inpainted image.
[0097] In some embodiments, the diffusion module 208 receives an incomplete object as input and outputs a complete object. The diffusion module 208 can also add the complete object to a second location in the inpainted image by blending pixels of the complete object that correspond to the complete object with the inpainted pixels.
[0098] In some embodiments, if an object is captured at the edge of an image and a portion of the complete object is missing, the diffusion module 208 may fill in the missing portion of the object before performing the blending. For example, if a woman is wearing a long dress and a portion of the dress is missing, the diffusion module 208 may output the complete dress based on the incomplete dress.
[0099] 4A shows an exemplary initial image 400 of a child 405 sitting on a bench 410 and holding a balloon 415 that is partially cut off by the boundary of the initial image 400. In this example, the user interface module 202 provides a user interface with options for the user to select objects. The user selects the child 405, the bench 410, and the balloon 415 in a first position. The balloon 415 represents an incomplete image. The segmenter 204 segments the child 405, the bench 410, and the balloon 415 to separate the objects from the initial image 400.
[0100] The user interface module 202 includes an option to move the selected object to another location. The user selects a second location. The segmenter 204 removes the selected object from the initial image. The inpainting module 206 generates an inpainting image that replaces object pixels corresponding to the removed object with inpainting pixels.
[0101] The diffusion module 208 receives as input the selected object and the coordinates of the second position and outputs the complete object: the balloon and the longer bench. Figure 4B shows an example rectified image 450 in which the child 455, bench 460, and balloon 465 have been moved to the second position. In this example, the diffusion module 208 outputs a rectified image that uses an object mask to blend one or more versions of the child 455, bench 460, and balloon 465 with one or more versions of the inpainted image.
[0102] In some embodiments, the user interface module 202 receives a command to uncrop an image from a user interface. The command to uncrop an image may occur on a modified image, an initial image, etc. The command to uncrop an image may be a button that is part of the user interface that is a suggestion to help center an object, such as a person in the center of the image. In some embodiments, the command may be based on the user specifying new boundaries for the image by directly extending the boundaries of the image, or on the movement of a selected object that extends the boundaries of the image.
[0103] The inpainter module 206 receives as input an uncropped image and the dimensions of the uncropped image, and outputs an uncropped image that replaces the boundary between the image and the uncropped image with inpainting pixels that match the image. For example, if the boundary is water, the inpainter module 206 may use the water pixels in the uncropped portion of the image.
[0104] 5A shows an example user interface 500 for an initial image 504, including a button 412 for modifying the boundary of the initial image 504. In this example, the user interface 500 includes a first button 510 for editing the image and a second button 512 for modifying the boundary of the image. Other mechanisms for providing a command to uncrop an image are possible. For example, a user may select an edge 506 of the initial image 504 and drag and drop the edge 506 to indicate where the user wants the new boundary to end.
[0105] The user may change the initial image boundaries because the person 508 in the image is not centered in the image and cropping the image to reduce the image on the right would result in the image being too narrow.
[0106] 5B shows an example user interface 515 with an indicator 521 used to expand the boundaries of the initial image 517. In this example, the user clicks and drags the indicator 521 to expand the left boundary of the initial image 517. The expanded region 519 is shown with pixelated content while the restorer module 206 generated the uncropped image.
[0107] In some embodiments, the restorer module 206 receives as input the uncropped image and dimensions for the uncropped image, for example, the dimensions include the length and width of a new border on the left side of the initial image.
[0108] The inpainter module 206 outputs an uncropped image including inpainted pixels between the uncropped and expanded boundaries of the modified image based on the dimensions. In this case, the inpainter module 206 copies pixels of the bushes, rocks, water, and flowers. Figure 5C shows an example user interface 530 of the uncropped image 532 output based on the initial image.
[0109] 5D shows an alternative exemplary user interface 540 that uses a selected person 544 to extend the uncropped boundary of initial image 542. In this alternative example, instead of using an indicator, such as indicator 521 in FIG. 5B , or an edge of the boundary, such as edge 506 in FIG. 5A , to extend the uncropped boundary of initial image 542, the user can select an object in the image and move the object to extend the boundary. In FIG. 5D , the user has moved person 544 to the edge of initial image 542. The new position of person 544 defines a new edge to be generated for the uncropped image, where 555 corresponds to the extended area.
[0110] Exemplary Methods 6 shows an example flowchart of a method 600 for generating a corrected image of a complete object from an incomplete object according to some embodiments described herein. Method 600 may be performed by computing device 200 of FIG. 2. In some embodiments, method 600 is performed by user device 115, media server 101, or performed partially on user device 115 and partially on media server 101.
[0111] 6 may begin at block 602. In block 602, it is determined whether permission to access the initial image is granted by the user. If permission is not granted, method 600 ends. If permission is granted, block 602 may be followed by block 604.
[0112] In block 604, a selection of an incomplete object in the initial image is received, the incomplete object being associated with a first location in the initial image, with an omitted portion of the incomplete object being clipped by a boundary of the initial image or obscured by another object. For example, the incomplete object may be a car clipped by an edge of the initial image. A user may select the incomplete object by clicking on the object, circling the object, accepting a suggestion to modify the object generated by the user interface, etc. Block 604 may be followed by block 606.
[0113] An object mask is generated at block 606, including incomplete object pixels associated with the incomplete object. The incomplete object pixels may be determined by segmenting the initial image to identify pixels associated with the incomplete object within the initial object. Block 606 may be followed by block 608.
[0114] In block 608, the incomplete object pixels associated with the incomplete object are removed from the initial image. Block 608 may be followed by block 610.
[0115] In block 610, an inpainting image is generated that replaces the defective object pixels with inpainting pixels. Block 610 may be followed by block 612.
[0116] In block 612, the object mask, the incomplete object, and the inpainting image are provided as inputs to a diffusion model. Block 612 may be followed by block 614.
[0117] In block 614, the diffusion model outputs the complete object. For example, the diffusion model may receive the incomplete object, a second position where the complete object will be placed in the modified image, and dimensions of the complete object, including resized dimensions if the user resized the incomplete object, or the change from the first position to the second position will cause the incomplete object to be resized. Block 614 may be followed by block 616.
[0118] At block 616, a corrected image is generated by blending one or more versions of the complete object with one or more versions of the inpainted image using the object mask. The complete object is placed in a second location in the corrected image that is different from its first location in the initial image. Continuing with the example above, the complete image may include a version of a complete car. The car may be resized based on being moved from the foreground to the background, and may be reduced in size to account for the reflected distance by being placed in the background.
[0119] In some embodiments, the modified image is a first modified image, and the method further includes receiving a request to remove a selected object from the first modified image, providing the first modified image and the request as input to an object removal model associated with the diffusion model, and using the object removal model to output a second modified image without the selected object. In some embodiments, the modified image is the first modified image, and the method further includes receiving a request to move the selected object from a third position to a fourth position, providing the first modified image and the request as input to an object insertion model associated with the diffusion model, and using the object insertion model to output a second modified image with the selected object at the fourth position based on the request.
[0120] In some embodiments, the method further includes receiving a request to add an additional object to the initial image, outputting the additional object using a diffusion model, and outputting the modified image by blending one or more versions of the additional object with one or more versions of the repaired image using the diffusion model. The additional object is positioned at a third position in the repaired image that is different from the second position in the repaired image. In some embodiments, the request to add the additional object includes a text prompt describing the additional object, and the diffusion model outputs the additional object using generative artificial intelligence.
[0121] In some embodiments, the selected person in the repair image may be resized to account for being moved forward or backward from a first position. For example, the selected person may be smaller or larger than the person in the initial image. In some embodiments, the diffusion model resizes the complete object based on the change from a first position in the initial image to a second position in the repair image.
[0122] Shadows corresponding to the selected people in different positions may also be generated, with the shadows matching the direction of light in the image. In some embodiments, a first object in the initial image is replaced with a second object from the initial image, and the modified image includes the first object replaced with the second object. In some embodiments, the lighting of the repaired image is also modified. For example, the sky may be darkened, such as by brightening the sky or thickening the clouds to reduce illumination. In some embodiments, the method further includes modifying the lighting of the modified image based on the direction of lighting in the modified image and adding shadows to the complete objects.
[0123] 7 shows an example flowchart of a method 700 for outputting an uncropped image from an initial image according to some embodiments described herein. Method 700 may be performed by computing device 200 of FIG. 2. In some embodiments, method 700 is performed by user device 115, media server 101, or partially on user device 115 and partially on media server 101.
[0124] 7 may begin at block 702, where an initial image is displayed in a user interface. Block 702 may be followed by block 704.
[0125] At block 704, a command to uncrop the modified image is received to extend an uncrop boundary of the modified image to an extended boundary, the command being based on at least one action selected from the group consisting of selecting an uncrop button, moving an indicator to define the extended boundary, moving an edge of the initial image to define the extended boundary, moving a selected object to define the extended boundary, and combinations thereof. Block 704 may be followed by block 706.
[0126] At block 706, based on the command, an uncropped image is output that includes the repair pixels between the uncropped boundary and the extended boundary of the modified image.
[0127] 8 shows an example flowchart of a method 800 for training an object removal model according to some embodiments described herein. Method 800 may be performed by computing device 200 of FIG. 2. In some embodiments, method 800 is performed by user device 115, media server 101, or partially on user device 115 and partially on media server 101.
[0128] Method 800 may begin at block 802. In block 802, for each counterfactual image pair, counterfactual training data is generated by capturing a factual image that includes an object in a scene, physically removing the object while avoiding camera motion, lighting changes, or movement of other objects, capturing a counterfactual image of the scene without the object, and segmenting the factual image to create an object mask. Block 802 may be followed by block 804.
[0129] In block 804, for each counterfactual image pair, a combined image is created that includes the factual image and the object mask and the counterfactual image. Block 804 may be followed by block 806.
[0130] At block 806, a diffusion model is trained, including an object removal model, based on using counterfactual image pairs.
[0131] 9 shows an example flowchart of a method 900 for training an object insertion model according to some embodiments described herein. Method 900 may be performed by computing device 200 of FIG. 2. In some embodiments, method 900 is performed by user device 115, media server 101, or partially on user device 115 and partially on media server 101.
[0132] Method 900 may begin at block 902. In block 902, for each counterfactual image pair, counterfactual training data is generated by capturing a factual image that includes an object in a scene, physically removing the object while avoiding camera motion, lighting changes, or movement of other objects, capturing a counterfactual image of the scene without the object, and segmenting the factual image to create an object mask. Block 902 may be followed by block 904.
[0133] In block 904, for each counterfactual image pair, a combined image is created that includes the factual image and the object mask and the counterfactual image. Block 904 may be followed by block 906.
[0134] In block 906, a diffusion model is trained, including an object removal model, based on using counterfactual image pairs. Block 906 may be followed by block 908.
[0135] In block 908, for each synthetic image pair, synthetic training data is generated by selecting an original image that contains an object, using an object removal model to output a modified image from the original image without the object, generating an input image by inserting the object into the modified image, and segmenting the original image to create an object mask. Block 908 may be followed by block 910.
[0136] For each composite pair, a second combined image is created that includes the original image, the object mask, and the input image in block 910. Block 910 may be followed by block 912.
[0137] At block 912, a diffusion model is pre-trained to include an object insertion model based on using synthetic image pairs. Block 912 may be followed by block 914.
[0138] At block 914, the diffusion model is refined to include an object insertion model based on using counterfactual image pairs.
[0139] In addition to the above, a user may be provided with controls that allow the user to choose both whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information regarding the user's social networks, social actions, or activities, occupation, user preferences, or the user's current location), as well as whether the user is sent content or communications from the server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, a user's identity may be processed such that personally identifiable information cannot be determined about the user, or if location information is obtained (e.g., to the city, zip code, or state level), the user's geographic location may be generalized so that the user's specific location cannot be determined. Thus, a user may control what information is collected about them, how that information is used, and what information is provided to them.
[0140] In the foregoing description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the present specification. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In some instances, structures and devices are shown in block diagram form to avoid obscuring the description. For example, the present embodiments may be described above primarily with reference to a user interface and specific hardware. However, the embodiments may be applied to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.
[0141] A reference herein to "some embodiments" or "some examples" means that a particular feature, structure, or characteristic described in connection with an embodiment or example may be included in at least one embodiment of the description. The appearances of the phrase "in some embodiments" in various places in the specification do not necessarily all refer to the same embodiment.
[0142] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. These steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0143] It should be recognized, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. As will become apparent from the discussion that follows, unless specifically stated otherwise, throughout this specification, discussions utilizing terms including "processing" or "calculating" or "computing" or "determining" or "displaying" will be understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities in the computer system's registers and memory into other data that are similarly represented as physical quantities in the computer system's memory or registers or other such information storage, transmission, or display device.
[0144]
[0013] Embodiments herein may also relate to a processor for performing one or more steps of the above-described methods. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored on a non-transitory computer-readable storage medium, including, but not limited to, an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory including a USB key having non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0145] This specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, this specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.
[0146] Furthermore, the descriptions may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer-usable or computer-readable medium may be any device that can store, preserve, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0147] A data processing system suitable for storing or executing program code includes at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.
Claims
1. 1. A computer-implemented method comprising: receiving a selection of an incomplete object in an initial image, the incomplete object being associated with a first location in the initial image, an omitted portion of the incomplete object being clipped by a boundary of the initial image or obscured by another object, the method further comprising: generating an object mask including incomplete object pixels associated with the incomplete object; removing the incomplete object pixels associated with the incomplete object from the initial image; generating an inpainting image in which the imperfect object pixels corresponding to the imperfect object are replaced with inpainting pixels; providing the object mask, the incomplete object, and the inpainting image as inputs to a diffusion model; outputting a complete object using the diffusion model; and generating a modified image by blending one or more versions of the complete object with one or more versions of the inpainted image using the object mask, wherein the complete object is positioned at a second position in the modified image that is different from the first position in the initial image.
2. the modified image is a first modified image; receiving a request to remove a selected object from the first modified image; providing the first modified image and the request as inputs to an object removal model associated with the diffusion model; outputting a second modified image that does not include the selected object using the object removal model; and The method of claim 1 further comprising:
3. the modified image is a first modified image; receiving a request to move the selected object from a third position to a fourth position; providing the first modified image and the request as inputs to an object insertion model associated with the diffusion model; outputting a second modified image including the selected object at the fourth location based on the request using the object insertion model; The method of claim 1 further comprising:
4. the modified image is a first modified image; receiving a request to add an additional object to the initial image; outputting the additional object using the diffusion model; outputting a second modified image by blending one or more versions of the additional object with one or more versions of the inpainted image using the diffusion model; The method of claim 1 further comprising:
5. The method of claim 4 , wherein the request to add the additional object includes a text prompt that describes the additional object.
6. The method of claim 1 , wherein the complete object is resized based on a change from the first position in the initial image to the second position in the modified image.
7. receiving a command to uncrop the modified image such that an uncrop boundary of the modified image is extended to an extended boundary; outputting an uncropped image containing inpainted pixels between the uncropped boundary and the extended boundary of the modified image based on the command; The method of claim 1 further comprising:
8. The command to uncrop the repaired image comprises: Select the Uncrop button and a command to directly extend the uncropped boundary of the modified image to the extended boundary, or a movement of the complete object that extends the uncropped boundary of the modified image to the extended boundary; The method of claim 7, comprising:
9. modifying the illumination of the modified image; adding a shadow to the complete object based on the direction of the lighting in the modified image; The method of claim 1 further comprising:
10. 1. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: receiving a selection of an incomplete object in an initial image, the incomplete object being associated with a first location in the initial image, an omitted portion of the incomplete object being clipped by a boundary of the initial image or obscured by another object; generating an object mask including incomplete object pixels associated with the incomplete object; removing the incomplete object pixels associated with the incomplete object from the initial image; generating an inpainting image in which the imperfect object pixels corresponding to the imperfect object are replaced with inpainting pixels; providing the object mask, the incomplete object, and the inpainting image as inputs to a diffusion model; outputting a complete object using the diffusion model; and generating a modified image by blending one or more versions of the complete object with one or more versions of the inpainted image using the object mask, wherein the complete object is positioned at a second position in the modified image that is different from the first position in the initial image.
11. The modified image is a first modified image, and the operation comprises: receiving a request to remove a selected object from the first modified image; providing the first modified image and the request as inputs to an object removal model associated with the diffusion model; outputting a second modified image that does not include the selected object using the object removal model; and The non-transitory computer-readable medium of claim 10 further comprising:
12. The modified image is a first modified image, and the operation comprises: receiving a request to move the selected object from a third position to a fourth position; providing the first modified image and the request as inputs to an object insertion model associated with the diffusion model; outputting a second modified image including the selected object at the fourth location based on the request using the object insertion model; The non-transitory computer-readable medium of claim 10 further comprising:
13. The modified image is a first modified image, and the operation comprises: receiving a request to add an additional object to the initial image, the request including a text prompt describing the additional object, the operation further comprising: outputting the additional object using the diffusion model; outputting a second modified image by blending one or more versions of the additional object with one or more versions of the inpainted image using the diffusion model; The non-transitory computer-readable medium of claim 10 further comprising:
14. The operation is receiving a command to uncrop the modified image such that an uncrop boundary of the modified image is extended to an extended boundary; outputting an uncropped image containing inpainted pixels between the uncropped boundary and the extended boundary of the modified image based on the command; The non-transitory computer-readable medium of claim 10 further comprising:
15. 1. A system comprising: a processor; a memory coupled to the processor and storing instructions that, when executed by the processor, cause the processor to perform operations, the operations including: receiving a selection of an incomplete object in an initial image, the incomplete object being associated with a first location in the initial image, an omitted portion of the incomplete object being clipped by a boundary of the initial image or obscured by another object, the operations comprising: generating an object mask including incomplete object pixels associated with the incomplete object; removing the incomplete object pixels associated with the incomplete object from the initial image; generating an inpainting image in which the imperfect object pixels corresponding to the imperfect object are replaced with inpainting pixels; providing the object mask, the incomplete object, and the inpainting image as inputs to a diffusion model; outputting a complete object using the diffusion model; and generating a modified image by blending one or more versions of the complete object with one or more versions of the inpainted image using the object mask, wherein the complete object is positioned at a second position in the modified image that is different from the first position in the initial image.
16. The modified image is a first modified image, and the operation comprises: receiving a request to remove a selected object from the first modified image; providing the first modified image and the request as inputs to an object removal model associated with the diffusion model; outputting a second modified image that does not include the selected object using the object removal model; and The system of claim 15 further comprising:
17. The modified image is a first modified image, and the operation comprises: receiving a request to move the selected object from a third position to a fourth position; providing the first modified image and the request as inputs to an object insertion model associated with the diffusion model; outputting a second modified image including the selected object at the fourth location based on the request using the object insertion model; The system of claim 15 further comprising:
18. The modified image is a first modified image, and the operation comprises: receiving a request to add an additional object to the initial image, the request including a text prompt describing the additional object, the operation further comprising: outputting the additional object using the diffusion model; outputting a second modified image by blending one or more versions of the additional object with one or more versions of the inpainted image using the diffusion model; The system of claim 15 further comprising:
19. The system of claim 15 , wherein the complete object is resized based on a change from the first position in the initial image to the second position in the modified image.
20. The operation is receiving a command to uncrop the modified image such that an uncrop boundary of the modified image is extended to an extended boundary; outputting an uncropped image containing inpainted pixels between the uncropped boundary and the extended boundary of the modified image based on the command; The system of claim 15 further comprising:
Citation Information
Patent Citations
Denoising diffusion generative adversarial networks
US20230095092A1
Image-to-Image Mapping by Iterative De-Noising
US20230103638A1