Rearranging, replacing, and generating objects within an image.
The diffusion model addresses the challenges of object repositioning in images by generating and blending complete objects, ensuring high-quality, realistic results without pixel misidentification or empty spaces.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2026-03-27
AI Technical Summary
Existing image editing techniques fail to seamlessly move objects within an image without causing pixel misidentification, creating out-of-place empty spaces, or resulting in poorly edited images.
Utilizing a diffusion model to generate a complete object by removing incomplete object pixels, replacing them with restored pixels, and blending the complete object with a modified image, allowing for seamless repositioning and resizing while maintaining a realistic appearance.
The diffusion model effectively corrects images with incomplete objects, producing high-quality, error-free results by ensuring the complete object fits seamlessly into its new position, maintaining a realistic appearance.
Smart Images

Figure 0007836940000003 
Figure 0007836940000004 
Figure 0007836940000005
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 465,230, filed on 9 May 2023, entitled “Repositioning Objects in an Image,” and to U.S. Provisional Patent Application No. 63 / 562,634, filed on 7 March 2024, entitled “Performing Scene Impact Editing Tasks Using Diffusion Neural Networks,” and these U.S. Provisional Patent Applications are incorporated herein by reference as a whole. [Background technology]
[0002] Users may capture images where objects are in undesirable positions. For example, an object may be cut off by the image's boundaries or by other objects. Techniques exist for moving objects within an image. However, attempts to move objects to different positions within an image can yield disastrous results. For example, pixels associated with an object may be improperly identified so that part of the object remains in its original position while the rest is moved to a different position (e.g., a chicken's body is moved while its legs remain). In another example, the empty space created by deleting pixels associated with the moved object may be filled with pixels that look out of place. In yet another example, pixels surrounding the moved object may appear different from the background, resulting in an image that looks poorly edited.
[0003] The background art provided herein is for general purposes only to illustrate the context of this disclosure. The inventors' works currently attributed are not expressly or implicitly recognized as prior art to this disclosure, to the extent described in this background art section, as are the aspects of this specification that may not be considered prior art at the time of filing. [Overview of the project]
[0004] A computer implementation method includes receiving a selection of incomplete objects in an initial image, wherein the incomplete objects are associated with a first position in the initial image, and the missing portion of the incomplete objects is cut off by the boundaries of the initial image or obscured by other objects. The method further includes generating an object mask containing the incomplete object pixels associated with the incomplete objects. The method further includes removing the incomplete object pixels associated with the incomplete objects from the initial image. The method further includes generating a restored image in which the incomplete object pixels corresponding to the incomplete objects are replaced with restored pixels. The method further includes providing the object mask, the incomplete objects, and the restored image as input to a diffusion model. The method further includes outputting a complete object using the diffusion model. The method further includes generating a modified image by blending one or more versions of the complete object with one or more versions of the restored image using the object mask, wherein the complete object is located at a second position in the modified image different from a first position in the initial image.
[0005] In some embodiments, the modified image is a first modified image, and the method further includes receiving a request to remove a selected object from the first modified image, providing the first modified image and the request as input to an object removal model associated with a diffusion model, and using the object removal model to output a second modified image that does not include the selected object. In some embodiments, the modified image is a first modified image, and the method further includes receiving a request to move a selected object from a third position to a fourth position, providing the first modified image and the request as input to an object insertion model associated with a diffusion model, and using the object insertion model to output a second modified image that includes the selected object at the fourth position based on the request.
[0006] In some embodiments, the modified image is a first modified image, and the method further includes receiving a request to add additional objects to the initial image, outputting the additional objects using a diffusion model, and outputting a second modified image by blending one or more versions of the additional objects with one or more versions of the restored image using a diffusion model. In some embodiments, the request to add additional objects includes a text prompt describing the additional objects.
[0007] In some embodiments, the complete object is resized based on a change from a first position in the initial image to a second position in the modified image. In some embodiments, the method further includes receiving a command to uncrop the modified image so that the uncrop boundary of the modified image extends to the extended boundary, and outputting an uncropped image containing repaired pixels between the uncrop boundary and the extended boundary of the modified image based on the command. In some embodiments, the command to uncrop the modified image includes selecting an uncrop button and either a command to directly extend the uncrop boundary of the modified image to the extended boundary, or moving the complete object to extend the uncrop boundary of the modified image to the extended boundary. In some embodiments, the method further includes correcting the lighting of the modified image and adding shadows to the complete object based on the direction of the lighting of the modified image.
[0008] In some embodiments, a non-temporary computer-readable medium that stores instructions causing one or more processors to perform an operation when executed by one or more processors. The operation includes receiving a selection of incomplete objects in an initial image, wherein the incomplete objects are associated with a first position in the initial image, and the missing portion of the incomplete objects is cut off by the boundary of the initial image or obscured by other objects; generating an object mask containing the incomplete object pixels associated with the incomplete objects; removing the incomplete object pixels associated with the incomplete objects from the initial image; generating a restored image in which the incomplete object pixels corresponding to the incomplete objects are replaced with restored pixels; providing the object mask, the incomplete objects, and the restored image as input to a diffusion model; outputting a complete object using the diffusion model; and generating a modified image by blending one or more versions of the complete object with one or more versions of the restored image using the object mask, wherein the complete object is located at a second position in the modified image different from a first position in the initial image.
[0009] In some embodiments, the modified image is a first modified image, and the operation further includes receiving a request to remove a selected object from the first modified image, providing the first modified image and the request as input to an object removal model associated with the diffusion model, and using the object removal model to output a second modified image that does not include the selected object. In some embodiments, the modified image is a first modified image, and the operation further includes receiving a request to move a selected object from a third position to a fourth position, providing the first modified image and the request as input to an object insertion model associated with the diffusion model, and using the object insertion model to output a second modified image that includes the selected object at the fourth position based on the request.
[0010] In some embodiments, the corrected image is a first corrected image, and the operation further includes receiving a request to add additional objects to an initial image, the request including a text prompt describing the additional objects; outputting the additional objects using a diffusion model; and outputting a second corrected image by blending one or more versions of the additional objects with one or more versions of the repaired image using a diffusion model. In some embodiments, the complete objects are resized based on a change from a first position in the initial image to a second position in the corrected image. In some embodiments, the operation further includes receiving a command to uncrop the corrected image so that the uncropped boundary of the corrected image is extended to an extended boundary; and outputting an uncropped image containing repaired pixels between the uncropped boundary and the extended boundary of the corrected image based on the command.
[0011] In some embodiments, the system includes a processor and a memory coupled to the processor, the memory storing instructions, and when an instruction is executed by the processor, the memory causes the processor to perform an operation. The operation includes receiving a selection of incomplete objects in an initial image, wherein the incomplete objects are associated with a first position in the initial image, and the missing portion of the incomplete objects is cut off by the boundary of the initial image or obscured by other objects; generating an object mask containing the incomplete object pixels associated with the incomplete objects; removing the incomplete object pixels associated with the incomplete objects from the initial image; generating a restored image in which the incomplete object pixels corresponding to the incomplete objects are replaced with restored pixels; providing the object mask, the incomplete objects, and the restored image as input to a diffusion model; outputting a complete object using the diffusion model; and generating a modified image by blending one or more versions of the complete object with one or more versions of the restored image using the object mask, wherein the complete object is located at a second position in the modified image different from a first position in the initial image.
[0012] In some embodiments, the modified image is a first modified image, and the operation further includes receiving a request to remove a selected object from the first modified image, providing the first modified image and the request as input to an object removal model associated with the diffusion model, and using the object removal model to output a second modified image that does not include the selected object. In some embodiments, the modified image is a first modified image, and the operation further includes receiving a request to move a selected object from a third position to a fourth position, providing the first modified image and the request as input to an object insertion model associated with the diffusion model, and using the object insertion model to output a second modified image that includes the selected object at the fourth position based on the request.
[0013] In some embodiments, the corrected image is a first corrected image, and the operation further includes receiving a request to add additional objects to an initial image, the request including a text prompt describing the additional objects; outputting the additional objects using a diffusion model; and outputting a second corrected image by blending one or more versions of the additional objects with one or more versions of the repaired image using a diffusion model. In some embodiments, the complete objects are resized based on a change from a first position in the initial image to a second position in the corrected image. In some embodiments, the operation further includes receiving a command to uncrop the corrected image so that the uncropped boundary of the corrected image is extended to an extended boundary; and outputting an uncropped image containing repaired pixels between the uncropped boundary and the extended boundary of the corrected image based on the command. [Brief explanation of the drawing]
[0014] [Figure 1] This is a block diagram of an exemplary network environment according to some embodiments described herein. [Figure 2]This is a block diagram of an exemplary computing device according to some embodiments described herein. [Figure 3A] Exemplary initial images are shown of some embodiments described herein. [Figure 3B] Figure 3A shows an exemplary initial image in which two objects were selected for modification, according to some embodiments described herein. [Figure 3C] The following are exemplary initial images showing an object surrounded by a mask, according to some embodiments described herein. [Figure 3D] The following are exemplary restored images, in which bystander objects have been removed, subject objects have been moved, and the images have been resized, according to some embodiments described herein. [Figure 3E] The following are exemplary restoration images, according to some embodiments described herein, in which roads are replaced with grass, first type trees are replaced with second type trees, and cloudy skies are replaced with clear skies. [Figure 4A] An exemplary initial image of a child sitting on a bench and holding a balloon, partially cropped by the boundary of the initial image, is shown according to some embodiments described herein. [Figure 4B] The following are exemplary modified images showing the child, bench, and balloon moved to a second position according to some embodiments described herein. [Figure 5A] This specification shows an exemplary user interface for an initial image, including buttons for changing the boundaries of the initial image, according to several embodiments described herein. [Figure 5B] This specification describes an exemplary user interface with an indicator used to broaden the boundaries of an initial image, according to several embodiments described herein. [Figure 5C] This specification shows an exemplary user interface for an uncropped image output based on an initial image, according to several embodiments described herein. [Figure 5D]An alternative exemplary user interface is shown in which a selected object is used to expand the boundaries of an initial image, according to some embodiments described herein. [Figure 6] An exemplary flowchart of a method for generating a corrected image of a complete object from an incomplete object is shown, according to some embodiments described herein. [Figure 7] An exemplary flowchart of a method for outputting an uncropped image from an initial image is shown, according to some embodiments described herein. [Figure 8] An exemplary flowchart of a method for training an object removal model is shown, according to some embodiments described herein. [Figure 9] An exemplary flowchart of a method for training an object insertion model is shown, according to some embodiments described herein. **DETAILED DESCRIPTION OF THE INVENTION**
[0015] A user may capture an image in which an object is in an undesirable position. For example, the object may be cut off by the boundaries of the image, or by other objects. There are techniques for moving an object within an image. However, attempts to move an object to a different position within the image can result in various undesirable outcomes. For example, pixels associated with the object may be inappropriately identified such that part of the object remains in its original position while the remaining part of the object is moved to a different position (e.g., the body of a chicken is moved while the chicken's legs remain). In other examples, empty space created by deleting pixels associated with the moved object may be filled with out-of-place pixels. In still other examples, pixels surrounding the moved object may appear different from the background and result in an image that appears inadequately edited.
[0016] The technique described below advantageously solves these problems by providing incomplete objects as input to a diffusion machine learning model, referred to herein as a diffusion model, and outputting complete objects. An incomplete object is a partial representation of an object in an image. The object is partially present (not completely present) in the image. The portion of an object that is not present in the image is referred herein to as the "omitted portion" of the incomplete object. The user can select an incomplete object in the initial image and move its position.
[0017] The space left by the incomplete object is restored with restoration pixels to form a restored image. The complete object is a complete representation of the object, including the incomplete object and the missing parts of the incomplete object. The restored image differs from the initial image in that the incomplete object pixels associated with the incomplete object are removed from the restored image, and the incomplete object pixels are replaced with restoration pixels that can be selected based on their proximity to surrounding pixels, selected from a reference image including background pixels, etc.
[0018] The diffusion model ensures that the perfect object fits into its new position. For example, if an object is moved from the background to the foreground, the diffusion model increases the size of the moved object. The diffusion model uses an object mask to blend one or more versions of the perfect object with one or more versions of the restored image, outputting a modified image in which the perfect object is seamlessly merged with the restored image. For example, the diffusion model can blend progressively noisier versions of the perfect object with the corresponding noisier versions of the restored image, but it can also generate denoised versions of the perfect object and corresponding denoised versions of the restored image. Noisier versions of the perfect object are created by increasing the entropy of the image, and the more noisy they are, the less discernible the details of the perfect object in the image become. Similarly, noisier versions of the restored image are created by increasing the entropy of the restored image, and the more noisy they are, the less discernible the details of the restored image become.
[0019] By using a diffusion model instead of other machine learning models, media applications can maintain a realistic appearance of corrected images under a wide variety of circumstances. The techniques described below enable the correction of flawed images, i.e., images containing incomplete objects, in an efficient manner. Complete objects created by utilizing a diffusion model are of high quality and free from the aforementioned errors. The image processing described herein effectively and efficiently corrects images with respect to incomplete objects present within them.
[0020] Exemplary Environment 100 Figure 1 shows a block diagram of an exemplary environment 100. In some embodiments, the environment 100 includes a media server 101, user device 115a, and user device 115n, all coupled to a network 105. Users 125a and 125n may be associated with their respective user devices 115a and 115n. In some embodiments, the environment 100 may include other servers or devices not shown in Figure 1. In Figure 1 and the remaining figures, letters following a reference number, such as "115a," indicate a reference to the element having that particular reference number. Reference numbers in the text without following letters, such as "115," indicate a general reference to embodiments of the element prefixed with that reference number.
[0021] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicably coupled to the network 105 via signal lines 102. The signal lines 102 may be a wired connection such as Ethernet®, coaxial cable, or fiber optic cable, or a wireless connection such as Wi-Fi®, Bluetooth®, or other wireless technology. In some embodiments, the media server 101 sends and receives data to and from one or more user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.
[0022] Database 199 may store machine learning models, training datasets, images, etc. Database 199 may also store social network data associated with user 125, user preferences of user 125, etc.
[0023] The user device 115 may be a computing device that includes memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or other electronic device that can access the network 105.
[0024] In the illustrated embodiment, user device 115a is connected to network 105 via signal line 108, and user device 115n is connected to network 105 via signal line 110. Media application 103 may be stored as media application 103b on user device 115a and / or as media application 103c on user device 115n. Signal lines 108 and 110 may be wired connections such as Ethernet, coaxial cable, or fiber optic cable, or wireless connections such as Wi-Fi®, Bluetooth®, or other wireless technologies. User devices 115a and 115n are accessed by users 125a and 125n, respectively. User devices 115a and 115n in Figure 1 are used as examples. Although Figure 1 shows two user devices, 115a and 115n, this disclosure applies to system architectures having one or more user devices 115.
[0025] The media application 103 may be stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are performed on the media server 101 or the user device 115. In some embodiments, some operations may be performed on the media server 101 and some on the user device 115. The execution of operations is subject to user settings. For example, user 125a may specify that operations are performed on each device 115a and not on the media server 101. Such a setting would cause the operations described herein to be performed entirely on the user device 115a and not on the media server 101. Furthermore, user 125a may specify that user images and / or other data be stored only locally on the user device 115a and not on the media server 101. Such a setting would cause user data not to be sent to or stored on the media server 101. The transmission of user data to the media server 101, any temporary or persistent storage of such data by the media server 101, and the performance of actions on such data by the media server 101 are performed only if the user consents to the transmission, storage, and performance of actions by the media server 101. The user is provided with the option to change settings at any time, for example, so that the user can enable or disable the use of the media server 101.
[0026] Machine learning models (e.g., diffusion models, neural networks, or other types of models) are stored locally on the user device 115 and used when used for one or more operations, with the permission of the specific user. Server-side models are used only when permitted by the user. Furthermore, trained models may be provided for use on the user device 115. During such use, on-device training of the model may be performed if permitted by user 125. Updated model parameters may be sent to the media server 101, for example, to enable federative learning, if permitted by user 125. Model parameters do not include any user data.
[0027] The media application 103 receives an initial image. For example, the media application 103 receives an initial image from a camera which is part of the user device 115, or the media application 103 receives an initial image via the network 105. The media application 103 receives a selection of an incomplete object in the initial image. The incomplete object is associated with a first position in the initial image, and the missing portion of the incomplete object is cut off by the boundary of the initial image or covered by other objects. An incomplete object may be selected, for example, when the user 125 taps an object, draws a shape (e.g., a circle) around the object, and acknowledges a suggestion from the media application 103 to modify the object.
[0028] The media application 103 generates an object mask containing the incomplete object pixels associated with the incomplete object, and removes the incomplete object pixels associated with the incomplete object from the initial image. The media application 103 generates a restored image in which the incomplete object pixels corresponding to the incomplete object are replaced with restored pixels.
[0029] Media application 103 outputs a complete object using a diffusion model. For example, if an incomplete object is cut off by the edges of the initial image and the incomplete object is moved to the center of the initial image, the diffusion model outputs a complete object that fills in the missing portion of the incomplete object. Media application 103 outputs a modified image by using the diffusion model and an object mask to blend one or more versions of the complete object with one or more versions of the modified image. The complete object is positioned at a second location in the modified image, different from a first location in the initial image. In some embodiments, the modified image may include a watermark or other indicator to identify that the modified image was generated using a machine learning model.
[0030] In some embodiments, the media application 103 can be implemented using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, the media application 103a can be implemented using a combination of hardware and software.
[0031] Exemplary computing device 200 Figure 2 is a block diagram of an exemplary computing device 200 that may be used to implement one or more features described herein. The computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a media server 101 used to implement a media application 103a. In another example, the computing device 200 is a user device 115.
[0032] In some embodiments, the computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, display 241, camera 243, and storage device 245, all coupled via a bus 218. The processor 235 may be coupled to the bus 218 via a signal line 222, the memory 237 may be coupled to the bus 218 via a signal line 224, the I / O interface 239 may be coupled to the bus 218 via a signal line 226, the display 241 may be coupled to the bus 218 via a signal line 228, the camera 243 may be coupled to the bus 218 via a signal line 230, and the storage device 245 may be coupled to the bus 218 via a signal line 232.
[0033] The processor 235 may be one or more processors and / or processing circuits that execute program code and control the basic operation of the computing device 200. "Processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. The processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., single-core, dual-core, or multi-core configuration), multiple processing units (e.g., a multi-processor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a composite programmable logic device (CPLD), a dedicated circuit for realizing a function, a dedicated processor for implementing processing based on a neural network model, a neural circuit, a processor optimized for matrix calculations (e.g., matrix multiplication), or a system having other systems. In some embodiments, the processor 235 may include one or more coprocessors that perform neural network processing. In some embodiments, the processor 235 may be a processor that processes data to produce a probabilistic output, for example, the output produced by the processor 235 may be inaccurate or accurate within a range from an expected output. Processing does not need to be limited to a specific geographical location or have temporal constraints. For example, a processor can perform its functions in real time, offline, or batch mode. Parts of the processing can be performed by different (or the same) processing systems at different times and in different locations. A computer can be any processor that communicates with memory.
[0034] Memory 237 is provided within the computing device 200 for access by the processor 235 and can be any suitable processor-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), or flash memory, which is suitable for storing instructions for execution by the processor or a set of processors, and is located separately from and / or integrally with the processor 235. Memory 237 can store software that runs on the computing device 200 by the processor 235, including media applications 103.
[0035] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, and the like. One or more methods disclosed herein may operate in several environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application having web pages, as a mobile application ("App") that runs on a mobile computing device, and the like.
[0036] Application data 266 may be data generated by other applications 264 or the hardware of computing device 200. For example, application data 266 may include images used by an image library application and user actions identified by other applications 264 (e.g., a social networking application).
[0037] The I / O interface 239 can provide functionality that enables the computing device 200 to interface with other systems and devices. Interfaced devices may be included as part of the computing device 200, or they may be separate and communicate with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices can communicate via the I / O interface 239. In some embodiments, the I / O interface 239 can connect to interface devices such as input devices (keyboards, pointing devices, touchscreens, microphones, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, monitors, etc.).
[0038] Some examples of interface devices that can be connected to the I / O interface 239 include a display 241 that can be used to display content, such as images, videos, and / or user interfaces of output applications described herein, and to receive touch (or gesture) input from the user. For example, the display 241 may be used to display a user interface, including graphical guides, on a viewfinder. The display 241 may include any suitable display device such as a liquid crystal display (LCD), light-emitting diode (LED), or plasma display screen, cathode ray tube (CRT), television, monitor, touchscreen, three-dimensional display screen, or other visual display device. For example, the display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen on a computer device.
[0039] Camera 243 can be any type of image capture device capable of capturing images and / or video. In some embodiments, camera 243 captures images or video that the I / O interface 239 transmits to the media application 103.
[0040] The storage device 245 stores data related to the media application 103. For example, the storage device 245 can store a training dataset that includes labeled images, machine learning models, and outputs from the machine learning models.
[0041] Figure 2 shows an exemplary media application 103, which includes a user interface module 202, a segmenter 204, a repairer module 206, and a diffusion module 208, all stored in memory 237.
[0042] The user interface module 202 generates graphic data for displaying a user interface that includes an image. In some embodiments, the user interface module 202 receives an initial image. The initial image may be received from the camera 243 of the computing device 200 or from the media server 101 via the I / O interface 239. The initial image includes a subject, such as a person. In some embodiments, the user interface includes options for selecting various people and other objects in the initial image. For example, the user may select a person by tapping on them, circling an object, or painting an object with a brush. In some embodiments, the user interface generates suggestions for modifying the image, such as displaying text asking if the user wants to remove bystanders from the objects.
[0043] In some embodiments, when a user selects an object, the user interface module 202 updates the graphic data to include a highlighted version of the selected object. The user can modify the selected object. For example, the user may drag and drop the selected object from a first position to a second position, resize an image, or select a button to erase an image.
[0044] In some embodiments, the segmentator 204 generates a segmentation score that reflects the quality of identification of pixels associated with selected objects in the initial image. The user interface may include various options for modifying the selected objects based on the segmentation score. For example, if the segmentation score exceeds a threshold, the user interface module 202 provides options to move the selected objects, replace the selected objects with different objects, or erase the selected objects. In other examples, if the segmentation score does not exceed a threshold, the user interface module 202 does not provide an option to move the selected objects, but provides options to replace or erase the selected objects.
[0045] In some embodiments, the user interface module 202 generates graphic data for displaying the restored image, moves selected objects from a first position to a second position in the image, resizes the selected image, adds additional objects, and so on. The user interface may also include options for editing the restored image, sharing the restored image, adding the restored image to a photo album, and so on.
[0046] The segmenter 204 segments the selected objects from the initial image by identifying the pixels corresponding to the selected objects. In some embodiments, the segmenter 204 uses an alpha map as part of a technique to distinguish the foreground and background of the initial image during segmentation. The segmenter 204 may also identify the texture of the selected objects within the foreground of the initial image. In some embodiments, the segmenter 204 generates a segmentation map that identifies the pixels associated with one or more objects in the initial image. For example, the segmentation map may include identification of the pixels associated with the selected objects.
[0047] The segmentator 204 can perform segmentation by detecting objects in the initial image. Objects may include people, animals, cars, buildings, etc. People may be subjects of the initial image or not subjects of the initial image (i.e., bystanders). Bystanders may include people walking, running, riding bicycles, standing behind subjects, or otherwise entering the initial image. In different examples, bystanders may be in the foreground (e.g., someone crossing in front of the camera), at the same depth as a subject (e.g., someone standing next to a subject), or in the background. In some examples, there may be multiple bystanders in the initial image. Bystanders may be people in any pose, e.g., standing, sitting, crouching, lying down, jumping, etc. Bystanders may be facing the camera, at an angle to the camera, or with their faces turned away from the camera.
[0048] The segmentator 204 can perform object recognition to identify the expected shape of the object in order to determine whether a pixel is associated with the selected object or with the background, and can detect the type of the object by comparing the object to a preceding object such as a person, vehicle, or building. The segmentator 204 can generate a region of interest of the selected object, such as a bounding box with x, y coordinates and scale.
[0049] The segmentator 204 generates one or more object masks for one or more selected objects in the initial image. These object masks represent regions of interest. Object masks are described in more detail below with reference to the diffusion model.
[0050] In some embodiments, one or more object masks are generated based on generating superpixels in an image and matching the centroids of the superpixels with depth map values (e.g., values obtained by camera 243 using a depth sensor, or by deriving depth from pixel values) with depth-based cluster detection. More specifically, a depth range may be determined using the depth values of the masked region, and superpixels that fall within the depth range may be identified. Other techniques for generating masks include weighting depth values based on how close the depth values are to an object mask represented by a distance transformation map.
[0051] In some embodiments, the segmenter 204 segments an initial image and generates an object mask using a machine learning algorithm such as a neural network, or more specifically, a convolutional neural network. The segmenter 204 can specify a circuit configuration (e.g., for a programmable processor, for a field-programmable gate array (FPGA), etc.) that enables the processor 235 to apply the machine learning model. In some embodiments, the segmenter 204 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the segmenter 204 may provide an application programming interface (API) that can be used by an operating system 262 and / or other applications 264 to call the segmenter 204, for example, to apply a machine learning model to application data 266 and output the object mask.
[0052] The segmentator 204 uses the training data to generate a trained machine learning model. For example, the training data may include pairs of initial images with one or more objects and output images with one or more corresponding object masks.
[0053] Training data may be obtained from any source, such as a data repository marked specifically for training, or data that has been granted permission to be used as training data for machine learning. In some embodiments, training may be performed on a media server 101 that directly provides training data to the user device 115, training may be performed locally on the user device 115, or a combination of both.
[0054] In some embodiments, the segmenter 204 uses unedited / untransferred weights obtained from other applications. For example, in these embodiments, the trained model may be generated on a different device and provided as part of the segmenter 204. In various embodiments, the trained model may be provided as a data file containing the model structure or format (e.g., defining the number and type of neural network nodes, connectivity between nodes, and the organization of nodes into multiple layers) and associated weights. The segmenter 204 may read the data file of the trained model and implement a neural network with node connectivity, layers, and weights based on the model structure or format specified in the trained model.
[0055] A trained machine learning model may include one or more model forms or structures. For example, the model forms or structures can include any type of neural network, such as linear networks, deep learning neural networks that implement multiple layers (for example, with a "hidden layer" between the input and output layers, where each layer is a linear network), convolutional neural networks (for example, a network that divides or partitions input data into multiple parts or tiles, processes each tile individually using one or more neural network layers, and aggregates the results of processing each tile), and sequence-to-sequence neural networks (for example, a network that takes sequential data such as words in a sentence or frames in a video as input and outputs a resulting sequence).
[0056] The model format or structure can specify the connectivity between various nodes and the organization of the nodes into layers. For example, the nodes in the first layer (e.g., the input layer) may receive data as input data or application data. Such data may include, for example, one or more pixels per node, if the trained model is used, for example, to analyze an initial image. Subsequent intermediate layers may receive the outputs of the nodes in the previous layer as input, according to the connectivity specified in the model format or structure. These layers are sometimes also called hidden layers. For example, the first layer may output a segmentation between the foreground and background. The final layer (e.g., the output layer) produces the output of the machine learning model. For example, the output layer may receive a segmentation of the initial image into foreground and background and output whether a pixel is part of an object mask or the rest of the initial image. In some embodiments, the model format or structure also specifies the number and / or type of nodes in each layer.
[0057] In different embodiments, a trained model may include one or more models. One or more of the models may include multiple nodes arranged in layers according to the model structure or form. In some embodiments, a node may be a memoryless computation node configured to process one unit of input and produce one unit of output. The computation performed by the node may include, for example, multiplying each of the multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output. In some embodiments, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, computations by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multicore processor, individual processing units of a graphics processing unit (GPU), or dedicated neural circuits. In some embodiments, a node may include memory, for example, to store and use one or more previous inputs when processing subsequent inputs. For example, a node with memory may include a Long Short-Term Memory (LSTM) node. An LSTM node can use memory to maintain a "state" that allows the node to operate like a finite state machine (FSM).
[0058] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may start as a group of nodes organized into layers, as specified by the model format or model structure. In initialization, each weight may be applied to the connections between each pair of nodes connected according to the model format, for example, to nodes in a series of layers of a neural network. For example, each weight may be assigned randomly or initialized to a default value. The model can then be trained, for example, using training data, to produce results.
[0059] Training may involve applying supervised learning techniques. In supervised learning, training data can include multiple inputs (e.g., initial image, object mask, object mask) and corresponding ground truth outputs for each input (e.g., ground truth masks that correctly identify objects in each image). Based on a comparison of the model's output with the ground truth output, the weight values are automatically adjusted, for example, to increase the probability that the model generates a ground truth output for an image.
[0060] In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In some embodiments, the trained model may include a set of fixed weights downloaded, for example, from a server that provides weights. In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In embodiments where data is omitted, the segmenter 204 may generate a trained model based on pre-training, such as pre-training by the developer of the segmenter 204, or pre-training by a third party. In some embodiments, the trained model may include a set of fixed weights downloaded, for example, from a server that provides weights.
[0061] In some embodiments, the trained machine learning model receives an initial image containing one or more selected objects. In some embodiments, the trained machine learning model outputs one or more object masks containing one or more objects.
[0062] After one or more object masks are output by the segmentator 204 (for example, from a machine learning model), the segmentator 204 removes one or more selected objects from the initial image.
[0063] The restorer module 206 generates a restored image in which object pixels corresponding to one or more objects are replaced with restorer pixels. The restorer pixels may be based on pixels from a reference image at the same location without the object. Alternatively, the restorer module 206 may identify the restorer pixels to replace the object to be removed based on their proximity to other pixels surrounding the object. The restorer module 206 may determine the properties of the restorer pixels using the gradient of neighboring pixels. For example, if a bystander was standing on the ground, the restorer module 206 would replace the restorer pixels with ground pixels. Other restoration techniques are possible, including machine learning-based restoration techniques that output restorer pixels based on training data containing images with similar structures.
[0064] In an embodiment where the user chooses to erase the selected object, the user interface module 202 can display a repaired image in which the selected object has been removed and the pixels of the selected object have been replaced with repaired pixels.
[0065] In embodiments where the user chooses to move, replace, or add selected objects to an image, the diffusion module 208 uses a diffusion model to blend the objects with the object mask and the restored image. The diffusion model may receive an object mask, an incomplete object, and a restored image as input. The diffusion model may also receive additional inputs such as a text request, the number of pixels to fill in to output a complete object based on the incomplete object, and the dimensions of the restored image.
[0066] A diffusion model includes a forward process in which the diffusion model adds noise to the data, and a backward process in which the diffusion model learns to recover data from the noise. For example, if a selected object is moved from a first position to a second position, the diffusion module 208 applies the diffusion model by blending the selected object with an increasingly noisy version of the restored image, and then with an increasingly denoised version. In some embodiments, an object stitch diffusion model is used to move an object from a first position to a second position. In some embodiments, a generation diffusion model is used when the object is an incomplete object and only a portion of the object is generated, and / or for new objects generated from a text prompt.
[0067] Object Stitch Diffusion Model In some embodiments, an object-stitch diffusion model is used when an object is moved from a first position to a second position. In some embodiments, the diffusion module 208 includes an object image encoder that extracts semantic features from selected objects, a diffusion model that blends the objects with an image, and a content adapter that translates a set of visual tokens into a set of text tokens to fill the domain gap between the image and the text. In some embodiments, the diffusion module 208 trains the diffusion model using self-supervision based on training data that includes pairs of images and text. In some embodiments, the diffusion model is trained on synthetic data that simulates real-world scenarios. The diffusion model may also be trained using data augmentation generated by introducing random shifts and crop augmentation during training, while simultaneously ensuring that foreground objects are contained within the crop window.
[0068] The content adapter is trained during the first stage to maintain high-level semantics of objects using image and text pairs, and during the second stage, the content adapter is trained in the context of a diffusion model to encode important identity features of objects by facilitating a visual reconstruction of the objects in the original images. The diffusion model may be trained with embeddings generated by the content adapter via a cross-attention block.
[0069] The diffusion model uses an object mask to blend the restored image with the object. The diffusion model may denoise the masked area. The content adapter may convert visual features from the object image encoder into text features (tokens) to be used as conditioning for the diffusion model.
[0070] Generative diffusion model In some embodiments, the generative diffusion model is used to output a complete object based on an incomplete object, or to generate a new object. The diffusion module 208 trains the generative diffusion model based on training data. The training data may include image-text pairs used to create an embedding space for images and text. The image-text pairs may include images associated with corresponding text, such as an image of a dog and the text containing "pit bull". The diffusion module 208 may be trained with a loss that reflects the cosine distance between the text prompt embedding and the estimated clean image (i.e., no text-generated object).
[0071] The Diffusion Module 208 can use training data to perform text conditioning, which describes the process of outputting objects conditioned by text prompts. The Diffusion Module 208 can train a neural network to output objects based on text prompts provided by the user or a media application. For example, the text prompt could be a suggestion generated by a media application based on the context of an initial image (for example, if the initial image is a beach, the text prompt could be about a beach ball, a turtle, etc.).
[0072] In some embodiments, the diffusion module 208, based on receiving an incomplete object as input data, can output at least a portion of the missing part of the object, the position in the image where the incomplete object is being moved, and the dimensions of the output (e.g., the original dimensions or modified dimensions if the object has been resized). For example, if a user selects an object in a user interface that is partially cut off by a boundary and moves that object from a first position to a second position, and the second position also cuts off a portion of the object, the diffusion module 208 may output a modified object that includes more of the visible object based on the movement of the object in the image. In some embodiments, the diffusion module 208 may output a complete object based on an incomplete object selected by the user. For example, if a user selects a beach ball that is partially obscured by other objects, the diffusion model may be trained to output a complete beach ball.
[0073] In some embodiments, the diffusion module 208 generates a version of the complete object with progressively increased noise compared to a previous version, and generates a version of the restored image with progressively increased noise compared to a previous version. For example, the forward Markov noising process generates a series of noisy restored images by gradually adding Gaussian noise until nearly isotropic Gaussian noise samples are obtained. The forward noising process defines transitions of image manifolds, each manifold consisting of noisy images.
[0074] The diffusion module 208 can use an object mask to spatially blend a noisy version of a perfect object with the corresponding noisy version of the restored image. For example, the diffusion module 208 can use an object mask to blend each noisy version of a perfect object with each corresponding noisy version of the restored image, and the object mask defines the boundaries of the perfect object so that the object mask defines the area to be corrected during the blending process. In some embodiments, the diffusion process may include local perfect-object-guided diffusion, where the image generation loss determined during the training process is used under the object mask during positional object generation diffusion.
[0075] The diffusion module 208 can perform diffusion steps to denoise a latent space in a direction dependent on a text prompt. The diffusion module 208 generates a version of the complete object that is progressively denoised compared to the previous version, and a version of the restored image that is progressively denoised compared to the previous version. For example, an inverse Markov process transforms a Gaussian noisy sample by iteratively denoising the restored image using a learned posterior. Each step of the denoising diffusion process projects the noisy image onto the next, less noisy manifold.
[0076] The diffuse module 208 restores consistency by performing a denoising diffuse step after each blend and projecting it onto the next manifold. Once spatial blending is complete, the diffuse module 208 preserves the background by replacing the area outside the object mask with the corresponding area from the restored image.
[0077] In some embodiments, the diffusion module 208 injects contextual information into objects using cross-domain synthesis to apply an iterative refinement scheme, so as to match the objects to the style of the restored image. For example, if an object is generated for an indoor setting and added to an outdoor restored image, the object may be modified to be brighter to match the restored image. In another example, if an object is in a first location in the shadow and a second location is completely in the sun, the object may be modified to match the brightness of the second location.
[0078] Object Removal Model In some embodiments, instead of using a segmenter 204 to remove objects and a restorer module 206 to add pixels to the removed areas of the initial image, the diffusion model 208 is trained to include an object removal model.
[0079] The diffusion module 208 generates counterfactual training data to train the diffusion model to include an object removal model. For each pair of counterfactual images, the diffusion module 208 captures a factual image containing the object in the scene, physically removes the object while avoiding camera movement, lighting changes, or the movement of other objects, captures a counterfactual image of the scene without the object, and segments the factual image to create an object mask. Segmenting the factual image creates a segmented map (M) of the object O removed from the factual image X. o This includes creating ).
[0080] The diffusion module 208 creates a combined image for each image pair that includes the factual image, an object mask, and a counterfactual image. The object mask is a binary object mask (M o (X)) may be a counterfactual image pair, and the factual image and the binary object mask (X,M o Input pair with (X), and output counterfactual image (X cf It may be written as follows:
[0081]
number
[0082] Once the diffusion model, including the object removal model, is trained, the user interface module 202 can receive a request to remove selected objects from a first modified image. The initial image and the request are provided as input to the object removal model, which outputs a modified image that does not contain the selected objects.
[0083] Object Insertion Model In some embodiments, instead of using a segmenter 204 to remove objects, a restorer module 206 is used to add pixels to the removed areas of the initial image, and a diffuse model 208 is used to blend the objects with the pixels at the new locations, and the diffuse model 208 is trained to include an object insertion model.
[0084] In some embodiments, the object insertion model is trained on a number of image pairs that exceeds the number of available counterfactual image pairs. As a result, the diffusion module 208 generates synthetic training data. For each synthetic image pair, the diffusion module 208 selects an original image containing an object, uses an object removal model to output a corrected image from the object-free original image, generates an input image by inserting the object into the corrected image, and segments the original image to create an object mask. The corrected image lacking the object is represented by z i is referred to as.
[0085]
Number
[0086] The synthetic image pair is (y i , M o (x i )), and the corresponding target is the original image x i . Both the input image and the output image contain the object o, but the input image does not include the effect of the object on the scene, while the output image includes the effect of the object on the scene. In some embodiments, the diffusion module 208 trains the object insertion model using the diffusion target presented in Equation 1.
[0087] For each synthetic image pair, the diffusion module 208 creates a second combined image including the original image, the object mask, and the input image. The diffusion module 208 pre-trains the diffusion model to include the object insertion model based on using the synthetic image pair, and fine-tunes the diffusion model to include the object insertion model based on using the counterfactual image pairs used to train the object removal model.
[0088] In some embodiments, the user interface module 202 generates graphic data for displaying a user interface that provides the user with options to specify the position of an object and to resize an object. The diffusion module 208 adds the selected objects removed from the initial image to the new positions. In some embodiments, the diffusion module 208 provides the selected objects, as well as the positions in which the selected objects are located within the modified image, as input to a diffusion model, and outputs a modified image in which the selected objects are blended with the restored image. For example, the diffusion module 208 can spatially blend a noisy version of the restored image with a noisy version of the selected objects.
[0089] In some embodiments, the diffusion module 208 may add a shadow to a selected object at a new position. The shadow may match the direction of light in the image. For example, if the sun casts rays from the upper left corner of the image, the shadow may appear to the right of the person and / or object. In some embodiments, the diffusion module 208 uses a machine learning model to output a shadow mask used to generate the shadow to be cast on the object.
[0090] Once the selected object is added to the restored image, the user interface module 202 may include additional functions for modifying the restored image, such as an option to change the lighting of the restored image.
[0091] Figure 3A shows an exemplary initial image 300. The initial image includes a person 301, the subject of the initial image 300, a bystander 302, grass 303, a road 304, trees 305, and a cloudy sky 306 in the foreground of the initial image 300.
[0092] Figure 3B shows an exemplary initial image 310 in which two objects from Figure 3A have been selected for modification. The two objects are displayed along with the outlines 311, 312 of how the user selected the two objects. The user may have selected the two objects by using their finger, mouse, or other object, such as by circling the two objects, painting with a brush, or double-tapping. The user interface module 202 can identify the objects in the initial image 310 using object selection tools, lasso tools, artificial intelligence segmentation tools, etc. In some embodiments, the user interface module 202 may have suggested selecting the two objects, and in response to the user confirming the selection, the two objects were highlighted.
[0093] Once two objects are selected, the segmenter 204 segments the objects from the initial image 310 and generates object masks. Figure 3C shows an exemplary initial image 320 with object masks 321 and 322 surrounding the two objects. The person is surrounded by the first object mask 321, and the bystander is surrounded by the second object mask 322. The segmenter 204 removes the person and the bystander from the initial image 320.
[0094] Once the person and bystander are removed from the initial image 320 in Figure 3C, the restorer module 206 generates a restored image in which object pixels corresponding to the removed objects are replaced with restored pixels. Figure 3D shows an exemplary restored image 330 in which the bystander has been removed from the initial image 320 and the person 331 has been moved and resized.
[0095] In some embodiments, the diffusion module 208 resizes an object to be larger or smaller than the object in the initial image. For example, an object may be resized to be larger when moved forward and smaller when moved backward. In this example, the user provides input to resize person 331 to be smaller, and the diffusion module 208 resizes person 331 and further blends the person with the new position. In some embodiments, moving person 331 may be a separate action from resizing, or both moving and resizing may be part of the same action.
[0096] Additional modifications may be made to the restored image. Figure 3E shows an example of a restored image 340 in which the road 332 from Figure 3D is replaced with grass 341. Restored image 340 also includes a first type of tree replaced with a second type of tree 342. Restored image 340 also includes a cloudy sky from Figure 3D replaced with a cheerful sky with sunlight 343 coming from the upper right corner of the sky. Due to the direction of the sun 343, the diffuse module 208 outputs a shadow 344 that matches the person 345. In some embodiments, the road is replaced with grass 341 by using the diffuse module 208 to output grass and blending the grass with the restored image.
[0097] In some embodiments, the diffusion module 208 receives an incomplete object as input and outputs a complete object. The diffusion module 208 can also add a complete object to a second position in the restored image by blending the pixels of the complete object corresponding to the complete object with the restored pixels.
[0098] In some embodiments, if an object is captured at the edge of an image and a portion of the complete object is missing, the diffusion module 208 can create the missing portion of the object before performing blending. For example, if a woman is wearing a long dress and a portion of the dress is missing, the diffusion module 208 may output the complete dress based on the incomplete dress.
[0099] Figure 4A shows an exemplary initial image 400 of a child 405 sitting on a bench 410 and holding a balloon 415 partially cut off by the boundary of the initial image 400. In this example, the user interface module 202 provides the user interface with options for the user to select objects. The user selects the child 405, the bench 410, and the balloon 415 in the first position. The balloon 415 represents an incomplete image. The segmenter 204 segments the child 405, the bench 410, and the balloon 415 to separate the objects from the initial image 400.
[0100] The user interface module 202 includes an option to move the selected object to another location. The user selects a second location. The segmenter 204 removes the selected object from the initial image. The repair module 206 generates a repaired image in which the object pixels corresponding to the removed object are replaced with repair pixels.
[0101] The diffusion module 208 receives the coordinates of the selected object and the second position as input and outputs the balloon and the longer bench, which are the complete objects. Figure 4B shows an exemplary corrected image 450 in which the child 455, bench 460, and balloon 465 have been moved to the second position. In this example, the diffusion module 208 outputs a corrected image in which one or more versions of the child 455, bench 460, and balloon 465 are blended with one or more versions of the corrected image using an object mask.
[0102] In some embodiments, the user interface module 202 receives an image uncrop command from the user interface. The image uncrop command may occur in the modified image, the initial image, etc. The image uncrop command may be a button that is part of the user interface, which is a suggestion to help center an object, such as a person in the center of the image. In some embodiments, the command may be based on the user specifying a new boundary for the image by directly extending the image boundary, or on the movement of a selected object that extends the image boundary.
[0103] The restorer module 206 receives an uncropped image and the dimensions of the uncropped image as input, and outputs an uncropped image in which the boundary between the image and the uncropped image is replaced with restoration pixels that match the image. For example, if the boundary is water, the restorer module 206 may use water pixels for the uncropped portion of the image.
[0104] Figure 5A shows an exemplary user interface 500 for an initial image 504, including a button 412 for changing the boundaries of the initial image 504. In this example, the user interface 500 includes a first button 510 for editing the image and a second button 512 for changing the boundaries of the image. Other mechanisms are possible to provide a command to uncrop the image. For example, the user may select an edge 506 of the initial image 504 and drag and drop the edge 506 to indicate where the user wants to end the new boundary.
[0105] Users may change the initial image boundaries because the person in the image (person 508) is not centered, and cropping the image to reduce the right side results in an excessively narrow image.
[0106] Figure 5B shows an exemplary user interface 515 with an indicator 521 used to widen the boundary of the initial image 517. In this example, the user clicks and drags the indicator 521 to widen the left boundary of the initial image 517. The widened area 519 is shown with pixelated content, while the restorer module 206 generates an uncropped image.
[0107] In some embodiments, the repair module 206 receives an uncropped image and dimensions of the uncropped image as input. For example, the dimensions include the length and width of the new boundary to the left of the initial image.
[0108] The restorer module 206 outputs an uncropped image containing the restored pixels between the uncropped and extended boundaries of the corrected image, based on the dimensions. In this case, the restorer module 206 copies the pixels of the shrubs, rocks, water, and flowers. Figure 5C shows an exemplary user interface 530 of the uncropped image 532 output based on the initial image.
[0109] Figure 5D shows an alternative exemplary user interface 540 that extends the uncropped boundary of the initial image 542 using a selected person 544. In this alternative example, instead of extending the uncropped boundary of the initial image 542 using indicators such as indicator 521 in Figure 5B or boundary edges such as edge 506 in Figure 5A, the user can select an object in the image and move the object to expand the boundary. In Figure 5D, the user has moved person 544 to an edge of the initial image 542. The new position of person 544 defines a new edge that will be generated for the uncropped image, where 555 corresponds to the expanded region.
[0110] Exemplary Method Figure 6 shows an exemplary flowchart of a method 600 for generating a corrected image of a complete object from an incomplete object, according to some embodiments described herein. Method 600 may be performed by the computing device 200 of Figure 2. In some embodiments, method 600 is performed by a user device 115, a media server 101, or partially on the user device 115 and partially on the media server 101.
[0111] Method 600 in Figure 6 can begin with block 602. Block 602 determines whether permission has been granted by the user to access the initial image. If permission is not granted, method 600 terminates. If permission is granted, block 604 may follow block 602.
[0112] In block 604, a selection of an incomplete object in the initial image is received, the incomplete object is associated with a first position in the initial image, and the missing portion of the incomplete object is either cut off by the boundary of the initial image or obscured by another object. For example, the incomplete object could be a car cut off by the edge of the initial image. The user can select the incomplete object by clicking on the object, circling the object, or accepting a suggestion to modify the object generated by the user interface. Block 606 may follow block 604.
[0113] In block 606, an object mask is generated that includes the incomplete object pixels associated with the incomplete object. The incomplete object pixels may be determined by segmenting the initial image and identifying the pixels associated with the incomplete object within the initial object. Block 608 may follow block 606.
[0114] In block 608, the incomplete object pixels associated with the incomplete object are removed from the initial image. Block 610 may follow block 608.
[0115] In block 610, a repaired image is generated by replacing incomplete object pixels with repaired pixels. Block 612 may follow block 610.
[0116] In block 612, the object mask, the incomplete object, and the repaired image are provided as input to the diffusion model. Block 614 may follow block 612.
[0117] In block 614, the diffusion model outputs a complete object. For example, the diffusion model may receive an incomplete object, a second position where the complete object will be placed in the modified image, and dimensions of the complete object including the resized dimensions if the user resized the incomplete object, or the incomplete object will be resized due to a change from the first position to the second position. Block 616 may follow block 614.
[0118] In block 616, the corrected image is generated by using an object mask to blend one or more versions of the complete object with one or more versions of the corrected image. The complete object is placed in a second position in the corrected image, different from its first position in the initial image. Continuing the example above, the complete image could include a complete version of the car. The car may be resized based on its movement from foreground to background, and by being placed in the background, its size may be reduced to account for the distance it reflects.
[0119] In some embodiments, the modified image is a first modified image, and the method further includes receiving a request to remove a selected object from the first modified image, providing the first modified image and the request as input to an object removal model associated with a diffusion model, and using the object removal model to output a second modified image that does not include the selected object. In some embodiments, the modified image is a first modified image, and the method further includes receiving a request to move a selected object from a third position to a fourth position, providing the first modified image and the request as input to an object insertion model associated with a diffusion model, and using the object insertion model to output a second modified image that includes the selected object at the fourth position based on the request.
[0120] In some embodiments, the method further includes receiving a request to add an additional object to an initial image, outputting the additional object using a diffusion model, and outputting a modified image by blending one or more versions of the additional object with one or more versions of the modified image using the diffusion model. The additional object is positioned at a third location in the modified image, different from a second location in the modified image. In some embodiments, the request to add an additional object includes a text prompt describing the additional object, and the diffusion model outputs the additional object using generative artificial intelligence.
[0121] In some embodiments, a selected person in the restored image may be resized to account for movement forward or backward from a first position. For example, the selected person may be smaller or larger than the person in the initial image. In some embodiments, the diffusion model resizes the entire object based on the change from a first position in the initial image to a second position in the modified image.
[0122] Shadows corresponding to selected people in different positions may also be generated, and the shadows match the direction of light in the image. In some embodiments, a first image in the initial image is replaced with a second object from the initial image, and the modified image includes the first object that is replaced by the second object. In some embodiments, the lighting of the modified image is also changed. For example, the sky may be made brighter, the clouds thickened to reduce illumination, or the sky may be darkened. In some embodiments, the method further includes correcting the lighting of the modified image based on the direction of light in the modified image and adding shadows to the complete object.
[0123] Figure 7 shows an exemplary flowchart of a method 700 for outputting an uncropped image from an initial image, according to some embodiments described herein. Method 700 may be performed by the computing device 200 of Figure 2. In some embodiments, method 700 is performed by a user device 115, a media server 101, or partially on the user device 115 and partially on the media server 101.
[0124] Method 700 in Figure 7 can begin with block 702. In block 702, an initial image is displayed within the user interface. Block 704 may follow block 702.
[0125] In block 704, a command is received to uncrop the modified image so that the uncrop boundary of the modified image is extended to an extended boundary. The command is based on selecting an uncrop button and at least one action selected from a group of actions including moving an indicator to define the extended boundary, moving the edges of the initial image to define the extended boundary, moving selected objects to define the extended boundary, and combinations thereof. Block 704 may be followed by block 706.
[0126] In block 706, based on the command, an uncropped image is output that includes the repaired pixels between the uncropped boundary and the extended boundary of the corrected image.
[0127] Figure 8 shows an illustrative flowchart of a method 800 for training an object removal model according to some embodiments described herein. Method 800 may be performed by the computing device 200 of Figure 2. In some embodiments, method 800 is performed by a user device 115, a media server 101, or partially on the user device 115 and partially on the media server 101.
[0128] Method 800 can begin with block 802. In block 802, counterfactual training data is generated for each counterfactual image pair by capturing a factual image containing objects in the scene, physically removing the objects while avoiding camera movement, changes in lighting, or movement of other objects, capturing a counterfactual image of the scene without the objects, and segmenting the factual image to create an object mask. Block 804 may follow block 802.
[0129] In block 804, for each counterfactual image pair, a combined image is created containing the factual image and the object mask and counterfactual image. Block 806 may follow block 804.
[0130] In block 806, a diffusion model is trained to include an object removal model based on the use of counterfactual image pairs.
[0131] Figure 9 shows an illustrative flowchart of a method 900 for training an object insertion model according to several embodiments described herein. Method 900 may be performed by the computing device 200 of Figure 2. In some embodiments, method 900 is performed by a user device 115, a media server 101, or partially on the user device 115 and partially on the media server 101.
[0132] Method 900 can begin with block 902. In block 902, counterfactual training data is generated for each counterfactual image pair by capturing a factual image containing objects in the scene, physically removing the objects while avoiding camera movement, changes in lighting, or movement of other objects, capturing a counterfactual image of the scene without the objects, and segmenting the factual image to create an object mask. Block 904 may follow block 902.
[0133] In block 904, for each counterfactual image pair, a combined image is created containing the factual image, the object mask, and the counterfactual image. Block 906 may follow block 904.
[0134] In block 906, a diffusion model is trained to include an object removal model based on the use of counterfactual image pairs. Block 908 may follow block 906.
[0135] In block 908, for each composite image pair, the composite training data is generated by selecting the original image containing the object, using an object removal model to output a modified image from the original image without the object, generating an input image by inserting the object into the modified image, and segmenting the original image to create an object mask. Block 910 may follow block 908.
[0136] In block 910, for each composite pair, a second combined image is created that includes the original image, object mask, and input image. Block 912 may follow block 910.
[0137] In block 912, a diffusion model is pre-trained to include an object insertion model based on the use of composite image pairs. Block 914 may follow block 912.
[0138] In block 914, the diffusion model is fine-tuned to include an object insertion model based on the use of counterfactual image pairs.
[0139] In addition to the foregoing, the System, Program, or Functions described herein may provide the user with controls that allow the user to choose whether and when it may enable the collection of user information (e.g., information about the user's social networks, social actions, or activities, occupation, user preferences, or user's current location), and whether the user receives content or communications from the Server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, a user's identity may be processed so that personally identifiable information cannot be determined about the user, or the user's geographical location may be generalized so that the user's specific location cannot be determined if location information is obtained (e.g., at the city, zip code, or state level). Thus, the user can control what information is collected about them, how that information is used, and what information is provided to them.
[0140] In the above description, many specific details have been given for illustrative purposes to provide a complete understanding of this specification. However, it will be apparent to those skilled in the art that this disclosure can be implemented without these specific details. In some cases, structures and devices are shown in block diagram form to avoid ambiguity in the description. For example, this embodiment may be described above with reference primarily to the user interface and specific hardware. However, the embodiment can be applied to any type of computing device capable of receiving data and commands, and any peripheral device that provides services.
[0141] Any reference in this specification to “some embodiments” or “some examples” means that certain features, structures, or characteristics described in relation to the embodiments or examples may be included in at least one embodiment of the description. The phrase “in some embodiments” appearing in various parts of this specification does not necessarily refer to the same embodiment in all instances.
[0142] Some of the detailed explanations above are presented in terms of algorithms and symbolic representations of operations on data bits in computer memory. These descriptions and representations of algorithms are means used by those skilled in the art to most effectively communicate the content of their work to others skilled in the art. Here, and also generally, an algorithm is considered to be a self-consistent set of steps that lead to a desired result. These steps are steps that require the physical manipulation of physical quantities. Usually, though not essential, these quantities take the form of electrical or magnetic data that can be stored, transferred, combined, compared, and otherwise manipulated. It has sometimes proven convenient to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc., mainly for reasons of common usage.
[0143] However, it should be recognized that all these terms and similar terms should correspond to appropriate physical quantities and are merely convenient labels applied to those quantities. As will be evident from the following discussion, unless otherwise stated, throughout this specification, discussions using terms such as “process,” “calculate,” “calculate,” “determine,” or “display” are understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulate and convert data represented as physical (electronic) quantities in the registers and memory of the computer system into other data similarly represented as physical quantities in the memory or registers or other such information storage, transmission, or display devices of the computer system.
[0144] Embodiments of this specification may also relate to a processor for performing one or more steps of the methods described above. The processor may be a dedicated processor that is selectively started or reconfigured by a computer program stored in the computer. Such a computer program may be stored in non-temporary computer-readable storage media, including but not limited to optical discs, ROMs, CD-ROMs, magnetic disks, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memory including USB keys with non-volatile memory, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.
[0145] This specification may take the form of several entirely hardware embodiments, several entirely software embodiments, or several embodiments that include both hardware and software elements. In some embodiments, this specification is implemented in software, including but not limited to firmware, resident software, and microcode.
[0146] Furthermore, the description may take the form of a computer program product accessible from a computer-enabled medium or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, the computer-enabled medium or computer-readable medium may be any device that can store, save, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0147] A data processing system suitable for storing or executing program code includes at least one processor directly or indirectly coupled to a memory element via a system bus. The memory element may include local memory used during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage for at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during execution.
Claims
1. A method performed by a computer, The method includes receiving a selection of an incomplete object in an initial image, wherein the incomplete object is associated with a first position in the initial image, and the missing portion of the incomplete object is cut off by the boundary of the initial image or obscured by another object, and the method further includes To generate an object mask that includes the incomplete object pixels associated with the aforementioned incomplete object, Removing the incomplete object pixels associated with the incomplete object from the initial image, The process involves generating a repaired image in which the pixels of the incomplete object corresponding to the incomplete object are replaced with repaired pixels, The object mask, the incomplete object, and the repaired image are provided as input to the diffusion model. Using the aforementioned diffusion model, output a complete object, A method comprising generating a modified image by blending one or more versions of the complete object with one or more versions of the restored image using the object mask, wherein the complete object is positioned at a second position in the modified image that is different from the first position in the initial image.
2. The aforementioned modified image is the first modified image, Receiving a request to remove the selected object from the first modified image, The first modified image and the request are provided as inputs to the object removal model associated with the diffusion model. Using the object removal model, output a second modified image that does not include the selected object. The method according to claim 1, further comprising:
3. The aforementioned modified image is the first modified image, Receiving a request to move the selected object from the third position to the fourth position, The first modified image and the request are provided as inputs to the object insertion model associated with the diffusion model. Using the object insertion model, a second modified image including the selected object is output to the fourth position based on the request, The method according to claim 1, further comprising:
4. The aforementioned modified image is the first modified image, Receiving a request to add an additional object to the initial image, Using the aforementioned diffusion model, the additional object is output, Using the diffusion model, a second modified image is output by blending one or more versions of the additional object with one or more versions of the restored image. The method according to claim 1, further comprising:
5. The method according to claim 4, wherein the request to add the additional object includes a text prompt describing the additional object.
6. The method according to claim 1, wherein the complete object is resized based on a change from the first position in the initial image to the second position in the modified image.
7. The system receives a command to uncrop the modified image so that the uncropping boundary of the modified image is extended to the expanded boundary, Based on the command, an uncropped image is output that includes repaired pixels between the uncropped boundary and the extended boundary of the corrected image. The method according to claim 1, further comprising:
8. The command to uncrop the repaired image is: Select the uncrop button, A command to directly extend the uncropped boundary of the modified image to the extended boundary, or a movement of the complete object to extend the uncropped boundary of the modified image to the extended boundary, The method according to claim 7, including the method described in claim 7.
9. Correcting the lighting in the aforementioned corrected image, Based on the direction of the lighting in the modified image, a shadow is added to the complete object. The method according to claim 1, further comprising:
10. A program that, when executed by one or more processors, causes the one or more processors to perform an operation, wherein the operation is: The operation includes receiving a selection of an incomplete object in an initial image, the incomplete object being associated with a first position in the initial image, the missing portion of the incomplete object being cut off by the boundary of the initial image or obscured by other objects, and further To generate an object mask that includes the incomplete object pixels associated with the aforementioned incomplete object, Removing the incomplete object pixels associated with the incomplete object from the initial image, The process involves generating a repaired image in which the pixels of the incomplete object corresponding to the incomplete object are replaced with repaired pixels, The object mask, the incomplete object, and the repaired image are provided as input to the diffusion model. Using the aforementioned diffusion model, output a complete object, A program that generates a modified image by blending one or more versions of the complete object with one or more versions of the restored image using the object mask, wherein the complete object is positioned at a second position in the modified image that is different from the first position in the initial image.
11. The aforementioned modified image is the first modified image, and the operation is as follows: Receiving a request to remove the selected object from the first modified image, The first modified image and the request are provided as inputs to the object removal model associated with the diffusion model. Using the object removal model, output a second modified image that does not include the selected object. The program according to claim 10, further comprising:
12. The aforementioned modified image is the first modified image, and the operation is as follows: Receiving a request to move the selected object from the third position to the fourth position, The first modified image and the request are provided as inputs to the object insertion model associated with the diffusion model. Using the object insertion model, a second modified image including the selected object is output to the fourth position based on the request, The program according to claim 10, further comprising:
13. The aforementioned modified image is the first modified image, and the operation is as follows: The operation further includes receiving a request to add an additional object to the initial image, the request including a text prompt describing the additional object, and the operation is: Using the aforementioned diffusion model, the additional object is output, Using the diffusion model, a second modified image is output by blending one or more versions of the additional object with one or more versions of the restored image. The program according to claim 10, further comprising:
14. The aforementioned operation is, The system receives a command to uncrop the modified image so that the uncropping boundary of the modified image is extended to the expanded boundary, Based on the command, an uncropped image is output that includes repaired pixels between the uncropped boundary and the extended boundary of the corrected image. The program according to claim 10, further comprising:
15. It is a system, Processor and The system comprises a memory coupled to the processor that stores instructions, and when an instruction is executed by the processor, it causes the processor to perform an operation, and the operation is The process includes receiving a selection of an incomplete object in an initial image, wherein the incomplete object is associated with a first position in the initial image, and the missing portion of the incomplete object is cut off by the boundary of the initial image or obscured by another object, and the operation is performed To generate an object mask that includes the incomplete object pixels associated with the aforementioned incomplete object, Removing the incomplete object pixels associated with the incomplete object from the initial image, The process involves generating a repaired image in which the pixels of the incomplete object corresponding to the incomplete object are replaced with repaired pixels, The object mask, the incomplete object, and the repaired image are provided as input to the diffusion model. Using the aforementioned diffusion model, output a complete object, A system comprising generating a modified image by blending one or more versions of the complete object with one or more versions of the restored image using the object mask, wherein the complete object is positioned at a second position in the modified image that is different from the first position in the initial image.
16. The aforementioned modified image is the first modified image, and the operation is as follows: Receiving a request to remove the selected object from the first modified image, The first modified image and the request are provided as inputs to the object removal model associated with the diffusion model. Using the object removal model, output a second modified image that does not include the selected object. The system according to claim 15, further comprising:
17. The aforementioned modified image is the first modified image, and the operation is as follows: Receiving a request to move the selected object from the third position to the fourth position, The first modified image and the request are provided as inputs to the object insertion model associated with the diffusion model. Using the object insertion model, a second modified image including the selected object is output to the fourth position based on the request, The system according to claim 15, further comprising:
18. The aforementioned modified image is the first modified image, and the operation is as follows: The operation further includes receiving a request to add an additional object to the initial image, the request including a text prompt describing the additional object, and the operation is: Using the aforementioned diffusion model, the additional object is output, Using the diffusion model, a second modified image is output by blending one or more versions of the additional object with one or more versions of the restored image. The system according to claim 15, further comprising:
19. The system according to claim 15, wherein the complete object is resized based on a change from the first position in the initial image to the second position in the modified image.
20. The aforementioned operation is, The system receives a command to uncrop the modified image so that the uncropping boundary of the modified image is extended to the expanded boundary, Based on the command, an uncropped image is output that includes repaired pixels between the uncropped boundary and the extended boundary of the corrected image. The system according to claim 15, further comprising:
Citation Information
Patent Citations
Denoising diffusion generative adversarial networks
US20230095092A1
Image-to-Image Mapping by Iterative De-Noising
US20230103638A1