User input based distraction removal in media items - Patents.com
Generate bounding boxes through user input areas and use segmented machine learning models to remove interfering objects in visual media, solving the problem of automatic removal of objects in the prior art, and achieving a more accurate and automated removal process.
Patent Information
- Application Number
- JP2023561358
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-18
- Filing Date
- 2022-10-18
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2042-10-18
AI Technical Summary
The prior art is difficult to automatically and effectively remove interfering objects from visual media, and problems of error deletion or partial residues are prone to occur.
By receiving user input, the object removal area is generated, and the clipping area of the media item is provided to the segmentation machine learning model, the segmentation mask and corresponding segmentation score are output to perform object removal.
A more accurate and automated removal process is achieved, reducing false deletion and residual problems, and improving media quality.
Smart Images

Figure 0007673233000001 
Figure 0007673233000002 
Figure 0007673233000003
Abstract
Description
[Technical field]
[0001] REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 257,111, filed October 18, 2021, entitled "Translating User Annotation for Distraction Removal in Media Items," which is incorporated herein by reference in its entirety. [Background technology]
[0002] background The user-perceived quality of visual media items such as images (static images, images with selective motion, etc.) and videos may be improved by removing certain objects that distract from the focus of the media item or affect the visual appeal of the media item. For example, a user may capture a photo or video that includes a windmill, people in the background, a fence, or other objects that are not part of the main subject that the user intends to capture. For example, a photo may be intended to capture a foreground individual, tree, building, landscape, etc., but one or more distracting objects may be in the foreground (e.g., a fence, traffic light, or other object closer to the camera than the object of interest), in the background (e.g., a person in the background, power lines above the object of interest, or other objects farther away from the camera than the object of interest), or in the same plane (e.g., a person with their back to the camera but at the same distance to the camera as the object of interest). Summary of the Invention [Problem to be solved by the invention]
[0003] Users can use manual image or video editing techniques to remove distracting objects. However, this task can be tedious and incomplete. Furthermore, automatically removing distracting objects is difficult because it can result in false positives where additional objects or parts of objects are also removed, or because imperfect segmentation can result in parts of removed objects still being visible.
[0004] The discussion of the background art provided herein is for purposes of generally presenting the context of the present disclosure. The work of the presently named inventors is not expressly or impliedly admitted as prior art to the present disclosure, as are aspects of the description that may not qualify as prior art at the time of filing, to the extent that they are described in this background art section. [Means for solving the problem]
[0005] overview A computer-implemented method includes receiving user input indicating one or more objects to be removed from a media item. The method further includes converting the user input into a bounding box. The method further includes providing a crop of the media item based on the bounding box to a segmentation machine learning model. The method further includes outputting, with the segmentation machine learning model, a segmentation mask for the one or more segmented objects within the crop of the media item and a corresponding segmentation score indicative of a quality of the segmentation mask.
[0006] In some embodiments, the bounding box is an axis-aligned bounding box or an oriented bounding box. In some embodiments, the user input includes one or more strokes made with reference to the media item. In some embodiments, the bounding box is an oriented bounding box, and a direction of the oriented bounding box matches a direction of at least one of the one or more strokes. In some embodiments, prior to providing the crop of the media item, the segmentation machine learning model is trained using training data including a plurality of training images and a ground truth segmentation mask. In some embodiments, the method further includes determining that the segmentation mask is invalid based on one or more of a corresponding segmentation score not meeting a threshold score, a number of valid mask pixels below a threshold number of pixels, a segmentation mask size below a threshold size, or the segmentation mask being greater than a threshold distance from a region indicated by the user input, and generating a different mask based on a region in the user input in response to determining that the segmentation mask is invalid. In some embodiments, the method further includes inpainting portions of the media item that match the segmentation mask to obtain an output media item, where one or more objects are absent from the output media item. In some embodiments, the inpainting is performed using an inpainting machine learning model, and the media item and the segmentation mask are provided as inputs to the inpainting machine learning model. In some embodiments, the method further includes providing a user interface that includes the output media item.
[0007] In some embodiments, a non-transitory computer-readable medium having instructions stored thereon, when the instructions are executed by one or more computers, causes the one or more computers to perform the following operations: receiving user input indicating one or more objects to be removed from a media item, converting the user input into a bounding box, providing a crop of the media item based on the bounding box to a segmentation machine learning model, and using the segmentation machine learning model to output a segmentation mask for one or more segmented objects within the crop of the media item and a corresponding segmentation score indicative of a quality of the segmentation mask.
[0008] In some embodiments, the bounding box is an axis-aligned bounding box or an oriented bounding box. In some embodiments, the user input includes one or more strokes made with reference to the media item. In some embodiments, the bounding box is an oriented bounding box, and the orientation of the oriented bounding box matches the orientation of at least one of the one or more strokes. In some embodiments, prior to providing a crop of the media item, the segmentation machine learning model is trained using training data including a plurality of training images and a ground truth segmentation mask.
[0009] In some embodiments, a computing device comprises one or more processors and a memory coupled to the one or more processors and storing instructions that, when executed by the processor, cause the processor to perform a number of operations including receiving user input indicating one or more objects to be removed from a media item, converting the user input into a bounding box, providing a crop of the media item based on the bounding box to a segmentation machine learning model, and using the segmentation machine learning model to output a segmentation mask for one or more segmented objects within the crop of the media item and a corresponding segmentation score indicative of a quality of the segmentation mask.
[0010] In some embodiments, the bounding box is an axis-aligned bounding box or an oriented bounding box. In some embodiments, the user input includes one or more strokes made with reference to the media item. In some embodiments, the bounding box is an oriented bounding box, and a direction of the oriented bounding box matches a direction of at least one of the one or more strokes. In some embodiments, prior to providing the crop of the media item, the segmentation machine learning model is trained using training data including a plurality of training images and a ground truth segmentation mask. In some embodiments, the operations further include determining that the segmentation mask is invalid based on one or more of a corresponding segmentation score not meeting a threshold score, a number of valid mask pixels below a threshold number of pixels, a segmentation mask size below a threshold size, or the segmentation mask being greater than a threshold distance from a region indicated by the user input, and generating a different mask based on a region within the user input in response to determining that the segmentation mask is invalid.
[0011] The techniques described herein advantageously allow a media application to determine user intent associated with a user input, for example, when a user encircles a portion of an image, the media application determines the particular object the user is requesting to be removed. [Brief description of the drawings]
[0012] [Figure 1] FIG. 1 is a block diagram of an example network environment for removing objects from an image, according to some embodiments described herein. [Diagram 2] 1 is a block diagram of an example computing device for removing an object from an image, according to some embodiments described herein. [Figure 3A] 1A-1C illustrate example images with user input for removing an object, consistent with certain embodiments described herein. [Figure 3B] FIG. 2 illustrates an example image with an axis-aligned bounding box, consistent with certain embodiments described herein. [Figure 3C] 1A-1C are diagrams illustrating example images with different segmentation masks, in accordance with certain embodiments described herein. [Figure 3D] 1A-1C are diagrams illustrating example images with objects removed, according to certain embodiments described herein. [Figure 4A] 13A-13C are diagrams illustrating example images of a goat with user input to remove a portion of a fence, according to certain embodiments described herein. [Figure 4B] FIG. 2 illustrates an example image with an inaccurate bounding box, in accordance with certain embodiments described herein. [Figure 4C] 1 illustrates an exemplary image in which a goat is removed from a media item, according to certain embodiments described herein. [Figure 4D]FIG. 2 illustrates an example image with an oriented bounding box that properly identifies a fence as a target for removal, consistent with certain embodiments described herein. [Figure 4E] 1 illustrates an example image in which portions of the fence have been correctly removed, consistent with certain embodiments described herein. [Diagram 5] FIG. 2 illustrates a flowchart of an example method for generating a segmentation mask according to some embodiments described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] Detailed Description Exemplary Environment 100 FIG. 1 shows a block diagram of an exemplary environment 100. In some embodiments, the environment 100 includes a media server 101, a user device 115a, and a user device 115n coupled to a network 105. The users 125a, 125n may be associated with the respective user devices 115a, 115n. In some embodiments, the environment 100 may include other servers or devices not shown in FIG. 1. In FIG. 1 and the remaining figures, a letter following a reference number, e.g., "115a," represents a reference to the element with that particular reference number. A reference number in text without a following letter, e.g., "115," represents a general reference to an embodiment of the element with that reference number.
[0014] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively connected to a network 105 via signal line 102. The signal line 102 may be a wired connection, such as Ethernet, coaxial cable, fiber optic cable, or a wireless connection, such as Wi-Fi, Bluetooth, or other wireless technology. In some embodiments, the media server 101 transmits data to or receives data from one or more of the user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.
[0015] The database 199 may store machine learning models, training data sets, images, etc. Upon receiving user consent, the database 199 may store social network data associated with the user 125, user preferences of the user 125, etc.
[0016] The user device 115 may be a computing device that includes a memory coupled to a hardware processor. For example, the user device 115 may include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device capable of accessing the network 105.
[0017] In the illustrated implementation, user device 115a is coupled to network 105 via signal line 108, and user device 115n is coupled to network 105 via signal line 110. Media application 103 may be stored as media application 103b on user device 115a and / or media application 103c on user device 115n. Signal lines 108 and 110 may be wired connections such as Ethernet, coaxial cable, fiber optic cable, etc., or wireless connections such as Wi-Fi, Bluetooth, or other wireless technologies, etc. User devices 115a, 115n are accessed by users 125a, 125n, respectively. User devices 115a, 115n in FIG. 1 are used as an example. Although FIG. 1 shows two user devices 115a and 115n, the present disclosure applies to a system architecture having one or more user devices 115.
[0018] The media application 103 may be stored on the media server 101 and / or the user device 115. In some embodiments, the operations described herein are executed on the media server 101 or the user device 115. In some embodiments, some operations may be executed on the media server 101 and some operations may be executed on the user device 115. The execution of the operations is subject to user settings. For example, the user 125a may specify settings that operations should be executed on the respective device 115a and not on the server 101. In such settings, the operations described herein are executed entirely on the user device 115a and not on the media server 101. Additionally, the user 125a may specify that the user's images and / or other data are stored locally only on the user device 115a and not on the media server 101. In such settings, the user data is not transmitted to or stored in the media server 101. The transmission of user data to the media server 101, the temporary or permanent storage of such data by the media server 101, and the performance of operations on such data by the media server 101 are performed only if the user consents to the transmission, storage, and performance of operations by the media server 101. The user is provided with the option to change the settings at any time, for example to enable or disable the use of the media server 101.
[0019] Machine learning models (e.g., neural networks or other types of models) are stored locally on the user device 115 and utilized with specific user authorization when utilized for one or more operations. Server-side models are utilized only if authorized by the user. Model training is performed using a synthesized dataset, as described below with reference to FIG. 5. Furthermore, trained models may be provided for use on the user device 115. During such use, on-device training of the model may be performed, if authorized by the user 125. Updated model parameters may be transmitted to the media server 101, for example, to enable federated learning, if authorized by the user 115. Model parameters do not include any user data.
[0020] The media application 103 receives a media item. For example, the media application 103 receives the media item from a camera that is part of the user device 115, or the media application 103 receives the media item over the network 105. The media application 103 receives user input indicating one or more objects to be erased from the media item. For example, the user input is a circle surrounding the object to be removed. The media application 103 converts the user input into a bounding box. The media application 103 provides a crop of the media item based on the bounding box to a segmentation machine learning model. The segmentation machine learning model outputs a segmentation mask for one or more segmented objects within the crop of the media item and a corresponding segmentation score indicating the quality of the segmentation mask. In some embodiments, the media application 103 inpaints a portion of the media item that matches the segmentation mask to obtain an output media item, where one or more objects are not present in the output media item.
[0021] In some embodiments, the media application 103 may be implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a machine learning processor / co-processor, any other type of processor, or a combination thereof. In some embodiments, the media application 103a may be implemented using a combination of hardware and software.
[0022] Exemplary Computing Device 200 2 is a block diagram of an example computing device 200 that may be used to implement one or more features described herein. The computing device 200 may be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a media server 101 used to implement a media application 103a. In another example, the computing device 200 is a user device 115.
[0023] In some embodiments, computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, a display 241, a camera 243, and storage 245, all coupled via a bus 218. Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, display 241 may be coupled to bus 218 via signal line 228, camera 243 may be coupled to bus 218 via signal line 230, and storage 245 may be coupled to bus 218 via signal line 232.
[0024] The processor 235 may be one or more processors and / or processing circuits for executing program code and controlling basic operations of the computing device 200. A "processor" includes any suitable hardware system, mechanism, or component for processing data, signals, or other information. A processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multiprocessor configuration), graphics processing units (GPUs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), complex programmable logic devices (CPLDs), systems having dedicated circuits for achieving functionality, dedicated processors for implementing neural network model-based processing, neural circuits, processors optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, the processor 235 may include one or more co-processors for implementing neural network processing. In some embodiments, the processor 235 may be a processor that processes data to generate a probabilistic output, e.g., the output generated by the processor 235 may be inaccurate or may be accurate within a range from an expected output. The processing need not be limited to a particular geographic location, nor need it have time limitations. For example, the processor may perform its functions in real-time, offline, in batch mode, etc. Parts of the processing may be performed by different (or the same) processing systems at different times, in different locations, etc. A computer may be any processor in communication with a memory.
[0025] Memory 237 is provided within computing device 200 for access by processor 235 and may be any suitable processor-readable storage medium, such as random access memory (RAM), read only memory (ROM), electrically erasable read only memory (EEPROM), flash memory, etc., suitable for storing instructions for execution by the processor or set of processors, and may be located separately from and / or integrated with processor 235. Memory 237 may store software operated on computing device 200 by processor 235, including media application 103.
[0026] Memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, etc. One or more methods disclosed herein may operate in a number of environments and platforms, for example, as a standalone computer program that may run on any type of computing device, as a web application having a web page, as a mobile application ("app") running on a mobile computing device, etc.
[0027] Application data 266 may be data generated by other applications 264 or hardware of computing device 200. For example, application data 266 may include images used by an image library application, user actions identified by other applications 264 (e.g., social networking applications), and the like.
[0028] The I / O interface 239 can provide functionality that allows the computing device 200 to interface with other systems and devices. The interfaced devices can be included as part of the computing device 200 or can be separate and in communication with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage 245), and input / output devices can communicate through the I / O interface 239. In some embodiments, the I / O interface 239 can connect to interface devices such as input devices (keyboards, pointing devices, touch screens, microphones, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, monitors, etc.).
[0029] Some examples of interface connection devices that can be connected to I / O interface 239 may include a display 241 that may be used to display content, e.g., images, videos, and / or user interfaces of output applications as described herein, and to receive touch (or gesture) input from a user. For example, display 241 may be utilized to display a user interface including a graphical guide on a viewfinder. Display 241 may include any suitable display device, such as a liquid crystal display (LCD), light emitting diode (LED), or plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touch screen, a three-dimensional display screen, or other visual display device. For example, display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in a glasses form factor or headset device, or a monitor screen for a computing device.
[0030] The camera 243 may be any type of image capture device capable of capturing media items, including images and / or video. In some embodiments, the camera 243 captures images or video that the I / O interface 239 provides to the media application 103.
[0031] Storage 245 stores data related to media application 103. For example, storage 245 can store training data sets that include labeled images, machine learning models, output from the machine learning models, and the like.
[0032] FIG. 2 illustrates an example media application 103 stored in memory 237 that includes a bounding box module 202 , a segmentation machine learning module 204 , an inpainting module 206 , and a user interface module 208 .
[0033] The bounding box module 202 generates a bounding box. In some embodiments, the bounding box module 202 includes a set of instructions executable by the processor 235 to generate a bounding box. In some embodiments, the bounding box module 202 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0034] In some embodiments, the bounding box module 202 receives a media item, which may be received from the camera 243 of the computing device 200, from application data 266, or from the media server 101 via I / O interface 239. In various embodiments, the media item may be an image, a video, a series of images (e.g., GIF), or the like.
[0035] In some implementations, a media item includes a user input indicating one or more objects to be erased from the media item. In some implementations, the user input may be received at the client device 110 as touch input via a touch screen, input via a mouse / trackpad / other pointing device, or other suitable input mechanism. In some implementations, the user input is received with reference to a particular media item. In some embodiments, the user input is a manually drawn stroke that surrounds or is on the object to be erased from the media item. For example, the user input may be a circle that surrounds the object, a line or series of lines that are on the object, a square that surrounds the object, etc. The user input may be provided on the computing device 200 by a user drawing on a touch screen using a finger or stylus, by mouse or pointer input, by gesture input (e.g., detected by a camera), etc.
[0036] 3A, an exemplary image 300 with user inputs for removing objects is illustrated. In this example, the media is an image of a dandelion field with windmills in the background. The user inputs include roughly circular shapes 305, 310, and 315 that surround the objects to be removed. User input 305 surrounds the first windmill, user input 310 surrounds two windmills, and user input 315 surrounds the fourth windmill.
[0037] In some embodiments, the bounding box module 202 converts the user input into a bounding box. The bounding box module 202 identifies an object associated with the user input. For example, in FIG. 3A, the bounding box module 202 identifies that the user input 305 is associated with a windmill surrounded by the user input 305. In some embodiments, if the user input may include multiple objects, the bounding box module 202 identifies a percentage of the object associated with the user input. For example, the user input 310 encloses almost all of the pixels of the image corresponding to two windmills. As a result, the bounding box module 202 associates the user input 310 with the two windmills. In some embodiments, if the user input does not enclose all of the object, the bounding box module 202 determines whether the amount of the user input associated with the object exceeds a threshold percentage (e.g., measured in terms of pixels) of the object. For example, the user input 315 includes all of the windmill except one of the blades, and the percentage is 85%, which exceeds the threshold percentage of 70%.
[0038] In some embodiments, the bounding box module 202 identifies an object associated with the user input and compares the object's identity to a list of commonly removed objects to determine whether the user input includes a particular object. For example, the list of commonly removed objects may include a person, a power line, a scooter, a trash can, etc. If the user input encloses both a person and a portion of a tree in the background, the bounding box module 202 may determine that the user input corresponds to a person and not a tree because only the person, and not the tree, is part of the list of commonly removed objects.
[0039] The bounding box module 202 generates a bounding box that contains one or more objects. In some embodiments, the bounding box is a rectangular bounding box that encloses all pixels of the one or more objects. In some embodiments, the bounding box module 202 uses a suitable machine learning algorithm, such as a neural network, or more specifically, a convolutional neural network, to identify the one or more objects and generate the bounding box. The bounding box is associated with the x and y coordinates of the media item (image or video).
[0040] In some embodiments, the bounding box module 202 converts the user input into an axis-aligned bounding box or an oriented bounding box. The axis-aligned bounding box is aligned with the x-axis and y-axis of the media item. In some embodiments, the axis-aligned bounding box fits snugly around the stroke so that the edges of the bounding box touch the widest part of the stroke. The axis-aligned bounding box is the smallest box that contains the object indicated by the user input. Referring to FIG. 3B, an exemplary image 310 having axis-aligned bounding boxes is shown. Each of the bounding boxes 325, 330, and 335 contains one or more respective objects, and the bounding boxes 325, 330, and 335 enclose the corresponding user input stroke.
[0041] In Figure 3B, three strokes of user input were converted into three bounding boxes, but other embodiments are possible, such as four bounding boxes, each corresponding to a respective object, that fit snugly around the strokes except in the area where the objects are separated. For example, the bounding box 330 may be divided into two boxes with the outermost lines of the strokes aligned with the bounding box, and one or more additional lines in the middle to indicate the separation between the objects.
[0042] In some embodiments, the bounding box module 202 generates an oriented bounding box whose direction matches the direction of the stroke. For example, the oriented bounding box may be applied by the bounding box module 202 when the user input is in one direction, such as when the user provides one or more lines on the media item. In some embodiments, the bounding box module 202 generates an oriented bounding box that fits snugly around the stroke, which may be rotated about the image axes. In some embodiments, an oriented bounding box is any bounding box whose faces and edges are not parallel to the edges of the media item.
[0043] In some embodiments, the bounding box module 202 generates a crop of the bounding box based on the bounding box, for example, the bounding box module 202 generates a crop that uses the coordinates of the bounding box to generate a crop that includes one or more objects within the bounding box.
[0044] In some embodiments, the segmentation machine learning module 204 includes (and optionally performs the training of) a trained model, referred to herein as a segmentation machine learning model. In some embodiments, the segmentation machine learning module 204 is configured to apply the machine learning model to input data, such as application data 266 (e.g., media items captured by user device 115), and output a segmentation mask. In some embodiments, the segmentation machine learning module 204 may include code executed by the processor 235. In some embodiments, the segmentation machine learning module 204 may be stored in memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0045] In some embodiments, the segmentation machine learning module 204 may specify circuitry (e.g., for a programmable processor, for a field programmable gate array (FPGA), etc.) that enables the processor 235 to apply the machine learning model. In some embodiments, the segmentation machine learning module 204 may include software instructions, hardware instructions, or a combination. In some embodiments, the segmentation machine learning module 204 may provide an application programming interface (API) that may be used by the operating system 262 and / or other applications 264 to invoke the segmentation machine learning module 204, e.g., to apply the machine learning model to application data 266 and output a segmentation mask.
[0046] The segmentation machine learning module 204 uses training data to generate a trained segmentation machine learning model. For example, the training data may include training images and ground truth segmentation masks. The training images may be crops of manually segmented bounding boxes and / or crops of bounding boxes of synthetic images. In some embodiments, the segmentation machine learning module 204 trains the segmentation machine learning model using axis-aligned or oriented bounding boxes.
[0047] In some embodiments, the training data may include synthetic data generated for training purposes, such as data not based on activity in the training context, e.g., data generated from simulated or computer-generated images / videos. The training data may include synthetic images of bounding box crops of synthetic images. In some embodiments, the synthetic images are generated by overlaying a 2D or 3D object on a background image. The 3D object may be rendered from a particular view to convert the 3D object to a 2D object.
[0048] The training data may be obtained from any source, e.g., a data repository specifically marked for training, data for which permission is provided for use as training data for machine learning, etc. In some embodiments, the training may occur on the media server 101 providing the training data directly to the user device 115, the training may occur locally on the user device 115, or a combination of both.
[0049] In some embodiments, the segmentation machine learning module 204 uses weights taken from another application and unedited / transferred. For example, in these embodiments, the trained model may be generated, for example, on a different device and provided as part of the media application 103. In various embodiments, the trained model may be provided as a data file that includes a model structure or morphology (e.g., defining the number and type of neural network nodes, connectivity between the nodes and organization of the nodes into layers) and associated weights. The segmentation machine learning module 204 may read the data file for the trained model and realize a neural network with node connectivity, layers, and weights based on the model structure or morphology specified in the trained model.
[0050] The trained machine learning model may include one or more model formats or structures. For example, the model formats or structures may include any type of neural network, such as a linear network, a deep learning neural network that implements multiple layers (e.g., a "hidden layer" between the input layer and the output layer, where each layer is a linear network), a convolutional neural network (e.g., a network that splits or partitions input data into multiple parts or tiles, processes each tile separately using one or more neural network layers, and aggregates the results from the processing of each tile), or a sequence-to-sequence neural network (e.g., a network that receives as input a sequence of data, such as words in a sentence, frames in a video, etc., and produces as output an array of results).
[0051] The model format or structure may specify the connectivity between various nodes and the organization of the nodes into layers. For example, the nodes of the first layer (e.g., input layer) may receive data as input data or application data. Such data may include, for example, one or more pixels per node, for example, when the trained model is used, for example, for analysis of an initial image. Subsequent intermediate layers may receive as input the output of the nodes of the previous layer according to the connectivity specified in the model format or structure. These layers may also be called hidden layers. For example, the first layer may output a segmentation between foreground and background. The final layer (e.g., output layer) generates the output of the machine learning model. For example, the output layer may receive the segmentation of the initial image into foreground and background and output whether the pixel is part of the segmentation mask. In some embodiments, the model format or structure also specifies the number and / or type of nodes in each layer.
[0052] In different embodiments, the trained model may include one or more models. One or more of the models may include multiple nodes arranged in layers per model structure or morphology. In some embodiments, the multiple nodes may be memoryless computational nodes configured, for example, to process a unit of input and generate a unit of output. The computation performed by a node may include, for example, multiplying each of the multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to generate a node output. In some embodiments, the computation performed by a node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a non-linear function. In various embodiments, such computation may include operations such as matrix multiplication. In some embodiments, the computation by the multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, using individual processing units of a graphics processing unit (GPU), or using dedicated neural circuitry. In some embodiments, a node may include memory, for example, to store and use one or more previous inputs in processing a subsequent input. For example, a node with memory may include a long short-term memory (LSTM) node, which may use memory to maintain a "state" that allows the node to behave like a finite state machine (FSM).
[0053] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may begin as a number of nodes organized into layers as specified by the model format or structure. At initialization, respective weights may be applied to the connections between each pair of nodes that are connected according to the model format, e.g., nodes in successive layers of a neural network. For example, the respective weights may be randomly assigned or initialized to default values. The model may then be trained, e.g., using training data, to generate results.
[0054] Training may include applying supervised learning techniques. In supervised learning, training data may include multiple inputs (e.g., manually annotated segments and synthesized media items) and corresponding ground truth outputs for each input (e.g., ground truth segmentation masks that accurately identify one or more objects to be removed from each stroke of the media item). Based on a comparison of the model's output to the ground truth outputs, the values of the weights are, for example, automatically adjusted to increase the probability that the model generates a ground truth output for the media item.
[0055] In some embodiments, during training, the segmentation machine learning module 204 outputs the segmentation mask along with a segmentation score that indicates the quality of the segmentation mask in identifying objects to be erased in the media item. The segmentation score may reflect the intersection of the set of unions (loU) between the segmentation mask output by the segmentation machine learning model and the ground truth segmentation mask.
[0056] In various embodiments, the trained model includes a set of weights or embeddings that correspond to the model structure. In some embodiments, the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights. In various embodiments, the trained model includes a set of weights or embeddings that correspond to the model structure. In embodiments where data is omitted, the segmentation machine learning module 204 may generate a trained model that is based on previous training, e.g., by a developer of the segmentation machine learning module 204, by a third party, etc. In some embodiments, the trained model may include a set of weights that are fixed, e.g., downloaded from a server that provides the weights.
[0057] In some embodiments, the segmentation machine learning module 204 receives a crop of a media item. The segmentation machine learning module 204 provides the crop of the media item as an input to a trained machine learning model. In some embodiments, the trained machine learning model outputs a segmentation mask for one or more segmented objects within the crop of the media item and a corresponding segmentation score indicating the quality of the segmentation mask. In some embodiments, the segmentation score is based on a segmentation score generated during training of the machine learning model that reflects the loU between the segmentation mask output by the machine learning model and a ground truth segmentation mask. In some embodiments, the segmentation score is a number out of a total number, such as 40 / 100. Other representations of the segmentation score are possible.
[0058] In some embodiments, the segmentation machine learning model outputs a confidence value for each segmentation mask output by the trained machine learning model. The confidence value may be expressed as a percentage, a number between 0 and 1, etc. For example, the machine learning model outputs a confidence value of 85% for the confidence that the segmentation mask correctly covered the object identified in the user input.
[0059] In some embodiments, the segmentation machine learning module 204 determines that the segmentation mask was not successfully generated. For example, the segmentation score may not meet a threshold score. In another example, the segmentation machine learning module 204 may determine a number of valid mask pixels and determine that the number is below a threshold number of pixels. In another example, the segmentation machine learning module 204 may determine a size of the segmentation mask and determine that the size of the segmentation mask is below a threshold size. In yet another example, the segmentation machine learning module 204 may determine a distance between the segmentation mask and a region indicated by a user input and that the distance is greater than a threshold distance. In one or more of these examples, the segmentation machine learning module 204 outputs different segmentation masks based on the regions in the user input.
[0060] 3C, an example image 340 is shown with different segmentation masks 345, 350, 355. In this example, the segmentation machine learning module 204 outputs different segmentation masks that contain pixels that correspond to regions in the user input.
[0061] The repair module 206 generates an output media item in which one or more objects are absent (erased from the source media item). In some embodiments, the repair module 206 includes a set of instructions executable by the processor 235 to generate the output media item. In some embodiments, the repair module 206 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0062] In some embodiments, the inpainting module 206 receives the segmentation mask from the segmentation machine learning module 204. The inpainting module 206 performs inpainting of portions of the media item that match the segmentation mask. For example, the inpainting module 206 replaces pixels in the segmentation mask with pixels that match the background in the media item. In some embodiments, the pixels that match the background may be based on another media item in the same location. Figure 3D illustrates an exemplary inpainted image 360 in which no object is present in the output media item after inpainting.
[0063] In some embodiments, the repair module 206 receives a media item and a segmentation mask as input from the segmentation machine learning module 204 and trains a repair machine learning model to output an output media item in which one or more objects are absent from the output media item.
[0064] The user interface module 208 generates a user interface. In some embodiments, the user interface module 208 includes a set of instructions executable by the processor 235 to generate a user interface. In some embodiments, the user interface module 208 may be stored in the memory 237 of the computing device 200 and accessible and executable by the processor 235.
[0065] The user interface module 208 generates a user interface that asks the user for permission to access the user's media items before performing any of the steps performed by the modules of FIG. 2 and described in FIG.
[0066] The user interface module 208 generates a user interface that includes the media item and accepts user input to identify one or more objects for removal. For example, the user interface accepts touch input of a stroke. The user input indicates a distracting (or otherwise problematic) object that the user indicates for removal from the media item. For example, the media item may be an image of a family at a beach and the distracting object may be two people walking along the edge of the beach in the background. The user may use the user interface to circle the two people walking along the edge of the beach.
[0067] The user interface module 208 generates a user interface that includes the restored output media item. Continuing with the example, the media item is a family at the beach, except for two people walking along the edge of the beach. In some embodiments, the output media item may be labeled (visually) or marked (with a code, e.g., steganographically) to indicate that the media item has been edited to erase one or more objects. In some embodiments, the user interface includes options for editing the output media item, sharing the output media item, adding the output media item to a photo album, etc. Options for editing the output media item may include the ability to undo the erasure of objects.
[0068] In some embodiments, the output media item may be labeled (visually) or marked (with a code, e.g., with stenography) to indicate that the media item has been edited to erase one or more objects.
[0069] In some embodiments, the user interface module 208 receives feedback from users on the user devices 115. The feedback may take the form of users posting output media items, users deleting output media items, users sharing output media items, etc.
[0070] Example Oriented Bounding Box 4A illustrates an example image 400 of a goat with a user input to remove a segment of a fence, according to some embodiments described herein. The bounding box module 202 receives the user input and generates an oriented bounding box. The orientation of the oriented bounding box is determined based on the direction of the user input. In FIG. 4A, the user input 405 is a stroke along the diagonal of the chain link fence. The bounding box 407 is an axis-aligned bounding box.
[0071] 4B illustrates an example image 410 with an incorrect bounding box, according to some embodiments described herein. Because an axis-aligned bounding box is a rectangular box whose sides are aligned with the x- and y-axes, the bounding box incorrectly identifies a goat as the object for removal, instead of the chain link fence that was identified for removal by user input.
[0072] FIG. 4C shows an example image 420 in which the goat has been removed from the media item. 4D illustrates an example image 430 in which the bounding box module 202 uses an oriented bounding box to properly identify the fence as a target for removal. As shown in FIG. 4D, when the segmentation machine learning module 204 receives a cropped version of the oriented bounding box, the resulting segmentation mask more closely captures the user's intent to remove a portion of the chain link fence than when it receives a cropped version of the axis-aligned bounding box, which is when the segmentation machine learning module 204 misinterpreted the user's intent as selecting the goat behind the chain link fence.
[0073] FIG. 4E illustrates an example image 440 in which a fence segment has been correctly removed, according to certain embodiments described herein.
[0074] Exemplary Method 500 Figure 5 shows a flowchart of an example method 500 for generating a segmentation mask. The method 500 of Figure 5 may begin at block 502. The method 500 shown in the flowchart may be performed by the computing device 200 of Figure 2. In some embodiments, the method 500 is performed by the user device 115, the media server 101, or partially on the user device 115 and partially on the media server 101.
[0075] At block 502, user permission to perform method 500 is received. For example, a user may load an application to provide user input by circling an object in a media item, but before the media item is displayed, a user interface asks for user permission to access the media item associated with the user. The user interface may also ask for permission to modify the media item, to allow the user to allow access to only certain media items, to ensure that media items are not stored or transferred to a server without user permission, etc. Block 502 may be followed by block 504.
[0076] At block 504, it is determined whether user permission has been received. If user permission has not been received, block 504 is followed by block 506, which stops the method 500. If user permission has been received, block 504 is followed by block 508.
[0077] At block 508, user input is received indicating one or more objects to be erased from the media item. For example, the image may include a trash can in the background and the user input is a circle around the trash can. Block 508 may be followed by block 510.
[0078] In block 510, the user input is converted into a bounding box. For example, the bounding box may be an axis-aligned bounding box or an oriented bounding box. Block 510 may be followed by block 512.
[0079] The crop of the media item is provided to a segmentation machine learning model based on the bounding box at block 512. Block 506 may be followed by block 514.
[0080] At block 514, a segmentation mask is output using the trained segmentation machine learning model for the one or more segmented objects within the crop of the media item and a corresponding segmentation score indicating the quality of the segmentation mask.
[0081] In addition to the above, a user may be provided with controls that allow the user to select both whether and when the systems, programs, or features described herein may enable collection of user information (e.g., information about the user's media items, including images and / or videos, social networks, social actions or activities, occupation, the user's preferences (e.g., regarding objects in images), or the user's current location), as well as whether the user is sent content or communications from the server. Additionally, certain data may be treated in one or more ways before being stored or used such that personally identifiable information is removed. For example, the user's identity may be treated such that no personally identifiable information can be determined for that user, or location information is obtained (e.g., to the city, ZIP code, or state level) such that the user's geographic location is generalized and the user's specific location cannot be determined. Thus, the user may have control over what information is collected about the user, how that information is used, and what information is provided to the user.
[0082] In the above description, for the purpose of explanation, numerous specific details are set forth to provide a thorough understanding of the present specification. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details. In some instances, structures and devices are shown in block diagram form to avoid obscuring the description. For example, the embodiments may be described above primarily with reference to user interfaces and specific hardware. However, the embodiments may be applied to any type of computing device capable of receiving data and commands, as well as any peripheral device that provides services.
[0083] Reference herein to "some embodiments" or "some instances" means that a particular feature, structure, or characteristic described in connection with an embodiment or instance may be included in at least one implementation of the description. The appearances of the phrase "in some embodiments" in various places in the specification are not necessarily all referring to the same embodiments.
[0084] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0085] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specified, as will be apparent from the discussion that follows, throughout the description, discussions utilizing terms including "processing" or "calculating" or "computing" or "determining" or "displaying" and the like will be understood to refer to operations and processes of a computer system or similar electronic computing device that manipulates or converts data represented as physical (electronic) quantities in the computer system's registers and memory into other data represented as physical quantities in the computer system's memory or registers or other such information storage, transmission, or display device.
[0086] The embodiments herein may also relate to a processor for performing one or more steps of the above-mentioned methods. The processor may be a dedicated processor selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, including any type of disk including, but not limited to, an optical disk, a ROM, a CD-ROM, a magnetic disk, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory including a USB key having a non-volatile memory, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0087] This specification can take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements, hi some embodiments, this specification is implemented in software, including but not limited to firmware, resident software, microcode, etc.
[0088] Furthermore, the descriptions may take the form of a computer program product accessible from a computer usable or computer readable medium providing program code for use by or in connection with a computer or any instruction execution system. For purposes of this description, a computer usable or computer readable medium may be any apparatus that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device.
[0089] A data processing system suitable for storing or executing program code includes at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory used during actual execution of the program code, mass storage devices, and cache memory that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from mass storage devices during execution.
Claims
1. 1. A computer-implemented method comprising: receiving user input indicating one or more objects to be removed from the media item; converting the user input into a bounding box; providing a crop of the media item based on the bounding box to a segmentation machine learning model; and outputting, using the segmentation machine learning model, a segmentation mask for one or more segmented objects within the crop of the media item and a corresponding segmentation score indicative of a quality of the segmentation mask; the user input comprises one or more strokes made with reference to the media item; A method wherein the bounding box is an oriented bounding box, the orientation of the oriented bounding box coinciding with a direction of at least one of the one or more strokes.
2. The method of claim 1 , wherein the oriented bounding box has faces and edges that are not parallel to edges of the media item.
3. The method of claim 1 , wherein prior to providing the crop of the media item, the segmentation machine learning model is trained using training data comprising a plurality of training images and a ground truth segmentation mask.
4. 1. A computer-implemented method comprising: receiving user input indicating one or more objects to be removed from the media item; converting the user input into a bounding box; providing a crop of the media item based on the bounding box to a segmentation machine learning model; using the segmentation machine learning model to output a segmentation mask for one or more segmented objects within the crop of the media item and a corresponding segmentation score indicative of a quality of the segmentation mask; determining that a segmentation mask is invalid based on one or more of: a corresponding segmentation score not meeting a threshold score, a number of valid mask pixels below a threshold number of pixels, a segmentation mask size below a threshold size, or the segmentation mask being greater than a threshold distance from a region indicated by the user input; in response to determining that the segmentation mask is invalid, generating a different mask based on regions within the user input.
5. The method of claim 1 , further comprising inpainting a portion of a media item that matches the segmentation mask to obtain an output media item, wherein the one or more objects are not present in the output media item.
6. The method of claim 5 , wherein the inpainting is performed using an inpainting machine learning model, and the media items and the segmentation mask are provided as inputs to the inpainting machine learning model.
7. The method of claim 5 , further comprising providing a user interface that includes the output media items.
8. A program for causing one or more processors to execute the method according to any one of claims 1 to 7.
9. 1. A computing device comprising: A processor; and a memory coupled to the processor having instructions stored therein, the instructions, when executed by the processor, causing the processor to perform the following operations: receiving user input indicating one or more objects to be removed from the media item; converting the user input into a bounding box; providing a crop of the media item based on the bounding box to a segmentation machine learning model; using the segmentation machine learning model to output a segmentation mask for one or more segmented objects within the crop of the media item and a corresponding segmentation score indicative of a quality of the segmentation mask; the user input comprises one or more strokes made with reference to the media item; A computing device, wherein the bounding box is an oriented bounding box, a direction of the oriented bounding box coinciding with a direction of at least one of the one or more strokes.
10. The computing device of claim 9 , wherein the oriented bounding box has faces and edges that are not parallel to edges of the media item.
11. 10. The computing device of claim 9, wherein prior to providing the crop of the media item, the segmentation machine learning model is trained using training data comprising a plurality of training images and a ground truth segmentation mask.
12. 1. A computing device comprising: A processor; and a memory coupled to the processor having instructions stored therein, the instructions, when executed by the processor, causing the processor to perform the following operations: receiving user input indicating one or more objects to be removed from the media item; converting the user input into a bounding box; providing a crop of the media item based on the bounding box to a segmentation machine learning model; using the segmentation machine learning model to output a segmentation mask for one or more segmented objects within the crop of the media item and a corresponding segmentation score indicative of a quality of the segmentation mask; determining that a segmentation mask is invalid based on one or more of: a corresponding segmentation score not meeting a threshold score, a number of valid mask pixels below a threshold number of pixels, a segmentation mask size below a threshold size, or the segmentation mask being greater than a threshold distance from a region indicated by the user input; and in response to determining that the segmentation mask is invalid, generating a different mask based on regions within the user input.
Citation Information
Patent Citations
Object detection model training method and device
CN112241675A
Method and apparatus for tilt adjustment and layout of photograph, and recording medium
JP2000228722A
Segmenting Objects In Video Sequences
US20200143171A1