Afterside and affiliated shadow removal

By combining the side part segmentation model and the shadow segmentation model with repair technology, the problem of the side part shadow in the image is difficult to accurately remove, and efficient erasing of side part and shadow is achieved, and image quality is improved.

CN120303689APending Publication Date: 2025-07-11GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085120.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-10
Filing Date
2023-11-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately remove the shadows of the side in the image, resulting in poor image editing effects, especially when the side does not project the shadows or the shadows mistakenly identify the shadows of other objects.

Method used

The image is analyzed using the side-segmentation model and the shadow segmentation model, and the side-segmentation and shadow mask are generated. The pixel values are updated through repair technology to erase side-segment and shadows. The convolutional neural network is trained using supervised learning to optimize the model accuracy.

Benefits of technology

Effectively remove the side and shadows in the image, improving the visual effect of the image, ensuring the accuracy of shadow identification and the integrity of side removal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303689A_ABST
    Figure CN120303689A_ABST
Patent Text Reader

Abstract

A media application derives a partner mask from an image by analyzing the image with a partner segmentation model, where the image depicts a partner and the partner mask identifies a plurality of first pixels associated with the partner in the image. The media derives a shadow mask of the partner by analyzing the image with a shadow segmentation model, where the image and the partner mask are provided as inputs to the shadow segmentation model, and where the shadow mask identifies a plurality of second pixels in the image associated with a shadow of the partner. The media application modifies the image to update pixel values of the plurality of first pixels and the plurality of second pixels such that bypasses and shadows are erased from the image.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims priority to U.S. Patent Application No. 18 / 195,401, titled "BYSTANDER AND ATTACHED SHADOW REMOVAL", filed on May 10, 2023, the entire content of which is incorporated herein by reference in its entirety. Background Art

[0002] The attractiveness of visual media items such as images (static images, images with selective motion, etc.) and videos can be enhanced by removing bystanders that distract the focus of the media item. However, when a bystander has an attached shadow, the pixels associated with the bystander can be identified and modified to remove the bystander, while the pixels associated with the shadow are retained. Prior art has attempted to address this problem by generating a shadow mask, but it is difficult to accurately generate a shadow mask. For example, the process of identifying the shadow may result in false positives in the shadow mask candidates, in which case a region larger than the shadow is removed. Additionally, in a scenario where a person in the image does not cast a shadow, the process of identifying the shadow may result in identifying a shadow that belongs to other objects. When object classification is used for image - editing purposes (e.g., replacing the bystander with pixels matching the background), the resulting image may be unsatisfactory due to the presence of pixels associated with the shadow that remain in the image even after the bystander has been erased.

[0003] The background art description provided herein is for the purpose of generally presenting the background of the disclosure. The work of the currently named inventors (to the extent it is described in this background art section) and aspects of the specification that may not otherwise count as prior art at the time of filing are neither expressly nor impliedly admitted to be prior art to the present disclosure. Summary of the Invention

[0004] A computer - implemented method includes: deriving a bystander mask from an image by analyzing the image with a bystander segmentation model, where the image depicts a bystander and the bystander mask identifies a plurality of first pixels in the image that are associated with the bystander. The method further includes: deriving a shadow mask of the bystander by analyzing the image with a shadow segmentation model, where the image and the bystander mask are provided as inputs to the shadow segmentation model, and where the shadow mask identifies a plurality of second pixels in the image that are associated with the shadow of the bystander. The method further includes: modifying the image to update the pixel values of the plurality of first pixels and the plurality of second pixels such that the bystander and the shadow are erased from the image.

[0005] In some embodiments, the method further includes: using a shadow classifier model based on the image and the bystander mask to determine the likelihood of the presence of a shadow, wherein obtaining the bystander's shadow mask is performed when the likelihood of the presence of the shadow meets a threshold. In some embodiments, modifying the image includes: applying an inpainting technique to update the pixel values of the plurality of first pixels and the plurality of second pixels. In some embodiments, the inpainting technique is performed by an inpainting model. In some embodiments, the method further includes: merging the bystander mask and the shadow mask before modification. In some embodiments, the shadow segmentation model includes a convolutional neural network trained using supervised learning, wherein training the convolutional neural network is performed using a training data set that includes a plurality of training images, a segmentation mask associated with a person in each of the plurality of training images, and a ground truth shadow mask for each of the plurality of training images, and wherein training includes, for each of the plurality of training images: obtaining a predicted shadow mask based on the training image and the segmentation mask associated with the person in the training image, calculating a loss value based on a comparison of the predicted shadow mask and the ground truth shadow mask for the image, and updating the weights of one or more nodes of the convolutional neural network based on the loss value. In some embodiments, the training data set is generated by including candidate images selected from the following group in the plurality of training images: the person is associated with a body bounding box that has less than a threshold overlap value with a vehicle bounding box, the body bounding box is associated with an aspect ratio that is less than a threshold aspect ratio, the segmentation mask is outside the vehicle bounding box, the intersection of the body bounding box and a mobile object bounding box is less than a threshold overlap value, the ground truth shadow mask is between a threshold first size and a threshold second size, the plurality of training images meet an illumination threshold, the candidate images include shadows originating from a direction associated with the person's feet, and combinations thereof. In some embodiments, one or more of the plurality of training images are associated with an empty ground truth shadow mask. In some embodiments, one or more of the plurality of training images are generated from a series of candidate images of a scene including a person captured at different times by: generating a clean-plate image containing the static elements of the scene, comparing each of the series of candidate images with the clean-plate image to identify the dynamic parts of the scene, generating a segmentation mask for the person in the series of images, and determining that the pixels of the ground truth shadow mask are adjacent to the pixels corresponding to the segmentation mask.In some embodiments, the shadow segmentation model includes a convolutional neural network trained using supervised learning, where training the convolutional neural network is performed with a training dataset that includes a plurality of training images and a ground truth shadow mask for each of the plurality of training images, and where training includes, for each of the plurality of training images: obtaining a predicted shadow mask based on the training image, where one or more pixels of the training image associated with a person have values associated with a specific color that does not occur in natural images, calculating a loss value based on a comparison of the predicted shadow mask and the ground truth shadow mask for the image, and updating the weights of one or more nodes of the convolutional neural network based on the loss value. In some embodiments, the image depicts two or more bystanders, and the deriving, the obtaining, and the modifying are performed for each of the two or more bystanders. In some embodiments, the image is a single frame of a video. In some embodiments, at least one pixel of a first plurality of pixels of the bystander mask is adjacent to at least one pixel of a second plurality of pixels of the shadow mask.

[0006] In some embodiments, a non-transitory computer-readable medium stores instructions that, when executed by one or more computers, cause the one or more computers to perform operations including: deriving a bystander mask from an image by analyzing the image with a bystander segmentation model, where the image depicts a bystander and the bystander mask identifies a plurality of first pixels in the image associated with the bystander; obtaining a shadow mask of the bystander by analyzing the image with a shadow segmentation model, where the image and the bystander mask are provided as inputs to the shadow segmentation model, and where the shadow mask identifies a plurality of second pixels in the image associated with the shadow of the bystander; and modifying the image to update the pixel values of the plurality of first pixels and the plurality of second pixels such that the bystander and the shadow are erased from the image.

[0007] In some embodiments, the operations further include determining a likelihood of the presence of a shadow using a shadow classifier model based on the image and the bystander mask, where obtaining the shadow mask of the bystander is performed when the likelihood of the presence of the shadow meets a threshold. In some embodiments, modifying the image includes: applying an inpainting technique to update the pixel values of the plurality of first pixels and the plurality of second pixels. In some embodiments, the operations further include: merging the bystander mask and the shadow mask before modifying.

[0008] In some embodiments, a computing device includes: one or more processors; and a memory coupled to the one or more processors, the memory storing instructions that, when executed by the processors, cause the processors to perform operations. The operations may include: deriving a bystander mask from an image by analyzing the image with a bystander segmentation model, where the image depicts a bystander and the bystander mask identifies a plurality of first pixels in the image associated with the bystander; deriving a shadow mask of the bystander by analyzing the image with a shadow segmentation model, where the image and the bystander mask are provided as inputs to the shadow segmentation model and where the shadow mask identifies a plurality of second pixels in the image associated with the shadow of the bystander; and modifying the image to update the pixel values of the plurality of first pixels and the plurality of second pixels such that the bystander and the shadow are erased from the image.

[0009] In some embodiments, the operations further include determining a likelihood of the presence of a shadow using a shadow classifier model based on the image and the bystander mask, where deriving the shadow mask of the bystander is performed when the likelihood of the presence of the shadow meets a threshold. In some embodiments, modifying the image includes: applying an inpainting technique to update the pixel values of the plurality of first pixels and the plurality of second pixels.

[0010] The techniques described in the specification advantageously address the problem of identifying shadows attached to bystanders to be removed from an image. The shadow segmentation model takes a person mask as input and predicts the mask of their shadow. The bystander segmentation model is updated to use state-of-the-art person detection and segmentation instead of trying to make the shadow segmentation model detect people as accurately as a state-of-the-art person detector. The techniques further include ways of generating a training data set for training a machine learning model, the training data set reducing false positives and including examples where the bystander does not cast a shadow. Finally, the techniques include a segmentation pipeline with different classifiers for predicting the shadow of a bystander. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a block diagram of an example network environment for identifying shadows attached to bystanders according to some embodiments described herein.

[0012] Figure 2 is a block diagram of an example computing device for identifying shadows attached to bystanders according to some embodiments described herein.

[0013] Figure 3 Illustrates an example image according to some embodiments described herein, an example image where the bystander is removed but the shadow is not removed, and an example image where both the bystander and the shadow are removed.

[0014] Figure 4 Illustrates an example image excluded from the training data according to some embodiments described herein.

[0015] Figure 5 Illustrates example images excluded based on the size of a candidate shadow mask being too small or too large according to some embodiments described herein.

[0016] Figure 6 Illustrates images excluded by examples where clouds block the projected shadow and example images with shadows according to some embodiments described herein.

[0017] Figure 7 Illustrates example images where a person depicted in the image projects a shadow according to some embodiments.

[0018] Figure 8 Illustrates example candidate images with a ground truth shadow mask as part of training data according to some embodiments described herein.

[0019] Figure 9 Illustrates an example flowchart of a method for training a machine learning model to output a shadow mask according to some embodiments described herein.

[0020] Figure 10A Illustrates an example flowchart of a method for outputting a combined bystander and shadow mask according to some embodiments described herein.

[0021] Figure 10B Illustrates an example flowchart of another method for outputting a combined bystander and shadow mask according to some embodiments described herein.

[0022] Figure 11 Illustrates an example flowchart of a method for modifying an image to erase bystanders and their shadows from the image based on a derived bystander mask and a derived shadow mask according to some embodiments described herein. Detailed Description

[0023] Example Environment 100

[0024] Figure 1 Illustrates a block diagram of an example environment 100 for identifying shadows attached to bystanders. Environment 100 includes a media server 101, user devices 115a and 115n coupled to a network 105. Users 125a, 125n may be associated with the respective user devices 115a, 115n. In some embodiments, environment 100 may include Figure 1 other servers or devices not shown. In Figure 1 and the remaining figures, the letter following a reference numeral, e.g., "115a", indicates a reference to an element having that particular reference numeral. A reference numeral in the text without a following letter, e.g., "115", indicates a general reference to embodiments of an element bearing that reference numeral.

[0025] The media server 101 includes a processor, a memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively coupled to the network 105 via the signal line 102. The signal line 102 can be a wired connection, such as Ethernet, coaxial cable, fiber optic cable, etc., or a wireless connection, such as Wi-Fi®, Bluetooth®, or other wireless technologies. In some embodiments, the media server 101 sends data to and receives data from one or more of the user devices 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0026] The database 199 can store machine learning models, training data sets, images, etc. The database 199 can also store social network data associated with the user 125, user preferences of the user 125, etc.

[0027] The user device 115 is a computing device including a memory coupled to a hardware processor. For example, the user device 115 can include a mobile device, a tablet computer, a mobile phone, a wearable device, a head-mounted display, a mobile email device, a portable game player, a portable music player, a reader device, or another electronic device capable of accessing the network 105.

[0028] In the illustrated implementation, the user device 115a is coupled to the network 105 via the signal line 108, and the user device 115n is coupled to the network 105 via the signal line 110. The media application 103 can be stored on the user device 115a as the media application 103b and / or on the user device 115n as the media application 103c. The signal lines 108 and 110 can be a wired connection, such as Ethernet, coaxial cable, fiber optic cable, etc., or a wireless connection, such as Wi-Fi®, Bluetooth®, or other wireless technologies. The user devices 115a, 115n are accessed by the users 125a, 125n respectively. Figure 1 The user devices 115a, 115n in are used only as examples. Although Figure 1 two user devices 115a and 115n are illustrated, the present disclosure is applicable to a system architecture having one or more user devices 115.

[0029] The media application 103 is stored on the media server 101 or the user device 115. In some embodiments, the operations described herein are performed on the media server 101 or the user device 115. In some embodiments, some operations may be performed on the media server 101 and some operations may be performed on the user device 115. The execution of the operations is based on user settings. For example, user 125a may specify the following setting: the operations are to be performed on their respective devices 115a rather than on the server 101. Under such a setting, the operations described herein (e.g., referring to Figures 9 to 11 ) are performed entirely on the user device 115a and no operations are performed on the media server 101. Further, user 125a may specify that the user's image and / or other data are to be stored only locally on the user device 115a and not on the media server 101. Under such a setting, no user data is transmitted to or stored on the media server 101. The transmission of user data to the media server 101, any temporary or permanent storage of such data on the media server 101, and the execution of operations on such data by the media server 101 are performed only if the user has consented to the transmission, storage, and operation execution via the media server 101. The user is provided with the option to change the settings at any time, such as enabling or disabling the use of the media server 101.

[0030] A machine learning model (e.g., a neural network or other type of model) is stored and utilized locally on the user device 115 with specific user permission for one or more operations. The server-side model is used only with user permission. Further, a trained model may be provided for use on the user device 115. During such use, on-device training of the model may be performed with the permission of user 125. The updated model parameters may be transmitted to the media server 101 with the permission of user 115, e.g., to enable federated learning. The model parameters do not include any user data.

[0031] In some embodiments, the media application 103 may be implemented using hardware, which includes a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, the media application 103a may be implemented using software or a combination of hardware and software.

[0032] The media application 103 receives an image depicting a bystander. For example, the media application 103 receives the image from a camera that is part of the user device 115, or the media application 103 receives the image via the network 105. The media application 103 derives a bystander mask from the image by analyzing the image with a bystander segmentation model. The bystander mask identifies a plurality of first pixels in the image that are associated with the bystander.

[0033] In some embodiments, the image and the bystander mask are provided as inputs to a shadow segmentation model. In some embodiments, a shadow classifier model is applied to the image to determine whether the likelihood of the presence of a shadow in the image meets a threshold. In these embodiments, if the likelihood of the presence of a shadow in the image meets the threshold, the shadow segmentation model is used; otherwise, bystander removal is performed without using the shadow segmentation model. The shadow segmentation model derives a shadow mask of the bystander by analyzing the image. The shadow mask identifies a plurality of second pixels in the image that are associated with the shadow of the bystander.

[0034] The media application 103 modifies the image to update the pixel values of the plurality of first pixels and the plurality of second pixels such that the bystander and the shadow are erased from the image. In some embodiments, the media application 103 includes a inpainter model that applies inpainting techniques to update the pixel values.

[0035] Example computing device 200

[0036] Figure 2 is a block diagram of an example computing device 200 that can be used to implement one or more features described herein. The computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a media server 101 for implementing the media application 103a. In another example, the computing device 200 is the user device 115.

[0037] In some embodiments, the computing device 200 includes a processor 235, a memory 237, an input / output (I / O) interface 239, a display 241, a camera 243, and a storage device 245, all of which are coupled via a bus 218. The processor 235 can be coupled to the bus 218 via a signal line 222, the memory 237 can be coupled to the bus 218 via a signal line 224, the I / O interface 239 can be coupled to the bus 218 via a signal line 226, the display 241 can be coupled to the bus 218 via a signal line 228, the camera 243 can be coupled to the bus 218 via a signal line 230, and the storage device 245 can be coupled to the bus 218 via a signal line 232.

[0038] The processor 235 can be one or more processors and / or processing circuits for executing program code and controlling the basic operations of the computing device 200. A "processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. The processor can include a system having: a general-purpose central processing unit (CPU) with one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), application-specific circuitry for implementing functionality, a dedicated processor for implementing neural network model-based processing, a neural circuit, a processor optimized for matrix calculations (e.g., matrix multiplication), or other systems. In some embodiments, the processor 235 can include one or more coprocessors for implementing neural network processing. In some embodiments, the processor 235 can be a processor that processes data to produce a probabilistic output. For example, the output produced by the processor 235 can be imprecise, or can be accurate within a range from the expected output. The processing is not limited to a particular geographical location or have a time limit. For example, the processor can perform its functions in real time, offline, in batch mode, etc. Portions of the processing can be performed at different times and in different locations by different (or the same) processing systems. A computer can be any processor that communicates with a memory.

[0039] The memory 237 is typically provided in the computing device 200 for access by the processor 235 and can be any suitable processor-readable storage medium suitable for storing instructions for execution by the processor or multiple sets of processors and is located separately from and / or integrated with the processor 235, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc. The memory 237 can store software operated by the processor 235 on the computing device 200, which includes the media application 103.

[0040] The memory 237 can include an operating system 262, other applications 264, and application data 266. The other applications 264 can include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, etc. One or more of the methods disclosed herein can operate in several environments and platforms, such as as a stand-alone computer program that can run on any type of computing device, as a web application with a web page, as a mobile application ("app") that runs on a mobile computing device, etc.

[0041] The application data 266 can be data generated by other applications 264 or the hardware of the computing device 200. For example, the application data 266 can include images used by an image gallery application and user actions identified by other applications 264 (e.g., a social networking application), etc.

[0042] The I / O interface 239 can provide functionality for enabling the computing device 200 to interface with other systems and devices. The interfaced devices can be included as part of the computing device 200 or can be separate and communicate with the computing device 200. For example, a network communication device, a storage device (e.g., the memory 237 and / or the storage device 245), and input / output devices can communicate via the I / O interface 239. In some embodiments, the I / O interface 239 can be connected to interface devices such as input devices (keyboard, pointing device, touch screen, microphone, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.).

[0043] Some examples of interfaced devices that can be connected to the I / O interface 239 can include a display 241, which can be used to display content (e.g., images, videos, and / or the user interface of an output application as described herein) and receive touch (or gesture) input from a user. For example, the display 241 can be used to display a user interface including an image in which bystanders and their shadows have been removed. The display 241 can include any suitable display device, such as a liquid crystal display (LCD), a light emitting diode (LED), or a plasma display panel, a cathode ray tube (CRT), a television, a monitor, a touch screen, a three-dimensional display, or other visual display devices. For example, the display 241 can be a flat display screen provided on a mobile device, a plurality of display screens embedded in a glasses form factor or a head-mounted device, or a monitor screen of a computer device.

[0044] The camera 243 can be any type of image capture device capable of capturing images and / or videos. In some embodiments, the camera 243 captures an image or video, and the I / O interface 239 transmits the image or video to the media application 103.

[0045] The storage device 245 stores data related to the media application 103. For example, the storage device 245 can store a training data set including training data such as a plurality of training images, a bystander segmentation model, a shadow segmentation model, a restorer model, etc.

[0046] Figure 2Illustrates an example media application 103 that includes a bounding box module 202, a bystander splitter 204, a shadow splitter 206, a retoucher module 208, and a user interface module 210. In some embodiments, each of the components includes a set of instructions executable by a processor 235 to perform steps discussed in more detail below. In some embodiments, each of the components is stored in a memory 237 of the computing device 200 and can be accessed and executed by the processor 235.

[0047] The components of the media application 103 identify bystanders in an image, identify shadows associated with the bystanders, and update pixel values from the image associated with the bystanders and the shadows to erase the bystanders and the shadows. Figure 3 Illustrates three example images. In a first example image 300, a bystander 305 and a shadow 310 associated with the bystander are illustrated (the main subject of the image is a female facing a camera that captured the image 300). In a second example image 325, the bystander has been removed, but the shadow 330 remains. This example is distracting and unnatural because the shadow 330 is not connected to an object and provides an unsatisfactory result of erasing the bystander 305. In a third example image 350, both the bystander and the shadow have been removed, and the image is satisfactory. Although the following description describes a single bystander and a single shadow, the description also applies to two or more bystanders associated with two or more corresponding shadows.

[0048] Turning to the bounding box module 202, the bounding box module 202 receives an image. The image can be received from a camera 243 of the computing device 200 or from a media server 101 via an I / O interface 239. In some embodiments, the image is part of a video, e.g., a frame of a video.

[0049] The image includes a subject such as a person and one or more bystanders. The bounding box module 202 detects bystanders in the image. A bystander is a person who is not the subject of the image, such as a person walking, running, cycling, standing behind the subject, or otherwise within the image. In different examples, the bystander may be in the foreground (e.g., a person visible in front of an object, e.g., passing between the camera 243 and the object), at the same depth as the subject (e.g., a person standing to the side of the subject), or in the background (behind the object). A bystander can be a human in any pose (e.g., standing, sitting, squatting, lying, jumping, etc.). The bystander can face the camera 243, be at an angle to the camera 243, or can have their back to the camera 243.

[0050] In some embodiments, the bounding box module 202 generates a body bounding box for any person and / or bystander in the image. For example, Figure 3Includes a person and bystanders who are the subjects of the image. In another example, the training images discussed in more detail below may include a person and no bystanders.

[0051] In some embodiments, the body bounding box is a rectangular bounding box that encloses all the pixels of the person and / or bystanders. The bounding box module 202 can detect the person and / or bystanders by performing object recognition, comparing the object with a prior object of a person, and discarding objects that are not a person. In some embodiments, the bounding box module 202 uses a machine learning algorithm - such as a neural network or more specifically, a convolutional neural network - to identify the person and / or bystanders and generate the body bounding box. The body bounding box is associated with the x and y coordinates of the bounding box.

[0052] The bounding box module 202 generates an object bounding box that encloses the objects that overlap with the person and / or bystanders in the image. In some embodiments, the bounding box module 202 generates an object bounding box for all the objects in the image and then identifies the object bounding boxes of the objects that overlap with the bystanders in the image. The intersecting object bounding boxes include, for example, a person on a bicycle, a person on a scooter, a person in a vehicle, etc.

[0053] The bystander segmenter 204 analyzes the image with a bystander segmentation model to derive a bystander mask from the image. A bystander mask that encloses the bystanders is generated. The bystander segmenter 204 identifies a plurality of first pixels in the image that are associated with the bystanders. In some embodiments, the bystander segmenter 204 identifies the plurality of first pixels in the image based on analyzing the pixels within the body bounding box generated by the bounding box module 202.

[0054] In some embodiments, the bystander mask is generated based on: generating superpixels of the image and matching the superpixel centroids with depth map values (e.g., obtained by the camera 243 using a depth sensor or by deriving depth from pixel values) to cluster the detections based on depth. More specifically, the depth values in the masked region can be used to determine a depth range, and the superpixels that fall within the depth range can be identified. However, the depth of the objects attached to the bystanders is the same as the ground, which can result in an over-inclusive bystander mask.

[0055] Another technique for generating the bystander mask includes weighting the depth values based on the proximity of the depth values to the bystander mask, where the weights are represented by a distance transform map. However, the bystander mask can be both over-inclusive with respect to the ground and under-inclusive with respect to some of the objects.

[0056] The shadow segmenter 206 receives an image and a bystander mask as inputs to a shadow segmentation model. The shadow segmentation model analyzes the image and derives a bystander's shadow mask, where the shadow mask identifies a plurality of second pixels in the image associated with the bystander's shadow.

[0057] In some embodiments, the shadow segmentation model is a trained machine learning model. In some embodiments, the shadow segmenter 206 is configured to apply a bystander segmentation model to input data such as application data 266 (e.g., an image captured by user device 115) to output a bystander mask.

[0058] The shadow segmenter 206 uses training data to generate the shadow segmentation model. Model generation can be performed offline (e.g., on media server 101 or another server), and the shadow segmentation model is provided as part of the shadow segmenter 206. For example, the training data can include training images (positive examples) with a subject (e.g., a person, a building, a garden, etc.) and a shadow attached to the subject, and training images (negative examples) with a person who does not cast a shadow. In some embodiments, the training data further includes a body bounding box around the bystander and an object bounding box of an object in the image.

[0059] The training images can be red, green, blue (RGB) images, where the color of each pixel is characterized by three components R, G, and B. In some embodiments, one or more of the RGB images are associated with a segmentation mask corresponding to a person in the RGB image. In some embodiments, one or more of the RGB images are associated with a body bounding box around the person. In some embodiments, one or more of the RGB images include pixels of a person having values associated with a specific color that does not occur in natural images. For example, the shadow segmenter 206 can use a flood fill algorithm on the segmentation mask to identify adjacent values and / or change adjacent values to be associated with an uncommon color based on the similarity of the adjacent values to an initial seed point. In some embodiments, the flood fill algorithm is used to identify the location of a person in the image using an uncommon color instead of disturbing the person pixels of the RGB image.

[0060] The training data can be obtained from any source, e.g., a data repository specifically labeled for training, data for which a license is provided for use as training data for machine learning, etc. In some embodiments, the training can occur on media server 101 that directly provides the training data to user device 115, locally on user device 115, or a combination of both.

[0061] In some embodiments, training data can be obtained from a series of images of a scene captured at different times (e.g., time-lapse photography or an image burst). For example, the series of images can be obtained from security footage. The shadow segmenter 206 can extract three-dimensional camera pose information, organize the images based on the same three-dimensional camera pose, and apply a median filter (or mode filter) across time to pixels to generate a clean-bottom image based on the median or mode of the pixel values. The median filter removes dynamic objects from the image, such as moving vehicles, people, animals, etc. Thus, the clean-bottom image represents the static elements of the scene, such as buildings, sidewalks, and roads.

[0062] The shadow segmenter 206 compares each individual frame in the series of images with the clean-bottom image to identify the dynamic portions of the scene that include people and their shadows. The shadow segmenter 206 then uses the identification of the dynamic portions of the scene to identify potential people in each image by performing object recognition. The shadow segmenter 206 generates a segmentation mask for the person. The shadow segmenter 206 uses a connected components algorithm (e.g., breadth-first search) on the difference between the clean-bottom image and the individual frame to identify all dynamic pixels adjacent to the pixels corresponding to the segmentation mask. These dynamic pixels form a high-quality mask of the person's shadow, which is referred to as the ground truth shadow mask as part of the training data.

[0063] In some embodiments, the shadow segmenter 206 uses criteria to determine whether a candidate training image is included in or excluded from the training images. For example, the shadow segmenter 206 identifies candidate images to be included as part of the training images based on the intersection between the body bounding box and the object bounding box.

[0064] The shadow segmenter 206 can select candidate images to be included in the training images where the person is associated with a body bounding box having a less than threshold overlap value with the vehicle bounding box or a ground truth person mask located within the vehicle bounding box. For example, Figure 4 The first example image 400 and the second example image 410 are illustrated, where the body bounding boxes 402, 412 are located within the object boxes 404, 424 of the vehicle, and thus the images 400, 410 are excluded from the training images. In this example, the threshold overlap value can be 100%, but other values are possible, such as values indicating that the person is too close to the vehicle, partially inside the vehicle, etc.

[0065] The shadow segmenter 206 can select candidate images in which the bystander is associated with a body bounding box having an aspect ratio less than a threshold aspect ratio to include as part of the training images. The shadow segmenter 206 determines the aspect ratio by dividing the width of the body bounding box by the height of the body bounding box. When the aspect ratio exceeds the threshold aspect ratio, the candidate image is likely to be occluded by an object in the candidate image.

[0066] The shadow segmenter 206 can select candidate images to include as part of the training images, where the bystander in the candidate image is associated with a body bounding box having an intersection with the bounding box of the actionable object less than a threshold overlap value. The actionable object can be a bicycle, scooter, tricycle, unicycle, skateboard, etc. Figure 4 The third example image 450 and the fourth example image 460 are illustrated, where the intersections of the body bounding boxes 452, 462 with the bicycle bounding boxes 454, 464 are greater than the threshold overlap value, and thus the images 450, 460 are excluded from the training images.

[0067] The shadow segmenter 206 can select candidate images in which the ground truth shadow mask is between a first threshold size and a second threshold size to include as training images. The first threshold size can be for a ground truth shadow mask that is too small, and the second threshold size can be for a ground truth shadow mask that is too large. For example, Figure 5 The first image 500 and the second image 525 are illustrated. Both the first image and the second image include ground truth shadow masks 505, 530. Since the shadow masks are too small, the first image and the second image are excluded from the training images for failing to exceed the first threshold size. Figure 5 The third image 550 and the fourth image 575 are further illustrated. Both the third image and the fourth image include ground truth shadow masks 555, 580. Since the shadow masks 555 and 580 are larger than the second threshold, the images 550 and 575 are excluded from the training images for exceeding the second threshold size.

[0068] The shadow segmenter 206 can select candidate images in which the candidate images meet an illumination threshold to include as training images. The use of the illumination threshold ensures that images taken under certain conditions - such as during cloudy days - are not used because these conditions prevent shadows from being formed. Figure 6 The first image 600 is illustrated, where clouds have caused the objects in the scene not to cast shadows, and the first image 600 fails to meet the illumination threshold. As a result, the shadow segmenter 206 excludes the first image 600 from the training images. In contrast, Figure 6 the second image 625 in includes sufficient illumination to meet the illumination threshold, where the shadow 630 associated with the tree is very prominent.

[0069] The shadow segmenter 206 may select candidate images in which the candidate image includes a shadow originating from a direction associated with the feet of a person in the image to include as training images. In some embodiments, the shadow segmenter 206 determines the direction from which the shadow should originate based on the geographical location on Earth and the date and time corresponding to the metadata of the candidate image. For example, at a particular date, time, and location, the sun projects a shadow at a particular angle, and the shadow segmenter 206 can determine whether the shadow is associated with a person or other object (such as a nearby object) in the image.

[0070] The shadow segmenter 206 may select candidate images in which the candidate image does not include a shadow to include as training images. A person may not cast a shadow for several reasons, including: the person is standing in a larger shadow, such as a shadow cast by a building; the shadow is not visible, for example because the person's feet are occluded; or the lighting of the scene is shady, resulting in no direct sunlight. Figure 7 Example images that do not include shadows are illustrated, where the first image 700 includes a person walking in a larger shadow 705 that covers a portion of the street, the second image 725 includes a scene that is too shady for the person to cast a shadow, and in the third image 750, a vehicle 755 occludes the person's feet.

[0071] In some embodiments, the shadow segmenter 206 may determine that the calculated shadow fails to include a number of pixels that meets a threshold number of pixels (e.g., at least one pixel, at least five pixels, or another threshold) by running a shadow extraction algorithm, and select candidate images that do not include a shadow to include as training images. In some embodiments, the shadow segmenter 206 determines that the candidate image does not include a shadow by, for example, determining whether the scene was captured during cloudy weather conditions based on applying an illumination model. In some embodiments, the shadow segmenter 206 determines that the candidate image does not include a shadow by requiring a human evaluator to identify applicable images by labeling the image with a label indicating that the image does not include a shadow.

[0072] In some embodiments, the shadow segmenter 206 selects candidate images to include as training images based on human feedback. For example, a human can evaluate the images on a scale (e.g., 1 - 4, good or bad, etc.). In some embodiments, after the shadow segmenter 206 has selected candidate images to include as training images using the above criteria, the human evaluator receives the candidate images. Turning to Figure 8 , example candidate images 800, 825, 850 are illustrated, where the ground truth shadow masks 805, 830, 855 are included as part of the training data.

[0073] The trained machine learning model trained by the shadow splitter 206 may include one or more model forms or structures. For example, the model form or structure may include any type of neural network, such as a linear network, a deep learning neural network implementing multiple layers (e.g., "hidden layers" between an input layer and an output layer, where each layer is a linear network), a convolutional neural network (e.g., a network that splits or partitions input data into multiple parts or tiles, uses one or more neural network layers to process each tile individually, and aggregates the results of the processing from each tile), a sequence-to-sequence neural network (e.g., a network that receives sequential data such as words in a sentence, frames in a video, etc. as input and produces a result sequence as output), and so on.

[0074] The model form or structure may specify the connectivity between individual nodes and the organization of nodes into layers. For example, the nodes of the first layer (e.g., the input layer) may receive data as input data or application data 266. For example, when using the trained model for, e.g., the analysis of an image, such data may include, e.g., one or more pixels per node. Subsequent intermediate layers may receive the outputs of the nodes of the previous layer as input according to the connectivity specified in the model form or structure. These layers may also be referred to as hidden layers. The final layer (e.g., the output layer) produces the output of the machine learning model. In some implementations, the model form or structure also specifies the number and / or type of nodes in each layer.

[0075] In different implementations, the trained shadow segmentation model may include one or more models. One or more of the models in the model may include multiple nodes, which are arranged into layers according to the model structure or form. In some implementations, a node may be, for example, a memoryless computational node configured to process a unit of input to produce a unit of output. The computation performed by the node may include, for example, multiplying each of the multiple node inputs by a weight, obtaining a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce the node output. In some implementations, the computation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some implementations, the step / activation function may be a non-linear function. In various implementations, such computations may include operations such as matrix multiplication. In some implementations, the computations performed by multiple nodes may be executed in parallel, e.g., using multiple processor cores of a multi-core processor, using individual processing units of a graphics processing unit (GPU), or using dedicated neural circuitry. In some implementations, a node may include memory, e.g., may be capable of storing one or more earlier inputs and using the one or more earlier inputs when processing subsequent inputs. For example, a node with memory may include a long short-term memory (LSTM) node. The LSTM node may use the memory to maintain a "state" that permits the node to act like a finite state machine (FSM).

[0076] In some implementations, the trained model can include embeddings or weights for individual nodes. For example, the model can start as a number of nodes organized into layers as specified by the model form or structure. At initialization, the corresponding weights can be applied to the connections between each pair of nodes connected according to the model form - for example, the nodes in successive layers of a neural network. For example, the corresponding weights can be randomly assigned or initialized to default values. Then, the model can be trained (e.g., using training data) to produce results.

[0077] Training can include applying supervised learning techniques. In supervised learning, the training data can include multiple inputs (e.g., multiple training images) and corresponding ground truth outputs for each input (e.g., the ground truth shadow mask for each of the multiple training images). Based on the comparison of the output of the model (e.g., the predicted shadow mask) with the ground truth output (e.g., the ground truth shadow mask), the values of the weights are automatically adjusted, e.g., in a way that increases the probability that the model produces the ground truth shadow output of the image.

[0078] In various implementations, the trained model includes a set of weights or embeddings corresponding to the model structure. In some implementations, the trained shadow segmentation model can include an initial set of weights downloaded, for example, from a server providing the weights. In various implementations, the trained shadow segmentation model includes a set of weights or embeddings corresponding to the model structure. In implementations where data is omitted, the shadow segmenter 206 can generate a trained shadow segmentation model based on, for example, prior training performed by the developer of the shadow segmenter 206, a third party, etc.

[0079] In some embodiments, in the case where the shadow segmentation model includes a convolutional neural network trained using supervised learning, the training of the shadow segmentation model can include, for each of a number of training images, obtaining a predicted shadow mask based on that training image. The shadow segmentation model can calculate a loss value based on the comparison of the predicted shadow mask for the image and the ground truth shadow mask (included in the training data). The shadow segmentation model can update the weights of one or more nodes of the convolutional neural network based on the loss value (e.g., in such a way that, after adjusting and running another training cycle, the loss value is reduced until the loss value is below a threshold).

[0080] Figure 9An example flowchart illustrating a method 900 for training a machine learning model to output a shadow mask. An RGB image and a segmentation mask 905 of a person are provided as inputs to a machine learning model 910, such as the shadow segmentation model described above. In some embodiments, the machine learning model 910 is a convolutional neural network (CNN), such as UNet. The machine learning model 910 outputs a predicted shadow mask 915. The predicted shadow mask 915 is compared with a ground truth shadow mask 920 to determine a loss value 925. In some embodiments, the loss value 925 is calculated using binary cross-entropy loss, which is a comparison of each predicted probability in the predicted probabilities with the actual occlusion output of 0 or 1. The loss value 925 is calculated as a score reflecting the distance between the predicted probabilities and the actual values. This score is used to update the weights of the model during training. This process can be continued with additional training images until the score meets a threshold score (or the training data / computation budget is exhausted) and training is complete.

[0081] In some embodiments, in the case where the shadow segmentation model includes a convolutional neural network trained using supervised learning, the training of the shadow segmentation model can include, for each of a plurality of training images, obtaining a predicted shadow mask based on the training image, wherein one or more pixels of the training image associated with a person have values associated with a specific color that does not occur in natural images. The shadow segmentation model can calculate a loss value based on a comparison of the predicted shadow mask and the ground truth shadow mask for the image. The shadow segmentation model can update the weights of one or more nodes of the convolutional neural network based on the loss value.

[0082] In some embodiments, the shadow segmenter 206 receives an image and a bystander mask. The shadow segmenter 206 provides the image and the bystander mask as inputs to the shadow segmentation model. In some embodiments, the shadow segmentation model outputs a shadow mask.

[0083] Figure 10A An example flowchart illustrating a method 1000 for outputting a combined bystander and shadow mask. At step 1005, bystander detection is performed. For example, in this case, a girl 1010 is detected to the right of a person 1015 who is the subject of an image. For each bystander, a person segmenter 1020 is applied and a shadow segmenter 1025 is applied. The person segmenter 1020 outputs a bystander mask 1022, and the shadow segmenter 1025 outputs a shadow mask 1027. At step 1030, the shadow mask and the bystander mask are combined.

[0084] Figure 10BAn example flowchart illustrating another method 1050 for outputting a combined bystander and shadow mask. At step 1055, bystander detection is performed. For example, in this case, a girl 1060 is detected to the right of a person 1065 who is the subject of the image. For each bystander, a person segmenter 1070 is applied, a shadow classifier 1075 determines whether the bystander casts a shadow, and if the bystander does cast a shadow, a shadow segmenter 1080 is applied. The person segmenter 1070 outputs a bystander mask 1072, the shadow classifier 1075 determines that the bystander casts a shadow, and the shadow segmenter 1080 outputs a shadow mask 1082. At step 1085, the shadow mask and the bystander mask are combined. If the shadow classifier 1075 determines that the bystander does not cast a shadow, the bystander mask is used as-is for further tasks.

[0085] In some embodiments, a restorer module 208 modifies the image to update pixel values of a plurality of first pixels and a plurality of second pixels such that the bystander and the shadow are erased from the image. The restorer module 208 may apply inpainting techniques to update pixel values of the plurality of first pixels and the plurality of second pixels. In some embodiments, the restorer module 208 uses a restorer model to apply inpainting techniques.

[0086] In some embodiments, the restorer module 208 generates a restored image that updates pixel values of pixels within the bystander mask and the shadow mask to match the background in the image such that the bystander and the shadow are erased from the image. The restorer module 208 may update pixel values of a plurality of first pixels that are part of the bystander mask independently of pixel values of the plurality of second pixels that are part of the shadow mask. In some embodiments, the restorer module 208 combines the bystander mask and the shadow mask to form a combined mask, and the restorer module 208 updates pixel values of pixels within the combined mask such that the bystander and the shadow are erased from the image.

[0087] Pixels matching the background may be based on another image of the same location without the subject and / or bystander. Alternatively, the restorer module 208 may match pixels removed from the bystander mask based on pixels surrounding the pixels included in the bystander mask. For example, in the case where the bystander is standing on the ground, such as in Figure 3 the third image 350 in, the restorer module 208 replaces the pixel with pixels of the ground. Other inpainting techniques are possible, including machine learning-based inpainting techniques.

[0088] In some embodiments, the above techniques can be applied to video. For example, components of media application 103 can derive a bystander mask and a shadow mask for a first image and then apply temporal smoothing between the images to keep the mask pixels consistent, or slowly change the mask over multiple frames so that there are no discontinuities across frames. In some embodiments, optical flow associated with a person is used to determine the direction of movement, and the mask pixels are updated accordingly across the frames of the video. In some embodiments, the bystander mask and the shadow mask are determined for a subset of the frames and interpolation is used to extend the mask across the entire video.

[0089] The user interface module 210 generates a user interface. In some embodiments, the user interface module 210 includes a set of instructions executable by the processor 235 to generate the user interface. In some embodiments, the user interface module 210 is stored in the memory 237 of the computing device 200 and is accessible and executable by the processor 235.

[0090] The user interface module 210 generates a user interface including the repaired image. In some embodiments, the user interface includes options for editing the repaired image, sharing the repaired image, adding the repaired image to an album, etc.

[0091] Example flowchart

[0092] Figure 11 An example flowchart illustrating method 1100 for modifying an image to erase bystanders and their shadows from the image based on the derived bystander mask and the derived shadow mask. Method 1100 can be executed by Figure 2 the computing device 200 in Figure 1 In some embodiments, method 1100 is executed by the user device 115, the media server 101, or partially on the user device 115 and partially on the media server 101.

[0093] Figure 11 Method 1100 of can begin at block 1102. At block 1102, a bystander mask is derived from the image by analyzing the image with a bystander segmentation model. The image depicts a bystander, and the bystander mask identifies multiple first pixels in the image associated with the bystander. Block 1102 can be followed by block 1104.

[0094] At block 1104, an optional step is to use a shadow classifier model to determine the likelihood of the presence of a shadow based on the image and the bystander mask. If the likelihood of the presence of a shadow meets a threshold, block 1104 can be followed by block 1106.

[0095] At block 1106, a shadow mask of the bystander is derived by analyzing the image with a shadow segmentation model. The image is provided as input to the shadow segmentation model. The shadow mask identifies a plurality of second pixels in the image associated with the bystander's shadow.

[0096] In some embodiments, the shadow segmentation model includes a convolutional neural network trained using supervised learning, where training the convolutional neural network is performed with a training data set that includes a plurality of training images and a ground truth shadow mask for each of the plurality of training images. Training can include, for each of the plurality of training images: obtaining a predicted shadow mask based on the training image and a ground truth shadow mask associated with the person in the training image, calculating a loss value based on a comparison of the predicted shadow mask and the ground truth shadow mask for the image, and updating the weights of one or more nodes of the convolutional neural network based on the loss value. In some embodiments, training includes, for each of the plurality of training images: obtaining a predicted shadow mask based on the training image, where one or more pixels of the training image associated with the person have values associated with a specific color that does not occur in natural images, calculating a loss value based on a comparison of the predicted shadow mask and the ground truth shadow mask for the image, and updating the weights of one or more nodes of the convolutional neural network based on the loss value.

[0097] In some embodiments, the plurality of training images further includes a segmentation mask corresponding to the person in each of the plurality of training images, and the training data set is generated by including candidate images selected from the group consisting of: the person is associated with a body bounding box that has a vehicle bounding box with a less than threshold overlap value, the body bounding box is associated with an aspect ratio less than a threshold aspect ratio, the segmentation mask is outside the vehicle bounding box, the intersection of the body bounding box and the actionable object bounding box is less than a threshold overlap value, the ground truth shadow mask is between a threshold first size and a threshold second size, the plurality of training images satisfy an illumination threshold, the candidate image includes a shadow originating from a direction associated with the person's feet, and combinations thereof. In some embodiments, one or more of the plurality of training images are associated with an empty ground truth shadow mask. In some embodiments, an empty ground truth shadow mask is a mask that identifies zero pixels as part of the shadow cast by the person, thereby ensuring that the training images include negative examples where there is no shadow (e.g., images taken at noon, images where the person is within the shadow of a larger object such as a building, etc.).

[0098] In some embodiments, one or more of the training images are generated from a series of candidate images of a scene including a person captured at different times by generating a clean background image containing the static elements of the scene, comparing each candidate image in the series of candidate images with the clean background image to identify the dynamic parts of the scene, generating a segmentation mask for the person in the series of images, and determining that pixels of a ground truth shadow mask are adjacent to pixels corresponding to the segmentation mask.

[0099] In some embodiments, the image depicts two or more bystanders, and the deriving, the obtaining, and the modifying are performed for each of the two or more bystanders. In some embodiments, the image is a single frame of a video. After block 1106 can be block 1108.

[0100] At block 1108, an optional step is to merge the bystander mask and the shadow mask. After block 1108 can be block 1110. At least one pixel of a first plurality of pixels of the bystander mask may be adjacent to at least one pixel of a second plurality of pixels of the shadow mask.

[0101] At block 1110, modify the image (e.g., using inpainting techniques) to update the pixel values of the plurality of first pixels and the plurality of second pixels such that the bystander and the shadow are erased from the image. After block 1110 can be block 1112.

[0102] At block 1112, an optional step is to apply inpainting techniques to update the pixel values of the plurality of first pixels and the plurality of second pixels. The inpainting techniques may be performed by an inpainting model.

[0103] In addition to the above description, controls may be provided to the user that allow the user to make choices regarding whether and when the systems, procedures, or features described herein may enable the collection of user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current location) and whether to send content or communications from the server to the user. Additionally, before storing or using certain data, it may be processed in one or more ways such that personally identifiable information is removed. For example, the user's identity may be processed such that the user's personally identifiable information cannot be determined, or the user's geographical location may be generalized (such as to a city, zip code, or state level) if location information is obtained such that the user's specific location cannot be determined. Thus, the user has control over what information is collected about the user, how that information is used, and what information is provided to the user.

[0104] In the foregoing description, numerous specific details have been set forth for purposes of explanation in order to provide a thorough understanding of the present specification. It will be apparent, however, to one skilled in the art that the present disclosure may be practiced without these specific details. In some instances, well-known structures and devices are shown in block diagram form in order to avoid obscuring the description. For example, the above may describe embodiments mainly with reference to user interfaces and specific hardware. However, the embodiments may be applied to any type of computing device that can receive data and commands, as well as any peripheral device that provides services.

[0105] References in this specification to "some embodiments" or "some examples" mean that a particular feature, structure, or characteristic described in connection with the embodiments or examples may be included in at least one implementation of the description. The phrase "in some embodiments" appearing in various places in this specification does not necessarily refer to the same embodiment.

[0106] Some portions of the detailed descriptions above are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulation of physical quantities. Although not necessarily, these quantities typically take the form of electrical or magnetic data capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these data as bits, values, elements, symbols, characters, terms, numbers, etc.

[0107] However, it should be borne in mind that all of these and similar terms are to be associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, as will be apparent from the following discussion, it should be understood that throughout the description, discussions using terms including "processing" or "computing" or "operating" or "determining" or "displaying" or the like refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities within the registers and memories of the computer system and transforms that data into other data similarly represented as physical quantities within the memories or registers or other such information storage, transmission, or display devices of the computer system.

[0108] Embodiments of this specification may also relate to a processor for performing one or more steps of the above method. The processor may be a dedicated processor selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, which includes but is not limited to: any type of disk (including optical disks), ROM, CD-ROM, magnetic disks, RAM, EPROM, EEPROM, magnetic cards or optical cards, flash memory (including USB keys with non-volatile memory), or any type of medium suitable for storing electronic instructions, each of which is coupled to the computer system bus.

[0109] This specification may take the form of some fully hardware embodiments, some fully software embodiments, or some embodiments that include both hardware elements and software elements. In some embodiments, this specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.

[0110] In addition, this specification may take the form of a computer program product that can be accessed from a computer-usable or computer-readable medium, which provides program code for use by or in conjunction with a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium may be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0111] A data processing system suitable for storing or executing program code will include at least one processor directly or indirectly coupled to memory elements through a system bus. The memory elements may include local memory, mass storage, and cache memory employed during the actual execution of the program code, and the cache memory provides temporary storage of at least some of the program code to reduce the number of times code must be retrieved from mass storage during execution.

Claims

1. A computer-implemented method, comprising: Deriving a bystander mask from the image by analyzing the image with a bystander segmentation model, wherein the image depicts a bystander and the bystander mask identifies a plurality of first pixels in the image associated with the bystander; Deriving a shadow mask of the bystander by analyzing the image with a shadow segmentation model, wherein the image and the bystander mask are provided as inputs to the shadow segmentation model, and wherein the shadow mask identifies a plurality of second pixels in the image associated with the shadow of the bystander; and Modifying the image to update pixel values of the plurality of first pixels and the plurality of second pixels such that the bystander and the shadow are erased from the image.

2. The computer-implemented method according to claim 1, further comprising: Determining a likelihood of the presence of the shadow using a shadow classifier model based on the image and the bystander mask, wherein deriving the shadow mask of the bystander is performed when the likelihood of the presence of the shadow meets a threshold.

3. The computer-implemented method according to claim 1, wherein, Modifying the image includes: applying an inpainting technique to update the pixel values of the plurality of first pixels and the plurality of second pixels.

4. The computer-implemented method according to claim 3, wherein, The inpainting technique is performed by an inpainting model.

5. The computer-implemented method according to claim 1, further comprising: Merging the bystander mask and the shadow mask before the modification.

6. The computer-implemented method according to claim 1, wherein, The shadow segmentation model includes a convolutional neural network trained using supervised learning, wherein training the convolutional neural network is performed with a training data set that includes a plurality of training images, a segmentation mask associated with a person in each of the plurality of training images, and a ground truth shadow mask for each of the plurality of training images, and wherein the training includes, for each of the plurality of training images: Obtaining a predicted shadow mask based on the training image and the segmentation mask associated with the person in the training image; Calculating a loss value based on a comparison of the predicted shadow mask for the image and the ground truth shadow mask; and Updating weights of one or more nodes of the convolutional neural network based on the loss value.

7. The computer-implemented method according to claim 6, wherein: The training data set is generated by including in the plurality of training images candidate images selected from the group consisting of: the person is associated with a body bounding box that has a vehicle bounding box overlap value less than a threshold, the body bounding box is associated with an aspect ratio less than a threshold aspect ratio, the segmentation mask is outside the vehicle bounding box, the intersection of the body bounding box and an actionable object bounding box is less than the threshold overlap value, the ground truth shadow mask is between a threshold first size and a threshold second size, the plurality of training images meet an illumination threshold, the candidate images include shadows originating from a direction associated with the person's feet, and combinations thereof.

8. The computer-implemented method according to claim 6, wherein, One or more of the plurality of training images are associated with an empty ground truth shadow mask.

9. The computer-implemented method according to claim 6, wherein, One or more of the plurality of training images are generated from a series of candidate images of a scene including a person captured at different times by: Generate a clean background image that includes static elements of the scene; Compare each candidate image in the series of candidate images with the clean background image to identify the dynamic part of the scene; Generate the segmentation mask for the person in the series of images; And Determine that pixels of the ground truth shadow mask corresponding to the segmentation mask are adjacent to pixels corresponding to the segmentation mask.

10. The computer-implemented method according to claim 1, wherein, The shadow segmentation model includes a convolutional neural network trained using supervised learning, wherein training the convolutional neural network is performed using a training data set that includes a plurality of training images and a ground truth shadow mask for each of the plurality of training images, and wherein the training includes, for each of the plurality of training images: Obtain a predicted shadow mask based on the training image, wherein one or more pixels of the training image associated with a person have values associated with a specific color that does not occur in natural images; Calculate a loss value based on a comparison of the predicted shadow mask and the ground truth shadow mask for the image; and Update weights of one or more nodes of the convolutional neural network based on the loss value.

11. The computer-implemented method according to claim 1, wherein, The image depicts two or more bystanders, and wherein the deriving, the obtaining, and the modifying are performed for each of the two or more bystanders.

12. The computer-implemented method according to claim 1, wherein, The image is a single frame of a video.

13. The computer-implemented method according to claim 1, wherein, At least one pixel of the first plurality of pixels of the bystander mask is adjacent to at least one pixel of the second plurality of pixels of the shadow mask.

14. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations that include: Derive a bystander mask from the image by analyzing the image with a bystander segmentation model, wherein the image depicts a bystander and the bystander mask identifies a plurality of first pixels in the image associated with the bystander; Derive a shadow mask of the bystander from the image by analyzing the image with a shadow segmentation model, wherein the image and the bystander mask are provided as inputs to the shadow segmentation model, and wherein the shadow mask identifies a plurality of second pixels in the image associated with the shadow of the bystander; and Modify the image to update pixel values of the plurality of first pixels and the plurality of second pixels such that the bystander and the shadow are erased from the image.

15. The non-transitory computer-readable medium according to claim 14, wherein, The operations further include: determining a likelihood of the presence of the shadow using a shadow classifier model based on the image and the bystander mask, wherein deriving the shadow mask of the bystander is performed when the likelihood of the presence of the shadow meets a threshold.

16. The non-transitory computer-readable medium according to claim 14, wherein, Modifying the image includes: applying an inpainting technique to update the pixel values of the plurality of first pixels and the plurality of second pixels.

17. The non-transitory computer-readable medium according to claim 14, wherein, The operations further include: merging the bystander mask and the shadow mask before the modifying.

18. A computing device, comprising: A processor; And A memory coupled to the processor, with instructions stored on the memory, which when executed by the processor cause the processor to perform operations, the operations including: Deriving a bystander mask from the image by analyzing the image with a bystander segmentation model, wherein the image depicts a bystander, and the bystander mask identifies a plurality of first pixels in the image associated with the bystander; Deriving a shadow mask of the bystander by analyzing the image with a shadow segmentation model, wherein the image and the bystander mask are provided as inputs to the shadow segmentation model, and wherein the shadow mask identifies a plurality of second pixels in the image associated with the shadow of the bystander; and Modifying the image to update the pixel values of the plurality of first pixels and the plurality of second pixels such that the bystander and the shadow are erased from the image.

19. The computing device according to claim 18, wherein, The operations further include: determining a likelihood of the presence of the shadow using a shadow classifier model based on the image and the bystander mask, wherein deriving the shadow mask of the bystander is performed when the likelihood of the presence of the shadow meets a threshold.

20. The computing device according to claim 18, wherein, Modifying the image includes: applying an inpainting technique to update the pixel values of the plurality of first pixels and the plurality of second pixels.