Method, system and computer readable medium for obtaining foreground video

By receiving the depth data and color data of the video frame, downsampling is performed to generate an initial segmentation mask, calculating pixel weights and fine segmentation. The head enclosure box and three-point graph optimize the segmentation process is solved, and the challenges in image and video segmentation are achieved, and high-quality promising video segmentation is achieved.

CN114072850BActive Publication Date: 2025-06-06GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080044658.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-15
Filing Date
2020-04-15
Publication Date
2025-06-06
Estimated Expiration
2040-04-15

AI Technical Summary

Technical Problem

The prior art faces challenges such as color disguise, moving backgrounds, color shadows, foreground scenes and inaccurate depth data in image and video segmentation, resulting in inaccurate foreground and background segmentation.

Method used

By receiving the depth data and color data of the video frame, the initial segmentation mask is generated after downsampling, the pixel weight is calculated and fine segmentation is performed. The segmentation process is optimized using the head enclosure box and the three-point graph, and the segmentation accuracy is improved by combining the graph cutting technology and the time low-pass filter.

Benefits of technology

It realizes high-quality segmentation of foreground video, can process and provide clear foreground video in real time, reducing background interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114072850B_ABST
    Figure CN114072850B_ABST
Patent Text Reader

Abstract

Embodiments described herein relate to methods, systems, and computer-readable media for obtaining foreground video. In some embodiments, a method includes receiving a plurality of video frames including depth data and color data. The method further includes downsampling frames of the video. The method further includes, for each frame, generating an initial segmentation mask that classifies each pixel of the frame as a foreground pixel or a background pixel. The method further includes determining a tripartite map that classifies each pixel of the frame as a known background, a known foreground, or an unknown. The method further includes, for each pixel classified as unknown, calculating a weight and storing the weight in a weight map. The method further includes performing fine segmentation to obtain a binary mask for each frame. The method further includes upsampling multiple frames based on the binary mask for each frame to obtain a foreground video.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Segmenting an image into foreground and background parts is used in many image and video applications. For example, in video applications such as video conferencing, remote presentation in virtual reality (VR) or augmented reality (AR), it may be necessary to remove the background and replace it with a new background. In another example, segmentation is often used for portrait mode and bokeh, where the background part of the image is blurred.

[0002] There are many challenges in the segmentation of images or videos. One challenge is color camouflage, which occurs when a video or image includes a foreground object whose color is similar to that in the background. Other challenges include moving or changing backgrounds, color shading, scenes without foreground parts (e.g., a video conference without people present), and lighting changes in the scene during the time the image or video was captured.

[0003] Some image capture devices (e.g., desktop cameras, cameras in mobile devices, etc.) can capture depth data along with color data of an image or video. Such depth data can be used for segmentation. However, such depth data is often inaccurate, for example, due to the quality of the depth sensor, lighting conditions and objects in the captured scene, etc. Inaccurate data is another challenge when using depth data to segment images or videos.

[0004] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. The work of the presently named inventors described in this background section, as well as aspects of the description that may not otherwise be qualified as prior art at the time of filing, should neither be explicitly nor implicitly identified as prior art for the present disclosure. Summary of the invention

[0005] Embodiments described herein relate to methods, systems, and computer-readable media for obtaining a foreground video. In some embodiments, a computer-implemented method includes receiving a plurality of frames of a video. Each frame of the video may include depth data and color data for a plurality of pixels. The method further includes downsampling each of the plurality of frames of the video. The method further includes, after downsampling, and for each frame, generating an initial segmentation mask that classifies each pixel of the frame as a foreground pixel or a background pixel based on the depth data. The method further includes, for each frame, determining a tripartite map that classifies each pixel of the frame as one of a known background, a known foreground, or an unknown. The method further includes, for each frame and for each pixel of a frame classified as unknown, calculating a weight for the pixel, and storing the weight in a weight map. The method further includes, for each frame, performing a fine segmentation based on the color data, the tripartite map, and the weight map to obtain a binary mask for the frame. The method further includes upsampling a plurality of frames based on the binary mask of each corresponding frame to obtain a foreground video.

[0006] In some embodiments, generating the initial segmentation mask may include setting the pixel as a foreground pixel if the depth value associated with the pixel is within the depth range, and setting the pixel as a background pixel if the depth value associated with the pixel is outside the depth range. In some embodiments, generating the initial segmentation mask may further include performing one or more of a morphological opening process or a morphological closing process.

[0007] In some embodiments, the method may further include detecting a head bounding box based on one or more of the color data or the initial segmentation mask. In some embodiments, detecting the head bounding box may include converting the frame to grayscale and performing histogram equalization. The method may further include detecting one or more faces in the frame by Haar cascade face detection. Each of the one or more faces may be associated with a face region including facial pixels of the face.

[0008] In some embodiments, the method may further include determining whether each of the one or more faces is valid. In some embodiments, a face is determined to be valid if it is verified that a threshold proportion of pixels of the facial region of the face are classified as foreground pixels in the initial segmentation mask and at least a threshold percentage of pixels of the facial region of the face meet the skin color criteria. In some embodiments, the method may further include, for each face determined to be valid, expanding the facial region of each face to obtain a head region corresponding to the face. The head bounding box may include a head region for each face determined to be valid. In some embodiments, if no face is determined to be valid, the method may further include analyzing the initial segmentation mask to detect a head, and determining whether the head is valid based on the head skin color verification. If the head is determined to be valid, the method may further include selecting a bounding box associated with the head as the head bounding box.

[0009] In some embodiments, generating the initial segmentation mask may include assigning a mask value to each pixel. In the initial segmentation mask, each foreground pixel may be assigned a mask value of 255, and each background pixel may be assigned a mask value of 0. In these embodiments, determining the tri-map may include, for each pixel of the frame that is not in the head bounding box, calculating the L1 distance between the pixel position of the pixel and a mask boundary of the initial segmentation mask, wherein the mask boundary includes the position where at least one foreground pixel in the initial segmentation mask is adjacent to at least one background pixel. If the L1 distance meets a foreground distance threshold and the pixel is classified as a foreground pixel, the method may further include classifying the pixel as a known foreground. If the L1 distance meets a background distance threshold and the pixel is classified as a background pixel, the method may further include classifying the pixel as a known background, and if the pixel is not classified as a known foreground and is not classified as a known background, the method may further include classifying the pixel as unknown.

[0010] Determining the triplicate map may further include, for each pixel in the head bounding box, identifying whether the pixel is a known foreground, a known background, or an unknown. Identifying whether the pixel is a known foreground, a known background, or an unknown may include classifying the pixel as a known foreground if the pixel is within an inner mask determined for the head bounding box, classifying the pixel as a known background if the pixel is outside an outer mask determined for the head bounding box, and classifying the pixel as unknown if the pixel is not classified as a known foreground and a known background.

[0011] In some embodiments, the method may further include detecting whether there is a uniform bright background near the hair region of the head in the head bounding box, and if the uniform bright background is detected, extending the hair region of the head based on the head bounding box, the color data and the initial segmentation mask. In these embodiments, after the hair region is extended, the expansion size of the outer mask is increased.

[0012] In some embodiments, the method may further include maintaining a background image of the video, wherein the background image is a color image having the same size as each frame of the video. The method may further include updating the background image based on the trigraph before performing the fine segmentation. In these embodiments, calculating the weight of the pixel may include calculating the Euclidean distance between the color of the pixel and the background color of the background image, determining the probability that the pixel is a background pixel based on the Euclidean distance, and assigning a background weight to the pixel in the weight map if the probability meets a background probability threshold.

[0013] In some embodiments, the method may further include identifying one or more skin regions in the frame based on skin color detection. The one or more skin regions exclude a facial region. In these embodiments, the method may further include, for each pixel of the frame within the one or more skin regions, classifying the pixel as unknown and assigning a zero weight to the pixel in a weight map. The method may further include assigning a background weight to the pixel in the weight map if the color of the pixel and the background color of the background image meet a similarity threshold. The method may further include assigning a foreground weight to the pixel in the weight map if the color of the pixel is a skin color. The method may further include assigning a foreground weight to the pixel in the weight map if the color of the pixel and the background color of the background image meet a dissimilarity threshold.

[0014] In some embodiments, multiple frames of the video may be in a sequence. In these embodiments, the method may further include, for each frame, comparing the initial segmentation mask to a previous frame binary mask of an immediately previous frame in the sequence to determine a proportion of pixels of the frame that are classified similar to pixels of the previous frame. The method may further include calculating a global coherence weight based on the proportion. In these embodiments, calculating the weight of the pixel and storing the weight in a weight map may include determining the weight based on the global coherence weight and a distance between the pixel and a mask boundary of the previous frame binary mask. In some embodiments, the weight of the pixel is positive if the corresponding pixel is classified as a foreground pixel in the previous frame binary mask, and the weight is negative if the corresponding pixel is not classified as a foreground pixel in the previous frame binary mask.

[0015] In some implementations, performing fine segmentation can include applying a graph cut technique to the frame, wherein the graph cut technique is applied to pixels classified as unknown.

[0016] In some embodiments, the method may further include, after performing the fine segmentation, applying a temporal low pass filter to the binary mask. Applying the temporal low pass filter may update the binary mask based on similarities between one or more previous frames.

[0017] Some embodiments may include a non-transitory computer-readable medium having instructions stored thereon. When the instructions are executed by one or more hardware processors, the instructions cause the processor to perform operations including receiving multiple frames of a video. Each frame of the video may include depth data and color data for multiple pixels. The operation further includes downsampling each of the multiple frames of the video. The operation further includes, after downsampling, and for each frame, generating an initial segmentation mask that classifies each pixel of the frame as a foreground pixel or a background pixel based on the depth data. The operation further includes, for each frame, determining a tripartite map that classifies each pixel of the frame as one of a known background, a known foreground, or an unknown. The operation further includes, for each frame and for each pixel of a frame classified as unknown, calculating a weight for the pixel and storing the weight in a weight map. The operation further includes, for each frame, performing fine segmentation based on the color data, the tripartite map, and the weight map to obtain a binary mask for the frame. The operation further includes upsampling multiple frames based on the binary mask of each corresponding frame to obtain a foreground video.

[0018] In some embodiments, the non-transitory computer-readable medium may include further instructions that, when executed by one or more hardware processors, cause the processor to perform operations including maintaining a background image of the video, the background image being a color image of the same size as each frame of the video. The operation may further include updating the background image based on the trigraph before performing fine segmentation. In these embodiments, the operation of calculating the weight of the pixel may include calculating the Euclidean distance between the color of the pixel and the background color of the background image, determining the probability that the pixel is a background pixel based on the Euclidean distance, and assigning a background weight to the pixel in the weight map if the probability meets a background probability threshold.

[0019] Some embodiments may include a system comprising one or more hardware processors coupled to a memory. The memory may include instructions stored thereon. When the instructions are executed by the one or more hardware processors, the instructions cause the processor to perform an operation, the operation comprising receiving a plurality of frames of a video. Each frame of the video may include depth data and color data of a plurality of pixels. The operation further comprises downsampling each of the plurality of frames of the video. The operation further comprises, after downsampling, and for each frame, generating an initial segmentation mask that classifies each pixel of the frame as a foreground pixel or a background pixel based on the depth data. The operation further comprises, for each frame, determining a tripartite map that classifies each pixel of the frame as one of a known background, a known foreground, or an unknown. The operation further comprises, for each frame and for each pixel of a frame classified as unknown, calculating a weight for the pixel, and storing the weight in a weight map. The operation further comprises, for each frame, performing a fine segmentation based on the color data, the tripartite map, and the weight map to obtain a binary mask for the frame. The operation further comprises upsampling a plurality of frames based on the binary mask of each corresponding frame to obtain a foreground video. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The patent or application file contains at least one drawing drawn in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0021] Figure 1 is a block diagram of an example network environment that can be used with one or more implementations described herein.

[0022] Figure 2 is a flow chart illustrating an example method of determining a foreground mask according to some embodiments.

[0023] Figure 3 is a flow chart illustrating an example method of detecting a head bounding box according to some implementations.

[0024] Figure 4 is a flow chart illustrating an example method of generating a tripartite map of a head region according to some implementations.

[0025] Figure 5 An example video frame and the corresponding initial segmentation mask are shown.

[0026] Figure 6 Two example images with separated foreground and background without using a tripartite map are shown.

[0027] Figure 7 An example image with identified tripartite portions is shown.

[0028] Figure 8 Four frames of an input video and corresponding output frames of an output video including a foreground generated by segmenting the input video are shown according to some embodiments.

[0029] Fig. 9 is a block diagram of an example computing device that may be used with one or more implementations described herein. DETAILED DESCRIPTION

[0030] The embodiments described herein generally relate to segmenting a video or image into background and foreground portions. In particular, the embodiments relate to obtaining the foreground portion of a video or image.

[0031] One or more embodiments described herein include methods, devices, and computer-readable media with instructions for obtaining a foreground video. In some embodiments, the foreground video may have a blank background. In addition, some embodiments may generate a composite video that includes a foreground video segmented from a captured scene and overlaid on a background different from the original captured background. In some embodiments, the segmentation of the foreground video is performed in real time, for example, to provide the foreground video in a video conference.

[0032] Figure 1 1 shows a block diagram of an example network environment 100 that can be used in some embodiments described herein. In some embodiments, the network environment 100 includes one or more server systems, such as, Figure 1 102 in the server system. For example, the server system 102 can communicate with the network 130. The server system 102 may include a server device 104 and a database 106 or other storage device. In some embodiments, the server device 104 may provide a video application 152b, such as a video call application, an augmented reality application, a virtual reality application, etc.

[0033] The network environment 100 may also include one or more client devices, such as client devices 120, 122, 124, and 126, which may communicate with each other and / or with the server system 102 via a network 130. The network 130 may be any type of communication network, including one or more of the Internet, a local area network (LAN), a wireless network, a switch or hub connection, etc. In some embodiments, the network 130 may include, for example, a wireless network using a peer-to-peer wireless protocol (e.g., An example of peer-to-peer communication between two client devices 120 and 122 is shown by arrow 132.

[0034] For ease of explanation, Figure 1One box for server system 102, server device 104, database 106 is shown, and four boxes for client devices 120, 122, 124, and 126 are shown. Server boxes 102, 104, and 106 can represent multiple systems, server devices, and network databases, and the boxes can be provided in a different configuration than shown. For example, server system 102 can represent multiple server systems that can communicate with other server systems via network 130. In some embodiments, server system 102 can include, for example, a cloud hosting server. In some examples, a database 106 and / or other storage devices separate from server device 104 can be provided in the server system box and can communicate with server device 104 and other server systems via network 130.

[0035] Moreover, there can be any number of client devices. Each client device can be any type of electronic device, for example, a desktop computer, a notebook computer, a portable or mobile device, a mobile phone, a smart phone, a tablet computer, a television, a TV set-top box or an entertainment device, a wearable device (for example, display glasses or goggles, a wristwatch, an earphone, an armband, jewelry, etc.), a personal digital assistant (PDA), a media player, a gaming device, etc. Some client devices can also have a local database similar to database 106 or other storages. In some embodiments, network environment 100 may not have all the components shown, and / or may have other elements, including elements of other types other than or in addition to those elements described herein.

[0036] In various embodiments, end users U1, U2, U3, and U4 may communicate with server system 102 and / or with each other using respective client devices 120, 122, 124, and 126. In some examples, users U1, U2, U3, and U4 may interact with each other via applications running on respective client devices and / or server system 102, and / or via a network service, such as a social networking service or other type of network service implemented on server system 102. For example, respective client devices 120, 122, 124, and 126 may communicate data to and from one or more server systems (e.g., system 102).

[0037] In some embodiments, the server system 102 can provide appropriate data to the client device so that each client device can receive the communication content or shared content uploaded to the server system 102 and / or the network service. In some examples, users U1-U4 can interact via audio / video calls, audio, video or text chats, or other communication modes or applications. The network services implemented by the server system 102 may include systems that allow users to perform various communications, form links and associations, upload and publish shared content such as images, text, video, audio and other types of content, and / or perform other functions. For example, the client device can display received data, such as content posts that are sent or streamed to the client device and originate from different client devices (or directly from different client devices) or from server systems and / or network services via servers and / or network services. In some embodiments, client devices can communicate directly with each other, for example, using peer-to-peer communications between client devices as described above. In some embodiments, "users" may include one or more programs or virtual entities, as well as people combined with a system or network.

[0038] In some implementations, any of the client devices 120, 122, 124, and / or 126 may provide one or more applications. Figure 1 As shown in FIG. 1 , client device 120 may provide video application 152a and one or more other applications 154. Client devices 122-126 may also provide similar applications.

[0039] For example, the video application 152 can provide the user of the corresponding client device (e.g., user U1-U4) with the ability to participate in a video call with one or more other users. In a video call, with the user's permission, the client device can transmit the locally captured video to other devices participating in the video call. For example, such a video can include a live video captured using a camera (e.g., a front camera, a rear camera, and / or one or more other cameras) of the client device. In some embodiments, the camera can be separated from the client device and can be coupled to the client device, for example, via a network, via a hardware port of the client device, etc. The video application 152 can be a software application executed on the client device 120. In some embodiments, the video application 152 can provide a user interface. For example, the user interface can enable a user to place a video call to one or more other users, receive a video call from other users, leave a video message for other users, view video messages from other users, etc.

[0040] As reference Fig. 9As described, the video application 152a may be implemented using hardware and / or software of the client device 120. In various implementations, the video application 152a may be a standalone client application, for example, executed on any of the client devices 120-124, or may work in conjunction with a video application 152b provided on the server system 102. The video application 152a and the video application 152b may provide video calling (including video calling with two or more participants) functionality, audio or video messaging functionality, address book functionality, and the like.

[0041] In some implementations, the client device 120 may include one or more other applications 154. For example, the other applications 154 may be applications that provide various types of functionality, such as calendars, address books, email, web browsers, shopping, transportation (e.g., taxi, train, airline reservations, etc.), entertainment (e.g., music players, video players, game applications, etc.), social networking (e.g., messaging or chat, audio / video calls, sharing images / videos, etc.), image capture and editing (e.g., image or video capture, video editing, etc.), etc. In some implementations, the one or more other applications 154 may be stand-alone applications executed on the client device 120. In some implementations, the one or more other applications 154 may access a server system that provides data and / or functionality for the application 154.

[0042] The user interface on the client device 120, 122, 124 and / or 126 can enable the display of user content and other content, including images, videos, data and other content as well as communications, privacy settings, notifications and other data. Such user interfaces can be displayed using software on the client device, software on the server device, and / or a combination of client software and server software executed on the server device 104, such as application software or client software that communicates with the server system 102. The user interface can be displayed by a display device (e.g., a touch screen or other display screen, a projector, etc.) of the client device or server device. In some embodiments, an application running on the server system can communicate with the client device to receive user input at the client device and output data such as visual data, audio data, etc. at the client device.

[0043] Other implementations of the features described herein can use any type of system and / or service. For example, other networked services (e.g., connected to the Internet) can be used instead of or in addition to social networking services. Any type of electronic device can use the features described herein. Some embodiments can provide one or more features described herein on one or more client or server devices that are disconnected or intermittently connected to a computer network. In some examples, a client device including or connected to a display device can display content posts stored on a storage device local to the client device, for example, previously received over a communication network.

[0044] Figure 2 2 is a flow chart showing an example of a method 200 for obtaining a foreground video according to some embodiments. In some embodiments, for example, the method 200 may be implemented in Figure 1 In some implementations, some or all of method 200 may be implemented on a server system 102 as shown in FIG. Figure 1 , implemented on one or more client devices 120, 122, 124, or 126 shown in , implemented on one or more server devices, and / or implemented on both server devices and client devices. In the described examples, the implementation system includes one or more digital processors or processing circuits ("processors"), and one or more storage devices (e.g., database 106 or other storage). In some embodiments, different components of one or more servers and / or clients can perform different blocks or other portions of method 200. In some examples, the first device is described as a block that performs method 200. Some embodiments may have one or more blocks of method 200 performed by one or more other devices (e.g., other client devices or server devices), which one or more other devices may send results or data to the first device.

[0045] In some embodiments, method 200 or part of the method can be automatically initiated by the system. In some embodiments, the implementation system is the first device. For example, the method (or part thereof) can be performed periodically, or based on one or more specific events or conditions, for example, an application (for example, a video call application) initiated by a user, a camera of a user device being activated to capture a video, a video editing application being started, and / or one or more other conditions that can be specified in the settings read by the method occur. In some embodiments, this condition can be specified by the user in the user's custom preferences stored.

[0046] In one example, the first device can be a camera, a mobile phone, a smart phone, a tablet, a wearable device, or other client device that can capture video, and method 200 can be performed. In another example, a server device can perform method 200 for video, for example, a client device can capture a video frame processed by a server device. Some embodiments can initiate method 200 based on user input. A user (e.g., an operator or an end user) can, for example, select the initiation of method 200 from a displayed user interface (e.g., an application user interface or other user interface). In some embodiments, method 200 can be implemented by a client device. In some embodiments, method 200 can be implemented by a server device.

[0047] As mentioned herein, a video may include a sequence of image frames (also referred to as frames). Each image frame may include color data and depth data for a plurality of pixels. For example, the color data may include a color value for each pixel, and the depth data may include a depth value, for example, a distance from a camera that captures the video. For example, the video may be a high-definition video, wherein each image frame of the video has a size of 1920×1080, for a total of 1,958,400 pixels. The techniques described herein may be used for other video resolutions, for example, standard definition video, 4K video, 8K video, and the like. For example, a video may be captured by a mobile device camera (e.g., a smartphone camera, a tablet camera, a wearable camera, and the like). In another example, a video may be captured by a computer camera (e.g., a laptop camera, a desktop camera, and the like). In yet another example, a video may be captured by a video call appliance or device (such as a smart speaker, a smart home appliance, a dedicated video call device, and the like). In some embodiments, the video may also include audio data. Method 200 may start at block 202.

[0048] In block 202, a check is made to see whether user consent (e.g., user permission) has been obtained to use user data in embodiments of method 200. For example, the user data may include video captured by the user using a client device, video stored or accessed by the user (e.g., using a client device), video metadata, user data related to use of a video calling application, user preferences, etc. In some embodiments, one or more blocks of the methods described herein may use such user data.

[0049] If user consent has been obtained from the relevant user that the user data of the relevant user may be used in method 200, then in block 204, it is determined that it is reasonable to use the user data as described for the blocks of the method herein to implement those blocks, and the method continues to block 210. If user consent has not been obtained, then in block 206, it is determined that the blocks are not implemented using user data, and the method continues to block 210. In some embodiments, if user consent has not been obtained, the blocks are implemented using synthetic data and / or generic or publicly accessible and publicly usable data instead of user data. In some embodiments, if user consent has not been obtained, method 200 is not performed.

[0050] In block 210 of method 200, a plurality of video frames of a video are received. For example, the plurality of video frames may be captured by a client device. In some embodiments, the plurality of video frames may be captured during a time when the client device is part of a live video call, such as via a video calling application. In some embodiments, the plurality of video frames may be previously recorded, such as being part of a recorded video. The video may include a plurality of frames in a sequence. In some embodiments, each frame may include color data and depth data for a plurality of pixels.

[0051] In some embodiments, color data can be captured at 30 frames per second and depth data can be captured at 15 frames per second, such that only half of the captured frames include depth data when captured. For example, depth data can be captured for alternating frames. In embodiments where one or more captured frames do not have depth data, depth data of an adjacent frame (e.g., a previous frame) can be used as the depth data. For example, when a video frame is received, it can be determined whether one or more frames do not have depth data. If a frame is missing depth data, depth data of an adjacent frame (e.g., the immediately previous frame in a sequence of video frames) can be used as the depth data for the frame. Block 210 may be followed by block 212.

[0052] In block 212, the video is downsampled. In some embodiments, downsampling includes adjusting the video to a smaller size. For example, if the video is 1920×1080, the video can be adjusted to one-fourth of the size, i.e., 480×270. Resizing reduces the computational complexity of performing method 200, for example, because the total data to be processed is also reduced to one-fourth. For example, in some embodiments, after downsampling, the processing time for implementing method 200 can be about 90ms per frame. In some embodiments, for example, when the device implementing method 200 has high computing power, downsampling may not be performed. In some embodiments, the size of the downsampled video can be selected based on the size of the received video, the computing power of the device implementing method 200, etc. In various embodiments, the color data and depth data of the frame are directly downsampled. In various embodiments, the bit depth of the color data and the depth data does not change after downsampling. The selection of the downsampling ratio (e.g., the ratio of the number of pixels of the downsampled frame to the number of pixels of the original video frame) can be based on a tradeoff between processing time and segmentation quality. In some implementations, the downsampling ratio may be selected based on the available processing power of the device on which method 200 is implemented to achieve a suitable compromise. In some implementations, the downsampling ratio may be determined heuristically. Block 212 may be followed by block 213.

[0053] In block 213, a frame of the downsampled video is selected. Block 213 may be followed by block 214.

[0054] In block 214, an initial segmentation mask is generated that classifies each pixel of the selected frame or image as a foreground pixel or a background pixel. In some embodiments, the initial segmentation mask is based on a depth range. The depth value of each pixel can be compared to the depth range to determine whether the depth value is within the depth range. If the depth value is within the depth range, the pixel is classified as a foreground pixel. If the depth value is outside the depth range, the pixel is classified as a background pixel. The initial segmentation mask thus generated includes the value of each pixel of the selected frame, indicating whether the pixel is a foreground pixel or a background pixel. In some embodiments, the depth range can be 0.5 meters to 1.5 meters. For example, when method 200 is performed for a video call application or other application, where one or more users participating in the video call are close to a camera that captures video, such as in a conference room, at a table, etc., the range may be appropriate. In different applications, different depth ranges can be used, for example, selected based on the typical distance between the foreground object and the camera of those applications.

[0055] In some embodiments, the initial segmentation mask may include a mask value for each pixel of the frame. For example, if a pixel is determined to be a background pixel, a mask value of 0 may be assigned to the pixel, and if the pixel is determined to be a foreground pixel, a mask value of 255 may be assigned to the pixel.

[0056] In some embodiments, generating the initial segmentation mask may further include performing a morphological opening process. The morphological opening process may remove noise from the initial segmentation mask. In some embodiments, segmenting the image may further include performing a morphological closing process. The morphological closing process may fill one or more holes in the initial segmentation mask.

[0057] Noise and / or holes in the segmentation mask may be caused by various reasons. For example, when a video frame is captured by a client device using a camera with depth capabilities, the depth value of one or more pixels may be inaccurately determined. This inaccuracy may be caused, for example, due to the lighting conditions of the captured frame, due to sensor errors, due to features in the captured scene, etc. For example, if the camera does not capture the depth value of one or more pixels, a hole may be generated. For example, if the camera uses a reflection-based sensor to measure depth, if no reflected light is detected from the scene, one or more pixels may have an infinite depth value. Such pixels may cause holes in the segmentation mask. Box 214 may be followed by box 216.

[0058] In block 216, a head bounding box may be detected. The head bounding box may specify an area of ​​the frame that includes or may include a head. For example, when the video includes a single person, the head bounding box may specify pixels of the frame that correspond to the head of the single person. In some embodiments, for example, when the video includes multiple persons, the head bounding box may specify multiple areas of the frame, each area corresponding to a specific person. In some embodiments, the head bounding box may be detected based on color data, an initial segmentation mask, or both. Any suitable method may be used for head bounding box detection. Reference Figure 3 An example method of detecting a head bounding box is described. Block 216 may be followed by block 218 .

[0059] In block 218, a trimap is generated for the non-head portion of the frame (e.g., the portion of the frame outside the head bounding box). In some embodiments, the trimap may be generated based on the initial segmentation mask. The trimap may classify each pixel of the image as one of a known foreground, a known background, and an unknown.

[0060] In some embodiments, generating the initial segmentation mask may include calculating an L1 distance between a pixel position of each pixel of the frame that is a non-head portion (outside the head bounding box) and a mask boundary of the initial segmentation mask. The mask boundary may correspond to a position, for example, expressed in pixel coordinates, where at least one foreground pixel in the initial segmentation mask is adjacent to at least one background pixel.

[0061] In some embodiments, if the L1 distance satisfies the foreground distance threshold and the pixel is classified as a foreground pixel in the initial segmentation mask, then the pixel is classified as a known foreground in the trigraph. For example, in some embodiments, the foreground distance threshold may be set to 24. In these embodiments, pixels with an L1 distance greater than 24 from the mask boundary and a mask value of 255 (foreground pixels) are classified as known foreground. In different embodiments, different foreground distance thresholds may be used.

[0062] In addition, if the L1 distance satisfies the background distance threshold and the pixel is classified as a background pixel in the initial segmentation mask, the pixel is classified as known background in the trigraph. For example, in some embodiments, the background distance threshold can be set to 8. In these embodiments, pixels with an L1 distance greater than 8 from the mask boundary and a mask value of 0 (background pixels) are classified as known background. In different embodiments, different background distance thresholds can be used. In some embodiments, the threshold can be based on the quality of the initial segmentation mask and / or the downsampling ratio. In some embodiments, when the quality of the initial segmentation mask is low, a higher threshold can be selected, and when the quality of the initial segmentation mask is high, a lower threshold can be selected. In some embodiments, the threshold can be reduced in proportion to the downsampling ratio.

[0063] In addition, each pixel that is not classified as a known foreground and is not classified as a known background is classified as unknown in the triplicate map. The generated triplicate map can indicate the known foreground, known background, and unknown areas of the portion of the frame outside the head bounding box. Block 218 can be followed by block 220.

[0064] In block 220, a triplane is generated for the head portion of the frame (eg, the portion of the frame within the head bounding box). In some implementations, generating the triplane may include, for each pixel in the head bounding box, identifying whether the pixel is a known foreground, a known background, or an unknown.

[0065] In some embodiments, identifying a pixel as known foreground may include classifying the pixel as known foreground if the pixel is within an inner mask determined for the head bounding box. Further, in these embodiments, identifying a pixel as known background may include classifying the pixel as known background if the pixel is outside an outer mask determined for the head bounding box. Further, in these embodiments, identifying a pixel as unknown may include classifying pixels that are not classified as known foreground and not classified as known background as unknown. The generated tripartite map may indicate known foreground, known background, and unknown areas of the portion of the frame that is within the head bounding box. In different embodiments, any suitable technique may be used to obtain the inner mask and the outer mask. Reference Figure 4 Example methods of obtaining inner and outer masks are described.

[0066] After generating the tripartite map of the head portion of the frame, it can be merged with the tripartite map of the non-head portion of the image to obtain a tripartite map of the entire frame. In this way, the generated tripartite map classifies each pixel of the frame as known background (BGD), known foreground (FGD) or unknown. Block 220 can be followed by block 222.

[0067] For example, by utilizing inner and outer masks for triplanogram generation, separate generation of a triplanogram of a head portion may provide improved segmentation results due to identification and incorporation of head-specific features in triplanogram generation of the head portion.

[0068] In block 222, the generated tripartite map is refined and a weight map is calculated. For example, a weight map may be calculated for pixels classified as unknown in the tripartite map. The weight map may represent the level or degree to which each pixel classified as unknown is tilted toward the foreground or background in the frame. A background image is determined and maintained for the video. For example, in many cases, such as when a user is participating in a video call (or video game) from a conference room, desktop computer, or mobile device, the device camera may be static and the scene may be static, e.g., the background portion of the frame may not change from frame to frame.

[0069] The background image can be determined based on a binary mask of one or more frames of the video. The retained background image can be a color image with the same size as each frame of the video (or downsampled video). For example, the retained background image can include the color value of each pixel identified as the background in the binary mask of one or more previous frames of the video. Therefore, the retained background image includes information indicating the background color of the scene at each pixel position. In some embodiments, the moving average of each pixel can be determined based on the previous frames in the frame sequence (e.g., for 2 previous frames, for 5 previous frames, for 10 previous frames, etc.). In these embodiments, the Gaussian model can be used to estimate the possibility that the pixel in the current frame is the background based on the moving average from the previous frame.

[0070] In some embodiments, maintaining the background image may include updating the background image based on the tripartite map, for example, pixels classified as background (BGD) in the tripartite map of the frame. For example, the background image may be updated using the following formula:

[0071] Maintained background = 0.8 * previous background + 0.2 * new background Wherein, the previous background is the color value of the pixel in the background before the update, the new background is the color value of the corresponding pixel in the trigraph, and the maintained background is the updated background image. The coefficients of the previous background (0.8) and the new background (0.2) can be selected based on the application. In some embodiments, the coefficients can be selected based on a previous estimate of the stability of the camera. For example, for a fixed camera, a coefficient value of 0.8 can be selected for the previous background, and a coefficient value of 0.2 can be selected for the new background. In another example, for example, for a handheld camera or other camera that undergoes movement, for example, because the historical data may be of lower value due to the movement of the camera, a coefficient value of 0.5 can be selected for the previous background, and a coefficient value of 0.5 can be selected for the new background.

[0072] The maintained background image can be used to determine a weight map that assigns weights to each pixel in the trigraph. In some embodiments, determining the weight map can include calculating the weight of each pixel of the frame classified as unknown in the trigraph. In some embodiments, calculating the weight of the pixel can include calculating the Euclidean distance between the color of the pixel and the background color of the background image. Calculating the weight can further include determining the probability that the pixel is a background pixel based on the Euclidean distance. In some embodiments, the probability (p) can be calculated by using the following formula:

[0073] p = exp(-0.01*|pixel color - preserved background color| 2 )

[0074] In some embodiments, the pixel color and the background color maintained can be in a red-green-blue (RGB) color space. Calculating the weight based on probability further includes determining whether the probability meets a background probability threshold. For example, the background probability threshold can be set to 0.5 so that pixels with p>0.5 meet the background probability threshold. If the pixel meets the background probability threshold, a background weight (e.g., a negative value) is assigned to the pixel in the weight map. In some embodiments, the weight value can be an estimate based on camera stability, for example, a higher weight value can be used when the camera is stable, and a lower weight value can be used when the camera has movement during capturing video.

[0075] Additionally, skin color detection may be performed to identify one or more skin regions in a frame, excluding a facial region. For example, a facial region may be excluded by performing skin color detection on a frame except for a portion of the frame that is within a head bounding box. For example, the one or more skin regions may correspond to a hand, an arm, or other portion of a body depicted in a frame.

[0076] In some embodiments, the color value of a pixel in the color data of a frame may be an RGB value. In some embodiments, skin color detection may be performed on each pixel to determine whether it may be a skin color by using the following formula:

[0077] Skin = R>95AND G>40AND B>20AND (RG)>15AND R>B

[0078] Here, R, G, and B refer to the red, green, and blue channel values ​​of the pixel.

[0079] Based on the pixels identified as being likely to be skin color, one or more skin regions may be identified. For example, one or more skin regions may be identified as regions that include pixels within a threshold distance of skin color pixels. For example, the threshold distance may be 40.

[0080] In some embodiments, after identifying one or more skin regions, the tripartite map can be updated to set each pixel in the one or more skin regions to unknown. In addition, a zero weight can be assigned to each such pixel. In addition, the color of the pixel can be compared to the background color of the background image. If the pixel color and the background color meet a similarity threshold, a background weight (e.g., a negative value) is assigned to the pixel in the weight map. For example, the similarity threshold can be a probability threshold as described above (e.g., p>0.5).

[0081] If the pixel is a skin color pixel (as identified using skin color detection), then a foreground weight (e.g., a positive value) is assigned to the pixel in the weight map. Additionally, if the color of the pixel and the color of the background meet a dissimilarity threshold, then a foreground weight (e.g., a positive value) is assigned to the pixel in the weight map. For example, the dissimilarity threshold may be a probability threshold (e.g., p<0.0025). Other pixels in the skin region may retain a zero weight. Block 222 may be followed by block 224.

[0082] In block 224, an incoherence penalty weight may be calculated for pixels classified as unknown in the tripartite map. In some embodiments, the initial segmentation mask may be compared to a previous frame binary mask of an immediately previous frame in a sequence of frames to determine the proportion of pixels of the frame that are classified similarly to the pixels of the previous frame. The proportion may be defined as the similarity between frames. For example, when there is significant motion in the scene, the similarity may be low, and for most static scenes, the similarity may be high. A global coherence weight may be calculated based on the similarity. For example, in some embodiments, the global coherence weight may be calculated using the following formula:

[0083] w = A*2 / 1+exp(50*(1-similarity)))

[0084] Where w is the global coherence weight and A is a predefined constant.

[0085] This formula ensures that when the similarity is low, e.g., close to zero, the global coherence weight decreases exponentially. In this case, the global coherence weight does not affect the weight map. On the other hand, when the similarity is high, the global coherence weight may be higher. In this way, the global coherence weight described in this article is a function of frame similarity.

[0086] In addition, the weight of pixels classified as unknown in the tripartite map can be calculated based on the global coherence weight. In some embodiments, the weight of the pixel can be determined based on the global coherence weight and the distance between the pixel and the mask boundary of the binary mask of the previous frame. When the corresponding pixel in the binary mask is classified as a foreground pixel in the binary mask of the previous frame, the weight calculated for the pixel is positive. When the corresponding pixel in the binary mask is classified as a background pixel in the binary mask of the previous frame, the weight calculated for the pixel is negative. In some embodiments, the weight can be proportional to the distance. In some embodiments, the global coherence weight can be used as a cutoff value for the weight, for example, when the distance is equal to or greater than the cutoff distance value, the value of the weight can be set equal to the global coherence weight. The calculated weight can be stored in a weight map. In some embodiments, the cutoff distance value can be experimentally determined, for example, based on segmentation results obtained for a large number of videos. In some embodiments, a higher cutoff distance value can correspond to a different possibility of segmentation of adjacent frames, corresponding to weaker coherence between consecutive frames.

[0087] Calculating weights for pixels classified as unknown in the tripartite map in this manner and storing the weight map can ensure consistency between segmentations of consecutive frames, e.g., when the frames are similar, corresponding pixels in consecutive frames are more likely to have similar categories in the binary mask. This can reduce the visual effect of flickering that can occur when such pixels have different categories in the binary masks of consecutive frames. Block 224 can be followed by block 226.

[0088] In block 226, a binary mask of the frame is obtained by performing fine segmentation based on the color data, the tripartite map, and the weight map. In some embodiments, performing fine segmentation may include applying a graph cut technique to the frame. When applying the graph cut technique, the color data, the tripartite map, and the weight map may be provided as input.

[0089] In the graph cutting technique, color data is used to globally create background and foreground color models. For example, a Gaussian mixture model (GMM) can be used to construct such a color model. GMM utilizes the initial marking of foreground and background pixels obtained, for example, via user input. GMM generates a new pixel distribution that marks unknown pixels as possible background or possible foreground based on the similarity level of the color value (e.g., RGB value) of each unknown pixel and the pixel that has been marked as foreground or background in the initial marking.

[0090] In the graph cutting technique, a graph including nodes corresponding to each pixel of the image is generated. The graph further includes two additional nodes: a source node connected to each pixel marked as foreground and a sink node connected to each pixel marked as background. In addition, the graph cutting technique also includes calculating a weight for each pixel, which indicates the possibility that the pixel is background or foreground. The weight is assigned to the edge connecting the pixel to the source node / sink node. In the graph cutting technique, the weight between pixels is defined by edge information or pixel similarity (color similarity). If there is a large difference in the pixel color of two pixels, the edge connecting the two pixels is assigned a lower weight. By removing edges, for example, by minimizing a cost function, iterative cutting is performed to separate the foreground and background. For example, the cost function can be the sum of the weights of the edges that are cut.

[0091] In some embodiments, the graph cut technique is applied to pixels of the frame that are classified as unknown in the tripartite map. In these embodiments, pixels identified as known foreground and known background are excluded from the graph cut. In some embodiments, the global color model is disabled when the graph cut technique is applied. Disabling the global color model can save computing resources, for example, by eliminating the need to build a global color model for the foreground and background as described above, and instead using the categories from the tripartite map.

[0092] Applying the graph cut technique to a small portion of the frame (the unknown portion of the triplanar map) and excluding the known foreground and known background of the triplanar map can improve the performance of the segmentation because the pixels of the known foreground and known background are not added to the graph. For example, the size of the graph processed to obtain the binary mask can be smaller than the size of the graph when the color model-based graph cut technique as described above is used. In an example, excluding the known foreground and known background can result in the computational load of the graph cut being approximately 33% of the computational load when the known foreground and known background are included. Block 226 can be followed by block 228.

[0093] In box 228, a temporal low-pass filter may be applied to a binary mask, for example, a binary mask generated by a graph cutting technique. Applying a temporal low-pass filter may be performed as part of performing a fine segmentation. Even if the scene captured in the video is static, the depth value of the corresponding pixel may vary between consecutive frames. This may occur due to imperfect depth data being captured by the sensor. The temporal low-pass filter updates the binary mask based on the similarity between one or more previous frames and the current frame. For example, if the scene captured in multiple video frames is static, consecutive frames may include similar depth values ​​for corresponding pixels. If there is a change in the depth value of the corresponding pixel and the scene is static, such depth value may be wrong, and a temporal low-pass filter is used to update such depth value. If the similarity between one or more previous frames and the current frame is high, applying a temporal low-pass filter causes the segmentation of the current frame to be consistent with the segmentation of one or more previous frames. When the similarity is low, for example, when the scene is not static, the consistency produced by the temporal low-pass filter is weak. Box 228 may be followed by box 230.

[0094] In block 230, a Gaussian filter may be applied to the binary mask as part of performing fine segmentation. The Gaussian filter may smooth segmentation boundaries in the binary mask. Applying the Gaussian filter may provide alpha matting, for example, to ensure that the binary mask separates hairy or blurred foreground objects from the background. Block 230 may be followed by block 232.

[0095] In block 232, it is determined whether there are more frames of the video for which a binary mask is to be determined. If it is determined that there is another frame to be processed, block 232 may be followed by block 213 to select that frame, e.g., the next frame in the sequence of frames. If there are no remaining frames (the entire video has been processed), block 232 may be followed by block 234.

[0096] In block 234, the corresponding binary mask obtained for each frame of the plurality of frames is upsampled, for example, to the size of the original video and used to obtain a foreground video. For example, a foreground mask for each frame of the video may be determined using the corresponding upsampled binary mask, and the foreground mask may be used to identify pixels of the video to be included in the foreground video. Block 234 may be followed by block 236.

[0097] In block 236, the foreground video may be rendered. In some embodiments, rendering the foreground video may include generating a plurality of frames including the foreground segmented using a binary mask. In some embodiments, rendering may further include displaying a user interface including the foreground video. For example, the foreground video may be displayed in a video call application or other application. In some embodiments, the foreground video may be displayed without a background (e.g., a blank background). In some embodiments, a background different from the original background of the video may be provided together with the foreground video.

[0098] By subtracting the background and obtaining the foreground video, any suitable or user-preferred background may be provided for the video. For example, in a video calling application, a user may indicate a preference to replace the background with a particular scene, and such background may be displayed with the foreground video. For example, replacing the background in this manner may allow a participant in a video call to remove video clutter in a room where the participant has joined the video call by replacing a portion of the background.

[0099] In some embodiments, the method 200 can be implemented in a multi-threaded manner, for example, multiple threads implementing the method 200 or parts thereof can be executed simultaneously, for example, on a multi-core processor, a graphics processor, etc. In addition, the threads of the method 200 can be executed simultaneously with other threads (e.g., one or more threads for capturing video and / or one or more threads for displaying the segmented video). In some embodiments, the segmentation is performed in real time. Some embodiments can perform real-time segmentation at a rate of 30 frames per second.

[0100] Detecting the head bounding box can improve the quality of the segmentation of the frame. In many applications, such as video call applications, the head of a participant in a video call is the focus of attention of other participants. Therefore, accurately detecting the boundaries of the head is valuable for providing a high-quality foreground video. For example, a high-quality foreground video can include all (or nearly all) pixels containing the participant's head while excluding background pixels. In a high-quality foreground video, fine areas such as hair and neck are accurately segmented. Detecting the head bounding box can make the foreground video a high-quality video.

[0101] In some embodiments, it is possible to combine Figure 2. For example, box 218 and box 220 may be combined or performed in parallel. In another example, box 222 may be combined with box 224. In some embodiments, one or more boxes may not be performed. For example, in some embodiments, box 224 is not performed. In these embodiments, the incoherence penalty weight is not calculated. In another example, in some embodiments, box 228 may not be performed. In some embodiments, box 232 may not be performed, so that the binary mask obtained after applying the Gaussian filter is directly used in box 234 to obtain the foreground video.

[0102] In some embodiments, the blocks of method 200 may be performed in parallel or in parallel. Figure 2 . For example, in some embodiments, the received video can be divided into a plurality of video segments, each of which includes a subset of video frames. Then, method 200 can be used to process each video segment to obtain a foreground segment. In these embodiments, different video segments can be processed in parallel, and the obtained foreground segments can be combined to form a foreground video.

[0103] In some embodiments, the foreground video can be rendered in real time, e.g., so that there is little or no perceptible lag between the capture of the video and the rendering or display of the foreground video. In some embodiments, upsampling and foreground video rendering (blocks 234 and 236) for a portion of the video can be performed in parallel with blocks 213-232 for subsequent portions of the video.

[0104] Carrying out method 200 or its part in parallel to render foreground video can make it possible to display foreground video in real time without user-perceivable lag.In addition, parallel execution (e.g., using multithreading method) can advantageously use available hardware resources, e.g., multiple cores of multi-core processors, graphics processors, etc.

[0105] Method 200 may be performed by a client device (e.g., any one of client devices 120-126) and / or a server device (e.g., server device 104). For example, in some embodiments, a client device may capture video and perform method 200 to render a foreground video locally. For example, when the client device has suitable processing hardware (e.g., a dedicated graphics processing unit (GPU) or another image processing unit (e.g., ASIC, FPGA, etc.)), method 200 may be performed locally. In another example, in some embodiments, a client device may capture video and send the video to a server device, which performs method 200 to render a foreground video. For example, when a client device lacks the processing power to perform method 200, or in other cases, such as when the battery power available on the client device is below a threshold, method 200 may be performed by a server device. In some embodiments, method 200 may be performed by a client device other than the device that captured the video. For example, a sending device in a video call may capture a video frame and send the video frame to a receiving device. The receiving device may then perform method 200 to render a foreground video. This implementation may be advantageous when the sending device lacks the ability to perform method 200 in real time.

[0106] Figure 3 2 is a flow chart illustrating an example method 300 for detecting a head bounding box according to some implementations. For example, the method 300 may be used to detect a head bounding box of a video frame in block 216 .

[0107] The method 300 may begin at block 302. In block 302, a color image (e.g., a frame of a video) and a corresponding segmentation mask may be received. For example, the segmentation mask may be an initial segmentation mask determined based on depth data of the video frame (e.g., as determined in block 214 of the method 200). The initial segmentation mask may be binary, e.g., the initial segmentation mask may classify each pixel in the color image as a foreground pixel or a background pixel. Block 302 may be followed by block 304.

[0108] In block 304, the received image (frame) is converted to grayscale. Block 304 may be followed by block 306.

[0109] In block 306 , a histogram equalization is performed on the grayscale image. Block 306 may be followed by block 308 .

[0110] In block 308, Haar cascade face detection is performed to detect one or more faces in the image. For each detected face, a face region including facial pixels corresponding to the face is identified. Block 308 may be followed by block 310.

[0111] In block 310, the detected faces are verified using the initial segmentation mask. In some implementations, for each detected face, it is determined whether at least a threshold proportion of pixels of the face region are classified as foreground pixels in the initial segmentation mask. Block 310 may be followed by block 312.

[0112] In block 312, it is also determined whether the detected face has sufficient skin area. Determining whether the detected face has sufficient skin area may include determining whether at least a threshold percentage of pixels in the face area have skin color. For example, whether a pixel has skin color may be determined based on a color value of the pixel. In some embodiments, the color value of a pixel in the color data of the frame may be an RGB value. In some embodiments, skin color detection may be performed on each pixel to determine whether it is likely to have skin color by using a skin color criterion given by the following formula:

[0113] Skin = R>95AND G>40AND B>20AND (RG)>15AND R>B

[0114] Here, R, G, and B refer to the red, green, and blue channel values ​​of the pixel.

[0115] In some implementations, blocks 310 and 312 may be performed for each detected face in the image. Block 312 may be followed by block 314.

[0116] In block 314, a determination is made as to whether the image includes at least one valid face. For example, block 314 may be performed for each detected face. In some embodiments, a face is determined to be valid if it is verified that the facial region of the face includes at least a threshold proportion of pixels classified as foreground pixels and at least a threshold percentage of pixels of the facial region of the face have skin color. In some embodiments, the threshold proportion of pixels classified as foreground pixels may be 0.6 (60%), and the threshold percentage of pixels of the facial region of the face having skin color may be 0.2 (20%). If at least a valid face is detected in block 314, block 314 may be followed by block 316. If no face is detected, block 314 may be followed by block 320.

[0117] In block 316, the facial region of each valid face is enlarged to cover the head region. In some embodiments, the enlargement may be performed to enlarge the facial bounding box of the valid face by a specific percentage. Block 316 may be followed by block 330.

[0118] In block 330, a head bounding box may be obtained. For example, the head bounding box may include a face region and additionally include a hair region, a neck region, or a neckline region. For example, if the image includes multiple valid faces, the head bounding box may identify multiple different regions of the image.

[0119] In block 320, the initial segmentation mask is analyzed to detect the head in the image. In some embodiments, horizontal scan lines are first calculated based on the initial segmentation mask. The connections of the scan lines are analyzed, and based on the connections, the position and / or size are determined to detect the position of the head region. Block 320 may be followed by block 322.

[0120] In block 322, a determination is made as to whether the detected head is valid. For example, the determination may be based on verifying the skin color of the head, for example, similar to the verification of the skin color of the face in block 312. Block 322 may be followed by block 324.

[0121] In block 324, a determination is made as to whether a valid head is detected in the image. If a valid head is detected, block 324 may be followed by block 330. If no head is detected, block 324 may be followed by block 326.

[0122] In block 326, the head bounding box may be set to empty or null value so that no pixel of the image is within the head bounding box. An empty head bounding box may indicate that no head is detected in the image. When no head is detected, the triplicate map of the non-head portion is the triplicate map of the frame. In some embodiments, head detection may be turned off so that all regions of the image are treated similarly in the triplicate map.

[0123] Method 300 may provide a number of technical benefits. For example, using the Haar cascade face detection technique may detect multiple heads, for example, when multiple people appear in a frame. In addition, as described with reference to blocks 310 and 312, verification of the face may ensure that false positives (non-face regions that were mistakenly identified as face regions during the Haar cascade face detection of block 308) are eliminated. In addition, as described with reference to block 316, face region enlargement may ensure that regions such as hair, neck, neckline, etc. are identified as known foreground in the triplicate map, thereby achieving high quality segmentation.

[0124] Additionally, as described with reference to blocks 320 and 322, if a face is not detected by Haar face detection (or if a detected face is not verified), a head bounding box can be constructed by mask analysis and verification of the initial segmentation mask. Thus, this technique can compensate for false negatives, e.g., situations where Haar cascade face detection fails to detect a face. In some embodiments, mask analysis-based head detection can provide a high detection rate (e.g., higher than Haar cascade face detection) and has a lower computational cost.

[0125] Figure 4 is a flow chart illustrating an example method 400 of generating a tripartite map of a head region according to some implementations. The method 400 may begin at block 402 .

[0126] In block 402, a color image, a depth mask (e.g., an initial segmentation mask), and a head bounding box may be received. For example, the color image may correspond to a frame of a video and may include color data for pixels of the frame. Block 402 may be followed by block 404.

[0127] In box 404, it is detected whether there is a background near the head region (e.g., as identified by the head bounding box). For example, it can be detected whether the pixels of the image near the head region (e.g., near the hair region) are bright and uniform. Such a background may generally cause the depth sensor of the camera to be unable to detect the depth of the hair region or to detect an incorrect depth of the hair region. If a uniform bright background is detected near the head region, an extension of the hair region is performed. The extension of the hair region causes an increase in the expansion size to create an outer mask. Box 404 may be followed by box 406.

[0128] In block 406, an inner mask reduction of the neck and / or shoulder region is performed to obtain an inner mask. In the inner mask, the region around the neck (or shoulder) is erased. This erasure can compensate for erroneous or unreliable depth data near the neck region. Block 406 may be followed by block 408.

[0129] In block 408, inner mask dilation is performed. For example, pixels near the inner mask having skin color may be analyzed to determine whether the pixel has skin color. If the pixel has skin color, such pixel is added to the inner mask, which may avoid over-erosion, such as may occur when performing block 406. Block 408 may be followed by block 410.

[0130] In block 410, the collar region of the inner mask may be enlarged. This may also avoid over-etching, such as may occur when, for example, block 406 is performed on the collar region. Block 410 may be followed by block 412.

[0131] In block 412, a mask morphological dilation is performed to obtain an outer mask. Block 412 may be followed by block 414. Method 400 may improve the triplanar map in head regions, such as in the neck and neckline regions, where depth data captured by the camera is typically not as reliable as in other portions. Using skin color to identify foreground regions may improve the triplanar map.

[0132] Figure 5 An example video frame (502) with color portions and a corresponding mask (504) are shown. Figure 5 As can be seen in FIG. 5 , the foreground of the video frame ( 502 ) includes a person with his arms raised. The background of the video frame includes an office environment with a whiteboard and another person at a workstation, who has their back to the person in the foreground.

[0133] The mask (504) may be an initial segmentation mask. For example, the mask (504) may be a segmentation mask that classifies each pixel of the frame (502) as a foreground pixel or a background pixel. Figure 5 In , the foreground pixels of the mask are white, while the background pixels are black. As can be seen, the segmentation mask does not accurately classify the pixels of the frame.

[0134] For example, multiple white pixels are seen in the upper left quadrant of the image. These pixels correspond to the background portion of the image, but are classified as foreground in the mask. In another example, portions of a white board near the glass window and the left arm of the person in the foreground are incorrectly classified as foreground. Also, although individual fingers of the person can be seen in the color image, the mask does not accurately depict the finger regions. Therefore, using Figure 5 The foreground frames that can be obtained with masks are of low quality due to these segmentation errors.

[0135] Figure 6 Two example images with separated foreground and background are shown without using a triangulation as described above. In the first image (602), it can be seen that while the rest of the image correctly separates the background (in light gray) from the foreground including the person, the area near the person's neck (604) is incorrectly identified as the foreground. In the second image (612), it can be seen that while the rest of the image correctly separates the background (in light gray) from the foreground including the person, the area near the person's raised hand (614) is incorrectly identified as the foreground. Therefore, in each image (602, 612), a portion of the foreground (e.g., the person's neck region (604) and the fingers and hand region (614)) is not correctly segmented because a portion of the background is incorrectly identified as the foreground.

[0136] Figure 7 An example image (702) is shown with identified trigraph parts. Figure 7, pixels of the image having modified colors that are different from the pixel colors of the original image are classified as unknown in the tripartite map. Three different colors are used to illustrate the parts of the tripartite map having different weights. The red part of image 702 has a foreground weight in the weight map. The blue part of image 702 has a background weight in the weight map. The green part of image 702 has a neutral (e.g., zero) weight in the weight map and is classified as unknown.

[0137] like Figure 7 As can be seen in the figure, the red part is closer to the person's body than the other parts of the triplicate map, for example, red pixels can be seen near the right arm and the inner half of the right area compared to the green and blue parts. In addition, the green part that is closer to the person's body can be seen, for example, on the outer part of the left arm of the person in the image. The blue part classified as the background in the triplicate map is seen farther away from the person's body.

[0138] When provided as input to the graph cut algorithm, the triplanar and weight map include weights that penalize classification of foreground areas (external parts of the arms or other parts of the person's body) as background and enable the background to be correctly removed via the background weights. Using a specially generated triplanar that includes head-specific optimizations, for example, by using a head bounding box to generate a head portion triplanar, and skin region detection (e.g., hand regions), the output of graph cut is enabled to provide an improved segmentation of the image. For example, a custom implementation of graph cut is utilized that enables graph cut to be run only on pixels that are classified as unknown. Because graph cut is only applied to a small portion of the image (e.g., pixels classified as unknown in the triplanar ( Figure 7 The modified color part)) and other pixels are not added to the graph, so the computational cost of graph cutting is reduced. In some embodiments, the processing time of graph cutting using the trigraph can be less than about one-third of the processing time of the entire image.

[0139] Figure 8 Four frames (802, 812, 822, 832) of an input video and corresponding output frames (804, 814, 824, 834) of an output video including a foreground generated by segmenting the input video are shown in accordance with some embodiments. As can be seen, in each output frame, the foreground is separated from the background of the conference room, which has been replaced by a mountain view. In particular, the segmentation is accurate in the presence of motion. For example, the person in the foreground moves between frames 802 and 812; raises a hand with fingers spread apart in frame 822; and rotates the hand in frame 832. In each case, the corresponding output frame correctly segments the foreground because no portion (or only a minimal portion) of the original background is seen in the corresponding output frame.

[0140] Fig. 9 is a block diagram of an example device 900 that may be used to implement one or more features described herein. In one example, the device 900 may be used to implement a client device, such as, Figure 1 . Alternatively, device 900 may implement a server device, such as server system 102 or server device 104. In some embodiments, device 900 may be used to implement a client device, a server device, or both a client device and a server device. Device 900 may be any suitable computer system, server, or other electronic or hardware device as described above.

[0141] One or more of the methods described herein may be run in a standalone program that can be executed on any type of computing device, in a program that runs on a web browser, in a mobile application ("app") that runs on a mobile computing device (e.g., a cell phone, a smart phone, a tablet, a wearable device (a watch, an armband, jewelry, headwear, virtual reality goggles or glasses, augmented reality goggles or glasses, a head-mounted display, etc.), a laptop, etc.). In one example, a client / server architecture may be used, for example, a mobile computing device (as a client device) sends user input data to a server device, and receives final output data from the server for output (e.g., for display). In another example, all computations may be performed within a mobile app (and / or other apps) on a mobile computing device. In another example, computations may be split between a mobile computing device and one or more server devices.

[0142] In some embodiments, the device 900 includes a processor 902, a memory 904, an input / output (I / O) interface 906, and a camera 914. The processor 902 can be one or more processors and / or processing circuits to execute program code and control the basic operations of the device 900. A "processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. The processor may include a general-purpose central processing unit (CPU) having one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), a plurality of processing units (e.g., in a multi-processor configuration), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), a dedicated circuit for implementing a function, a dedicated processor for implementing a neural network model-based processing, a neural circuit, a system of processors optimized for matrix calculations (e.g., matrix multiplication), or other systems.

[0143] In some embodiments, the processor 902 may include a CPU and a GPU (or other parallel processor). In an embodiment, the GPU or parallel processor may include multiple processing cores that can perform calculations in parallel, for example, 100 cores, 1000 cores, etc. In addition, the GPU or parallel processor may include a GPU memory separate from the main memory 904. The GPU memory can be accessed by each GPU core. An interface may be provided to enable data to be transferred between the main memory 904 and the GPU memory.

[0144] In some embodiments, a GPU can be used to implement method 200, 300 or 400 or portions thereof. In particular, a GPU can be used to render video frames based on segmentation of background and foreground, for example, to render foreground video after subtracting the background. The GPU can also replace the background with a different background. In some embodiments, the color data and depth data can be stored in a GPU memory (also referred to as a GPU buffer). In these embodiments, the color data and depth data can be processed by the GPU, which can be faster than using a CPU to process the data.

[0145] In some embodiments, the processor 902 may include one or more coprocessors that implement neural network processing. In some embodiments, the processor 902 may be a processor that processes data to produce a probabilistic output, for example, the output produced by the processor 902 may be imprecise or may be accurate within a range from the expected output. Processing does not need to be limited to a specific geographic location or have time constraints. For example, the processor can perform its functions in "real time", "offline", in "batch mode", etc. Parts of the processing can be performed at different times and in different locations, by different (or the same) processing systems. The computer can be any processor that communicates with a memory.

[0146] The memory 904 is typically provided in the device 900 for access by the processor 902 and may be any suitable processor-readable storage medium, such as a random access memory (RAM), a read-only memory (ROM), an electrically erasable read-only memory (EEPROM), flash memory, etc., suitable for storing instructions executed by the processor and provided separately from and / or integrated with the processor 902. The memory 904 may store software operated by the processor 902 on the server device 900, including an operating system 908, a video call application 910, and application data 912. One or more other applications may also be stored in the memory 904. For example, the other applications may include applications such as a data display engine, a web hosting engine, an image display engine, a notification engine, a social networking engine, an image / video editing application, a media sharing application, and the like. In some embodiments, the video call application 910 and / or other applications may each include instructions that enable the processor 902 to perform the functions described herein, such as, for example, Figure 2 , 3 or some or all of the methods of 4. One or more of the methods disclosed herein can operate in multiple environments and platforms, for example, as a stand-alone computer program that can run on any type of computing device, as a web application with a web page, as a mobile application ("app") running on a mobile computing device, etc.

[0147] The application data 912 may include a video, for example, a sequence of video frames. In particular, the application data 912 may include color data and depth data for each of a plurality of video frames of the video.

[0148] Any software in memory 904 may alternatively be stored in any other suitable storage location or computer-readable medium. In addition, memory 904 (and / or other connected storage devices) may store one or more messages, one or more classification standards, electronic encyclopedias, dictionaries, thesauruses, knowledge bases, message data, grammars, user preferences, and / or other instructions and data used in the features described herein. Memory 904 and any other type of storage (disk, optical disk, tape, or other tangible medium) may be considered "storage" or "storage devices."

[0149] The I / O interface 906 can provide functionality to enable the device 900 to be interfaced with other systems and devices. The interface device can be included as a part of the device 900, or can be separated from the device 900 and communicate with the device 900. For example, a network communication device, a storage device (e.g., a memory and / or a database 106), and an input / output device can communicate via the I / O interface 906. In some embodiments, the I / O interface can be connected to an interface device such as an input device (keyboard, pointing device, touch screen, microphone, camera, scanner, sensor, etc.) and / or an output device (display device, speaker device, printer, motor, etc.).

[0150] Some examples of engagement devices that can be connected to the I / O interface 906 can include one or more display devices 930 that can be used to display content (e.g., images, videos, and / or user interfaces of output applications as described herein). The display device 930 can be connected to the device 900 via a local connection (e.g., a display bus) and / or via a network connection, and can be any suitable display device. The display device 930 can include any suitable display device, such as an LCD, LED (including OLED) or plasma display screen, CRT, television, monitor, touch screen, 3D display screen, or other visual display device. For example, the display device 930 can be a flat display screen provided on a mobile device, multiple display screens provided in goggles or a head-mounted device, or a monitor screen of a computer device.

[0151] The I / O interface 906 can be coupled to other input and output devices. Some examples include a camera 932 that can capture images and / or video. In particular, the camera 932 can capture color data and depth data for each video frame of the video. Some embodiments may provide a microphone for capturing sound (e.g., as part of a captured image, voice commands, etc.), an audio speaker device for outputting sound, or other input and output devices.

[0152] For ease of explanation, Fig. 9A frame for each of processor 902, memory 904, I / O interface 906, software block 908 and 910 and application data 912 is shown. These frames can represent one or more processors or processing circuits, operating systems, memories, I / O interfaces, applications and / or software modules. In other embodiments, equipment 900 may not have all components shown, and / or may have other elements, including elements that are not or other types except those elements shown in this article. Although some components are described as carrying out the frame and operation described in some embodiments as herein, any suitable component or combination of components of environment 100, equipment 900, similar systems, or any suitable one or more processors associated with such systems may carry out the frame and operation described.

[0153] The method described herein can be implemented by computer program instructions or codes, which can be executed on a computer. For example, the code can be implemented by one or more digital processors (e.g., microprocessors or other processing circuits), and can be stored on a computer program product including a non-temporary computer-readable medium (e.g., storage medium), such as a magnetic, optical, electromagnetic or semiconductor storage medium, including semiconductor or solid-state memory, tape, removable computer disk, random access memory (RAM), read-only memory (ROM), flash memory, rigid disk, optical disk, solid-state memory drive, etc. The program instruction can also be contained in an electronic signal and provided as an electronic signal, for example in the form of software as a service (SaaS) transmitted from a server (e.g., a distributed system and / or a cloud computing system). Alternatively, one or more methods can be implemented in hardware (logic gates, etc.) or in a combination of hardware and software. Example hardware can be a programmable processor (e.g., a field programmable gate array (FPGA), a complex programmable logic device), a general-purpose processor, a graphics processor, an application-specific integrated circuit (ASIC), etc. One or more of the methods may be performed as part or component of an application running on the system, or as an application or software running in conjunction with other applications and an operating system.

[0154] Although the description has been described with respect to specific embodiments thereof, these specific embodiments are merely illustrative and not restrictive. The concepts shown in the examples may be applied to other examples and embodiments.

[0155] In the case where certain embodiments discussed herein may collect or use personal information about a user (e.g., user data, information about a user's social network, a user's location and time at that location, a user's biometric information, a user's activities, and demographic information), one or more opportunities are provided to the user to control whether to collect information, whether to store personal information, whether to use personal information, and how to collect, store, and use information about the user. That is, in particular, when explicit authorization is received from the relevant user to collect, store, and / or use user personal information, the system and method discussed herein collect, store, and / or use user personal information. For example, the user is provided with control over whether a program or feature collects user information about the particular user or other users associated with the program or feature. One or more options are presented to each user whose personal information is to be collected to allow control over the collection of information associated with the user, to provide permission or authorization regarding whether to collect information and regarding which parts of the information to be collected. For example, one or more such control options may be provided to the user via a communication network. In addition, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information is removed. As an example, the identity of the user may be processed so that no personally identifiable information may be determined. As another example, the geographic location of a user device may be generalized to a larger area so that the user's specific location may not be determined.

[0156] Note that, as known to those skilled in the art, the function blocks, operations, features, methods, devices and systems described in this disclosure can be integrated or divided into different combinations of systems, devices and function blocks. Any suitable programming language and programming techniques can be used to implement the routine of a specific embodiment. Different programming techniques can be adopted, such as being process-oriented or object-oriented. The routine can be performed on a single processing device or multiple processors. Although steps, operations or calculations can be presented in a specific order, in different specific embodiments, the order can be changed. In some embodiments, multiple steps or operations shown in order in this specification can be performed simultaneously.

Claims

1. A computer-implemented method for obtaining a foreground video, It is characterized in that The method comprises: Receiving a plurality of frames of a video, wherein each frame includes depth data and color data for a plurality of pixels; downsampling each of the plurality of frames of the video; After the described downsampling, for each frame: Based on the depth data, generating an initial segmentation mask that classifies each pixel of the frame as a foreground pixel or a background pixel; detecting a head bounding box based on one or more of the color data or the initial segmentation mask; Determine a tripartite map that classifies each pixel of the frame as one of known background, known foreground, or unknown, wherein determining the tripartite map comprises: generating a first tripartite map for a non-head portion of the frame, wherein the non-head portion excludes pixels within the head bounding box, generating a second tripartite map for a head portion of the frame, wherein the head portion excludes pixels outside the head bounding box, and merging the first third image and the second third image; For each pixel of the frame classified as unknown in the trigraph, calculating a weight for the pixel and storing the weight in a weight map; and performing fine segmentation based on the color data, the triplicate map and the weight map to obtain a binary mask of the frame; and The plurality of frames are upsampled based on the binary mask of each corresponding frame to obtain the foreground video.

2. The computer-implemented method of claim 1, It is characterized in that Generating the initial segmentation mask includes setting a pixel as a foreground pixel if a depth value associated with the pixel is within a depth range, and setting the pixel as a background pixel if the depth value associated with the pixel is outside the depth range.

3. The computer-implemented method of claim 2, It is characterized in that Generating the initial segmentation mask further includes performing one or more of a morphological opening process or a morphological closing process.

4. The computer-implemented method of claim 1, It is characterized in that Detecting the head bounding box includes: converting the frame to grayscale; After said conversion, performing histogram equalization; and After the histogram equalization, one or more faces in the frame are detected by Haar cascade face detection, wherein each of the one or more faces is associated with a face region including facial pixels of the face.

5. The computer-implemented method of claim 4, It is characterized in that Further comprising determining whether each of the one or more faces is valid, wherein the face is determined to be valid if it is verified that a threshold proportion of pixels of the facial area of ​​the face are classified as foreground pixels in the initial segmentation mask and at least a threshold percentage of pixels of the facial area of ​​the face meet a skin color criterion.

6. The computer-implemented method of claim 5, It is characterized in that Further comprising, for each face determined to be valid, expanding the face region of each face to obtain a head region corresponding to the face, and wherein the head bounding box includes the head region of each face determined to be valid.

7. The computer-implemented method of claim 5, It is characterized in that Further including, if no face is determined to be valid: analyzing the initial segmentation mask to detect a head; Determining whether the head is valid based on head skin color verification; as well as If the header is valid, the bounding box associated with the header is selected as the header bounding box.

8. The computer-implemented method of claim 1, It is characterized in that Generating the initial segmentation mask comprises assigning a mask value to each pixel, wherein each foreground pixel is assigned the mask value of 255 and each background pixel is assigned the mask value of 0, and wherein determining the tri-map comprises, for each pixel of the frame that is not within the head bounding box: calculating an L1 distance between a pixel position of the pixel and a mask boundary of the initial segmentation mask, wherein the mask boundary includes a position where at least one foreground pixel in the initial segmentation mask is adjacent to at least one background pixel; If the L1 distance satisfies the foreground distance threshold, and the pixel is classified as a foreground pixel, classifying the pixel as a known foreground; If the L1 distance satisfies the background distance threshold, and the pixel is classified as a background pixel, classifying the pixel as known background; and If the pixel is not classified as known foreground and is not classified as known background, the pixel is classified as unknown.

9. The computer-implemented method of claim 8, It is characterized in that Determining the tripartite map further includes, for each pixel in the head bounding box, identifying whether the pixel is a known foreground, a known background, or an unknown pixel, wherein the identifying includes: If the pixel is within the inner mask determined for the head bounding box, classifying the pixel as known foreground; If the pixel is outside the outer mask determined for the head bounding box, classifying the pixel as known background; and If the pixel is not classified as known foreground and known background, the pixel is classified as unknown.

10. The computer-implemented method of claim 9, It is characterized in that Further comprising, prior to said identifying: Detecting whether there is a uniform bright background near the hair region of the head in the head bounding box; and If a uniform bright background is detected, a hair region extension is performed on the head based on the head bounding box, the color data and the initial segmentation mask, wherein after the hair region extension is performed, the expansion size of the outer mask is increased.

11. The computer-implemented method of claim 1 , It is characterized in that Further including: maintaining a background image of the video, wherein the background image is a color image having the same size as each frame of the video; and Before performing the fine segmentation, updating the background image based on the tripartite map, and Wherein, calculating the weight of the pixel comprises: Calculating the Euclidean distance between the color of the pixel and the background color of the background image; determining a probability that the pixel is a background pixel based on the Euclidean distance; and If the probability satisfies a background probability threshold, a background weight is assigned to the pixel in the weight map.

12. The computer-implemented method of claim 11, It is characterized in that Further including: identifying one or more skin regions in the frame based on skin color detection, wherein the one or more skin regions exclude a facial region; and For each pixel of the frame within the one or more skin regions: classifying the pixel as unknown and assigning a zero weight to the pixel in the weight map; If the color of the pixel and the background color of the background image satisfy a similarity threshold, assigning a background weight to the pixel in the weight map; If the color of the pixel is skin color, assigning a foreground weight to the pixel in the weight map; and If the color of the pixel and the background color of the background image satisfy a dissimilarity threshold, then a foreground weight is assigned to the pixel in the weight map.

13. The computer-implemented method of claim 1, It is characterized in that The plurality of frames are in a sequence, the method further comprising, for each frame: comparing the initial segmentation mask to a previous frame binary mask of an immediately previous frame in the sequence to determine a proportion of the pixels of the frame that are classified similarly to pixels of the previous frame; Based on the ratio, calculating a global coherence weight; as well as Wherein, calculating the weight of the pixel and storing the weight in the weight map comprises determining the weight based on the global coherence weight and a distance between the pixel and a mask boundary of the previous frame binary mask.

14. The computer-implemented method of claim 13, It is characterized in that The weight of the pixel is positive if the corresponding pixel is classified as a foreground pixel in the previous frame binary mask, and is negative if the corresponding pixel is not classified as a foreground pixel in the previous frame binary mask.

15. The computer-implemented method of claim 1, It is characterized in that Performing fine segmentation includes applying a graph cut technique to the frame, wherein the graph cut technique is applied to pixels classified as unknown.

16. The computer-implemented method of claim 1, It is characterized in that Further comprising, after performing the fine segmentation, applying a temporal low pass filter to the binary mask, wherein the temporal low pass filter updates the binary mask based on similarities between one or more previous frames and the frame.

17. A non-transitory computer readable medium, It is characterized in that Instructions are stored thereon. When the instructions are executed by one or more hardware processors, the instructions cause the one or more hardware processors to perform operations, including: Receiving a plurality of frames of a video, wherein each frame includes depth data and color data for a plurality of pixels; downsampling each of the plurality of frames of the video; After the described downsampling, for each frame: Based on the depth data, generating an initial segmentation mask that classifies each pixel of the frame as a foreground pixel or a background pixel; detecting a head bounding box based on one or more of the color data or the initial segmentation mask; Determine a tripartite map that classifies each pixel of the frame as one of known background, known foreground, or unknown, wherein determining the tripartite map comprises: generating a first tripartite map for a non-head portion of the frame, wherein the non-head portion excludes pixels within the head bounding box, generating a second tripartite map for a head portion of the frame, wherein the head portion excludes pixels outside the head bounding box, and merging the first third image and the second third image; For each pixel of the frame classified as unknown in the trigraph, calculating a weight for the pixel and storing the weight in a weight map; and performing fine segmentation based on the color data, the triplicate map and the weight map to obtain a binary mask of the frame; and The plurality of frames are upsampled based on the binary mask of each corresponding frame to obtain a foreground video.

18. The non-transitory computer readable medium of claim 17, It is characterized in that Instructions are further stored thereon. When the instructions are executed by the one or more hardware processors, the instructions cause the one or more hardware processors to perform operations, the operations including: maintaining a background image of the video, wherein the background image is a color image having the same size as each frame of the video; and Before performing the fine segmentation, updating the background image based on the tripartite map, and Wherein, calculating the weight of the pixel comprises: Calculating the Euclidean distance between the color of the pixel and the background color of the background image; determining a probability that the pixel is a background pixel based on the Euclidean distance; and If the probability satisfies a background probability threshold, a background weight is assigned to the pixel in the weight map.

19. A system for obtaining a foreground video, It is characterized in that The system comprises: one or more hardware processors; and A memory, the memory being coupled to the one or more hardware processors and having instructions thereon, and when the instructions are executed by the one or more hardware processors, operations are performed, the operations comprising: Receiving a plurality of frames of a video, wherein each frame includes depth data and color data for a plurality of pixels; downsampling each of the plurality of frames of the video; After the described downsampling, for each frame: Based on the depth data, generating an initial segmentation mask that classifies each pixel of the frame as a foreground pixel or a background pixel; detecting a head bounding box based on one or more of the color data or the initial segmentation mask; Determine a tripartite map that classifies each pixel of the frame as one of known background, known foreground, or unknown, wherein determining the tripartite map comprises: generating a first tripartite map for a non-head portion of the frame, wherein the non-head portion excludes pixels within the head bounding box, generating a second tripartite map for a head portion of the frame, wherein the head portion excludes pixels outside the head bounding box, and merging the first third image and the second third image; For each pixel of the frame classified as unknown in the trigraph, calculating a weight for the pixel and storing the weight in a weight map; and performing fine segmentation based on the color data, the triplicate map and the weight map to obtain a binary mask of the frame; and The plurality of frames are upsampled based on the binary mask of each corresponding frame to obtain the foreground video.

Citation Information

Patent Citations

  • Image segmentation using color & depth information

    US20160171706A1

  • Temporal saliency map

    US20170200279A1

  • Semi-automatic image segmentation

    US9443316B1