Removal of distortions from real-time video using masked frames
Through the video processing manager, the foreground and background in the video clip are generated and applied, and the foreground and background in the video clip are solved, and the foreground blur problem caused by motion blur in the prior art is achieved, thereby achieving higher quality video output.
Patent Information
- Application Number
- CN202380081665.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-28
- Filing Date
- 2023-09-26
- Publication Date
- 2025-07-01
AI Technical Summary
When the prior art applies motion blur, it is easy to cause blurred foreground in the video clip, and the foreground and background cannot be effectively separated, resulting in a decline in video quality.
By using the video processing manager to receive the body mask, motion vector and predicted mask of the current frame, generate the final mask and apply it to the current frame to segment the foreground and background, edit the masked frame to remove distortion, and generate the output frame.
Effectively separate the foreground and background in video clips, avoid foreground blur, improve video quality, reduce distortion, and provide clearer video output.
Smart Images

Figure CN120239872A_ABST
Abstract
Description
Cross - reference to related disclosures
[0001] This application claims the priority of U.S. Provisional Patent Application No. 63 / 377,484, filed on September 28, 2022, the entire disclosure of which is incorporated herein by reference. Background Art
[0002] Many video applications apply modifications to video segments in a global manner. That is, for a video segment that includes a foreground and a background, the modification is applied equally to both the foreground and the background. For example, motion blur can be globally applied to a video segment to hide jitter artifacts caused by three - two pulldown or a fast - pan shot. However, for a video segment that includes a prominent foreground, globally applying motion blur results in the prominent foreground being blurred, which may be undesirable. Summary of the Invention
[0003] This document describes systems and techniques for using masked frames to remove distortions from live video. In aspects, an image capture device having a video processing manager is configured to receive a video segment that includes a sequence of frames. The frame sequence includes at least a current frame that has a foreground and a background. The video processing manager receives a subject mask, a motion vector, and a predicted mask for the current frame. The video processing manager generates a final mask for the current frame based on the subject mask, the motion vector, and the predicted mask. The video processing manager then applies the final mask to the current frame to segment the foreground from the background and provide a masked frame. The video processing manager edits the masked frame to remove distortions to generate an output frame and outputs the output frame. By repeating the described method for each frame in the frame sequence, the video processing manager provides an improved video segment with less distortion.
[0004] Details of one or more aspects of using masked frames to remove distortions from live video will be set forth in the drawings and the following description. Other features and advantages will be apparent from the drawings, the detailed description, and the claims. The summary of the invention is provided to introduce the subject matter further described in the detailed description and the drawings. Accordingly, the summary of the invention neither describes essential features nor limits the scope of the claimed subject matter. Brief Description of the Drawings
[0005] This specification describes systems and techniques for using masked frames to remove distortions from live video with reference to the following drawings, in which like numerals are used throughout the drawings to refer to like features and components: Figure 1FIG. 0 illustrates an example environment in which aspects of an image capture device can implement using masked frames to remove distortions from live video; Figure 2 FIG. 1 more particularly illustrates an example implementation of an image capture device from Figure 1 ; Figure 3 FIG. 2 illustrates an example method for extracting a foreground from a video clip according to one or more aspects; and Figure 4 FIG. 3 illustrates an example method for segmenting a current frame of a frame sequence to produce a segmentation result. DETAILED DESCRIPTION OVERVIEW
[0006] Since the advent of video cameras, engineers have been working to improve the functionality of video cameras so that the videos captured by video cameras look realistic. To this end, engineers have worked to improve resolution and frame rate capabilities. Different from a display, reality is not limited to a finite resolution consisting of a finite number of individual pixels or light sources. Instead, reality can be considered to have infinite resolution. Therefore, a camera capable of capturing a video of a scene at a high resolution is very important for a realistic representation of the scene. In addition, different from a display, reality is not limited to a finite frame rate consisting of a finite number of frames or static images that are continuously and rapidly displayed. Instead, reality can be considered to have an infinite frame rate. Therefore, a camera capable of capturing a video of a scene at an extremely high frame rate is very important for a realistic representation of the scene. Given infinite resources, space, and time, designing a camera with such capabilities is very simple. However, in the absence of infinite resources, space, and time, engineers have developed alternative solutions.
[0007] As an example, a parent attends a child's track and field meet, in which the child plans to participate in the 100-meter (100 m) dash. They line up at the starting line and shake their legs back and forth to get ready. Before the starting gun goes off, the parent turns on the camera app on their smartphone and selects a video mode that is set to capture 1080p recordings at 30 frames per second (fps). When the starting gun goes off, the child sprints down the track like a rocket. At the same time, the parent starts using the smartphone to capture a video clip. As the child runs past the parent towards the finish line, the parent quickly pans the camera to keep their child in focus and centered in the frame. As they approach the finish line, they lunge forward with their shoulders and head and cross the finish line by a hair's breadth.
[0008] Once the child recovers from the 100-meter sprint race, the parent and the child will view video clips of the child's performance. In many conventional methods, the video will have defects, such as the foreground of the child being included and the background including the stands and other athletes becoming blurred. This may be because the real-time video processing pipeline of the parent's smartphone globally applies motion blur to the foreground and background of the video clip to hide the jitter caused by the parent quickly panning the smartphone. In the absence of motion blur, the background of the video clip will show jitter or stuttering artifacts due to the quick panning motion of the frame. However, although the use of motion blur reduces the apparent jitter, the foreground of the child including the parent also becomes blurred. This video quality is an undesirable result for both the parent and the child.
[0009] This document describes systems and techniques for using masked frames to remove distortions from real-time video. The disclosed systems and techniques can address the problem of foreground blurring in video clips due to the global application of motion blur. The systems and techniques extract the foreground from the video clip, thereby separating the foreground from the background and enabling the editing of the background separated from the foreground. The following discussion describes the operating environment, the techniques that can be employed in the operating environment, and example methods. Although systems and techniques for using masked frames to remove distortions from real-time video are described, the subject matter in the appended claims is not limited to the specific features or methods described. Instead, the specific features and methods are disclosed as example implementations and are referenced only by way of example to the operating environment. Example Environment
[0010] Figure 1 An example environment 100 is shown in which aspects of an image capture device 102 for using masked frames to remove distortions from real-time video can be implemented. The image capture device 102 includes a display 104, a camera 106, one or more processors 108, and a video processing manager 110 configured to extract the foreground of a video clip in real time. In one example, a user 112 wishes to record a video of an athlete 114 sprinting past a tree 116. The user 112 takes out the image capture device 102, opens a camera application (not shown) installed on the image capture device 102, selects the video mode (not shown), and taps the shutter button (not shown) to start recording the moment when the athlete 114 sprints past the tree 116.
[0011] In response to user 112 tapping the shutter button, video processing manager 110 captures a video clip including a sequence of frames that includes a previous frame and a current frame received immediately after the first frame. Video processing manager 110 receives a body mask of the current frame. For example, a machine learning (ML) model on image capture device 102 can be used to generate the body mask. Video processing manager 110 receives a motion vector of the current frame, which can be generated by an optical flow measurement tool on image capture device 102. The optical flow measurement tool can use the previous frame and the current frame to generate the motion vector. Video processing manager 110 receives a predicted mask of the current frame. The predicted mask can be generated based on the motion vector and the previous frame. For example, video processing manager 110 can modify (e.g., translate, scale, rotate) the mask of the previous frame according to the motion vector of the current frame. Video processing manager 110 generates a final mask for the current frame based on the body mask, the motion vector, and the predicted mask. Then, video processing manager 110 applies the final mask to the current frame to provide a masked frame for which the foreground and the background are segmented from each other. Video processing manager 110 edits the masked frame to remove distortion from the masked frame to generate an output frame. As an example, the distortion can be jitter in the background of the masked frame, and the editing applied by video processing manager 110 can be motion blur. Video processing manager 110 can apply motion blur to the background of the masked frame to hide the jitter. Next, video processing manager 110 outputs the output frame.
[0012] Video processing manager 110 can repeat the method described herein for a second current frame, relative to which the current frame is a second previous frame. As an example, the sequence of frames includes a first frame, a second frame, a third frame, and a fourth frame. For the first iteration of the method, the first frame is the previous frame and the second frame is the current frame. Video processing manager can implement one or more portions of the method to generate a final mask for the second frame. For the second iteration of the method, the second frame is the second previous frame and the third frame is the second current frame. In the context of the second iteration, the second frame is the previous frame and the third frame is the current frame. For the third iteration of the method, the third frame is the second previous frame and the fourth frame is the second current frame. That is, in the context of the third iteration, the third frame is the previous frame and the fourth frame is the current frame. Although four frames are described in this example, the sequence of frames can include any number of frames, and video processing manager can iterate through the previous frame, the current frame, the second previous frame, and the second current frame until a final mask is generated for each frame in the sequence of any number of frames.
[0013] After recording a video clip of athlete 114 sprinting past tree 116 in the background in the foreground, user 112 views the video clip. As Figure 1As shown, display 104-1 shows the first frame in a sequence of three or more frames. Sprinter 114 is centered in the foreground of the first frame. Tree 116-1 is located to the right of the background of the first frame. As shown, tree 116-1 includes a dark front and two light-colored sides on the left. The two light-colored sides of tree 116-1 represent motion blur applied to the background of the first frame.
[0014] Display 104-2 further shows the second frame in a sequence of three or more frames. Again, athlete 114 centered in the foreground of the second frame sprints past tree 116-2 centered in the background of the second frame. Tree 116-2 includes a dark front and two light-colored sides on the left to represent motion blur applied to the background of the second frame.
[0015] Display 104-3 further shows the third frame in a sequence of three or more frames. Athlete 114 remains centered in the foreground of the third frame. Athlete 114 continues to sprint past tree 116-3 on the left side of the background of the third frame. Again, tree 116-3 includes a dark front and two light-colored left sides to represent motion blur applied to the background of the third frame. User 112 is satisfied with the video clip because video processing manager 110 extracts the foreground from the video clip in real time and applies motion blur only to the background to hide jitter in the background. Although jitter is described, the distortion can be any distortion. Similarly, although editing the background of the video clip is described, video processing manager 110 can implement any of the disclosed systems and techniques to edit the foreground of the video clip. Example implementation
[0016] Figure 2 Shows in more detail from Figure 1Example implementation 200 of the image capture device 102. The image capture device 102 is shown as a variety of example devices, including consumer electronic devices. As a non-limiting example, the image capture device 102 can be a smartphone 102-1, a tablet computer 102-2, a laptop computer 102-3, a desktop computer 102-4, a smartwatch 102-5, a pair of smart glasses 102-6, a game controller 102-7, a speaker 102-8, or a microwave appliance 102-9. Although not shown, the image capture device 102 can also be implemented as an audio recording device, a health monitoring device, a home automation system, a home security system, a game console, a personal media device, a personal assistant device, a drone, a household appliance, etc. It should be noted that the image capture device 102 can be wearable, non-wearable but mobile, or relatively immobile (e.g., a desktop computer, a household appliance). It should also be noted that the image capture device 102 can be used with or embedded in many image capture devices 102 or peripheral devices, such as in a motor vehicle or as an attachment to a personal computer. The image capture device 102 can include additional components and interfaces that are omitted for clarity from Figure 2 Additional components and interfaces that are omitted.
[0017] As Figure 2 shown, the image capture device 102 includes a display 104, a camera 106, and a processor 108. The display 104 can be any of a variety of displays, including a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, an in-plane switching (IPS) display, a twisted nematic (TN) display, and so on. The display can be referred to as a "screen" such that content (e.g., images, videos) can be "displayed on the screen". The camera 106 can include one or more image sensors, one or more lenses, one or more autofocus motors, a flash, an image stabilization component, and so on. The camera 106 can be configured to capture video at various resolutions (e.g., 1080p, 2k, 4k) and frame rates (e.g., 30 fps, 60 fps, 120 fps). The camera 106 can include an associated application with which the user can interact to adjust capture settings (e.g., resolution, frame rate) and view the captured images and videos. The processor 108 can include one or more of a suitable single-core or multi-core processor (such as a graphics processing unit (GPU) or a central processing unit (CPU)).
[0018] The image capture device 102 includes also in Figure 2The computer-readable medium (CRM) 202 shown in [figure]. The CRM 202 includes a memory medium 204 and a storage medium 206. The memory medium 204 and the storage medium 206 may include one or more non-transitory storage devices, such as random access memory (RAM), dynamic RAM (DRAM), solid state drive (SSD), magnetic rotating hard disk drive (HDD), or any type of storage medium suitable for storing electronic instructions, each of which is coupled to a data bus. The term "coupled" may refer to two or more elements in direct contact (e.g., physical, electrical, magnetic, optical), or to two or more elements that do not directly contact each other but still cooperate or interact with each other.
[0019] The CRM 202 further includes an operating system (OS) 208, an application 210, and a video processing manager 110. The OS 208, the application 210, and the video processing manager 110 may be implemented as computer-readable instructions on the CRM 202, which can be executed by a processor 108 to provide some or all of the functions described herein. For example, the processor 108 may execute specific computing tasks of the OS 208, which are designed to remove distortions from live video using masked frames. The application 210 may include a power management application, a camera application, a background service application, a communication application (e.g., audio call, video call), and so on.
[0020] In various aspects, the implementation of the video processing manager 110 may include one or more integrated circuits (ICs), system-on-chips (SoCs), secure key stores, hardware embedded with firmware stored in read-only memory (ROM), printed circuit boards (PCBs) with various hardware components, or any combination thereof. As described herein, a system for removing distortions from live video using masked frames may include one or more components of an image capture device 102 configured to remove distortions from live video using masked frames, as Figure 1 and Figure 2 shown. In additional implementations, a system for removing distortions from live video using masked frames may be implemented as the image capture device 102.
[0021] As Figure 2As further shown, the image capture device 102 includes an input / output (I / O) port 212. The I / O port 212 enables the image capture device 102 to interact with other devices or users via peripherals, thereby transmitting any combination of digital, analog, and radio frequency signals. The I / O port 212 may include any combination of internal or external ports, such as universal serial bus (USB) ports, audio ports, video ports, dual in-line memory module (DIMM) card slots, peripheral component interconnect express (PCIe) slots, and so on. Various peripheral devices may be operatively coupled to the I / O port 212, such as human input devices (HIDs), external CRMs, speakers, displays, keyboards, mice, or other peripheral devices. Although not shown, the image capture device 102 may also include a system bus, interconnect, or data transfer system coupled to various components within the image capture device 102. The system bus or interconnect may include any one or combination of different bus structures, such as a memory bus, a peripheral bus, a USB, a local bus, or a processor bus utilizing one of multiple bus architectures.
[0022] In addition, the image capture device 102 includes one or more sensors 214, as Figure 2 shown. The sensors 214 may be disposed anywhere on or within the image capture device 102. Additionally or alternatively, the sensors 214 may be disposed on or within a peripheral device (e.g., wirelessly, wired) connected to the image capture device 102. The sensors 214 may include any one of a variety of sensing components, such as audio sensors (e.g., microphones), touch input sensors (e.g., touchscreens), image sensors (e.g., phase detection autofocus sensors, cameras, or a part of a camera system), ambient light sensors (e.g., photodetectors), acceleration sensors (e.g., accelerometers), proximity sensors (e.g., laser detection autofocus sensors), or pressure sensors (e.g., barometers). The sensing components may be disposed within the housing of the image capture device 102. In an implementation, the image capture device may include any one or more than one of the sensing components. Example methods
[0023] In the following sections, Figure 1 and Figure 2The video processing manager 110 in [the device] can perform example methods for implementing various aspects of removing distortion from live video using masked frames. The method is shown as a set of blocks that specify operations or actions performed by the video processing manager 110, the processor 108, the sensor 214, or other unspecified components of the image capture device. The method is not limited to the order or combination of blocks shown for performing operations in the respective blocks. Additionally, any one or more of the operations can be repeated, combined, reorganized, or linked to provide additional or alternative methods. In the following discussion, reference can be made, for example, only to Figure 1 and Figure 2 the example implementations and entities detailed therein.
[0024] Figure 3 An example method 300 for extracting a foreground from a video clip according to one or more aspects is shown. At 302, the video processing manager (e.g., video processing manager 110) receives a video clip. The video processing manager can use an image capture device (e.g., image capture device 102), its components (e.g., camera 106, sensor 214), or a combination thereof to capture the video clip. For example, the video processing manager can utilize the camera of the image capture device to capture the video clip. The video clip consists of a sequence of frames that includes a previous frame and a current frame received immediately after the previous frame. As an example, the video clip can include ten frames, where the first frame is the previous frame. Thus, the second frame immediately after the first frame is the current frame.
[0025] At 304, the video processing manager receives a body mask of the current frame. The body mask can be generated using a machine learning (ML) model. The ML model can be trained using labeled bodies such as humans, vehicles, pets, or other subjects of interest that can reside in the foreground of the video clip. Additionally, the ML model can reside on the image capture device, for example, as computer-readable instructions stored on the CRM (e.g., CRM 202) of the image capture device. At 306, the video processing manager receives a motion vector of the current frame. The motion vector can be generated by an optical flow measurement tool (e.g., an inverse compositional implementation of the Lucas-Kanade method) using the previous frame and the current frame. The optical flow measurement tool can be stored as computer-readable instructions on the CRM of the image capture device. The motion vector generated by the optical flow measurement tool describes the change in position of pixels or groups of pixels from the previous frame to the current frame. For example, the motion vector can be stored on the CRM (e.g., CRM 202) of the image capture device as a heatmap or another suitable encoding. At 308, the video processing manager receives a predicted mask of the current frame. The predicted mask is generated based on the motion vector and the previous frame. For example, the mask of the previous frame can be aligned to the current frame (e.g., rotated, translated, scaled to the current frame) using the motion vector.
[0026] At 310, the video processing manager generates a final mask for the current frame. The final mask is based on the body mask, the motion vector, and the predicted mask. For example, the video processing manager can generate the final mask by combining the body mask and the predicted mask. Although not shown, the video processing manager can apply a sharpening operation to the final mask before proceeding to 312. The sharpening operation can be based on a luminance (e.g., grayscale) version of the current frame, a bilateral grid, and the final mask. The sharpening operation can sharpen the edges of the final mask to prevent the final mask from becoming dull or rough.
[0027] At 312, the video processing manager applies the final mask to the current frame to provide a masked frame. The final mask divides the foreground of the masked frame from the background of the masked frame. At 314, the video processing manager edits the masked frame to remove distortion from the masked frame to generate an output frame. As an example, the distortion can be jitter in the background of the masked frame. The jitter can be caused by 3:2 pulldown, video footage at a low frame rate (e.g., 24 fps, 30 fps), a fast-panning video footage, or a combination thereof. By using the final mask to divide the foreground and background of the masked frame, the video processing manager enables separate editing of the foreground and background of the masked frame. Thus, the video processing manager can apply motion blur only to the background of the masked frame to hide the jitter. At 316, the video processing manager outputs the output frame. In this example, the output frame includes motion blur applied to the background but no editing applied to the foreground, thereby hiding the jitter in the background and maintaining a sharp foreground. In an implementation, the video processing manager can apply editing to the foreground, the background, both the foreground and the background, or neither the foreground nor the background. Further, the editing can include one or more of a variety of edits, including but not limited to cropping, color adjustment, highlight adjustment, filter application, Gaussian blur, or motion blur.
[0028] Figure 4 An example method 400 for segmenting a current frame of a frame sequence to produce a segmentation result is shown. Any of the blocks shown in example method 400 can be repeated, combined, reorganized, or linked to a set of blocks in example method 400 or example method 300. For example, example method 400 can utilize the motion vector received at 306 of example method 300.
[0029] At 402, the video processing manager quantizes the motion vector of the current frame into two or more bins. The two or more bins can each include one or more motion vectors. The bins group similar motion vectors together, which can be used in step 404.
[0030] At 404, the video processing manager calculates the average motion vector of the interval that contains the majority of the motion vectors among two or more intervals. As an example, refer to Figure 1 Example environment 100 of. Since user 112 translated the image capture device 102 to keep athlete 114 centered in the foreground of the video clip, the motion vectors can be grouped into two intervals. The first interval contains the motion vectors of the background, while the second interval contains the motion vectors of the foreground. Further, since athlete 114 occupies less space in each frame of the video clip, the background motion vector interval can be the interval that contains the majority of the motion vectors. Thus, the video processing manager can calculate the average motion vector of the background interval.
[0031] At 406, the video processing manager compares the motion vectors of two or more intervals with the average motion vector to produce a comparison result. In this example, the average motion vector of the background interval is greater than any of the motion vectors in the foreground interval. This difference is due to user 112 translating the image capture device 102 to keep athlete 114 centered in the foreground. Relative to the image capture device 102, athlete 114 does not move. In other words, the foreground motion vectors are close to zero. Different from athlete 114 in the foreground, the tree 116 in the background moves relative to the image capture device 102. In other words, the background motion vectors are greater than zero. In this example, the comparison result can indicate that the foreground motion vectors are less than the average motion vector of the background motion vectors.
[0032] At 408, the video processing manager classifies one or more of the motion vectors as outliers based on the comparison result exceeding a threshold. The threshold can be an integer, a fraction, a percentage, a difference relative to another value (e.g., the average motion vector), or another quantifier that can be compared with the motion vectors. In this example, the video processing manager can classify the foreground motion vectors as outliers based on them being a certain difference (e.g., ten percent, fifteen percent) smaller than the average motion vector. As an additional example, if a motion vector is a certain difference (e.g., ten percent, twenty-five percent) larger than the average motion vector, the video processing manager can classify it as an outlier. As a further example, if a motion vector is close to zero, close to infinity, close to another integer, or close to another independent value, the video processing manager can classify it as an outlier.
[0033] At 410, the video processing manager segments the current frame based on outliers to produce a segmentation result of the current frame. Continuing with this example, the segmentation result can include two segments, one for the foreground and one for the background. The video processing manager can separately edit the background from the foreground, or the foreground from the background, or a combination of both, based on the foreground segment and the background segment. Further, when combined with example method 300, the video processing manager can generate a final mask by combining the segmentation result with the predicted mask.
[0034] In some aspects, the video processing manager can utilize distance information from an additional sensor (e.g., sensor 214) of the image capture device. For example, the video processing manager can utilize distance information from a proximity sensor (e.g., sonar, radar, lidar) to more quickly or accurately identify the foreground or background of a previous frame or the current frame. The distance information can include distance measurements of the foreground (e.g., 6 m, 15 m) and distance measurements of the background (e.g., 25 m, 31 m). The video processing manager can segment the foreground of the current frame from the background of the current frame based on the distance information from the proximity sensor. As another example, the video processing manager can utilize a second camera having a different viewpoint from the first camera to identify the foreground or background of the current frame. The background of the current frame may look similar from the viewpoint of the first camera and the viewpoint of the second camera. The foreground of the current frame may look different from the viewpoint of the first camera and the viewpoint of the second camera. The video processing manager can segment the foreground of the current frame from the background based on the difference in the foregrounds or the similarity of the backgrounds from different viewpoints.
[0035] Throughout the discussion, examples are provided of a video processing manager editing the background of frames in a sequence of frames of a video clip. However, the systems and techniques described herein are not limited to editing the background of frames. In various aspects, the systems and techniques can also be implemented by a video processing manager to edit the foreground of frames. Additionally or alternatively, the systems and techniques described herein can be implemented by a video processing manager in a long exposure photo application. For example, assume a user wants to take a photo of a subject in low light. To do so, the user composes a shot of the subject using an image capture device having a video processing manager configured to remove distortion from a live video using masked frames. The video processing manager can capture multiple frames of the subject in low light using a long exposure time. The long exposure time provides sufficient time for the image sensor of the image capture device to capture sufficient light for each frame. If the user's hand shakes when taking multiple frames at a long exposure time, the subject may be blurry. However, the video processing manager can implement the techniques and systems described herein to segment the background and foreground of the multiple frames. For example, the video processing manager can use motion vectors to perform the segmentation. The video processing manager can also use motion vectors to stabilize the foreground of the long exposure frames in real time (e.g., by aligning the foreground mask to the current frame using the motion vectors), resulting in a clear foreground. For example, multiple frames can be combined (e.g., superimposed) into a single output photo with a clear foreground. Additional Examples
[0036] In the following sections, additional examples are provided.
[0037] Example 1: A method, comprising: receiving a video clip that includes a sequence of frames that includes a previous frame and a current frame, the current frame being immediately sequenced after the previous frame; receiving a subject mask for the current frame, the subject mask being generated using a machine learning (ML) model; receiving a motion vector for the current frame, the motion vector being generated by an optical flow measurement tool using the previous frame and the current frame; receiving a predicted mask for the current frame, the predicted mask being generated based on the motion vector and the previous frame; generating a final mask for the current frame, the final mask being based on the subject mask, the motion vector, and the predicted mask; applying the final mask to the current frame to provide a masked frame; editing the masked frame to remove distortion from the masked frame to generate an output frame; and outputting the output frame.
[0038] Example 2: The method of Example 1, wherein: the video clip is captured by and received from a camera of an image capture device; the ML model is located on the image capture device; and the optical flow measurement tool is located on the image capture device.
[0039] Example 3: The method as described in Example 1 further includes: quantizing the motion vectors of the current frame into two or more intervals; calculating the average motion vector of the interval that contains most of the motion vectors among the two or more intervals; comparing the motion vectors of the two or more intervals with the average motion vector to generate a comparison result; classifying one or more of the motion vectors as outliers based on the comparison result exceeding a threshold; and segmenting the current frame based on the outliers to generate a segmentation result of the current frame.
[0040] Example 4: The method as described in Example 3, wherein the final mask is generated by combining the segmentation result of the current frame with the predicted mask.
[0041] Example 5: The method as described in Example 1, wherein the predicted mask is generated by aligning the final mask of the previous frame to the current frame using the motion vectors.
[0042] Example 6: The method as described in Example 1 further includes: performing a sharpening process on the final mask before applying the final mask to the current frame.
[0043] Example 7: The method as described in Example 6, wherein: the sharpening process is performed by an edge sharpening tool; and the edge sharpening tool is located on the image capture device from which the video clip is received.
[0044] Example 8: The method as described in Example 1, wherein the previous frame and the current frame include a background, a foreground in front of the background, and a subject of interest in the foreground.
[0045] Example 9: The method as described in Example 8, wherein editing the masked frame to generate the output frame further includes: editing the background of the masked frame.
[0046] Example 10: The method as described in Example 8, wherein editing the masked frame to generate the output frame further includes: editing the foreground of the masked frame.
[0047] Example 11: The method as described in Example 8, wherein editing the masked frame to generate the output frame further includes: editing the foreground and the background of the masked frame.
[0048] Example 12: The method as described in Example 8 further includes: receiving distance information of the current frame, the distance information being captured by a sensor of the image capture device; and segmenting the foreground of the current frame from the background of the current frame based on the distance information.
[0049] Example 13: The method as described in Example 8 further includes: receiving the viewpoint information of the current frame, where the viewpoint information is captured by different image capture devices; and segmenting the foreground of the current frame from the background of the current frame based on the viewpoint information.
[0050] Example 14: An image capture device includes: at least one camera; one or more sensors; one or more processors; and a memory that stores instructions that, when executed by the one or more processors, cause the one or more processors to implement a video processing manager to provide video processing by using the at least one camera and the one or more processors by executing the method as described in any one of the preceding claims.
[0051] Example 15: A computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to execute the method as described in any one of claims 1 to 13. Conclusion
[0052] Unless the context otherwise requires, the use of the word "or" herein may be regarded as an "inclusive or" or the use of a term that permits the inclusion or application of one or more of the items joined by the word "or" (e.g., the phrase "A or B" may be interpreted to permit only "A", only "B", or both "A" and "B"). Further, as used herein, a phrase referring to "at least one" in a list of items means any combination of those items, including a single member. For example, "at least one of a, b, or c" may cover a, b, c, a - b, a - c, b - c, and a - b - c, as well as any combination having multiples of the same element (e.g., a - a, a - a - a, a - a - b, a - a - c, a - b - b, a - c - c, b - b, b - b - b, b - b - c, c - c, and c - c - c, or any other ordering of a, b, and c). Additionally, the items represented in the drawings and the terms discussed herein may refer to one or more items or terms, and thus the items and terms in the written description may be referred to in their singular or plural forms interchangeably.
[0053] Although systems and techniques for using masked frames to remove distortions from live video and implementations of devices enabling the use of masked frames to remove distortions from live video have been described in language specific to certain features and / or methods, the subject matter of the appended claims is not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed as example implementations of using masked frames to remove distortions from live video.
Claims
1. A method, comprising: Receiving a video clip, the video clip including a sequence of frames, the sequence of frames including a previous frame and a current frame, the current frame being immediately sequenced after the previous frame; Receiving a body mask of the current frame, the body mask being generated using a machine learning (ML) model; Receiving a motion vector of the current frame, the motion vector being generated by an optical flow measurement tool using the previous frame and the current frame; Receiving a predicted mask of the current frame, the predicted mask being generated based on the motion vector and the previous frame; Generating a final mask of the current frame, the final mask being based on the body mask, the motion vector, and the predicted mask; Applying the final mask to the current frame to provide a masked frame; Editing the masked frame to remove distortion from the masked frame to generate an output frame; And Outputting the output frame.
2. The method of claim 1, wherein: The video clip is captured by and received from a camera of an image capture device; The ML model is stored on a computer-readable medium (CRM) of the image capture device; and The optical flow measurement tool is stored on the CRM of the image capture device.
3. The method of claim 1, further comprising: Quantizing the motion vector of the current frame into two or more intervals; Calculating an average motion vector of an interval among the two or more intervals that contains most of the motion vectors; Comparing the motion vectors of the two or more intervals with the average motion vector to produce a comparison result; Classifying one or more of the motion vectors as outliers based on the comparison result exceeding a threshold; And Segmenting the current frame based on the outliers to produce a segmentation result of the current frame.
4. The method of claim 3, wherein the final mask is generated by combining the segmentation result of the current frame with the predicted mask.
5. The method of claim 1, wherein the predicted mask is generated by aligning the final mask of the previous frame to the current frame using the motion vector.
6. The method of claim 1, further comprising: Performing a sharpening process on the final mask before applying the final mask to the current frame.
7. The method of claim 6, wherein: The sharpening process is performed by an edge sharpening tool; and The edge sharpening tool is located on the image capture device from which the video clip is received.
8. The method of claim 1, wherein the previous frame and the current frame include a background, a foreground in front of the background, and an interested body in the foreground.
9. The method of claim 8, wherein editing the masked frame to generate the output frame further comprises: Editing the background of the masked frame.
10. The method of claim 8, wherein editing the masked frame to generate the output frame further comprises: Editing the foreground of the masked frame.
11. The method according to claim 8, wherein editing the masked frame to generate the output frame further comprises: Editing the foreground and the background of the masked frame.
12. The method according to claim 8, further comprising: Receiving distance information of the current frame, the distance information being captured by a sensor of an image capture device; And Segmenting the foreground of the current frame from the background of the current frame based on the distance information.
13. The method according to claim 8, further comprising: Receiving viewpoint information of the current frame, the viewpoint information being captured by a different image capture device; And Segmenting the foreground of the current frame from the background of the current frame based on the viewpoint information.
14. An image capture device, comprising: At least one camera; One or more sensors; One or more processors; And A memory storing instructions that, when executed by the one or more processors, cause the one or more processors to implement a video processing manager to provide video processing by using the at least one camera, the one or more sensors, and the one or more processors by executing the method according to any one of the preceding claims.
15. A computer-readable medium (CRM) comprising instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to any one of claims 1 to 13.