Removing Distortion from Real-Time Video Using Masked Frames

By segmenting foreground from background using masked frames, the video processing manager applies motion blur only to the background, addressing the issue of global blur and improving video clarity.

JP2025532908APending Publication Date: 2025-10-03GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025518296
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-28
Filing Date
2023-09-26
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Conventional video processing methods apply motion blur globally, leading to undesirable blurring of significant foreground objects in video segments, while attempting to hide judder artifacts in the background.

Method used

A video processing manager uses masked frames to segment foreground from background, applying motion blur only to the background to hide judder, while maintaining a sharp foreground.

Benefits of technology

This approach effectively reduces judder in the background while preserving the clarity of foreground objects, enhancing video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532908000001_ABST
    Figure 2025532908000001_ABST
Patent Text Reader

Abstract

Described herein are systems and techniques for removing distortion from real-time video using masked frames. In an aspect, an image capture device having a video processing manager is configured to capture a video segment including a sequence of frames. The sequence of frames includes at least a current frame having a foreground and a background. The video processing manager receives an object mask, a motion vector, and a predicted mask for the current frame. The video processing manager generates a final mask for the current frame based on the object mask, the motion vector, and the predicted mask. The video processing manager applies the final mask to the current frame to segment the foreground from the background and provide a masked frame. The video processing manager edits the masked frames to remove distortion and generate an output frame, and outputs the output frame. By repeating the described method for each frame in the sequence of frames, the video processing manager provides an improved video segment.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference to related disclosures This application claims priority to U.S. Provisional Patent Application No. 63 / 377,484, filed September 28, 2022, which is incorporated herein by reference in its entirety. [Background technology]

[0002] Many video applications apply modifications to a video segment in a global manner. That is, for a video segment that includes a foreground and a background, the modifications are applied equally to both the foreground and the background. For example, motion blur can be applied globally to a video segment to hide judder artifacts resulting from 3-2 pulldown or rapidly panning action shots. However, for a video segment that includes a significant foreground, applying motion blur globally can result in the significant foreground being blurred, which may be undesirable. Summary of the Invention

[0003] Described herein are systems and techniques for removing distortion from real-time video using masked frames. In an aspect, an image capture device having a video processing manager is configured to receive a video segment including a sequence of frames. The sequence of frames includes at least a current frame having a foreground and a background. The video processing manager receives an object mask, a motion vector, and a predicted mask for the current frame. The video processing manager generates a final mask for the current frame based on the object mask, the motion vector, and the predicted mask. The video processing manager then applies the final mask to the current frame to segment the foreground from the background and provide a masked frame. The video processing manager edits the masked frames to remove distortion and generate an output frame, and outputs the output frame. By repeating the described method for each frame in the sequence of frames, the video processing manager provides an improved video segment with less distortion.

[0004] The details of one or more aspects of using masked frames to remove distortion from real-time video are set forth in the accompanying drawings and the description below. Other features and advantages will be apparent from the drawings, the description, and the claims. This Summary is provided to introduce the subject matter that is further described in the Detailed Description and the Drawings. As such, this Summary is not intended to describe essential features or to limit the scope of the claimed subject matter.

[0005] This specification describes systems and techniques for removing distortion from real-time video using masked frames with reference to the following drawings, in which like numbers are used throughout to reference like features and components: [Brief explanation of the drawings]

[0006] [Figure 1]1 illustrates an exemplary environment in which an image capture device may implement an aspect of using masked frames to remove distortion from real-time video. [Figure 2] 2 illustrates an exemplary implementation of the image capture device of FIG. 1 in more detail. [Figure 3] 1 illustrates an example method for extracting foreground from a video segment, according to one or more aspects. [Figure 4] 1 illustrates an exemplary method for segmenting a current frame of a sequence of frames to generate a segmentation result. DETAILED DESCRIPTION OF THE INVENTION

[0007] overview Since the advent of video cameras, engineers have strived to improve their capabilities so that the video they capture appears lifelike. To do so, engineers have focused on improving resolution and frame rate capabilities. Unlike a display, reality is not limited to a finite resolution consisting of a finite number of individual pixels or light sources. Rather, reality can be thought of as having infinite resolution. Therefore, cameras that can capture video of a scene at high resolution are important for a lifelike representation of a scene. Also, unlike a display, reality is not limited to a finite frame rate consisting of a finite number of frames or still images displayed in rapid succession. Rather, reality can be thought of as having an infinite frame rate. Therefore, cameras that can capture video of a scene at very high frame rates are important for a lifelike representation of a scene. Given unlimited resources, space, and time, designing a camera with such capabilities is trivial. However, because unlimited resources, space, and time do not exist, engineers have developed alternative solutions.

[0008] As an example, a parent attends their child's track meet, where the child is scheduled to compete in the one hundred meter (100m) race. The child lines up at the starting line and swings their legs as they warm up. Before the starting pistol is fired, the parent opens the camera application on their smartphone and selects video mode set to capture 1080p recording at 30 frames per second (fps). Once the starting pistol is fired, the child takes off onto the track like a rocket. Meanwhile, the parent begins capturing a video segment using the smartphone. The parent quickly pans the camera to keep the child focused and centered in the frame as the child swiftly passes the parent and heads toward the finish line. As the child approaches the finish line, they thrust their shoulders and head forward, just in time to cross the line.

[0009] When a child returns from the 100-meter race, the parent and child review a video segment of the child's performance. In many conventional approaches, the video contains imperfections, such as blurring, in the foreground, including the child, and in the background, including the stands and other athletes. This can be due to the parent's smartphone's real-time video processing pipeline globally applying motion blur to the foreground and background of the video segment to hide judder resulting from the parent quickly panning the smartphone. Without motion blur, the background of the video segment may contain judder, or stuttering artifacts, resulting from a quickly panned action shot. However, while motion blur reduces noticeable judder, the parent's foreground, including the child, is also blurred. This video quality is undesirable for both the parent and the child.

[0010] Described herein are systems and techniques for removing distortion from real-time video using masked frames. The disclosed systems and techniques may address foreground blurring in video segments resulting from the global application of motion blur. The systems and techniques extract the foreground from a video segment, thereby separating the foreground from the background and enabling editing of the background separately from the foreground. The following discussion describes an operating environment, techniques that may be employed within the operating environment, and example methods. While systems and techniques directed to removing distortion from real-time video using masked frames are described, the subject matter of the appended claims is not limited to the particular features or methods described. Rather, the particular features and methods are disclosed as exemplary implementations, and reference is made to the operating environment by way of example only.

[0011] Example Environment 1 illustrates an exemplary environment 100 in which an image capture device 102 may implement an embodiment for removing distortion from real-time video using masked frames. The image capture device 102 includes a display 104, a camera 106, one or more processors 108, and a video processing manager 110 configured to extract the foreground of a video segment in real time. In one example, a user 112 wishes to capture video of an athlete 114 sprinting past a tree 116. The user 112 takes out the image capture device 102, opens a camera application (not shown) installed on the image capture device 102, selects video mode (not shown), and taps a shutter button (not shown) to begin recording the athlete 114 sprinting past the tree 116.

[0012] In response to the user 112 tapping the shutter button, the video processing manager 110 captures a video segment including a sequence of frames, including a previous frame and a current frame received immediately after the first frame. The video processing manager 110 receives an object mask for the current frame. The object mask may be generated, for example, using a machine learning (ML) model in the image capture device 102. The video processing manager 110 receives a motion vector for the current frame, which may be generated by an optical flow measurement tool in the image capture device 102. The optical flow measurement tool may generate the motion vector using, for example, the previous frame and the current frame. The video processing manager 110 receives a predicted mask for the current frame. The predicted mask may be generated from the motion vector and the previous frame. For example, the video processing manager 110 may modify (e.g., translate, scale, rotate) the mask for the previous frame according to the motion vector for the current frame. The video processing manager 110 generates a final mask for the current frame based on the object mask, the motion vector, and the predicted mask. The video processing manager 110 then applies the final mask to the current frame to provide a masked frame in which the foreground and background are segmented from each other. The video processing manager 110 edits the masked frame to remove distortion from the masked frame and generate an output frame. As an example, the distortion may be judder in the background of the masked frame, and the edit applied by the video processing manager 110 may be motion blur. The video processing manager 110 may apply motion blur to the background of the masked frame to hide the judder. The video processing manager 110 then outputs the output frame.

[0013] The video processing manager 110 may repeat the method described herein for a second current frame, for which the current frame is the second previous frame. As an example, the sequence of frames includes a first frame, a second frame, a third frame, and a fourth frame. For the first iteration of the method, the first frame is the previous frame and the second frame is the current frame. The video processing manager may implement one or more portions of the method to generate a final mask for the second frame. For the second iteration of the method, the second frame is the second previous frame and the third frame is the second current frame. Within the context of the second iteration, the second frame is the previous frame and the third frame is the current frame. For the third iteration of the method, the third frame is the second previous frame and the fourth frame is the second current frame. That is, within the context of the third iteration, the third frame is the previous frame and the fourth frame is the current frame. Although four frames are described in this example, the sequence of frames may include any number of frames, and the video processing manager may iterate through the previous frame, the current frame, the second previous frame, and the second current frame until a final mask is generated for each frame in the series of any number of frames.

[0014] After recording a video segment of an athlete 114 in the foreground sprinting past a tree 116 in the background, a user 112 reviews the video segment. As shown in FIG. 1 , display 104-1 shows a first frame of a sequence of three or more frames. The sprinting athlete 114 is centered in the foreground of the first frame. A tree 116-1 is located on the right side in the background of the first frame. As shown, tree 116-1 includes a dark front and two light sides on the left side. The two light sides of tree 116-1 represent motion blur applied to the background of the first frame.

[0015] Also shown by display 104-2 is a second frame of the sequence of three or more frames. Again, athlete 114, centered in the foreground of the second frame, sprints past tree 116-2, centered in the background of the second frame. Tree 116-2 includes a dark front and two light sides to the left to represent motion blur applied to the background of the second frame.

[0016] Display 104-3 further shows a third frame of the sequence of three or more frames. Athlete 114 remains centered in the foreground of the third frame. Athlete 114 continues sprinting past tree 116-3 on the left side of the background of the third frame. Again, tree 116-3 includes a dark front and two light left faces to represent the motion blur applied to the background of the third frame. User 112 is satisfied with the video segment because video processing manager 110 extracted the foreground from the video segment in real time and applied motion blur to the background only to hide the judder in the background. While judder is described, distortion can be any distortion. Similarly, while editing the background of a video segment is described, video processing manager 110 may implement any one of the disclosed systems and techniques for editing the foreground of a video segment.

[0017] Exemplary Implementations FIG. 2 illustrates an example implementation 200 of the image capture device 102 of FIG. 1 in more detail. The image capture device 102 is shown as various example devices, including consumer electronic devices. By way of non-limiting example, the image capture device 102 may be a smartphone 102-1, a tablet 102-2, a laptop computer 102-3, a desktop computer 102-4, a smartwatch 102-5, smart glasses 102-6, a game controller 102-7, a speaker 102-8, or a microwave appliance 102-9. While not shown, the image capture device 102 may also be implemented as an audio recording device, a health monitoring device, a home automation system, a home security system, a game console, a personal media device, a personal assistant device, a drone, a consumer electronics device, and the like. Note that the image capture device 102 may be wearable, non-wearable but mobile, or relatively stationary (e.g., a desktop computer, a consumer electronics device). Also, note that image capture device 102 can be used with or incorporated within many image capture devices 102 or peripherals, such as in an automobile or as an attachment to a personal computer. Image capture device 102 may include additional components and interfaces that are omitted from FIG. 2 for clarity.

[0018] As shown in FIG. 2 , image capture device 102 includes display 104, camera 106, and processor 108. Display 104 can be any one of a variety of displays, including a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, an in-plane switching (IPS) display, a twisted nematic (TN) display, etc. The display may be referred to as a “screen” so that content (e.g., images, video) may be displayed “on the screen.” Camera 106 may include one or more image sensors, one or more lenses, one or more autofocus motors, a flash, image stabilization components, etc. Camera 106 may be configured to capture video at various resolutions (e.g., 1080p, 2k, 4k) and frame rates (e.g., 30 fps, 60 fps, 120 fps). Camera 106 may include associated applications with which a user may interact to adjust capture settings (e.g., resolution, frame rate) and review captured images and video. The processor 108 may include one or more suitable single-core or multi-core processors, such as a graphics processing unit (GPU) or a central processing unit (CPU).

[0019] 2, the image capture device 102 includes a computer-readable medium (CRM) 202. The CRM 202 includes a memory medium 204 and a storage medium 206. The memory medium 204 and the storage medium 206 may each include one or more non-transitory storage devices, such as random access memory (RAM), dynamic RAM (DRAM), a solid-state drive (SSD), a magnetic hard drive disk (HDD), or any other type of storage medium suitable for storing electronic instructions, coupled with a data bus. The term "coupled" may refer to two or more elements in direct contact (e.g., physically, electrically, magnetically, optically) or two or more elements that are not in direct contact with each other but still cooperate or interact with each other.

[0020] CRM 202 further includes an operating system (OS) 208, applications 210, and a video processing manager 110. OS 208, applications 210, and video processing manager 110 may be implemented as computer-readable instructions on CRM 202 that can be executed by processor 108 to provide some or all of the functionality described herein. For example, processor 108 may perform specific computational tasks of OS 208 directed to removing distortion from real-time video using masked frames. Applications 210 may include power management applications, camera applications, background service applications, communication applications (e.g., audio calling, video calling), etc.

[0021] In aspects, an implementation of the video processing manager 110 may include one or more integrated circuits (ICs), a system-on-chip (SoC), a secure key store, hardware with embedded firmware stored in read-only memory (ROM), a printed circuit board (PCB) with various hardware components, or any combination thereof. As described herein, a system for removing distortion from real-time video using masked frames may include one or more components of an image capture device 102 configured to remove distortion from real-time video using masked frames, as shown in Figures 1 and 2. In additional embodiments, a system for removing distortion from real-time video using masked frames may be implemented as an image capture device 102.

[0022] As further shown in FIG. 2 , image capture device 102 includes input / output (I / O) ports 212. I / O ports 212 allow image capture device 102 to interact with other devices or users through peripheral devices and transmit any combination of digital, analog, or radio frequency signals. I / O ports 212 may include any combination of internal or external ports, such as Universal Serial Bus (USB) ports, audio ports, video ports, dual in-line memory module (DIMM) card slots, and Peripheral Component Interconnect Express (PCIe) slots. Various peripherals may be operatively coupled to I / O ports 212, such as a human input device (HID), external CRM, speakers, a display, a keyboard, a mouse, or other peripherals. Although not shown, image capture device 102 may also include a system bus, interconnect, or data transfer system that couples various components within image capture device 102. The system bus or interconnect may include any one or combination of different bus structures such as a memory bus, a peripheral bus, USB, a local bus, or a processor bus utilizing one of a variety of bus architectures.

[0023] 2 , the image capture device 102 includes one or more sensors 214. The sensors 214 may be located anywhere on or in the image capture device 102. Additionally or alternatively, the sensors 214 may be located on or in a peripheral device connected (e.g., wirelessly, wired) to the image capture device 102. The sensors 214 may include any of a variety of sensing components, such as an audio sensor (e.g., a microphone), a touch input sensor (e.g., a touchscreen), an image sensor (e.g., a phase-difference detection autofocus sensor, a camera or part of a camera system), an ambient light sensor (e.g., a photodetector), an acceleration sensor (e.g., an accelerometer), a proximity sensor (e.g., a laser-detection autofocus sensor), or a pressure sensor (e.g., a barometer). The sensing components may be located within the housing of the image capture device 102. In an embodiment, the image capture device may include more than one of any one or more sensing components.

[0024] Exemplary Methods The following sections describe exemplary methods that the video processing manager 110 of FIGS. 1 and 2 may execute to implement aspects of removing distortion from real-time video using masked frames. The methods are illustrated as a set of blocks that specify operations or actions to be performed by the video processing manager 110, the processor 108, the sensor 214, or other components of the image capture device not mentioned. The methods are not limited to the order or combination of the set of blocks shown to perform the operations by the respective blocks. Furthermore, any one or more operations may be repeated, combined, rearranged, or linked to provide additional or alternative methods. The following discussion may refer, by way of example only, to the exemplary implementations and entities detailed in FIGS. 1 and 2.

[0025] 3 illustrates an example method 300 for extracting foreground from a video segment according to one or more aspects. At 302, a video processing manager (e.g., video processing manager 110) receives a video segment. The video processing manager may capture the video segment using an image capture device (e.g., image capture device 102), its components (e.g., camera 106, sensor 214), or a combination thereof. For example, the video processing manager may utilize the camera of the image capture device to capture the video segment. The video segment is composed of a sequence of frames including a previous frame and a current frame received immediately after the previous frame. As an example, the video segment may include 10 frames, the first of which is the previous frame. Thus, the second frame immediately following the first frame is the current frame.

[0026] At 304, the video processing manager receives an object mask for the current frame. The object mask may be generated using a machine learning (ML) model. The ML model may be trained using marked objects, such as humans, vehicles, pets, or other objects of interest that may be present in the foreground of the video segment. Furthermore, the ML model may reside on the image capture device, for example, as computer-readable instructions stored in the CRM (e.g., CRM 202) of the image capture device. At 306, the video processing manager receives motion vectors for the current frame. The motion vectors may be generated by an optical flow measurement tool (e.g., a reverse synthesis implementation of the Lucas-Kanade method) using the previous frame and the current frame. The optical flow measurement tool may be stored, for example, as computer-readable instructions in the CRM (e.g., CRM 202) of the image capture device. The motion vectors generated by the optical flow measurement tool describe the change in position of pixels or groups of pixels from the previous frame to the current frame. The motion vectors may be stored, for example, as a heat map or other suitable encoding in the CRM (e.g., CRM 202) of the image capture device. At 308, the video processing manager receives a predicted mask for the current frame. The predicted mask is generated from the motion vectors and the previous frame. For example, the mask for the previous frame may be aligned (e.g., rotated, translated, scaled) to the current frame using the motion vectors.

[0027] At 310, the video processing manager generates a final mask for the current frame. The final mask is based on the object mask, the motion vectors, and the predicted mask. For example, the video processing manager may generate the final mask by taking a combination of the object mask and the predicted mask. Although not shown, the video processing manager may apply a sharpening operation to the final mask before proceeding to 312. The sharpening operation may be based on a luminance (e.g., grayscale) version of the current frame, a bilateral grid, and the final mask. The sharpening operation may sharpen the edges of the final mask to avoid a blurry or rough final mask.

[0028] At 312, the video processing manager applies the final mask to the current frame to provide a masked frame. The final mask segments the foreground of the masked frame from the background of the masked frame. At 314, the video processing manager edits the masked frame to remove distortion from the masked frame and generate an output frame. By way of example, the distortion may be judder in the background of the masked frame. The judder may result from 3-2 pulldown, a low frame rate (e.g., 24 fps, 30 fps) video shot, a quickly panned video shot, or a combination thereof. By segmenting the foreground from the background of the masked frame using the final mask, the video processing manager enables separate editing of the foreground and background of the masked frame. Thus, the video processing manager may apply motion blur only to the background of the masked frame to hide the judder. At 316, the video processing manager outputs the output frame. In this example, the output frame includes motion blur applied to the background and no edits applied to the foreground, thereby hiding judder in the background and maintaining a sharp foreground. In embodiments, the video processing manager may apply edits to the foreground, the background, both the foreground and background, or neither the foreground nor the background. Furthermore, the edits may include any one or more of a variety of edits, including, but not limited to, cuts, color adjustments, highlight adjustments, filter application, Gaussian blur, or motion blur.

[0029] 4 illustrates an example method 400 for segmenting a current frame of a sequence of frames to generate a segmentation result. Any one of the blocks illustrated in example method 400 can be repeated, combined, rearranged, or linked with the set of blocks of example method 400 or example method 300. For example, example method 400 can utilize the motion vectors received at 306 of example method 300.

[0030] At 402, the video processing manager quantizes the motion vectors for the current frame into two or more bins, which may contain one or more motion vectors per bin. The bins group similar motion vectors together, which may be used in step 404.

[0031] At 404, the video processing manager calculates an average motion vector for one of the two or more bins that contains the majority of the motion vectors. For example, see the exemplary environment 100 of FIG. 1. Because the user 112 panned the image capture device 102 to keep the contestant 114 centered in the foreground of the video segment, the motion vectors may be grouped into two bins. The first bin contains motion vectors for the background, and the second bin contains motion vectors for the foreground. Furthermore, because the contestant 114 occupies less space in each frame of the video segment, the background motion vector bin may be the bin that contains the majority of the motion vectors. Thus, the video processing manager may calculate an average motion vector for the background bin.

[0032] At 406, the video processing manager compares the motion vectors of two or more bins with the average motion vector to generate a comparison result. In this example, the average motion vector for the background bin is greater than any of the motion vectors in the foreground bin. This difference is due to the user 112 panning the image capture device 102 to keep the athlete 114 centered in the foreground. The athlete 114 does not move relative to the image capture device 102. In other words, the motion vector of the foreground is close to zero. Unlike the athlete 114 in the foreground, the tree 116 in the background moves relative to the image capture device 102. In other words, the motion vector of the background is greater than zero. In this example, the comparison result may indicate that the motion vector of the foreground is less than the average motion vector of the background motion vectors.

[0033] At 408, the video processing manager classifies one or more of the motion vectors as outliers based on the comparison results exceeding a threshold. The threshold may be an integer, a fraction, a percentage, a difference relative to another value (e.g., the average motion vector), or other quantifier to which the motion vectors may be compared. In this example, the video processing manager may classify foreground motion vectors as outliers based on them being smaller than the average motion vector by a difference (e.g., 10 percent, 15 percent). As an additional example, the video processing manager may classify a motion vector as an outlier if it is larger than the average motion vector by a difference (e.g., 10 percent, 25 percent). As a further example, the video processing manager may classify a motion vector as an outlier if it is close to zero, close to infinity, close to another integer, or close to another standalone value.

[0034] At 410, the video processing manager segments the current frame based on the outliers to generate a segmentation result for the current frame. Continuing with the present example, the segmentation result may include two segments, one for the foreground and one for the background. Based on the foreground and background segments, the video processing manager may edit the background separately from the foreground, edit the foreground separately from the background, or a combination of both. Further, when combined with exemplary method 300, the video processing manager may generate a final mask by combining the segmentation result with a predicted mask.

[0035] In some aspects, the video processing manager may utilize distance information from additional sensors (e.g., sensor 214) of the image capture device. For example, the video processing manager may utilize distance information from a proximity sensor (e.g., sonar, radar, lidar) to more quickly or accurately identify the foreground or background of a previous or current frame. The distance information may include a distance measurement to the foreground (e.g., 6 m, 15 m) and a distance measurement to the background (e.g., 25 m, 31 m). The video processing manager may segment the foreground of the current frame from the background of the current frame based on the distance information from the proximity sensor. As another example, the video processing manager may utilize a second camera having a different perspective from the first camera to identify the foreground or background of the current frame. The background of the current frame may appear similar from the perspective of the first camera and the perspective of the second camera. The foreground of the current frame may appear different from the perspective of the first camera and the perspective of the second camera. The video processing manager may segment the foreground from the background of the current frame based on differences in the foreground or similarities in the background from different viewpoints.

[0036] Throughout this discussion, examples of a video processing manager editing the background of one frame of a sequence of frames of a video segment are provided. However, the systems and techniques described herein are not limited to editing the background of a frame. In aspects, the systems and techniques may also be implemented by a video processing manager to edit the foreground of a frame. Additionally or alternatively, the systems and techniques described herein may be implemented by a video processing manager in a long-exposure photography application. For example, assume a user wants to take a photo of an object in low light. To do so, the user creates frames of the shot of the object using an image capture device having a video processing manager configured to remove distortion from real-time video using masked frames. The video processing manager may capture multiple frames of the object in low light using a long exposure time. The long exposure time provides enough time for sufficient light to be captured by the image sensor of the image capture device for each frame. If the user's hands shake when capturing multiple frames with a long exposure time, the object may be blurred. However, the video processing manager may implement the techniques and systems described herein to segment the background from the foreground of multiple frames. The video processing manager may use the motion vectors, for example, to perform segmentation. The video processing manager may also use the motion vectors to stabilize the foreground of long-exposure frames in real time (e.g., by aligning a foreground mask with the motion vectors to the current frame), resulting in a clear foreground. Multiple frames may be merged (e.g., overlaid) into a single output photo with a clear foreground, for example.

[0037] Further Examples Further examples are provided in the following sections.

[0038] receiving a video segment, the video segment comprising a sequence of frames, the sequence of frames comprising a previous frame and a current frame, the current frame being sequenced immediately after the previous frame; further comprising receiving a subject mask for the current frame, the subject mask being generated using a machine learning (ML) model; further comprising receiving a motion vector for the current frame, the motion vector being generated by an optical flow measurement tool using the previous frame and the current frame; further comprising receiving a predicted mask for the current frame, the predicted mask being generated from the motion vector and the previous frame;

[0039] Example 2: The method of example 1, wherein the video segments are captured by and received from a camera of an image capture device, the ML model is on the image capture device, and the optical flow measurement tool is on the image capture device.

[0040] Example 3: The method of Example 1, further comprising: quantizing the motion vectors for the current frame into two or more bins; calculating an average motion vector for one of the two or more bins that includes a majority of the motion vectors; comparing the motion vectors of the two or more bins with the average motion vector to generate a comparison result; classifying one or more of the motion vectors as outliers based on the comparison result exceeding a threshold; and segmenting the current frame based on the outliers to generate a segmentation result for the current frame.

[0041] Example 4: The method of example 3, wherein the final mask is generated by combining the segmentation result of the current frame with the predicted mask.

[0042] Example 5: The method of example 1, wherein the predicted mask is generated by aligning a final mask of the previous frame to the current frame using the motion vector.

[0043] Example 6: The method of example 1, further comprising performing a sharpening process on the final mask before applying the final mask to the current frame.

[0044] Example 7: The method of example 6, wherein the sharpening process is performed by an edge sharpening tool, the edge sharpening tool being on an image capture device from which the video segment is received.

[0045] Example 8: The method of example 1, wherein the previous frame and the current frame include a background, a foreground in front of the background, and a subject of interest in the foreground.

[0046] Example 9: The method of Example 8, wherein editing the masked frame to generate the output frame further comprises editing the background of the masked frame.

[0047] Example 10: The method of Example 8, wherein editing the masked frame to generate the output frame further comprises editing the foreground of the masked frame.

[0048] Example 11: The method of Example 8, wherein editing the masked frame to generate the output frame further comprises editing the foreground and the background of the masked frame.

[0049] Example 12: The method of Example 8, further comprising receiving distance information for the current frame, the distance information being captured by a sensor of an image capture device, and further comprising segmenting the foreground of the current frame from the background of the current frame based on the distance information.

[0050] Example 13: The method of Example 8, further comprising receiving viewpoint information for the current frame, the viewpoint information being captured by a different image capture device, and further comprising segmenting the foreground of the current frame from the background of the current frame based on the viewpoint information.

[0051] Example 14: An image capture device comprising at least one camera, one or more sensors, one or more processors, and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to execute a video processing manager to provide video processing utilizing the at least one camera and the one or more processors by performing a method according to any one of the preceding claims.

[0052] Example 15: A computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 13.

[0053] conclusion Unless the context dictates otherwise, the use of the word "or" herein may be considered an "inclusive or" or the use of a term permitting the inclusion or application of one or more items linked by the word "or" (e.g., the phrase "A or B" may be interpreted as permitting only "A," permitting only "B," or permitting both "A" and "B"). Also, as used herein, a phrase referring to "at least one" of a list of items refers to any combination of those items, including single members. For example, "at least one of a, b, c" can cover not only a, b, c, ab, ac, bc, abc, but also any combination containing multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, ccc, or any other order of a, b, c). Furthermore, items depicted in the accompanying drawings and terms described herein may refer to one or more items or terms, and therefore, the singular or plural forms of items and terms in this description may be referred to interchangeably.

[0054] Although embodiments of systems and techniques for removing distortion from real-time video using masked frames, and embodiments of apparatus enabling such removal, have been described in language specific to particular features and / or methods, the subject matter of the appended claims is not necessarily limited to the particular features or methods described. Rather, the particular features and methods are disclosed as exemplary embodiments for removing distortion from real-time video using masked frames.

Claims

1. 1. A method comprising: receiving a video segment, the video segment including a sequence of frames, the sequence of frames including a previous frame and a current frame, the current frame being sequenced immediately after the previous frame, the method further comprising: receiving an object mask for the current frame, the object mask being generated using a machine learning (ML) model, the method further comprising: receiving a motion vector for the current frame, the motion vector being generated by an optical flow measurement tool using the previous frame and the current frame, the method further comprising: receiving a predicted mask for the current frame, the predicted mask being generated from the motion vector and the previous frame, the method further comprising: generating a final mask for the current frame, the final mask based on the object mask, the motion vectors, and the predicted mask, the method further comprising: applying the final mask to the current frame to provide a masked frame; editing the masked frame to remove distortion from the masked frame and generate an output frame; outputting the output frame; A method comprising:

2. the video segments are captured by and received from a camera of an image capture device; the ML model is stored on a computer readable medium (CRM) of the image capture device; The method of claim 1 , wherein the optical flow measurement tool is stored in the CRM of the image capture device.

3. quantizing the motion vectors for the current frame into two or more bins; calculating an average motion vector for one of the two or more bins that contains a majority of the motion vectors; comparing the motion vectors of the two or more bins with the average motion vector to generate a comparison result; classifying one or more of the motion vectors as outliers based on the comparison result exceeding a threshold; segmenting the current frame based on the outliers to generate a segmentation result for the current frame; The method of claim 1 further comprising:

4. The method of claim 3 , wherein the final mask is generated by combining the segmentation result of the current frame with the predicted mask.

5. The method of claim 1 , wherein the predicted mask is generated by aligning a final mask of the previous frame to the current frame using the motion vectors.

6. The method of claim 1 , further comprising: performing a sharpening process on the final mask before applying the final mask to the current frame.

7. The sharpening process is performed by an edge sharpening tool; The method of claim 6 , wherein the edge sharpening tool is on an image capture device from which the video segments are received.

8. The method of claim 1 , wherein the previous frame and the current frame include a background, a foreground in front of the background, and a subject of interest in the foreground.

9. Editing the masked frames to generate the output frames includes: The method of claim 8 , further comprising editing the background of the masked frame.

10. Editing the masked frames to generate the output frames includes: The method of claim 8 , further comprising editing the foreground of the masked frame.

11. Editing the masked frames to generate the output frames includes: The method of claim 8 , further comprising editing the foreground and the background of the masked frame.

12. further comprising receiving distance information for the current frame, the distance information being captured by a sensor of an image capture device; The method of claim 8 , further comprising segmenting the foreground of the current frame from the background of the current frame based on the distance information.

13. further comprising receiving viewpoint information for the current frame, the viewpoint information being captured by a different image capture device; The method of claim 8 , further comprising segmenting the foreground of the current frame from the background of the current frame based on the viewpoint information.

14. 1. An image capture device, comprising: at least one camera; one or more sensors; one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to execute a video processing manager to provide video processing utilizing the at least one camera, the one or more sensors, and the one or more processors by performing a method according to any one of the preceding claims. Image capture device.

15. A computer readable medium (CRM) comprising instructions that, when executed by one or more processors, cause said one or more processors to perform the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Image processing method, apparatus and program thereof

    JP2006050070A

  • Depth-based video background subtraction

    JP2022528294A

  • Methods and apparatus for applying motion blur to overcaptured content

    US10997697B1

  • Segmentation for image effects

    US11276177B1