Video Matting via Sparse Low-Rank Representation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video matting methods fail to ensure temporal and spatial consistency of extracted foreground objects due to poor representative ability and inconsistency in alpha matte values, particularly when using sparse representation and non-local priors.
Innovation Solution
A method for video matting via sparse and low-rank representation, which involves determining known and unknown pixels, training dictionaries with foreground and background samples, obtaining reconstruction coefficients, setting non-local and Laplace matrices, and extracting foreground objects using these matrices to ensure temporal and spatial consistency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If sparse representation is used for image matting, then the method can automatically select appropriate sample points to reconstruct the original image, but it fails to guarantee temporal and spatial consistency of video alpha matte
Solution Approach 1:
The patent segments the video processing into keyframe processing and non-keyframe processing. Keyframes undergo dictionary training and sparse representation, while non-keyframes use the trained dictionary for faster processing. This segmentation maintains consistency across frames while enabling automatic sample point selection where needed.
Solution Approach 2:
The patent performs dictionary training on keyframes in advance (preliminary action) before processing all video frames. This pre-trained dictionary ensures temporal and spatial consistency is established beforehand, allowing subsequent frames to maintain consistency without repeating the full training process.
2Device complexity
If only foreground pixels are used as dictionary for sparse representation, then the processing is simpler, but the representative ability is poor leading to poor extraction quality
Solution Approach 1:
The patent creates a composite dictionary structure combining foreground pixels and background pixels. This composite dictionary enhances representative ability by incorporating both foreground and background information, leading to improved extraction quality while maintaining reasonable processing complexity.
3Ease of operation
If a fixed number of sample points are selected for each pixel, then the processing is straightforward, but selecting fewer points misses good samples while selecting more points introduces noise
Solution Approach 1:
The patent dynamically adjusts the number of sample points selected for each pixel based on local characteristics and reconstruction needs. Rather than using a fixed number, the method adapts the sample count to optimize reconstruction quality while avoiding noise introduction, balancing operational simplicity with precision.
4Manufacturing precision
If non-local structure is constructed for video alpha matte, then extraction quality is improved, but it is difficult to construct consistent non-local structure for pixels with similar characteristics resulting in temporal and spatial inconsistency
Solution Approach 1:
The patent introduces a Laplace matrix as an intermediary to enforce consistency in the non-local structure. The Laplace matrix acts as a mediator that ensures pixels with similar characteristics maintain consistent alpha values across space and time, resolving the inconsistency problem while preserving the quality improvements from non-local processing.
Data Source
AI summary
The present invention provides a method for video matting via sparse and low-rank representation, which firstly selects frames which represent video characteristics in input video as keyframes, then trains a dictionary according to known pixels in the keyframes, next obtains a reconstruction coefficient satisfying the restriction of low-rank, sparse and non-negative according to the dictionary, and sets the non-local relationship matrix between each pixel in the input video according to the reconstruction coefficient, meanwhile sets the Laplace matrix between multiple frames, obtains a video alpha matte of the input video, according to α values of the known pixels of the input video and α values of sample points in the dictionary, the non-local relationship matrix and the Laplace matrix; and finally extracts a foreground object in the input video according to the video alpha matte, therefore improving quality of the extracted foreground object.

