Video Matting via Sparse Low-Rank Representation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video matting methods fail to ensure temporal and spatial consistency of extracted foreground objects due to poor representative ability and inconsistency in alpha matte values, particularly when using sparse representation and non-local priors.

Innovation Solution

A method for video matting via sparse and low-rank representation, which involves determining known and unknown pixels, training dictionaries with foreground and background samples, obtaining reconstruction coefficients, setting non-local and Laplace matrices, and extracting foreground objects using these matrices to ensure temporal and spatial consistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If sparse representation is used for image matting, then the method can automatically select appropriate sample points to reconstruct the original image, but it fails to guarantee temporal and spatial consistency of video alpha matte

Engineering Contradiction:
Improveautomatic sample point selectionVSAvoidtemporal and spatial consistency
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the video processing into keyframe processing and non-keyframe processing. Keyframes undergo dictionary training and sparse representation, while non-keyframes use the trained dictionary for faster processing. This segmentation maintains consistency across frames while enabling automatic sample point selection where needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs dictionary training on keyframes in advance (preliminary action) before processing all video frames. This pre-trained dictionary ensures temporal and spatial consistency is established beforehand, allowing subsequent frames to maintain consistency without repeating the full training process.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If only foreground pixels are used as dictionary for sparse representation, then the processing is simpler, but the representative ability is poor leading to poor extraction quality

Engineering Contradiction:
Improvedictionary construction simplicityVSAvoidforeground object extraction quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent creates a composite dictionary structure combining foreground pixels and background pixels. This composite dictionary enhances representative ability by incorporating both foreground and background information, leading to improved extraction quality while maintaining reasonable processing complexity.

Inventive Principle:
Principle #40Composite materials

3Ease of operation

If a fixed number of sample points are selected for each pixel, then the processing is straightforward, but selecting fewer points misses good samples while selecting more points introduces noise

Engineering Contradiction:
Improvesample point selection processVSAvoidreconstruction quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent dynamically adjusts the number of sample points selected for each pixel based on local characteristics and reconstruction needs. Rather than using a fixed number, the method adapts the sample count to optimize reconstruction quality while avoiding noise introduction, balancing operational simplicity with precision.

Inventive Principle:
Principle #15Dynamics

4Manufacturing precision

If non-local structure is constructed for video alpha matte, then extraction quality is improved, but it is difficult to construct consistent non-local structure for pixels with similar characteristics resulting in temporal and spatial inconsistency

Engineering Contradiction:
Improveforeground object extraction qualityVSAvoidtemporal and spatial consistency
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent introduces a Laplace matrix as an intermediary to enforce consistency in the non-local structure. The Laplace matrix acts as a mediator that ensures pixels with similar characteristics maintain consistent alpha values across space and time, resolving the inconsistency problem while preserving the quality improvements from non-local processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10235571B2Method for video matting via sparse and low-rank representation
Publication Date: 2019.03.19 BEIHANG UNIV
  • US10235571B2 patent drawing
  • US10235571B2 patent drawing

AI summary

The present invention provides a method for video matting via sparse and low-rank representation, which firstly selects frames which represent video characteristics in input video as keyframes, then trains a dictionary according to known pixels in the keyframes, next obtains a reconstruction coefficient satisfying the restriction of low-rank, sparse and non-negative according to the dictionary, and sets the non-local relationship matrix between each pixel in the input video according to the reconstruction coefficient, meanwhile sets the Laplace matrix between multiple frames, obtains a video alpha matte of the input video, according to α values of the known pixels of the input video and α values of sample points in the dictionary, the non-local relationship matrix and the Laplace matrix; and finally extracts a foreground object in the input video according to the video alpha matte, therefore improving quality of the extracted foreground object.