Guidance Map-Free Video Matting for XR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video matting techniques for extended reality (XR) platforms require auxiliary guidance maps, increasing computational costs and limiting high-resolution, real-time alpha matte generation, especially on power-constrained devices.

Innovation Solution

The use of trained machine learning (ML) models that operate on low-resolution inputs and auxiliary signals to generate high-resolution alpha mattes, allowing for low-latency, domain-specific, and guidance map-free video matting, which reduces computational costs and enhances efficiency without compromising visual quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current video matting techniques use auxiliary guidance maps, then alpha matte generation accuracy is improved, but computational cost and device complexity increase

Engineering Contradiction:
Improvealpha matte generation accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the requirement for auxiliary guidance maps from the video matting process. By training domain-specific ML models on targeted data, the system achieves accurate alpha matte generation without needing external guidance inputs, thereby reducing computational overhead and device complexity while maintaining precision.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the operational parameters of the matting system by transitioning from guidance-map-dependent algorithms to ML models that operate directly on input video frames. This parameter change enables the system to achieve high accuracy through learned domain-specific features rather than through complex auxiliary processing.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If high-resolution alpha mattes are generated in real-time, then video matting speed is improved, but power consumption and thermal output increase

Engineering Contradiction:
Improvevideo matting speedVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary action by training domain-specific ML models in advance on curated datasets. This pre-training enables the models to achieve high-resolution, real-time alpha matte generation with optimized computational efficiency during deployment, reducing power consumption and thermal output on power-constrained devices while maintaining high productivity.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If domain-specific ML models are used, then alpha matte generation accuracy is improved, but model training complexity increases

Engineering Contradiction:
Improvealpha matte generation accuracyVSAvoidmodel training complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the ML model training into domain-specific components. Instead of training general-purpose models, the system trains specialized models for specific video matting domains using targeted datasets. This segmentation approach improves generation accuracy while managing training complexity through focused, domain-specific training rather than comprehensive general training.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240104686A1Low-Latency Video Matting
Publication Date: 2024.03.28 APPLE INC
  • US20240104686A1 patent drawing
  • US20240104686A1 patent drawing
  • US20240104686A1 patent drawing

AI summary

Techniques are disclosed herein for implementing a novel, low latency, guidance map-free video matting system, e.g., for use in extended reality (XR) platforms. The techniques may be designed to work with low resolution auxiliary inputs (e.g., binary segmentation masks) and to generate alpha mattes (e.g., alpha mattes configured to segment out any object(s) of interest, such as human hands, from a captured image) in near real-time and in a computationally efficient manner. Further, in a domain-specific setting, the system can function on a captured image stream alone, i.e., it would not require any auxiliary inputs, thereby reducing computational costs—without compromising on visual quality and user comfort. Once an alpha matte has been generated, various alpha-aware graphical processing operations may be performed on the captured images according to the generated alpha mattes (e.g., background replacement operations, synthetic shallow depth of field (SDOF) rendering operations, and/or various XR environment rendering operations).