Live Video Avatar Extraction Using Depth Map and Spatial Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing systems for virtual and augmented reality environments face challenges in achieving high image quality, particularly at the edges of live video avatars, due to optical effects and measurement inaccuracies, making it difficult to transplant images effectively across different backgrounds without requiring monochromatic screens.

Innovation Solution

The system employs a depth sensor to create a depth map of a live video avatar in a heterogeneous environment, combining it with spatial filtering to enhance edge detection and image quality, allowing for the extraction and transplantation of avatars across diverse backgrounds, while minimizing data processing and storage burdens by using edge data from depth sensing to set a mask for subsequent spatial filtering operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If spatial filtering is used to extract live video avatars from heterogeneous backgrounds, then image transplantation across diverse environments is enabled, but edge detection quality deteriorates due to optical effects and measurement inaccuracies

Engineering Contradiction:
Improveimage transplantation capabilityVSAvoidedge detection quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent combines depth sensing technology with spatial filtering to create a hybrid extraction system. The depth sensor provides accurate edge detection data that compensates for the weaknesses of spatial filtering alone, while the spatial filtering handles the background removal. This merging of two different technologies resolves the contradiction by maintaining both adaptability for diverse backgrounds and precision for edge quality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The depth map serves as an intermediary element that bridges the gap between the live video feed and the final extracted avatar. The depth sensor creates a depth map that acts as a mediator, providing accurate edge information that guides the spatial filtering process and corrects edge artifacts, thereby improving edge detection quality without limiting background versatility.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If depth sensing is used to improve edge detection performance, then image quality at edges is improved, but computational load and data processing requirements increase

Engineering Contradiction:
Improveedge detection performanceVSAvoidcomputational load
Core Design Contradiction:
Manufacturing precisionVSPower

Solution Approach 1:

The system applies depth sensing selectively to edge regions rather than processing the entire image at high resolution. By focusing computational resources on the critical edge areas where precision is most needed, the system improves edge detection performance without proportionally increasing the overall computational load. This local quality approach optimizes the power-to-precision ratio.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses depth sensing to provide slightly more data than strictly necessary for basic extraction, creating a depth map with higher resolution or detail than the minimum required. This excessive action in data collection allows for more accurate edge detection and better correction of edge artifacts, while the processing is optimized to use only the necessary portion of this data, balancing computational load with quality improvement.

Inventive Principle:
Principle #16Partial or excessive action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach improves edge detection performance and image quality for live video avatars, enabling seamless transplantation across environments without the need for homogenous backgrounds, providing enhanced visual fidelity and reduced computational load.

Implementation Method 1

a depth sensor for creating a depth map based first live video avatar

Methodology Applied
Scientific EffectTime of Flight: Time of Flight

Data Source

PatentUS11218669B1System and method for extracting and transplanting live video avatar images
Publication Date: 2022.01.04 BENMAN WILLIAM J
  • US11218669B1 patent drawing
  • US11218669B1 patent drawing
  • US11218669B1 patent drawing

AI summary

A system for extracting and transplanting live video avatar images including a depth sensor for creating a depth map based first live video avatar of a user or object disposed in a heterogeneous first environment with an arbitrary background; a processor coupled to the depth sensor; code fixed in a tangible medium for execution by the processor for extracting the depth map from the first environment to provide an extracted depth map based live video avatar; and a display system coupled to the processor for showing the extracted depth map based live video avatar in a second environment diverse from the first environment. In a second embodiment, the system includes a camera coupled to the processor to provide live video images of the user in the first environment and code for spatially filtering the images to provide a spatially filtered extracted second live video avatar. This embodiment further includes code for combining the first live video avatar with the second live video avatar to provide an enhanced extracted depth map based third live video avatar. Images from multiple cameras and or depth sensors are combined simultaneously to provide the third live video avatar using the spatially enhanced extracted depth map. A routing server is included for receiving streams from multiple users and sending to each user the live video avatar images from other users based on their locations in a shared space or for use in a local user's AR environment.