2D Video Depth Simulation Using Face-Pose Parallax

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video-conferencing systems fail to provide a realistic sense of depth in two-dimensional (2D) video streams, limiting the immersive experience for participants.

Innovation Solution

Implementing a parallax effect by removing the background from a 2D video of a remote speaker and combining it with a background image, adjusting the orientation based on the viewer's face pose using feature detection, either through an out-of-band channel or an in-band channel.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a conventional 2D video stream is used in video-conferencing systems, then the system complexity is low and transmission is simple, but the sense of depth and immersion is insufficient

Engineering Contradiction:
Improvesense of depthVSAvoidvideo processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The video is segmented into a foreground subject (person) and background, with the background further divided into multiple layers at different depths. This segmentation allows the system to process and display depth information selectively, creating a three-dimensional effect while maintaining manageable complexity in the video processing pipeline.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from a two-dimensional video representation to a three-dimensional representation by adding depth information through multiple background layers at different distances. This dimensional change enables viewers to perceive depth and spatial relationships without requiring complex 3D video capture and transmission infrastructure.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If multiple background layers at different depths are added to create depth perception, then the sense of immersion is improved, but the video transmission bandwidth and processing load increase

Engineering Contradiction:
Improveimmersive experienceVSAvoidvideo data volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The background is segmented into multiple discrete layers, each representing a different depth plane. This segmentation allows the system to transmit only the necessary background information for each layer rather than entire high-resolution 3D scenes, reducing overall data volume while maintaining immersive effect.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different background layers are processed and transmitted with appropriate quality levels based on their importance and computational cost. The system can allocate bandwidth and processing resources dynamically, focusing higher quality on layers that contribute most to the immersive experience while using lower quality for less critical layers.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250247502A1Simulating Depth In A Two-Dimensional Video Using Feature Detection And Parallax Effect With Multilayer Video And An In-Band Channel
Publication Date: 2025.07.31 ZOOM COMMUNICATIONS INC
  • US20250247502A1 patent drawing
  • US20250247502A1 patent drawing
  • US20250247502A1 patent drawing

AI summary

A video-conferencing system that simulates depth in a two-dimensional video of a remote speaker via a parallax effect. The background of the video of the remote speaker is removed and the resulting backgroundless video is combined with a background image according to poses of a viewing participant face captured by a camera. As the poses change, the orientation of the backgroundless video and the background image are changed proportionally to yield a parallax effect. The backgroundless video and the background image are combined at the client device of the remote speaker and is transferred to the client device of the viewing participant as a multilayer video.