2D Video Background Extraction for 3D Virtual Conference Avatars

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video communication systems struggle to effectively integrate users from video communications platforms into virtual environments, particularly in terms of seamlessly presenting user representations without backgrounds.

Innovation Solution

The system employs a video extraction module to determine the boundary between a user and their background in a video stream, allowing for the extraction and processing of the user representation without the background, which can then be rendered in a virtual environment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If video streams with backgrounds are used in virtual environments, then users can maintain their original appearance and context, but the virtual environment becomes cluttered and less immersive

Engineering Contradiction:
Improveuser representation fidelityVSAvoidbackground interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts the user from the video stream by determining a boundary between the user and background, then provides only the user representation to the virtual environment. This separation removes the harmful background interference while preserving the user's visual identity, directly resolving the contradiction between maintaining user appearance and eliminating background clutter.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system introduces an intermediary processing layer that includes boundary determination and user extraction modules. This intermediary layer acts as a mediator between the original video stream and the virtual environment, filtering out background elements while preserving user information, thus enabling clean integration without direct background interference.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If 2D video representations are used in 3D virtual environments, then integration is simpler, but the immersion and spatial presence are reduced

Engineering Contradiction:
Improverepresentation integration complexityVSAvoidvirtual environment immersion
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent transforms 2D video representations into 3D volumetric representations by processing extracted user data through a volumetric reconstruction module. This dimensional transformation enables the user to be rendered in three-dimensional space within the virtual environment, significantly enhancing immersion and spatial presence while maintaining manageable integration complexity through automated processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of operation

If automated user extraction is implemented, then background removal is achieved without manual intervention, but processing time and computational resources increase

Engineering Contradiction:
Improveautomatic background removalVSAvoidvideo processing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system implements self-service automation where the boundary determination module automatically analyzes video frames and extracts user representations without requiring manual annotation or intervention. The process uses automated algorithms that continuously process video streams in real-time, eliminating the need for manual operations while managing processing time through efficient computational methods.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250071159A1Integrating 2D And 3D Participant Representations In A Virtual Video Conference Environment
Publication Date: 2025.02.27 ZOOM COMMUNICATIONS INC
  • US20250071159A1 patent drawing
  • US20250071159A1 patent drawing
  • US20250071159A1 patent drawing

AI summary

A virtual environment that is three-dimensional and that includes digital representations of video conference participants is provided in a video conference session. A representation of a first participant is provided in the virtual environment as an augmented or virtual reality (AR/VR) participant in three dimensions. A two-dimensional video stream of a second participant who is not in AR/VR is received. A boundary around the second participant within the two-dimensional video stream is defined to separate an interior depiction of the second participant from an exterior background. The interior depiction of the second participant and the representation of the first participant are displayed within the virtual environment.