2D Video Background Extraction for 3D Virtual Conference Avatars
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video communication systems struggle to effectively integrate users from video communications platforms into virtual environments, particularly in terms of seamlessly presenting user representations without backgrounds.
Innovation Solution
The system employs a video extraction module to determine the boundary between a user and their background in a video stream, allowing for the extraction and processing of the user representation without the background, which can then be rendered in a virtual environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If video streams with backgrounds are used in virtual environments, then users can maintain their original appearance and context, but the virtual environment becomes cluttered and less immersive
Solution Approach 1:
The patent extracts the user from the video stream by determining a boundary between the user and background, then provides only the user representation to the virtual environment. This separation removes the harmful background interference while preserving the user's visual identity, directly resolving the contradiction between maintaining user appearance and eliminating background clutter.
Solution Approach 2:
The system introduces an intermediary processing layer that includes boundary determination and user extraction modules. This intermediary layer acts as a mediator between the original video stream and the virtual environment, filtering out background elements while preserving user information, thus enabling clean integration without direct background interference.
2Device complexity
If 2D video representations are used in 3D virtual environments, then integration is simpler, but the immersion and spatial presence are reduced
Solution Approach 1:
The patent transforms 2D video representations into 3D volumetric representations by processing extracted user data through a volumetric reconstruction module. This dimensional transformation enables the user to be rendered in three-dimensional space within the virtual environment, significantly enhancing immersion and spatial presence while maintaining manageable integration complexity through automated processing.
3Ease of operation
If automated user extraction is implemented, then background removal is achieved without manual intervention, but processing time and computational resources increase
Solution Approach 1:
The system implements self-service automation where the boundary determination module automatically analyzes video frames and extracts user representations without requiring manual annotation or intervention. The process uses automated algorithms that continuously process video streams in real-time, eliminating the need for manual operations while managing processing time through efficient computational methods.
Data Source
AI summary
A virtual environment that is three-dimensional and that includes digital representations of video conference participants is provided in a video conference session. A representation of a first participant is provided in the virtual environment as an augmented or virtual reality (AR/VR) participant in three dimensions. A two-dimensional video stream of a second participant who is not in AR/VR is received. A boundary around the second participant within the two-dimensional video stream is defined to separate an interior depiction of the second participant from an exterior background. The interior depiction of the second participant and the representation of the first participant are displayed within the virtual environment.


