Immersive Remote Conferencing via Depth-Map 3D Scene Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current videoconferencing technologies fail to effectively convey non-verbal social signals such as accurate eye gaze and gesture direction, leading to an unnatural experience and loss of valuable information.
Innovation Solution
Processing depth and video data to place remote conference participants into a common scene, allowing each user to choose a virtual environment, with head tracking for motion parallax compensation and spatial audio adjustments, providing a realistic and immersive experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If desktop video conferencing is used, then cost is reduced and accessibility is improved, but the ability to convey non-verbal social signals (eye gaze, gesture direction) deteriorates
Solution Approach 1:
The patent creates virtual copies of remote participants using depth maps and video data, placing them into a common 3D scene that preserves their spatial relationships and non-verbal cues. This copying approach allows standard video conferencing equipment to deliver immersive experiences with accurate eye gaze and gesture direction information.
Solution Approach 2:
The patent transitions from 2D video feeds to 3D spatial representations by generating depth maps and placing participants into a common scene with proper spatial positioning. This dimensional enhancement preserves non-verbal social signals while maintaining compatibility with standard video conferencing infrastructure.
2Loss of information
If high-end room conferencing systems are used, then the ability to convey non-verbal social signals is improved, but size and cost increase making usage very limited
Solution Approach 1:
The patent replaces complex mechanical room conferencing systems with software-based processing of standard video and depth data. By using computational methods to generate 3D scenes and preserve non-verbal cues, the system achieves high-end functionality without requiring expensive specialized hardware.
Solution Approach 2:
The patent creates virtual representations of participants using depth maps and video feeds from standard cameras, eliminating the need for expensive specialized conferencing equipment. This copying approach delivers high-quality non-verbal signal preservation using accessible technology.
3Reliability
If photo-realistic representations with head tracking are used, then immersion and realism are improved, but processing complexity and computational requirements increase
Solution Approach 1:
The patent performs preliminary processing of video and depth data to create optimized 3D scene representations before rendering. By pre-processing spatial relationships and participant positions, the system reduces real-time computational requirements while maintaining photo-realistic quality and immersive experience.
Data Source
AI summary
The subject disclosure is directed towards an immersive conference, in which participants in separate locations are brought together into a common virtual environment (scene), such that they appear to each other to be in a common space, with geometry, appearance, and real-time natural interaction (e.g., gestures) preserved. In one aspect, depth data and video data are processed to place remote participants in the common scene from the first person point of view of a local participant. Sound data may be spatially controlled, and parallax computed to provide a realistic experience. The scene may be augmented with various data, videos and other effects/animations.


