Merged Reality Scene Streaming via Multi-View Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional media player devices face limitations in providing immersive merged reality scenes due to the need for preloading data and the inability to scale with increasing complexity, especially for real-time experiences like live events, as they struggle to process and stream detailed 3D models of multiple interacting virtual and real-world objects.
Innovation Solution
A system and method for generating a merged reality scene by capturing and streaming color and depth video data from multiple vantage points, allowing for the dynamic rendering of virtual and real-world objects without preloading 3D models, using a network of 3D capture devices and server-side rendering engines to create a transport stream that includes entity description data for virtual and real-world objects, enabling real-time experience of complex scenes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data representative of merged reality scene is preloaded prior to user experience, then users can experience merged reality scenes with maximum flexibility, but it precludes real-time experiences of live events and places significant limitations on certain types of experiences
Solution Approach 1:
The system performs preliminary actions by capturing and processing video data from multiple vantage points in advance, creating a foundation of preprocessed visual information that can be rapidly assembled and streamed in real-time without requiring full scene preloading
Solution Approach 2:
The merged reality scene is segmented into multiple independent video data streams from different vantage points, allowing each segment to be processed and streamed independently, enabling real-time assembly without preloading the entire scene
2Quantity of substance
If 3D models for multiple objects are streamed to media player device, then larger and more detailed merged reality scenes can be presented, but provider system processing burdens cannot scale
Solution Approach 1:
Instead of streaming complex 3D models that require heavy processing, the system creates simplified video data copies from multiple camera vantage points that can be rendered directly, significantly reducing the computational burden on both provider and client systems while maintaining visual fidelity
Solution Approach 2:
The system replaces the mechanical process of generating and streaming complex 3D models with a video-based approach where pre-captured video frames are assembled and rendered, substituting computationally intensive 3D rendering with more efficient video processing and composition
Data Source
Figure 1
Figure 2
Figure 3A~3C
AI summary
An exemplary merged reality scene capture system (system) receives a first frame set of surface data frames from a plurality of three-dimensional (3D) capture devices disposed with respect to a real-world scene so as to have a plurality of different vantage points of the real-world scene. Based on the first frame set, the system generates a transport stream that includes color and depth video data streams for each of the 3D capture devices. Based on the transport stream, the system generates entity description data representative of a plurality of entities included within a 3D space of a merged reality scene. The plurality of entities includes a virtual object, the real-world object, and virtual viewpoints into the 3D space from which a second frame set of surface data frames are to be rendered representing color and depth data for both the virtual and the real-world objects.