Encoding Free Viewpoint Data in Movie Containers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating real-time, immersive augmented reality experiences with virtual objects in real-world scenes are hindered by the difficulty in preprocessing and rendering 3D items within real-time video streams, particularly due to limitations in available computing resources and the need for precise coordinate translation between real-world and virtual scenes.
Innovation Solution
A method of encoding arbitrary data, such as free-view data defining 3D models, into a standard movie data container like an MPEG container, allowing for the insertion of non-video data into video streams with predefined labels, enabling efficient processing and rendering of 3D models with varying frame rates and resolutions within the same video stream or separate streams.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If non-video data (free-view data) is inserted into video streams for efficient processing, then the portability and wearability of AR systems is improved, but the complexity of data container structure increases
Solution Approach 1:
The patent combines non-video free-view data with video streams by inserting them into the same movie data container (MP4 format). This merging approach allows both video and 3D model data to be processed together in the video stream, enabling efficient real-time rendering on portable devices with limited resources.
Solution Approach 2:
The patent uses a universal video stream format (MP4 container) to carry multiple types of data - both traditional video frames and non-video free-view 3D model data. This multi-functional approach allows a single data container to serve multiple purposes, reducing the need for separate processing pipelines and improving system portability.
2Productivity
If free-view data is encoded into standard movie containers for efficient processing, then real-time rendering performance is improved, but the precision of coordinate translation between real-world and virtual scenes may be compromised
Solution Approach 1:
The patent introduces a coordinate translation mechanism that acts as an intermediary between the real-world camera coordinates and the virtual 3D model coordinates. This mediator ensures accurate spatial alignment while allowing the data to be processed through the efficient video rendering pipeline, maintaining both precision and performance.
Solution Approach 2:
The patent transforms free-view 3D model data into a format compatible with video stream encoding by changing its parameter representation. The 3D models are encoded with metadata including position, orientation, and scale parameters that enable accurate coordinate translation while being processed through standard video decoding pipelines for real-time performance.
3Manufacturing precision
If multiple video streams with different frame rates and resolutions are used for 3D models, then the quality of virtual object rendering is improved, but the device complexity increases
Solution Approach 1:
The patent segments different 3D model data into separate video streams with optimized frame rates and resolutions based on their specific requirements. This segmentation allows each virtual object to be rendered at appropriate quality levels while maintaining manageable stream complexity through standardized container formatting.
Data Source
AI summary
Multiple Holocam Orbs observe a real-life environment and generate an artificial reality representation of the real-life environment. Depth image data is cleansed of error due to LED shadow by identifying the edge of a foreground object in an (near infrared light) intensity image, identifying an edge in a depth image, and taking the difference between the start of both edges. Depth data error due to parallax is identified noting when associated text data in a given pixel row that is progressing in a given row direction (left-to-right or right-to-left) reverses order. Sound sources are identified by comparing results of a blind audio source localization algorithm, with the spatial 3D model provided by the Holocam Orb. Sound sources that corresponding to identifying 3D objects are associated together. Additionally, types of data supported by a standard movie data container, such as an MPEG container, is expanding to incorporate free viewpoint data (FVD) model data. This is done by inserting FVD data of different individual 3D objects at different sample rates into a single video stream. Each 3D object is separately identified by a separately assigned ID.


