Spatial Random Access Volumetric Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for capturing and displaying volumetric video in virtual reality face challenges with immense data volumes, leading to storage and bandwidth issues, as well as high decoding complexity, which hinders low-latency and high-fidelity playback.
Innovation Solution
A spatial random access coding and viewing scheme is implemented, allowing arbitrary viewing of a field-of-view within a volumetric video stream by dividing the viewing volume into a three-dimensional sampling grid, using inter-vantage and inter-spatial layer predictions to reduce bandwidth and storage requirements, and employing multi-spatial resolution layers for error resilience and flexibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If volumetric video data is captured and stored in full resolution for arbitrary viewpoint access, then viewing flexibility and immersion are improved, but data volume and storage requirements increase exponentially
Solution Approach 1:
The volumetric video data is divided into multiple spatial layers representing different depth planes. Each layer contains viewpoint images at specific depth positions, allowing the system to segment the complete volumetric data into manageable slices that can be selectively decoded and transmitted based on viewer position and device capabilities.
Solution Approach 2:
Different spatial layers are encoded at different resolutions and quality levels. Layers corresponding to depths within the viewer's current field of view are encoded at higher quality, while layers outside the immediate viewing cone use lower quality or are omitted entirely, optimizing bandwidth usage while maintaining perceived visual fidelity.
2Ease of operation
If complete volumetric video data is transmitted to enable random access viewing, then arbitrary viewpoint access is enabled, but network bandwidth requirements become prohibitive
Solution Approach 1:
Volumetric video data is pre-encoded into multiple spatial layers with varying resolutions and quality levels before transmission. This preliminary segmentation and encoding allows the receiver to selectively decode only the necessary layers based on current viewing conditions, avoiding the need to transmit and decode the complete volumetric dataset for every viewpoint change.
Solution Approach 2:
The system transmits a subset of the complete volumetric data by selecting specific spatial layers and viewpoint images that are most relevant to the current viewing situation. This partial transmission approach provides sufficient data for high-quality rendering of the current field of view without the overhead of transmitting all possible viewpoint data.
3Measurement precision
If high-resolution volumetric video is decoded in real-time for virtual reality playback, then visual fidelity is improved, but processing complexity and latency increase
Solution Approach 1:
The decoded volumetric data is organized into separate spatial layers that can be processed independently. This segmentation allows the decoding pipeline to handle smaller, more manageable data units in parallel, reducing the computational burden on real-time rendering systems while maintaining the ability to reconstruct high-fidelity images.
Solution Approach 2:
The system prioritizes decoding and rendering of spatial layers that correspond to the viewer's current field of view at high resolution, while using lower resolution or skipping decoding of layers outside the immediate viewing area. This selective processing approach maintains visual fidelity where it matters most while reducing overall computational complexity.
4Reliability
If all spatial layers are transmitted and decoded for error-free playback, then reliability is improved, but data transmission time and bandwidth usage increase
Solution Approach 1:
The system transmits multiple spatial layers with different quality levels and redundancy characteristics. Critical layers that contain essential viewing information are transmitted with higher reliability and redundancy, while less critical layers use more efficient, less redundant encoding. This differential approach maintains playback reliability for the current field of view while reducing overall data transmission requirements.
Data Source
AI summary
An environment may be displayed from a viewpoint. According to one method, volumetric video data may be acquired depicting the environment, for example, using a tiled camera array. A plurality of vantages may be distributed throughout a viewing volume from which the environment is to be viewed. The volumetric video data may be used to generate video data for each vantage, representing the view of the environment from that vantage. User input may be received designating a viewpoint within the viewing volume. From among the plurality of vantages, a subset nearest to the viewpoint may be identified. The video data from the subset may be retrieved and combined to generate viewpoint video data depicting the environment from the viewpoint. The viewpoint video data may be displayed for the viewer to display a view of the environment from the viewpoint selected by the user.


