Spatial Random Access Volumetric Video Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for capturing and displaying volumetric video in virtual reality face challenges with immense data volumes, leading to storage and bandwidth issues, as well as high decoding complexity, which hinders low-latency and high-fidelity playback.

Innovation Solution

A spatial random access coding and viewing scheme is implemented, allowing arbitrary viewing of a field-of-view within a volumetric video stream by dividing the viewing volume into a three-dimensional sampling grid, using inter-vantage and inter-spatial layer predictions to reduce bandwidth and storage requirements, and employing multi-spatial resolution layers for error resilience and flexibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If volumetric video data is captured and stored in full resolution for arbitrary viewpoint access, then viewing flexibility and immersion are improved, but data volume and storage requirements increase exponentially

Engineering Contradiction:
Improveviewing flexibilityVSAvoiddata volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The volumetric video data is divided into multiple spatial layers representing different depth planes. Each layer contains viewpoint images at specific depth positions, allowing the system to segment the complete volumetric data into manageable slices that can be selectively decoded and transmitted based on viewer position and device capabilities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different spatial layers are encoded at different resolutions and quality levels. Layers corresponding to depths within the viewer's current field of view are encoded at higher quality, while layers outside the immediate viewing cone use lower quality or are omitted entirely, optimizing bandwidth usage while maintaining perceived visual fidelity.

Inventive Principle:
Principle #3Local quality

2Ease of operation

If complete volumetric video data is transmitted to enable random access viewing, then arbitrary viewpoint access is enabled, but network bandwidth requirements become prohibitive

Engineering Contradiction:
Improverandom access capabilityVSAvoidbandwidth consumption
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

Volumetric video data is pre-encoded into multiple spatial layers with varying resolutions and quality levels before transmission. This preliminary segmentation and encoding allows the receiver to selectively decode only the necessary layers based on current viewing conditions, avoiding the need to transmit and decode the complete volumetric dataset for every viewpoint change.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system transmits a subset of the complete volumetric data by selecting specific spatial layers and viewpoint images that are most relevant to the current viewing situation. This partial transmission approach provides sufficient data for high-quality rendering of the current field of view without the overhead of transmitting all possible viewpoint data.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If high-resolution volumetric video is decoded in real-time for virtual reality playback, then visual fidelity is improved, but processing complexity and latency increase

Engineering Contradiction:
Improvevisual fidelityVSAvoiddecoding complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The decoded volumetric data is organized into separate spatial layers that can be processed independently. This segmentation allows the decoding pipeline to handle smaller, more manageable data units in parallel, reducing the computational burden on real-time rendering systems while maintaining the ability to reconstruct high-fidelity images.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system prioritizes decoding and rendering of spatial layers that correspond to the viewer's current field of view at high resolution, while using lower resolution or skipping decoding of layers outside the immediate viewing area. This selective processing approach maintains visual fidelity where it matters most while reducing overall computational complexity.

Inventive Principle:
Principle #3Local quality

4Reliability

If all spatial layers are transmitted and decoded for error-free playback, then reliability is improved, but data transmission time and bandwidth usage increase

Engineering Contradiction:
Improveplayback reliabilityVSAvoidtransmission data
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system transmits multiple spatial layers with different quality levels and redundancy characteristics. Critical layers that contain essential viewing information are transmitted with higher reliability and redundancy, while less critical layers use more efficient, less redundant encoding. This differential approach maintains playback reliability for the current field of view while reducing overall data transmission requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10341632B2Spatial random access enabled video system with a three-dimensional viewing volume
Publication Date: 2019.07.02 GOOGLE LLC
  • US10341632B2 patent drawing
  • US10341632B2 patent drawing
  • US10341632B2 patent drawing

AI summary

An environment may be displayed from a viewpoint. According to one method, volumetric video data may be acquired depicting the environment, for example, using a tiled camera array. A plurality of vantages may be distributed throughout a viewing volume from which the environment is to be viewed. The volumetric video data may be used to generate video data for each vantage, representing the view of the environment from that vantage. User input may be received designating a viewpoint within the viewing volume. From among the plurality of vantages, a subset nearest to the viewpoint may be identified. The video data from the subset may be retrieved and combined to generate viewpoint video data depicting the environment from the viewpoint. The viewpoint video data may be displayed for the viewer to display a view of the environment from the viewpoint selected by the user.