Adaptive Streaming for 6 DoF VR Using Virtual Camera Views

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in efficiently encoding and streaming high-data-rate virtual reality (VR) scenes for six degrees of freedom (6 DoF) viewing, as existing video codecs are not optimized for hardware-accelerated decoding of volumetric VR content, leading to bottlenecks in delivery and suboptimal immersive experiences.

Innovation Solution

The method involves generating multiple two-dimensional virtual camera views, optimized for color and depth, which are encoded and streamed adaptively based on viewer selection, using existing video codecs, and decoded efficiently in hardware decoders to create a real-time 3D volumetric view, with parameters like pose, resolution, and field of view encoded for each stream.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple virtual camera views are generated and encoded for 6 DoF VR viewing, then the immersive experience and adaptability are improved, but the data transmission volume and processing complexity increase

Engineering Contradiction:
ImproveadaptabilityVSAvoiddata transmission volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent segments the complete 360-degree VR scene into multiple discrete virtual camera views (e.g., 6 or more views covering different directions). Instead of transmitting the entire spherical scene data, only the selected virtual camera views corresponding to the viewer's current field of view are transmitted. This segmentation enables adaptive streaming while significantly reducing data transmission volume.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary encoding of multiple virtual camera views during content preparation, storing them as pre-encoded streams. When streaming, the system can quickly select and transmit only the required views based on viewer orientation, avoiding real-time encoding overhead and reducing processing complexity during delivery.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If existing video codecs are used for encoding VR content, then device compatibility and ease of operation are improved, but decoding efficiency and processing speed deteriorate

Engineering Contradiction:
Improveease of operationVSAvoiddecoding efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent uses existing video codecs to encode virtual camera views into standard video stream formats. This copying approach allows the VR content to be processed by conventional video decoding hardware, maintaining broad device compatibility while enabling efficient hardware-accelerated decoding instead of requiring specialized software decoders.

Inventive Principle:
Principle #26Copying

3Measurement precision

If full 360-degree scene data is transmitted, then measurement precision and reliability are improved, but loss of time and productivity worsen

Engineering Contradiction:
Improvemeasurement precisionVSAvoidloss of time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts only the necessary portion of the 360-degree scene data corresponding to the viewer's current field of view and transmits only those selected virtual camera views. This extraction approach maintains sufficient measurement precision for the visible scene while significantly reducing transmission time and data volume compared to sending complete spherical scene data.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11218683B2Method and an apparatus and a computer program product for adaptive streaming
Publication Date: 2022.01.04 NOKIA TECHNOLOGIES OY
  • US11218683B2 patent drawing
  • US11218683B2 patent drawing
  • US11218683B2 patent drawing

AI summary

The invention relates to a method and technical equipment for implementing the method. The method comprises generating a three-dimensional segment of a scene of a content; generating more than one two-dimensional views of the three-dimensional segment, each two-dimensional view representing a virtual camera view; generating multi-view streams by encoding each of the two-dimensional views; encoding parameters of a virtual camera to the respective stream of the multi-view stream; receiving a selection of one or more streams of the multi-view stream; and streaming only the selected one or more streams.