Volumetric Video Decoding via Adaptive Virtual Camera Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies for encoding and decoding volumetric video face challenges in efficiently transmitting and rendering 6DoF video content due to high bandwidth requirements, limiting user immersion and parallax rendering, especially in scenarios with varying network throughput and user navigation behaviors.
Innovation Solution
The method involves encoding and decoding volumetric video by selecting and transmitting color and depth data patches associated with virtual cameras, allowing adaptive selection based on network conditions and user behavior, enabling flexible parallax rendering and improved immersion by optimizing the number and location of virtual cameras used.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If complete 6DoF volumetric video content is transmitted to provide full immersion and parallax rendering, then user immersion and rendering quality are improved, but network bandwidth requirements increase significantly
Solution Approach 1:
The volumetric video content is divided into multiple virtual camera views, each represented as separate data streams. The system segments the complete 6DoF content into discrete camera positions and uses selective transmission of only the required segments based on user navigation state, reducing overall bandwidth while maintaining immersion quality.
Solution Approach 2:
Instead of transmitting all possible virtual camera views simultaneously, the system transmits only the partial set of views needed for the current user position and navigation behavior. This partial action approach reduces bandwidth consumption while providing sufficient parallax rendering for the user's actual viewing requirements.
2Adaptability or versatility
If multiple virtual cameras are used to enable parallax rendering and 6DoF navigation, then rendering quality and user experience are improved, but device complexity and data processing requirements increase
Solution Approach 1:
The system dynamically adjusts the number and positioning of virtual cameras based on real-time user navigation behavior. The set of active virtual cameras is updated adaptively, allowing the system to maintain high rendering quality when needed while reducing processing complexity during periods of limited user interaction or when fewer views are required.
Solution Approach 2:
The system changes parameters such as the number of virtual cameras, their positions, and data transmission rates based on user navigation state and device capabilities. By adjusting these parameters dynamically, the system optimizes the balance between rendering quality and processing requirements without fixed complexity.
3Loss of energy
If adaptive selection of virtual camera views is implemented based on user behavior, then bandwidth efficiency is improved, but system complexity and decoding requirements increase
Solution Approach 1:
The system incorporates feedback mechanisms that monitor user navigation behavior and adjust the selection of virtual camera views accordingly. By continuously adapting to user actions, the system optimizes bandwidth efficiency while managing decoding complexity through intelligent prediction of required views based on observed user patterns.
Data Source
AI summary
A method and an apparatus for decoding a volumetric video are disclosed. Such a method comprises receiving a data stream representative of a file comprising information for selecting, according to a rendering viewpoint, at least one atlas comprising color and depth data patches associated with a viewpoint in said volumetric video, said color and depth data patches being generated with respect to depth and color reference data acquired from a reference viewpoint in said volumetric video.


