Volumetric Video Navigation Metadata for 3D Viewing Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing immersive video technologies, such as 3DoF and 3DoF+, fail to provide adequate navigation freedom and can cause dizziness due to incomplete head translations, while 6DoF videos require restrictions to ensure consistent visual feedback and prevent users from leaving the 3D scene, necessitating a solution for signaling navigation restrictions.
Innovation Solution
Encoding metadata in a data stream that includes a viewing bounding box, curvilinear path, and viewing direction ranges to restrict navigation within a 3D space, ensuring consistent visual feedback and preventing users from leaving the 3D scene.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If users are allowed to navigate freely within volumetric video content, then the immersive experience and freedom of movement are improved, but the quality of visual data deteriorates when users move outside the captured 3D scene
Solution Approach 1:
The patent applies preliminary action by pre-defining valid viewing regions (viewing bounding boxes) and navigation paths within the volumetric video content during the encoding phase. These constraints are embedded in the metadata so that the playback device can guide users along predetermined paths where high-quality visual data is guaranteed to exist, preventing users from navigating to regions without captured data.
Solution Approach 2:
The patent uses metadata as an intermediary between the volumetric video content and the user navigation system. The metadata contains encoded information about valid viewing regions and navigation paths, acting as a mediator that translates the physical constraints of the captured scene into guidance instructions for the playback device, thereby maintaining visual quality while enabling navigation.
2Reliability
If viewing constraints are imposed on users, then the quality of immersive experience is maintained, but the freedom of navigation is reduced
Solution Approach 1:
The patent applies dynamics by making the navigation constraints adaptive rather than static. The system dynamically adjusts the guidance based on the user's current position and the predefined valid regions. The playback device provides real-time feedback and guidance to steer users back toward valid viewing areas, making the constraints feel more like dynamic guidance than rigid restrictions.
Solution Approach 2:
The patent implements feedback by having the playback device monitor user navigation and provide guidance information when users approach or enter invalid regions. The system feedbacks navigation suggestions to users, indicating which directions lead to valid viewing areas, thereby maintaining visual quality while preserving user autonomy through informative guidance rather than强制 restrictions.
3Adaptability or versatility
If complete free navigation is provided, then user immersion is enhanced, but data completeness for all possible viewing positions cannot be guaranteed
Solution Approach 1:
The patent applies segmentation by dividing the 3D space into distinct regions: valid viewing regions (viewing bounding boxes) where complete visual data exists, and invalid regions where data is incomplete or unavailable. By segmenting the navigation space and providing guidance to stay within valid regions, the system ensures data completeness is maintained while still allowing extensive navigation within those regions.
Solution Approach 2:
The patent transitions from thinking about navigation freedom in three-dimensional space to adding a fourth dimension of information quality. By overlaying the dimension of data validity (represented by viewing bounding boxes and paths), the system creates a navigable volume where users can move freely within constraints, effectively transforming the navigation problem into a multi-dimensional solution space.
Data Source
AI summary
Methods, devices and data stream are provided for signaling and decoding information representative of restrictions of navigation in a volumetric video. The data stream comprises metadata associated to video data representative of the volumetric video. The metadata comprise data representative of a viewing bounding box, data representative of a curvilinear path in the 3D space of said volumetric video; and data representative of at least one viewing direction range associated with a point on the curvilinear path.


