360-Degree Video Metadata Encoding for Viewing Space Shape Types
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current VR systems face inefficiencies in transmitting 360-degree video data, particularly in providing interactive experiences with accurate viewing position and space information, which limits the quality of 3DoF+ content consumption.
Innovation Solution
A method and apparatus for processing and transmitting 360-degree video data that includes acquiring, encoding, and transmitting video data along with metadata containing viewing space information, enabling efficient storage and transmission of 3DoF+ content, and displaying a user interface that reflects viewing positions and spaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 360-degree video data is transmitted with detailed viewing position and space information, then the quality of 3DoF+ content consumption is improved, but the data transmission efficiency deteriorates
Solution Approach 1:
The patent segments viewing space information into discrete shape types (spherical, cubic, cylindrical) and encodes them using compact metadata structures. This segmentation allows the system to transmit precise viewing position information without proportionally increasing data volume, as the viewing space geometry is represented through categorized parameters rather than continuous coordinate data for every point in space.
Solution Approach 2:
Instead of transmitting complete 3D coordinate data for all viewing positions, the patent inverts the approach by transmitting compact metadata that describes the viewing space geometry (shape type, boundaries, key parameters). The receiver then reconstructs the viewing position information from this compressed metadata, achieving high precision with reduced transmission data.
2Adaptability or versatility
If metadata including viewing space information is transmitted, then the interactivity of VR content is improved, but the transmission data volume increases
Solution Approach 1:
The patent extracts only the essential elements needed for interactivity (viewing space shape type, boundaries, and key geometric parameters) from the complete 3D environment data. This selective extraction creates a compact metadata subset that enables interactive 3DoF+ experiences without transmitting the full complexity of the virtual environment, thus reducing overall data volume while maintaining interactivity.
Solution Approach 2:
The patent changes the parameter representation from continuous 3D coordinates to discrete geometric shape categories with bounded parameters. By representing viewing spaces as standardized shapes (sphere, cube, cylinder) with defined parameters, the system reduces data volume while preserving the adaptability needed for interactive VR content delivery.
3Manufacturing precision
If viewing space information is encoded and transmitted, then the rendering accuracy at different viewing positions is improved, but the processing complexity increases
Solution Approach 1:
The patent performs preliminary encoding of viewing space information into standardized metadata formats during content creation. By pre-processing and categorizing the viewing space geometry into defined shape types with standardized parameters, the system reduces real-time rendering complexity while maintaining high rendering precision across different viewing positions.
Solution Approach 2:
The patent combines multiple types of information (geometry type, boundaries, viewing position data) into a unified composite metadata structure. This composite approach integrates various rendering parameters into a single coherent data format, simplifying the decoding and rendering processes while preserving the precision needed for accurate 3DoF+ content delivery.
Data Source
AI summary
A 360-degree video data processing method performed by a 360-degree video reception apparatus, according to the present invention, comprises the steps of: receiving 360-degree video data; deriving metadata and information on an encoded picture for a specific viewing position in specific viewing space based on the 360-degree video data; decoding the encoded picture based on the information on the encoded picture; rendering the decoded picture based on the metadata; and displaying a viewport in a 360-degree video generated based on the rendering, wherein the viewport includes a mini-map representing the specific viewing space, and wherein a shape of the mini-map is derived as a shape corresponding to a shape type of the specific viewing space.


