Directional Encoding of 3D Video Using Head-Tracking Regions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual reality systems face challenges in efficiently streaming high-quality 360° video content due to massive data requirements and bandwidth constraints, often resulting in lag or lower quality when updating content based on user direction, and prediction errors can further degrade the experience.
Innovation Solution
A method for generating optimal segment parameters for three-dimensional video using head-tracking data to determine directional encoding formats, identify regions of interest, and re-encode video segments, allowing for higher resolution display of interest areas while reducing overall bitrate, by projecting latitudes and longitudes onto a plane and minimizing cost functions for optimal segment parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If high-quality content is generated for every pixel in the sphere, then visual quality is improved, but data requirements become massive and bandwidth constraints are exceeded
Solution Approach 1:
The patent applies local quality by encoding different regions of the 360° video at different resolutions. Specifically, it identifies regions of interest (such as where users are likely to look based on head-tracking data) and allocates higher bitrate and resolution to these regions, while using lower resolution for less important areas. This resolves the contradiction by maintaining high visual quality where needed while reducing overall data requirements.
2Manufacturing precision
If virtual reality content is updated based on user direction, then visual quality is improved, but lag occurs due to continuous updates required
Solution Approach 1:
The patent uses head-tracking data to predict where users are likely to look and pre-encodes multiple possible regions of interest into the video stream. This preliminary action allows the system to have high-quality content ready in advance for anticipated viewing directions, eliminating the need for continuous real-time updates and reducing lag while maintaining visual quality.
3Loss of energy
If prediction of user gaze direction is implemented, then bandwidth usage is reduced, but prediction errors degrade content quality and stability
Solution Approach 1:
The patent implements a dynamic encoding approach where multiple encoding configurations are prepared based on different predicted gaze directions. The system dynamically selects and switches between these pre-encoded configurations based on actual user head movements, rather than relying on a single static prediction. This reduces bandwidth usage while maintaining content stability by having backup encoding options ready.
Data Source
AI summary
A method includes receiving head-tracking data that describes one or more positions of one or more people while the one or more people are viewing a three-dimensional video. The method further includes generating video segments from the three-dimensional video. The method further includes, for each of the video segments: determining a directional encoding format that projects latitudes and longitudes of locations of a surface of a sphere onto locations on a plane, determining a cost function that identifies a region of interest on the plane based on the head-tracking data, and generating optimal segment parameters that minimize a sum-over position for the region of interest.


