Probabilistic Gaze Model for 3D Video Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual reality systems face bandwidth constraints and lag due to generating high-quality content for a 360° environment, as most pixels are not viewed by the user, leading to inefficient data transmission and lower quality when predicting user gaze incorrectly.
Innovation Solution
A method using probabilistic models, such as heat maps, to determine optimal segment parameters for three-dimensional video, blurring regions based on user gaze probability, allowing for re-encoding and transmission with lower bitrate, and displaying regions of interest at higher resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If high-quality content is generated for every pixel in the 360° sphere, then visual quality is improved, but bandwidth consumption increases significantly
Solution Approach 1:
The patent applies local quality by differentiating the rendering quality across different regions of the 360° video sphere. Regions where users are likely to look (front hemisphere, areas near the gaze direction) are rendered at high quality, while regions where users are unlikely to look (back hemisphere, peripheral areas) are rendered at lower quality or blurred. This resolves the contradiction by maintaining high visual quality only where needed, thereby reducing overall bandwidth consumption while preserving user experience.
2Loss of information
If content is transmitted for all directions in the 360° environment, then completeness of content is improved, but transmission efficiency deteriorates
Solution Approach 1:
The patent extracts and transmits only the essential portions of the 360° video content that are likely to be viewed by the user. By using gaze estimation and probability maps, the system identifies and extracts the front hemisphere and regions of interest for transmission, while omitting or heavily compressing the back hemisphere and low-probability regions. This resolves the contradiction by maintaining content completeness for viewed areas while improving transmission efficiency through selective extraction.
3Adaptability or versatility
If the system updates virtual reality content for different user gaze directions, then responsiveness to user movement is improved, but lag increases due to processing time
Solution Approach 1:
The patent implements preliminary action by pre-rendering and pre-transmitting the front hemisphere and high-probability regions of the 360° video before the user actually looks in those directions. By anticipating likely gaze directions and preparing content in advance, the system reduces the processing time required when the user moves their gaze, thereby improving responsiveness while minimizing lag. The blurred back hemisphere serves as a placeholder that can be quickly replaced if needed.
4Quantity of substance
If prediction of user gaze direction is used to reduce bandwidth, then bandwidth consumption is reduced, but content quality and stability deteriorate when prediction is wrong
Solution Approach 1:
The patent uses parameter changes by dynamically adjusting the quality parameter based on the predicted gaze direction and probability. When gaze prediction confidence is high, the system transmits high-quality content for the predicted region. When confidence is low or the prediction is likely wrong, the system transitions to transmitting a blurred version of the back hemisphere or uses probability maps to adjust quality distribution. This resolves the contradiction by adaptively changing quality parameters to maintain reliability while reducing bandwidth consumption.
Data Source
AI summary
A method includes receiving head-tracking data that describe one or more positions of people while the people are viewing a three-dimensional video. The method further includes generating a probabilistic model of the one or more positions of the people based on the head-tracking data, wherein the probabilistic model identifies a probability of a viewer looking in a particular direction as a function of time. The method further includes generating video segments from the three-dimensional video. The method further includes, for each of the video segments: determining a directional encoding format that projects latitudes and longitudes of locations of a surface of a sphere onto locations on a plane, determining a cost function that identifies a region of interest on the plane based on the probabilistic model, and generating optimal segment parameters that minimize a sum-over position for the region of interest.


