VR Video Encoding Using Eye Gaze Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional VR systems fail to provide high visual quality and low latency, which are essential for a continuous realistic virtual reality experience, due to limitations in processing time, memory bandwidth, and inefficient encoding processes.
Innovation Solution
The use of tracking information, such as user position and eye gaze point, to predict user viewpoints and guide motion searching and encoding decisions, allowing for efficient video compression and encoding without comparing coding modes or evaluating rate distortion costs, thereby reducing processing time and memory bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional video encoding processes are used in VR systems, then encoding completeness is maintained, but processing time increases and latency increases
Solution Approach 1:
The system performs preliminary actions by predicting the user's future viewpoint using tracking information before the actual encoding process. This allows the encoder to pre-determine motion vectors and encoding parameters, significantly reducing the time required during actual frame encoding and transmission.
Solution Approach 2:
The encoding process is segmented into different regions based on predicted user attention areas. Only regions of interest are encoded with high quality while other areas use lower quality encoding, reducing overall processing time and computational requirements while maintaining perceived visual quality.
2Productivity
If conventional video encoding processes are used in VR systems, then encoding completeness is maintained, but memory bandwidth increases
Solution Approach 1:
The system applies local quality by encoding different regions of the video frame with different quality levels based on predicted user gaze and attention areas. High-quality encoding is applied only to regions of interest while peripheral areas use lower quality, reducing overall memory bandwidth requirements.
Solution Approach 2:
Motion vectors and encoding parameters are determined in advance using tracking information, allowing the system to avoid loading and processing complete reference frames during encoding. This preliminary determination significantly reduces memory bandwidth consumption during the actual encoding process.
3Loss of time
If tracking information is used to predict user viewpoint and guide encoding, then processing time decreases, but system complexity increases
Solution Approach 1:
The system introduces tracking information as an intermediary element that bridges user behavior and encoding parameters. This intermediary provides a straightforward mechanism to predict user viewpoint and guide encoding decisions without requiring complex real-time analysis of user actions.
Solution Approach 2:
User viewpoint prediction is performed in advance using tracking information, creating a simplified workflow where encoding parameters are predetermined based on predicted gaze patterns. This preliminary action reduces the need for complex real-time processing during frame encoding.
4Manufacturing precision
If high visual quality video data is transmitted in VR systems, then visual experience realism is improved, but latency increases
Solution Approach 1:
The system performs preliminary encoding and compression of video data before transmission, using predicted user viewpoint information to optimize compression parameters. This allows high visual quality to be maintained in regions of interest while achieving lower overall bitrates and reduced transmission latency.
Solution Approach 2:
Different quality levels are applied to different regions of the video frame based on predicted user attention. High visual quality is maintained only in regions where the user is likely to look, while peripheral areas use lower quality encoding, reducing overall data transmission requirements and latency.
Data Source
AI summary
Systems, methods and apparatuses of processing data of a VR system are disclosed that comprise receiving tracking information which includes at least one of user position information and eye gaze point information. One or more processors may be used to predict, based on the user tracking information, a user viewpoint of a next frame of a sequence of frames of video data to be displayed. Using the prediction, a portion of the next frame of video data to be displayed is rendered at an estimated location in the next frame. A corresponding matching portion in a previously encoded frame is determined based on the estimated location of the portion in the next frame and the portion of the next frame of video data is encoded.


