Foveated Rendering Reconstruction via Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies for high-fidelity visual rendering and compression in head-mounted displays face challenges in efficiently managing computational resources and transmission costs due to the need for high-resolution, high-frame-rate visuals across the entire field of view, while human visual acuity significantly decreases in peripheral vision.
Innovation Solution
A machine-learning approach that uses foveated rendering and compression, where only a sparse subset of pixels is rendered or transmitted based on human visual acuity, with a machine-learning model reconstructing complete images from sparse sample datasets, leveraging optical flow data and natural video statistics to minimize computational costs and artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-resolution rendering is applied across the entire field of view, then visual quality is improved, but computational cost increases significantly
Solution Approach 1:
The patent applies different rendering qualities to different regions of the visual field. High-resolution rendering is concentrated in the foveal region where visual acuity is highest, while peripheral regions use lower resolution. This local differentiation maintains overall visual quality while significantly reducing computational cost by avoiding uniform high-resolution rendering across the entire field of view.
Solution Approach 2:
The patent segments the visual field into multiple regions based on visual acuity characteristics, specifically dividing it into a foveal region (center) and peripheral regions. This segmentation allows the system to apply computationally intensive high-resolution rendering only to the foveal region while using simpler rendering for peripheral regions, thus resolving the contradiction between visual quality and computational cost.
2Productivity
If foveated compression is applied to reduce image quality in peripheral vision, then computational savings are achieved, but noticeable artifacts appear in the periphery
Solution Approach 1:
The patent performs preliminary action by pre-rendering or pre-computing high-quality reference images of the peripheral regions before the actual display or transmission. These pre-computed reference images are stored and later used to correct artifacts in the compressed peripheral regions, allowing the system to achieve computational savings through compression while maintaining visual quality by applying the pre-computed corrections.
Solution Approach 2:
The patent introduces an intermediary mechanism - a machine learning model or correction algorithm - that acts as a mediator between the compressed peripheral regions and the final displayed image. This intermediary processes the compressed peripheral data, identifies artifacts, and reconstructs or corrects them using information from the high-quality foveal region and pre-computed references, thus eliminating visible artifacts while preserving computational savings.
3Reliability
If conservative foveated rendering is used to maintain quality, then artifacts are minimized, but computational savings are only modest
Solution Approach 1:
The patent replaces the traditional mechanical or algorithmic approach of uniform quality rendering with a machine learning-based system. This substitution enables the system to achieve both artifact reduction and significant computational savings by using learned patterns and predictions to intelligently allocate rendering resources, rather than relying on conservative uniform quality settings that prevent artifacts but consume excessive computational resources.
Data Source
AI summary
In one embodiment, a computing system configured to generate a current frame may access a current sample dataset having incomplete pixel information of a current frame in a sequence of frames. The system may access a previous frame in the sequence of frames with complete pixel information. The system may further access a motion representation indicating pixel relationships between the current frame and the previous frame. The previous frame may then be transformed according to the motion representation. The system may generate the current frame having complete pixel information by processing the current sample dataset and the transformed previous frame using a first machine-learning model.


