Gaze-Based Super-Resolution for XR Foveated Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Low-resolution cameras in extended reality (XR) systems adversely impact image content resolution, necessitating techniques to improve image quality while reducing power consumption and system costs.
Innovation Solution
A computing device worn by a user determines the XR environment context using eye characteristics, generates a foveated map, and applies gaze-based super-resolution reconstruction using machine-learning models like CNN, DCNN, ViT, or GAN to enhance image resolution, particularly in the central 30°×30° field of view, and applies multi-frame temporal super-resolution reconstruction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If low-resolution cameras are used in XR systems, then power consumption and system costs are reduced, but image content resolution deteriorates
Solution Approach 1:
The image content is divided into multiple regions based on eye gaze characteristics, with the foveal region (central vision) receiving super-resolution processing while peripheral regions maintain lower resolution. This segmentation allows selective application of computational resources to only the most visually important areas, resolving the contradiction between reduced power consumption and maintained image quality in critical regions.
Solution Approach 2:
Different resolution qualities are applied to different regions of the image based on their importance to human vision. The foveal region receives high-resolution super-resolution reconstruction while peripheral regions maintain lower resolution, matching the non-uniform sensitivity of human visual acuity across the visual field and thereby reducing overall computational load while preserving perceived image quality.
2Manufacturing precision
If super-resolution reconstruction is applied to the entire image, then image resolution is improved, but computational complexity and processing time increase
Solution Approach 1:
The image is segmented into foveal and peripheral regions, with super-resolution reconstruction applied only to the foveal region. This reduces the total number of pixels requiring computationally intensive processing while maintaining high resolution in the most visually important area, thereby reducing computational complexity and processing time.
Solution Approach 2:
Instead of applying super-resolution reconstruction to the entire image (excessive action), the method applies it only to the necessary foveal region (partial action). This partial application of the computational process achieves sufficient image quality for the user's actual visual needs while significantly reducing computational complexity and processing requirements.
3Manufacturing precision
If high-resolution cameras are used, then image content resolution is improved, but power consumption and system costs increase
Solution Approach 1:
Instead of capturing high-resolution images directly with hardware (which increases power consumption and cost), the system captures low-resolution images and creates a computational copy through super-resolution reconstruction algorithms. This software-based copying approach achieves high-resolution output without the power and cost penalties of high-resolution hardware cameras.
Solution Approach 2:
The system changes the resolution parameter dynamically based on eye gaze data, applying super-resolution processing only to the foveal region where high resolution is needed. This parameter change approach allows the system to achieve high effective resolution in critical areas while maintaining lower overall resolution, thereby reducing power consumption compared to uniformly high-resolution capture.
Data Source
AI summary
A method implemented by a computing device includes rendering on displays of a computing device an extended reality (XR) environment, and determining a context of the XR environment with respect to a user. Determining the context includes determining characteristics associated with an eye of the user with respect to content displayed. The method includes generating, based on the characteristics associated with the eye, a foveated map including a plurality of foveal regions. The plurality of foveal regions includes a plurality of zones each corresponding to a low-resolution area of the content for the respective zone. The method includes inputting one or more of the plurality of zones into a machine-learning model trained to generate a super-resolution reconstruction of the foveated map based on regions of interest identified within the one or more of the plurality of zones, and outputting, by the machine-learning model, the super-resolution reconstruction of the foveated map.


