Coded Aperture Mask Gaze Tracking for AR Miniaturization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality (AR) and virtual reality (VR) devices face challenges in miniaturizing gaze-tracking camera systems without compromising performance, as they often require bulky lens assemblies and struggle to efficiently extract feature points from coded images for accurate gaze direction tracking.
Innovation Solution
An electronic device comprising a light source, a pattern mask, and an image sensor that outputs and receives light to generate a coded image, which is then processed to extract feature points, including pupil and glint information, using an AI model to determine gaze direction and optionally reconstruct the original image for authentication purposes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a camera system with lens assembly is used for gaze tracking, then gaze direction tracking capability is achieved, but device size increases
Solution Approach 1:
The patent replaces the traditional mechanical lens-based optical system with a computational imaging approach using a coded aperture mask and AI-based image reconstruction. This substitution eliminates bulky lenses while maintaining gaze tracking capability through algorithmic processing of light patterns captured by a smaller sensor array.
Solution Approach 2:
The invention changes the fundamental parameters of the optical system by using a coded aperture mask with specific transmission patterns instead of conventional lenses. The AI model learns to interpret these modified light patterns, enabling gaze tracking with a compact sensor configuration that would be impossible with traditional optical components.
2Volume of moving object
If a coded image is used for feature point extraction, then device size is reduced, but feature point extraction accuracy may be degraded
Solution Approach 1:
The AI model serves as an intermediary that bridges the coded aperture mask and the image sensor. It learns to decode the complex light patterns created by the mask, transforming the distorted coded images into accurate feature point measurements. This intermediary processing layer enables precise gaze tracking despite the non-traditional optical path.
Solution Approach 2:
The system performs preliminary encoding of the light pattern using the coded aperture mask before detection. The AI model is pre-trained to recognize and decode these encoded patterns, enabling accurate feature point extraction from the modified images. This preliminary encoding action allows the use of smaller sensors while maintaining measurement precision.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution enables the miniaturization of gaze-tracking systems while maintaining performance by extracting feature points from coded images, allowing for accurate gaze direction tracking and user authentication, and reduces the size of camera modules, making them suitable for integration into AR glasses and other devices.
Implementation Method 1
obtain a coded image that is phase-modulated based on light transmitted through the pattern mask
Implementation Method 2
receive light that is output from the light source, reflected by an eye, and transmitted through the pattern mask
Implementation Method 3
a time of flight (TOF) sensor configured to obtain depth information based on received light
Data Source
AI summary
An electronic device includes a light source configured to output light, a pattern mask configured to change a path of light transmitted through a pattern of the pattern mask, an image sensor configured to receive light that is output from the light source, reflected by an eye, and transmitted through the pattern mask, and at least one processor configured to obtain a coded image that is phase-modulated based on light transmitted through the pattern mask, obtain a feature point of the eye from the coded image, and obtain gaze information of a user based on the feature point.


