Lensless Camera Eye Tracking via Optical Mask and Neural Network

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing eye and hand tracking systems in head-mounted display (HMD) devices face challenges in efficiently and accurately tracking user gaze and hand movements due to the need for complex image reconstruction and high computational requirements.

Innovation Solution

The implementation of lensless camera systems using optical masks with point spread functions (PSFs) that convolve optical images of eyes and hands, combined with machine learning systems like convolutional neural networks (CNNs), which extract body features directly from coded representations without deconvolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional lens-based camera systems are used for eye and hand tracking, then image quality and recognition accuracy are improved, but device complexity and computational requirements increase

Engineering Contradiction:
Improvetracking accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the lens component from the traditional camera system, retaining only the optical mask and sensor. This removes the complex image formation mechanism while preserving the ability to capture coded information directly from reflected light, thereby reducing device complexity while maintaining tracking functionality

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical/optical image formation system (lens-based focusing) with a computational approach using coded aperture masks and machine learning algorithms. The optical encoding replaces traditional optical focusing, and neural networks replace conventional image processing pipelines, reducing mechanical complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If conventional lens-based camera systems with full image reconstruction are used, then detailed body feature information is obtained, but power consumption and processing time increase

Engineering Contradiction:
Improvebody feature extraction accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential body feature information (pupil position, hand joint locations) directly from coded representations without performing full image reconstruction. This selective extraction approach maintains tracking accuracy while avoiding the computational overhead of reconstructing complete images

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The optical mask pre-encodes the incoming light patterns with spatially varying point spread functions before detection. This preliminary optical encoding embeds body feature information directly into the captured signal, allowing the neural network to extract features without requiring intensive post-processing or full image reconstruction

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If deconvolution processes are performed to reconstruct original images from coded representations, then image quality is improved, but computational intensity and processing time increase

Engineering Contradiction:
Improveimage qualityVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The optical mask with spatially varying PSFs performs preliminary optical encoding that directly maps body features to distinct patterns in the coded representation. This pre-encoding structure allows neural networks to extract features through pattern recognition without requiring computationally intensive deconvolution or image reconstruction processes

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the computational deconvolution process with a neural network-based direct feature extraction approach. Instead of mathematically reversing the convolution operation to reconstruct images, the neural network learns to map coded representations directly to body feature coordinates, significantly reducing computational intensity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables accurate and efficient body tracking with reduced computational intensity, lower power consumption, and faster processing times, thereby enhancing the performance and battery life of HMD devices.

Implementation Method 1

lensless camera systems using optical masks with point spread functions (PSFs) that convolve optical images of eyes and hands

Methodology Applied
Scientific EffectOptical convolution:

Implementation Method 2

optical mask is described by a point spread function (PSF) and may be implemented using diffractive optical elements such as coded apertures, amplitude masks, phase masks, diffusers, holographic diffraction grating films, or metasurfaces

Methodology Applied
Scientific EffectDiffraction: Diffraction

Implementation Method 3

an illumination system for flooding diffuse illumination to the HMD device user's eye to produce reflective glints

Methodology Applied
Scientific EffectReflection: Reflection

Data Source

PatentUS20250199614A1Eye and hand tracking utilizing lensless camera and machine learning
Publication Date: 2025.06.19 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250199614A1 patent drawing
  • US20250199614A1 patent drawing
  • US20250199614A1 patent drawing

AI summary

Eye and hand tracking systems in head-mounted display (HMD) devices are arranged with lensless camera systems using optical masks as encoding elements that apply convolutions to optical images of body parts (e.g., eyes or hands) of HMD device users. The convolved body images are scrambled or coded representations that are captured by a sensor in the system, but are not human-recognizable. A machine learning system such as a neural network is configured to extract body features directly from the coded representation without performance of deconvolutions conventionally utilized to reconstruct the original body images in human-recognizable form. The extracted body features are utilized by the respective eye or hand tracking systems to output relevant tracking data for the user's eyes or hands which may be utilized by the HMD device to support various applications and user experiences. The lensless camera and machine learning system are jointly optimizable on an end-to-end basis.