Gaze Estimation Neural Network RGB Camera Feature Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing RGB-based eye-tracking systems face challenges in accurately estimating gaze direction due to limitations in hardware and computational resources, particularly in real-world environments with poor illumination and extreme head movements, and are not suitable for commercial end-user devices like mobile phones or tablets.

Innovation Solution

A neural network-based method that uses an RGB camera to estimate gaze direction by extracting feature representations from face and eye images, fusing them, and outputting a gaze vector, which can be implemented on standard end-user devices with limited resources, including mobile devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If IR camera-based eye tracking is used, then gaze tracking accuracy is improved, but hardware cost and complexity increase

Engineering Contradiction:
Improvegaze tracking accuracyVSAvoidhardware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses a neural network model trained on IR camera data to copy and replicate the accuracy benefits of IR cameras using only RGB camera inputs. The network learns to map RGB image features to gaze directions, effectively copying the performance of more complex IR systems without requiring the specialized hardware.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the physical IR camera hardware system with a computational neural network model. Instead of using specialized IR sensors and illumination equipment, the system uses standard RGB cameras combined with deep learning algorithms to achieve similar or better gaze estimation accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If IR camera-based eye tracking is used, then gaze tracking accuracy is improved, but accessibility to real-world environments decreases

Engineering Contradiction:
Improvegaze tracking accuracyVSAvoidreal-world adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal gaze estimation system that works with standard RGB cameras found in mobile devices and consumer electronics, making the technology accessible across multiple platforms and environments. The neural network model can be deployed on various devices without requiring specialized IR hardware, enabling real-world applications in diverse settings.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If complex neural network models are used, then gaze estimation performance is improved, but computational resource requirements increase

Engineering Contradiction:
Improvegaze estimation performanceVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent optimizes the neural network architecture by changing parameters such as the number of layers, filter sizes, and activation functions to achieve an efficient balance between performance and computational cost. The model uses depth-wise separable convolutions and optimized parameter configurations that reduce computational burden while maintaining accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12175014B2Methods and devices for gaze estimation
Publication Date: 2024.12.24 HUAWEI TECH CO LTD
  • US12175014B2 patent drawing
  • US12175014B2 patent drawing
  • US12175014B2 patent drawing

AI summary

Methods and systems for estimating a gaze direction of an individual using a trained neural network. Inputs to the neural network include a face image and an image of a visually significant eye in the face image. Feature representations are extracted for the face image and significant eye image and feature fusion is performed on the feature representations to generate a fused feature representation. The fused feature representation is input into a trained gaze estimator to output a gaze vector including gaze angles, the gaze vector representing a gaze direction. The disclosed network may enable gaze estimation performance on user devices typically having limited hardware and computational resources such as mobile devices.