Single Sensor Head Pose Estimation Using Eye Region Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing eye tracking and head tracking systems face challenges in integrating both functionalities into a single device, often requiring sacrifices in sensor frame rate, power consumption, processing resources, and increasing costs due to the need for multiple sensors and complex coordination.

Innovation Solution

A system utilizing a single sensor to capture image data optimized for both eye tracking and head pose determination, leveraging artificial intelligence techniques like machine learning and deep learning to process eye region data, excluding non-essential facial features, and determining head pose information based on eye features such as irises, pupils, and eyelids.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single sensor system is used to perform both eye tracking and head tracking, then device complexity and cost are reduced, but measurement precision and frame rate may be compromised

Engineering Contradiction:
Improvesensor system complexityVSAvoideye tracking precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the facial region into distinct zones (eye region, nose region, mouth region) and processes each zone separately using dedicated algorithms. The eye region is processed with high-resolution algorithms for precise gaze tracking, while the nose and mouth regions are processed with algorithms optimized for head pose estimation. This segmentation allows a single sensor to simultaneously provide both eye tracking precision and head tracking accuracy without compromising either function.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If multiple sensors are used to achieve both eye tracking and head tracking functions, then measurement precision is improved, but device complexity, power consumption, and cost increase

Engineering Contradiction:
Improvehead pose measurement precisionVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal processing system where a single sensor captures images that are subsequently processed to serve dual purposes: eye tracking and head pose estimation. The system uses a unified image processing pipeline that extracts both gaze information and head pose information from the same captured images, eliminating the need for separate sensors for each function while maintaining measurement precision through specialized algorithms.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If image data is processed at high resolution for eye tracking, then measurement precision is improved, but processing resources and power consumption increase

Engineering Contradiction:
Improvegaze information precisionVSAvoidprocessing power consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality processing by analyzing different regions of the facial image with different levels of computational intensity. The eye region, which requires high precision for gaze tracking, is processed with high-resolution algorithms, while the nose and mouth regions are processed with algorithms optimized for head pose estimation that require less computational resources. This localized approach maintains measurement precision for critical functions while reducing overall processing power consumption.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11715231B2Head pose estimation from local eye region
Publication Date: 2023.08.01 TOBII TECH AB
  • US11715231B2 patent drawing
  • US11715231B2 patent drawing
  • US11715231B2 patent drawing

AI summary

Head pose information may be determined using information describing a fixed gaze and image data corresponding to a user's eyes. The head pose information may be determined in a manner that is disregards facial features with the exception of the user's eyes. The head pose information may be usable to interact with a user device.