360-Degree Gaze Tracking for Object Identification in Unconstrained Spaces

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing gesture control and gaze tracking systems are limited to constrained environments and cannot accurately detect human gaze or gestures in unconstrained spaces, failing to identify objects of attention due to limited camera views and inability to observe the same environment as the user.

Innovation Solution

A system utilizing a 360-degree monocular camera and infrared depth sensors to capture and process 2D and 3D data, generating saliency maps and directionality vectors to accurately locate and identify objects based on human gaze and gestures in unconstrained environments, with machine learning for knowledge expansion and interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If narrow vision cameras are used for gaze tracking, then system complexity is reduced, but the ability to detect objects in unconstrained environments is lost

Engineering Contradiction:
Improvecamera system complexityVSAvoidenvironmental adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent transitions from 2D camera views to 360-degree spatial coverage by deploying multiple cameras at different positions (including ceiling-mounted cameras) to capture the entire environment. This dimensional expansion allows the system to observe objects from multiple angles and create a comprehensive 3D understanding of the space, enabling accurate gaze and gesture detection in unconstrained environments.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If the camera view is limited to a fixed area, then measurement precision for gaze direction is improved, but the ability to identify objects outside the camera view is lost

Engineering Contradiction:
Improvegaze detection precisionVSAvoidenvironmental information loss
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent divides the environment into multiple zones of interest and assigns different cameras to monitor specific regions. By segmenting the visual field and processing data from multiple camera perspectives, the system maintains precise gaze detection within each zone while simultaneously tracking objects across the entire environment. This segmentation approach allows the system to handle large spatial areas without sacrificing measurement precision.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If the system is designed for constrained environments, then ease of operation is improved, but applicability to general environments is reduced

Engineering Contradiction:
Improvesystem operation simplicityVSAvoidenvironmental versatility
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent designs a multi-functional system that can operate in both constrained and unconstrained environments by integrating multiple detection modalities (gaze tracking, gesture recognition, object detection) with a unified processing framework. The system automatically adapts its operation mode based on the environment type, maintaining ease of operation while achieving universal applicability across diverse settings.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3781896B1System for locating and identifying an object in unconstrained environments
Publication Date: 2024.02.14 TOYOTA RESEARCH INSTITUTE INC
  • EP3781896B1 patent drawingFigure 1
  • EP3781896B1 patent drawingFigure 2A~2B
  • EP3781896B1 patent drawingFigure 3A~3B

AI summary

A system for gaze and gesture detection in unconstrained environments includes a 360-degree (omnidirectional) camera system, one or more depth sensors, and associated memory, processors and programming instructions to determine an object of a human user's attention in the unconstrained environment. The illustrative system may identify the object using eye gaze, gesture detection, and/or speech recognition. The system may generate a saliency map and identify areas of interest. A directionality vector may be projected on the saliency map to find intersecting areas of interest. The system may identify the object of attention once the object of attention is located.