Voice Recognition System Using Eye Gaze Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition systems face challenges in accurately interpreting voice commands, especially in noisy environments and when multiple commands are given simultaneously, due to limitations in distinguishing user intentions without additional contextual information.

Innovation Solution

A voice recognition system that integrates a user interface, a camera for image capture, and a microphone to filter voice commands based on eye gaze and facial recognition, narrowing the search field for more accurate speech-to-text translation by correlating user visual focus and voice commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice recognition systems process voice commands without additional contextual information, then the system operation is simple, but the accuracy of voice command interpretation deteriorates in noisy environments and when multiple commands are given simultaneously

Engineering Contradiction:
Improveaccuracy of voice command interpretationVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines voice recognition with eye gaze tracking and facial recognition systems into a unified interface control system. The controller integrates multiple input sources (microphone for voice, camera for gaze and facial recognition) to process commands more accurately by cross-referencing visual and auditory data, thereby improving interpretation accuracy while managing system complexity through coordinated multi-sensor operation

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If the system uses eye gaze tracking to narrow down the search field for voice commands, then the accuracy of speech to text translation is improved, but the device complexity increases due to additional sensors and processing requirements

Engineering Contradiction:
Improveaccuracy of speech to text translationVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The controller acts as an intermediary that receives and processes data from both the camera (eye gaze tracking) and microphone (voice commands). It correlates the visual focus data with the voice command data to determine the intended target or function, thereby improving translation accuracy without requiring direct complex interaction between all system components

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary eye gaze tracking to identify the user's area of interest before processing the voice command. By pre-determining the contextual focus area through visual tracking, the system narrows down the possible interpretations of the voice command, improving accuracy while distributing processing loads across different time stages

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3563373B1Voice recognition system
Publication Date: 2022.11.30 HARMAN INT IND INC
  • EP3563373B1 patent drawingFigure 1~2
  • EP3563373B1 patent drawingFigure 3~4
  • EP3563373B1 patent drawingFigure 5~6

AI summary

A voice recognition system is provided with a user interface to display content, a camera to provide a first signal indicative of an image of a user viewing the content, and a microphone to provide a second signal indicative of a voice command that corresponds to a requested action. The voice recognition system is further provided with a controller that is programmed to receive the first and second signals, filter the voice command based on the image, and perform the requested action based on the filtered voice command.