Voice Recognition System Using Eye Gaze Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems face challenges in accurately interpreting voice commands, especially in noisy environments and when multiple commands are given simultaneously, due to limitations in distinguishing user intentions without additional contextual information.
Innovation Solution
A voice recognition system that integrates a user interface, a camera for image capture, and a microphone to filter voice commands based on eye gaze and facial recognition, narrowing the search field for more accurate speech-to-text translation by correlating user visual focus and voice commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice recognition systems process voice commands without additional contextual information, then the system operation is simple, but the accuracy of voice command interpretation deteriorates in noisy environments and when multiple commands are given simultaneously
Solution Approach 1:
The patent combines voice recognition with eye gaze tracking and facial recognition systems into a unified interface control system. The controller integrates multiple input sources (microphone for voice, camera for gaze and facial recognition) to process commands more accurately by cross-referencing visual and auditory data, thereby improving interpretation accuracy while managing system complexity through coordinated multi-sensor operation
2Measurement precision
If the system uses eye gaze tracking to narrow down the search field for voice commands, then the accuracy of speech to text translation is improved, but the device complexity increases due to additional sensors and processing requirements
Solution Approach 1:
The controller acts as an intermediary that receives and processes data from both the camera (eye gaze tracking) and microphone (voice commands). It correlates the visual focus data with the voice command data to determine the intended target or function, thereby improving translation accuracy without requiring direct complex interaction between all system components
Solution Approach 2:
The system performs preliminary eye gaze tracking to identify the user's area of interest before processing the voice command. By pre-determining the contextual focus area through visual tracking, the system narrows down the possible interpretations of the voice command, improving accuracy while distributing processing loads across different time stages
Data Source
Figure 1~2
Figure 3~4
Figure 5~6
AI summary
A voice recognition system is provided with a user interface to display content, a camera to provide a first signal indicative of an image of a user viewing the content, and a microphone to provide a second signal indicative of a voice command that corresponds to a requested action. The voice recognition system is further provided with a controller that is programmed to receive the first and second signals, filter the voice command based on the image, and perform the requested action based on the filtered voice command.