Speech Recognition Device Using Camera for Gesture Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies fail to detect and interpret hand gestures that may accompany speech input, leading to incomplete or inaccurate speech recognition.
Innovation Solution
A speech recognition system that integrates a camera unit to capture hand gestures and a processor to link these gestures with speech input, using a database to match the gestures with intended meanings, enhancing recognition accuracy when speech alone is not sufficient.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition relies solely on speech input, then the system is simple, but recognition accuracy deteriorates when speech is ambiguous
Solution Approach 1:
The patent combines speech input processing with hand gesture recognition into a unified recognition system. The speech acquiring unit and image acquiring unit work together, with the processor integrating both inputs to determine intended meaning, thereby improving accuracy while managing complexity through integrated design
Solution Approach 2:
The recognition system is designed to handle multiple types of input (speech and hand gestures) through a single processor that can interpret both modalities. The database stores meanings associated with both speech and gesture patterns, allowing the system to universally process different input types for speech recognition
2Measurement precision
If hand gestures are added to speech input processing, then recognition accuracy improves, but device complexity increases
Solution Approach 1:
The patent merges speech processing and gesture recognition into a single integrated system where the processor handles both input types. The database unifies the storage of meanings for both speech and gestures, reducing overall system complexity through consolidation while improving recognition accuracy
Solution Approach 2:
The processor is designed with multi-functionality to interpret both speech inputs and hand gesture images. The database universally stores meanings that can be accessed regardless of whether the input is speech or gesture-based, allowing the system to maintain simplicity while handling multiple input modalities
Data Source
AI summary
A speech recognition device includes a speech acquiring unit, a speech outputting unit, a camera unit, and a processor. The processor obtains speech input acquired by the speech acquiring unit, obtains images acquired by the camera unit, chronologically links the obtained speech and the obtained images together, compares the obtained speech to a speech database to confirm matching speech and a confidence level of the matching speech, determines whether the confidence level of the matching speech exceeds a predetermined confidence level, and outputs the matching speech when the confidence level of the matching speech exceeds the predetermined confidence level.


