Speech Recognition Device Using Camera for Gesture Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies fail to detect and interpret hand gestures that may accompany speech input, leading to incomplete or inaccurate speech recognition.

Innovation Solution

A speech recognition system that integrates a camera unit to capture hand gestures and a processor to link these gestures with speech input, using a database to match the gestures with intended meanings, enhancing recognition accuracy when speech alone is not sufficient.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition relies solely on speech input, then the system is simple, but recognition accuracy deteriorates when speech is ambiguous

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines speech input processing with hand gesture recognition into a unified recognition system. The speech acquiring unit and image acquiring unit work together, with the processor integrating both inputs to determine intended meaning, thereby improving accuracy while managing complexity through integrated design

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The recognition system is designed to handle multiple types of input (speech and hand gestures) through a single processor that can interpret both modalities. The database stores meanings associated with both speech and gesture patterns, allowing the system to universally process different input types for speech recognition

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If hand gestures are added to speech input processing, then recognition accuracy improves, but device complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges speech processing and gesture recognition into a single integrated system where the processor handles both input types. The database unifies the storage of meanings for both speech and gestures, reducing overall system complexity through consolidation while improving recognition accuracy

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The processor is designed with multi-functionality to interpret both speech inputs and hand gesture images. The database universally stores meanings that can be accessed regardless of whether the input is speech or gesture-based, allowing the system to maintain simplicity while handling multiple input modalities

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10714088B2Speech recognition device and method of identifying speech
Publication Date: 2020.07.14 HON HAI PRECISION INDUSTRY CO LTD
  • US10714088B2 patent drawing
  • US10714088B2 patent drawing
  • US10714088B2 patent drawing

AI summary

A speech recognition device includes a speech acquiring unit, a speech outputting unit, a camera unit, and a processor. The processor obtains speech input acquired by the speech acquiring unit, obtains images acquired by the camera unit, chronologically links the obtained speech and the obtained images together, compares the obtained speech to a speech database to confirm matching speech and a confidence level of the matching speech, determines whether the confidence level of the matching speech exceeds a predetermined confidence level, and outputs the matching speech when the confidence level of the matching speech exceeds the predetermined confidence level.