Voice Recognition Device Using Visual Trigger Events
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition devices face challenges in accurately determining the desired utterance section in noisy environments, as they struggle to differentiate between the user's voice and ambient noise, and rely on unreliable methods such as lip motion analysis or user operation, which fail when the user is not directly interacting with the device.
Innovation Solution
A voice recognition system that uses a combination of microphone arrays and camera inputs to determine the voice source direction and section by analyzing phase differences in sound signals and visual triggers, such as face direction and posture, to enhance the target sound and reduce ambient noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If noise reduction techniques such as beam forming or echo cancellation are used, then voice recognition accuracy is improved, but it is still difficult to achieve sufficient voice recognition accuracy in noisy environments
Solution Approach 1:
The patent introduces image information from a camera as an intermediary to assist in determining the voice section. The image processing unit detects face direction and posture, which serve as visual cues to identify when the user is actively speaking. This visual intermediary complements the audio-based beam forming technique, allowing the system to more accurately distinguish the user's voice from ambient noise by cross-referencing both audio and visual data.
2Measurement precision
If lip motion detection is used to determine utterance section, then voice section identification is improved, but inaccurate detection occurs when unrelated motions such as gum chewing are made
Solution Approach 1:
The system employs feedback by continuously monitoring both audio and visual data streams and adjusting the voice section determination based on multiple indicators. The voice section determination unit receives input from both the audio processing unit (which analyzes sound characteristics) and the image processing unit (which analyzes face direction and posture). This feedback mechanism allows the system to distinguish between genuine speaking motions and unrelated motions like gum chewing by looking for consistent patterns across both modalities.
3Measurement precision
If user operation is required to determine voice section, then accurate voice section identification is achieved, but the method becomes unusable when the user is apart from the device
Solution Approach 1:
The system implements self-service by automatically detecting the voice section through analysis of audio and visual data without requiring any user operation. The voice section determination unit autonomously identifies when the user is speaking by analyzing sound characteristics from the microphone array and visual cues from the camera, such as face direction and posture changes. This eliminates the need for users to manually press buttons or interact with the device, allowing accurate voice section identification even when the user is at a distance.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly improves voice recognition accuracy by accurately identifying the voice source direction and section, even in noisy conditions, reducing the impact of ambient noise and eliminating the need for direct user interaction.
Implementation Method 1
analyzing phase differences in sound signals
Implementation Method 2
acquisition sound acquired by a microphone includes various kinds of noises
Data Source
Figure 1
Figure 2
Figure 3
AI summary
By recognizing visual trigger events to determine start points and/or end points of voice data signals, the negative effects of noise on voice recognition may be significantly minimized. The visual trigger events may be predetermined gestures and/or predetermined postures of a user captured by a camera, which allow a system to appropriately focus attention on a user to optimize the receipt of a voice command in a noisy environment. This may be accomplished through the assistance of visual feedback complementing the voice feedback provided to the system by the user. Since the visual trigger events are predetermined gestures and/or postures, the system may be able to distinguish which sounds produced by a user are voice commands and which sounds produced by the user is noise that in unrelated to the operation of the system.