Voice Processing Apparatus Using Visual-Audio Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition techniques fail to accurately distinguish user voice instructions from ambient noise, leading to recognition errors, especially when ambient sounds like reflections and passer-by voices interfere with the signal.
Innovation Solution
A voice processing apparatus and method that utilizes a combination of sound and video signals to position the voice source, employing a processor, sound receiver, and camera to detect a voice onset time, human face, and mouth contour changes, verifying preset conditions such as time differences and angle alignments to isolate the user's voice from ambient noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech recognition is performed on sound signal received by microphone, then voice instruction can be recognized in case of no ambient noise, but recognition accuracy deteriorates when ambient noise exists
Solution Approach 1:
The patent segments the sound signal processing into multiple independent verification stages: (1) voice onset time detection, (2) face detection, (3) mouth contour change detection, and (4) temporal synchronization verification. By dividing the recognition process into these separate modules, the system can independently verify each condition and only proceed with recognition when all conditions are satisfied, thereby filtering out ambient noise interference while maintaining accurate voice instruction recognition.
2Measurement precision
If multiple verification conditions are added to distinguish user voice from ambient noise, then recognition accuracy improves, but system complexity increases
Solution Approach 1:
The patent employs a camera that serves multiple functions: it detects the user's face position, tracks mouth contour changes, and provides temporal synchronization data. This single device performs multiple verification tasks that would otherwise require separate sensors, thereby improving voice source identification accuracy while minimizing the increase in system complexity.
Solution Approach 2:
The patent replaces complex acoustic analysis methods with a simpler visual-based verification system. Instead of using sophisticated signal processing to distinguish voice sources acoustically, the system uses camera-based face and mouth detection combined with temporal synchronization to verify the sound source, substituting a more straightforward optical measurement approach for complex mechanical/acoustic analysis.
Data Source
AI summary
An apparatus and a corresponding method for voice processing are provided. The apparatus includes a sound receiver, a camera, and a processor. The sound receiver receives a sound signal. The camera takes a video. The processor is coupled to the sound receiver and the camera. The processor obtains a voice onset time (VOT) of the sound signal, detects a human face in the video, detects a change time of a mouth contour of the human face, and verifies at least one preset condition. When all of the preset conditions are true, the processor performs speech recognition on the sound signal. The at least one preset condition includes that a difference between the VOT and the change time is smaller than a threshold value.


