Voice Processing Apparatus Using Visual-Audio Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition techniques fail to accurately distinguish user voice instructions from ambient noise, leading to recognition errors, especially when ambient sounds like reflections and passer-by voices interfere with the signal.

Innovation Solution

A voice processing apparatus and method that utilizes a combination of sound and video signals to position the voice source, employing a processor, sound receiver, and camera to detect a voice onset time, human face, and mouth contour changes, verifying preset conditions such as time differences and angle alignments to isolate the user's voice from ambient noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech recognition is performed on sound signal received by microphone, then voice instruction can be recognized in case of no ambient noise, but recognition accuracy deteriorates when ambient noise exists

Engineering Contradiction:
Improvevoice instruction recognition accuracyVSAvoidambient noise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the sound signal processing into multiple independent verification stages: (1) voice onset time detection, (2) face detection, (3) mouth contour change detection, and (4) temporal synchronization verification. By dividing the recognition process into these separate modules, the system can independently verify each condition and only proceed with recognition when all conditions are satisfied, thereby filtering out ambient noise interference while maintaining accurate voice instruction recognition.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If multiple verification conditions are added to distinguish user voice from ambient noise, then recognition accuracy improves, but system complexity increases

Engineering Contradiction:
Improvevoice source identification accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent employs a camera that serves multiple functions: it detects the user's face position, tracks mouth contour changes, and provides temporal synchronization data. This single device performs multiple verification tasks that would otherwise require separate sensors, thereby improving voice source identification accuracy while minimizing the increase in system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces complex acoustic analysis methods with a simpler visual-based verification system. Instead of using sophisticated signal processing to distinguish voice sources acoustically, the system uses camera-based face and mouth detection combined with temporal synchronization to verify the sound source, substituting a more straightforward optical measurement approach for complex mechanical/acoustic analysis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9520131B2Apparatus and method for voice processing
Publication Date: 2016.12.13 WISTRON CORP
  • US9520131B2 patent drawing
  • US9520131B2 patent drawing
  • US9520131B2 patent drawing

AI summary

An apparatus and a corresponding method for voice processing are provided. The apparatus includes a sound receiver, a camera, and a processor. The sound receiver receives a sound signal. The camera takes a video. The processor is coupled to the sound receiver and the camera. The processor obtains a voice onset time (VOT) of the sound signal, detects a human face in the video, detects a change time of a mouth contour of the human face, and verifies at least one preset condition. When all of the preset conditions are true, the processor performs speech recognition on the sound signal. The at least one preset condition includes that a difference between the VOT and the change time is smaller than a threshold value.