Voice Signal Processing Device for Acoustic Echo Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional hands-free call devices struggle to accurately determine whether a user is speaking, leading to ineffective suppression of acoustic echoes due to incorrect mouth movement detection and erroneous voice signal amplification.
Innovation Solution
A voice signal processing device that combines image acquisition from a camera with microphone array data to estimate whether a near-end talker is uttering by matching the talker's direction and voice arrival direction, adjusting the amplification factor accordingly to suppress acoustic echoes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If mouth movement detection is used to determine whether a user is speaking, then the system can identify talker presence, but the detection accuracy is insufficient leading to erroneous voice signal amplification
Solution Approach 1:
The patent combines camera-based image processing and microphone array-based voice signal processing into a unified talker detection system. The image acquisition unit captures facial images, the voice signal acquisition unit collects audio signals, and the talker determination unit integrates both data streams to make accurate talker presence decisions, resolving the contradiction between detection accuracy and echo suppression reliability
Solution Approach 2:
The talker determination unit acts as an intermediary that mediates between the image processing results and voice signal processing results. It receives both types of data, compares them, and produces a unified talker presence determination that is more reliable than either method alone, thereby improving both detection accuracy and echo suppression reliability
2Productivity
If voice signal amplification is applied without accurate talker detection, then the system can transmit audio, but acoustic echoes are not effectively suppressed
Solution Approach 1:
The system implements feedback control where the talker determination unit continuously monitors both image and voice signal data, and the voice signal processing unit adjusts amplification in real-time based on this feedback. When talker presence is confirmed through both modalities, amplification is applied; when not confirmed, amplification is reduced or suppressed, effectively controlling acoustic echoes while maintaining transmission efficiency
Solution Approach 2:
The voice signal processing unit dynamically changes the amplification parameter based on the talker determination results. The amplification factor is adjusted from high values (when talker is detected) to low or zero values (when talker is not detected), thereby controlling the generation and suppression of acoustic echoes while maintaining voice transmission quality
Data Source
AI summary
A voice signal processing device detects a near-end talker included in an image, specifies a direction in which the near-end talker is located, detects a voice arrival direction based on a voice signal collected by a microphone array, estimates whether the near-end talker is uttering, based on the direction in which the near-end talker is located and the voice arrival direction, sets an amplification factor of the voice signal to a value equal to or greater than 1 in a case where the near-end talker is estimated to be uttering, sets the amplification factor of the voice signal to a value smaller than 1 in a case where the near-end talker is estimated not to be uttering, and adjusts a level of the voice signal based on the set amplification factor, and outputs the adjusted voice signal as a transmission signal to be transmitted to a far-end talker.


