Voice Signal Processing Device for Acoustic Echo Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional hands-free call devices struggle to accurately determine whether a user is speaking, leading to ineffective suppression of acoustic echoes due to incorrect mouth movement detection and erroneous voice signal amplification.

Innovation Solution

A voice signal processing device that combines image acquisition from a camera with microphone array data to estimate whether a near-end talker is uttering by matching the talker's direction and voice arrival direction, adjusting the amplification factor accordingly to suppress acoustic echoes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If mouth movement detection is used to determine whether a user is speaking, then the system can identify talker presence, but the detection accuracy is insufficient leading to erroneous voice signal amplification

Engineering Contradiction:
Improvetalker detection accuracyVSAvoidacoustic echo suppression reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent combines camera-based image processing and microphone array-based voice signal processing into a unified talker detection system. The image acquisition unit captures facial images, the voice signal acquisition unit collects audio signals, and the talker determination unit integrates both data streams to make accurate talker presence decisions, resolving the contradiction between detection accuracy and echo suppression reliability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The talker determination unit acts as an intermediary that mediates between the image processing results and voice signal processing results. It receives both types of data, compares them, and produces a unified talker presence determination that is more reliable than either method alone, thereby improving both detection accuracy and echo suppression reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If voice signal amplification is applied without accurate talker detection, then the system can transmit audio, but acoustic echoes are not effectively suppressed

Engineering Contradiction:
Improvevoice transmission efficiencyVSAvoidacoustic echo
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The system implements feedback control where the talker determination unit continuously monitors both image and voice signal data, and the voice signal processing unit adjusts amplification in real-time based on this feedback. When talker presence is confirmed through both modalities, amplification is applied; when not confirmed, amplification is reduced or suppressed, effectively controlling acoustic echoes while maintaining transmission efficiency

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The voice signal processing unit dynamically changes the amplification parameter based on the talker determination results. The amplification factor is adjusted from high values (when talker is detected) to low or zero values (when talker is not detected), thereby controlling the generation and suppression of acoustic echoes while maintaining voice transmission quality

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240386901A1Voice signal processing device, voice signal processing method, and non-transitory computer readable recording medium storing voice signal processing program
Publication Date: 2024.11.21 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US20240386901A1 patent drawing
  • US20240386901A1 patent drawing
  • US20240386901A1 patent drawing

AI summary

A voice signal processing device detects a near-end talker included in an image, specifies a direction in which the near-end talker is located, detects a voice arrival direction based on a voice signal collected by a microphone array, estimates whether the near-end talker is uttering, based on the direction in which the near-end talker is located and the voice arrival direction, sets an amplification factor of the voice signal to a value equal to or greater than 1 in a case where the near-end talker is estimated to be uttering, sets the amplification factor of the voice signal to a value smaller than 1 in a case where the near-end talker is estimated not to be uttering, and adjusts a level of the voice signal based on the set amplification factor, and outputs the adjusted voice signal as a transmission signal to be transmitted to a far-end talker.