Spatial Double-Talk Detection Using Microphone Array Beamforming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Audio systems, particularly in vehicles, face challenges in distinguishing between user voice and acoustic echoes or noise, leading to suboptimal performance in echo reduction and noise suppression, as existing double-talk detection methods are not robust enough to handle complex environments with multiple audio sources and users.

Innovation Solution

The implementation of a double-talk detection system that uses a microphone array for spatial discrimination, employing beamforming techniques and spectral analysis to differentiate between user speech and other acoustic sources, thereby providing an accurate indication of user voice activity to control adaptive echo reduction algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional double-talk detection methods are used, then the system is simple to implement, but the detection accuracy is insufficient in complex environments with multiple audio sources

Engineering Contradiction:
Improvedetection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio signal processing into multiple independent components: beamforming module for spatial filtering, spectral analysis module for frequency-domain processing, and double-talk detection module for voice activity detection. Each module processes specific aspects of the audio signal independently, improving detection accuracy while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatial dimension analysis by using a microphone array to capture audio signals from multiple positions. The beamforming module processes signals in the spatial domain to create directional sensitivity, adding a spatial dimension to the traditional temporal and spectral analysis. This multi-dimensional approach significantly improves detection accuracy in complex acoustic environments.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If adaptive components are continuously updated, then the echo reduction performance is optimized, but the user voice signal is distorted when the user speaks

Engineering Contradiction:
Improveecho reduction performanceVSAvoidvoice signal distortion
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent implements a feedback mechanism where the double-talk detection module continuously monitors the audio signal and provides control feedback to the adaptive echo reduction components. When user speech is detected, the system receives feedback to pause adaptive updates, preventing distortion of the user's voice signal while maintaining optimized echo reduction performance during non-speech periods.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent makes the adaptive components dynamic by enabling them to switch between update and pause states based on real-time detection conditions. The system dynamically adjusts its behavior: continuously updating adaptive filters during non-speech periods for optimal echo reduction, and pausing updates during speech periods to preserve voice signal integrity. This dynamic adaptation resolves the contradiction between performance optimization and signal protection.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3692704B1Spatial double-talk detector
Publication Date: 2023.09.06 BOSE CORP
  • EP3692704B1 patent drawingFigure 1
  • EP3692704B1 patent drawingFigure 2
  • EP3692704B1 patent drawingFigure 3

AI summary

Double-talk detection systems and methods are provided that include a plurality of microphones and array processing to detect when a vehicle occupant is speaking. One or more array processors combine the microphone signals to provide a primary signal and a reference signal. The responses of the primary signal and the reference signal are different in the direction of an occupant location. A comparison of energy content is made between the primary signal and reference signal to selectively indicate double-talk based on the comparison.