Directional Speech Separation via Time-Delay Cross-Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques for directional speech separation in electronic devices face challenges in accurately isolating desired speech from undesired speech and noise, particularly when dealing with localized sources like wireless loudspeakers, due to limitations in beamforming and acoustic echo cancellation methods.

Innovation Solution

The system dynamically determines directions of interest for audio sources by analyzing time delays between microphones, generates energy data, and uses cross-correlation to select relevant directions, creating time-frequency mask data for isolating specific audio sources, thereby improving speech separation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If beamforming and acoustic echo cancellation methods are used, then speech separation can be performed, but accuracy in isolating desired speech from undesired speech and noise deteriorates

Engineering Contradiction:
Improvespeech separation accuracyVSAvoidisolation effectiveness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the audio spectrum into multiple frequency bands and applies separate processing to each band. By dividing the speech separation task into frequency-specific segments, the system can more accurately isolate desired speech from noise and echo in each band, then combine the results to achieve overall improved separation accuracy and reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a frequency dimension to the speech separation process by analyzing and processing different frequency bands separately. This dimensional approach allows the system to exploit spectral characteristics of speech and noise, enabling more effective separation by treating each frequency band as an independent channel for applying beamforming and echo cancellation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If conventional speech separation techniques are used, then processing can be performed, but the ability to handle multi-source environments deteriorates

Engineering Contradiction:
Improvemulti-source environment handlingVSAvoidspeech isolation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements dynamic adaptation to multi-source environments by continuously analyzing the acoustic scene and adjusting processing parameters for each frequency band. The system dynamically identifies the number and characteristics of active speech sources, then adapts the beamforming and echo cancellation parameters accordingly, enabling versatile handling of varying multi-source scenarios while maintaining high isolation accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes processing parameters based on the detected acoustic environment and frequency band characteristics. By adjusting parameters such as beamforming weights, echo cancellation coefficients, and frequency band boundaries according to the specific multi-source scenario, the system achieves both adaptability to different environments and precision in speech isolation.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach effectively isolates audio data from individual sources, reducing noise and echo, and enhances the accuracy of speech recognition in multi-source environments.

Implementation Method 1

The system dynamically determines directions of interest for audio sources by analyzing time delays between microphones

Methodology Applied
Scientific EffectTime delay analysis:

Implementation Method 2

generates energy data, and uses cross-correlation to select relevant directions, creating time-frequency mask data

Methodology Applied
Scientific EffectCross-correlation:

Data Source

PatentUS11749294B2Directional speech separation
Publication Date: 2023.09.05 AMAZON TECH INC
  • US11749294B2 patent drawing
  • US11749294B2 patent drawing
  • US11749294B2 patent drawing

AI summary

A system configured to perform directional speech separation. The system may dynamically associate direction-of-arrivals with one or more audio sources in order to generate output audio data that separates each of the audio sources. The system identifies a target direction for each audio source, dynamically determines directions that are correlated with the target direction, and generates output signals for each audio source. The system may associate individual frequency bands with specific directions based on a time delay detected by two or more microphones. The system may determine a cross-correlation between each direction and the target direction and select directions with strong correlation. The system may generate time-frequency mask data indicating frequency bands corresponding to the directions associated with a particular audio source. Using the mask data, the system generates output audio data specific to the audio source, resulting in directional speech separation between different audio sources.