Sound Source Separation via Complex Conjugate Beamforming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sound source separation technologies face challenges in effectively separating target sound sources from multiple sound sources in environments with uneven microphone sensitivity, leading to degraded speech recognition performance, especially in automotive settings where multiple speakers and background noise are present.

Innovation Solution

A sound source separation device using at least two microphones performs beamforming processing with complex conjugate coefficients to attenuate sound sources from specific directions, computes power spectrum information, and extracts target sound spectrum information based on differences, thereby compensating for uneven microphone sensitivities and improving separation performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple microphones are used to separate sound sources, then speech recognition accuracy improves, but the cost and complexity of the system increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex hardware systems with software-based signal processing. Specifically, it uses beamforming algorithms and spectral subtraction techniques to achieve sound source separation using only 2 microphones, substituting the need for complex acoustic hardware with computational methods. This resolves the contradiction by maintaining high speech recognition accuracy while minimizing device complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the processing domain from time-domain to frequency-domain analysis. By performing Fast Fourier Transform (FFT) and processing signals in the frequency domain, the system can effectively separate sound sources using simple 2-microphone arrays. This parameter change enables high accuracy speech recognition without requiring complex hardware configurations.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If low-cost microphones are used, then system cost decreases, but uneven sensitivity characteristics degrade performance

Engineering Contradiction:
Improvesystem costVSAvoidsound source separation performance
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent converts the harmful effect of uneven microphone sensitivity into a beneficial feature. By using spectral subtraction and adaptive filtering, the system identifies and exploits the consistent phase relationships between microphones, even when their sensitivities differ. This allows low-cost microphones with uneven frequency characteristics (±3 dB) to achieve high performance in sound source separation.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent introduces signal processing algorithms as intermediaries between the microphones and the speech recognition system. The beamforming and spectral subtraction processing acts as a mediator that compensates for microphone imperfections, allowing low-cost microphones to deliver high-quality speech recognition performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Object-affected harmful factors

If adaptive beamforming is used to suppress noise, then noise reduction improves, but target signals from unexpected directions are also suppressed

Engineering Contradiction:
Improvenoise reductionVSAvoidtarget signal preservation
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent employs dynamic adaptive filtering that continuously adjusts its parameters based on the acoustic environment. The system uses least mean squares (LMS) adaptation to dynamically update filter coefficients, allowing it to track moving speakers and maintain reliable target signal preservation while effectively suppressing stationary noise. This dynamic approach resolves the contradiction by making the noise suppression adaptive rather than fixed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements feedback mechanisms where the output of the beamforming process is fed back into the adaptive filter for continuous refinement. This feedback loop allows the system to learn from its performance and adjust its noise suppression characteristics in real-time, ensuring that target signals from unexpected directions are preserved while maintaining effective noise reduction.

Inventive Principle:
Principle #23Feedback

4Measurement precision

If many microphones are installed in car cabin, then sound source separation performance improves, but space requirements and cost increase

Engineering Contradiction:
Improvesound source separation performanceVSAvoidinstallation space
Core Design Contradiction:
Measurement precisionVSArea of stationary object

Solution Approach 1:

The patent extracts and utilizes only the essential spatial information needed for sound source separation. By focusing on the phase and amplitude relationships between just 2 microphones and using frequency-domain processing, the system achieves high performance without requiring multiple microphones distributed throughout the car cabin. This extraction approach eliminates the need for extensive installation space while maintaining separation performance.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS8112272B2Sound source separation device, speech recognition device, mobile telephone, sound source separation method, and program
Publication Date: 2012.02.07 ASAHI KASEI KOGYO KABUSHIKI KAISHA
  • US8112272B2 patent drawing
  • US8112272B2 patent drawing
  • US8112272B2 patent drawing

AI summary

A sound source signal from a target sound source is allowed to be separated from a mixed sound which consists of sound source signals emitted from a plurality of sound sources without being affected by uneven sensitivity of microphone elements. A beamformer section 3 of a source separation device 1 performs beamforming processing for attenuating sound source signals arriving from directions symmetrical with respect to a perpendicular line to a straight line connecting two microphones 10 and 11 respectively by multiplying output signals from the microphones 10 and 11 after spectrum analysis by weighted coefficients which are complex conjugate to each other. Power computation sections 40 and 41 compute power spectrum information, and target sound spectrum extraction sections 50 and 51 extract spectrum information of a target sound source based on a difference between the power spectrum information.