Microphone Array Voice Separation Using Arrival-Direction Phase Masks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice separation devices using the Expectation Maximization (EM) algorithm for updating probability distribution models struggle with inaccurate voice separation.
Innovation Solution
A sound collection device that performs Fourier transforms on input signals from multiple microphones, estimates arrival directions, calculates cross-spectrum phases, determines mask coefficients using a pre-generated database, and applies inverse Fourier transforms to separate target voices with high accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the EM algorithm is used for updating the parameter of the probability distribution model, then the device can perform voice separation, but the voice separation accuracy deteriorates in certain cases
Solution Approach 1:
The invention changes the fundamental parameters of the approach by abandoning the probabilistic EM algorithm in favor of deterministic time-frequency analysis. Specifically, it transforms voice separation from a statistical parameter optimization problem to a phase-difference-based signal processing problem, where mask coefficients are derived directly from cross-spectrum phase relationships rather than through iterative probability model updates
Solution Approach 2:
The invention substitutes the mechanical/iterative EM algorithm process with a direct computational approach based on Fourier transform and phase analysis. Instead of repeatedly updating probability distribution parameters, the system calculates mask coefficients through deterministic mathematical operations on the time-frequency representation of the signals, replacing an iterative optimization mechanism with a direct computational formula
2Measurement precision
If complex probability distribution models are used for voice separation, then separation capability is achieved, but computational complexity increases
Solution Approach 1:
The invention extracts only the essential feature needed for voice separation—the phase difference in the time-frequency domain—while discarding the complex probabilistic modeling framework. By focusing solely on phase relationship extraction from cross-spectra and using this to generate mask coefficients, the system achieves voice separation without the computational burden of maintaining and updating full probability distribution models
Solution Approach 2:
The invention segments the voice separation problem into distinct frequency bins and time frames, processing each independently through Fourier transform. This segmentation allows the system to apply simple phase-difference-based mask calculations to each segment rather than attempting a complex global probabilistic model, reducing overall computational complexity while maintaining separation capability
Data Source
AI summary
A sound collection device performs Fourier transform on a first sound reception signal outputted from a first microphone and outputs a first signal, performs Fourier transform on a second sound reception signal outputted from a second microphone and outputs a second signal, estimates an arrival direction of the voice, calculates a phase of a cross-spectrum of the first signal and the second signal, determines a mask coefficient based on an arrival direction phase table indicating a relationship between the phase and the arrival direction regarding each frequency band, the calculated phase, and the estimated arrival direction, separates a signal from the first signal or the second signal by using the mask coefficient, and performs inverse Fourier transform on the separated signal and outputs a signal of a target voice.


