Microphone Array Voice Separation Using Arrival-Direction Phase Masks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice separation devices using the Expectation Maximization (EM) algorithm for updating probability distribution models struggle with inaccurate voice separation.

Innovation Solution

A sound collection device that performs Fourier transforms on input signals from multiple microphones, estimates arrival directions, calculates cross-spectrum phases, determines mask coefficients using a pre-generated database, and applies inverse Fourier transforms to separate target voices with high accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the EM algorithm is used for updating the parameter of the probability distribution model, then the device can perform voice separation, but the voice separation accuracy deteriorates in certain cases

Engineering Contradiction:
Improvevoice separation accuracyVSAvoidseparation accuracy consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The invention changes the fundamental parameters of the approach by abandoning the probabilistic EM algorithm in favor of deterministic time-frequency analysis. Specifically, it transforms voice separation from a statistical parameter optimization problem to a phase-difference-based signal processing problem, where mask coefficients are derived directly from cross-spectrum phase relationships rather than through iterative probability model updates

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The invention substitutes the mechanical/iterative EM algorithm process with a direct computational approach based on Fourier transform and phase analysis. Instead of repeatedly updating probability distribution parameters, the system calculates mask coefficients through deterministic mathematical operations on the time-frequency representation of the signals, replacing an iterative optimization mechanism with a direct computational formula

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If complex probability distribution models are used for voice separation, then separation capability is achieved, but computational complexity increases

Engineering Contradiction:
Improvevoice separation capabilityVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The invention extracts only the essential feature needed for voice separation—the phase difference in the time-frequency domain—while discarding the complex probabilistic modeling framework. By focusing solely on phase relationship extraction from cross-spectra and using this to generate mask coefficients, the system achieves voice separation without the computational burden of maintaining and updating full probability distribution models

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The invention segments the voice separation problem into distinct frequency bins and time frames, processing each independently through Fourier transform. This segmentation allows the system to apply simple phase-difference-based mask calculations to each segment rather than attempting a complex global probabilistic model, reducing overall computational complexity while maintaining separation capability

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12477272B2Sound collection device, sound collection method, and storage medium storing sound collection program
Publication Date: 2025.11.18 MITSUBISHI ELECTRIC CORP
  • US12477272B2 patent drawing
  • US12477272B2 patent drawing
  • US12477272B2 patent drawing

AI summary

A sound collection device performs Fourier transform on a first sound reception signal outputted from a first microphone and outputs a first signal, performs Fourier transform on a second sound reception signal outputted from a second microphone and outputs a second signal, estimates an arrival direction of the voice, calculates a phase of a cross-spectrum of the first signal and the second signal, determines a mask coefficient based on an arrival direction phase table indicating a relationship between the phase and the arrival direction regarding each frequency band, the calculated phase, and the estimated arrival direction, separates a signal from the first signal or the second signal by using the mask coefficient, and performs inverse Fourier transform on the separated signal and outputs a signal of a target voice.