Neural Network Speech Separation Using Phase-Aware Masks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated speech separation systems struggle to effectively separate audio signals from multiple speakers, especially in scenarios where the number of speakers changes or speech overlaps, without prior knowledge of the number of speakers.

Innovation Solution

A method and system that utilize multiple microphones to capture mixed speech signals, extract features, and input them into a speech separation model to generate time-frequency masks, allowing for the separation of speaker-specific signals without prior knowledge of the number of speakers, using a neural network or other machine-learning models that incorporate phase information and inter-microphone phase differences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If automated speech separation is performed without prior knowledge of the number of speakers, then the system can handle dynamic scenarios where speakers enter or leave, but the accuracy of speaker attribution deteriorates

Engineering Contradiction:
Improveability to handle dynamic speaker scenariosVSAvoidspeaker attribution accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the mixed audio signal into multiple speaker-specific signals by training the neural network to output a specified number of separated signals. The system divides the complex task of speech separation into manageable components by processing each speaker's signal separately through the network, enabling accurate attribution even when the number of speakers varies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of the neural network output to produce a fixed number of speaker-specific signals regardless of the actual number of speakers in the input. By adjusting the network configuration to output a predetermined number of signals, the system maintains measurement precision while adapting to dynamic speaker scenarios.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If multiple microphones are used to capture speech signals, then the robustness of speech separation in overlapping speech scenarios is improved, but the device complexity increases

Engineering Contradiction:
Improvespeech separation robustnessVSAvoidmicrophone array complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges the signals from multiple microphones into a unified processing pipeline. By combining the inputs from the microphone array and processing them together through the neural network, the system achieves robust speech separation without requiring complex post-processing for each individual microphone signal.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal speech separation system that handles multiple functions: capturing signals from multiple microphones, separating overlapping speech, attributing speakers, and processing dynamic speaker scenarios. The neural network is designed to perform all these functions through a single integrated architecture, reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If neural network models incorporating phase information are used for speech separation, then the accuracy of separating overlapping speech is improved, but the computational processing time increases

Engineering Contradiction:
Improveoverlapping speech separation accuracyVSAvoidsignal processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing of the audio signals by transforming them into the frequency domain and extracting phase information before inputting to the neural network. By preparing the phase information in advance, the system enables accurate overlapping speech separation while optimizing the processing time through efficient pre-computation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3776535B1Multi-microphone speech separation
Publication Date: 2023.06.14 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3776535B1 patent drawingFigure 1
  • EP3776535B1 patent drawingFigure 2
  • EP3776535B1 patent drawingFigure 3

AI summary

This document relates to separation of audio signals into speaker-specific signals. One example obtains features reflecting mixed speech signals captured by multiple microphones. The features can be input a neural network and masks can be obtained from the neural network. The masks can be applied one or more of the mixed speech signals captured by one or more of the microphones to obtain two or more separate speaker-specific speech signals, which can then be output.