Neural Network Speech Separation Using Phase-Aware Masks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated speech separation systems struggle to effectively separate audio signals from multiple speakers, especially in scenarios where the number of speakers changes or speech overlaps, without prior knowledge of the number of speakers.
Innovation Solution
A method and system that utilize multiple microphones to capture mixed speech signals, extract features, and input them into a speech separation model to generate time-frequency masks, allowing for the separation of speaker-specific signals without prior knowledge of the number of speakers, using a neural network or other machine-learning models that incorporate phase information and inter-microphone phase differences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If automated speech separation is performed without prior knowledge of the number of speakers, then the system can handle dynamic scenarios where speakers enter or leave, but the accuracy of speaker attribution deteriorates
Solution Approach 1:
The patent segments the mixed audio signal into multiple speaker-specific signals by training the neural network to output a specified number of separated signals. The system divides the complex task of speech separation into manageable components by processing each speaker's signal separately through the network, enabling accurate attribution even when the number of speakers varies.
Solution Approach 2:
The patent changes the parameter of the neural network output to produce a fixed number of speaker-specific signals regardless of the actual number of speakers in the input. By adjusting the network configuration to output a predetermined number of signals, the system maintains measurement precision while adapting to dynamic speaker scenarios.
2Reliability
If multiple microphones are used to capture speech signals, then the robustness of speech separation in overlapping speech scenarios is improved, but the device complexity increases
Solution Approach 1:
The patent merges the signals from multiple microphones into a unified processing pipeline. By combining the inputs from the microphone array and processing them together through the neural network, the system achieves robust speech separation without requiring complex post-processing for each individual microphone signal.
Solution Approach 2:
The patent creates a universal speech separation system that handles multiple functions: capturing signals from multiple microphones, separating overlapping speech, attributing speakers, and processing dynamic speaker scenarios. The neural network is designed to perform all these functions through a single integrated architecture, reducing overall system complexity.
3Measurement precision
If neural network models incorporating phase information are used for speech separation, then the accuracy of separating overlapping speech is improved, but the computational processing time increases
Solution Approach 1:
The patent performs preliminary processing of the audio signals by transforming them into the frequency domain and extracting phase information before inputting to the neural network. By preparing the phase information in advance, the system enables accurate overlapping speech separation while optimizing the processing time through efficient pre-computation.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
This document relates to separation of audio signals into speaker-specific signals. One example obtains features reflecting mixed speech signals captured by multiple microphones. The features can be input a neural network and masks can be obtained from the neural network. The masks can be applied one or more of the mixed speech signals captured by one or more of the microphones to obtain two or more separate speaker-specific speech signals, which can then be output.