Mask Calculation Neural Network for Target Speaker Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional techniques fail to effectively extract the speech of a target speaker from observed speech that includes multiple speakers, as they incorrectly identify non-target speaker speech as noise due to similar spectro-temporal characteristics.
Innovation Solution
A mask calculation device and method that includes a feature extractor, a mask calculator, and an object signal calculator, using a neural network with clustered layers to calculate weights and masks for extracting the target speaker's signal from observed speech, employing error backpropagation and weight updates to optimize the extraction process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional techniques treat speech other than target speaker speech as noise, then extraction can work for single speaker plus noise scenarios, but the techniques fail when multiple speakers are present because their spectro-temporal characteristics are similar
Solution Approach 1:
The patent segments the speech signal processing into distinct functional modules: a feature extractor that extracts spectro-temporal features from the mixed speech signal, a mask calculator that computes separation masks based on these features, and an object signal calculator that reconstructs the target speaker speech. This segmentation allows each module to specialize in specific tasks, enabling accurate separation even when multiple speakers have similar characteristics.
Solution Approach 2:
The patent transforms the speech signal from the time domain to the frequency domain through Fourier transformation, extracting spectro-temporal features in the frequency-time representation. By operating in this transformed parameter space rather than the original time domain, the system can better distinguish between multiple speakers with similar temporal characteristics but different spectral patterns.
2Reliability
If the system uses spectro-temporal characteristics to differentiate speakers, then it works for speakers with distinct characteristics, but fails when speakers have similar spectro-temporal characteristics
Solution Approach 1:
The patent extends the feature extraction beyond basic spectro-temporal characteristics to include additional dimensional information. The mask calculator utilizes not only frequency and time dimensions but also introduces a speaker-specific dimension through the use of speaker embeddings or identity features. This multi-dimensional approach allows the system to differentiate between speakers even when their spectro-temporal characteristics are similar in the traditional frequency-time plane.
3Ease of manufacture
If conventional techniques assume noise has different spectro-temporal characteristics from speech, then simple noise suppression works, but the assumption breaks down when other speakers' speech is present
Solution Approach 1:
The patent introduces a mask as an intermediary element between the feature extractor and the object signal calculator. This mask acts as a selective filter that is computed based on the extracted features and then applied to the frequency-domain representation of the mixed speech. The mask calculator serves as a mediator that translates the extracted features into a separation mechanism, enabling precise control over which frequency-time components are attributed to the target speaker versus other speakers or background noise.
Data Source
AI summary
Features are extracted from an observed speech signal including at least speech of multiple speakers including a target speaker. A mask is calculated for extracting speech of the target speaker based on the features of the observed speech signal and a speech signal of the target speaker serving as adaptation data of the target speaker. The signal of the speech of the target speaker is calculated from the observed speech signal based on the mask. Speech of the target speaker can be extracted from observed speech that includes speech of multiple speakers.


