Speaker Separation Using Transformer Attractors for Single-Channel Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in accurately separating individual speech signals from a mixture signal recorded with multiple speakers, especially when dealing with single-channel audio inputs, leading to poor performance in speech recognition systems.
Innovation Solution
A speaker separation method and device utilizing an encoder-decoder separation model, which maps the mixture signal to an N-dimensional latent representation, employs a dual-path processing block, a transformer decoder-based attractor calculation module, and a triple-path processing block to model spectrotemporal patterns and inter-speaker relations, enabling separation of an unknown number of speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If existing speaker separation technologies are used with single-channel audio inputs, then device complexity is reduced, but speaker separation performance deteriorates sharply
Solution Approach 1:
The patent transforms the single-channel audio signal into an N-dimensional latent representation through the encoder, where N is greater than or equal to 2. This dimensional transformation enables the model to capture spectrotemporal patterns and inter-speaker relationships that are not directly available in the original single-channel signal, thereby achieving high-quality speaker separation without requiring multiple audio channels.
2Adaptability or versatility
If the number of speakers to be separated is unknown, then adaptability is improved, but device complexity increases due to additional processing modules
Solution Approach 1:
The patent employs a dynamic separator architecture with a transformer decoder-based attractor calculation module that can adaptively determine the number of speakers based on the input mixture signal. The system dynamically adjusts the number of attractors and corresponding speaker representations without requiring prior knowledge or fixed configuration of the number of speakers, enabling flexible handling of varying speaker counts.
Solution Approach 2:
The separator is designed with universal components including a dual-path processing block for modeling spectrotemporal patterns and a triple-path processing block for modeling inter-speaker relations, which can handle any number of speakers from 2 to infinity. This universal design allows the same separator structure to adapt to different speaker counts without requiring separate models or complex reconfiguration.
3Measurement precision
If complex spectrotemporal patterns and inter-speaker relations are modeled, then speaker separation performance is improved, but processing time increases
Solution Approach 1:
The patent divides the processing into distinct segments: the encoder maps the mixture signal to latent representations, the separator processes these representations through dual-path and triple-path blocks for different types of pattern modeling, and the decoder reconstructs the separated speaker signals. This segmentation allows each component to focus on specific aspects of the problem, improving overall efficiency while maintaining comprehensive modeling capabilities.
Data Source
AI summary
Provided are a speaker separation method and device. A speaker separation method for separating an unknown number of speakers from a recorded mixture signal based on an encoder-decoder separation model, including: mapping the mixture signal to an N-dimensional latent representation using an encoder of the encoder-decoder separation model; inputting the N-dimensional latent representation into a separator of the encoder-decoder separation model, wherein the separator includes a dual-path processing block for modeling spectrotemporal patterns, a transformer decoder-based attractor (TDA) calculation module for handling an unknown number of speakers, and a triple-path processing block for modeling inter-speaker relations; performing speaker estimation for speaker separation using the separator to obtain source representations corresponding to the number of separated speakers; and outputting an audio signal for each source representation corresponding to the number of separated speakers using a decoder of the encoder-decoder separation model.


