Speaker Separation Using Transformer Attractors for Single-Channel Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in accurately separating individual speech signals from a mixture signal recorded with multiple speakers, especially when dealing with single-channel audio inputs, leading to poor performance in speech recognition systems.

Innovation Solution

A speaker separation method and device utilizing an encoder-decoder separation model, which maps the mixture signal to an N-dimensional latent representation, employs a dual-path processing block, a transformer decoder-based attractor calculation module, and a triple-path processing block to model spectrotemporal patterns and inter-speaker relations, enabling separation of an unknown number of speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If existing speaker separation technologies are used with single-channel audio inputs, then device complexity is reduced, but speaker separation performance deteriorates sharply

Engineering Contradiction:
Improveaudio channel configurationVSAvoidspeaker separation performance
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transforms the single-channel audio signal into an N-dimensional latent representation through the encoder, where N is greater than or equal to 2. This dimensional transformation enables the model to capture spectrotemporal patterns and inter-speaker relationships that are not directly available in the original single-channel signal, thereby achieving high-quality speaker separation without requiring multiple audio channels.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If the number of speakers to be separated is unknown, then adaptability is improved, but device complexity increases due to additional processing modules

Engineering Contradiction:
Improvehandling unknown number of speakersVSAvoidseparator structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent employs a dynamic separator architecture with a transformer decoder-based attractor calculation module that can adaptively determine the number of speakers based on the input mixture signal. The system dynamically adjusts the number of attractors and corresponding speaker representations without requiring prior knowledge or fixed configuration of the number of speakers, enabling flexible handling of varying speaker counts.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The separator is designed with universal components including a dual-path processing block for modeling spectrotemporal patterns and a triple-path processing block for modeling inter-speaker relations, which can handle any number of speakers from 2 to infinity. This universal design allows the same separator structure to adapt to different speaker counts without requiring separate models or complex reconfiguration.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If complex spectrotemporal patterns and inter-speaker relations are modeled, then speaker separation performance is improved, but processing time increases

Engineering Contradiction:
Improvespeaker separation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the processing into distinct segments: the encoder maps the mixture signal to latent representations, the separator processes these representations through dual-path and triple-path blocks for different types of pattern modeling, and the decoder reconstructs the separated speaker signals. This segmentation allows each component to focus on specific aspects of the problem, improving overall efficiency while maintaining comprehensive modeling capabilities.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250140276A1Speaker separation method and device
Publication Date: 2025.05.01 42DOT INC
  • US20250140276A1 patent drawing
  • US20250140276A1 patent drawing
  • US20250140276A1 patent drawing

AI summary

Provided are a speaker separation method and device. A speaker separation method for separating an unknown number of speakers from a recorded mixture signal based on an encoder-decoder separation model, including: mapping the mixture signal to an N-dimensional latent representation using an encoder of the encoder-decoder separation model; inputting the N-dimensional latent representation into a separator of the encoder-decoder separation model, wherein the separator includes a dual-path processing block for modeling spectrotemporal patterns, a transformer decoder-based attractor (TDA) calculation module for handling an unknown number of speakers, and a triple-path processing block for modeling inter-speaker relations; performing speaker estimation for speaker separation using the separator to obtain source representations corresponding to the number of separated speakers; and outputting an audio signal for each source representation corresponding to the number of separated speakers using a decoder of the encoder-decoder separation model.