Target Speaker Separation via Multi-Cue Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech separation technologies struggle in real-world environments with unknown or dynamically changing numbers of speakers, relying on restrictive assumptions and limited cue types, leading to over- or under-separation issues and subjective evaluation challenges.

Innovation Solution

A multi-cue driven target speaker separation system integrating spatial, dynamic, and steady-state cues through a masked pre-training-based auditory cue inference method, using semi-supervised learning with simulated and real data sets to enhance robustness and adaptability, and generating pseudo-clean reference speech for objective evaluation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional speech separation methods specify the number of speakers in advance, then the model can be trained effectively on benchmark datasets, but the system cannot adapt to real-world scenarios where the number of speakers is unknown or dynamically changes

Engineering Contradiction:
Improveadaptability to dynamic speaker scenariosVSAvoidmodel training complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies dynamics by making the speaker separation model adaptive to dynamic numbers of speakers through iterative separation. The system processes mixed speech through multiple separation stages, where each stage can handle a different number of speakers, allowing the model to adapt to unknown or changing speaker counts without requiring retraining or prior specification of speaker quantity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent segments the speech separation task into multiple independent stages or iterations. Each stage separates one or more speakers from the mixture, and the output of one stage becomes the input for the next. This segmentation allows the system to handle arbitrary numbers of speakers by simply adding or removing separation stages, rather than requiring a fixed configuration.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If speech separation models are trained on simulated mixed data with clean speech labels, then manufacturing precision of separation results is improved, but reliability on real mixed data deteriorates due to domain mismatch

Engineering Contradiction:
Improveseparation accuracy on simulated dataVSAvoidgeneralization to real data
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent introduces an intermediary evaluation mechanism that uses automatic speech recognition (ASR) as a bridge between the speech separation model and real-world performance assessment. By evaluating separation quality through ASR transcription accuracy rather than relying on clean reference labels, the system can be trained on simulated data while being evaluated reliably on real data, bridging the domain gap.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback through iterative evaluation and refinement. The system separates speech, evaluates the quality using ASR on real test data, and uses this feedback to adjust and improve the separation model. This closed-loop feedback mechanism allows the model to adapt to real data characteristics while maintaining the benefits of simulated training data.

Inventive Principle:
Principle #23Feedback

3Device complexity

If single-cue or limited-cue auditory models are used, then device complexity is reduced, but adaptability to different auditory scenarios and robustness deteriorate

Engineering Contradiction:
Improvemodel structure simplicityVSAvoidrobustness across auditory scenarios
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by designing a speech separation model that can handle multiple types of auditory cues and scenarios within a single unified framework. The model is capable of processing spatial cues, spectral cues, and temporal cues simultaneously, and can adapt to different listening scenarios (e.g., far-field, near-field, reverberant environments) without requiring separate specialized models for each scenario.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11978470B2Target speaker separation system, device and storage medium
Publication Date: 2024.05.07 INST OF AUTOMATION CHINESE ACAD OF SCI
  • US11978470B2 patent drawing
  • US11978470B2 patent drawing
  • US11978470B2 patent drawing

AI summary

Disclosed are a target speaker separation system, an electronic device and a storage medium. The system includes: first, performing, jointly unified modeling on a plurality of cues based a masked pre-training strategy, to boost the inference capability of a model for missing cues and enhance the representation accuracy of disturbed cues; and second, constructing a hierarchical cue modulation module. A spatial cue is introduced into a primary cue modulation module for directional enhancement of a speech of a speaker; in an intermediate cue modulation module, the speech of the speaker is enhanced on the basis of temporal coherence of a dynamic cue and an auditory signal component; a steady-state cue is introduced into an advanced cue modulation module for selective filtering; and finally, the supervised learning capability of simulation data and the unsupervised learning effect of real mixed data are sufficiently utilized.