Target Speaker Separation via Multi-Cue Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech separation technologies struggle in real-world environments with unknown or dynamically changing numbers of speakers, relying on restrictive assumptions and limited cue types, leading to over- or under-separation issues and subjective evaluation challenges.
Innovation Solution
A multi-cue driven target speaker separation system integrating spatial, dynamic, and steady-state cues through a masked pre-training-based auditory cue inference method, using semi-supervised learning with simulated and real data sets to enhance robustness and adaptability, and generating pseudo-clean reference speech for objective evaluation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional speech separation methods specify the number of speakers in advance, then the model can be trained effectively on benchmark datasets, but the system cannot adapt to real-world scenarios where the number of speakers is unknown or dynamically changes
Solution Approach 1:
The patent applies dynamics by making the speaker separation model adaptive to dynamic numbers of speakers through iterative separation. The system processes mixed speech through multiple separation stages, where each stage can handle a different number of speakers, allowing the model to adapt to unknown or changing speaker counts without requiring retraining or prior specification of speaker quantity.
Solution Approach 2:
The patent segments the speech separation task into multiple independent stages or iterations. Each stage separates one or more speakers from the mixture, and the output of one stage becomes the input for the next. This segmentation allows the system to handle arbitrary numbers of speakers by simply adding or removing separation stages, rather than requiring a fixed configuration.
2Manufacturing precision
If speech separation models are trained on simulated mixed data with clean speech labels, then manufacturing precision of separation results is improved, but reliability on real mixed data deteriorates due to domain mismatch
Solution Approach 1:
The patent introduces an intermediary evaluation mechanism that uses automatic speech recognition (ASR) as a bridge between the speech separation model and real-world performance assessment. By evaluating separation quality through ASR transcription accuracy rather than relying on clean reference labels, the system can be trained on simulated data while being evaluated reliably on real data, bridging the domain gap.
Solution Approach 2:
The patent implements feedback through iterative evaluation and refinement. The system separates speech, evaluates the quality using ASR on real test data, and uses this feedback to adjust and improve the separation model. This closed-loop feedback mechanism allows the model to adapt to real data characteristics while maintaining the benefits of simulated training data.
3Device complexity
If single-cue or limited-cue auditory models are used, then device complexity is reduced, but adaptability to different auditory scenarios and robustness deteriorate
Solution Approach 1:
The patent applies universality by designing a speech separation model that can handle multiple types of auditory cues and scenarios within a single unified framework. The model is capable of processing spatial cues, spectral cues, and temporal cues simultaneously, and can adapt to different listening scenarios (e.g., far-field, near-field, reverberant environments) without requiring separate specialized models for each scenario.
Data Source
AI summary
Disclosed are a target speaker separation system, an electronic device and a storage medium. The system includes: first, performing, jointly unified modeling on a plurality of cues based a masked pre-training strategy, to boost the inference capability of a model for missing cues and enhance the representation accuracy of disturbed cues; and second, constructing a hierarchical cue modulation module. A spatial cue is introduced into a primary cue modulation module for directional enhancement of a speech of a speaker; in an intermediate cue modulation module, the speech of the speaker is enhanced on the basis of temporal coherence of a dynamic cue and an auditory signal component; a steady-state cue is introduced into an advanced cue modulation module for selective filtering; and finally, the supervised learning capability of simulation data and the unsupervised learning effect of real mixed data are sufficiently utilized.


