Teacher-Student PIT Framework for Multi-Talker Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for multi-talker speech recognition, such as deep learning models, face challenges with high word error rates due to label ambiguity and permutation issues, especially in scenarios with multiple speakers and single-channel mixed speech.
Innovation Solution
The implementation of a Teacher-Student permutation invariant training (PIT) framework, which transfers knowledge from single-talker ASR models to multi-talker ASR models using soft labels and a progressive training scheme, along with data augmentation and domain adaptation, to improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If deep learning models are used for multi-talker speech recognition, then recognition capability is improved, but word error rate increases due to label ambiguity and permutation issues
Solution Approach 1:
The patent segments the multi-talker speech recognition problem by introducing separate output streams for different talkers, allowing the model to process mixed speech through multiple specialized pathways rather than a single ambiguous stream
Solution Approach 2:
The patent introduces an intermediary assignment mechanism that maps output streams to talkers based on permutation invariant training, serving as a mediator between the deep learning model's outputs and the ground truth labels to resolve label ambiguity
2Measurement precision
If speaker adaptation is applied to reduce mismatch between training and test speakers, then WER improves for single-talker cases, but it cannot be directly applied to multi-talker scenarios
Solution Approach 1:
The patent creates a universal PIT training framework that can handle both single-talker and multi-talker scenarios, making the adaptation mechanism versatile across different speech recognition contexts by treating talker assignment as a permutation-invariant problem
3Ease of manufacture
If traditional training methods are used for multi-talker ASR, then model training is simpler, but recognition accuracy remains low due to label permutation problems
Solution Approach 1:
The patent enables the model to self-resolve the label permutation problem through permutation invariant training, where the training process automatically learns to match output streams to talkers without requiring external intervention or complex preprocessing
Data Source
AI summary
Provided are a speech recognition training processing method and an apparatus including the same. The speech recognition training processing method includes acquiring a multi-talker mixed speech signal from a plurality of speakers, performing permutation invariant training (PIT) model training on the multi-talker mixed speech signal based on knowledge from a single-talker speech recognition model and updating a multi-talker speech recognition model based on a result of the PIT model training.


