Teacher-Student Speech Separation Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies face challenges in accurately separating target speech from interfering sound sources in complex environments, lacking robustness and generalization due to the inefficiency of supervised learning methods that require labeled high-quality training samples, which are time-consuming and impractical to cover all application scenarios.

Innovation Solution

A speech signal processing method involving a student model and a teacher model, where the models are iteratively trained using a mixed speech signal, with accuracy and consistency information used to adjust their parameters, employing permutation invariant training and salience-based selection mechanisms to improve speech separation accuracy and consistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If supervised learning methods are used for speech separation training, then the model can achieve accurate separation performance, but extensive labeled high-quality training samples are required which are time-consuming and impractical to cover all application scenarios

Engineering Contradiction:
Improvespeech separation accuracyVSAvoidtraining data preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces a teacher model as an intermediary that generates pseudo-labels for training the student model. The teacher model processes mixed speech signals and produces separation results that serve as training targets, eliminating the need for extensive manually labeled data while maintaining training effectiveness

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs self-service by having the teacher model generate its own training data (pseudo-labels) from the mixed speech signals. This self-generated labeling process eliminates the need for external manual annotation, making the training process autonomous and efficient

Inventive Principle:
Principle #25Self-service

2Measurement precision

If supervised learning methods are used for speech separation training, then the model can achieve accurate separation performance, but the method lacks robustness and generalization in complex and variable input environments

Engineering Contradiction:
Improvespeech separation accuracyVSAvoidgeneralization capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The teacher model acts as an adaptive intermediary that can handle diverse input scenarios. By using the teacher model's pseudo-labels for training, the student model learns from a broader range of scenarios including those not covered in the original training data, improving generalization to complex and variable environments

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent employs parameter changes by dynamically adjusting model parameters based on the teacher model's outputs. The student model's parameters are updated iteratively using pseudo-labels from the teacher model, allowing the system to adapt to different input conditions and improve robustness across varying environments

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If a teacher model and student model architecture is used, then speech separation accuracy and consistency are improved, but the device complexity increases

Engineering Contradiction:
Improveseparation accuracyVSAvoidmodel architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses copying by creating a teacher model that replicates the student model's architecture. This allows the teacher model to generate pseudo-labels using the same structural principles, simplifying the overall system design while maintaining accuracy through the teacher-student training paradigm

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12106768B2Speech signal processing method and speech separation method
Publication Date: 2024.10.01 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12106768B2 patent drawing
  • US12106768B2 patent drawing
  • US12106768B2 patent drawing

AI summary

This application provides a speech signal processing method performed by a computer device. Through an iterative training process, a teacher speech separation model can play a smooth role in the training of a student speech separation model based on the accuracy of separation results of the student speech separation model of outputting a target speech signal from a mixed speech signal and the consistency between separation results obtained by the teacher speech separation model of outputting the target speech signal from the mixed speech signal and the student speech separation model of performing the same task, thereby maintaining the separation stability while improving the separation accuracy of the student speech separation model as a trained speech separation model, and greatly improving the separation capability of the trained speech separation model.