Joint Training of Feature Enhancement and Speaker Recognition Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speaker recognition systems face performance issues in noisy environments due to the separate training of feature enhancement and I-vector extraction models, which can distort speaker feature information, especially with low signal-to-noise ratios.
Innovation Solution
A combined learning method using deep neural network-based feature enhancement and a modified loss function for joint training of feature enhancement and speaker feature vector extraction models, where the models are connected and trained together using a single loss function to enhance acoustic features and improve speaker recognition performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If separate training of feature enhancement model and I-vector extraction model is used, then training flexibility is improved, but speaker recognition performance in noisy environments deteriorates
Solution Approach 1:
The patent combines the feature enhancement model and the I-vector extraction model into a unified joint training framework. The two models are trained simultaneously with a shared loss function that optimizes both feature enhancement quality and speaker recognition performance together, rather than training them separately and then combining them.
Solution Approach 2:
The joint training loss function is segmented into multiple components: a feature enhancement loss term that preserves speaker characteristics, and an I-vector extraction loss term that optimizes recognition performance. This segmented approach allows each component to be optimized independently while contributing to the overall joint optimization.
2Object-affected harmful factors
If deep neural network-based feature enhancement is applied, then noise removal capability is improved, but speaker feature information distortion increases
Solution Approach 1:
The patent modifies the loss function parameters to include a feature enhancement component that explicitly preserves speaker characteristics. By adjusting the loss function to balance noise removal and feature preservation, the system optimizes the enhancement process to minimize information loss while maximizing noise removal effectiveness.
Solution Approach 2:
The joint training framework establishes a feedback loop where the I-vector extraction model's performance on enhanced features provides feedback to the feature enhancement model. This feedback mechanism allows the enhancement model to learn which transformations preserve speaker information while removing noise, continuously optimizing the balance between the two objectives.
Data Source
AI summary
Presented are a combined learning method and device using a transformed loss function and feature enhancement based on a deep neural network for speaker recognition that is robust in a noisy environment. A combined learning method using a transformed loss function and feature enhancement based on a deep neural network, according to one embodiment, can comprise the steps of: learning a feature enhancement model based on a deep neural network; learning a speaker feature vector extraction model based on the deep neural network; connecting an output layer of the feature enhancement model with an input layer of the speaker feature vector extraction model; and considering the connected feature enhancement model and speaker feature vector extraction model as one mode and performing combined learning for additional learning.


