Joint Training of Feature Enhancement and Speaker Recognition Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker recognition systems face performance issues in noisy environments due to the separate training of feature enhancement and I-vector extraction models, which can distort speaker feature information, especially with low signal-to-noise ratios.

Innovation Solution

A combined learning method using deep neural network-based feature enhancement and a modified loss function for joint training of feature enhancement and speaker feature vector extraction models, where the models are connected and trained together using a single loss function to enhance acoustic features and improve speaker recognition performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If separate training of feature enhancement model and I-vector extraction model is used, then training flexibility is improved, but speaker recognition performance in noisy environments deteriorates

Engineering Contradiction:
Improvetraining flexibilityVSAvoidspeaker recognition performance
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent combines the feature enhancement model and the I-vector extraction model into a unified joint training framework. The two models are trained simultaneously with a shared loss function that optimizes both feature enhancement quality and speaker recognition performance together, rather than training them separately and then combining them.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The joint training loss function is segmented into multiple components: a feature enhancement loss term that preserves speaker characteristics, and an I-vector extraction loss term that optimizes recognition performance. This segmented approach allows each component to be optimized independently while contributing to the overall joint optimization.

Inventive Principle:
Principle #1Segmentation

2Object-affected harmful factors

If deep neural network-based feature enhancement is applied, then noise removal capability is improved, but speaker feature information distortion increases

Engineering Contradiction:
Improvenoise removal capabilityVSAvoidspeaker feature information
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent modifies the loss function parameters to include a feature enhancement component that explicitly preserves speaker characteristics. By adjusting the loss function to balance noise removal and feature preservation, the system optimizes the enhancement process to minimize information loss while maximizing noise removal effectiveness.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The joint training framework establishes a feedback loop where the I-vector extraction model's performance on enhanced features provides feedback to the feature enhancement model. This feedback mechanism allows the enhancement model to learn which transformations preserve speaker information while removing noise, continuously optimizing the balance between the two objectives.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12067989B2Combined learning method and apparatus using deepening neural network based feature enhancement and modified loss function for speaker recognition robust to noisy environments
Publication Date: 2024.08.20 INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
  • US12067989B2 patent drawing
  • US12067989B2 patent drawing
  • US12067989B2 patent drawing

AI summary

Presented are a combined learning method and device using a transformed loss function and feature enhancement based on a deep neural network for speaker recognition that is robust in a noisy environment. A combined learning method using a transformed loss function and feature enhancement based on a deep neural network, according to one embodiment, can comprise the steps of: learning a feature enhancement model based on a deep neural network; learning a speaker feature vector extraction model based on the deep neural network; connecting an output layer of the feature enhancement model with an input layer of the speaker feature vector extraction model; and considering the connected feature enhancement model and speaker feature vector extraction model as one mode and performing combined learning for additional learning.