Speaker Identification Model Voice Quality Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker identification techniques face challenges in achieving high accuracy due to limitations in increasing the variety of voice data without being restricted by content and language, which affects the model's ability to generalize effectively.

Innovation Solution

A training method that involves voice quality conversion of first voice data to generate second voice data, allowing the speaker identification model to be trained using a diverse set of voice data, thereby increasing the number of training data without content and language restrictions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speaker identification techniques are used, then the model can be trained with available voice data, but the accuracy of speaker identification is limited due to insufficient variety in training data

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidvariety of training data
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates synthetic voice data by copying and transforming existing voice recordings. A voice quality conversion model generates artificial voice data of target speakers by converting source speaker recordings, effectively creating copies with modified characteristics. This allows the training dataset to be expanded without requiring additional real-world recordings, directly addressing the limitation of insufficient training data variety while maintaining speaker identification accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies parameter changes by modifying voice quality attributes such as pitch, timbre, and spectral characteristics through the voice quality conversion model. By changing these acoustic parameters, the system generates diverse training examples from limited source data, enabling the speaker identification model to learn robust features across different voice qualities without being restricted by content and language

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If the training dataset is expanded with more diverse voice data, then speaker identification accuracy improves, but the complexity of data collection and processing increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoiddata collection and processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Instead of manually collecting and processing diverse voice data from multiple sources, the system automatically generates synthetic training data by copying and transforming existing recordings. The voice quality conversion model automates the creation of diverse voice samples, eliminating the need for complex data collection campaigns and manual processing workflows, thus improving accuracy while reducing operational complexity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-service by automatically generating its own training data through the voice quality conversion model. Rather than requiring external data collection efforts, the model synthesizes diverse training examples from available source data, enabling the training process to be self-sufficient and reducing the complexity of data preparation infrastructure

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11580989B2Training method of a speaker identification model based on a first language and a second language
Publication Date: 2023.02.14 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US11580989B2 patent drawing
  • US11580989B2 patent drawing
  • US11580989B2 patent drawing

AI summary

A training method of training a speaker identification model which receives voice data as an input and outputs speaker identification information for identifying a speaker of an utterance included in the voice data is provided. The training method includes: performing voice quality conversion of first voice data of a first speaker to generate second voice data of a second speaker; and performing training of the speaker identification model using, as training data, the first voice data and the second voice data.