Global Style Token Embeddings for Data-Efficient Speaker Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speaker adaptation methods require large amounts of data and fine-tuning of the entire model, which is inefficient and resource-intensive.

Innovation Solution

A speaker adaptation system using a global style token mechanism to generate and predict speaker embeddings, allowing for efficient representation of a speaker's tone without extensive fine-tuning, by constructing a voice conversion model, extracting variance, and predicting a final speaker embedding through similarity comparison.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speaker adaptation methods are used, then speaker tone representation can be achieved, but large amounts of data and extensive fine-tuning are required

Engineering Contradiction:
Improvespeaker tone representation accuracyVSAvoiddata amount required
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts speaker tone characteristics by generating multiple speaker embeddings from a single speaker embedding through the voice conversion model with global style token mechanism. This extraction approach isolates the essential tone representation without requiring large amounts of additional speaker data, directly resolving the contradiction between accurate tone representation and data quantity requirements

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary generation of multiple speaker embeddings before the adaptation process. By pre-generating a plurality of speaker embeddings that represent different tones of a speaker, the system prepares diverse tone representations in advance, eliminating the need for extensive fine-tuning and large data sets during actual speaker adaptation

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If conventional speaker adaptation methods are used, then speaker tone representation can be achieved, but the entire model needs to be fine-tuned

Engineering Contradiction:
Improvespeaker tone representation accuracyVSAvoidmodel fine-tuning complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the necessary speaker embedding representations through the voice conversion model and global style token mechanism, rather than fine-tuning the entire model. This selective extraction approach maintains accurate speaker tone representation while avoiding the complexity of full model fine-tuning

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system segments the speaker adaptation process into distinct components: generating multiple speaker embeddings using the voice conversion model, selecting appropriate embeddings through similarity comparison, and combining with pitch embeddings. This segmentation allows tone representation without requiring complex full-model fine-tuning

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If conventional speaker adaptation methods are used, then speaker adaptation can be performed, but resource requirements are high

Engineering Contradiction:
Improvespeaker adaptation capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system generates a plurality of speaker embeddings (excessive action) from a single speaker embedding through the voice conversion model. This partial generation approach provides sufficient tone diversity for adaptation without requiring computationally expensive full-model fine-tuning on large data sets, thus reducing overall resource consumption while maintaining adaptability

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250292776A1Speaker embedding-based speaker adaptation method and system generated by using global style tokens and prediction model
Publication Date: 2025.09.18 INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
  • US20250292776A1 patent drawing
  • US20250292776A1 patent drawing
  • US20250292776A1 patent drawing

AI summary

Disclosed are a speaker embedding-based speaker adaptation method and system generated by using global style tokens and a prediction model. The speaker adaptation method performed by the speaker adaptation system, according to an embodiment, may comprise the steps of: generating a plurality of speaker embeddings representing the tone of a speaker from a speaker embedding by using a voice transformation model including a global style token mechanism; and predicting the final speaker embedding representing a new speaker through similarity comparison between a new speaker embedding predicted by using a prediction model for predicting a speaker embedding and the plurality of generated speaker embeddings.