Speech Synthesis Embedding Dimensionality Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis methods struggle to effectively modify and express desired utterance features in synthesized speech, particularly due to the complexity of high-dimensional embedding vectors, making it difficult to directly adjust component values and understand the meaning of each element.

Innovation Solution

A method and device that generate an initial embedding vector, reduce its dimensionality using a predetermined technique, adjust component values based on user input, and restore the dimensionality to create a modified embedding vector, allowing for the synthesis of speech with modified utterance features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If high-dimensional embedding vectors are used to represent utterance features, then the speech synthesis system can capture comprehensive speaker characteristics, but it becomes difficult to directly modify component values and understand the meaning of each element

Engineering Contradiction:
Improveutterance feature expression capabilityVSAvoidease of modifying embedding components
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent transforms the high-dimensional embedding vector into a low-dimensional embedding vector through dimensionality reduction. This creates a new dimension of representation where each component has more interpretable meaning and can be easily modified. The low-dimensional vector maintains the essential speaker characteristics while enabling practical manipulation and understanding of individual components.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If high-dimensional embedding vectors are used to represent utterance features, then comprehensive speaker characteristics can be captured, but the system complexity increases and requires extensive additional components

Engineering Contradiction:
Improveutterance feature expression capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts the essential information from the high-dimensional embedding vector by projecting it into a lower-dimensional space. This extraction process retains the most important speaker characteristics while eliminating redundant dimensions, thereby reducing system complexity and the number of additional components needed for manipulation and interpretation.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If direct modification of high-dimensional embedding vector components is attempted, then utterance features can be adjusted, but the meaning of each element becomes unclear and processing becomes difficult

Engineering Contradiction:
Improveutterance feature modification capabilityVSAvoiddifficulty of understanding element meaning
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent creates a low-dimensional representation where each component corresponds to a more specific and interpretable aspect of speaker characteristics. This dimensional transformation makes it easier to detect, measure, and understand the meaning of individual components while maintaining the ability to modify utterance features effectively.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20240428776A1Method and device for synthesizing speech with modified utterance features
Publication Date: 2024.12.26 XINAPSE CO LTD
  • US20240428776A1 patent drawing
  • US20240428776A1 patent drawing
  • US20240428776A1 patent drawing

AI summary

Provided are a method and device for synthesizing speech with modified utterance features. The method of synthesizing speech with modified utterance features may include generating an initial embedding vector based on predetermined utterance information, generating a low-dimensional embedding vector by reducing dimensionality of the initial embedding vector by using a predetermined dimensionality reduction technique, adjusting a component value of the low-dimensional embedding vector based on a user input, and generating a modified embedding vector by restoring dimensionality of the low-dimensional embedding vector of which the component value is adjusted.