Speech Synthesis Training Using Perception Representation Acoustic Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis technologies, particularly those using Hidden Markov Models (HMM), face challenges in accurately capturing and replicating the unique voice characteristics and tone of individual speakers, leading to inconsistencies in voice quality and perception.

Innovation Solution

A training apparatus and method that utilizes an average voice model, perception representation score information, and multiple regression Hidden Semi-Markov Models to train perception representation acoustic models, which are then used to control and refine the voice characteristics in speech synthesis, ensuring consistency between acoustic and perceptual spaces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If statistical training of acoustic models using HMM is used for speech synthesis, then the synthesis process can be automated and scaled, but the accuracy in capturing individual speaker characteristics and voice quality deteriorates

Engineering Contradiction:
Improveautomation of speech synthesis trainingVSAvoidaccuracy of speaker characteristics
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent segments the speech synthesis model into multiple independent acoustic models, each trained on specific perception representation scores (e.g., pitch, timbre, rhythm). This segmentation allows the system to maintain automation while improving accuracy by focusing on specific voice quality attributes individually rather than treating all characteristics as a single statistical model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the approach by changing from traditional HMM parameters to perception representation scores that directly represent human-perceivable voice qualities. By using perceptual parameters (pitch contour, timbre characteristics, rhythm patterns) instead of purely acoustic features, the system maintains automation while significantly improving the accuracy of speaker characteristics.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If traditional acoustic models are used, then the model structure remains simple and computationally efficient, but the consistency and quality of synthesized speech deteriorates

Engineering Contradiction:
Improvesimplicity of model structureVSAvoidconsistency of voice quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent divides the speech synthesis system into multiple specialized acoustic models, each responsible for a specific perception representation (pitch, timbre, rhythm). This segmentation maintains computational efficiency by keeping individual models simple while improving overall reliability through specialized focus on each voice quality attribute.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal framework where multiple acoustic models work together to produce consistent speech synthesis. Each model is trained on perception representation scores and contributes to the overall synthesis process, ensuring that voice quality consistency is maintained across different speech conditions and speakers.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If acoustic data from multiple speakers is aggregated into an average voice model, then the training data quantity increases and model generalization improves, but the uniqueness of individual speaker characteristics is lost

Engineering Contradiction:
Improveamount of training dataVSAvoidindividual speaker characteristics
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by training separate acoustic models for different perception representations (pitch, timbre, rhythm) rather than creating a single average model. This allows the system to utilize aggregated training data from multiple speakers while preserving the unique characteristics of individual speakers in each specialized model.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

By segmenting the training process into multiple perception-specific models, the patent enables the system to process large quantities of training data from multiple speakers while maintaining the ability to capture and reproduce individual speaker characteristics through the combined output of specialized models.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10540956B2Training apparatus for speech synthesis, speech synthesis apparatus and training method for training apparatus
Publication Date: 2020.01.21 KK TOSHIBA
  • US10540956B2 patent drawing
  • US10540956B2 patent drawing
  • US10540956B2 patent drawing

AI summary

According to one embodiment, a training apparatus for speech synthesis includes a storage device and a hardware processor in communication with the storage device. The storage stores an average voice model, training speaker information representing a feature of speech of a training speaker and perception representation information represented by scores of one or more perception representations related to voice quality of the training speaker, the average voice model constructed by utilizing acoustic data extracted from speech waveforms of a plurality of speakers and language data. The hardware processor, based at least in part on the average voice model, the training speaker information, and the perception representation score, train one or more perception representation acoustic models corresponding to the one or more perception representations.