Text-to-Speech Conversion Parameters for Speaker Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech technologies based on hidden Markov models (HMM) face challenges in accurately reproducing the context-dependent locally-appearing characteristics of a target speaking style, such as pitch and pause variations, and struggle to precisely replicate the voice quality of a target speaker due to the need for extensive voice recording and phonetic labeling, or require complex cluster adaptive training frameworks.

Innovation Solution

A text-to-speech device that includes a context acquirer, acoustic model parameter acquirer, conversion parameter acquirer, converter, and waveform generator, which analyzes input text to acquire context sequences and convert acoustic model parameters from a standard speaking style to a target speaking style using decision tree-based conversion parameters, effectively reflecting context dependency and voice quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If extensive voice recording and phonetic labeling are performed to train HMM for target speaker and speaking style, then voice quality and speaking style accuracy are improved, but recording cost and time consumption increase

Engineering Contradiction:
Improvevoice quality accuracyVSAvoidrecording time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the voice synthesis task into independent components: a standard speaking style HMM that is pre-trained once, and speaker-specific conversion parameters that are computed individually for each target speaker. This segmentation allows the heavy computational work to be done once for the standard style, while individual speaker adaptation is performed efficiently through parameter conversion rather than retraining, thus reducing overall time consumption while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a copy of the standard speaking style HMM for each target speaker by computing conversion parameters that transform the standard HMM into speaker-specific versions. Instead of training a complete HMM from scratch for each speaker (which would require extensive recording), the system copies the standard HMM structure and applies speaker-specific conversions, significantly reducing the data requirements and time needed for speaker adaptation.

Inventive Principle:
Principle #26Copying

2Ease of manufacture

If standard speaking style HMM is used without speaker adaptation, then training cost is reduced, but voice quality accuracy of target speaker deteriorates

Engineering Contradiction:
Improvetraining costVSAvoidvoice quality accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent changes the parameters of the standard speaking style HMM by computing conversion parameters for each target speaker. These conversion parameters adjust the HMM parameters (such as mean and covariance of phonetic features) to match the target speaker's voice characteristics. This parameter transformation allows the system to maintain the benefits of a single pre-trained standard HMM while achieving accurate voice quality reproduction for each target speaker through efficient parameter conversion rather than complete retraining.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If context dependency for each speaking style is not considered, then model complexity is reduced, but locally-appearing characteristics of speaking style deteriorate

Engineering Contradiction:
Improvemodel complexityVSAvoidspeaking style characteristic accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent applies local quality by considering context dependency specifically for phonetic features that are sensitive to speaking style changes. Instead of treating all phonetic parameters uniformly, the system identifies and adjusts local characteristics (such as pitch contours, pause durations, and stress patterns) that are particularly important for reproducing speaking style nuances. This targeted approach maintains model efficiency while improving the accuracy of locally-appearing speaking style characteristics.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9830904B2Text-to-speech device, text-to-speech method, and computer program product
Publication Date: 2017.11.28 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9830904B2 patent drawing
  • US9830904B2 patent drawing
  • US9830904B2 patent drawing

AI summary

According to an embodiment, a text-to-speech device includes a context acquirer, an acoustic model parameter acquirer, a conversion parameter acquirer, a converter, and a waveform generator. The context acquirer is configured to acquire a context sequence affecting fluctuations in voice. The acoustic model parameter acquirer is configured to acquire an acoustic model parameter sequence that corresponds to the context sequence and represents an acoustic model in a standard speaking style of a target speaker. The conversion parameter acquirer is configured to acquire a conversion parameter sequence corresponding to the context sequence to convert an acoustic model parameter in the standard speaking style into one in a different speaking style. The converter is configured to convert the acoustic model parameter sequence using the conversion parameter sequence. The waveform generator is configured to generate a voice signal based on the acoustic model parameter sequence acquired after conversion.