Speech Synthesis Training Using Perception Representation Acoustic Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies, particularly those using Hidden Markov Models (HMM), face challenges in accurately capturing and replicating the unique voice characteristics and tone of individual speakers, leading to inconsistencies in voice quality and perception.
Innovation Solution
A training apparatus and method that utilizes an average voice model, perception representation score information, and multiple regression Hidden Semi-Markov Models to train perception representation acoustic models, which are then used to control and refine the voice characteristics in speech synthesis, ensuring consistency between acoustic and perceptual spaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If statistical training of acoustic models using HMM is used for speech synthesis, then the synthesis process can be automated and scaled, but the accuracy in capturing individual speaker characteristics and voice quality deteriorates
Solution Approach 1:
The patent segments the speech synthesis model into multiple independent acoustic models, each trained on specific perception representation scores (e.g., pitch, timbre, rhythm). This segmentation allows the system to maintain automation while improving accuracy by focusing on specific voice quality attributes individually rather than treating all characteristics as a single statistical model.
Solution Approach 2:
The patent transforms the approach by changing from traditional HMM parameters to perception representation scores that directly represent human-perceivable voice qualities. By using perceptual parameters (pitch contour, timbre characteristics, rhythm patterns) instead of purely acoustic features, the system maintains automation while significantly improving the accuracy of speaker characteristics.
2Device complexity
If traditional acoustic models are used, then the model structure remains simple and computationally efficient, but the consistency and quality of synthesized speech deteriorates
Solution Approach 1:
The patent divides the speech synthesis system into multiple specialized acoustic models, each responsible for a specific perception representation (pitch, timbre, rhythm). This segmentation maintains computational efficiency by keeping individual models simple while improving overall reliability through specialized focus on each voice quality attribute.
Solution Approach 2:
The patent creates a universal framework where multiple acoustic models work together to produce consistent speech synthesis. Each model is trained on perception representation scores and contributes to the overall synthesis process, ensuring that voice quality consistency is maintained across different speech conditions and speakers.
3Quantity of substance
If acoustic data from multiple speakers is aggregated into an average voice model, then the training data quantity increases and model generalization improves, but the uniqueness of individual speaker characteristics is lost
Solution Approach 1:
The patent applies local quality by training separate acoustic models for different perception representations (pitch, timbre, rhythm) rather than creating a single average model. This allows the system to utilize aggregated training data from multiple speakers while preserving the unique characteristics of individual speakers in each specialized model.
Solution Approach 2:
By segmenting the training process into multiple perception-specific models, the patent enables the system to process large quantities of training data from multiple speakers while maintaining the ability to capture and reproduce individual speaker characteristics through the combined output of specialized models.
Data Source
AI summary
According to one embodiment, a training apparatus for speech synthesis includes a storage device and a hardware processor in communication with the storage device. The storage stores an average voice model, training speaker information representing a feature of speech of a training speaker and perception representation information represented by scores of one or more perception representations related to voice quality of the training speaker, the average voice model constructed by utilizing acoustic data extracted from speech waveforms of a plurality of speakers and language data. The hardware processor, based at least in part on the average voice model, the training speaker information, and the perception representation score, train one or more perception representation acoustic models corresponding to the one or more perception representations.


