Continuous Speech Parameter Synthesis for Natural Text-to-Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional text-to-speech systems use step-wise parameter approximation, which does not mimic the natural continuous flow of speech, resulting in unnatural and unintelligible speech synthesis.
Innovation Solution
A system and method that generate continuous parameters using a speech model to process text, incorporating variance scaling and Hidden Markov Model techniques to create a smoother parameter trajectory, mimicking the natural flow of speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If step-wise parameter approximation is used in text-to-speech synthesis, then the computational process is simplified and easier to implement, but the speech output becomes unnatural and unintelligible
Solution Approach 1:
The patent applies continuity by generating speech parameters as a continuous stream rather than discrete step-wise values. The system produces a continuous parameter trajectory that mimics natural speech flow, where parameters evolve smoothly over time without abrupt transitions. This continuous parameter generation resolves the contradiction by maintaining both computational feasibility and natural speech characteristics.
Solution Approach 2:
The system dynamically adjusts parameters along a continuous trajectory rather than using static step-wise values. The parameter generation process incorporates temporal dynamics by predicting how parameters should evolve continuously over time, allowing the speech synthesis to adapt smoothly to natural speech patterns while maintaining computational efficiency.
2Reliability
If continuous parameter generation is used to mimic natural speech flow, then speech intelligibility and naturalness are improved, but the computational complexity increases
Solution Approach 1:
The patent replaces complex mechanical-like step-wise parameter generation with a statistical parametric model that naturally produces continuous parameters. By using probabilistic models and statistical methods instead of deterministic step-wise approaches, the system achieves continuous parameter generation with computational efficiency, resolving the contradiction between naturalness and complexity.
Solution Approach 2:
The system changes the fundamental approach to parameter generation by using statistical parametric models that inherently produce continuous parameter trajectories. This parameter transformation approach allows the system to generate smooth, natural-sounding speech parameters without requiring computationally intensive methods, as the statistical model naturally captures the continuous nature of speech.
Data Source
AI summary
A system and method are presented for the synthesis of speech from provided text. Particularly, the generation of parameters within the system is performed as a continuous approximation in order to mimic the natural flow of speech as opposed to a step-wise approximation of the feature stream. Provided text may be partitioned and parameters generated using a speech model. The generated parameters from the speech model may then be used in a post-processing step to obtain a new set of parameters for application in speech synthesis.


