Continuous Speech Parameter Synthesis for Natural Text-to-Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional text-to-speech systems use step-wise parameter approximation, which does not mimic the natural continuous flow of speech, resulting in unnatural and unintelligible speech synthesis.

Innovation Solution

A system and method that generate continuous parameters using a speech model to process text, incorporating variance scaling and Hidden Markov Model techniques to create a smoother parameter trajectory, mimicking the natural flow of speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If step-wise parameter approximation is used in text-to-speech synthesis, then the computational process is simplified and easier to implement, but the speech output becomes unnatural and unintelligible

Engineering Contradiction:
Improveease of implementationVSAvoidspeech intelligibility
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent applies continuity by generating speech parameters as a continuous stream rather than discrete step-wise values. The system produces a continuous parameter trajectory that mimics natural speech flow, where parameters evolve smoothly over time without abrupt transitions. This continuous parameter generation resolves the contradiction by maintaining both computational feasibility and natural speech characteristics.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system dynamically adjusts parameters along a continuous trajectory rather than using static step-wise values. The parameter generation process incorporates temporal dynamics by predicting how parameters should evolve continuously over time, allowing the speech synthesis to adapt smoothly to natural speech patterns while maintaining computational efficiency.

Inventive Principle:
Principle #15Dynamics

2Reliability

If continuous parameter generation is used to mimic natural speech flow, then speech intelligibility and naturalness are improved, but the computational complexity increases

Engineering Contradiction:
Improvespeech naturalnessVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical-like step-wise parameter generation with a statistical parametric model that naturally produces continuous parameters. By using probabilistic models and statistical methods instead of deterministic step-wise approaches, the system achieves continuous parameter generation with computational efficiency, resolving the contradiction between naturalness and complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the fundamental approach to parameter generation by using statistical parametric models that inherently produce continuous parameter trajectories. This parameter transformation approach allows the system to generate smooth, natural-sounding speech parameters without requiring computationally intensive methods, as the statistical model naturally captures the continuous nature of speech.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10733974B2System and method for synthesis of speech from provided text
Publication Date: 2020.08.04 GENESYS CLOUD SERVICES INC
  • US10733974B2 patent drawing
  • US10733974B2 patent drawing
  • US10733974B2 patent drawing

AI summary

A system and method are presented for the synthesis of speech from provided text. Particularly, the generation of parameters within the system is performed as a continuous approximation in order to mimic the natural flow of speech as opposed to a step-wise approximation of the feature stream. Provided text may be partitioned and parameters generated using a speech model. The generated parameters from the speech model may then be used in a post-processing step to obtain a new set of parameters for application in speech synthesis.