Prosody Models for Natural Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-to-speech synthesis systems lack natural variability in prosody, resulting in flat and artificial-sounding speech, as they fail to replicate the rich emotional and contextual variations present in human speech.

Innovation Solution

The development of methods and systems for training and applying prosody models that generate and apply natural, varying prosody to speech synthesis, using speech recognition engines to annotate text with prosody information, translating formats as needed, and selecting multiple models based on context and confidence indicators to produce nuanced pitch, rate, and volume changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single prosody style is used in TTS synthesis, then the system is simple to implement, but the speech lacks natural variability and sounds flat and artificial

Engineering Contradiction:
Improveprosody variabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments prosody into multiple independent styles (e.g., newscaster, conversational, emphatic) and trains separate prosody models for each style. This allows the system to select and apply appropriate prosody styles based on context, achieving natural variability without requiring a completely complex unified model. Each style model is trained on specific annotated speech data characteristic of that style.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal prosody model framework that can handle multiple prosody styles through a common architecture. The system uses a single text-to-speech engine that can be directed to apply different prosody models, making the system multi-functional while maintaining a unified base structure. This reduces overall complexity compared to having entirely separate systems for each style.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multiple prosody models are applied to a single piece of text, then the speech achieves richer natural variation, but conflicts in prosody directives must be arbitrated

Engineering Contradiction:
Improveprosody style rangeVSAvoidmodel coordination complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms where prosody models provide confidence indicators with their directives. The arbitration system uses this feedback to resolve conflicts between multiple models, selecting the prosody directives with highest confidence or those most appropriate for the given context. This systematic feedback-based arbitration reduces the complexity of coordinating multiple models.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the parameter of model selection by using context analysis (keywords, subject matter, desired emotional tone) to dynamically select which prosody models to apply and how to weight their directives. This parameter-based control mechanism simplifies the coordination of multiple models by providing clear rules for their interaction based on the synthesis context.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9070365B2Training and applying prosody models
Publication Date: 2015.06.30 MORPHISM LLC
  • US9070365B2 patent drawing
  • US9070365B2 patent drawing
  • US9070365B2 patent drawing

AI summary

Techniques for training and applying prosody models for speech synthesis are provided. A speech recognition engine processes audible speech to produce text annotated with prosody information. A prosody model is trained with this annotated text. After initial training, the model is applied during speech synthesis to generate speech with non-standard prosody from input text. Multiple prosody models can be used to represent different prosody styles.