Voice Style Sheet Framework for Expressive Text-to-Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Statistical Parametric Speech Synthesis Systems lack the ability to author and manipulate multiple low-level voice parameters concurrently, limiting their expressiveness and emotional depth in speech generation.

Innovation Solution

A system with a Voice Style Sheet (VSS) framework that allows independent control of pitch, speed, vocal tract length, and other parameters in real-time, enabling the creation and application of detailed animation controls for speech synthesis, similar to computer graphic animation, using a scheduler to ensure precise parameter manipulation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional Statistical Parametric Speech Synthesis Systems are used, then speech generation is achieved, but the ability to author and manipulate multiple low-level voice parameters concurrently is limited

Engineering Contradiction:
Improveparameter manipulation capabilityVSAvoidsystem structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the speech synthesis system into distinct functional modules: a text and labels module for phonetic description, a parameter generation module for creating speech parameters, an audio generation module for synthesizing audio samples, and a scheduler for coordinating operations. This segmentation enables independent control and manipulation of multiple voice parameters concurrently while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic parameter control through the scheduler, which monitors and schedules parameter generation in real-time. The system transitions from static parameter generation to dynamic adjustment of pitch, speed, vocal tract length, and other parameters during speech synthesis, enabling expressive and emotive speech output while managing complexity through time-based parameter modulation.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If detailed markup is added to control speech rendering at the phoneme level, then duration and pitch control is improved, but the system still lacks comprehensive control over other voice parameters

Engineering Contradiction:
Improvespeech control precisionVSAvoidparameter control range
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal parameter generation framework that handles multiple voice parameters (pitch, duration, speed, vocal tract length, and other low-level parameters) through a single integrated system. The parameter generation module generates comprehensive speech parameters that can be manipulated concurrently, providing both precise phoneme-level control and broader parameter versatility without requiring separate markup systems for each parameter type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If speech synthesis parameters are statically generated over the complete input sentence, then processing is simplified, but variable control of a movable portion of the input sentence is lost

Engineering Contradiction:
Improveprocessing complexityVSAvoidreal-time parameter control
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements periodic action through the scheduler, which generates and updates speech synthesis parameters in continuous time intervals rather than statically for the entire sentence. The scheduler monitors parameter generation requests, schedules parameter creation in periodic cycles, and enables real-time modification of parameters for movable portions of the input sentence, achieving both simplified processing through systematic scheduling and real-time adaptability.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS12020686B2System providing expressive and emotive text-to-speech
Publication Date: 2024.06.25 D & M HOLDINGS INC
  • US12020686B2 patent drawing
  • US12020686B2 patent drawing
  • US12020686B2 patent drawing

AI summary

A speech to text system includes a text and labels module receiving a text input and providing a text analysis and a label with a phonetic description of the text. A label buffer receives the label from the text and labels module. A parameter generation module accesses the label from the label buffer and generates a speech generation parameter. A parameter buffer receives the parameter from the parameter generation module. An audio generation module receives the text input, the label, and/or the parameter and generates a plurality of audio samples, A scheduler monitors and schedules the text and label module, the parameter generation module, and/or the audio generation module. The parameter generation module is further configured to initialize a voice identifier with a Voice Style Sheet (VSS) parameter, receive an input indicating a modification to the VSS parameter, and modify the VSS parameter according to the modification.