Voice Style Sheet Framework for Expressive Text-to-Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Statistical Parametric Speech Synthesis Systems lack the ability to author and manipulate multiple low-level voice parameters concurrently, limiting their expressiveness and emotional depth in speech generation.
Innovation Solution
A system with a Voice Style Sheet (VSS) framework that allows independent control of pitch, speed, vocal tract length, and other parameters in real-time, enabling the creation and application of detailed animation controls for speech synthesis, similar to computer graphic animation, using a scheduler to ensure precise parameter manipulation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional Statistical Parametric Speech Synthesis Systems are used, then speech generation is achieved, but the ability to author and manipulate multiple low-level voice parameters concurrently is limited
Solution Approach 1:
The patent segments the speech synthesis system into distinct functional modules: a text and labels module for phonetic description, a parameter generation module for creating speech parameters, an audio generation module for synthesizing audio samples, and a scheduler for coordinating operations. This segmentation enables independent control and manipulation of multiple voice parameters concurrently while maintaining manageable system complexity through modular architecture.
Solution Approach 2:
The patent introduces dynamic parameter control through the scheduler, which monitors and schedules parameter generation in real-time. The system transitions from static parameter generation to dynamic adjustment of pitch, speed, vocal tract length, and other parameters during speech synthesis, enabling expressive and emotive speech output while managing complexity through time-based parameter modulation.
2Ease of operation
If detailed markup is added to control speech rendering at the phoneme level, then duration and pitch control is improved, but the system still lacks comprehensive control over other voice parameters
Solution Approach 1:
The patent creates a universal parameter generation framework that handles multiple voice parameters (pitch, duration, speed, vocal tract length, and other low-level parameters) through a single integrated system. The parameter generation module generates comprehensive speech parameters that can be manipulated concurrently, providing both precise phoneme-level control and broader parameter versatility without requiring separate markup systems for each parameter type.
3Device complexity
If speech synthesis parameters are statically generated over the complete input sentence, then processing is simplified, but variable control of a movable portion of the input sentence is lost
Solution Approach 1:
The patent implements periodic action through the scheduler, which generates and updates speech synthesis parameters in continuous time intervals rather than statically for the entire sentence. The scheduler monitors parameter generation requests, schedules parameter creation in periodic cycles, and enables real-time modification of parameters for movable portions of the input sentence, achieving both simplified processing through systematic scheduling and real-time adaptability.
Data Source
AI summary
A speech to text system includes a text and labels module receiving a text input and providing a text analysis and a label with a phonetic description of the text. A label buffer receives the label from the text and labels module. A parameter generation module accesses the label from the label buffer and generates a speech generation parameter. A parameter buffer receives the parameter from the parameter generation module. An audio generation module receives the text input, the label, and/or the parameter and generates a plurality of audio samples, A scheduler monitors and schedules the text and label module, the parameter generation module, and/or the audio generation module. The parameter generation module is further configured to initialize a voice identifier with a Voice Style Sheet (VSS) parameter, receive an input indicating a modification to the VSS parameter, and modify the VSS parameter according to the modification.


