Automated SSML Tag Generation for Natural Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Text-to-Speech (TTS) systems sound robotic and monotonic due to the lack of fine-tuned control over speech synthesis, with manual curation of Speech Synthesis Markup Language (SSML) tags being impractical for large datasets, leading to decreased user engagement.
Innovation Solution
An SSML-enabled TTS system automatically generates SSML tags using natural language understanding models to determine prosodic values such as pitch, rate, and volume, allowing for dynamic control over speech synthesis, eliminating the need for manual curation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual curation of SSML tags is performed, then speech synthesis quality is improved, but time consumption and labor cost increase significantly
Solution Approach 1:
The system enables automatic generation of SSML tags through machine learning models that analyze text content and autonomously determine appropriate prosodic parameters, eliminating the need for manual curation while maintaining speech synthesis quality
Solution Approach 2:
The patent replaces the manual mechanical process of SSML tag curation with an automated computational system using natural language understanding models and sequence-to-sequence architectures to generate tags programmatically
2Manufacturing precision
If manual curation of SSML tags is performed, then speech synthesis quality is improved, but productivity decreases
Solution Approach 1:
The system enables automatic generation of SSML tags through machine learning models that analyze text content and autonomously determine appropriate prosodic parameters, eliminating the need for manual curation while maintaining speech synthesis quality
Solution Approach 2:
The patent transforms the discrete manual tagging process into a continuous automated parameter generation process, where prosodic values are computed based on text analysis features, enabling parallel processing and significant productivity improvement
3Manufacturing precision
If fine-tuned control over speech synthesis is added, then speech naturalness is improved, but system complexity increases
Solution Approach 1:
The system segments the complex speech synthesis control into distinct functional modules: text preprocessing, feature extraction, prosodic parameter prediction, and SSML tag generation, where each module handles a specific aspect of the transformation process
Solution Approach 2:
The patent introduces intermediate prosodic feature representations that bridge the gap between raw text input and final SSML output, using these features as mediators to guide the generation of natural-sounding speech parameters without requiring direct complex control
Data Source
AI summary
In particular embodiments, an apparatus comprises a non-transitory computer-readable storage media and a processor coupled to the media executes instructions to: access a plurality of text, generate, using one or more natural language understanding (NLU) models, one or more scores for at least a portion of the plurality of text. The apparatus determines, based on the scores, one or more prosodic values corresponding to the portion of the plurality of text. The apparatus determines, based on the one or more prosodic values, one or more speech synthesis markup language (SSML) tags. The apparatus then generates, based on the prosodic values, SSML-tagged data comprising each determined SSML tag and that tag's location in the plurality of text.


