Automated SSML Tag Generation for Natural Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Text-to-Speech (TTS) systems sound robotic and monotonic due to the lack of fine-tuned control over speech synthesis, with manual curation of Speech Synthesis Markup Language (SSML) tags being impractical for large datasets, leading to decreased user engagement.

Innovation Solution

An SSML-enabled TTS system automatically generates SSML tags using natural language understanding models to determine prosodic values such as pitch, rate, and volume, allowing for dynamic control over speech synthesis, eliminating the need for manual curation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual curation of SSML tags is performed, then speech synthesis quality is improved, but time consumption and labor cost increase significantly

Engineering Contradiction:
Improvespeech synthesis qualityVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system enables automatic generation of SSML tags through machine learning models that analyze text content and autonomously determine appropriate prosodic parameters, eliminating the need for manual curation while maintaining speech synthesis quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of SSML tag curation with an automated computational system using natural language understanding models and sequence-to-sequence architectures to generate tags programmatically

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If manual curation of SSML tags is performed, then speech synthesis quality is improved, but productivity decreases

Engineering Contradiction:
Improvespeech synthesis qualityVSAvoidprocessing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system enables automatic generation of SSML tags through machine learning models that analyze text content and autonomously determine appropriate prosodic parameters, eliminating the need for manual curation while maintaining speech synthesis quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the discrete manual tagging process into a continuous automated parameter generation process, where prosodic values are computed based on text analysis features, enabling parallel processing and significant productivity improvement

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If fine-tuned control over speech synthesis is added, then speech naturalness is improved, but system complexity increases

Engineering Contradiction:
Improvespeech naturalnessVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system segments the complex speech synthesis control into distinct functional modules: text preprocessing, feature extraction, prosodic parameter prediction, and SSML tag generation, where each module handles a specific aspect of the transformation process

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate prosodic feature representations that bridge the gap between raw text input and final SSML output, using these features as mediators to guide the generation of natural-sounding speech parameters without requiring direct complex control

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11380300B2Automatically generating speech markup language tags for text
Publication Date: 2022.07.05 SAMSUNG ELECTRONICS CO LTD
  • US11380300B2 patent drawing
  • US11380300B2 patent drawing
  • US11380300B2 patent drawing

AI summary

In particular embodiments, an apparatus comprises a non-transitory computer-readable storage media and a processor coupled to the media executes instructions to: access a plurality of text, generate, using one or more natural language understanding (NLU) models, one or more scores for at least a portion of the plurality of text. The apparatus determines, based on the scores, one or more prosodic values corresponding to the portion of the plurality of text. The apparatus determines, based on the one or more prosodic values, one or more speech synthesis markup language (SSML) tags. The apparatus then generates, based on the prosodic values, SSML-tagged data comprising each determined SSML tag and that tag's location in the plurality of text.