Emoticon-Based Text-to-Speech Expressivity Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-speech systems struggle to accurately convey emotion or audible expressivity, relying on limited analysis of punctuation and word arrangement, which often fails to capture the intended mood of the text, especially when the composer's mood varies.

Innovation Solution

The use of emoticons as contextual cues to enhance the expressivity of text-to-speech synthesis, where identified emoticons in a text string are used to modify audio output characteristics such as intonation, prosody, speed, and pauses, thereby improving the perception of emotional tone.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional TTS systems analyze only punctuation and word arrangement to determine mood, then the system complexity remains low, but the accuracy of emotional expression deteriorates

Engineering Contradiction:
Improvemood detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary identification and tagging of emoticons in the text before the main text-to-speech conversion process. By pre-processing the text to detect emoticons and attach expressivity tags, the system prepares emotional context information in advance, enabling more accurate mood detection without significantly increasing overall system complexity during the conversion phase

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary component (emoticon identifier and expressivity tagger) that bridges the gap between simple text input and emotional expression in speech output. This intermediary layer processes emoticons and generates expressivity tags that guide the TTS engine, thereby improving mood detection accuracy while isolating the complexity to a specific module rather than the entire system

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If TTS systems use basic punctuation analysis to infer emotion, then the processing speed remains fast, but the reliability of emotional conveyance deteriorates

Engineering Contradiction:
Improveemotional conveyance accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system segments the text processing task into distinct components: emoticon identification, expressivity tag generation, and text-to-speech conversion with expressivity guidance. By dividing the processing into these segments, the system can efficiently handle emoticon detection as a separate step while maintaining overall processing speed, and improve reliability through specialized handling of emotional cues

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If TTS systems ignore emoticons or speak their literal names, then the system simplicity is maintained, but the expressivity and naturalness of speech output deteriorates

Engineering Contradiction:
Improvespeech expressivityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system changes the parameter of text representation by introducing expressivity tags that encode emotional information from emoticons. Instead of ignoring emoticons or speaking them literally, the system transforms them into actionable parameters (expressivity tags) that modify the speech synthesis process, thereby enhancing speech expressivity while managing complexity through parameter-based control

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9767789B2Using emoticons for contextual text-to-speech expressivity
Publication Date: 2017.09.19 CERENCE OPERATING CO
  • US9767789B2 patent drawing
  • US9767789B2 patent drawing
  • US9767789B2 patent drawing

AI summary

Techniques disclosed herein include systems and methods that improve audible emotional characteristics used when synthesizing speech from a text source. Systems and methods herein use emoticons identified from a source text to provide contextual text-to-speech expressivity. In general, techniques herein analyze text and identify emoticons included within the text. The source text is then tagged with corresponding mood indicators. For example, if the system identifies an emoticon at the end of a sentence, then the system can infer that this sentence has a specific tone or mood associated with it. Depending on whether the emoticon is a smiley face, angry face, sad face, laughing face, etc., the system can infer use or mood from the various emoticons and then change or modify the expressivity of the TTS output such as by changing intonation, prosody, speed, pauses, and other expressivity characteristics.