Emoticon-Based Text-to-Speech Expressivity Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text-to-speech systems struggle to accurately convey emotion or audible expressivity, relying on limited analysis of punctuation and word arrangement, which often fails to capture the intended mood of the text, especially when the composer's mood varies.
Innovation Solution
The use of emoticons as contextual cues to enhance the expressivity of text-to-speech synthesis, where identified emoticons in a text string are used to modify audio output characteristics such as intonation, prosody, speed, and pauses, thereby improving the perception of emotional tone.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional TTS systems analyze only punctuation and word arrangement to determine mood, then the system complexity remains low, but the accuracy of emotional expression deteriorates
Solution Approach 1:
The system performs preliminary identification and tagging of emoticons in the text before the main text-to-speech conversion process. By pre-processing the text to detect emoticons and attach expressivity tags, the system prepares emotional context information in advance, enabling more accurate mood detection without significantly increasing overall system complexity during the conversion phase
Solution Approach 2:
The patent introduces an intermediary component (emoticon identifier and expressivity tagger) that bridges the gap between simple text input and emotional expression in speech output. This intermediary layer processes emoticons and generates expressivity tags that guide the TTS engine, thereby improving mood detection accuracy while isolating the complexity to a specific module rather than the entire system
2Reliability
If TTS systems use basic punctuation analysis to infer emotion, then the processing speed remains fast, but the reliability of emotional conveyance deteriorates
Solution Approach 1:
The system segments the text processing task into distinct components: emoticon identification, expressivity tag generation, and text-to-speech conversion with expressivity guidance. By dividing the processing into these segments, the system can efficiently handle emoticon detection as a separate step while maintaining overall processing speed, and improve reliability through specialized handling of emotional cues
3Adaptability or versatility
If TTS systems ignore emoticons or speak their literal names, then the system simplicity is maintained, but the expressivity and naturalness of speech output deteriorates
Solution Approach 1:
The system changes the parameter of text representation by introducing expressivity tags that encode emotional information from emoticons. Instead of ignoring emoticons or speaking them literally, the system transforms them into actionable parameters (expressivity tags) that modify the speech synthesis process, thereby enhancing speech expressivity while managing complexity through parameter-based control
Data Source
AI summary
Techniques disclosed herein include systems and methods that improve audible emotional characteristics used when synthesizing speech from a text source. Systems and methods herein use emoticons identified from a source text to provide contextual text-to-speech expressivity. In general, techniques herein analyze text and identify emoticons included within the text. The source text is then tagged with corresponding mood indicators. For example, if the system identifies an emoticon at the end of a sentence, then the system can infer that this sentence has a specific tone or mood associated with it. Depending on whether the emoticon is a smiley face, angry face, sad face, laughing face, etc., the system can infer use or mood from the various emoticons and then change or modify the expressivity of the TTS output such as by changing intonation, prosody, speed, pauses, and other expressivity characteristics.


