Text-to-Visual Speech System for Emotion-Conveying Instant Messaging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-to-visual speech systems fail to effectively convey emotions in text-based communication, as they rely on typed emoticons or video technology, which are limited by small screen sizes and high bandwidth requirements, and do not accurately represent the intended emotional state of the sender.
Innovation Solution
A method and system that analyze text messages to generate phoneme and wave data, mapping it to viseme data to create an animated image that conveys emotions, using a Text-to-Speech engine, phoneme bigram table, and animation generation components to produce synchronized face and lip frames based on the emotional content of the message.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If video technology is used to convey emotions clearly, then emotional expression is improved, but network bandwidth consumption and data processing resources increase significantly
Solution Approach 1:
The patent creates a visual speech system that generates synthetic talking head images from text messages. Instead of using actual video feeds, the system synthesizes visual representations of speech by mapping phoneme data to viseme data, which controls animated face and lip frames. This copying approach replicates the essential visual elements of speech without requiring bandwidth-intensive video transmission.
Solution Approach 2:
The patent extracts only the essential visual components needed for emotional expression - specifically face and lip frames that correspond to phonemes. By taking out only these critical elements rather than transmitting full video streams, the system achieves effective emotional communication while minimizing network bandwidth consumption and processing requirements.
2Loss of information
If emoticons are used to personalize text messages, then emotional expression is improved, but screen space for text messages is reduced
Solution Approach 1:
The patent transitions from two-dimensional text-based emoticons to a more immersive visual dimension by generating talking head images. Instead of using text symbols like :-) that occupy screen space, the system creates animated visual representations that convey emotion and speech simultaneously, allowing text messages to remain prominent while emotional expression is enhanced through visual imagery.
3Ease of operation
If text-to-visual speech systems use typed text input, then message creation is simplified, but emotional accuracy of the output is reduced
Solution Approach 1:
The patent incorporates emotion tag detection and feedback mechanisms that analyze the sent message and adjust the generated visual speech accordingly. The system detects emotion tags in the input text, maps them to appropriate viseme data, and generates corresponding animated expressions. This feedback loop ensures that the visual output accurately reflects the intended emotional state while maintaining the simplicity of text-based input.
Data Source
AI summary
Emotions can be expressed in the user interface for an instant messaging system based on the content of a received text message. The received text message is analyzed using a text-to-speech engine to generate phoneme data and wave data based on the text content. Emotion tags embedded in the message by the sender are also detected. Each emotion tag indicates the sender's intent to change the emotion being conveyed in the message. A mapping table is used to map phoneme data to viseme data. The number of face/lip frames required to represent viseme data is determined based on at least the length of the associated wave data. The required number of face/lip frames is retrieved from a stored set of such frames and used in generating an animation. The retrieved face/lip frames and associated wave data are presented in the user interface as synchronized audio/video data.


