Animated Message Delivery System Using Text-to-Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing communication methods fail to effectively transform electronic messages into moving images of selected characters that utter the message, lacking the ability to convert text into voice and display animated characters on screens.

Innovation Solution

Implementing a system that uses text-to-speech synthesis and computer animation to convert electronic messages into speech and generate moving images of selected animation characters, allowing users to compose messages that are presented as animated videos on devices like smartphones and computer networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text-to-speech synthesis and computer animation are implemented to convert electronic messages into speech and generate moving images, then message engagement and user interaction are improved, but device complexity and processing requirements increase

Engineering Contradiction:
Improvemessage engagementVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system divides the message delivery process into separate functional modules: text-to-speech synthesis module, animation generation module, and video rendering module. This segmentation allows each component to be optimized independently and facilitates distributed processing across different devices, reducing the complexity burden on any single device while maintaining high engagement through integrated animated message delivery

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A cloud-based processing server acts as an intermediary between the user's device and the final animated message output. The server receives text input, processes it through text-to-speech synthesis, generates animated character sequences, and outputs video files. This intermediary approach offloads complex processing from user devices, enabling high-engagement animated messages without requiring sophisticated local hardware

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If text is converted into speech and animated character images are generated and transmitted, then message impact and user interaction are enhanced, but transmission time and processing duration increase

Engineering Contradiction:
Improvemessage impactVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system pre-processes and generates animated character images and speech audio separately before final assembly. Text-to-speech synthesis is performed in advance to create audio tracks, and animation frames are pre-generated. This preliminary action allows for optimized processing pipelines where time-critical components can be prepared ahead of time, reducing overall delivery latency while maintaining high message impact through quality animated presentations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The animation generation process uses periodic rendering of character frames based on synthesized speech timing. Rather than generating all frames simultaneously, the system renders animation frames periodically as the speech audio is generated, allowing parallel processing of audio synthesis and visual animation. This periodic action reduces total processing time while maintaining the synchronized impact of the final animated message

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS9667574B2Animated delivery of electronic messages
Publication Date: 2017.05.30 MITII
  • US9667574B2 patent drawing
  • US9667574B2 patent drawing
  • US9667574B2 patent drawing

AI summary

An electronic message is transformed into moving images uttering the content of the electronic message. Methods of the present invention may be implemented on devices such as smart phones to enable users to compose text and select an animation character which may include cartoons, persons, animals, or avatars. The recipient is presented with an animation or video of the animation character with a voice that speaks the words of the text. The user may further select and include a catch-phrase associated with the character. The user may further select a background music identifier and a background music associated with the background music identifier is played back while the animated text is being presented. The user may further select a type of animation and the animation character will be animated according to the type of animation.