Lip Animation Synthesis Using Phonetic Symbol Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current animation display systems for robots lack diversity and realism in simulating lip movements during speech, as they primarily rely on simple mouth opening and closing without sophisticated lip animation synchronization.
Innovation Solution
An animation display system comprising a processor, storage, and display that converts input text into phonetic symbols and timestamps, matches these symbols with corresponding lip movements using a phonetic-symbol lip-motion matching database, and synthesizes lip animations using a lip motion synthesis database, enabling synchronized speech-lip animation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If simple mouth opening and closing animation is used, then device complexity is reduced, but animation realism and diversity deteriorate
Solution Approach 1:
The patent introduces an intermediary lip animation synthesis system that mediates between the speech output and the visual display. This system uses phonetic symbol databases and motion synthesis algorithms to generate realistic lip movements that correspond to the spoken words, rather than using simple mouth opening and closing animations. The intermediary system processes phonetic information and transforms it into corresponding visual lip motions, thereby improving animation realism without requiring complex direct control mechanisms.
2Reliability
If phonetic-symbol lip-motion matching database is used, then animation diversity and realism are improved, but device complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-building comprehensive phonetic-symbol lip-motion matching databases before actual speech animation is needed. During system operation, the database simply needs to retrieve pre-computed lip motion sequences corresponding to phonetic symbols, rather than computing animations in real-time. This pre-processing approach stores complex motion data in advance, allowing the system to quickly generate diverse and realistic lip animations during actual speech without requiring complex real-time computation.
3Measurement precision
If lip motion synthesis module generates synchronized animations, then speech-lip synchronization accuracy is improved, but processing time increases
Solution Approach 1:
The patent uses preliminary action by pre-computing and storing lip motion sequences in the phonetic-symbol lip-motion matching database before actual speech synthesis is needed. During real-time operation, the system retrieves pre-computed motion sequences based on phonetic symbols and timestamps, rather than computing synchronized animations in real-time. This approach maintains high synchronization accuracy between speech and lip movements while significantly reducing processing time during actual speech output.
Data Source
AI summary
An animation display system is provided. The animation display system includes a display; a storage configured to store a language model database, a phonetic-symbol lip-motion matching database and a lip motion synthesis database; and a processor electronically connected to the storage and the display, respectively. The processor includes a speech conversion module, a phonetic-symbol lip-motion matching module, and a lip motion synthesis module. A lip animation display method is also provided.


