Avatar Animation via Texture Wrapping and Phoneme Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current mobile communication systems face challenges in implementing real-time video conferencing due to the high bandwidth required for audio and video data transmission, particularly in displaying a representation of a cellular phone user across a network, with no existing low-bandwidth or lightweight method available.
Innovation Solution
The method involves an animation module on mobile devices that displays an avatar representing a participant by receiving a text source, identifying images, selecting a generic animation template, texture wrapping images over the template, and synchronously playing audio speech signals with mouth position alterations to animate the avatar's speech, using phonemes and voice inflections to simulate speaking.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real-time video conferencing is implemented between cellular phones across a network, then communication quality and user representation are improved, but bandwidth consumption increases significantly
Solution Approach 1:
The patent creates a simplified copy (avatar) of the user instead of transmitting actual video footage. The avatar is a lightweight graphical representation that can be animated locally on the receiving device, eliminating the need to transmit continuous video streams while still providing visual representation and communication quality
Solution Approach 2:
The patent segments the visual representation into discrete graphical elements (avatar components, animations, expressions) that can be transmitted efficiently. Instead of transmitting continuous video data, the system transmits segmented graphical assets and animation parameters, significantly reducing bandwidth requirements
2Quantity of substance
If lightweight avatar animation is used instead of real-time video, then bandwidth requirements are reduced, but visual representation fidelity decreases
Solution Approach 1:
The patent implements dynamic animation of the avatar through programmatically generated movements, expressions, and gestures. The avatar can exhibit various states (talking, listening, emoting) through animation sequences, providing a lively and engaging representation that compensates for the simplified visual model
Solution Approach 2:
The system uses parameter-based animation control where communication metadata (speaker identification, speech activity, emotional state) drives changes in avatar parameters (mouth position, eye movement, facial expression). This allows the avatar to dynamically reflect the communication state without requiring high-fidelity video
Data Source
AI summary
Animating speech of an avatar representing a participant in a mobile communication including selecting one or more images; selecting a generic animation template; fitting the one or more images with the generic animation template; texture wrapping the one more images over the generic animation template; and displaying the one or more images texture wrapped over the generic animation template. Receiving an audio speech signal; identifying a series of phonemes; and for each phoneme: identifying a new mouth position for the mouth of the generic animation template; altering the mouth position to the new mouth position; texture wrapping a portion of the one or more images corresponding to the altered mouth position; displaying the texture wrapped portion of the one or more images corresponding to the altered mouth position of the mouth of the generic animation template; and playing the portion of the audio speech signal represented by the phoneme.


