Avatar Animation via Texture Wrapping and Phoneme Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current mobile communication systems face challenges in implementing real-time video conferencing due to the high bandwidth required for audio and video data transmission, particularly in displaying a representation of a cellular phone user across a network, with no existing low-bandwidth or lightweight method available.

Innovation Solution

The method involves an animation module on mobile devices that displays an avatar representing a participant by receiving a text source, identifying images, selecting a generic animation template, texture wrapping images over the template, and synchronously playing audio speech signals with mouth position alterations to animate the avatar's speech, using phonemes and voice inflections to simulate speaking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-time video conferencing is implemented between cellular phones across a network, then communication quality and user representation are improved, but bandwidth consumption increases significantly

Engineering Contradiction:
Improvecommunication qualityVSAvoidbandwidth consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates a simplified copy (avatar) of the user instead of transmitting actual video footage. The avatar is a lightweight graphical representation that can be animated locally on the receiving device, eliminating the need to transmit continuous video streams while still providing visual representation and communication quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent segments the visual representation into discrete graphical elements (avatar components, animations, expressions) that can be transmitted efficiently. Instead of transmitting continuous video data, the system transmits segmented graphical assets and animation parameters, significantly reducing bandwidth requirements

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If lightweight avatar animation is used instead of real-time video, then bandwidth requirements are reduced, but visual representation fidelity decreases

Engineering Contradiction:
Improvebandwidth requirementsVSAvoidvisual representation fidelity
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent implements dynamic animation of the avatar through programmatically generated movements, expressions, and gestures. The avatar can exhibit various states (talking, listening, emoting) through animation sequences, providing a lively and engaging representation that compensates for the simplified visual model

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses parameter-based animation control where communication metadata (speaker identification, speech activity, emotional state) drives changes in avatar parameters (mouth position, eye movement, facial expression). This allows the avatar to dynamically reflect the communication state without requiring high-fidelity video

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8125485B2Animating speech of an avatar representing a participant in a mobile communication
Publication Date: 2012.02.28 ACTIVISION PUBLISHING INC
  • US8125485B2 patent drawing
  • US8125485B2 patent drawing
  • US8125485B2 patent drawing

AI summary

Animating speech of an avatar representing a participant in a mobile communication including selecting one or more images; selecting a generic animation template; fitting the one or more images with the generic animation template; texture wrapping the one more images over the generic animation template; and displaying the one or more images texture wrapped over the generic animation template. Receiving an audio speech signal; identifying a series of phonemes; and for each phoneme: identifying a new mouth position for the mouth of the generic animation template; altering the mouth position to the new mouth position; texture wrapping a portion of the one or more images corresponding to the altered mouth position; displaying the texture wrapped portion of the one or more images corresponding to the altered mouth position of the mouth of the generic animation template; and playing the portion of the audio speech signal represented by the phoneme.