Personalized Avatar Animation via Audio-Driven ML Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing avatar animation systems that rely solely on audio input fail to accurately represent a user's characteristic movements, such as facial expressions and head pose, leading to a degraded user experience and reduced ability to express or perceive personality during communication.

Innovation Solution

A computer-implemented method that involves training a personalized machine learning model using video data from a user to predict their movements based on audio input, allowing for the generation of authentic avatar animations that reflect the user's characteristic movements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If a generic avatar is used with audio-driven animation, then bandwidth and processing requirements are reduced, but the ability to express and perceive user personality is degraded

Engineering Contradiction:
Improvebandwidth and processing requirementsVSAvoiduser personality expression
Core Design Contradiction:
Loss of energyVSLoss of information

Solution Approach 1:

The system performs preliminary action by capturing and storing user-specific movement characteristics during a training phase before actual communication. Video data is collected and used to train a machine learning model that learns the user's characteristic movements, head poses, and facial expressions. This pre-trained model then enables personalized avatar animation during communication without requiring real-time video transmission, thus reducing bandwidth while preserving personality expression.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If video data is transmitted for real-time avatar animation, then user personality and characteristic movements are accurately represented, but bandwidth and processing requirements increase

Engineering Contradiction:
Improveuser personality expressionVSAvoidbandwidth and processing requirements
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The system creates a digital copy of the user's movement characteristics through a trained machine learning model. Instead of transmitting actual video data, the system uses the trained model to generate avatar animations that replicate the user's characteristic movements. This copying approach preserves personality expression while avoiding the high bandwidth requirements of real-time video transmission.

Inventive Principle:
Principle #26Copying

3Device complexity

If audio-driven animation uses simple mouth movement only, then processing requirements are reduced, but the authenticity of user representation is degraded

Engineering Contradiction:
Improveprocessing requirementsVSAvoidauthenticity of user representation
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system performs preliminary action by capturing and storing user-specific movement characteristics during a training phase before actual communication. Video data is collected and used to train a machine learning model that learns the user's characteristic movements, head poses, and facial expressions. This pre-trained model then enables personalized avatar animation during communication without requiring real-time video transmission, thus reducing bandwidth while preserving personality expression.

Inventive Principle:
Principle #10Preliminary action

4Loss of information

If full motion avatar animation is implemented, then user personality expression is improved, but processing requirements and complexity increase

Engineering Contradiction:
Improveuser personality expressionVSAvoidprocessing requirements
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system replaces the mechanical system of real-time video capture and transmission with a machine learning-based prediction system. A trained model predicts avatar movements from audio input alone, substituting the need for complex real-time video processing. This approach enables full motion avatar animation with reduced processing requirements by using learned patterns instead of real-time computational geometry and physics.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250069308A1Method, apparatus and computer program
Publication Date: 2025.02.27 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250069308A1 patent drawing
  • US20250069308A1 patent drawing
  • US20250069308A1 patent drawing

AI summary

A computer-implemented method comprising: receiving, from a user device, video data from a user; training a first machine learning model based on the video data to provide a second machine learning model, the second machine learning model being personalized to the user, wherein the second machine learning model is trained to predict movement of the user based on audio data; receiving further audio data from the user; determining predicted movements of the user based on the further audio data and the second machine learning model; using the predicted movements of the user to generate animation of an avatar of the user.