3D Avatar Video Generation for Real-Time Expression Sync

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional avatar generation methods lack flexibility and fail to reflect real-time changes in facial expressions, leading to identity theft and network security issues due to the lack of unique user identity representation.

Innovation Solution

A method and system for generating a two-dimensional avatar image based on a reference image and a video frame, transforming it into a three-dimensional avatar using multi-modality features, including motion, audio, and text, to create a customized three-dimensional avatar video that aligns with the user's unique features and expressions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional avatar generation methods are used, then the process is simple, but the avatar lacks flexibility and cannot reflect real-time changes in facial expressions

Engineering Contradiction:
Improvereal-time expression reflectionVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-extracting facial feature points and establishing correspondence between 2D avatar images and 3D models before real-time processing. The facial feature point extraction and matching algorithms are prepared in advance, enabling rapid real-time expression tracking without complex processing during actual avatar generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements dynamics by continuously tracking facial feature points across video frames and dynamically updating the 3D avatar's facial expressions in real-time. The facial feature correspondence relationship allows dynamic mapping of 2D expression changes to 3D avatar movements, enabling flexible real-time adaptation.

Inventive Principle:
Principle #15Dynamics

2Reliability

If conventional avatar generation methods are used, then the implementation is straightforward, but user privacy cannot be protected due to lack of unique identity representation

Engineering Contradiction:
Improveprivacy protectionVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system applies local quality by extracting and protecting specific local facial feature points rather than processing entire facial images. By focusing on key facial landmarks and their correspondence relationships, the system maintains unique identity representation for privacy protection while avoiding comprehensive biometric data processing.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system creates a stylized 2D avatar image that copies essential facial characteristics from the user's reference image. This avatar copy serves as a unique identity representation that protects privacy by not storing or transmitting original biometric data, while still enabling reliable user identification.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If detailed three-dimensional transformation is performed, then the avatar representation becomes more vivid, but the processing time and computational resources increase

Engineering Contradiction:
Improveavatar detail precisionVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system segments the complex 3D transformation process into distinct stages: 2D avatar generation, facial feature point extraction, feature correspondence establishment, and 3D model animation. This segmentation allows each module to be optimized independently, achieving detailed avatar representation while managing processing time through modular computation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial action by focusing computational resources on key facial feature points rather than processing entire facial surfaces. By tracking only essential landmarks (eyes, nose, mouth corners, etc.) and their correspondence, the system achieves vivid expression representation with reduced computational overhead compared to full-face 3D reconstruction.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12482163B2Method, device, and computer program product for processing video
Publication Date: 2025.11.25 DELL PROD LP
  • US12482163B2 patent drawing
  • US12482163B2 patent drawing
  • US12482163B2 patent drawing

AI summary

Methods, devices and computer program products for processing video are disclosed herein. A method includes: generating, based on a reference image and a first frame of a video comprising an object, a two-dimensional avatar image of the object; and generating a base three-dimensional avatar of the object by performing a three-dimensional transformation on the two-dimensional avatar image and the object in the first frame. The method further includes: generating a three-dimensional avatar video corresponding to the video based on the base three-dimensional avatar and features of the video, the features comprising differences of the object between adjacent frames of the video. This solution enables the generation of a customized three-dimensional avatar video for an object in a video, where the avatar can move in synchronization with the object and retain the unique features of the object, and can provide a more detailed and vivid representation than a two-dimensional avatar.