Audio-Visual Avatar Synthesis With Voice-Video Custom Cloning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in creating virtual avatars with high naturalness and realism, particularly in terms of voice and visual appearance, due to computational complexity and time inefficiencies in processing multiple input sources, which hinders their effectiveness in applications like tutoring.

Innovation Solution

A method and system for creating audio-visual avatars involves training a synthesis module with audio and video datasets, followed by customization for a target person, using voice and video custom synthesis modules to generate voice and video clones that mimic the target's speech and physical characteristics, including lip movements, gestures, and body postures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If multiple input sources are processed to create realistic avatars, then the naturalness and realism of the avatar is improved, but the computational complexity and time consumption increase

Engineering Contradiction:
Improveavatar realismVSAvoidcomputation complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system segments the avatar creation process into distinct modules: audio processing module, video processing module, and synthesis module. Each module handles specific aspects (audio training data processing, video training data processing, and final avatar synthesis respectively), allowing parallel development and optimization without increasing overall system complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-processing and organizing training data before the actual avatar synthesis. Audio training data and video training data are separately processed and prepared in advance, creating ready-to-use datasets that reduce computational burden during the final synthesis phase

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If multiple input sources are processed to create realistic avatars, then the naturalness and realism of the avatar is improved, but the time consumption increases

Engineering Contradiction:
Improveavatar realismVSAvoidcreation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system divides the data processing into separate audio and video streams that can be processed independently and in parallel. This segmentation allows simultaneous processing of multiple input sources without sequential delays, reducing total creation time while maintaining realism

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Training data is pre-processed and organized into structured formats before synthesis begins. Audio data is pre-analyzed for speech patterns and video data is pre-analyzed for visual characteristics, so that the actual avatar generation requires minimal additional processing time

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12561876B2System and method for an audio-visual avatar creation
Publication Date: 2026.02.24 SIT AUTONOMOUS AG
  • US12561876B2 patent drawing
  • US12561876B2 patent drawing
  • US12561876B2 patent drawing

AI summary

The present disclosure relates to an avatar generator to generate an audio-visual avatar specific to an application, such as tutoring. The avatar generator includes a general synthesizer to receive a training dataset. The general synthesizer includes a voice synthesis module and a video synthesis module trained by the training dataset. The avatar generator includes a customized synthesizer consisting of a voice custom synthesis module and a video custom synthesis module trained on the audio-video samples of the target person. The avatar generator further includes a video generator to create an audio-visual avatar and is configured to synthesize a voice clone using an input text, process the voice clone, synthesize a video clone based on the video synthesis module and the video custom synthesis module, and apply the voice clone to the video clone.