Audio-Visual Avatar Synthesis With Voice-Video Custom Cloning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in creating virtual avatars with high naturalness and realism, particularly in terms of voice and visual appearance, due to computational complexity and time inefficiencies in processing multiple input sources, which hinders their effectiveness in applications like tutoring.
Innovation Solution
A method and system for creating audio-visual avatars involves training a synthesis module with audio and video datasets, followed by customization for a target person, using voice and video custom synthesis modules to generate voice and video clones that mimic the target's speech and physical characteristics, including lip movements, gestures, and body postures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If multiple input sources are processed to create realistic avatars, then the naturalness and realism of the avatar is improved, but the computational complexity and time consumption increase
Solution Approach 1:
The system segments the avatar creation process into distinct modules: audio processing module, video processing module, and synthesis module. Each module handles specific aspects (audio training data processing, video training data processing, and final avatar synthesis respectively), allowing parallel development and optimization without increasing overall system complexity
Solution Approach 2:
The system performs preliminary actions by pre-processing and organizing training data before the actual avatar synthesis. Audio training data and video training data are separately processed and prepared in advance, creating ready-to-use datasets that reduce computational burden during the final synthesis phase
2Manufacturing precision
If multiple input sources are processed to create realistic avatars, then the naturalness and realism of the avatar is improved, but the time consumption increases
Solution Approach 1:
The system divides the data processing into separate audio and video streams that can be processed independently and in parallel. This segmentation allows simultaneous processing of multiple input sources without sequential delays, reducing total creation time while maintaining realism
Solution Approach 2:
Training data is pre-processed and organized into structured formats before synthesis begins. Audio data is pre-analyzed for speech patterns and video data is pre-analyzed for visual characteristics, so that the actual avatar generation requires minimal additional processing time
Data Source
AI summary
The present disclosure relates to an avatar generator to generate an audio-visual avatar specific to an application, such as tutoring. The avatar generator includes a general synthesizer to receive a training dataset. The general synthesizer includes a voice synthesis module and a video synthesis module trained by the training dataset. The avatar generator includes a customized synthesizer consisting of a voice custom synthesis module and a video custom synthesis module trained on the audio-video samples of the target person. The avatar generator further includes a video generator to create an audio-visual avatar and is configured to synthesize a voice clone using an input text, process the voice clone, synthesize a video clone based on the video synthesis module and the video custom synthesis module, and apply the voice clone to the video clone.


