Voice-Driven Lip Syncing via Neural Network Audio Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current lip syncing methods in applications like virtual assistants and video conferencing suffer from synchronization deficiencies between voice and lip movements, necessitating more realistic and efficient audio-driven video avatars and text-driven video solutions.
Innovation Solution
A system and method for voice-driven animation using artificial neural networks (ANNs) that samples input videos, extracts voice-related features from audio data, generates expression representations, and combines these with auxiliary image data to produce synchronized lip and face animations, including a training phase to refine the animation process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If commonly used lip syncing methods are employed, then the implementation is simple and fast, but the synchronization between voice and lip movement is deficient and unrealistic
Solution Approach 1:
The patent replaces traditional mechanical or rule-based lip syncing methods with an artificial neural network system. The ANN learns the complex mapping between audio features and facial expressions from training data, substituting manual synchronization rules with an intelligent model that automatically generates realistic lip movements matching the audio input.
Solution Approach 2:
The system transforms audio data into multiple parameter representations including spectrogram features, acoustic word embeddings, and voice-related features. These parameter transformations enable the neural network to capture nuanced audio characteristics and map them to corresponding facial expression parameters, improving synchronization accuracy through multi-parameter analysis.
2Reliability
If audio-driven facial video synthesis is used, then lip syncing can be achieved, but the synchronization between voice and lip movement remains deficient
Solution Approach 1:
The neural network is trained using feedback from paired audio-visual data where the ground truth facial expressions are compared against generated expressions. The loss function calculates the discrepancy between predicted and actual lip movements, providing feedback that guides the network to improve synchronization reliability and lip movement accuracy through iterative optimization.
Solution Approach 2:
The system performs preliminary action by pre-processing audio data into multiple feature representations (spectrogram, acoustic embeddings, voice features) before feeding them to the neural network. This preliminary transformation of audio into structured parameters enables more reliable and accurate generation of lip movement sequences.
Data Source
AI summary
A system and method voice driven animation of an object in an image by sampling an input video, depicting a puppet object, to obtain an image, receiving audio data, extracting voice related features from the audio data, producing an expression representation based on the voice related features, wherein the expression representation is related to a region of interest, obtaining from the image, auxiliary data related to the image and generating a target image based on the expression representation and the auxiliary data.


