Dubbed Audio Lip-Viseme Correlation for Lip-Sync Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Dubbed audio often fails to match facial expressions in video, leading to a confusing user experience due to synchronization issues between audio and video.
Innovation Solution
Analyze lip poses in the source video and expected visemes based on dubbed audio, compare their similarity, and adjust the dubbing process or the original video to align lip poses with dubbed audio using computer vision techniques and machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated dubbing is used to provide media presentation quickly and reliably, then productivity and reliability are improved, but synchronization between lip poses and dubbed audio deteriorates
Solution Approach 1:
The system extracts lip pose features from video frames and audio features from dubbed audio, compares them using a neural network model, and uses the synchronization score as feedback to identify and correct mismatched segments, enabling continuous improvement of dubbing quality
Solution Approach 2:
The patent replaces manual lip-sync checking with an automated system using computer vision to extract lip pose features, neural networks to predict visemes, and machine learning models to calculate synchronization scores, substituting mechanical/manual processes with intelligent automated analysis
2Manufacturing precision
If manual dubbing with human speakers is used to improve lip-sync accuracy, then synchronization accuracy is improved, but productivity and cost efficiency deteriorate
Solution Approach 1:
The system performs self-evaluation by automatically analyzing its own dubbed output, comparing extracted lip poses with audio visemes, and identifying synchronization issues without requiring external manual review, enabling autonomous quality control
Solution Approach 2:
The patent introduces an intermediate automated analysis layer between the dubbing process and final output, using feature extraction, neural network prediction, and synchronization scoring as mediators to objectively assess and improve lip-sync quality
3Productivity
If dubbing audio is translated to match source audio timing, then productivity is improved, but lip pose matching deteriorates due to language-specific pronunciation differences
Solution Approach 1:
The system changes the parameter being optimized from temporal alignment to visual-auditory feature matching, using neural networks to predict visemes from audio and compare them with extracted lip pose features, accounting for language-specific pronunciation characteristics
Data Source
AI summary
Methods and apparatus are described for evaluating dubbing of media content. Phonemes in dubbed audio are extracted and mapped to visemes. Lip poses in video frames of the media content corresponding to the phonemes of the dubbed audio are compared to the visemes determined from the dubbed audio. A notification may be generated based on the comparison that indicates synchronization of the dubbed audio to lip poses of the video.


