Automatic Viseme Detection for Animatable Puppet Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer animation methods require manual creation and modification of characters to depict visemes, which is time-consuming and can result in lower-quality animations, especially for users who struggle to perform specific mouth shapes and indicate key frames accurately.
Innovation Solution
A computing device automatically detects video frames depicting visemes by aligning audio data with annotated reference audio datasets, allowing for the extraction and tagging of relevant frames to generate animatable puppets, reducing the need for manual input and improving animation quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual creation and modification of characters is used to depict visemes, then users can create customized characters, but the process is time-consuming and may result in lower-quality animations
Solution Approach 1:
The system captures real facial viseme images from video frames and uses them as templates to generate corresponding animated puppet visemes. This copying approach transfers the accuracy of real human facial expressions directly to the animated character, eliminating manual creation while preserving high quality
Solution Approach 2:
The manual mechanical process of creating and adjusting viseme frames is replaced with an automated computer vision system that detects, extracts, and synchronizes viseme images programmatically. This substitution eliminates the time-consuming manual operations while maintaining or improving animation quality through consistent automated processing
2Manufacturing precision
If users manually perform specific mouth shapes and indicate key frames, then viseme data can be captured, but users who struggle with this process produce lower-quality animations
Solution Approach 1:
The system performs automatic viseme detection and extraction without requiring user intervention. The computer vision algorithm independently identifies mouth shapes in video frames, extracts the relevant images, and synchronizes them with audio data, eliminating the need for users to manually perform or indicate visemes
Solution Approach 2:
The manual user operation of performing mouth shapes and indicating key frames is replaced with an automated image recognition and audio synchronization system. This substitution removes the skill barrier and operational difficulty while improving accuracy through consistent automated detection
3Productivity
If automatic viseme detection is implemented, then animation quality and efficiency are improved, but the system complexity increases
Solution Approach 1:
The system integrates multiple functions into a unified automated pipeline: video frame processing, viseme detection, image extraction, audio analysis, and synchronization all occur in sequence without requiring separate manual operations. This multi-functionality improves productivity while containing complexity within a single integrated system
Data Source
AI summary
Certain embodiments involve automatically detecting video frames that depict visemes and that are usable for generating an animatable puppet. For example, a computing device accesses video frames depicting a person performing gestures usable for generating a layered puppet, including a viseme gesture corresponding to a target sound or phoneme. The computing device determines that audio data including the target sound or phoneme aligns with a particular video frame from the video frames that depicts the person performing the viseme gesture. The computing device creates, from the video frames, a puppet animation of the gestures, including an animation of the viseme corresponding to the target sound or phoneme that is generated from the particular video frame. The computing device outputs the puppet animation to a presentation device.


