Song-Specific Viseme-Timed Facial Effects for AR Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Lip-synching to a song while capturing a video is difficult, especially for new songs, as it requires precise timing and knowledge of lyrics, making it challenging for users to achieve a synchronized video performance.
Innovation Solution
An animation system that animates a user's face in a video stream based on song-specific data files containing viseme identifiers and timestamps, aligning facial expressions with the song's lyrics using viseme analysis and computer vision techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If users manually lip-synch to a song while capturing video, then they can create video performances, but it requires precise timing and knowledge of lyrics which makes it difficult and complex
Solution Approach 1:
The system automatically performs lip-synching by detecting the user's facial movements and automatically synchronizing them with the audio track's visemes. The user simply records their natural speech or singing, and the system handles the complex alignment of visual and audio elements without requiring manual timing or lyric knowledge, thus making the operation easy while eliminating complexity
Solution Approach 2:
The patent replaces the manual mechanical process of timing and coordinating lip movements with audio playback with an automated computer vision and machine learning system. The system uses facial landmark detection, viseme identification, and temporal alignment algorithms to automatically synchronize the user's face with the audio track, substituting complex manual coordination with intelligent automated processing
2Adaptability or versatility
If users lip-synch without knowing lyrics, then accessibility improves, but synchronization accuracy deteriorates
Solution Approach 1:
The system continuously monitors the user's facial movements during recording and compares them against the expected viseme sequence from the audio track. It uses real-time feedback from facial landmark detection to adjust and refine the synchronization, ensuring accurate alignment even when the user doesn't know the lyrics beforehand. The system feedback loop maintains precision while allowing user versatility
Solution Approach 2:
The system performs preliminary analysis of the audio track to extract viseme timestamps and sequences before the user records. This pre-processing creates a reference framework that guides the synchronization process, allowing users to record naturally without lyric knowledge while the system ensures accurate alignment by comparing their facial movements against the pre-analyzed viseme sequence
3Productivity
If real-time facial animation is applied, then user engagement improves, but processing time and computational resources increase
Solution Approach 1:
The system segments the facial animation process into distinct stages: facial landmark detection, viseme identification, temporal alignment, and rendering. By dividing the complex real-time processing into manageable segments that can be processed in parallel or pipelined, the system achieves real-time performance without overwhelming computational resources or excessive processing time
Solution Approach 2:
The system uses periodic action by processing video frames at regular intervals and updating facial animations at key viseme transition points rather than continuously. This periodic processing approach maintains real-time user engagement while reducing overall computational load and processing time compared to continuous frame-by-frame analysis
Data Source
AI summary
An audio track with vocals is played back using a device with a display screen that displays a video feed from a camera. A location of a mouth depicted in the video feed is detected. A timestamp of playback of the audio track is compared to viseme-timestamp data for the audio track to identify a viseme corresponding to the timestamp of the audio playback a viseme is positioned at the detected location of the mouth in the video feed.


