Lip-Sync Animation Model Learning via Acoustic Feature Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current lip-sync animation generation technologies lack stability and naturalness in their output, failing to effectively utilize voice data to create realistic character animations.
Innovation Solution
A model learning apparatus comprising a voice model and a rig model, where the voice model extracts acoustic feature values through signal processing and transformation, and the rig model extracts frame feature values to output character control information for controlling facial expressions and animations, using techniques like convolutional neural networks and transformer encoders for improved feature extraction and animation synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional lip-sync animation generation technology is used, then the process is simple, but the output lacks stability and naturalness
Solution Approach 1:
The system segments the animation generation process into distinct functional modules: a voice model that extracts acoustic feature values from voice data, and a rig model that generates animation based on these features. This segmentation allows each module to be optimized independently, improving overall reliability while managing complexity through modular architecture.
Solution Approach 2:
The patent introduces acoustic feature values as an intermediary representation between raw voice data and animation output. The voice model transforms voice data into acoustic feature values, which then serve as input to the rig model. This intermediary layer enables more stable and natural animation generation by decoupling the complexity of voice processing from animation synthesis.
2Manufacturing precision
If acoustic feature value extraction is implemented, then naturalness of animation improves, but processing time increases
Solution Approach 1:
The voice model performs preliminary extraction of acoustic feature values from voice data before animation generation. By pre-processing the voice data into meaningful acoustic features (such as pitch, timbre, and temporal characteristics), the system prepares optimized input for the rig model, improving animation quality while enabling efficient processing through staged computation.
Solution Approach 2:
The patent replaces direct mechanical mapping from voice data to animation with a learned transformation system. Instead of using rigid rule-based approaches, the voice model learns to extract acoustic feature values that capture essential vocal characteristics, enabling more natural and efficient animation generation that better balances quality and processing time.
Data Source
AI summary
Embodiments of the present disclosure provide methods, systems and non-transitory computer readable media of performing voice model learning and rig model learning. The voice model learning includes extracting an acoustic feature value by executing predetermined acoustic signal processing with respect to voice data including human voice and extracting a voice feature value by executing first transformation processing with respect to first input information including the extracted acoustic feature value. The rig model learning includes extracting a frame feature value by executing second transformation processing with respect to second input information including the extracted voice feature value and outputting character control information for controlling a character from the extracted frame feature value.


