Facial Animation Generation Using Pinyin Sequence Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating animation of Chinese mouth shapes based on audio inputs face challenges due to variations in speaking style and face shape in training data, making it difficult to accurately map audio to facial expressions, especially when the audio is not present in the training set or differs significantly from trained data.
Innovation Solution
A method that processes input material to generate a normalized text, converts it into a Chinese pinyin sequence, and uses this sequence along with a reference audio to obtain animation of facial expressions, leveraging pre-annotated dictionaries and pinyin-expression coefficient dictionaries to stitch face image elements and generate reliable facial animations without requiring extensive annotation of audio and facial expressions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If deep learning is used to directly learn the audio-facial expression coefficient mapping relationship, then the animation generation can be achieved, but the model cannot accurately handle audio not present in training set or differing significantly from trained data due to variations in speaking style and face shape
Solution Approach 1:
The patent introduces an intermediary Pinyin sequence as a bridge between audio and facial expressions. Instead of directly mapping audio to facial expressions, the system converts audio to Pinyin sequence first, then maps Pinyin to facial expression coefficients. This intermediary layer generalizes the mapping relationship, enabling accurate animation generation for unseen audio content by leveraging the universal Pinyin representation rather than direct audio-specific patterns.
2Reliability
If all sentences and audio variables are included in training data to ensure comprehensive coverage, then the model accuracy improves, but the amount of training data required becomes excessively large and impossible to collect
Solution Approach 1:
The patent extracts the essential phonetic information from audio by converting it to a Pinyin sequence. This extraction process removes unnecessary audio variations (speaking style, tone, amplitude) while retaining the core phonetic content that determines mouth shapes. By working with the extracted Pinyin representation, the system achieves comprehensive coverage of all possible sentences with a manageable training set, as only the phonetic patterns need to be learned rather than all audio variations.
Solution Approach 2:
The Pinyin sequence serves as a universal representation that can map to multiple audio variations and speaking styles. A single Pinyin-to-facial-expression mapping can handle different tones, speeds, and emphases of the same phonetic content. This universality allows the model to generalize across diverse audio data, reducing the need for extensive training data while maintaining comprehensive coverage of all possible sentences.
3Reliability
If extensive annotation of audio and facial expressions is performed to improve mapping accuracy, then the generation reliability improves, but the development cost and time required increases significantly
Solution Approach 1:
The system performs self-service by automatically converting audio to Pinyin sequences using standard phonetic conversion algorithms, eliminating the need for manual annotation of audio-to-facial-expression mappings. The Pinyin-to-facial-expression mapping can be trained using automatically extracted Pinyin data rather than requiring manual annotation of corresponding facial expressions. This automation significantly reduces the complexity and cost of data preparation while maintaining mapping accuracy.
Data Source
AI summary
The present disclosure discloses a method and apparatus for generating animation. An implementation of the method may include: processing a to-be-processed material to generate a normalized text; analyzing the normalized text to generate a Chinese pinyin sequence of the normalized text; generating a reference audio based on the to-be-processed material; and obtaining a animation of facial expressions corresponding to the timing sequence of the reference audio based on the Chinese pinyin sequence and the reference audio.


