Facial Animation Generation Using Pinyin Sequence Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating animation of Chinese mouth shapes based on audio inputs face challenges due to variations in speaking style and face shape in training data, making it difficult to accurately map audio to facial expressions, especially when the audio is not present in the training set or differs significantly from trained data.

Innovation Solution

A method that processes input material to generate a normalized text, converts it into a Chinese pinyin sequence, and uses this sequence along with a reference audio to obtain animation of facial expressions, leveraging pre-annotated dictionaries and pinyin-expression coefficient dictionaries to stitch face image elements and generate reliable facial animations without requiring extensive annotation of audio and facial expressions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If deep learning is used to directly learn the audio-facial expression coefficient mapping relationship, then the animation generation can be achieved, but the model cannot accurately handle audio not present in training set or differing significantly from trained data due to variations in speaking style and face shape

Engineering Contradiction:
Improveadaptability to unseen audioVSAvoidaccuracy of facial expression mapping
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an intermediary Pinyin sequence as a bridge between audio and facial expressions. Instead of directly mapping audio to facial expressions, the system converts audio to Pinyin sequence first, then maps Pinyin to facial expression coefficients. This intermediary layer generalizes the mapping relationship, enabling accurate animation generation for unseen audio content by leveraging the universal Pinyin representation rather than direct audio-specific patterns.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If all sentences and audio variables are included in training data to ensure comprehensive coverage, then the model accuracy improves, but the amount of training data required becomes excessively large and impossible to collect

Engineering Contradiction:
Improvecompleteness of audio-facial mappingVSAvoidvolume of training data
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts the essential phonetic information from audio by converting it to a Pinyin sequence. This extraction process removes unnecessary audio variations (speaking style, tone, amplitude) while retaining the core phonetic content that determines mouth shapes. By working with the extracted Pinyin representation, the system achieves comprehensive coverage of all possible sentences with a manageable training set, as only the phonetic patterns need to be learned rather than all audio variations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The Pinyin sequence serves as a universal representation that can map to multiple audio variations and speaking styles. A single Pinyin-to-facial-expression mapping can handle different tones, speeds, and emphases of the same phonetic content. This universality allows the model to generalize across diverse audio data, reducing the need for extensive training data while maintaining comprehensive coverage of all possible sentences.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If extensive annotation of audio and facial expressions is performed to improve mapping accuracy, then the generation reliability improves, but the development cost and time required increases significantly

Engineering Contradiction:
Improveaccuracy of audio-facial mappingVSAvoidcomplexity of data annotation process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically converting audio to Pinyin sequences using standard phonetic conversion algorithms, eliminating the need for manual annotation of audio-to-facial-expression mappings. The Pinyin-to-facial-expression mapping can be trained using automatically extracted Pinyin data rather than requiring manual annotation of corresponding facial expressions. This automation significantly reduces the complexity and cost of data preparation while maintaining mapping accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11948236B2Method and apparatus for generating animation, electronic device, and computer readable medium
Publication Date: 2024.04.02 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11948236B2 patent drawing
  • US11948236B2 patent drawing
  • US11948236B2 patent drawing

AI summary

The present disclosure discloses a method and apparatus for generating animation. An implementation of the method may include: processing a to-be-processed material to generate a normalized text; analyzing the normalized text to generate a Chinese pinyin sequence of the normalized text; generating a reference audio based on the to-be-processed material; and obtaining a animation of facial expressions corresponding to the timing sequence of the reference audio based on the Chinese pinyin sequence and the reference audio.