Voice-Driven Lip Animation Using Preconfigured Pronunciation Model Library

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video-based animation technologies for voice-driven lip animation have complex algorithms and high calculation costs, making them inefficient for changing lip shapes effectively.

Innovation Solution

A method and apparatus that obtain audio signals to calculate the motion extent proportion of lip shapes, generate a motion extent value, and use a preconfigured lip pronunciation model library to simplify the algorithm and reduce costs, enabling efficient lip shape changes and animations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If a Machine Learning mode is used to map audio signals to lip shape parameters, then the lip animation quality is improved, but the algorithm complexity and calculation cost increase

Engineering Contradiction:
Improvelip animation qualityVSAvoidalgorithm complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent uses a pre-configured lip pronunciation model library that stores predetermined lip shape parameters for different phonemes. Instead of using complex Machine Learning to generate lip shapes in real-time, the system copies pre-computed lip shape data from the library based on audio signal analysis, significantly reducing algorithm complexity while maintaining animation quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The lip pronunciation model library is pre-configured with lip shape parameters for various phonemes before runtime. This preliminary preparation allows the system to simply retrieve and apply appropriate lip shapes during animation generation, avoiding the need for complex real-time calculations

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If a Machine Learning mode is used to map audio signals to lip shape parameters, then the lip animation quality is improved, but the calculation cost increases

Engineering Contradiction:
Improvelip animation qualityVSAvoidcalculation cost
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The system retrieves pre-computed lip shape parameters from the lip pronunciation model library rather than performing expensive Machine Learning calculations during runtime. This copying approach dramatically reduces calculation cost and energy consumption while preserving animation quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

Lip shape parameters are pre-computed and stored in the model library during an offline phase. During actual animation generation, the system only needs to retrieve these pre-computed values, significantly reducing online calculation costs and energy usage

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If a complex Machine Learning algorithm is used for lip shape mapping, then the animation accuracy is improved, but the processing time increases

Engineering Contradiction:
Improveanimation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system copies pre-determined lip shape parameters from the model library based on phoneme recognition, eliminating the need for time-consuming Machine Learning inference while maintaining accurate lip synchronization with the audio input

Inventive Principle:
Principle #26Copying

Data Source

PatentUS8350859B2Method and apparatus for changing lip shape and obtaining lip animation in voice-driven animation
Publication Date: 2013.01.08 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US8350859B2 patent drawing
  • US8350859B2 patent drawing
  • US8350859B2 patent drawing

AI summary

The present invention discloses a method and apparatus for changing lip shape and obtaining a lip animation in a voice-driven animation, and relate to computer technologies. The method for changing lip shape includes: obtaining audio signals and obtaining motion extent proportion of lip shape according to characteristics of the audio signals; obtaining an original lip shape model inputted by a user and generating a motion extent value of the lip shape according to the original lip shape model and the obtained motion extent proportion of the lip shape; generating a lip shape grid model set according to the obtained motion extent value of the lip shape and a preconfigured lip pronunciation model library. The method for changing lip shape in a voice-driven animation includes an obtaining module, a first generating module and a second generating module. The solutions provided by the present invention have a simple algorithm and low cost.