Lip-Sync Animation Model Learning via Acoustic Feature Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current lip-sync animation generation technologies lack stability and naturalness in their output, failing to effectively utilize voice data to create realistic character animations.

Innovation Solution

A model learning apparatus comprising a voice model and a rig model, where the voice model extracts acoustic feature values through signal processing and transformation, and the rig model extracts frame feature values to output character control information for controlling facial expressions and animations, using techniques like convolutional neural networks and transformer encoders for improved feature extraction and animation synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional lip-sync animation generation technology is used, then the process is simple, but the output lacks stability and naturalness

Engineering Contradiction:
Improvestability of animation outputVSAvoidcomplexity of model learning system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the animation generation process into distinct functional modules: a voice model that extracts acoustic feature values from voice data, and a rig model that generates animation based on these features. This segmentation allows each module to be optimized independently, improving overall reliability while managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces acoustic feature values as an intermediary representation between raw voice data and animation output. The voice model transforms voice data into acoustic feature values, which then serve as input to the rig model. This intermediary layer enables more stable and natural animation generation by decoupling the complexity of voice processing from animation synthesis.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If acoustic feature value extraction is implemented, then naturalness of animation improves, but processing time increases

Engineering Contradiction:
Improvequality of lip-sync animationVSAvoidprocessing time for feature extraction
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The voice model performs preliminary extraction of acoustic feature values from voice data before animation generation. By pre-processing the voice data into meaningful acoustic features (such as pitch, timbre, and temporal characteristics), the system prepares optimized input for the rig model, improving animation quality while enabling efficient processing through staged computation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces direct mechanical mapping from voice data to animation with a learned transformation system. Instead of using rigid rule-based approaches, the voice model learns to extract acoustic feature values that capture essential vocal characteristics, enabling more natural and efficient animation generation that better balances quality and processing time.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240078996A1Model learning system, model learning method, a non-transitory computer-readable recording medium, an animation generation system, and an animation generation method
Publication Date: 2024.03.07 SQUARE ENIX HLDG CO LTD
  • US20240078996A1 patent drawing
  • US20240078996A1 patent drawing
  • US20240078996A1 patent drawing

AI summary

Embodiments of the present disclosure provide methods, systems and non-transitory computer readable media of performing voice model learning and rig model learning. The voice model learning includes extracting an acoustic feature value by executing predetermined acoustic signal processing with respect to voice data including human voice and extracting a voice feature value by executing first transformation processing with respect to first input information including the extracted acoustic feature value. The rig model learning includes extracting a frame feature value by executing second transformation processing with respect to second input information including the extracted voice feature value and outputting character control information for controlling a character from the extracted frame feature value.