2D Human Speaker Animation via MPEG4 Key Point Trajectories

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current animation techniques, particularly in computer animation, face challenges in generating highly accurate and efficient facial animations of human speakers, as they are often prohibitively expensive and difficult to dynamically adjust movement and sound speeds accurately.

Innovation Solution

A method and system that utilize key points defined by the MPEG4 standard to generate a table of trajectories for 2D animation, mapping lip and teeth movements over time, which are then used to create a frame list that can render an animated image in real-time, adjustable to speech rates, allowing for accurate and dynamic facial animation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If highly accurate facial animation is created using 3D methods, then animation accuracy is improved, but cost and computational resources become prohibitively expensive

Engineering Contradiction:
Improvefacial animation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent uses 2D images as simplified copies of 3D facial structures, mapping key points from these 2D representations to create animation trajectories. This approach captures essential facial movement characteristics without requiring full 3D models, significantly reducing computational complexity while maintaining visual accuracy for the intended application.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent employs lightweight 2D image-based representations instead of resource-intensive 3D models. These 2D images serve as disposable, computationally inexpensive substitutes that can be processed efficiently on mobile devices, sacrificing some geometric fidelity for substantial gains in performance and accessibility.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Adaptability or versatility

If traditional animation techniques are used, then animation can be created, but dynamic adjustment of movement and sound speeds is difficult

Engineering Contradiction:
Improvedynamic speed adjustmentVSAvoidanimation system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates dynamically adjustable animation by generating trajectories from actual video recordings of speech. These trajectories can be scaled and adjusted in real-time to match varying speech rates, allowing the animation system to adapt dynamically to different speaking speeds without requiring complex manual re-animation for each scenario.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system allows flexible adjustment of animation parameters including playback speed and timing by modifying the trajectory execution parameters. The key point trajectories can be replayed at different speeds while maintaining the natural correspondence between lip movements and speech content, enabling versatile adaptation to various speech rates.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If detailed facial animation is generated, then animation accuracy is improved, but processing time and power consumption increase

Engineering Contradiction:
Improveanimation detail accuracyVSAvoidpower consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential facial movement information needed for accurate speech animation by identifying and tracking key points on facial features. This selective extraction approach captures the critical motion data required for lip-sync accuracy while discarding unnecessary computational overhead, enabling efficient processing on power-constrained mobile devices.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The facial animation is segmented into discrete key points that can be independently tracked and animated. This segmentation allows the system to process only the necessary facial regions and movements, reducing overall computational load and power consumption while maintaining accuracy in the animated speech output.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS7388586B2Method and apparatus for animation of a human speaker
Publication Date: 2008.06.17 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • US7388586B2 patent drawing
  • US7388586B2 patent drawing
  • US7388586B2 patent drawing

AI summary

Methods and apparatus for representing speech in an animated image. In one embodiment, key point on the object to be animated are defined, and a table of trajectories is generated to map positions of the key points over time as the object performs defined actions accompanied by corresponding sounds. In another embodiment, the table of trajectories and a sound rate of the video are used to generate a frame list that includes information to render an animated image of the object in real time at a rate determined by the sound rate. In still another embodiment, a 2D animation of a human speaker is produced. Key points are selected from the Motion Picture Expert Group 4 (MPEG4) defined points for human lips and teeth.