2D Human Speaker Animation via MPEG4 Key Point Trajectories
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current animation techniques, particularly in computer animation, face challenges in generating highly accurate and efficient facial animations of human speakers, as they are often prohibitively expensive and difficult to dynamically adjust movement and sound speeds accurately.
Innovation Solution
A method and system that utilize key points defined by the MPEG4 standard to generate a table of trajectories for 2D animation, mapping lip and teeth movements over time, which are then used to create a frame list that can render an animated image in real-time, adjustable to speech rates, allowing for accurate and dynamic facial animation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If highly accurate facial animation is created using 3D methods, then animation accuracy is improved, but cost and computational resources become prohibitively expensive
Solution Approach 1:
The patent uses 2D images as simplified copies of 3D facial structures, mapping key points from these 2D representations to create animation trajectories. This approach captures essential facial movement characteristics without requiring full 3D models, significantly reducing computational complexity while maintaining visual accuracy for the intended application.
Solution Approach 2:
The patent employs lightweight 2D image-based representations instead of resource-intensive 3D models. These 2D images serve as disposable, computationally inexpensive substitutes that can be processed efficiently on mobile devices, sacrificing some geometric fidelity for substantial gains in performance and accessibility.
2Adaptability or versatility
If traditional animation techniques are used, then animation can be created, but dynamic adjustment of movement and sound speeds is difficult
Solution Approach 1:
The patent creates dynamically adjustable animation by generating trajectories from actual video recordings of speech. These trajectories can be scaled and adjusted in real-time to match varying speech rates, allowing the animation system to adapt dynamically to different speaking speeds without requiring complex manual re-animation for each scenario.
Solution Approach 2:
The system allows flexible adjustment of animation parameters including playback speed and timing by modifying the trajectory execution parameters. The key point trajectories can be replayed at different speeds while maintaining the natural correspondence between lip movements and speech content, enabling versatile adaptation to various speech rates.
3Manufacturing precision
If detailed facial animation is generated, then animation accuracy is improved, but processing time and power consumption increase
Solution Approach 1:
The patent extracts only the essential facial movement information needed for accurate speech animation by identifying and tracking key points on facial features. This selective extraction approach captures the critical motion data required for lip-sync accuracy while discarding unnecessary computational overhead, enabling efficient processing on power-constrained mobile devices.
Solution Approach 2:
The facial animation is segmented into discrete key points that can be independently tracked and animated. This segmentation allows the system to process only the necessary facial regions and movements, reducing overall computational load and power consumption while maintaining accuracy in the animated speech output.
Data Source
AI summary
Methods and apparatus for representing speech in an animated image. In one embodiment, key point on the object to be animated are defined, and a table of trajectories is generated to map positions of the key points over time as the object performs defined actions accompanied by corresponding sounds. In another embodiment, the table of trajectories and a sound rate of the video are used to generate a frame list that includes information to render an animated image of the object in real time at a rate determined by the sound rate. In still another embodiment, a 2D animation of a human speaker is produced. Key points are selected from the Motion Picture Expert Group 4 (MPEG4) defined points for human lips and teeth.


