Real-Time Face Animation Synthesis in AR Messaging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current face animation synthesis techniques are either fast but non-photorealistic or time-consuming and not suitable for real-time processing on mobile devices, and existing messaging systems face challenges in efficiently processing and rendering augmented reality content on power and resource-constrained devices.
Innovation Solution
The proposed messaging system employs advanced image processing operations and neural networks to enable efficient face animation synthesis and augmented reality content generation, allowing for real-time processing and rendering of augmented reality content on mobile devices, including face animation and image manipulation, while reducing latency and power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If advanced face animation synthesis techniques are used to achieve photorealistic results, then the quality of face animation is improved, but the processing time increases making it unsuitable for real-time mobile processing
Solution Approach 1:
The system segments the face animation synthesis process into distinct components: face detection, landmark identification, expression parameter extraction, and animation rendering. This segmentation allows each component to be optimized independently, enabling photorealistic quality while reducing overall processing time for real-time mobile execution
Solution Approach 2:
The system performs preliminary actions by pre-processing and storing face models, expression parameters, and animation templates in advance. These pre-computed resources are then rapidly applied during real-time face animation synthesis, maintaining high quality output while minimizing processing time on mobile devices
2Manufacturing precision
If complex image processing operations are performed to achieve photorealistic face animation, then the quality of augmented reality content is improved, but the power consumption increases on mobile devices
Solution Approach 1:
The system applies partial action by selectively processing only the necessary facial regions and parameters required for photorealistic animation, rather than processing the entire image. This targeted approach maintains high quality output while significantly reducing computational load and power consumption on mobile devices
Solution Approach 2:
The system replaces traditional computationally intensive mechanical image processing methods with optimized algorithms and neural network-based approaches. This substitution maintains photorealistic quality while reducing the computational overhead and power consumption associated with complex image processing operations
3Speed
If real-time face animation synthesis is implemented on mobile devices, then the processing speed is improved, but the device complexity increases
Solution Approach 1:
The system implements a universal face animation framework that can operate across different mobile devices with varying capabilities. The unified architecture handles face detection, landmark identification, and animation synthesis through standardized processes, enabling real-time performance without requiring device-specific complex implementations
Solution Approach 2:
The system introduces intermediary components such as pre-trained neural networks and optimization layers that mediate between the input face image and the final animated output. These intermediaries streamline the processing pipeline, achieving real-time speeds while keeping the overall system architecture manageable and not excessively complex
Data Source
AI summary
The subject technology receives a selection of a selectable graphical item to initiate generating augment reality content including facial synthesis, the selection being received by a third party application, the third party application being executed by a computing device separate from a first party application and a messaging server system. The subject technology captures image data by the client device. The subject technology generates, by the one or more hardware processors and based at least in part on frames of a source media content, sets of source pose parameters. The subject technology generates, based at least in part on sets of the source pose parameters, an output media content using an interface communicating with the messaging server system. The subject technology provides augmented reality content based at least in part on the output media content for display on the computing device.


