Deferred Neural Rendering for Actor Replacement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multimedia distribution services face challenges in efficiently customizing video content for regional markets, particularly in synchronizing lip movements and replacing actors, which requires significant computing resources and can result in artifacts and poor temporal cohesion.
Innovation Solution
A video synthesis system that generates a synthesized video by creating 3D models of actors' faces and bodies, merging facial and body parameters from original and dubbed content, and using deferred neural rendering to ensure accurate lip synchronization and replacement, reducing computing resources and artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If video content is re-shot with regional actors to customize content for particular regions, then content accessibility and regional relevance are improved, but computing resources and production costs increase significantly
Solution Approach 1:
The patent creates digital twins (3D models) of actors that can be copied and reused across multiple video customizations. Instead of re-shooting with different actors, the system generates a digital replica that can be instantiated multiple times with different parameters (clothing, accessories, background) to create region-specific content, dramatically reducing the need for repeated physical production
Solution Approach 2:
The system customizes video content by modifying parameters of the digital twin model including facial features, body shape, clothing, and background elements. By changing these parameters rather than re-filming entire scenes, the system achieves regional customization with minimal computational overhead compared to traditional re-shooting methods
2Ease of manufacture
If traditional video processing methods are used for actor replacement and lip synchronization, then implementation is straightforward, but artifacts and poor temporal cohesion occur
Solution Approach 1:
The patent replaces traditional frame-by-frame manual editing and mechanical lip-sync methods with an end-to-end neural network model. This neural approach automatically generates temporally coherent lip movements and facial expressions that match the dubbed audio, eliminating the artifacts and discontinuities inherent in traditional processing methods while maintaining implementation simplicity through automated processing
3Manufacturing precision
If significant computing resources are allocated for video re-shooting and processing, then content customization quality improves, but processing time and energy consumption increase
Solution Approach 1:
The system performs preliminary action by pre-processing video frames to extract and store intermediate representations (feature maps, depth information, normal maps) that capture essential visual content. These pre-extracted features are cached and reused during the neural rendering process, avoiding redundant computations and significantly reducing processing time while maintaining high customization quality
Data Source
AI summary
Techniques are disclosed for performing video synthesis of audiovisual content. In an example, a computing system may determine first parameters of a face and body of a source person from a first frame in a video shot. The system also determines second parameters of a face and body of a target person. The system determines that the target person is a replacement for the source person in the first frame. The system generates third parameters of the target person based on merging the first parameters with the second parameters. The system then performs deferred neural rendering of the target person based on a neural texture that corresponds to a texture space of the video shot. The system then outputs a second frame that shows the target person as the replacement for the source person.


