3D Face Alignment and Neural Rendering for Photorealistic Dubbing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional dubbing techniques result in noticeable disparities between facial movements and audio in different languages due to time-consuming and non-photorealistic rendering, often requiring complex facial capture systems.
Innovation Solution
A method involving 3D tracking and neural rendering using machine learning models to align and render photorealistic faces, allowing for efficient dubbing without the need for facial capture systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional graphics rendering engines are used to render facial images, then the rendering process can be performed without facial capture systems, but the rendering time is considerable and the images are not photorealistic
Solution Approach 1:
The patent replaces conventional graphics rendering engines with a neural rendering model based on machine learning. This substitution transforms the rendering process from a computational graphics approach to an AI-driven synthesis approach, achieving photorealistic results without requiring facial capture systems and reducing rendering time significantly.
Solution Approach 2:
The patent changes the fundamental parameters of the rendering process by using trained machine learning models that have learned photorealistic facial representations during training. The model takes as input 3D geometry, texture maps, and lighting maps, and outputs photorealistic facial images directly, bypassing the time-consuming conventional rendering pipeline.
2Ease of manufacture
If conventional graphics rendering engines are used to render facial images, then dubbing can be performed without facial capture systems, but the rendered faces do not look photorealistic and resemble video game characters
Solution Approach 1:
The patent replaces conventional graphics rendering engines with a neural rendering model based on machine learning. This substitution transforms the rendering process from a computational graphics approach to an AI-driven synthesis approach, achieving photorealistic results without requiring facial capture systems and reducing rendering time significantly.
Solution Approach 2:
The patent changes the fundamental parameters of the rendering process by using trained machine learning models that have learned photorealistic facial representations during training. The model takes as input 3D geometry, texture maps, and lighting maps, and outputs photorealistic facial images directly, bypassing the time-consuming conventional rendering pipeline.
3Manufacturing precision
If careful word selection is used to match facial movements with dubbed audio, then language disparities can be reduced, but noticeable disparities between facial movements and audio remain
Solution Approach 1:
The patent performs preliminary action by training the neural rendering model on large datasets of facial expressions and corresponding audio before actual dubbing. This pre-training enables the model to automatically learn the complex mappings between audio signals and facial movements, eliminating the need for manual word selection and achieving natural synchronization without noticeable disparities.
Data Source
AI summary
In various embodiments, a dubbing application performs three-dimensional (3D) tracking of (1) the face of an actor within video frames of a first media content item to generate 3D geometry representing the face of the actor, and (2) the face of a dubber within video frames of a second media content item to generate 3D geometry representing the face of the dubber. The dubbing application also tracks the texture and lighting of the face of the actor in the first media content item. The dubbing application aligns the 3D geometry of the face of the dubber with the 3D geometry of the face of the actor. Then, the dubbing application performs neural rendering to generate dubbed video frames using a trained machine learning model, the aligned 3D geometry of the dubber, the texture and lighting of the face of the actor, and the video frames of the first media content.


