3D Face Alignment and Neural Rendering for Photorealistic Dubbing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional dubbing techniques result in noticeable disparities between facial movements and audio in different languages due to time-consuming and non-photorealistic rendering, often requiring complex facial capture systems.

Innovation Solution

A method involving 3D tracking and neural rendering using machine learning models to align and render photorealistic faces, allowing for efficient dubbing without the need for facial capture systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If conventional graphics rendering engines are used to render facial images, then the rendering process can be performed without facial capture systems, but the rendering time is considerable and the images are not photorealistic

Engineering Contradiction:
Improveelimination of facial capture system requirementVSAvoidrendering time
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The patent replaces conventional graphics rendering engines with a neural rendering model based on machine learning. This substitution transforms the rendering process from a computational graphics approach to an AI-driven synthesis approach, achieving photorealistic results without requiring facial capture systems and reducing rendering time significantly.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the rendering process by using trained machine learning models that have learned photorealistic facial representations during training. The model takes as input 3D geometry, texture maps, and lighting maps, and outputs photorealistic facial images directly, bypassing the time-consuming conventional rendering pipeline.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If conventional graphics rendering engines are used to render facial images, then dubbing can be performed without facial capture systems, but the rendered faces do not look photorealistic and resemble video game characters

Engineering Contradiction:
Improveelimination of facial capture system requirementVSAvoidphotorealism quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent replaces conventional graphics rendering engines with a neural rendering model based on machine learning. This substitution transforms the rendering process from a computational graphics approach to an AI-driven synthesis approach, achieving photorealistic results without requiring facial capture systems and reducing rendering time significantly.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the rendering process by using trained machine learning models that have learned photorealistic facial representations during training. The model takes as input 3D geometry, texture maps, and lighting maps, and outputs photorealistic facial images directly, bypassing the time-consuming conventional rendering pipeline.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If careful word selection is used to match facial movements with dubbed audio, then language disparities can be reduced, but noticeable disparities between facial movements and audio remain

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidnaturalness of dubbed content
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The patent performs preliminary action by training the neural rendering model on large datasets of facial expressions and corresponding audio before actual dubbing. This pre-training enables the model to automatically learn the complex mappings between audio signals and facial movements, eliminating the need for manual word selection and achieving natural synchronization without noticeable disparities.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250209759A1Techniques for generating dubbed media content items
Publication Date: 2025.06.26 NETFLIX INC
  • US20250209759A1 patent drawing
  • US20250209759A1 patent drawing
  • US20250209759A1 patent drawing

AI summary

In various embodiments, a dubbing application performs three-dimensional (3D) tracking of (1) the face of an actor within video frames of a first media content item to generate 3D geometry representing the face of the actor, and (2) the face of a dubber within video frames of a second media content item to generate 3D geometry representing the face of the dubber. The dubbing application also tracks the texture and lighting of the face of the actor in the first media content item. The dubbing application aligns the 3D geometry of the face of the dubber with the 3D geometry of the face of the actor. Then, the dubbing application performs neural rendering to generate dubbed video frames using a trained machine learning model, the aligned 3D geometry of the dubber, the texture and lighting of the face of the actor, and the video frames of the first media content.