Deferred Neural Rendering for Facial Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multimedia distribution services face challenges in efficiently customizing video content for regional audiences, particularly in synchronizing lip movements and replacing actors with regional talent, which requires significant computing resources and often results in artifacts and jitter.

Innovation Solution

A video synthesis system that generates a synthesized video by creating 3D models of actors' faces and bodies, synchronizing lip movements with dubbed audio, and replacing actors using neural textures and deferred neural rendering, while maintaining temporal cohesion and minimizing artifacts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If re-shooting video to include new actors and/or audio tracks, then content accessibility to particular regions is improved, but computing resources required increase significantly

Engineering Contradiction:
Improvecontent accessibilityVSAvoidcomputing resources
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent creates a digital copy of the original video content and modifies it through neural rendering techniques. Instead of re-shooting the entire video with new actors and audio tracks, the system generates a synthesized version by rendering new facial textures and lip movements onto the original video frames, thereby reducing computing resource requirements while maintaining content accessibility for regional audiences

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of physical re-shooting with a digital neural rendering system. By using neural networks to generate and synthesize facial expressions, lip movements, and textures directly from audio tracks and original video frames, the system eliminates the need for actual re-filming operations, significantly reducing computing resources while achieving the same adaptability for regional content distribution

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If traditional video synthesis methods are used, then customization for regional distribution is achieved, but artifacts and jitter are introduced

Engineering Contradiction:
Improvecustomization capabilityVSAvoidvideo quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent replaces traditional mechanical video editing and compositing methods with neural rendering technology. By using neural networks to directly generate and synthesize facial textures, expressions, and lip movements at the pixel level, the system achieves high-quality customization without the artifacts and jitter characteristic of traditional frame-by-frame editing approaches

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of video synthesis by operating at the neural network level rather than the pixel or frame level. By training and rendering neural representations of faces and expressions, the system can smoothly transition between different actors, languages, and regional variations without introducing temporal discontinuities or visual artifacts, thereby improving manufacturing precision while maintaining customization capability

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11581020B1Facial synchronization utilizing deferred neural rendering
Publication Date: 2023.02.14 AMAZON TECH INC
  • US11581020B1 patent drawing
  • US11581020B1 patent drawing
  • US11581020B1 patent drawing

AI summary

Techniques are disclosed for performing video synthesis of audiovisual content. In an example, a computing system may determine first facial parameters of a face of a particular person from a first frame in a video shot, whereby the video shot shows the particular person speaking a message. The system may determine second facial parameters based on an audio file that corresponds to the message being spoken in a different way from the video shot. The system may generate third facial parameters by merging the first and the second facial parameters. The system may identify a region of the face that is associated with a difference between the first and second facial parameters, render the region of the face based on a neural texture of the video shot, and then output a new frame showing the face of the particular person speaking the message in the different way.