Synthetic Lip Synchronization Texture Warping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for generating synthetic lip synchronization for videos in different languages face challenges such as the need for large source libraries, high computational costs, and the inability to maintain faithful surface textures, often resulting in unnatural and unappealing video content.

Innovation Solution

The method involves generating source and target face models with corresponding facial keypoints, applying two-dimensional texture warping to maintain surface textures, and blending matching mouth interiors to create synthetic lip synchronization that matches the target audio track while preserving the original video's surface features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional methods are used to generate synthetic lip synchronization, then lip synchronization can be achieved, but surface textures become distorted and image quality deteriorates

Engineering Contradiction:
Improvesurface texture fidelityVSAvoidimage quality
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The method segments the face model into multiple triangular regions with vertices at facial keypoints. By applying texture warping independently to each triangle based on keypoint displacements, the system maintains local texture fidelity while achieving global lip synchronization. This segmentation approach prevents distortion that would occur with global transformation methods.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by using two-dimensional texture warping that preserves surface features in specific regions. The texture coordinates of each triangle are transformed based on the displacement between source and target keypoints, ensuring that local surface textures (such as skin patterns, wrinkles, and pores) maintain their original appearance while adapting to new mouth positions.

Inventive Principle:
Principle #3Local quality

2Productivity

If three-dimensional transformation methods are used, then lip synchronization is achieved, but computational complexity and processing time increase significantly

Engineering Contradiction:
Improveprocessing speedVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The method extracts only the essential two-dimensional texture transformation components needed for lip synchronization, discarding complex three-dimensional transformation operations. By focusing on 2D texture coordinate warping based on keypoint displacements, the system achieves the desired lip synchronization effect with significantly reduced computational complexity and faster processing speeds.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If conventional texture mapping is applied, then lip movements match target audio, but original surface features and textures are lost

Engineering Contradiction:
Improvelip synchronization accuracyVSAvoidsurface texture preservation
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent transitions from three-dimensional model transformation to two-dimensional texture coordinate transformation. By working in the 2D texture space rather than 3D model space, the system can warp texture coordinates to match target lip positions while preserving the original surface texture information. This dimensional change allows independent control of position and texture appearance.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12198241B1Systems and methods to generate synthetic lip synchronization having faithful textures
Publication Date: 2025.01.14 AMAZON TECH INC
  • US12198241B1 patent drawing
  • US12198241B1 patent drawing
  • US12198241B1 patent drawing

AI summary

Systems and methods to generate synthetic lip synchronization may generate source facial keypoints based on a source video, generate target facial keypoints based on a target audio, determine distances between the source facial keypoints and target facial keypoints, and transform or warp the source facial keypoints and associated surfaces to the target facial keypoints. In this manner, target video having synthetic lip synchronization that matches the target audio may be generated, and the target video may substantially preserve or maintain surface textures or features from the source video in the target video, thereby generating natural and believable synthetic lip synchronization corresponding to the target audio.