Audio-Driven Lip Synchronization Using Facial Landmark Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to achieve accurate and realistic lip synchronization in audiovisual environments, particularly across different languages and speakers, often resulting in distorted facial movements and artifacts.

Innovation Solution

A device for synchronizing features of digital objects with predefined audio contents, utilizing a processing unit with a control unit that includes encoders, a generator, and discriminators to extract and align audio and visual features, ensuring accurate and high-quality lip movements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional deep learning technologies are used for audio-visual synchronization, then the resolution of synchronized images is improved, but the original quality of facial shapes is distorted and image quality deteriorates

Engineering Contradiction:
Improvesynchronization resolutionVSAvoidfacial shape quality
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent segments the facial region into multiple landmarks and uses selective area processing. Instead of processing the entire face image uniformly, it identifies specific facial landmarks (eyes, nose, mouth, chin) and applies synchronization transformations only to these critical regions, preserving the original quality of non-critical facial areas while achieving accurate lip synchronization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities to different regions of the face. Critical lip regions receive high-precision synchronization transformation based on audio input, while other facial regions maintain their original appearance quality. This local differentiation allows the system to improve synchronization accuracy without compromising overall facial shape quality.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If speaker-specific neural voice puppetry techniques are used, then automated re-dubbing quality is improved, but the system cannot scale to general cases and requires additional pre-processing data per identity

Engineering Contradiction:
Improvere-dubbing qualityVSAvoidspeaker independence
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal synchronization system that works across different speakers and languages without requiring speaker-specific training data. The system uses a general audio-driven transformation model that can adapt to any speaker's characteristics in real-time, eliminating the need for extensive pre-processing for each identity while maintaining high re-dubbing quality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the approach from learning speaker-specific parameters to learning speaker-invariant parameters. Instead of training models on specific speakers, the system learns the fundamental relationship between audio and facial movements that is common across all speakers, allowing it to generate accurate synchronizations for any new speaker without additional training data.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If GAN-based speaker independent solutions are used, then speaker independence is achieved, but artifacts are generated and quality is compromised at higher resolutions

Engineering Contradiction:
Improvespeaker independenceVSAvoidgeneration quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent replaces the GAN-based generative approach with a direct audio-driven transformation approach. Instead of using GANs to generate facial images from audio (which introduces artifacts), the system directly transforms existing facial images based on audio-driven landmark movements, eliminating the source of artifacts while maintaining speaker independence.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent uses copying and transformation of existing facial features rather than generating new features. It extracts facial landmarks from the input image and copies their spatial relationships, then transforms these landmark positions based on audio input to create synchronized facial movements, avoiding the artifact generation problem of synthetic GAN-based approaches.

Inventive Principle:
Principle #26Copying

4Device complexity

If lip poses are generated directly from audio signals, then the process is simplified, but temporal effects such as co-articulation are ignored and actual facial dynamics are not addressed

Engineering Contradiction:
Improveprocessing complexityVSAvoidsynchronization accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent performs preliminary extraction of facial landmark positions and temporal segmentation of audio signals before generating lip poses. It pre-processes both visual and audio data to identify phonetic segments and corresponding facial regions, then uses this pre-extracted information to generate accurate lip movements that respect temporal dynamics and co-articulation effects.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent incorporates dynamic temporal modeling to capture the time-varying nature of speech and facial movements. Instead of generating static lip poses, the system models the dynamic evolution of facial landmarks throughout the audio signal, accounting for co-articulation and other temporal effects that are crucial for accurate lip synchronization.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250173935A1Device for synchronization of features of digital objects with audio contents
Publication Date: 2025.05.29 NEURALGARAGE PTE LTD
  • US20250173935A1 patent drawing
  • US20250173935A1 patent drawing
  • US20250173935A1 patent drawing

AI summary

Disclosed is a device for synchronization of features of digital objects with audio contents (100). The device of the present invention synchronizes the features of digital objects and audio contents of the audio-visual environment. The device (100) includes a processing unit (105), an input unit (110), and an output unit (115). The device (100) is removably connectable to a host device (120) and a power supply unit (125). The processing unit (105) is defined by a microcontroller and it is configured with various modules that are responsible for synchronization of features of digital objects with audio contents. The device of the present invention advantageously transforms the feature of the digital object in the input video to be in synchronization with audio content irrespective of the identity of object and language of audio.