Lip Movement Modification for Localized Accent Sync

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing lip-synced videos, especially after translation, lack synchronization between lip movements and audio, making them appear unnatural and harder to understand, particularly for viewers with different accents or cultural backgrounds.

Innovation Solution

A computer-implemented method that detects cultural context and accents in video streams, applies accent tags to translated text, and modifies lip movements to match the target language and accent, resulting in a translated video stream with synchronized and culturally accurate lip movements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If lip-syncing is applied to translated videos, then the video can be understood by viewers who do not speak the original language, but the lip movements do not match the translated audio, making the video appear unnatural and harder to understand

Engineering Contradiction:
Improvelanguage accessibilityVSAvoidlip-sync accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies different lip movement characteristics to different parts of the speech (phonemes) based on their acoustic properties. By analyzing the acoustic features of each phoneme in the translated audio and mapping them to corresponding lip movements, the system achieves localized accuracy in lip-syncing rather than applying a uniform transformation to the entire speech sequence.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system transforms lip movement parameters by analyzing acoustic features (frequency, duration, amplitude) of the translated speech and mapping them to corresponding visual lip movement parameters. This parameter transformation enables the lip movements to reflect the characteristics of the target language while maintaining synchronization with the translated audio.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If traditional lip-syncing methods are used, then translation to target language is achieved, but cultural context and accents are lost, reducing naturalness and understandability

Engineering Contradiction:
Improvelanguage translation capabilityVSAvoidcultural context and accent information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system incorporates feedback loops where the acoustic characteristics of the translated speech are continuously analyzed and used to adjust the lip movement generation. By comparing the acoustic features of the target language speech with the generated lip movements, the system iteratively refines the output to preserve accent and cultural characteristics.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces acoustic feature analysis as an intermediary between the translated audio and lip movement generation. This intermediary layer extracts and preserves cultural and accent information from the translated speech, which then guides the synthesis of culturally appropriate lip movements rather than directly mapping from source to target without intermediate analysis.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If lip movements are modified to match translated language, then synchronization improves, but the process becomes more complex requiring detection of cultural context and application of accent tags

Engineering Contradiction:
Improvelip-sync accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech into individual phonemes and processes each phoneme separately through the lip-syncing pipeline. This segmentation allows the system to apply specific transformations to each phoneme based on its acoustic properties, making the overall complex process more manageable and systematic.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis of the translated audio to extract acoustic features and determine accent characteristics before generating lip movements. By preparing the acoustic feature representation in advance, the system simplifies the subsequent lip movement synthesis process and avoids needing to handle full complexity during real-time generation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12278999B2Generation of video stream having localized lip-syncing with personalized characteristics
Publication Date: 2025.04.15 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12278999B2 patent drawing
  • US12278999B2 patent drawing
  • US12278999B2 patent drawing

AI summary

A computer-implemented method, in accordance with one embodiment, includes detecting cultural context and accents of speakers portrayed in a video stream and/or an audience of the video stream. Accent tags are selected for the speakers according to the cultural context and accents of the speakers and/or the audience of the video stream. A textual representation of spoken words of the speakers is translated from a source language to a target language. The accent tags are applied to the textual representation of the spoken words in the target language according to the speakers corresponding to the textual representation of the spoken words in the target language. Speech lip movements of the speakers portrayed in the video stream are modified to match the target language and the locale accent tags. A translated video stream having the speakers appearing to speak in the target language with the modified lip movements is output.