Lip Syncing Foreign Language Video Streams

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current language translation methods fail to effectively preserve the emotional and tonal characteristics of spoken words when converting video/audio streams from one language to another, particularly in sermons, where these characteristics are critical for meaning.

Innovation Solution

A method that separates source language video/audio streams into independent streams, transcribes and translates the audio, and synchronizes the translated audio with pre-generated viseme and phoneme morphing targets for lip syncing, ensuring the emotional and tonal characteristics are maintained in the target language, using commercially available software like 3ds Max for rendering and compositing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional translation methods are used to convert video/audio streams, then language translation is achieved, but emotional and tonal characteristics are lost

Engineering Contradiction:
Improveemotional and tonal characteristicsVSAvoidtranslation accuracy
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The video stream is separated into independent video and audio streams, allowing separate processing of visual and auditory information. The audio stream is further segmented into speech content and emotional/tonal characteristics, enabling selective preservation of emotional elements while translating the linguistic content.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A digitally rendered mouth serves as an intermediary element that bridges the source language audio and target language audio. The rendered mouth visually represents the emotional and tonal characteristics of the original speaker while synchronizing with the translated speech, allowing emotional information to be transmitted through visual cues when the translated audio lacks corresponding emotional qualities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If audio is translated to target language, then language conversion is achieved, but lip synchronization is lost

Engineering Contradiction:
Improvelanguage conversion capabilityVSAvoidlip sync accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

Instead of attempting to modify the original speaker's mouth movements to match translated speech, the system creates a digital copy/rendering of the mouth that can be independently animated. This rendered mouth copy is then synchronized with the target language audio, achieving lip synchronization without compromising the accuracy of the language translation.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the parameter of mouth representation from a static or naturally captured element to a dynamically renderable digital model. By parameterizing the mouth movements and enabling independent animation of the rendered mouth, the system can synchronize visual lip movements with translated audio while preserving the original speaker's visual appearance.

Inventive Principle:
Principle #35Parameter changes

3Speed

If real-time processing is implemented for live video streams, then immediate translation is achieved, but processing complexity increases

Engineering Contradiction:
Improvetranslation speedVSAvoidprocessing system complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-separating video and audio streams and preparing the audio for translation before the actual translation occurs. For live streams, this means establishing the processing pipeline and rendering engine in advance, allowing the system to process incoming audio frames and generate corresponding rendered mouth animations with minimal latency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces traditional mechanical or manual translation processes with automated digital signal processing and computer-generated rendering. By substituting manual lip-sync adjustment and emotional tone analysis with automated algorithms and rendered graphics, the system reduces processing complexity while maintaining real-time performance capabilities.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10657972B2Method of translating and synthesizing a foreign language
Publication Date: 2020.05.19 MTEC GRP INC
  • US10657972B2 patent drawing
  • US10657972B2 patent drawing
  • US10657972B2 patent drawing

AI summary

A method to interactively convert a source language video/audio stream into one or more target languages in high definition video format using a computer. The spoken words in the converted language are synchronized with synthesized movements of a rendered mouth. Original audio and video streams from pre-recorded or live sermons are synthesized into another language with the original emotional and tonal characteristics. The original sermon could be in any language and be translated into any other language. The mouth and jaw are digitally rendered with viseme and phoneme morphing targets that are pre-generated for lip synching with the synthesized target language audio. Each video image frame has the simulated lips and jaw inserted over the original. The new audio and video image then encoded and uploaded for internee viewing or recording to a storage medium.