Multimedia Reanimation for Dubbed Lip-Sync Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multimedia content, such as videos, often experiences visual inconsistencies when dubbed into different languages due to mismatched mouth movements of characters with the dubbed vocal content, leading to inconsistent playback durations and synchronization issues.

Innovation Solution

A method and apparatus for reanimating multimedia content by accessing primary visual and audio content, identifying and modifying visual frames to match the playback duration of dubbed vocal content, and reanimating character mouth movements using computer-generated imagery (CGI) to synchronize with dubbed vocal sounds, ensuring consistent and synchronized visual and audio playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional dubbing methods are used to replace vocalized content in multimedia, then users can view content in their native language, but visual inconsistencies occur due to mismatched mouth movements and playback duration differences

Engineering Contradiction:
Improvelanguage adaptabilityVSAvoidvisual synchronization precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system dynamically adjusts the number of video frames to match the playback duration of dubbed audio content. When the dubbed audio duration differs from the original video duration, the system automatically adds or removes frames proportionally, maintaining visual-speech synchronization while accommodating language translation timing differences

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the frame rate parameter of the video content to align with the dubbed audio playback duration. By adjusting this fundamental parameter, the video playback duration is modified to match the audio duration, ensuring that lip movements remain synchronized with the dubbed speech without requiring manual frame-by-frame editing

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If manual reanimation of character mouth movements is performed to match dubbed vocal content, then visual synchronization is achieved, but production time and complexity increase significantly

Engineering Contradiction:
Improvelip-sync precisionVSAvoidproduction time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system replaces manual mechanical frame-by-frame animation work with an automated computational process. The processor automatically determines the frame adjustment quantity based on duration differences and applies the necessary frame additions or removals algorithmically, eliminating the need for time-consuming manual reanimation while maintaining precise lip-sync accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically detecting duration mismatches between dubbed audio and original video, calculating the required frame adjustments, and applying the modifications without human intervention. This automated self-correction process maintains synchronization precision while dramatically reducing production time compared to manual methods

Inventive Principle:
Principle #25Self-service

3Duration of action of moving object

If the number of frames in video content is modified to match dubbed audio duration, then playback duration synchronization is achieved, but frame rate consistency may be affected

Engineering Contradiction:
Improveplayback durationVSAvoidframe rate stability
Core Design Contradiction:
Duration of action of moving objectVSStability of the object's composition

Solution Approach 1:

The system deliberately changes the frame rate parameter as a controlled adjustment to achieve duration synchronization. By calculating the exact frame rate modification needed based on the ratio between dubbed audio duration and original video duration, the system maintains temporal consistency throughout the video, ensuring that all visual elements remain synchronized with the audio at the new frame rate

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9324340B2Methods and apparatuses for use in animating video content to correspond with audio content
Publication Date: 2016.04.26 SONY INTERACTIVE ENTERTAINMENT LLC
  • US9324340B2 patent drawing
  • US9324340B2 patent drawing
  • US9324340B2 patent drawing

AI summary

Some embodiments provide methods of reanimating multimedia content, comprising: accessing multimedia content; accessing a plurality of dubbed vocalized content; determining a playback duration of a first dubbed vocalized content is different than a first primary vocalized content; identifying a first portion of the primary visual content corresponding to the first primary vocalized content; modifying the first portion of the primary visual content such that a number of frames in the first portion of the primary visual content is changed and has a playback duration that is more consistent with the playback duration of the first dubbed vocalized content; identifying a character movement corresponding to each distinct vocal sound within the first dubbed vocalized content; and reanimating a first character such that reanimated character movements of the first character are consistent and synchronized with the identified character movement corresponding to each of the distinct vocal sounds.