Multimedia Authoring Syllable-Level Voice Motion Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multimedia authoring tools fail to achieve natural synchronization between voice and motion, particularly at the syllable level, leading to unnatural voice playback and limitations in applying voice synchronization to pre-made voice clips.

Innovation Solution

A multimedia authoring technique that allows for syllable-specific playback time point extraction through voice recognition, enabling users to edit and synchronize motion clips with voice clips on a timeline, allowing for natural voice and motion synchronization by altering voice clips on a syllable basis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional multimedia authoring tools are used for voice and motion synchronization, then the basic playback function is achieved, but syllable-level precise synchronization cannot be implemented

Engineering Contradiction:
Improvesynchronization precisionVSAvoidauthoring tool complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The voice clip is segmented into individual syllables, each represented as a separate block on the timeline. This segmentation enables precise syllable-level synchronization by allowing independent manipulation of each syllable's timing and positioning, resolving the contradiction between achieving high synchronization precision and managing system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The timeline serves as an intermediary visual interface that maps syllable boundaries to specific time positions. By visualizing syllables as discrete blocks on the timeline, the system provides an intuitive mediator between the voice audio track and motion animation track, enabling precise synchronization without complex manual adjustment.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If video-driven voice synchronization is applied, then lip movement synchronization is achieved, but the voice playback becomes unnatural and cannot be applied to pre-made voice clips

Engineering Contradiction:
Improvevoice playback naturalnessVSAvoidapplicability to pre-made voice clips
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

Instead of driving voice synthesis from video lip movements (video-driven approach), the system inverts the approach by extracting syllable timing information from the voice clip itself and using that to control motion playback. This inversion maintains the naturalness of pre-made voice clips while achieving synchronization, as the voice audio remains the primary reference rather than being synthesized from visual data.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The system copies the temporal structure and syllable boundaries from the original voice clip to control the motion playback timing. By replicating the voice clip's syllable timing pattern, the system preserves the natural characteristics of pre-made voice clips while synchronizing them with corresponding motion segments.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If syllable-level synchronization is implemented, then natural voice and motion synchronization is achieved, but the authoring process becomes more complex

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidauthoring operation simplicity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

By segmenting the voice clip into syllable blocks on the timeline, the system transforms the complex continuous synchronization problem into discrete manageable units. Each syllable block can be independently positioned and timed, making the authoring process more intuitive and easier to control while maintaining high synchronization accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The timeline display provides visual feedback showing the alignment between voice syllables and motion segments. This feedback mechanism allows authors to quickly adjust and verify synchronization accuracy, simplifying the authoring process by providing immediate visual confirmation of the synchronization state.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10885943B1Multimedia authoring apparatus with synchronized motion and voice feature and method for the same
Publication Date: 2021.01.05 ARTIFICIAL INTELLIGENCE RES INST
  • US10885943B1 patent drawing
  • US10885943B1 patent drawing
  • US10885943B1 patent drawing

AI summary

Disclosed is a technique for a multimedia authoring tool embodied in a computer program. A voice clip is displayed on a timeline, and a playback time point for each syllable and pronunciation information of the corresponding syllable are also displayed. Also, a motion clip may be edited in synchronization with the voice clip on the basis of the playback time point for each syllable. By moving a syllable of the voice clip along the timeline, a portion of the voice clip may be altered.