Audio Synthesis Synchronizing Speech with Video Characteristics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis technologies do not consider video environments, making it cumbersome and time-consuming to synchronize speech with videos for dubbing purposes, requiring manual processes.

Innovation Solution

An audio synthesis method that receives video and text inputs, extracts time-series characteristics from the video and phoneme characteristics from the text, using AI models to generate audio spectrum characteristics that correlate with video characteristics, enabling automated synchronization of speech with video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If speech synthesis is performed without considering video characteristics, then the synthesis process is simple, but the speech cannot be synchronized with video for dubbing purposes

Engineering Contradiction:
Improvespeech-video synchronization precisionVSAvoidsynthesis process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system extracts video characteristics (mouth opening degree, tongue position, etc.) in advance and stores them as reference data before speech synthesis. This preliminary extraction enables the synthesis process to directly reference pre-analyzed video features, achieving synchronized speech output without adding complex real-time analysis during synthesis

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediate processing layer that extracts and analyzes video characteristics separately, creating a bridge between raw video input and speech synthesis. This intermediary analysis of mouth shapes, tongue positions, and other visual features enables precise synchronization while keeping the core synthesis process manageable

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If manual operation is used to convert speech and synchronize with video, then speech-video synchronization can be achieved, but the process is cumbersome and time-consuming

Engineering Contradiction:
Improvespeech-video synchronization precisionVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs automatic extraction of video characteristics and automated speech synthesis with built-in synchronization capability. The processor autonomously analyzes video frames, extracts relevant features, and generates synchronized speech without requiring manual intervention, thereby eliminating the time-consuming manual conversion process while maintaining high synchronization precision

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of speech conversion and video synchronization with an automated computational system. The processor executes algorithms that automatically align speech with video characteristics, substituting human manual operations with efficient machine-based automated processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If speech is added to video without considering video characteristics, then the process is simple, but the speech is not contextually relevant to the video

Engineering Contradiction:
Improvespeech adaptability to video contextVSAvoidanalysis process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system extracts specific local characteristics from video frames such as mouth opening degree, tongue position, and facial expressions rather than analyzing the entire video globally. This localized feature extraction focuses computational resources on relevant visual cues that directly impact speech synthesis quality, achieving contextual adaptability without excessive complexity

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the video analysis process into distinct feature extraction components (mouth shape, tongue position, facial expression) that can be independently processed and combined. This segmentation allows the system to adapt speech synthesis to specific video contexts by selectively processing relevant visual features

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10923106B2Method for audio synthesis adapted to video characteristics
Publication Date: 2021.02.16 KOREA ELECTRONICS TECH INST
  • US10923106B2 patent drawing
  • US10923106B2 patent drawing
  • US10923106B2 patent drawing

AI summary

An audio synthesis method adapted to video characteristics is provided. The audio synthesis method according to an embodiment includes: extracting characteristics x from a video in a time-series way; extracting characteristics p of phonemes from a text; and generating an audio spectrum characteristic St used to generate an audio to be synthesized with a video at a time t, based on correlations between an audio spectrum characteristic St-1, which is used to generate an audio to be synthesized with a video at a time t−1, and the characteristics x. Accordingly, an audio can be synthesized according to video characteristics, and speech according to a video can be easily added.