Audio Synthesis Synchronizing Speech with Video Characteristics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies do not consider video environments, making it cumbersome and time-consuming to synchronize speech with videos for dubbing purposes, requiring manual processes.
Innovation Solution
An audio synthesis method that receives video and text inputs, extracts time-series characteristics from the video and phoneme characteristics from the text, using AI models to generate audio spectrum characteristics that correlate with video characteristics, enabling automated synchronization of speech with video.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If speech synthesis is performed without considering video characteristics, then the synthesis process is simple, but the speech cannot be synchronized with video for dubbing purposes
Solution Approach 1:
The system extracts video characteristics (mouth opening degree, tongue position, etc.) in advance and stores them as reference data before speech synthesis. This preliminary extraction enables the synthesis process to directly reference pre-analyzed video features, achieving synchronized speech output without adding complex real-time analysis during synthesis
Solution Approach 2:
The patent introduces an intermediate processing layer that extracts and analyzes video characteristics separately, creating a bridge between raw video input and speech synthesis. This intermediary analysis of mouth shapes, tongue positions, and other visual features enables precise synchronization while keeping the core synthesis process manageable
2Manufacturing precision
If manual operation is used to convert speech and synchronize with video, then speech-video synchronization can be achieved, but the process is cumbersome and time-consuming
Solution Approach 1:
The system performs automatic extraction of video characteristics and automated speech synthesis with built-in synchronization capability. The processor autonomously analyzes video frames, extracts relevant features, and generates synchronized speech without requiring manual intervention, thereby eliminating the time-consuming manual conversion process while maintaining high synchronization precision
Solution Approach 2:
The patent replaces the manual mechanical process of speech conversion and video synchronization with an automated computational system. The processor executes algorithms that automatically align speech with video characteristics, substituting human manual operations with efficient machine-based automated processing
3Adaptability or versatility
If speech is added to video without considering video characteristics, then the process is simple, but the speech is not contextually relevant to the video
Solution Approach 1:
The system extracts specific local characteristics from video frames such as mouth opening degree, tongue position, and facial expressions rather than analyzing the entire video globally. This localized feature extraction focuses computational resources on relevant visual cues that directly impact speech synthesis quality, achieving contextual adaptability without excessive complexity
Solution Approach 2:
The patent segments the video analysis process into distinct feature extraction components (mouth shape, tongue position, facial expression) that can be independently processed and combined. This segmentation allows the system to adapt speech synthesis to specific video contexts by selectively processing relevant visual features
Data Source
AI summary
An audio synthesis method adapted to video characteristics is provided. The audio synthesis method according to an embodiment includes: extracting characteristics x from a video in a time-series way; extracting characteristics p of phonemes from a text; and generating an audio spectrum characteristic St used to generate an audio to be synthesized with a video at a time t, based on correlations between an audio spectrum characteristic St-1, which is used to generate an audio to be synthesized with a video at a time t−1, and the characteristics x. Accordingly, an audio can be synthesized according to video characteristics, and speech according to a video can be easily added.


