Text-Driven Video Synthesis Using Phoneme-Pose Dictionary
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech-to-video models are not robust to different speakers, require large training datasets, and are less flexible in using speech as input, with limitations in manipulating motion output and handling speaker variability.
Innovation Solution
A text-based approach that generates talking-head videos from text input using a phoneme-pose dictionary and a generative adversarial network (GAN), employing forced alignment to extract phonemes and their timestamps, and interpolating pose sequences to produce photorealistic videos.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech-to-video models are trained on large amounts of data to handle speaker variability, then robustness to different speakers is improved, but training data requirements increase and device complexity increases
Solution Approach 1:
The patent introduces an intermediate representation system that converts speech into phoneme sequences and visual feature sequences, which then drive the video generation process. This intermediary approach allows the system to handle speaker variability without requiring extensive speaker-specific training data, as the phoneme-based representation is speaker-agnostic
Solution Approach 2:
The system transforms the input speech into a different parameter space (phoneme sequences and visual features) rather than directly mapping speaker-specific acoustic features to video. This parameter transformation enables generalization across speakers with minimal training data
2Ease of operation
If speech-to-video models use LSTM to learn audio information, then audio processing capability is improved, but flexibility in manipulating motion output decreases and the network becomes a black box
Solution Approach 1:
The patent segments the speech-to-video generation process into distinct components: speech-to-phoneme conversion, phoneme-to-visual-feature mapping, and video generation from visual features. This segmentation allows each component to be independently controlled and manipulated, providing flexibility in motion output while maintaining manageable system complexity
Solution Approach 2:
The introduction of phoneme sequences and visual feature sequences as intermediate representations provides interpretable control points in the system. These intermediaries allow users to manipulate specific aspects of the output (such as lip-sync accuracy or facial expressions) without dealing with the complexity of the underlying neural network
3Adaptability or versatility
If speech-to-video models are designed to handle different speakers, then adaptability to speaker variability is improved, but training data requirements increase
Solution Approach 1:
The patent creates a universal phoneme-based representation system that functions across different speakers and languages. The visual feature extraction module is designed to be speaker-agnostic, extracting only the essential visual characteristics needed for video generation without encoding speaker-specific information, thereby achieving broad adaptability with minimal speaker-specific training data
Data Source
AI summary
Presented herein are novel approaches to synthesize video of the speech from text. In a training phase, embodiments build a phoneme-pose dictionary and train a generative neural network model using a generative adversarial network (GAN) to generate video from interpolated phoneme poses. In deployment, the trained generative neural network in conjunction with the phoneme-pose dictionary convert an input text into a video of a person speaking the words of the input text. Compared to audio-driven video generation approaches, the embodiments herein have a number of advantages: 1) they only need a fraction of the training data used by an audio-driven approach; 2) they are more flexible and not subject to vulnerability due to speaker variation; and 3) they significantly reduce the preprocessing, training, and inference times.


