Text-Driven Video Synthesis Using Phoneme-Pose Dictionary

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech-to-video models are not robust to different speakers, require large training datasets, and are less flexible in using speech as input, with limitations in manipulating motion output and handling speaker variability.

Innovation Solution

A text-based approach that generates talking-head videos from text input using a phoneme-pose dictionary and a generative adversarial network (GAN), employing forced alignment to extract phonemes and their timestamps, and interpolating pose sequences to produce photorealistic videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech-to-video models are trained on large amounts of data to handle speaker variability, then robustness to different speakers is improved, but training data requirements increase and device complexity increases

Engineering Contradiction:
Improverobustness to different speakersVSAvoidtraining data requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent introduces an intermediate representation system that converts speech into phoneme sequences and visual feature sequences, which then drive the video generation process. This intermediary approach allows the system to handle speaker variability without requiring extensive speaker-specific training data, as the phoneme-based representation is speaker-agnostic

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system transforms the input speech into a different parameter space (phoneme sequences and visual features) rather than directly mapping speaker-specific acoustic features to video. This parameter transformation enables generalization across speakers with minimal training data

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If speech-to-video models use LSTM to learn audio information, then audio processing capability is improved, but flexibility in manipulating motion output decreases and the network becomes a black box

Engineering Contradiction:
Improveflexibility in manipulating motion outputVSAvoidnetwork complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the speech-to-video generation process into distinct components: speech-to-phoneme conversion, phoneme-to-visual-feature mapping, and video generation from visual features. This segmentation allows each component to be independently controlled and manipulated, providing flexibility in motion output while maintaining manageable system complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The introduction of phoneme sequences and visual feature sequences as intermediate representations provides interpretable control points in the system. These intermediaries allow users to manipulate specific aspects of the output (such as lip-sync accuracy or facial expressions) without dealing with the complexity of the underlying neural network

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If speech-to-video models are designed to handle different speakers, then adaptability to speaker variability is improved, but training data requirements increase

Engineering Contradiction:
Improvehandling speaker variabilityVSAvoidtraining data requirements
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent creates a universal phoneme-based representation system that functions across different speakers and languages. The visual feature extraction module is designed to be speaker-agnostic, extracting only the essential visual characteristics needed for video generation without encoding speaker-specific information, thereby achieving broad adaptability with minimal speaker-specific training data

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11587548B2Text-driven video synthesis with phonetic dictionary
Publication Date: 2023.02.21 BAIDU USA LLC
  • US11587548B2 patent drawing
  • US11587548B2 patent drawing
  • US11587548B2 patent drawing

AI summary

Presented herein are novel approaches to synthesize video of the speech from text. In a training phase, embodiments build a phoneme-pose dictionary and train a generative neural network model using a generative adversarial network (GAN) to generate video from interpolated phoneme poses. In deployment, the trained generative neural network in conjunction with the phoneme-pose dictionary convert an input text into a video of a person speaking the words of the input text. Compared to audio-driven video generation approaches, the embodiments herein have a number of advantages: 1) they only need a fraction of the training data used by an audio-driven approach; 2) they are more flexible and not subject to vulnerability due to speaker variation; and 3) they significantly reduce the preprocessing, training, and inference times.