Visual Speech Recognition Using Unsupervised GAN Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional visual speech recognition systems are inflexible, inaccurate, and computationally inefficient due to their reliance on supervised machine learning models that require large amounts of labeled training data, limiting their ability to recognize speech across various digital video domains and consuming excessive computing resources.

Innovation Solution

The implementation of an unsupervised machine learning model using a generative adversarial neural network (GAN) that generates self-supervised deep visual speech representations from unlabeled digital videos, allowing for the recognition of viseme sequences and subsequent electronic transcription or audio generation without the need for annotated data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If supervised machine learning models are used for visual speech recognition, then the system can achieve reasonable accuracy on training data, but the system becomes inflexible and unable to recognize speech in videos not represented by the labeled training data

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidrecognition scope flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system uses self-supervised learning where the model learns from unlabeled video data by creating its own training signals through data augmentation and consistency regularization. The model generates pseudo-labels from augmented views of the same video, enabling it to learn speech representations without human-annotated data, thus achieving both accuracy and flexibility

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the fundamental parameter of training data supervision from labeled to unlabeled data. By using contrastive learning objectives and data augmentation strategies on unlabeled videos, the model learns robust speech representations that generalize across different video domains without requiring labeled training data

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If large annotated training data sets are used to improve speech recognition accuracy, then the system achieves better performance, but the computing resources and training time required increase significantly

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system eliminates the need for expensive human annotation by using self-supervised learning on unlabeled data. The model creates its own learning signals through data augmentation and consistency constraints, achieving comparable accuracy to supervised methods without the time-consuming annotation process

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system extracts and removes the dependency on labeled training data from the speech recognition pipeline. By using only unlabeled videos with data augmentation and self-supervised objectives, the system achieves accurate speech recognition without requiring annotated datasets, significantly reducing data preparation time

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If large annotated training data sets are stored and processed, then the system can train robust models, but excessive computing resources are consumed for data storage and processing

Engineering Contradiction:
Improvemodel robustnessVSAvoidcomputing resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system uses self-supervised learning to train robust models on unlabeled data, eliminating the need to store and process large annotated datasets. The model learns effective representations through data augmentation and consistency regularization, achieving model robustness with significantly reduced computing resources

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Instead of using labeled data to train models, the system inverts the approach by using unlabeled data with self-supervised objectives. This inversion removes the computational burden of data annotation and storage while maintaining model robustness through unsupervised learning signals

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS20230252993A1Visual speech recognition for digital videos utilizing generative adversarial learning
Publication Date: 2023.08.10 ADOBE INC
  • US20230252993A1 patent drawing
  • US20230252993A1 patent drawing
  • US20230252993A1 patent drawing

AI summary

This disclosure describes one or more implementations of systems, non-transitory computer-readable media, and methods that recognize speech from a digital video utilizing an unsupervised machine learning model, such as a generative adversarial neural network (GAN) model. In one or more implementations, the disclosed systems utilize an image encoder to generate self-supervised deep visual speech representations from frames of an unlabeled (or unannotated) digital video. Subsequently, in one or more embodiments, the disclosed systems generate viseme sequences from the deep visual speech representations (e.g., via segmented visemic speech representations from clusters of the deep visual speech representations) utilizing the adversarially trained GAN model. Indeed, in some instances, the disclosed systems decode the viseme sequences belonging to the digital video to generate an electronic transcription and/or digital audio for the digital video.