Visual Speech Recognition Using Unsupervised GAN Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional visual speech recognition systems are inflexible, inaccurate, and computationally inefficient due to their reliance on supervised machine learning models that require large amounts of labeled training data, limiting their ability to recognize speech across various digital video domains and consuming excessive computing resources.
Innovation Solution
The implementation of an unsupervised machine learning model using a generative adversarial neural network (GAN) that generates self-supervised deep visual speech representations from unlabeled digital videos, allowing for the recognition of viseme sequences and subsequent electronic transcription or audio generation without the need for annotated data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If supervised machine learning models are used for visual speech recognition, then the system can achieve reasonable accuracy on training data, but the system becomes inflexible and unable to recognize speech in videos not represented by the labeled training data
Solution Approach 1:
The system uses self-supervised learning where the model learns from unlabeled video data by creating its own training signals through data augmentation and consistency regularization. The model generates pseudo-labels from augmented views of the same video, enabling it to learn speech representations without human-annotated data, thus achieving both accuracy and flexibility
Solution Approach 2:
The system changes the fundamental parameter of training data supervision from labeled to unlabeled data. By using contrastive learning objectives and data augmentation strategies on unlabeled videos, the model learns robust speech representations that generalize across different video domains without requiring labeled training data
2Measurement precision
If large annotated training data sets are used to improve speech recognition accuracy, then the system achieves better performance, but the computing resources and training time required increase significantly
Solution Approach 1:
The system eliminates the need for expensive human annotation by using self-supervised learning on unlabeled data. The model creates its own learning signals through data augmentation and consistency constraints, achieving comparable accuracy to supervised methods without the time-consuming annotation process
Solution Approach 2:
The system extracts and removes the dependency on labeled training data from the speech recognition pipeline. By using only unlabeled videos with data augmentation and self-supervised objectives, the system achieves accurate speech recognition without requiring annotated datasets, significantly reducing data preparation time
3Reliability
If large annotated training data sets are stored and processed, then the system can train robust models, but excessive computing resources are consumed for data storage and processing
Solution Approach 1:
The system uses self-supervised learning to train robust models on unlabeled data, eliminating the need to store and process large annotated datasets. The model learns effective representations through data augmentation and consistency regularization, achieving model robustness with significantly reduced computing resources
Solution Approach 2:
Instead of using labeled data to train models, the system inverts the approach by using unlabeled data with self-supervised objectives. This inversion removes the computational burden of data annotation and storage while maintaining model robustness through unsupervised learning signals
Data Source
AI summary
This disclosure describes one or more implementations of systems, non-transitory computer-readable media, and methods that recognize speech from a digital video utilizing an unsupervised machine learning model, such as a generative adversarial neural network (GAN) model. In one or more implementations, the disclosed systems utilize an image encoder to generate self-supervised deep visual speech representations from frames of an unlabeled (or unannotated) digital video. Subsequently, in one or more embodiments, the disclosed systems generate viseme sequences from the deep visual speech representations (e.g., via segmented visemic speech representations from clusters of the deep visual speech representations) utilizing the adversarially trained GAN model. Indeed, in some instances, the disclosed systems decode the viseme sequences belonging to the digital video to generate an electronic transcription and/or digital audio for the digital video.


