Pseudo Audio-Visual Speech Reconstruction for Corrupted Multimedia

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems are limited in handling multiple forms of corrupted audio, requiring separate systems for different types of corruption, which hampers their ability to universally enhance speech in multimedia files.

Innovation Solution

A pseudo audio-visual speech recognition system that utilizes visual cues, such as mouth movements, in conjunction with self-supervised learning tokenizers to encode and synthesize clean speech, allowing for the reconstruction of speech in the presence of various forms of corrupted audio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate systems are constructed for different forms of corrupted audio, then each system can be optimized for its specific form, but the overall system complexity increases and adaptability to multiple corruption types decreases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem architecture
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a universal speech recognition system that can handle multiple forms of corrupted audio (noise, reverberation, mixed corruption) through a single integrated architecture. The system uses a corruption-agnostic front end that extracts features robust to various corruption types, followed by a back end that models speaker and acoustic variability. This multi-functional design eliminates the need for separate specialized systems while maintaining high recognition accuracy across different corruption scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system segments the speech recognition task into distinct functional components: a front end for corruption-robust feature extraction and a back end for speaker/acoustic modeling. This segmentation allows each component to be optimized independently while working together within a unified framework, reducing overall system complexity compared to having separate complete systems for each corruption type.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If a single universal system is used for all forms of corrupted audio, then system complexity is reduced, but the ability to handle specific corruption types effectively may be compromised

Engineering Contradiction:
Improvesystem architectureVSAvoidspeech recognition accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent applies local quality by making the front end corruption-agnostic (robust to various corruption types) while allowing the back end to model specific speaker and acoustic characteristics. This localized optimization ensures that the universal system maintains high accuracy for specific corruption types without requiring separate specialized systems, as each part of the system has tailored properties where needed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system handles different corruption types by changing parameters in the acoustic model rather than changing the overall system architecture. The corruption-agnostic front end extracts features that remain stable across corruption types, and the back end adapts to specific conditions through parameter adjustments in speaker and acoustic modeling, maintaining accuracy without increasing structural complexity.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If models are trained for particular forms of corrupted audio, then performance on those specific forms is improved, but the models become unsuitable for handling other forms of corrupted audio

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidhandling multiple corruption types
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent trains a universal model that can handle multiple corruption types through corruption-agnostic feature extraction. The front end is designed to extract features that are robust to various corruption types (noise, reverberation, mixed corruption) without requiring separate training for each type. This enables the single model to generalize across different corruption scenarios, achieving both high accuracy and broad adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The corruption-agnostic front end automatically adapts to different corruption types through self-supervised learning and data augmentation techniques. The system uses available training data to learn robust feature representations that inherently handle various corruption types without requiring explicit corruption-type-specific training, enabling the model to serve multiple corruption scenarios effectively.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240282291A1Speech Reconstruction System for Multimedia Files
Publication Date: 2024.08.22 META PLATFORMS TECHNOLOGIES LLC
  • US20240282291A1 patent drawing
  • US20240282291A1 patent drawing
  • US20240282291A1 patent drawing

AI summary

A speech recognition system may determine speech in the presence of multiple, different forms of corrupted audio. The system may obtain audio-visual data including visual data associated with a person and audio data associated with the person. The system may also determine, based on the visual data, pronunciation data associated with speech by the person. The system may also convert the speech to encoded data. The system may also synthesize, based on the encoded data, the speech to obtain synthesized speech.