Multimodal Voice Recognition With Lip Video–Audio Feature Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice recognition technologies face challenges in environments with overlapping utterances or low signal-to-noise ratios, leading to reduced performance when relying on single modality inputs, such as audio or video alone.
Innovation Solution
A multi-modality voice recognition system using a single deep learning model that integrates video and audio inputs, incorporating a video encoder, audio encoder, modality reconstructor, random selector, and video-audio decoder to enhance recognition performance across various conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single modality-based voice recognition model is used, then the device complexity is low, but the recognition reliability deteriorates in environments with overlapping utterances or low signal-to-noise ratios
Solution Approach 1:
The patent combines video and audio modalities into a unified deep learning model that processes both input types simultaneously. The model integrates visual features from lip videos and acoustic features from audio signals through shared encoder layers, enabling the system to leverage complementary information from both modalities to improve recognition reliability in challenging environments while maintaining a relatively compact architecture.
Solution Approach 2:
The patent creates a universal deep learning model that can handle multiple input conditions - it processes both video-only and audio-only inputs through the same architecture. The model's encoders and shared layers are designed to be modality-agnostic, allowing it to adaptively process different input types and maintain reliable performance across various service environments without requiring separate specialized models.
2Manufacturing precision
If conventional technology is optimized for a specific modality input condition, then the manufacturing precision of the model is high for that condition, but the adaptability to different environments deteriorates
Solution Approach 1:
The patent designs a universal model architecture that processes both video and audio inputs through shared encoder layers, enabling it to adapt to different service environments. The model maintains high performance for specific conditions (video-only, audio-only, or combined inputs) while simultaneously being adaptable to various environmental conditions such as overlapping utterances and low signal-to-noise ratio scenarios through its multi-modality design.
Data Source
AI summary
A multi-modality voice recognition device is provided. The multi-modality voice recognition device includes a video encoder trained to receive a lip video to a video encoder to extract visual feature information for voice recognition, an audio encoder trained to receive a voice to extract voice feature information for voice recognition, a modality reconstructor trained to reconstruct the voice feature information from the visual feature information to generate reconstruction voice feature information, a random selector configured to randomly output one of the voice feature information and the reconstruction voice feature information, and a video-audio decoder trained to receive a multi-modality feature, where the visual feature information is connected to an output of the random selector, to output a character string which is a voice recognition result.


