Speech-to-text conversion using audio and video fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition techniques suffer from inaccuracies due to background noise, leading to a maximum accuracy of 80%, which compromises the reliability of applications and can result in security vulnerabilities.
Innovation Solution
A method and system that utilizes both audio and video inputs to generate raw texts, compares them for errors, and applies rules based on domain-specific databases, conversation context, and prior communication history to correct errors, thereby improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing speech recognition techniques are applied, then speech-to-text conversion can be performed, but background noise is captured resulting in loss of words and misinterpretation, causing accuracy to decline to maximum 80 percent
Solution Approach 1:
The patent segments the speech signal processing into multiple independent components: audio signal processing, video signal processing, and their integration. By dividing the audio signal into frequency bands and processing them separately before combining, the system can selectively enhance speech components while suppressing background noise in each band, thereby improving overall accuracy beyond the 80% limitation of conventional single-channel approaches
Solution Approach 2:
The patent merges audio and video signal processing paths to create a unified speech recognition system. The audio processor and video processor work together, with their outputs combined in an integrator that produces a final output signal. This combination allows the system to leverage complementary information from both modalities, where video data helps distinguish speech from background noise that confuses audio-only processors, thereby improving accuracy while maintaining robustness against background interference
2Reliability
If existing speech recognition techniques are applied, then speech processing can be performed, but reliability of applications is compromised due to inaccuracy and inability to differentiate between false positives and false negatives
Solution Approach 1:
The patent implements feedback mechanisms where the output from one processing path is used to refine the other. The video processor provides feedback to the audio processor to help identify and suppress background noise, while the audio processor's speech detection results guide the video processing focus. This cross-modal feedback loop continuously refines the recognition accuracy, reducing false positives and false negatives, thereby improving application reliability for security-critical systems
Solution Approach 2:
The patent introduces an integrator as an intermediary component that receives processed signals from both audio and video processors and combines them to produce the final output. This intermediary performs sophisticated fusion that weighs evidence from both modalities, allowing the system to resolve ambiguities that would cause false positives or negatives in either modality alone, thereby enhancing reliability for critical applications
Data Source
AI summary
This disclosure relates generally to speech recognition, and more particularly to system and method for speech-to-text conversion using audio as well as video input. In one embodiment, a method is provided for performing speech to text conversion. The method comprises receiving an audio data and a video data of a user while the user is speaking, generating a first raw text based on the audio data via one or more audio-to-text conversion algorithms, generating a second raw text based on the video data via one or more video-to-text conversion algorithms, determining one or more errors by comparing the first raw text and the second raw text, and correcting the one or more errors by applying one or more rules. The one or more rules employ at least one of a domain specific word database, a context of conversation, and a prior communication history.


