Audio-Visual Speech Recognition Using Camera Frames
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic speech recognition (ASR) systems face challenges in accurately distinguishing foreground speech from background interference, leading to reduced recognition accuracy in noisy environments.
Innovation Solution
The integration of visual information from imaging devices, such as cameras, to identify and separate speech from the foreground speaker by analyzing visual features in conjunction with audio frames, and transmitting only relevant audio frames to the ASR engine for processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional ASR systems process all audio input, then speech recognition can be performed, but recognition accuracy deteriorates in noisy environments with background interference
Solution Approach 1:
The system segments the audio input by dividing it into multiple audio frames and further into sub-frames, allowing individual processing and analysis of different time segments to identify foreground speech versus background noise
Solution Approach 2:
The system extracts visual information from camera feeds and integrates it with audio data to identify and isolate foreground speaker speech from background interference, effectively taking out the relevant speech signal for processing
2Measurement precision
If visual information processing is added to identify foreground speech, then speech separation accuracy improves, but system complexity increases
Solution Approach 1:
The system uses a unified audio-visual processing framework that handles both audio and visual data streams through integrated neural networks, allowing multi-functional processing that reduces overall system complexity despite the added visual processing capabilities
Data Source
AI summary
Methods and apparatus for using visual information to facilitate a speech recognition process. The method comprises dividing received audio information into a plurality of audio frames, determining for each of the plurality of audio frames, whether the audio information in the audio frame comprises speech from the foreground speaker, wherein the determining is based, at least in part, on received visual information, and transmitting the audio frame to an automatic speech recognition (ASR) engine for speech recognition when it is determined that the audio frame comprises speech from the foreground speaker.


