Audio-Visual Speech Recognition Using Camera Frames

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automatic speech recognition (ASR) systems face challenges in accurately distinguishing foreground speech from background interference, leading to reduced recognition accuracy in noisy environments.

Innovation Solution

The integration of visual information from imaging devices, such as cameras, to identify and separate speech from the foreground speaker by analyzing visual features in conjunction with audio frames, and transmitting only relevant audio frames to the ASR engine for processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional ASR systems process all audio input, then speech recognition can be performed, but recognition accuracy deteriorates in noisy environments with background interference

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidbackground noise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system segments the audio input by dividing it into multiple audio frames and further into sub-frames, allowing individual processing and analysis of different time segments to identify foreground speech versus background noise

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts visual information from camera feeds and integrates it with audio data to identify and isolate foreground speaker speech from background interference, effectively taking out the relevant speech signal for processing

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If visual information processing is added to identify foreground speech, then speech separation accuracy improves, but system complexity increases

Engineering Contradiction:
Improveforeground speech identification accuracyVSAvoidaudio-visual processing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses a unified audio-visual processing framework that handles both audio and visual data streams through integrated neural networks, allowing multi-functional processing that reduces overall system complexity despite the added visual processing capabilities

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10109277B2Methods and apparatus for speech recognition using visual information
Publication Date: 2018.10.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10109277B2 patent drawing
  • US10109277B2 patent drawing
  • US10109277B2 patent drawing

AI summary

Methods and apparatus for using visual information to facilitate a speech recognition process. The method comprises dividing received audio information into a plurality of audio frames, determining for each of the plurality of audio frames, whether the audio information in the audio frame comprises speech from the foreground speaker, wherein the determining is based, at least in part, on received visual information, and transmitting the audio frame to an automatic speech recognition (ASR) engine for speech recognition when it is determined that the audio frame comprises speech from the foreground speaker.