Speech Recognition False Positive Reduction via Acoustic-Visual Cross-Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition systems face challenges in distinguishing intended speech inputs from background noise and unintended speech sources, leading to false positives, especially in dynamic environments with multiple users.

Innovation Solution

The use of identity information, combining acoustic locational data from a microphone array and visual locational data from a depth-sensing camera, to adjust confidence levels in speech recognition systems, ensuring that only intended speech inputs from users within the field of view are recognized and processed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If speech recognition systems accept all speech inputs without verification, then the system responds quickly to user commands, but false positive recognition increases due to background noise and unintended speech

Engineering Contradiction:
Improveresponse speedVSAvoidrecognition accuracy
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent introduces an intermediary verification mechanism that cross-references acoustic locational data with visual locational data from a camera. This intermediary check determines whether the speech source is a detected user before processing the speech input, thereby reducing false positives while maintaining quick response times for legitimate commands

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system uses multiple verification methods including camera verification, then recognition accuracy improves, but system complexity increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges the acoustic processing pipeline with visual processing by integrating camera data verification into the existing speech recognition flow. The system combines microphone array data with camera detection data in a unified verification process, reducing the need for separate complex verification subsystems while improving recognition accuracy

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If the system verifies each speech input through multiple data sources, then false positives are reduced, but processing time increases

Engineering Contradiction:
Improvefalse positive reductionVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by continuously monitoring and pre-processing visual data from the camera to detect users in the field of view before speech inputs occur. This preliminary visual verification is ready in advance, so when speech is detected, the system can quickly cross-reference with pre-available visual data rather than performing full verification processing after speech detection

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8676581B2Speech recognition analysis via identification information
Publication Date: 2014.03.18 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8676581B2 patent drawing
  • US8676581B2 patent drawing
  • US8676581B2 patent drawing

AI summary

Embodiments are disclosed that relate to the use of identity information to help avoid the occurrence of false positive speech recognition events in a speech recognition system. One embodiment provides a method comprising receiving speech recognition data comprising a recognized speech segment, acoustic locational data related to a location of origin of the recognized speech segment as determined via signals from the microphone array, and confidence data comprising a recognition confidence value, and also receiving image data comprising visual locational information related to a location of each person in an image. The acoustic locational data is compared to the visual locational data to determine whether the recognized speech segment originated from a person in the field of view of the image sensor, and the confidence data is adjusted depending on this determination.