Speech Recognition False Positive Reduction via Acoustic-Visual Cross-Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition systems face challenges in distinguishing intended speech inputs from background noise and unintended speech sources, leading to false positives, especially in dynamic environments with multiple users.
Innovation Solution
The use of identity information, combining acoustic locational data from a microphone array and visual locational data from a depth-sensing camera, to adjust confidence levels in speech recognition systems, ensuring that only intended speech inputs from users within the field of view are recognized and processed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If speech recognition systems accept all speech inputs without verification, then the system responds quickly to user commands, but false positive recognition increases due to background noise and unintended speech
Solution Approach 1:
The patent introduces an intermediary verification mechanism that cross-references acoustic locational data with visual locational data from a camera. This intermediary check determines whether the speech source is a detected user before processing the speech input, thereby reducing false positives while maintaining quick response times for legitimate commands
2Reliability
If the system uses multiple verification methods including camera verification, then recognition accuracy improves, but system complexity increases
Solution Approach 1:
The patent merges the acoustic processing pipeline with visual processing by integrating camera data verification into the existing speech recognition flow. The system combines microphone array data with camera detection data in a unified verification process, reducing the need for separate complex verification subsystems while improving recognition accuracy
3Reliability
If the system verifies each speech input through multiple data sources, then false positives are reduced, but processing time increases
Solution Approach 1:
The system performs preliminary actions by continuously monitoring and pre-processing visual data from the camera to detect users in the field of view before speech inputs occur. This preliminary visual verification is ready in advance, so when speech is detected, the system can quickly cross-reference with pre-available visual data rather than performing full verification processing after speech detection
Data Source
AI summary
Embodiments are disclosed that relate to the use of identity information to help avoid the occurrence of false positive speech recognition events in a speech recognition system. One embodiment provides a method comprising receiving speech recognition data comprising a recognized speech segment, acoustic locational data related to a location of origin of the recognized speech segment as determined via signals from the microphone array, and confidence data comprising a recognition confidence value, and also receiving image data comprising visual locational information related to a location of each person in an image. The acoustic locational data is compared to the visual locational data to determine whether the recognized speech segment originated from a person in the field of view of the image sensor, and the confidence data is adjusted depending on this determination.


