Image-Guided Voice Recognition for Multi-User False Trigger Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition technologies face challenges in adapting to multiple-user environments, leading to increased instances of false positives and reduced efficiency in voice command recognition.
Innovation Solution
The use of image data to capture and analyze the vicinity of a device, adjusting voice recognition parameters such as trigger thresholds and algorithms based on the number of individuals detected, and employing gaze detection to authenticate authorized users and filter out unintended voice commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice recognition uses a fixed trigger threshold, then the device is simple to operate, but it produces false positives in multi-user environments
Solution Approach 1:
The patent implements dynamic adjustment of voice recognition parameters including trigger thresholds and algorithms based on the detected number of individuals in the vicinity. The system transitions from static, fixed parameters to dynamic parameters that automatically adapt to environmental conditions, resolving the contradiction between reliability and complexity by making the system adaptive rather than manually configurable
Solution Approach 2:
The voice recognition system performs self-adjustment by automatically detecting the number of people nearby and modifying its own parameters without user intervention. This self-service mechanism eliminates the need for manual parameter tuning while maintaining high recognition accuracy across different user scenarios
2Reliability
If the device adapts voice recognition parameters based on image data, then voice recognition accuracy improves, but processing time increases
Solution Approach 1:
The system captures image data and determines the number of individuals in advance before voice recognition processing begins. This preliminary action allows the appropriate voice recognition parameters to be pre-configured, eliminating delays during actual voice command processing and reducing overall time loss
3Reliability
If gaze detection is used to authenticate users, then false positives are reduced, but device complexity increases
Solution Approach 1:
The patent combines gaze detection with existing voice recognition functionality to create an integrated authentication system. By merging these two methods, the system achieves improved user authentication accuracy without requiring completely separate systems, thus limiting the increase in overall device complexity
Data Source
AI summary
A device performs a method for using image data to aid voice recognition. The method includes the device capturing image data of a vicinity of the device and adjusting, based on the image data, a set of parameters for voice recognition performed by the device. The set of parameters for the device performing voice recognition include, but are not limited to: a trigger threshold of a trigger for voice recognition; a set of beamforming parameters; a database for voice recognition; and/or an algorithm for voice recognition. The algorithm may include using noise suppression or using acoustic beamforming.


