Image-Aided Voice Recognition for Multi-User Command Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition technologies face challenges in adapting to multiple-user environments, leading to increased instances of false positives and reduced efficiency in voice command recognition.
Innovation Solution
The use of image data to capture and analyze the vicinity of a device, adjusting voice recognition parameters such as trigger thresholds and algorithms based on the presence of individuals, and employing gaze detection to authenticate and prioritize voice inputs from authorized users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice recognition is performed without image data, then the system is simpler and faster, but accuracy in multiple-user environments deteriorates
Solution Approach 1:
The patent combines voice recognition with image recognition by merging the voice processing module and image processing module into an integrated system. The voice recognition module processes audio input while the image recognition module processes camera input, and both modules work together to identify the active user through multiple modalities, thereby improving accuracy without significantly increasing perceived system complexity.
Solution Approach 2:
The patent introduces an active user identification module that acts as an intermediary between the voice recognition module and the device control system. This module uses both voice and image data to determine which user is actively speaking and intends to interact with the device, then provides this information to adjust voice recognition parameters accordingly.
2Adaptability or versatility
If voice recognition parameters are fixed, then the system is more stable, but adaptability to multiple-user environments deteriorates
Solution Approach 1:
The patent implements dynamic adjustment of voice recognition parameters based on real-time identification of the active user. When a different user is detected through image or voice analysis, the system automatically adjusts parameters such as sensitivity thresholds, accepted voice commands, and language models to match the newly identified user's profile, enabling seamless adaptation while maintaining operational stability.
Solution Approach 2:
The system changes multiple parameters including voice recognition sensitivity, acceptable dialect ranges, and command interpretation thresholds based on the identified user's characteristics. These parameter changes allow the same device to serve multiple users with different voice patterns and preferences while maintaining stable performance for each individual user.
3Measurement precision
If image processing is added to aid voice recognition, then user differentiation improves, but processing time increases
Solution Approach 1:
The system performs preliminary actions by continuously capturing and pre-processing image data in the background even before voice input is received. Facial recognition features are extracted and stored in advance, so when voice recognition is triggered, the system can quickly match the current speaker against pre-processed user profiles without performing full image analysis from scratch, significantly reducing processing time.
Solution Approach 2:
The patent implements a selective processing approach where the system skips full image analysis in favor of quicker verification methods when possible. For example, if basic facial recognition quickly identifies the active user, the system skips more time-consuming detailed analysis and proceeds directly to voice recognition parameter adjustment, thereby maintaining accuracy while minimizing time loss.
Data Source
AI summary
A device performs a method for using image data to aid voice recognition. The method includes the device capturing image data of a vicinity of the device and adjusting, based on the image data, a set of parameters for voice recognition performed by the device. The set of parameters for the device performing voice recognition include, but are not limited to: a trigger threshold of a trigger for voice recognition; a set of beamforming parameters; a database for voice recognition; and/or an algorithm for voice recognition. The algorithm may include using noise suppression or using acoustic beamforming.


