User-Dedicated ASR with Spatial Filtering and Gesture Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled systems using automatic speech recognition (ASR) face challenges in distinguishing between multiple users, leading to interference and decreased performance for unintended speakers due to spatial positioning and beamforming direction, which limits effective multi-user interactions.
Innovation Solution
A user-dedicated, multi-mode voice-controlled interface that employs acoustic and visual information to selectively focus on a specific speaker using activation words, gestures, and image processing, switching between broad and selective listening modes to ensure only the designated user can control the system, with adaptive beamforming and vocabulary adjustments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If beamforming spatial filtering is used to improve ASR performance for a desired speaker, then recognition accuracy for the target speaker is improved, but ASR performance for other speakers deteriorates
Solution Approach 1:
The beamforming spatial filter direction is dynamically adjusted based on detected user positions and activation gestures. The system transitions from a fixed steering direction to a dynamic one that follows the active user's location, allowing the same hardware configuration to serve different users by changing the spatial filtering direction in real-time
Solution Approach 2:
The system performs preliminary detection of user positions and activation gestures before initiating speech recognition. By detecting which user has activated the system through gesture recognition and position tracking, the system pre-configures the beamforming spatial filter to focus on the correct user's direction before processing their speech, ensuring accurate recognition from the start
2Ease of operation
If a push-to-talk button or activation word is used to trigger ASR, then the system can respond to any user's speech, but it cannot distinguish between multiple users, leading to interference
Solution Approach 1:
The system introduces gesture recognition and user position detection as intermediary mechanisms between the user and the ASR system. Instead of directly responding to any speech input, the system first detects which user has performed the activation gesture and is positioned in the detection zone, using this intermediate information to determine whether to process the speech input
Solution Approach 2:
The system segments the speech recognition process by creating separate detection zones for different users and requiring user-specific activation gestures. Each user has their own spatial segment and gesture sequence requirements, which divides the overall speech control function into user-specific segments, preventing cross-user interference
3Reliability
If the ASR system is dedicated to a single user with directional beamforming, then interference from other users is suppressed, but the system cannot handle multiple users simultaneously
Solution Approach 1:
The system dynamically switches between serving different users based on real-time detection of activation gestures and user positions. When one user activates the system, the beamforming spatial filter is directed toward that user; when another user activates, the filter direction changes accordingly. This dynamic reconfiguration allows the single-channel ASR system to serve multiple users sequentially without interference
Solution Approach 2:
The system uses feedback from gesture detection and position tracking to control the beamforming spatial filter direction. The detection of user gestures and positions provides feedback that automatically adjusts the spatial filtering to focus on the active user, creating a closed-loop system that adapts to multi-user scenarios while maintaining user-dedicated recognition accuracy
Data Source
AI summary
A multi-mode voice controlled user interface is described. The user interface is adapted to conduct a speech dialog with one or more possible speakers and includes a broad listening mode which accepts speech inputs from the possible speakers without spatial filtering, and a selective listening mode which limits speech inputs to a specific speaker using spatial filtering. The user interface switches listening modes in response to one or more switching cues.


