User-Dedicated ASR with Spatial Filtering and Gesture Switching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controlled systems using automatic speech recognition (ASR) face challenges in distinguishing between multiple users, leading to interference and decreased performance for unintended speakers due to spatial positioning and beamforming direction, which limits effective multi-user interactions.

Innovation Solution

A user-dedicated, multi-mode voice-controlled interface that employs acoustic and visual information to selectively focus on a specific speaker using activation words, gestures, and image processing, switching between broad and selective listening modes to ensure only the designated user can control the system, with adaptive beamforming and vocabulary adjustments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If beamforming spatial filtering is used to improve ASR performance for a desired speaker, then recognition accuracy for the target speaker is improved, but ASR performance for other speakers deteriorates

Engineering Contradiction:
ImproveASR recognition accuracyVSAvoidmulti-user recognition capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The beamforming spatial filter direction is dynamically adjusted based on detected user positions and activation gestures. The system transitions from a fixed steering direction to a dynamic one that follows the active user's location, allowing the same hardware configuration to serve different users by changing the spatial filtering direction in real-time

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary detection of user positions and activation gestures before initiating speech recognition. By detecting which user has activated the system through gesture recognition and position tracking, the system pre-configures the beamforming spatial filter to focus on the correct user's direction before processing their speech, ensuring accurate recognition from the start

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If a push-to-talk button or activation word is used to trigger ASR, then the system can respond to any user's speech, but it cannot distinguish between multiple users, leading to interference

Engineering Contradiction:
Improvevoice control accessibilityVSAvoiduser-specific recognition reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system introduces gesture recognition and user position detection as intermediary mechanisms between the user and the ASR system. Instead of directly responding to any speech input, the system first detects which user has performed the activation gesture and is positioned in the detection zone, using this intermediate information to determine whether to process the speech input

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments the speech recognition process by creating separate detection zones for different users and requiring user-specific activation gestures. Each user has their own spatial segment and gesture sequence requirements, which divides the overall speech control function into user-specific segments, preventing cross-user interference

Inventive Principle:
Principle #1Segmentation

3Reliability

If the ASR system is dedicated to a single user with directional beamforming, then interference from other users is suppressed, but the system cannot handle multiple users simultaneously

Engineering Contradiction:
Improveuser-dedicated recognition accuracyVSAvoidmulti-user system capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system dynamically switches between serving different users based on real-time detection of activation gestures and user positions. When one user activates the system, the beamforming spatial filter is directed toward that user; when another user activates, the filter direction changes accordingly. This dynamic reconfiguration allows the single-channel ASR system to serve multiple users sequentially without interference

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from gesture detection and position tracking to control the beamforming spatial filter direction. The detection of user gestures and positions provides feedback that automatically adjusts the spatial filtering to focus on the active user, creating a closed-loop system that adapts to multi-user scenarios while maintaining user-dedicated recognition accuracy

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP2817801B1User dedicated automatic speech recognition
Publication Date: 2017.02.22 NUANCE COMMUNICATIONS INC
  • EP2817801B1 patent drawing
  • EP2817801B1 patent drawing
  • EP2817801B1 patent drawing

AI summary

A multi-mode voice controlled user interface is described. The user interface is adapted to conduct a speech dialog with one or more possible speakers and includes a broad listening mode which accepts speech inputs from the possible speakers without spatial filtering, and a selective listening mode which limits speech inputs to a specific speaker using spatial filtering. The user interface switches listening modes in response to one or more switching cues.