Depth Camera Gesture Context for Speech Recognition Ambiguity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice recognition systems in vehicles face ambiguity due to shared verbal commands across different devices, leading to unintended operations and decreased user experience, as the context of user commands is not properly identified.

Innovation Solution

A computer-implemented method using a depth camera to capture images of a user's pose or gesture, processing this information to select appropriate verbal commands associated with targeted devices, thereby enhancing the accuracy of speech recognition by determining the context of user inputs based on the location of their hands or forearms relative to the camera.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition is used to operate vehicle devices, then hands-free operation is improved, but command ambiguity increases due to shared verbal commands across different devices

Engineering Contradiction:
Improvehands-free operationVSAvoidcommand accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces gesture recognition as an intermediary mechanism between the user and speech recognition systems. By detecting hand gestures and their spatial positions, the system determines contextual information about which device the user intends to control. This intermediary layer resolves the ambiguity of shared verbal commands by providing device-specific context, thereby maintaining hands-free operation while improving command accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple devices share similar verbal commands, then system versatility is improved, but command disambiguation becomes more difficult

Engineering Contradiction:
Improvedevice compatibilityVSAvoidcontext identification
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent adds a spatial dimension to command recognition by incorporating gesture position data. Instead of relying solely on verbal commands, the system detects the three-dimensional position of the user's hand gestures and uses this spatial information to determine which device context is intended. This dimensional addition allows multiple devices to share verbal commands while maintaining clear disambiguation through spatial context.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If gesture recognition is added to speech recognition, then command accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem architecture
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the command recognition system into distinct functional modules: gesture detection module, speech recognition module, and command disambiguation module. Each module handles a specific aspect of the recognition process independently. The gesture detection module captures hand position data, the speech recognition module processes verbal commands, and the disambiguation module integrates both inputs to determine the intended device. This segmentation improves accuracy while managing system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP2862125B1Depth based context identification
Publication Date: 2017.02.22 HONDA MOTOR CO LTD
  • EP2862125B1 patent drawing
  • EP2862125B1 patent drawing
  • EP2862125B1 patent drawing

AI summary

A method or system for selecting or pruning applicable verbal commands associated with speech recognition based on a user's motions detected from a depth camera. Depending on the depth of the user's hand or arm, the context of the verbal command is determined and verbal commands corresponding to the determined context are selected. Speech recognition is then performed on an audio signal using the selected verbal commands. By using an appropriate set of verbal commands, the accuracy of the speech recognition is increased.