Multimodal Command Recognition via Verb-Noun Parsing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech and gesture recognition technologies in multimodal environments are static and brittle, struggling with accurate transcription of terse utterances and anaphora, leading to poor command recognition quality.

Innovation Solution

A method and system that utilize a hardware-based recognizer to parse and recognize user utterances and gestures by comparing verb and noun parts with sample utterances, enabling dynamic learning and execution of user commands through interactive machine learning, capable of processing anaphora and deixis in real-time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If rule-based approaches are used for multimodal command recognition, then the system structure is simple and easy to implement, but the system becomes static and brittle with poor transcription quality for terse utterances

Engineering Contradiction:
Improveease of implementationVSAvoidtranscription quality
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent transitions from static rule-based systems to dynamic machine learning models that can adapt and learn from data. The system uses trained models to parse utterances into verb and noun parts, enabling dynamic adjustment to different speech patterns and improving transcription accuracy for terse utterances without requiring complex manual rule configuration.

Inventive Principle:
Principle #15Dynamics

2Ease of manufacture

If rule-based systems are used for anaphora resolution, then the implementation is straightforward, but the system fails to handle ambiguous references and pronouns effectively

Engineering Contradiction:
Improveease of implementationVSAvoidanaphora resolution capability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent introduces machine learning models as intermediaries between raw speech input and command execution. These models serve as mediators that can interpret ambiguous references and pronouns by learning contextual patterns, effectively resolving anaphora without requiring explicit rule-based definitions for every possible scenario.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If accurate transcription of terse utterances is achieved through improved rules, then command recognition quality improves, but the system remains brittle and failure-prone

Engineering Contradiction:
Improvetranscription accuracyVSAvoidsystem robustness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the fundamental parameters of the recognition system by transitioning from rule-based logic to machine learning models. This allows the system to achieve both high transcription accuracy and robustness by learning from diverse training data, enabling it to handle variations in speech patterns, accents, and utterance styles without becoming brittle.

Inventive Principle:
Principle #35Parameter changes

4Adaptability or versatility

If a learning-based approach is implemented for multimodal command recognition, then the system becomes dynamic and capable of resolving anaphora, but the device complexity increases

Engineering Contradiction:
Improvelearning capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the complex recognition task into distinct components: utterance parsing into verb and noun parts, gesture recognition, and command execution. This segmentation allows the use of specialized machine learning models for each component, managing overall system complexity while maintaining learning capabilities for resolving anaphora and improving adaptability.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10656909B2Learning intended user actions
Publication Date: 2020.05.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10656909B2 patent drawing
  • US10656909B2 patent drawing
  • US10656909B2 patent drawing

AI summary

A method and system are provided. The method includes receiving, by a microphone and camera, user utterances indicative of user commands and associated user gestures for the user utterances. The method further includes parsing, by a hardware-based recognizer, sample utterances and the user utterances into verb parts and noun parts. The method also includes recognizing, by a hardware-based recognizer, the user utterances and the associated user gestures based on the sample utterances and descriptions of associated supporting gestures for the sample utterances. The recognizing step includes comparing the verb parts and the noun parts from the user utterances individually and as pairs to the verb parts and the noun parts of the sample utterances. The method additionally includes selectively performing a given one of the user commands responsive to a recognition result.