Multi-Modal Voice Interaction Processing via Intent Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice user interfaces are cumbersome, requiring specific command-and-control sequences, leading to user frustration due to inaccurate speech recognition and limitations in engaging users in cooperative dialogue, and fail to integrate information across devices and applications, restricting users from accessing desired content or services seamlessly.

Innovation Solution

A natural language voice services environment that processes multi-modal interactions using a voice-click module, allowing users to interact with devices in a free-form manner by combining voice and non-voice inputs, such as voice recognition and gestures, to request content or services across multiple devices and applications, leveraging constellation models and natural language processing components for intent determination and context understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice recognition software is used to simplify user interaction, then ease of operation improves, but device complexity increases due to the need for sophisticated natural language processing and multi-modal integration systems

Engineering Contradiction:
Improveuser interaction simplicityVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces a natural language processing system as an intermediary layer between the user and the device's command structure. This mediator translates free-form speech into structured commands, shielding users from the underlying system complexity while enabling sophisticated multi-modal interactions across multiple devices and applications.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The voice user interface is designed to be universal, handling multiple types of interactions (voice, gestures, clicks) and supporting access to content across different devices and applications. This multi-functional approach consolidates various interaction modes into a single unified system, improving ease of operation without proportionally increasing perceived complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If existing voice user interfaces require specific command-and-control sequences, then measurement precision of user intent improves, but ease of operation deteriorates due to learning curves and user frustration

Engineering Contradiction:
Improveuser intent recognition accuracyVSAvoidinteraction simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system dynamically adapts its command structure based on the user's speech patterns and context. Rather than requiring rigid pre-established commands, the natural language processing system flexibly interprets varied user inputs, maintaining precision in intent recognition while significantly improving ease of operation by eliminating mandatory command sequences.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of voice interaction from fixed command structures to variable natural language parameters. The system processes speech as continuous natural language rather than discrete commands, allowing users to express intent in multiple ways while maintaining accurate interpretation through contextual analysis and multi-modal input integration.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If voice user interfaces are constrained to finite sets of applications and devices, then reliability of speech recognition improves, but adaptability deteriorates, restricting users from accessing desired content or services seamlessly

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcross-device access capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The voice user interface is designed to universally access content across multiple devices and applications simultaneously. The system maintains reliable speech recognition by processing voice inputs through a standardized natural language interpretation layer, while the underlying architecture enables flexible routing to various devices and applications, achieving both reliability and adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The natural language processing system serves as an intermediary that decouples speech recognition from device-specific applications. This mediator layer maintains consistent, reliable interpretation of user intent while enabling flexible access to diverse devices and content, resolving the contradiction between recognition reliability and cross-device adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Device complexity

If existing voice user interfaces use cumbersome interfaces with complex menus, then device complexity is reduced, but ease of operation deteriorates, preventing quick and focused interaction

Engineering Contradiction:
Improveinterface structure simplicityVSAvoidinteraction speed
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The patent replaces the mechanical navigation of complex menu structures with a natural language processing system. Instead of requiring users to physically navigate through hierarchical menus, the system substitutes this mechanical interaction with speech-based direct commands, significantly improving interaction speed while maintaining or reducing overall interface complexity through the unified natural language layer.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP2399255B1System and method for processing multi-modal device interactions in a natural language voice services environment
Publication Date: 2020.07.08 VB ASSETS LLC
  • EP2399255B1 patent drawingFigure 1
  • EP2399255B1 patent drawingFigure 2
  • EP2399255B1 patent drawingFigure 3

AI summary

A system and method for processing multi-modal device interactions in a natural language voice services environment may be provided. In particular, one or more multi-modal device interactions may be received in a natural language voice services environment that includes one or more electronic devices. The multi-modal device interactions may include a non-voice interaction with at least one of the electronic devices or an application associated therewith, and may further include a natural language utterance relating to the non-voice interaction. Context relating to the non-voice interaction and the natural language utterance may be extracted and combined to determine an intent of the multi-modal device interaction, and a request may then be routed to one or more of the electronic devices based on the determined intent of the multi-modal device interaction.