Multi-modal Input Fusion for Intelligent App Command Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current electronic devices can only support one input method (voice or touch) for interacting with intelligent apps, leading to longer utterance times and potential misinterpretation of user commands when voice input is used without essential parameters.

Innovation Solution

A system and method that combines a microphone and touchscreen display with a processor to receive user utterances, display a user interface, and accept touch or gesture inputs, allowing the identification of items and parameters to provide a response based on user intent and input methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If only voice input method is supported for intelligent apps, then the interaction can be hands-free and convenient, but the utterance time becomes longer and the utterance content becomes complicated

Engineering Contradiction:
Improvehands-free convenienceVSAvoidutterance time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent combines voice input and touch input methods into a unified interface for intelligent apps. The system detects both voice commands and touch gestures (such as taps, slides, or clicks on the touchscreen) and processes them together to execute app operations. This merging allows users to use whichever input method is more convenient for each specific task, reducing the time needed for utterance while maintaining hands-free convenience when needed.

Inventive Principle:
Principle #5Merging (Combining)

2Device complexity

If only one input method (voice or touch) is activated, then the system is simpler to implement, but the other input method is inactivated causing inconvenience to the user

Engineering Contradiction:
Improveinput method supportVSAvoiduser convenience
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The patent implements a universal input interface that supports multiple input methods simultaneously. The system includes a voice recognition module for processing voice commands and a touch input module for processing touchscreen gestures. Both modules operate independently but converge in the command execution system, allowing the interface to adapt to different user preferences and task requirements. This multi-functionality ensures that users can switch between voice and touch inputs without any inconvenience, while the system remains manageable in complexity through modular design.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of time

If voice command is used without essential parameters, then the input process is faster, but the electronic device requires additional input or performs operation different from user intent

Engineering Contradiction:
Improveinput speedVSAvoidcommand accuracy
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent introduces an intermediary processing layer that receives both voice commands and touch input, then integrates them to determine the final user intent. When a voice command is given without essential parameters (such as 'open the map app'), the system detects accompanying touch gestures (such as tapping on a specific map location or selecting a destination from a list) and uses these as supplementary information to accurately interpret the user's intended action. This intermediary integration ensures that commands are executed with high accuracy while maintaining fast input speed, as the system processes multiple input streams in parallel rather than sequentially.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11144175B2Rule based application execution using multi-modal inputs
Publication Date: 2021.10.12 SAMSUNG ELECTRONICS CO LTD
  • US11144175B2 patent drawing
  • US11144175B2 patent drawing
  • US11144175B2 patent drawing

AI summary

A system is provided. The system includes a microphone, a touchscreen display, at least one processor operatively connected to the microphone and the display, at least one memory operatively connected to the processor. The memory stores instructions that, when executed, cause the processor to receive a user utterance via the microphone, to display a user interface (UI) on the display, to receive a touch or gesture input associated with the UI via the display, to identify at least one item associated with the user interface, based at least partly on the touch or gesture input, to identify an intent based at least partly on the user utterance, to identify at least one parameter using at least part of the at least one item, and to provide a response, based at least partly on the intent and the at least one parameter.