Multi-modal Input Fusion for Intelligent App Command Execution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current electronic devices can only support one input method (voice or touch) for interacting with intelligent apps, leading to longer utterance times and potential misinterpretation of user commands when voice input is used without essential parameters.
Innovation Solution
A system and method that combines a microphone and touchscreen display with a processor to receive user utterances, display a user interface, and accept touch or gesture inputs, allowing the identification of items and parameters to provide a response based on user intent and input methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If only voice input method is supported for intelligent apps, then the interaction can be hands-free and convenient, but the utterance time becomes longer and the utterance content becomes complicated
Solution Approach 1:
The patent combines voice input and touch input methods into a unified interface for intelligent apps. The system detects both voice commands and touch gestures (such as taps, slides, or clicks on the touchscreen) and processes them together to execute app operations. This merging allows users to use whichever input method is more convenient for each specific task, reducing the time needed for utterance while maintaining hands-free convenience when needed.
2Device complexity
If only one input method (voice or touch) is activated, then the system is simpler to implement, but the other input method is inactivated causing inconvenience to the user
Solution Approach 1:
The patent implements a universal input interface that supports multiple input methods simultaneously. The system includes a voice recognition module for processing voice commands and a touch input module for processing touchscreen gestures. Both modules operate independently but converge in the command execution system, allowing the interface to adapt to different user preferences and task requirements. This multi-functionality ensures that users can switch between voice and touch inputs without any inconvenience, while the system remains manageable in complexity through modular design.
3Loss of time
If voice command is used without essential parameters, then the input process is faster, but the electronic device requires additional input or performs operation different from user intent
Solution Approach 1:
The patent introduces an intermediary processing layer that receives both voice commands and touch input, then integrates them to determine the final user intent. When a voice command is given without essential parameters (such as 'open the map app'), the system detects accompanying touch gestures (such as tapping on a specific map location or selecting a destination from a list) and uses these as supplementary information to accurately interpret the user's intended action. This intermediary integration ensures that commands are executed with high accuracy while maintaining fast input speed, as the system processes multiple input streams in parallel rather than sequentially.
Data Source
AI summary
A system is provided. The system includes a microphone, a touchscreen display, at least one processor operatively connected to the microphone and the display, at least one memory operatively connected to the processor. The memory stores instructions that, when executed, cause the processor to receive a user utterance via the microphone, to display a user interface (UI) on the display, to receive a touch or gesture input associated with the UI via the display, to identify at least one item associated with the user interface, based at least partly on the touch or gesture input, to identify an intent based at least partly on the user utterance, to identify at least one parameter using at least part of the at least one item, and to provide a response, based at least partly on the intent and the at least one parameter.


