Multi-Modal Communication System for Transaction Interfaces
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current user interfaces for electronic transactions, such as web pages and IVR systems, are limited to a single modality, making interactions cumbersome and inefficient, particularly for tasks like airline reservations that require multiple selections.
Innovation Solution
A multi-modal communication system that allows users to interact using both mechanical motion and audio input/output modalities simultaneously, enabling the use of voice commands to input information and visual displays for output, with a voice server and data server facilitating the conversion and processing of these modalities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a single modality interface (web page or IVR) is used, then the interface is simple to implement, but the user interaction efficiency is poor
Solution Approach 1:
The patent combines multiple input modalities (voice recognition, keyboard, mouse) and multiple output modalities (audio output, visual display) into a unified interface system. The system integrates these different communication channels to work together, allowing users to interact through their preferred modality while maintaining system coherence and managing complexity through unified architecture.
Solution Approach 2:
The interface is designed to support multiple functions across different modalities simultaneously. A single interface can process both voice commands and keyboard input, display visual information and provide audio feedback, making it universally adaptable to various user preferences and task requirements without requiring separate specialized systems.
2Ease of operation
If mechanical motion input is used for all selections, then the interface is consistent, but the ease of operation deteriorates for complex tasks
Solution Approach 1:
The patent replaces mechanical motion input (keyboard, mouse) with voice recognition input for certain tasks. The voice recognition system captures spoken commands and converts them into actionable inputs, eliminating the need for repeated mechanical actions such as navigating through multiple drop-down menus, thereby reducing both effort and time for complex selection tasks.
3Adaptability or versatility
If visual output only is used, then the interface is simple to implement, but the adaptability to different user needs is limited
Solution Approach 1:
The system merges visual display output with audio output capabilities into a unified interface architecture. This integration allows the system to adapt to different user needs by providing information through appropriate modalities while maintaining system coherence through unified design and management.
Solution Approach 2:
The interface is designed with universal output capabilities that can simultaneously provide both visual and audio information. This multi-functional output system adapts to different user preferences, accessibility needs, and task contexts without requiring separate specialized output systems.
4Ease of operation
If IVR with audio input/output is used, then accessibility is improved, but the productivity for complex tasks deteriorates
Solution Approach 1:
The system merges the accessibility benefits of audio-based IVR interaction with the efficiency of visual display and multiple input modalities. By integrating voice recognition with visual feedback and combining it with keyboard/mouse input capabilities, the system maintains high accessibility for users who prefer or require audio interaction while enabling faster task completion through multi-modal input options.
Data Source
AI summary
Systems and methods for handling dual modality communication between at least one user device and at least one server. The modalities comprise audio modalities and mechanical motion modalities. The server may be simultaneously connected to the user device via a data network and a voice network and simultaneously receive audio-based input and mechanical motion-based input.


