Automated Voice Interface Generation for Small Display Usability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing user interfaces for electronic devices face challenges in transitioning between graphical user interfaces (GUIs) and voice user interfaces (VUIs, particularly in complex interactions and on devices with small displays, where speech recognition is difficult and input methods like virtual keyboards are cumbersome.
Innovation Solution
A method that identifies and aggregates GUI screens, analyzes navigational flow paths, and generates data structures to create a voice user interface (VUI) by determining select and prompt objects using natural language processing, allowing users to interact with electronic devices through voice commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice user interface (VUI) is implemented for user interaction, then hands-free operation and simplified interaction are achieved, but speech recognition accuracy deteriorates in complex interactions
Solution Approach 1:
The patent segments the user interface into multiple modalities (GUI and VUI) that can be used independently or in combination. The system divides complex interactions into smaller units that can be handled by appropriate interface elements, with GUI objects providing visual confirmation and VUI handling voice commands, thereby maintaining both hands-free operation and recognition accuracy.
Solution Approach 2:
The patent introduces an intermediary layer that translates between GUI and VUI modalities. This intermediary processes voice commands, matches them with corresponding GUI objects, and provides visual feedback, thereby improving speech recognition accuracy while maintaining hands-free operation capability.
2Measurement precision
If graphical user interface (GUI) with visual indicators is used, then precise visual feedback is provided, but usability deteriorates on devices with small displays
Solution Approach 1:
The patent segments the interface elements into hierarchical levels, with main navigation options presented as larger, more accessible elements and detailed information available upon selection. This segmentation allows precise visual feedback to be provided without requiring all elements to be simultaneously visible on small displays.
Solution Approach 2:
The patent utilizes temporal dimension by implementing progressive disclosure, where visual feedback is provided in stages as users navigate through the interface. This allows comprehensive visual information to be conveyed on small displays without requiring all elements to be present simultaneously, thereby improving usability while maintaining feedback precision.
3Quantity of substance
If virtual keyboard is used for input on mobile devices, then text input capability is provided, but ease of operation deteriorates due to small display size
Solution Approach 1:
The patent introduces voice recognition as an intermediary input method that translates spoken text into digital input. This intermediary system handles text input capability without requiring users to manually type on small virtual keyboards, thereby maintaining text input functionality while dramatically improving ease of operation on mobile devices.
Solution Approach 2:
The patent replaces the mechanical interaction of typing on a virtual keyboard with acoustic interaction through voice commands. This substitution eliminates the need for precise manual manipulation on small displays while maintaining full text input capability, thereby resolving the contradiction between input capability and ease of operation.
Data Source
AI summary
Techniques are disclosed for generating a voice user interface (VUI) modality within an application that includes graphical user interface (GUI) screens. A GUI screen parser analyzes the GUI screens to determine the various navigational GUI screen paths that are associated with edge objects within multiple GUI screens. Some edge objects are identified as select objects or prompt objects. A natural language processing system generates a select object synonym data structure and a prompt object data structure that may be utilized by a VUI generator to generate VUI data structures that give the application VUI modality.


