GUI Language Model Audio Output for Visually Impaired
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computer systems with graphical user interfaces (GUIs) pose challenges for visually impaired users due to their visual nature, making it difficult to create effective audio output analogs that are easy to use and efficient.
Innovation Solution
Converting the directed graph representation of a GUI into a formatted text string using a user-appropriate language model, which is then parsed and output as an audio signal, allowing for concise and intuitive navigation and interaction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If graphical user interfaces are used to recreate computer experience mirroring visual appearance, then user interface intuitiveness is improved, but accessibility for visually impaired users deteriorates
Solution Approach 1:
The patent substitutes visual mechanical interaction (GUI elements requiring sight) with acoustic interaction (audio descriptions and speech recognition). The system converts visual interface elements into audio representations that can be perceived and manipulated through sound, replacing the visual-mechanical paradigm with an acoustic one that is accessible to visually impaired users.
Solution Approach 2:
The patent introduces speech recognition technology and text-to-speech conversion as intermediary layers between the user and the computer system. These intermediaries translate visual interface concepts into audio formats, allowing visually impaired users to interact with the system through voice commands and audio feedback rather than direct visual engagement with GUI elements.
2Loss of information
If detailed GUI descriptions are provided to visually impaired users, then information completeness is improved, but complexity of audio output increases
Solution Approach 1:
The patent segments GUI information into hierarchical levels of detail, providing comprehensive information only when needed while offering simplified summaries for routine interactions. The system divides complex interface descriptions into manageable audio segments that can be presented on-demand, allowing users to access detailed information about specific elements without being overwhelmed by complete system descriptions.
Solution Approach 2:
The patent implements partial action by providing audio descriptions at appropriate levels of detail based on context and user needs. Rather than describing every GUI element continuously, the system provides partial descriptions focused on currently relevant elements, reducing overall audio complexity while maintaining information completeness for tasks requiring it.
Data Source
AI summary
Method and apparatus for allowing visually impaired users to easily interact with GUI applications is provided. The method and apparatus may utilize a directed graph of the GUI and a language model to describe the GUI in a brief but concise and descriptive manner.


