Screen Reader Concurrent Audio Delivery for Accessibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing screen readers present information serially, requiring users to wait for data processing, which is inefficient given the human brain's ability to process multiple simultaneous information sources.
Innovation Solution
A screen reader software that delivers textual and non-textual information simultaneously through a combination of speech synthesis and additional audio layers or tactile Braille displays, allowing users to customize the auditory interface with pitch, speed, volume, and emphasis adjustments, and using aural tags to convey non-textual information concurrently with synthesized text.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If information is presented serially through speech synthesis, then information is delivered to the user, but data processing time increases and efficiency decreases
Solution Approach 1:
The patent segments information into multiple parallel audio streams, with each stream handling different types of information (textual content, non-textual information, navigation cues). This segmentation allows the screen reader to process and deliver different information types simultaneously through separate synthesis engines, rather than serially processing all information through a single engine.
Solution Approach 2:
The patent merges multiple speech synthesis engines to operate concurrently, each engine dedicated to specific information streams. Multiple audio outputs are combined and delivered to the user simultaneously, enabling parallel information delivery that reduces overall processing time while maintaining information completeness.
2Productivity
If multiple information streams are delivered simultaneously, then data processing time is reduced, but system complexity increases
Solution Approach 1:
The system segments functionality into separate synthesis engines, each responsible for specific information streams. This modular segmentation reduces individual engine complexity while enabling parallel operation, as each engine handles a focused subset of information types rather than requiring one complex engine to handle all information sequentially.
Solution Approach 2:
The patent implements universal audio output interfaces that can handle multiple information streams through standardized protocols. The audio system is designed to accept and process multiple concurrent streams using consistent formatting and delivery mechanisms, reducing the complexity overhead that would otherwise arise from managing disparate information delivery channels.
3Adaptability or versatility
If speech synthesis is used to read textual information, then accessibility is provided, but non-textual information delivery is limited
Solution Approach 1:
The patent segments information delivery into dedicated audio streams: one stream handles textual content through speech synthesis, while separate streams handle non-textual information such as audio descriptions, navigation cues, and interface feedback. This segmentation ensures that non-textual information receives dedicated processing resources and is not lost or oversubscribed in the textual information stream.
Solution Approach 2:
The system merges multiple information types into concurrent audio streams that are delivered simultaneously to the user. Textual information, non-textual descriptions, and interface feedback are combined in a multi-channel audio output that preserves the completeness of each information type while delivering them in parallel, preventing information loss that would occur in serial delivery.
Data Source
AI summary
A screen reader software product for low-vision users, the software having a reader module collecting textual and non-textual display information generated by a web browser or word processor. Font styling, interface layout information and the like are communicated to the end user by sounds broadcast simultaneously rather than serially with the synthesized speech to improve the speed and efficiency in which information may be digested by the end user.


