Screen Reading Engine for Multimodal Graphic Content

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multimodal content, which combines text and images, is difficult to translate into non-visual formats accessible to blind and visually impaired individuals due to its complex composition and interdependence of elements, making it challenging to convey the full context and narrative of in-text graphics.

Innovation Solution

A screen reading system that includes a processor-executable engine and an accessible content database, allowing users to interact with multimodal document content through different modes: Global Narrative Mode, Narrative Grammar Mode, and Free Exploration Mode, using audio files and haptic feedback to provide a sequential and location-specific presentation of text and images, guiding users through the correct panel sequencing and enabling detailed exploration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If multimodal content uses multiple visual methods of communication simultaneously (text, pictures, relationships), then the information sharing capability is enhanced, but the difficulty to translate to non-visual accessible forms increases

Engineering Contradiction:
Improveinformation sharing capabilityVSAvoidtranslation complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments multimodal content into distinct panels with defined locations, allowing each panel to be processed and described separately. This segmentation enables the complex multimodal content to be broken down into manageable units that can be translated to non-visual forms individually, reducing the overall translation complexity while preserving information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that receives multimodal content, analyzes the relationships between elements, and generates structured audio descriptions. This intermediary system acts as a mediator between the visual multimodal content and non-visual accessible forms, managing the complexity of translation through systematic processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If comic books use panels of text and images organized in sequential order, then the narrative communication is enhanced, but the accessibility translation difficulty increases due to page and panel composition complexity

Engineering Contradiction:
Improvenarrative communication effectivenessVSAvoidpage and panel composition complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-defining panel locations and establishing narrative sequencing before the actual translation process. This preliminary structuring of panel relationships and sequential order enables the translation system to follow a predetermined path, reducing the complexity of translating the composition while maintaining narrative effectiveness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transitions the two-dimensional visual panel layout into a temporal dimension through sequential audio playback. By adding the time dimension to the translation process, the system preserves the narrative structure and panel relationships without requiring direct translation of the spatial composition, thereby reducing translation complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of operation

If a screen reading system provides multiple interaction modes (global narrative, sequential narrative, location-specific), then the accessibility and user control are improved, but the system complexity increases

Engineering Contradiction:
Improveuser control and accessibilityVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system implements dynamic interaction modes that can switch between global narrative, sequential narrative, and location-specific descriptions based on user input. This dynamic capability allows the system to adapt its complexity level to user needs, providing detailed control when required while offering simpler global narratives when appropriate, thereby improving ease of operation without permanently increasing system complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12061777B2Systems and methods for conveying multimoldal graphic content
Publication Date: 2024.08.13 THE WICHITA STATE UNIV
  • US12061777B2 patent drawing
  • US12061777B2 patent drawing

AI summary

A screen reading system for providing accessibility to multimodal document content includes a memory storing: a document file defining a displayable document page containing text and images presented in panels to define a graphic narrative; an accessible content database stores audio files for the text and images, defined page locations for the audio files, panel locations for panels, and narrative sequencing for at least some of the audio files and the panels; and a processor-executable screen reading engine that when executed by a processor displays the document page on a display, receives user input to the display, and plays the audio files in response to the user input. The screen reading engine facilitates multiple modes of screen reading and can provide haptic feedback.