Graphical Artifact Accessibility via Multi-Modal Semantic Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies are limited in providing non-visual access to graphical information, especially mathematical and scientific graphs, due to a lack of datasets and appropriate semantic information, leading to inadequate accessibility and understanding of such content.
Innovation Solution
A system utilizing deep learning models, including convolutional neural networks and vision-language transformer-based models, to classify and extract semantic information from graphical artifacts, converting it into accessible forms like braille, audio, or haptic representations for delivery.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing AI models are used to extract information from graphical artifacts, then textual and non-textual information can be combined for recognition and captioning, but the models fail to accurately understand mathematical and scientific graphs due to lack of appropriate datasets and semantic information
Solution Approach 1:
The system segments the graphical artifact into distinct components: visual elements (graphs, charts, diagrams) and textual elements (captions, labels, annotations). Each component is processed separately through specialized models - vision-language transformers for visual understanding and text processing models for textual information extraction. This segmentation allows each component to be analyzed with appropriate semantic frameworks, resolving the issue of lost mathematical and scientific information.
Solution Approach 2:
The system introduces an intermediary processing layer that bridges visual and textual information. A unified model integrates the extracted visual features and textual components, creating a comprehensive representation that preserves mathematical and scientific semantics. This intermediary layer translates graphical artifacts into structured data formats that maintain semantic integrity, preventing information loss during the conversion process.
2Ease of operation
If haptic and auditory feedback technologies are used for non-visual access, then individuals can access graphical content, but the feedback is limited to aesthetic features only and does not convey mathematical or scientific semantics
Solution Approach 1:
The system creates a universal accessibility layer that serves multiple functions simultaneously. The same processing pipeline handles various graphical artifact types (graphs, charts, diagrams) and converts them into multiple accessible formats (haptic feedback, auditory descriptions, textual summaries). This multi-functional approach ensures that mathematical and scientific semantics are preserved across all output modalities, not just aesthetic features.
Solution Approach 2:
The system replaces direct visual perception with alternative sensing mechanisms. Instead of requiring visual processing, the system substitutes visual information extraction with computational analysis that produces haptic patterns, auditory signals, and textual representations. This substitution maintains semantic integrity by processing mathematical and scientific information through algorithms rather than mechanical visual interpretation.
3Reliability
If deep learning models are deployed to extract semantic information from graphical artifacts, then accessibility can be improved, but the system complexity increases due to multiple models and processing stages
Solution Approach 1:
The system merges multiple deep learning models into a coordinated pipeline where each model handles a specific aspect of graphical artifact analysis. Vision-language transformers combine visual and textual processing in a unified framework, while separate models specialize in extracting mathematical and scientific semantics. This merging approach maintains high reliability through specialized processing while organizing complexity into modular, manageable stages.
Solution Approach 2:
The system performs preliminary processing actions to simplify subsequent analysis. Before the main semantic extraction, the system pre-processes graphical artifacts by detecting visual elements, extracting textual components, and organizing data into structured formats. This preliminary action reduces the complexity of subsequent processing stages and improves the efficiency of the multi-model pipeline.
Data Source
AI summary
A method for providing non-visual access to graphical artifacts available in digital content includes classifying a graphical artifact into known and/or unknown categories using a deep neural network. The method further includes identifying semantically connected visual and textual components of the graphical artifact, using a deep learning-based object detection model. Furthermore, the method includes extracting the visual and the textual components in a unified framework with predefined semantics associated with each component, using a pre-trained large multi-modal model fine-tuned to extract both the visual and the textual components from an image in the graphical artifact. The method further includes filtering out the predefined semantics through extraction and converting the predefined semantics into accessible representations. Also, the method includes delivering the accessible representations in conformance with requirements of a delivery system.


