Graphical Artifact Accessibility via Multi-Modal Semantic Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies are limited in providing non-visual access to graphical information, especially mathematical and scientific graphs, due to a lack of datasets and appropriate semantic information, leading to inadequate accessibility and understanding of such content.

Innovation Solution

A system utilizing deep learning models, including convolutional neural networks and vision-language transformer-based models, to classify and extract semantic information from graphical artifacts, converting it into accessible forms like braille, audio, or haptic representations for delivery.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing AI models are used to extract information from graphical artifacts, then textual and non-textual information can be combined for recognition and captioning, but the models fail to accurately understand mathematical and scientific graphs due to lack of appropriate datasets and semantic information

Engineering Contradiction:
Improveaccuracy of graphical artifact understandingVSAvoidloss of mathematical and scientific semantic information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system segments the graphical artifact into distinct components: visual elements (graphs, charts, diagrams) and textual elements (captions, labels, annotations). Each component is processed separately through specialized models - vision-language transformers for visual understanding and text processing models for textual information extraction. This segmentation allows each component to be analyzed with appropriate semantic frameworks, resolving the issue of lost mathematical and scientific information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary processing layer that bridges visual and textual information. A unified model integrates the extracted visual features and textual components, creating a comprehensive representation that preserves mathematical and scientific semantics. This intermediary layer translates graphical artifacts into structured data formats that maintain semantic integrity, preventing information loss during the conversion process.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If haptic and auditory feedback technologies are used for non-visual access, then individuals can access graphical content, but the feedback is limited to aesthetic features only and does not convey mathematical or scientific semantics

Engineering Contradiction:
Improveaccessibility of graphical contentVSAvoidloss of mathematical and scientific semantic information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system creates a universal accessibility layer that serves multiple functions simultaneously. The same processing pipeline handles various graphical artifact types (graphs, charts, diagrams) and converts them into multiple accessible formats (haptic feedback, auditory descriptions, textual summaries). This multi-functional approach ensures that mathematical and scientific semantics are preserved across all output modalities, not just aesthetic features.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system replaces direct visual perception with alternative sensing mechanisms. Instead of requiring visual processing, the system substitutes visual information extraction with computational analysis that produces haptic patterns, auditory signals, and textual representations. This substitution maintains semantic integrity by processing mathematical and scientific information through algorithms rather than mechanical visual interpretation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If deep learning models are deployed to extract semantic information from graphical artifacts, then accessibility can be improved, but the system complexity increases due to multiple models and processing stages

Engineering Contradiction:
Improveaccuracy of semantic information extractionVSAvoidcomplexity of multi-model processing system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges multiple deep learning models into a coordinated pipeline where each model handles a specific aspect of graphical artifact analysis. Vision-language transformers combine visual and textual processing in a unified framework, while separate models specialize in extracting mathematical and scientific semantics. This merging approach maintains high reliability through specialized processing while organizing complexity into modular, manageable stages.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs preliminary processing actions to simplify subsequent analysis. Before the main semantic extraction, the system pre-processes graphical artifacts by detecting visual elements, extracting textual components, and organizing data into structured formats. This preliminary action reduces the complexity of subsequent processing stages and improves the efficiency of the multi-model pipeline.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240362942A1Method and system for providing non-visual access to graphical artifacts available in digital content
Publication Date: 2024.10.31 UNAR LABS LLC
  • US20240362942A1 patent drawing
  • US20240362942A1 patent drawing
  • US20240362942A1 patent drawing

AI summary

A method for providing non-visual access to graphical artifacts available in digital content includes classifying a graphical artifact into known and/or unknown categories using a deep neural network. The method further includes identifying semantically connected visual and textual components of the graphical artifact, using a deep learning-based object detection model. Furthermore, the method includes extracting the visual and the textual components in a unified framework with predefined semantics associated with each component, using a pre-trained large multi-modal model fine-tuned to extract both the visual and the textual components from an image in the graphical artifact. The method further includes filtering out the predefined semantics through extraction and converting the predefined semantics into accessible representations. Also, the method includes delivering the accessible representations in conformance with requirements of a delivery system.