Screen Reader Image Description Contextual Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current screen reader applications poorly integrate image descriptions into surrounding text, leading to confusion for blind or visually impaired users as images are not always described in contextually relevant proximity to the related text, disrupting the user experience.

Innovation Solution

A mechanism that analyzes images and text using image recognition and natural language processing to generate a natural language description of images, then compares this description to surrounding sentences using cosine similarity and ontology mapping to determine the most relevant sentence for integration, ensuring the image description is read in close proximity to the contextually related text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If image descriptions are inserted at fixed positions in the document, then the screen reader operation is simple, but the image descriptions are not contextually relevant to the surrounding text

Engineering Contradiction:
Improvescreen reader operationVSAvoidcontextual relevance
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent applies dynamics by making the insertion position of image descriptions flexible rather than fixed. The system dynamically determines the optimal insertion position by comparing the image description with surrounding text using cosine similarity and ontology mapping, allowing the description to be placed where it is most contextually relevant while maintaining simple screen reader operation through automated position selection

Inventive Principle:
Principle #15Dynamics

2Reliability

If image descriptions are inserted at every image location, then all images are described, but the text becomes redundant and less coherent

Engineering Contradiction:
Improveimage description completenessVSAvoidtext coherence
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by making the insertion of image descriptions selective rather than universal. The system evaluates each image description against the surrounding text using cosine similarity and ontology mapping, inserting descriptions only at locations where they add contextual value. This ensures comprehensive image description coverage while maintaining text coherence by avoiding redundant insertions

Inventive Principle:
Principle #3Local quality

3Productivity

If image descriptions are inserted without analyzing text relevance, then the processing is fast, but the descriptions may be misplaced and confuse users

Engineering Contradiction:
Improveprocessing speedVSAvoiddescription placement accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by performing text relevance analysis before inserting image descriptions. The system pre-processes the document by comparing each image description with surrounding text using cosine similarity and ontology mapping to determine the optimal insertion position. This preliminary analysis ensures accurate placement while maintaining processing efficiency through automated relevance evaluation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10540445B2Intelligent integration of graphical elements into context for screen reader applications
Publication Date: 2020.01.21 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10540445B2 patent drawing
  • US10540445B2 patent drawing
  • US10540445B2 patent drawing

AI summary

A mechanism is provided for intelligently integrating descriptions of images into surrounding text for a screen reader. A natural language understanding image description is determined for an image in a document. For each sentence of a set of sentences in the text of the document, a relatedness score between the sentence and the natural language understanding image description is determined thereby forming a set of relatedness scores. A highest relatedness score is determined from the set of relatedness scores. The natural language image description is inserted in close proximity to a sentence associated with the highest relatedness score, such that, when the text is read out by the screen reader, the natural language image description of the image is read out in close proximity to the sentence.