OCR Text Recognition Accuracy via Contextual Meaning Disambiguation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Optical Character Recognition (OCR) processes often produce inaccurate text from images due to issues like lack of focus, contrast, and incomplete textual strings, leading to insufficient information for further processing or user action.

Innovation Solution

The system enhances OCR by identifying significant words or phrases in images, using topical information and database entries to determine meanings, and retrieving additional relevant data, such as author or publication information, to provide a more immersive experience for users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional OCR processes are used to identify text in images, then text recognition can be performed, but the accuracy and completeness of the recognized text is insufficient due to image quality issues

Engineering Contradiction:
ImproveOCR text recognition accuracyVSAvoidtext data completeness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system uses feedback mechanisms where the processed image and recognized text are fed back into the system to generate corrections and refinements. The processor identifies errors in the initially recognized text by comparing against the image data and generates corrected text output, continuously improving accuracy through iterative feedback loops.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

An intermediary correction mechanism is introduced between the OCR processing and final output. The system includes a correction module that acts as an intermediary, analyzing the relationship between image data and recognized text to identify and correct errors before the final text output is generated, thereby improving both accuracy and completeness.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If additional information is retrieved and displayed for selected words or phrases, then user interaction and information accessibility are improved, but system complexity and processing time increase

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidsystem processing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system extracts and isolates specific words or phrases from the recognized text for which additional information needs to be retrieved. By taking out only the relevant portions rather than processing the entire text, the system reduces unnecessary processing complexity while maintaining the ability to provide enhanced user interaction and information accessibility where needed.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The information retrieval and display process is segmented into discrete units corresponding to individual words or phrases. The system processes and displays additional information for each selected unit separately, which divides the complex task into manageable segments, reducing overall system complexity while preserving ease of operation for users.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If the system processes and displays additional information for recognized text, then the usefulness of OCR output is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveinformation usefulnessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system applies partial action by processing and retrieving additional information only for selected words or phrases rather than the entire text. This selective approach provides sufficient information usefulness for user needs while significantly reducing processing time and computational resources compared to processing all text content.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary identification and selection of words or phrases that are most likely to benefit from additional information retrieval. By pre-selecting and prioritizing relevant terms before executing the full information retrieval process, the system minimizes unnecessary processing time while ensuring that the most useful information is obtained.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10216989B1Providing additional information for text in an image
Publication Date: 2019.02.26 AMAZON TECH INC
  • US10216989B1 patent drawing
  • US10216989B1 patent drawing
  • US10216989B1 patent drawing

AI summary

Disclosed are techniques for providing additional information for text in an image. In some implementations, a computing device receives an image including text. Optical character recognition (OCR) is performed on the image to produce recognized text. A word or a phrase is selected from the recognized text for providing additional information. One or more potential meanings of the selected word or phrase are determined. One of the potential meanings is selected based on other text in the image. A source of additional information corresponding to the selected meaning is selected for providing the additional information to a user's device.