OCR Series Image Median String Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing optical character recognition (OCR) systems face challenges in accurately processing series of images with varying image quality, distortion, and geometry inconsistencies due to defects like noise, poor focus, and different imaging conditions, leading to uncertainties in recognizing symbols and determining the median string geometry.

Innovation Solution

The method involves receiving a series of images, performing optical symbol recognition, associating symbol sequences using coordinate transformations, merging sequences to identify a median string, calculating transformation of symbol sequence quadrangles, determining distances, and displaying OCR text using a median symbol sequence quadrangle to represent the original document, thereby improving accuracy by normalizing and weighting distances and applying projective transformations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If OCR is performed on individual images with varying quality and geometry, then processing speed is maintained, but recognition accuracy deteriorates due to noise, poor focus, and distortion

Engineering Contradiction:
Improvesymbol recognition accuracyVSAvoidimage processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The method segments the document processing into multiple image captures and processes each image separately through OCR, then combines the results. This allows individual images to be processed independently while maintaining overall accuracy through aggregation of multiple observations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges OCR results from multiple images by associating symbol sequences across images using coordinate transformations and identifying median strings. This combining process consolidates information from multiple sources to overcome individual image deficiencies.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If multiple images are processed to improve accuracy, then recognition reliability improves, but processing time increases

Engineering Contradiction:
ImproveOCR result reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by capturing multiple images beforehand and performing OCR on each image before final result consolidation. This preliminary processing of multiple images enables reliable median string identification while organizing work in advance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The method creates multiple copies of the document view through sequential image captures and processes each copy independently. By working with copies rather than requiring perfect single images, the system achieves reliability without excessive time penalty through parallelizable processing.

Inventive Principle:
Principle #26Copying

3Stability of the object's composition

If coordinate transformations are applied to associate symbol sequences across images, then geometry consistency improves, but computational complexity increases

Engineering Contradiction:
Improvegeometry consistencyVSAvoidcoordinate transformation complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent introduces coordinate transformations as an intermediary mechanism to bridge symbol sequences from different images. This mediator enables association of corresponding symbols across images by mapping their coordinates, achieving geometry consistency without direct complex comparison of all image data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10977511B2Optical character recognition of series of images
Publication Date: 2021.04.13 ABBYY DEVELOPMENT INC
  • US10977511B2 patent drawing
  • US10977511B2 patent drawing
  • US10977511B2 patent drawing

AI summary

Systems and methods for performing OCR of a series of images depicting text symbols. An example method comprises performing OCR a series of images to produce a current symbol sequence and corresponding symbol sequence quadrangle; associating the current symbol sequence with a previous symbol sequence for a previously received image; identifying a median string; determining a median symbol sequence quadrangle; and displaying, using the median symbol sequence quadrangle, a resulting OCR text representing at least a portion of the original document.