Character Recognition Disambiguation via Contextual Nesting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Optical character recognition (OCR) systems often struggle to accurately identify suspect characters due to lack of context, leading to potential errors in recognition, as human operators may misinterpret characters in a visual carpet format without understanding their original context.

Innovation Solution

The method involves selecting characters based on confidence levels and fragmentation, displaying them with context information in a grid or carpet format to allow human operators to select suspect characters accurately, using neighboring characters and varying context data to enhance recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If characters are displayed in a visual carpet format for easy identification, then ease of operation is improved, but loss of information occurs because context information is lost

Engineering Contradiction:
Improveease of identificationVSAvoidcontext information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent implements a nested display structure where character images are displayed within their contextual environment. The character carpet format is nested within the original document layout, allowing operators to see both the extracted characters and their surrounding text context simultaneously. This nested presentation resolves the contradiction by preserving context information while maintaining the efficient carpet-style verification interface.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Measurement precision

If all suspect characters are displayed for review, then measurement precision is improved, but loss of time increases due to reviewing more characters

Engineering Contradiction:
Improverecognition accuracyVSAvoidreview time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the character verification process by grouping suspect characters into distinct visual clusters based on their recognition confidence levels and similarity metrics. High-confidence characters are separated from low-confidence ones, and characters with similar visual patterns are clustered together. This segmentation allows operators to quickly identify and focus on the most problematic characters, improving review efficiency while maintaining high recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality enhancement by providing different levels of context information for different character groups. Characters that are more ambiguous or have lower confidence scores receive more detailed contextual display (showing more surrounding text), while high-confidence characters receive minimal context. This differentiated approach reduces overall review time while ensuring thorough verification of problematic characters.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9008428B2Efficient verification or disambiguation of character recognition results
Publication Date: 2015.04.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US9008428B2 patent drawing
  • US9008428B2 patent drawing
  • US9008428B2 patent drawing

AI summary

Machines, systems and methods for character recognition disambiguation are provided. The method comprises selecting a first set of characters that match a first visual profile based on results of a character recognition process applied to target content; selecting a subset of the first set based on criteria associated with at least one of confidence level with which characters grouped in the subset are recognized or fragmentation associated with the characters grouped in the subset; and disambiguating recognition results for the characters grouped in the subset by displaying the characters along with context information, wherein reviewing two or more of the characters on a display screen along with context information associated with said two or more characters allows a human operator to select one or more suspect characters from among the two or more characters.