Document Image Character Detection with Visual Correspondence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Operators face challenges in identifying and correcting character recognition results, especially when dealing with unfamiliar document formats, as they need to ascertain which item is shown and where it is located within the document.

Innovation Solution

An image-processing device and method that detects character strings in a document image using pre-learned feature amounts and outputs correspondence relations between images, assisting operators in identifying items even when the document format is unknown.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If operators manually check and correct character recognition results in unfamiliar document formats, then correction accuracy can be maintained, but the time required to ascertain item locations and identification increases significantly

Engineering Contradiction:
Improvecharacter recognition correction accuracyVSAvoidtime to ascertain item location and identification
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces a correspondence display unit that acts as an intermediary between the document image and the character recognition results. This unit displays visual correspondences (such as lines or highlights) connecting recognized character strings to their locations in the original document image, helping operators quickly identify and verify items without manually searching through unfamiliar formats. The correspondence display serves as a mediator that bridges the gap between automated recognition and human verification.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a visual copy or representation of the document layout overlaid with recognition results. The correspondence display unit generates a visual replica of the document structure with annotated positions of recognized items, allowing operators to verify results by comparing this visual copy against the original document image without having to mentally map positions in unfamiliar formats.

Inventive Principle:
Principle #26Copying

2Productivity

If automated character recognition is performed on documents with unknown formats, then processing speed increases, but the ability to accurately identify and locate items deteriorates

Engineering Contradiction:
Improvecharacter recognition processing speedVSAvoiditem location and identification information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent performs preliminary extraction and organization of recognition results before presenting them to operators. The character string detection unit pre-processes the recognition output by extracting individual character strings and their positional information, and the correspondence display unit pre-arranges this information in a visually organized manner that preserves location data. This preliminary action ensures that location information is not lost during automated processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms two-dimensional spatial information from the document image into a different dimensional representation in the correspondence display. By mapping character string positions to visual indicators in the correspondence display unit, the system preserves location information in a new dimensional format that is easier to interpret, maintaining the spatial relationships while presenting them in a more accessible manner.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11605219B2Image-processing device, image-processing method, and storage medium on which program is stored
Publication Date: 2023.03.14 NEC CORP
  • US11605219B2 patent drawing
  • US11605219B2 patent drawing
  • US11605219B2 patent drawing

AI summary

An image-processing device includes: a character string detection unit configured to detect a character string of a specific item in a first document image based on a feature amount of the displayed first document image among feature amounts which are recorded in advance based on a result of learning using a plurality of document images and indicate features of the character string of the item for each kind of document image and each specific item; and an output unit configured to output information regarding a correspondence relation indicating the same specific item between the first document image and a second document image displayed to correspond to the first document image.