Invisible Junctions for Document Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods fail to effectively identify and retrieve the corresponding electronic document from a printed document using low-quality images, especially due to challenges with text recognition, language limitations, and varying viewing conditions, leading to inefficiencies in bridging the gap between physical and virtual document worlds.

Innovation Solution

The system employs 'invisible junctions' – unique local features based on the intrinsic skeleton of document pages – to match low-quality images with electronic documents, providing stability and discrimination across different languages and conditions, enabling fast and accurate retrieval of the corresponding electronic document, page, and viewing region.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If SIFT key points are used for document recognition, then local features can be extracted, but the method is not suitable for text documents and is not stable in noisy environments

Engineering Contradiction:
Improverecognition accuracyVSAvoidstability in noisy environments
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the document page into an intrinsic skeleton structure that captures the hierarchical layout of text blocks, paragraphs, and lines. This segmentation transforms the continuous image into a discrete structural representation that is invariant to noise and photometric changes, enabling reliable recognition in noisy environments while maintaining high accuracy for text documents.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intrinsic skeleton as an intermediary representation between the raw image and the recognition system. This skeleton acts as a mediator that filters out noise and photometric variations while preserving the essential structural information, thereby improving both stability in noisy environments and recognition accuracy for text documents.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If prior art recognition methods are used, then document pages can be matched, but they return large number of candidate matches with no ranking or ranking that provides too many false positive matches

Engineering Contradiction:
Improverecognition speedVSAvoiddiscrimination capability
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent employs a feedback mechanism where the intrinsic skeleton matching results are iteratively refined by comparing hierarchical structural elements. The system provides feedback by ranking candidate matches based on the degree of skeleton correspondence, eliminating false positives while maintaining high recognition speed through efficient hierarchical comparison.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies local quality by focusing on discriminative local features within the intrinsic skeleton, such as junction points and branch patterns, rather than processing the entire document uniformly. This localized approach improves discrimination capability by highlighting unique structural characteristics while maintaining overall recognition efficiency.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If geometric features of text block are used, then layout information can be captured, but the method is not suitable for Asian or ideographic languages

Engineering Contradiction:
Improvelanguage compatibilityVSAvoidlayout feature extraction
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent creates a universal intrinsic skeleton representation that functions across different language types including Asian and ideographic languages. The skeleton extraction method is language-agnostic, capturing the hierarchical structure of any text layout, thereby achieving both broad language compatibility and precise layout feature extraction through its multi-functional design.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Ease of operation

If low quality images from hand-held cameras are used, then document access becomes portable, but the input image quality is insufficient for electronic document retrieval

Engineering Contradiction:
ImproveportabilityVSAvoidimage quality
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent extracts the essential intrinsic skeleton structure from low-quality images, separating the critical structural information from the noisy visual data. This extraction process removes photometric variations and noise while retaining the document's hierarchical layout, enabling accurate retrieval from portable low-quality images without sacrificing measurement precision.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a simplified copy of the document structure through its intrinsic skeleton representation. This skeletal copy captures the essential layout information needed for retrieval while being invariant to image quality issues, allowing portable access with low-quality images to achieve the same retrieval accuracy as high-quality scans.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP2015166B1Recognition and tracking using invisible junctions
Publication Date: 2018.09.05 RICOH CO LTD
  • EP2015166B1 patent drawingFigure 1
  • EP2015166B1 patent drawingFigure 2
  • EP2015166B1 patent drawingFigure 3

AI summary

The present invention uses invisible junctions which are a set of local features unique to every page of the electronic document to match the captured image to a part of an electronic document. The present invention includes: an image capture devise, a feature extraction and recognition system and database. When an electronic document is printed, the feature extraction and recognition system captures an image of the document page. The features in the captured image are then extracted, indexed and stored m the database. Given a query image, usually a small patch of some document page captured by a low resolution image capture device, the features in the query image are extracted and compared against those stored in the database to identify the query image. The present invention also includes methods for recognizing and tracking the viewing region and look at point corresponding to the input query image. This a information is combined with a rendering of the original input document to generate a new graphical user interface to the user. This user interface can be displayed on a conventional browser or even on the display of an image capture device.