Invisible Junctions for Document Image Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods fail to effectively identify and retrieve electronic documents from printed documents using low-quality images, especially due to challenges with text recognition, language limitations, and integration with both western and eastern languages, as well as text and image combinations, and lack effective interfaces for bridging the physical and virtual worlds.

Innovation Solution

A system utilizing 'invisible junctions' for image-based document patch recognition, which includes an image capture device, feature extraction and recognition system, and database, capable of extracting and indexing unique local features from printed documents to match low-resolution images, providing fast and accurate retrieval across various languages and image types, and generating a graphical user interface for navigation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If SIFT key points are used for document recognition, then local features can be extracted, but the method is not suitable for text documents and performs poorly in noisy environments

Engineering Contradiction:
Improvefeature recognition accuracyVSAvoidstability in noisy environments
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies local quality by making different parts of the document have different functional properties. Text regions use skeleton-based junction features while image regions use traditional feature detection. This allows each region to be processed with the most appropriate method for its characteristics, resolving the contradiction between text suitability and noise robustness.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses composite materials by combining multiple feature extraction approaches into a unified system. It integrates skeleton-based features for text with traditional image-based features for graphics, creating a composite feature representation that leverages the strengths of both methods while mitigating their individual weaknesses.

Inventive Principle:
Principle #40Composite materials

2Adaptability or versatility

If geometric layout features are used for recognition, then western language documents can be processed, but the method is not suitable for Asian or ideographic languages

Engineering Contradiction:
Improvelanguage compatibilityVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies universality by creating a multi-functional feature extraction system that can handle both western and Asian languages. The skeleton-based junction detection method works universally across different writing systems and languages, making the recognition system adaptable to diverse linguistic contexts while maintaining high precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If traditional recognition methods are used, then document pages can be identified, but there is no method for indicating the viewing region and camera look-at point on the electronic document

Engineering Contradiction:
Improvespatial information retrievalVSAvoiduser interface capability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent uses an intermediary approach by introducing a coordinate transformation system that maps between camera space and document space. This intermediary mechanism enables the system to not only identify document pages but also to indicate viewing regions and camera look-at points, bridging the gap between physical capture and digital representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If paper documents are linked to the virtual world, then electronic document access is enabled, but there are no effective interfaces for working with both paper and electronic documents simultaneously

Engineering Contradiction:
Improvemixed media integrationVSAvoidinterface complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies merging by combining paper document interaction with electronic document manipulation into a unified interface. The system merges the physical paper document workspace with the virtual electronic document environment, allowing users to work with both media types simultaneously through an integrated interface that reduces rather than increases complexity.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP2015227B8User interface for three-dimensional navigation
Publication Date: 2019.09.18 RICOH CO LTD

AI summary

The present invention uses invisible junctions which are a set of local features unique to every page of the electronic document to match the captured image to a part of an electronic document. The present invention includes: an image capture device, a feature extraction and recognition system and database. When an electronic document is printed, the feature extraction and recognition system captures an image of the document page. The features in the captured image are then extracted, indexed and stored in the database. Given a query image, usually a small patch of some document page captured by a low resolution image capture device, the features in the query image are extracted and compared against those stored in the database to identify the query image. The present invention also includes methods for recognizing and tracking the viewing region and look at point corresponding to the input query image. This information is combined with a rendering of the original input document to generate a new graphical user interface to the user. This user interface can be displayed on a conventional browser or even on the display of an image capture device.