Document Recognition Using Invisible Junctions and Geometric Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods fail to effectively identify and retrieve electronic documents from printed documents using low-quality images, especially due to challenges like small image portions, similar document pages, varying viewing conditions, and noise, which affects recognition accuracy and efficiency.
Innovation Solution
The system employs 'invisible junctions' as local features unique to each document page for image-based document patch recognition, using a feature extraction and recognition system that includes units for feature extraction, indexing, retrieval, and geometric estimation to match low-quality images with electronic documents, working with both western and eastern languages and text-image combinations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If low-quality images from hand-held cameras are used for document recognition, then accessibility and ease of operation improve, but recognition accuracy and measurement precision deteriorate
Solution Approach 1:
The patent extracts invisible junction features from the image data - these are geometric constraints derived from the spatial relationships between visible text elements. By separating these structural features from the noisy visual appearance, the system achieves accurate document identification even from low-quality camera images
Solution Approach 2:
The patent pre-computes and stores invisible junction features for all documents in the database during an indexing phase. When a query image is received, the system only needs to compare features rather than perform full document analysis, enabling real-time recognition despite image quality limitations
2Ease of operation
If small portions of document pages are captured, then portability and ease of operation improve, but the ability to uniquely identify documents deteriorates
Solution Approach 1:
The patent computes invisible junction features locally at each detected text element position within the captured image region. Each feature encapsulates the geometric relationship between nearby text lines and elements, creating a unique local signature that identifies the document even when only a small portion is visible
Solution Approach 2:
The patent transforms 2D image coordinates into a normalized geometric feature space that is invariant to scale, rotation, and perspective distortion. This dimensional transformation allows small image portions to be reliably matched regardless of their size or orientation in the captured frame
3Adaptability or versatility
If traditional feature extraction methods like SIFT are used, then general applicability improves, but discrimination capability and measurement precision deteriorate for text documents
Solution Approach 1:
The patent segments the document image into individual text elements (words, characters) and computes geometric features based on their relative positions. This segmentation approach creates features that are inherently discriminative for text documents, unlike SIFT which treats the entire image as a continuous signal
Solution Approach 2:
The patent replaces the gradient-based detection mechanism of SIFT with a text-structure-based geometric constraint system. Instead of detecting extrema in scale space, the system identifies invisible junctions formed by text line intersections, which are uniquely characteristic of document layouts
4Device complexity
If prior art recognition methods are used, then implementation simplicity improves, but the number of false positive matches increases
Solution Approach 1:
The patent introduces invisible junction features as an intermediary representation between the image data and the document identification process. These features serve as a discriminative filter that eliminates false positives by capturing the unique geometric structure of each document's text layout
Solution Approach 2:
The system uses geometric estimation to verify candidate matches by checking whether the computed invisible junctions are consistent with the expected document structure. This feedback mechanism filters out false positives by requiring that matched features satisfy geometric constraints
Data Source
AI summary
The present invention uses invisible junctions which are a set of local features unique to every page of the electronic document to match the captured image to a part of an electronic document. The present invention includes: an image capture device, a feature extraction and recognition system and database. When an electronic document is printed, the feature extraction and recognition system captures an image of the document page. The features in the captured image are then extracted, indexed and stored in the database. Given a query image, usually a small patch of some document page captured by a low resolution image capture device, the features in the query image are extracted and compared against those stored in the database to identify the query image. The present invention advantageously uses geometric estimation to reduce the query results to a single one or a few candidate matches. In one embodiment, the two separate geometric estimations are used to rank and verify matching candidates.


