Digital Document Recognition via Text Search and Feature Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current form recognition processes are slow and inaccurate due to the need to compare digital documents with millions of standard documents, often resulting in numerous false matches and requiring manual verification.

Innovation Solution

An apparatus and method that extracts textual content from a digital document, performs a text search on an indexed master document database, generates a candidate list of matching documents, and compares features to identify the correct document, significantly reducing processing time and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional form recognition compares digital documents with millions of standard documents using layout-based or image-based comparisons, then comprehensive document matching is achieved, but processing time becomes excessively long and accuracy decreases

Engineering Contradiction:
Improvedocument identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts and compares specific textual features and structural elements from documents rather than performing comprehensive layout-based or image-based comparisons with millions of standard documents. This selective extraction of key identifying features significantly reduces processing time while maintaining identification accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the document comparison process into distinct feature extraction and comparison stages, analyzing specific textual content and structural elements separately rather than performing monolithic document-wide comparisons. This segmentation enables faster processing by focusing computational resources on discriminative features.

Inventive Principle:
Principle #1Segmentation

2Reliability

If traditional form recognition performs comprehensive comparisons with all standard documents, then all possible matches are identified, but the number of false matches increases requiring manual verification

Engineering Contradiction:
Improvematch accuracyVSAvoidmanual verification effort
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies local quality by focusing comparison efforts on specific discriminative features and textual elements that are most characteristic of document types, rather than uniformly comparing all document aspects. This targeted approach improves match accuracy while reducing false positives that would require manual verification.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs partial comparison by analyzing only the most relevant textual features and structural elements necessary for accurate document identification, rather than exhaustively comparing all possible document attributes. This partial action suffices for high-accuracy matching while eliminating unnecessary computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If feature-based comparison is performed on all extracted features of digital documents, then comprehensive document characterization is achieved, but computational complexity and processing time increase significantly

Engineering Contradiction:
Improvedocument feature analysis accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and compares only the most discriminative textual features and structural elements from documents, rather than analyzing all possible document features. This selective feature extraction maintains comprehensive document characterization accuracy while significantly reducing computational complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments feature comparison into hierarchical levels, analyzing critical identifying features first and only proceeding to more detailed feature comparison when necessary. This segmented approach achieves comprehensive characterization when needed while reducing average computational complexity across all document comparisons.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10579653B2Apparatus, method, and computer-readable medium for recognition of a digital document
Publication Date: 2020.03.03 APRYSE SOFTWARE CORP
  • US10579653B2 patent drawing
  • US10579653B2 patent drawing
  • US10579653B2 patent drawing

AI summary

Described herein are an apparatus, method, and computer-readable medium. The apparatus including processing circuitry configured to extract a textual content included within a digital document, perform a text search using the extracted textual content on an indexed master document database to identify one or more master documents that are similar, within a pre-determined threshold, to the digital document, generate a candidate master document list using the one or more master documents identified based on the text search, extract a plurality of features of the digital document, perform a comparison, after performing the text search, of the plurality of features of the digital document with features of the one or more master documents in the candidate master document, and identify a master document of the one or more master documents that matches the digital document based on the comparison of the features.