Partial Document Recognition for Complete Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine recognition systems such as OCR and ASR are prone to errors due to imperfect scanning, processing, and flawed algorithms, making it difficult to accurately recognize and retrieve complete documents, especially when portions are damaged or of poor quality.

Innovation Solution

A system and method that uses machine recognition outputs as input for search engines to locate complete documents by scanning or recognizing portions of text or speech, allowing refinement of recognition results and enabling retrieval of entire documents from databases or networks, even when only partial information is available.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If machine recognition systems (OCR/ASR) are used to recognize text or speech from damaged or poor quality sources, then document retrieval is enabled, but recognition accuracy deteriorates due to errors from imperfect scanning, processing, and flawed algorithms

Engineering Contradiction:
Improvedocument retrieval capabilityVSAvoidrecognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary search system that bridges the gap between imperfect machine recognition outputs and complete document retrieval. The search system acts as a mediator that takes the erroneous recognized portions as input queries, searches the database for matching documents, and returns complete copies. This intermediary layer enables document retrieval functionality while circumventing the accuracy limitations of the underlying machine recognition systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by using the search results (complete document copies) to refine and correct the original machine recognition outputs. The recognized portions are compared against the retrieved complete documents, allowing error correction and improvement of recognition accuracy through iterative feedback from the search and retrieval process.

Inventive Principle:
Principle #23Feedback

2Productivity

If only partial portions of documents are scanned for recognition, then processing time and resource consumption are reduced, but completeness of retrieved information deteriorates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddocument completeness
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies segmentation by dividing the document processing into two distinct stages: (1) scanning and recognizing only partial portions of documents for efficient querying, and (2) retrieving complete document copies from the database based on search results. This segmentation allows the system to maintain high processing efficiency while eliminating information loss, as the complete documents are obtained from the database rather than attempting to recognize entire documents directly.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If machine recognition algorithms are refined to reduce errors, then recognition accuracy improves, but system complexity and computational resources increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the complexity burden from the machine recognition algorithms by separating the recognition function from the retrieval function. Instead of relying on complex algorithms to achieve high accuracy across entire documents, the system extracts only the essential function of obtaining partial recognized text, which serves as a query input. The complex task of achieving complete and accurate document recognition is transferred to the search and retrieval process, which operates on the extracted partial information.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS8442964B2Information retrieval based on partial machine recognition of the same
Publication Date: 2013.05.14 IDENTITY DIGITAL INC
  • US8442964B2 patent drawing
  • US8442964B2 patent drawing
  • US8442964B2 patent drawing

AI summary

A system and method for capturing and recognizing at least a portion of a source document, whether written or audible, then searching for information, or other documents, that correspond to the captured and recognized portion of the source document. Various techniques for adding translation and/or searching are also disclosed. In some instances, an iterative machine learning process is applied to improve the performance of an aspect of the system.