Fuzzy Document Matching for OCR-Based Record Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of correlating third-party documents, such as those from governmental entities, with bank database entries is complicated by inconsistent and unreliable information provided by third parties, leading to difficulties in locating correct customer records due to misspellings, varying descriptions, and lack of unique identifiers, necessitating manual human intervention.
Innovation Solution
A computing device processes images of third-party documents using Optical Character Recognition (OCR) and a fuzzy matching algorithm to identify associations with database entries, allowing automatic generation of electronic transactions based on trustworthiness assessments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual evaluation and searching is used to identify correct database entries, then accuracy can be maintained, but time consumption and operational complexity increase significantly
Solution Approach 1:
The patent replaces manual human evaluation and searching processes with an automated image processing system that uses optical character recognition (OCR) to extract text from document images and fuzzy matching algorithms to identify corresponding database entries. This substitution of mechanical manual operations with automated computational processes directly addresses the time consumption issue while maintaining accuracy through the fuzzy matching approach that handles inconsistencies in third-party documents.
2Extent of automation
If third-party documents are used for identification, then process automation can be achieved, but reliability decreases due to inconsistencies and inaccuracies in the documents
Solution Approach 1:
The patent introduces an intermediary processing layer between the unreliable third-party documents and the database entries. This intermediary system uses OCR to extract information from document images and fuzzy matching algorithms to bridge the gaps caused by inconsistencies, misspellings, and inaccurate descriptions in third-party documents. The intermediary effectively filters and reconciles the unreliable data, enabling automated processing while maintaining reliability through the fuzzy matching approach that can handle variations in the data.
3Adaptability or versatility
If fuzzy matching algorithm is used to compare document content with database entries, then ability to handle inconsistencies improves, but computational complexity increases
Solution Approach 1:
The patent segments the complex fuzzy matching process into distinct manageable steps: first extracting text from document images using OCR, then processing and filtering the extracted text to identify key information, and finally comparing this processed information with database entries using fuzzy matching algorithms. This segmentation of the computational task into separate functional modules makes the overall system more manageable and optimizes computational resources while maintaining the adaptability needed to handle inconsistencies in third-party documents.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Automatically correlates document images with database records despite inconsistencies, reducing the need for human intervention and improving efficiency in processing unclaimed property transactions.
Implementation Method 1
The image may be then be processed using Optical Character Recognition (OCR) (and, for instance, an OCR algorithm) to identify textual content
Data Source
AI summary
Systems, methods, and apparatuses are described for automatically processing images of printed documents transmitted by third parties to identify associations between those printed documents and database entries. A computing device may store, in a database, records of various individuals. Those records may correspond to different individuals and may comprise one or more expected properties of a document o be received concerning a given individual. The computing device may later receive a document image and process it using Optical Character Recognition to identify textual content. A fuzzy matching algorithm may be used to corelate the textual content with the database records to identify a first record associated with an e-mail account. Based on a trust level associated with that e-mail account, funds may be transmitted.


