Automated Document Entity Assignment via Text Block Structure
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current processes for assigning documents to database entities lack an automated approach, leading to challenges in associating and managing information stored in different formats, particularly with scanned documents, which can result in missed connections to individuals and non-compliance with regulatory requirements for data reporting.
Innovation Solution
A method that groups documents by similarity based on their structure, retrieves text block values, assigns attributes to these values, and automatically assigns documents to matching entities in a database using similarity-based matching scores, ensuring accurate association and compliance with regulatory demands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If documents are stored in image format (scanned forms), then document storage capacity increases, but information extraction difficulty increases
Solution Approach 1:
The patent replaces manual information extraction from scanned documents with an automated computer vision system that uses image processing algorithms to detect, recognize, and extract text and data from document images, transforming the mechanical process of manual extraction into an automated digital process
Solution Approach 2:
The patent introduces an intermediary automated extraction system that acts as a bridge between stored document images and the database, enabling automatic population of database fields from scanned documents without requiring direct human intervention in the extraction process
2Adaptability or versatility
If information is stored in different formats (scanned documents and structured databases), then data storage flexibility increases, but data association accuracy decreases
Solution Approach 1:
The patent creates a universal data association system that can handle multiple document formats and database structures through a single automated process, using extracted document data to match and associate with corresponding database entities regardless of the original storage format
Solution Approach 2:
The patent implements a feedback mechanism where extracted document data is compared against existing database records to verify correct association, allowing the system to learn from matching results and improve the precision of future document-to-entity associations
3Productivity
If automated document processing is implemented, then processing efficiency increases, but system complexity increases
Solution Approach 1:
The patent divides the automated document processing system into distinct functional modules including image preprocessing, text detection, information extraction, and database association components, allowing each segment to be optimized independently while maintaining overall processing efficiency
Solution Approach 2:
The patent performs preliminary actions by pre-processing document images (such as noise reduction, contrast enhancement, and orientation correction) before the main extraction process, and by pre-defining database schemas and association rules, thereby simplifying the subsequent processing steps
Data Source
AI summary
In an approach, a processor groups documents into a plurality of groups based on similarity, where: documents of each group have a same document structure; and the document structure is defined by coordinates of text blocks. A processor, for each group of the plurality of groups and for each document of the respective group: retrieves a value of each text block of the respective document in accordance with a document structure of the group; and assigns to each text block of the respective document an attribute that represents the retrieved value of the text block. A processor assigns a first document of the documents to an entity of a database that matches the first document based on the group of text block values and the assigned attributes of the document.


