Automated Document Entity Assignment via Text Block Structure

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current processes for assigning documents to database entities lack an automated approach, leading to challenges in associating and managing information stored in different formats, particularly with scanned documents, which can result in missed connections to individuals and non-compliance with regulatory requirements for data reporting.

Innovation Solution

A method that groups documents by similarity based on their structure, retrieves text block values, assigns attributes to these values, and automatically assigns documents to matching entities in a database using similarity-based matching scores, ensuring accurate association and compliance with regulatory demands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If documents are stored in image format (scanned forms), then document storage capacity increases, but information extraction difficulty increases

Engineering Contradiction:
Improvedocument storage capacityVSAvoidinformation extraction difficulty
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent replaces manual information extraction from scanned documents with an automated computer vision system that uses image processing algorithms to detect, recognize, and extract text and data from document images, transforming the mechanical process of manual extraction into an automated digital process

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary automated extraction system that acts as a bridge between stored document images and the database, enabling automatic population of database fields from scanned documents without requiring direct human intervention in the extraction process

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If information is stored in different formats (scanned documents and structured databases), then data storage flexibility increases, but data association accuracy decreases

Engineering Contradiction:
Improvedata storage flexibilityVSAvoiddata association accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent creates a universal data association system that can handle multiple document formats and database structures through a single automated process, using extracted document data to match and associate with corresponding database entities regardless of the original storage format

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements a feedback mechanism where extracted document data is compared against existing database records to verify correct association, allowing the system to learn from matching results and improve the precision of future document-to-entity associations

Inventive Principle:
Principle #23Feedback

3Productivity

If automated document processing is implemented, then processing efficiency increases, but system complexity increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the automated document processing system into distinct functional modules including image preprocessing, text detection, information extraction, and database association components, allowing each segment to be optimized independently while maintaining overall processing efficiency

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-processing document images (such as noise reduction, contrast enhancement, and orientation correction) before the main extraction process, and by pre-defining database schemas and association rules, thereby simplifying the subsequent processing steps

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11593417B2Assigning documents to entities of a database
Publication Date: 2023.02.28 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11593417B2 patent drawing
  • US11593417B2 patent drawing
  • US11593417B2 patent drawing

AI summary

In an approach, a processor groups documents into a plurality of groups based on similarity, where: documents of each group have a same document structure; and the document structure is defined by coordinates of text blocks. A processor, for each group of the plurality of groups and for each document of the respective group: retrieves a value of each text block of the respective document in accordance with a document structure of the group; and assigns to each text block of the respective document an attribute that represents the retrieved value of the text block. A processor assigns a first document of the documents to an entity of a database that matches the first document based on the group of text block values and the assigned attributes of the document.