Automated Accounting Data Extraction Using Entity Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for processing digitized accounting source documents, such as invoices and receipts, are labor-intensive and inefficient due to format variations and the need for extensive computation resources, particularly in identifying vendor names and processing specific accounting information like inventory items, which requires different techniques.
Innovation Solution
A computer-assisted method that utilizes an entity database and a digital template library to match entity identifiers in digitized documents with corresponding processing templates, allowing for efficient extraction and processing of accounting data, including the use of numerical, text, or URL identifiers, and automatic verification of data accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual data inputting is used for processing invoices and receipts, then data accuracy can be maintained, but processing time and labor intensity increase significantly
Solution Approach 1:
The system enables self-service by automatically extracting data from digitized invoices and receipts using OCR and NLP technologies. The accounting system autonomously identifies vendor names, document types, and accounting information without human intervention, allowing the system to serve itself in data extraction and population tasks
Solution Approach 2:
The patent replaces the mechanical manual data entry process with an automated digital system. Optical character recognition (OCR) converts scanned documents into machine-readable text, and natural language processing (NLP) algorithms automatically extract and categorize accounting information, substituting human manual operations with computational processes
2Measurement precision
If extensive computation resources are used for automatic context parsing, then data extraction accuracy improves, but processing efficiency decreases
Solution Approach 1:
The system performs preliminary actions by pre-processing digitized documents through OCR to convert images into searchable text before the main extraction process. This preliminary step prepares the data in advance, making subsequent NLP-based extraction faster and more accurate without requiring excessive computational resources during the actual processing phase
Solution Approach 2:
The patent segments the complex document processing task into distinct phases: OCR text recognition, vendor name identification, document type classification, and accounting information extraction. Each segment handles a specific aspect of the processing, allowing optimized computational approaches for each sub-task rather than applying intensive computation uniformly across the entire process
3Adaptability or versatility
If diverse document formats are processed using generic techniques, then system simplicity is maintained, but processing accuracy for specific document types deteriorates
Solution Approach 1:
The system achieves universality by implementing a multi-functional processing architecture that can handle various document types (invoices, receipts, bills) and formats (PDF, images, scanned documents) through a single integrated platform. The NLP-based extraction engine adapts to different document structures while maintaining consistent accuracy across diverse formats
Data Source
AI summary
To generate data for populating/updating accounting databases based on digitized accounting source documents, access to an entity database comprising identifiers of entities associated with an accounting database and to a digital template library comprising processing templates for processing digitized accounting source documents is provided. Each entity in the entity database is associated with one processing template. A processor receives digitized data representing a digitized accounting source document; determines if the digitized data comprises an entity identifier that matches a particular identifier of a particular entity in the entity database; and in response to determining that the entity identifier matches the particular identifier of the particular entity in the entity database, retrieves from the template library a particular processing template associated with the particular entity; and processes the digitized data to generate processed data, according to the particular processing template, for populating/updating the accounting database.


