Automated OCR Template Generation for Accounting Documents
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for digitizing accounting supporting documents are limited in their ability to recognize and process various types of documents, often requiring manual template settings and increasing costs due to the need for OCR template setup, and struggle with handling both printed and handwritten data.
Innovation Solution
An information processing device and method that includes a reception unit for image data, a character recognition unit, an identification unit to classify the purpose of the document, an allocation unit to identify standard items, and an association unit to format the data, which automatically generates a template for character recognition, capable of distinguishing between printed and handwritten areas and associating characters with standard items.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If OCR template setting is performed in advance by a user, then character recognition accuracy is improved, but operation complexity and cost increase
Solution Approach 1:
The system automatically performs template setting and optimization without requiring user intervention. The template setting unit autonomously analyzes document images and configures appropriate OCR templates, eliminating the need for manual template configuration while maintaining high recognition accuracy
Solution Approach 2:
The system pre-registers multiple types of templates for different document formats in advance. When processing a document, the appropriate pre-registered template is automatically selected and applied, avoiding the need for real-time manual template creation while ensuring accurate character recognition
2Measurement precision
If manual input of document details is performed, then data accuracy is improved, but productivity decreases
Solution Approach 1:
The system compares OCR-recognized data with pre-registered standard items and provides feedback for validation. The comparison unit checks whether recognized characters match expected patterns, and the system automatically corrects or flags discrepancies, ensuring data accuracy without requiring manual verification of every field
Solution Approach 2:
The system handles multiple document types (invoices, receipts, bankbooks) using a unified automated processing framework. The same OCR and validation pipeline processes various formats by selecting appropriate pre-registered templates, maintaining both accuracy and high productivity across different document types
3Productivity
If fixed templates are used for character recognition, then processing efficiency is improved, but adaptability to various document types deteriorates
Solution Approach 1:
The system dynamically selects and switches between different pre-registered templates based on the input document type. The template setting unit analyzes document characteristics and automatically configures the appropriate template, enabling efficient processing of various document formats without requiring fixed universal templates
Solution Approach 2:
The system divides document processing into separate handling paths for different document types (invoices, receipts, bankbooks). Each document type has its own optimized template and processing rules, allowing efficient specialized processing while maintaining overall system versatility through modular architecture
Data Source
AI summary
The information processing device includes: a reception unit configured to receive image data regarding an accounting supporting document; a first acquisition unit configured to recognize characters in the image data and acquire the recognized characters as character data; a first identification unit configured to identify a purpose classification on the basis of the character data acquired by the first acquisition unit and the image data; an allocation unit configured to identify a standard item regarding the accounting supporting document on the basis of the character data and allocate characters to the standard item; an association unit configured to associate the characters with other characters adjacent to the characters; and an output unit configured to output the purpose classification identified by the first identification unit and the standard item and the other characters in a format in which the standard item and the other characters are associated by the association unit.


