PDF Text Recognition Accuracy via Dual Data Source Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for processing PDF files, such as recognizing characters in images, often result in incorrect text data due to inaccuracies in character recognition, leading to incorrect text data being obtained.
Innovation Solution
An information processing system that determines whether to use pre-existing text data or newly generated text data from character recognition, utilizing a server that acquires and processes PDF files, and employs multiple types of ledger sheet definition information to improve recognition accuracy by associating items with their values.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If character recognition is performed on images in PDF files, then text data can be extracted from scanned documents, but the text data may be incorrectly recognized
Solution Approach 1:
The system performs preliminary actions by acquiring both image data and corresponding text data from PDF files before the recognition process. This allows the system to have reference text data available for comparison with the recognized results, enabling correction of recognition errors and improving overall reliability.
Solution Approach 2:
The system implements feedback by comparing the text data obtained through character recognition with the original text data that was extracted from the PDF file. This feedback mechanism allows the system to identify and correct incorrect recognitions, thereby improving the reliability of the text extraction process.
2Adaptability or versatility
If only character recognition on images is used, then text data can be obtained from any document format, but the accuracy of the extracted text data decreases
Solution Approach 1:
The system merges two approaches: using both character recognition on images and extraction of text data from the PDF file structure. By combining these methods, the system maintains adaptability to process any document format while improving text recognition accuracy through the use of the extracted text data as a reference.
Solution Approach 2:
The system creates a composite approach by integrating multiple data sources (image data and text data from PDF structure) into a unified text extraction process. This composite method leverages the strengths of both approaches to achieve both versatility in handling different formats and high accuracy in text extraction.
3Productivity
If pre-converted text data from PDF files is used, then processing speed is improved, but the text data may contain errors from the conversion process
Solution Approach 1:
The system uses feedback by comparing the pre-converted text data with the original image data and re-recognized text. This allows the system to quickly process documents using the pre-converted text while still being able to detect and correct errors, thereby maintaining both high productivity and reliability.
Solution Approach 2:
The system performs preliminary extraction of text data from the PDF file structure before the final recognition process. This preliminary action speeds up processing by having text data ready in advance, while the subsequent comparison with image-based recognition ensures accuracy, thus resolving the contradiction between speed and reliability.
Data Source
AI summary
An information processing apparatus includes processing circuitry that acquires an electronic file containing first text data, and determines, based on the acquired electronic file, whether to use the first text data or second text data to perform a process. The second text data is generated through character recognition performed on an image contained in the acquired electronic file.


