PDF Text Recognition Accuracy via Dual Data Source Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for processing PDF files, such as recognizing characters in images, often result in incorrect text data due to inaccuracies in character recognition, leading to incorrect text data being obtained.

Innovation Solution

An information processing system that determines whether to use pre-existing text data or newly generated text data from character recognition, utilizing a server that acquires and processes PDF files, and employs multiple types of ledger sheet definition information to improve recognition accuracy by associating items with their values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If character recognition is performed on images in PDF files, then text data can be extracted from scanned documents, but the text data may be incorrectly recognized

Engineering Contradiction:
Improvetext data extractionVSAvoidrecognition accuracy
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system performs preliminary actions by acquiring both image data and corresponding text data from PDF files before the recognition process. This allows the system to have reference text data available for comparison with the recognized results, enabling correction of recognition errors and improving overall reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback by comparing the text data obtained through character recognition with the original text data that was extracted from the PDF file. This feedback mechanism allows the system to identify and correct incorrect recognitions, thereby improving the reliability of the text extraction process.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If only character recognition on images is used, then text data can be obtained from any document format, but the accuracy of the extracted text data decreases

Engineering Contradiction:
Improvedocument format compatibilityVSAvoidtext recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system merges two approaches: using both character recognition on images and extraction of text data from the PDF file structure. By combining these methods, the system maintains adaptability to process any document format while improving text recognition accuracy through the use of the extracted text data as a reference.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system creates a composite approach by integrating multiple data sources (image data and text data from PDF structure) into a unified text extraction process. This composite method leverages the strengths of both approaches to achieve both versatility in handling different formats and high accuracy in text extraction.

Inventive Principle:
Principle #40Composite materials

3Productivity

If pre-converted text data from PDF files is used, then processing speed is improved, but the text data may contain errors from the conversion process

Engineering Contradiction:
Improvetext processing speedVSAvoidtext data accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system uses feedback by comparing the pre-converted text data with the original image data and re-recognized text. This allows the system to quickly process documents using the pre-converted text while still being able to detect and correct errors, thereby maintaining both high productivity and reliability.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary extraction of text data from the PDF file structure before the final recognition process. This preliminary action speeds up processing by having text data ready in advance, while the subsequent comparison with image-based recognition ensures accuracy, thus resolving the contradiction between speed and reliability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11113559B2Information processing apparatus for improving text data recognition, information processing method, and non-transitory recording medium
Publication Date: 2021.09.07 RICOH CO LTD
  • US11113559B2 patent drawing
  • US11113559B2 patent drawing
  • US11113559B2 patent drawing

AI summary

An information processing apparatus includes processing circuitry that acquires an electronic file containing first text data, and determines, based on the acquired electronic file, whether to use the first text data or second text data to perform a process. The second text data is generated through character recognition performed on an image contained in the acquired electronic file.