PDF Text Extraction via DPI Rendering and Line Templates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting data from Adobe PDF documents are cumbersome and error-prone, requiring manual re-entry of data into computer applications due to the inability to manipulate PDF documents effectively.

Innovation Solution

A user interface that renders Adobe PDF documents into DPI images, allowing text extraction and transformation into editable formats like Microsoft Excel, using line insertion tools to define templates and a text extraction parser to export renderable text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual data re-entry from PDF documents is used, then data can be transferred to computer applications, but the process is cumbersome and error-prone

Engineering Contradiction:
Improvedata transfer efficiencyVSAvoiddata accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent uses optical character recognition (OCR) technology to create digital copies of text from PDF documents. The system captures the visual appearance of text in the PDF and converts it into machine-readable text data, eliminating the need for manual re-entry while preserving accuracy through automated recognition processes.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the manual mechanical process of reading and re-typing data with an automated computer vision system. The OCR engine automatically detects, recognizes, and extracts text from PDF documents, substituting human manual operations with computational processes that are both faster and more accurate.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If PDF documents are rendered into DPI images, then text extraction becomes possible in Java regions, but the original PDF structure is lost

Engineering Contradiction:
Improvecompatibility with Java regionsVSAvoidPDF document structure
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent performs preliminary rendering of the PDF document into a high-resolution DPI image before text extraction. This pre-processing step converts the PDF into a format that is compatible with Java regions, enabling subsequent OCR processing to work effectively while preserving the visual text information needed for accurate extraction.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a DPI image as an intermediary format between the original PDF document and the final extracted text. This intermediate representation maintains the visual fidelity of the original document while being compatible with Java-based processing systems, allowing text extraction without direct loss of the original PDF structure.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If line insertion tools are used to define templates, then text extraction accuracy improves, but the operation becomes more complex

Engineering Contradiction:
Improvetext extraction accuracyVSAvoidtemplate definition complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements automated template generation that analyzes the PDF document structure and creates extraction templates without requiring manual line insertion. The system automatically detects text regions, tables, and structural elements, generating appropriate extraction patterns that improve accuracy while eliminating the complexity of manual template definition.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent automatically adjusts template parameters based on the characteristics of the PDF document being processed. The system dynamically modifies extraction parameters such as region boundaries, text selection criteria, and formatting rules to match the specific document structure, thereby maintaining high accuracy without requiring complex manual configuration.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10146763B2Renderable text extraction tool
Publication Date: 2018.12.04 BANK OF AMERICA CORP
  • US10146763B2 patent drawing
  • US10146763B2 patent drawing
  • US10146763B2 patent drawing

AI summary

Methods for converting an Adobe™ PDF document into an editable document is provided. Methods may receive an Adobe™ PDF document and displaying the Adobe™ PDF document. Methods may enable a user to create a plurality of horizontal lines and a plurality of vertical lines on the document. The horizontal and vertical lines may create rows and columns. Methods may create an editable document upon receipt of at least one row and at least one column on the document. The editable document may correspond to the rows and columns within the created horizontal and vertical lines. The editable document may be a Microsoft Excel™ spreadsheet or any other suitable document. Methods may create a horizontal line or vertical line at a location of a cursor when a corresponding click is received.