Random Document Generation for OCR Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The manual extraction of data from paper documents for electronic data processing systems is labor-intensive, time-consuming, and prone to errors, especially when dealing with non-standardized documents containing various formats, artifacts, and extraneous markings, which hinders the efficient testing of automatic document processing systems.

Innovation Solution

A method and system for automatically generating a large number of varied test samples of documents, including standardized and semi-structured forms, populated with random or real data, and incorporating artifacts such as creases and handwriting, to simulate real-world conditions and improve the testing of document processing systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual extraction of data from paper documents is performed, then data can be transferred to electronic systems, but the process is labor-intensive and time-consuming

Engineering Contradiction:
Improvedata extraction speedVSAvoidtime for manual processing
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical data extraction with an automated optical document processing system that uses image scanning and computer vision algorithms to automatically extract data from paper documents, eliminating the need for manual typing and data entry while significantly improving processing speed and accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables documents to process themselves through automated optical character recognition and data extraction algorithms that can independently identify, extract, and validate information from various document formats without human intervention, allowing the system to serve itself in the data extraction task

Inventive Principle:
Principle #25Self-service

2Extent of automation

If optical character recognition is used to extract data from documents, then data can be read automatically, but the system struggles with non-standardized documents containing various formats and artifacts

Engineering Contradiction:
Improveautomatic data extractionVSAvoidhandling of document variations
Core Design Contradiction:
Extent of automationVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic document processing system that can adapt its extraction algorithms based on the specific characteristics of each document encountered. The system dynamically adjusts its approach to handle different document formats, layouts, and artifacts by using machine learning models that learn from various document types and improve their adaptability over time

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters such as image preprocessing techniques, recognition algorithms, and validation rules based on the detected document type and characteristics. By dynamically adjusting these parameters, the system maintains high automation levels while effectively handling non-standardized documents with various formats and artifacts

Inventive Principle:
Principle #35Parameter changes

3Reliability

If a large number of test samples are generated for system testing, then accurate performance predictions can be made, but the generation process becomes time-consuming and resource-intensive

Engineering Contradiction:
Improveaccuracy of performance predictionsVSAvoidtesting efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent uses template-based document generation that creates realistic test samples by copying and populating predefined document templates with varied data. This approach efficiently generates large numbers of test documents with consistent formatting and realistic content variations, enabling comprehensive system testing without manually creating each document from scratch

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary preparation by creating reusable document templates that contain the structural characteristics of target documents. These pre-prepared templates can be rapidly instantiated with different data to generate numerous test samples, significantly reducing the time and resources required for comprehensive testing while maintaining high reliability in performance predictions

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7840890B2Generation of randomly structured forms
Publication Date: 2010.11.23 OPEN TEXT CORP
  • US7840890B2 patent drawing
  • US7840890B2 patent drawing
  • US7840890B2 patent drawing

AI summary

A computer system and method for generating a plurality of unique, randomly structured forms, such as invoices, that may be populated with data to produce test forms for testing automatic document processing systems. The forms may have major blocks such as a header, a body and a footer, and the major blocks may have randomly selected sizes. For each block, the locations of data fields and the ordering of the data fields within the block are defined randomly, the data locations within the fields and the data formats are randomly defined, and a blank image and an XML file of the randomly structured form are produced. The forms generated may be populated with data for the testing document processing system.