3D Document Simulation for OCR Image Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional optical character recognition (OCR) systems face challenges in accurately recognizing text from images captured by digital cameras and mobile devices due to image distortions and the limited availability of high-quality reference images, leading to inefficient data processing and biased results.

Innovation Solution

A computer-implemented method using a three-dimensional modeling engine to simulate the capture of documents under various circumstances, generating a large set of high-quality reference images to train a computer model that determines optimal parameters for adjusting image characteristics, such as lighting and camera pose, to improve OCR accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional OCR systems use images captured by digital cameras and mobile devices, then image acquisition becomes more accessible and convenient, but image quality deteriorates due to distortions such as blur, skew, rotation, and shadow marks

Engineering Contradiction:
Improveimage acquisition convenienceVSAvoidimage quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system performs preliminary image processing operations including de-skewing, de-rotating, and shadow removal on captured images before OCR recognition. By pre-processing the distorted images to correct geometric distortions and remove artifacts, the system maintains accessibility of mobile capture while improving image quality for accurate text recognition

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary image processing system that acts as a mediator between the mobile capture device and the OCR engine. This intermediary layer applies corrective transformations and filtering to bridge the gap between convenient mobile capture and the high-quality images required for accurate OCR, without requiring changes to the capture device or the OCR engine

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If conventional systems use a limited source image set for OCR accuracy determination, then processing time is reduced, but measurement precision deteriorates due to bias from incidental characteristics of available images

Engineering Contradiction:
Improveprocessing speedVSAvoidOCR accuracy determination
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system transitions from using a limited two-dimensional set of actual captured images to a three-dimensional simulated image space that can generate unlimited reference images with controlled variations in lighting, angle, distance, and document type. This dimensional expansion allows comprehensive training data without the constraints of manual image collection, improving accuracy determination while maintaining efficient processing through automated generation

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

Instead of relying on a limited set of original captured images, the system creates numerous synthetic copies through 3D rendering of document models under various simulated conditions. These copied and transformed reference images provide diverse training data for accurate OCR determination without requiring proportional increases in manual image collection effort

Inventive Principle:
Principle #26Copying

3Reliability

If manual pre-processing is performed to redact confidential information from customer images, then data security is improved, but productivity deteriorates due to the time-consuming nature of manual redaction

Engineering Contradiction:
Improvedata securityVSAvoidimage processing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system employs automated image processing algorithms that independently identify and redact confidential information from customer images without requiring manual human intervention. The automated redaction system maintains data security while dramatically improving processing efficiency by handling multiple images simultaneously through computer vision and machine learning techniques

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10943107B2Simulating image capture
Publication Date: 2021.03.09 INTUIT INC
  • US10943107B2 patent drawing
  • US10943107B2 patent drawing
  • US10943107B2 patent drawing

AI summary

The present disclosure relates to simulating the capture of images. In some embodiments, a document and a camera are simulated using a three-dimensional modeling engine. In certain embodiments, a plurality of images are captured of the simulated document from a perspective of the simulated camera, each of the plurality of images being captured under a different set of simulated circumstances within the three-dimensional modeling engine. In some embodiments, a model is trained based at least on the plurality of images which determines at least a first technique for adjusting a set of parameters in a separate image to prepare the separate image for optical character recognition (OCR).