Document Image Text Recognition via Orientation Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional online storage systems are inflexible and inefficient in managing and searching digital images of documents, particularly those captured by smartphones, due to rigid metadata usage, inaccuracies like skewing and blurring, and the inability to accurately identify text within distorted images.

Innovation Solution

The implementation of a digital image character recognition system utilizing an orientation neural network and text prediction neural network to detect documents, rectify images, and generate searchable text, enabling flexible search and management of digital images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional optical character recognition algorithms are used to identify text from scanned documents, then text identification accuracy is improved for sterile scanned documents, but accuracy deteriorates for user-captured digital images with imperfections and distortions

Engineering Contradiction:
Improvetext identification accuracyVSAvoidhandling capability for distorted images
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms physical image parameters (skew, rotation, blur, shading) into correctable digital parameters through automated image processing. The system detects distortion parameters and applies corresponding correction transformations, converting imperfect captured images into standardized formats suitable for accurate OCR processing.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediate image processing stage between image capture and text recognition. This intermediary system automatically detects and corrects distortions, serving as a mediator that bridges the gap between imperfect user-captured images and the requirements of accurate text identification algorithms.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If conventional systems generate and provide thousands of thumbnails for users to review and search, then comprehensive search coverage is improved, but computational efficiency and time consumption deteriorate

Engineering Contradiction:
Improvesearch coverageVSAvoidsearch efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent performs preliminary text extraction and indexing from digital images before user search operations. By pre-processing images to identify and extract text content, the system creates searchable indexes in advance, eliminating the need for users to manually review thousands of thumbnails and dramatically improving search efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts text content from digital images and separates it from the visual image data. This extraction creates independent searchable text indexes that can be queried without requiring users to view actual images, significantly reducing computational overhead and improving search productivity.

Inventive Principle:
Principle #2Taking out (Extraction)

3Device complexity

If rigid metadata such as user given title, date of creation, and technical specifications are used for storing and searching digital photos, then data organization structure is improved, but search flexibility deteriorates

Engineering Contradiction:
Improvedata organization structureVSAvoidsearch flexibility
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements a multi-functional metadata system that combines traditional rigid metadata (title, date, specifications) with extracted content-based metadata (text content, document type, entities). This universal metadata structure serves multiple purposes: maintaining organized data storage while enabling flexible content-based searching and querying across different document types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent adds a new dimension to metadata by incorporating extracted text content and semantic information alongside traditional structural metadata. This creates a multi-dimensional metadata space that allows searching along both organizational dimensions (file structure, timestamps) and content dimensions (text keywords, document semantics), greatly enhancing search flexibility.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Ease of operation

If digital photos of documents are frequently skewed, blurred, shaded, rotated or otherwise distorted, then ease of capture is improved, but text identification accuracy deteriorates

Engineering Contradiction:
Improveease of document captureVSAvoidtext identification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent converts the harmful effects of distortion (skew, rotation, blur, shading) into beneficial processing opportunities. By detecting these distortions and automatically applying corrective transformations, the system turns captured image imperfections into corrected, OCR-ready images, maintaining ease of capture while restoring text identification accuracy.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS11645826B2Generating searchable text for documents portrayed in a repository of digital images utilizing orientation and text prediction neural networks
Publication Date: 2023.05.09 DROPBOX INC
  • US11645826B2 patent drawing
  • US11645826B2 patent drawing
  • US11645826B2 patent drawing

AI summary

The present disclosure relates to generating computer searchable text from digital images that depict documents utilizing an orientation neural network and/or text prediction neural network. For example, one or more embodiments detect digital images that depict documents, identify the orientation of the depicted documents, and generate computer searchable text from the depicted documents in the detected digital images. In particular, one or more embodiments train an orientation neural network to identify the orientation of a depicted document in a digital image. Additionally, one or more embodiments train a text prediction neural network to analyze a depicted document in a digital image to generate computer searchable text from the depicted document. By utilizing the identified orientation of the depicted document before analyzing the depicted document with a text prediction neural network, the disclosed systems can efficiently and accurately generate computer searchable text for a digital image that depicts a document.