Document Image Text Recognition via Orientation Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional online storage systems are inflexible and inefficient in managing and searching digital images of documents, particularly those captured by smartphones, due to rigid metadata usage, inaccuracies like skewing and blurring, and the inability to accurately identify text within distorted images.
Innovation Solution
The implementation of a digital image character recognition system utilizing an orientation neural network and text prediction neural network to detect documents, rectify images, and generate searchable text, enabling flexible search and management of digital images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional optical character recognition algorithms are used to identify text from scanned documents, then text identification accuracy is improved for sterile scanned documents, but accuracy deteriorates for user-captured digital images with imperfections and distortions
Solution Approach 1:
The patent transforms physical image parameters (skew, rotation, blur, shading) into correctable digital parameters through automated image processing. The system detects distortion parameters and applies corresponding correction transformations, converting imperfect captured images into standardized formats suitable for accurate OCR processing.
Solution Approach 2:
The patent introduces an intermediate image processing stage between image capture and text recognition. This intermediary system automatically detects and corrects distortions, serving as a mediator that bridges the gap between imperfect user-captured images and the requirements of accurate text identification algorithms.
2Loss of information
If conventional systems generate and provide thousands of thumbnails for users to review and search, then comprehensive search coverage is improved, but computational efficiency and time consumption deteriorate
Solution Approach 1:
The patent performs preliminary text extraction and indexing from digital images before user search operations. By pre-processing images to identify and extract text content, the system creates searchable indexes in advance, eliminating the need for users to manually review thousands of thumbnails and dramatically improving search efficiency.
Solution Approach 2:
The patent extracts text content from digital images and separates it from the visual image data. This extraction creates independent searchable text indexes that can be queried without requiring users to view actual images, significantly reducing computational overhead and improving search productivity.
3Device complexity
If rigid metadata such as user given title, date of creation, and technical specifications are used for storing and searching digital photos, then data organization structure is improved, but search flexibility deteriorates
Solution Approach 1:
The patent implements a multi-functional metadata system that combines traditional rigid metadata (title, date, specifications) with extracted content-based metadata (text content, document type, entities). This universal metadata structure serves multiple purposes: maintaining organized data storage while enabling flexible content-based searching and querying across different document types.
Solution Approach 2:
The patent adds a new dimension to metadata by incorporating extracted text content and semantic information alongside traditional structural metadata. This creates a multi-dimensional metadata space that allows searching along both organizational dimensions (file structure, timestamps) and content dimensions (text keywords, document semantics), greatly enhancing search flexibility.
4Ease of operation
If digital photos of documents are frequently skewed, blurred, shaded, rotated or otherwise distorted, then ease of capture is improved, but text identification accuracy deteriorates
Solution Approach 1:
The patent converts the harmful effects of distortion (skew, rotation, blur, shading) into beneficial processing opportunities. By detecting these distortions and automatically applying corrective transformations, the system turns captured image imperfections into corrected, OCR-ready images, maintaining ease of capture while restoring text identification accuracy.
Data Source
AI summary
The present disclosure relates to generating computer searchable text from digital images that depict documents utilizing an orientation neural network and/or text prediction neural network. For example, one or more embodiments detect digital images that depict documents, identify the orientation of the depicted documents, and generate computer searchable text from the depicted documents in the detected digital images. In particular, one or more embodiments train an orientation neural network to identify the orientation of a depicted document in a digital image. Additionally, one or more embodiments train a text prediction neural network to analyze a depicted document in a digital image to generate computer searchable text from the depicted document. By utilizing the identified orientation of the depicted document before analyzing the depicted document with a text prediction neural network, the disclosed systems can efficiently and accurately generate computer searchable text for a digital image that depicts a document.


