Document Type Detection for Adaptive Top-Bottom Orientation Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing systems face challenges in accurately determining the top-bottom orientation of documents scanned in incorrect orientations, particularly when multiple document types are scanned simultaneously, leading to decreased accuracy and increased processing time due to the lack of adaptive switching between OCR and AI-based top-bottom determination methods.
Innovation Solution
An image processing apparatus that determines the type of document based on OCR setting information and selects the appropriate top-bottom determination method, switching between OCR and AI methods dynamically to optimize accuracy for each document type, ensuring high accuracy in top-bottom orientation correction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single top-bottom determination method is used for all document types, then the system is simple to operate, but the accuracy decreases when scanning mixed document types
Solution Approach 1:
The patent implements dynamic switching between OCR-based and AI-based top-bottom determination methods based on document type detection. The system automatically selects the appropriate method for each document type, transitioning from a static single-method approach to a dynamic multi-method approach, thereby improving accuracy without requiring manual user intervention to switch methods.
Solution Approach 2:
The patent changes the determination method parameter based on document type characteristics. By detecting document type and selecting different determination methods (OCR for text documents, AI for images/photos), the system adapts the processing parameters to match the document characteristics, resolving the contradiction between accuracy and complexity.
2Productivity
If OCR method is used for all document types, then text documents are processed accurately, but processing time increases for image documents
Solution Approach 1:
The patent applies different determination methods to different document types locally. Text documents use OCR-based determination for high accuracy, while image documents use AI-based determination for faster processing. This localized approach ensures each document type receives the most suitable method, optimizing both speed and accuracy for their respective categories.
Solution Approach 2:
The patent applies OCR processing only partially to text documents rather than all documents. By detecting document type first, the system applies OCR only where appropriate (text documents) and uses AI methods for image documents, avoiding excessive OCR processing on images that would waste time and resources while maintaining accuracy where needed.
3Loss of time
If AI method is used for all document types, then processing time is reduced, but accuracy decreases for text documents
Solution Approach 1:
The system dynamically selects between AI-based and OCR-based methods based on document type detection. For image documents, it uses AI methods for fast processing; for text documents, it switches to OCR methods for high accuracy. This dynamic adaptation resolves the contradiction by matching the method to the document characteristics rather than using a fixed approach.
4Measurement precision
If multiple determination methods are available but not switched adaptively, then the system has high accuracy potential, but processing time increases due to unnecessary method evaluation
Solution Approach 1:
The patent performs preliminary document type detection before selecting the determination method. By classifying documents into text, image, or other types in advance, the system pre-determines which method to apply, avoiding unnecessary method evaluation and switching during the actual processing. This preliminary classification action reduces processing time while maintaining high accuracy.
Data Source
AI summary
An image processing apparatus includes circuitry to determine a type of a document read by a scanner, set a top-bottom determination method based on the type of the document, and determine a top-bottom orientation of a target image by the top-bottom determination method. The target image is obtained by reading the document with the scanner.


