Image Orientation Detection Using Trained Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image processing apparatuses struggle to correctly identify the top and bottom of documents that contain only images, as they rely on optical character recognition (OCR) techniques which are ineffective for text-less documents.
Innovation Solution
An image processing system that includes a trained model based on image data without text, using machine learning and neural networks to estimate the orientation of images, allowing for the identification and correction of top and bottom orientation in documents containing only images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If OCR technique is used to detect top and bottom of document, then text-containing documents can be correctly oriented, but text-less documents cannot be identified
Solution Approach 1:
A trained machine learning model serves as an intermediary between the image input and orientation detection. The model has been pre-trained on diverse document images including text-less documents, enabling it to generalize orientation detection across different document types without requiring document-specific processing rules
Solution Approach 2:
The system changes the detection parameters by using pixel intensity distributions, gradient patterns, and spatial frequency features instead of relying on text-specific OCR parameters. This allows the system to detect orientation in text-less documents by analyzing visual patterns that exist in all document types
2Productivity
If conventional OCR-based method is used, then text document processing is efficient, but image-only document processing fails
Solution Approach 1:
The machine learning model performs multiple functions by analyzing various visual features (edges, gradients, textures, patterns) simultaneously. This universal approach allows the same model to process text-containing documents, text-less documents, and mixed-content documents with consistent reliability
Solution Approach 2:
The system replaces the mechanical OCR text-recognition process with a neural network-based image analysis system. This substitution enables the processing of image-only documents while maintaining efficiency through automated feature extraction and classification
Data Source
AI summary
An image processing apparatus includes a display device configured to display information, a reading device configured to read a document, and one or more controllers configured to function as a unit configured to input an image read by the reading device to a trained model trained based on an image that does not contain text and orientation information about the image that does not contain text, and a unit configured to display information about the image read by the reading device on the display device based on at least an output result from the trained model.


