Image Orientation Detection Using Trained Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image processing apparatuses struggle to correctly identify the top and bottom of documents that contain only images, as they rely on optical character recognition (OCR) techniques which are ineffective for text-less documents.

Innovation Solution

An image processing system that includes a trained model based on image data without text, using machine learning and neural networks to estimate the orientation of images, allowing for the identification and correction of top and bottom orientation in documents containing only images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If OCR technique is used to detect top and bottom of document, then text-containing documents can be correctly oriented, but text-less documents cannot be identified

Engineering Contradiction:
Improveorientation detection accuracyVSAvoiddocument type compatibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

A trained machine learning model serves as an intermediary between the image input and orientation detection. The model has been pre-trained on diverse document images including text-less documents, enabling it to generalize orientation detection across different document types without requiring document-specific processing rules

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the detection parameters by using pixel intensity distributions, gradient patterns, and spatial frequency features instead of relying on text-specific OCR parameters. This allows the system to detect orientation in text-less documents by analyzing visual patterns that exist in all document types

Inventive Principle:
Principle #35Parameter changes

2Productivity

If conventional OCR-based method is used, then text document processing is efficient, but image-only document processing fails

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddetection reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The machine learning model performs multiple functions by analyzing various visual features (edges, gradients, textures, patterns) simultaneously. This universal approach allows the same model to process text-containing documents, text-less documents, and mixed-content documents with consistent reliability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system replaces the mechanical OCR text-recognition process with a neural network-based image analysis system. This substitution enables the processing of image-only documents while maintaining efficiency through automated feature extraction and classification

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12113938B2Image processing system, image processing apparatus, control method
Publication Date: 2024.10.08 CANON KK
  • US12113938B2 patent drawing
  • US12113938B2 patent drawing
  • US12113938B2 patent drawing

AI summary

An image processing apparatus includes a display device configured to display information, a reading device configured to read a document, and one or more controllers configured to function as a unit configured to input an image read by the reading device to a trained model trained based on an image that does not contain text and orientation information about the image that does not contain text, and a unit configured to display information about the image read by the reading device on the display device based on at least an output result from the trained model.