Document Title Judgment Using Selective Character Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for classifying documents scanned continuously face challenges such as converting separate documents into a single image and misclassifying continuous documents into different images per page, with labor-intensive solutions like inserting white sheets being necessary.

Innovation Solution

An information processing apparatus with a control device that extracts character areas, calculates feature quantities, generates a machine learning model for title judgment, calculates title reliability levels, performs character recognition, and judges titles based on pre-stored candidates, thereby automatically classifying documents without manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If all character areas are converted to text data for title judgment, then title judgment accuracy is improved, but processing time and computational load increase significantly

Engineering Contradiction:
Improvetitle judgment accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by selectively converting only certain character areas to text data rather than all character areas. The system identifies and converts only those character areas with high title reliability scores (above a threshold) into text data for title judgment, while leaving other character areas unconverted. This selective approach maintains title judgment accuracy for relevant areas while significantly reducing overall processing time and computational load.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If machine learning model processes all character areas, then title detection accuracy is improved, but processing load increases

Engineering Contradiction:
Improvetitle detection accuracyVSAvoidprocessing load
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The system applies partial action by having the machine learning model process only character areas that meet specific criteria (high title reliability scores) rather than all character areas uniformly. This selective processing maintains title detection accuracy for relevant character areas while significantly reducing the overall processing load on the system.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent applies local quality by assigning different processing treatments to different character areas based on their individual title reliability scores. Character areas with high title reliability undergo full machine learning processing and text conversion, while character areas with low title reliability are excluded from intensive processing. This localized quality approach optimizes resource allocation while maintaining detection accuracy where it matters most.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250029416A1Information processing apparatus and image reading apparatus judging title of read document
Publication Date: 2025.01.23 KYOCERA DOCUMENT SOLUTIONS INC
  • US20250029416A1 patent drawing
  • US20250029416A1 patent drawing
  • US20250029416A1 patent drawing

AI summary

In an image reading apparatus, a character area extractor extracts character areas from a document image in units of rows. A title reliability calculator calculates a title reliability level of each character area using a feature quantity data set and a machine learning model. A character recognizer converts character areas of which title reliability levels exceed a threshold into text data. A title judger collates text data with a title candidate. In a case in which one piece of text data coinciding with a title candidate is judged and detected, the title judger sets the text data as a title of the document image. In a case in which a plurality of pieces of coinciding text data are judged and detected, the title judger sets text data of which a title reliability level is the highest among the detected pieces of text data as a title of the document image.