Document Characteristic Extraction Using Noise-Resistant Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current information processing systems face challenges in accurately extracting and distinguishing elements representing characteristics from multiple documents, particularly in the presence of noise and variations, which can lead to incorrect document classification.

Innovation Solution

An information processing apparatus comprising an acquiring unit, an extraction unit, and a selection unit that acquires candidates for document elements, extracts common and unique elements, and determines these elements as representing characteristics, even in noisy conditions, by using character recognition and ruled-line recognition methods, and generating multiple images with varying noise levels to improve similarity analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If character recognition and ruled-line recognition are used to extract document elements, then document classification accuracy is improved, but the system becomes more complex and computationally intensive

Engineering Contradiction:
Improvedocument classification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The document analysis process is divided into distinct segments: character recognition extracts text elements, ruled-line recognition extracts structural elements, and common element extraction identifies characteristic patterns. Each segment handles a specific aspect of document analysis, improving overall accuracy while maintaining manageable system complexity through modular processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary common element extraction process that bridges character recognition and ruled-line recognition results. This intermediary step identifies elements common to multiple documents, serving as a mediator that consolidates information from different recognition processes and reduces computational burden on subsequent classification stages

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If multiple images with varying noise levels are generated for similarity analysis, then element extraction robustness is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveelement extraction robustnessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

Multiple images with varying noise levels are generated in advance before the actual similarity analysis. This preliminary action creates a robust dataset that accounts for potential noise variations in real document processing, ensuring reliable element extraction while allowing the main processing stage to work with pre-prepared data, thus reducing overall processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system varies noise parameters when generating multiple images for similarity analysis. By changing the noise level parameter across different images, the system tests element extraction robustness under different conditions without requiring completely separate processing pipelines, thereby managing computational resources efficiently while improving reliability

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If common elements are extracted from multiple documents to represent characteristics, then document classification accuracy is improved, but the ability to identify unique document features may be reduced

Engineering Contradiction:
Improvedocument classification accuracyVSAvoidunique feature information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system extracts common elements that appear across multiple documents of the same type, separating these characteristic features from unique document-specific variations. This extraction process identifies the essential commonalities that define document categories while allowing unique features to be preserved as differentiating characteristics, thus improving classification accuracy without losing important information

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10049269B2Information processing apparatus, information processing method, and non-transitory computer readable medium
Publication Date: 2018.08.14 FUJIFILM BUSINESS INNOVATION CORP
  • US10049269B2 patent drawing
  • US10049269B2 patent drawing
  • US10049269B2 patent drawing

AI summary

An information processing apparatus includes an acquiring unit, an extraction unit, and a selection unit. The acquiring unit acquires, for multiple documents, candidates for elements representing characteristics of each of the multiple documents. The extraction unit extracts, from the candidates acquired by the acquiring unit, common elements common to two or more of the multiple documents. The selection unit extracts, from the multiple documents, a document including two or more common elements among the common elements, and determines the two or more common elements included in the extracted document to be elements representing characteristics of the document.