Document Processing System Linking Offline Images to Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document processing systems fail to effectively link offline documents, such as paper documents or facsimile documents, with their associated metadata and relation data to online semantic networks, leading to disconnection and inefficient management and search capabilities.

Innovation Solution

A document processing system that includes a storage unit for document data with metadata and relation information, an input unit for image data, and a control unit to specify and store related documents based on metadata, enabling the linking of input image data with relevant document data and maintaining relation information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If document processing systems store and manage document data with metadata and relation information, then document management and search capabilities are improved, but the system complexity increases

Engineering Contradiction:
Improvedocument management efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a document processing apparatus as an intermediary component that sits between document storage systems and user applications. This apparatus automatically extracts metadata and relation information from documents, manages the semantic networks, and provides standardized interfaces for document retrieval and management, thereby reducing overall system complexity while improving management efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables documents to be automatically processed and indexed without manual intervention. The document processing apparatus autonomously extracts metadata, establishes relation information, and integrates documents into semantic networks automatically, reducing the need for complex manual management processes

Inventive Principle:
Principle #25Self-service

2Ease of operation

If offline documents are disconnected from online semantic networks, then system simplicity is maintained, but document retrieval efficiency deteriorates

Engineering Contradiction:
Improvesystem simplicityVSAvoiddocument retrieval efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system performs preliminary processing of offline documents by automatically extracting metadata and establishing relation information before they are needed for retrieval. This pre-processing creates the semantic network connections in advance, enabling efficient document retrieval without requiring complex real-time processing or maintaining overly complex system structures

Inventive Principle:
Principle #10Preliminary action

3Productivity

If relation information between documents is stored and managed, then document organization is improved, but storage requirements increase

Engineering Contradiction:
Improvedocument organization efficiencyVSAvoidstorage requirements
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent applies local quality by storing relation information and metadata selectively at the most appropriate levels (document level, collection level, or network level) rather than uniformly throughout the system. This targeted approach maintains effective document organization while minimizing unnecessary storage requirements by only retaining essential relation data where it provides maximum organizational value

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9026564B2Document processing system and control method thereof, program, and storage medium
Publication Date: 2015.05.05 CANON KK
  • US9026564B2 patent drawing
  • US9026564B2 patent drawing
  • US9026564B2 patent drawing

AI summary

This invention is directed to a document processing system and control method thereof. The system stores a plurality of items of document data each containing metadata pertaining to the contents of each item of document data, and relation information representing the relations between the plurality of items of document data. When scanned image data or facsimile-received image data is input, document data related to the input image data is specified among the plurality of items of stored document data, based on the metadata contained in each item of document data. Relation information representing the relation between the input image data and the specified related document data is stored. Even document data obtained from a paper document is able to be stored as document data subjected to search processing.