Document Tag Generation via OCR and LDA Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document storage and retrieval systems face challenges in efficiently organizing and searching large collections of documents, as file names do not adequately convey document content, making it difficult to locate specific documents, especially as the number of documents increases.

Innovation Solution

An image forming apparatus equipped with a control device that includes a character recognizer and a tag generator, which analyzes documents using OCR and LDA to generate tags representing the document's features, allowing for improved organization and retrieval by displaying tags alongside file names, with options to limit tag number and prioritize important tags.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If document collection size increases, then storage capacity is improved, but document searchability deteriorates

Engineering Contradiction:
Improvedocument collection sizeVSAvoiddocument searchability
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The system performs preliminary OCR recognition and tag generation on documents during storage, rather than when search is needed. This advance preparation of metadata (tags representing document features) enables fast retrieval even as document collection grows, resolving the contradiction between storage capacity and searchability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Tags serve as an intermediary between the full document content and the search interface. Instead of searching through all document text directly, the system uses generated tags as intermediate indicators that represent document features, making search efficient regardless of collection size

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If all document features are displayed, then information completeness is improved, but information overload increases

Engineering Contradiction:
Improveinformation completenessVSAvoidinformation overload
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system extracts only the most salient features from documents and displays them as tags, rather than showing all document content or metadata. This selective extraction maintains information completeness for search purposes while preventing information overload in the user interface

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Different documents receive different numbers of tags based on their characteristics and importance. The system applies local quality by varying the level of detail presented for different documents, showing more tags for important documents and fewer for less critical ones, optimizing the balance between completeness and overload

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11968342B2Image reading device capable of generating tag about document and image forming apparatus with the same
Publication Date: 2024.04.23 KYOCERA DOCUMENT SOLUTIONS INC
  • US11968342B2 patent drawing
  • US11968342B2 patent drawing
  • US11968342B2 patent drawing

AI summary

An image forming apparatus includes a storage device and a control device. The storage device stores a document. The control device includes a processor and functions, through the processor executing a control program, as a character recognizer and a tag generator. The character recognizer analyzes the document stored in the storage device and recognizes characters contained in the document. The tag generator analyzes, based on a recognition result of the character recognizer, a sequence of characters contained in the document and generates a tag expressing a feature of descriptive content of the document.