Document Analysis Using Integrated Topic and Format Feature Values

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional document analysis technologies face challenges in achieving high accuracy, particularly when dealing with new topics or documents that require input of additional reference documents for background topic word extraction.

Innovation Solution

An information processing device extracts both topic and format information from the same document data, calculating feature values for each and integrating them to enhance analysis accuracy without needing additional documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If background topic words are extracted from reference documents, then document analysis accuracy is improved, but device complexity and processing time increase due to requiring additional documents

Engineering Contradiction:
Improvedocument analysis accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the necessary topic words from the input document itself, rather than requiring separate reference documents. The topic word extraction unit identifies and extracts topic words directly from the input document, eliminating the need for additional reference documents while maintaining analysis accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses the input document itself as the source for extracting topic words, making the document self-sufficient for the analysis process. The document provides its own topic words without needing external reference materials, simplifying the overall system architecture

Inventive Principle:
Principle #25Self-service

2Measurement precision

If multiple documents are processed to extract topic words, then analysis accuracy is improved, but processing time increases

Engineering Contradiction:
Improveanalysis accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts topic words directly from the input document in a single processing step, rather than requiring multiple documents to be processed separately. This direct extraction approach reduces processing time while maintaining accuracy by focusing on the essential topic words within the given document

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system extracts only the necessary topic words from the document rather than processing all content in detail. By selectively extracting only the relevant topic words needed for analysis, the system achieves accurate results more efficiently

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12118313B2Information processing device, information processing method, and computer program product
Publication Date: 2024.10.15 KK TOSHIBA
  • US12118313B2 patent drawing
  • US12118313B2 patent drawing
  • US12118313B2 patent drawing

AI summary

An information processing device includes at least one hardware processor. The hardware processor selects one or more pieces of partial document data from document data. The hardware processor extracts, from the partial document data, first information being a word or a phrase for specifying a first attribute of the partial document data. The hardware processor extracts, from the partial document data, second information being a word or a phrase for specifying a second attribute of the partial document data. The hardware processor calculates a first feature value representing a feature of the first information. The hardware processor calculates a second feature value representing a feature of the second information. The hardware processor analyzes the document data on the basis of the first feature value and the second feature value.