Document classification apparatus, method, and program

The document classification device enhances classification accuracy by constructing and updating vocabulary embedding spaces based on common vocabulary across logical elements, addressing the challenge of structural information handling in semi-structured documents.

JP7864604B2Active Publication Date: 2026-05-25KK TOSHIBA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP Β· JP
Patent Type
Patents
Current Assignee / Owner
KK TOSHIBA
Filing Date
2022-09-14
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Existing document classification methods struggle with semi-structured documents due to the inability to actively handle structural information of logical elements, leading to decreased classification performance and increased manual verification costs.

Method used

A document classification device that includes a logical element extraction unit, element selection unit, element analysis unit, and embedding space update unit to construct and update vocabulary embedding spaces based on common vocabulary across multiple logical elements, enhancing classification accuracy.

Benefits of technology

Improves classification performance by actively utilizing the characteristics of logical elements and maintaining unique vocabulary occurrence distributions, reducing manual verification costs and improving resolution and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007864604000001
    Figure 0007864604000001
  • Figure 0007864604000002
    Figure 0007864604000002
  • Figure 0007864604000003
    Figure 0007864604000003
Patent Text Reader

Abstract

To provide a document classification device, method, and program which perform document classification by adding a feature of a logical element to document data in a semi-structured format.SOLUTION: A document classification device 100 for analyzing and classifying an input document comprises: a logical element extraction unit; a logical element selection unit; an element analysis unit; an embedded space update unit; and a classification output unit. The logical element extraction unit acquires text contents for each logical element with respect to document data in a semi-structured format. The logical element selection unit generates a logical element set including one or more logical elements. The element analysis unit analyzes the text contents of a plurality of logical element sets and constructs a vocabulary embedded space. The embedded space update unit selects a first vocabulary embedded space and a second vocabulary embedded space having a common vocabulary and updates the first vocabulary embedded space on the basis of the similarity to the common vocabulary in the second vocabulary embedded space. The classification output unit outputs a classification result of the document data by using the first vocabulary embedded space.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art