Dynamic Hierarchical Clustering for Document Group Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document management systems face challenges in dynamically changing the hierarchical structure of items and database for new search perspectives, which is costly and difficult to implement, especially in advanced stages of operations, and are limited by automatic classification methods that restrict the types of items that can be used.

Innovation Solution

An information processing device performs hierarchical clustering of key phrases, divides them into candidate clusters, calculates utility scores for selected items, and extracts sub-items to present expansion images representing information volume, allowing for flexible classification and presentation of document groups without pre-defined item structures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the hierarchical structure of items and database is designed in advance for facet search, then the document classification and search functionality is established, but the cost and difficulty of changing the structure for new search perspectives increases substantially

Engineering Contradiction:
Improveadaptability to new search perspectivesVSAvoidcomplexity of changing hierarchical structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic item structure where the hierarchical classification structure is not fixed in advance but can be changed flexibly at any time. The system allows users to add, delete, and modify items and their hierarchical relationships dynamically without requiring database restructuring, enabling adaptation to new search perspectives while maintaining system functionality.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary clustering analysis of document groups to automatically generate candidate items and their hierarchical structures before user interaction. This preliminary action provides a ready-to-use classification framework that can be easily adjusted later, reducing the initial setup cost and complexity while maintaining adaptability.

Inventive Principle:
Principle #10Preliminary action

2Ease of manufacture

If automatic classification using clustering is used to generate items, then the cost of designing item structure is reduced, but the types of items that can be used are restricted to those with hierarchy information described in documents

Engineering Contradiction:
Improvecost of designing item structureVSAvoidvariety of usable item types
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary mechanism between automatic clustering and item generation. The system performs clustering analysis to identify document groups, then uses user-defined templates and metadata to generate diverse item types including those without explicit hierarchy information in documents. This intermediary step expands the variety of usable item types beyond what pure clustering can provide.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements a universal item structure that can accommodate multiple types of classification items (hierarchical, flat, metadata-based, user-defined) within a single framework. This multi-functional item structure allows the system to use various item types for different classification needs without being restricted to only hierarchical items with structure described in documents.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If the information volume for each item is presented in detail, then the user can refer to document features accurately, but the complexity of presenting and processing the information increases

Engineering Contradiction:
Improveaccuracy of document feature informationVSAvoidcomplexity of information presentation
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the information volume presentation into hierarchical levels. The system divides document groups into item groups, then into individual items, with each level displaying aggregated information volume. This segmentation allows users to view information at appropriate detail levels without overwhelming complexity, maintaining accuracy while reducing presentation burden.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system adds a visual dimension to information volume presentation using graphical representations such as bar charts, pie charts, or heat maps. Instead of presenting raw numerical data in tabular form, the system transforms the information into visual formats that convey document feature accuracy while simplifying the presentation complexity through intuitive visual cues.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10740378B2Method for presenting information volume for each item in document group
Publication Date: 2020.08.11 KK TOSHIBA
  • US10740378B2 patent drawing
  • US10740378B2 patent drawing
  • US10740378B2 patent drawing

AI summary

An information processing device according to an embodiment includes one or more processors. The processors perform hierarchical clustering of a key phrase group. The processors divide the key phrase group into candidate clusters. The processors receive a selectin operation of one item from predetermined items for classifying the document group. The processors calculate, for each candidate cluster, a score indicating utility with respect to the selected item. The processors decide, as a reference cluster, a candidate cluster for which the score has a predetermined ranking. The processors divide the reference cluster into sub-clusters. The processors extract predetermined sub-items in the lower levels of the selected item. And the processors control presentation of an expansion image for expressing the information volume of the documents for each sub-item and each sub-cluster.