Dynamic Hierarchical Clustering for Document Group Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document management systems face challenges in dynamically changing the hierarchical structure of items and database for new search perspectives, which is costly and difficult to implement, especially in advanced stages of operations, and are limited by automatic classification methods that restrict the types of items that can be used.
Innovation Solution
An information processing device performs hierarchical clustering of key phrases, divides them into candidate clusters, calculates utility scores for selected items, and extracts sub-items to present expansion images representing information volume, allowing for flexible classification and presentation of document groups without pre-defined item structures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the hierarchical structure of items and database is designed in advance for facet search, then the document classification and search functionality is established, but the cost and difficulty of changing the structure for new search perspectives increases substantially
Solution Approach 1:
The patent implements dynamic item structure where the hierarchical classification structure is not fixed in advance but can be changed flexibly at any time. The system allows users to add, delete, and modify items and their hierarchical relationships dynamically without requiring database restructuring, enabling adaptation to new search perspectives while maintaining system functionality.
Solution Approach 2:
The system performs preliminary clustering analysis of document groups to automatically generate candidate items and their hierarchical structures before user interaction. This preliminary action provides a ready-to-use classification framework that can be easily adjusted later, reducing the initial setup cost and complexity while maintaining adaptability.
2Ease of manufacture
If automatic classification using clustering is used to generate items, then the cost of designing item structure is reduced, but the types of items that can be used are restricted to those with hierarchy information described in documents
Solution Approach 1:
The patent introduces an intermediary mechanism between automatic clustering and item generation. The system performs clustering analysis to identify document groups, then uses user-defined templates and metadata to generate diverse item types including those without explicit hierarchy information in documents. This intermediary step expands the variety of usable item types beyond what pure clustering can provide.
Solution Approach 2:
The system implements a universal item structure that can accommodate multiple types of classification items (hierarchical, flat, metadata-based, user-defined) within a single framework. This multi-functional item structure allows the system to use various item types for different classification needs without being restricted to only hierarchical items with structure described in documents.
3Loss of information
If the information volume for each item is presented in detail, then the user can refer to document features accurately, but the complexity of presenting and processing the information increases
Solution Approach 1:
The patent segments the information volume presentation into hierarchical levels. The system divides document groups into item groups, then into individual items, with each level displaying aggregated information volume. This segmentation allows users to view information at appropriate detail levels without overwhelming complexity, maintaining accuracy while reducing presentation burden.
Solution Approach 2:
The system adds a visual dimension to information volume presentation using graphical representations such as bar charts, pie charts, or heat maps. Instead of presenting raw numerical data in tabular form, the system transforms the information into visual formats that convey document feature accuracy while simplifying the presentation complexity through intuitive visual cues.
Data Source
AI summary
An information processing device according to an embodiment includes one or more processors. The processors perform hierarchical clustering of a key phrase group. The processors divide the key phrase group into candidate clusters. The processors receive a selectin operation of one item from predetermined items for classifying the document group. The processors calculate, for each candidate cluster, a score indicating utility with respect to the selected item. The processors decide, as a reference cluster, a candidate cluster for which the score has a predetermined ranking. The processors divide the reference cluster into sub-clusters. The processors extract predetermined sub-items in the lower levels of the selected item. And the processors control presentation of an expansion image for expressing the information volume of the documents for each sub-item and each sub-cluster.


