Cluster Analysis Method Using Time-Based Document Sets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for classifying and analyzing large numbers of documents, such as academic papers, are time-consuming and vary in accuracy due to human expertise, and lack the ability to effectively generate clusters based on different time axes and understand relationships between clusters.
Innovation Solution
A computer-based cluster analysis method that extracts document sets based on specific conditions, calculates inter-document and inter-cluster similarities, and generates association information to link relevant clusters across sets, enabling the classification of documents into clusters and understanding their relationships, including time-series relationships.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual classification of documents is performed by human workers, then documents can be classified by content, but the analysis takes time and accuracy varies depending on worker expertise
Solution Approach 1:
The patent replaces the mechanical system of manual human document classification with an automated computer-based system. The system performs morphological analysis, calculates inter-document similarities using vector space models, and automatically clusters documents without human intervention, thereby eliminating time loss and ensuring consistent accuracy.
Solution Approach 2:
The system enables documents to be classified automatically through self-service mechanisms. The computer autonomously performs similarity calculations, cluster formation, and even generates summaries of clustered documents without requiring human workers, making the classification process self-sufficient and efficient.
2Productivity
If manual classification is performed by workers without specialized knowledge, then classification can be done quickly, but accuracy decreases due to lack of expertise
Solution Approach 1:
The patent replaces human workers regardless of their expertise level with an automated computer system that objectively analyzes document content through morphological analysis and vector space modeling. This substitution ensures both high productivity and consistent accuracy without being constrained by human knowledge limitations.
3Adaptability or versatility
If traditional cluster analysis groups documents by similarity, then documents are classified into clusters, but the system cannot generate clusters based on different time axes or understand relationships between clusters
Solution Approach 1:
The patent extends traditional cluster analysis by adding temporal dimensions. It extracts time information from documents, creates time-series displays of clusters, and enables analysis across multiple time axes. This dimensional extension allows the system to track how clusters evolve over time and understand relationships between clusters at different time points without losing temporal relationship information.
Data Source
AI summary
A server 4 executes a set extracting step (S1) of extracting a set from a plurality of documents according to a condition using time information, an inter-document similarity calculation step (S2) of calculating inter-document similarity between content of one document and content of another document included in the set, a cluster classifying step (S3) of classifying documents that are similar based on the inter-document similarity in the set into a plurality of clusters, an inter-cluster similarity calculation step (S6) of calculating inter-cluster similarity between clusters of a plurality of sets, and a cluster associating step (S7) of generating association information in which clusters that are relevant are linked to each other over sets based on the inter-cluster similarity.


