Cluster Analysis Method Using Time-Based Document Sets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for classifying and analyzing large numbers of documents, such as academic papers, are time-consuming and vary in accuracy due to human expertise, and lack the ability to effectively generate clusters based on different time axes and understand relationships between clusters.

Innovation Solution

A computer-based cluster analysis method that extracts document sets based on specific conditions, calculates inter-document and inter-cluster similarities, and generates association information to link relevant clusters across sets, enabling the classification of documents into clusters and understanding their relationships, including time-series relationships.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification of documents is performed by human workers, then documents can be classified by content, but the analysis takes time and accuracy varies depending on worker expertise

Engineering Contradiction:
Improveclassification accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical system of manual human document classification with an automated computer-based system. The system performs morphological analysis, calculates inter-document similarities using vector space models, and automatically clusters documents without human intervention, thereby eliminating time loss and ensuring consistent accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables documents to be classified automatically through self-service mechanisms. The computer autonomously performs similarity calculations, cluster formation, and even generates summaries of clustered documents without requiring human workers, making the classification process self-sufficient and efficient.

Inventive Principle:
Principle #25Self-service

2Productivity

If manual classification is performed by workers without specialized knowledge, then classification can be done quickly, but accuracy decreases due to lack of expertise

Engineering Contradiction:
Improveclassification speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent replaces human workers regardless of their expertise level with an automated computer system that objectively analyzes document content through morphological analysis and vector space modeling. This substitution ensures both high productivity and consistent accuracy without being constrained by human knowledge limitations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If traditional cluster analysis groups documents by similarity, then documents are classified into clusters, but the system cannot generate clusters based on different time axes or understand relationships between clusters

Engineering Contradiction:
Improvecluster analysis capabilityVSAvoidrelationship information between clusters
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent extends traditional cluster analysis by adding temporal dimensions. It extracts time information from documents, creates time-series displays of clusters, and enables analysis across multiple time axes. This dimensional extension allows the system to track how clusters evolve over time and understand relationships between clusters at different time points without losing temporal relationship information.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11989222B2Cluster analysis method, cluster analysis system, and cluster analysis program
Publication Date: 2024.05.21 AIXS INC
  • US11989222B2 patent drawing
  • US11989222B2 patent drawing
  • US11989222B2 patent drawing

AI summary

A server 4 executes a set extracting step (S1) of extracting a set from a plurality of documents according to a condition using time information, an inter-document similarity calculation step (S2) of calculating inter-document similarity between content of one document and content of another document included in the set, a cluster classifying step (S3) of classifying documents that are similar based on the inter-document similarity in the set into a plurality of clusters, an inter-cluster similarity calculation step (S6) of calculating inter-cluster similarity between clusters of a plurality of sets, and a cluster associating step (S7) of generating association information in which clusters that are relevant are linked to each other over sets based on the inter-cluster similarity.