Cluster Stability Evaluation With Similarity-Based Subsampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current indicator values fail to accurately evaluate the stability and rationality of cluster results, leading to subjective perceptions of cluster quality, and lack of a unified reference for assessing cluster uncertainty.
Innovation Solution
A method and system for evaluating cluster stability by uniformly down-sampling raw data, calculating similarities using statistical tests, clustering sub-data, organizing cluster label models, and calculating a cluster stability indicator to provide a unified reference for evaluating cluster results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current indicator values (e.g., contour coefficients) are used to evaluate cluster results, then the evaluation process is simple, but the evaluation accuracy and reliability are insufficient
Solution Approach 1:
The patent segments the evaluation process into multiple independent modules: data sampling module, similarity calculation module, clustering module, and indicator calculation module. Each module processes a specific aspect of cluster evaluation independently, improving overall accuracy while maintaining manageable complexity through modular architecture.
Solution Approach 2:
The patent performs preliminary actions by uniformly down-sampling the raw data into multiple groups of sub-data before clustering. This pre-processing step ensures that the subsequent clustering and evaluation processes work with manageable data subsets, improving evaluation accuracy without proportionally increasing computational complexity.
2Reliability
If subjective perception is used to judge cluster quality, then the evaluation is flexible and adaptable, but the evaluation lacks objectivity and consistency
Solution Approach 1:
The patent implements a feedback mechanism where the cluster stability indicator is calculated based on the similarity between original data and sub-data, providing an objective feedback signal that reflects cluster quality. This feedback loop enables consistent, repeatable evaluations that are not dependent on subjective human judgment, while the standardized calculation process remains easy to operate.
3Adaptability or versatility
If multiple cluster algorithms are applied to the same data, then the analysis comprehensiveness is improved, but the comparison and evaluation become more difficult
Solution Approach 1:
The patent creates a universal evaluation framework that can accommodate multiple cluster algorithms through a common set of tools: uniform down-sampling, similarity calculation, and stability indicator computation. This universal framework allows different algorithms to be evaluated using the same standardized process, improving adaptability while simplifying comparison through consistent evaluation criteria.
Data Source
AI summary
The present invention is an indicator evaluation method of a cluster stability, which is executed by an indicator evaluation system of a cluster stability. The indicator evaluation system includes a processing device. The processing device uniformly down-samples a raw data to be clustered to generate sub-data. The processing device calculates similarities of the sub-data according to a statistical test, and keeps the sub-data with the similarities greater than a similarity threshold as sub-data to be analyzed. The processing device clusters the sub-data to be analyzed to generate sub-data cluster results. The processing device organizes cluster label models of the sub-data cluster results, and generates organized sub-data cluster results according to organized cluster label models. The processing device further calculates a cluster stability indicator according to the organized sub-data cluster results. The present invention provides a reference indicator for evaluating stability of cluster results, and misleading results can be reduced.


