Clustering Evaluation Score Calculation Method

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The evaluation score VRC tends to excessively classify larger-area or higher-density data groups, leading to an increase in the number of clusters, as it favors higher scores for these groups when they are not classified into one cluster but rather into multiple clusters.

Innovation Solution

A method for calculating an evaluation score based on the degree of internal compactness and external separation, where the degree of internal compactness is normalized by the number of data within each cluster, and the degree of external separation is normalized by the number of clusters, using specific index values to suppress excessive classification and determine the optimal number of clusters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the evaluation score VRC is used to determine the number of clusters, then the clustering evaluation can be performed, but larger-area or higher-density data groups are excessively classified into multiple clusters, increasing the number of clusters

Engineering Contradiction:
Improveclustering evaluation accuracyVSAvoidnumber of clusters
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent changes the parameters of the evaluation score by introducing a penalty term based on cluster area and data density. The modified evaluation score formula incorporates these parameters to suppress excessive classification of large-area or high-density data groups, thereby resolving the contradiction between evaluation accuracy and cluster number inflation.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If clustering is repeated by changing the number of clusters to determine optimal clusters, then the optimal number can be found, but the calculation complexity and time increase

Engineering Contradiction:
Improveoptimal cluster determination accuracyVSAvoidcalculation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-calculating or pre-estimating the penalty terms related to cluster area and data density before the main clustering evaluation. This allows the modified evaluation score to be computed more efficiently, reducing the time loss associated with repeated clustering calculations while maintaining determination accuracy.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If the evaluation score favors higher scores for larger-area or higher-density data groups when classified into multiple clusters, then the evaluation score increases, but the clustering quality deteriorates due to excessive classification

Engineering Contradiction:
Improveevaluation score valueVSAvoidclustering quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies preliminary anti-action by introducing a penalty mechanism that counteracts the tendency of the evaluation score to favor excessive classification of large-area or high-density data groups. The penalty term is designed to reduce the evaluation score when such excessive classification occurs, thereby preventing clustering quality deterioration while maintaining score differentiation capability.

Inventive Principle:
Principle #9Preliminary anti-action

Data Source

PatentUS11610083B2Method for calculating clustering evaluation value, and method for determining number of clusters
Publication Date: 2023.03.21 TOHOKU UNIV
  • US11610083B2 patent drawing
  • US11610083B2 patent drawing
  • US11610083B2 patent drawing

AI summary

Provided is a method for calculating an evaluation score of clustering quality, based on the number of clusters into which a plurality of data is clustered. The calculating the evaluation score includes: calculating a degree of internal compactness that is a sum of values, each being defined by normalizing a first index value by a first value that is based on a number of data within each cluster, the first index value indicating a degree of dispersion of data within each cluster; calculating a degree of external separation defined by normalizing a sum of a second index value for each cluster by a second value that is based on the number of clusters, the second index value indicating an index of a distance between the clusters; and calculating the evaluation score according to a predetermined formula having, as variables, the degree of internal compactness and the degree of external separation.