Data Classification by Cluster Merging and Bias Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing clustering methods fail to provide relationships between data clusters, hiding real reasons for special situations in industrial systems, and thus hinder effective data analysis.
Innovation Solution
A method and apparatus that utilize neural network models to determine bias degrees of classification and re-classification, merging related data clusters to improve accuracy and uncover hidden links between clusters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If existing clustering methods are used to group data, then data can be divided into multiple clusters, but the relationships between clusters cannot be provided and real reasons of special situations cannot be learned
Solution Approach 1:
The patent merges multiple data clusters into a super-cluster when they share common characteristics or relationships. By combining related clusters, the system preserves relationship information while reducing the complexity of analyzing individual cluster relationships. The merging process identifies overlapping features and consolidates clusters that should be analyzed together, thus preventing information loss about cluster relationships.
Solution Approach 2:
The patent introduces an intermediary mechanism that analyzes and identifies relationships between clusters. This intermediary layer examines feature overlaps, similarity metrics, and contextual connections between clusters to determine which clusters should be merged. It acts as a mediator that preserves relationship information by systematically evaluating cluster connections before merging decisions are made.
2Productivity
If multiple data clusters are obtained through clustering, then data analysis can be performed on clusters, but the number of clusters increases computation load and hides real reasons
Solution Approach 1:
The patent applies merging by consolidating multiple related data clusters into fewer super-clusters based on shared characteristics. This reduction in the number of clusters directly decreases computation load while maintaining the ability to perform meaningful data analysis. The merging process preserves essential information by grouping clusters that have relationships, thus improving productivity without losing critical insights.
Solution Approach 2:
The patent segments the analysis process into two stages: first performing initial clustering to identify potential groups, then applying a second-stage merging process to combine related clusters. This segmented approach allows the system to initially capture detailed cluster structures and then consolidate them to reduce computation load, achieving both thorough analysis and efficiency.
3Measurement precision
If clustering is performed without considering relationships between clusters, then clustering can be completed quickly, but real reasons of special situations are hidden
Solution Approach 1:
The patent performs preliminary action by identifying and marking potential cluster relationships during the initial clustering process. Instead of performing exhaustive relationship analysis after clustering, the system pre-identifies clusters that likely have relationships based on feature similarity and proximity metrics. This preliminary identification reduces the time required for subsequent relationship analysis while maintaining measurement precision.
Solution Approach 2:
The patent replaces the mechanical approach of exhaustive pairwise cluster comparison with a more efficient system that uses feature-based similarity metrics and predefined relationship criteria. By substituting the brute-force mechanical comparison with a smarter system that leverages cluster features and characteristics, the patent reduces analysis time while preserving the accuracy needed to identify real reasons behind special situations.
Data Source
AI summary
A method and an apparatus are for classifying data. In an embodiment, the method includes: classifying at least two pieces of data, to obtain at least two data clusters; determining a bias degree of classification; re-classifying the at least two pieces of data by merging any several of the at least two data clusters; determining a bias degree of re-classification; and determining, by comparing the bias degree of first classification and the bias degree of re-classification, which classification is more accurate. By way of the method and apparatus, a related data cluster can be found from multiple data clusters for better data analysis.


