Data Classification Model Using Cluster Parameter Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data classification methods face challenges in accurately classifying data with ambiguous labels, particularly when the label is represented by a degree of class membership, leading to low reliability in classification results.
Innovation Solution
An apparatus and method that cluster data vectors based on degrees of class membership, using optimized cluster parameters to generate a classification model, and adjust the influence of class membership to improve classification accuracy, incorporating a verifier to ensure the model's accuracy and adjust parameters as needed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If data is classified based on degree of class membership alone, then the classification process is simple, but the classification reliability is low
Solution Approach 1:
The patent segments the classification process into multiple stages: first clustering data based on degree of class membership, then performing secondary classification based on cluster characteristics and additional features. This segmentation allows the system to handle ambiguous labels by breaking down the complex classification task into manageable steps, improving reliability while maintaining operational simplicity.
Solution Approach 2:
The patent applies preliminary clustering action before final classification. By pre-grouping data points based on degree of class membership and then using this preliminary structure to guide subsequent classification, the system prepares the data in advance, making the final classification more reliable without adding significant complexity to the overall process.
2Reliability
If cluster parameters are optimized to improve classification accuracy, then classification reliability improves, but device complexity increases
Solution Approach 1:
The patent implements feedback mechanisms where the system evaluates classification results and adjusts cluster parameters accordingly. The verifier component provides feedback on classification accuracy, and the cluster parameter determiner uses this feedback to optimize parameters automatically. This feedback loop enables the system to improve accuracy while managing complexity through automated parameter adjustment.
Solution Approach 2:
The patent dynamically changes cluster parameters based on data characteristics and classification performance. The cluster parameter determiner adjusts parameters such as number of clusters, distance thresholds, and similarity metrics based on feedback from verification processes. This parameter adaptation allows the system to achieve high accuracy without requiring manually configured complex parameter sets.
Data Source
AI summary
Provided are an apparatus and method for classifying data and a system for collecting data. The method includes clustering vectors, each of which consists of at least one attribute value, for a plurality of pieces of target data including degrees of class membership and the vectors in view of the degrees of class membership, labeling the plurality of pieces of target data according to a result of the clustering, and generating a classification model using the labeled pieces of target data.


