Data Classification Model Using Cluster Parameter Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data classification methods face challenges in accurately classifying data with ambiguous labels, particularly when the label is represented by a degree of class membership, leading to low reliability in classification results.

Innovation Solution

An apparatus and method that cluster data vectors based on degrees of class membership, using optimized cluster parameters to generate a classification model, and adjust the influence of class membership to improve classification accuracy, incorporating a verifier to ensure the model's accuracy and adjust parameters as needed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If data is classified based on degree of class membership alone, then the classification process is simple, but the classification reliability is low

Engineering Contradiction:
Improveclassification process simplicityVSAvoidclassification reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the classification process into multiple stages: first clustering data based on degree of class membership, then performing secondary classification based on cluster characteristics and additional features. This segmentation allows the system to handle ambiguous labels by breaking down the complex classification task into manageable steps, improving reliability while maintaining operational simplicity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary clustering action before final classification. By pre-grouping data points based on degree of class membership and then using this preliminary structure to guide subsequent classification, the system prepares the data in advance, making the final classification more reliable without adding significant complexity to the overall process.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If cluster parameters are optimized to improve classification accuracy, then classification reliability improves, but device complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidparameter optimization complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms where the system evaluates classification results and adjusts cluster parameters accordingly. The verifier component provides feedback on classification accuracy, and the cluster parameter determiner uses this feedback to optimize parameters automatically. This feedback loop enables the system to improve accuracy while managing complexity through automated parameter adjustment.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent dynamically changes cluster parameters based on data characteristics and classification performance. The cluster parameter determiner adjusts parameters such as number of clusters, distance thresholds, and similarity metrics based on feedback from verification processes. This parameter adaptation allows the system to achieve high accuracy without requiring manually configured complex parameter sets.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9582736B2Apparatus and method for classifying data and system for collecting data
Publication Date: 2017.02.28 SAMSUNG SDS CO LTD
  • US9582736B2 patent drawing
  • US9582736B2 patent drawing
  • US9582736B2 patent drawing

AI summary

Provided are an apparatus and method for classifying data and a system for collecting data. The method includes clustering vectors, each of which consists of at least one attribute value, for a plurality of pieces of target data including degrees of class membership and the vectors in view of the degrees of class membership, labeling the plurality of pieces of target data according to a result of the clustering, and generating a classification model using the labeled pieces of target data.