Information Processing Device Supervised Clustering Data Homogeneity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Classical methods for data analysis, such as the Mahalanobis-Taguchi method, require high-quality data but lack techniques to improve data quality and do not function well without specific knowledge of the task, especially when large amounts of labeled data are unavailable.
Innovation Solution
An information processing device and method that uses supervised clustering to determine the homogeneity of a data set by extracting features from digital data and applying labels unrelated to homogeneity, allowing for the assessment of clustering possibility and identifying data inhomogeneities without requiring specific task knowledge.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If classical methods (e.g., Mahalanobis-Taguchi method) are used for data analysis, then the method requires only a small amount of data for learning, but the method does not function unless the quality of the data is high and lacks techniques to improve data quality
Solution Approach 1:
The patent changes the parameter of data quality assessment by introducing homogeneity determination through supervised clustering. Instead of assuming high data quality, the system evaluates homogeneity using clustering possibility with labels unrelated to homogeneity, transforming the approach from quality assumption to quality evaluation through parameter transformation.
Solution Approach 2:
The patent introduces an intermediary mechanism (supervised clustering with unrelated labels) to assess data homogeneity. This intermediary process acts as a mediator between the raw data and the analysis task, providing objective feedback on data quality without requiring specific task knowledge or manual quality assessment.
2Adaptability or versatility
If classical methods are used, then the method requires specific knowledge of the task to be performed, but this limits the generality and applicability of the method
Solution Approach 1:
The patent applies universality by designing a homogeneity determination method that works across different tasks and data types. The supervised clustering approach with unrelated labels serves as a universal assessment mechanism that does not require task-specific knowledge, making the method adaptable to various applications while maintaining simplicity.
3Measurement precision
If deep learning methods are used, then the system can automatically find latent structures and achieve high generalization performance, but the system does not function in situations in which a large amount of labeled data is unavailable
Solution Approach 1:
The patent applies partial action by using only a small amount of labeled data (unrelated to homogeneity) for the homogeneity assessment task, rather than requiring large amounts of task-specific labeled data. This partial use of labels enables the system to evaluate data quality without the excessive data requirements of deep learning methods.
Data Source
AI summary
An information processing device (100) includes a memory and processing circuitry. The memory stores a data set (DG) including multiple items of digital data (DD) and a label set (RG) including multiple labels. Each of the multiple labels are added to each of the multiple items of digital data (DD). The processing circuitry generates a feature vector set (BG) by extracting a predetermined feature from each of the multiple items of digital data (DD) and generating feature vectors indicating the extracted features. The feature vector set includes the feature vectors. The processing circuitry determines homogeneity of the data set (DG) by performing a trial of supervised clustering on the feature vector set (BG) by using the label set (RG) and determining the possibility of the clustering.


