Dual-Threshold Data Grouping for Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data grouping methods often result in inaccurate groupings due to errors in similarity calculations, leading to undesired results where data is incorrectly included or excluded from groups.
Innovation Solution
An information processing apparatus and method that employs two threshold values to accurately group data, where data with a degree of similarity greater than the first threshold value is included in a group and new representative data is selected from data with a degree of similarity less than the second threshold value, allowing for precise classification and representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single threshold value is used for data grouping, then the grouping process is simple, but the accuracy of grouping deteriorates due to errors in similarity calculation
Solution Approach 1:
The patent divides the single threshold value into two distinct threshold values: a first threshold value for determining group membership and a second threshold value for selecting representative data. This segmentation allows the system to handle similarity calculation errors by using different criteria for different purposes, thereby improving grouping accuracy without significantly increasing overall process complexity.
Solution Approach 2:
The patent changes the parameter structure by introducing a second threshold value that is lower than the first threshold value. This parameter change creates a buffer zone that accommodates similarity calculation errors, allowing the system to maintain high grouping accuracy while managing complexity through structured parameter differentiation.
2Quantity of substance
If a lower threshold value is used to include more data in groups, then the coverage of grouping improves, but the reliability of group classification deteriorates due to incorrect inclusions
Solution Approach 1:
The patent segments the threshold functionality into two parts: the first threshold value controls group membership inclusion to ensure reliability, while the second threshold value controls representative data selection to maintain coverage. This segmentation resolves the contradiction by assigning different roles to different threshold values.
Solution Approach 2:
The patent applies different quality standards locally: the first threshold value applies a stricter standard for group membership to ensure reliability, while the second threshold value applies a more lenient standard for representative data selection to maintain coverage. This local differentiation of quality standards resolves the contradiction between coverage and reliability.
Data Source
AI summary
An information processing apparatus (100) includes an input unit (102) that inputs a first threshold value and a second threshold value which are threshold values related to a degree of similarity of a feature value of each of a plurality of pieces of data, the first threshold value being for regarding data as belonging to an identical group and the second threshold value being smaller than the first threshold value, and a grouping unit (104) that groups the data by using the degree of similarity, the first threshold value, and the second threshold value, in which the grouping unit (104) causes data of which the degree of similarity with representative data is greater than the first threshold value to be included in the same group, and selects new representative data from among pieces of data of which the degree of similarity with the representative data already present is less than the second threshold value.


