Data Clustering via Human Supervision and Disentanglement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data clustering techniques are resource-intensive and often produce inaccurate results due to the number of analysis iterations required and unnecessary use of power, communication, processing, and storage resources, failing to optimally disentangle and classify unstructured heterogeneous data effectively.

Innovation Solution

A computer-implemented method that processes heterogeneous data using a data clustering algorithm to obtain an initial data cluster, then refines it through human supervision by comparing machine-obtained and human-obtained results, modifying the algorithm to improve accuracy and efficiency, and applying disentanglement matrices to subsequent data corpora for improved clustering and learning speed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional data clustering techniques are used to process heterogeneous data, then data can be organized into clusters, but the process requires excessive computational resources and produces inaccurate results due to multiple analysis iterations

Engineering Contradiction:
Improveclustering accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by performing human supervision and validation on a subset of data before full-scale clustering. Human analysts pre-validate clustering results on sample data to establish accurate ground truth, which then guides and improves automated clustering algorithms on the complete dataset, reducing the need for multiple resource-intensive iterations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where human supervision results are fed back into the clustering algorithm to continuously improve its performance. The system compares automated clustering results with human-validated results, uses the discrepancies to refine algorithm parameters, and iterates until convergence, thereby improving accuracy while reducing overall computational resource consumption.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If multiple analysis iterations are performed to improve clustering accuracy, then more accurate clusters are obtained, but resource consumption including power, communication, processing, and storage increases unnecessarily

Engineering Contradiction:
Improveclustering accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

Human supervision is performed preliminarily on a representative subset of data to establish accurate clustering ground truth before processing the complete dataset. This preliminary human validation creates a reference framework that guides subsequent automated processing, eliminating the need for multiple full-scale iterations and thereby improving processing efficiency while maintaining high accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by performing human supervision only on a subset of data rather than the entire dataset. This selective approach provides sufficient guidance for accurate clustering without the excessive resource consumption that would result from applying the same level of analysis to all data points, thus optimizing the balance between accuracy and productivity.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If conventional clustering algorithms process heterogeneous data without human supervision, then processing speed is maintained, but the ability to disentangle and classify unstructured data effectively is insufficient

Engineering Contradiction:
Improveprocessing speedVSAvoiddata classification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces human analysts as an intermediary between raw heterogeneous data and automated clustering algorithms. These human intermediaries validate and annotate sample data, creating a bridge that translates unstructured heterogeneous data into structured training examples that guide the automated algorithms, thereby improving classification accuracy without significantly compromising processing speed.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements self-service by using human-validated sample data to automatically train and refine clustering algorithms that then process the complete dataset independently. Once the algorithms are trained on human-validated samples, they autonomously perform high-speed clustering on the remaining data without requiring continuous human supervision, thus maintaining processing speed while improving accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11222059B2Data clustering
Publication Date: 2022.01.11 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11222059B2 patent drawing
  • US11222059B2 patent drawing
  • US11222059B2 patent drawing

AI summary

A set of data comprising heterogeneous data is processed in accordance with a data clustering algorithm so as to obtain an initial data cluster comprising homogeneous data. A supervised data cluster generated based on a human analysis of the set of data is obtained and compared with the initial data cluster to obtain a comparison result. The data clustering algorithm is modified based on the comparison result.