Multi-stage clustering for compliance data labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Evaluating large data sets for compliance with changing regulations is resource-intensive and prone to human error, as existing supervised classification algorithms require significant manual labeling and are not well-suited for varying data characteristics.

Innovation Solution

Multi-stage clustering techniques that iteratively group similar data elements into tightly knit clusters, using unsupervised learning models and similarity scoring to reduce manual effort and resource consumption, allowing for prioritization and auto-labeling of data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If supervised classification algorithms are used for compliance evaluation, then classification accuracy can be achieved, but manual labeling effort and resource consumption increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidmanual labeling effort
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing unsupervised clustering before supervised classification. The system pre-processes data by automatically grouping similar items into clusters using unsupervised learning models, which prepares the data structure in advance for more efficient supervised classification. This preliminary clustering reduces the manual labeling burden by organizing data into coherent groups that can be evaluated more systematically.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism by using unsupervised clustering as a bridge between raw data and supervised classification. The clustering process acts as an intermediate step that organizes data into clusters, which then serve as the input for supervised classification. This intermediary structure reduces the direct burden of manual labeling while maintaining classification accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If traditional clustering methods are used, then data grouping can be achieved, but computational resources and time consumption increase for large data sets

Engineering Contradiction:
Improvedata grouping capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The patent applies segmentation by dividing the large data set into multiple smaller clusters through iterative unsupervised learning. Instead of processing all data at once, the system segments the data into manageable groups in multiple stages, where each stage processes a portion of the data and produces intermediate clusters. This segmentation reduces computational resource consumption by breaking down the complex task of clustering large data sets into smaller, more efficient sub-tasks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamics by using iterative clustering where the clustering process is performed in multiple stages with progressively refined groupings. The system dynamically adjusts cluster formations across iterations, starting with broader groupings and progressively creating more granular clusters. This dynamic approach optimizes computational efficiency by adapting the clustering granularity to the specific characteristics of the data at each stage.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If manual compliance review is performed, then detailed assessment can be conducted, but human error and inconsistency increase

Engineering Contradiction:
Improveassessment detailVSAvoidconsistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies feedback by implementing an iterative clustering process where the results of each clustering stage are evaluated and used to inform subsequent clustering operations. The system incorporates feedback loops that assess cluster quality and adjust clustering parameters accordingly. This feedback mechanism reduces human error and inconsistency by systematically evaluating and refining cluster formations based on measured performance criteria rather than subjective human judgment.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11782955B1Multi-stage clustering
Publication Date: 2023.10.10 AMAZON TECH INC
  • US11782955B1 patent drawing
  • US11782955B1 patent drawing
  • US11782955B1 patent drawing

AI summary

Systems and techniques are generally described for multi-stage text-based clustering. In various examples, a first data set comprising a first set of elements may be received. Each element of the first set of elements may include text data. A clustering algorithm may be used to determine a first set of clusters of the first set of elements based at least in part on similarities in the text data. A second set of clusters, for each cluster of the first set of clusters, may be determined using the clustering algorithm. A number of resources for labeling the first set of elements may be determined based at least in part on the first set of clusters. The number of resources may be assigned to label clusters of the second set of clusters based at least in part on a number of elements in each cluster of the second set of clusters.