Weighted Clustering for Balanced Dataset Sampling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The exponential growth of unstructured data in organizations poses scalability issues for processing, as training models on large datasets is challenging due to imbalanced domain-specific datasets, making it difficult to select representative subsets for classification and data management.

Innovation Solution

A method and system for content and context-aware data classification using weighted clustering and deep learning techniques to sample representative subsets from unstructured data, balancing datasets for efficient processing and classification, while prioritizing business-critical data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all documents in the data repositories are processed, then complete data coverage is achieved, but processing time and computational resources become unmanageable

Engineering Contradiction:
Improvedata coverage completenessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the large dataset into multiple subsets through clustering and sampling. The documents are divided into different clusters based on their characteristics, and then samples are selected from each cluster to form manageable subsets for processing, maintaining representativeness while reducing overall processing scope

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of processing all documents, the patent applies partial action by selecting representative samples from each cluster. This partial processing approach achieves sufficient data coverage for training purposes without the excessive computational burden of processing the entire dataset

Inventive Principle:
Principle #16Partial or excessive action

2Productivity

If representative subsets are selected from the data, then processing efficiency improves, but data distribution balance deteriorates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddata distribution balance
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

The patent segments the data processing into two stages: first clustering all documents to identify different groups, then performing balanced sampling from each cluster. This segmentation ensures that the representative subset maintains the original data distribution across different document types and categories

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the sampling parameters by applying different sampling strategies to different clusters. By adjusting sampling rates and selection criteria based on cluster characteristics, the system maintains data distribution balance while achieving efficient processing of representative subsets

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If domain-specific imbalanced datasets are used for training, then model training becomes feasible, but classification accuracy for minority classes deteriorates

Engineering Contradiction:
Improvemodel training feasibilityVSAvoidclassification accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies local quality by ensuring that each cluster in the sampled subset maintains the same quality characteristics and distribution as the original data. This localized preservation of data properties across different document types ensures that minority classes are adequately represented in the training samples, improving classification accuracy for all classes

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11675926B2Systems and methods for subset selection and optimization for balanced sampled dataset generation
Publication Date: 2023.06.13 DATHENA SCI PTE LTD
  • US11675926B2 patent drawing
  • US11675926B2 patent drawing
  • US11675926B2 patent drawing

AI summary

Methods and systems for data management of documents in one or more data repositories in a computer network or cloud infrastructure are provided. The method includes sampling the documents in the one or more data repositories and formulating representative subsets of the sampled documents. The method further includes generating sampled data sets of the sampled documents and balancing the sampled data sets for further processing of the sampled documents. The formulation of the representative subsets is performed for identification of some of the representative subsets for initial processing.