Weighted Clustering for Balanced Dataset Sampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The exponential growth of unstructured data in organizations poses scalability issues for processing, as training models on large datasets is challenging due to imbalanced domain-specific datasets, making it difficult to select representative subsets for classification and data management.
Innovation Solution
A method and system for content and context-aware data classification using weighted clustering and deep learning techniques to sample representative subsets from unstructured data, balancing datasets for efficient processing and classification, while prioritizing business-critical data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all documents in the data repositories are processed, then complete data coverage is achieved, but processing time and computational resources become unmanageable
Solution Approach 1:
The patent segments the large dataset into multiple subsets through clustering and sampling. The documents are divided into different clusters based on their characteristics, and then samples are selected from each cluster to form manageable subsets for processing, maintaining representativeness while reducing overall processing scope
Solution Approach 2:
Instead of processing all documents, the patent applies partial action by selecting representative samples from each cluster. This partial processing approach achieves sufficient data coverage for training purposes without the excessive computational burden of processing the entire dataset
2Productivity
If representative subsets are selected from the data, then processing efficiency improves, but data distribution balance deteriorates
Solution Approach 1:
The patent segments the data processing into two stages: first clustering all documents to identify different groups, then performing balanced sampling from each cluster. This segmentation ensures that the representative subset maintains the original data distribution across different document types and categories
Solution Approach 2:
The patent changes the sampling parameters by applying different sampling strategies to different clusters. By adjusting sampling rates and selection criteria based on cluster characteristics, the system maintains data distribution balance while achieving efficient processing of representative subsets
3Ease of manufacture
If domain-specific imbalanced datasets are used for training, then model training becomes feasible, but classification accuracy for minority classes deteriorates
Solution Approach 1:
The patent applies local quality by ensuring that each cluster in the sampled subset maintains the same quality characteristics and distribution as the original data. This localized preservation of data properties across different document types ensures that minority classes are adequately represented in the training samples, improving classification accuracy for all classes
Data Source
AI summary
Methods and systems for data management of documents in one or more data repositories in a computer network or cloud infrastructure are provided. The method includes sampling the documents in the one or more data repositories and formulating representative subsets of the sampled documents. The method further includes generating sampled data sets of the sampled documents and balancing the sampled data sets for further processing of the sampled documents. The formulation of the representative subsets is performed for identification of some of the representative subsets for initial processing.


