Representative Data Subset Selection for Unstructured Dataset Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large datasets, especially those containing unstructured machine-generated data, pose challenges in efficient analysis and extraction of field values due to their size and complexity, leading to difficulties in determining correct extraction rules and performing actions on the data.
Innovation Solution
A method for selecting a variable representative sampling of data as a subset from a larger dataset, using unsupervised clustering approaches to generate clusters and identify a sufficient number of cluster types, which allows for the creation of extraction rules and subsequent analysis, thereby conserving time and resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a larger dataset is analyzed to ensure comprehensive field value extraction, then extraction rule accuracy is improved, but analysis time and computational resources increase
Solution Approach 1:
The patent segments the large dataset into multiple smaller subsets through sampling. Each subset is processed independently to generate extraction rules, which are then validated against the full dataset. This segmentation allows comprehensive analysis to be performed on manageable portions of data while maintaining overall accuracy.
Solution Approach 2:
The patent performs preliminary sampling and analysis on a representative subset of the data before committing to full dataset processing. By identifying extraction rules early from the sample, the system can validate whether these rules work correctly before applying them to the entire dataset, saving time and computational resources.
2Measurement precision
If a larger dataset is analyzed to ensure comprehensive field value extraction, then extraction rule accuracy is improved, but computational resources increase
Solution Approach 1:
The patent divides the computationally intensive task of analyzing the full dataset into smaller segments by processing sampled subsets. This reduces the immediate computational burden while still achieving accurate extraction rules through iterative validation against the complete dataset.
Solution Approach 2:
The patent performs analysis on a partial subset of the data that is sufficient to generate accurate extraction rules, rather than processing the entire dataset. This partial action approach consumes fewer computational resources while achieving the necessary accuracy for effective field value extraction.
3Loss of information
If unstructured data is processed directly to extract field values, then data completeness is improved, but processing complexity increases
Solution Approach 1:
The patent performs preliminary processing on a sampled subset of the unstructured data to identify patterns and generate extraction rules. These rules are then applied to the complete dataset, avoiding the need to directly process and analyze every piece of unstructured data individually, thus reducing complexity while maintaining completeness.
4Productivity
If a representative subset is generated to reduce data complexity, then processing efficiency is improved, but data representativeness may be compromised
Solution Approach 1:
The patent implements feedback mechanisms where extraction rules generated from the representative subset are validated against the full dataset. This feedback loop ensures that the subset accurately represents the complete data, allowing the system to maintain high processing efficiency while verifying data representativeness through iterative validation.
Data Source
AI summary
Embodiments are directed towards generating a representative sampling as a subset from a larger dataset that includes unstructured data. A graphical user interface enables a user to provide various data selection parameters, including specifying a data source and one or more subset types desired, including one or more of latest records, earliest records, diverse records, outlier records, and/or random records. Diverse and/or outlier subset types may be obtained by generating clusters from an initial selection of records obtained from the larger dataset. An iteration analysis is performed to determine whether a sufficient number of clusters and/or cluster types have been generated that exceed at least one threshold and when not exceeded, additional clustering is performed on additional records. From the resultant clusters, and/or other subtype results, a subset of records is obtained as the representative sampling subset.


