Representative Data Subset Selection for Unstructured Dataset Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large datasets, especially those containing unstructured machine-generated data, pose challenges in efficient analysis and extraction of field values due to their size and complexity, leading to difficulties in determining correct extraction rules and performing actions on the data.

Innovation Solution

A method for selecting a variable representative sampling of data as a subset from a larger dataset, using unsupervised clustering approaches to generate clusters and identify a sufficient number of cluster types, which allows for the creation of extraction rules and subsequent analysis, thereby conserving time and resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a larger dataset is analyzed to ensure comprehensive field value extraction, then extraction rule accuracy is improved, but analysis time and computational resources increase

Engineering Contradiction:
Improveextraction rule accuracyVSAvoidanalysis time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the large dataset into multiple smaller subsets through sampling. Each subset is processed independently to generate extraction rules, which are then validated against the full dataset. This segmentation allows comprehensive analysis to be performed on manageable portions of data while maintaining overall accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary sampling and analysis on a representative subset of the data before committing to full dataset processing. By identifying extraction rules early from the sample, the system can validate whether these rules work correctly before applying them to the entire dataset, saving time and computational resources.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If a larger dataset is analyzed to ensure comprehensive field value extraction, then extraction rule accuracy is improved, but computational resources increase

Engineering Contradiction:
Improveextraction rule accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent divides the computationally intensive task of analyzing the full dataset into smaller segments by processing sampled subsets. This reduces the immediate computational burden while still achieving accurate extraction rules through iterative validation against the complete dataset.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs analysis on a partial subset of the data that is sufficient to generate accurate extraction rules, rather than processing the entire dataset. This partial action approach consumes fewer computational resources while achieving the necessary accuracy for effective field value extraction.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of information

If unstructured data is processed directly to extract field values, then data completeness is improved, but processing complexity increases

Engineering Contradiction:
Improvedata completenessVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent performs preliminary processing on a sampled subset of the unstructured data to identify patterns and generate extraction rules. These rules are then applied to the complete dataset, avoiding the need to directly process and analyze every piece of unstructured data individually, thus reducing complexity while maintaining completeness.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If a representative subset is generated to reduce data complexity, then processing efficiency is improved, but data representativeness may be compromised

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddata representativeness
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where extraction rules generated from the representative subset are validated against the full dataset. This feedback loop ensures that the subset accurately represents the complete data, allowing the system to maintain high processing efficiency while verifying data representativeness through iterative validation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11775548B1Selection of representative data subsets from groups of events
Publication Date: 2023.10.03 CISCO TECHNOLOGY INC
  • US11775548B1 patent drawing
  • US11775548B1 patent drawing
  • US11775548B1 patent drawing

AI summary

Embodiments are directed towards generating a representative sampling as a subset from a larger dataset that includes unstructured data. A graphical user interface enables a user to provide various data selection parameters, including specifying a data source and one or more subset types desired, including one or more of latest records, earliest records, diverse records, outlier records, and/or random records. Diverse and/or outlier subset types may be obtained by generating clusters from an initial selection of records obtained from the larger dataset. An iteration analysis is performed to determine whether a sufficient number of clusters and/or cluster types have been generated that exceed at least one threshold and when not exceeded, additional clustering is performed on additional records. From the resultant clusters, and/or other subtype results, a subset of records is obtained as the representative sampling subset.