Field Extraction Wizard for Dynamic Data Schemas

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in developing effective field extraction rules for machine-generated data, where the data format is often unknown at the time of collection, making it difficult to formulate and refine schemas and extraction rules, especially as the data is dynamic and diverse.

Innovation Solution

The technology introduces a wizard-guided process for formulating and refining field extraction rules, using example events and sampling tools to identify and validate field extraction, allowing users to create and save extraction rules for later use, and incorporating these rules into a late binding schema for query-time application.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If field extraction rules are developed for machine-generated data with unknown formats, then data extraction capability is improved, but the complexity of formulating and refining schemas increases

Engineering Contradiction:
Improvedata extraction capabilityVSAvoidschema formulation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically generating field extraction rules through sampling machine-generated data and identifying patterns. The extractor autonomously formulates schemas without requiring manual programming, thereby improving data extraction capability while reducing schema formulation complexity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback by continuously sampling extracted fields and using them to refine and improve extraction rules. This iterative process allows the system to adapt to diverse data formats while maintaining manageable complexity through automated learning from actual data patterns

Inventive Principle:
Principle #23Feedback

2Measurement precision

If extraction rules are refined using sampling tools and example events, then extraction accuracy is improved, but the time required for rule development increases

Engineering Contradiction:
Improveextraction accuracyVSAvoidrule development time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-sampling machine-generated data and pre-identifying extraction patterns before actual extraction operations. This advance preparation creates reusable extraction rules that improve accuracy while reducing the time needed for rule development during operational phases

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses copying by creating template extraction rules from sampled example events. Once a rule is developed from a sample, it can be copied and applied to similar data formats, improving extraction accuracy across multiple data types while reducing the time investment required for developing rules for each new data format

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10783318B2Facilitating modification of an extracted field
Publication Date: 2020.09.22 CISCO TECHNOLOGY INC
  • US10783318B2 patent drawing
  • US10783318B2 patent drawing
  • US10783318B2 patent drawing

AI summary

The technology disclosed relates to formulating and refining field extraction rules that are used at query time on raw data with a late-binding schema. The field extraction rules identify portions of the raw data, as well as their data types and hierarchical relationships. These extraction rules are executed against very large data sets not organized into relational structures that have not been processed by standard extraction or transformation methods. By using sample events, a focus on primary and secondary example events help formulate either a single extraction rule spanning multiple data formats, or multiple rules directed to distinct formats. Selection tools mark up the example events to indicate positive examples for the extraction rules, and to identify negative examples to avoid mistaken value selection. The extraction rules can be saved for query-time use, and can be incorporated into a data model for sets and subsets of event data.