Field Extraction Wizard for Dynamic Data Schemas
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in developing effective field extraction rules for machine-generated data, where the data format is often unknown at the time of collection, making it difficult to formulate and refine schemas and extraction rules, especially as the data is dynamic and diverse.
Innovation Solution
The technology introduces a wizard-guided process for formulating and refining field extraction rules, using example events and sampling tools to identify and validate field extraction, allowing users to create and save extraction rules for later use, and incorporating these rules into a late binding schema for query-time application.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If field extraction rules are developed for machine-generated data with unknown formats, then data extraction capability is improved, but the complexity of formulating and refining schemas increases
Solution Approach 1:
The system performs self-service by automatically generating field extraction rules through sampling machine-generated data and identifying patterns. The extractor autonomously formulates schemas without requiring manual programming, thereby improving data extraction capability while reducing schema formulation complexity
Solution Approach 2:
The system implements feedback by continuously sampling extracted fields and using them to refine and improve extraction rules. This iterative process allows the system to adapt to diverse data formats while maintaining manageable complexity through automated learning from actual data patterns
2Measurement precision
If extraction rules are refined using sampling tools and example events, then extraction accuracy is improved, but the time required for rule development increases
Solution Approach 1:
The system performs preliminary action by pre-sampling machine-generated data and pre-identifying extraction patterns before actual extraction operations. This advance preparation creates reusable extraction rules that improve accuracy while reducing the time needed for rule development during operational phases
Solution Approach 2:
The system uses copying by creating template extraction rules from sampled example events. Once a rule is developed from a sample, it can be copied and applied to similar data formats, improving extraction accuracy across multiple data types while reducing the time investment required for developing rules for each new data format
Data Source
AI summary
The technology disclosed relates to formulating and refining field extraction rules that are used at query time on raw data with a late-binding schema. The field extraction rules identify portions of the raw data, as well as their data types and hierarchical relationships. These extraction rules are executed against very large data sets not organized into relational structures that have not been processed by standard extraction or transformation methods. By using sample events, a focus on primary and secondary example events help formulate either a single extraction rule spanning multiple data formats, or multiple rules directed to distinct formats. Selection tools mark up the example events to indicate positive examples for the extraction rules, and to identify negative examples to avoid mistaken value selection. The extraction rules can be saved for query-time use, and can be incorporated into a data model for sets and subsets of event data.


