Late-Binding Schema for Dynamic Data Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in developing effective schema and extraction rules for machine-generated data, as the format of this data is often unknown at the time of collection, making it difficult to analyze and process efficiently.
Innovation Solution
The technology employs a late-binding schema and field extraction rules that are formulated and refined at query time, using a wizard-guided process to select and validate extraction rules, allowing for the extraction of values from raw machine data without prior transformation, and enabling the handling of diverse data formats through tools like Splunk Enterprise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If schema and extraction rules are formulated at data collection time, then data processing efficiency is improved, but the system cannot handle diverse and unknown data formats from machine-generated data
Solution Approach 1:
The patent implements dynamic schema formulation by delaying schema and extraction rule creation until query time rather than data collection time. The system adapts schemas based on the specific query requirements and data formats encountered, allowing flexible handling of diverse machine-generated data formats while maintaining processing efficiency through on-demand schema generation.
Solution Approach 2:
The system performs preliminary data collection and storage without formal schema definition, then formulates extraction rules when queries are executed. This preliminary action allows the system to collect diverse data formats initially, then apply appropriate extraction rules at query time based on the specific data formats and query requirements.
2Adaptability or versatility
If data is collected without determining format, then data collection flexibility is improved, but schema and extraction rule development becomes a moving target requiring continuous refinement
Solution Approach 1:
The system allows data to be collected in diverse formats without pre-defined schemas, then uses query-time analysis to automatically determine appropriate extraction rules. The system serves itself by generating schemas based on actual query requirements and data patterns, reducing the need for continuous manual schema refinement.
Solution Approach 2:
The system performs preliminary data collection in flexible formats, then formulates schemas and extraction rules at query time. This preliminary data collection without format determination allows maximum flexibility, while the subsequent query-time schema formulation addresses the complexity by generating rules based on actual needs rather than requiring continuous refinement.
3Adaptability or versatility
If extraction rules are formulated at query time, then adaptability to different data formats is improved, but data processing time increases
Solution Approach 1:
The system dynamically formulates extraction rules at query time based on the specific query requirements and data formats encountered. This dynamic approach allows the system to adapt to different data formats for each query while minimizing processing time by generating only the necessary extraction rules for the current query rather than pre-processing all possible formats.
Data Source
AI summary
The technology disclosed relates to formulating and refining field extraction rules that are used at query time on raw data with a late-binding schema. The field extraction rules identify portions of the raw data, as well as their data types and hierarchical relationships. These extraction rules are executed against very large data sets not organized into relational structures that have not been processed by standard extraction or transformation methods. By using sample events, a focus on primary and secondary example events help formulate either a single extraction rule spanning multiple data formats, or multiple rules directed to distinct formats. Selection tools mark up the example events to indicate positive examples for the extraction rules, and to identify negative examples to avoid mistaken value selection. The extraction rules can be saved for query-time use, and can be incorporated into a data model for sets and subsets of event data.


