Technology Add-On Content UI for Adaptive Event Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing and searching massive quantities of machine-generated data from diverse sources is time-consuming and challenging due to varying data types and formats, necessitating improved data intake and query systems that facilitate flexible schema application and efficient data retrieval.
Innovation Solution
The SPLUNKĀ® ENTERPRISE system employs a late-binding schema that applies extraction rules to event data during search time, enabling flexible schema development and efficient data processing, indexing, and retrieval, even from disparate data sources, using forwarders, indexers, and search heads to manage and analyze machine-generated data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional data intake systems process and index machine-generated data from diverse sources, then data retrieval becomes possible, but the process becomes time-consuming and complex due to varying data types and formats
Solution Approach 1:
The system performs preliminary actions by collecting and storing minimally processed machine data in its raw form before any analysis is requested. This allows the data to be readily available for immediate retrieval and analysis without requiring complex preprocessing operations at query time, thus improving productivity while managing complexity.
Solution Approach 2:
The data intake system is designed to handle multiple types and formats of machine-generated data from diverse sources through a universal collection mechanism. By creating a unified data model that can represent heterogeneous data sources, the system achieves multi-functionality without proportionally increasing complexity, enabling efficient retrieval across different data types.
2Productivity
If extraction rules are applied to machine data during ingestion, then data processing becomes efficient, but flexibility in schema development is reduced
Solution Approach 1:
The system implements dynamic schema application where extraction rules and data models can be modified, added, or refined at any time without reprocessing historical data. This dynamic approach allows the schema to adapt to changing analytical needs while maintaining efficient processing of existing data through the stored minimally processed format.
Solution Approach 2:
By storing minimally processed machine data with preserved original structure and metadata during the ingestion phase, the system performs a preliminary action that enables future schema developments. This preliminary preservation of data flexibility allows extraction rules to be applied or modified later without losing access to the original data characteristics.
3Adaptability or versatility
If minimally processed machine data is retained for later analysis, then comprehensive data analysis is enabled, but data storage requirements increase
Solution Approach 1:
The system extracts and preserves only the essential metadata and structural information needed for future analysis while retaining the minimally processed data in a compact format. This selective extraction approach enables comprehensive data analysis capability by preserving necessary data characteristics without storing unnecessary redundant information, thus managing storage requirements.
Solution Approach 2:
The system changes the parameter of data processing by storing data in a minimally processed state rather than fully processing it during ingestion. This parameter change allows the data to maintain its analytical value while occupying less storage space than fully processed and normalized data would require, enabling comprehensive analysis with reduced storage demands.
4Adaptability or versatility
If a late-binding schema is implemented for flexible schema development, then schema refinement is improved, but data processing time during search increases
Solution Approach 1:
The system performs preliminary actions by organizing and indexing the minimally processed data with preserved metadata during ingestion, so that when late-binding schema application is needed during search, the data is already in a state that facilitates quick filtering and retrieval. This preliminary organization reduces the time penalty associated with flexible schema refinement.
Solution Approach 2:
The late-binding schema implementation allows extraction rules to be dynamically applied or modified during search operations based on analytical needs. This dynamic approach enables schema refinement without requiring system reconfiguration, and the preserved minimally processed data structure allows these dynamic schema applications to execute efficiently with minimal time loss.
Data Source
AI summary
Determining a set of extraction rules include clustering event segments into at least a first group of event segments, and determining, using first field data in the first group of event segments, a first set of extraction rules for extracting the first field data from each event segment of the first group of event segments. A determination is made that the first set of extraction rules fails to successfully extract all of the first field data. Responsive to the determination, the event segments are re-clustered into at least a second group of event segments and a third group of event segments until a successful set of extraction rules are identified. The successful set of extraction rules are stored in computer memory.


