Event-Based Data Intake System With Flexible Schema
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern data centers face challenges in analyzing and searching massive quantities of machine-generated data due to its unstructured nature and diverse formats, which complicates the application of semantic meaning and efficient processing.
Innovation Solution
An event-based data intake and query system with a flexible schema, allowing for late-binding schema application during search time, enables the collection, indexing, and retrieval of machine data from various sources, using extraction rules and configuration files to identify and extract relevant information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a rigid pre-defined schema is used for data collection and indexing, then data processing efficiency is improved, but adaptability to diverse data formats from multiple sources deteriorates
Solution Approach 1:
The patent implements a dynamic schema system that evolves from rigid pre-defined schemas to flexible late-binding schemas. The schema is applied dynamically during search time rather than being fixed during data collection, allowing the system to adapt to diverse data formats while maintaining processing efficiency through indexed fields and extracted information.
2Adaptability or versatility
If unstructured machine-generated data is collected without pre-defined formats, then adaptability to diverse sources is improved, but difficulty in applying semantic meaning and performing indexing operations increases
Solution Approach 1:
The patent performs preliminary extraction of structured fields from unstructured data during the indexing phase. By extracting and organizing data into standardized fields (host, source, event type, etc.) before search operations, the system prepares semantic meaning in advance, making subsequent searching and analysis operations efficient despite the diverse origins of the data.
3Speed
If data is indexed and searched using traditional methods, then processing speed is maintained, but ability to perform intelligent analysis and derive insights from large volumes of data deteriorates
Solution Approach 1:
The patent introduces an intermediary search head component that coordinates between data sources, indexers, and users. The search head receives queries, distributes them to appropriate indexers, collects results, and presents them to users. This intermediary layer enables intelligent analysis by orchestrating complex search operations across multiple data sources while maintaining processing speed through distributed architecture.
Data Source
AI summary
Machine data reflecting operation of a monitored system is ingested and made available for search by a data intake and query system (DIQS). Monitoring includes obtaining a subset of ordered events that are assigned to a task. In a graphical user interface on a display, a chart for the task is displayed. The chart includes an event identifier for each event of the subset of the ordered events, a confidence level value related to each event identifier of each event of the subset of ordered events, the confidence level value indicating the confidence level that the event is in the task. The chart further includes a time reference value identifying a time of each event.


