Late-binding schema for flexible machine data analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing and searching massive quantities of machine data generated from diverse sources in IT environments is challenging due to the vast amount of different types and formats of data, which is time-consuming and often requires pre-processing that discards significant amounts of data, limiting analysis flexibility.
Innovation Solution
An event-based data intake and query system that uses a late-binding schema to store and process machine data as events with flexible schema, allowing extraction rules to be applied at search time, enabling field-searchability and retention of minimally processed data for flexible analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If pre-processing is applied to reduce the amount of data, then data management becomes easier, but significant amounts of data are discarded limiting analysis flexibility
Solution Approach 1:
The system performs preliminary actions by extracting and indexing specific fields from machine data at ingestion time, creating a structured representation that enables efficient queries. This preliminary processing reduces the complexity of managing raw data while preserving the ability to analyze the original data when needed through the maintained mapping between extracted fields and source data.
Solution Approach 2:
The patent extracts key fields from machine data and stores them in a structured format for efficient retrieval and analysis. By taking out only the essential information needed for common queries while maintaining references to the original data, the system reduces data management complexity without sacrificing analysis flexibility for specialized queries.
2Adaptability or versatility
If all machine data is stored for later analysis, then analysis flexibility is improved, but the amount of data to manage becomes massive
Solution Approach 1:
The system extracts and stores only the essential fields from machine data in a structured format, significantly reducing the quantity of data that needs to be managed. The extracted fields include those most commonly queried, allowing efficient analysis while maintaining flexibility for specialized queries through the preserved mapping to original data.
Solution Approach 2:
The structured representation of extracted fields serves multiple functions: it enables efficient common queries, provides a simplified data model for management, and maintains references to original data for specialized analysis. This multi-functionality reduces the need to manage all raw data while preserving analysis flexibility.
3Productivity
If data is processed and structured immediately, then retrieval efficiency is improved, but data flexibility for later analysis is reduced
Solution Approach 1:
The system performs preliminary extraction and indexing of fields at data ingestion time, creating a structured representation that enables efficient retrieval. This preliminary action does not finalise the data structure but creates a maintainable mapping that allows later analysis flexibility when needed through queries against the structured data while preserving access to original data.
Data Source
AI summary
Security related anomalies in the data related to network entities are identified, and a risk score is assigned to each entity based on the anomalies. Visualization data is generated for a color-coded interactive visualization. Generating the visualization data includes assigning each entity to a separate polygon to be displayed concurrently on a display screen; selecting a size of each polygon to indicate one of: a number of security related anomalies associated with the entity, or a risk level assigned to the entity, where the risk level is based on the risk score of the entity, and selecting a color of each polygon to indicate the other one of: the number of security related anomalies associated with the entity, or the risk level assigned to the entity; and causing, the color-coded interactive visualization to be displayed on a display device based on the visualization data.


