Late-Binding Schema for Machine Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing and searching massive quantities of diverse machine data generated from various sources, such as system logs, network packets, and sensors, is time-consuming and challenging due to the vast types and formats of data, leading to inefficiencies in data retrieval and analysis.
Innovation Solution
A data intake and query system that uses a late-binding schema to process and store machine data as events, allowing flexible schema definition and extraction rules application at search time, enabling field-searchability and efficient retrieval of specific data items across disparate data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If massive quantities of diverse machine data are stored for later analysis, then data flexibility and insight potential are improved, but data retrieval and analysis time increase
Solution Approach 1:
The patent applies preliminary action by pre-processing machine data during ingestion to extract and store metadata, tags, and structured fields before analysis is needed. This advance preparation enables rapid retrieval and filtering operations later without requiring full data scanning, thus resolving the contradiction between storing comprehensive data and maintaining fast retrieval performance
Solution Approach 2:
The patent segments machine data into structured components (metadata, tags, fields) and unstructured portions, storing them in optimized formats suitable for different query types. This segmentation allows the system to efficiently handle diverse data types while maintaining fast retrieval performance for specific data elements
2Productivity
If pre-processing is applied to reduce data volume, then data retrieval efficiency is improved, but data analysis flexibility deteriorates
Solution Approach 1:
The patent applies local quality by applying different processing levels to different portions of data based on their intended use. Critical structured fields and metadata are pre-processed with high detail for fast retrieval, while less frequently accessed data maintains more original form, thus achieving both retrieval efficiency and analysis flexibility
Solution Approach 2:
The patent implements dynamic data processing where the level of pre-processing and schema application adapts based on query requirements. The system can apply schemas dynamically at search time rather than requiring fixed pre-processing, allowing flexibility to adjust processing depth based on actual analysis needs
3Loss of information
If diverse data types from multiple sources are analyzed, then insight potential is improved, but analysis complexity increases
Solution Approach 1:
The patent applies universality by creating a unified data model and common schema framework that can represent diverse machine data types from multiple sources using consistent structures. This universal approach enables diverse data to be analyzed together without requiring separate complex processing pipelines for each data type
Solution Approach 2:
The patent introduces metadata and tags as intermediary layers between raw diverse data and analysis queries. These intermediaries provide a standardized interface for querying diverse data types, reducing analysis complexity by abstracting away the heterogeneity of source data while preserving insight potential
Data Source
AI summary
A method comprises acquiring anomaly data including a plurality of anomalies detected from streaming data, wherein each of the anomalies relates to an entity on or associated with a computer network. The method determines a risk score of each of the anomalies, and adjusts the risk score of an anomaly according to a set of factors. The method further determines, for each of a plurality of sliding time windows of different lengths, an entity score of the entity in relation to the sliding time window, based on an aggregation of risk scores of all anomalies related to the entity that were detected within the sliding time window, where the entity score corresponds to a risk level associated with the entity. An action to prevent the entity from performing an operation can be determined and caused to occur based on the entity score.


