Anomaly Detection in Machine Data via Predictive Text Indexing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Analyzing and searching massive quantities of machine-generated data from diverse sources is challenging due to its unstructured nature and varying formats, making it difficult to extract relevant information efficiently, especially in large datasets like terabytes, which hampers timely analysis and insights.
Innovation Solution
An event-based data intake and query system, such as the SPLUNKĀ® ENTERPRISE system, employs a late-binding schema and automated data generation using a deep-learning engine to process and index machine data, allowing for flexible schema development and anomaly detection, enabling efficient search and analysis across disparate data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional data analysis methods are used on unstructured machine data, then data processing can be performed, but analysis efficiency and speed deteriorate significantly when dealing with terabytes of data
Solution Approach 1:
The patent applies preliminary action by pre-processing and structuring machine data during the data intake phase. The system automatically parses, normalizes, and indexes data as it is ingested, preparing it for rapid retrieval and analysis later. This upfront preparation eliminates the need for time-consuming processing during analysis, directly resolving the contradiction between analysis efficiency and time consumption.
Solution Approach 2:
The patent replaces traditional mechanical data processing methods with automated computational systems. Machine learning algorithms and automated parsing systems substitute manual or conventional processing approaches, enabling the system to handle terabytes of unstructured data efficiently. This substitution dramatically improves productivity while reducing analysis time.
2Adaptability or versatility
If data is stored in unstructured formats to maintain flexibility, then data retrieval flexibility is preserved, but search and analysis speed deteriorates
Solution Approach 1:
The patent applies segmentation by dividing unstructured machine data into structured components during ingestion. The system parses data into discrete fields, extracts meaningful entities, and organizes them into indexed structures. This segmentation maintains adaptability for various query types while enabling rapid search through organized data segments, resolving the contradiction between flexibility and speed.
Solution Approach 2:
The patent changes the structural parameters of data from unstructured to semi-structured formats during processing. By transforming data into standardized schemas with defined fields and relationships, the system maintains retrieval flexibility through configurable query parameters while achieving fast search speeds through indexed access to structured data.
3Reliability
If comprehensive data processing is performed on all machine data, then complete analysis coverage is achieved, but processing complexity and resource requirements increase
Solution Approach 1:
The patent extracts only the most relevant and valuable data elements during processing, rather than analyzing every piece of data comprehensively. The system identifies and extracts key entities, events, and metrics that are most likely to provide actionable insights, discarding redundant information. This extraction approach maintains reliable analysis coverage of critical data while reducing processing complexity.
Solution Approach 2:
The patent applies partial action by performing comprehensive processing only on data that meets specific criteria or contains particular patterns. The system uses filtering and prioritization to apply full processing power selectively to high-value data subsets, achieving reliable coverage of critical information without the excessive complexity of processing all data uniformly.
Data Source
AI summary
Described herein is a technology that facilitates the production of and the use of automated datagens for event-based systems. A datagen (i.e., data-generator or data generation system) is a component, module, or subsystem of computer systems that searches, monitors, and analyzes machine data. Existing datagens are not capable of detecting an anomaly in machine data. An anomaly is a variance in the input data stream that exceeds some acceptable amount of deviation from the norm (i.e., standard, expectation, etc.). An embodiment of datagen, in accordance with the technology described herein, detects anomalies in the input machine data.


