Customizable Data Streams and Processing Pipelines for Real-Time Machine Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current tools lack the capability to efficiently and intuitively search and analyze large sets of raw machine data from diverse sources, making it challenging to derive insights due to the complexity and volume of data generated in IT environments.
Innovation Solution
A data intake and query system that utilizes a flexible schema to process and store machine data as events, allowing for real-time search and analysis through a pipelined search language and query system, enabling users to define data streams and processing pipelines for efficient data routing and retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is pre-processed to extract specified data items for efficient retrieval, then retrieval efficiency is improved, but data flexibility and completeness are reduced
Solution Approach 1:
The system performs preliminary actions by ingesting and storing all raw machine data without pre-processing or filtering. The data intake system captures complete data streams from multiple sources and stores them in a searchable format, enabling both efficient retrieval and flexible analysis later without having to pre-select specific data items
Solution Approach 2:
The system implements dynamic data processing pipelines that can be configured and modified at runtime. Users can dynamically select different data sources, transformation operations, and destinations without reprocessing historical data, allowing the system to adapt to changing analysis requirements while maintaining efficient access to all raw data
2Adaptability or versatility
If massive quantities of raw data are stored for later retrieval, then data analysis flexibility is improved, but data management complexity increases
Solution Approach 1:
The system segments data management into distinct modular components: data intake, data storage, data processing pipelines, and query systems. Each component handles specific aspects of raw data management independently, reducing overall system complexity while enabling flexible analysis of massive data quantities
Solution Approach 2:
The patent introduces user-defined data streams as intermediary abstractions that simplify access to massive raw data. These data streams provide a standardized interface between the complex underlying data storage system and user queries, masking the complexity of managing massive data quantities while enabling flexible retrieval and analysis
3Productivity
If tools are used to search data systems separately and collect results over a network, then data retrieval capability is improved, but analysis efficiency and user experience are reduced
Solution Approach 1:
The system merges multiple separate data systems and search tools into a unified data intake and query system. It consolidates data ingestion, storage, processing, and search capabilities into a single integrated platform, eliminating the need to use separate tools and manually collect results over a network, thereby improving both retrieval capability and analysis efficiency
Data Source
AI summary
Systems and methods are described for customizable data streams in a streaming data processing system. Routing criteria for the customizable data streams are defined by a user, an automated process, or any other process. The routing criteria can be defined using graphical controls. The streaming data processing system uses the routing criteria to determine data that should be used to populate a particular data stream. Further, processing pipelines are customized such that a particular processing pipeline can obtain data from a particular user defined data stream and write data to a particular user defined data stream. Data is routed through the user defined data streams and customized processing pipelines based on a data route. A data route for a set of data may include multiple user defined data streams and multiple processing pipelines. The data route can include a loop of processing pipelines and data streams.


