Sampling-Based Storage Estimate for Data Intake Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users of data intake and query systems lack an accurate method to estimate the volume of new data sources and required storage space, leading to uncertainty and potential avoidance of adding valuable data sources due to concerns about exceeding storage limits.
Innovation Solution
A data estimation technique that allows users to specify a data source and sampling rate, acquiring and computing metadata to estimate storage requirements, indexing only a small portion of the data, and providing visual estimates of storage needs and license requirements, enabling users to predict appropriate license levels and transition data sources from estimated to fully indexed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If users store massive quantities of minimally processed machine data for later retrieval and analysis, then flexibility and analytical capability are improved, but storage requirements and system complexity increase
Solution Approach 1:
The patent applies partial action by implementing sampling-based data collection. Instead of indexing all machine data, the system collects and indexes only a sampled portion (e.g., 1% or 10% of events) from data sources. This sampled subset provides sufficient information for storage estimation and analysis while dramatically reducing storage requirements compared to full data retention.
2Productivity
If users pre-process machine data based on anticipated analysis needs, then data retrieval efficiency is improved, but data flexibility and potential analysis capabilities are reduced
Solution Approach 1:
The system performs preliminary action by collecting and storing sampled data in advance before full analysis is needed. The sampling process pre-processes data by extracting key events and metadata, storing them in an optimized format that enables fast retrieval. This preliminary sampling action balances efficiency (fast retrieval of sampled data) with flexibility (ability to perform multiple types of analysis on the sampled subset).
3Adaptability or versatility
If users add new data sources to the system, then data comprehensiveness and analytical value are improved, but uncertainty about storage requirements and licensing increases
Solution Approach 1:
The system implements self-service by automatically performing storage estimation when new data sources are added. The sampling mechanism automatically collects data from new sources, computes storage requirements based on the sampled subset, and provides licensing information without requiring manual calculation or external tools. This self-service estimation capability eliminates uncertainty about storage requirements while enabling comprehensive data source addition.
4Quantity of substance
If users discard portions of machine data during pre-processing, then storage requirements are reduced, but analytical capability and data flexibility are lost
Solution Approach 1:
The system applies copying by creating a sampled copy of the original machine data rather than discarding all data. The sampling process generates a representative subset (copy) of the full data set that retains the essential characteristics and analytical value needed for storage estimation and analysis. This copied sampled data provides sufficient analytical capability while dramatically reducing storage requirements compared to retaining all original data.
Data Source
AI summary
Disclosed herein is a data estimation technique for a data intake and query system. The system receives user inputs indicative that a first data source is to be the subject of a storage related estimate. The system receives a first plurality of events generated by the first data source. The system indexes only a sample of the received first plurality of events, based on a sampling criterion, where the sample is fewer than all of the first plurality of events. The system generates the storage related estimate based on at least some of the first plurality of events, and causes an indication of the estimate to be output to a user.


