Time-Windowed Statistical Compression for Network Telemetry Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The exponential growth of data storage demands an efficient and cost-effective method to process and store large volumes of data, as organizations face increasing storage costs with petabytes and exabytes of data, necessitating innovative compression techniques.
Innovation Solution
Lossy statistical data compression encapsulates time-bounded data into mathematically modeled empirical functions, allowing for the deletion of redundant data and reducing storage needs by representing data through statistical empirical distribution analysis, performed in real-time or on-demand, using time windows and variance thresholds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data is stored in its original form, then data completeness and accuracy are maintained, but storage costs and storage space requirements increase exponentially
Solution Approach 1:
The patent extracts only the essential statistical characteristics and patterns from the original data, separating the valuable analytical information from the redundant raw data. By performing statistical empirical distribution analysis on time-windowed data segments, the system extracts distribution parameters and patterns that capture the essential behavior of the data without requiring storage of every individual data point.
Solution Approach 2:
The patent transforms the data from its original raw form into statistical parameters and empirical distribution functions. By changing the representation from individual data points to aggregated statistical measures (such as mean, variance, and empirical distribution functions), the system maintains the essential information needed for analysis while dramatically reducing the quantity of data that must be stored.
2Measurement precision
If all empirical data is retained for analysis, then analytical precision and insight quality are maintained, but processing time and computational resources increase
Solution Approach 1:
The patent divides the continuous stream of empirical data into discrete time windows or segments. By segmenting the data temporally, the system can perform statistical analysis on manageable chunks rather than processing the entire dataset at once. This segmentation enables incremental processing and reduces the computational burden while maintaining analytical precision through consistent statistical methods applied to each segment.
Solution Approach 2:
The patent performs preliminary statistical analysis on data as it arrives in time windows, computing empirical distribution functions and statistical parameters in advance. This preliminary processing prepares the data for future analysis by pre-computing aggregated statistics, so that when analytical queries are executed, the system can quickly retrieve pre-computed results rather than reprocessing raw data, significantly reducing query response time.
3Quantity of substance
If data is compressed using traditional methods, then storage efficiency improves, but data quality and analytical value deteriorate
Solution Approach 1:
The patent creates statistical copies or representations of the original data through empirical distribution functions. Rather than compressing raw data points, the system creates mathematical models that copy the essential statistical behavior and patterns of the original data. These empirical distribution function copies can be stored efficiently while accurately representing the underlying data characteristics needed for analytical purposes.
Data Source
AI summary
A method performed in real-time includes receiving and storing time-based data over a specific time period and dividing the specific time period into a plurality of time windows. The method further includes determining that data associated with two or more proximate time windows are within a predetermined variance of one another and responsive to the determination: generating a mathematical function representative of the data associated with the two or more proximate time windows, deleting the data associated with the two or more proximate time windows, and generating a representation of the deleted data from the mathematical function. In certain embodiments, the data comprises empirical network telemetry data.


