Metric Data Compression via Linear Approximation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The exponential growth in complexity and data volume within distributed computing systems leads to a computational bottleneck in storing and processing metric data, necessitating more efficient methods for data storage and compression to facilitate automated administration and management.
Innovation Solution
The method involves approximating time-associated metric data with one or more linear functions, using a running variability metric to control variation and convert non-numeric data into numeric data for efficient compression, thereby reducing physical data storage overheads in configuration-management databases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If metric data is stored in traditional formats, then data completeness is maintained, but storage requirements and processing overhead increase exponentially
Solution Approach 1:
The patent transforms metric data from raw numerical values to a mathematical representation using linear functions characterized by parameters (slope, intercept, time points). This parameter transformation enables compact storage while preserving the essential characteristics of the original data through controlled approximation.
Solution Approach 2:
Instead of storing complete raw data and compressing it, the patent inverts the approach by storing only the essential parameters needed to reconstruct approximate data. The system stores minimal representation data (linear function parameters) rather than the full dataset, achieving compression by storing what is necessary rather than what exists.
2Measurement precision
If all metric data points are stored, then measurement precision is maintained, but device complexity and processing overhead increase
Solution Approach 1:
The patent applies partial action by storing only sufficient data to achieve acceptable approximation accuracy rather than complete precision. Linear functions are used to represent data segments, storing parameters for each segment rather than all individual data points, achieving a balance between accuracy and complexity.
Solution Approach 2:
The patent segments the metric data into multiple time-based segments, each represented by a linear function. This segmentation allows the system to manage complexity by breaking down large datasets into smaller, manageable segments that can be processed and stored independently with reduced overhead.
3Productivity
If data compression is applied to metric data, then storage efficiency improves, but data retrieval and interpretation complexity increases
Solution Approach 1:
The patent creates a simplified copy of the original metric data using linear function approximations. This copied representation maintains the essential trends and patterns of the original data while occupying minimal storage space. The linear function parameters serve as a compact copy that can be easily stored and transmitted.
Solution Approach 2:
The patent transforms the data into a different parameter space where storage is more efficient. By representing data as linear function parameters (slope, intercept, time boundaries) rather than raw values, the system achieves compact representation that is actually simpler to process mathematically than the original high-volume raw data.
Data Source
AI summary
The current document is directed to methods and subsystems within computing systems, including distributed computing systems that efficiently store metric data by approximating a sequence of time-associated data values with one or more linear functions. In a described implementation, a running variability metric is used to control variation within the metric data with respect to the approximating linear functions, with a variation threshold employed to maximize the number of data points represented by a given linear function while ensuring that the variation of the data with respect to the given linear function does not exceed a threshold value. In one implementation, the metric data occurs within a graph-like configuration-management-database representation of the current state of a computer system.


