Aggregated Data Models for Mass Metric Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large-scale data processing systems generate vast amounts of log data, making it difficult to store and diagnose issues due to the massive quantity of metrics produced, which increases storage needs and complicates problem-solving processes.
Innovation Solution
The system aggregates metrics into data models that represent each time period, allowing for reduced storage space while maintaining system monitoring capabilities, where original metrics are discarded after aggregation, and the data models are stored with dimensions that can be dynamically expanded based on received metrics, providing flexible storage and control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all original metrics are stored for monitoring and diagnostics, then complete system monitoring capability is maintained, but storage requirements become prohibitively large
Solution Approach 1:
The patent extracts only the essential diagnostic information from the original metrics by creating aggregated data models. Instead of storing all raw metrics, the system extracts key performance indicators and system state information into condensed data models that retain monitoring capability while dramatically reducing storage requirements.
Solution Approach 2:
The patent inverts the traditional approach by not storing original metrics directly, but rather storing aggregated data models that represent the essential information. The system works backwards from the need for diagnostics to determine what minimum information is required, storing only that essential data in an inverted hierarchy where summaries precede details.
2Loss of information
If vast amounts of log data are generated to capture all system metrics, then comprehensive monitoring information is obtained, but data review and problem diagnosis become significantly more difficult
Solution Approach 1:
The patent segments the vast log data into organized data models grouped by system components, time periods, and metric types. This segmentation structures the information hierarchically, allowing operators to navigate from high-level system summaries to specific metric details only when needed, making data review manageable despite comprehensive coverage.
Solution Approach 2:
The patent adds dimensional organization to the data by structuring it in multiple hierarchical levels and categories. Data models are organized across dimensions such as system component, time period, metric type, and severity level, transforming the flat overwhelming log data into a multi-dimensional structured format that enables efficient navigation and analysis.
3Ease of manufacture
If fixed data storage structures are used for metrics, then storage organization is simple, but the system cannot adapt to new metric types or changing monitoring requirements
Solution Approach 1:
The patent implements dynamic data models that can adapt their structure based on the type of metric being stored and the monitoring requirements. The data model schema is not fixed but can be extended and modified to accommodate new metric types, system components, or changing diagnostic needs while maintaining organized storage through standardized model templates.
Data Source
AI summary
A plurality of data models are generated in a server from a stream of metrics describing a state of at least one system. Each of the data models represents a time grouping of a subset of the metrics. One or more dimensions are associated with each of the metrics. The data models are stored in association with respective ones of the dimensions in a memory. The dimensions with which the data models are associated in the memory are increased based upon an appearance of at least one previously non-existing dimension associated with a metric in the stream.


