Aggregated Data Models for Mass Metric Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large-scale data processing systems generate vast amounts of log data, making it difficult to store and diagnose issues due to the massive quantity of metrics produced, which increases storage needs and complicates problem-solving processes.

Innovation Solution

The system aggregates metrics into data models that represent each time period, allowing for reduced storage space while maintaining system monitoring capabilities, where original metrics are discarded after aggregation, and the data models are stored with dimensions that can be dynamically expanded based on received metrics, providing flexible storage and control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all original metrics are stored for monitoring and diagnostics, then complete system monitoring capability is maintained, but storage requirements become prohibitively large

Engineering Contradiction:
Improvesystem monitoring capabilityVSAvoidstorage requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential diagnostic information from the original metrics by creating aggregated data models. Instead of storing all raw metrics, the system extracts key performance indicators and system state information into condensed data models that retain monitoring capability while dramatically reducing storage requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent inverts the traditional approach by not storing original metrics directly, but rather storing aggregated data models that represent the essential information. The system works backwards from the need for diagnostics to determine what minimum information is required, storing only that essential data in an inverted hierarchy where summaries precede details.

Inventive Principle:
Principle #13The other way round (Inversion)

2Loss of information

If vast amounts of log data are generated to capture all system metrics, then comprehensive monitoring information is obtained, but data review and problem diagnosis become significantly more difficult

Engineering Contradiction:
Improvemonitoring information completenessVSAvoiddata review difficulty
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments the vast log data into organized data models grouped by system components, time periods, and metric types. This segmentation structures the information hierarchically, allowing operators to navigate from high-level system summaries to specific metric details only when needed, making data review manageable despite comprehensive coverage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds dimensional organization to the data by structuring it in multiple hierarchical levels and categories. Data models are organized across dimensions such as system component, time period, metric type, and severity level, transforming the flat overwhelming log data into a multi-dimensional structured format that enables efficient navigation and analysis.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of manufacture

If fixed data storage structures are used for metrics, then storage organization is simple, but the system cannot adapt to new metric types or changing monitoring requirements

Engineering Contradiction:
Improvestorage organization simplicityVSAvoiddimension management flexibility
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic data models that can adapt their structure based on the type of metric being stored and the monitoring requirements. The data model schema is not fixed but can be extended and modified to accommodate new metric types, system components, or changing diagnostic needs while maintaining organized storage through standardized model templates.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8032797B1Storage of mass data for monitoring
Publication Date: 2011.10.04 AMAZON TECH INC
  • US8032797B1 patent drawing
  • US8032797B1 patent drawing
  • US8032797B1 patent drawing

AI summary

A plurality of data models are generated in a server from a stream of metrics describing a state of at least one system. Each of the data models represents a time grouping of a subset of the metrics. One or more dimensions are associated with each of the metrics. The data models are stored in association with respective ones of the dimensions in a memory. The dimensions with which the data models are associated in the memory are increased based upon an appearance of at least one previously non-existing dimension associated with a metric in the stream.