Entropy-Based Anomaly Detection in Data Center Metric Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data centers with horizontally and vertically organized components generate vast amounts of metric data, and existing threshold-based methods struggle to effectively detect anomalies in dynamic environments with diverse workload patterns and online management of virtual machines.

Innovation Solution

The Entropy-based Anomaly Testing (EbAT) framework analyzes metric distributions across utility clouds, using entropy as a measurement to capture dispersal or concentration, generating entropy time series that are processed using tools like spike detection, signal processing, or subspace methods to identify anomalies at each hierarchy level.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If threshold-based methods are used for anomaly detection, then anomaly detection capability is provided, but the methods fail to effectively detect anomalies in dynamic environments with diverse workload patterns

Engineering Contradiction:
Improveanomaly detection capabilityVSAvoidadaptability to dynamic environments
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transforms the anomaly detection approach by changing from fixed threshold parameters to dynamic entropy-based parameters. Instead of using static threshold values that cannot adapt to changing conditions, the system calculates entropy values that dynamically reflect the current state of metric distributions, enabling reliable detection across diverse and dynamic workload patterns

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces entropy as an intermediary measurement that bridges the gap between raw metric data and anomaly detection. Rather than directly comparing metrics against fixed thresholds, the system uses entropy to capture the dispersal or concentration of metric distributions, providing an adaptive intermediate representation that works effectively in dynamic environments

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If entropy calculation is performed on raw metric data, then computational accuracy is maintained, but processing time and computational complexity increase significantly

Engineering Contradiction:
Improvecomputational accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary normalization and binning actions to metric data before entropy calculation. By pre-processing the data to scale values to a standard range and group them into bins, the system prepares the data in advance, reducing the computational burden during actual entropy calculation while maintaining the accuracy needed for effective anomaly detection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments continuous metric values into discrete bins before calculating entropy. This segmentation transforms continuous data into categorical distributions, simplifying the entropy calculation process while preserving the essential information about metric distribution patterns needed for accurate anomaly detection

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8843422B2Cloud anomaly detection using normalization, binning and entropy determination
Publication Date: 2014.09.23 HEWLETT PACKARD ENTERPRISE DEV LP
  • US8843422B2 patent drawing
  • US8843422B2 patent drawing
  • US8843422B2 patent drawing

AI summary

Illustrated is a system and method for anomaly detection in data centers and across utility clouds using an Entropy-based Anomaly Testing (EbAT), the system and method including normalizing sample data through transforming the sample data into a normalized value that is based, in part, on an identified average value for the sample data. Further, the system and method includes binning the normalized value through transforming the normalized value into a binned value that is based, in part, on a predefined value range for a bin such that a bin value, within the predefined value range, exists for the sample data. Additionally, the system and method includes identifying at least one vector value from the binned value. The system and method also includes generating an entropy time series through transforming the at least one vector value into an entropy value to be displayed as part of a look-back window.