Telemetry Data Mining for Virtual Machine Performance Diagnosis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge in datacenters is the strain on resources due to the large volumes of telemetry data generated from virtual machines (VMs), which is difficult to store, process, and analyze effectively, especially given the high sample rates and the need to identify performance issues and similarities between VMs.

Innovation Solution

The system and method involve transforming raw telemetry data into fingerprints that characterize VM performance, using clustering techniques and statistical machine learning to identify similar VMs, diagnose performance issues, and group VMs based on resource consumption patterns, thereby providing insights into performance variations and anomalies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high sample rates are used to monitor VM performance, then measurement precision is improved, but data volume increases causing storage and processing strain

Engineering Contradiction:
ImproveVM performance monitoring precisionVSAvoidtelemetry data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the most relevant and distinctive features from the high-volume telemetry data using dimensionality reduction techniques. By identifying and extracting key performance indicators that characterize VM behavior patterns, the system maintains measurement precision while reducing data volume for storage and analysis.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transforms raw telemetry parameters into transformed parameters through statistical processing and dimensionality reduction. This parameter transformation converts high-dimensional raw data into a lower-dimensional feature space that preserves essential performance characteristics while reducing storage requirements.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If comprehensive telemetry data is collected from all VMs, then analysis completeness is improved, but processing complexity increases

Engineering Contradiction:
Improveperformance analysis completenessVSAvoiddata processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the comprehensive telemetry data into individual VM-specific data sets, allowing independent processing and analysis of each virtual machine. This segmentation enables parallel processing and reduces overall system complexity while maintaining complete coverage of all VM performance metrics.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies dimensionality reduction to transform complex multi-dimensional telemetry parameters into a simplified feature space. This parameter transformation maintains the essential information needed for complete performance analysis while reducing computational complexity for processing and comparison.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If raw telemetry data is stored and processed without transformation, then data fidelity is improved, but storage and computational resources are strained

Engineering Contradiction:
Improvetelemetry data fidelityVSAvoidcomputational resource consumption
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The patent performs preliminary transformation of raw telemetry data into a compressed feature representation before storage and analysis. By pre-processing the data to extract essential characteristics and reduce dimensionality, the system preserves critical information fidelity while significantly reducing the computational resources needed for subsequent processing and storage operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9213565B2Methods and systems for mining datacenter telemetry data
Publication Date: 2015.12.15 VMWARE INC
  • US9213565B2 patent drawing
  • US9213565B2 patent drawing
  • US9213565B2 patent drawing

AI summary

This disclosure is directed to systems and methods for mining streams of telemetry data in order to identify virtual machines (“VMs”), discover relationships between groups of VMs, and evaluate VM performance problems. The systems and methods transform streams of raw telemetry data consisting of resource usage and VM-related metrics into information that may be used to identify each VM, determine which VMs are similar based on their telemetry data patterns, and determine which VMs are similar based on their patterns of resource consumption. The similarity patterns can be used to group VMs that run the same applications and diagnose and debug VM performance.