Telemetry Data Mining for Virtual Machine Performance Diagnosis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in datacenters is the strain on resources due to the large volumes of telemetry data generated from virtual machines (VMs), which is difficult to store, process, and analyze effectively, especially given the high sample rates and the need to identify performance issues and similarities between VMs.
Innovation Solution
The system and method involve transforming raw telemetry data into fingerprints that characterize VM performance, using clustering techniques and statistical machine learning to identify similar VMs, diagnose performance issues, and group VMs based on resource consumption patterns, thereby providing insights into performance variations and anomalies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high sample rates are used to monitor VM performance, then measurement precision is improved, but data volume increases causing storage and processing strain
Solution Approach 1:
The patent extracts only the most relevant and distinctive features from the high-volume telemetry data using dimensionality reduction techniques. By identifying and extracting key performance indicators that characterize VM behavior patterns, the system maintains measurement precision while reducing data volume for storage and analysis.
Solution Approach 2:
The system transforms raw telemetry parameters into transformed parameters through statistical processing and dimensionality reduction. This parameter transformation converts high-dimensional raw data into a lower-dimensional feature space that preserves essential performance characteristics while reducing storage requirements.
2Loss of information
If comprehensive telemetry data is collected from all VMs, then analysis completeness is improved, but processing complexity increases
Solution Approach 1:
The patent segments the comprehensive telemetry data into individual VM-specific data sets, allowing independent processing and analysis of each virtual machine. This segmentation enables parallel processing and reduces overall system complexity while maintaining complete coverage of all VM performance metrics.
Solution Approach 2:
The system applies dimensionality reduction to transform complex multi-dimensional telemetry parameters into a simplified feature space. This parameter transformation maintains the essential information needed for complete performance analysis while reducing computational complexity for processing and comparison.
3Loss of information
If raw telemetry data is stored and processed without transformation, then data fidelity is improved, but storage and computational resources are strained
Solution Approach 1:
The patent performs preliminary transformation of raw telemetry data into a compressed feature representation before storage and analysis. By pre-processing the data to extract essential characteristics and reduce dimensionality, the system preserves critical information fidelity while significantly reducing the computational resources needed for subsequent processing and storage operations.
Data Source
AI summary
This disclosure is directed to systems and methods for mining streams of telemetry data in order to identify virtual machines (“VMs”), discover relationships between groups of VMs, and evaluate VM performance problems. The systems and methods transform streams of raw telemetry data consisting of resource usage and VM-related metrics into information that may be used to identify each VM, determine which VMs are similar based on their telemetry data patterns, and determine which VMs are similar based on their patterns of resource consumption. The similarity patterns can be used to group VMs that run the same applications and diagnose and debug VM performance.


