ML Workload Identification for Storage Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional time series analysis approaches face accuracy and scalability challenges in identifying individual workloads contributing to performance anomalies in storage systems, due to their limitations in processing dynamic workload metrics.
Innovation Solution
Implementing a machine learning algorithm that calculates similarity measurements between primary and candidate time series using covariance, dynamic time warping, and shape-based distance metrics, and assigns weights to identify contributing workloads through a weighted majority vote process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional time series analysis approaches are used to identify workloads contributing to performance anomalies, then the analysis can be performed with simple methods, but the accuracy and scalability are insufficient
Solution Approach 1:
The patent introduces an intermediary machine learning system that mediates between the raw time series data and the identification of contributing workloads. This intermediary layer processes the complex patterns in time series data using multiple similarity measurements (covariance, dynamic time warping, shape-based distance) and weighted majority voting, thereby achieving high accuracy without requiring direct complex analysis of individual workload contributions.
Solution Approach 2:
The patent replaces conventional mechanical/time-based analysis methods with machine learning algorithms. Instead of using traditional statistical or rule-based approaches to analyze time series data, the system employs ML techniques including multiple similarity measurements and ensemble methods, substituting the mechanical analysis process with a more sophisticated computational approach that achieves both accuracy and scalability.
2Productivity
If conventional time series analysis approaches are used, then the system complexity remains low, but the scalability to handle multiple workloads and metrics is limited
Solution Approach 1:
The patent creates a universal machine learning framework that can handle multiple workloads, metrics, and time series data simultaneously. The system uses a unified approach with multiple similarity measurements and weighted majority voting that works across different scales and types of data, making it scalable from single-workload to multi-workload scenarios without requiring separate analysis systems for each case.
Solution Approach 2:
The patent employs multiple copies of similarity measurement mechanisms (covariance, dynamic time warping, shape-based distance) that work in parallel to analyze time series data. By using multiple independent similarity measurements that can be applied to any workload or metric, the system achieves scalability through replication of the analysis approach rather than through complex hierarchical structures.
3Measurement precision
If multiple similarity measurements are calculated using machine learning techniques, then the accuracy of identifying contributing workloads improves, but the computational complexity increases
Solution Approach 1:
The patent implements a weighted majority voting mechanism that allows partial contribution from multiple similarity measurements. Instead of requiring all measurements to reach a threshold, the system uses weighted votes where each similarity measurement contributes proportionally to the final identification. This partial action approach maintains accuracy while reducing computational burden compared to requiring all measurements to be processed to full precision.
Solution Approach 2:
The patent changes the parameter of similarity measurement from a single complex calculation to multiple simpler measurements with different weights. By transforming the problem from one highly accurate measurement to multiple measurements with varying importance, the system achieves comparable or better accuracy while distributing computational load across different types of calculations that can be optimized independently.
Data Source
AI summary
Methods, apparatus, and processor-readable storage media for automatic identification of workloads contributing to behavioral changes in storage systems using machine learning techniques are provided herein. An example computer-implemented method includes obtaining a primary time series and a set of candidate time series; calculating, using machine learning techniques, similarity measurements between the primary time series and each candidate time series in the set; for each similarity measurement, assigning weights to the candidate time series based on similarity values; generating, for each candidate time series, a similarity score based on the assigned weights; automatically identifying, based on the similarity scores, a candidate time series as contributing to an anomaly exhibited in the primary time series; and outputting identifying information of the at least one identified candidate time series for use in one or more automated actions associated with the storage system.


