ML Workload Identification for Storage Anomaly Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional time series analysis approaches face accuracy and scalability challenges in identifying individual workloads contributing to performance anomalies in storage systems, due to their limitations in processing dynamic workload metrics.

Innovation Solution

Implementing a machine learning algorithm that calculates similarity measurements between primary and candidate time series using covariance, dynamic time warping, and shape-based distance metrics, and assigns weights to identify contributing workloads through a weighted majority vote process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional time series analysis approaches are used to identify workloads contributing to performance anomalies, then the analysis can be performed with simple methods, but the accuracy and scalability are insufficient

Engineering Contradiction:
Improveaccuracy of identifying contributing workloadsVSAvoidcomplexity of analysis system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary machine learning system that mediates between the raw time series data and the identification of contributing workloads. This intermediary layer processes the complex patterns in time series data using multiple similarity measurements (covariance, dynamic time warping, shape-based distance) and weighted majority voting, thereby achieving high accuracy without requiring direct complex analysis of individual workload contributions.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces conventional mechanical/time-based analysis methods with machine learning algorithms. Instead of using traditional statistical or rule-based approaches to analyze time series data, the system employs ML techniques including multiple similarity measurements and ensemble methods, substituting the mechanical analysis process with a more sophisticated computational approach that achieves both accuracy and scalability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If conventional time series analysis approaches are used, then the system complexity remains low, but the scalability to handle multiple workloads and metrics is limited

Engineering Contradiction:
Improvescalability of analyzing multiple workloadsVSAvoidcomplexity of processing system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates a universal machine learning framework that can handle multiple workloads, metrics, and time series data simultaneously. The system uses a unified approach with multiple similarity measurements and weighted majority voting that works across different scales and types of data, making it scalable from single-workload to multi-workload scenarios without requiring separate analysis systems for each case.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs multiple copies of similarity measurement mechanisms (covariance, dynamic time warping, shape-based distance) that work in parallel to analyze time series data. By using multiple independent similarity measurements that can be applied to any workload or metric, the system achieves scalability through replication of the analysis approach rather than through complex hierarchical structures.

Inventive Principle:
Principle #26Copying

3Measurement precision

If multiple similarity measurements are calculated using machine learning techniques, then the accuracy of identifying contributing workloads improves, but the computational complexity increases

Engineering Contradiction:
Improveaccuracy of workload identificationVSAvoidcomputational resources required
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent implements a weighted majority voting mechanism that allows partial contribution from multiple similarity measurements. Instead of requiring all measurements to reach a threshold, the system uses weighted votes where each similarity measurement contributes proportionally to the final identification. This partial action approach maintains accuracy while reducing computational burden compared to requiring all measurements to be processed to full precision.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes the parameter of similarity measurement from a single complex calculation to multiple simpler measurements with different weights. By transforming the problem from one highly accurate measurement to multiple measurements with varying importance, the system achieves comparable or better accuracy while distributing computational load across different types of calculations that can be optimized independently.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11175829B2Automatic identification of workloads contributing to behavioral changes in storage systems using machine learning techniques
Publication Date: 2021.11.16 EMC IP HLDG CO LLC
  • US11175829B2 patent drawing
  • US11175829B2 patent drawing
  • US11175829B2 patent drawing

AI summary

Methods, apparatus, and processor-readable storage media for automatic identification of workloads contributing to behavioral changes in storage systems using machine learning techniques are provided herein. An example computer-implemented method includes obtaining a primary time series and a set of candidate time series; calculating, using machine learning techniques, similarity measurements between the primary time series and each candidate time series in the set; for each similarity measurement, assigning weights to the candidate time series based on similarity values; generating, for each candidate time series, a similarity score based on the assigned weights; automatically identifying, based on the similarity scores, a candidate time series as contributing to an anomaly exhibited in the primary time series; and outputting identifying information of the at least one identified candidate time series for use in one or more automated actions associated with the storage system.