Storage Capacity Forecasting via Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional storage management approaches rely on rule-based capacity threshold alerts, leading to false positives and false negatives in predicting capacity exhaustion, particularly at the pool level, due to oversubscription across storage objects.

Innovation Solution

The implementation of machine learning techniques to forecast capacity for storage objects, aggregating these forecasts to determine potential capacity exhaustion at the pool level, and performing automated actions based on these determinations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If rule-based capacity threshold alerts are used at the pool level, then storage management is simplified, but false positives and false negatives increase

Engineering Contradiction:
Improvestorage management simplicityVSAvoidcapacity exhaustion prediction accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the capacity prediction problem from pool-level to storage object-level forecasts. Instead of applying a single rule-based threshold at the pool level, the system divides the pool into individual storage objects (LUNs, file systems) and generates capacity forecasts for each object separately using machine learning. These individual forecasts are then aggregated to determine pool-level capacity status, thereby improving prediction accuracy while maintaining manageable complexity through automated processing.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If machine learning techniques are applied at the storage object level, then prediction accuracy improves, but system complexity increases

Engineering Contradiction:
Improvecapacity exhaustion prediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements self-service through automated machine learning models that independently forecast capacity for each storage object without requiring manual configuration or intervention. The system automatically collects historical capacity data, trains forecasting models, generates predictions, and aggregates results. This automation eliminates the need for manual rule configuration at multiple levels, allowing the system to handle the increased complexity of object-level forecasting while maintaining ease of operation through self-service capabilities.

Inventive Principle:
Principle #25Self-service

3Productivity

If pool-level capacity monitoring is used, then management overhead is reduced, but false predictions occur due to oversubscription

Engineering Contradiction:
Improvemanagement efficiencyVSAvoidcapacity prediction reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent transitions from monitoring capacity at a single pool level to forecasting capacity at the individual storage object level, effectively adding a dimensional breakdown to the monitoring approach. By examining capacity trends for each LUN and file system separately and then aggregating these forecasts, the system captures the nuanced impact of oversubscription on individual objects. This dimensional change allows the system to maintain management efficiency through automation while improving prediction reliability by accounting for object-specific capacity behaviors.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11513938B2Determining capacity in storage systems using machine learning techniques
Publication Date: 2022.11.29 EMC IP HLDG CO LLC
  • US11513938B2 patent drawing
  • US11513938B2 patent drawing
  • US11513938B2 patent drawing

AI summary

Methods, apparatus, and processor-readable storage media for determining capacity in storage systems using machine learning techniques are provided herein. An example computer-implemented method includes obtaining capacity-related data from a storage system; forecasting, for a given temporal period, capacity of one or more storage objects of the storage system by applying machine learning techniques to at least a portion of the capacity-related data; aggregating the forecasted capacity for at least portions of the one or more storage objects; determining, based on the aggregated forecasted capacity of the storage objects, whether at least a portion of the storage system will run out of capacity in connection with the given temporal period; and performing one or more automated actions based at least in part on the determination as to whether the at least a portion of the at least one storage system will run out of capacity.