Machine Learning Data Archiving for Storage Object Access Forecasting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for identifying storage objects for long-term archiving are not accurate and do not consider the inverse process of identifying active storage objects for tiering, leading to inefficiencies in storage management.
Innovation Solution
A classification-based machine learning model is used to divide storage objects into activity classes based on historical IO requests, forecasting the next access time for each object, and archiving or removing objects based on this forecast, considering factors like storage system capacity, archiving cost, and retrieval cost.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If simple statistical methods are used to forecast storage object temperature, then the implementation is simpler, but the accuracy of temperature forecasting is insufficient
Solution Approach 1:
The patent replaces simple statistical methods with a machine learning-based forecasting system that uses classification models (such as random forests, gradient boosting, or neural networks) to predict storage object temperature. This substitution enables significantly higher accuracy in forecasting storage object activity patterns, allowing for more effective data tiering and caching decisions while managing the increased computational complexity through automated training and deployment pipelines.
2Measurement precision
If the inverse process of identifying active storage objects is used for archiving, then the approach is simpler, but the accuracy of archiving decisions is insufficient
Solution Approach 1:
The patent inverts the conventional approach by not simply reversing the tiering identification process, but by implementing a dedicated archiving identification process that uses the same machine learning temperature forecasting models. This inverted approach allows for more accurate archiving decisions by explicitly modeling cold storage object characteristics and applying them through a separate optimization process that considers archiving costs, retrieval costs, and storage capacity constraints.
Solution Approach 2:
The patent changes the parameters used for identification by introducing multiple factors beyond just temperature forecasting, including archiving costs, retrieval costs, storage system capacity, and performance tradeoffs. This parameter expansion enables more accurate and nuanced archiving decisions that balance multiple objectives simultaneously.
3Quantity of substance
If more storage objects are archived to optimize capacity, then storage system capacity utilization improves, but the complexity of managing archived objects increases
Solution Approach 1:
The patent implements self-service through automated machine learning-based identification and archiving processes that continuously monitor storage object temperature and automatically make archiving decisions. The system uses trained models to predict which objects should be archived based on their access patterns, eliminating manual intervention and reducing management complexity while optimizing capacity utilization.
Solution Approach 2:
The patent incorporates feedback mechanisms where the archiving decisions are continuously monitored and adjusted based on actual storage object access patterns. The machine learning models receive feedback about archiving effectiveness and refine their predictions accordingly, enabling adaptive management that optimizes capacity while controlling complexity through automated learning and adjustment.
Data Source
AI summary
A method, computer program product, and computing system for processing a plurality of historical input/output (IO) requests associated with a plurality of storage objects of a storage system from a plurality of time intervals. The plurality of storage objects may be divided into a plurality of storage activity classes using a classification-based machine learning model and the plurality of historical IO requests. A next access time for each storage object may be forecasted based upon, at least in part, the plurality of storage activity classes.


