Workload Affinity Grouping for Storage Tier Forecasting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data storage systems face challenges in predicting and adapting to changes in data access frequency, leading to inefficient storage tier allocation, as they assume constant access rates and fail to anticipate increased access frequency of infrequently accessed data.

Innovation Solution

A system that forecasts workload activity by selecting metrics, grouping data based on workload affinity, and establishing precursor/target relationships to anticipate and respond to changes in access patterns, using a non-transitory computer-readable medium with executable code to optimize storage tier allocation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If current access frequency is used as predictor of future access, then data currently accessed frequently can be moved to faster storage tiers, but access frequency changes are not anticipated leading to suboptimal storage allocation

Engineering Contradiction:
Improvestorage allocation efficiencyVSAvoidresponse to access pattern changes
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary actions by forecasting future workload activity before it actually occurs. It analyzes historical access patterns, identifies trends and seasonality, and predicts future access frequencies in advance, allowing data to be proactively moved to appropriate storage tiers before the access surge happens, rather than merely reacting to current access patterns

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system makes storage allocation dynamic by continuously monitoring access patterns, recalculating forecasts, and adapting storage tier assignments in real-time. The workload forecasting mechanism adjusts to changing access behaviors, seasonal variations, and emerging trends, transforming the static storage allocation into a dynamic system that automatically responds to evolving data access requirements

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If workload forecasting is performed for all data portions, then accurate predictions are achieved, but computational resources and time are excessively consumed

Engineering Contradiction:
Improveworkload prediction accuracyVSAvoidcomputational resource requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the data population into distinct groups based on workload affinity and access pattern characteristics. By dividing data into segments with similar behavior patterns, the system can apply forecasting methods selectively to representative samples from each segment rather than processing every single data portion, thereby maintaining prediction accuracy while reducing computational complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs partial action by forecasting workload for a subset of representative data portions rather than all data portions. It selects key representatives from each workload affinity group to model and predict behavior, which is sufficient to infer the access patterns of the entire group, achieving acceptable prediction accuracy with significantly reduced computational effort

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10671431B1Extent group workload forecasts
Publication Date: 2020.06.02 EMC IP HLDG CO LLC
  • US10671431B1 patent drawing
  • US10671431B1 patent drawing
  • US10671431B1 patent drawing

AI summary

Forecasting workload activity for data stored on a data storage device includes selecting at least one metric for measuring workload activity, providing at least one grouping of portions of the data according to a workload affinity determination provided for each of the portions at a subset of a plurality of time steps, where the workload affinity determination is based on each of the data portions in the group experiencing above-average workload activity during same ones of the subset of the plurality of time steps, the subset corresponding to at least one business cycle for accessing the data, and forecasting workload activity for all of the portions of data in the group based on forecasting workload activity for a subset of the data portions that is less than all of the data portions.