Cost-Aware Data Storage Tiering with Predictive Risk Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data tiering solutions rely on rule or policy-based approaches, leading to inefficient storage costs as they wait for data segments to become inactive before tiering, resulting in excessive usage costs when the data is accessed again, and fail to intelligently manage sporadically accessed data segments.
Innovation Solution
A method that detects data segments, gathers access and usage information, calculates a tiering threshold based on activity and pricing, identifies inactive segments, calculates risk scores, and generates a tiering list to proactively store data in cost-effective storage options using predictive analytics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If rule or policy-based tiering approaches are used, then data storage management is simplified, but storage costs increase due to waiting for data inactivity and excessive access costs
Solution Approach 1:
The system performs preliminary actions by calculating risk scores and predicting future data activity patterns before data becomes inactive. This allows proactive tiering decisions to be made in advance, moving data to appropriate storage tiers before inactivity is confirmed, thereby reducing the waiting period and associated storage costs while maintaining simplified management through automated predictions
Solution Approach 2:
The system implements feedback mechanisms by continuously monitoring data access patterns and using this information to refine risk score calculations. This feedback loop enables the system to learn from actual data behavior, improving prediction accuracy over time and optimizing tiering decisions to reduce storage costs while maintaining manageable complexity through adaptive policy adjustment
2Loss of energy
If data is tiered based on current inactivity status, then storage costs are reduced, but access costs increase when data is frequently retrieved after tiering
Solution Approach 1:
The system performs preliminary risk assessment and prediction before making tiering decisions. By calculating risk scores that incorporate predicted future access patterns, the system proactively identifies data unlikely to be accessed soon, enabling confident tiering decisions that reduce storage costs while minimizing the risk of expensive access operations for frequently needed data
Solution Approach 2:
The system changes the decision-making parameter from simple current inactivity status to a comprehensive risk score that incorporates multiple factors including predicted future activity, data criticality, and access patterns. This parameter transformation enables more accurate prediction of data behavior, allowing optimal tiering decisions that balance storage cost reduction with access cost minimization
3Ease of manufacture
If traditional tiering methods are used, then implementation is straightforward, but sporadically accessed data segments are not intelligently managed
Solution Approach 1:
The system implements self-service capabilities by enabling automated risk score calculation and prediction through machine learning models that continuously learn from data patterns. This self-improving mechanism allows the system to intelligently adapt to various data access patterns including sporadic access without requiring complex manual configuration, maintaining implementation ease while significantly enhancing adaptability to different data behaviors
Solution Approach 2:
The system performs preliminary analysis and prediction using trained machine learning models before making tiering decisions. This advance preparation enables the system to recognize and adapt to sporadic access patterns and other complex data behaviors, providing intelligent management capabilities while maintaining straightforward implementation through pre-trained predictive models
Data Source
AI summary
An embodiment for a method of cost-aware tiering for data storage. The embodiment may detect data to be stored in one or more storage systems and gathering access and usage information. The embodiment may maintain a first data structure including periods of activity below a defined threshold for the detected data, and a common data structure to identify and track application-wide data patterns. The embodiment may gather static pricing information and context pricing information for the one or more storage systems and calculate a tiering threshold corresponding to a continuous period of activity below a second defined threshold required for which local storage costs exceed a tiering cost. The embodiment may calculate for inactive data segments, probabilities of the inactive data segments remaining inactive for a duration that exceeds the calculated tiering threshold. The embodiment may calculate a risk score and generate a tiering list for storing the data.


