Data Storage Tier Optimization via Affinity-Based Grouping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems face challenges in efficiently managing and optimizing data distribution across multiple storage tiers based on workload and performance characteristics, leading to suboptimal storage resource utilization and performance.
Innovation Solution
A method for grouping data portions based on similarity in behavior, using affinity measurements and activity ratios to determine data movement between storage tiers, with a computer-readable medium and data storage system that performs data storage optimizations and models workload performance, allowing for automatic movement of data between tiers to optimize performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data is stored across multiple storage tiers with varying technologies and performance characteristics, then storage system capacity and versatility are improved, but data placement optimization and performance efficiency deteriorate
Solution Approach 1:
The patent changes the parameters used for data placement from simple sequential numbering to behavior-based affinity measurements. By monitoring workload characteristics, access patterns, and performance metrics over time, the system dynamically adjusts placement decisions based on actual data behavior rather than static tier definitions, resolving the contradiction between multi-tier versatility and placement efficiency
Solution Approach 2:
The system implements feedback loops that continuously monitor data workload characteristics and access patterns, then use this information to adjust data placement decisions. The affinity measurement mechanism incorporates temporal aspects by observing behavior over multiple time intervals, allowing the system to learn from past performance and optimize future placement, thereby improving productivity while maintaining adaptability across storage tiers
2Ease of operation
If sequential numbering is used for data placement across storage tiers, then ease of operation is improved, but storage resource utilization and performance optimization deteriorate
Solution Approach 1:
The system enables data portions to effectively select their own optimal storage location through the affinity measurement mechanism. By automatically monitoring workload characteristics and comparing them against tier-specific behavior profiles, the system performs self-optimization without requiring manual intervention or complex administrative operations, thus maintaining ease of operation while dramatically improving resource utilization
Solution Approach 2:
The patent transforms the data placement parameter from a simple sequential counter to a multi-dimensional affinity measurement that incorporates workload characteristics, access patterns, and temporal behavior. This parameter change enables intelligent placement decisions that optimize storage resource utilization while the automated nature of the process maintains operational simplicity
3Device complexity
If data placement is based on sequential numbering, then device complexity is reduced, but storage system performance and optimization capabilities deteriorate
Solution Approach 1:
The patent introduces an intermediary affinity measurement mechanism that sits between the simple sequential numbering system and the complex performance optimization requirements. This intermediary layer translates complex workload characteristics and performance metrics into a standardized affinity score, which then guides data placement decisions. This approach maintains relatively simple placement logic while achieving sophisticated performance optimization through the intermediary's ability to process and synthesize multiple performance parameters
Data Source
AI summary
Techniques for grouping data portions are disclosed. Each group includes data portions determined to exhibit similar behavior. The techniques may include determining whether an affinity measurement with respect to two groups exceeds an affinity threshold; merging the two groups into a single group responsive to the affinity measurement exceeding the affinity threshold; modeling movement of at least one data portion of the single group between two storage tiers at a particular time of day using predicted workload metrics; and performing the data movement of the at least one data portion between the two storage tiers. Predicted workload metrics may be determined by revising first modeled workload metrics using a bias value, where bias values are associated with different times of day, and the bias value is selected based on the particular time of day that the predicted workload metrics are modeling.


