Data Storage Tier Optimization via Affinity-Based Grouping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data storage systems face challenges in efficiently managing and optimizing data distribution across multiple storage tiers based on workload and performance characteristics, leading to suboptimal storage resource utilization and performance.

Innovation Solution

A method for grouping data portions based on similarity in behavior, using affinity measurements and activity ratios to determine data movement between storage tiers, with a computer-readable medium and data storage system that performs data storage optimizations and models workload performance, allowing for automatic movement of data between tiers to optimize performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data is stored across multiple storage tiers with varying technologies and performance characteristics, then storage system capacity and versatility are improved, but data placement optimization and performance efficiency deteriorate

Engineering Contradiction:
Improvestorage tier compatibilityVSAvoiddata placement efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent changes the parameters used for data placement from simple sequential numbering to behavior-based affinity measurements. By monitoring workload characteristics, access patterns, and performance metrics over time, the system dynamically adjusts placement decisions based on actual data behavior rather than static tier definitions, resolving the contradiction between multi-tier versatility and placement efficiency

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system implements feedback loops that continuously monitor data workload characteristics and access patterns, then use this information to adjust data placement decisions. The affinity measurement mechanism incorporates temporal aspects by observing behavior over multiple time intervals, allowing the system to learn from past performance and optimize future placement, thereby improving productivity while maintaining adaptability across storage tiers

Inventive Principle:
Principle #23Feedback

2Ease of operation

If sequential numbering is used for data placement across storage tiers, then ease of operation is improved, but storage resource utilization and performance optimization deteriorate

Engineering Contradiction:
Improvedata placement simplicityVSAvoidstorage resource utilization
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system enables data portions to effectively select their own optimal storage location through the affinity measurement mechanism. By automatically monitoring workload characteristics and comparing them against tier-specific behavior profiles, the system performs self-optimization without requiring manual intervention or complex administrative operations, thus maintaining ease of operation while dramatically improving resource utilization

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the data placement parameter from a simple sequential counter to a multi-dimensional affinity measurement that incorporates workload characteristics, access patterns, and temporal behavior. This parameter change enables intelligent placement decisions that optimize storage resource utilization while the automated nature of the process maintains operational simplicity

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If data placement is based on sequential numbering, then device complexity is reduced, but storage system performance and optimization capabilities deteriorate

Engineering Contradiction:
Improveplacement logic complexityVSAvoidstorage system performance
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent introduces an intermediary affinity measurement mechanism that sits between the simple sequential numbering system and the complex performance optimization requirements. This intermediary layer translates complex workload characteristics and performance metrics into a standardized affinity score, which then guides data placement decisions. This approach maintains relatively simple placement logic while achieving sophisticated performance optimization through the intermediary's ability to process and synthesize multiple performance parameters

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9753987B1Identifying groups of similar data portions
Publication Date: 2017.09.05 EMC IP HLDG CO LLC
  • US9753987B1 patent drawing
  • US9753987B1 patent drawing
  • US9753987B1 patent drawing

AI summary

Techniques for grouping data portions are disclosed. Each group includes data portions determined to exhibit similar behavior. The techniques may include determining whether an affinity measurement with respect to two groups exceeds an affinity threshold; merging the two groups into a single group responsive to the affinity measurement exceeding the affinity threshold; modeling movement of at least one data portion of the single group between two storage tiers at a particular time of day using predicted workload metrics; and performing the data movement of the at least one data portion between the two storage tiers. Predicted workload metrics may be determined by revising first modeled workload metrics using a bias value, where bias values are associated with different times of day, and the bias value is selected based on the particular time of day that the predicted workload metrics are modeling.