ML-Based Storage Tier Switching for Distributed Data Access Patterns
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large-scale distributed data storage systems face challenges in efficiently managing different storage tiers with varying cost and access cost profiles, leading to inefficient data placement and resource utilization due to the lack of automated switching between data storage schemes based on access patterns.
Innovation Solution
A machine learning-based approach is employed to predict future content item access patterns, allowing for dynamic and automatic switching between hotter and colder storage schemes, optimizing data placement based on predicted access times and patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual management of storage tiers is used, then data placement can be optimized for specific workloads, but system complexity and management burden increase significantly
Solution Approach 1:
The system employs machine learning models that automatically analyze access patterns and make data placement decisions without human intervention. The ML-based policy autonomously predicts future access patterns and switches data between storage tiers, eliminating the need for manual management while maintaining optimized data placement.
Solution Approach 2:
The patent replaces manual mechanical management processes with automated machine learning-based decision-making. The ML models process access pattern data and automatically determine optimal storage tier placements, substituting human-operated mechanical systems with intelligent automated systems.
2Power
If data is stored in hotter storage schemes, then data access cost is reduced, but data storage cost increases
Solution Approach 1:
The system dynamically switches data between hotter and colder storage schemes based on predicted access patterns. Data that is predicted to be accessed frequently remains in hotter storage, while data predicted to be accessed infrequently is moved to colder storage, optimizing the balance between access cost and storage cost in real-time.
Solution Approach 2:
The patent changes the storage scheme parameters (hotter vs. colder storage) based on predicted access patterns. The system adjusts which storage tier data resides in by changing the accessibility and cost parameters of the storage scheme, rather than maintaining fixed storage assignments.
3Productivity
If automated machine learning-based switching is implemented, then resource utilization improves, but system complexity increases
Solution Approach 1:
The system performs preliminary analysis of access patterns using machine learning models to predict future data access behavior. By anticipating access patterns in advance, the system proactively switches data between storage tiers before actual access occurs, improving resource utilization while automating the process.
Solution Approach 2:
The patent implements a feedback loop where the system continuously monitors actual access patterns, compares them with predicted patterns, and uses this information to refine ML model predictions and adjust data placement decisions. This automated feedback mechanism improves resource utilization while managing system complexity through closed-loop control.
Data Source
AI summary
Systems and methods for dynamic and automatic data storage scheme switching in a distributed data storage system. A machine learning-based policy for computing probable future content item access patterns based on historical content item access patterns is employed to dynamically and automatically switch the storage of content items (e.g., files, digital data, photos, text, audio, video, streaming content, cloud documents, etc.) between different data storage schemes. The different data storage schemes may have different data storage cost and different data access cost characteristics. For example, the different data storage schemes may encompass different types of data storage devices, different data compression schemes, and/or different data redundancy schemes.


