ML-Based Storage Tier Switching for Distributed Data Access Patterns

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large-scale distributed data storage systems face challenges in efficiently managing different storage tiers with varying cost and access cost profiles, leading to inefficient data placement and resource utilization due to the lack of automated switching between data storage schemes based on access patterns.

Innovation Solution

A machine learning-based approach is employed to predict future content item access patterns, allowing for dynamic and automatic switching between hotter and colder storage schemes, optimizing data placement based on predicted access times and patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual management of storage tiers is used, then data placement can be optimized for specific workloads, but system complexity and management burden increase significantly

Engineering Contradiction:
Improvedata placement optimizationVSAvoidmanagement complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system employs machine learning models that automatically analyze access patterns and make data placement decisions without human intervention. The ML-based policy autonomously predicts future access patterns and switches data between storage tiers, eliminating the need for manual management while maintaining optimized data placement.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical management processes with automated machine learning-based decision-making. The ML models process access pattern data and automatically determine optimal storage tier placements, substituting human-operated mechanical systems with intelligent automated systems.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Power

If data is stored in hotter storage schemes, then data access cost is reduced, but data storage cost increases

Engineering Contradiction:
Improvedata access costVSAvoiddata storage cost
Core Design Contradiction:
PowerVSLoss of energy

Solution Approach 1:

The system dynamically switches data between hotter and colder storage schemes based on predicted access patterns. Data that is predicted to be accessed frequently remains in hotter storage, while data predicted to be accessed infrequently is moved to colder storage, optimizing the balance between access cost and storage cost in real-time.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the storage scheme parameters (hotter vs. colder storage) based on predicted access patterns. The system adjusts which storage tier data resides in by changing the accessibility and cost parameters of the storage scheme, rather than maintaining fixed storage assignments.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated machine learning-based switching is implemented, then resource utilization improves, but system complexity increases

Engineering Contradiction:
Improveresource utilizationVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of access patterns using machine learning models to predict future data access behavior. By anticipating access patterns in advance, the system proactively switches data between storage tiers before actual access occurs, improving resource utilization while automating the process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a feedback loop where the system continuously monitors actual access patterns, compares them with predicted patterns, and uses this information to refine ML model predictions and adjust data placement decisions. This automated feedback mechanism improves resource utilization while managing system complexity through closed-loop control.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11422721B2Data storage scheme switching in a distributed data storage system
Publication Date: 2022.08.23 DROPBOX INC
  • US11422721B2 patent drawing
  • US11422721B2 patent drawing
  • US11422721B2 patent drawing

AI summary

Systems and methods for dynamic and automatic data storage scheme switching in a distributed data storage system. A machine learning-based policy for computing probable future content item access patterns based on historical content item access patterns is employed to dynamically and automatically switch the storage of content items (e.g., files, digital data, photos, text, audio, video, streaming content, cloud documents, etc.) between different data storage schemes. The different data storage schemes may have different data storage cost and different data access cost characteristics. For example, the different data storage schemes may encompass different types of data storage devices, different data compression schemes, and/or different data redundancy schemes.