ML-Based Data Sampling for Storage Deduplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing storage systems face challenges in efficiently managing and optimizing data reduction across distributed storage environments, particularly in terms of scalability and performance.
Innovation Solution
The use of machine learning and data sampling techniques to optimize data reduction in storage systems, allowing for more efficient data management and improved performance across distributed environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of substance
If traditional data reduction methods are used in distributed storage environments, then data reduction is achieved, but scalability and performance deteriorate
Solution Approach 1:
The patent replaces traditional mechanical data reduction methods with machine learning-based intelligent algorithms. The system uses trained machine learning models to analyze data characteristics and determine optimal reduction strategies, substituting rule-based mechanical processing with adaptive intelligent processing that maintains high performance while achieving data reduction
Solution Approach 2:
The system dynamically changes data processing parameters based on learned patterns. By using machine learning to identify optimal compression ratios, deduplication thresholds, and reduction strategies tailored to specific data types and access patterns, the system adapts parameters to maintain performance while maximizing data reduction
2Loss of substance
If traditional data reduction methods are used in distributed storage environments, then data reduction is achieved, but scalability deteriorates
Solution Approach 1:
The patent replaces centralized mechanical data reduction control with distributed machine learning inference. Each storage node independently applies trained models to local data, eliminating the need for centralized coordination and enabling the system to scale to distributed environments without performance degradation
Solution Approach 2:
The system segments the data reduction function into independent machine learning models deployed at each storage node. This segmentation allows each node to autonomously perform data reduction on its local data, enabling horizontal scaling of the distributed storage system without creating bottlenecks
Data Source
AI summary
Using machine learning and data sampling to optimize data reduction, including: collecting a plurality of data samples from data stored in a storage system; calculating, based on the plurality of data samples, a deduplication ratio; calculating, based on the plurality of data samples, a compression ratio; and performing a data reduction resource allocation in the storage system based on the deduplication ratio and the compression ratio.


