ML-Based Data Sampling for Storage Deduplication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing storage systems face challenges in efficiently managing and optimizing data reduction across distributed storage environments, particularly in terms of scalability and performance.

Innovation Solution

The use of machine learning and data sampling techniques to optimize data reduction in storage systems, allowing for more efficient data management and improved performance across distributed environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of substance

If traditional data reduction methods are used in distributed storage environments, then data reduction is achieved, but scalability and performance deteriorate

Engineering Contradiction:
Improvedata reductionVSAvoidperformance
Core Design Contradiction:
Loss of substanceVSProductivity

Solution Approach 1:

The patent replaces traditional mechanical data reduction methods with machine learning-based intelligent algorithms. The system uses trained machine learning models to analyze data characteristics and determine optimal reduction strategies, substituting rule-based mechanical processing with adaptive intelligent processing that maintains high performance while achieving data reduction

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system dynamically changes data processing parameters based on learned patterns. By using machine learning to identify optimal compression ratios, deduplication thresholds, and reduction strategies tailored to specific data types and access patterns, the system adapts parameters to maintain performance while maximizing data reduction

Inventive Principle:
Principle #35Parameter changes

2Loss of substance

If traditional data reduction methods are used in distributed storage environments, then data reduction is achieved, but scalability deteriorates

Engineering Contradiction:
Improvedata reductionVSAvoidscalability
Core Design Contradiction:
Loss of substanceVSAdaptability or versatility

Solution Approach 1:

The patent replaces centralized mechanical data reduction control with distributed machine learning inference. Each storage node independently applies trained models to local data, eliminating the need for centralized coordination and enabling the system to scale to distributed environments without performance degradation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system segments the data reduction function into independent machine learning models deployed at each storage node. This segmentation allows each node to autonomously perform data reduction on its local data, enabling horizontal scaling of the distributed storage system without creating bottlenecks

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250165172A1Using Machine Learning And Data Sampling To Optimize Data Reduction
Publication Date: 2025.05.22 PURE STORAGE INC
  • US20250165172A1 patent drawing
  • US20250165172A1 patent drawing
  • US20250165172A1 patent drawing

AI summary

Using machine learning and data sampling to optimize data reduction, including: collecting a plurality of data samples from data stored in a storage system; calculating, based on the plurality of data samples, a deduplication ratio; calculating, based on the plurality of data samples, a compression ratio; and performing a data reduction resource allocation in the storage system based on the deduplication ratio and the compression ratio.