Storage System Deduplication and Compression Configuration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage systems face inefficiencies when attempting to deduplicate or compress data that is not compressible or deduplicable, leading to degraded performance, as current methods lack the ability to determine the suitability of data for these processes on a per-data-block basis.
Innovation Solution
A method and system that analyze input/output patterns to identify application types and select optimal deduplication and compression configurations accordingly, allowing for intelligent decision-making on when to apply deduplication and compression to enhance storage system performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data deduplication and compression are applied to all data blocks, then storage space savings are improved, but system performance deteriorates due to processing non-compressible or non-deduplicable data
Solution Approach 1:
The patent applies different quality treatments to different parts of the data based on their compressibility and deduplicability characteristics. By analyzing each data block individually and applying deduplication or compression only where beneficial, the system achieves local optimization rather than uniform processing, thereby improving storage efficiency without unnecessarily degrading performance on non-compressible data
Solution Approach 2:
The system dynamically changes the processing parameters (deduplication ratio, compression level) based on the characteristics of each data block. By adjusting these parameters according to the actual data properties, the system optimizes the balance between storage space savings and processing performance, avoiding the application of intensive processing to data that cannot benefit from it
2Productivity
If deduplication and compression settings are optimized per application type, then system performance is improved, but device complexity increases due to pattern analysis and configuration selection
Solution Approach 1:
The system performs self-service by automatically analyzing I/O patterns, identifying application types, and selecting optimal deduplication and compression configurations without requiring manual intervention. This automation reduces the operational complexity for users while maintaining sophisticated performance optimization capabilities
Solution Approach 2:
The system implements feedback mechanisms by continuously monitoring I/O patterns and using this information to adaptively adjust deduplication and compression settings. The feedback loop allows the system to learn from actual data characteristics and performance outcomes, automatically optimizing configurations based on real-world usage patterns
3Quantity of substance
If manual analysis of data compressibility is performed, then storage efficiency is improved, but loss of time occurs due to processing delays
Solution Approach 1:
The system performs preliminary analysis of I/O patterns and data characteristics in advance, before actual deduplication or compression operations are applied. By pre-identifying compressible and deduplicable data blocks through pattern recognition, the system avoids time-consuming trial processing and directly applies optimization only where needed, reducing overall processing time while maintaining storage efficiency
Data Source
AI summary
Implementations are provided herein for systems, methods, and a non-transitory computer product configured to analyze an input/output (IO) pattern for a data storage system, to identify an application type based on the IO pattern, and to select optimal deduplication and compression configurations based on the application type. The teachings herein facilitate machine learning of various metrics and the interrelations between these metrics, such as past IO patterns, application types, deduplication configurations, compression configurations, and overall system performance. These metrics and interrelations can be stored in a data lake. In some embodiments, data objects can be segmented in order to optimize configurations with more granularity. In additional embodiments, predictive techniques are used to select deduplication and compression configurations.


