Storage System Clustering for Data Pattern Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current storage systems lack an efficient mechanism for sharing data patterns across multiple storage systems, which hinders the effectiveness of data deduplication processes, leading to suboptimal storage utilization and performance.
Innovation Solution
A monitoring and analytics platform is implemented to collect and cluster storage systems based on data patterns, creating data pattern sharing clusters, where a subset of identified patterns is shared among the systems to enhance deduplication efficiency using inline pattern detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data patterns are shared across multiple storage systems through clustering, then data deduplication performance is improved, but system complexity increases
Solution Approach 1:
The patent segments storage systems into distinct clusters based on their data pattern characteristics. Each cluster is independently managed with its own deduplication optimization strategies, allowing the system to handle complexity in a modular fashion while improving overall deduplication performance through targeted pattern sharing within clusters.
Solution Approach 2:
The patent introduces a monitoring and analytics platform as an intermediary that collects data patterns from storage systems, performs clustering analysis, and distributes identified patterns back to appropriate systems. This intermediary handles the complex analysis and coordination work, isolating the complexity from individual storage systems while enabling improved deduplication across the cluster.
2Productivity
If data patterns are collected and analyzed across multiple storage systems, then deduplication efficiency is enhanced, but processing time and resources increase
Solution Approach 1:
The patent implements preliminary action by proactively collecting and analyzing data patterns from storage systems in advance, before deduplication operations are performed. The monitoring platform continuously gathers pattern data and performs clustering analysis ahead of time, so that when deduplication is needed, the system can immediately apply pre-identified patterns without performing time-consuming analysis during the actual deduplication process.
3Quantity of substance
If clustering is implemented to share data patterns, then storage utilization is optimized, but implementation complexity increases
Solution Approach 1:
The patent implements self-service by enabling storage systems to automatically receive and apply data patterns from their own cluster without requiring manual configuration or intervention. The monitoring and analytics platform autonomously performs pattern collection, clustering analysis, identification of shareable patterns, and distribution to member systems, allowing the system to optimize its own storage utilization automatically.
Data Source
AI summary
An apparatus comprises at least one processing device configured to collect, from a plurality of storage systems, data patterns for data stored in the plurality of storage systems and to cluster the plurality of storage systems into one or more data pattern sharing clusters based at least in part on the collected data patterns, a given one of the one or more data pattern sharing clusters comprising two or more of the plurality of storage systems. The at least one processing device is also configured to identify, for the given data pattern sharing cluster, a subset of the collected data patterns and to provide, to the two or more storage systems of the given data pattern sharing cluster, the identified subset of the data patterns, wherein the identified subset of the collected data patterns are utilized by the two or more storage systems in performing data deduplication.


