Storage System Clustering for Data Pattern Sharing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current storage systems lack an efficient mechanism for sharing data patterns across multiple storage systems, which hinders the effectiveness of data deduplication processes, leading to suboptimal storage utilization and performance.

Innovation Solution

A monitoring and analytics platform is implemented to collect and cluster storage systems based on data patterns, creating data pattern sharing clusters, where a subset of identified patterns is shared among the systems to enhance deduplication efficiency using inline pattern detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data patterns are shared across multiple storage systems through clustering, then data deduplication performance is improved, but system complexity increases

Engineering Contradiction:
Improvedata deduplication performanceVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments storage systems into distinct clusters based on their data pattern characteristics. Each cluster is independently managed with its own deduplication optimization strategies, allowing the system to handle complexity in a modular fashion while improving overall deduplication performance through targeted pattern sharing within clusters.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a monitoring and analytics platform as an intermediary that collects data patterns from storage systems, performs clustering analysis, and distributes identified patterns back to appropriate systems. This intermediary handles the complex analysis and coordination work, isolating the complexity from individual storage systems while enabling improved deduplication across the cluster.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If data patterns are collected and analyzed across multiple storage systems, then deduplication efficiency is enhanced, but processing time and resources increase

Engineering Contradiction:
Improvededuplication efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by proactively collecting and analyzing data patterns from storage systems in advance, before deduplication operations are performed. The monitoring platform continuously gathers pattern data and performs clustering analysis ahead of time, so that when deduplication is needed, the system can immediately apply pre-identified patterns without performing time-consuming analysis during the actual deduplication process.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If clustering is implemented to share data patterns, then storage utilization is optimized, but implementation complexity increases

Engineering Contradiction:
Improvestorage utilizationVSAvoidimplementation complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling storage systems to automatically receive and apply data patterns from their own cluster without requiring manual configuration or intervention. The monitoring and analytics platform autonomously performs pattern collection, clustering analysis, identification of shareable patterns, and distribution to member systems, allowing the system to optimize its own storage utilization automatically.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11775483B2Clustering storage systems for sharing of data patterns used for deduplication
Publication Date: 2023.10.03 DELL PROD LP
  • US11775483B2 patent drawing
  • US11775483B2 patent drawing
  • US11775483B2 patent drawing

AI summary

An apparatus comprises at least one processing device configured to collect, from a plurality of storage systems, data patterns for data stored in the plurality of storage systems and to cluster the plurality of storage systems into one or more data pattern sharing clusters based at least in part on the collected data patterns, a given one of the one or more data pattern sharing clusters comprising two or more of the plurality of storage systems. The at least one processing device is also configured to identify, for the given data pattern sharing cluster, a subset of the collected data patterns and to provide, to the two or more storage systems of the given data pattern sharing cluster, the identified subset of the data patterns, wherein the identified subset of the collected data patterns are utilized by the two or more storage systems in performing data deduplication.