Hierarchical Cluster Management for Storage Nodes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large-scale dynamic scale-out storage systems face challenges in efficient cluster management and communication, leading to reduced performance as the number of storage nodes increases, due to the reliance on a single primary management node that becomes overwhelmed with processing and communication load.

Innovation Solution

Implementing a distributed hierarchical cluster management system that partitions storage nodes into subclusters, each with a local management subsystem communicating with a global management system, thereby distributing management responsibilities and reducing network and processing load on the global node.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single primary management node is used to manage the storage cluster, then the system structure is simple and easy to implement, but the management node becomes overwhelmed with processing and communication load as the number of storage nodes increases

Engineering Contradiction:
Improvemanagement system structureVSAvoidcluster management performance
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent divides the storage cluster into multiple subclusters, each managed by a local management subsystem. This segmentation distributes the management load from a single centralized node to multiple distributed nodes, allowing each local subsystem to handle management tasks for its specific subcluster independently, thereby improving overall cluster management performance while maintaining manageable complexity through modular organization

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If additional storage nodes are added to increase storage capacity, then the storage system scales to meet data demands, but the communication and management overhead increases significantly

Engineering Contradiction:
Improvestorage capacityVSAvoidcommunication and processing load
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

By partitioning the storage cluster into multiple subclusters with dedicated local management subsystems, the patent enables storage capacity to scale across more nodes while containing communication overhead within each subcluster. Local management reduces the need for every node to communicate with a centralized manager, thereby reducing overall communication and processing load as the system scales

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical dimension to the management structure, with local management subsystems operating at the subcluster level and a global management system operating at the cluster level. This dimensional organization allows storage capacity to expand horizontally across multiple subclusters while maintaining efficient local management, reducing the communication overhead that would otherwise scale linearly with the total number of nodes

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If a distributed hierarchical management system is implemented to reduce load on the global management node, then cluster management performance improves, but the system complexity increases

Engineering Contradiction:
Improvecluster management performanceVSAvoidmanagement system structure
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The hierarchical management system is segmented into distinct local and global components with clearly defined responsibilities. Local management subsystems handle routine subcluster operations independently, while the global management system coordinates across subclusters. This segmentation improves performance by distributing processing load while controlling complexity through functional separation and standardized interfaces between levels

Inventive Principle:
Principle #1Segmentation

4Loss of energy

If local management subsystems are deployed in each subcluster, then the communication load on the global network is reduced, but the number of management components in the system increases

Engineering Contradiction:
Improvenetwork communication loadVSAvoidnumber of management components
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

Local management subsystems are deployed on existing storage nodes within each subcluster, enabling those nodes to self-manage their local environment without requiring dedicated external management hardware. This self-service approach reduces network communication load by handling management traffic locally while avoiding the proliferation of separate management components, as the local subsystems leverage the computational resources already present in the storage nodes

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12099719B2Cluster management in large-scale storage systems
Publication Date: 2024.09.24 DELL PROD LP
  • US12099719B2 patent drawing
  • US12099719B2 patent drawing
  • US12099719B2 patent drawing

AI summary

Techniques are provided for implementing a distributed hierarchical cluster management system. A system comprises a data storage system and a cluster management system. The data storage system comprises a cluster of storage nodes that is partitioned into a plurality of subclusters of storage nodes. The cluster management system is deployed on at least some of the storage nodes of the data storage system, and comprises a global management system and a plurality of local management subsystems. Each local management subsystem is configured to manage a respective subcluster of the plurality of subclusters of storage nodes, and communicate with the global management system to provide subcluster status information to the global management system regarding a current state and configuration of the respective subcluster of storage nodes. The global management system is configured to manage the cluster of storage nodes using the subcluster status information provided by the local management subsystems.