Hierarchical Control Planes for Distributed Storage Reliability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing storage systems face challenges in efficiently managing data across distributed computing environments, particularly in terms of data persistence, redundancy, and scalability.

Innovation Solution

The implementation of hierarchical control planes that manage distributed computing environments, utilizing multiple storage array controllers and non-volatile solid state storage units to enhance data management, redundancy, and scalability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single control plane manages the distributed computing environment, then device management is centralized and simple, but scalability and reliability are limited

Engineering Contradiction:
Improvecontrol plane structureVSAvoidsystem reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The control plane is segmented into multiple independent control plane nodes that collectively manage the distributed computing environment. Each node can independently handle control plane functions, eliminating the single point of failure and enabling load distribution across multiple nodes for improved reliability and scalability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The control plane architecture transitions from a single-dimensional centralized structure to a multi-dimensional distributed structure. Control plane nodes are distributed across multiple dimensions (physical locations, network segments), enabling the system to scale horizontally while maintaining management capabilities through spatial distribution.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If data is stored across multiple storage array controllers, then redundancy and reliability are improved, but device complexity and management difficulty increase

Engineering Contradiction:
Improvedata redundancyVSAvoidstorage management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Each storage array controller is designed with universal functionality to perform all control plane operations independently. Any controller can assume the role of lead controller, eliminating the need for specialized roles and simplifying management while maintaining redundancy through functional equivalence across multiple controllers.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The distributed control plane nodes automatically self-manage through peer-to-peer communication and automated failover mechanisms. When a controller fails, remaining controllers automatically detect and redistribute responsibilities without external intervention, reducing management complexity while maintaining redundancy.

Inventive Principle:
Principle #25Self-service

3Speed

If NVRAM is used for quick data access, then data access speed is improved, but data persistence and reliability depend on additional protection mechanisms

Engineering Contradiction:
Improvedata access speedVSAvoiddata persistence
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

Data written to NVRAM is protected by beforehand cushioning mechanisms including immediate replication to persistent storage media and confirmation protocols. Before data is considered fully persisted, redundant copies are created and verified, cushioning against potential NVRAM failures while maintaining the speed benefits of NVRAM for active data access.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS20250077295A1Using Hierarchical Control Planes to Manage Distributed Computing Environments
Publication Date: 2025.03.06 PURE STORAGE INC
  • US20250077295A1 patent drawing
  • US20250077295A1 patent drawing
  • US20250077295A1 patent drawing

AI summary

An illustrative method includes a global control plane managing a plurality of local control planes that manage a respective plurality of distributed computing environments, each of the plurality of local control planes configured to manage its respective computing environment in accordance with a mode of operation that is selected from a plurality of modes of operation based on a respective connectivity status with the global control plane, wherein managing the plurality of local control planes comprises selectively deploying an update to one or more local control planes of the plurality of local control planes based on the connectivity statuses of the plurality of local control planes with the global control plane. In some embodiments, the global control plane receives log data associated with the update and deploys, based on the log data, a second update to the one or more local control planes.