Distributed Storage Pool Selection for Fail-In-Place Data Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data storage systems lack effective management and operation strategies, particularly when components approach failure or reach storage capacity, leading to potential data loss and inefficiencies in cloud storage networks.

Innovation Solution

A dispersed storage network (DSN) with error encoding and decoding capabilities, utilizing a Decentralized Agreement Protocol (DAP) to distribute data across multiple storage units, ensuring data integrity and availability even in the presence of component failures, by encoding data into multiple slices and strategically selecting storage pools based on available capacity and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If data is stored in a centralized storage system, then storage management is simplified, but the system reliability decreases when components approach failure or reach storage capacity

Engineering Contradiction:
Improvestorage managementVSAvoiddata availability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments data into multiple encoded slices and distributes them across multiple storage units in a dispersed storage network. This segmentation allows the system to maintain reliability even when individual storage units fail or reach capacity, as data can be retrieved from remaining units. The encoding scheme ensures that a threshold number of slices are sufficient to reconstruct the original data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a decentralized agreement protocol as an intermediary mechanism that enables storage units to autonomously negotiate and manage data placement, retrieval, and failure recovery. This protocol mediates interactions between storage units, allowing them to coordinate without centralized control, thereby maintaining both reliability and operational efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If data is distributed across multiple storage units, then system reliability improves, but storage management complexity increases

Engineering Contradiction:
Improvedata availabilityVSAvoidstorage management
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service mechanisms where storage units autonomously manage their own data operations through the decentralized agreement protocol. Storage units can independently handle data placement, retrieval, and failure recovery without requiring complex centralized management, thereby reducing overall system complexity while maintaining high reliability.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The decentralized agreement protocol serves multiple functions simultaneously: it manages data placement, handles failure recovery, coordinates capacity allocation, and enables autonomous storage unit operation. This multi-functionality reduces the need for separate management systems and simplifies the overall architecture despite the distributed nature of the storage network.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If storage capacity of individual units is increased, then fewer units are needed, but the risk of data loss increases when that unit fails

Engineering Contradiction:
Improvestorage capacityVSAvoiddata loss risk
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent divides data into multiple encoded slices and distributes them across multiple storage units, ensuring that no single unit holds all data. This segmentation means that even if a large-capacity unit fails, only a portion of the data is affected, and the remaining units can still reconstruct the original data through the encoding scheme.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs error encoding schemes that transform the original data into multiple slices with specific redundancy properties. By changing the parameter of data representation from raw storage to encoded slices, the system achieves tolerance against storage unit failures while efficiently utilizing available capacity across units.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If error correction encoding is applied to data, then data integrity is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvedata integrityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies error correction encoding to data before distribution, performing the computationally intensive encoding operation in advance during the data ingestion phase. This preliminary action ensures data integrity is built-in from the start, and subsequent read operations can proceed more quickly without requiring real-time error correction processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses erasure coding schemes where only a threshold number of encoded slices are needed to reconstruct the original data, rather than requiring all slices. This partial action approach allows the system to tolerate failures and perform error correction more efficiently by working with subsets of the encoded data rather than the complete set.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10275185B2Fail-in-place supported via decentralized or Distributed Agreement Protocol (DAP)
Publication Date: 2019.04.30 PURE STORAGE INC
  • US10275185B2 patent drawing
  • US10275185B2 patent drawing
  • US10275185B2 patent drawing

AI summary

A computing device includes an interface configured to interface and communicate with a dispersed storage network (DSN), a memory that stores operational instructions, and processing circuitry operably coupled to the interface and to the memory. The processing circuitry is configured to execute the operational instructions to perform various operations and functions. The computing device receives a data access request and determines a DSN address associated therewith. The computing device identifies available of storage unit (SU) pools, then, for each available SU pool, updates a corresponding weighting level and determines corresponding ranked scoring information based on the DSN address and a corresponding updated weighting level(s) in accordance with system configuration(s) of a Decentralized, or Distributed, Agreement Protocol (DAP). The computing device selects an available SU pool and issues resource access request(s) thereto to process the data access request.