Storage Unit Testing During Erasure-Coded Data Availability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current distributed storage systems face challenges in maintaining data integrity and availability in the face of failures within the storage network, particularly in ensuring reliable and secure storage and retrieval of encoded data slices across geographically dispersed locations.

Innovation Solution

A distributed computing system that employs dispersed error encoding and decoding schemes, allowing data to be encoded into multiple slices stored across different locations, with a decentralized agreement protocol for managing storage and task processing, enabling fault tolerance and secure data retrieval without the need for redundant copies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is stored using traditional redundant copying methods, then data availability is improved, but storage efficiency and scalability deteriorate

Engineering Contradiction:
Improvedata availabilityVSAvoidstorage efficiency
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments data into multiple slices and distributes them across different storage units in a decentralized network. Instead of storing complete redundant copies of entire datasets, the system divides data into manageable slices that can be independently stored and retrieved, improving both availability and storage efficiency simultaneously

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses error correction encoding to create encoded slices that can reconstruct the original data. Rather than simple redundant copying, the system generates multiple encoded versions where any sufficient subset can recover the original data, achieving high availability without proportional increases in storage requirements

Inventive Principle:
Principle #26Copying

2Reliability

If distributed storage is implemented across geographically dispersed locations, then system reliability is improved, but network complexity and data retrieval overhead increase

Engineering Contradiction:
Improvesystem reliabilityVSAvoidnetwork complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service mechanisms where storage units autonomously manage their own encoding, storage, and verification operations. The decentralized agreement protocol enables nodes to independently reach consensus on data integrity without centralized coordination, reducing network complexity while maintaining distributed reliability

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback mechanisms where storage units continuously verify data integrity and report status to the network. This automated feedback loop enables the decentralized protocol to maintain consistency and reliability across geographically dispersed locations without requiring complex manual coordination

Inventive Principle:
Principle #23Feedback

3Reliability

If error correction encoding is used to protect data integrity, then data security is improved, but processing time and computational overhead increase

Engineering Contradiction:
Improvedata integrityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies error correction encoding during the initial data storage phase rather than during retrieval. By performing the computationally intensive encoding operation beforehand, the system ensures data integrity is built into the stored slices, eliminating the need for time-consuming verification and correction operations during data access

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11956312B2Testing a storage unit in a storage network
Publication Date: 2024.04.09 PURE STORAGE INC
  • US11956312B2 patent drawing
  • US11956312B2 patent drawing
  • US11956312B2 patent drawing

AI summary

A method for execution by one or more computing devices of a storage network includes identifying a storage unit of a set of storage units for testing, where a data segment of data is error encoded into a set of encoded data slices that is stored in the set of storage units. The method further includes determining whether a threshold number of favorably performing other storage units of the set of storage units will be available during the testing. When the threshold number of favorably performing other storage units will be available, the method further includes initiating the testing of the storage unit and setting a status of the storage unit to unavailable. When the testing has been completed, the method further includes updating the status of the storage unit to available. The method further includes generating a testing report regarding the testing of the storage unit.