Multi-Stage Slice Recovery for Corrupt DSN Encoded Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

RAID systems face challenges with increasing disk failures, maintenance costs, data security, and vulnerability to natural disasters due to the need for manual replacement and multiple data copies, which can lead to unauthorized access and complete data loss.

Innovation Solution

A dispersed storage network (DSN) with a managing unit, integrity processing unit, and computing devices that use error encoding and decoding techniques like Cauchy Reed-Solomon to distribute data across multiple storage units, ensuring data integrity and security without the need for redundant copies, and allowing for recovery from multiple storage unit failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If RAID systems store multiple copies of data across disks, then data security and redundancy are improved, but the risk of unauthorized access increases and maintenance costs rise due to manual replacement requirements

Engineering Contradiction:
Improvedata securityVSAvoidunauthorized access risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments data into multiple encoded slices distributed across different storage units. Instead of storing complete redundant copies like RAID, the data is divided and encoded such that a threshold number of slices are needed for reconstruction. This segmentation approach maintains data security while reducing unauthorized access risk, as no single storage unit contains a complete usable copy of the data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the fundamental parameter of data representation from redundant copies to encoded slices with threshold-based recovery. By using error correction coding schemes, the system transforms data into a form where reliability is achieved through mathematical redundancy rather than physical duplication, thereby improving data security without increasing unauthorized access vulnerability.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If more disks are added to RAID array to improve storage capacity, then storage capacity increases, but the probability of disk failure rises and maintenance costs increase

Engineering Contradiction:
Improvestorage capacityVSAvoiddisk failure probability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent divides data into encoded slices distributed across multiple storage units, allowing the system to scale storage capacity by adding more units without proportionally increasing failure risk. The error correction coding ensures that the system can tolerate a certain number of failures while maintaining data integrity, thus improving reliability as the system scales.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements monitoring and recovery mechanisms that detect storage unit failures and automatically initiate data reconstruction. This feedback loop allows the system to maintain high reliability even as storage capacity and the number of storage units increase, by continuously monitoring system health and responding to failures before they result in data loss.

Inventive Principle:
Principle #23Feedback

3Reliability

If RAID systems require manual disk replacement to maintain data integrity, then data security is maintained, but productivity decreases due to manual intervention requirements

Engineering Contradiction:
Improvedata integrityVSAvoidmaintenance efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent enables the system to automatically detect storage unit failures, retrieve remaining encoded slices, and reconstruct lost data without manual intervention. The error correction coding scheme allows the system to self-heal by mathematically recovering missing slices from available slices, thereby maintaining data integrity while eliminating the need for manual disk replacement and significantly improving maintenance productivity.

Inventive Principle:
Principle #25Self-service

4Reliability

If data is copied to multiple RAID devices to prevent data loss, then reliability improves, but device complexity increases and security risks arise from multiple access points

Engineering Contradiction:
Improvedata loss preventionVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments data into encoded slices distributed across storage units, creating a simpler system architecture compared to traditional RAID copying mechanisms. Each storage unit holds only a portion of the encoded data, and the system requires a threshold number of slices for reconstruction. This approach prevents data loss while reducing system complexity and minimizing security risks by eliminating multiple complete data copies.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10025665B2Multi-stage slice recovery in a dispersed storage network
Publication Date: 2018.07.17 PURE STORAGE INC
  • US10025665B2 patent drawing
  • US10025665B2 patent drawing
  • US10025665B2 patent drawing

AI summary

A method for use by a computing device in a dispersed storage network (DSN) to recover corrupt encoded data slices. In response to a request to storage units of the DSN for encoded data slices corresponding to a data segment, the computing device of a receives less than a decode threshold number of valid encoded data slices and at least one integrity error message that provides an indication of a corrupt encoded data slice. The computing device requests and receives at least one corrupt encoded data slice corresponding to the integrity error message(s). Utilizing at least one correction approach involving stored integrity data, the computing device then corrects the corrupt slice(s) to produce a decode threshold number of encoded data slices in order to decode the corresponding data segment. A variety of correction approaches may be employed, including a multi-stage approach that utilizes data from both valid and invalid slices.