Distributed Storage Encoding for Multi-Node Local Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data loss restoration techniques in distributed storage systems are limited in guaranteeing jointly optimized locality for the loss of two or more nodes, requiring complex processes and higher storage capacity.

Innovation Solution

A method of encoding that divides target data into blocks, generates information and parity symbols using a generator matrix with specific Hamming weights, and stores these symbols across nodes to ensure efficient recovery of lost nodes with reduced parity symbols and lower complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing data loss restoration encoding techniques are used, then locality for a specific number of node losses is guaranteed, but jointly optimized locality for two or more specific numbers of nodes cannot be guaranteed

Engineering Contradiction:
Improvelocality guaranteeVSAvoidjointly optimized locality
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The encoding process is segmented into multiple stages: dividing target data into data blocks, generating information symbols, selecting two different data blocks, and generating parity symbols. This segmentation allows the system to handle different node loss scenarios independently and combine them for jointly optimized locality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects two different data blocks from the divided data blocks based on the encoding requirements. This dynamic selection enables the encoding process to adapt to different node loss scenarios and achieve jointly optimized locality for multiple specific numbers of nodes

Inventive Principle:
Principle #15Dynamics

2Reliability

If repetition encoding method is used to restore lost data, then data recovery is possible when one node is lost, but the number of nodes required increases and storage capacity increases

Engineering Contradiction:
Improvedata recovery capabilityVSAvoidnumber of nodes
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The invention extracts only the necessary parity symbols needed for data recovery instead of creating full repetitions of data blocks. By selecting two different data blocks and generating parity symbols from them, the system recovers lost data with fewer nodes and reduced storage capacity

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the encoding parameters by using a specific encoding process that generates parity symbols from selected data blocks rather than creating complete copies. This parameter change reduces the number of nodes required while maintaining data recovery capability

Inventive Principle:
Principle #35Parameter changes

3Reliability

If existing encoding techniques are used to guarantee locality, then data recovery is possible, but device complexity increases

Engineering Contradiction:
Improvedata restorationVSAvoidencoding process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The encoding process performs partial action by selecting only two different data blocks out of multiple divided data blocks to generate parity symbols. This partial action reduces the complexity of the encoding process while still achieving effective data recovery for multiple node losses

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9864550B2Method and apparatus of recovering and encoding for data recovery in storage system
Publication Date: 2018.01.09 IND ACADEMIC COOP FOUND YONSEI UNIV
  • US9864550B2 patent drawing
  • US9864550B2 patent drawing
  • US9864550B2 patent drawing

AI summary

Provided are a method and an apparatus of recovering and encoding for data recovery in a storage system and a distributed storage system of supporting, when a node storing data is lost in a distributed storage environment, a function to recover the lost node. According to exemplary embodiments of the present invention, in a method of encoding for recovering data loss in a distributed storage system, it is possible to guarantee jointly optimized locality with respect to a loss of two or more specific numbers of nodes and recover data of lost nodes in the distributed storage system by using a smaller number of nodes while using a smaller storage capacity.