Parity Group Migration via Non-Deterministic Addressing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage architectures for high-performance computing lack efficient methods for migrating Parity Groups between high-performance compute nodes and long-term storage, particularly in handling unstructured data and ensuring fault tolerance and redundancy.
Innovation Solution
A data migration system utilizing a Burst Buffer tier with Non-Volatile Memory (NVM) arrays and a Distributed Hash Table, which supports the construction, ingestion, and orderly egress of Parity Groups, employing Parity Group Information descriptors for non-deterministic data addressing and ensuring coherency and fault tolerance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If data is stored in an unstructured manner in the intermediate storage tier for fast ingestion, then data ingestion speed is improved, but data retrieval and reclamation becomes complex and inefficient
Solution Approach 1:
The system segments data into Parity Groups with structured organization within the unstructured storage tier. Each Parity Group is assigned a unique identifier and metadata that enables efficient retrieval without requiring full structure enforcement during ingestion. This allows fast unstructured ingestion while maintaining organized access paths through the parity group abstraction.
Solution Approach 2:
The patent introduces an intermediate indexing layer that mediates between the unstructured storage tier and retrieval operations. This layer maintains metadata about Parity Groups including location information, enabling efficient data location and reclamation without imposing structure on the underlying storage mechanism. The intermediary layer handles the complexity of retrieval while allowing flexible ingestion.
2Reliability
If Parity Groups are distributed throughout the intermediate storage tier for fault tolerance, then data reliability is improved, but system complexity for tracking and managing distributed data increases
Solution Approach 1:
The Parity Group structure serves multiple functions simultaneously: it provides fault tolerance through distribution, enables efficient retrieval through structured metadata, and simplifies management through unified identification. Each Parity Group acts as a universal container that can be stored unstructured yet retrieved efficiently, managing both reliability and complexity requirements.
Solution Approach 2:
The system implements feedback mechanisms through metadata tracking that monitors the location and status of distributed Parity Groups. This feedback enables the system to track distributed data without manual intervention, automatically updating location information and managing the complexity of distributed storage while maintaining fault tolerance.
3Productivity
If non-deterministic write methods are used for fast data ingestion, then ingestion performance is improved, but data location and consistency becomes difficult to ensure
Solution Approach 1:
The system performs preliminary actions by assigning unique identifiers and metadata to Parity Groups before they are stored in the unstructured tier. This preliminary tagging enables deterministic location of data even though the physical storage location is non-deterministic. The metadata is prepared in advance, allowing fast ingestion without sacrificing location accuracy.
Solution Approach 2:
The patent replaces mechanical positioning systems with information-based location methods. Instead of relying on fixed physical locations for data storage, the system uses metadata and identifiers to logically track data positions. This substitution allows non-deterministic physical storage while maintaining precise logical location tracking through software-based addressing.
Data Source
AI summary
The present invention is directed to data migration, and particularly, Parity Group migration, between high performance data generating entities and data storage structure in which distributed NVM arrays are used as a single intermediate logical storage which requires a global registry/addressing capability that facilitates the storage and retrieval of the locality information (metadata) for any given fragment of unstructured data and where Parity Group Identifier and Parity Group Information (PGI) descriptors for the Parity Groups' members tracking, are created and distributed in the intermediate distributed NVM arrays as a part of the non-deterministic data addressing system to ensure coherency and fault tolerance for the data and the metadata. The PGI descriptors act as collection points for state describing the residency and replay status of members of the Parity Groups.


