Pipeline Planning for Low Latency Archival Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in archival storage systems is to reduce costs while maintaining high storage density, often requiring the sacrifice of computing capabilities and storage access bandwidth, as conventional solutions struggle to manage large volumes of infrequently accessed data effectively.
Innovation Solution
The implementation of multiple-data-storage-devices cartridges with low-cost, low-quality data storage devices that are designed for limited lifespans, combined with a data range API and parallel processing capabilities, allows for efficient storage and access of 'cold' data through a scalable and cost-effective archival storage system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional archival storage solutions use low-cost, low-quality data storage devices with limited lifespans, then storage cost is reduced and storage density is increased, but computing capabilities and storage access bandwidth are sacrificed
Solution Approach 1:
The system segments storage operations into separate modules: data reception module, pipeline planning module, and data storage module. This segmentation allows each module to be optimized independently, enabling the system to maintain high storage density with low-cost devices while preserving access bandwidth through efficient pipeline management and parallel processing capabilities.
2Quantity of substance
If conventional archival storage solutions use low-cost, low-quality data storage devices with limited lifespans, then storage cost is reduced and storage density is increased, but computing capabilities are sacrificed
Solution Approach 1:
The pipeline planning module performs preliminary actions by pre-planning storage operations and preparing data processing pipelines before actual storage operations occur. This allows the system to optimize computing resource usage in advance, maintaining computing capabilities while using low-cost storage devices, as the complex processing is orchestrated beforehand rather than requiring high-performance computing during storage operations.
Data Source
AI summary
A write request including payload data is received. The payload data of the write request is stored in a staging area of a storage manager. A transformation pipeline is determined based, at least in part, on an attribute of the write request. The transformation pipeline is queued for execution. Data fragments are generated based, at least in part, on the payload data and the transformation pipeline. The data fragments are transmitted to a plurality of enclosures.


