Multi-Resolution File Replication in Parallel Storage Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing parallel storage systems in high-performance computing environments face inefficiencies in storing replica copies with different resolutions, leading to increased data processing and transfer bandwidth costs and disk space usage.
Innovation Solution
The method involves generating and storing files and their replicas in a parallel computing system using semantic information to create sub-files with varying resolutions, allowing for reduced data processing and transfer bandwidth costs while preserving disk space by selectively storing relevant data subsets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple replica copies of the same data are stored with the same resolution, then reliability and fault-tolerance are improved, but disk space consumption increases
Solution Approach 1:
The patent segments the data storage into multiple resolution levels (full-resolution and reduced-resolution replicas). Instead of storing identical full-resolution copies for all replicas, the system divides replicas into different resolution segments, allowing selective storage based on access requirements and reducing overall disk space consumption while maintaining reliability through multi-resolution redundancy.
Solution Approach 2:
The patent applies local quality by assigning different resolution qualities to different replica copies based on their intended use. Frequently accessed or critical data retains full resolution, while less critical or backup replicas use reduced resolution. This localized quality differentiation maintains necessary reliability while optimizing disk space utilization across the storage system.
2Measurement precision
If full-resolution replica copies are stored, then data accuracy is maintained, but data transfer bandwidth costs increase
Solution Approach 1:
The patent implements dynamic resolution selection where the system can adaptively choose to retrieve full-resolution or reduced-resolution replicas based on real-time access patterns and requirements. This dynamic approach allows the system to maintain data accuracy when necessary while reducing bandwidth consumption during routine operations, optimizing the trade-off between precision and energy costs.
Solution Approach 2:
The patent changes the resolution parameter of replica copies stored in the system. By storing replicas at different resolution levels (parameter variation), the system can satisfy accuracy requirements for specific operations while reducing overall bandwidth consumption for data transfer, particularly for backup and less critical access scenarios.
3Quantity of substance
If reduced-resolution replicas are stored, then disk space is preserved, but data processing complexity increases
Solution Approach 1:
The patent performs preliminary action by pre-processing data during the storage phase to create reduced-resolution replicas. This upfront processing effort is performed once during data ingestion or replication, rather than repeatedly during access operations. The preliminary resolution reduction simplifies subsequent data retrieval and processing operations while maintaining acceptable space efficiency.
Data Source
AI summary
Techniques are provided for storing files in a parallel computing system using different resolutions. A method is provided for storing at least one file generated by a distributed application in a parallel computing system. The file comprises one or more of a complete file and a sub-file. The method comprises the steps of obtaining semantic information related to the file; generating a plurality of replicas of the file with different resolutions based on the semantic information; and storing the file and the plurality of replicas of the file in one or more storage nodes of the parallel computing system. The different resolutions comprise, for example, a variable number of bits and/or a different sub-set of data elements from the file. A plurality of the sub-files can be merged to reproduce the file.


