Cluster File System Data Mover Modules for Storage Quota Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Lustre file systems face challenges in balancing storage capacity and IO throughput, leading to suboptimal performance and excessive costs, particularly in high-performance computing environments where existing storage solutions fail to meet performance requirements.
Innovation Solution
Implementing a cluster file system with a front-end and back-end file system configuration, utilizing data mover modules and a quota manager to control data archiving between the two systems based on user quotas, optimizing hierarchical storage management and avoiding unnecessary archiving of temporary files.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If highly cost effective storage such as scale-out network attached storage is used, then storage capacity and cost are improved, but IO throughput and performance characteristics deteriorate
Solution Approach 1:
The storage system is divided into multiple storage tiers with different performance characteristics. The patent segments storage into high-performance storage for frequently accessed data and high-capacity cost-effective storage for less frequently accessed data, allowing the system to achieve both high IO throughput for active workloads and high storage capacity at lower cost.
Solution Approach 2:
The patent introduces an intermediary layer (such as a cache or buffer) between the compute nodes and the back-end storage arrays. This intermediary enables data to be pre-fetched or cached in high-performance storage before being accessed by compute nodes, thereby maintaining high IO throughput while utilizing cost-effective back-end storage for bulk capacity.
2Quantity of substance
If data is archived to back-end storage arrays, then storage capacity utilization is improved, but IO operation performance deteriorates
Solution Approach 1:
The system performs preliminary actions by pre-fetching data from back-end storage arrays to front-end high-performance storage before compute nodes need to access it. This anticipatory data movement ensures that when compute nodes request data, it is already available in high-speed storage, maintaining fast IO operation speeds while still utilizing back-end storage for capacity.
Solution Approach 2:
The patent implements dynamic data placement and movement between storage tiers based on access patterns and workload demands. Frequently accessed data is automatically moved to or cached in high-performance front-end storage, while less frequently accessed data resides in back-end storage, allowing the system to dynamically optimize both capacity utilization and IO operation speed.
3Adaptability or versatility
If storage devices are not well matched to current system needs, then system flexibility is improved, but performance and cost-effectiveness deteriorate
Solution Approach 1:
The patent creates a universal storage architecture that can accommodate multiple types of storage devices with different performance characteristics. The system is designed to work with various front-end and back-end storage technologies, allowing flexible matching of storage devices to current system needs while maintaining high performance through the multi-tiered architecture that can be optimized for different workload requirements.
Data Source
AI summary
A cluster file system comprises a front-end file system, a back-end file system, data mover modules arranged between the front-end and back-end file systems, and a quota manager associated with at least a given one of data mover modules. The data mover modules are configured to control archiving of data between the front-end file system and the back-end file system for respective users based at least in part on respective user quotas established by the quota manager and identifying corresponding portions of the back-end file system available to the users. The front-end file system may comprise archive directories configured for respective ones of the users, with the data mover modules being configured to archive a given file from the front-end file system in the back-end file system responsive to a link to that file being stored in a corresponding one of the archive directories of the front-end file system.


