Metadata-Driven Data Placement in Distributed Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed data storage systems, existing methods globally optimize data availability across all servers, which can lead to inefficient data retrieval times, especially for real-time data like video, and lack control over data placement and migration, limiting optimization for specific business needs and server utilization.
Innovation Solution
A metadata-driven approach that allows clients to bind data files to specific server classes based on defined criteria, enabling localized control over data placement and migration within a distributed storage system, using a class database to manage server selection and replication, allowing for dynamic partitioning and optimization of data storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is stored globally optimized across all servers, then data availability is improved, but data retrieval time deteriorates for real-time data
Solution Approach 1:
The patent applies local quality by allowing different data to be placed on different server classes based on specific requirements. Real-time data can be directed to high-performance servers while less time-sensitive data can be stored on standard servers, optimizing both availability and retrieval time for different data types simultaneously
Solution Approach 2:
The patent segments the server infrastructure into multiple server classes with different performance characteristics. This segmentation allows the system to match specific data requirements with appropriate server capabilities, resolving the contradiction between global optimization and specialized performance needs
2Reliability
If global optimization is used for data placement, then overall system availability is improved, but control over data placement and migration deteriorates
Solution Approach 1:
The patent implements dynamic data placement where data can be automatically migrated between server classes based on changing requirements. This dynamic approach maintains system availability while providing flexible control, as the system can adapt data placement in real-time based on performance metrics and business needs
Solution Approach 2:
The system uses feedback mechanisms to monitor data access patterns and server performance, automatically adjusting data placement decisions. This feedback loop maintains high availability while providing controlled optimization, as the system learns from actual usage patterns rather than relying on static global optimization
3Device complexity
If data is placed without directed control, then system simplicity is maintained, but server utilization optimization deteriorates
Solution Approach 1:
The patent implements self-service data placement where the system automatically directs data to appropriate server classes based on predefined criteria and metadata. This self-service approach maintains system simplicity from the user perspective while achieving optimized server utilization through automated intelligent placement decisions
Data Source
AI summary
A data processing apparatus, comprising a metadata store storing information about files that are stored in a distributed data storage system, and comprising a class database; one or more processing units; logic configured for receiving and storing in the class database a definition of a class of data storage servers comprising one or more subclasses each comprising one or more server selection criteria; associating the class with one or more directories of the data storage system; in response to a data client storing a data file in a directory, binding the class to the data file, determining and storing a set of identifiers of one or more data storage servers in the system that match the server selection criteria, and providing the set of identifiers to the data client.


