Storage Load Balancing via Virtual Synthetics Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing load balancing techniques in storage clusters fail to consider efficient storage space utilization on destination nodes when redistributing data, leading to inefficient use of storage capacity.
Innovation Solution
A method that uses virtual synthetics metadata to generate a relationship graph, identifying subsets of files that deduplicate well and migrates them from a source storage node to a destination node, optimizing storage space usage by eliminating redundant data blocks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is redistributed across storage nodes using existing load balancing techniques, then storage utilization is improved, but storage space efficiency deteriorates due to failure to consider deduplication opportunities
Solution Approach 1:
The system performs preliminary analysis of virtual synthetics metadata before data migration to identify files with high deduplication potential. By pre-generating relationship graphs and calculating deduplication metrics, the system prepares migration candidates in advance, ensuring that subsequent migration operations achieve both load balancing and storage space efficiency goals.
Solution Approach 2:
The patent introduces virtual synthetics metadata and relationship graphs as intermediary structures that mediate between source and destination storage nodes. These intermediaries enable the system to evaluate deduplication potential and file relationships without direct inspection of actual data blocks, facilitating intelligent migration decisions that optimize both utilization and space efficiency.
2Productivity
If all files are migrated from source to destination node, then load balancing is achieved, but storage overhead increases due to migration of redundant data blocks
Solution Approach 1:
The system extracts and analyzes virtual synthetics metadata to identify and separate files with high deduplication potential from the general file set. By taking out only the necessary files for migration based on relationship graph analysis, the system avoids migrating redundant data blocks, thereby achieving load balancing while minimizing storage overhead.
Solution Approach 2:
The patent changes the parameter selection criteria for migration by incorporating deduplication metrics and relationship graph analysis. Instead of migrating files based solely on storage utilization thresholds, the system uses multiple parameters including deduplication potential, file relationships, and destination node characteristics to optimize migration decisions and reduce storage overhead.
3Quantity of substance
If storage utilization monitoring triggers migration, then storage capacity is optimized, but migration efficiency decreases without consideration of file relationships
Solution Approach 1:
The system performs preliminary generation of relationship graphs and deduplication metric calculations before migration execution. By pre-processing virtual synthetics metadata and identifying migration candidates in advance, the system reduces the time required for actual migration operations while still achieving optimal storage capacity utilization.
Solution Approach 2:
The patent uses relationship graphs as intermediary structures that encapsulate file relationship information. These graphs serve as mediators between storage utilization monitoring and migration execution, enabling efficient migration decisions without requiring deep analysis of actual file relationships during the migration process itself.
Data Source
AI summary
A method and system for storage load balancing based on virtual synthetics metadata. When storing data onto a storage cluster, data submitted thereto may often be distributed unevenly across the constituent storage nodes thereof. To address the issue, some form of load balancing (or re-distribution of data) across the storage nodes may be implemented. Existing load balancing techniques, however, tend to migrate data between storage nodes without consideration for the efficient utilization of available storage space on the storage node where the data ends up (or destination storage node). Accordingly, the disclosed method and system propose a load balancing mechanism whereby the migrated data dedupes well, thereby securing the efficient consumption of storage space on the destination storage node.


