Deduplication Database Segmentation for Heterogeneous Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large deduplication systems face performance reduction due to the inefficiencies in managing heterogeneous data types, leading to increased storage costs and reduced access speeds as the database grows, necessitating improved scalability and performance.
Innovation Solution
Implementing a deduplicated storage system that assigns deduplication databases based on client device types and automatically creates new databases when critical thresholds are reached, with further partitioning of databases based on data block distribution policies to enhance efficiency and speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a single deduplication database stores all data types, then storage utilization is improved, but access speed and efficiency deteriorate as the database grows
Solution Approach 1:
The patent divides the single deduplication database into multiple separate databases, with each database dedicated to storing signatures of a specific data type (e.g., application files, databases, virtual machines). This segmentation allows each database to remain smaller and more manageable, maintaining faster access speeds while collectively providing comprehensive storage utilization across all data types.
2Device complexity
If heterogeneous data types are stored in the same deduplication database, then device complexity is reduced, but productivity deteriorates due to reduced deduplication efficiency
Solution Approach 1:
The system segments the deduplication database by data type, creating separate databases for different heterogeneous data types. This segmentation improves deduplication efficiency by allowing optimized processing for each data type while the overall system complexity is managed through automated database creation and assignment mechanisms.
Solution Approach 2:
The system implements self-service mechanisms where the deduplication database automatically creates new databases when thresholds are reached and automatically assigns databases to client devices based on their data types. This automation reduces manual management complexity while maintaining high deduplication efficiency across heterogeneous data types.
3Quantity of substance
If the deduplication database approaches disk space capacity, then storage utilization is maximized, but I/O operation speed substantially reduces
Solution Approach 1:
By segmenting the large deduplication database into multiple smaller databases organized by data type, the system prevents any single database from approaching disk space capacity. This segmentation maintains I/O operation speeds at acceptable levels while collectively maximizing storage capacity utilization across all data types through the distributed database structure.
Data Source
AI summary
A deduplicated storage system is provided according to certain embodiments that uses one or more mechanisms to assign the deduplication databases based on the type of the client device and automatically create a new deduplication database when critical thresholds are reached. In other embodiments, deduplication databases are further split into multiple database partitions. Based on a data block distribution policy, each data block is then further assigned to a particular database partition within the deduplication database to further improve efficiency and speed of the deduplication process.


