Deduplication Database Segmentation for Heterogeneous Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large deduplication systems face performance reduction due to the inefficiencies in managing heterogeneous data types, leading to increased storage costs and reduced access speeds as the database grows, necessitating improved scalability and performance.

Innovation Solution

Implementing a deduplicated storage system that assigns deduplication databases based on client device types and automatically creates new databases when critical thresholds are reached, with further partitioning of databases based on data block distribution policies to enhance efficiency and speed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a single deduplication database stores all data types, then storage utilization is improved, but access speed and efficiency deteriorate as the database grows

Engineering Contradiction:
Improvestorage utilizationVSAvoidaccess speed
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The patent divides the single deduplication database into multiple separate databases, with each database dedicated to storing signatures of a specific data type (e.g., application files, databases, virtual machines). This segmentation allows each database to remain smaller and more manageable, maintaining faster access speeds while collectively providing comprehensive storage utilization across all data types.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If heterogeneous data types are stored in the same deduplication database, then device complexity is reduced, but productivity deteriorates due to reduced deduplication efficiency

Engineering Contradiction:
Improvedatabase management complexityVSAvoiddeduplication efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The system segments the deduplication database by data type, creating separate databases for different heterogeneous data types. This segmentation improves deduplication efficiency by allowing optimized processing for each data type while the overall system complexity is managed through automated database creation and assignment mechanisms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements self-service mechanisms where the deduplication database automatically creates new databases when thresholds are reached and automatically assigns databases to client devices based on their data types. This automation reduces manual management complexity while maintaining high deduplication efficiency across heterogeneous data types.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If the deduplication database approaches disk space capacity, then storage utilization is maximized, but I/O operation speed substantially reduces

Engineering Contradiction:
Improvestorage capacity utilizationVSAvoidI/O operation speed
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

By segmenting the large deduplication database into multiple smaller databases organized by data type, the system prevents any single database from approaching disk space capacity. This segmentation maintains I/O operation speeds at acceptable levels while collectively maximizing storage capacity utilization across all data types through the distributed database structure.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11663178B2Efficient implementation of multiple deduplication databases in a heterogeneous data storage system
Publication Date: 2023.05.30 COMMVAULT SYSTEMS INC
  • US11663178B2 patent drawing
  • US11663178B2 patent drawing
  • US11663178B2 patent drawing

AI summary

A deduplicated storage system is provided according to certain embodiments that uses one or more mechanisms to assign the deduplication databases based on the type of the client device and automatically create a new deduplication database when critical thresholds are reached. In other embodiments, deduplication databases are further split into multiple database partitions. Based on a data block distribution policy, each data block is then further assigned to a particular database partition within the deduplication database to further improve efficiency and speed of the deduplication process.