Deduplicated Data Packing via Similarity Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The inefficiencies in replicating deduplicated data to remote sites due to high bandwidth requirements and resource strain, as well as the need for proportional physical cartridges to user data rather than deduplicated size, hinder the performance and efficiency of data center operations.
Innovation Solution
Calculating a similarity score between deduplicated data files to group them into subsets for packaging into finite-sized containers, utilizing transitive closures to optimize storage and reduce bandwidth demands, allowing for efficient data export and backup.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If deduplicated data is replicated to remote sites, then data backup and replication are achieved, but bandwidth requirements increase and network resources are strained
Solution Approach 1:
The patent segments deduplicated data into distinct subsets based on similarity scores, allowing selective replication of only the most similar data groups to remote sites. This segmentation enables targeted data transmission that reduces overall bandwidth consumption while maintaining backup reliability.
Solution Approach 2:
The patent changes the parameter of data selection by using similarity scores to identify and prioritize which deduplicated data subsets should be replicated. By transforming the selection criterion from random or sequential to similarity-based, the system optimizes bandwidth usage while ensuring critical data is backed up.
2Quantity of substance
If physical cartridges are allocated proportional to user data size, then adequate storage capacity is provided, but storage resources are wasted due to lack of deduplication consideration
Solution Approach 1:
The patent performs preliminary grouping of deduplicated data into similarity-based subsets before allocating physical cartridges. This preliminary action allows the system to understand the actual deduplicated data volume and structure, enabling accurate cartridge allocation that matches real storage needs rather than raw user data size.
Solution Approach 2:
The patent creates a virtual representation of deduplicated data organized by similarity scores, which serves as a blueprint for physical cartridge allocation. This virtual copy allows the system to plan storage resources efficiently based on actual deduplicated content rather than original user data volumes.
3Reliability
If rehydration process is used to export deduplicated data, then complete data restoration is achieved, but data center resources and bandwidth are stretched
Solution Approach 1:
The patent extracts only the necessary deduplicated data subsets that meet similarity threshold criteria for export, rather than performing full rehydration of all deduplicated data. This extraction approach maintains data restoration reliability for critical information while avoiding the resource-intensive process of reconstructing entire data sets.
4Loss of energy
If similarity scoring and grouping is implemented for data export, then storage and bandwidth efficiency are improved, but computational overhead increases
Solution Approach 1:
The patent applies partial action by calculating similarity scores only for data subsets that are candidates for export, rather than computing similarities for all possible data pairs. This partial computation approach reduces computational overhead while still achieving the bandwidth efficiency benefits of similarity-based grouping.
Data Source
AI summary
Deduplicated data is packed into finite-sized containers. A similarity score is calculated between files that are similarly of the deduplicated data. The similarity score is used for grouping the similarly compared files of the deduplicated data into subsets for destaging each of the subsets from a deduplication system to one a finite-sized container. The similarity score is used for grouping the similarly compared files of the deduplicated data into subsets for destaging each of the subsets from a deduplication system to one of the finite-sized containers. An indication is received by a user of which of the similarly compared files are to be grouped into the subsets for destaging each of the subsets from a deduplication system to one of the finite-sized containers. Transitive closures are used for assisting with using the similarity score for grouping the similarly compared files of the deduplicated data into the subsets.


