Deduplicated Data Packing via Similarity Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The inefficiencies in replicating deduplicated data to remote sites due to high bandwidth requirements and resource strain, as well as the need for proportional physical cartridges to user data rather than deduplicated size, hinder the performance and efficiency of data center operations.

Innovation Solution

Calculating a similarity score between deduplicated data files to group them into subsets for packaging into finite-sized containers, utilizing transitive closures to optimize storage and reduce bandwidth demands, allowing for efficient data export and backup.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If deduplicated data is replicated to remote sites, then data backup and replication are achieved, but bandwidth requirements increase and network resources are strained

Engineering Contradiction:
Improvedata backupVSAvoidbandwidth usage
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent segments deduplicated data into distinct subsets based on similarity scores, allowing selective replication of only the most similar data groups to remote sites. This segmentation enables targeted data transmission that reduces overall bandwidth consumption while maintaining backup reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of data selection by using similarity scores to identify and prioritize which deduplicated data subsets should be replicated. By transforming the selection criterion from random or sequential to similarity-based, the system optimizes bandwidth usage while ensuring critical data is backed up.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If physical cartridges are allocated proportional to user data size, then adequate storage capacity is provided, but storage resources are wasted due to lack of deduplication consideration

Engineering Contradiction:
Improvestorage capacityVSAvoidstorage efficiency
Core Design Contradiction:
Quantity of substanceVSLoss of substance

Solution Approach 1:

The patent performs preliminary grouping of deduplicated data into similarity-based subsets before allocating physical cartridges. This preliminary action allows the system to understand the actual deduplicated data volume and structure, enabling accurate cartridge allocation that matches real storage needs rather than raw user data size.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a virtual representation of deduplicated data organized by similarity scores, which serves as a blueprint for physical cartridge allocation. This virtual copy allows the system to plan storage resources efficiently based on actual deduplicated content rather than original user data volumes.

Inventive Principle:
Principle #26Copying

3Reliability

If rehydration process is used to export deduplicated data, then complete data restoration is achieved, but data center resources and bandwidth are stretched

Engineering Contradiction:
Improvedata restorationVSAvoidresource utilization
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts only the necessary deduplicated data subsets that meet similarity threshold criteria for export, rather than performing full rehydration of all deduplicated data. This extraction approach maintains data restoration reliability for critical information while avoiding the resource-intensive process of reconstructing entire data sets.

Inventive Principle:
Principle #2Taking out (Extraction)

4Loss of energy

If similarity scoring and grouping is implemented for data export, then storage and bandwidth efficiency are improved, but computational overhead increases

Engineering Contradiction:
Improvebandwidth usageVSAvoidcomputational complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent applies partial action by calculating similarity scores only for data subsets that are candidates for export, rather than computing similarities for all possible data pairs. This partial computation approach reduces computational overhead while still achieving the bandwidth efficiency benefits of similarity-based grouping.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11079953B2Packing deduplicated data into finite-sized containers
Publication Date: 2021.08.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11079953B2 patent drawing
  • US11079953B2 patent drawing
  • US11079953B2 patent drawing

AI summary

Deduplicated data is packed into finite-sized containers. A similarity score is calculated between files that are similarly of the deduplicated data. The similarity score is used for grouping the similarly compared files of the deduplicated data into subsets for destaging each of the subsets from a deduplication system to one a finite-sized container. The similarity score is used for grouping the similarly compared files of the deduplicated data into subsets for destaging each of the subsets from a deduplication system to one of the finite-sized containers. An indication is received by a user of which of the similarly compared files are to be grouped into the subsets for destaging each of the subsets from a deduplication system to one of the finite-sized containers. Transitive closures are used for assisting with using the similarity score for grouping the similarly compared files of the deduplicated data into the subsets.