Out-of-Core Similarity Matching for Delta Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Delta encoding in data storage systems faces inefficiencies due to the challenge of selecting which data portion to encode and relative to which portion, leading to redundant storage and reassembly requirements.

Innovation Solution

The system divides data into chunks, assigns unique fingerprints or representative values using hash functions, and employs similarity matching to identify base data chunks for delta encoding, allowing efficient storage of relative differences and reducing redundant data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If delta encoding is implemented to compress data, then storage space is reduced, but data reassembly complexity increases

Engineering Contradiction:
Improvestorage spaceVSAvoiddata reassembly complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing fingerprints for all data chunks before delta encoding. These fingerprints serve as pre-prepared indexes that enable efficient matching and reassembly without complex computations during the reassembly phase. The recipe file is also prepared in advance with all necessary reconstruction information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces fingerprints as an intermediary element between original data chunks and delta-encoded representations. These fingerprints act as mediators that enable efficient similarity matching and identification of base chunks during reassembly, simplifying the overall process by providing a straightforward lookup mechanism rather than requiring complex direct comparisons.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If data is divided into chunks for delta encoding, then compression efficiency improves, but processing overhead increases

Engineering Contradiction:
Improvecompression efficiencyVSAvoidprocessing overhead
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies segmentation by dividing data into fixed-size chunks with overlapping regions. This segmentation enables parallel processing of multiple chunks independently while maintaining compression efficiency through the overlap regions that capture similarities between adjacent chunks. The segmented approach allows the system to process large files by working with manageable chunk-sized units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements partial action by using a fixed number of overlap regions (e.g., 3 overlap regions) between chunks rather than computing similarities for all possible chunk pairs. This partial approach to similarity computation significantly reduces processing overhead while still achieving effective compression by capturing the most relevant similarities in the overlapping regions.

Inventive Principle:
Principle #16Partial or excessive action

3Quantity of substance

If similarity matching is performed to identify base chunks, then compression ratio improves, but computational complexity increases

Engineering Contradiction:
Improvecompression ratioVSAvoidcomputational complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical system of direct data comparison with a hash-based fingerprint matching system. Instead of computationally intensive byte-by-byte comparison to identify similar chunks, the system uses hash functions to generate fingerprints and performs efficient fingerprint matching. This substitution dramatically reduces computational complexity while maintaining accurate identification of similar chunks for effective delta encoding.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9727573B1Out-of core similarity matching
Publication Date: 2017.08.08 EMC IP HLDG CO LLC
  • US9727573B1 patent drawing
  • US9727573B1 patent drawing
  • US9727573B1 patent drawing

AI summary

A method for storing data in a data storage system by partitioning the data into a plurality of data chunks and generating representative data for each of the plurality of chunks by applying a predetermined algorithm to each chunk of the plurality of chunks. Subsequently, the representative data is compared and sorted. Representative data for base data chunks and representative data for other data chunks that can be stored relative to the base data chunks are identified by evaluating the sorted set of representative data. Finally, each of the other data chunks identified as those that can be stored relative to a base data chunk are stored in the data storage system as the difference between the data chunk and a base data chunk.