Out-of-Core Similarity Matching for Delta Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Delta encoding in data storage systems faces inefficiencies due to the challenge of selecting which data portion to encode and relative to which portion, leading to redundant storage and reassembly requirements.
Innovation Solution
The system divides data into chunks, assigns unique fingerprints or representative values using hash functions, and employs similarity matching to identify base data chunks for delta encoding, allowing efficient storage of relative differences and reducing redundant data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If delta encoding is implemented to compress data, then storage space is reduced, but data reassembly complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing fingerprints for all data chunks before delta encoding. These fingerprints serve as pre-prepared indexes that enable efficient matching and reassembly without complex computations during the reassembly phase. The recipe file is also prepared in advance with all necessary reconstruction information.
Solution Approach 2:
The patent introduces fingerprints as an intermediary element between original data chunks and delta-encoded representations. These fingerprints act as mediators that enable efficient similarity matching and identification of base chunks during reassembly, simplifying the overall process by providing a straightforward lookup mechanism rather than requiring complex direct comparisons.
2Productivity
If data is divided into chunks for delta encoding, then compression efficiency improves, but processing overhead increases
Solution Approach 1:
The patent applies segmentation by dividing data into fixed-size chunks with overlapping regions. This segmentation enables parallel processing of multiple chunks independently while maintaining compression efficiency through the overlap regions that capture similarities between adjacent chunks. The segmented approach allows the system to process large files by working with manageable chunk-sized units.
Solution Approach 2:
The patent implements partial action by using a fixed number of overlap regions (e.g., 3 overlap regions) between chunks rather than computing similarities for all possible chunk pairs. This partial approach to similarity computation significantly reduces processing overhead while still achieving effective compression by capturing the most relevant similarities in the overlapping regions.
3Quantity of substance
If similarity matching is performed to identify base chunks, then compression ratio improves, but computational complexity increases
Solution Approach 1:
The patent replaces the mechanical system of direct data comparison with a hash-based fingerprint matching system. Instead of computationally intensive byte-by-byte comparison to identify similar chunks, the system uses hash functions to generate fingerprints and performs efficient fingerprint matching. This substitution dramatically reduces computational complexity while maintaining accurate identification of similar chunks for effective delta encoding.
Data Source
AI summary
A method for storing data in a data storage system by partitioning the data into a plurality of data chunks and generating representative data for each of the plurality of chunks by applying a predetermined algorithm to each chunk of the plurality of chunks. Subsequently, the representative data is compared and sorted. Representative data for base data chunks and representative data for other data chunks that can be stored relative to the base data chunks are identified by evaluating the sorted set of representative data. Finally, each of the other data chunks identified as those that can be stored relative to a base data chunk are stored in the data storage system as the difference between the data chunk and a base data chunk.


