Application Aware Deduplication for Random Access to Compressed Files

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deduplication file systems experience high latency when handling random access requests to compressed files, particularly in primary storage systems where fast and efficient access is crucial, due to inefficient chunking and compression algorithms that fail to deduplicate similar files effectively.

Innovation Solution

The system maintains compressed files in a decompressed format on the deduplication file system, using metadata files to map the structure and location of subfiles, allowing for efficient random access and high deduplication ratios by treating files as sets of separate subfiles rather than a single entity, and employing native compression algorithms for global compression.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If deduplication is applied to compressed files in a primary file system, then storage efficiency is improved, but random access latency increases

Engineering Contradiction:
Improvestorage efficiencyVSAvoidrandom access latency
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent segments compressed files into multiple chunks and stores them separately in the deduplication file system. Each chunk can be independently accessed and deduplicated, allowing random access to specific portions of files without requiring decompression of the entire file. This segmentation enables both high deduplication ratios and fast random access performance.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary compression of files before storing them in the deduplication file system. By pre-compressing files into manageable chunks and storing metadata about their locations and structures, the system prepares the data in advance for efficient random access operations, eliminating the need for full decompression during access operations.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If compressed files are stored as single entities, then file integrity is maintained, but deduplication efficiency decreases

Engineering Contradiction:
Improvefile integrityVSAvoiddeduplication efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent divides compressed files into multiple independent chunks that can be separately stored and deduplicated. Each chunk maintains its integrity through individual checksums and metadata, while the overall file integrity is preserved through the ability to reconstruct the original file from its chunks. This segmentation dramatically improves deduplication efficiency by allowing fine-grained comparison and matching of similar file portions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces metadata as an intermediary layer between the compressed file chunks and the storage system. This metadata contains information about chunk boundaries, offsets, and integrity checks, enabling the system to maintain file integrity while treating chunks as separate deduplication units. The metadata acts as a mediator that preserves the logical structure of the original file while enabling efficient physical storage and retrieval.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11061867B2Application aware deduplication allowing random access to compressed files
Publication Date: 2021.07.13 EMC IP HLDG CO LLC
  • US11061867B2 patent drawing
  • US11061867B2 patent drawing
  • US11061867B2 patent drawing

AI summary

A file is received from a client for storage at a deduplication file system. The file is in an archive file format that is used by an application on the client. The file includes subfiles compressed together in the file according to the archive file format, local headers corresponding to the subfiles, and a central directory used by the application to locate information stored in the file. The file is decompressed to store the subfiles separately. A metadata file is created that describes a structure of the file. The metadata file includes the local headers, central directory, pointers to the subfiles, but does not include the subfiles. The file is presented to the client as a single file having the archive file format. A request from the client is received to read the file and the metadata file is read to return data responsive to the request.