Application Aware Deduplication for Random Access to Compressed Files
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current deduplication file systems experience high latency when handling random access requests to compressed files, particularly in primary storage systems where fast and efficient access is crucial, due to inefficient chunking and compression algorithms that fail to deduplicate similar files effectively.
Innovation Solution
The system maintains compressed files in a decompressed format on the deduplication file system, using metadata files to map the structure and location of subfiles, allowing for efficient random access and high deduplication ratios by treating files as sets of separate subfiles rather than a single entity, and employing native compression algorithms for global compression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If deduplication is applied to compressed files in a primary file system, then storage efficiency is improved, but random access latency increases
Solution Approach 1:
The patent segments compressed files into multiple chunks and stores them separately in the deduplication file system. Each chunk can be independently accessed and deduplicated, allowing random access to specific portions of files without requiring decompression of the entire file. This segmentation enables both high deduplication ratios and fast random access performance.
Solution Approach 2:
The system performs preliminary compression of files before storing them in the deduplication file system. By pre-compressing files into manageable chunks and storing metadata about their locations and structures, the system prepares the data in advance for efficient random access operations, eliminating the need for full decompression during access operations.
2Reliability
If compressed files are stored as single entities, then file integrity is maintained, but deduplication efficiency decreases
Solution Approach 1:
The patent divides compressed files into multiple independent chunks that can be separately stored and deduplicated. Each chunk maintains its integrity through individual checksums and metadata, while the overall file integrity is preserved through the ability to reconstruct the original file from its chunks. This segmentation dramatically improves deduplication efficiency by allowing fine-grained comparison and matching of similar file portions.
Solution Approach 2:
The patent introduces metadata as an intermediary layer between the compressed file chunks and the storage system. This metadata contains information about chunk boundaries, offsets, and integrity checks, enabling the system to maintain file integrity while treating chunks as separate deduplication units. The metadata acts as a mediator that preserves the logical structure of the original file while enabling efficient physical storage and retrieval.
Data Source
AI summary
A file is received from a client for storage at a deduplication file system. The file is in an archive file format that is used by an application on the client. The file includes subfiles compressed together in the file according to the archive file format, local headers corresponding to the subfiles, and a central directory used by the application to locate information stored in the file. The file is decompressed to store the subfiles separately. A metadata file is created that describes a structure of the file. The metadata file includes the local headers, central directory, pointers to the subfiles, but does not include the subfiles. The file is presented to the client as a single file having the archive file format. A request from the client is received to read the file and the metadata file is read to return data responsive to the request.


