Inline Wire Speed Data Deduplication Engine
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data compression techniques, particularly lossless compression methods, are not well-suited for high-speed, low-latency applications and face performance issues in systems using thin provisioning with sparse mapping, leading to file fragmentation and reduced caching efficiency.
Innovation Solution
The implementation of an inline wire speed data deduplication system that includes an inline deduplication engine capable of processing data at rates of at least 4 Gigabytes per second, utilizing hardware-assisted deduplication blades for efficient data deduplication and compression, which identifies and eliminates duplicate data chunks by storing only one instance and using pointers for repeated data sequences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If software-based lossless compression techniques are used, then data storage efficiency is improved, but processing speed and latency are degraded
Solution Approach 1:
The patent replaces software-based compression mechanisms with hardware-based compression circuits that operate in parallel, achieving wire-speed compression without the sequential processing limitations of software implementations
Solution Approach 2:
The compression system is divided into independent parallel processing units that can simultaneously compress different data streams, with each unit handling a portion of the total data load to maintain high throughput
2Ease of manufacture
If hardware-based compression processes one byte at a time, then processing speed is limited to clock frequency, but implementation simplicity is improved
Solution Approach 1:
The data stream is segmented into parallel lanes, with each lane having its own compression unit operating simultaneously, thereby multiplying the effective processing bandwidth beyond single-clock-frequency limitations
Solution Approach 2:
Multiple parallel compression units are merged into a unified system that processes data across multiple lanes simultaneously, combining their individual capacities to achieve aggregate throughput exceeding any single unit's capability
3Adaptability or versatility
If data is distributed over different areas within storage devices using thin provisioning and sparse mapping, then storage flexibility is improved, but caching efficiency is degraded due to file fragmentation
Solution Approach 1:
The patent introduces an intermediary caching layer that sits between the distributed storage system and the host, pre-fetching and caching data blocks that are likely to be accessed, thereby compensating for the fragmentation caused by sparse mapping
Solution Approach 2:
The system performs preliminary data retrieval and caching operations before actual host access patterns are fully established, proactively positioning data in the cache to minimize future access latency despite distributed storage layout
Data Source
AI summary
Systems for performing inline wire speed data deduplication are described herein. Some embodiments include a device for inline data deduplication that includes one or more input ports for receiving an input data stream containing duplicates, one or more output ports for providing a data deduplicated output data stream, and an inline data deduplication engine coupled to said one or more input ports and said one or more output ports to process input data containing duplicates into output data which is data deduplicated, said inline data deduplication engine having an inline data deduplication bandwidth of at least 4 Gigabytes per second.


