Data Protection Appliance In-Line Deduplication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data replication systems face inefficiencies due to the need for post-processing compression and decompression, which burdens CPU resources and introduces latency, especially during data replication across distances.

Innovation Solution

Implementing in-line deduplication and compression using XIO arrays and Quick Assist Technology (QAT) within a 'friendly zone' of storage arrays, where the Data Protection Appliance (DPA) offloads data processing to specialized hardware, optimizing data access and reducing resource consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If post-processing compression and decompression are used in data replication systems, then data can be stored and transmitted, but CPU resources are burdened and latency is introduced

Engineering Contradiction:
Improvedata replication efficiencyVSAvoidCPU resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent extracts the compression and decompression functions from the main data replication path by using a separate compression engine and dedicated hardware modules. This allows compression to occur in parallel or independently, preventing CPU bottlenecks during data replication operations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a compression engine as an intermediary component between the data source and storage system. This intermediary handles all compression and decompression operations, shielding the main replication system from CPU-intensive operations and reducing overall system latency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If post-processing compression and decompression are used in data replication systems, then data can be stored and transmitted, but latency is introduced

Engineering Contradiction:
Improvedata replication efficiencyVSAvoidreplication latency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent implements preliminary compression of data before replication operations begin. By pre-compressing data and preparing compression buffers in advance, the system eliminates time-consuming compression operations from the critical replication path, thereby reducing overall latency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces software-based compression with hardware-based compression modules and dedicated compression engines. This substitution moves compression operations from the software CPU layer to specialized hardware, significantly reducing processing time and replication latency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If hardware-based compression and decompression are implemented, then CPU load is reduced and throughput is improved, but device complexity increases

Engineering Contradiction:
Improvedata replication throughputVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent designs the compression engine and hardware modules to serve multiple functions: compression, decompression, data validation, and replication coordination. By making these components multi-functional, the patent reduces the need for separate dedicated hardware for each function, thereby managing system complexity while maintaining high throughput.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10235064B1Optimized data replication using special NVME protocol and running in a friendly zone of storage array
Publication Date: 2019.03.19 EMC IP HLDG CO LLC
  • US10235064B1 patent drawing
  • US10235064B1 patent drawing
  • US10235064B1 patent drawing

AI summary

A storage system comprises a storage array in operable communication with and responsive to instructions from a host. The storage array comprises a logical unit (LU), a privileged zone, a data protection appliance (DPA), and an intermediary device. The privileged zone is configured within the storage array and permits at least one predetermined process to have access to the LU. The DPA is configured to operate within the privileged zone. The intermediary device is in operable communication with the host, LU, and DPA and is configured to receive read and write instructions from the DPA and to ensure that I/O's passing through the intermediary device to at least one of the LU and the DPA, in response to the reads and writes, are formatted to a first predetermined standard, for at least one of the LU and the DPA.