NVMe Parity Offload via PCIe Accelerator Engine

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for generating parities, such as XOR parity, are CPU and memory intensive, leading to performance bottlenecks, especially when scaling with high-performance NVMe drives, and there is a need for more efficient methods to offload these operations without burdening the central processing unit.

Innovation Solution

Utilizing an accelerator engine connected via PCIe endpoint and NVMe transport protocol to generate parities, with direct memory access (DMA) for communication, offloading the parity generation process from the host processor.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional CPU instructions or AVX instructions are used for parity generation, then parity calculation can be performed, but host CPU and host DRAM bandwidth are significantly consumed

Engineering Contradiction:
Improveparity generation throughputVSAvoidhost CPU and memory controller bandwidth consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent extracts the parity generation function from the host CPU and host DRAM system, creating a separate accelerator engine with its own PCIe interface and memory space. This allows parity calculations to be performed independently without consuming host CPU cycles or host DRAM bandwidth, directly resolving the resource consumption problem while maintaining high throughput capability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a PCIe endpoint as an intermediary device between the host system and the parity generation function. The accelerator engine communicates with the host through PCIe transactions, acting as a mediator that handles parity generation requests without burdening the host CPU or memory controller, thus solving the bandwidth consumption issue

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional HW RAID architecture is used, then parity generation can be performed, but it becomes a performance bottleneck when scaling with high-performance NVMe drives

Engineering Contradiction:
ImproveRAID scalability with NVMe drivesVSAvoidhost processor and memory controller involvement
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the RAID system into distinct functional components: the host system for I/O management and the accelerator engine for parity generation. This segmentation allows the host to focus on high-level RAID operations while the accelerator handles computation-intensive parity calculations, enabling better scalability with NVMe drives without increasing host processor involvement

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces the traditional mechanical/CPU-based parity generation mechanism with a dedicated hardware accelerator engine. This substitution eliminates the performance bottleneck by using specialized hardware optimized for parity calculations, allowing the system to scale effectively with high-performance NVMe drives without being constrained by host processor capabilities

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12511195B2Non-volatile memory express transport protocol messaging for parity generation
Publication Date: 2025.12.30 MICROCHIP TECHNOLOGY INC
  • US12511195B2 patent drawing
  • US12511195B2 patent drawing
  • US12511195B2 patent drawing

AI summary

Methods, devices and systems to communicate an instruction to generate parities from a host processor to an accelerator engine via a nonvolatile memory transport protocol, generate parities from source data via the accelerator engine based on the instruction, and store the generated parities. Methods, devices and systems to generate parities via an accelerator engine on source data from a central processing unit using non-volatile memory transport protocol to communicate a parity generation instruction.