An NVMeoF Hardware Offloading Device and Method
By implementing NVMeoF hardware offloading in FPGAs, data is directly stored into NVMe SSD, solving the bottleneck of block I/O layer and CPU processing performance in NVMeoF remote data access, achieving more efficient access performance and reducing latency.
Patent Information
- Application Number
- CN202510329574.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In the prior art, the block I/O layer performance and CPU processing performance that provide remote data access services to multiple computing boards have become bottlenecks for NVMeoF remote data access.
The hardware offload of RNIC and NVMeoF is implemented in the FPGA on the target side. Only RDMA drivers, NVMeoF drivers and NVMe drivers are implemented on the CPU. The data is directly stored in the NVMe SSD without sending the CPU DDR to the CPU, thereby unloading the CPU computing power on the storage board.
Improve access performance, reduce access latency, and solve the problem that block I/O layer performance and CPU processing performance become bottlenecks.
Smart Images

Figure CN119854364B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and particularly to a hardware offloading device and method for NVMeoF. Background Art
[0002] In today's big data era, it is necessary to process a vast amount of data in real time, which poses more stringent requirements on the bandwidth and throughput of storage systems. To meet the new requirements for processing massive data, the performance of solid state drives (SSDs) has been continuously improved. However, the theoretical bandwidth of the traditional SATA (Serial Advanced Technology Attachment) 3.0 interface, which is 600MB / s, has gradually become a performance bottleneck. In contrast, the PCIe (Peripheral Component Interconnection Express) system bus features high bandwidth and low latency, making the combination of PCIe and SSD an inevitable trend. The NVMe (Non-volatile Memory Express) protocol proposed for PCIe SSDs has achieved great optimizations in terms of capacity, power consumption, scalability, and error management, and has become the most important storage method for data-intensive applications such as data centers or supercomputing clouds.
[0003] NVMe is designed for the use and transfer of flash memory on local nodes and is targeted at PCIe. Although PCIe, as a high-speed interconnection bus, can fully utilize the many advantages of the PCIe interface when used locally, as an internal bus, PCIe is restricted by the CPU and motherboard, and the number of available PCIe channels is very limited. The CPU also needs to reserve resources for the channels, further exacerbating the difficulty of expansion and management. Moreover, PCIe is an internal bus and lacks the ability to become a general-purpose network, and there is no switch available, making it unsuitable for building large-scale storage networks and long-distance data transmission. Under such limitations, NVMe is only applicable to the architecture where the computing node and the storage node are integrated, and the CPU directly accesses the SSD through the internal system bus. The NVMeoF (NVMe over Fabric) network storage protocol based on the NVMe protocol can be used to communicate with remote NVMe devices between the host and the storage system through various network structures. NVMeoF extends the NVMe protocol over different network protocols, so that in the NVMe storage system with an interconnection structure, it can also have the advantages of high performance, low latency, and low protocol overhead like the NVMe protocol in a single-machine system, ensuring that the storage performance of the NVMe device after being extended by the NVMeoF protocol will not be significantly reduced. For the remote host side, it will occupy fewer CPU resources, reduce the high requirements for the CPU to handle a large amount of I / O, and the saved CPU resources can be used to process storage-related algorithms; for the backend storage system, since it is independent of the computing node, the network node and the storage device are scalable, no longer restricted by the PCIe channels, and more SSDs can be extended and connected, also reducing the requirements for deployment costs.
[0004] In the NVMeoF remote data access architecture of the current data center, the processing of network data packets is completed by the independent network card at the storage target end. The NVMe and NVMeoF protocol processing are respectively completed by the NVMe protocol processing module and the NVMeoF protocol processing module in the operating system running on the CPU. The NVMe memory card, network card, and CPU are interconnected via a PCIe bridge chip. In I / O-intensive applications, the processing performance of the CPU operating system block I / O layer will become the performance bottleneck for business applications to access peripherals. Especially in an asymmetric, many-to-one storage structure, the performance of the block I / O layer of the storage board for providing remote data access services to multiple computing boards is prone to becoming the bottleneck of remote data access; on the other hand, since the operating system internally schedules each thread according to time slices, in long-term frequent remote data access, the processing threads of the block I / O layer will be periodically scheduled by the operating system to a non-executing state. At this time, a large number of remote data accesses will wait for system scheduling, greatly increasing the latency and processing uncertainty of remote access. The processing performance of the CPU becomes the bottleneck of NVMeoF remote data access. Summary of the Invention
[0005] The purpose of the present invention is to provide a hardware offloading device for NVMeoF. The device has a simple structure, is safe, effective, reliable, and easy to operate. It can implement the hardware offloading of RNIC and NVMeoF in the FPGA at the target end. Only the RDMA driver, NVMeoF driver, and NVMe driver are implemented on the CPU. Data is directly stored into the NVMe SSD via the FPGA without being sent to the DDR of the CPU, which improves the access performance while reducing the access latency, offloads the computing power of the storage board CPU, and solves the problem that the performance of the block I / O layer for providing remote data access services to multiple computing boards and the processing performance of the CPU become the bottleneck of NVMeoF remote data access.
[0006] Based on the above purpose, the technical solution provided by the present invention is as follows:
[0007] A hardware offloading device for NVMeoF, comprising: a hardware offloading engine, a cache module, a network interface control module, and an NVMe storage module;
[0008] The cache module is respectively connected to the hardware offloading engine and the network interface control module;
[0009] The network interface control module is connected to the remote host end;
[0010] The hardware offloading engine is connected to the NVMe storage module through a PCIe interface;
[0011] The hardware offloading engine is used to perform encapsulation, format checking, address translation, reverse address mapping, direct access, and management of data transmission operations on NVMeoF commands at the remote host end in sequence;
[0012] The cache module is used to cache RDMA read messages and RDMA write messages of the network interface control module, NVMe commands after address translation and reverse address mapping, read operation data, and write operation data;
[0013] The network interface control module is used to extract and process virtual descriptors, establish descriptor extraction tasks, cache root endpoint mode requests, construct request packets and response packets, verify request packets and convert their formats, and verify and cache response packets;
[0014] The network interface control module is also used for encapsulation and construction of ROCE packets, submission of completion messages, control of transmission rate, configuration of transmission parameters, and network communication with the remote host end.
[0015] Preferably, the hardware offloading engine includes a VP queue processing unit, an NVMeoF command analysis and mapping unit, an arbitration processing unit, an NVMeoF command processing unit, a data transmission unit, and a PCIe EP unit;
[0016] The VP queue processing unit is used to interact with the network interface control module, receive NVMeoF command encapsulation from the remote host end, send the NVMeoF command to the submission queue of NVMeoF, and after the access ends, send the corresponding response encapsulation to the remote host end;
[0017] The NVMeoF command analysis and mapping unit is used to extract NVMeoF commands from the submission queue of NVMeoF, check whether the NVMeoF commands conform to a preset format, and if so, perform address translation processing to form the NVMe commands and store them in the submission queue of the NVMe storage module;
[0018] The NVMeoF command analysis and mapping unit is also used to perform reverse address mapping on the submission queue of the NVMe storage module to obtain the NVMe commands and store them in the submission queue of NVMeoF;
[0019] The arbitration processing unit is used to process the NVMe commands and then send them to the queue of NVMeoF;
[0020] The NVMeoF command processing unit is used to process the corresponding queue of NVMeoF and interact with the NVMe storage module;
[0021] The data transmission processing unit is used to manage the data transmission between the hardware offloading engine and the cache module.
[0022] Preferably, the cache module includes: a VP cache unit, an NVMe storage module queue cache unit, an NVMeoF queue cache unit, and a read / write data cache unit;
[0023] The VP cache unit is used to cache the RDMA read message and the RDMA write message in the network interface control module;
[0024] The NVMe storage module queue cache unit is used to cache the NVMe commands after address conversion by the NVMeoF command analysis and mapping unit;
[0025] The NVMeoF queue cache unit is used to cache the NVMe commands after reverse address mapping by the NVMeoF command analysis and mapping unit;
[0026] The read / write data cache unit is used to cache the read operation data and the write operation data generated during the RDMA read operation and the write operation in the network interface control module.
[0027] Preferably, the network interface control module includes: a descriptor communication processing unit, a reliable transmission control unit, a message construction unit, a message processing unit, and an address conversion unit;
[0028] The descriptor communication processing module is used to extract the request descriptor and the response descriptor from the VP queue processing unit, and after scheduling, caching, parsing, and verification in sequence, establish a descriptor extraction task triggered by the retransmission descriptor;
[0029] The reliable transmission control unit is used to cache the root endpoint mode request before the root endpoint mode confirmation is completed;
[0030] The message construction unit is used to construct a request message and / or a response message according to the request descriptor and / or the response descriptor;
[0031] The message processing unit is used to verify the request message and screen the request message;
[0032] The message processing unit is also used to check the response message, determine whether retransmission is required, and transmit the response message that needs to be retransmitted to the read / write data cache unit;
[0033] The address conversion unit is used to convert the reaction chamber component of NVMeoF into the RDMA VP format.
[0034] Preferably, the network interface control module further includes: a message construction and parsing unit, a CQE processing unit, a congestion control unit, a configuration processing unit, and a network interface;
[0035] The message construction and parsing unit is configured to perform a legality check on the Ethernet input message based on RDMA, and encapsulate and construct the Ethernet output message based on RDMA;
[0036] The CQE processing unit is configured to submit a completion message for the check or encapsulation and construction;
[0037] The congestion control unit is configured to control the transmission rate of the network communication at the remote host end;
[0038] The configuration processing unit is configured to set and store the transmission parameter configuration between the CPU at the remote host end and the network interface control module;
[0039] The network interface is configured to perform network communication with the remote host end.
[0040] An NVMeoF hardware offloading method is implemented based on the NVMeoF hardware offloading device described in any one of the above. The NVMeoF hardware offloading method includes a write operation of NVMeoF based on RDMA, and specifically includes the following steps:
[0041] The remote host end encapsulates the NVMeoF command, encapsulates it into an RDMA message and sends it to the network interface control module;
[0042] The network interface control module completes the offloading of the RDMA protocol, and strips the NVMeoF command encapsulation from the RDMA message;
[0043] Write the NVMeoF command encapsulation into the submission queue of the cache module;
[0044] The hardware offloading engine retrieves the NVMeoF command encapsulation in the submission queue of the cache module and completes the protocol offloading work of command parsing and address conversion.
[0045] Preferably, after completing the protocol offloading work of command parsing and address conversion, the following steps are further included:
[0046] The hardware offloading engine allocates data space for the read / write data cache unit in the cache module, and writes the RDMA read message into the VP queue processing unit;
[0047] The remote host reads receipts from the memory according to the RDMA read request initiated by the network interface control module, writes the read data into the data space of the read / write data cache unit allocated in the cache module, and notifies the hardware offloading engine;
[0048] The NVMe storage module retrieves the data to be written from the data space of the read / write data cache unit allocated in the cache module according to the doorbell signal sent by the hardware offloading engine;
[0049] After the data transfer is completed, the NVMe storage module writes the completion command into the queue cache unit of NVMeoF in the cache module. After being mapped and converted by the hardware offloading engine, it is encapsulated into the response capsule component of NVMeoF and written into the completion queue of NVMeoF in the cache module;
[0050] The network interface control module converts the response capsule component of NVMeoF in the NVMeoF completion queue into the RDMAVP format;
[0051] The network interface control module sends the completion information in the VP queue processing unit to the remote host;
[0052] Preferably, the hardware offloading method of NVMeoF includes the read operation of NVMeoF based on RDMA, specifically the following steps:
[0053] The remote host encapsulates the NVMeoF command, encapsulates it into an RDMA message and sends it to the network interface control module;
[0054] The network interface control module completes the offloading of the RDMA protocol and strips the NVMeoF command encapsulation from the RDMA message;
[0055] Write the NVMeoF command encapsulation into the submission queue of the cache module;
[0056] The hardware offloading engine takes out the NVMeoF command encapsulation in the submission queue of the cache module and completes the protocol offloading work of command parsing and address conversion;
[0057] The NVMe storage module takes out the NVMe command from the submission queue of the NVMe storage module in the cache module according to the doorbell signal sent by the hardware offloading engine. After execution, it writes the data into the read / write data cache unit in the cache module;
[0058] After the data transmission is completed, the NVMe storage module writes the completed command into the queue cache unit of NVMeoF in the cache module. After being mapped and converted by the hardware offloading engine, it is encapsulated into a reaction chamber component of NVMeoF and written into the completion queue of NVMeoF in the cache module.
[0059] The network interface control module initiates an RDMA write request to write the data in the read / write data cache unit in the cache module into the remote host memory.
[0060] The network interface control module converts the reaction chamber component of NVMeoF into the RDMA VP format.
[0061] The network interface control module sends the completion information in the VP queue processing unit to the remote host.
[0062] The hardware offloading device of NVMeoF provided by the present invention is provided with a hardware offloading engine, a cache module, a network interface control module and an NVMe storage module. The remote host sends an NVMeoF command, and the hardware offloading engine performs operations such as encapsulation, format checking, address conversion, reverse address mapping, direct access and management of data transmission on the command; the cache module caches the RMDA read message and write message of the network interface control module, the NVMe command after address conversion and reverse address mapping, and the read operation and write operation data; through the network interface control module, virtual descriptors are extracted and processed, a descriptor extraction task is established, a cache root node mode request is cached, and request messages and response messages are constructed and verified; through the network interface control module, the ROCE message is encapsulated and transformed, a completion message is submitted, the transmission rate is controlled, transmission parameters are configured and communication with the remote host is performed.
[0063] Compared with the prior art, the present invention realizes the hardware offloading of RNIC and NVMeoF in the FPGA at the target end. Only the RDMA driver, NVMeoF driver and NVMe driver are implemented on the CPU. The data is directly stored in the NVMe SSD via the FPGA without being sent to the DDR of the CPU, which improves the access performance while reducing the access latency, offloads the computing power of the storage board CPU, and solves the problem that the performance of the block I / O layer for providing remote data access services to multiple computing boards and the processing performance of the CPU become bottlenecks in NVMeoF remote data access.
[0064] The present invention also provides a hardware offloading method for NVMeoF. Since it belongs to the same inventive concept as the system and solves the same technical problems, it should have the same beneficial effects and will not be elaborated here. Brief Description of the Drawings
[0065] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0066] Figure 1 Schematic diagram of a hardware offloading device for NVMeoF provided by an embodiment of the present invention;
[0067] Figure 2 Schematic diagram of the hardware offloading engine provided by an embodiment of the present invention;
[0068] Figure 3 Schematic diagram of the cache module provided by an embodiment of the present invention;
[0069] Figure 4 Schematic diagram of the network interface control module provided by an embodiment of the present invention;
[0070] Figure 5 Flowchart of the write operation of NVMeoF based on RDMA in a hardware offloading method for NVMeoF provided by an embodiment of the present invention;
[0071] Figure 6 Flowchart after step S4 provided by an embodiment of the present invention;
[0072] Figure 7 Flowchart of the read operation of NVMeoF based on RDMA in a hardware offloading method for NVMeoF provided by an embodiment of the present invention. Detailed implementation manners
[0073] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0074] The embodiments of the present invention are written in a progressive manner.
[0075] The embodiments of the present invention provide a hardware offloading device and method for NVMeoF. It mainly solves the technical problem that in the prior art, the performance of the block I / O layer for providing remote data access services to multiple computing boards and the processing performance of the CPU have become bottlenecks for NVMeoF remote data access.
[0076] Such as Figure 1As shown in the figure, an NVMeoF hardware offloading device includes: a hardware offloading engine, a cache module, a network interface control module, and an NVMe storage module;
[0077] The cache module is respectively connected to the hardware offloading engine and the network interface control module;
[0078] The network interface control module is connected to the remote host side;
[0079] The hardware offloading engine is connected to the NVMe storage module through a PCIe interface;
[0080] The hardware offloading engine is used to sequentially perform encapsulation, format checking, address conversion, reverse address mapping, direct access, and management of data transfer operations on NVMeoF commands from the remote host side;
[0081] The cache module is used to cache RDMA read messages and RDMA write messages of the network interface control module, NVMe commands after address conversion and reverse address mapping, read operation data, and write operation data;
[0082] The network interface control module is used to extract and process virtual descriptors, establish descriptor extraction tasks, cache root endpoint mode requests, construct request packets and response packets, verify and convert the format of request packets, and verify and cache response packets;
[0083] The network interface control module is also used for encapsulation and construction of ROCE packets, submission of completion messages, control of transmission rate, configuration of transmission parameters, and network communication with the remote host side.
[0084] In the actual application process, an NVMeoF hardware offloading engine, a cache module Buffer, a network interface control module RNIC, and an NVMe storage module SSD are provided in the NVMeoF hardware offloading device; among them, RNIC is connected to the remote host side through a network interface, and the NVMeoF hardware offloading engine is connected to the internal PCIe interface through a PCIe EP for data transmission with the NVMe SSD storage module.
[0085] It should be noted that NVMe-oF (NVMe over Fabrics) is non-volatile memory based on an architecture. NVMe-oF supports data transmission between a host and a solid-state storage device or system through a network. NVMe-oF inherits all the advantages of NVMe, including a lightweight and efficient command set, multi-core awareness, and protocol parallelism. NVMe-oF is truly network-independent because it supports all common fabrics, including Fibre Channel, InfiniBand, and Ethernet;
[0086] RDMA bypasses the kernel and directly accesses the RDMA enabled NIC from the user space through a proprietary RDMA network interface card RNIC (RDMA-aware Network Interface Controller). RDMA provides a proprietary Verbs Interface instead of the traditional TCP / IP Socket Interface. To use RDMA, first, a data path from RDMA to the application memory needs to be established. These data paths can be established through the proprietary Verbs Interace interface of RDMA. Once the data paths are established, the user space buffer can be directly accessed.
[0087] As Figure 2 shown, preferably, the hardware offloading engine includes a VP queue processing unit, an NVMeoF command analysis and mapping unit, an arbitration processing unit, an NVMeoF command processing unit, a data transfer unit, and a PCIe EP unit;
[0088] The VP queue processing unit is used to interact with the network interface control module, receive the NVMeoF command encapsulation from the remote host side, send the NVMeoF command to the submission queue of NVMeoF, and after the access ends, send the corresponding response encapsulation to the remote host side;
[0089] The NVMeoF command analysis and mapping unit is used to extract the NVMeoF command from the submission queue of NVMeoF, check whether the NVMeoF command conforms to the preset format. If it conforms, it performs address conversion processing to form an NVMe command and stores it in the submission queue of the NVMe storage module;
[0090] The NVMeoF command analysis and mapping unit is also used to perform reverse address mapping on the submission queue of the NVMe storage module to obtain the NVMe command and store it in the submission queue of NVMeoF;
[0091] The arbitration processing unit is used to process the NVMe command and then send it to the queue of NVMeoF;
[0092] The NVMeoF command processing unit is used to process the corresponding queue of NVMeoF and interact with the NVMe storage module;
[0093] The data transfer processing unit is used to manage the data transfer between the hardware offloading engine and the cache module.
[0094] In the actual operation process, a virtual port VP queue processing unit, an NVMeoF command analysis and mapping unit, an arbitration processing unit, and a data transmission processing unit are provided in the hardware offloading engine; the VP queue processing unit is used to interact with the network interface control module RNIC, receive the NVMeoF command encapsulation from the remote host side, submit the NVMeoF command to the submission queue SQ of NVMeoF, and after the data access ends, encapsulate and send the corresponding response to the remote host side; the NVMeoF command analysis and mapping unit retrieves the NVMeoF command from the submission queue SQ of NVMeoF, checks whether the command meets the preset format requirements, and if not, feeds back an error; if it meets the requirements, the NVMeoF command analysis and mapping unit performs address conversion processing on the correct NVMeoF command to form an NVMe command; after being arbitrated by the arbitration processing component, the NVMe command is stored in the submission queue SQ of NVMeoF; through the NVMeoF command analysis and mapping unit, reverse address mapping is performed on the submission queue of the NVMe storage module to obtain the NVMe command, which is stored in the SQ / CQ of NVMeoF; the NVMe command processing unit is responsible for interacting with the SSD controller of the NVMe storage module, processing the corresponding SQ / CQ of NVMe, and completing the P2P direct access to the NVMe storage module SSD. The data transmission processing component is responsible for managing the data transmission between the NVMe hardware acceleration engine and the cache module Buffer.
[0095] As Figure 3 shown, preferably, the cache module includes: a VP cache unit, an NVMe storage module queue cache unit, an NVMeoF queue cache unit, and a read / write data cache unit;
[0096] The VP cache unit is used to cache the RDMA read message and the RDMA write message in the network interface control module;
[0097] The NVMe storage module queue cache unit is used to cache the NVMe command after address conversion by the NVMeoF command analysis and mapping unit;
[0098] The NVMeoF queue cache unit is used to cache the NVMe command after reverse address mapping by the NVMeoF command analysis and mapping unit;
[0099] The read / write data cache unit is used to cache the read operation data and the write operation data generated during the RDMA execution of the read operation and the write operation in the network interface control module.
[0100] During the actual operation process, the cache module Buffer is provided with a VP cache unit VP Buffer, an NVMe storage module queue cache unit SSD SQ / CQ, an NVMeoF queue cache unit NVMeoF SQ / CQ Buffer, and a read / write data cache unit; the RMDA read message and the RMDA write message in the RNIC module are cached through the VP Buffer; the NVMe command after reverse address mapping by the NVMeoF command analysis and mapping unit is cached through the NVMeoF SQ / CQ Buffer; the NVMe command after address conversion by the NVMeoF command analysis and mapping unit is cached through the NVMe storage module queue cache unit SSD SQ / CQ; the read / write data cache unit specifically includes a read data sub-cache unit and a write data sub-cache unit, and the read operation data generated during the execution of the read operation is cached through the read data cache unit, and the write operation data generated during the execution of the write operation is cached through the write data sub-cache unit.
[0101] As Figure 4 shown, preferably, the network interface control module includes: a descriptor communication processing unit, a reliable transmission control unit, a message construction unit, a message processing unit, and an address conversion unit;
[0102] The descriptor communication processing module is used to extract the request descriptor and the response descriptor from the VP queue processing unit, and after scheduling, caching, parsing, and verification in sequence, establish a descriptor extraction task triggered by the retransmission descriptor;
[0103] The reliable transmission control unit is used to cache the root endpoint mode request before the root endpoint mode confirmation is completed;
[0104] The message construction unit is used to construct a request message and / or a response message according to the request descriptor and / or the response descriptor;
[0105] The message processing unit is used to verify the request message and screen the request message;
[0106] The message processing unit is also used to check the response message, determine whether retransmission is required, and transmit the response message that needs to be retransmitted to the read / write data cache unit;
[0107] The address conversion unit is used to convert the reaction chamber component of NVMeoF into the RDMA VP format.
[0108] During the actual operation process, a descriptor communication processing unit, a reliable transmission control unit, a message construction unit, a message processing unit, and an address conversion unit are provided in the network interface control module RNIC. The descriptor communication processing unit extracts corresponding request descriptors and response descriptors from the VP buffer unit VP Buffer, schedules and caches them in sequence, performs preliminary parsing and verification, and after completion, recommends a descriptor extraction task triggered by a retransmission descriptor; the reliable transmission control unit caches root endpoint mode requests before the root endpoint RC mode confirmation is completed; the message construction unit constructs a request message or a response message according to the request descriptor / response descriptor; the message processing unit verifies the request message to filter out duplicate messages or invalid messages and leaves normal messages; the address conversion unit converts the NVMeoF reaction chamber component of the normal request message into the RDMA VP format; the message processing unit checks the response message to determine whether retransmission is required and transmits the response message that needs to be retransmitted to the read / write data cache unit.
[0109] As Figure 4 shown, preferably, the network interface control module further includes: a message construction and parsing unit, a CQE processing unit, a congestion control unit, a configuration processing unit, and a network interface;
[0110] The message construction and parsing unit is used to perform a legality check on the RDMA-based Ethernet input message and encapsulate and construct the RDMA-based Ethernet output message;
[0111] The CQE processing unit is used to submit a completion message for inspection or encapsulation and construction;
[0112] The congestion control unit is used to control the transmission rate of network communication on the remote host side;
[0113] The configuration processing unit is used to set and store the transmission parameter configuration between the CPU on the remote host side and the network interface control module;
[0114] The network interface is used to perform network communication with the remote host side.
[0115] During the actual operation process, a message construction and parsing unit, a CQE processing unit, a congestion control unit, a configuration processing unit, and a network interface are also provided in the network interface control module RNIC; the message construction and parsing unit performs a legality check on the received ROCE message and encapsulates and constructs the sent ROCE message; the CQE processing unit submits a completion message; the congestion control component controls the transmission rate of network communication on the remote host side; the configuration processing unit sets and stores the transmission parameter configuration between the CPU on the remote host side and the network interface control module; the network interface performs network communication with the remote host side.
[0116] It should be noted that both ROCE and InfiniBand are network protocol stacks defined by the InfiniBand Trade Association (IBTA). Among them, Infiniband is a high-performance network specifically designed for RDMA, which ensures the reliability of data transmission at the hardware level. To further leverage the advantages of RDMA, the IBTA defined ROCE in 2010. ROCE is the integration of Infiniband and Ethernet technology. While maintaining the core advantages of Infiniband, it achieves compatibility with existing Ethernet infrastructure. Specifically, ROCE is different from Infiniband at the link layer * and network layer, but at the transport layer and RDMA protocol aspects, ROCE inherits the essence of Infiniband.
[0117] CQE (Capacity Quality Estimator): In some network architectures, especially for more refined control of network resource allocation and performance optimization, dedicated CQE algorithms or tools may be used. These tools can estimate the capacity and quality of the network based on multiple parameters (such as SINR, interference level, user density, etc.).
[0118] As Figure 5 shown, a hardware offloading method for NVMeoF is implemented based on any one of the above-mentioned hardware offloading devices for NVMeoF. The hardware offloading method for NVMeoF includes the write operation of NVMeoF based on RDMA, specifically the following steps:
[0119] S1. The remote host encapsulates the NVMeoF command, encapsulates it into an RDMA message and sends it to the network interface control module;
[0120] S2. The network interface control module completes the offloading of the RDMA protocol and strips the NVMeoF command encapsulation from the RDMA message;
[0121] S3. Write the NVMeoF command encapsulation into the submission queue of the cache module;
[0122] S4. The hardware offloading engine retrieves the NVMeoF command encapsulation in the submission queue of the cache module and completes the protocol offloading work of command parsing and address translation.
[0123] In steps S1 to S4, the hardware offloading method of NVMeoF includes a write operation of NVMeoF based on RDMA. Specifically: The remote host encapsulates the command capsule of NVMeoF into an RDMA message and sends it to the remote RNIC in the way of RDMA Send; The network interface control module completes the offloading of the RDMA protocol, strips the command capsule of NVMeoF from the RDMA message, and writes the stripped command capsule of NVMeoF into the submission queue of the cache module; The hardware offloading engine takes out the command capsule of NVMeoF in the submission queue of this cache module to complete the protocol offloading work of command parsing and address conversion.
[0124] As Figure 6 shown, preferably, after step S4, the following steps are further included:
[0125] A1. The hardware offloading engine allocates the data space of the read-write data cache unit in the cache module and writes the RDMA read message into the VP queue processing unit;
[0126] A2. According to the RDMA read request initiated by the network interface control module, the remote host reads the receipt from the memory, writes the read data into the data space of the read-write data cache unit allocated in the cache module, and notifies the hardware offloading engine;
[0127] A3. The NVMe storage module retrieves the data to be written from the data space of the read-write data cache unit allocated in the cache module according to the doorbell signal sent by the hardware offloading engine;
[0128] A4. After the data transmission is completed, the NVMe storage module writes the completion command into the NVMeoF queue cache unit in the cache module. After being mapped and converted by the hardware offloading engine, it is encapsulated into an NVMeoF response capsule component and written into the NVMeoF completion queue in the cache module;
[0129] A5. The network interface control module converts the NVMeoF response capsule component in the NVMeoF completion queue into the RDMA VP format;
[0130] A6. The network interface control module sends the completion information in the VP queue processing unit to the remote host side.
[0131] In the actual operation process, after completing the protocol offloading work of command parsing and address conversion, the hardware offloading engine also allocates the write data sub-buffer unit RX Buffer data space in the read / write data buffer unit in the cache module Buffer, and writes the RDMA Read message into the RDMA VP; the RNIC initiates an RDMA Read to the Host, reads data from the Host memory; the Host returns the data, and RDMA writes the received read data into the RX Buffer data space allocated by the Buffer, and notifies the NVMeoF hardware offloading engine; the NVMeoF hardware offloading engine sends a Doorbell to the NVMe SSD to notify the device that there is a command waiting to be processed; the NVMe SSD retrieves the data to be written from the RX Buffer in the Buffer; after the data transmission is completed, the NVMe SSD writes the completed command into the SSD CQ of the Buffer, which is mapped and converted by the NVMeoF hardware offloading engine and encapsulated as an NVMeoF response capsule and written into the NVMeoF CQ of the Buffer; the RNIC converts the response capsule of the NVMeoF CQ into the format of the RDMA VP; the RNIC uses RDMA Send to send the completion information in the RDMA VP to the remote host side Host, indicating the end of the current write data process.
[0132] As Figure 7 shown, preferably, the hardware offloading method of NVMeoF includes the read operation of NVMeoF based on RDMA, specifically the following steps:
[0133] B1. The remote host side encapsulates the NVMeoF command, encapsulates it as an RDMA message and sends it to the network interface control module;
[0134] B2. The network interface control module completes the offloading of the RDMA protocol, and strips the NVMeoF command encapsulation from the RDMA message;
[0135] B3. Write the NVMeoF command encapsulation into the submission queue of the cache module;
[0136] B4. The hardware offloading engine takes out the NVMeoF command encapsulation in the submission queue of the cache module and completes the protocol offloading work of command parsing and address conversion;
[0137] B5. The NVMe storage module takes out the NVMe command from the submission queue of the NVMe storage module of the cache module according to the Doorbell signal sent by the hardware offloading engine. After execution, the data is written into the read / write data buffer unit in the cache module;
[0138] B6. After the data transmission is completed, the NVMe storage module writes the completed command into the queue cache unit of NVMeoF in the cache module. After being mapped and converted by the hardware offloading engine, it is encapsulated into the response capsule component of NVMeoF and written into the completion queue of NVMeoF in the cache module.
[0139] B7. The network interface control module initiates an RDMA write request to write the data in the read / write data cache unit in the cache module into the remote host-side memory.
[0140] B8. The network interface control module converts the response capsule component of NVMeoF into the RDMA VP format.
[0141] B9. The network interface control module sends the completion information in the VP queue processing unit to the remote host side.
[0142] In the actual application process, the NVMeoF hardware offloading method includes the read operation of NVMeoF based on RDMA. Specifically: the remote host side encapsulates the NVMeoF command into a command capsule and sends it to the remote RNIC in the form of an RDMA Send; the RNIC completes the offloading of the RDMA protocol, strips the command capsule from the RDMA message; writes the command capsule into the submission queue NVMeoF SQ in the storage cache Buffer; the NVMeoF hardware offloading engine takes out the command capsule from the submission queue NVMeoF SQ in the Buffer and completes protocol offloading work such as command parsing and address conversion; the NVMeoF hardware offloading engine sends a Doorbell signal to the NVMe SSD to notify the device that there is a command waiting to be processed; the NVMe SSD takes out the NVMe command from the SSD SQ in the Buffer, and after execution, writes the data into the read data sub-cache unit TX Buffer in the read / write data cache unit of the Buffer; after the data transmission is completed, the NVMe SSD writes the completed command into the SSD CQ in the Buffer; the NVMeoF hardware offloading engine maps and converts the completed command in the SSD CQ, encapsulates it into the response capsule of NVMeoF, and writes it into the CQ of NVMeoF in the Buffer; the RNIC initiates an RDMA Write to write the data in the TX Buffer of the Buffer into the Host-side memory; the RNIC converts the response capsule of NVMeoF into the RDMA VP format; the RNIC sends the completion information in the RDMA VP to the remote host side Host in the form of an RDMA Send, indicating the end of this read data process.
[0143] In the embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0144] In addition, in the embodiments of the present invention, each functional module can be all integrated in a processor, or each module can be separately used as a device alone, or two or more modules can be integrated in a device; each functional module in the embodiments of the present invention can be implemented in the form of hardware, or can be implemented in the form of a combination of hardware and software functional units.
[0145] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed through program instructions and related hardware. The foregoing program instructions can be stored in a computer-readable storage medium. When the program instructions are executed, the steps including the above method embodiments are executed; and the foregoing storage media include: various media such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs that can store program codes.
[0146] It should be understood that in the present application, if the terms "system", "device", "unit" and / or "module" are used, they are only a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the term can be replaced by other expressions.
[0147] As shown in the present application and the claims, unless the context clearly indicates an exception, the words "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "including" and "comprising" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements. The element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0148] If a flowchart is used in this application, the flowchart is used to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the operations before or after may not necessarily be executed precisely in order. On the contrary, the steps may be processed in reverse order or simultaneously. At the same time, other operations may also be added to these processes, or one or more steps may be removed from these processes.
[0149] The above has introduced in detail a hardware offloading device and method for NVMeoF provided by the present invention. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A hardware unloading device of NVMeoF, which realizes the hardware unloading of RNIC and NVMeoF in the FPGA of the target end, characterized in that: include: Hardware offload engine, cache module, network interface control module and NVMe storage module; The cache module is connected to the hardware offload engine and the network interface control module respectively; The network interface control module is connected to the remote host end; The hardware offload engine is connected to the NVMe storage module via a PCIe interface; The hardware offload engine is used to sequentially perform encapsulation, format checking, address translation, reverse address mapping, direct access and management data transmission operations on remote host-side NVMeoF commands; The cache module is used to cache the RDMA read message and RDMA write message of the network interface control module, the NVMe command after address conversion and reverse address mapping, the read operation data and the write operation data; The network interface control module is used to extract and process virtual descriptors, establish descriptor extraction tasks, cache root endpoint mode requests, construct request messages and response messages, verify request messages and convert formats, and verify and cache response messages; The network interface control module is also used to encapsulate and construct ROCE messages, submit completion messages, control transmission rates, configure transmission parameters, and perform network communications with remote host terminals; The hardware offload engine includes a VP queue processing unit, an NVMeoF command analysis mapping unit, an arbitration processing unit, an NVMeoF command processing unit, a data transmission unit and a PCIe EP unit; The VP queue processing unit is used to interact with the network interface control module, receive the NVMeoF command package of the remote host, send the NVMeoF command to the NVMeoF submission queue, and send the corresponding response package to the remote host after the access is completed; The NVMeoF command analysis and mapping unit is used to extract the NVMeoF command from the NVMeoF submission queue, check whether the NVMeoF command conforms to a preset format, and if so, perform address conversion processing to form the NVMe command, and store it in the submission queue of the NVMe storage module; The NVMeoF command analysis and mapping unit is further used to perform reverse address mapping on the submission queue of the NVMe storage module to obtain the NVMe command and store it in the submission queue of the NVMeoF; The arbitration processing unit is used to process the NVMe command and then send it to the NVMeoF queue; The NVMeoF command processing unit is used to process the corresponding NVMeoF queue and interact with the NVMe storage module; The data transmission processing unit is used to manage the data transmission between the hardware offload engine and the cache module.
2. The NVMeoF hardware unloading device according to claim 1 implements the hardware unloading of RNIC and NVMeoF in the FPGA of the target end, characterized in that: The cache module includes: a VP cache unit, an NVMe storage module queue cache unit, an NVMeoF queue cache unit and a read-write data cache unit; The VP cache unit is used to cache the RDMA read message and the RDMA write message in the network interface control module; The NVMe storage module queue cache unit is used to cache the NVMe commands after address conversion by the NVMeoF command analysis mapping unit; The NVMeoF queue cache unit is used to cache the NVMe commands after the NVMeoF command analysis mapping unit undergoes reverse address mapping; The read-write data cache unit is used to cache the read operation data and the write operation data generated when the RDMA in the network interface control module performs the read operation and the write operation.
3. The NVMeoF hardware unloading device as claimed in claim 2 implements the hardware unloading of RNIC and NVMeoF in the FPGA of the target end, characterized in that: The network interface control module includes: a descriptor communication processing unit, a reliable transmission control unit, a message construction unit, a message processing unit and an address conversion unit; The descriptor communication processing module is used to extract the request descriptor and the response descriptor from the VP queue processing unit, and to establish a descriptor extraction task based on the retransmission descriptor trigger after scheduling, caching, parsing and checking in sequence; The reliable transmission control unit is used to cache the root endpoint mode request before the root endpoint mode confirmation is completed; The message construction unit is used to construct a request message and / or a response message according to the request descriptor and / or the response descriptor; The message processing unit is used to verify the request message and filter the request message; The message processing unit is further used to check the response message, determine whether it needs to be retransmitted, and transmit the response message that needs to be retransmitted to the read-write data cache unit; The address conversion unit is used to convert the NVMeoF reaction cabin component into the RDMA VP format.
4. The NVMeoF hardware unloading device as claimed in claim 3 implements the hardware unloading of RNIC and NVMeoF in the FPGA of the target end, characterized in that: The network interface control module further includes: a message construction and parsing unit, a CQE processing unit, a congestion control unit, a configuration processing unit and a network interface; The message construction and parsing unit is used to perform a legality check on the RDMA-based Ethernet input message and to encapsulate and construct the RDMA-based Ethernet output message; The CQE processing unit is used to submit the completion message of the inspection or packaging and construction; The congestion control unit is used to control the transmission rate of the network communication of the remote host end; The configuration processing unit is used to set and store the transmission parameter configuration between the CPU of the remote host and the network interface control module; The network interface is used for network communication with the remote host.
5. A method for hardware unloading of NVMeoF, which implements hardware unloading of RNIC and NVMeoF in the FPGA of the target end, and is implemented based on the hardware unloading device of NVMeoF described in any one of claims 1 to 4, characterized in that: The NVMeoF hardware offloading method includes a write operation of NVMeoF based on RDMA, which is specifically the following steps: The remote host encapsulates the NVMeoF command into an RDMA message and sends it to the network interface control module; The network interface control module completes the unloading of the RDMA protocol and strips the NVMeoF command encapsulation from the RDMA message; Writing the NVMeoF command encapsulation into a submission queue of a cache module; The hardware offload engine takes out the NVMeoF command encapsulation in the submission queue of the cache module, and completes the protocol offload work of command parsing and address conversion.
6. The NVMeoF hardware unloading method as claimed in claim 5, wherein the hardware unloading of RNIC and NVMeoF is implemented in the FPGA of the target end, wherein: After completing the protocol unloading work of command parsing and address translation, the following steps are also included: The hardware offload engine allocates data space of the read and write data cache unit in the cache module, and writes the RDMA read message into the VP queue processing unit; The remote host reads the data from the memory according to the RDMA read request initiated by the network interface control module, writes the read data into the data space of the read-write data cache unit allocated in the cache module, and notifies the hardware offload engine; The NVMe storage module retrieves the data to be written from the data space of the read-write data cache unit allocated in the cache module according to the doorbell signal sent by the hardware offload engine; After the data transmission is completed, the NVMe storage module writes the completion command into the NVMeoF queue cache unit in the cache module, which is mapped and converted by the hardware offload engine, encapsulated as an NVMeoF reaction cabin component, and written into the NVMeoF completion queue in the cache module; The network interface control module converts the NVMeoF reaction cabin component in the NVMeoF completion queue into the RDMA VP format; The network interface control module sends the completion information in the VP queue processing unit to the remote host end.
7. The NVMeoF hardware unloading method as claimed in claim 5, wherein the hardware unloading of RNIC and NVMeoF is implemented in the FPGA of the target end, wherein: The hardware offloading method of NVMeoF includes a read operation of NVMeoF based on RDMA, which is specifically the following steps: The remote host encapsulates the NVMeoF command into an RDMA message and sends it to the network interface control module; The network interface control module completes the unloading of the RDMA protocol and strips the NVMeoF command encapsulation from the RDMA message; Writing the NVMeoF command encapsulation into a submission queue of a cache module; The hardware offload engine takes out the NVMeoF command encapsulation in the submission queue of the cache module and completes the protocol offload work of command parsing and address conversion; The NVMe storage module takes out the NVMe command from the submission queue of the NVMe storage module of the cache module according to the doorbell signal sent by the hardware offload engine, and writes the data into the read-write data cache unit in the cache module after execution; After the data transmission is completed, the NVMe storage module writes the completion command into the NVMeoF queue cache unit in the cache module, which is mapped and converted by the hardware offload engine, encapsulated as an NVMeoF reaction cabin component, and written into the NVMeoF completion queue in the cache module; The network interface control module initiates an RDMA write request to write the data in the read-write data cache unit in the cache module into the remote host memory; The network interface control module converts the NVMeoF reaction cabin component into an RDMA VP format; The network interface control module sends the completion information in the VP queue processing unit to the remote host end.
Citation Information
Patent Citations
Protocol conversion method, gateway and equipment and readable storage medium
CN112291259A
NVMeoF response instruction transmission method based on RDMA completion event
CN118672953A
Cited By
Dynamic queue binding method and system based on NVMe over Fabrics (NVMeoF)
CN122018808A