RDMA packet aggregation processing method and network card device
By dynamically aggregating packets with the same message through parsing the BTH header field of RDMA packets, and employing a metadata-driven strategy and hardware zero-copy technology, the problem of limited processing efficiency of RDMA network cards is solved, thereby improving bandwidth and processing capabilities.
Patent Information
- Application Number
- CN202510627152.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Existing RDMA network cards fail to design differentiated processing for different types of RDMA packets, resulting in limited processing efficiency and a limitation on the network card's PPS processing capability, thus failing to improve bandwidth.
By parsing fields such as OpCode, Destination QP, and PSN in the BTH header of RDMA packets, consecutive packets of the same message are dynamically aggregated, reducing the number of downstream processing steps. A metadata-driven aggregation strategy is adopted to ensure that the aggregation conforms to the RDMA protocol segmentation rules, and zero-copy aggregation is achieved through smart network interface card hardware.
It increases the bandwidth of the network card for processing RDMA packets, reduces the number of processing steps of downstream modules, reduces the CPU load, and improves processing efficiency, especially in high-throughput scenarios to meet the high PPS requirement.
Smart Images

Figure CN120434317B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data communication, and particularly relates to an RDMA message aggregation processing method and a network card device. BACKGROUND
[0002] In the RDMA protocol and the derived RoCEv1 and RoCEv2 protocols, the operation of RDMA (such as Send, RDMAWrite, RDMA Read Response) is limited by the PMTU size (PMTU supports 256 bytes / 512 bytes / 1024 bytes / 2048 bytes / 4096 bytes) and the MTU size of the actual network (default 1500, at this time the PMTU can only use 1024 bytes), resulting in a large number of small-size messages for a Send Message (maximum 2GB).
[0003] The traditional network card processes each received RDMA message independently, and is limited by the PPS (Packet Per Second) of the processed message, thereby reducing the overall processing bandwidth of the network card (processing bandwidth = processing PPS * message length).
[0004] Existing message aggregation technologies are mostly based on the TCP / IP protocol stack (such as GRO / LRO), but RoCEv2 uses UDP / IP protocol and relies on hardware offloading, so the general aggregation method cannot be used.
[0005] The existing RDMA network card does not design differential processing measurement for different types of RDMA messages, resulting in limited processing efficiency.
[0006] RDMA messages are sent according to messages as a request (Work Request), and a message can support 0 bytes to 2GB. When RDMA is carried on Ethernet (RDMA Over Ethernet), it is limited by the MTU of the limited Ethernet, resulting in a message being split into multiple small messages for sending. The opposite end network card receives the small message and directly processes the message according to the RDMA protocol. When the message is large, one message will be split into multiple small messages, causing the PPS number of the opposite end network card to increase greatly. The network card hardware is limited by its own processing PPS capability, resulting in the inability to improve the bandwidth of the network card.
[0007] The existing network card does not attempt to perform stateless aggregation processing on the RDMA message when processing the RDMA message, and first aggregates the messages of the same type that meet the continuous condition into a large packet, and then sends the aggregated large packet to the subsequent module for final processing according to the RDMA protocol. The aggregation adopts a best-effort strategy. Once the received message does not meet the aggregation condition, the aggregation processing is ended, and the subsequent module continues to process, without affecting the final processing effect, but reducing the number of messages processed by the subsequent module. Especially for the soft core processing used by the subsequent module, such as using a general-purpose CPU for processing, since the CPU frequency is fixed, the processing period consumed by a message is fixed. When the message length is large, the number of messages processed per second is small, and the message bandwidth processed by a single CPU is large.
[0008] For RDMA messages, the network card needs to process each received RDMA message according to the RDMA protocol. If each received message is processed packet by packet, it is easy to cause queue congestion of the subsequent module, increase processing delay and reduce overall bandwidth efficiency, especially in high-throughput scenarios. At this time, the network card requires a very high number of RDMA messages processed per second, and the conventional network card cannot meet such processing capability.
[0009] By aggregating messages, the number of processed messages can be effectively reduced, but existing aggregation techniques are mainly designed for TCP / IP protocol stacks, such as LRO / GRO, which cannot directly adapt to the requirements of the RDMA protocol layer semantics.
[0010] The present application relates to the technical field of computer networks, and specifically relates to a method and device for improving processing bandwidth by analyzing and aggregating specific types of RDMA messages in an intelligent network card supporting the RDMA protocol, suitable for high-performance computing, distributed storage and AI training scenarios.
[0011] Chinese patent CN117278396B discloses a DPU network card configuration method and device for network communication transmission.
[0012] However, the above-mentioned disclosed scheme has the following disadvantages: the above-mentioned scheme does not design differentiated processing measures for different types of RDMA messages in actual use, which limits the processing efficiency.
[0013] The present application proposes an RDMA message aggregation processing method and network card device to solve the above-mentioned problems. SUMMARY
[0014] The application aims to improve the bandwidth of the network card in processing RDMA message in the implementation of RDMA remote memory access based on RDMA protocol. The method comprises the following steps: the network card receives the RDMA message sent by the sending end, determines whether the received RDMA message can be aggregated according to the OpCode, Destination QP, PSN and other fields in the BTH header of the RDMA message, aggregates the received multiple RDMA messages into one larger RDMA message, and sends the larger RDMA message to the downstream module for processing, thereby reducing the number of messages processed by the downstream module, improving the bandwidth of the network card in processing RDMA message, and overcoming the problems in the above background art.
[0015] Based on the above technical idea, the technical scheme adopted by the application is as follows:
[0016] An RDMA message aggregation processing method and a network card device, comprising the following steps:
[0017] S1, an RDMA message analysis step, comprising the following steps: hardware level message capture, BTH header analysis, message classification and metadata generation;
[0018] S2, an RDMA cache step, comprising the following steps: message descriptor recording and cache strategy;
[0019] S3, an RDMA message aggregation step, comprising the following steps: aggregation context table, aggregation flow, new aggregation flow and existing context aggregation flow;
[0020] S4, a zero-copy aggregation step, comprising the following steps: metadata transfer, DMA delay processing and host memory writing;
[0021] S5, a dynamic control strategy step, comprising the following steps: dynamic threshold adjustment, exception flow processing and multi-QP concurrent support;
[0022] S6, a downstream protocol processing optimization step, comprising the following steps: the aggregated message is processed as a single logical unit, the protocol module is updated based on the aggregated Total_Payload_Size and Coalescing_Num to check the state, and the number of protocol stack processing (such as ACK generation, memory writing operation) is reduced.
[0023] In furtherance of the above technical solutions, the S1 RDMA packet analysis step, the hardware-level packet capture link includes a network card DMA engine receiving an original Ethernet packet, identifying a UDP target port number (RoCEv2 default 4791), filtering non-RDMA traffic, checking an IP / UDP header checksum, and directly discarding error packets; the BTH header analysis link includes locating the starting position of the BTH header (0th byte after the Ethernet header + IP header + UDP header), and analyzing the following fields: OpCode (5 bits): determining the packet type (such as SEND First, RDMA WRITE Middle, etc.); Destination QP (24 bits): target queue pair identification, used for flow classification; PSN (24 bits): packet sequence number, checking continuity; A (1 bit): whether ACK confirmation is required; SE (1 bit): whether it belongs to the same Message fragment; the packet classification and metadata generation link includes classification rules: First class: OpCode is SEND First, RDMA WRITE First, RDMA READ Response First; Middle class: packet with OpCode containing Middle; Last class: packet with OpCode containing Last; Others class: independent packet (such as SEND Only, ACK) or type that cannot be aggregated.
[0024] In furtherance of the above technical solutions, the S2 RDMA cache step, the packet descriptor record link includes a hardware ring buffer (Ring Buffer) storing packet metadata, and a fixed number of descriptors (such as 4096) are allocated during initialization to support multi-threaded lock-free access; the cache strategy link includes a write process, after the UserData is generated by the analysis module, the original packet DMA address and length are written to the idle descriptor, and the descriptor state is marked as “Pending Coalescing” (Pending Coalescing); the cache strategy link also includes timeout release, if the descriptor is not processed by the aggregation module within a predetermined time (such as 10 μs), it is forcibly released and notified to the downstream packet-by-packet processing.
[0025] In furtherance of the above technical solution, the S3 RDMA packet aggregation step, the aggregation context table link in the step includes table item capacity, supports configurable size (such as 16K-1M items), and query is accelerated through TCAM or hash table, the aggregation flow link includes aggregation enable check and context matching, the aggregation enable check includes reading a global configuration register of a network card, if the aggregation function is closed, all packets are directly transmitted to a downstream module, and the context matching includes using a UserData.qpn+UserData.function_id combination key to query the aggregation context table: hit (Existing Context): performing aggregation continuity check; miss (New Context): triggering new aggregation flow creation.
[0026] In furtherance of the above technical solution, the new aggregation flow creation link includes admission condition and context application, the admission condition allows only First or Middle type packets to create a new aggregation, the context application allocates a table item from an idle context pool, if the pool is empty, an LRU eviction strategy is triggered, and if the packet is a Last type packet but has no context, the packet is marked as an abnormal flow, and aggregation is prohibited; the existing context aggregation flow link includes PSN continuity check, checking whether a current packet PSN is equal to ctx.expected_psn, if not, the aggregation is terminated, and abnormal processing is triggered, packet type check, if the current packet is a First type packet but the context already exists, flow conflict is determined, the current aggregation is terminated, and the context is reset, aggregation update, the current packet descriptor is added to ctx.pd_list, and fields are updated.
[0027] In furtherance of the above technical solution, the S4 zero-copy aggregation step, the metadata transmission link in the step includes that the aggregation module only transmits aggregated_user_data and pd_list to a downstream protocol processing module, and does not involve packet content copying, the DMA delay processing link includes that a downstream module acquires a physical address of an original packet according to pd_list, splices on demand, deletes a header: skips a BTH header (such as deleting 40 bytes) of each packet according to edit_hdr_info, splices loads: logically splices loads of multiple packets into continuous virtual addresses (through a Scatter-Gather DMA descriptor), and the host memory writing link includes that a hardware DMA engine directly writes into a target memory area according to a logical structure (continuous load) after aggregation, and only triggers PCIe transmission once.
[0028] In further limitation of the technical solution, the S5 dynamic control strategy step, the dynamic threshold adjustment link in the step includes monitoring the network card load (such as PPS, cache utilization), dynamically adjusting the aggregation threshold, reducing the threshold (such as reducing the maximum aggregation number from 64 to 16) in high load, reducing the processing delay, increasing the threshold in low load, and maximizing the bandwidth utilization, the abnormal flow processing link includes PSN disorder: forcibly terminating aggregation, discarding the disordered message and triggering retransmission (informing the opposite end through NACK), context leakage: timer scanning timeout context (such as not updating for more than 100 μs), forcibly releasing and recording log, and the multi-QP concurrent support link includes dispersing the contexts of different QPs to multiple hardware tables through Hash Partitioning to avoid lock competition.
[0029] In further limitation of the technical solution, the S6 downstream protocol processing optimization step, the step further includes an on-demand enabling aggregation link and a threshold dynamic adjustment link, the on-demand enabling aggregation link includes supporting a global switch (such as NIC register configuration) and a QP-level strategy (dynamically controlled through Verbs API), and the aggregation is closed for delay-sensitive small messages (such as control signaling) and enabled for large flows (such as storage block transmission), and the threshold dynamic adjustment link includes a configurable upper limit of the number of aggregated messages (such as 1-64), a maximum aggregation length (such as 64 KB), a timeout time (such as 10 μs), and self-adaptive network load: automatically reducing the aggregation threshold in a high congestion scenario to preferentially guarantee real-time performance.
[0030] In further limitation of the technical solution, the S6 downstream protocol processing optimization step, the step further includes a hardware offloading aggregation logic link and a quantified benefit link, the hardware offloading aggregation logic link includes realizing message analysis, context management and zero-copy aggregation through intelligent network card hardware to completely release the host CPU, and the quantified benefit link includes CPU utilization optimization: in the NVMe over Fabrics scenario, the CPU occupancy rate is reduced from 30% to less than 5%, and the energy efficiency ratio is improved: the hardware aggregation power consumption is only 1 / 10 (actual measurement data) of the software scheme.
[0031] Compared with the prior art, the beneficial effects of the present application are:
[0032] 1. Dynamic aggregation strategy, based on OpCode and PSN in BTH header, the continuous messages (First / Middle / Last) of the same Message are aggregated into a logical unit to reduce the number of downstream processing times;
[0033] 2. Metadata-driven aggregation, only passing User Data (message type, PSN, load length, etc.) and message descriptor, and preserving the original message storage position;
[0034] 3. Protocol-aware aggregation, identify message type by OpCode (e.g. Send First / Middle), ensure aggregation complies with RDMA protocol segmentation rules. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0036] Figure 1 A flowchart of the RDMA message aggregation processing method and the network card device of the present application;
[0037] Figure 2 A schematic diagram of the RDMA message content of the present application;
[0038] Figure 3 A schematic diagram of the reorganization information of the present application;
[0039] Figure 4 A processing flowchart of the newly created aggregation context information in the present application;
[0040] Figure 5 A processing flowchart of the existing aggregation flow processing in the present application;
[0041] Figure 6 A schematic diagram of the module of the present application;
[0042] Figure 7 A schematic diagram of the Message aggregation of the present application;
[0043] Figure 8 A schematic diagram of the Send First aggregation information of the present application;
[0044] Figure 9 A schematic diagram of the Send Middle aggregation of the present application;
[0045] Figure 10 A schematic diagram of the Send Last message aggregation of the present application;
[0046] Figure 11 A schematic diagram of the structure in the RDMA message aggregation processing method and the network card device of the present application;
[0047] Figure 12 A schematic diagram of the UserData format of the present application;
[0048] Figure 13 A schematic diagram of the aggregation context format of the present application. DETAILED DESCRIPTION
[0049] The application is further described below in conjunction with the accompanying drawings of which: Figures 1-13 The application is further described below in conjunction with the accompanying drawings of which:
[0050] Embodiment 1: This embodiment provides a RDMA packet aggregation processing method and a network card device, as shown in the following figure, including the following steps: Figures 1-13
[0051] S1 RDMA packet analysis step, including hardware level packet capture, BTH header analysis and packet classification and metadata generation links;
[0052] S2 RDMA cache step, including packet descriptor record and cache strategy links;
[0053] S3 RDMA packet aggregation step, including aggregation context table, aggregation flow, new aggregation flow and existing context aggregation flow links;
[0054] S4 zero-copy aggregation step, including metadata transfer, DMA delay processing and host memory writing links;
[0055] S5 dynamic control strategy step, including dynamic threshold adjustment, exception flow processing and multi-QP concurrent support links;
[0056] S6 downstream protocol processing optimization step, including processing the aggregated packet as a single logical unit, checking the protocol module based on the aggregated Total_Payload_Size and Coalescing_Num to update the state, and reducing the protocol stack processing times (such as ACK generation, memory writing operation).
[0057] The S1 RDMA message analysis step, the hardware level message capture link includes a network card DMA engine receiving original Ethernet message, identifying UDP target port number (RoCEv2 default 4791), filtering non-RDMA flow, checking IP / UDP header checksum, and directly discarding error messages; the BTH header analysis link includes locating the BTH header starting position (0th byte after the Ethernet header + IP header + UDP header), and analyzing the following fields: OpCode (5 bits): determining the message type (such as SEND First, RDMAWRITE Middle, etc.); Destination QP (24 bits): target queue pair identification, used for flow classification; PSN (24 bits): message sequence number, checking continuity; A (1 bit): whether ACK confirmation is required; SE (1 bit): whether it belongs to the same Message fragment; the message classification and metadata generation link includes classification rules, First class: OpCode is SEND First, RDMAWRITE First, RDMA READ Response First; Middle class: message containing OpCode Middle; Last class: message containing OpCode Last; Others class: independent message (such as SEND Only, ACK) or message of an aggregation type that cannot be aggregated.
[0058] The S2 RDMA cache step, the message descriptor record link includes a hardware ring buffer (RingBuffer) storing message metadata, and a fixed number of descriptors (such as 4096) are allocated during initialization to support multi-threaded lock-free access; the cache strategy link includes a write process, after the UserData is generated by the analysis module, the original message DMA address and length are written into the idle descriptor, and the descriptor state is marked as “Pending Coalescing” (Pending Coalescing); the cache strategy link also includes timeout release, if the descriptor is not processed by the aggregation module within a predetermined time (such as 10 μs), it is forced to release and notify the downstream packet-by-packet processing.
[0059] The S3 RDMA message aggregation step, wherein the aggregation context table component includes table capacity, supports configurable size (such as 16K-1M entries), and the query is accelerated by TCAM or hash table, the aggregation process component includes aggregation enable check and context matching, the aggregation enable check includes reading the global configuration register of the network card, if the aggregation function is closed, all messages are directly transmitted to the downstream module, and the context matching includes using the UserData.qpn+UserData.function_id combination key to query the aggregation context table: hit (Existing Context): performing aggregation continuity check; miss (New Context): triggering new aggregation process.
[0060] The new aggregation process component includes admission condition and context application, the admission condition only allows First or Middle type messages to create a new aggregation, the context application allocates table entries from the idle context pool, if the pool is empty, triggers the LRU eviction strategy, and the exception handling marks as an abnormal flow if the message is of the Last type but has no context, and prohibits aggregation; the existing context aggregation process component includes PSN continuity check, checking whether the current message PSN is equal to ctx.expected_psn, if not, terminating the aggregation, triggering exception handling, message type check, if the current message is of the First type but the context already exists, determining as flow conflict, terminating the current aggregation and resetting the context, aggregation update, adding the current message descriptor to ctx.pd_list, and updating the field.
[0061] The S4 zero-copy aggregation step, wherein the metadata transmission component includes that the aggregation module only transmits aggregated_user_data and pd_list to the downstream protocol processing module, and does not involve message content copying, the DMA delay processing component includes that the downstream module acquires the original message physical address according to pd_list, as needed, splicing, header deletion: skipping the BTH header (such as deleting 40 bytes) of each message according to edit_hdr_info, load splicing: logically splicing the load parts of multiple messages into continuous virtual addresses (by using Scatter-Gather DMA descriptor), and the host memory writing component includes that the hardware DMA engine directly writes into the target memory area according to the logical structure (continuous load) after aggregation, and triggers PCIe transmission only once.
[0062] The S5 dynamic control strategy step, in which the dynamic threshold adjustment link includes monitoring the network card load (such as PPS, cache utilization), dynamically adjusting the aggregation threshold, reducing the threshold (such as reducing the maximum aggregation number from 64 to 16) in high load to reduce processing delay, increasing the threshold in low load to maximize bandwidth utilization, the abnormal flow processing link includes PSN disorder: forcibly terminating aggregation, discarding the out-of-order message and triggering retransmission (by NACK informing the opposite end), context leakage: timer scanning timeout context (such as not updating for more than 100 μs), forced release and log recording, and the multi-QP concurrent support link includes dispersing the contexts of different QPs to multiple hardware tables through hash partitioning to avoid lock contention.
[0063] The S6 downstream protocol processing optimization step further includes an on-demand aggregation enabling link and a threshold dynamic adjustment link, the on-demand aggregation enabling link includes supporting a global switch (such as a NIC register configuration) and a QP-level strategy (dynamically controlled through a Verbs API), and aggregation is closed for delay-sensitive small messages (such as control signaling) and aggregation is enabled for large flows (such as storage block transmission), and the threshold dynamic adjustment link includes a configurable upper limit of the number of aggregated messages (such as 1-64), a maximum aggregation length (such as 64 KB), a timeout time (such as 10 μs), and adaptive network load: automatically reducing the aggregation threshold in a high congestion scenario to prioritize real-time performance.
[0064] The S6 downstream protocol processing optimization step further includes a hardware offloading aggregation logic link and a quantitative benefit link, the hardware offloading aggregation logic link includes realizing message parsing, context management and zero-copy aggregation through intelligent network card hardware to completely release the host CPU, and the quantitative benefit link includes CPU utilization optimization: in the NVMe over Fabrics scenario, the CPU occupancy rate is reduced from 30% to less than 5%, and energy efficiency ratio is improved: the hardware aggregation power consumption is only 1 / 10 (actual measurement data) of the software solution.
[0065] Embodiment 2: The embodiment provides an RDMA message aggregation processing method and a network card device, as shown in the following figure, including the following operation steps: Figures 1-13
[0066] Receiving a 64 KB RDMA Write Message (split into 64 1 KB messages);
[0067] First message processing (First message):
[0068] The analysis module identifies it as an RDMA WRITE First, generates User data and creates a new aggregation context.
[0069] Middle message processing (Middle message):
[0070] After consecutive successful PSN verifications and a cumulative payload_size of 63KB, the context for handling last packet (Last packet) is updated.
[0071] Trigger the aggregation termination condition and generate aggregated UserData (total length 64KB, 64 messages).
[0072] Zero-copy output:
[0073] The DMA engine writes the payload of 64 packets directly to the target memory in just one transfer (instead of 64).
[0074] Example 3: This example provides an RDMA packet aggregation processing method and network interface card (NIC) device, such as... Figures 1-13 As shown, the following operational steps are also included:
[0075] RDMA message parsing module.
[0076] like Figure 2 As shown, according to the RoCEv2 protocol, RDMA messages contain a BTH (Base Transport Header). The message type can be identified through the OpCode field of the BTH header, and whether the messages belong to the same flow can be identified through the Destination QP field. Only messages belonging to the same flow can participate in aggregation, also known as aggregation candidate messages.
[0077] like Figure 3 As shown, messages belonging to the same Message can be distinguished by OpCode. For example, for Send messages in RC mode, if the Message is very large, it will be split into different OpCodes such as Send First, Send Middle, and Send Last when it is sent. The Message can be reassembled by the OpCode.
[0078] The parsing module identifies all types of RDMA messages and classifies received RDMA messages into four categories: First, Middle, Last, and Others. This facilitates stateless aggregation by the subsequent aggregation module. Simultaneously, it extracts the necessary data from the aggregation message, including fields such as Destination QP, PSN, A, SE, and M, to form UserData information, which is then sent to downstream modules. Directly transmitting information through UserData simplifies processing by downstream modules.
[0079] RDMA cache module
[0080] The RDMA message analysis module transmits the UserData information to the RDMA message aggregation module for aggregation processing.
[0081] The RDMA message aggregation module.
[0082] The RDMA message aggregation module receives the UserData information and the message descriptor sent by the upstream module, and checks whether the network card enables the RDMA message aggregation function. If the network card enables the RDMA message aggregation function, the aggregation operation is performed; otherwise, the received RDMA message is processed in the traditional way. If the RDMA message aggregation function is enabled, the aggregation context table is searched according to the UserData information to determine whether the RDMA message belongs to an existing aggregation. The process can be divided into two situations: a new aggregation process and an existing aggregation process.
[0083] Specifically, the Destination QP and the Function Num in the UserData are used to identify an aggregation context.
[0084] New aggregation process
[0085] When the aggregation context is missed by searching the aggregation context by using the Destination QP and the Function Num as the KEY, the new aggregation process is entered. In the new aggregation process, whether the message is the First and Middle message is determined according to the type field in the UserData. If the message is the First and Middle message, the aggregation context is applied; if not, the non-RDMA aggregation message processing process is directly entered.
[0086] As shown in Figure 4 If the aggregation context table is full and has no remaining space, the oldest created aggregation context is forcibly ended, and the space is used to store the newly created aggregation context information.
[0087] Existing aggregation process
[0088] As shown in Figure 5 When the aggregation context is hit by searching the aggregation context by using the Destination QP and the Function Num as the KEY, the existing aggregation process is entered. The existing aggregation process determines, according to the received UserData information and the hit aggregation context, whether the newly received RDMA message continues to participate in the aggregation or terminates the current aggregation.
[0089] Aggregation end strategy
[0090] In order to reduce the occupation of the network card resources, the number of aggregated messages is limited, and the aggregation is actively ended when the configured threshold is exceeded.
[0091] Zero-copy aggregated message
[0092] The required information is extracted as User Data and sent to the downstream module for processing. The downstream module only needs to record the message descriptor and indicate the length of the deleted message header in the User Data. During the entire processing process, no copy of the RDMA message is involved. Only when the message needs to be moved to the host at the end, DMA moving is performed according to the message descriptor and the length of the deleted message indicated in the User Data. During the entire aggregation processing process, no message moving is involved.
[0093] Embodiment 4: The embodiment provides an RDMA message aggregation processing method and a network card device, as shown in Figures 1-13 , and further comprising
[0094] An example of aggregation implementation for RDMA Send message is as follows:
[0095] 1. The Send Message is composed of two messages of Send First and Send Last. After the network card receives the two messages, the network card aggregates the two messages into a complete Message and sends the Message to the downstream RDMA message protocol processing module for protocol processing, as shown in Figure 7 .
[0096] 2. The Send Message is composed of one Send First, multiple Send Middle, and one Send Last. After the network card receives the messages, the network card aggregates the messages into a complete Message and sends the Message to the downstream RDMA message protocol processing module for protocol processing, as shown in Figure 8 .
[0097] 3. The Send Message is composed of one Send First, multiple Send Middle, and one Send Last. However, because the number of messages is too large, the Send First and the multiple Send Middle cannot be aggregated into a complete Message at one time. The Send First and the multiple Send Middle are aggregated into a larger Send First message, which is sent to the downstream RDMA message protocol processing module for protocol processing, as shown in Figure 9 .
[0098] 4、Send Message is composed of 1 Send First and multiple Send Middle and 1 Send Last message, but because the Message is too much, it cannot be aggregated into a complete Message at one time, several Send Middle are aggregated into a larger Send Middle message, which is sent to the downstream RDMA message protocol processing module for protocol processing, as shown in Figure 10
[0099] 5、Send Message is composed of 1 Send First and multiple Send Middle and 1 Send Last message, but because the Message is too much, it cannot be aggregated into a complete Message at one time, several Send Middle and the last 1 Send Last are aggregated into a larger Send Last message, which is sent to the downstream RDMA message protocol processing module for protocol processing, as shown in Figure 11
[0100] 6、It also includes a possible UserData format, as shown in Figure 12
[0101] 7、It also includes a possible aggregation context format, as shown in Figure 13
[0102] The above is a further detailed description of the present application in combination with specific preferred embodiments, which facilitates the understanding and application of the present application by those skilled in the art, and cannot be regarded as limiting the specific implementation of the present application to these descriptions.
Claims
1. A method and network interface card (NIC) for RDMA packet aggregation processing, characterized in that, Includes the following steps: The S1 RDMA message parsing steps include hardware-level message capture, BTH header parsing, and message classification and metadata generation. The S2 RDMA caching step includes message descriptor recording and caching strategy steps. The S3 RDMA message aggregation step includes the aggregation context table, aggregation process, new aggregation process, and existing context aggregation process steps. The S4 zero-copy aggregation step includes metadata transfer, DMA latency processing, and host memory write. S5 Dynamic Control Strategy Steps, which include dynamic threshold adjustment, abnormal flow handling, and multi-QP concurrency support; S6 downstream protocol processing optimization steps include processing the aggregated message as a single logical unit, checking the protocol module update status based on the aggregated Total_Payload_Size and Coalescing_Num, and reducing the number of protocol stack processing steps. The S5 dynamic control strategy steps include a dynamic threshold adjustment step that monitors network card load, dynamically adjusts the aggregation threshold, lowers the threshold under high load to reduce processing latency, and raises the threshold under low load to maximize bandwidth utilization. The abnormal flow processing step includes PSN out-of-order handling: forcibly terminating aggregation, discarding out-of-order packets and triggering retransmission, and handling context leakage: scanning timed-out contexts, forcibly releasing and logging them. The multi-QP concurrency support step includes distributing the contexts of different QPs to multiple hardware tables through hash sharding to avoid lock contention. The S6 downstream protocol processing optimization step further includes an on-demand aggregation activation step and a dynamic threshold adjustment step. The on-demand aggregation activation step includes support for global switching and QP-level policies, disabling aggregation for latency-sensitive small packets and enabling aggregation for large flows. The dynamic threshold adjustment step includes configurable upper limit for the number of aggregated packets, maximum aggregation length, and timeout time, and adaptive network load: automatically reducing the aggregation threshold in high congestion scenarios to prioritize real-time performance. The S6 downstream protocol processing optimization step also includes a hardware offloading aggregation logic step and a quantification benefit step. The hardware offloading aggregation logic step includes packet parsing, context management, and zero-copy aggregation through smart network interface card hardware, completely freeing up the host CPU. The quantification benefit step includes CPU utilization optimization: in the NVMe over Fabrics scenario, the CPU utilization rate is reduced from 30% to below 5%, and the energy efficiency ratio is improved: the power consumption of hardware aggregation is only 1 / 10 of that of the software solution.
2. The RDMA packet aggregation processing method and network card device according to claim 1, characterized in that, The S1RDMA packet parsing step includes a hardware-level packet capture stage where the network card DMA engine receives raw Ethernet packets, identifies the UDP destination port number, filters non-RDMA traffic, verifies the IP / UDP header checksum, and discards erroneous packets. The BTH header parsing stage includes locating the start position of the BTH header and parsing it according to the following fields: OpCode: determines the packet type; Destination QP: target queue pair identifier, used for flow classification; PSN: packet sequence number, used to verify continuity. A: Whether ACK confirmation is required; SE: Whether it belongs to the same Message fragment; The message classification and metadata generation process includes classification rules: First class: OpCode is SEND First, RDMA WRITE First, RDMA READ Response First; Middle class: Messages with OpCode containing Middle; Last class: Messages with OpCode containing Last; Others class: Independent messages or types that cannot be aggregated.
3. The RDMA packet aggregation processing method and network card device according to claim 2, characterized in that, The S2RDMA caching step includes a message descriptor recording stage that stores message metadata in a hardware circular buffer, allocating a fixed number of descriptors during initialization to support multi-threaded lock-free access. The caching strategy stage includes a writing process where, after the parsing module generates UserData, the original message DMA address and length are written to an idle descriptor, and the descriptor status is marked as "pending aggregation." The caching strategy stage also includes timeout release; if a descriptor is not processed by the aggregation module within a predetermined time, it is forcibly released and downstream devices are notified to process it packet by packet.
4. The RDMA packet aggregation processing method and network card device according to claim 3, characterized in that, The S3RDMA packet aggregation step includes an aggregation context table with configurable entry size, accelerated by TCAM or hash table lookups. The aggregation process includes aggregation enable checks and context matching. The aggregation enable check involves reading the network interface card's global configuration register; if aggregation is disabled, all packets pass directly to downstream modules. Context matching involves querying the aggregation context table using the UserData.qpn + UserData.function_id key combination: if a match is found, aggregation continuity verification is performed; otherwise, a new aggregation process is triggered.
5. The RDMA packet aggregation processing method and network interface card device according to claim 4, characterized in that, The process of creating a new aggregation includes admission conditions and context request. Admission conditions allow only First or Middle class messages to create new aggregations. Context request allocates entries from the idle context pool. If the pool is empty, the LRU eviction policy is triggered. Exception handling marks the Last class message as an abnormal flow if it has no context, and aggregation is prohibited.
6. The RDMA packet aggregation processing method and network interface card device according to claim 5, characterized in that, The existing context aggregation process includes PSN continuity verification, checking whether the current packet PSN is equal to ctx.expected_psn. If they are not equal, the aggregation is terminated and exception handling is triggered. Packet type verification, if the current packet is of the First class but the context already exists, it is determined to be a flow conflict, the current aggregation is terminated and the context is reset, and the aggregation is updated by adding the current packet descriptor to ctx.pd_list and updating the fields.
7. The RDMA packet aggregation processing method and network interface card device according to claim 5, characterized in that, The S4 zero-copy aggregation step includes a metadata transmission stage where the aggregation module only transmits aggregated_user_data and pd_list to the downstream protocol processing module, without copying the packet content. The DMA delay processing stage includes the downstream module obtaining the original packet physical address based on pd_list, concatenating them as needed, deleting the header (BTH header of each packet based on edit_hdr_info), and concatenating the payload portion of multiple packets into a continuous virtual address. The host memory writing stage includes the hardware DMA engine directly writing to the target memory region according to the aggregated logical structure, triggering only one PCIe transfer.
Citation Information
Patent Citations
DPU network card configuration method and device
CN117278396B
Network interface supporting remote data direct access protocol
CN116722884A