RDMA protocol efficient analysis and forwarding method based on RISC-V architecture
By introducing hardware acceleration modules and customized instruction sets on the RISC-V architecture, combined with the dynamic traffic scheduling of the hardware scheduling module, the problem of CPU performance bottlenecks in high-frequency and large-flow data processing is solved, and efficient packet analysis and forwarding is achieved.
Patent Information
- Application Number
- CN202510053378.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing high-frequency and large-stream data, there are performance bottlenecks in traditional CPUs for protocol stack analysis and processing.
The hardware acceleration module and customized instruction set based on the RISC-V architecture are adopted to parse and verify the RDMA protocol, and dynamic traffic scheduling and load balancing are performed through the hardware scheduling module.
It significantly improves the packet parsing speed, reduces CPU load, avoids overloading of a single computing node, and ensures efficient forwarding of data packets and system stability.
Smart Images

Figure CN119996539A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of power systems, and more specifically, to an efficient parsing and forwarding method for RDMA protocol based on RISC-V architecture. Background Art
[0002] With the continuous development of smart grids and modern energy management systems, power grids are gradually moving towards digitalization, informatization and intelligence. The introduction of these technologies enables power grids to have the capabilities of real-time monitoring, remote control and optimized dispatching. However, the complexity of power grids has also increased, especially when dealing with large amounts of real-time data and high-frequency communications. How to effectively parse and forward power grid data streams has become an important issue.
[0003] The communication system in the smart grid usually adopts a distributed architecture, and various smart devices and control units (such as substations, distribution networks, load control devices, etc.) need to be interconnected through high-speed networks. In order to improve communication efficiency and reduce latency, many modern power grid communication systems use remote direct memory access (RDMA) technology, which can provide high-bandwidth, low-latency data transmission, reduce processor intervention in the data transmission process, and improve data processing efficiency.
[0004] Data transmission in the power grid mostly relies on network protocols such as TCP / IP, UDP or customized power grid protocols.
[0005] The traditional approach is to parse and process data packets by running complex protocol stacks on the CPU. These protocol stacks are responsible for receiving data packets from the network, performing protocol parsing, verification, and data processing, and determining the forwarding path for data packets. Although this method is common, it is limited by CPU performance, especially when processing high-frequency, high-volume data, where performance bottlenecks are more obvious. Summary of the invention
[0006] The present invention provides an efficient parsing and forwarding method for the RDMA protocol based on the RISC-V architecture, which is intended to solve the technical problem that the performance bottleneck is currently limited by the CPU performance, especially when processing high-frequency and large-flow data.
[0007] The efficient parsing and forwarding method of the RDMA protocol based on the RISC-V architecture includes the following steps:
[0008] Step 1: The data packet obtained by the edge node reaches the local computing node through the RDMA network interface;
[0009] Step 2: The data packet is parsed and verified by the hardware acceleration module for the RDMA protocol. The header information is extracted and CRC verification is performed through the customized RISC-V instruction set.
[0010] Step 3: The parsed and verified data packets are transmitted to the flow module according to the target information. The local computing node determines the forwarding path of the data packet based on the routing table and topology information;
[0011] Step 4: Based on the determined forwarding path, the data packet is forwarded to the target processing unit or network interface, wherein if the forwarding target is a local computing node, the data packet is stored in the cache; if the forwarding target is a remote computing node, the data packet continues to be forwarded through the network interface;
[0012] Step 5: During packet forwarding, the hardware scheduling module dynamically adjusts traffic distribution according to the current load conditions and controls the processing load of each processing unit to avoid overload.
[0013] The present invention proposes an efficient parsing and forwarding method for the RDMA protocol based on the RISC-V architecture, which provides an effective solution to the performance bottleneck problem caused by the CPU performance limitation in the current high-frequency and large-flow data processing. By introducing a hardware acceleration module and a customized RISC-V instruction set, the method realizes the hardware acceleration of the protocol parsing and verification process, thereby greatly improving the parsing speed and reducing the CPU load. In the protocol parsing process of the data packet, this scheme optimizes the instruction set for the special needs of the RDMA protocol by utilizing the flexibility and customizability of the RISC-V architecture, and can efficiently extract the data packet header information and perform CRC verification. At the same time, through dynamic traffic scheduling and load balancing, the hardware scheduling module can monitor the load of each computing unit in real time, intelligently adjust the traffic distribution, avoid the overload problem of a single computing node, and ensure the efficient forwarding of the data packet. In addition, based on the customized processing capability of the RISC-V architecture, the forwarding path of the data packet can be flexibly adjusted according to the real-time network topology and routing information, which improves the adaptability and processing capability of the entire network, thereby effectively solving the problem that the traditional processing scheme cannot meet the performance requirements under high traffic and high concurrency conditions.
[0014] Preferably, the format of the data packet includes a protocol header, a data payload and a tail;
[0015] The protocol header includes a source address, a destination address, a protocol type, and a data length;
[0016] The data payload includes data content to be transmitted;
[0017] The tail contains CRC check information of the data packet.
[0018] Preferably, step 2 comprises the following steps:
[0019] Data packets are stored in DMA buffer: After the hardware interface RDMA NIC receives the data packet, the data is stored in a temporary DMA buffer;
[0020] Customized RISC-V instruction set and protocol parsing: Parse the source address, destination address, protocol type, and data length of the data packet based on the customized RISC-V instruction set;
[0021] Based on the RISC-V instruction set, a set of instructions are defined to quickly calculate the CRC check value:
[0022]
[0023] Where: D = d0, d1, ..., d n―1 Indicates the content of the data packet and calculates CRC bit by bit; Indicates bitwise XOR; InitialCRC indicates the initial value of CRC;
[0024] By initializing the CRC value to 0xFFFFFFFF, an XOR operation is performed on each byte of each data packet, and the data is shifted and updated by table lookup. After the calculation is completed, the CRC value is reversed and compared with the check value in the CRC check information. If they are consistent, the check is passed, and if they are inconsistent, the check fails.
[0025] Information extraction and error handling: Extract the protocol header fields and CRC check results. If the check fails, the data packet is discarded.
[0026] Preferably, step 3 comprises the following steps:
[0027] Routing table lookup and path calculation: Match the target IP address with each entry in the routing table, select the entry that matches the longest prefix, obtain the next hop information of the matching entry, and determine the forwarding path of the data packet based on the matching result;
[0028] The routing table information stores routing information through a hardware Trie tree to achieve fast search.
[0029] Preferably, the steps of determining the forwarding path based on the hardware Trie tree are as follows:
[0030] Pass the input target IP address bit by bit to the Trie tree;
[0031] Starting from the root node of the Trie tree, match each bit of the target IP address:
[0032] If the current bit of the target IP is 0, continue searching along the left subtree;
[0033] If the current value of the target IP is 1, continue searching along the right subtree;
[0034] Whenever a node is queried, the routing information of the node is recorded;
[0035] After querying all bits, the routing information corresponding to the longest matching prefix is returned.
[0036] Preferably, the step 5 comprises the following steps:
[0037] Load detection and real-time monitoring: The hardware scheduling module periodically monitors the current load of each processing unit to obtain the real-time load data of each processing unit, including the current queue length, CPU usage, memory usage, and network bandwidth usage;
[0038] Threshold judgment: If the load of a processing unit exceeds the preset upper threshold, the processing unit is considered to be overloaded and tasks need to be reallocated;
[0039] If the load is below a preset lower threshold, the processing unit is considered to be in an idle state;
[0040] Task migration: Migrate data packets from overloaded processing units to idle processing units through hash algorithms;
[0041] Traffic segmentation: The hardware scheduling module distributes traffic according to the load of each processing unit to achieve load balancing:
[0042]
[0043] Where: P in Indicates the number of packets to be forwarded; P out,i represents the number of packets assigned to the i-th processing unit; L i represents the load value of the i-th processing unit; L max represents the maximum load value; N represents the number of processing units.
[0044] Preferably, the steps of implementing task migration through the hash algorithm are as follows:
[0045] Load-aware value calculation: A load-aware value is calculated for each processing unit:
[0046] LAHV i =w1×CPU i +w2×QueueLength i +w3×Memory i +w4×Bandwidth i ;
[0047] Where: LAHV i represents the load perception value of the i-th processing unit; CPUi Indicates the CPU usage of the i-th processing unit; QueueLength i Indicates the queue length of the i-th processing unit; Memory i Indicates the memory usage of the i-th processing unit; Bandwidth i represents the bandwidth utilization rate of the i-th processing unit; w1, w2, w3, and w4 represent weight coefficients respectively;
[0048] Calculate hash value: Apply hash function to the load-aware value of each processing unit and map the load-aware value to a position in the hash ring:
[0049] Hash i =HashFunction(LAHV i );
[0050] Where: Hash i represents the hash value of the i-th processing unit;
[0051] Ring construction: Map each processing unit to a position in the hash ring;
[0052] Calculate the hash value of the data packet: Calculate the hash value using an identifier of the data packet:
[0053] TaskHash=HashFuntion(PacketIdentifier);
[0054] Where: TaskHash represents the hash value of the data packet; PacketIdentifier represents the identifier of the data packet;
[0055] Search for a processing unit on the hash ring: Start from the hash value on the hash ring and search for the next processing unit in a clockwise direction. According to the size of the hash value, find the processing unit with the lightest load and matching hash value in the hash ring to migrate the task.
[0056] After the data packet is migrated to the target processing unit, the load perception value of the target processing unit needs to be updated. By periodically calculating the load perception value of each processing unit, the node position on the hash ring is dynamically adjusted.
[0057] The beneficial effects of the present invention include:
[0058] The present invention proposes an efficient parsing and forwarding method for the RDMA protocol based on the RISC-V architecture, which provides an effective solution to the performance bottleneck problem caused by the CPU performance limitation in the current high-frequency and large-flow data processing. By introducing a hardware acceleration module and a customized RISC-V instruction set, the method realizes the hardware acceleration of the protocol parsing and verification process, thereby greatly improving the parsing speed and reducing the CPU load. In the protocol parsing process of the data packet, this scheme optimizes the instruction set for the special needs of the RDMA protocol by utilizing the flexibility and customizability of the RISC-V architecture, and can efficiently extract the data packet header information and perform CRC verification. At the same time, through dynamic traffic scheduling and load balancing, the hardware scheduling module can monitor the load of each computing unit in real time, intelligently adjust the traffic distribution, avoid the overload problem of a single computing node, and ensure the efficient forwarding of the data packet. In addition, based on the customized processing capability of the RISC-V architecture, the forwarding path of the data packet can be flexibly adjusted according to the real-time network topology and routing information, which improves the adaptability and processing capability of the entire network, thereby effectively solving the problem that the traditional processing scheme cannot meet the performance requirements under high traffic and high concurrency conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0060] Figure 1 An overall step block diagram provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0061] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0062] See also Figure 1 As shown, the best embodiment of the present invention is further described;
[0063] The efficient parsing and forwarding method of the RDMA protocol based on the RISC-V architecture includes the following steps:
[0064] Step 1: The data packets obtained by the edge node reach the local computing node through the RDMA network interface; for example, in a smart substation, a large number of sensors are installed to monitor current, voltage, power and other parameters. These data are transmitted to the central control system through the RDMA protocol. The control system detects in real time that the current in a certain area is abnormally high, and immediately sends a control command to adjust the load in the area and start the backup power supply. At the same time, the system also transmits real-time data to the dispatcher through the RDMA protocol for further analysis and processing.
[0065] Step 2: The data packet is parsed and verified by the hardware acceleration module for the RDMA protocol. The header information is extracted and CRC verification is performed through the customized RISC-V instruction set.
[0066] Data packet reception and preliminary analysis:
[0067] After the data packet arrives at the local computing node through the RDMA network interface, it is directly transmitted to the hardware acceleration module (such as a dedicated RDMA network interface card, DMA controller, etc.) for preliminary processing.
[0068] Data packet format: First, the basic format of a data packet includes a protocol header, a data payload, and a tail (which usually contains checksum information).
[0069] Protocol header: contains key information such as source address, destination address, protocol type, data length, etc.
[0070] Data payload: carries the data content to be transmitted.
[0071] Trailer: Contains the CRC check information of the data packet.
[0072] Packets are stored in DMA buffer: After the hardware interface (such as RDMA NIC) receives a packet, the data is first stored in a temporary DMA buffer for further processing.
[0073] Customized RISC-V instruction set and protocol analysis:
[0074] Through the customized RISC-V instruction set, the hardware acceleration module can efficiently extract the protocol header information in the RDMA data packet and perform CRC check. The specific implementation process is as follows:
[0075] Protocol header parsing: The customized RISC-V instruction set can directly access the memory location of the data packet and quickly parse out the various fields of the protocol header. The following are typical steps for parsing the protocol header:
[0076] Source address resolution: Extract the source address field in the data packet, which is usually a 16-byte IPv6 address or a 4-byte IPv4 address.
[0077] Destination address parsing: Extract the destination address field.
[0078] Protocol type parsing: Extract the protocol type field (such as RDMA write, RDMA read, RDMA send, etc.).
[0079] Data length analysis: Extract the length information of the data payload.
[0080] Example of custom instructions: In RISC-V, a dedicated instruction can be designed to extract the 16-byte source address directly from the packet buffer:
[0081] ld r1,[r0] / / Load 16 bytes of data (source address) from memory address r0 to register r1;
[0082] CRC check: CRC check of data packets is an important step to ensure data integrity and correctness. The RISC-V instruction set can customize a set of instructions to quickly calculate CRC check values and reduce CPU processing overhead.
[0083] Assuming that the CRC check is based on the standard CRC-32 (a common 32-bit CRC algorithm), the calculation formula is as follows:
[0084]
[0085] Where: D = d0, d1, ..., d n―1 Indicates the content of the data packet and calculates CRC bit by bit; Indicates bitwise XOR; knitialCRC indicates the initial value of CRC;
[0086] By initializing the CRC value to 0xFFFFFFFF, an XOR operation is performed on each byte of each data packet, and the data is shifted and updated by table lookup. After the calculation is completed, the CRC value is reversed and compared with the check value in the CRC check information. If they are consistent, the check is passed, and if they are inconsistent, the check fails.
[0087] Custom instructions can be used to complete this step through a dedicated hardware module. Using hardware-accelerated CRC calculations can significantly improve efficiency. For example, the following CRC registers and table lookup modules may be included in the hardware module:
[0088] crc32 r2,r1,0xFFFFFFFF / / Use the hardware acceleration module to calculate the CRC32 checksum and return it to register r2; further processing after data packet parsing
[0089] After the protocol parsing and CRC check are completed, the information of the data packet can be used for further flow processing. The key to this process is to extract the parsed information and transfer it to the flow module so as to further determine the forwarding path of the data packet based on the target information.
[0090] Extracted information:
[0091] Protocol header fields (source address, destination address, protocol type, etc.)
[0092] CRC check result (whether it passed the check)
[0093] Error handling: If the CRC check fails (i.e. the integrity of the data packet does not pass), the data packet will be discarded to prevent the erroneous data from entering the subsequent processing chain.
[0094] RISC-V custom instruction example:
[0095] Data loading instructions:
[0096] ld r1,[r0] / / load 16-byte source address from memory to register r1
[0097] CRC calculation instructions:
[0098] crc32 r2,r1,0xFFFFFFFF / / Calculate the CRC32 checksum of the data in register r1, the initial value is 0xFFFFFFFF;
[0099] The design ideas of the customized RISC-V instruction set are as follows:
[0100] Customized instruction sets usually include hardware-level optimizations for specific operations, such as data extraction, checksum calculations, etc. These operations are critical for RDMA protocol processing. We will design customized instruction sets in the following aspects:
[0101] Data Extraction Instructions: Efficiently extract protocol header fields from RDMA packets.
[0102] CRC check instruction: accelerates CRC calculation to ensure data integrity.
[0103] Endian and memory manipulation instructions: Simplify endian conversion and memory access.
[0104] Data extraction instruction design:
[0105] Extracting header information from RDMA packets is a key step in the parsing process. The header of a packet usually contains fields such as source address, destination address, and protocol type. The purpose of designing customized instructions is to directly access memory and efficiently extract fields through hardware acceleration.
[0106] Source address extraction instructions:
[0107] In the RDMA protocol, the source address is usually a 16-byte (IPv6 address) or 4-byte (IPv4 address) field. We need a customized instruction to quickly extract the address information.
[0108] For example, assuming the source address is stored somewhere in the packet buffer, the instruction set could include the following:
[0109] ld64 r1,[r0] / / Starting from address r0, load 64 bits (8 bytes) of data into register r1;
[0110] Here, r0 stores the starting memory location of the source address. With the ld64 instruction, we can load 64 bits of data into register r1 at once. If the complete source address needs to be extracted, it can be loaded multiple times. For example, if it is an IPv6 address, two ld64 operations may be required to extract the complete 16-byte source address.
[0111] ld128: If the source address is IPv6, a longer load instruction (such as ld128) can be used to load 128 bits of data at a time.
[0112] Target address extraction instruction:
[0113] The extraction of the target address is similar to the source address. Assuming that the target address is also stored somewhere in the data packet, customized instructions can also improve efficiency by optimizing memory loading:
[0114] ld64 r2,[r0+16] / / Load 64-bit data from address r0 offset by 16 bytes to r2 (target address);
[0115] This instruction will load the first 8 bytes of the target address starting at 16 bytes offset from the source address. If the target address is 4 bytes (IPv4 address), only 32 bits of data need to be loaded and the instruction can be optimized similarly.
[0116] CRC check instruction design:
[0117] CRC check is an indispensable part of data packet verification, especially in network protocols, where it is crucial to ensure data integrity. In order to speed up CRC check, special hardware acceleration instructions can be designed.
[0118] CRC-32 calculation:
[0119] CRC-32 is a commonly used 32-bit checksum algorithm, which is widely used in network transmission protocols. In order to speed up CRC calculation, hardware-accelerated CRC calculation instructions can be designed:
[0120]
[0121] Where: D = d0, d1, ..., d n―1 Indicates the content of the data packet and calculates CRC bit by bit; Indicates bitwise XOR; InitialCRC indicates the initial value of CRC;
[0122] In the case of hardware acceleration, the CRC calculation can be quickly implemented by table lookup. We assume that a hardware CRC module is designed to handle CRC-32 calculations and is called directly by instructions.
[0123] Instruction format: crc32 r2,r1,0xFFFFFFFF / / Calculate the CRC32 check of the data in the r1 register, the initial value is 0xFFFFFFFF;
[0124] The above directive works as follows:
[0125] r1 stores the data to be checked (possibly the payload or header of a data packet).
[0126] 0xFFFFFFFF is the initial value of the CRC check, which is usually all 1s.
[0127] r2 stores the final CRC check value.
[0128] The hardware acceleration module will perform CRC calculation based on the table lookup method, which greatly improves the speed of CRC calculation. After each CRC calculation, the result can be stored in a register for subsequent processing.
[0129] CRC check instruction design ideas:
[0130] Calculate the check value: Perform CRC calculation directly in hardware and use a fast table lookup method to speed up the CRC process.
[0131] Data processing pipeline: The data stream is transmitted through the pipeline so that CRC checks can be calculated in parallel for multiple data blocks, thereby improving throughput.
[0132] Byte order and memory access optimization:
[0133] In network protocol processing, endianness issues are a common challenge. To optimize endianness conversion and memory access, we can design specific endianness conversion instructions.
[0134] Byte order conversion instructions:
[0135] Network protocols usually use big-endian byte order, but many processors use little-endian byte order. To handle the byte order conversion problem, we designed a hardware-supported byte order conversion instruction:
[0136] bswap r1,r2 / / Reverse the byte order of the data in r2 and store the result in r1;
[0137] This instruction will reverse the byte order of the data stored in r2 and store it in r1. For example, convert 32-bit data from little-endian to big-endian.
[0138] Step 3: The parsed and verified data packets are transmitted to the flow module according to the target information. The local computing node determines the forwarding path of the data packet based on the routing table and topology information;
[0139] In step 2, the header of the data packet has been extracted and CRC checked by the hardware acceleration module. The target information of the data packet (such as the target IP address or the target MAC address) will be extracted and stored in the register or memory. At this point, the target information of the data packet will be used as the input of the routing decision.
[0140] Assume the target address is stored in register r1 and has been resolved. The target information may include:
[0141] Destination IP address (IPv4 or IPv6);
[0142] Destination MAC address (if forwarding at the data link layer);
[0143] Service type (e.g. QoS, priority), etc.
[0144] Routing table lookup and path calculation:
[0145] In order to achieve efficient forwarding of data packets, the system first needs to search the local routing table. The routing table stores the mapping relationship between each destination address and the next hop route. The routing table search is usually implemented through the Longest Prefix Match (LPM) algorithm.
[0146] The process of routing table lookup is as follows:
[0147] Match the destination IP address with each entry in the routing table, selecting the entry that matches the longest prefix.
[0148] Get the next hop information of the matching entry (such as the IP address of the next hop router, outbound interface, etc.).
[0149] The forwarding path of the data packet is determined based on the matching results.
[0150] To speed up this process, a customized RISC-V instruction set can be used for hardware acceleration.
[0151] Custom RISC-V instructions to speed up routing table lookups
[0152] In the RISC-V architecture, in order to speed up the routing table lookup process, specific instructions or hardware acceleration modules can be designed to optimize LPM lookups. For example, a hardware Trie tree or hash table can be used to store routing information to achieve fast lookups.
[0153] Custom directives:
[0154] LPM lookup instructions: Assuming we use the hardware-supported LPM lookup module, we can use custom instructions to speed up the prefix matching process:
[0155] lpm_lookup r2,r1,routing_table / / Look up the target address r1 in routing_table and store the result in r2;
[0156] This instruction will perform the longest prefix match in routing_table through the hardware module, quickly obtain the routing information corresponding to the target address (such as the IP address of the next-hop router, outbound interface, etc.), and store the search result in the r2 register.
[0157] The routing information is stored through a hardware-supported Trie tree (prefix tree), and prefix matching (LPM) is performed in combination with customized RISC-V instructions. The essence of the Trie tree is to perform efficient searches based on prefix matching, which is suitable for the longest prefix match (LPM) algorithm in the IP routing table. We will implement efficient prefix matching based on the hardware-accelerated Trie tree structure and further optimize the LPM search process.
[0158] The basic structure of a Trie tree:
[0159] Trie tree is a data structure based on dictionary tree, each node of which represents a possible prefix and represents different paths through branches. In order to adapt to the longest prefix match (LPM) of IP addresses, each node of Trie tree can save a binary bit, representing 0 or 1 of a certain bit.
[0160] For IPv4 addresses, 32-bit IP addresses can be stored in a Trie tree. Each node represents a binary bit (0 or 1) of the IP address. The binary value of 0 and 1 on the path from the root node to a leaf node represents an IP prefix.
[0161] In hardware implementation, the Trie tree can be represented as a fixed-size storage array, where each node corresponds to a position in the array, and the branches (child nodes) of the node can be implemented through hardware pointers or indexes. Each node contains the following information:
[0162] Branch pointer (i.e., pointer to the left child node or the right child node, indicating 0 or 1);
[0163] The matching result, that is, whether the node is the endpoint of a certain prefix.
[0164] Routing information, such as the next-hop address, interface information, etc.
[0165] Since the Trie tree is a data structure arranged in prefix levels, its query process is to sequentially search for each binary bit. Hardware can speed up this search process by parallelizing instructions and accelerating storage mechanisms.
[0166] Customized RISC-V instructions and hardware acceleration modules:
[0167] Assuming that the custom instruction **lpm_lookup** supported by the RISC-V architecture can directly call the hardware-accelerated Trie tree module and accelerate LPM search through instructions, the following is a technical solution combining hardware and instructions:
[0168] Custom instruction: lpm_lookup
[0169] Through the lpm_lookup instruction provided by the hardware acceleration module, the RISC-V processor can directly interact with the hardware Trie tree to perform prefix matching lookups;
[0170] Instruction format: lpm_lookup r2,r1,routing_trie / / r1: input target IP address, routing_trie: Trie tree structure storing routing table, r2: output matching routing information;
[0171] The execution flow of this instruction is as follows:
[0172] The input destination IP address (32 bits for IPv4, for example) is stored in register r1.
[0173] The hardware Trie tree module reads the IP address of r1 and searches bit by bit starting from the root node of the Trie tree.
[0174] After finding the longest matching prefix of the target address, the relevant routing information (such as the next hop address, outgoing interface, etc.) is output and stored in register r2.
[0175] At the hardware level, the process of searching the Trie tree can be parallelized to increase the search speed. For example, multiple binary bits can be scanned simultaneously using multi-channel parallel scanning technology.
[0176] Hardware module design:
[0177] The hardware Trie tree module can be composed of the following key parts:
[0178] Address input interface: receives the target IP address and parses it into a series of binary bits.
[0179] Trie tree storage unit: stores the node information of the Trie tree through hardware, including the child node pointer of each node (branch of 0 or 1).
[0180] Parallel search engine: Uses parallel scanning technology to speed up the search process and accelerates the longest prefix match through multi-bit parallel search.
[0181] Output interface: Returns the matched routing information (such as the next hop address, outbound interface, service type, etc.) through the r2 register.
[0182] Query process:
[0183] During the query process, the hardware performs an LPM lookup through the following steps:
[0184] The input target IP address (e.g., 32 bits of IPv4) is passed bit by bit to the Trie tree.
[0185] Starting from the root node of the Trie tree, match each bit of the target IP address (from the highest bit to the lowest bit):
[0186] If the current bit of the target IP is 0, continue searching along the left subtree (pointer is 0);
[0187] If the current bit of the target IP is 1, continue searching along the right subtree (pointer is 1).
[0188] Whenever a node is queried, the routing information of the node is recorded (if the node is a valid prefix endpoint).
[0189] After querying all bits, the routing information corresponding to the longest matching prefix is returned.
[0190] Traditional LPM searches are performed bit by bit, but in hardware, multiple bits can be searched at the same time using parallel technology, greatly speeding up the search. For example, multiple bits can be searched at the same time in each layer of the Trie tree search; supporting the search of the Trie tree with dedicated hardware circuits can eliminate multiple data access delays in software execution and greatly improve performance. The hardware acceleration module can provide low-latency searches and support high-concurrency requests; the hardware-implemented Trie tree has a fixed storage layout, avoiding dynamic memory allocation and complex data structure operations in traditional software implementations.
[0191] Step 4: Based on the determined forwarding path, the data packet is forwarded to the target processing unit or network interface, wherein if the forwarding target is a local computing node, the data packet is stored in the cache; if the forwarding target is a remote computing node, the data packet continues to be forwarded through the network interface;
[0192] Step 5: During packet forwarding, the hardware scheduling module dynamically adjusts traffic distribution according to the current load conditions and controls the processing load of each processing unit to avoid overload.
[0193] The step 5 comprises the following steps:
[0194] Load detection and real-time monitoring: The hardware scheduling module periodically monitors the current load of each processing unit to obtain the real-time load data of each processing unit, including the current queue length, CPU usage, memory usage, and network bandwidth usage;
[0195] Threshold judgment: If the load of a processing unit exceeds the preset upper threshold, the processing unit is considered to be overloaded and tasks need to be reallocated;
[0196] If the load is below a preset lower threshold, the processing unit is considered to be in an idle state;
[0197] Task migration: Migrate data packets from overloaded processing units to idle processing units through hash algorithms;
[0198] Traffic segmentation: The hardware scheduling module distributes traffic according to the load of each processing unit to achieve load balancing:
[0199]
[0200] Where: P in Indicates the number of packets to be forwarded; P out,i represents the number of packets assigned to the i-th processing unit; L i represents the load value of the i-th processing unit; L max represents the maximum load value; N represents the number of processing units.
[0201] Preferably, the steps of implementing task migration through the hash algorithm are as follows:
[0202] Load-aware value calculation: A load-aware value is calculated for each processing unit:
[0203] LAHV i =w1×CPU i +w2×QueueLength i +w3×Memory i +w4×Bandwidth i ;
[0204] Where: LAHV i represents the load perception value of the i-th processing unit; CPU i Indicates the CPU usage of the i-th processing unit; QueueLength i Indicates the queue length of the i-th processing unit; Memory i Indicates the memory usage of the i-th processing unit; Bandwidth i represents the bandwidth utilization rate of the i-th processing unit; w1, w2, w3, and w4 represent weight coefficients respectively;
[0205] Calculate hash value: Apply hash function to the load-aware value of each processing unit and map the load-aware value to a position in the hash ring:
[0206] Hash i =HashFunction(LAHV i );
[0207] Where: Hash i represents the hash value of the i-th processing unit;
[0208] Ring construction: Map each processing unit to a position in the hash ring;
[0209] Calculate the hash value of the data packet: Calculate the hash value using an identifier of the data packet:
[0210] TaskHash=HashFunction(PacketIdentifier);
[0211] Where: TaskHash represents the hash value of the data packet; PacketIdentifier represents the identifier of the data packet;
[0212] Search for a processing unit on the hash ring: Start from the hash value on the hash ring and search for the next processing unit in a clockwise direction. According to the size of the hash value, find the processing unit with the lightest load and matching hash value in the hash ring to migrate the task.
[0213] After the data packet is migrated to the target processing unit, the load perception value of the target processing unit needs to be updated. By periodically calculating the load perception value of each processing unit, the node position on the hash ring is dynamically adjusted.
[0214] By monitoring the load of each processing unit in real time and dynamically adjusting the flow distribution of data packets, the system can still run stably under high load conditions. By adopting load balancing algorithms, task migration and flow splitting strategies, overload can be effectively avoided and the performance and stability of the system can be maintained.
[0215] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. An efficient parsing and forwarding method for RDMA protocol based on RISC-V architecture, characterized in that: The following steps are involved: Step 1: The data packet obtained by the edge node reaches the local computing node through the RDMA network interface; Step 2: The data packet is parsed and verified by the hardware acceleration module for the RDMA protocol. The header information is extracted and CRC verification is performed through the customized RISC-V instruction set. Step 3: The parsed and verified data packets are transmitted to the flow module according to the target information. The local computing node determines the forwarding path of the data packet based on the routing table and topology information; Step 4: Based on the determined forwarding path, the data packet is forwarded to the target processing unit or network interface, wherein if the forwarding target is a local computing node, the data packet is stored in the cache; if the forwarding target is a remote computing node, the data packet continues to be forwarded through the network interface; Step 5: During packet forwarding, the hardware scheduling module dynamically adjusts traffic distribution according to the current load conditions and controls the processing load of each processing unit to avoid overload.
2. According to the RISC-V architecture-based RDMA protocol efficient parsing and forwarding method according to claim 1, it is characterized in that: The format of the data packet includes a protocol header, a data payload and a tail; The protocol header includes a source address, a destination address, a protocol type, and a data length; The data payload includes data content to be transmitted; The tail contains CRC check information of the data packet.
3. The RDMA protocol efficient parsing and forwarding method based on RISC-V architecture according to claim 2 is characterized in that: The step 2 comprises the following steps: Data packets are stored in DMA buffer: After the hardware interface RDMA NIC receives the data packet, the data is stored in a temporary DMA buffer; Customized RISC-V instruction set and protocol parsing: Parse the source address, destination address, protocol type, and data length of the data packet based on the customized RISC-V instruction set; Based on the RISC-V instruction set, a set of instructions are defined to quickly calculate the CRC check value: Where: D = d0, d1, ..., d n―1 Indicates the content of the data packet and calculates CRC bit by bit; Indicates bitwise XOR; InitialCRC indicates the initial value of CRC; By initializing the CRC value to 0xFFFFFFFF, an XOR operation is performed on each byte of each data packet, and the data is shifted and updated by table lookup. After the calculation is completed, the CRC value is reversed and compared with the check value in the CRC check information. If they are consistent, the check is passed, and if they are inconsistent, the check fails. Information extraction and error handling: Extract the protocol header fields and CRC check results. If the check fails, the data packet is discarded.
4. The RDMA protocol efficient parsing and forwarding method based on RISC-V architecture according to claim 1 is characterized in that: The step 3 comprises the following steps: Routing table lookup and path calculation: Match the target IP address with each entry in the routing table, select the entry that matches the longest prefix, obtain the next hop information of the matching entry, and determine the forwarding path of the data packet based on the matching result; The routing table information stores routing information through a hardware Trie tree to achieve fast search.
5. The RDMA protocol efficient parsing and forwarding method based on RISC-V architecture according to claim 4 is characterized in that: The steps to determine the forwarding path based on the hardware Trie tree are as follows: Pass the input target IP address bit by bit to the Trie tree; Starting from the root node of the Trie tree, match each bit of the target IP address: If the current bit of the target IP is 0, continue searching along the left subtree; If the current value of the target IP is 1, continue searching along the right subtree; Whenever a node is queried, the routing information of the node is recorded; After querying all bits, the routing information corresponding to the longest matching prefix is returned.
6. The RDMA protocol efficient parsing and forwarding method based on RISC-V architecture according to claim 1 is characterized in that: The step 5 comprises the following steps: Load detection and real-time monitoring: The hardware scheduling module periodically monitors the current load of each processing unit to obtain the real-time load data of each processing unit, including the current queue length, CPU usage, memory usage, and network bandwidth usage; Threshold judgment: If the load of a processing unit exceeds the preset upper threshold, the processing unit is considered to be overloaded and tasks need to be reallocated; If the load is below a preset lower threshold, the processing unit is considered to be in an idle state; Task migration: Migrate data packets from overloaded processing units to idle processing units through hash algorithms; Traffic segmentation: The hardware scheduling module distributes traffic according to the load of each processing unit to achieve load balancing: Where: P in Indicates the number of packets to be forwarded; P out,i represents the number of packets assigned to the i-th processing unit; L i represents the load value of the i-th processing unit; L max represents the maximum load value; N represents the number of processing units.
7. The RDMA protocol efficient parsing and forwarding method based on RISC-V architecture according to claim 6 is characterized in that: The steps to implement task migration through the hash algorithm are as follows: Load-aware value calculation: A load-aware value is calculated for each processing unit: LAV i =w1×CPU i +w2×QueueLength i +w3×Memory i +w4×Bandwidth i ; Where: LAHV i represents the load perception value of the i-th processing unit; CPU i Indicates the CPU usage of the i-th processing unit; QueueLength i Indicates the queue length of the i-th processing unit; Memory i Indicates the memory usage of the i-th processing unit; Bandwidth i represents the bandwidth utilization rate of the i-th processing unit; w1, w2, w3, and w4 represent weight coefficients respectively; Calculate hash value: Apply hash function to the load-aware value of each processing unit and map the load-aware value to a position in the hash ring: Hash i =HashFunction(LAHV i ); Where: Hash i represents the hash value of the i-th processing unit; Ring construction: Map each processing unit to a position in the hash ring; Calculate the hash value of the data packet: Calculate the hash value using an identifier of the data packet: TaskHash=HashFunction(PacketIdentifier); Where: TaskHash represents the hash value of the data packet; PacketIdentifier represents the identifier of the data packet; Search for a processing unit on the hash ring: Start from the hash value on the hash ring and search for the next processing unit in a clockwise direction. According to the size of the hash value, find the processing unit with the lightest load and matching hash value in the hash ring to migrate the task. After the data packet is migrated to the target processing unit, the load perception value of the target processing unit needs to be updated. By periodically calculating the load perception value of each processing unit, the node position on the hash ring is dynamically adjusted.
Citation Information
Patent Citations
Method and device for data processing
CN105812258A
Distributed-type metadata management method with dynamic equilibrium load
CN106161120A
Edge computing hardware architecture based on RISC-V
CN110007961A
System and method for facilitating efficient load balancing in a network interface controller (NIC)
CN113692725A
FPGA acceleration board card and market data processing method thereof
CN114328348A
Cited By
Distributed adapter dynamic management method based on equipment allocation list
CN121262232A