Infiniband protocol conversion method and system
The Infiniband protocol conversion method implemented by FPGA uses hop count counters and pointers for routing decisions to convert Infiniband packets into Ethernet packets. This solves the network latency and bandwidth loss problems in high-performance computing and large-scale data transmission in data centers, and realizes efficient RDMA communication across data centers.
Patent Information
- Application Number
- CN202511036756.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-07-28
AI Technical Summary
Traditional TCP/IP technology cannot meet the needs of high-performance computing and large-scale data transmission in data centers. RDMA has bottlenecks in long-distance transmission, and there are compatibility issues between Infiniband and Ethernet protocol conversion, resulting in network latency and bandwidth loss.
An FPGA-based Infiniband protocol conversion method is adopted, which uses hop count counters and hop count pointers for routing decisions to convert Infiniband packets into Ethernet packets. A hardware-level protocol conversion engine is used to achieve lossless conversion, and combined with a hardware-accelerated link layer subnet management interface and flow control mechanism, it supports long-distance RDMA communication across data centers.
It achieves lossless conversion between Infiniband and Ethernet protocols, supports long-distance RDMA communication across data centers, meets the requirements for high-performance data transmission, solves the bottlenecks of traditional TCP/IP technology and the challenges of long-distance RDMA transmission, and is compatible with existing infrastructure.
Smart Images

Figure CN120528991B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer network communication, in particular to an Infiniband protocol conversion method and system. BACKGROUND
[0002] At present, the demand for high-performance and low-latency networks continues to grow in data center applications, especially when processing compute and data-intensive workloads such as artificial intelligence and scientific simulations. Large-scale artificial intelligence model training requires synchronization of tens of billions of parameter models across data centers, and high-energy physics experiments generate massive amounts of data that need to be transmitted in real time to supercomputing centers for analysis. Traditional TCP / IP technology has been unable to meet the requirements of data centers in high-performance computing and large-scale data transmission due to its inherent network latency, throughput limitations, and high CPU resource consumption.
[0003] In the traditional network communication architecture, data transmission needs to go through multiple layers of protocol stack processing, resulting in unnecessary data replication between the operating systems of the sender and the receiver, reducing transmission efficiency. In addition, in high-bandwidth networks, the TCP / IP protocol stack processing protocol can cause high CPU usage, limiting the CPU cycles available for application workloads.
[0004] To solve these problems, Remote Direct Memory Access (RDMA) technology has emerged, initially developed for high-performance computing to enable high-speed network interconnection between computers. However, there are core bottlenecks in large-scale deployment of RDMA. The specific reasons are:
[0005] (1) The long-distance transmission problem of RDMA stems from its strong dependence on lossless networks due to its zero-copy architecture, while the physical delay (millisecond level) and dynamic congestion in cross-domain scenarios can trigger protocol timeouts and retransmission storms, resulting in a cliff-like drop in throughput.
[0006] (2) Compatibility issues with existing infrastructure, which is essentially a conflict between the innovative design of RDMA (kernel bypass / hardware flow control) and traditional network devices (old switches / firewalls / operations systems), typically manifested as flow control failure and security policy interception.
[0007] (3) Infiniband and Ethernet protocol conversion difficulties, focusing on three major gaps: protocol semantic differences (connection-oriented vs. connectionless), flow control mechanism fragmentation, and address system mismatch, causing more than 30% bandwidth loss and state synchronization risks.
[0008] These three factors together constitute the core bottleneck of large-scale deployment of RDMA. SUMMARY
[0009] In order to solve the above problems, the application provides an Infiniband protocol conversion method and system, which realizes lossless conversion of InfiniBand and Ethernet protocols and supports long-distance RDMA communication across data centers.
[0010] In order to achieve the above object, the application adopts the following technical scheme:
[0011] In the first aspect, the application provides an Infiniband protocol conversion method, which comprises the following steps:
[0012] receiving an SMP data packet and initializing a hop count counter and a hop pointer;
[0013] making a routing decision based on the values of the hop count counter and the hop pointer to forward the SMP data packet;
[0014] In the first aspect, the application provides an Infiniband protocol conversion method, which comprises the following steps:
[0015] sending the SMP data packet to an Ethernet transceiver by the IB link layer subnet management agent node to convert the SMP data packet into an Ethernet data packet for forwarding.
[0016] As an optional implementation, the forwarding process from the starting point of the forwarding path to the last hop before the target point comprises:
[0017] When the hop count counter is not zero and the hop pointer is zero, the SMP data packet is pressed into the initial path array by the hop pointer, the initial path array is indexed by the hop pointer to obtain the sending port number, and then the initial path array forwards the SMP data packet to the sending port, without involving the record of the return path array.
[0018] As an optional implementation, the forwarding process from the starting point of the forwarding path to the last hop before the target point comprises:
[0019] When the hop count counter Hop Count is not zero and 1≤ Hop Count ≤ Hop Pointer; and Hop Count==Hop Pointer and Hop Pointer is not zero;
[0020] First, the sending port number is obtained by indexing the initial path array with the hop pointer; then, the sending port number of the current SMP packet is stored in the return path array by the switch; finally, the SMP packet is pushed into the initial path array by the hop pointer, and then the SMP packet is forwarded to the sending port by the initial path array.
[0021] As an alternative embodiment, the Infiniband protocol conversion method further comprises: in the return path, the initialization process of the IB link layer subnet manager agent node comprises:
[0022] When sending data from the outside to the subnet manager, the SMP packet is generated, and the hop pointer, the initial path array, the hop count and the return path array are copied and sent to the IB link layer subnet manager agent node, and then the obtained packet is transmitted to the subnet management interface;
[0023] The routing decision of the subnet management interface comprises: if the obtained packet is from the RD routing part, further processing; otherwise, direct output.
[0024] As an alternative embodiment, in the return path, the processing process of the subnet management interface comprises:
[0025] When the hop count is not zero and the hop pointer Hop Pointer == Hop Count + 1, it indicates that the SMP packet is in a position after the end of the forwarding path, then the packet is sent to the designated sending port through the subnet management interface;
[0026] When the hop count is not zero and 2 ≤ Hop Pointer < Hop Count, it indicates that the SMP packet is in the middle hop of the forwarding path, and the hop pointer is in the valid range, then the SMP packet is sent to the sending port designated by the switch through the subnet management interface, and bypasses the initial path array Initial Path and the return path array ReturnPath;
[0027] When Hop Pointer == 1, it indicates that the SMP packet is in the first hop of the forwarding path, then it is judged whether the target link layer address is equal to the permitted link layer address; if equal, it is transmitted to the subnet manager for processing by the IB link layer subnet management interface; if not equal, it is directly output by the B link layer subnet management interface;
[0028] When Hop Pointer == 0, it indicates that the SMP packet is at the absolute starting point of the forwarding path and has not yet entered the forwarding path, then the SMP packet is directly sent to the subnet manager through the IB link layer subnet management interface.
[0029] As an alternative implementation, the Infiniband protocol conversion method further comprises:
[0030] a link control process, specifically: the receiving end allocates data buffers for each virtual channel, and periodically notifies the sending end of the current remaining buffer size through a credit value data packet; the sending end limits the amount of data sent according to the credit value, and deducts the corresponding credit for each packet sent; the receiving end triggers credit update after releasing the data buffer;
[0031] a VL buffer design, specifically: data buffers are configured for each virtual channel, a VL field is designed in the header of the RDMA data packet, when writing data, if the packet header character is detected as SDP, the received data is selectively allocated to the data buffer corresponding to the VL value according to the value in the VL field; when reading data, the data packet is extracted from the corresponding data buffer according to the value of the VL field.
[0032] In a second aspect, the application provides an Infiniband protocol conversion system, comprising:
[0033] a receiving module configured to receive an SMP data packet and initialize a hop count counter and a hop count pointer;
[0034] a forwarding module configured to make a routing decision based on the values of the hop count counter and the hop count pointer to forward the SMP data packet;
[0035] wherein, in the forwarding process from the starting point of the forwarding path to the last hop before the target point, the initial path array is indexed by the hop count pointer to obtain a sending port number, and the SMP data packet is pushed into the initial path array by the hop count pointer, and the SMP data packet is forwarded to the sending port by the initial path array, at this time the hop count counter is decremented by one and the hop count pointer is incremented by one; and when in the middle hop of the forwarding path, the sending port number is stored in the return path array until the SMP data packet is forwarded to the last hop before the target point, and the SMP data packet is sent to the IB link layer subnet management agent node;
[0036] a conversion module configured to send the SMP data packet to an Ethernet transceiver by the IB link layer subnet management agent node to convert the SMP data packet into an Ethernet data packet for forwarding.
[0037] In a third aspect, the application provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and running on the processor, when the computer instructions are run by the processor, the method of the first aspect is completed.
[0038] In a fourth aspect, the present application provides a computer readable storage medium for storing computer instructions, which, when executed by a processor, complete the method of the first aspect.
[0039] In a fifth aspect, the present application provides a computer program product comprising a computer program, which, when executed by a processor, implements the method of the first aspect.
[0040] Compared with the prior art, the present application has the following beneficial effects:
[0041] The application provides an Infiniband protocol conversion method and system based on FPGA (Field Programmable Gate Array), which makes routing decisions based on the values of hop count counters and hop count pointers to forward subnet management packets (SMP), and finally sends the SMP packets to an Ethernet transceiver by an IB link layer subnet management agent to convert the SMP packets into Ethernet packets for forwarding. The application realizes lossless conversion between InfiniBand and Ethernet protocols, supports long-distance RDMA communication across data centers, meets the high-performance data transmission requirements in the fields of artificial intelligence and scientific computing, and solves the problems of power and network fusion requirements, traditional TCP / IP technology bottlenecks, RDMA long-distance transmission difficulties and existing infrastructure compatibility.
[0042] The method of the present application also has the following technical advantages:
[0043] 1. Hardware-level protocol conversion engine: InfiniBand and Ethernet lossless protocol conversion is implemented in FPGA, LRH (Local Route Header) and BTH (Base Transport Header) of InfiniBand are mapped to Ethernet VLAN (Virtual Local Area Network) tag and custom header, while key RDMA (Remote Direct Memory Access) semantics such as atomic operation and memory window are preserved; The solution is embodied in hardware-level field mapping (LRH / BTH→VLAN tag+custom header) compression header but key opcode and flow identification are preserved; Atomic operation and memory window are implemented by FPGA proxy mechanism to pass through semantics; VL (Virtual Link) isolation buffer + dynamic credit flow control ensures lossless transmission across protocols. The solution is compatible with carrier-grade network infrastructure, without modifying the switch configuration. By hard-coding the end-to-end path in FPGA, the initialization process of traditional subnet manager is skipped.
[0044] 2. Hardware-accelerated IB link layer subnet management interface (Subnet Management Interface, SMI) and IB link layer subnet management agent (Subnet Management Agent, SMA): Pre-stored cross-data center path mapping table, link layer address (Link Interface Device, LID) and routing information can be directly configured through FPGA-based SMI / SMA components, eliminating the need for dynamic topology discovery.
[0045] 3. Long-distance cache-aware flow control design: Combining dynamic link state awareness with absolute credit mechanism to adjust flow control credit in real time. Independent buffer is allocated for each virtual lane VL to prevent cross-flow congestion propagation.
[0046] Advantages of additional aspects of the present application will be given in part in the following description, some will become apparent from the following description, or will be learned by practice of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.
[0048] Figure 1 Infiniband protocol conversion method flow chart provided for embodiment 1 of the present application;
[0049] Figure 2 Protocol conversion schematic diagram provided for embodiment 1 of the present application;
[0050] Figure 3 Routing architecture implementation schematic diagram provided for embodiment 1 of the present application;
[0051] Figure 4 SMI processing flow schematic diagram of initial path provided for embodiment 1 of the present application;
[0052] Figure 5 SMA initialization processing flow schematic diagram of return path provided for embodiment 1 of the present application;
[0053] Figure 6 SMI processing flow schematic diagram of return path provided for embodiment 1 of the present application;
[0054] Figure 7 Link control schematic diagram provided for embodiment 1 of the present application;
[0055] Figure 8 VL buffer schematic diagram provided for embodiment 1 of the present application;
[0056] Figure 9 RTT statistical index distribution diagram provided for embodiment 1 of the present application;
[0057] Figure 10 RTT cumulative distribution diagram provided for embodiment 1 of the present application;
[0058] Figure 11 RTT stability trend schematic diagram provided for embodiment 1 of the present application. DETAILED DESCRIPTION
[0059] The present application will be further described below in conjunction with the accompanying drawings and embodiments.
[0060] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0061] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments in accordance with the present application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.
[0062] The embodiments in the present application and the features in the embodiments can be combined with each other in the case of no conflict.
[0063] InfiniBand (IB) is a computer network communication standard for high-performance computing, which has extremely high throughput and extremely low latency, and is used for data interconnection between computers. InfiniBand is also used as direct or switched interconnection between servers and storage systems, and interconnection between storage systems.
[0064] The method of the present application provides a hardware solution for converting InfiniBand to Ethernet protocol based on FPGA, realizing long-distance RDMA transmission. It introduces a dedicated switch at the server end for sending and receiving. By connecting the InfiniBand network to the switch, long-distance transmission of the InfiniBand network can be realized without modifying the existing infrastructure. It is a technology for realizing lossless conversion between InfiniBand and Ethernet protocol through hardware acceleration, which is used to support long-distance RDMA communication across data centers, and is suitable for scenarios such as distributed training of artificial intelligence models, scientific computing, and weather prediction, which require low-latency and high-throughput data synchronization.
[0065] Embodiment 1
[0066] The embodiment provides an Infiniband protocol conversion method, as shown in Figure 1 , comprising:
[0067] S101: receiving an SMP data packet, and initializing a hop count counter and a hop count pointer;
[0068] S102: making a routing decision based on the values of the hop count counter and the hop count pointer to forward the SMP data packet;
[0069] S103: sending the SMP data packet to an Ethernet transceiver by an IB link layer subnet management agent node to convert the SMP data packet into an Ethernet data packet for forwarding.
[0070] In step S102, specifically including: from the start of the forwarding path to the last hop before the target point, the forwarding process is indexed by the initial path array to obtain the sending port number, and the SMP data packet is pushed into the initial path array by the hop number pointer, and the SMP data packet is forwarded to the sending port by the initial path array, at this time the hop number counter is reduced by one, and the hop number pointer is increased by one; and when the intermediate hop of the forwarding path, the sending port number is stored in the return path array, until the SMP data packet is forwarded to the last hop before the target point, the SMP data packet is sent to the IB link layer subnet management agent node.
[0071] The method of the embodiment will be described in detail below.
[0072] 1. Protocol conversion mechanism.
[0073] Based on the InfiniBand network protocol and using FPGA related IP design, IB physical layer EDR (Enhanced Data Rate) and IB link layer SMA / SMI components are developed to realize the interconnection and intercommunication of the underlying IB network.
[0074] IB physical layer protocol specifies different encoding rules for different rates: current SDR (Single Data Rate), DDR (Double Data Rate) and QDR (Quad Data Rate) use 8B / 10B encoding rule, and FDR (Fourteen Data Rate) and EDR use 64B / 66B encoding rule.
[0075] To ensure rate compatibility, the IB physical layer protocol defines a rate negotiation mechanism, which is the most critical part of the physical layer protocol implementation. In the transmission process of IB data packet, a series of key steps are taken to ensure its conversion into Ethernet data packet. In the initial stage, IB data packet is managed by IB link layer SMP data packet to ensure its correct transmission in the network. Then, SMP data packet is processed by SMI, which is responsible for managing devices and components within the IB network. After these two layers of processing, SMP data packet is passed to SMA for further processing and management. Finally, SMP data packet is sent to the Ethernet transceiver to convert SMP data packet into Ethernet data packet; Ethernet data packet is segmented for further splitting into smaller data units to realize efficient transmission in Ethernet, as shown in Figure 2 The whole process ensures seamless conversion of data packets between different network protocols, thus realizing efficient data transmission.
[0076] 2. Delay optimization routing architecture.
[0077] In the IB link layer, link layer addresses are assigned by the Subnet Manager (SM) during IB subnet initialization, which differs from the pre-configuration of addresses in Ethernet. The SM is the core component for IB subnet initialization and configuration, responsible for discovering the subnet physical topology, assigning link layer addresses (LIDs) to end nodes / switches / routers, establishing paths, monitoring node changes, scanning subnets, and managing topology changes.
[0078] During subnet initialization, the primary switch manager (SM) obtains and sets switch node information via SMP packets to complete LID allocation. To enable interaction with the SM, the SMA and SMI components of the IB link layer protocol need to be implemented. The SMI provides routing functionality for SMP during subnet initialization, while the SMA, deployed in all IB devices, is responsible for monitoring, managing, and configuring nodes for the SM. During subnet initialization, the SMA communicates with the SM via the SMI to complete the initial configuration of each node.
[0079] Therefore, based on IB's unique subnet management model, this embodiment designs a low-latency optimized routing architecture.
[0080] (a) Routing architecture implementation.
[0081] like Figure 3 As shown, the implementation of the routing architecture relies on modifying and comparing three key identifiers: the Initial Path array, the Return Path array, and the HopPointer, which indicates the hop count pointer for forwarding SMP packets. At the same time, the Hop Counter is used to record the path value, reflecting the path length between the destination address and the source address.
[0082] Based on the comparison process of architecture direction and identifier modification, four key processing nodes can be identified: First, when the SM sends data outward, two key processing nodes are distinguished: the SM initialization processing node and the SMI processing node. Second, on the return path, two more key processing nodes are identified: the SMA initialization processing node and the SMI processing node.
[0083] Since the FPGA does not actively send subnet management data packets, the SM initialization processing node can be omitted. After completing a series of subnet management initializations, the FPGA is recognized by the main SM as a dual-port, one-to-one forwarding switch node.
[0084] (b) SMI processing node of the initial path.
[0085] When the FPGA receives the SMP packet (D = 0), first compare the value of the hop pointer Hop Pointer with the hop counter Hop Count, and perform subsequent processing steps according to the comparison result, as shown in Figure 4 . If the packet enters a branch flow not covered in the Figure 4 , discard the packet. Wherein, D is used to indicate the direction of the SMP packet, D = 0 represents sending from the SM to the outside, or sending from the outside to the SM under the trap method (trap is a command used in script programming to handle signals, allowing a specified operation to be performed when a specific signal is received); D ≤ 1 represents sending from the outside to the SM, or sending from the SM to the outside under the trap method.
[0086] As shown in Figure 4 , the process of making a routing decision based on the value of Hop Count and the hop pointer includes:
[0087] (1) When the hop counter Hop Count!= 0 and the hop pointer Hop Pointer == 0;
[0088] Indicates: the packet is at the starting point of the forwarding path (the first hop), and Hop Count is not zero, indicating that the forwarding path has not been completed; Hop Pointer = 0 indicates that the first element of the initial path array Initial Path is used for forwarding.
[0089] Thus, the data packet is pushed into the initial path array Initial Path by the hop pointer Hop Pointer, and the initial path array Initial Path is indexed by the hop pointer Hop Pointer to obtain the sending port number, and then the data packet is forwarded to the specified sending port by the initial path array Initial Path, at this time, the return path array Return Path is not involved in the record.
[0090] Routing role: this condition processes the initialization stage of the forwarding path. For example, when the packet is sent from the source host, Hop Pointer is initialized to 0 and Hop Count is the total number of hops. This ensures that the packet can correctly enter the first hop of the forwarding path, and the return path is not recorded, because the starting point usually does not need reverse path information.
[0091] (2) Hop Count!= 0 and 1 ≤ Hop Count ≤ Hop Pointer;
[0092] Indicates that the packet is in the middle hop of the forwarding path (not the start point and not the end point), and the remaining hop count (HopCount) is in the valid range (at least 1 hop, and not more than the current pointer position). 1≤ Hop Count≤ Hop Pointer ensures that the packet has not completed forwarding, and Hop Pointer has been advanced to a sufficient position (1≤ Hop Pointer), allowing the return path to be recorded.
[0093] Then, first, the sending port number is obtained by indexing the initial path array Initial Path with the hop count pointer Hop Pointer; then, the sending port number of the current data packet to be sent by the switch (i.e., the sending port number obtained by indexing the initial path array Initial Path) is stored in the return path array Return Path; finally, the data packet is pushed into the initial path array Initial Path by the hop count pointer Hop Pointer, and the data packet is forwarded to the sending port by the initial path array Initial Path.
[0094] Routing role: This condition handles most intermediate forwarding hops. The return path array Return Path is recorded to build a reverse path, for example, for fault recovery, response packet routing, or network diagnosis. The constraints of Hop Count and Hop Pointer (1≤ Hop Count ≤ Hop Pointer) ensure that only the middle of the forwarding path is recorded, avoiding repeated recording at the start point or end point.
[0095] (3) Hop Count == Hop Pointer and Hop Pointer!= 0;
[0096] Indicates that the packet is in the middle or near the end of the forwarding path, and the remaining hop count is equal to the current pointer position (Hop Count == Hop Pointer). Hop Pointer!= 0 indicates that this is not the start point. This condition overlaps with condition (2) (when Hop Count == Hop Pointer, naturally satisfies 1≤ Hop Count ≤ Hop Pointer), but is listed separately to emphasize the specific state.
[0097] So, first, the sending port number is obtained by indexing the initial path array Initial Path with the hop pointer Hop Pointer; then, the sending port number of the current packet to be sent (i.e., the sending port number obtained by indexing the initial path array Initial Path) is stored in the return path array Return Path by the switch; finally, the packet is pushed into the initial path array Initial Path by the hop pointer Hop Pointer, and the packet is forwarded to the sending port by the initial path array Initial Path.
[0098] Routing role: This condition strengthens the consistency of path recording and forwarding. When Hop Count == HopPointer, it means that the packet may be close to the path endpoint (for example, when Hop Pointer is large), but the forwarding still needs to continue. Recording the Return Path ensures the completeness of the path information and prepares for possible endpoint processing. The redundant design with condition (2) is used for specific optimization of code implementation.
[0099] (4) Hop Count+1==Hop Pointer;
[0100] Meaning: The packet has reached the penultimate hop or equivalent position of the forwarding path, i.e., it will complete the forwarding. HopCount + 1 == Hop Pointer means that the remaining hop count (Hop Count) plus one equals the current pointer position, which usually corresponds to the endpoint processing point of the path.
[0101] So, the packet does not perform regular forwarding, but is directly sent to the IB link layer subnet management agent (SMA) node for subsequent processing.
[0102] Routing role: This condition identifies the termination state of the forwarding path. The processing includes: verifying path integrity (using the initial path array Initial Path and the return path array Return Path), generating a response packet (using the return path array Return Path to build a reverse path), resource release or statistical collection. For example, when the packet reaches the last hop before the target, this condition triggers to avoid unnecessary forwarding.
[0103] (5) All other cases discard the packet.
[0104] In IB protocol, the transit switch is required not to modify the packet header (to maintain end-to-end semantics), and only the target node needs to modify the LRH / BTH header through SMA to adapt to Ethernet. The trigger condition Hop count+1 == Hop pointer accurately locates the target device, avoiding invalid calls to high-delay SMA for transit packets; Hop Pointer is used as a pointer to index the current byte number in the Initial Path / Return Path field; Hop Counter is used to indicate the number of valid bytes in the Initial Path / Return Path field. The entire process updates Hop Count and Hop Pointer, and guides the forwarding of data packets through the initial path and return path.
[0105] (c) The SMA initialization processing node of the return path.
[0106] The switch based on FPGA can also be a routing target node during subnet management initialization, so it needs to generate a response packet. The SMA initialization processing process of the return path is as shown in Figure 5
[0107] (1) Condition judgment; when D≤1, that is, the reply to a read / write request or stop repeating trap, SMP packet is generated.
[0108] Indicates the meaning: SMP packet (D≤1) is processed by SMA / SMI exclusively, and does not occupy data forwarding resources. RD (Route Distinguisher, route part) only affects SMP packet, ensuring that data forwarding is not disturbed by policy changes.
[0109] (2) Copy the key attributes such as Hop Pointer, Initial Path array, Hop Counter, and Return Path array to generate data packets for SMA, and then transmit the data packets to SMI.
[0110] Indicates the meaning: Copy the attributes such as Initial Path / Hop Pointer, which supports the back-and-forth of management packets in complex paths (such as SM→external→response return). Key attribute copying solves the context loss problem in stateless routing.
[0111] (3) SMI routing decision: if the data packet is from the RD route part, further processing; otherwise, directly output.
[0112] (d) The SMI processing node of the return path.
[0113] Similar to the SMI processing of the initial path, the FPGA switch needs to participate in the SMI initialization process of the return path. When D=1 is encountered, the value of the hop count pointer is compared with the value of the hop count counter, and the next operation is determined based on the comparison result, as shown in Figure 6 The specific rules are as follows:
[0114] (1) Hop Count!= 0 and Hop Pointer == Hop Count + 1;
[0115] Indicates the meaning: The data packet is in a position after the end of the forwarding path (Hop Count is greater than the remaining hop count HopCount by 1). This usually means that the data packet has completed the normal forwarding path, but needs additional processing.
[0116] Then, the data packet is sent to the specified sending port through SMI (without going through the normal path array).
[0117] Routing role: Handle special operations after the end of the forwarding path (for example: send an acknowledgement packet, generate a management response).
[0118] (2) Hop Count!= 0 and 2 <= Hop Pointer < Hop Count;
[0119] Indicates the meaning: The data packet is in the middle of the forwarding path, and the hop count pointer is in the valid range (from the second hop to the last hop).
[0120] Then, the data packet is sent to the sending port specified by the switch through SMI (bypassing the Initial Path / Return Path mechanism).
[0121] Routing role: Support dynamic path decision (non-predefined path), and the forwarding port is determined by the switch in real time.
[0122] (3) Hop Pointer == 1;
[0123] Indicates the meaning: Hop Pointer == 1 corresponds to the first hop after the start of the forwarding path, i.e. the data packet is in the first hop of the forwarding path.
[0124] Then, check whether the target link layer address DLID (DLID is a kind of data link layer address, mainly used for communication identification between network devices) is equal to the permissive link layer address (Permissive LID); if equal, it is transmitted to the SM by SMI; if not equal, the data packet is directly output by SMI (normal forwarding).
[0125] Routing role: permission control and security management: Permissive LID identifies high-privilege targets (such as management nodes), and such packets need to be processed by SM. Non-privileged targets are normally forwarded (possibly using predefined paths).
[0126] (4) Hop Pointer == 0;
[0127] Indicates that the data packet is at the absolute starting point of the forwarding path and has not yet entered the forwarding path.
[0128] Then, send the data packet directly to the SM through SMI.
[0129] Routing role: handle locally generated management packets, such as control messages initiated by the switch itself.
[0130] (5) Other cases: discard the data packet.
[0131] 3. Flow control mechanism of IB network protocol.
[0132] The flow control mechanism of IB network protocol aims to prevent data packet loss due to overflow of link receiving end buffer. Its design is based on the concept of "absolute credit", and the sender maintains "sending permission credit" in real time, and the receiver updates this credit through flow control packets regularly. IB network protocol specifies that the maximum buffer size represented by the "absolute credit" field in the flow control packet is 128 KB. As a key component of network communication technology, the link layer flow control mechanism ensures the reliability and stability of data transmission by preventing buffer overflow and data packet loss due to limited receiving buffer capacity.
[0133] Link control process.
[0134] As shown in Figure 7 , where FCCL (Flow Control Credit Limit) is the upper limit of flow control credit, FCTBS (Flow Control Token Bucket Size) is the size of flow control token bucket, and ABR (Available Bit Rate) is the effective bit rate. The core of the link control mechanism lies in the specific information carried in the flow control packet. With these information, the sender monitors and calculates the buffer occupancy of the receiver in real time, and decides whether to continue sending data packets. This mechanism ensures the dynamic balance of data transmission and effectively prevents data packet loss due to buffer overflow.
[0135] In addition, the receiver sends flow control packets to the sender in real time according to the buffer occupancy. These packets contain the current buffer state information of the receiver, allowing the sender to adjust the data transmission strategy to match the processing capacity of the receiver.
[0136] Through this mechanism, link flow control realizes dynamic coordination between the sending end and the receiving end, ensuring the efficiency and reliability of data transmission.
[0137] Thus, the InfiniBand credit-based flow control mechanism realizes lossless transmission through real-time credit synchronization: the receiving end allocates independent data buffers for each virtual lane VL, and periodically notifies the sending end of the current remaining buffer size through a credit value data packet; the sending end strictly limits the amount of data sent according to the credit value, deducts the corresponding credit for each packet sent, and only allows sending when the credit is sufficient, thus fundamentally preventing data buffer overflow; and stores the sending result in the data buffer; the receiving end triggers credit update immediately after releasing the data buffer.
[0138] VL buffer design.
[0139] As shown in Figure 8 In view of the independence of the virtual lane VL and the strict requirements of the link flow control mechanism, this embodiment configures a dedicated data buffer for each VL. This design ensures that the data streams of different VLs do not interfere with each other, thereby enhancing the stability of the framework and the reliability of data transmission. In addition, the header of all RDMA data packets is specially designed with a VL field to ensure accurate identification and processing of data on different virtual lanes.
[0140] Based on the established design concept, this embodiment proposes an innovative data packet processing strategy. Specifically, when writing data, if the packet header character of the data packet is identified as SDP (a standard for describing information format), the received data will be selectively allocated to a specific buffer corresponding to the VL value in the VL field. In the data reading stage, the data packet is still extracted from the corresponding buffer according to the value of the VL field.
[0141] RTT stability and link capacity evaluation.
[0142] To evaluate the reliability of long-distance links, the statistical indicators of RTT (Round-Trip Time) under different data packet sizes (10-100k) were measured, including the minimum value Min, the maximum value Max, the average value Avg, and the absolute deviation ADev. ADev was used to measure the fluctuation degree of RTT, and the lower the ADev value, the higher the network stability. In addition, a link capacity model was established to evaluate the throughput capacity of the link in long-distance transmission.
[0143] As shown in Figures 9-11As shown, the minimum RTT of Ud protocol decreases from 0.114 ms at 10 packets to 0.033 ms at 100k packets, reducing by 71.1%. The absolute deviation ADev of Ud protocol remains in the range of 0.012 to 0.015 ms, which is 20% to 40% lower than TCP, indicating that its flow control mechanism effectively suppresses fluctuations. In the simulated transmission scenario, the effective bandwidth of Ud protocol is 36.4% higher than traditional TCP. The cumulative probability of RTT in the range of 0.1 to 0.1 ms, Ud protocol reaches 95%, while TCP is only 68%, which indicates that protocol conversion significantly improves the reliability of low-delay transmission. At 0.3 ms, the cumulative probability of Ud protocol is 99.8%, while TCP is 88%, highlighting the significant optimization of tail delay. At the 100k packet scale, the ADev of direct routing is 0.015 ms, which is 46.4% lower than traditional routing.
[0144] Embodiment 2
[0145] The embodiment provides an Infiniband protocol conversion system, comprising:
[0146] A receiving module configured to receive an SMP packet and initialize a hop count counter and a hop pointer;
[0147] A forwarding module configured to make a routing decision based on the values of the hop count counter and the hop pointer to forward the SMP packet;
[0148] Wherein, in the forwarding process from the starting point of the forwarding path to the last hop before the target point, the initial path array is indexed by the hop pointer to obtain a sending port number, and the SMP packet is pushed into the initial path array by the hop pointer, and the SMP packet is forwarded to the sending port by the initial path array, at this time the hop count counter is decremented by one and the hop pointer is incremented by one; and when in the middle hop of the forwarding path, the sending port number is stored in the return path array, until the SMP packet is forwarded to the last hop before the target point, the SMP packet is sent to the IB link layer subnet management agent node;
[0149] A conversion module configured to send the SMP packet to an Ethernet transceiver by the IB link layer subnet management agent node to convert the SMP packet into an Ethernet packet for forwarding.
[0150] It should be noted that the above modules correspond to the steps described in Embodiment 1, and the above modules have the same examples and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules can be executed in a computer system such as a set of computer executable instructions as part of the system.
[0151] In more embodiments, it is also provided that:
[0152] An electronic device includes a memory and a processor and computer instructions stored on the memory and running on the processor, when the computer instructions are run by the processor, the method described in embodiment 1 is completed. For brevity, it will not be described here.
[0153] It should be understood that in the embodiments, the processor can be a central processing unit CPU, and the processor can also be other general-purpose processors, digital signal processors DSPs, application-specific integrated circuits ASICs, ready-to-program gate arrays FPGA or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0154] The memory can include read-only memory and random access memory, and provide instructions and data to the processor, and a part of the memory can also include non-volatile random access memory. For example, the memory can also store device type information.
[0155] A computer readable storage medium for storing computer instructions, when the computer instructions are executed by the processor, the method described in embodiment 1 is completed.
[0156] The method in embodiment 1 can be directly embodied as hardware processor execution completion, or executed by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory, and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0157] A computer program product includes a computer program, which is executed by the processor to realize the method described in embodiment 1.
[0158] The present application also provides at least one computer program product tangibly stored on a non-transitory computer readable storage medium. The computer program product includes computer executable instructions, such as instructions included in program modules, which are executed in devices on real or virtual processors of targets to perform processes / methods as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of the program modules can be combined or divided as needed among the program modules. Machine executable instructions for program modules can be executed within local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.
[0159] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages. The computer program code can execute entirely on a computer, a special purpose computer, or other programmable apparatus to produce the functions / acts specified in the flow diagrams and / or block diagrams. The program code can execute entirely on a computer, a special purpose computer, or other programmable apparatus, as a stand-alone software package, partly on the computer and partly on a remote computer, or entirely on the remote computer or server.
[0160] In the context of the present application, the computer program code or related data can be carried by any suitable carrier, to enable the device, apparatus or processor to perform the various processes and operations described above. Examples of carriers include signals, computer readable media, and the like. Examples of signals can include electrical, optical, radio, sound or other forms of propagated signals, such as carrier waves, infrared signals, and the like.
[0161] Those skilled in the art can realize that the units and algorithm steps of the examples described in conjunction with the present embodiments can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0162] The above describes the specific embodiments of the present application in conjunction with the accompanying drawings, but is not a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications or variations made by those skilled in the art on the basis of the technical solutions of the present application without inventive labor are still within the scope of protection of the present application.
Claims
1. An Infiniband protocol conversion method, characterized in that, include: Receive SMP data packets and initialize the hop count counter and hop count pointer; Routing decisions are made based on the values of the hop count counter and hop count pointer to forward SMP packets; During the forwarding process from the starting point of the forwarding path to the last hop before the target point, the hop count pointer indexes the initial path array to obtain the sending port number, and the hop count pointer pushes the SMP data packet into the initial path array. The initial path array then forwards the SMP data packet to the sending port. At this time, the hop count counter is decremented by one, and the hop count pointer is incremented by one. When the forwarding path is in the middle hop, the sending port number is stored in the return path array. This process continues until the SMP data packet is forwarded to the last hop before the target point, at which point the SMP data packet is sent to the IB link layer subnet management agent node. The IB link layer subnet management agent node sends SMP packets to the Ethernet transceiver to convert the SMP packets into Ethernet packets for forwarding; The forwarding process from the starting point of the forwarding path to the last hop before the target point includes: When the Hop Count counter is not zero and 1 ≤ Hop Count ≤ Hop Pointer; and when Hop Count == Hop Pointer and Hop Pointer is not zero; First, the hop count pointer indexes the initial path array to obtain the sending port number; then, the switch stores the sending port number of the SMP data packet to be sent into the return path array; finally, the SMP data packet is pushed into the initial path array by the hop count pointer, and then the initial path array forwards the SMP data packet to the sending port.
2. The Infiniband protocol conversion method as described in claim 1, characterized in that, The forwarding process from the starting point of the forwarding path to the last hop before the destination also includes: When the hop count counter is not zero and the hop count pointer is zero, the SMP data packet is pushed into the initial path array by the hop count pointer. The initial path array is then indexed by the hop count pointer to obtain the sending port number. The SMP data packet is then forwarded to the sending port by the initial path array. At this time, the record of the return path array is not involved.
3. The Infiniband protocol conversion method as described in claim 1, characterized in that, The Infiniband protocol conversion method further includes: in the return path, the initialization process of the IB link layer subnet management agent node includes: When sending data from the outward subnet manager, an SMP packet is generated, and the hop count pointer, initial path array, hop count counter and return path array are copied and sent to the IB link layer subnet management agent node. The resulting packet is then transmitted to the subnet management interface. The routing decision of the subnet management interface includes further processing if the received data packet originates from the RD routing section; otherwise, it is output directly.
4. The Infiniband protocol conversion method as described in claim 3, characterized in that, In the return path, the subnet management interface processing includes: When the hop count counter (Hop Count) is not zero and the hop pointer (Hop Pointer) is equal to the hop count plus 1, it means that the SMP packet is located after the end of the forwarding path. In this case, the packet is sent to the specified sending port through the subnet management interface. When the hop count counter Hop Count is not zero and 2 ≤ Hop Pointer < Hop Count, it indicates that the SMP data packet is in the middle hop of the forwarding path and the hop pointer is within the valid range. Then, the SMP data packet is sent to the specified sending port of the switch through the subnet management interface, bypassing the Initial Path array and the ReturnPath array; When Hop Pointer == 1; it indicates that the SMP data packet is in the first hop of the forwarding path. Then, it is judged whether the target link layer address is equal to the permitted link layer address; if they are equal, it is transmitted to the subnet manager for processing through the IB link layer subnet management interface; if they are not equal, it is directly output through the B link layer subnet management interface; When Hop Pointer == 0, it indicates that the SMP data packet is at the absolute starting point of the forwarding path and has not entered the forwarding path yet. Then, the SMP data packet is directly sent to the subnet manager through the IB link layer subnet management interface.
5. The Infiniband protocol conversion method as described in claim 1, characterized in that, The Infiniband protocol conversion method further includes: A link control process; specifically: the receiving end allocates data buffers for each virtual channel and periodically notifies the sending end of the current remaining buffer size through credit value data packets; the sending end restricts the data volume to be sent based on the credit value, and deducts the corresponding credit for each packet sent; the receiving end triggers credit update after releasing the data buffer; VL buffer design, specifically: configure data buffers for each virtual channel, design a VL field in the header of the RDMA data packet, and when writing data, when it is detected that the header character identifier is SDP, selectively allocate the received data to the data buffer corresponding to the VL value according to the value in the VL field; when reading data, extract the data packet from the corresponding data buffer according to the value of the VL field.
6. An Infiniband protocol conversion system, employing an Infiniband protocol conversion method as described in any one of claims 1-5, characterized in that, It includes: A receiving module, configured to receive the SMP data packet and initialize the hop count counter and the hop pointer; A forwarding module, configured to make a routing decision based on the values of the hop count counter and the hop pointer to forward the SMP data packet; Among them, during the forwarding process from the starting point to the last hop before the destination point of the forwarding path, the initial path array is indexed by the hop pointer to obtain the sending port number, and the SMP data packet is pushed into the initial path array by the hop pointer. The initial path array forwards the SMP data packet to the sending port. At this time, the hop count counter is decremented by one and the hop pointer is incremented by one; and when in the middle hop of the forwarding path, the sending port number is stored in the return path array until the SMP data packet is forwarded to the last hop before the destination point, and the SMP data packet is sent to the IB link layer subnet management agent node; A conversion module, configured to send the SMP data packet to the Ethernet transceiver by the IB link layer subnet management agent node to convert the SMP data packet into an Ethernet data packet for forwarding.
7. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method according to any one of claims 1-5 is completed.
8. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-5.
9. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-5.
Citation Information
Patent Citations
Method and device for forwarding IB network direct routing management message
CN117729166A