Method for forwarding a computing message in a forwarding network, forwarding node and computer storage medium
By using intranet computing identifiers for hash routing and RoCE transmission in forwarding nodes, the problem of static configuration of intranet computing message paths in distributed computing is solved, enabling flexible, secure, and efficient intranet computing flow forwarding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2023-01-19
- Publication Date
- 2026-07-31
AI Technical Summary
In distributed computing scenarios, the static configuration of forwarding paths for computing packets within the network in existing technologies is labor-intensive, lacks flexibility, requires reconfiguration when the network changes, and has insufficient transmission reliability and security.
By using hash routing based on the intranet computation identifier of the intranet computation message, the forwarding node dynamically determines the forwarding path. Combined with the RoCE transmission protocol, it improves path consistency and transmission reliability, and optimizes the forwarding table through weighted smoothing processing.
It enables flexible forwarding of computing packets within the network, improves path consistency and transmission security, reduces network maintenance costs, and enhances processing efficiency and reliability.
Smart Images

Figure CN118368230B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method for forwarding network-computed messages, a forwarding node, and a computer storage medium. Background Technology
[0002] In distributed computing scenarios, multiple computing nodes communicate with each other to exchange data, thereby enabling the execution of the same computing task across multiple nodes. Furthermore, to improve the execution efficiency of computing tasks, parts of the task can be distributed to forwarding nodes located between the computing nodes. That is, the forwarding nodes perform further computation on the data while forwarding it between different computing nodes; this technology is known as intra-network computing. In intra-network computing, a message exchanged between two computing nodes can be called an intra-network computing message, and intra-network computing messages exchanged by a large number of computing nodes for the same computing task can be collectively referred to as an intra-network computing message.
[0003] In related technologies, different computing nodes typically interact with each other through multiple layers of forwarding nodes. In this scenario, after receiving intra-network computing messages from computing nodes and performing aggregation calculations, the underlying forwarding nodes use a pre-configured static path when forwarding the aggregated messages. This ensures that different underlying forwarding nodes forward aggregated messages belonging to the same intra-network computing message to the same upper-layer forwarding node, so that the upper-layer forwarding node can continue to perform aggregation calculations on different intra-network computing messages belonging to the same intra-network computing message, thus satisfying the path consistency characteristic of intra-network computing technology.
[0004] In the above methods, when the number of computing nodes and forwarding nodes in a distributed computing scenario is large, statically configuring paths requires a significant amount of effort. Furthermore, if the network changes, the paths need to be reconfigured statically, which also requires considerable effort. Therefore, this method of forwarding packets has poor flexibility. Summary of the Invention
[0005] This application provides a method for forwarding network-computed packets, a forwarding node, and a computer storage medium, which can improve the flexibility of packet forwarding. The technical solution is as follows:
[0006] Firstly, a method for forwarding intra-network computation messages is provided. In this method, a forwarding node receives multiple intra-network computation messages; the forwarding node determines an intra-network computation identifier for each intra-network computation message among the multiple intra-network computation messages, the intra-network computation identifier indicating the intra-network computation message to which the corresponding intra-network computation message belongs; based on the intra-network computation identifier of each intra-network computation message among the multiple intra-network computation messages, the forwarding node performs aggregate computation on intra-network computation messages belonging to the same intra-network computation message among the multiple intra-network computation messages to obtain at least one aggregated message, the at least one aggregated message corresponding one-to-one with at least one intra-network computation identifier; the forwarding node forwards each aggregated message based on the intra-network computation identifier corresponding to each aggregated message in the at least one aggregated message using a hash routing algorithm.
[0007] In this embodiment, after aggregating intranet computing packets, the forwarding node does not need to forward them through a pre-configured static path. Instead, the forwarding node performs hash routing based on the intranet computing identifier of the intranet computing packets. This ensures that intranet computing packets belonging to the same intranet computing message are routed to the same next hop. In other words, while using dynamic routing, the path consistency requirement of intranet computing traffic is met, which improves the flexibility of forwarding intranet computing packets.
[0008] Furthermore, since no static path configuration is required, the routing policies of forwarding nodes are transparent to the compute nodes, enhancing the security of compute stream transmission within the network. Additionally, the elimination of static path configuration eliminates the need for a separate centralized management node, simplifying network design. Even if the network undergoes changes, network maintenance costs are relatively low.
[0009] Based on the method provided in the first aspect, in one possible implementation, the forwarding node stores an aggregation information mapping relationship, which includes at least one intranet computing identifier and a data stream identifier corresponding to the at least one intranet computing identifier.
[0010] In this scenario, the process by which a forwarding node determines the intranet computing identifier for each intranet computing packet among multiple intranet computing packets can be as follows: For the first intranet computing packet among multiple intranet computing packets, determine the data flow identifier corresponding to the first intranet computing packet to obtain the first data flow identifier. The first intranet computing packet can be any one of the multiple intranet computing packets. If the first data flow identifier exists in the aggregation information mapping relationship, then obtain the intranet computing identifier corresponding to the first data flow identifier from the aggregation information mapping relationship to obtain the first intranet computing identifier, and use the first intranet computing identifier as the intranet computing identifier of the first intranet computing packet. Correspondingly, if the first data flow identifier does not exist in the aggregation information mapping relationship, the forwarding node obtains the first data flow identifier and the first intranet computing identifier from the packet header of the first intranet computing packet; and adds the correspondence between the first intranet computing identifier and the first data flow identifier to the aggregation information mapping relationship.
[0011] To avoid excessive bandwidth consumption by intra-network computing flows, the intra-network computing identifier is carried in the first packet of an intra-network computing flow, but not in subsequent packets. In this scenario, when a forwarding node receives the first packet of each intra-network computing flow, it can obtain the data flow identifier and the intra-network computing identifier to which the flow belongs from the first packet. Then, it records the mapping relationship between the obtained data flow identifier and the intra-network computing identifier in the aggregated information mapping relationship. This facilitates the determination of the intra-network computing identifier of subsequent non-first packets from the aggregated information mapping relationship.
[0012] Based on the method provided in the first aspect, in one possible implementation, the process of the forwarding node performing aggregation calculation on intra-network computing messages belonging to the same intra-network computing message based on the intra-network computing identifier of each intra-network computing message in multiple intra-network computing messages can be as follows: binding the first intra-network computing message with the first data stream identifier in the aggregation information mapping relationship; performing aggregation calculation on the intra-network computing messages that are respectively bound to the multiple data stream identifiers corresponding to the first intra-network computing identifier in the aggregation information mapping relationship to obtain the aggregated message corresponding to the first intra-network computing identifier.
[0013] Since the aggregation information relationship stores the data flow identifier corresponding to each intranet computing identifier, aggregation calculation can be performed based on the intranet computing packets bound to the data flow identifiers corresponding to the same intranet computing identifier in the aggregation information mapping relationship, thereby improving the processing efficiency of the forwarding node.
[0014] Based on the method provided in the first aspect, in one possible implementation, the process of performing aggregation calculation on intra-network computing packets that are respectively bound to multiple data flow identifiers corresponding to the first intra-network computing identifier in the aggregation information mapping relationship can be as follows: For intra-network computing packets that are respectively bound to multiple data flow identifiers corresponding to the first intra-network computing identifier, determine the packet sequence number of each bound intra-network computing packet; based on the first packet sequence number corresponding to each data flow identifier among the multiple data flow identifiers corresponding to the first intra-network computing identifier, and the packet sequence number of each bound intra-network computing packet, determine the offset of each bound intra-network computing packet; based on the offset of each bound intra-network computing packet, perform aggregation calculation on intra-network computing packets with the same offset.
[0015] The above method enables the aggregation and calculation of intra-network computing messages with the same offset in different intra-network computing streams belonging to the same intra-network computing message.
[0016] Based on the method provided in the first aspect, in one possible implementation, the data stream identifier includes the first packet sequence number, the communication operation identifier, and the data stream length.
[0017] By using the above method, the aforementioned fields can be extended in the network-computed message to serve as data flow identifiers, thereby improving the application flexibility of the embodiments of this application.
[0018] Based on the method provided in the first aspect, in one possible implementation, the intranet computing identifier includes a communication group identifier and an intranet computing message identifier, wherein the communication group identifier indicates the communication group currently performing intranet computing, and the intranet computing message identifier indicates an intranet computing message transmitted within the communication group.
[0019] Considering that the intra-network computing message identifiers transmitted in different communication groups may be the same, in this embodiment of the application, the intra-network computing message identifier and the communication group identifier can be combined as an intra-network computing identifier to uniquely identify an intra-network computing message.
[0020] Based on the method provided in the first aspect, in one possible implementation, multiple intranet computing packets are Remote Direct Memory Access (RoCE) packets based on converged Ethernet. The forwarding node stores at least one Remote Direct Memory Access (RDMA) connection identifier, which indicates an RDMA connection between a first computing node and a second computing node. The first computing node is the computing node accessed by the forwarding node.
[0021] In this scenario, before the forwarding node determines the intranet computing identifier of each intranet computing packet among multiple intranet computing packets, for the second intranet computing packet among multiple intranet computing packets, the forwarding node obtains the RDMA connection identifier carried by the second intranet computing packet. The second intranet computing packet is any one of the multiple intranet computing packets. If the RDMA connection identifier carried by the second intranet computing packet exists among at least one stored RDMA connection identifier, then the operation of determining the corresponding intranet computing identifier of the second intranet computing packet is performed.
[0022] In this embodiment, intra-network computing flows can be transmitted via Remote Direct Memory Access over Converged Ethernet (RoCE) instead of unreliable UDP, improving the reliability of intra-network flow transmission. In this scenario, the information carried in the intra-network computing packet can be used to determine whether it is an uplink intra-network computing packet. If it is determined to be an uplink intra-network computing packet, the method provided in this embodiment is then executed, improving the processing efficiency of the forwarding node.
[0023] Based on the method provided in the first aspect, in one possible implementation, the RDMA connection identifier includes the source Internet Protocol IP address, the source port identifier, and the destination queue number QPN.
[0024] The above fields can uniquely identify an RDMA connection, improving the flexibility of the embodiments of this application.
[0025] Based on the method provided in the first aspect, in one possible implementation, the forwarding node receives an announcement message from the uplink communication link; the forwarding node obtains the target RDMA connection identifier carried in the announcement message, the announcement message being used to announce that the RDMA connection indicated by the target RDMA connection identifier is used to transmit intra-network computational messages; and the forwarding node stores the target RDMA connection identifier.
[0026] In this way, during network initialization, when the uplink computing node of a forwarding node creates an RDMA connection with other computing nodes, the forwarding node can automatically store at least one RDMA connection identifier based on the created RDMA connection.
[0027] Based on the method provided in the first aspect, in one possible implementation, the process of the forwarding node obtaining the target RDMA connection identifier carried in the announcement message can be as follows: the forwarding node obtains the type information carried in the announcement message; if the type information indicates that the announcement message is an announcement message for intra-network computed packets, then the operation of obtaining the target RDMA connection identifier carried in the announcement message is performed.
[0028] The announcement message is a control layer message. In order to distinguish the control layer message in this embodiment from other control messages, type information can be carried in the announcement message. The type information indicates whether the announcement message is an announcement message for intranet computing messages, so as to improve the processing efficiency of forwarding nodes.
[0029] Based on the method provided in the first aspect, in one possible implementation, the process by which a forwarding node forwards each aggregated packet using a hash routing algorithm based on the intranet computation identifier corresponding to each aggregated packet in at least one aggregated packet can be as follows: For a first aggregated packet in at least one aggregated packet, a first hash factor is determined, the first hash factor including the intranet computation identifier corresponding to the first aggregated packet, and the first aggregated packet being any one of the at least one aggregated packets; a first routing identifier is determined based on the first hash factor using a hash routing algorithm, the first routing identifier indicating a forwarding path; and the first aggregated packet is forwarded based on the first routing identifier.
[0030] When the intranet computation identifier corresponding to the aggregated packet is used as a hash factor for routing, for two different aggregated packets aggregated by two different underlying forwarding nodes, regardless of where the intranet computation packet corresponding to the aggregated packet comes from, as long as the intranet identifiers corresponding to the different aggregated packets are the same, based on the method provided in the embodiments of this application, both underlying forwarding nodes will forward the aggregated packets to the same next hop, thereby ensuring the path consistency of the intranet computation flow.
[0031] Based on the method provided in the first aspect, in one possible implementation, the first hash factor also includes the protocol version number and destination port number carried by the network-internal computation message corresponding to the first aggregated message.
[0032] The protocol version number and destination port number can be used to indicate the communication protocol used by the current intranet computing message. When the hash factor includes the intranet computing identifier, protocol version number, and destination port number, it can enable messages based on the same communication protocol and belonging to the same intranet computing message to be forwarded to the same forwarding node.
[0033] Based on the method provided in the first aspect, in one possible implementation, the process of a forwarding node forwarding each aggregated packet based on the intranet computing identifier corresponding to each aggregated packet in at least one aggregated packet using a hash routing algorithm can be as follows: For a second aggregated packet in at least one aggregated packet, based on the intranet computing identifier corresponding to the second aggregated packet and the forwarding table, the second aggregated packet is forwarded using a hash routing algorithm, wherein the second aggregated packet is any one of the at least one aggregated packet; wherein the forwarding table includes multiple target forwarding table entries, the next hop in the multiple target forwarding table entries is the same, and the total number of the multiple target forwarding table entries indicates the intranet computing capability of the corresponding next hop, and the multiple target forwarding table entries are scattered in the forwarding table.
[0034] Since the forwarding table is set based on the network computing capacity of the next hop, the number of forwarding table entries that include the same next hop is relatively large for forwarding table entries with a high-configuration forwarding node as the next hop. This increases the probability of selecting these forwarding table entries and correspondingly increases the traffic forwarded to the high-configuration forwarding node, thereby fully utilizing the network computing capacity of the high-configuration forwarding node.
[0035] Furthermore, if target forwarding entries that share the same next hop are concentrated in the same location within the forwarding table, assuming a large number of such entries, the probability of a forwarding node selecting other forwarding entries will be very low. This can easily lead to network congestion at the next hop corresponding to the target forwarding entry, while the next hops corresponding to other forwarding entries experience zero load. Therefore, in this embodiment, multiple target entries that share the same next hop are distributed across the forwarding table to ensure sufficient load on the next hops included in other forwarding entries.
[0036] Secondly, a forwarding node is provided, which has the functionality to implement the method behavior of forwarding intra-network computational packets as described in the first aspect above. The forwarding node includes at least one module for implementing the method of forwarding intra-network computational packets provided in the first aspect above.
[0037] Thirdly, a forwarding node is provided, the forwarding node comprising a processor and a memory. The memory stores programs that support the forwarding node in executing the method for forwarding intra-network computational messages provided in the first aspect, and stores data related to implementing the method for forwarding intra-network computational messages provided in the first aspect. The processor is configured to execute the programs stored in the memory. The operating means of the storage device may further include a communication bus for establishing a connection between the processor and the memory.
[0038] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the method for forwarding network-computed messages as described in the first aspect.
[0039] Fifthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to execute the method for forwarding network-wide computational messages described in the first aspect.
[0040] The technical effects achieved by the second, third, fourth, and fifth aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect, and will not be repeated here. Attached Figure Description
[0041] Figure 1 This is a schematic diagram illustrating a path consistency feature provided in an embodiment of this application;
[0042] Figure 2 This is a schematic diagram illustrating another scenario of path consistency characteristics provided in an embodiment of this application;
[0043] Figure 3 This is a schematic diagram of a scenario with different sources and different hosts provided in an embodiment of this application;
[0044] Figure 4 This is a schematic diagram of two multi-path routing schemes provided in an embodiment of this application;
[0045] Figure 5 This is a schematic diagram of the architecture of a communication system provided in an embodiment of this application;
[0046] Figure 6 This is a schematic diagram of the architecture of another communication system provided in an embodiment of this application;
[0047] Figure 7 This is a flowchart of a method for forwarding network-computed packets provided in an embodiment of this application;
[0048] Figure 8 This is a schematic diagram of the format of a RoCEv2 message provided in an embodiment of this application;
[0049] Figure 9 This is a schematic diagram illustrating the format of some fields of an announcement message provided in an embodiment of this application;
[0050] Figure 10 This is a schematic diagram illustrating the format of some fields of an intranet computing message provided in an embodiment of this application;
[0051] Figure 11 This is a schematic diagram of a weighting process provided in an embodiment of this application;
[0052] Figure 12 This is a schematic diagram of a smoothing process provided in an embodiment of this application;
[0053] Figure 13 This is a schematic diagram of a process for generating a weighted smooth routing table provided in an embodiment of this application;
[0054] Figure 14 This is a schematic diagram of a process for forwarding network-computed packets provided in an embodiment of this application;
[0055] Figure 15 This is a schematic diagram of the architecture of a forwarding node provided in an embodiment of this application;
[0056] Figure 16 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0057] Figure 17 This is a schematic diagram of a forwarding node device provided in an embodiment of this application. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0059] Before explaining the embodiments of this application, the application scenarios of the embodiments of this application will be introduced first.
[0060] As data volumes grow, the computing power of a single computing node can no longer meet the demands of large-scale data processing. Therefore, distributed computing, which utilizes multiple computing nodes to collaboratively compute the same task, has become a trend. For example, sharing the computing power of clusters or the cloud can be used to train models for the same artificial intelligence (AI) task.
[0061] Distributed computing can significantly shorten the execution time of computing tasks. However, after a computing task is distributed across multiple computing nodes, these nodes need to communicate and exchange data during the execution of the task. This communication time between computing nodes increases the execution time of the computing task, making communication time between computing nodes a bottleneck problem in distributed computing.
[0062] Intra-network computing technology offloads some computational tasks to forwarding nodes within the network, such as switches with computing capabilities. Since some computation is performed by the forwarding nodes, the load on the central processing unit (CPU) of the computing nodes is reduced. Furthermore, intra-network computing enables aggregated computation, where forwarding nodes compress multiple data sources from different computing nodes into a single data set, thus saving network bandwidth and accelerating data transmission. For example, high-performance computing (HPC) tasks involve data aggregation across multiple computing nodes, and AI training requires parameter aggregation across multiple computing nodes. In either case, intra-network computing technology can reduce task completion time.
[0063] In addition, the traffic of computing tasks implemented through intranet computing technology (hereinafter referred to as intranet computing flow) has the following characteristics:
[0064] 1) Path consistency feature.
[0065] Intra-network computing flows belonging to the same intra-network computing message on different computing nodes will converge to the same forwarding node to complete the aggregation computing. For example, when using hierarchical intra-network computing, that is, when multiple layers of forwarding nodes are deployed between different computing nodes, the different forwarding nodes at the bottom layer need to converge different intra-network computing flows belonging to the same intra-network computing message to the same next-layer forwarding node so that the next-layer forwarding node can perform aggregation computing on the different intra-network computing flows belonging to the same intra-network computing message.
[0066] Figure 1 and Figure 2 These are schematic diagrams illustrating a path consistency feature provided in the embodiments of this application. Figure 1 The scenario shown includes two top-of-rack (ToR) rack switches and six compute nodes. Figure 1 The two ToR switches are labeled ToR1 and ToR2, and the six compute nodes are labeled w1 to w6. Figure 1 The scenario includes two scenarios: Scenario 1, which is a scenario where intra-network computing can be performed using ToR, and Scenario 2, which is a scenario where intra-network computing cannot be performed using ToR.
[0067] like Figure 1 As shown, in Scenario 1 where intra-network computation can be performed using ToR, when w1 to w6 execute the same computation task, each intra-network computation message (i.e., intra-network computation message belonging to the same intra-network computation message) for that computation task is sent to ToR1 through the communication link shown by the solid line. Then ToR1 can perform aggregate computation on the intra-network computation messages sent by w1 to w6 that belong to the same intra-network computation message. Alternatively, the intra-network computation messages belonging to the same intra-network computation message are sent to ToR2 through the communication link shown by the dashed line. Then ToR2 can perform aggregate computation on the intra-network computation messages sent by w1 to w6 that belong to the same intra-network computation message.
[0068] like Figure 1 As shown, in Scenario 2 where intranet computation cannot be completed using ToR, when w1 to w6 are performing the same computation task, intranet computation messages exchanged between w1 and w2, and between w2 and w3, are forwarded through ToR1. However, intranet computation messages exchanged between w3 and w4, between w4 and w5, between w5 and w6, and between w6 and w1 are forwarded through ToR2. Thus, it is impossible to aggregate and compute all intranet computation messages belonging to the same intranet computation message through the same forwarding node.
[0069] Figure 2The scenario shown includes two top-of-rack (ToR) rack switches, two leaf (Spine) switches, and six compute nodes. Figure 2 The two ToR switches are labeled ToR1 and ToR2, the two leaf switches are labeled S1 and S2, and the six compute nodes are labeled w1 to w6. Figure 2 The system also includes two scenarios: Scenario 1, which is a scenario where intra-network computation can be performed using leaf switches, and Scenario 2, which is a scenario where intra-network computation cannot be performed using leaf switches.
[0070] like Figure 2 As shown, in Scenario 1 where intranet computation can be completed using leaves, when w1 to w6 execute the same computation task, any one of w1 to w3 will send all intranet computation messages (i.e., intranet computation messages belonging to the same intranet computation message) for that computation task to ToR1, and any one of w4 to w6 will send all intranet computation messages belonging to the same intranet computation message to ToR2. ToR1 and ToR2 will then forward all intranet computation messages belonging to that intranet computation message through... Figure 2 The link shown by the solid line sends data to S1, which then performs aggregation calculations on the intra-network calculation messages sent by w1 to w6 that belong to the same intra-network calculation message.
[0071] like Figure 2 As shown, in Scenario 2 where intranet computation cannot be completed using leaf switches, when w1 to w6 are performing the same computation task, intranet computation messages exchanged between w1 and w2, and between w2 and w3, are forwarded via ToR1 and S1. However, intranet computation messages exchanged between w3 and w4 are forwarded via ToR1, S1, and ToR2; intranet computation messages exchanged between w4 and w5, and between w5 and w6, are forwarded via ToR2 and S2; and intranet computation messages exchanged between w6 and w1 are forwarded via ToR2, S2, and ToR1. Thus, it is impossible to aggregate and compute all intranet computation messages belonging to the same intranet computation message sent by w1 to w6 through the same leaf switch.
[0072] 2) Traffic volume is limited by the computing resources of the network's computing switches (i.e., forwarding nodes).
[0073] Since computational tasks require computational resources (such as registers) on the forwarding nodes to complete the computation, the size of the intranet computational flow that a link can accommodate is limited by the amount of computational resources available on the forwarding nodes.
[0074] Based on the above two characteristics, the routing paths of intranet computing flows are usually statically fixed, that is, the forwarding paths of intranet computing flows are statically configured. However, when intranet computing flows and traditional traffic (that is, forwarding nodes only forward packets and do not perform aggregation calculations) coexist in the network, this static configuration method has the following problems.
[0075] (1) If the path of the intranet computing flow is not a lightly loaded network link, the network transmission time of the intranet computing flow will increase, thereby increasing the completion time of the computing task.
[0076] (2) Since the intranet computing flow is usually large-scale, the blocking effect of the intranet computing flow on the traditional traffic on the same link is also not negligible. Therefore, the transmission time of traditional traffic and the completion time of traditional computing tasks are also increased.
[0077] Furthermore, in certain scenarios, such as the cloud, multiple tenants typically conduct AI training (i.e., AI tasks). For the small-scale traffic generated by these AI tasks and the traffic generated by computationally intensive AI tasks, even using in-network computing technology will not significantly reduce task completion time. Therefore, this traffic is still forwarded using traditional methods, and is referred to as traditional traffic. However, for the large-scale traffic generated by communication-intensive AI tasks, using in-network computing technology would significantly shorten task completion time. Therefore, this traffic can utilize in-network computing technology, and is referred to as in-network computing flow. In this scenario, as the traditional traffic load on the same link increases, the transmission latency of the in-network computing flow will increase, ultimately causing the actual completion time of the in-network computing flow to increase rather than decrease.
[0078] Therefore, in some scenarios, network-wide computing flows also require load balancing. Currently, traditional Transmission Control Protocol / Internet Protocol (TCP / IP) networks typically employ equal-cost multiple path (ECMP) routing distribution technology based on hashing. Specifically, when a switch port receives a packet, it extracts the protocol version number, source IP address, source port number, destination IP address, and destination port number from the packet header to form a 5-tuple. This 5-tuple is then used as a hash factor to calculate a hash value (Hash Key). This hash value is then used for a modulo-N operation to determine the next-hop forwarding route.
[0079] However, ECMP is not suitable for multi-path forwarding of intra-network computing flows because routing based on the current 5-tuple cannot satisfy the path consistency characteristic of intra-network computing flows. For example, when deploying multi-layer forwarding nodes among different computing nodes, multiple lower-level forwarding nodes need to route packets from different intra-network computing flows belonging to the same intra-network computing message to the same higher-level forwarding node to complete the aggregation computation. However, the 5-tuples carried by the packets in the intra-network computing flows on different lower-level forwarding nodes are not the same, belonging to heterogeneous sources and destinations. Therefore, the next-hop forwarding route selected by the hash calculation is also inconsistent, thus failing to forward to the same higher-level forwarding node, and consequently failing to complete the aggregation computation of the higher-level forwarding node.
[0080] Figure 3 This is a schematic diagram of a heterogeneous source and heterogeneous host scenario provided in an embodiment of this application. Figure 3 In this diagram, 'w' represents a computing node, and w1 to w6 represent a group used to perform the same computing task. 'L' represents a lower-level forwarding node, 'S' represents a higher-level forwarding node, 'load' indicates the number of packets transmitted on the link, and 'w1->w2' indicates that the current data packet is sent from w1 (source) to w2 (destination). Other similar concepts can be found in this explanation, and will not be illustrated further here.
[0081] Let's take L1 and L2 forwarded messages as an example for illustration. Figure 3 As shown, L1 receives packets including w1->w2 and w2->w3, while L2 receives packets including w3->w4 and w4->w5. After the packets sent by w1 and w2 are aggregated on L1, when L1 forwards the aggregated packets to the next-hop route, the IP header carried in the packets may be w1->w2. After the packets sent by w3 and w4 are aggregated on L2, when L2 forwards the aggregated packets to the next-hop route, the IP header carried in the packets may be w3->w4. In this case, according to ECMP technology, since the packets sent by L1 and L2 to the higher-level forwarding nodes are from different sources and destinations, the next-hop route used by L1 to forward the aggregated packets will be different from the next-hop route used by L2. Therefore, it will be impossible to send the aggregated packets from L1 and L2 to the same next-hop route for further aggregation.
[0082] Based on this, currently, statically configured multipathing is commonly used to forward intra-network computing flows belonging to different intra-network computing messages to specified paths within multiple paths. For example, in PANAMA technology (an intra-network computing technology), different paths can be distinguished using different IP multicast addresses. Because the paths are statically configured, the path consistency characteristic of intra-network computing flows belonging to the same intra-network computing message is also guaranteed while multipath forwarding is implemented for intra-network computing flows belonging to different intra-network computing messages.
[0083] However, the above static configuration method still has the following drawbacks:
[0084] 1) In PANAMA technology, different paths are distinguished by IP multicast addresses. This requires the compute nodes to pre-configure the relevant IP multicast addresses for the paths. In other words, the routing decision of the compute flow within the network is not transparent to the compute nodes (end side), which can easily lead to malicious attacks on the compute flow within the network.
[0085] 2) The above static configuration method requires an additional centralized management node to statically configure the path, which is complex in design. When the number of packets in the intranet computing flow belonging to a certain intranet computing message is large and the number changes, it is usually necessary to reconfigure the path for the intranet computing flow belonging to that intranet computing message, which is costly in terms of configuration and maintenance.
[0086] 3) Currently, in PANAMA technology, the intranet computing stream is transmitted based on the unreliable Layer 4 transport protocol User Datagram Protocol (UDP), which results in no reliability guarantee for the transmission of messages in the intranet computing stream, making it a best-effort service.
[0087] 4) The multipath method described above using static configuration is an equivalent multipath method. When different forwarding nodes impose different limits on the maximum capacity of intranet computing traffic on the link, that is, when different forwarding nodes have different intranet computing capabilities, the equivalent multipath method cannot fully utilize the intranet computing performance of high-configuration forwarding nodes.
[0088] Figure 4 This is a schematic diagram of two multi-path routing schemes provided in an embodiment of this application. The maximum capacity of S1 for intra-network computing flows is twice that of S2 for intra-network computing flows. When using the equal-cost multi-path method, both ToR1 and ToR2 forward half of the intra-network computing flows to S1 and the other half to S2. At this time, the traffic received by S1 and S2 is the same, but since the maximum capacity of S1 is twice that of S2, half of S1's maximum capacity is wasted.
[0089] However, if a non-equivalent multipath method is used, such as both ToR1 and ToR2 forwarding 2 / 3 of the intranet computing flow to S1 and the other 1 / 3 to S2, then the size of the intranet computing flow received by S1 is twice the size of the intranet computing flow received by S2. Therefore, the intranet computing capacity of S1 can be fully utilized.
[0090] Based on the problems existing in the static path configuration method, this application provides a method for forwarding network-computed packets. For ease of subsequent understanding, the overall concept of this application embodiment will be explained first.
[0091] (1) In this embodiment, there is no need to pre-configure paths statically. Instead, the forwarding node performs hash routing based on the intra-network computing identifier of the intra-network computing message. This ensures that intra-network computing messages belonging to the same intra-network computing message are routed to the same next hop, thus satisfying the path consistency requirement of intra-network computing traffic while dynamically selecting paths. Furthermore, since there is no need to statically configure paths, the routing strategy of the forwarding node is transparent to the computing node, enhancing the transmission security of intra-network computing flows.
[0092] (2) Since no static path configuration is required, there is no need to configure an additional centralized management node for static path configuration, simplifying the network design. Even if the network changes in the future, the cost of network maintenance is relatively low.
[0093] (3) In this embodiment of the application, the intranet computing stream can be transmitted via Remote Direct Memory Access over Converged Ethernet (RoCE) instead of via unreliable UDP, thereby improving the transmission reliability of the intranet computing stream.
[0094] (4) The embodiments of this application can also perform weighted smoothing processing on the forwarding table. By performing weighted smoothing processing on the forwarding table, the intra-network computing flow received by each forwarding node after load balancing through the forwarding table can be matched as closely as possible to its own intra-network computing capabilities.
[0095] The communication architecture involved in the embodiments of this application will be explained below.
[0096] Figure 5 This is a schematic diagram of the architecture of a communication system provided in an embodiment of this application. Figure 5 As shown, the communication system includes multiple computing nodes, which exchange data through a multi-layer forwarding network. Each layer of the forwarding network includes at least two forwarding nodes to meet the multi-path requirements of intra-network computing traffic.
[0097] In this embodiment, the number of layers in the forwarding network is not limited. Figure 5 This example uses a two-layer forwarding network. The embodiments of this application do not limit the number of forwarding nodes or computing nodes included in each layer of the forwarding network. Figure 5 This example illustrates the concept of a forwarding network where each layer consists of two forwarding nodes.
[0098] Figure 6This is a schematic diagram of the architecture of another communication system provided in an embodiment of this application. For example... Figure 6 As shown, this communication system typically includes six computing nodes, a first-layer forwarding network, and a second-layer forwarding network. Among them, Figure 6 The six computing nodes are labeled w1 to w6. Additionally, as... Figure 6 As shown, the first-layer forwarding network example includes three top-of-rack (ToR) rack switches, labeled ToR1, ToR2, and ToR3, respectively. The second-layer forwarding network example includes two leaf switches, labeled S1 and S2, respectively.
[0099] based on Figure 5 or Figure 6 The following explains the method for forwarding network-computed messages provided in the embodiments of this application, based on the communication system shown.
[0100] Figure 7 This is a flowchart illustrating a method for forwarding network-computed packets according to an embodiment of this application. Figure 7 As shown, the method includes the following steps 701 to 704.
[0101] Step 701: The forwarding node receives multiple intranet computing packets.
[0102] For example, a forwarding node can be Figure 5 The diagram shows the underlying forwarding nodes in a communication network. For example, a forwarding node could be... Figure 6 The ToR1, ToR2, or ToR3 in the model.
[0103] In addition, the method provided in this application is applicable to scenarios where a forwarding node receives an uplink intranet computing message from a computing node. For scenarios where a forwarding node receives a downlink result message from another forwarding node or a scenario where a forwarding node receives traditional traffic, the forwarding node can forward the message in the normal forwarding manner.
[0104] Among them, uplink intra-network computation messages refer to intra-network computation messages sent from the computing node to the higher-layer forwarding node, while downlink result messages refer to intra-network computation messages sent from the higher-layer forwarding node after aggregation to the computing node. For example... Figure 6 As shown, intra-network computing messages transmitted on the link from the computing node (any one of w1 to w6) to the ToR switch and then to the leaf switch (S1 or S2) are called uplink intra-network computing messages, and intra-network computing messages transmitted on the link from the leaf switch (S1 or S2) to the ToR switch and then to the computing node (any one of w1 to w6) are called downlink result messages.
[0105] In some embodiments, to improve message transmission reliability, intra-network computing messages can be Remote Direct Memory Access over Converged Ethernet (RoCE) messages. That is, the RoCE protocol is used to transmit intra-network computing messages within the network.
[0106] RoCE is a Remote Direct Memory Access (RDMA) protocol based on converged Ethernet. RoCE allows remote hosts to directly access each other's memory over a network. The following example illustrates RoCE-based message delivery using RoCE version 2 (RoCEv2).
[0107] RoCEv2 is a network layer protocol. RoCEv2-based packets (i.e., RoCEv2 messages) carry routing information in their headers, allowing forwarding nodes to forward them. Therefore, RoCEv2 is widely used in distributed computing applications (such as HPC and AI applications). The intra-network computing packets provided in this embodiment can, for example, be RoCEv2 messages.
[0108] Figure 8 This is a schematic diagram illustrating the format of a RoCEv2 message provided in an embodiment of this application. For example... Figure 8 As shown, RoCEv2 messages include two layers of Ethernet header (ETH L2), IP header, UDP header, InfiniBand BaseTransport Header (IB BTH), IB payload, cyclic redundancy check (CRC), and frame check sequence (FCS).
[0109] The ETH L2 header carries the source and destination media access control (MAC) addresses. The IP header carries the source and destination IP addresses. The UDP header carries the source and destination port numbers; in this case, the destination port number for the RoCEv2 packet is 4791. CRC and FCS correspond to redundancy detection and frame check, respectively.
[0110] IB TH is used to carry key fields for intelligent traffic analysis, such as Figure 8As shown, the IB TH includes fields such as Opcode, Pad Count, Destination Queue Pair (QPN), and Packet Sequence Number (PSN). The Opcode indicates the type of RoCEv2 packet; for details on RoCEv2 packet types, please refer to relevant standards. The Pad Count indicates how many extra bytes are padded into the IB payload. The QPN identifies a RoCEv2 stream. The PSN identifies the sequence number of the RoCEv2 packet; for example, checking the continuity of PSNs can help determine if there are any lost packets.
[0111] about Figure 8 Other fields in the RoCEv2 message shown can be found in the definitions in the relevant standards, and will not be explained in detail here.
[0112] When transmitting intra-network computation messages via RoCEv2, the compute node organizes the data to be transmitted into multiple messages, and then uses the RDMA network card to divide each message into multiple RDMA messages according to the IB payload size limit for transmission. The data to be transmitted is also a single intra-network computation stream, which can be, for example, model parameters or gradients, such as tensors.
[0113] The above explanation uses RoCEv2 messages as an example. Optionally, other versions of the RoCE protocol can also be used for intra-network computing messages, which will not be illustrated here.
[0114] When the intra-network computation message is a RoCE message, it can be determined whether the intra-network computation message is an uplink intra-network computation message based on the information carried in the message. In this scenario, the forwarding node stores at least one RDMA connection identifier, and each RDMA connection identifier indicates an RDMA connection between a first computing node and a second computing node, where the first computing node is the computing node accessed by the forwarding node. That is, each RDMA connection identifier indicates an RDMA connection from one uplink computing node of the forwarding node to other computing nodes.
[0115] At this point, for the second intra-network computing packet among multiple intra-network computing packets, the forwarding node can obtain the RDMA connection identifier carried by the second intra-network computing packet, which can be any one of the multiple intra-network computing packets. If the RDMA connection identifier carried by the second intra-network computing packet is present in at least one of the stored RDMA connection identifiers, it indicates that the second intra-network computing packet is an uplink intra-network computing packet, and the second intra-network computing packet is forwarded using the method provided in this application embodiment.
[0116] The RDMA connection identifier may, for example, include the source IP address, source port identifier (or source port number), and QPN. Additionally, at least one RDMA connection identifier may be stored in the aggregated IP mapping relationship. In this scenario, when the source IP address, source port number, and QPN in the header of the second intra-network packet are all recorded in the aggregated IP mapping relationship, it indicates that the second intra-network computation packet is transmitted through an RDMA connection between an uplink computation node and another computation node; therefore, the second RDMA packet is an uplink intra-network computation packet.
[0117] Optionally, the RDMA connection identifier example may also include Figure 8 The additional information in the message header shown will not be illustrated here.
[0118] In addition, at least one RDMA connection identifier can be generated based on the RDMA connections created between compute nodes during network initialization. That is, during network initialization, when an uplink compute node of a forwarding node creates an RDMA connection with other compute nodes, the forwarding node can automatically store at least one RDMA connection identifier based on the created RDMA connection.
[0119] Based on this, in some embodiments, the forwarding node receives an announcement message from the uplink communication link; the forwarding node obtains the target RDMA connection identifier carried in the announcement message, and the announcement message is used to announce that the RDMA connection indicated by the target RDMA connection identifier is used to transmit intra-network computational packets; the forwarding node stores the target RDMA connection identifier. For example, the target RDMA connection identifier can be stored in the above-mentioned aggregated IP mapping relationship.
[0120] In this embodiment, the announcement message is a control layer message. To distinguish the control layer message from other control messages, type information can be carried in the announcement message. This type information indicates whether the announcement message is for an intra-network computation message. In this scenario, when the forwarding node receives the announcement message, it can also obtain the type information carried in the announcement message. If the type information indicates that the announcement message is for an intra-network computation message, then the operation of obtaining the target RDMA connection identifier carried in the announcement message is performed. Correspondingly, if the type information indicates that the announcement message is not for an intra-network computation message, then the operation of obtaining the target RDMA connection identifier carried in the announcement message does not need to be performed. In this way, the efficiency of the forwarding node in storing at least one RDMA connection identifier can be improved.
[0121] For example, a field in the notification message can be expanded to carry type information. For instance, when the notification message format is... Figure 8When using the RoCEv2 message shown, the IB payload field of the RoCEv2 message can be extended to allow the IB payload field to carry this type of information.
[0122] Figure 9 This is a schematic diagram illustrating the format of some fields in a notification message provided in an embodiment of this application. The format of the notification message is as follows: Figure 8 The RoCEv2 message shown, and the IB payload field of this notification message includes Figure 9 The information shown. For example... Figure 9 As shown, the IB payload fields of the notification message include the communication group ID (CommID), the true communication group ID (trueCommID), the communication protocol type (Comm Protocol Type), the communication operation type (Comm Op), the global group size (Global Group Size), and the local group size (Local Group Size).
[0123] Specifically, when the communication group identifier (CommID) is 0x0001 / 0x0002, it indicates that the announcement message is for intra-network compute packets. The target RDMA connection identifier carried in this announcement message can be obtained from... Figure 8 Retrieve other fields.
[0124] In this embodiment, the notification message is used to notify an RDMA connection established by an uplink computing node, and this RDMA connection is used to transmit intra-network computing messages. Therefore, the notification message can also carry some information about the intra-network computing messages transmitted by the RDMA connection. For example... Figure 9 As shown, since the communication group identifier (CommID) is used to carry type information, the announcement message also carries the true communication group identifier (true CommID), which indicates the true identifier of the communication group currently performing intranet calculations.
[0125] In addition, such as Figure 9 As shown, the announcement message also carries the communication protocol type, indicating the currently used protocol type. The announcement message also carries the communication operation type (Comm Op), which indicates the operation type required for the current operation. The announcement message also carries the global communication group size and the local communication group size. The global communication group size indicates the total number of compute nodes within the communication group, and the local communication group size indicates the number of compute nodes belonging to that communication group among the compute nodes accessed by this forwarding node.
[0126] In this scenario, two options, communication group identifier (CommID) and communication group size, can be added to the aggregated IP mapping relationship. This allows subsequent forwarding nodes to determine whether they have received all intra-network calculation packets that need to be aggregated based on the communication group identifier (CommID) and communication group size. The relevant content will be explained in detail in subsequent step 703, and will not be elaborated here.
[0127] Table 1 is a schematic diagram of an aggregated IP mapping relationship provided in an embodiment of this application. As shown in Table 1, the aggregated IP mapping relationship includes options such as source IP address, source port identifier (or source port number), QPN, communication group identifier, and communication group size (GroupSize). Among them, the source IP address, source port identifier (or source port number), and QPN indicate a unique RDMA connection.
[0128] Table 1
[0129]
[0130] Based on the aggregated IP mapping relationship shown in Table 1, when a forwarding node receives an advertisement message from the uplink communication link, it obtains the source IP address, source port identifier (or source port number), and QPN from the advertisement message and uses these three as an RDMA connection identifier. Then, it obtains the communication group identifier (CommID) from the IB payload of the advertisement message. If the communication group identifier (CommID) is 0x0001 / 0x0002, it adds a new entry to the aggregated IP mapping relationship. Then, it obtains the real communication group identifier and adds the mapping relationship between the real communication group and the RDMA connection identifier to the newly added entry.
[0131] Among them, 0x0001 and 0x0002 are used to distinguish whether the previous hop of the forwarding path of the announcement message is a calculated node.
[0132] For example, when the communication group identifier (CommID) is 0x0001, it indicates that the previous hop of the forwarding path of this advertisement message is a compute node. In this case, the forwarding node retrieves the local communication group size from the IB payload field of the advertisement message and adds it as a new entry to the communication group size table. Then, the forwarding node determines whether the local communication group size in the IB payload field of the advertisement message is less than the global communication group size, i.e., whether the Local Group Size is less than the Global GroupSize. If this condition is true, it indicates that multi-level intra-network computation is required, meaning there are higher-level forwarding nodes. Therefore, the forwarding node changes the communication group identifier (CommID) to 0x0002 and forwards the advertisement message to all higher-level forwarding nodes. If this condition is false, it indicates that higher-level intra-network computation is not required, and the normal forwarding process begins, meaning the forwarding node continues to forward the advertisement message downlink to other compute nodes.
[0133] When a forwarding node receives an advertisement message with a communication group identifier (CommID) of 0x0002, it indicates that the previous hop of the forwarding path of this advertisement message is another forwarding node, and the current forwarding node is a higher-layer forwarding node in the network. Since it may have received other advertisement messages for the same communication group from other forwarding nodes before the current time, the forwarding node first checks whether there is already an entry in the aggregated IP mapping relationship that includes the real communication group identifier in the advertisement message. If there is already an entry in the aggregated IP mapping relationship that includes the real communication group identifier in the advertisement message, then the communication group size in that entry is updated to the local communication group size carried in the advertisement message plus the original communication group size, that is, Group Size = Local Group Size + Group Size.
[0134] Accordingly, if there is no entry in the aggregated IP mapping relationship that includes the real communication group identifier in the announcement message, then a new entry is added to the aggregated IP mapping relationship, and the real communication group identifier and the local communication group size carried in the announcement message are added to the new entry.
[0135] Then the forwarding node determines the size of the communication group corresponding to the real communication group identifier in the announcement message in the aggregated IP mapping relationship, and judges the size relationship between the communication group size and the global communication group size.
[0136] If the communication group size is smaller than the global communication group size, that is, Group Size < Global Group Size, it indicates that not all the announcement messages published by the computing nodes within the communication group have been received yet, and there may still be other forwarding nodes in the current network that have not completed the deployment of the aggregated IP mapping relationship. Therefore, cache the current packet and wait for the announcement messages published by other computing nodes. The purpose of doing this is to ensure that when the in-network computing packets are received subsequently, it can be determined whether the in-network computing packet belongs to the upstream in-network computing packet from the complete aggregated IP mapping relationship, avoiding incorrect judgments. Among them, the complete aggregated IP mapping relationship refers to the RDMA connection identifiers corresponding to the announcement messages published by all the upstream computing nodes of the forwarding node.
[0137] Correspondingly, if the communication group size is equal to the global communication group size, that is, Group Size = Global Group Size, it indicates that all the forwarding nodes in the network have completed the deployment of the aggregated IP mapping relationship. Therefore, forward the announcement message with CommID = 0x0002 to other computing nodes normally.
[0138] When neither of the two conditions is met, it indicates that the current announcement message is incorrect, and the current announcement message can be discarded.
[0139] In addition, for the convenience of subsequent reference to each entry in Table 1, as shown in Table 1, each entry in the aggregated IP mapping relationship can be stored in a key-value manner, and each entry includes a key and a value. Among them, the key includes the source IP address, source port identifier (or source port number), QPN, and communication group size (Group Size), and the value includes the index generated based on the corresponding key and the communication group identifier.
[0140] It should be noted that the embodiments of the present application do not limit the storage method of the aggregated IP mapping relationship, and no further examples will be given here.
[0141] The above takes the computing node publishing the announcement message as an example to illustrate how to configure the aggregated IP mapping relationship on the forwarding node. Optionally, the aggregated IP mapping relationship can be statically configured on the forwarding node by a technician based on the RDMA connections in the network, or the control node can generate the aggregated IP mapping relationship based on the RDMA connections in the network and send it to the forwarding node.
[0142] Step 702: The forwarding node determines the in-network computing identifier of each in-network computing packet among multiple in-network computing packets, and the in-network computing identifier indicates the in-network computing message to which the corresponding in-network computing packet belongs.
[0143] In some embodiments, the intra-network computing identifier may include a communication group identifier and an intra-network computing message identifier. The communication group identifier indicates the communication group currently performing intra-network computing, and the intra-network computing message identifier indicates an intra-network computing message transmitted within the communication group. Considering that the intra-network computing message identifiers transmitted in different communication groups may be the same, in this embodiment, the intra-network computing message identifier and the communication group identifier can be combined as the intra-network computing identifier. Optionally, if the intra-network computing message identifiers transmitted in different communication groups are completely different, in this scenario, the intra-network computing message identifier can be directly used as the intra-network computing identifier.
[0144] In this embodiment, to ensure path consistency of intra-network computing flows through a dynamic hash routing algorithm, the format of intra-network computing packets can be extended to include some indicative information. Thus, when a forwarding node receives an intra-network computing packet, it can obtain the intra-network computing identifier based on the indicative information carried in the packet, facilitating subsequent hash routing using the intra-network computing identifier corresponding to the aggregated packet.
[0145] In some embodiments, each intranet computing packet may carry an intranet computing identifier. In this scenario, the forwarding node can directly obtain the intranet computing identifier from the packet header of the intranet computing packet.
[0146] Optionally, in other embodiments, to avoid excessive bandwidth consumption by intra-network computing flows, the intra-network computing identifier is carried in the first packet of an intra-network computing flow, but not in subsequent packets. In this scenario, when the forwarding node receives the first packet of each intra-network computing flow, it can obtain the data flow identifier and the intra-network computing identifier to which the intra-network computing flow belongs from the first packet, and then record the mapping relationship between the obtained data flow identifier and the intra-network computing identifier in the aggregated information mapping relationship. That is, the forwarding node will store the aggregated information mapping relationship, which includes at least one intra-network computing identifier and a data flow identifier corresponding to each of the at least one intra-network computing identifier.
[0147] In this scenario, the forwarding node can obtain the intra-network computing identifier for each intra-network computing packet as follows: For any intra-network computing packet among multiple intra-network computing packets, determine the data flow identifier of that intra-network computing packet, and then look up the intra-network computing identifier corresponding to that data flow identifier from the aggregation information mapping relationship. Conversely, if the intra-network computing identifier corresponding to that data flow identifier is not found in the aggregation information mapping relationship, it indicates that the intra-network computing packet is the first packet of an intra-network computing flow, and in this case, the intra-network computing identifier can be directly obtained from the packet header of that intra-network computing packet.
[0148] For example, for a first intra-network computing message among multiple intra-network computing messages, the data flow identifier corresponding to the first intra-network computing message is determined to obtain the first data flow identifier. The first intra-network computing message can be any one of the multiple intra-network computing messages. If the first data flow identifier exists in the aggregation information mapping relationship, the intra-network computing identifier corresponding to the first data flow identifier is obtained from the aggregation information mapping relationship to obtain the first intra-network computing identifier. The first intra-network computing identifier is the intra-network computing identifier of the first intra-network computing message. Correspondingly, if the first data flow identifier does not exist in the aggregation information mapping relationship, the first data flow identifier and the first intra-network computing identifier are obtained from the message header of the first intra-network computing message; the correspondence between the first intra-network computing identifier and the first data flow identifier is added to the aggregation information mapping relationship.
[0149] The data flow identifier may, for example, include a communication operation identifier (Tag), a data flow length, and a first packet sequence number. The data flow length may, for example, be the number of packets included in the intra-network computing flow; that is, the unit of data flow length is the number of packets. The communication operation identifier (Tag) indicates that the current computing node sends this intra-network computing flow operation. The first packet sequence number indicates the PSN of the first packet in the intra-network computing flow. By extending these fields into intra-network computing packets as data flow identifiers in this manner, the application flexibility of the embodiments of this application is improved.
[0150] In some embodiments, each packet in the intra-network computing flow carries a data flow identifier. In this scenario, the forwarding node can directly obtain the data flow identifier from the intra-network computing packets.
[0151] Optionally, in other embodiments, in scenarios where the intra-network computing message is a RoCE message, different intra-network computing flows are typically sent based on different RDMA connections. In this scenario, the data flow identifier can be carried in the first packet of an intra-network computing flow, while the data flow identifier is not carried in the subsequent packets of the intra-network computing flow. This avoids the intra-network computing flow consuming too much bandwidth and improves the transmission speed of the intra-network computing flow.
[0152] When a forwarding node receives the first packet, it records the mapping relationship between the data flow identifier, the intra-network computing identifier, and the RDMA connection identifier in the aggregation information mapping relationship. When a forwarding node receives a non-first packet, it determines the corresponding data flow identifier and the corresponding intra-network computing identifier based on the RDMA connection identifier carried by the non-first packet.
[0153] Table 2 shows an aggregated information mapping relationship provided in an embodiment of this application. As shown in Table 2, the aggregated information mapping relationship includes source IP address, source port identifier (or source port number), QPN, communication group identifier, communication operation identifier (Tag), network-internal computed message identifier (MsgID), data stream length (MsgLen), first packet PSN, etc.
[0154] Table 2
[0155]
[0156] As shown in Table 2, the source IP address, source port number, and QPN serve as RDMA connection identifiers to uniquely identify an RDMA connection. The communication group identifier and MsgID serve as intra-network computation identifiers to uniquely identify an intra-network computation message. Tag, MsgLen, and the first packet PSN serve as data stream identifiers to identify a data stream.
[0157] Based on the aggregation information mapping relationship shown in Table 2, in the scenario where the aggregation IP mapping relationship is stored at the forwarding node in step 701, the implementation method for the forwarding node to determine the intranet computing identifier of each intranet computing packet in multiple intranet computing packets can be as follows: For any intranet computing packet in multiple intranet computing packets, the forwarding node obtains the RDMA connection identifier carried by the intranet computing packet; if the RDMA connection identifier carried by the intranet computing packet exists in the aggregation IP mapping relationship, it indicates that the intranet computing packet is an uplink intranet computing packet. Then, the data flow identifier and intranet computing identifier of the intranet computing packet are searched from the aggregation information mapping relationship based on the RDMA connection identifier. If the data flow identifier and intranet computing identifier corresponding to the RDMA connection identifier are not found in the aggregation information mapping relationship, it indicates that the intranet computing packet is the first packet of an uplink intranet computing flow. At this time, the data flow identifier and intranet computing identifier are obtained from the packet header of the intranet computing packet, and the obtained data flow identifier, intranet computing identifier, and RDMA connection identifier are added to the aggregation information mapping relationship.
[0158] Optionally, in the scenario where Table 1 is stored using a key-value pair, to reduce the storage space of Table 2, Table 2 can also be stored using a key-value pair, with the first three columns of Table 2 stored using the corresponding communication group identifier from Table 1 plus an index. In this scenario, Table 2 can be converted into Table 3 as shown below.
[0159] In a scenario where Tables 1 and 3 are stored at the forwarding node, the implementation method for the forwarding node to determine the intra-network computing identifier of each intra-network computing packet among multiple intra-network computing packets can be as follows: For any intra-network computing packet among multiple intra-network computing packets, the forwarding node obtains the RDMA connection identifier carried by the intra-network computing packet; if the RDMA connection identifier carried by the intra-network computing packet exists in Table 1, it indicates that the intra-network computing packet is an uplink intra-network computing packet. Then, the index and communication group identifier are determined from Table 1, and the corresponding table entry is searched from Table 2 based on the index and communication group identifier. If both the data flow identifier and the intra-network computing identifier exist in the corresponding table entry, it indicates that the current intra-network computing packet is not the first packet. At this time, the data flow identifier and the intra-network computing identifier of the intra-network computing packet can be obtained from Table 3. If neither the data flow identifier nor the intranet computing identifier exists in the corresponding table entry, it indicates that the current intranet computing packet is the first packet. In this case, the data flow identifier and intranet computing identifier can be obtained from the intranet computing packet and added to the corresponding table entry.
[0160] Table 3
[0161]
[0162] In addition, the PSN of the first packet is stored as a value in Table 3 to facilitate subsequent aggregation calculations based on the PSN of the first packet. The relevant content will be explained in detail in subsequent embodiments, and will not be elaborated here.
[0163] The following example, using the format of the first packet of a certain intranet computing flow, illustrates how to carry the data flow identifier and intranet computing identifier in an intranet computing message.
[0164] Figure 10 This is a schematic diagram illustrating the format of some fields in an intranet computing message provided in an embodiment of this application. For example... Figure 10 As shown, it is possible to... Figure 8 The IB payload field of the intranet compute message is extended to carry the communication group identifier (CommID), communication operation identifier (Tag), intranet compute message identifier (MsgID), and data stream length (MsgLen) in the IB payload field.
[0165] Step 703: Based on the intranet computing identifier of each intranet computing message in multiple intranet computing messages, the forwarding node aggregates the intranet computing messages belonging to the same intranet computing message in multiple intranet computing messages to obtain at least one aggregated message. At least one aggregated message corresponds one-to-one with at least one intranet computing identifier.
[0166] In some embodiments, where an aggregation information mapping relationship is stored at the forwarding node, and the aggregation information mapping relationship stores at least one intranet computing identifier and a data flow identifier corresponding to each intranet computing identifier, the implementation method for the forwarding node to perform aggregation calculation on intranet computing packets belonging to the same intranet computing message based on the intranet computing identifier of each intranet computing packet in the plurality of intranet computing packets can be as follows: for any first intranet computing packet, the first intranet computing packet is bound to the first data flow identifier in the aggregation information mapping relationship; multiple data flow identifiers corresponding to the first intranet computing identifier in the aggregation information mapping relationship are determined; and aggregation calculation is performed on the intranet computing packets bound to the multiple data flow identifiers respectively to obtain the aggregated packet corresponding to the first intranet computing identifier.
[0167] Since the aggregation information relationship stores the data flow identifier corresponding to each intranet computing identifier, aggregation calculation can be performed based on the intranet computing packets bound to the data flow identifiers corresponding to the same intranet computing identifier in the aggregation information mapping relationship, thereby improving the processing efficiency of the forwarding nodes.
[0168] It should be noted that the forwarding node only performs aggregation computation upon receiving intra-network computation messages sent by all upstream computing nodes. Based on this, in some embodiments where the forwarding node stores the aggregation IP mapping relationship shown in Table 1, for the first intra-network computation identifier in the aggregation information mapping relationship, the forwarding node can determine the number of data stream identifiers corresponding to the first intra-network computation identifier. If the number of data stream identifiers corresponding to the first intra-network computation identifier is equal to the communication group size in the aggregation IP mapping relationship, it indicates that intra-network computation flows for the same intra-network computation message sent by all upstream computing nodes have been received, and aggregation computation can be performed in the following manner. Conversely, if the number of data stream identifiers corresponding to the first intra-network computation identifier is less than the communication group size in the aggregation IP mapping relationship, it indicates that intra-network computation flows for the same intra-network computation message sent by all upstream computing nodes have not yet been received, and the node can continue to wait.
[0169] In scenarios where the data flow identifier includes a first packet sequence number, the implementation method for aggregating and calculating the intra-network computing packets bound to the multiple data flow identifiers can be as follows: For the intra-network computing packets bound to the multiple data flow identifiers corresponding to the first intra-network computing identifier, determine the packet sequence number of each bound intra-network computing packet; based on the first packet sequence number in each data flow identifier among the multiple data flow identifiers corresponding to the first intra-network computing identifier, and the packet sequence number of each bound intra-network computing packet, determine the offset of each bound intra-network computing packet; based on the offset of each bound intra-network computing packet, perform aggregating and calculating the intra-network computing packets with the same offset.
[0170] That is, for different intra-network computing streams belonging to the same intra-network computing message, intra-network computing messages with the same offset in different intra-network computing streams are aggregated for computing.
[0171] In addition, not all packets in different intranet computing streams are sent to the forwarding node at the same time. Therefore, when the forwarding node aggregates intranet computing packets for a certain offset in different intranet computing streams, if the intranet computing packets for that offset in a certain intranet computing stream have not yet been received, it can continue to wait until all intranet computing packets for that offset in different intranet computing streams are received before performing aggregation calculation.
[0172] Alternatively, intra-network computation packets can be aggregated all at once, without waiting for all intra-network computation packets. It should be noted that in this scenario, the forwarding node will only continue to forward the aggregated packets after it has determined that all intra-network computation packets sent by uplink computing nodes have been aggregated, that is, after it has determined that the aggregation process is complete.
[0173] Step 704: The forwarding node forwards each aggregated packet using a hash routing algorithm based on the network-computed identifier corresponding to each aggregated packet in at least one aggregated packet.
[0174] In this embodiment of the application, in order to forward intranet computing messages belonging to the same intranet computing message to the same higher-layer forwarding node for continued aggregation computing, the forwarding node forwards the aggregation message based on the intranet computing identifier using a hash routing algorithm.
[0175] In some embodiments, step 704 can be implemented as follows: for a first aggregated packet in at least one aggregated packet, a first hash factor is determined, the first hash factor includes the network-internal calculation identifier corresponding to the first aggregated packet, and the first aggregated packet is any one of at least one aggregated packet; a first routing identifier is determined based on the first hash factor using a hash routing algorithm, the first routing identifier indicating a forwarding path; and the first aggregated packet is forwarded based on the first routing identifier.
[0176] When the intranet computation identifier corresponding to the aggregated packet is used as a hash factor for routing, for two different aggregated packets aggregated by two different underlying forwarding nodes, regardless of where the intranet computation packet corresponding to the aggregated packet comes from, as long as the intranet identifiers corresponding to the different aggregated packets are the same, based on the method provided in the embodiments of this application, both underlying forwarding nodes will forward the aggregated packets to the same next hop, thereby ensuring the path consistency of the intranet computation flow.
[0177] Optionally, in other embodiments, the first hash factor may include, in addition to the network computing identifier corresponding to the first aggregated message, the protocol version number and destination port number carried by the network computing message corresponding to the first aggregated message.
[0178] The protocol version number and destination port number can be used to indicate the communication protocol used by the current intranet computing message. When the hash factor includes the intranet computing identifier, protocol version number, and destination port number, it can enable messages based on the same communication protocol and belonging to the same intranet computing message to be forwarded to the same forwarding node.
[0179] For example, in a scenario where the intranet computation message is a RoCEv2 message, when the hash factor includes the intranet computation identifier, the protocol version number, and the destination port number, it is possible to forward intranet computation messages based on the reliable RoCEv2 protocol and belonging to the same intranet message to the same next hop, so as to ensure the path consistency of the intranet computation flow.
[0180] For example, in scenarios where the intranet computation identifier includes both a communication group identifier and an intranet computation message identifier, the hash factor provided in this application embodiment is a four-tuple, hence the name hash four-tuple. Table 4 shows a hash four-tuple provided in this application embodiment. As shown in Table 4, the hash four-tuple includes the protocol version number, destination port number, CommID, and MsgID. The protocol and destination port number identify that the intranet computation message is transmitted based on a reliable RoCEv2 connection, while CommID and MsgID identify the intranet computation message to which the current intranet computation message belongs.
[0181] Table 4
[0182]
[0183] When a forwarding node determines the routing identifier based on the new hash quadruple using a hash routing algorithm, it can do so using the following formula.
[0184] k=Hash(protocol, dstPort, commID, MsgID)
[0185] Where k represents the routing identifier, protocol represents the protocol version number, dstPort represents the destination port number, and key represents the routing identifier. Using the hash quadruple shown in Table 3, the routing identifier k does not change due to changes in the source IP and source port number of the intranet computation message, ensuring the consistency of the forwarding path for intranet computation messages with the same intranet computation message on different switches.
[0186] In addition, after determining the routing identifier k in the above manner, the forwarding node performs a modulo operation (Modulo-N, Mod-N) on the routing identifier k and the forwarding table length (N) using the following formula. The result (ID) of the modulo operation is used as the target forwarding table entry identifier to determine which forwarding table entry the aggregated packet should be forwarded from.
[0187] ID = k%N
[0188] Table 5 is a forwarding table provided in an embodiment of this application. As shown in Table 5, this forwarding table only lists all the optional forwarding entries, each corresponding to a forwarding path, that is, a next-hop forwarding node. In Table 5, the next-hop forwarding node is represented by the switch IP address. In this case, multi-path routing is randomized, and each forwarding entry has the same probability of being selected. Therefore, it is impossible to utilize the intranet computing power of high-configuration forwarding nodes in a heterogeneous environment.
[0189] Table 5
[0190] ID Switch IP 0 10.0.0.1 1 10.0.0.2 2 10.0.0.3
[0191] Based on this, in this embodiment of the application, in order to fully utilize the intra-network computing power of each forwarding node, the forwarding table configured on the forwarding node is improved. The improved forwarding table has multiple forwarding entries, each indicating a forwarding path. The multiple forwarding entries include multiple target forwarding entries, and the next hops in the multiple target forwarding entries are the same. The total number of the multiple target forwarding entries indicates the intra-network computing power of the corresponding next hop, and the multiple target entries are scattered throughout the forwarding table. The improved forwarding table can be called a weighted smooth routing table.
[0192] In this scenario, the implementation of forwarding nodes forwarding each aggregated packet based on the network computing identifier corresponding to each aggregated packet in at least one aggregated packet using a hash routing algorithm can be as follows: For the second aggregated packet in at least one aggregated packet, based on the network computing identifier corresponding to the second aggregated packet and the forwarding table, the second aggregated packet is forwarded using a hash routing algorithm, where the second aggregated packet is any one of the at least one aggregated packets.
[0193] Since the weighted smooth routing table sets the number of forwarding entries that include the same next hop based on the network's computing power of the next hop, for forwarding entries where the next hop is a high-configuration forwarding node, the number of forwarding entries that include that next hop will be relatively large. Therefore, the probability of selecting these forwarding entries can be increased, and the traffic forwarded to that high-configuration forwarding node can be increased accordingly, thereby giving full play to the network's computing power of that high-configuration forwarding node.
[0194] Furthermore, if target forwarding entries that share the same next hop are concentrated in the same location within the forwarding table, assuming a large number of such entries, the probability of a forwarding node selecting other forwarding entries will be very low. This can easily lead to network congestion at the next hop corresponding to the target forwarding entry, while the next hops corresponding to other forwarding entries experience zero load. Therefore, in this embodiment, multiple target entries that share the same next hop are distributed across the forwarding table to ensure sufficient load on the next hops included in other forwarding entries.
[0195] The following explains the process of generating the weighted smooth routing table.
[0196] The generation process of the above weighted smoothing routing table includes two steps: a weighting process and a smoothing process.
[0197] Figure 11 This is a schematic diagram of a weighting process provided in an embodiment of this application. For example... Figure 11 As shown, the forwarding node maintains an initial forwarding table at network initiation. Each entry (or routing entry) in the initial forwarding table contains a weight and the switch's IP address (as the next hop). Therefore, the initial forwarding table is also called the switch weight ID table. The weight represents the capacity limit of the corresponding switch's computing resources on the intranet computing flow, i.e., the intranet computing capability of the corresponding switch. Then, the forwarding table entries are expanded according to the weights, increasing the number of routing entries for the same next hop. This increases the probability of selecting the next hop with a higher weight, achieving the design goal of weighted traffic distribution. The resulting weighted expanded forwarding table is then obtained.
[0198] like Figure 11 As shown, the weight of next hop 10.0.0.1 is 3, therefore there are three forwarding entries with next hop 10.0.0.1 in the weighted expanded forwarding table. Correspondingly, the weight of next hop 10.0.0.2 is 2, therefore there are two forwarding entries with next hop 10.0.0.1 in the weighted expanded forwarding table. The weight of next hop 10.0.0.3 is 1, therefore there is one forwarding entry with next hop 10.0.0.1 in the weighted expanded forwarding table.
[0199] It's important to note that the Mod-N operation used to determine the routing identifier is a hash operation. The empty bucket rate and collision rate in hash calculations can disrupt load balancing between different forwarding nodes. The Mod-N routing process is equivalent to randomly placing m routing identifiers (keys) into N buckets (the total number of forwarding entries). Therefore, some forwarding nodes will not be selected (empty buckets), while others will be frequently selected (collisions). A higher empty bucket rate means that forwarding nodes with lower weights are less likely to receive traffic, resulting in an underloaded state and poor load balancing. A higher collision rate means that traffic is concentrated on high-weight forwarding nodes, which are overloaded, also resulting in poor load balancing.
[0200] The relationship between load factor (m / N), collision rate, and empty bucket rate is as follows: as the load factor increases, the empty bucket rate decreases, but the collision rate increases. Therefore, simply adjusting the values of m and N cannot simultaneously guarantee a low empty bucket rate and a low collision rate.
[0201] Based on this, this application's embodiments mitigate the negative impact of high empty bucket rates and high conflict rates on load balancing by adjusting the positions of forwarding table entries. Specifically, the positions of each forwarding entry in the weighted expanded forwarding table are adjusted so that forwarding entries including the same next hop are deployed in discontinuous positions within the forwarding table. This avoids empty bucket events concentrating on low-weight forwarding nodes, ensuring sufficient load distribution on low-weight links. Simultaneously, it avoids conflict events concentrating on high-weight forwarding nodes, preventing traffic overload on high-weight forwarding nodes.
[0202] Figure 12 This is a schematic diagram of a smoothing process provided in an embodiment of this application. For example... Figure 12 As shown, the three forwarding entries with the same next hop 10.0.0.1 in the smoothed forwarding table are distributed in a scattered manner, as are the two forwarding entries with the same next hop 10.0.0.2.
[0203] Figure 13 This is a schematic diagram of a process for generating a weighted smooth routing table provided in an embodiment of this application. Figure 13 For illustrative purposes only, the embodiments of this application do not limit the process of generating the weighted smooth routing table.
[0204] like Figure 13 As shown, the process of generating a weighted smooth routing table includes the following steps.
[0205] 1) Initialization:
[0206] a. Create an empty weighted smooth routing table and initialize ID = 0;
[0207] b. Copy a switch weight ID table, which is the initial forwarding table mentioned above, and generate a shadow table;
[0208] c. Calculate the sum of the weights of all switches in the initial forwarding table, denoted as Sum;
[0209] 2) Weighted smoothing process:
[0210] a. Select the switch k with the highest weight from the shadow table, insert it into the position pointed to by the current ID in the empty routing table, and then increment the ID value by 1;
[0211] b. Update the weight value of switch k in the shadow table: subtract Sum from the weight of switch k;
[0212] c. Determine if the weights of all switches in the shadow table are 0. If this condition is true, the weighted smoothing of the routing table has been completed; otherwise, proceed to step d.
[0213] d. Update the weight values of all switches in the shadow table: the weight of switch i = the current weight of switch i + the original weight of switch i, where the original weight of switch i is the weight of switch i in the switch weight ID table;
[0214] e. Repeat step ad.
[0215] In summary, the method provided in this application embodiment can achieve the following technical effects.
[0216] (1) In this embodiment, there is no need to pre-configure the path statically. Instead, the forwarding node performs hash routing based on the intra-network computing identifier of the intra-network computing message. This ensures that intra-network computing messages belonging to the same intra-network computing message are routed to the same next hop, thus satisfying the path consistency requirement of intra-network computing traffic. Furthermore, since there is no need to statically configure the path, the routing strategy of the forwarding node is transparent to the computing node, enhancing the transmission security of intra-network computing flows.
[0217] (2) Since no static path configuration is required, there is no need to configure an additional centralized management node for static path configuration, simplifying the network design. Even if the network changes in the future, the cost of network maintenance is relatively low.
[0218] (3) In this embodiment of the application, the intranet computing stream can be transmitted via Remote Direct Memory Access over Converged Ethernet (RoCE) instead of via unreliable UDP, thereby improving the transmission reliability of the intranet stream.
[0219] (4) The embodiments of this application can also perform weighted smoothing processing on the forwarding table. By smoothing the forwarding table, the intra-network computing flow received by each forwarding node can be matched as closely as possible to its own resource limitations after load balancing through the forwarding table.
[0220] The following is based on Figure 14 The method provided in the embodiments of this application will be illustrated using an example. It should be noted that... Figure 14 This is an optional implementation of the method provided in the embodiments of this application, and does not constitute a limitation on the method provided in the embodiments of this application.
[0221] like Figure 14 As shown, the method for calculating packets within this forwarding network includes the following steps:
[0222] Step 1: Update the switch weight ID table according to the switch weights, and perform weighted smoothing to generate a weighted smoothing routing table. Update the aggregated IP table (as shown in Table 1) based on the network-wide calculated packets from the control plane, recording True CommID, source IP, source port, QPN, and GroupSize in the aggregated IP table.
[0223] Step 2: When the switch receives the RoCEv2 message, it first determines whether the message belongs to the uplink network calculation message based on the aggregated IP table.
[0224] Step 2.1: If yes, then determine whether the current network-wide computation message is an uplink non-first packet message based on the information in the aggregation information table (such as the aggregation information mapping relationship shown in Table 2).
[0225] Step 2.1.1: If it is an uplink non-first packet, then proceed to the aggregation calculation process, and then go to step 3.
[0226] Step 2.1.2: If it is the first uplink packet, perform packet header parsing, record the CommID, Tag, MsgID, MsgLen and PSN information in the network calculation packet into the aggregation information table, enter the aggregation calculation processing flow, and then proceed to step 3.
[0227] Step 2.2: If not, determine if the current message is another type of message, such as a downlink result message or traditional traffic, and then forward it directly.
[0228] Step 3: For the network-wide computation message that has completed the aggregation computation process, parse the IP and UDP fields in the packet header, extract the protocol version number (protocol), destination port number (dstPort), and obtain the communication group identifier CommID and message identifier MsgID corresponding to the network-wide computation message from the aggregation information table, and form a new hash factor 4-tuple.
[0229] Step 4: Calculate a unique routing identifier based on the new hash factor.
[0230] Step 5: Based on the calculated routing identifier and the weighted smoothing routing table, perform Mod-N hash routing to determine the next routing forwarding port for the aggregated packet, and then proceed to the forwarding process.
[0231] Based on the methods in the foregoing embodiments, this application also provides an architecture for a forwarding node. Figure 15 This is a schematic diagram of the architecture of a forwarding node provided in an embodiment of this application. This forwarding node is used to implement the method for forwarding network-computed packets involved in the aforementioned embodiments.
[0232] like Figure 15 As shown, the forwarding node includes five components: a header parser, a packet processing module, a switching matrix, a CPU control plane, and a memory. The header parser, packet processing module, and switching matrix can be referred to as the data forwarding plane. The header parser is deployed at both the ingress and egress ports. The components involved in this embodiment include the CPU control plane, the ingress port header parser, and the packet processing module.
[0233] In some embodiments, all lower-level processing procedures for intra-network computation packets are first executed by the CPU control plane, and then the processing results are sent to the data forwarding plane for execution according to a preset procedure. For example, the intra-network computation identifier identification process, aggregation table update process, aggregation calculation process, hash factor extraction process for intra-network computation packets, and computation smoothing process can all be executed by the CPU control plane, and the execution results can then be announced to the data forwarding table, which will then process the data according to the predetermined procedure. Figure 15 The flowchart in the diagram completes the forwarding process of the entire network's computational packets.
[0234] In addition, such as Figure 15 As shown, the various tables processed by the CPU control plane, such as the aggregated IP table, aggregated information table, and weighted routing table, are directly stored in the memory.
[0235] Figure 16 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Any forwarding node in the foregoing embodiments can... Figure 16 This is implemented using the computer equipment shown. See also Figure 16 The computer device includes at least one processor 1601, a communication bus 1602, a memory 1603, and at least one communication interface 1604.
[0236] The processor 1601 may be a general-purpose central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application.
[0237] The communication bus 1602 may include a path for transmitting information between the aforementioned components.
[0238] Memory 1603 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory 1603 may exist independently and be connected to processor 1601 via communication bus 1602. Memory 1603 may also be integrated with processor 1601.
[0239] The memory 1603 stores program code for executing the scheme of this application, and its execution is controlled by the processor 1601. The processor 1601 executes the program code stored in the memory 1603. The program code may include one or more software modules. The forwarding node in the foregoing embodiment can determine the data used for application development through the processor 1601 and one or more software modules in the program code in the memory 1603.
[0240] The communication interface 1604 uses any transceiver-like device to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0241] In a specific implementation, as one example, a computer device may include multiple processors, for example... Figure 16The processors 1601 and 1605 are shown. Each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here can refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0242] In a specific implementation, as one embodiment, the computer device may further include an output device 1606 and an input device 1607. The output device 1606 communicates with the processor 1601 and can display information in various ways. For example, the output device 1606 may be a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 1607 communicates with the processor 1601 and can receive user input in various ways. For example, the input device 1607 may be a mouse, keyboard, touchscreen device, or sensing device, etc.
[0243] The aforementioned computer device can be a general-purpose computer device or a special-purpose computer device. In specific implementations, the computer device can be any communication device with data processing and forwarding capabilities, such as a switch with network-wide computing capabilities. This application does not limit the type of computer device.
[0244] Figure 17 This is a schematic diagram of a forwarding node device provided in an embodiment of this application. Figure 17 As shown, forwarding node 1700 includes the following modules.
[0245] Receiver module 1701 is used to receive multiple intranet computing messages;
[0246] Processing module 1702 is used to determine the intranet computing identifier of each intranet computing message in multiple intranet computing messages. The intranet computing identifier indicates the intranet computing message to which the corresponding intranet computing message belongs.
[0247] The processing module 1702 is also used to aggregate the intranet computing messages belonging to the same intranet computing message in the multiple intranet computing messages based on the intranet computing identifier of each intranet computing message in the multiple intranet computing messages, to obtain at least one aggregated message, and at least one aggregated message corresponds one-to-one with at least one intranet computing identifier;
[0248] The sending module 1703 is used to forward each aggregated message based on the network-computed identifier corresponding to each aggregated message in at least one aggregated message using a hash routing algorithm.
[0249] Optionally, the forwarding node stores an aggregation information mapping relationship, which includes at least one intra-network computing identifier and a data flow identifier corresponding to the at least one intra-network computing identifier.
[0250] The processing module is used for:
[0251] For the first intra-network computing message among multiple intra-network computing messages, determine the data flow identifier corresponding to the first intra-network computing message to obtain the first data flow identifier. The first intra-network computing message can be any one of the multiple intra-network computing messages.
[0252] If a first data flow identifier exists in the aggregation information mapping relationship, then the intranet computing identifier corresponding to the first data flow identifier is obtained from the aggregation information mapping relationship to obtain the first intranet computing identifier, and the first intranet computing identifier is used as the intranet computing identifier of the first intranet computing message.
[0253] Optionally, the processing module is also used for:
[0254] If the first data stream identifier does not exist in the aggregation information mapping relationship, then the first data stream identifier and the first intra-network computing identifier are obtained from the header of the first intra-network computing message;
[0255] Add the correspondence between the first intranet computing identifier and the first data stream identifier to the aggregated information mapping relationship.
[0256] Optionally, the processing module is used for:
[0257] Bind the first data stream identifier in the first network-internal computation message to the aggregation information mapping relationship;
[0258] Aggregate the intranet computing packets that are bound to multiple data flow identifiers corresponding to the first intranet computing identifier in the aggregation information mapping relationship to obtain the aggregated packet corresponding to the first intranet computing identifier.
[0259] Optionally, the processing module is used for:
[0260] For the intra-network computing packets that are respectively bound to multiple data flow identifiers corresponding to the first intra-network computing identifier, determine the packet sequence number of each bound intra-network computing packet;
[0261] Based on the first packet sequence number corresponding to each data flow identifier among the multiple data flow identifiers corresponding to the first intra-network computing identifier, and the packet sequence number of each bound intra-network computing message, the offset of each bound intra-network computing message is determined.
[0262] Based on the offset of each bound intranet computation packet, aggregate computation is performed on intranet computation packets with the same offset.
[0263] Optionally, the data stream identifier includes the first packet sequence number, the communication operation identifier, and the data stream length.
[0264] Optionally, the intranet computing identifier includes a communication group identifier and an intranet computing message identifier. The communication group identifier indicates the communication group currently performing intranet computing, and the intranet computing message identifier indicates an intranet computing message transmitted within the communication group.
[0265] Optionally, multiple intranet computing packets are Remote Direct Memory Access (RoCE) packets based on converged Ethernet. The forwarding node stores at least one Remote Direct Memory Access (RDMA) connection identifier. The RDMA connection identifier indicates an RDMA connection between the first computing node and the second computing node. The first computing node is the computing node accessed by the forwarding node.
[0266] The processing module is also used for:
[0267] For the second intra-network computing message among multiple intra-network computing messages, obtain the RDMA connection identifier carried by the second intra-network computing message. The second intra-network computing message can be any one of the multiple intra-network computing messages.
[0268] If at least one of the stored RDMA connection identifiers contains an RDMA connection identifier carried by the second intra-network computation message, then the operation of determining the corresponding intra-network computation identifier of the second intra-network computation message is performed.
[0269] Optionally, the RDMA connection identifier includes the source Internet Protocol IP address, the source port identifier, and the destination queue number (QPN).
[0270] Optionally, the receiving module is also configured to receive notification messages from the uplink communication link;
[0271] The processing module is also used to obtain the target RDMA connection identifier carried in the notification message. The notification message is used to notify that the RDMA connection indicated by the target RDMA connection identifier is used to transmit intra-network calculation messages.
[0272] The processing module is also used to store the target RDMA connection identifier.
[0273] Optionally, the processing module is used for:
[0274] Retrieve the type information carried in the notification message;
[0275] If the type information indicates that the announcement message is an announcement message for intranet compute packets, then the operation of obtaining the target RDMA connection identifier carried in the announcement message is performed.
[0276] Optionally, the sending module is used for:
[0277] For the first aggregated message in at least one aggregated message, a first hash factor is determined. The first hash factor includes the network-internal calculation identifier corresponding to the first aggregated message. The first aggregated message is any one of the at least one aggregated message.
[0278] The first routing identifier is determined by a hash routing algorithm based on the first hash factor, and the first routing identifier indicates a forwarding path;
[0279] The first aggregated message is forwarded based on the first routing identifier.
[0280] Optionally, the first hash factor may also include the protocol version number and destination port number carried in the network-based computational message corresponding to the first aggregated message.
[0281] Optionally, the sending module is used for:
[0282] For the second aggregated message in at least one aggregated message, the second aggregated message is forwarded through a hash routing algorithm based on the network-computed identifier and forwarding table corresponding to the second aggregated message. The second aggregated message is any one of the at least one aggregated message.
[0283] The forwarding table includes multiple target forwarding entries. The next hop in the multiple target forwarding entries is the same, and the total number of multiple target forwarding entries indicates the network computing capacity of the corresponding next hop. The multiple target forwarding entries are scattered in the forwarding table.
[0284] In this embodiment, after aggregating intranet computing packets, the forwarding node does not need to forward them through a pre-configured static path. Instead, the forwarding node performs hash routing based on the intranet computing identifier of the intranet computing packets. This ensures that intranet computing packets belonging to the same intranet computing message are routed to the same next hop. In other words, while using dynamic routing, the path consistency requirement of intranet computing traffic is met, which improves the flexibility of forwarding intranet computing packets.
[0285] Furthermore, since no static path configuration is required, the routing policies of forwarding nodes are transparent to the compute nodes, enhancing the security of compute stream transmission within the network. Additionally, the elimination of static path configuration eliminates the need for a separate centralized management node, simplifying network design. Even if the network undergoes changes, network maintenance costs are relatively low.
[0286] It should be noted that the above embodiments, when providing the forwarding node for packet calculation within the forwarding network, are only illustrative examples of the division of the aforementioned functional modules. In practical applications, the functions described above can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the forwarding node and the method embodiments for packet calculation within the forwarding network provided in the above embodiments belong to the same concept; the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0287] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0288] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0289] The above content is not intended to limit the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of the embodiments of this application.
Claims
1. A method for computing a packet in a forwarding network, characterized in that, The method includes: The forwarding node receives multiple intra-network computational packets; The forwarding node determines the intranet computing identifier of each intranet computing message among the plurality of intranet computing messages, and the intranet computing identifier indicates the intranet computing message to which the corresponding intranet computing message belongs; The forwarding node aggregates intra-network computing messages belonging to the same intra-network computing message based on the intra-network computing identifier of each intra-network computing message in the plurality of intra-network computing messages to obtain at least one aggregated message. The at least one aggregated message corresponds one-to-one with at least one intra-network computing identifier. The forwarding node forwards each aggregated message based on the network-computed identifier corresponding to each aggregated message in the at least one aggregated message using a hash routing algorithm.
2. The method of claim 1, wherein, The forwarding node stores an aggregation information mapping relationship, which includes at least one intra-network computing identifier and a data stream identifier corresponding to the at least one intra-network computing identifier. The forwarding node determines the intra-network computing identifier of each intra-network computing packet among the plurality of intra-network computing packets, including: For the first intra-network computing message among the plurality of intra-network computing messages, determine the data flow identifier corresponding to the first intra-network computing message to obtain the first data flow identifier. The first intra-network computing message is any one of the plurality of intra-network computing messages. If the first data flow identifier exists in the aggregation information mapping relationship, then the intranet computing identifier corresponding to the first data flow identifier is obtained from the aggregation information mapping relationship to obtain the first intranet computing identifier, and the first intranet computing identifier is used as the intranet computing identifier of the first intranet computing message.
3. The method of claim 2, wherein, The method further includes: If the first data stream identifier does not exist in the aggregation information mapping relationship, then the first data stream identifier and the first intra-network computing identifier are obtained from the header of the first intra-network computing message; Add the correspondence between the first intranet computing identifier and the first data stream identifier to the aggregation information mapping relationship.
4. The method of claim 2 or 3, wherein, The forwarding node, based on the intra-network computing identifier of each intra-network computing message in the plurality of intra-network computing messages, aggregates and calculates intra-network computing messages belonging to the same intra-network computing message, including: Bind the first intranet computed message to the first data stream identifier in the aggregation information mapping relationship; Aggregate the intranet computing packets that are respectively bound to multiple data flow identifiers corresponding to the first intranet computing identifier in the aggregation information mapping relationship to obtain the aggregated packet corresponding to the first intranet computing identifier.
5. The method as described in claim 4, characterized in that, The aggregation calculation of the intra-network computing packets bound to the multiple data flow identifiers corresponding to the first intra-network computing identifier in the aggregation information mapping relationship includes: For the intra-network computing packets that are respectively bound to the multiple data flow identifiers corresponding to the first intra-network computing identifier, determine the packet sequence number of each bound intra-network computing packet; Based on the first packet sequence number of each data flow identifier in the multiple data flow identifiers corresponding to the first intra-network computing identifier, and the packet sequence number of each bound intra-network computing message, the offset of each bound intra-network computing message is determined. Based on the offset of each bound intranet computation packet, aggregate computation is performed on intranet computation packets with the same offset.
6. The method as described in any one of claims 2-3 and 5, characterized in that, The data stream identifier includes the first packet sequence number, the communication operation identifier, and the data stream length.
7. The method as described in any one of claims 1-3 and 5, characterized in that, The intranet computing identifier includes a communication group identifier and an intranet computing message identifier. The communication group identifier indicates the communication group currently performing intranet computing, and the intranet computing message identifier indicates an intranet computing message transmitted within the communication group.
8. The method according to any one of claims 1-3, 5, characterized in that, The multiple intranet computing packets are Remote Direct Memory Access (RoCE) packets based on converged Ethernet. The forwarding node stores at least one Remote Direct Memory Access (RDMA) connection identifier. The RDMA connection identifier indicates an RDMA connection between a first computing node and a second computing node. The first computing node is the computing node accessed by the forwarding node. Before the forwarding node determines the intra-network computing identifier of each intra-network computing packet among the plurality of intra-network computing packets, the method further includes: For the second intra-network computing message among the plurality of intra-network computing messages, the forwarding node obtains the RDMA connection identifier carried by the second intra-network computing message, where the second intra-network computing message is any one of the plurality of intra-network computing messages; If the stored at least one RDMA connection identifier contains an RDMA connection identifier carried by the second intra-network computing message, then the operation of determining the corresponding intra-network computing identifier of the second intra-network computing message is performed.
9. The method as described in claim 8, characterized in that, The RDMA connection identifier includes the source Internet Protocol IP address, the source port identifier, and the destination queue number (QPN).
10. The method as described in claim 8, characterized in that, The method further includes: The forwarding node receives an announcement message from the uplink communication link; The forwarding node obtains the target RDMA connection identifier carried in the announcement message, and the announcement message is used to announce that the RDMA connection indicated by the target RDMA connection identifier is used to transmit intra-network computational packets. The forwarding node stores the target RDMA connection identifier.
11. The method as described in claim 10, characterized in that, The forwarding node obtains the target RDMA connection identifier carried in the announcement message, including: The forwarding node obtains the type information carried in the announcement message; If the type information indicates that the announcement message is an announcement message for an intranet compute message, then the operation of obtaining the target RDMA connection identifier carried in the announcement message is performed.
12. The method according to any one of claims 1-3, 5, and 9-11, characterized in that, The forwarding node forwards each aggregated packet based on the network-computed identifier corresponding to each aggregated packet in the at least one aggregated packet using a hash routing algorithm, including: For the first aggregated message in the at least one aggregated message, a first hash factor is determined. The first hash factor includes the network-internal calculation identifier corresponding to the first aggregated message. The first aggregated message is any one of the at least one aggregated messages. Based on the first hash factor, a first routing identifier is determined by the hash routing algorithm, and the first routing identifier indicates a forwarding path; The first aggregated message is forwarded based on the first routing identifier.
13. The method as described in claim 12, characterized in that, The first hash factor also includes the protocol version number and destination port number carried in the network-based computational message corresponding to the first aggregated message.
14. The method according to any one of claims 1-3, 5, 9-11, and 13, characterized in that, The forwarding node forwards each aggregated packet based on the network-computed identifier corresponding to each aggregated packet in the at least one aggregated packet using a hash routing algorithm, including: For the second aggregated packet in the at least one aggregated packet, the second aggregated packet is forwarded through a hash routing algorithm based on the network-internal calculated identifier and forwarding table corresponding to the second aggregated packet, wherein the second aggregated packet is any one of the at least one aggregated packets; The forwarding table includes multiple target forwarding entries, all of which have the same next hop. The total number of the multiple target forwarding entries indicates the network computing capacity of the corresponding next hop. The multiple target forwarding entries are scattered throughout the forwarding table.
15. A forwarding node, characterized in that, The forwarding nodes include: The receiving module is used to receive multiple intra-network computing messages; The processing module is used to determine the intranet computing identifier of each intranet computing message among the plurality of intranet computing messages, wherein the intranet computing identifier indicates the intranet computing message to which the corresponding intranet computing message belongs; The processing module is further configured to aggregate intra-network computing messages belonging to the same intra-network computing message based on the intra-network computing identifier of each intra-network computing message in the plurality of intra-network computing messages, to obtain at least one aggregated message, wherein the at least one aggregated message corresponds one-to-one with at least one intra-network computing identifier; The sending module is used to forward each aggregated message based on the network-internal calculation identifier corresponding to each aggregated message in the at least one aggregated message using a hash routing algorithm.
16. The forwarding node as described in claim 15, characterized in that, The forwarding node stores an aggregation information mapping relationship, which includes at least one intra-network computing identifier and a data stream identifier corresponding to the at least one intra-network computing identifier. The processing module is used for: For the first intra-network computing message among the plurality of intra-network computing messages, determine the data flow identifier corresponding to the first intra-network computing message to obtain the first data flow identifier. The first intra-network computing message is any one of the plurality of intra-network computing messages. If the first data flow identifier exists in the aggregation information mapping relationship, then the intranet computing identifier corresponding to the first data flow identifier is obtained from the aggregation information mapping relationship to obtain the first intranet computing identifier, and the first intranet computing identifier is used as the intranet computing identifier of the first intranet computing message.
17. The forwarding node as described in claim 16, characterized in that, The processing module is also used for: If the first data stream identifier does not exist in the aggregation information mapping relationship, then the first data stream identifier and the first intra-network computing identifier are obtained from the header of the first intra-network computing message; Add the correspondence between the first intranet computing identifier and the first data stream identifier to the aggregation information mapping relationship.
18. The forwarding node as described in claim 16 or 17, characterized in that, The processing module is used for: Bind the first intranet computed message to the first data stream identifier in the aggregation information mapping relationship; Aggregate the intranet computing packets that are respectively bound to multiple data flow identifiers corresponding to the first intranet computing identifier in the aggregation information mapping relationship to obtain the aggregated packet corresponding to the first intranet computing identifier.
19. The forwarding node as described in claim 18, characterized in that, The processing module is used for: For the intra-network computing packets that are respectively bound to the multiple data flow identifiers corresponding to the first intra-network computing identifier, determine the packet sequence number of each bound intra-network computing packet; Based on the first packet sequence number of each data flow identifier in the multiple data flow identifiers corresponding to the first intra-network computing identifier, and the packet sequence number of each bound intra-network computing message, the offset of each bound intra-network computing message is determined. Based on the offset of each bound intranet computation packet, aggregate computation is performed on intranet computation packets with the same offset.
20. The forwarding node as described in any one of claims 16-17 and 19, characterized in that, The data stream identifier includes the first packet sequence number, the communication operation identifier, and the data stream length.
21. The forwarding node as described in any one of claims 15-17, 19, characterized in that, The intranet computing identifier includes a communication group identifier and an intranet computing message identifier. The communication group identifier indicates the communication group currently performing intranet computing, and the intranet computing message identifier indicates an intranet computing message transmitted within the communication group.
22. The forwarding node as described in any one of claims 15-17, 19, characterized in that, The multiple intranet computing packets are Remote Direct Memory Access (RoCE) packets based on converged Ethernet. The forwarding node stores at least one Remote Direct Memory Access (RDMA) connection identifier. The RDMA connection identifier indicates an RDMA connection between a first computing node and a second computing node. The first computing node is the computing node accessed by the forwarding node. The processing module is also used for: For the second intra-network computing message among the plurality of intra-network computing messages, obtain the RDMA connection identifier carried by the second intra-network computing message, wherein the second intra-network computing message is any one of the plurality of intra-network computing messages; If the stored at least one RDMA connection identifier contains an RDMA connection identifier carried by the second intra-network computing message, then the operation of determining the corresponding intra-network computing identifier of the second intra-network computing message is performed.
23. The forwarding node as described in claim 22, characterized in that, The RDMA connection identifier includes the source Internet Protocol IP address, the source port identifier, and the destination queue number (QPN).
24. The forwarding node as described in claim 22, characterized in that, The receiving module is also used to receive notification messages from the uplink communication link; The processing module is further configured to obtain the target RDMA connection identifier carried in the notification message, wherein the notification message is used to notify that the RDMA connection indicated by the target RDMA connection identifier is used to transmit intra-network computing messages; The processing module is also used to store the target RDMA connection identifier.
25. The forwarding node as described in claim 24, characterized in that, The processing module is used for: Obtain the type information carried by the notification message; If the type information indicates that the announcement message is an announcement message for an intranet compute message, then the operation of obtaining the target RDMA connection identifier carried in the announcement message is performed.
26. The forwarding node as described in any one of claims 15-17, 19, and 23-25, characterized in that, The sending module is used for: For the first aggregated message in the at least one aggregated message, a first hash factor is determined. The first hash factor includes the network-internal calculation identifier corresponding to the first aggregated message. The first aggregated message is any one of the at least one aggregated messages. Based on the first hash factor, a first routing identifier is determined by the hash routing algorithm, and the first routing identifier indicates a forwarding path; The first aggregated message is forwarded based on the first routing identifier.
27. The forwarding node as described in claim 26, characterized in that, The first hash factor also includes the protocol version number and destination port number carried in the network-based computational message corresponding to the first aggregated message.
28. The forwarding node as described in any one of claims 15-17, 19, 23-25, and 27, characterized in that, The sending module is used for: For the second aggregated packet in the at least one aggregated packet, the second aggregated packet is forwarded through a hash routing algorithm based on the network-internal calculated identifier and forwarding table corresponding to the second aggregated packet, wherein the second aggregated packet is any one of the at least one aggregated packets; The forwarding table includes multiple target forwarding entries, all of which have the same next hop. The total number of the multiple target forwarding entries indicates the network computing capacity of the corresponding next hop. The multiple target forwarding entries are scattered throughout the forwarding table.
29. A forwarding node, characterized in that, The forwarding node includes a memory and a processor; The memory is used to store a program that supports the processor in executing the method according to any one of claims 1-14, and to store data related to implementing the method according to any one of claims 1-14; The processor is configured to execute programs stored in the memory.
30. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the method described in any one of claims 1-14.